I also believe different people think in different ways.
Some think visually, some think more abstract.
I remember an exchange on NOSTR some time ago, but it's also a commonly discussed theme:
Think about "coffee" for example.
Some people only hear the word, others start smelling the brew, others visualise a cup with liquid in it. Some do a mixture of these.
An interesting AI intersection exists here. LLMs are large language models. If you upload a photo, they have an image analyser that describes the image to them. I'm using ChatGPT to teach me photography, I regularly upload images for it to give critical analysis to.
It can give incredibly detailed analysis of anything relating to that image. It can describe the subject in ever increasing intricate detail, it can analyse composition, it can explain technical failings. This is a facsimile of seeing, not seeing itself.
