AI builds its own spatial model of the internet, in which information it sees frequently sits in the center and information that appears only once or twice is relegated to the edges.
In plain terms, the model reliably remembers what it saw often and forgets or misplaces what it saw rarely. Product details are the kind of information the web mentions only a handful of times.
Chatbots are more likely to make mistakes when answering questions about mid-sized and smaller brands
In September 2025, OpenAI’s research showed that a model cannot reliably recall facts it saw only once during training, which means virtually every model will hallucinate. In their words, “Our lower-bound for hallucinations is based on the fraction of prompts appearing just once in the training data.” 7
And the less often a product appears online, the less the model remembers about it.
For example, a famous fact like the capital of France appears everywhere online. But the spec sheet for one washing machine revision from a mid-tier brand might appear a handful of times, so models may not remember details like its water usage, its exact drum volume, and whether this year’s version kept last year’s motor.
Product.ai’s divergence study confirms this phenomenon. Asked for the best workout headphones under $150, ChatGPT recommended five separate products, across two of its five runs, that have all been discontinued or replaced by newer models, with no mention that any of them were out of date. 3
Similar products blur together
Models can also get confused, and may hallucinate, when products have similar-sounding names. In March 2025, Anthropic’s interpretability team showed that models may switch off their default caution, and commit to answering before they have anything true to say, if they recognize a name as familiar. 8
How does this play out in real life? A model that read thousands of pages about last year’s XM5 headphones may recognize the XM6 as familiar and mix features from both products in its answers about either.
In Product.ai’s divergence study, ChatGPT recommended the Sony WH-1000XM5 as the current top pick for frequent flying in all five of its test runs, though Sony had already replaced that model with the WH-1000XM6. Gemini named the same outdated model in three of its five runs, even with live web search turned on. Only Claude and Perplexity named the current model every time. 3
The model often knows more than it says
Models also struggle to give the right answer even when it exists in their training data. Independent studies keep finding a large gap between what a model detectably “knows” and what it actually says. In one 2025 study, models encoded on average 40% more factual knowledge internally than they expressed in their answers, with the gap ranging from 14% to 57% depending on the model. 9
So, of course, you should be skeptical of AI answers, especially when AI is talking about mid-sized or smaller brands. But what surprises many people is that models won’t tell you when they don’t know.