Product.ai
Product.ai / Research / Why you ask AI for product recommendations and spend hours researching anyway
Essay The Operator's Codex · No. 17

Why you ask AI for product recommendations and spend hours researching anyway

People don’t trust product recommendations from AI chatbots for good reason. Hallucination is baked into how AI works, publishers have banned agents from scraping product reviews, and the freshest product data now comes from sellers.

MQ
Michael Quoc
Founder & CEO
Published: August 25, 2026
Terminal-style graphic: one question, ‘which wireless headphones should I buy?’, fans out into four different AI answers, captioned ‘same question, four different answers, hours of research anyway.’

When you open any standard AI assistant, the first thing you see is some variant of this statement: “AI can make mistakes. Always double-check its answers.”

And the AI companies aren’t lying when they say this. This year’s Stanford AI Index found that the 26 top AI models have hallucination rates ranging from 22% to 94% on deliberately difficult factual questions. 1

Those rates come from hard questions rather than everyday use. But product questions are often complex, and when researchers measured product questions directly, the numbers were worse.

Standard AI product recommendations are unreliable

According to research from the University of California San Diego published in February 2026, AI review summaries hallucinated in responses to 60% of product questions and shifted the sentiment of the reviews they summarized in 26.5% of cases. 2

Those distortions influenced buying. In the same study, people who read the AI summaries said they’d buy the product 84% of the time, compared with 52% for people who read the human reviews the summaries were based on.

Product.ai’s research team also analyzed chatbots’ product answers for accuracy issues. In August 2026, Product.ai asked ChatGPT, Claude, Gemini, and Perplexity the same 200 real shopping questions, five times each, and compared the 4,000 answers. 78% of the questions produced at least one confirmed factual conflict, meaning a contradicted price, a discontinued product presented as current, or a spec claim the engines could not agree on. 3

In the Product.ai AI Shopping Divergence Study, 156 of 200 shopping questions, or 78 percent, produced at least one confirmed factual conflict across ChatGPT, Claude, Gemini, and Perplexity, and all 23 headphone questions conflicted
78% of 200 real shopping questions produced a confirmed factual conflict across four engines. Source: Product.ai AI Shopping Divergence Study.

Today, even TikTok accounts, like the one operated by tech writer Deana Burke, are showing that product recommendations from AI chatbots are unreliable and getting worse.

That’s because hallucinations aren’t the only problem. The quality of the data online is also to blame.

Internet data pollution is making it worse

Sources of reputable information are taking steps to stop their content from being crawled by LLMs while affiliate sellers are optimizing their content for them.

This means the AI companies’ training crawlers are locked out of some of the best-known review publishers, and AI training data is likely to include more biased sources and less journalism. In many cases the publishers also block web search crawlers like OpenAI’s and Perplexity’s, so a chatbot searching the web on your behalf is more likely to collect marketing copy, simply because it’s more available.

The reviewers are opting out

You can check this yourself. As of August 2026, the file Consumer Reports uses to tell crawlers what they may read blocks OpenAI’s GPTBot, Google’s AI training crawler, Meta’s crawlers, and Common Crawl. The New York Times, which owns Wirecutter, blocks a long list of AI crawlers, including the search-time ones from OpenAI, Perplexity, and Anthropic. 4

Live robots.txt captures from August 16, 2026: consumerreports.org disallows GPTBot, CCBot, Google-Extended, and Meta’s crawler; nytimes.com disallows OAI-SearchBot, PerplexityBot, anthropic-ai, ClaudeBot, and more
The review publishers’ own crawler instructions, captured August 16, 2026: access refused at the front door.

The sellers are optimizing what remains

Meanwhile, sellers are working hard to get their content ingested by frontier model bots. Optimizing content for AI answers is now an industry called generative engine optimization (GEO) or answer engine optimization (AEO), and one market projection puts it at roughly $365 million in the US for 2026. 5

And it works. The Princeton-led research that coined the GEO name showed that specific content changes can raise a source’s visibility in AI answers by up to 40%, visibility here meaning how often the source gets cited or quoted in the generated answer. 6

The optimized pages also get picked up, and nothing in that process checks whether a page is true.

This is not an easy problem to fix

Common sense suggests the obvious solution is for AI providers to fix models so they don’t hallucinate and to negotiate with publications so they can access better data. But it isn’t that simple.

Earlier this year I wrote about the Beige Singularity, the flood of same-sounding AI content spreading across the web. This article takes it one step further to explain how hallucinations are baked into AI’s very design and how product recommendations in particular are vulnerable to internet slop produced at scale.

Your instinct to double-check is correct

AI is a non-deterministic technology. In practice, this means it’s generating the most likely answer to any given product question, not the “best” one. It also means that AI often won’t give you the same answer to a question twice, so if you ask “what’s the best washing machine?” you won’t always see the same list of brands.

We measured this, too. In Product.ai’s divergence study, Gemini contradicted its own earlier answer to the identical question, with nothing new to justify the change, on 30% of the questions. ChatGPT did it on 20%. 3

Share of questions where an engine contradicted its own earlier answer to the identical question in the Product.ai AI Shopping Divergence Study: Gemini 30 percent, Claude 26 percent, ChatGPT 20 percent, Perplexity 8 percent
The same question, minutes apart: how often each engine contradicted its own earlier answer. Source: Product.ai AI Shopping Divergence Study.

How does this happen? AI models assemble each answer one word at a time from probabilities they absorbed during training. They don’t do any new research or check the finished sentence against known sources of truth.

AI builds a compressed photo of the internet, not a database

AI builds its own spatial model of the internet, in which information it sees frequently sits in the center and information that appears only once or twice is relegated to the edges.

In plain terms, the model reliably remembers what it saw often and forgets or misplaces what it saw rarely. Product details are the kind of information the web mentions only a handful of times.

Chatbots are more likely to make mistakes when answering questions about mid-sized and smaller brands

In September 2025, OpenAI’s research showed that a model cannot reliably recall facts it saw only once during training, which means virtually every model will hallucinate. In their words, “Our lower-bound for hallucinations is based on the fraction of prompts appearing just once in the training data.” 7

And the less often a product appears online, the less the model remembers about it.

For example, a famous fact like the capital of France appears everywhere online. But the spec sheet for one washing machine revision from a mid-tier brand might appear a handful of times, so models may not remember details like its water usage, its exact drum volume, and whether this year’s version kept last year’s motor.

Product.ai’s divergence study confirms this phenomenon. Asked for the best workout headphones under $150, ChatGPT recommended five separate products, across two of its five runs, that have all been discontinued or replaced by newer models, with no mention that any of them were out of date. 3

Similar products blur together

Models can also get confused, and may hallucinate, when products have similar-sounding names. In March 2025, Anthropic’s interpretability team showed that models may switch off their default caution, and commit to answering before they have anything true to say, if they recognize a name as familiar. 8

How does this play out in real life? A model that read thousands of pages about last year’s XM5 headphones may recognize the XM6 as familiar and mix features from both products in its answers about either.

In Product.ai’s divergence study, ChatGPT recommended the Sony WH-1000XM5 as the current top pick for frequent flying in all five of its test runs, though Sony had already replaced that model with the WH-1000XM6. Gemini named the same outdated model in three of its five runs, even with live web search turned on. Only Claude and Perplexity named the current model every time. 3

The model often knows more than it says

Models also struggle to give the right answer even when it exists in their training data. Independent studies keep finding a large gap between what a model detectably “knows” and what it actually says. In one 2025 study, models encoded on average 40% more factual knowledge internally than they expressed in their answers, with the gap ranging from 14% to 57% depending on the model. 9

So, of course, you should be skeptical of AI answers, especially when AI is talking about mid-sized or smaller brands. But what surprises many people is that models won’t tell you when they don’t know.

AI is taught to guess instead of admit it doesn’t know

The way models are trained and tested tends to reward them for guessing and punish them for admitting doubt.

First, most AI evaluations grade like a multiple-choice exam, so guessing is an efficient strategy. In OpenAI’s September 2025 paper, guessing produced better scores than admitting uncertainty, so models learn to always guess. 7

Second, Reinforcement Learning from Human Feedback also tends to push AI models away from hesitant answers. The people rating AI answers can’t fact-check a drum volume in thirty seconds, but they can tell which answer sounds confident and agreeable, so confident and agreeable is what the training rewards.

For example, in April 2025, OpenAI rolled back a GPT-4o update that had turned excessively flattering, and its postmortem said plainly that the training had overweighted thumbs-up feedback. 10

Frontier model providers are trying to address this by giving models tools to look up facts on the internet. But anyone who uses an AI assistant to find facts on the web knows this doesn’t always work.

AI web search is flawed

When an AI assistant searches the web, it brings back pages that rank highly on Google or Bing, and some of those pages are produced by sellers that are both biased and very good at SEO.

Also, if the model misunderstood your initial query, that misunderstanding is cycled into all of its sub-searches. For example, one prompt might fan out into roughly eight to twelve unique searches, each containing the original error. The model then reads the top-ranked pages from each search and quickly assembles its answer from them. It’s easy to see how these might include mistakes.

Citations are generated

The links under an AI answer suggest that answer was verified against sources. But that’s not actually the case. Models generate citations the same way they generate their sentences, and researchers have found that citations are sometimes attached after the answer has already been produced.

In May 2026, an audit by researchers at Washington University in St. Louis broke Google’s AI Overview answers into 98,020 individual claims and checked each one against the pages it cited. 11.0% of the claims weren’t supported by their own citations. And that audit only checked the sources the engine itself decided to show. 11

Each assistant searches a different version of the internet

Different assistants also search very different parts of the web. In an analysis of 680 million citations by the AI-visibility firm Profound, only 11% of the web domains cited by ChatGPT overlapped with the domains cited by Perplexity, and no pair of major engines overlapped more than about 16%. 12

Diagram of domains cited by ChatGPT and Perplexity across 680 million citations analyzed by Profound: only 11 percent of cited domains appear in both
Each assistant searches its own version of the internet: 11% domain overlap between ChatGPT and Perplexity. Source: Profound.

This is because each company built its own search system. In other words, they search different parts of the web, because they use different search indexes, prioritize different trusted sources, and set different retrieval policies. So, even if two models were trained on roughly the same data, they will provide different answers based on the information they can access from the web.

The sources each engine relies on change constantly as well. In Profound’s follow-up analysis, the set of domains an engine cited for identical questions changed by 40% to 60% within a month, with each engine shuffling its sources in its own way. 13

How much an engine searches changes the answer

The engines also differ in how much searching they do at all, sometimes on the same question. In Product.ai’s divergence study, ChatGPT looked up live information for none of its answers to “what’s the best noise-cancelling headphones,” and for all of its answers when we named two specific products to compare. An engine that skips the search is answering from its training memory, with all the memory problems described above. 3

And the web the engines search is no longer their only source of product data. The freshest information now comes from somewhere else entirely.

Sellers hand their data straight to the AI companies

As AI referrals make up an increasing share of search traffic, sellers have also learned how to gain preferential treatment from AI by writing lengthy answers in the form of structured data embedded in their web pages. Over the past year and a half, AI providers have started taking product data directly from sellers.

First came product feeds, then came ads

In April 2025, ChatGPT launched product cards and said the results were chosen independently and were “not ads.” 14 Merchants began submitting product feeds that push updates as often as every fifteen minutes.

A product feed is a data file the seller maintains, listing current prices, stock, and specifications for its products. The seller publishes updates and AI assistants take them in as-is. This means the freshest product data assistants get is from biased sources.

Instant Checkout followed in September 2025, then was scaled back around March 2026 after only a few dozen merchants ever went live, according to reporting from The Information. 15 But the march toward mixing answers and ads continued. Retailer catalog ads began in May 2026, and product carousel ads arrived in August 2026. 16

Timeline of ChatGPT shopping from April 2025 to August 2026: product cards billed as not ads, Instant Checkout launches, checkout scaled back to a few dozen merchants, retailer catalog ads begin, product carousel ads arrive
Sixteen months from ‘not ads’ to sponsored placements inside the shopping conversation. Sources: Digiday · The Information · OpenAI.

Today, paid placements are labeled as “sponsored,” but they appear inside conversations that users expect to provide neutral advice.

And it’s not just ChatGPT. Google is doing the same thing at a much larger scale. By its own account in January 2026, its Shopping Graph passed 50 billion product listings with more than 2 billion updated every hour, and the same month it launched a commerce protocol with partners including Shopify, Etsy, Wayfair, Target, and Walmart. 17 Similarly, Perplexity runs a PayPal-powered checkout program that is expected to make more than 5,000 merchants purchasable inside its answers. 18

Seller data is full of errors

While AI assistants get fresh data from sellers, that data is often inaccurate. Audits of seller-submitted product data routinely find double-digit error rates, and one marketplace-data company, Sellintu, reports that 78% of product listings are rejected for incomplete or incorrect data. 19

This means the newest product data inside AI shopping comes from brands and retailers, and this represents a conflict of interest when chatbots claim to provide objective advice. And virtually none of this seller data is ever checked for mistakes.

AI providers can’t afford to verify their answers

Chatbot providers allocate a tiny compute budget for every consumer query, especially on free tiers, and those small budgets are not enough to support a thorough analysis. This means the answers are very often somewhat or very incorrect.

Advanced reasoning models don’t solve this problem. A model that thinks longer using unverified inputs still ends up with unverified conclusions. The Stanford numbers from the top of this piece, 22% to 94% across top models on hard factual questions, show that smarter models don’t necessarily hallucinate less. 1

Also, a model’s average accuracy doesn’t necessarily predict how it will answer any given product question. The same model can rank near the top of one accuracy test and near the bottom of another, so the number you see depends on which test someone chose to run.

The same question, re-answered millions of times

Cheap, inaccurate answers are incinerating countless tokens with little upside for consumers or frontier model providers. OpenAI’s own usage research reports around 2.5 billion ChatGPT messages a day, with roughly 2% about purchasable products. That works out to about 50 million shopping questions a day, for one assistant alone. 20

Each one is answered from scratch. Because nothing learned in one answer is saved for the next, the same shallow research gets repeated millions of times and thrown away.

We need a verification layer for commerce

AI answers could be better if providers invested in building and storing verified answers to commonly asked questions. And doing this right would involve getting answers from multiple frontier models.

A possible approach is to verify product claims in advance and build a knowledge base of verified recommendations. Claims that can’t be validated are discarded, and those that pass are added to the knowledge base so they’re available to help answer future questions. Using multiple AI models to assess claims can make recommendations even more reliable by uncovering facts on which models disagree. These represent opportunities for further investigation.

At Product.ai, we’re building that verification layer for commerce. How it works is the subject of a separate article.

How to evaluate recommendations in the meantime

Until a verification layer exists, a product recommendation from a single frontier model should not be taken at face value. Look at each answer critically. Here is the checklist Product.ai applies when reading one.

1. Ask where each product claim came from. Then open the source and confirm the number appears there. If the citation doesn’t contain the claim, treat the claim as invented.

2. Distrust AI-generated specs for products that offer many similarly named models or a new version every year. There is a good chance your model will confuse these variants.

3. Double-check answers that aren’t tagged with dates. Undated information may be stale.

4. Ask a different AI model and compare the answers. When two or three agree and one doesn’t, the odd one out is usually, but not always, the wrong one. This cross-model agreement check is the same principle Product.ai’s verification layer runs at scale.

5. For big-ticket purchases, leave the chatbot and go to the manufacturer’s website. Lab tests, third-party reviews with tests, and primary spec sheets will help you understand the strengths and weaknesses of any product. Of course, parsing it all will take some work, but it’s worth it for more expensive buys.

Shoppers are already doing a lot of this. In Product.ai’s Trust in AI Commerce Report, 86% of shoppers who use AI for product research said they verify AI recommendations through another source before buying. 21 And, until a verification layer becomes a routine part of how AI answers are delivered, we should look at AI-generated shopping recommendations as fallible data that needs confirmation from another source.

Product.ai’s research on AI product answers is published as the AI Shopping Divergence Study, with the full question set and every answer available for download. 3

p.s. If an AI has ever suggested that you buy a product that doesn’t exist, I’d like to hear about it. Send me the screenshot.

Sources

  1. Stanford Institute for Human-Centered Artificial Intelligence, “Responsible AI,” The 2026 AI Index Report, Stanford University, 2026. https://hai.stanford.edu/ai-index/2026-ai-index-report/responsible-ai
  2. University of California San Diego, “How Much Does Chatbot Bias Influence Users? A Lot, It Turns Out,” UC San Diego Today, February 9, 2026. https://today.ucsd.edu/story/how-much-does-chatbot-bias-influence-users-a-lot-it-turns-out
  3. Sean Fisher and Dakota Nunley, “We Asked Four AI Engines the Same 200 Shopping Questions. 78% Produced a Factual Conflict” (AI Shopping Divergence Study), Product.ai Research, August 10, 2026. https://product.ai/research/ai-shopping-divergence-study/
  4. Consumer Reports robots.txt (https://www.consumerreports.org/robots.txt) and The New York Times robots.txt (https://www.nytimes.com/robots.txt), both accessed August 16, 2026.
  5. Dimension Market Research, “US Generative Engine Optimization Market,” 2026 (vendor estimate: $365.4 million for 2026). https://dimensionmarketresearch.com/report/us-generative-engine-optimization-market/
  6. Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande, “GEO: Generative Engine Optimization,” KDD 2024. https://arxiv.org/abs/2311.09735
  7. Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang, “Why Language Models Hallucinate,” OpenAI, September 2025, arXiv:2509.04664 (https://arxiv.org/abs/2509.04664; OpenAI’s announcement: https://openai.com/index/why-language-models-hallucinate/). Three of the four authors are OpenAI researchers; Vempala is at Georgia Tech. Builds on Kalai and Vempala, “Calibrated Language Models Must Hallucinate,” STOC 2024 (https://arxiv.org/abs/2311.14648).
  8. Jack Lindsey et al. (Anthropic), “On the Biology of a Large Language Model,” Transformer Circuits Thread, March 27, 2025. https://transformer-circuits.pub/2025/attribution-graphs/biology.html
  9. Zorik Gekhman et al., “Inside-Out: Hidden Factual Knowledge in LLMs,” COLM 2025, arXiv:2503.15299 (average relative gap of 40% between internally encoded and externally expressed factual knowledge; per-model gaps of 14% for Llama, 48% for Mistral, and 57% for Gemma). https://arxiv.org/abs/2503.15299. See also Hadas Orgad et al., “LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations,” ICLR 2025, arXiv:2410.02707 (https://arxiv.org/abs/2410.02707).
  10. OpenAI, “Sycophancy in GPT-4o: What Happened and What We’re Doing About It,” April 29, 2025 (https://openai.com/index/sycophancy-in-gpt-4o/); and “Expanding on What We Missed with Sycophancy,” May 2, 2025 (https://openai.com/index/expanding-on-sycophancy/).
  11. Haofei Xu, Umar Iqbal, and Jacob M. Montgomery, “Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact,” arXiv:2605.14021, May 2026 (preprint). https://arxiv.org/abs/2605.14021
  12. Profound, “Answer Engine Citation Overlap Strategy: How to Win at AI Visibility” (analysis of 680 million citations; 11.0% ChatGPT/Perplexity domain overlap; accessed August 16, 2026). https://www.tryprofound.com/blog/citation-overlap-strategy
  13. Profound, “AI Search Volatility: Why AI Search Results Keep Changing” (40-60% of cited domains change within one month for identical queries; accessed August 16, 2026). https://www.tryprofound.com/blog/ai-search-volatility
  14. OpenAI, “Shopping with ChatGPT Search,” OpenAI Help Center, feature announced April 28, 2025 (“Product results are selected independently by ChatGPT and are not ads”). https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search
  15. “OpenAI Scales Back Shopping Plans for ChatGPT,” The Information, March 4, 2026 (https://www.theinformation.com/articles/openai-scales-back-shopping-plans-chatgpt); see also Jason Goldberg, “Why OpenAI’s Checkout Retreat Spells Trouble for Its Commerce Strategy,” Forbes, March 10, 2026.
  16. Krystal Scanlon, “OpenAI Brings Product Carousels to ChatGPT Ads,” Digiday, August 6, 2026. https://digiday.com/marketing/openai-brings-product-carousels-to-chatgpt-ads/
  17. Sundar Pichai, “The AI Platform Shift and the Opportunity Ahead for Retail,” remarks at NRF 2026, Google - The Keyword, January 11, 2026. https://blog.google/company-news/inside-google/message-ceo/nrf-2026-remarks/
  18. PayPal Newsroom, “PayPal and Perplexity Launch Instant Buy Ahead of Black Friday,” November 25, 2025 (https://newsroom.paypal-corp.com/2025-11-PayPal-and-Perplexity-Launch-Instant-Buy); merchant count per CNBC coverage.
  19. Sellintu, “Product Data Quality: Your Key to Marketplace Success,” December 29, 2025. https://sellintu.com/blog/product-data-quality-marketplace-success/
  20. Aaron Chatterji, Tom Cunningham, David Deming, Zoe Hitzig, Christopher Ong, Carl Shan, and Kevin Wadman, “How People Use ChatGPT,” NBER Working Paper 34255, September 2025 (OpenAI Economic Research). https://www.nber.org/papers/w34255
  21. Product.ai, “The 2026 Trust in AI Commerce Report” (survey fielded April 2026; n=1,463 U.S. online shoppers, of whom 623 used AI for product research), June 23, 2026. https://product.ai/research/trust-in-ai-commerce-report/
M
Michael Quoc
Founder & CEO

Founded Product.ai (formerly Demand.io) in 2009 with a conviction that commerce data should be verified, not assumed. Built SimplyCodes into the leading coupon verification platform in the US - bootstrapped, profitable, and competing against billion-dollar acquisitions. Leads the company’s AI strategy, the Axiom Distillation Protocol, and the transition from coupon verification to full-spectrum commerce intelligence. Maintains 100% ownership because sovereignty and truth require the same thing: no conflicts of interest.

More from Research