- Product.ai found that 187 of 217 scored questions (86%) produced a confirmed conflict: a contradicted price, a stale product presented as current, or a spec claim engines couldn’t agree on. Each one is specific, checkable, and repeats across runs, never a one-off.
- When an engine got a price wrong, it wasn’t wrong by a little. Across 960 price answers checked against the sellers’ own pages the same day, 85% of the 913 we could verify were exactly right, and the answers that missed, missed by a median of $300. There was almost no middle ground between exactly right and badly wrong.
- Gemini has the highest rate of errors that would cost a shopper money or the wrong product, 56% of questions on the free tier, and paying doesn’t fix it. The paid tier runs 54%. Paying helps Claude a lot (44% free down to 21% paid). Perplexity posts the lowest rates on both tiers.
- Comparison questions produced a confirmed conflict 97% of the time, against 75% for plain spec lookups. A plain factual lookup is about 8 times more likely to come back clean than a head-to-head comparison.
- Asked about noise-cancelling headphones for flying, engines can’t agree what the current Sony flagship is. Gemini’s free tier brought up the superseded WH-1000XM5 in all five of its runs; the current XM6 has been Sony’s flagship for over a year.
- Gemini contradicted its own prior answer to the identical question, with nothing new to justify the change, in 29% of questions. That is the highest of the four engines, on both tiers.
- Gemini’s free tier did less live research when a question named two rivals head-to-head than when it asked for the best pick. On four of the eight matched question pairs we could check (sunscreen, laptop, TV, and mattress), Gemini’s free tier searched on most or all of its answers to the open “what’s the best” phrasing and on fewer, on two pairs none, of its answers to the head-to-head phrasing of the same rivalry; on the TV pair, both Gemini tiers searched 5 of 5 times for one phrasing and 0 of 5 for the other. Claude’s free tier showed a smaller gap on a fifth pair (moisturizer). ChatGPT searched on every answer to every pair. Five of the eight pairs failed the same-homework check, so no verdict difference on those pairs counts as a framing effect.
Key Findings
Executive Summary
Product.ai asked four AI engines (ChatGPT, Claude, Gemini, and Perplexity), on both their free and paid tiers, the same 220 real shopping questions, five times each, across nine product categories: headphones, sunscreen, skincare, supplements, smartphones, laptops, TVs, mattresses, and robot vacuums. That’s 8,794 captured answers, of which 8,680 (the 217 complete question groups) were checked against each other for confirmed conflicts, meaning a specific, checkable disagreement (not a difference of opinion) that replicates across at least 2 of 5 test runs. 187 of the 217 questions we could score met that bar, 86%. For 24 products, we also captured the seller’s own price the same day and scored every price answer against ground truth. Of the answers we could verify, the engines were exactly right 85% of the time. The rest of the time, the median miss was $300.
Why We Built This
An AI shopping answer sounds certain even when it’s wrong. Ask ChatGPT or Gemini “what’s the best noise-cancelling headphones,” and you get one confident, specific answer (a model name, a spec, a price) with nothing in the delivery to tell you whether that confidence is warranted. We wanted to know how often it actually is.
Hallucination isn’t unique to shopping. A Stanford RegLab and Human-Centered AI Institute study found large language models hallucinating on 69-88% of specific legal queries, depending on the model. We wanted our own number, for our own domain, using the questions people actually ask before they buy something instead of a synthetic benchmark.
So we picked 220 real shopping questions across nine categories, including a price battery built so every price answer could be checked against a seller’s own page the same day, and locked the list before running a single test. No question was added, removed, or reworded after we saw how the answers came in. Then we asked all four major AI engines the same questions on both their free and paid tiers, five times each, enough repetition to tell a real, repeatable pattern from a one-off fluke.
We built a scoring process, reviewed question by question, that only counts something as a confirmed conflict if it’s a specific, checkable claim (not a difference of opinion) and it holds up across multiple runs, not just one.
How We Define Our Terms
Three terms recur throughout this study, so we define them here rather than assume they’re obvious:
- A confirmed conflict is a factual disagreement, either between two engine configurations or within one configuration’s own five answers to the same question, about something specific and checkable: a product name, a spec figure, a price. Not a difference of opinion or writing style. And it has to show up in at least 2 of the 5 test runs on the affected side, so a single odd answer never becomes a headline claim on its own. For the record, looser cuts of the data run higher, with 95% of questions showing at least one checkable factual disagreement in some run, and every question showing at least some disagreement if differences of opinion are counted. We headline the strictest number.
- A likely-fabricated claim is a specific claim (a product name, a spec, a price) that’s contradicted by the weight of the other configurations’ answers, or that doesn’t match a product’s actual real-world status (discontinued, superseded, replaced by a newer model). We only use this label when there’s a concrete, checkable contradiction to point to, never just a claim that “seems off.”
- A wrong price is measured against ground truth rather than against other engines. Ground truth is the price a shopper would actually pay on the seller’s own page, captured by hand within 24 hours of the engine’s answer. An answer matching either the current selling price or the crossed-out list price counts as right (both are on the page). Where the question itself names no single product to price (7 of the 31 price questions, such as “a 65-inch LG C-series OLED” with no model year), it sits outside this check by construction (see Methodology); at the answer level, 47 answers of the 960 checked (4.9%) could not be verified and are excluded from the rate.
Full scoring detail, including how we handle genuine model/region variants so they don’t get mistaken for disagreement, is in Methodology below.
The Stale Sony Pick
Sony superseded the WH-1000XM5 with the WH-1000XM6 more than a year ago. Ask the engines for the best noise-cancelling headphones for frequent flying, and the XM5 comes back as a current pick.
Gemini’s free tier brought it up in all five of its runs on the flying-headphones question; our scoring flagged answers treating the XM5 and XM6 as near-equivalent current picks rather than a superseded model and its successor. “Stale” here does not mean vaporware. Sony still sells the XM5 at $198, less than half the XM6’s price. The problem is presenting last year’s flagship as if the succession never happened, without telling the shopper a newer model exists or that the old one is now the budget pick. The information that would change the purchase decision is the part that goes missing.
A Battery Spec That Can’t Be Two Numbers
Spec-sheet disagreements replicated cleanly enough to measure. On the flying-headphones question, engine configurations split on the Bose QuietComfort Ultra’s battery life: 24 hours against 30 hours, a split that showed up repeatedly across five different configurations’ answers. Both figures come from a real Bose spec sheet: 24 hours is the original model, 30 hours is the second generation Bose released last year. But the answers quoting 24 hours named the QuietComfort Ultra with no generation attached, while the answers quoting 30 hours mostly named the second generation. Each number is stated with full confidence. A shopper who gets the 24-hour answer is being told about the pair Bose replaced, presented as the current one.
Prices split the same way. On the Sony WH-1000XM6, some configurations quoted $398 and others $459.99. Neither number is invented. Sony charges $398 today, and $459.99 is the list price it’s crossed out from. The engines aren’t hallucinating; they’re reading different lines of the same price tag and presenting each as the price. That distinction, selling price versus list price, is why the price layer of this study scores an answer as right if it matches either line, and still found 15% of price answers matching neither.
Every Category Conflicts on at Least 79% of Questions
Every category produces a confirmed conflict on at least 79% of its questions. Laptops, mattresses, and TVs sit at the top of the table at 92%; skincare and supplements, the lowest, are at 79%.
| Category | Questions asked | Share with a confirmed conflict |
|---|---|---|
| Laptops | 24 | 92% |
| Mattresses | 24 | 92% |
| TVs / Monitors | 25 | 92% |
| Headphones | 25 | 88% |
| Sunscreen | 25 | 88% |
| Robot Vacuums | 23 | 87% |
| Smartphones | 22 | 82% |
| Skincare / Moisturizer | 24 | 79% |
| Supplements | 24 | 79% |
217 of the 220 questions asked were scored. Three were excluded for incomplete answer sets (see Methodology), and one scored question, a games-console price check, falls outside these nine categories, so the rows sum to 216.
One question from the price battery belongs to a tenth bucket (a standalone consumer-electronics price probe) and is excluded from this table rather than padding any category; it’s disclosed in Methodology.
The narrow spread is a bigger result than any single row. Slow-moving categories are not safer. Laptops and TVs, with multi-year product cycles, sit at the top of the table.
Shoppers trust AI most in electronics. In Product.ai’s companion Trust in AI Commerce Report, electronics ranked as the single most-trusted product category for AI shopping advice, of 18 categories measured (4.70/10, ahead of every other category including apparel’s 4.30 low). Headphones, laptops, and TVs, the heart of that most-trusted category, run 88-92% here.
Gemini’s Errors Are the Costliest, and Paying Doesn’t Fix Them
For every scored question, we checked whether each engine configuration made at least one likely-fabricated claim (a wrong product name or spec, or a claim contradicted by the weight of the other configurations), and separately, whether it made at least one claim that would cost a shopper something if they acted on it: money, the wrong product, or a busted budget.
Every engine ran on both its free tier and its paid tier. If paying for AI bought accuracy, it would show up here.
| Engine · tier | Questions with at least one likely-fabricated claim | Questions with at least one costly error |
|---|---|---|
| Gemini · free | 20% | 56% |
| Gemini · paid | 21% | 54% |
| Claude · free | 22% | 44% |
| Claude · paid | 12% | 21% |
| ChatGPT · free | 6% | 19% |
| ChatGPT · paid | 7% | 17% |
| Perplexity · free | 4% | 15% |
| Perplexity · paid | 3% | 14% |
Three things stand out.
Gemini owns the costly-error column, and money doesn’t buy its way out. It runs 56% on the free tier and 54% on the paid one. More than half of all questions, on either tier, produced at least one Gemini claim that would cost a shopper money or leave them with the wrong product.
Paying helps Claude more than anyone. Costly errors drop from 44% to 21%, and likely fabrications halve, from 22% to 12%. Claude is the one engine where the paid tier is a meaningfully different product on these measures.
Perplexity posts the lowest rates on both tiers, with ChatGPT close behind. One caveat applies. Perplexity looks up the web on every single answer by design, a structural advantage rather than proof it’s inherently more careful. The other engines searched on nearly every answer too, which keeps the comparison fair.
The Kind of Question That Breaks AI Consensus
Not every question type is equally risky. Product.ai’s questions split into four kinds: pure spec/fact lookups (“what’s the battery life on X”), head-to-head comparisons, open recommendations (“what’s the best X”), and price questions. The confirmed-conflict rate separates them:
| Question type | Questions asked | Share with a confirmed conflict |
|---|---|---|
| Comparisons (“X vs. Y”) | 63 | 97% |
| Recommendations (“what’s the best X”) | 35 | 89% |
| Price questions | 47 | 87% |
| Spec/fact lookups | 72 | 75% |
217 questions scored of the 220 asked. The 47 here counts price questions; it is unrelated to the 47 price answers, out of 960, that could not be verified in the price layer.
The moment a question asks an engine to render a verdict rather than retrieve a fact, conflict becomes near-certain. A plain factual lookup is roughly 8 times more likely to come back clean than a comparison (25% of spec lookups came back conflict-free, against 3% of comparisons). Even so, spec lookups conflicted 75% of the time.
When AI Contradicts Itself
Every finding so far has been about engines disagreeing with each other. A separate and more basic problem is how often an engine disagrees with its own answer from five minutes ago. Product.ai asked each configuration every question five times and checked whether its own five answers agreed with each other on the same specific factual claim, with no new information to justify a change.
| Engine · tier | Contradicted its own prior answer, at least once |
|---|---|
| Gemini · free | 29% of questions |
| Gemini · paid | 27% of questions |
| Claude · free | 26% of questions |
| ChatGPT · free | 22% of questions |
| ChatGPT · paid | 18% of questions |
| Claude · paid | 17% of questions |
| Perplexity · free | 17% of questions |
| Perplexity · paid | 13% of questions |
Nearly a third of the time, Gemini’s free tier gave itself two different answers to the identical question, asked minutes apart, with no new information in between. And self-contradiction is the one column where nobody gets under 13%: even the steadiest configuration in the study changes its own factual story at least one question in eight.
We Also Checked the Prices Against the Stores
Everything above measures engines against each other. The price layer measures them against reality. For 24 products where a single SKU and seller resolve cleanly, we captured the price a shopper would actually pay on the seller’s own page, by hand and within 24 hours of the engines answering, and scored all 960 price answers against it.
Of the 913 answers we could verify against a page, the engines were exactly right 85% of the time. The median error is $0.00, however the 47 unverifiable answers are counted.
The other 15% weren’t close at all. Among the answers that missed, the median miss was $300, and one in ten missed by $500 or more. There is almost no middle ground. An AI price answer is either the number on the seller’s page or it’s off by hundreds of dollars, stated with the same confidence either way.
| Engine · tier | Price answers exactly right | Stale prices quoted |
|---|---|---|
| Perplexity · free | 90% | 0 |
| Claude · paid | 88% | 1 |
| Claude · free | 87% | 0 |
| ChatGPT · free | 85% | 4 |
| Perplexity · paid | 85% | 0 |
| Gemini · free | 84% | 3 |
| ChatGPT · paid | 82% | 4 |
| Gemini · paid | 81% | 1 |
We built deliberate tests into the price battery, products whose prices had recently changed, where an engine leaning on training data instead of a live check would quote a dead number. The test that measured recency caught almost nobody. The Nintendo Switch 2’s price changed the very morning of our collection, and 39 of 40 answers already had the new number. The test that measured habit caught thirteen. Every stale price in the study was the 13-inch MacBook Air quoted at $1,099, a price Apple retired months ago after years at that number. Eight of the thirteen were ChatGPT’s. Live search has largely solved yesterday’s price change. It has not solved a price that stayed the same for years.
One more note on scoring. Sale prices and list prices are different numbers on the same page, and engines routinely pick different lines of the tag (the $398-vs-$459.99 Sony split above). We scored a match to either line as correct. Even with that generosity, 15% of price answers matched neither.
Gemini Searches Less When a Question Names Two Rivals
Nine matched pairs of questions asked the same rivalry two ways, as an open recommendation (“what’s the best X”) and as a named head-to-head (“X vs. Y”). The pairs test whether the wording of a question changes the verdict. A structural gate only counts a verdict flip as a framing effect if the engine did the same amount of live research on both sides, because an answer from memory and an answer from a live search are not comparable.
The gate caught Gemini. ChatGPT searched on every answer to every pair, on both tiers. Gemini did not. On four of the eight pairs we could score (one pair rides a question excluded for an incomplete answer set; see Methodology), Gemini’s free tier searched on most or all of its answers to the open recommendation phrasing and on fewer of its answers to the head-to-head phrasing, on two pairs none of the same rivalry: 4 of 5 against 0 of 5 on the sunscreen pair, 5 of 5 against 1 of 5 on the laptop pair, 5 of 5 against 3 of 5 on the mattress pair, and on the TV pair 5 of 5 against 0 of 5 on both the free and the paid tier. Claude’s free tier showed a smaller gap on a fifth, separate pair, the moisturizer questions: 3 of 5 against 1 of 5. In all, five of the eight pairs failed the same-homework check. That is what the gate is for. On those pairs, any difference between the two verdicts cannot be attributed to the question’s wording, because the engine did different homework for each. So we make no framing claim. The finding is the homework gap itself. The engines search on nearly every shopping answer, and the one that skips the search on a head-to-head question is Gemini.
What This Means If You’re Shopping With AI
Don’t treat a single AI’s answer as fact when it names a specific product or a specific number. Treat it as a claim to double-check, in every category now, not just the fast-moving ones.
When it names a price, check the seller’s page before you plan around it. The price layer of this study says an AI price answer is either exactly right or wrong by hundreds of dollars, with nothing in the delivery to tell you which one you got.
The fastest way to check a product claim is to ask the other engines the same question. If three of them agree with each other and the fourth doesn’t, the fourth one is usually wrong.
This isn’t only a Product.ai finding. A UC San Diego study found that shoppers who read an AI-generated review summary said they’d buy the product 84% of the time, versus 52% for shoppers who read the original human reviews, a 32-point jump in stated purchase intent. The same researchers also asked the AI fact-checkable follow-up questions outside the material it was summarizing, and it hallucinated an answer 60% of the time. Confidence and correctness aren’t the same thing, and right now shoppers have no reliable way to tell them apart from the outside. That instinct to double-check already exists in the market. Product.ai’s companion Trust in AI Commerce Report found that 86% of AI-assisted shoppers verify a recommendation through another source before buying. This study is a concrete reason that instinct is the right one.
Methodology
The question set: 220 pre-registered questions across 9 categories (73 spec/fact, 64 comparison, 36 recommendation, 47 price, including 9 paired questions for framing-reversal detection), locked and version-pinned before the first engine call. No question was added, removed, or reworded after the fact. (The question-type table above shows scored counts; the three excluded questions are one spec lookup, one comparison, and one recommendation.)
Collection: Each of four engines ran on two tiers, its free/consumer configuration and its paid configuration, for eight configurations total. Every configuration answered every question 5 times via its own API (8,800 answers fired; 8,794 captured), on 2026-09-01: OpenAI’s Responses API with web search (gpt-5.6-luna free / gpt-5.6-sol paid), Anthropic’s Messages API with server-side search (claude-sonnet-5 free / claude-opus-5 paid), Google’s Gemini API with search grounding (gemini-3.6-flash free / gemini-3.1-pro-preview paid), and Perplexity’s agent API on both tiers. Every model identity is verified per row in the published dataset by a post-collection audit that fails the run if any answer came from a model other than the one pinned.
What’s excluded, and why: Three of the 220 questions are excluded from conflict scoring because brief provider outages during collection left them short of the complete 8-configuration × 5-sample answer set the scoring requires (one question lost four samples to a capacity outage; two lost one sample each to timeouts). Their captured answers remain in the published dataset but contribute to no scored rate. All conflict rates in this piece use the 217 fully-scored questions as the denominator, and nothing is extrapolated to 220. One further question from the price battery (a standalone consumer-electronics price probe) belongs to no category and is excluded from the category table only.
Scoring: A rubric-driven judge model extracted structured claims from all 40 answers to a question together, then classified conflicts through a fixed taxonomy: a product-referent gate (catching genuine SKU/regional/generation differences before they’re mistaken for disagreement), cross-engine classification, within-engine stability checks, and a severity tier. No finding counts as a confirmed conflict without a checkable factual claim that replicates across at least 2 of 5 samples on the affected side, so a single odd answer doesn’t become a headline claim. This kind of cross-engine divergence is an active area of academic study too. A 2025 peer-reviewed study on generative AI as a shopping advisor examined how AI-generated reviews measurably shift consumer purchase decisions.
The price layer: For the 24 price-battery products where one SKU and one seller resolve cleanly, ground truth was captured by hand from the seller’s own page within 24 hours of collection, both the price a shopper pays and the list price it’s discounted from, each with a URL and timestamp. An answer matching either number (within a small tolerance for rounding) scores correct; conditional prices (trade-in, bundle, member pricing) never count as the price. Seven of the 31 price questions had no single reference price to grade against. Five asked about a product family with no single model and two named a product for which no single price could be resolved that day, so those seven sit outside the price check by construction. The 85% is computed on verified answers only, so it is identical however those seven are counted; counted as unverifiable, they would raise the share of unverifiable price answers from 4.9% of 960 to 26% of 1,240, which is why we state the scope of the check here rather than let a single rate stand for it.
Judge model disclosure: The scoring judge (Claude) is one of the engines under measurement. This is a real conflict of interest, not a solved problem. Engine identity was anonymized and re-shuffled per question for the judgment call itself, but writing style and citation habits can still leak identity even with labels swapped. Two mitigations ran: every confirmed conflict quoted in this piece requires an independent human fact-check before publication, with no exceptions, and a cross-vendor spot-check re-judged a sample of question groups with a non-Anthropic judge. On the conflict pairs both judges flagged, their classifications agreed 81% of the time; the outside judge flagged more conflicts than ours overall, which means the numbers above sit on the conservative side of the two reads. These are API-level findings. A hand-run check against the consumer apps (logged out, memory off) was outside the scope of this study.
Data: 220 questions × 4 engines × 2 tiers × 5 samples = 8,800 answers fired, 8,794 captured, 8,680 conflict-scored across the 217 complete questions. The full question set, every answer, and a column-by-column codebook are available in the downloadable dataset above.