Product.ai / Research / We Asked Four AI Engines the Same 200 Shopping Questions. 78% Produced a Factual Conflict.
Practitioner Account

We Asked Four AI Engines the Same 200 Shopping Questions. 78% Produced a Factual Conflict.

ChatGPT, Claude, Gemini, and Perplexity, five runs each, for 4,000 answers. Every headphone question in the set produced a conflict.

Author(s): Sean Fisher · AI Content Strategist, Dakota Nunley · Director of Content
Published: August 10, 2026
Cover image for the Product.ai AI Shopping Divergence Study, showing the four AI shopping engines tested, Gemini, ChatGPT, Claude, and Perplexity, as white logo marks on dark tiles over a halftone-dot field.
Research Artifacts

Key Findings

  • Product.ai found that 156 of 200 questions (78%) produced a confirmed conflict: a contradicted price, a discontinued product presented as current, or a spec claim an engine's own five answers couldn't agree on.
  • Every single headphone question in the study, 23 of 23 (100%), produced a confirmed conflict. No other category came close.
  • Asked for the best noise-cancelling headphones for flying, ChatGPT recommended the Sony WH-1000XM5 in all five test runs. Sony had already replaced that model with the WH-1000XM6. Gemini named the same stale model in three of its five runs, even with live search turned on. Only Claude and Perplexity got it right in every sample.
  • Across two of its five sample runs on "best workout headphones under $150," ChatGPT recommended five separate products that have all been discontinued or replaced by newer models, with no mention that any of them were out of date.
  • ChatGPT and Gemini tie for the highest rate of errors that would cost a shopper money or the wrong product, at 52%. ChatGPT posts that rate on the lowest likely-fabrication rate of the three general-purpose engines, at 18% (vs. 24% for Claude and Gemini).
  • Comparison and recommendation questions produced a confirmed conflict 94% of the time, nearly double the rate of plain spec lookups (52%).
  • Gemini contradicted its own prior answer to the identical question, with nothing new to justify the change, in 30% of questions. That is the highest self-contradiction rate of the four engines.
  • 22% of matched question pairs (2 of 9) showed an engine doing meaningfully more research for one phrasing of a question than another. Both times, it was ChatGPT.

Executive Summary

Product.ai asked four AI engines (ChatGPT, Claude, Gemini, and Perplexity) the same 200 real shopping questions, five times each, across nine product categories: headphones, sunscreen, skincare, supplements, smartphones, laptops, TVs, mattresses, and robot vacuums. That's 4,000 individual answers, checked against each other for confirmed conflicts: a specific, checkable disagreement (not a difference of opinion) that replicates across at least 2 of 5 test runs. 156 of the 200 questions met that bar, and headphones was the single worst category, hitting it on every question asked.

Why We Built This

An AI shopping answer sounds certain even when it's wrong. Ask ChatGPT or Gemini "what's the best noise-cancelling headphones," and you get one confident, specific answer (a model name, a spec, a price) with nothing in the delivery to tell you whether that confidence is earned. We wanted to know how often it actually is.

Hallucination isn't unique to shopping. A Stanford RegLab and Human-Centered AI Institute study found large language models hallucinating on 69-88% of specific legal queries, depending on the model. We wanted our own number, for our own domain, using the questions people actually ask before they buy something instead of a synthetic benchmark.

So we picked 200 real shopping questions across nine categories and locked the list before running a single test. No question was added, removed, or reworded after we saw how the answers came in. Then we asked all four major AI engines (ChatGPT, Claude, Gemini, and Perplexity) the same 200 questions, five times each: enough repetition to tell a real, repeatable pattern from a one-off fluke.

We built a scoring process, reviewed question by question, that only counts something as a confirmed conflict if it's a specific, checkable claim (not a difference of opinion) and it holds up across multiple runs, not just one.

We almost got the Framing-Reversal Check wrong. Our first instinct, reading ChatGPT's different verdicts on the same rivalry asked two ways, was to call it a finding about framing. It took hand-reading the raw answers to catch what was actually going on: ChatGPT hadn't looked anything up for one version of the question, and had for the other. We weren't measuring an opinion changing. We were measuring one answer coming from memory and the other from a live search. That near-miss is why the check below exists as a structural gate now, not something we have to remember to ask by hand each time.

Here's what we found. We're also not finished: two verification steps are still outstanding before we'd call this final, and we're saying so plainly rather than waiting to be asked (see Methodology).

How We Define Our Terms

Two words carry a lot of weight in this study, so we're defining them here instead of assuming they're obvious:

  • A confirmed conflict is a factual disagreement, either between two engines or within one engine's own five answers to the same question, about something specific and checkable: a product name, a spec figure, a price. Not a difference of opinion or writing style. And it has to show up in at least 2 of the 5 test runs on the affected side, so a single odd answer never becomes a headline claim on its own.
  • A likely-fabricated claim is a specific claim (a product name, a spec, a price) that's contradicted by the weight of the other three engines' answers, or that doesn't match a product's actual real-world status (discontinued, superseded, replaced by a newer model). We only use this label when there's a concrete, checkable contradiction to point to, never just a claim that "seems off."

Full scoring detail, including how we handle genuine model/region variants so they don't get mistaken for disagreement, is in Methodology below.

ChatGPT and Gemini Both Recommend a Discontinued Sony Headphone as the Current Top Pick

Across Product.ai's five test runs, ChatGPT told us every time that the Sony WH-1000XM5 is the current top pick for frequent flying, a model Sony has already replaced. Claude and Perplexity both correctly name the WH-1000XM6, the model Sony actually shipped as its current flagship, in every one of their samples. Gemini is the more interesting case: it names the same stale XM5 in three of its five runs, with live search turned on for all three. This isn't a training-data-cutoff problem for Gemini: it had the tools to check and still got it wrong most of the time.

It gets more specific than the model number. Claude cites the WH-1000XM6's battery life as 37 hours and 14 minutes, consistently across all 5 samples, attributed to RTINGS testing. Gemini's fifth sample puts it at "close to 32 hours." That's a five-hour gap on a spec sheet, not a judgment call. The two answers can't both be right.

When Five Stale Products Turn Up Across Two Runs

The Sony case is ChatGPT getting one product wrong. A different question, "what's the best workout headphones under $150," surfaced something worse: across two of ChatGPT's five sample runs on this same question, it recommended five separate products that have all been discontinued or superseded, with no acknowledgment that any of them had aged out.

One run named three:

  • The AfterShokz Aeropex: the brand renamed itself Shokz and replaced this model
  • The Bose SoundSport Wireless: discontinued by Bose entirely
  • The Sony WF-SP800N: a 2020 model with newer Sony sport earbuds since released

A separate run named two more:

  • The Jabra Elite Active 65t: released in 2018, several generations out of date
  • The Beats Powerbeats4: superseded by the Powerbeats Pro 2

None of the other three engines recommended any of these five products for the same question. Five different times, across two separate runs, a shopper following this answer would have been buying end-of-life hardware at full confidence.

Headphones Produced a Confirmed Conflict on Every Single Question: No Other Category Came Close

Divergence wasn't evenly spread. Headphones was the single worst category in this study: 23 of 23 questions (100%) produced a confirmed conflict, meaning at least one engine gave a specific, checkable answer that the others contradicted. Smartphones and supplements followed close behind at 82%. TVs and monitors were the most stable category at 64%, still nearly two-thirds, but the clear outlier on the low end.

CategoryQuestions askedShare with a confirmed conflict
Headphones23100%
Supplements2282%
Smartphones2282%
Robot Vacuums2277%
Mattresses2277%
Sunscreen2374%
Laptops2273%
Skincare / Moisturizer2273%
TVs / Monitors2264%

Product.ai AI Shopping Divergence Study: confirmed factual conflict rate in AI shopping recommendations by product category, across 200 questions and four AI engines. Headphones is highest at 100% (23 of 23). TVs and monitors is lowest at 64%. Seven other categories run 73% to 82% in between.

Headphones and phones are the two categories with the fastest model-refresh cycles in consumer electronics, with new flagships every 12-18 months. That's a plausible explanation, not a proven one: fast-moving categories are exactly where an engine's training-data cutoff shows up as a wrong answer, because "the current model" keeps changing underneath it.

Shoppers trust AI most in this exact category. In Product.ai's companion Trust in AI Commerce Report, electronics ranked as the single most-trusted product category for AI shopping advice, of 18 categories measured (4.70/10, ahead of every other category including apparel's 4.30 low). Headphones, a fast-moving corner of that same electronics category, is where this study found the least evidence behind that confidence.

Within headphones specifically, ChatGPT's answers included a likely-fabricated claim in 39% of the 23 questions, nearly 4 in 10 and the highest rate of any engine in the category. Perplexity had zero.

ChatGPT's Errors Are the Costliest, Even Though They Aren't the Most Frequent

For every one of the 200 questions, we checked whether each engine made at least one likely-fabricated claim (a wrong product name or spec, or a claim contradicted by the other three engines), and separately, whether it made at least one claim that would actually cost a shopper something if they acted on it: money, the wrong product, or a busted budget.

EngineQuestions with at least one likely-fabricated claimQuestions with at least one costly error
Gemini24%52%
Claude24%46%
ChatGPT18%52%
Perplexity5%16%

AI shopping accuracy by engine, comparing likely-fabrication rate against costly-error rate. ChatGPT and Gemini tie for the highest costly-error rate at 52%. ChatGPT has the lowest fabrication rate of the three general-purpose engines at 18%, where Gemini and Claude sit at 24%. Perplexity is lowest on both measures.

ChatGPT ties with Gemini for the highest rate of costly errors, at 52%, despite having the lowest likely-fabrication rate of the three general-purpose engines, at 18% versus 24% for both Claude and Gemini. Put plainly: when ChatGPT is wrong, it tends to be wrong in the way that costs a shopper money or gets them the wrong product, even though it's less likely than Claude or Gemini to simply get a fact wrong in the first place. Perplexity's likely-fabrication rate (5%) is roughly a quarter of the other three's average. One caveat: Perplexity looks up the web on every single answer by design, where the other three engines only sometimes do. That's a structural advantage, not proof Perplexity is inherently more careful, and it's exactly the asymmetry the Framing-Reversal Check below is built to isolate.

The Kind of Question That Breaks AI Consensus

Not every question type is equally risky. Product.ai's 200 questions split into four kinds: pure spec/fact lookups ("what's the battery life on X"), head-to-head comparisons, open recommendations ("what's the best X"), and price questions. The confirmed-conflict rate is dramatically different across them:

Question typeQuestions askedShare with a confirmed conflict
Comparisons ("X vs. Y")6494%
Recommendations ("what's the best X")3694%
Price questions2789%
Spec/fact lookups7352%

Confirmed conflict rate by shopping question type. Product comparisons and open recommendations both hit 94%. Price questions run 89%. Plain spec and fact lookups are lowest at 52%, roughly a coin flip.

A plain factual lookup, like "what's the battery life on the WH-1000XM6," is 8 times more likely to get a clean, agreed-upon answer than asking an AI to compare two products: 48% of spec/fact lookups came back clean, against just 6% of comparisons. The moment a question asks an engine to render a verdict, rather than just retrieve a fact, the odds of a confirmed conflict jump from roughly a coin flip to near-certain.

When AI Contradicts Itself

Every finding so far has been about engines disagreeing with each other. A separate, arguably more basic problem: how often does an engine disagree with its own answer from five minutes ago? Product.ai asked each engine every question five times and checked whether its own five answers agreed with each other on the same specific factual claim, with no new information to justify a change.

EngineContradicted its own prior answer, at least once
Gemini30% of questions
Claude26% of questions
ChatGPT20% of questions
Perplexity8% of questions

How often each AI engine contradicted its own prior shopping answer with no new information. Gemini is highest at 30%. Claude is 26%. ChatGPT is 20%. Perplexity is lowest at 8%.

Nearly a third of the time, Gemini gave itself two different answers to the identical question, asked minutes apart, with no new information in between.

ChatGPT's Apparent Framing Effect Is Really a Search-Effort Gap

We included nine matched pairs of questions: the same underlying rivalry, asked two ways. Once as an open recommendation ("what's the best X"), once as a head-to-head comparison between two named products. If an engine's answer flips between the two framings, that only counts as a fair "framing matters" finding if the engine actually did the same amount of homework both times, not if it looked something up for one version of the question and answered the other one from memory.

22% of the matched pairs (2 of 9) failed that check, and in both cases ChatGPT is the reason:

  • Sony WH-1000XM6 vs. Bose QuietComfort Ultra: ChatGPT looked up live information for 0% of its answers to the open recommendation question, but 100% of its answers to the named comparison question. Any verdict difference between those two questions can't be cleanly attributed to framing, because ChatGPT is answering one from memory and one from a live search.
  • iPhone 17 Pro vs. Samsung Galaxy S26 Ultra: ChatGPT again shows a shift (100% → 60% live lookups), and Gemini shows a smaller one (80% → 60%).

Two data points isn't a trend on their own. But it's a specific, checkable pattern worth watching as more waves of this study run: whichever engine's research effort swings hardest between framings is the one whose "disagreement" is most likely an artifact of effort, not opinion.

What This Means If You're Shopping With AI

Don't treat a single AI's answer as fact when it names a specific product or a specific number. Treat it as a claim to double-check, especially in fast-moving categories like headphones and phones, where the "current" model changes every year.

The fastest way to check it: ask the other three engines the same question. If three of them agree with each other and the fourth doesn't, the fourth one is usually wrong.

This isn't only a Product.ai finding. A UC San Diego study found that shoppers who read an AI-generated review summary said they'd buy the product 84% of the time, versus 52% for shoppers who read the original human reviews, a 32-point jump in stated purchase intent. The same researchers also asked the AI fact-checkable follow-up questions outside the material it was summarizing, and it hallucinated an answer 60% of the time. Confidence and correctness aren't the same thing, and right now shoppers have no reliable way to tell them apart from the outside. That instinct to double-check already exists in the market: Product.ai's companion Trust in AI Commerce Report found that 86% of AI-assisted shoppers verify a recommendation through another source before buying. This study is a concrete reason that instinct is the right one.

Methodology

The question set: 200 pre-registered questions across 9 categories (73 spec/fact, 64 comparison, 36 recommendation, 27 price), including 9 paired questions for framing-reversal detection. Locked before collection began: no question added, removed, or reworded after the fact.

Collection: Each engine answered every question 5 times (4,000 answers total) via its own API: OpenAI's Responses API with web search, Anthropic's Messages API with server-side search, Perplexity's native search API, and Google's Gemini API with search grounding. The exact model versions called were gpt-4o-2024-08-06, claude-sonnet-4-6, gemini-2.5-flash, and Perplexity sonar; every one is recorded per row in the published dataset.

Scoring: A rubric-driven judge model extracted structured claims from all 20 answers to a question together, then classified conflicts through a fixed taxonomy: a product-referent gate (catching genuine SKU/regional/generation differences before they're mistaken for disagreement), cross-engine classification, within-engine stability checks, and a severity tier (called "Tier-1" in the underlying scoring data, which is what this piece calls a confirmed conflict). No finding counts as a confirmed conflict without replicating across at least 2 of 5 samples on the affected side, so a single odd answer doesn't become a headline claim. This kind of cross-engine divergence is an active area of academic study too: a 2025 peer-reviewed study on generative AI as a shopping advisor examined how AI-generated reviews measurably shift consumer purchase decisions.

Judge model disclosure: The scoring judge (Claude) is one of the four engines under measurement. This is a real conflict of interest, not a solved problem. Engine identity was anonymized and re-shuffled per question for the judgment call itself, but writing style and citation habits can still leak identity even with labels swapped. Every confirmed conflict requires an independent human fact-check before publication, with no exceptions, enforced as a hard gate rather than a suggestion. Two verification steps are still outstanding before the topline number above is final: a cross-vendor re-judging pass on a sample of Claude's own answers, and a hand-run check of 15-20 questions against the actual consumer apps (logged out, memory off) to confirm these API-level findings match what a real shopper sees.

Data: 200 questions × 4 engines × 5 samples = 4,000 answers. The full question set, every answer, and a column-by-column codebook are available in the downloadable dataset above.

S
Sean Fisher
AI Content Strategist

Leads content initiatives and develops an overarching AI content strategy. Manages production and oversees content quality with both articles and video.

D
Dakota Nunley
Director of Content

Leads the ontological layer - defining the language and narrative architecture of Product.ai across every surface. Owns the Truth Graph’s editorial voice, thought leadership strategy, and the content engine that establishes category dominance.

More from Research