Back to News
InnovationAI Understanding briefing

Study finds high error rates in AI shopping assistants

A new study by Product.ai reveals that 86% of shopping queries to major AI models produced conflicting factual answers, highlighting significant reliability issues for holiday shoppers.

4 min readRead the original reporting
Source-provided image accompanying Study finds high error rates in AI shopping assistants
Attributed reportingSource recorded
Publisher
businessinsider.com
Source link
businessinsider.comhttps://www.businessinsider.com/study-shows-ai-errors-shopping-tools-struggle-accuracy-2026-9
Source type
Reporting by a news outlet — not a first-party document.

What we could not confirm independently: This claim is attributed to the named outlet. We did not verify it against a first-party document. (businessinsider.com)

ContextUnderstand this in 60 seconds

Start here

Key terms

Perplexity
A language-model metric measuring how surprised the model is by true next tokens.
Benchmark
A standardized test or dataset used to measure and compare model performance.
Test yourselfAI Agents Quiz

What happened

Product.ai published a study evaluating the accuracy of ChatGPT, Claude, Gemini, and for shopping tasks. The research found that 86% of 220 test questions resulted in repeatable factual conflicts, such as incorrect prices or specifications. While Perplexity showed the lowest error rate, all models struggled with consistency, and incorrect prices had a median deviation of $300.

Product.ai, a startup that verifies product claims, conducted a study testing the free and paid versions of ChatGPT, Claude, Gemini, and . The team used 220 shopping questions covering products like laptops, TVs, and robot vacuums, running each question five times per service to capture 8,794 responses.

The study found that 86% of the questions produced a repeatable factual conflict, defined as a checkable disagreement in price, model, or specification across multiple responses. Head-to-head comparison questions had a higher conflict rate of 97%, compared to 75% for straightforward factual queries.

Accuracy varied by model and tier. Gemini had the highest share of costly errors at 56% on its free tier and 54% on its paid tier. Claude’s paid version reduced costly errors to 21% from 44% on the free tier. recorded the lowest costly-error rate at 14% on its paid tier, followed by ChatGPT’s paid version at 17%.

When prices were incorrect, the median difference was $300. Product.ai’s head of search product, Dakota Nunley, noted that these errors are significant enough to warrant caution. He advised users to treat AI as a discovery tool rather than a definitive source for transactional data, recommending that shoppers verify information on seller websites or by comparing multiple AI outputs.

Source details: businessinsider.com ↗

Why it matters

This study provides concrete evidence that current AI models are not yet reliable enough for autonomous purchasing decisions. With AI agents increasingly integrated into retail ecosystems, the high rate of factual conflicts poses a direct risk to consumer financial safety. The findings suggest that users must continue to manually verify product details and prices before making purchases, limiting the practical utility of 'hands-off' AI shopping agents in the near term.

The study highlights a critical gap between the marketing of AI shopping agents and their actual reliability. As retailers and tech companies push for AI-driven commerce, the prevalence of factual conflicts undermines consumer trust and poses financial risks.

The findings suggest that 'hands-off' AI agents are not yet viable for high-stakes purchasing decisions. The median $300 price error indicates that relying solely on AI for price verification can lead to significant financial loss or missed deals.

This research provides a for the industry, showing that even leading models struggle with basic factual consistency in dynamic retail environments. It underscores the need for better integration of real-time, verified data sources in AI shopping tools.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Interactive Concept Check+10 Points
AI Agents Quiz

What most distinguishes an AI agent from a basic chatbot?

What to watch next

Monitor how retailers and AI developers respond to these accuracy benchmarks, particularly regarding the integration of real-time data verification in shopping agents. Watch for updates from , which claimed industry-leading accuracy, and observe if other providers implement stricter fact-checking protocols to reduce costly errors in consumer-facing applications.

Watch for responses from OpenAI, Anthropic, Google, and regarding these specific accuracy metrics, as they may influence future model updates or product features.

Monitor the development of AI shopping agents by retailers like Amazon and Walmart, which may need to implement additional verification layers to mitigate the risks identified in the study.

Observe if Product.ai or similar verification services become standard integrations in AI shopping platforms to address the accuracy gaps highlighted in the research.

Related guides & quizzes

AI AgentsAI Models ExplainedAI EthicsTest what you know — try a free AI quizLook up an AI term in our glossaryFollow the AI model release tracker
Found this useful?