What happened
Product.ai published a study evaluating the accuracy of ChatGPT, Claude, Gemini, and for shopping tasks. The research found that 86% of 220 test questions resulted in repeatable factual conflicts, such as incorrect prices or specifications. While Perplexity showed the lowest error rate, all models struggled with consistency, and incorrect prices had a median deviation of $300.
Product.ai, a startup that verifies product claims, conducted a study testing the free and paid versions of ChatGPT, Claude, Gemini, and . The team used 220 shopping questions covering products like laptops, TVs, and robot vacuums, running each question five times per service to capture 8,794 responses.
The study found that 86% of the questions produced a repeatable factual conflict, defined as a checkable disagreement in price, model, or specification across multiple responses. Head-to-head comparison questions had a higher conflict rate of 97%, compared to 75% for straightforward factual queries.
Accuracy varied by model and tier. Gemini had the highest share of costly errors at 56% on its free tier and 54% on its paid tier. Claude’s paid version reduced costly errors to 21% from 44% on the free tier. recorded the lowest costly-error rate at 14% on its paid tier, followed by ChatGPT’s paid version at 17%.
When prices were incorrect, the median difference was $300. Product.ai’s head of search product, Dakota Nunley, noted that these errors are significant enough to warrant caution. He advised users to treat AI as a discovery tool rather than a definitive source for transactional data, recommending that shoppers verify information on seller websites or by comparing multiple AI outputs.
Source details: businessinsider.com ↗
Why it matters
This study provides concrete evidence that current AI models are not yet reliable enough for autonomous purchasing decisions. With AI agents increasingly integrated into retail ecosystems, the high rate of factual conflicts poses a direct risk to consumer financial safety. The findings suggest that users must continue to manually verify product details and prices before making purchases, limiting the practical utility of 'hands-off' AI shopping agents in the near term.
The study highlights a critical gap between the marketing of AI shopping agents and their actual reliability. As retailers and tech companies push for AI-driven commerce, the prevalence of factual conflicts undermines consumer trust and poses financial risks.
The findings suggest that 'hands-off' AI agents are not yet viable for high-stakes purchasing decisions. The median $300 price error indicates that relying solely on AI for price verification can lead to significant financial loss or missed deals.
This research provides a for the industry, showing that even leading models struggle with basic factual consistency in dynamic retail environments. It underscores the need for better integration of real-time, verified data sources in AI shopping tools.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
crm_get_transaction(id='4092').What most distinguishes an AI agent from a basic chatbot?
What to watch next
Monitor how retailers and AI developers respond to these accuracy benchmarks, particularly regarding the integration of real-time data verification in shopping agents. Watch for updates from , which claimed industry-leading accuracy, and observe if other providers implement stricter fact-checking protocols to reduce costly errors in consumer-facing applications.
Watch for responses from OpenAI, Anthropic, Google, and regarding these specific accuracy metrics, as they may influence future model updates or product features.
Monitor the development of AI shopping agents by retailers like Amazon and Walmart, which may need to implement additional verification layers to mitigate the risks identified in the study.
Observe if Product.ai or similar verification services become standard integrations in AI shopping platforms to address the accuracy gaps highlighted in the research.