Back to News
InnovationAI Understanding briefing

Martin Cid reports a disputed AGI claim for OpenAI’s Astra

Martin Cid Magazine reports that OpenAI framed Astra as AGI while the benchmark’s reference evaluation scored it 62.7%, far below OpenAI’s reported 99.9% result. The account has not been independently confirmed from the supplied source.

4 min readRead the primary source
Source-provided image accompanying Martin Cid reports a disputed AGI claim for OpenAI’s Astra
Source referenceSource recorded
Publisher
martincid.com
Source link
martincid.comhttps://www.martincid.com/technology-sv/jensen-huang-agi-claim-benchmark-gap/
Source type
Linked source — primary-source status has not been established.

Story last revised

ContextUnderstand this in 60 seconds

Start here

Key terms

AGI (Artificial General Intelligence)
A hypothetical AI system that can perform most intellectual tasks at a human level across many domains.
Benchmark
A standardized test or dataset used to measure and compare model performance.
Test yourselfWhat is AI? Quiz

What happened

Martin Cid Magazine reports that OpenAI and Nvidia CEO Jensen Huang publicly framed OpenAI’s Astra model as evidence that artificial general intelligence has arrived. The outlet says OpenAI reported a 99.9% score on ARC-AGI-3, while the benchmark organization’s reference harness produced 62.7% for the same model. The source does not provide independent confirmation of either result or Astra’s availability.

Martin Cid Magazine reports that Jensen Huang posted on September 6 that “AGI has arrived” after a progression from ChatGPT to o1 to Astra. The outlet says OpenAI president Greg Brockman had used similar “AGI era” language three days earlier. These statements are attributed to the named outlet’s account; the supplied material does not independently verify the posts or provide links to primary statements.

The central reported discrepancy concerns ARC-AGI-3. According to Martin Cid Magazine, OpenAI said Astra scored 99.9% on the benchmark, alongside a reported 98% score on FrontierMath Tier 4. The outlet says the ARC-AGI-3 organization’s standard evaluation harness instead produced a 62.7% score. Martin Cid attributes the difference to distinct evaluation protocols and says the benchmark creators regard their reference condition as the valid basis for comparisons. The supplied source does not independently confirm the scores, test conditions, or methodology.

The article argues that 62.7% would still represent a strong result, but says the move from a high benchmark score to an “AGI has arrived” declaration requires an agreed definition of AGI. It reports that Astra was not yet visible in Artificial Analysis’s Intelligence Index or Arena.ai rankings as of publication, while noting that absence from those systems is not proof of poor performance. These are the outlet’s reported observations and interpretations, not independently established findings in the supplied material.

Source details: martincid.com

Why it matters

The reported gap illustrates why benchmark results cannot be separated from evaluation protocols, and why a company’s definition of AGI may not be accepted outside the company. A widely circulated AGI declaration based on an internal evaluation could influence investment, infrastructure spending, policy debates and public expectations even when external validation is incomplete. The source also places Huang’s statement in the commercial context of Nvidia supplying chips used to train Astra, while providing no evidence of wrongdoing.

A benchmark score is meaningful only in relation to the test version, harness, access conditions and scoring rules used to produce it. The reported 37.2-percentage-point difference between OpenAI’s result and the benchmark organization’s reference result therefore matters independently of whether either number is ultimately revised. Readers should not treat the higher figure as a field-wide measurement without comparable external testing.

The report also exposes a definitional problem. AGI is not governed by one universally accepted threshold, so a company can reach an internal milestone without establishing that the broader research community would recognize the same milestone. Martin Cid presents the declarations as potentially consequential market messaging, especially because it reports that Astra was trained on Nvidia hardware and that Huang publicly congratulated OpenAI. The source offers no evidence that the statements were intentionally misleading.

If the account is accurate, the practical implication is that customers, investors and policymakers should distinguish product capability claims from independently reproducible evidence. The supplied source does not establish Astra’s real-world reliability, general availability, pricing, safety performance or usefulness across the broad range of tasks implied by AGI.

What to watch next

Watch for the ARC-AGI-3 creators to publish or confirm a reproducible reference-harness result, for OpenAI to disclose its evaluation methodology, and for independent assessments of Astra to appear. The supplied source does not establish whether Astra is generally available, restricted to selected customers, or offered at a particular price. It also does not establish a field-wide definition of AGI.

The most important next evidence would be a public ARC-AGI-3 evaluation under the benchmark creators’ reference harness, with enough methodological detail for independent reproduction. A response from OpenAI explaining why its internal harness produced a materially different result would also clarify whether the gap reflects prompting, tool access, task selection, scoring, or another protocol difference.

Independent ranking and testing systems could provide additional evidence about Astra’s capabilities, but their absence at the article’s publication date should not be treated as a negative result. The source does not say when those systems might evaluate Astra or whether they have access to the model.

The source identifies no confirmed access route or price. It also does not establish whether OpenAI has formally changed its product status, documentation or policy based on the AGI language. Those practical details remain unknown and should be verified before describing Astra as generally available or AGI by consensus.

Related guides & quizzes

What is AI?AI Models ExplainedTransformersTest what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?