返回新聞
創新AI Understanding 簡報

Martin Cid 報告了 OpenAI 的 Astra 有爭議的 AGI 聲明

Martin Cid 雜誌報告稱,OpenAI 將 Astra 定義為 AGI,而基準測試的參考評估得分為 62.7%,遠低於 OpenAI 報告的 99.9% 的結果。該帳戶尚未得到所提供來源的獨立確認。

4 min readRead the linked source
Source-provided image accompanying Martin Cid reports a disputed AGI claim for OpenAI’s Astra
來源參考來源記錄
出版商
martincid.com
來源連結
martincid.comhttps://www.martincid.com/technology-sv/jensen-huang-agi-claim-benchmark-gap/
來源類型
連結來源-主要來源狀態尚未確定。
背景60 秒內了解這一點

從這裡開始

關鍵術語

AGI(通用人工智慧)
一個假設的人工智慧系統,可以在許多領域以人類層級執行大多數智力任務。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
測試一下自己什麼是人工智慧?測驗

發生了什麼事

Martin Cid Magazine reports that OpenAI and Nvidia CEO Jensen Huang publicly framed OpenAI’s Astra model as evidence that artificial general intelligence has arrived. The outlet says OpenAI reported a 99.9% score on ARC-AGI-3, while the organization’s reference harness produced 62.7% for the same model. The source does not provide independent confirmation of either result or Astra’s availability.

Martin Cid Magazine reports that Jensen Huang posted on September 6 that “AGI has arrived” after a progression from ChatGPT to o1 to Astra. The outlet says OpenAI president Greg Brockman had used similar “AGI era” language three days earlier. These statements are attributed to the named outlet’s account; the supplied material does not independently verify the posts or provide links to primary statements.

The central reported discrepancy concerns ARC-AGI-3. According to Martin Cid Magazine, OpenAI said Astra scored 99.9% on the , alongside a reported 98% score on FrontierMath Tier 4. The outlet says the ARC-AGI-3 organization’s standard evaluation harness instead produced a 62.7% score. Martin Cid attributes the difference to distinct evaluation protocols and says the benchmark creators regard their reference condition as the valid basis for comparisons. The supplied source does not independently confirm the scores, test conditions, or methodology.

The article argues that 62.7% would still represent a strong result, but says the move from a high score to an “AGI has arrived” declaration requires an agreed definition of AGI. It reports that Astra was not yet visible in Artificial Analysis’s Intelligence Index or Arena.ai rankings as of publication, while noting that absence from those systems is not proof of poor performance. These are the outlet’s reported observations and interpretations, not independently established findings in the supplied material.

來源詳情: martincid.com ↗

為什麼這很重要

The reported gap illustrates why results cannot be separated from evaluation protocols, and why a company’s definition of AGI may not be accepted outside the company. A widely circulated AGI declaration based on an internal evaluation could influence investment, infrastructure spending, policy debates and public expectations even when external validation is incomplete. The source also places Huang’s statement in the commercial context of Nvidia supplying chips used to train Astra, while providing no evidence of wrongdoing.

A score is meaningful only in relation to the test version, harness, access conditions and scoring rules used to produce it. The reported 37.2-percentage-point difference between OpenAI’s result and the benchmark organization’s reference result therefore matters independently of whether either number is ultimately revised. Readers should not treat the higher figure as a field-wide measurement without comparable external testing.

The report also exposes a definitional problem. AGI is not governed by one universally accepted threshold, so a company can reach an internal milestone without establishing that the broader research community would recognize the same milestone. Martin Cid presents the declarations as potentially consequential market messaging, especially because it reports that Astra was trained on Nvidia hardware and that Huang publicly congratulated OpenAI. The source offers no evidence that the statements were intentionally misleading.

If the account is accurate, the practical implication is that customers, investors and policymakers should distinguish product capability claims from independently reproducible evidence. The supplied source does not establish Astra’s real-world reliability, general availability, pricing, safety performance or usefulness across the broad range of tasks implied by AGI.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
What is AI? Quiz

A route planner searches possible journeys using explicit rules. What does this illustrate about AI?

接下來看什麼

Watch for the ARC-AGI-3 creators to publish or confirm a reproducible reference-harness result, for OpenAI to disclose its evaluation methodology, and for independent assessments of Astra to appear. The supplied source does not establish whether Astra is generally available, restricted to selected customers, or offered at a particular price. It also does not establish a field-wide definition of AGI.

The most important next evidence would be a public ARC-AGI-3 evaluation under the creators’ reference harness, with enough methodological detail for independent reproduction. A response from OpenAI explaining why its internal harness produced a materially different result would also clarify whether the gap reflects prompting, tool access, task selection, scoring, or another protocol difference.

Independent ranking and testing systems could provide additional evidence about Astra’s capabilities, but their absence at the article’s publication date should not be treated as a negative result. The source does not say when those systems might evaluate Astra or whether they have access to the model.

The source identifies no confirmed access route or price. It also does not establish whether OpenAI has formally changed its product status, documentation or policy based on the AGI language. Those practical details remain unknown and should be verified before describing Astra as generally available or AGI by consensus.

相關指引和測驗

什麼是人工智慧?人工智慧模型解釋變形金剛測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?