Back to News
IndustryAI Understanding briefing

OpenAI’s internal benchmark reveals Astra can code experiments, framing its 80% AGI claim

A TIME interview disclosed that OpenAI’s chief scientist Jakub Pachocki described a concrete internal benchmark – a system called Astra that writes, runs, and reports on experimental code – which underpins the company’s oft‑cited “80 percent to AGI” statement.

4 min readRead the linked source
Source-provided image accompanying OpenAI’s internal benchmark reveals Astra can code experiments, framing its 80% AGI claim
Source referenceSource recorded
Publisher
fourweekmba.com
Source link
fourweekmba.comhttps://fourweekmba.com/ai-openai-agi-benchmark-denominator-80-percent-claim/
Source type
Linked source — primary-source status has not been established.
ContextUnderstand this in 60 seconds

Start here

Key terms

AGI (Artificial General Intelligence)
A hypothetical AI system that can perform most intellectual tasks at a human level across many domains.
Benchmark
A standardized test or dataset used to measure and compare model performance.
Classification
A task where a model assigns an input to one or more predefined categories.
Test yourselfAI Models Explained Quiz

What happened

In a September 26, 2026 article, FourWeekMBA reproduced reporting by TIME journalist Alex Heath, who spent two weeks interviewing more than twenty OpenAI executives, investors and rivals. The piece reveals that Chief Research Officer Mark Chen told TIME the company is “80 percent of the way” to artificial general intelligence (AGI). Chief Scientist Jakub Pachocki clarified the denominator for that figure: an internal system named Astra that can take an experimental idea, generate the necessary code in OpenAI’s codebase, execute the experiment, and return the results. Sam Altman’s public phrasing was also captured – he said OpenAI expects to have “an internal system he would call AGI” by the end of the year, emphasizing that the label is his own . The article notes that the remaining 20 percent, the exact metric used, and any formal definition of AGI remain unspecified.

FourWeekMBA’s September 26 piece draws on Alex Heath’s TIME report, which is based on more than two weeks of interviews with over twenty OpenAI insiders, including executives, investors and rivals. The reporting captures three distinct statements: (1) Mark Chen’s estimate that OpenAI is “80 percent of the way” to AGI; (2) Sam Altman’s qualification that the company will have an internal system he would call AGI by year‑end; and (3) Jakub Pachocki’s description of a concrete internal – a system called Astra that can take an experimental idea, write the corresponding code in OpenAI’s codebase, run the experiment, and report the results.

The article emphasizes that the is described in a single, falsifiable sentence, making it a durable object that can be examined by outsiders, unlike the ambiguous percentage. However, the piece also notes that the exact composition of the remaining 20 percent, the methodology for arriving at the 80 percent figure, and any formal internal definition of AGI were not disclosed.

No additional technical details, scores, or release timelines for Astra were provided, and OpenAI has not published a formal definition of AGI or a public version of the benchmark.

Source details: fourweekmba.com ↗

Why it matters

The disclosure provides the first publicly documented, falsifiable for OpenAI’s internal progress toward AGI, a claim that has previously been expressed only as a vague percentage. By anchoring the 80 percent figure to a concrete task – automated code generation, execution, and reporting – analysts gain a tangible reference point to evaluate future statements from OpenAI. This is unusual in a field where many firms cite high‑level milestones without offering measurable criteria, making the benchmark a rare, verifiable artifact that could be examined by external researchers if OpenAI chooses to share more details. Moreover, the article highlights how media transmission can strip nuance from technical claims, underscoring the need for precise reporting when assessing AI capabilities that may have broad economic and safety implications.

The gives analysts a concrete yardstick to track OpenAI’s progress, which is rare in a sector where many capability claims are abstract. This could enable more rigorous external evaluation of OpenAI’s roadmap and help differentiate substantive advances from marketing hype.

By exposing how a numeric claim can lose its denominator through media chains, the article illustrates a broader communication challenge for AI governance: policymakers and the public need precise, verifiable metrics to assess claims about transformative technologies like AGI.

The disclosure may pressure OpenAI and other AI firms to provide clearer, more transparent metrics for future milestones, potentially shaping industry norms around reporting AI capability.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Model Parameter Size:8B Parameters
VRAM Required5.5 GBGPU memory footprint
Target HardwareMacBook / Single GPUDeployment tier
Privacy100% Air-GappedLocal device capability
Core takeaway: Small, quantized models (3B–8B) now run directly inside smartphones and laptops with complete data privacy, while mammoth 400B+ models remain the domain of datacenter clusters.
Interactive Concept Check+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

What to watch next

Future updates from OpenAI should be monitored for any expansion of the Astra , such as broader task coverage, performance metrics, or public release of evaluation data. Observers should also watch for how the company’s leadership continues to frame the “80 percent to AGI” narrative in investor briefings, regulatory filings, or product announcements, especially any clarification of the missing 20 percent. Finally, the reaction of competitors and policymakers to this newly visible internal metric could influence industry standards for transparency around AGI‑related milestones.

Any future OpenAI statements that expand on Astra’s capabilities, such as broader experimental domains, performance benchmarks, or public release of evaluation data.

Clarifications from OpenAI leadership about the missing 20 percent and whether the 80 percent figure reflects a formal internal metric or an informal estimate.

Regulatory or competitive responses that reference this as a standard for assessing AGI‑related progress.

Related guides & quizzes

AI Models ExplainedAI EthicsFuture of AITest what you know — try a free AI quizLook up an AI term in our glossaryFollow the AI funding tracker
Found this useful?