Back to News
InnovationAI Understanding briefing

OpenAI says coding agents now exceed human research labor in its labs

OpenAI reports that its researchers used 3.1 agent-workdays for every human workday by mid-August, while cautioning that the internal metrics are preliminary and do not directly measure overall research progress.

5 min readRead the primary source
Source-provided image accompanying OpenAI says coding agents now exceed human research labor in its labs
Primary-source documentSource recorded
Publisher
openai.com
Source link
openai.comhttps://openai.com/index/research-acceleration-view-inside-openai
Source type
Primary document — an official announcement, paper, filing, or first-party page we read directly.

Story last revised

ContextUnderstand this in 60 seconds

Start here

Key terms

API (Application Programming Interface)
A structured way for one software system to send requests to and receive responses from another system.
Benchmark
A standardized test or dataset used to measure and compare model performance.
Inference
The runtime phase where a trained model generates predictions or outputs.
Test yourselfAI Agents Quiz

What happened

OpenAI published an internal snapshot of how coding agents are being used by its research organization. The company says it has met its stated goal of an “automated research intern”: a system able to perform well-defined research tasks under human direction, including tasks that could take a skilled researcher several days. OpenAI says it is targeting an automated AI researcher by March 2028. According to the company’s measurements, agent use increased sharply during 2026. By mid-August, the median researcher was using coding agents daily at more than $600 per day in inference calculated at API prices, while the 90th-percentile user exceeded $7,000 per day. Across the research organization, OpenAI says agent runtime reached 3.1 standard eight-hour agent-workdays for every human workday. OpenAI also reports more experiments per active experimenter, growing use of agents for higher-level and longer-horizon tasks, and generally improving success rates from January through July on tasks with a known outcome. However, more than half of successful tasks estimated to require four to eight hours of human work involved at least one human intervention during the previous six months. The report describes these as OpenAI’s own preliminary measurements and internal impressions. It does not establish that agents caused all observed increases in research output: the company says available compute also grew significantly, and that bottlenecks such as compute, task selection, evaluation, safety checks and integration into core training runs can limit overall progress.

OpenAI says its researchers now use coding agents throughout the day, often in concurrent sessions, and that usage is growing faster than in other OpenAI teams. The company reports that the median researcher moved from modest use at the start of 2026 to daily use by mid-August. It gives internal spending estimates of more than $600 per day in inference at API prices for the median researcher and more than $7,000 per day for the 90th-percentile user. These are not prices offered to the public for an automated researcher.

The company measures total agent runtime against human labor using standard eight-hour workdays. It says that before June 2026, total agent runtime was below total human labor, but by mid-August the research organization used 3.1 agent-workdays for every human workday. The measurement includes agents launched directly by researchers and downstream subagents.

OpenAI says agent activity expanded across six stages of its AI research and development taxonomy: deciding, designing, building, running, analyzing and communicating. Research and infrastructure code remained the dominant category, while technical help and monitoring runs grew notably. High-level planning remained a small fraction of agent output tokens. OpenAI says agents have been particularly useful for troubleshooting internal research infrastructure, coinciding with lower attendance at some human-run technical-support sessions.

The report says coding-agent task success generally increased from January to July across several estimated human-time buckets, but agents still needed substantial steering as tasks became more complex. More than half of successful four-to-eight-hour tasks involved one or more interventions in the previous six months. OpenAI also says it paused reinforcement-learning training on some latest deployment models after a recent infrastructure incident, restored some workloads under stronger restrictions, and imposed additional restrictions on Astra-class experiments after preliminary evidence of critical cyber capabilities. It reports that Astra-class GPU allocation fell 59.2% in the following week while other model classes rose 17.2%, offsetting about 85% of the decline.

Source details: openai.com

Why it matters

The report offers an unusually concrete view into how a frontier AI lab is incorporating AI systems into the work of developing more capable AI. If the trend generalizes, coding agents could shift researchers’ time away from routine implementation and infrastructure troubleshooting toward tasks that remain harder to automate, potentially shortening parts of the model-development cycle. But the evidence is company-generated, preliminary and focused on selected operational metrics rather than independently verified gains in scientific quality or model capability.

The report matters because the AI systems are being used on the process that produces future AI systems, rather than merely assisting with routine office work. OpenAI’s figures suggest that agent capacity is becoming a meaningful part of internal research operations, with concurrent execution and higher task complexity changing how researchers allocate their time.

The practical implication is mixed. Agents may reduce delays caused by coding, infrastructure debugging and experiment setup, allowing researchers to run more trials. At the same time, faster execution can make evaluation, monitoring and safety review more important because errors may be produced and propagated more quickly. OpenAI itself says progress in specific metrics will not necessarily equal the same pace of overall research progress.

The report also makes human control a central condition. OpenAI says people still set priorities, judge which ideas and results to pursue, and decide whether to scale, pause or deploy systems. It acknowledges that alignment and safety progress may not keep pace with capability progress and that more capable systems may become harder to monitor.

No independent study, public benchmark or outside audit is supplied in the source. The reported spending, runtime, task-success and intervention figures therefore should be treated as OpenAI’s claims about its own organization, not as established evidence that AI research has been accelerated by a specific measured amount.

What to watch next

The key question is whether higher agent usage translates into durable, independently measurable improvements in research quality, safety and model development speed. Watch for clearer definitions of task success, external validation, evidence about errors and hidden human labor, and whether the remaining bottlenecks move toward compute, evaluation, alignment or decision-making. OpenAI has not documented public access to these internal workflows or any associated price, and the report does not show that a general-purpose automated researcher is available.

OpenAI says its measurement methods are still evolving. Future disclosures should clarify how tasks are sampled, how success is verified, how uncertain outcomes are handled, and how agent-generated code or experiments are distinguished from work that would otherwise have been done by humans.

The most important missing evidence is whether agent-assisted work produces better research outcomes, not just more code, experiments or tokens. Useful follow-up evidence would include reproducible examples, error rates, rework, failed experiments, safety findings and the amount of human review required at each stage.

The report does not state that the automated research intern is a publicly accessible product. It also does not provide a consumer or enterprise price, deployment requirements, eligibility criteria or release date. Access and pricing are therefore unknown.

OpenAI says it will continue reporting progress toward automated AI research while protecting security and proprietary information. Its future disclosures will show how much of the reported acceleration can be independently assessed and how research restrictions affect the distribution of available compute.

Related guides & quizzes

AI AgentsAI TrainingAI Models ExplainedFuture of AITest what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?