16 hrs ago
OpenAI Says AI Agents Now Supply More Research Work-Hours
OpenAI says its researchers are using AI coding agents for many hours every day.
By August, the agents were doing the equivalent of more than three workdays for each human researcher's workday.
Researchers use the agents to write code, run experiments, find problems, and study results.
The typical researcher’s use of these tools grew sharply during the year.
OpenAI also said the amount of work produced by a typical research employee rose greatly since December.
More researchers are running several agents at the same time.
The company said agents are becoming better at longer tasks, but people still need to help them often.
OpenAI cautioned that more code or experiments do not automatically mean better research.
By mid-August, OpenAI's research organization used the equivalent of 3.1 agent-workdays for every human researcher workday.
The organization crossed that threshold only after June, as coding-agent usage accelerated throughout the year.
The median researcher was generating more than $600 in daily inference usage at API prices, while the top 10% exceeded $7,000.
OpenAI reported a 124-fold increase in output for the typical research employee since December, alongside more code and experiments.
Agents are handling longer and harder tasks more often, but jobs estimated at four to eight human hours still required intervention more than half the time.
- Who
- OpenAI's research organization and its human researchers using AI coding agents.
- What
- The organization now relies on AI coding agents for more measured work-hours than human researchers contribute directly.
- Where
- Inside OpenAI's research organization.
- When
- The comparison was reported as of mid-August, with the threshold crossed since June and usage tracked from January onward.
- Why
- Agents are being used to accelerate research activities such as coding, experimentation, debugging, training jobs, analysis, and communication.
OpenAI's Productivity Case
OpenAI's Cautions
Research speed
OpenAI's Productivity Case
OpenAI says agents are speeding up coding, experimentation, debugging, training, analysis, and other parts of research.
OpenAI's Cautions
The company cautions that increased code volume and experiment counts do not necessarily translate directly into genuine research progress.
Agent capability
OpenAI's Productivity Case
Internal tracking indicates that agents are succeeding more often on harder and longer-duration tasks than they did at the start of the year.
OpenAI's Cautions
Tasks estimated to take a human four to eight hours still required a human to intervene more than half the time.
Limits on growth
OpenAI's Productivity Case
Wider adoption of OpenAI's Codex tool has accompanied several-fold growth in code shipped per active contributor and record experiment activity.
OpenAI's Cautions
OpenAI says compute availability, rather than agent capability alone, may increasingly determine how quickly experiments can run.
Key facts
- Agent workload
- 3.1 agent-workdays for every human researcher workday as of mid-August
- Usage threshold
- The 3.1-to-1 level was reached only since June
- Median daily inference cost
- More than $600 at API pricing by mid-August
- Top-user daily inference cost
- The highest-use 10% exceeded $7,000 at API pricing
- Reported output growth
- Output for the typical research employee rose 124-fold since December
- Experiment activity
- Experiments per active experimenter reached an all-time high in August
- Human intervention
- Tasks estimated at four to eight human hours still required intervention more than half the time







