OpenAI announced it reached a milestone set last autumn: researchers now deploy what the company terms an automated research intern, an agent capable of handling clearly specified tasks that would typically require a researcher several days to complete. Data reveals coding-agent usage rose steadily through 2026, with agents recording 3.1 agent-workdays for every human workday across the research organization by mid-August. The report examines how agents are reshaping research workflows — and where human oversight remains the limiting factor.
The median researcher spent more than $600 daily on inference at API pricing, while those at the 90th percentile exceeded $7,000 per day. OpenAI tracked agent activity across six categories — Decide, Design, Build, Run, Analyze, and Communicate — finding increases in all areas between January and August, though agents contributed relatively little to decisions about which research to pursue. Compute resources grew substantially as experiment volumes climbed. When the company evaluated agent performance using another model to assess success rates on tasks of varying difficulty, it found that humans still intervened in more than half of successful tasks that would have taken a person four to eight hours, despite improving success rates between January and July.
According to the report, an agent-workday and a human workday aren't equivalent. The company translates the time agents spend on tasks into standard eight-hour workdays, but because researchers can operate multiple agents simultaneously, the metric shows how long agents are active without necessarily indicating what they're accomplishing. OpenAI's definition of a research intern specifies it must complete well-defined research tasks that would take a skilled person several days, with a human still directing the work. The company's next target, an automated AI researcher, is one they aim to achieve by March 2028.
Much of the agent work is practical: writing research and infrastructure code, monitoring experiments, and delivering enough technical support that OpenAI reports attendance at debugging office hours has dropped, prompting one team to discontinue the sessions entirely. But more agent hours don't automatically yield more useful research. The report notes code output and experiment counts are relatively straightforward to measure, yet neither reveals how much progress agents actually made. Once engineers can operate several agents at once — with those agents launching subagents of their own — the challenge shifts to monitoring what they produce: catching runs that drift off course, reviewing code changes, and deciding what's ready to deploy or feed into a training run. Astra's persistent-agent features already allow researchers to delegate multi-day assignments, which intensifies this supervisory burden rather than easing it. The company acknowledges that as agents assume more execution work, the aspects of research hardest to automate will consume a greater share of an engineer's time, creating a practical ceiling on how much agent output one person can realistically oversee.
On July 20, a series of outages triggered by agents disrupted OpenAI's research infrastructure severely enough that the company took its training container service offline and later restored it with stricter controls. Nearly a month later on August 7, OpenAI tightened access again after early evidence suggested Astra could meet the "Critical" cybersecurity threshold in its Preparedness Framework, limiting the model to higher-security research zones and adding safeguards that developers may already be experiencing as unexpected API interruptions. Astra-class GPU allocation dropped 59.2% the following week, but that compute didn't remain idle. Researchers shifted much of the workload to other models, which saw GPU allocation rise 17.2% and compensated for roughly 85% of the decline in Astra usage. Rather than reducing the volume of work being executed, the restrictions pushed it to alternative models, demonstrating how readily workloads can migrate when one component of the system is restricted. OpenAI's researchers are delegating larger assignments to agents, operating more simultaneously and launching more experiments, but whether that translates into faster research is harder to quantify — and OpenAI is still determining how to price it. The infrastructure can absorb massive volumes of agent work, but the real constraint is human attention: supervision, not computation, now defines the ceiling for AI-assisted research. If agents continue to scale faster than oversight capacity, organizations may find themselves drowning in output they can't adequately review or trust.

