saved
Research acceleration: The view inside OpenAI
Hraness cites a source capture. The source author remains the source.
gist
OpenAI reports coding agents now dominate researcher workflows and that it hit its automated research intern goal: systems that complete well defined multi day research tasks under human direction, with an automated AI researcher targeted for March 2028. Metrics show median researchers spending over 600 dollars a day of API priced inference, 3.1 agent workdays per human workday, rising experiment velocity, and higher success on longer tasks that still need steering. The post pairs acceleration with pacing after the Hugging Face incident, including RL pauses and Astra restrictions, and argues public RSI measurement should become a shared norm.
ideas
- The research intern threshold is an operational claim. OpenAI says agents can already run well-scoped multi-day research tasks under human direction, and frames March 2028 as the automated AI researcher target.
- Agent labor now outweighs human labor in measured runtime. By mid-August the median researcher used agents daily at high token spend, and total agent-workdays exceeded human workdays inside research.
- Code and experiment throughput are rising, but bottlenecks move. More commits and experiments correlate with Codex adoption and more compute; least-automatable judgment and compute may become the binding constraints.
- Task mix is climbing the stack, not just Build. Tokens shift into help and monitoring as well as code, while high-level Decide/Design work remains mostly human; longer tasks succeed more often but still need interventions.
- Pacing is part of the product. After the Hugging Face compromise, OpenAI paused RL on deployable models, hardened research environments, and redirected Astra-class compute under stronger controls, treating safety gates as first-class schedule inputs.
quotes
“the median researcher was integrating agents daily into their work, using more than $600 per day”
“the research organization uses 3.1 agent-workdays of effort for every workday of human labor”
“over half of successful 4-8 hour tasks involved 1 or more interventions”
“High-level planning still remains a minimal fraction of agent output tokens”