saved
Free the models: harness design at the frontier
by Michele Catasta · X · published
Hraness republishes this public post from a saved copy. The post is the author’s own words.
President & Head of AI @Replit | 🇮🇹 @🇺🇸
Free the models: harness design at the frontier
Model routers are everywhere right now, but they have a fundamental limitation. No matter if based on advanced heuristics or a small model that reads each turn and picks which LLM to use, a router will always be less capable than the model it’s choosing for. Replit Agent lets the model decide instead.
The main agent, or core loop, chooses its subagents’ tier and effort, and adjusts its own as the task unfolds. Given that freedom, GPT-6 Astra hands routine implementation to less costly subagents and decides for itself where its tokens are worth spending. On both DeepSWE and Terminal-Bench, Replit Agent is Pareto-efficient against Astra on its own: no published Astra baseline costs less and scores higher. It also beats a sidekick architecture, the same setup with one long-lived worker, by 11 and 16 points.
For the full post with animations and footnotes, check https://replit.com/blog/free-the-models
Why we scaffold less
Every model release invalidates assumptions baked into the harness.
As models become stronger at long-horizon tasks, they don’t need as much scaffolding at the harness layer. In practice, we’ve observed them lean more towards delegation on their own: using subagents for context management and parallelism. Recent breakthroughs, Navier–Stokes among them, came in part from coordinating swarms of agents powered by frontier models [1].
But the frontier is jagged. The strongest coding model is not necessarily the strongest at designing UIs or making slides, nor the best at writing emails.
So we design our harness to let each model work its own way, with the guardrails it still needs and quality at minimum cost as the goal.Each new model sends us back to re-test what we held firmly, and to experiment fast with techniques that build on emergent behaviors. Freeing the model, then, means letting it decide how hard to think, when to hand work off, and who to hand it to.
Figure 1: the three decisions the core loop makes at every step.
Composable primitives for delegation
When we started experimenting with GPT-6 Astra [2], we found that the model delegates well. The GPT-6 family is also the first from OpenAI to support effort changes mid-turn without breaking the cache.
To use these capabilities, we refined four harness primitives. They give the core loop a small set of choices at each step: what kind of subagent to dispatch, at what size and effort, whether to return to one it has already briefed, and how hard to think:
- Domain-aware subagents. Alongside a general worker, the harness offers specialists: read-only explorers, browser testers, reviewers, and a design subagent for slides and UI, each with its own model and tooling. For now, the harness still decides which specialists exist; the core loop decides when and how to use them.
- Subagent tiers and effort. Small, standard, and large, each a step up in cost and capability, and an effort level within the tier. Both apply to every subagent, and the core loop picks them at each dispatch. For example, a mechanical rename goes to small at low effort, while generating hypotheses for a stubborn bug goes to large at high effort.
- Reusable subagents. The core loop can return to a subagent it has already briefed instead of starting over. There is no single sidekick kept alive for the session: [REDACTED] number of subagents stay warm across kinds and tiers, and it picks which to wake. A longer cache lifetime on OpenAI’s newer models keeps the cost of doing so down.
- Dynamic effort tuning. Now that changing effort mid-turn preserves the cache on some models, we trained an escalation system that checks the trajectory at each step and matches effort to task difficulty. Unlike a router, it acts mid-turn on the work in progress, not once on the request.
The code quality of Astra and Fable 5.1 [3] also let us use our code-review subagent less, with no drop in our eval scores. We’ve not seen this level of engineering quality from any model before.
Newer models delegate on their own
Frontier models like Astra and Fable cost more per token, which makes them look uneconomical next to smaller ones. We’ve observed them naturally delegate to less costly subagents, keeping their own tokens for the decisions that need them.
Replit Agent never forces the core loop to spawn subagents. Table 1 shows how three models handle that decision in production:
Table 1: Delegation in Replit Agent production, each model at medium reasoning effort.
All three models delegate, but each in its own way. At medium effort, Fable models rarely hand work to a general worker: they send out read-only explorers and reviewers and keep the implementation for themselves. Astra is the first model we’ve seen routinely delegate to general workers without being told to, and once it has briefed one it tends to go back to it rather than start over. This return rate has risen with every model generation.
Figure 2: A production turn on September 17, 2026, drawn from the trace. The core loop dispatched an explorer, two workers, and a tester; of its five worker dispatches, three were returns to a worker it had already briefed.
Results
We evaluated Replit Agent in Max mode, our highest-quality setting with Astra as the core loop, on two software engineering benchmarks: DeepSWE and Terminal-Bench. We compare against two baselines: Astra on its own in mini-swe-agent, as published on each leaderboard, and a sidekick architecture, the same configuration with one change: its subagent primitives replaced by a single long-lived worker. Each chart plots score against cost per task, so the most efficient configurations sit toward the top left.
On DeepSWE v1.1 [4], which tests long-horizon changes to active open-source repositories, Replit Agent scores 72% at $2.11 per task. Astra in mini-swe-agent at low effort scores 67% at $1.60, and at xhigh effort 74% at $4.43; the sidekick architecture scores 61% at $1.34. Terminal-Bench 4.0 [5] tests multi-step work done entirely from a shell. Replit Agent reaches 49% at $2.53 per task, against 42% at $2.25 for Astra at low effort and 60% at $5.86 at xhigh. The sidekick architecture manages 33% at $1.84.
Replit Agent beats the sidekick architecture on both benchmarks, by 11 and 16 points. The sidekick costs less, and gives up a sixth to a third of the score for it. Astra on its own scores higher only by spending more: its best settings sit 2 and 11 points above Replit Agent at more than twice the cost. Neither baseline wins on both cost and score. We ran Replit Agent exactly as it ships to users, with no changes to the prompting or harness.
The bitter lesson of harness design
We read these results as an instance of Sutton’s bitter lesson [6]. Baking human knowledge into an agent helps in the short term, plateaus in the long run, and is eventually overtaken by general methods that scale with computation. A rigid harness forces the model into one way of working; a composable one lets it choose. The smarter models get, the less the harness should decide for them.
Compared with a more prescribed architecture, this approach buys us three things:
- It bets on model scaling laws. Delegation that relies on the taste of the model improves with every release. Early previews of next-generation models continue the trend.
- It fits the task. The model spawns nothing for a small task, one explorer for a search, and a team when a build breaks into independent pieces.
- It reuses without persisting. A subagent keeps its context in case the model wants it back, and nothing persists unless it does.
In Sutton’s terms, the harness should let the model discover how to execute the work, not prescribe how we would have done it. Free the models.
Acknowledgements
Written by Daniel Furman, Jacky Zhao, Vaibhav Kumar, Ed Sioufi, and Michele Catasta. Thanks to James Austin, Toby Ho, Preeya Kirani, Zhen Li, Robin Newhouse, Devanshu Sen Pandey, Ibrahim Sheikh, Samuel Spitz, Peter Zhong, and the rest of the AI team at Replit for their contributions to this work. If you want to work on AI at Replit, my team is hiring; reach out to [email protected].
References
- On the Navier–Stokes Millennium Prize Problem
- Introducing GPT-6 Astra
- Introducing Claude Fable 5.1 and Claude Mythos 5.1
- DeepSWE v1.1
- Terminal-Bench 4.0
- The Bitter Lesson
- Idiosyncrasies in Large Language Models
- Design Arena
- Terminal-Bench 4.0 results, Artificial Analysis
Image transcription
Cover image (duplicate of the first inline diagram): “Freeing the model.” The harness offers the options and keeps the guardrails; at every step, the core loop decides.
Image 1: An abstract field of multicolored halftone dots and lines on a white background. No readable text.
Image 2: “Freeing the model.” “The harness offers the options and keeps the guardrails; at every step, the core loop decides.” “1 How hard to think.” “adjusted mid-turn.” Effort scale: “low,” “med,” “high,” “xhigh,” “max.” “Effort is set step by step and, on the latest models, changes mid-turn without a cache miss.” “2 When to hand work off.” “does it itself” and “or hands off, two at once.” “Hand-offs buy context management and parallelism; a small task spawns nothing at all.” “3 Who to hand it to.” Columns “A,” “B,” and “C.” Rows “coding,” “UI design,” “slides,” and “writing.” “The frontier is jagged, so each specialist runs on the model strongest at its job.”
Image 3: An abstract multicolored halftone field on a white background. No readable text.
Image 4: “DeepSWE: score against cost per task.” “Mean of 4 repetitions over 113 tasks. GPT-6 Astra alone in mini-swe-agent, public v1.1 leaderboard.” Chart labels: “80%,” “70%,” “60%,” “50%”; y-axis “Score”; x-axis “Cost per task”; x-axis ticks “$0,” “$2,” “$4,” “$6,” “$8,” “$10”; series labels “Replit Agent,” “Sidekick architecture,” and “GPT-6 Astra,” with effort labels “low,” “med,” “high,” “xhigh,” and “max.” “Terminal-Bench 4.0: score against cost per task.” “Mean of 4 repetitions over 63 tasks, GPU tasks excluded. GPT-6 Astra alone in mini-swe-agent, Artificial Analysis leaderboard.” Chart labels: “70%,” “60%,” “50%,” “40%,” “30%,” “20%”; y-axis “Score”; x-axis “Cost per task”; x-axis ticks “$0,” “$2,” “$4,” “$6,” “$8,” “$10”; series labels “Replit Agent,” “Sidekick architecture,” and “GPT-6 Astra,” with effort labels “low,” “med,” “high,” “xhigh,” and “max.”
Image 5: Table columns “Fable 5,” “Fable 5.1,” and “GPT-6 Astra.” Rows: “Turns that dispatch a subagent” — “32%,” “21%,” “36%.” “Turns that hand work to a general worker” — “0.9%,” “2.3%,” “20%.” “Dispatches that return to an existing subagent” — “17%,” “29%,” “42%.”
Image 6: Sequence diagram with columns “Agent,” “Explorer,” “Worker A,” “Worker B,” and “Tester.” Prompt: “When I add a service from a vendor's page, it never brings me back to the vendor. Can we fix this?” Visible labels include “Confirms the plan,” “Reads the flagged files,” “find where it drops,” “Scans the save flow,” “Reads the handler,” “found the handler,” “fix the return path,” “Tags vendor links,” “update the spec,” “Adds a return helper,” “Drafts the spec,” “spec drafted; needs evidence,” “Runs the suite,” “23 pass; type errors,” “which are real?”, “none; stale types,” “Rebuilds stale types,” “Restarts the app,” “run the journey,” “tests and diff,” “Cancel returns to vendor,” “Failed save keeps form,” “Folds in the diff,” “spec updated; wants evidence,” “Retry lands on vendor, row visible,” “3 of 3 journeys pass,” “browser evidence,” “Screenshots the result,” “spec final,” and “Fixed. Save returns to that vendor's Services & Costs section, showing the new service. Cancel or close returns to the same vendor.” Time markers: “0:00,” “1:43,” and “4:52.” “Replay.”




