hraness

saved

Introducing Fusion in Devin Desktop & CLI

by CognitionCognitionpublished

Hraness cites a source capture. The source author remains the source.

gist

Cognition ships Fusion in Devin Desktop and CLI: a two-agent harness where a frontier lead plans and reviews while a cheaper sidekick executes, claiming up to 39% lower cost than peer harnesses on Fable and Astra coding benchmarks. Parallel persistent contexts exchange briefs rather than full histories, so the lead stays in charge and can reclaim control. They argue price-per-task beats price-per-token, and that stronger (pricier) leads and sidekicks can make the whole system cheaper.

ideas

  • Lead plus sidekick beats one-shot routing. Difficulty is unknown until exploration, so Fusion keeps frontier intelligence reviewing while a cost-effective model does the work.
  • Briefs preserve cache and control. Agents exchange briefs, results, and feedback instead of whole conversations, with the lead able to take over when the sidekick is out of depth.
  • Price per task, not per token. Stronger leads and sidekicks can cut total session cost by delegating earlier, needing fewer correction rounds, and using tokens more efficiently.
  • Harness tuning is pair-specific. How detailed briefs are, whether the sidekick may push back, and what exploration is delegated all depend on the lead–sidekick pairing.
  • Recommended pairing: Fable 5.1 with SWE-2. Evaluations with Artificial Analysis and Vals AI show large cost cuts while scores stay near the frontier models alone.

quotes

Fusion works around the common pitfalls of routing with a key idea: running two parallel agents

Cognition, stating the Fusion architecture premise.

models (and model-harness combos) should be evaluated on price per task rather than price per token

Cognition, arguing for price-per-task evaluation.

using more expensive models can make the entire system cheaper

Cognition, on counterintuitive Fusion cost findings.