saved
Which tools do Claude Code, Codex and Cursor choose? We measured 16,893 sessions to find out.
Hraness cites a source capture. The source author remains the source.
gist
Armature ran 16,893 coding-agent sessions across Claude Code, Codex, and Cursor on 75 synthetic repositories with 1,163 prompt variants and four personas, then published 5,292 valid sessions. Agents often disagree on third-party tools: they agree in only 42% of cells, diverge by language and repo context, and mention incumbents far more often than they install them. Docs wording, bundled pricing, and a simulated human in the loop flip winners; traces and a sector leaderboard are public.
ideas
- Agents disagree on installs. Claude Code, Codex, and Cursor pick the same tool in only 42% of cells and differ on when they search the web versus use priors.
- Repository context steers winners. The same email or deploy ask yields different providers by language and stack; Vercel wins on TypeScript/Next.js while Render dominates Python.
- Mentions are not wins. PayPal, LangChain, Netlify, and Supabase appear often yet lose to Stripe, Neon, and leaner alternatives in the published sessions.
- Docs and packaging flip choice. Retention notes, BaaS bundling, cost framing, and a simulated human authorizing a vendor reduce in-house and cloud-native defaults.
- Method is measurable. Ephemeral sandboxes, persona prompts, fake lockfiles, and a Gemini judge over conversation plus code diffs make agent tool choice auditable.
quotes
“Cursor bases its decision on the web in 2/3 of the sessions.”
“All three agents pick the same tool in only 42% of the cells”
“In the payment service provider sector, Paypal is cited 139 times and never picked (Stripe won 124 of these 139 sessions).”
“With the exact same ask on 4 repositories in 4 different programming languages, we got 4 different email provider winners”