hraness
Theme
Appearance

saved

The Bitter Lesson for context management

by Rulin ShaoXpublished

Hraness republishes this public post from a saved copy. The post is the author’s own words.

Rulin Shao @RulinShao

PhD @UWNLP, visiting researcher @Meta.

‼The Bitter Lesson for context management: Giving LMs unrestricted control over their context beats human-designed SOTA!

Introducing 🩵Context Language Models (CLMs)🩵

  • Natively manage their own context
  • Treat context as a file
  • Learn policies in CLM weights, no harness

We implement CLMs with context as a file:

  • The live context is mirrored as an editable file that CLMs can manipulate arbitrarily
  • In multi-agent settings (e.g., agent swarms), multiple context files coexist and are managed by CLMs

Figure shows novel behaviors emerge from CLMs.

CLMs treat context management as a model skill and work out of box with existing LMs:

  • outperform Codex summary on 12hr EdgeBench and 24hr AgentWorld
  • beat specialized OpenEvolve on math problems
  • achieve the best performance at lower cost on deep research and coding

Context management is an intrinsic model behavior of CLMs. Therefore, different from harness-controlled strategies, CLMs can be steered and evolved in context.

  • Left: we steer CLMs with one sentence
  • Right: CLMs improve with in-context textual evolution

CLMs can also explore and internalize better context management strategies into weights through online RL.

We propose a success-gated efficiency advantage for stepwise GRPO. It improves Qwen3.5-9B on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs.

Arbitrary context edits challenge prefix-cache reuse in serving systems. We introduce Prefix-Reuse FLOPs (PF) to capture this cost. The above experiments all reported efficiency using RF, showing CLMs remain more efficient through better strategies.

We also develop Suffix Cache Reuse to further reduce PFLOPs without hurting performance.

Alongside this, we present ContextBench, a diagnostic benchmark that isolates context-management capabilities in selective verbatim retention, in-place surgical editing, and offloading & retrieval. ContextBench helps reveal limitations of existing baselines.

Check out our paper, code, and play with CLM!

Paper: https://arxiv.org/abs/2609.37725

Code: https://github.com/facebookresearch/context-language-models

Last thing: we provide Day 1 support for CLM + Pi-agent. Try it with pi install npm:@lolipopshock/pi-clm!

Image transcription

Photo 1. Poster titled "Context Language Models." Authors: Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang, Hamish Ivison, Radha Poovendran, Nathan Lambert, Teng Xiao, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, Pang Wei Koh. Affiliations: University of Washington, Meta Superintelligence Labs, MIT, Trillium Labs. Center diagram: context file C_t → f_θ^CLM ("any function of the context") → context file C_{t+1}; labels "context as a file." (A) Qualitative examples: "Writes loops to prune results" with code for t in turns: if irrelevant(t): t = "No relevant results."; "Defines and reuses functions" with def process_turns(s): ... reused at steps 109 and 142; "Tracks subagents in context" with ## STATE BEST SO FAR 0.9992 AGENTS 21 * 5 running. (B) Zero-shot application: BrowseComp-Plus 32K peak ctx scatter of Accuracy (%) vs Prefix-reuse PFLOPs / question with CLM highest near ~60% around 6 PFLOPs vs Qwen3.6-27B, Summary, MEM1, Self-Compact, RLM, Base, ACM; SoftwareWorld 24hr agent swarm chart of Held-out speedup vs Active hours with CLM swarm 1.044x vs Summary swarm 1.026x and GPT-5.6-Sol baseline. (C) Learn in context: Steer chart "Compact to 4k tokens once you reach Y tokens" with instructed first compaction tracking threshold Y along x=y; Evolve chart Accuracy (%) vs PFLOPs / task with start→later points along a Pareto frontier. (D) Learn in weights: Advantage for stepwise GRPO A_i = A_i^out + w_eff × A_i^eff (task advantage; success-gated efficiency advantage). BrowseComp-Plus Qwen3.5-9B bars Acc. (%) 28.8→42.5 and PFLOPs/Q 1.52→1.34 after RL. Footer: September 29, 2026; correspondence Rulin Shao rulin@cs.washington.edu; Code https://github.com/facebookresearch/context-language-models; Meta logo.

Photo 2. Panel of CLM in-context editing behaviors. (a) "CLM builds in-context scoreboards and trackers to orchestrate and monitor subagents with in-place editing." Snippets write ## STATE — Erdős Minimum Overlap Problem (compact) with LEDGER TOP 0.9992491, AGENTS 21 launched / 5 running, KEY FILES, FINDINGS plateau at 0.381157 (Erdős minimum overlap, step 455); and ## ORCHESTRATOR STATE (compact) Budget 7/100, slots BUSY, dead ends, next steps (Circle packing N = 26, steps 194, 198). (b) "CLM creates a new role alongside the original template roles for its own internal notes." Uses [CTX_TURN 4 role=notes] (BrowseComp-Plus, step 40). (c) "CLM writes for loops to remove past irrelevant search results or compact overly long outputs when creating a new view." Loop replaces turns with "No relevant results." and compacting long bcp_search / bcp_get_document bodies (BrowseComp-Plus, steps 1509 and 13). (d) "CLM defines and reuses a function to conveniently compact old results with a reference to its maintained note." Defines compact_turns and progress note (BrowseComp-Plus, step 109). (e) "CLM reproduces effective behaviors from existing baselines, preserving important facts and future TODOs in the summary." SUMMARY and EXPLORATION LEDGER with Best score 0.9931 and untried ideas (BrowseComp-Plus step 20; Circle packing N = 26, step 332).

Photo 3. Slide "Zero-shot CLMs: better and cheaper." EdgeBench 12hr run charts for Qwen3.6-27B and Claude 4.6 Sonnet: Best score so far vs Wall-clock hours for Base, Summary, CLM, CLMs (subagents). Qwen3.6-27B at 12h: CLMs (subagents) 44.6 / 179 PF; CLM 44.2 / 181 PF; Summary 42.3 / 437 PF; Base inset 21.8 / 3 PF. Claude 4.6 Sonnet: CLM 51.0; CLMs (subagents) 50.4; Summary 42.3; Base inset 24.2. Software World 24hr+ agent swarm: Society and Held-out library tables; GPT-5.6-Sol chart Held-out speedup vs Active hours with CLMs (agent swarm) 1.044 vs Summary (agent swarm) 1.026. Deep Research & Coding scatters Accuracy (%) vs Prefix-reuse PFLOPs / question for (a) BrowseComp-Plus avg ctx 97K, (b) TerminalBench 2.1 avg ctx 47K, (c) TBLite avg ctx 33K, with CLM points highest and leftward of Summary, MEM1, Self-Compact, ACM, Base. Mathematical Optimization table: Method | Circle packing (↑) | Heilbronn (↑) | Min-max/min-dist (↑) | Erdős overlap (↓). OpenEvolve 2.541 / 0.03127 / 0.07690 / 0.38123; OpenEvolve-Agent 2.525 / 0.03053 / 0.07724 / 0.38167; CLM 2.618 / 0.03653 / 0.07758 / 0.38094; CLM (Subagents) 2.636 / 0.03617 / 0.07758 / 0.38109.

Photo 4. Slide "CLMs Can Learn in Context." Left "Steer CLMs with one sentence": instruction "Compact to 4k tokens once you reach Y tokens" with instructed first compaction ≈16.0k, 23.6k, 30.9k vs no instruction ~40k; "Compact at sub-question boundaries" peak 0.77 at turn 1 and 0.88 within 2 turns vs 0.40 without; "Back up on disk before you compact" edits preceded by backup 0.0 without instruction vs 0.68 full +0.09 partial when instructed. Right "Evolve CLMs with textual evolution": assisted evolution Accuracy (%) vs PFLOPs / task moving start→later toward Pareto frontier; self-evolution Accuracy (%) vs USD / task improving from ~94% / ~$0.75 toward ~98% / ~$0.66.

Photo 5. Slide "CLMs Can Learn in Weights." Title "Stepwise GRPO with success-gated efficiency advantage." Formula A_i = A_i^out + w_eff × A_i^eff (task advantage; success-gated efficiency advantage). Legend CLM vs Summary; W/ FLOPs Reward vs W/o FLOPs Reward. (a) BrowseComp-Plus held-out 32K context: Accuracy (%) and PFLOPs / question vs Training step 0–70; CLM rises above 40% while Summary with FLOPs reward collapses near 0. (b) Accuracy ↑ vs FLOPs ↓ trajectories from step 0 to 70 with "better" toward top-left; CLM moves up/left; a red label marks "40 (1%)".

Photo 6. Slide "Suffix Cache Reuse." Diagram compares Standard Serving vs Suffix Cache Reuse after editing middle span B→B' in sequence A B C: both reuse prefix A; standard re-prefills B' and C; Suffix Cache Reuse prefills B' then reuses suffix C. Bars: Task accuracy (%) Standard and Suffix both 60.2%; PFLOPs per question Standard 10.98 vs Suffix 7.14 (Prefill vs Decode stacked); % of prompt tokens for All turns and Edited turns comparing Standard SGLang vs Suffix Cache Reuse with prefix-cache hit, reused by Suffix Cache Reuse, and prefilled shares (All turns Suffix: 73.9% prefix hits, 7.8% suffix reused, 18.3% prefilled; Edited turns Suffix: 29.9% / 28.2% / 41.8%).

Photo 7. ContextBench four-task figure. Columns: Needle Retention (selective verbatim retention) "keep needles"; Sudoku Sketchpad (in-place surgical editing) "edit cells"; KV Store (offloading & retrieval) "GET a7f3"; Log Triage (offloading & retrieval) grep -c ERROR → 2. Each plots Accuracy (0–1) vs Context pressure (log). Legend: Base, RLM, Summary, Context folding, Self-Compact, ACM, CLM (Ours). CLM (Ours) stays near 1.0 across pressures while baselines fall as pressure rises.

Research poster for Context Language Models with architecture diagram and benchmark panels.Five panels of CLM code snippets editing context files, scoreboards, and summaries.Zero-shot CLMs better and cheaper charts across EdgeBench, Software World, deep research, and math.CLMs Can Learn in Context slides for one-sentence steering and textual evolution.CLMs Can Learn in Weights slide for stepwise GRPO with success-gated efficiency advantage.Suffix Cache Reuse diagram and bars comparing accuracy and PFLOPs to standard SGLang.ContextBench accuracy versus context pressure across four diagnostic tasks.