saved
Language Models Can Control Their Own Attention
Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.
gist
Namgyu Ho and coauthors introduce Declarative Attention, a protocol that lets models declare which context to attend while decoding. Across 15 long-context tasks, DA cuts attended decode tokens by 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B, with accuracy drops of 1.27pp and 2.75pp that shrink as models scale. The mask comes from parseable chain-of-thought tags, not an extrinsic O(N) scorer, and runs zero-shot on off-the-shelf models.
ideas
- Declarative Attention makes the model name its attention span. Generation partitions into global (full context), focus (named ~2K-token magic chunks), and local (response-only) modes that the inference engine parses like tool calls.
- Savings come from the mask, not shorter generation. Ablations show DA-nm (same prompt, full attention) matches vanilla accuracy but attends more tokens; the dynamic mask drives reductions up to 71.1% versus that ablation.
- Cost falls while accuracy stays close. On Gemma-4-31B and Qwen-3.6-27B, average attended tokens drop 52.0% and 31.1% with 1.27pp and 2.75pp accuracy drops across 15 tasks.
- Scale closes the accuracy gap. From 4B to 31B, focus-call parsing succeeds more often and the accuracy gap to vanilla shrinks; absolute token savings grow with context length, up to about 21M tokens per response.
- vLLM can apply the mask without new kernels. Block-aligned KV masking keeps FlashAttention unchanged; roofline analysis projects decode wall-clock at 0.71× and 0.77× of vanilla on the two headline models when serving is well optimized.
quotes
“wouldn’t the model already know which parts of the context are relevant?”
“We extend this principle from dictating what to think to where to attend.”
“DA needs no auxiliary scorer and runs on off-the-shelf models with no training.”
“turning selective attention from a pattern inferred inside the network into one the model states out loud”