hraness
Theme
Appearance

saved

Language Models Can Control Their Own Attention

by Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster and Cicero Nogueira dos SantosarXivpublished

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Namgyu Ho and coauthors introduce Declarative Attention, a protocol that lets models declare which context to attend while decoding. Across 15 long-context tasks, DA cuts attended decode tokens by 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B, with accuracy drops of 1.27pp and 2.75pp that shrink as models scale. The mask comes from parseable chain-of-thought tags, not an extrinsic O(N) scorer, and runs zero-shot on off-the-shelf models.

ideas

  • Declarative Attention makes the model name its attention span. Generation partitions into global (full context), focus (named ~2K-token magic chunks), and local (response-only) modes that the inference engine parses like tool calls.
  • Savings come from the mask, not shorter generation. Ablations show DA-nm (same prompt, full attention) matches vanilla accuracy but attends more tokens; the dynamic mask drives reductions up to 71.1% versus that ablation.
  • Cost falls while accuracy stays close. On Gemma-4-31B and Qwen-3.6-27B, average attended tokens drop 52.0% and 31.1% with 1.27pp and 2.75pp accuracy drops across 15 tasks.
  • Scale closes the accuracy gap. From 4B to 31B, focus-call parsing succeeds more often and the accuracy gap to vanilla shrinks; absolute token savings grow with context length, up to about 21M tokens per response.
  • vLLM can apply the mask without new kernels. Block-aligned KV masking keeps FlashAttention unchanged; roofline analysis projects decode wall-clock at 0.71× and 0.77× of vanilla on the two headline models when serving is well optimized.

quotes

“wouldn’t the model already know which parts of the context are relevant?”

Namgyu Ho et al.

“We extend this principle from dictating what to think to where to attend.”

Namgyu Ho et al.

“DA needs no auxiliary scorer and runs on off-the-shelf models with no training.”

Namgyu Ho et al.

“turning selective attention from a pattern inferred inside the network into one the model states out loud”

Namgyu Ho et al.