saved
Language Model "Shape"
Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.
gist
Alex Zhang asks whether language-model shape should change to fit harnesses, not only harnesses to fit decoder-only Transformers. He treats shape as the model contract, cites Jev's prefill-only unit-interval outputs as an alternate tradeoff, and argues agent trajectories need dense local attention plus sparse older history rather than full-context compaction alone. With frontier capability now clearer, he says independent researchers can distill decoder-only models into structured shapes for specific tasks.
ideas
- Harnesses now absorb all shape risk. Since ChatGPT, frontier labs keep the proven autoregressive contract, so agent work mostly designs harnesses around that shape.
- Shape is the input/output contract. Changing it trades efficiency and long-context behavior; fitting the model to the harness can amortize better than endless harness patches.
- Jev shows a constrained output space. Typesafe AI's Jev limits outputs to the unit interval, enabling prefill-only sampling for calibrated fuzzy decisions.
- Agent history wants mixed density. Zhang argues decoder-only Transformers are not the natural final form for trajectories that need dense recent context and sparse older history.
- Distillation makes shape research viable. Open recipes and environments now let researchers distill decoder-only capability into alternate shapes without waiting on frontier lab incentives.
quotes
“whether it’s worth considering changing the shape of the language model to fit the harness.”
“what if we could assume some structure over the input / output shape of the model we are using?”
“I don’t think decoder-only Transformers are the natural or final form of how you’d want to process an agent trajectory”
“Jev, for example, is not meaningfully more intelligent than any frontier model, but will still likely become extraordinarily useful for fuzzy decision tasks.”