hraness
Theme
Appearance

saved

World Modeling in Transformers

by Pierre Beckmann, Matthieu Queloz and André FreitasarXivpublished

Hraness cites a source capture. The source author remains the source.

gist

Beckmann, Queloz, and Freitas revisit TaxiGPT, the Manhattan taxi transformer whose spaghetti-like behavioral maps looked like evidence of no world model. Mechanistic probes and causal edits show a faithful intersection/street map, position tracking, and a goal compass. Failures come from interference among superposed intersection features; affordance packing limits damage. Their indicators shift the question from whether a model “has” a world model to how its world-modeling capacities interact.

ideas

  • Behavior underdetermines the map. Reconstructing a street map from outputs cannot tell an incoherent map from broken localization on a good map.
  • Faithful map plus goal compass. Diff-means intersection features decode ~99.6% of nodes and support teleportation edits; a circular goal compass steers direction without being required for legality.
  • Superposition causes illegal moves. Weak or noisy writes let the wrong intersection feature win; affordance packing groups same-legal-move nodes so many slips stay on-graph.
  • Capacities emerge out of order. Legal-move and compass skills appear before reliable intersection identity, so early competence can look map-like without fine localization.
  • Indicators beat binary tests. Comparing architectures/objectives with mechanistic metrics separates map recovery, localization, and navigation better than behavioral legality alone.

quotes

Behavioral failures can make a transformer appear to lack a world model even when it has learned faithful representations of its environment.

Beckmann, Queloz, and Freitas, opening the abstract.

Affordance packing, which groups representations of intersections with the same legal moves, helps limit the consequences of these errors.

Beckmann, Queloz, and Freitas, naming the protective structure.

shift from asking whether a model has a world model to mechanistically studying its world modeling

Beckmann, Queloz, and Freitas, stating the methodological claim.