hraness

saved

The Emergent Symbolic Structure of Artificial Neural Networks

by R. Thomas McCoy, Paul Soulos, Tal Linzen and Paul SmolenskyarXivpublished

gist

McCoy, Soulos, Linzen, and Smolensky propose that neural-network vectors implicitly realize symbolic structure. Replacing a network's representation process with a closed-form symbolic equation leaves behavior largely unchanged in small list-manipulation models and in LLMs doing arithmetic, logic, code, and language. Precise interventions on the recovered structures then change those LLMs in the predicted ways, offering a reconciliation of symbolic intelligence with vector-based AI.

ideas

  • Vectors can still be symbols. Neural networks that store information as continuous vectors may nonetheless implement structured combinations of symbols inside those vectors.
  • A closed-form swap is the test. If replacing a network's representation-generating process with a symbolic equation leaves behavior largely unchanged, the network was already using that structure.
  • The result is not limited to toy nets. The same approximation holds for small list-manipulation models and for LLMs in arithmetic, logic, computer code, and language.
  • Causal edits show reliance. Changing the recovered internal structures changes LLM behavior in targeted, predicted ways rather than merely correlating with it.
  • This is a reconciliation. Symbolic accounts of intelligence remain useful because vector systems appear to implement them rather than replace them.

quotes

perhaps the internal representations of neural networks implicitly realize symbolic structure.

R. Thomas McCoy and coauthors, stating the core hypothesis.

we can replace the network's entire representation-generating process with a closed-form equation instantiating a symbolic structure, and the network's behavior remains largely unchanged.

R. Thomas McCoy and coauthors, describing the substitution test.

the LLM's behavior is reliant on the symbolic structures we have identified.

R. Thomas McCoy and coauthors, interpreting the intervention result.

This work provides a potential way to reconcile longstanding symbolic conceptions of intelligence with the vector-based nature of modern AI.

R. Thomas McCoy and coauthors, stating the broader claim.