hraness

saved

GPT-6 Astra, Looped Transformers, and Hidden Reasoning

by Sebastian RaschkaAhead of AIpublished

Hraness cites a source capture. The source author remains the source.

gist

Sebastian Raschka walks through GPT-6 Astra impressions, then explains looped transformers: weight-tied depth reuse from Nanbeige and Universal Transformers through Mixture-of-Recursions and SMELT. He judges Astra likely uses some looping, but credits training recipes more than the architecture rumor. Looped depth is not what hides chains of thought; shorter traces track capability, and Pachocki ties monitorability fragility to other causes. Computer use is the leap he expects to keep refining.

ideas

  • Astra's leap is computer use, not just scores. Raschka finds Astra the best model he has used so far, especially on 3D rendering, animation, and GUI demos. Independent Artificial Analysis indexes put it at the frontier without a huge gap; shared-harness evals may still understate its primary-harness strength. He also notes OpenAI's Mac Mini/Studio RL farm as environment machines for screenshot-and-action post-training.
  • Looped transformers reuse block weights across depth. Nanbeige runs 22 blocks twice; Universal Transformers and Mixture-of-Recursions vary loops per token via halting or routers; Ouro stacks more fixed passes. The point is more block applications without more distinct parameters. KV cache still grows with unrolled depth, and Nanbeige found shared caches hurt quality.
  • Astra likely loops a little; The Information overweights it. Reporting plus Pachocki's "within 2× of GPT-4" depth remark make looping plausible, but Raschka thinks better training data and recipes explain most of Astra's quality. Looping is a modest efficiency tweak, not the secret sauce.
  • Looping does not meaningfully hide CoT. OpenAI already hides most traces from users. Fewer output tokens at fixed accuracy can just mean fewer mistakes and less backtracking—as Luna versus Sol already showed. Pachocki says confused reporting should not spark an unmonitorability race; CoT monitoring remains fragile for reasons not contingent on architecture.
  • Recent papers refine when looping pays. Geiping-style latent recurrent depth adds test-time loops; Zhu et al. separate memorization (parameters) from multi-step reasoning (compute); SMELT estimates 6.8–18% less training compute for matched MoE looped transformers; full-bandwidth latent feedback can shorten traces in base models but not after instruction tuning.

quotes

A Looped Transformer is essentially an architectural tweak, with the main idea being to pass the intermediate representations through the same transformer blocks multiple times

Sebastian Raschka, defining weight-tied depth reuse.

I don’t think that looped transformers are significant contributors towards hiding or obscuring chains of thought.

Sebastian Raschka, rejecting the CoT-hiding rumor.

The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.

Jakub Pachocki, capping present frontier serial depth.

Looped transformers simply give better modeling performance at a fixed compute budget.

Sebastian Raschka, stating the practical payoff of looping.