hraness
Theme
Appearance

saved

Fable 5 — Median thinking declined in August

by Lon LundgrenXpublished

Hraness cites a source capture. The source author remains the source.

Lon Lundgren @Lon

After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why.

Measured five different ways, August delivered dramatically fewer thinking tokens than July.

This wasn't a one-time drop. Reasoning fell over the entire period, and fluctuated across multi-day episodes. Some of these fluctuations aligned with specific product announcements and releases. I started to see how the model could feel great one day, and terrible the next.

I was consistently using an xhigh or max effot level, but when I looked deeper, I found that most invocations to the model were receiving little to no thinking tokens at all. And when longer thinking runs did happen, they almost never reached published benchmark levels.

This unfolded across a six week data capture and analysis odyssey, and led to a number of surprising findings. Next time the model feels dumber, don't ask if the model was "nerfed". Ask about the inference regime you were served, instead.

Lollipop chart: Fable 5 — Median thinking declined in August. August drops of -21.9%, -23.1%, -35.1%, -46.2%, and -50.6% versus July across five weightings.Three-panel time series: Fable 5 — Thinking was unstable and declined over time, with product-event markers and a late-August majority-zero thinking interval.Histogram: Production rarely reached Fable 5's benchmark-scale reasoning; 96.3% of invocations under 4K thinking tokens versus benchmark means of 11.6K–51.7K.