Not an expert in LLMs, but this seems supported by the abstract of the paper cited in the above article:
it remains unclear to what extent these performance gains can be attributed to human-like task decomposition or simply the greater computation that additional tokens allow. [...] our results show that additional tokens can provide computational benefits independent of token choice. The fact that intermediate tokens can act as filler tokens raises concerns about large language models engaging in unauditable, hidden computations that are increasingly detached from the observed chain-of-thought tokens.
https://arxiv.org/html/2404.15758v1If you want to see it yourself: load up Qwen 3.8 in LM Studio and watch the CoT stumble around like a drunken sailor before miraculously jumping to the correct result.
If you want an example of subversion, Anthropic has some good ones:
https://transformer-circuits.pub/2025/attribution-graphs/bio...
https://transformer-circuits.pub/2025/attribution-graphs/bio...
But you don’t have to do that. You can skip straight to RL. If you do, the model will generate complete garbage reasoning traces before generating the correct answer. In fact, if you add a coherence reward to the reasoning trace, the model will perform worse (since you’re now diluting the correctness reward).
Depends on whether and how you want to rank stability in terms of better/worse. Models are diverging on this, which seems increasingly clear.. i.e. Fable isn't stable, but Opus isn't clever, and they hit different kinds of walls. So both the theory (diluting the correctness reward) and the practice (hard split on plan/implement/review work) seems to be pointing towards a strongly multi-model and highly agentic / harness-driven / complex-system kind of future instead of singleton monolithic super-smart models.
The do-everything model with solid reasoning AND solid results, and the honest/introspective helpful agent that doesn't actively resist governance may be at odds. Stable reasoning doesn't matter for pen-testing, and correct-answer with broken processes and fragile abstractions won't matter for math/science/coding.
I think the more intuitive mechanical explanation is, in RL when you are assigning rewards to a rollout you might give a reward for stable reasoning and another for correctness.
If you are just summing the two, a rollout with better correctness can score equivalently to a rollout with a better answer. So ultimately you can end up with worse answers.
> More interestingly, our experiments also show that models trained on corrupted traces, whose intermediate reasoning steps bear no relation to the problem they accompany, achieve performance largely comparable to those trained on correct traces.
> *TLDR*: Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)