September 12, 2026 · The Decoder
Study finds distinct internal signatures for each reasoning step
Researchers found that when AI models perform reasoning steps like calculation, formula retrieval, or deduction, each corresponds to a distinct, separable pattern in the model's internal states, especially in the middle layers.
Why it matters: This adds to evidence that a model's internal processing doesn't fully match what it writes in its visible chain-of-thought, reinforcing safety researchers' concerns that chain-of-thought monitoring alone may not reliably reveal what a model is actually doing.