Legibility is not interpretability: judged vs. actual step importance in chain-of-thought
Defines a reasoning step's true importance via Monte-Carlo rollouts — the change in expected reward when that step is present — and tests whether LLM judges can recover it from CoT text. Capable judges beat a prevalence baseline but fall well short of a noise ceiling; even fine-tuned step-level critics stay far from ceiling on correct responses.
Why it matters
COLM 2026 paper. A cautionary result for process reward modeling and LLM-judge error diagnosis: step importance is only partially recoverable from the reasoning trace, so readable CoT is not the same as interpretable CoT.
Importance: 3/5
Notable COLM 2026 paper with implications for process reward modeling