Legibility is not interpretability: judged vs. actual step importance in chain-of-thought

Research official 1 src. ~1 min

Defines a reasoning step's true importance via Monte-Carlo rollouts — the change in expected reward when that step is present — and tests whether LLM judges can recover it from CoT text. Capable judges beat a prevalence baseline but fall well short of a noise ceiling; even fine-tuned step-level critics stay far from ceiling on correct responses.

Why it matters

COLM 2026 paper. A cautionary result for process reward modeling and LLM-judge error diagnosis: step importance is only partially recoverable from the reasoning trace, so readable CoT is not the same as interpretable CoT.

Importance: 3/5

Notable COLM 2026 paper with implications for process reward modeling

Sources