LLMs Get Lost in Evolving User Intent
Microsoft Research
The paper introduces a framework that converts standard single-turn tasks into multi-turn conversations where the user's objective shifts mid-exchange, then tests how well LLMs track those shifts. It finds that strong single-turn performance does not transfer to this evolving-intent setting, with substantial accuracy drops across model families.
Why it matters
Highlights a blind spot in current evaluation methodology — models that look competent on static benchmarks can fail badly at tracking changing goals, a capability central to real-world collaborative agents.
Importance: 2/5
Notable eval-methodology paper; below the 100-upvote HF Daily bump threshold.