The Tasteful Agent: measuring 'taste' in long-horizon tasks
Microsoft
Introduces Taste-Bench: decision-fork questions auto-mined from real agent trajectories (parallel attempts and in-trajectory detours, no human annotation) measuring whether an agent picks the right branch before seeing outcomes. The best frontier model scores only 59.7%, forks with later-arriving evidence are much harder, and taste improves via distilling outcome-informed teacher judgment, lifting held-out SWE-bench Pro end-to-end success.
Why it matters
111 upvotes on HF Daily (09-23); a new axis for judging agent quality beyond final accuracy.
Importance: 4/5
Notable paper with 100+ HF Daily upvotes
Sources
secondary
HuggingFace Daily Papers 2026-09-23