The Tasteful Agent: measuring 'taste' in long-horizon tasks

Microsoft

Research official 2 src. ~1 min

Introduces Taste-Bench: decision-fork questions auto-mined from real agent trajectories (parallel attempts and in-trajectory detours, no human annotation) measuring whether an agent picks the right branch before seeing outcomes. The best frontier model scores only 59.7%, forks with later-arriving evidence are much harder, and taste improves via distilling outcome-informed teacher judgment, lifting held-out SWE-bench Pro end-to-end success.

Why it matters

111 upvotes on HF Daily (09-23); a new axis for judging agent quality beyond final accuracy.

Importance: 4/5

Notable paper with 100+ HF Daily upvotes

Sources