Last Translation Benchmark: human-authored tasks that still break frontier translation models
A large consortium (Vilem Zouhar, Niyati Bafna, Stella Biderman and others) releases LTBv1, a live collection of human-authored, peer-reviewed texts, images, audio and video that break leading machine translation models as standard benchmarks saturate. Each example ships with handcrafted verification rules describing concrete failure modes, enabling reliable, reward-hack-resistant, actionable evaluation.
Why it matters
An unsaturated, community-maintained eval with rule-based verification — a template for post-saturation benchmarking beyond MT.
Importance: 2/5
Notable community benchmark release
Sources
official
Last Translation Benchmark (arXiv)