LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
A harness that reformulates long-horizon agent execution as an explicit task-state management problem, keeping verified state outside the model's growing context instead of letting it accumulate unchecked assumptions. It lifts Qwen 3.7-Plus from 51.8% to 80.7% on WeaveBench and 2.8% to 8.3% on OSWorld 2.0, and raises Claude Opus 4.7 from 20.0% to 34.3% on an OSWorld 2.0 subset.
Why it matters
HuggingFace Daily Papers top pick for August 4, 2026 with 210 upvotes; large jumps on long-horizon computer-use benchmarks suggest state-management design, not just model scale, is a major lever for agent reliability.
Importance: 4/5
Notable paper with 210 HF Daily Papers upvotes (>=100 threshold), bumped +1.
Sources
official
HuggingFace Daily Papers