LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

Research official 2 src. ~1 min

A harness that reformulates long-horizon agent execution as an explicit task-state management problem, keeping verified state outside the model's growing context instead of letting it accumulate unchecked assumptions. It lifts Qwen 3.7-Plus from 51.8% to 80.7% on WeaveBench and 2.8% to 8.3% on OSWorld 2.0, and raises Claude Opus 4.7 from 20.0% to 34.3% on an OSWorld 2.0 subset.

Why it matters

HuggingFace Daily Papers top pick for August 4, 2026 with 210 upvotes; large jumps on long-horizon computer-use benchmarks suggest state-management design, not just model scale, is a major lever for agent reliability.

Importance: 4/5

Notable paper with 210 HF Daily Papers upvotes (>=100 threshold), bumped +1.

Sources