StateM: harness scaling pushes GPT-5.6 Sol to 95.3% on Terminal-Bench 2.1 and the same harness adapts to DeepSeek-V4 Flash for $15 total spend
An agent-native runtime that organizes execution around durable states, phase-local context, checked transitions, recoverable runbooks, and versioned practices. On Terminal-Bench 2.1 raises GPT-5.5 xhigh from 83.1% to 92.1%, pushes GPT-5.6 Sol xhigh to 95.3% raw accuracy, and adapts the same harness to DeepSeek-V4 Flash for under $38 with total API spend of $15 versus $574.68 for the GPT reference.
Why it matters
212 upvotes on HuggingFace Daily Papers (Aug 18 top). Argues long-horizon agent failures are harness problems, not model limits, and shows the same harness transfers between GPT-5.6 and DeepSeek-V4 with ~38x cost reduction — an actionable finding for anyone running coding agents.
Importance: 4/5
HF Daily 212 upvotes
Sources
official
StateM (HF Daily Papers)
media
arXiv listing