OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

NLP Group of The University of Hong Kong

Research official 1 src. ~1 min

Proposes a benchmark for judging computer-using agent trajectories with vision-language models, finds even state-of-the-art models show a systematic leniency bias that mislabels failures as successes, and releases the OS-Shepherd-100K dataset plus 9B/35B open reward models matching commercial alternatives at lower cost.

Why it matters

Second most-upvoted paper on HuggingFace Daily Papers for 2026-08-09 with 67 upvotes; exposes a concrete evaluation flaw affecting how computer-use agents are graded across the field.

Importance: 3/5

Notable HuggingFace Daily Paper (research) — Second most-upvoted paper on HuggingFace Daily Papers for 2026-08-09 with 67 upvotes; exposes a concrete evaluation flaw affecting how computer-use agents are graded across the field.

Sources