OSReward benchmarks cross-platform computer-use reward models

University of Hong Kong

Research official 1 src. ~1 min

A cross-university team built OSReward, a benchmark of 1,019 human-annotated computer-use trajectories across web, Windows, Ubuntu and mobile, to test VLM-as-judge reliability. Frontier judges score ~90% overall but drop to 70% on hard cases, mostly by wrongly accepting incomplete tasks as successes; the team also released open OS-Shepherd judge models (9B/35B) matching commercial judges at 30-60x lower cost.

Why it matters

Exposes a systematic blind spot (false-positive task completion) in using LLMs to grade computer-use agents, and ships an open, cheap alternative; 60 upvotes on Hugging Face Daily Papers.

Importance: 2/5

Notable benchmark and open judge-model release, below the 100-upvote bump threshold.

Sources