HarnessEval-W: Agentifying the Evaluation of Visual Worlds
NTU / MirroS-Lab consortium
Replaces brute-force scalar metrics for world-model rollouts with an agentified pipeline: a parent agent decomposes each evaluation into subproblems, spawns specialized sub-agents with tailored context and diagnostic tools, then validates and summarizes an evidence tree. Applied to 18 world models over 330 cases, its verdicts align with human preferences while exposing per-rollout reasoning.
Why it matters
HF: 30 upvotes. First world-model benchmark that ships a verifiable reasoning chain alongside the score.
Importance: 2/5
HF Daily 30 upvotes
Sources
official
arXiv:2608.16859