MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

Research official + media 2 src. ~1 min

Interactive, stateful, tool-centric benchmark for mobile planning agents, closing the gap between GUI-centric benchmarks (surface-level screen manipulation) and static function-calling benchmarks (offline API matching). Spans 13 functional domains and 212 realistic mobile tools, running on an executable sandbox with live databases and structured feedback. Evaluates three advanced dimensions: sub-agent collaboration, memory usage, and skill usage. Extensive experiments show current frontier LLMs remain unreliable in mobile settings — performance drops sharply under strict tool ordering, permission limits, and unexpected runtime errors.

Why it matters

HF Daily 36 upvotes on Aug 25. Doubles as both a diagnostic benchmark and an interactive foundation for agentic RL on mobile; 212 tools across 13 domains is the broadest public mobile agent eval surface to date.

Importance: 3/5

HF Daily 36 upvotes

Sources

official arXiv abstract