#tool-use
- MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks research
- LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks research
- IBM releases Granite 4.2 reasoning LLMs (3B/8B/30B, Apache 2.0, 512K context) trained on ~15T tokens with GB200 NVL72 GRPO RL IBM models-llm
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work Apodex AI research
- Prime Agent: A Self-Improving RLM Harness Prime Intellect research
- Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory research