MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

Alibaba Group

Research official 2 src. ~1 min

A 365-day, order-level simulated e-commerce environment grounded in nearly 99,000 real 1688 marketplace product records, with 26 interaction tools, designed to test whether LLM agents can preserve purposeful, adaptive behavior over long operational horizons.

Why it matters

Across 48 runs spanning eight LLMs and two agent frameworks, LLM agents show a substantial performance gap against a human baseline, highlighting how far current agents are from reliable long-term autonomous operation.

Importance: 2/5

Notable new long-horizon agent benchmark grounded in real marketplace data.

Sources