Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Qwen and Taobao release E-Commerce Bench: 18 models run…

Qwen and Taobao release E-Commerce Bench: 18 models run online stores for a simulated year

★★after cutoffbenchmarkAlibabaQwenTaobao & Tmall Groupconfidence: high

On 2026-09-03 Alibaba's Qwen team and Taobao & Tmall Group released E-Commerce Bench, a long-horizon agent benchmark with no natural stopping point. Models get ¥100,000 and a simulated year (365 days) of running online stores on desensitized Taobao data. GPT-5.6 Sol ended with ¥1.43M (14.3x), the top four finishers were all closed models, 10 of 90 episodes went bankrupt, and almost no model learned to buy more cheaply over the year.

Key facts

What happened

Most agent benchmarks end once a deliverable is produced. E-Commerce Bench instead scores sustained operation: an agent researches categories, haggles with suppliers, prices and lists goods, handles promotions, inventory, storage fees, returns and a three-account cash settlement, all over a simulated year. Runs are scored on seven axes: total assets, negotiation, fraud avoidance, cash flow, efficiency, execution and learning over time.

Why it matters

Results differed by orders of magnitude, and the top earner was not the most careful operator. The benchmark also found that current models barely learn within a long episode. It adds to the Vending-Bench-style evidence that long-horizon economic autonomy is still unreliable.

Some model names in the results (e.g. "Fable5", "Qwen3.8-Max-Preview") are copied exactly as Qwen wrote them.

Changelog

  • 2026-09-30: created

Sources (4)

id: 2026-09-03-qwen-e-commerce-bench · updated 2026-09-30 · open in the interactive timeline