Qwen and Taobao release E-Commerce Bench: 18 models run online stores for a simulated year
On 2026-09-03 Alibaba's Qwen team and Taobao & Tmall Group released E-Commerce Bench, a long-horizon agent benchmark with no natural stopping point. Models get ¥100,000 and a simulated year (365 days) of running online stores on desensitized Taobao data. GPT-5.6 Sol ended with ¥1.43M (14.3x), the top four finishers were all closed models, 10 of 90 episodes went bankrupt, and almost no model learned to buy more cheaply over the year.
Key facts
- Environment: 6,886 products in 60 categories, 576 suppliers (152 of them fraudulent), 12 store types, 10 market events, 8 promotions; each day has a 600-minute time budget that tool calls use up
- Suppliers use a deterministic negotiation kernel, with an LLM only rendering the dialogue, so agents cannot jailbreak prices below cost
- 18 models x 5 episodes: GPT-5.6 Sol ¥1.43M (14.31x); best open-weights model Qwen3.8-Max-Preview 4.16x; Qwen3.5-Plus averaged ¥1,100
- 10 of 90 episodes went bankrupt, including two GPT-5.5 runs that overstocked in January
- Negotiation: Claude Opus 4.7 scored 0.811 vs Kimi K2.6 0.596 (0.5 = accepting the opening quote)
- Fraud: share of spend going to fraudsters ranged from 0.12% (Claude Opus 4.7) to 20.11% (Qwen3.5-Plus); GPT-5.6 Sol ranked 16th
- Profit per tool call: Fable5 ¥479 vs GPT-5.6 Sol ¥363, with 59.9% fewer calls
- Across 8,647 repeat purchases only 2 of 18 models beat a random-order baseline on AnchorRatio; the median was 1.369, meaning prices drifted up
What happened
Most agent benchmarks end once a deliverable is produced. E-Commerce Bench instead scores sustained operation: an agent researches categories, haggles with suppliers, prices and lists goods, handles promotions, inventory, storage fees, returns and a three-account cash settlement, all over a simulated year. Runs are scored on seven axes: total assets, negotiation, fraud avoidance, cash flow, efficiency, execution and learning over time.
Why it matters
Results differed by orders of magnitude, and the top earner was not the most careful operator. The benchmark also found that current models barely learn within a long episode. It adds to the Vending-Bench-style evidence that long-horizon economic autonomy is still unreliable.
Some model names in the results (e.g. "Fable5", "Qwen3.8-Max-Preview") are copied exactly as Qwen wrote them.
Changelog
- 2026-09-30: created
Sources (4)
- officialQwen - E-Commerce Bench: Long-Horizon Operations, Multi-Dimensional Evaluation
- paperarXiv 2608.30730 - E-Commerce Bench
- codeGitHub - QwenLM/E-CommerceBench
- officialE-Commerce Bench project page
id: 2026-09-03-qwen-e-commerce-bench · updated 2026-09-30 · open in the interactive timeline