Companies
Companies

Alibaba test shows top AI model completed just 62% of commerce tasks

Claude Opus 5 finished 61.7% of 107 real merchant tasks in the CommerceAgentBench toolkit, with failures clustering on multi-step logistics and disputes.

Yuna · Sep 18, 2026 · 1 min

Copy linkShare

According to PYMNTS, Alibaba.com introduced CommerceAgentBench, an open toolkit that evaluated 107 real commerce tasks across 13 AI model families. The benchmark assessed agents’ ability to finish merchant duties ranging from product listings to logistics and after-sales support.

Claude Opus 5 recorded the best result, succeeding on 61.7% of the assignments. Data from the GitHub repository indicated the system used an average of 63 tool calls and roughly 10 minutes for each task. Outcomes differed by software environment: the leading model cleared 61.7% of cases in Alibaba’s Accio setup, 60.7% in Pi, and 56.1% in OpenClaw. Gemini 3 Flash recorded the lowest score, finishing 29% of the work across all three platforms.

Errors were most frequent in jobs requiring the retention of data over numerous stages, such as landed costs, return conflicts, and complex shipping paths. Alibaba stated, “No single model led across the board, reinforcing the case for task-level routing, the idea that different tasks are best handled by different AI models rather than relying on one general-purpose model for everything.”

In a Fortune opinion piece, Alibaba.com President Kuo Zhang described the peak performance as “high enough to be useful and low enough to be a warning.” The company said the benchmark was built using input from 10 million active small business accounts, 1.6 million dialogues, and 200,000 execution logs. A distinct PYMNTS Intelligence study revealed that 132 million U.S. adults had purchased retail goods with AI assistance, with Amazon accounting for 59% of those sales.

Source: PYMNTS

This story was produced by StreamSage's AI newsroom. Not financial advice.

More stories