Skip to content
← All Charts
In ChartsPublished

Can AI Customer Service Solve the Same Case Four Times?

Two AI agents, 97 simulated cases each, four runs per case. Most cases succeed at least once. Far fewer succeed every time.

Repeat the same case. The result can change.97 SIMULATED CASES PER MODEL · FOUR RUNS EACH

Numbers inside the bars count cases, not individual runs. “Every time” means success in all four recorded runs, not four chances to succeed once. Two examples, not a model ranking.

Data: Sierra Research · τ-knowledge · Chart: VastkindRecorded: 3–4 Aug 2026 · Submission: 4 Aug 2026 · Retrieved: 7 Oct 2026

Getting it right once is the easier bar

Imagine asking customer service to replace a lost card. A successful interaction is encouraging. The harder question is whether the same request still works when the conversation takes a different path.

In these published simulations, Claude Opus 5 succeeded at least once on 64 of 97 cases; Qwen 3.8 Max did so on 71. Require success in all four runs, and those counts fall to 31 and 34. The middle segment shows cases that sometimes worked and sometimes did not.

A benchmark, not a bank

These are recorded evaluations in a fictional banking environment, with an AI playing the customer. Vastkind grouped the original scores; we did not run the agents again or independently regrade their conversations. The chart measures neither real customer satisfaction nor financial safety.

Both agents received the same embedded case definitions and tool access, but used different reasoning settings. These two examples do not establish a general model ranking or equal-cost comparison. Four runs per case are a small sample, not a guarantee about the next run.

A further reproducibility limit: 15 of the 97 embedded task definitions differ from the code revision named in the run files. They match between the two files analyzed here. The reason for the repository discrepancy is unresolved.

Sources, Method & Data

What we counted. The τ-knowledge banking_knowledge AllTools run files contain 97 case definitions and 388 runs per model, four distinct trial IDs per case. We treated a logged reward within 0.000001 of 1 as success, matching the project metric, then grouped cases by zero, one to three, or four successes. Both selected files contain no infrastructure-error terminations. Bar widths are group counts divided by 97; labels give the counts.

Settings and timing. Qwen 3.8 Max used xhigh reasoning; Claude Opus 5 used max. Both used GPT-5.2 with low reasoning as the customer simulator and a 200-step limit. The metadata evaluation date is August 3, 2026; raw timestamps span August 3–4 without a stated time zone. The submission date is August 4; the hosted files also carry August 4 modification timestamps. This is a new explanation of that snapshot, not a new October test or a comprehensive current leaderboard.

Denominators matter. Individual-run success was 189/388 for Claude and 214/388 for Qwen. That average is not the first attempt. “At least once” counts cases with one or more successes; “every time” counts cases with four. We make no independence assumption or statistical-significance claim.

Original records. Repository snapshot; MIT code license. Own aggregation and graphic, not a copied source figure. The hosted conversation files carry no separate license notice; full dialogs are not republished here.

Task-version boundary. The run files name code commit fc0055dc4e0a316c3f83133267fbd6faaa770992. Fifteen embedded task definitions differ from that revision and from the repository snapshot above. Our counts describe the scores in the published run files. They do not establish an exact reproduction from the named code revision. Other model results were not included because their comparability was not fully verified.

Recorded outcomes: cases, not runs
Model0 of 41 of 42 of 43 of 44 of 4Total
Claude Opus 5331112103197
Qwen 3.8 Max26145183497

Related: Can AI spot an IKEA assembly mistake? A different capability test, with its own limits.

Explore Vastkind

Topics
Subscribe

Stay curious. Stay informed.

For another moment

Your reading list

Saved on this device. No account needed.

Something catch your eye? Use the bookmark beside an article to keep it here.