Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Tools
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status
  • AI Site Map

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
All benchmarks

τ²-Bench: Airline

τ²-Bench Airline tests whether a model can do an airline support agent's job: follow a policy manual, talk to a simulated customer, and call the right tools to search flights, change bookings, and issue refunds. The model doesn't need any domain knowledge; it's scored entirely on how well it executes tool calls across multi-step task trajectories. We run it continuously against the same provider endpoints that serve OpenRouter traffic, so a score reflects both the model and the provider running it. We use this benchmark because it has a high floor, so we can assess provider variance and not model capability. Our routing algorithm for tool call requests uses these same signals to send traffic to the best performing endpoints.

Last benchmark run Sep 29, 2026, 9:00 PM UTC

PaperGitHub

Which model leads the τ²-Bench Airline benchmark?

Google: Gemini 3.7 Flash leads τ²-Bench Airline at 80.6% as of Sep 29, 2026, 9:00 PM UTC. Airline tasks, policy, simulator and customer setup, step budget, and grading are held fixed across production provider endpoints, so the displayed results compare models and the providers serving them under the same evaluation conditions.

Tool-call errors before → after Auto Exacto

3.9% → 3.5%

Tool-call error rate on models enrolled in Auto Exacto, OpenRouter's automatic provider optimization for tool-calling requests.
Model comparisonCost efficiencyTool-call reliabilityLeaderboardWhy we run itWhat scores tell youHow tasks are scoredMethodologyAPI accessDownload and citeFAQ

Model comparison

Most Accurate

Favicon for google
Google: Gemini 3.7 Flash

80.6%

Best Value

Favicon for stepfun
StepFun: Step 3.7 Flash

$0.020/task

Fastest

Favicon for amazon
Amazon: Nova Micro 1.0

1.7m

Accuracy
Representative-run accuracy, best first.
Cost per task
Average cost per graded task, cheapest first.
Time per task
Average wall-clock time per task, fastest first; agents that loop or stall run long.

Cost efficiency

Accuracy vs. cost (Pareto frontier)
One point per model, using default routing (not pinned to a provider) when available. The line is the Pareto frontier: no model beats these on both accuracy and cost.

Tool-call reliability

Tool-call errors
Share of this benchmark's own requests where the model called a tool that doesn't exist, passed arguments that don't match the tool's schema, or emitted arguments that aren't valid JSON.

Leaderboard

Top-level rows use default routing where available; click a row to expand provider-pinned results.

τ²-Bench Airline leaderboard: model accuracy, cost, time, and output tokens per task by provider
#ModelStd dev
1
Favicon for google
Google: Gemini 3.7 Flash
Pareto
80.6%--$0.0772.0m14.7k
2
Favicon for anthropic
Anthropic: Claude Fable 5
79.6%±2.8pp$0.932.2m5.75k
3
Favicon for anthropic
Anthropic: Claude Fable 5.1
79.3%±1.4pp$0.691.9m6.93k
4
Favicon for anthropic
Anthropic: Claude Opus 5
78.8%±2.5pp$0.492.2m7.13k
5
Favicon for amazon
Amazon: Nova Micro 1.0
78.7%±2.0pp$1.281.7m5.69k
6
Favicon for z-ai
Z.ai: GLM 5
Pareto
78.0%±3.4pp$0.0463.4m9.39k
7
Favicon for google
Google: Gemini 3 Flash Preview
77.3%±2.0pp$0.122.5m20.5k
8
Favicon for stepfun
StepFun: Step 3.7 Flash
Pareto
77.3%--$0.0203.8m10.9k
9
Favicon for anthropic
Anthropic: Claude Opus 4.5
77.2%±2.3pp$0.532.9m10.5k
10
Favicon for qwen
Qwen: Qwen3.8 27B
77.1%±2.4pp$0.0754.4m12.8k
11
Favicon for deepseek
DeepSeek: DeepSeek V4 Pro 0813
76.9%±1.4pp$0.105.6m20.2k
12
Favicon for anthropic
Anthropic: Claude Sonnet 5.5
76.7%--$0.1880s6.16k
13
Favicon for openai
OpenAI: GPT-5.6 Sol Pro
76.7%--$1.865.6m22.9k
14
Favicon for deepseek
DeepSeek: DeepSeek V4 Pro 0423
76.7%±2.6pp$0.0404.1m12.2k
15
Favicon for openai
OpenAI: GPT-5.6 Sol
76.7%±0.7pp$0.212.1m5.14k
16
Favicon for google
Google: Gemma 4 31B
76.5%±4.2pp$0.0238.4m13.4k
17
Favicon for qwen
Qwen: Qwen3.5-122B-A10B
76.2%±2.6pp$0.1710.5m40.1k
18
Favicon for qwen
Qwen: Qwen3.5 397B A17B
76.2%±2.4pp$0.1611.0m26k
19
Favicon for z-ai
Z.ai: GLM 5.3
76.2%±2.5pp$0.0692.7m6.81k
20
Favicon for anthropic
Anthropic: Claude Opus 4.7
76.2%±2.1pp$0.421.6m5.63k
21
Favicon for anthropic
Anthropic: Claude Opus 4.8
76.1%±1.8pp$0.502.2m8.11k
22
Favicon for x-ai
SpaceXAI: Grok 4.6
76.0%--$0.354.1m11.6k
23
Favicon for deepseek
DeepSeek: DeepSeek V4.1 Flash
Pareto
75.7%±1.0pp$0.0184.0m15.1k
24
Favicon for xiaomi
Xiaomi: MiMo-V2.6-Pro
GMICloud
75.7%--$0.05416.8m38.2k
25
Favicon for google
Google: Gemini 3.5 Flash Lite
75.7%±1.7pp$0.111.7m25.1k
26
Favicon for openai
OpenAI: GPT-5.1
75.7%±4.3pp$0.438.7m37.6k
27
Favicon for openai
OpenAI: GPT-5.4
75.7%±0.3pp$0.302.5m10.8k
28
Favicon for anthropic
Anthropic: Claude Sonnet 5
75.4%±2.9pp$0.202.2m7.83k
29
Favicon for z-ai
Z.ai: GLM 5.1
75.3%±2.2pp$0.0784.3m9.88k
30
Favicon for openai
OpenAI: GPT-5.5
75.3%±2.7pp$0.502.6m8.28k
31
Favicon for qwen
Qwen: Qwen3.8 2.4T A95B
75.3%±1.4pp$0.206.0m12.6k
32
Favicon for google
Google: Gemini 3.1 Pro Preview
75.3%±2.0pp$0.362.3m12.8k
33
Favicon for z-ai
Z.ai: GLM 5.3 Flash
Pareto
75.0%±2.0pp$0.0062.4m4.22k
34
Favicon for anthropic
Anthropic: Claude Opus 4.6
74.8%±2.7pp$0.493.2m10k
35
Favicon for google
Google: Gemini 3.5 Flash
74.7%±0.7pp$0.422.8m28.3k
36
Favicon for deepseek
DeepSeek: DeepSeek V4 Flash 0423
74.5%±2.1pp$0.0093.7m14k
37
Favicon for anthropic
Anthropic: Claude Sonnet 4.6
74.4%±2.2pp$0.323.3m11.4k
38
Favicon for moonshotai
MoonshotAI: Kimi K2.6
74.2%±2.2pp$0.0765.5m12.9k
39
Favicon for meta
Meta: Muse Glimmer 30B
74.1%±3.3pp$0.0383.3m13.5k
40
Favicon for qwen
Qwen: Qwen3.6 27B
74.1%±4.4pp$0.1313.4m31k
41
Favicon for deepseek
DeepSeek: DeepSeek V4 Flash Vision Exp
73.8%±0.7pp$0.0293.2m18.4k
42
Favicon for minimax
MiniMax: MiniMax M3
73.7%±2.6pp$0.0232.5m7.29k
43
Favicon for xiaomi
Xiaomi: MiMo-V2.5-Pro
73.5%±3.2pp$0.0317.2m13.3k
44
Favicon for openai
OpenAI: GPT-5.6 Terra
73.3%±1.3pp$0.1001.6m4.17k
45
Favicon for deepseek
DeepSeek: DeepSeek V4 Flash 0731
73.3%±5.2pp$0.0085.4m15.6k
46
Favicon for anthropic
Anthropic: Claude Opus 5.5
73.3%--$0.342.0m7.08k
47
Favicon for openai
OpenAI: GPT-5.2
73.3%--$0.243.5m12.5k
48
Favicon for xiaomi
Xiaomi: MiMo-V2.6-Flash
73.3%--$0.02717.5m22.3k
49
Favicon for deepseek
DeepSeek: DeepSeek V3.2
73.2%±2.8pp$0.0318.1m14.4k
50
Favicon for google
Google: Gemini 3.6 Flash
73.0%±0.3pp$0.242.4m22.9k
51
Favicon for z-ai
Z.ai: GLM 5.2
72.8%±2.9pp$0.0322.7m8.64k
52
Favicon for tencent
Tencent: Hy4 preview
72.7%--$0.0673.8m13.1k
53
Favicon for google
Google: Gemini 3.1 Flash Lite
72.3%±1.0pp$0.0972.4m51.2k
54
Favicon for z-ai
Z.ai: GLM 4.7
72.1%±2.5pp$0.0615.3m9.73k
55
Favicon for z-ai
Z.ai: GLM 4.6
72.1%±1.9pp$0.0496.4m9.25k
56
Favicon for moonshotai
MoonshotAI: Kimi K2.7 Code
72.1%±2.2pp$0.0734.3m8.43k
57
Favicon for anthropic
Anthropic: Claude Sonnet 4.5
72.0%±3.5pp$0.333.4m10.3k
58
Favicon for inclusionai
inclusionAI: Ling 3.0 Flash VL
Pareto
Novita
72.0%--$0.0042.8m15.3k
59
Favicon for nvidia
NVIDIA: Nemotron 3 Ultra
72.0%--$0.151.8m13.1k
60
Favicon for qwen
Qwen: Qwen3.7 Max
AtlasCloud
72.0%--$0.403.3m16.3k
61
Favicon for moonshotai
MoonshotAI: Kimi K3
71.9%±2.8pp$0.307.0m9.3k
62
Favicon for openai
OpenAI: GPT-5.6 Luna Pro
71.3%±1.3pp$0.0645.8m34.2k
63
Favicon for deepseek
DeepSeek: DeepSeek V3.2 Exp
70.9%±1.7pp$0.09118.0m17.1k
64
Favicon for moonshotai
MoonshotAI: Kimi K2.5
70.8%±2.8pp$0.0405.4m9.1k
65
Favicon for qwen
Qwen: Qwen3.6 35B A3B
70.8%±7.5pp$0.0484.6m27.9k
66
Favicon for openrouterFavicon for openrouter
Auto Router (Beta)
70.7%--$0.242.5m17.9k
67
Favicon for openai
OpenAI: GPT-5.6 Luna
70.7%±0.0pp$0.0091.7m5.4k
68
Favicon for xiaomi
Xiaomi: MiMo-V2.5
70.4%±2.4pp$0.0085.1m13.7k
69
Favicon for deepseek
DeepSeek: DeepSeek V3.1 Terminus
70.4%±2.9pp$0.0598.9m14.2k
70
Favicon for z-ai
Z.ai: GLM 4.5 Air
70.0%±1.4pp$0.0184.9m10.3k
71
Favicon for openai
OpenAI: GPT-5 Mini
69.7%±3.0pp$0.1311.6m58.4k
72
Favicon for minimax
MiniMax: MiniMax M2.7
69.7%±3.3pp$0.0263.8m9.66k
73
Favicon for openai
OpenAI: GPT-6 Sol
68.7%--$0.131.9m4.99k
74
Favicon for qwen
Qwen: Qwen3.5-9B
68.6%±5.6pp$0.03015.8m40.4k
75
Favicon for openai
OpenAI: GPT-5
68.4%±3.8pp$0.6813.4m61.2k
76
Favicon for anthropic
Anthropic: Claude Sonnet 4
68.0%±2.0pp$0.332.9m8.95k
77
Favicon for openai
OpenAI: GPT-6 Astra
Amazon Bedrock
68.0%--$0.651.5m3.84k
78
Favicon for z-ai
Z.ai: GLM 4.7 Flash
67.9%±2.0pp$0.0094.1m10.6k
79
Favicon for openai
OpenAI: GPT-5.2 Chat
OpenAI
67.3%--$0.1180s4.28k
80
Favicon for openai
OpenAI: GPT-6 Luna Pro
67.3%--$0.0405.1m31.7k
81
Favicon for google
Google: Gemma 4 26B A4B
67.2%±3.9pp$0.0227.2m20.6k
82
Favicon for openai
OpenAI: GPT-5.4 Nano
67.0%±0.3pp$0.0332.5m17.5k
83
Favicon for thinkingmachines
Thinking Machines: Inkling
67.0%±3.1pp$0.07968s3.64k
84
Favicon for anthropic
Anthropic: Claude Haiku 4.5
66.7%±3.5pp$0.122.0m12.6k
85
Favicon for inclusionai
inclusionAI: Ling 3.0 Flash Fin
Novita
66.0%--$0.00573s17.9k
86
Favicon for openai
OpenAI: GPT-6 Luna
66.0%--$0.0081.5m8.32k
87
Favicon for minimax
MiniMax: MiniMax M2.5
65.7%±2.3pp$0.0224.8m10.4k
88
Favicon for minimax
MiniMax: MiniMax M2.1
64.9%±2.5pp$0.0252.8m11.3k
89
Favicon for nvidia
NVIDIA: Nemotron 3.5 Lightning
64.8%±1.3pp$0.0183.2m22k
90
Favicon for openai
OpenAI: gpt-oss-120b
63.2%±3.1pp$0.0214.8m22k
91
Favicon for qwen
Qwen: Qwen3.5-35B-A3B
63.1%±9.3pp$0.0689.8m38.2k
92
Favicon for openai
OpenAI: GPT-5.4 Mini
63.0%±5.0pp$0.0922.2m12.5k
93
Favicon for qwen
Qwen: Qwen3 235B A22B Thinking 2507
62.2%±1.4pp$0.1111.9m34.1k
94
Favicon for google
Google: Gemini 2.5 Pro
61.3%±0.7pp$0.223.2m15.1k
95
Favicon for thinkingmachines
Thinking Machines: Inkling Small
60.3%±2.9pp$0.0473.4m3.75k
96
Favicon for moonshotai
MoonshotAI: Kimi K2 Thinking
59.4%±22.1pp$0.0483.9m7.22k
97
Favicon for moonshotai
MoonshotAI: Kimi K2 0905
58.9%±3.6pp$0.1001.9m3.69k
98
Favicon for deepseek
DeepSeek: DeepSeek V3.1
58.0%±6.8pp$0.0537.2m10.2k
99
Favicon for google
Google: Gemini 2.5 Flash
57.7%±0.3pp$0.02979s8.17k
100
Favicon for inclusionai
inclusionAI: Ling 3.0 Flash
57.3%--$0.0122.0m16.8k
101
Favicon for qwen
Qwen: Qwen3 Coder Next
55.6%±9.8pp$0.0411.6m4.3k
102
Favicon for openai
OpenAI: gpt-oss-20b
52.9%±1.6pp$0.02729.5m143k
103
Favicon for nvidia
NVIDIA: Nemotron 3 Nano 30B A3B
52.3%±2.2pp$0.0287.4m57.4k
104
Favicon for deepseek
DeepSeek: R1 0528
51.8%±3.9pp$0.1425.6m30.9k
105
Favicon for openai
OpenAI: GPT-4.1
48.7%±0.7pp$0.1251s2.66k
106
Favicon for qwen
Qwen: Qwen3 32B
48.3%±3.0pp$0.02212.0m22.4k
107
Favicon for openai
OpenAI: GPT-5 Nano
47.6%±4.5pp$0.04311.6m99.7k
108
Favicon for qwen
Qwen: Qwen3 Next 80B A3B Instruct
47.3%±2.6pp$0.0252.1m3.8k
109
Favicon for google
Google: Gemini 2.5 Flash Lite
47.3%--$0.0242.9m45.1k
110
Favicon for qwen
Qwen: Qwen3 Coder 480B A35B
47.0%±3.2pp$0.111.7m2.76k
111
Favicon for openai
OpenAI: GPT-4o
46.3%±1.7pp$0.2052s2.08k
112
Favicon for openai
OpenAI: GPT-5.3 Chat
OpenAI
46.0%--$0.09965s2.75k
113
Favicon for qwen
Qwen: Qwen3 235B A22B Instruct 2507
45.3%±2.6pp$0.0283.5m3.69k
114
Favicon for openai
OpenAI: GPT-4o (2024-08-06)
OpenAI
44.7%±1.3pp$0.2147s2.22k
115
Favicon for openai
OpenAI: GPT-4.1 Mini
44.0%±1.3pp$0.02974s2.45k
116
Favicon for meta-llama
Meta: Llama 4 Maverick
44.0%±2.1pp$0.0462.1m2.04k
117
Favicon for deepseek
DeepSeek: DeepSeek V3 0324
43.7%±1.9pp$0.07713.1m19.6k
118
Favicon for mistralai
Mistral: Mistral Small 4
Mistral
43.7%±2.3pp$0.0141.6m10.2k
119
Favicon for qwen
Qwen: Qwen3 30B A3B
43.3%±3.4pp$0.0347.3m22.8k
120
Favicon for qwen
Qwen: Qwen3 Coder 30B A3B Instruct
43.1%±0.6pp$0.0243.6m3.15k
121
Favicon for qwen
Qwen: Qwen3 14B
42.3%±1.0pp$0.0308.7m20.1k
122
Favicon for qwen
Qwen: Qwen3 VL 235B A22B Instruct
40.5%±3.9pp$0.0595.0m7.41k
123
Favicon for openai
OpenAI: GPT-4o (2024-05-13)
40.0%--$1.1372s3.25k
124
Favicon for qwen
Qwen: Qwen3 30B A3B Instruct 2507
39.0%±2.3pp$0.0263.6m4.47k
125
Favicon for deepseek
DeepSeek: DeepSeek V3
38.3%±0.3pp$0.0714.7m7.51k
126
Favicon for meta-llama
Meta: Llama 3.3 70B Instruct
35.9%±3.8pp$0.0211.7m1.09k
127
Favicon for qwen
Qwen: Qwen3 VL 30B A3B Instruct
35.5%±2.4pp$0.0804.0m7.63k
128
Favicon for qwen
Qwen: Qwen3 VL 8B Instruct
35.3%±1.4pp$0.0391.6m4.61k
129
Favicon for openai
OpenAI: GPT-4o-mini
OpenAI
27.0%±3.0pp$0.02166s3.95k
130
Favicon for qwen
Qwen2.5 72B Instruct
26.0%±1.3pp$0.0799.8m5.67k
131
Favicon for mistralai
Mistral: Mistral Nemo
Mistral
23.3%±2.7pp$0.0296.3m5.49k
132
Favicon for meta-llama
Meta: Llama 3.1 8B Instruct
DeepInfra
21.2%±4.6pp$0.00546s1.35k
133
Favicon for qwen
Qwen: Qwen2.5 7B Instruct
Phala
16.7%--$0.01758s2.49k
134
Favicon for openai
OpenAI: GPT-4.1 Nano
10.3%±0.3pp$0.00761s2.32k

Why we run this benchmark

It's a tool-calling benchmark that is hard to game. Grading depends on live tool-call trajectories rather than memorized answers, so it resists training-data leakage better than Q&A-style evals. It exercises every tool-calling failure mode (wrong arguments, skipped policy checks, giving up, hallucinated confirmations) at a relatively low cost per run. The relative scores also carry more signal than the absolute ones. The same model can score differently across providers, and those deltas are what Exacto routing uses to pick higher-accuracy endpoints.

Each task is a simulated airline support conversation with a scripted user, a toolbox (flight search, booking changes, refunds, loyalty policies), and a gold reference solution. A task passes only if the final database state and the messages to the user match the reference; partial credit is not awarded.

What the scores can and can't tell you

There is still headroom. Top models fail roughly one in five tasks, and the airline domain is the hardest τ²-Bench split. Accuracy differences here separate models that follow multi-step policies from ones that merely chat well.

The floor is high, though. Many tasks reward inaction. A refusal task with an empty gold action list passes for any agent that changes nothing. Even weak models score well above zero, so the meaningful spread sits at the top of the range.

The benchmark is public, so tasks may appear in training corpora. Contamination inflates scores less here than in Q&A-style evals, though, since a leaked task still has to be executed correctly, step by step, against a live database.

Scoring fidelity has limits. The checker verifies two things: the final database hash and exact substring matches in the agent's messages. Each task's natural-language assertions ("agent should refuse the cancellation") are metadata, and no judge model reads the transcript. So a savings calculation fails if the agent says "$23,552.50" when the checker greps for "23553".

The user simulator matters too. We pin it to gemini-2.5-flash so agent scores stay comparable, but the sim is itself an LLM with failure modes of its own. It can stop the conversation before the agent finishes, leak its hidden task instructions, or keep a stuck agent looping until the 200-step ceiling kills the run. Swapping the sim model shifts absolute scores, which is why cross-paper τ²-Bench numbers rarely line up exactly.

How a task is scored

Every task ships a gold solution: a list of tool calls, strings the agent must say, and natural-language assertions. After the conversation ends, the checker replays the gold tool calls against a fresh database and compares hashes with the agent's final database. It then greps the agent's messages for each required string. The reward is the product of those two checks:

reward = db_match × communicate_met   // each ∈ {0, 1}
db_match        = hash(agent DB) == hash(gold DB)
communicate_met = every required string appears in an agent message
any run that hits MAX_STEPS instead of a clean stop scores 0 outright

The rollouts below are from real runs, with gemini-2.5-flash as the user simulator throughout.

reward = 1

Pass: three changes in one request, all three land

Task 17 agent: openai/gpt-5.1

For reservation FQ8APE: add 3 checked bags, swap the passenger to Omar Rossi, and upgrade basic economy to economy, paying with a gift card.

  • Database must match the gold state: update_reservation_flights (economy upgrade), update_reservation_passengers, and update_reservation_baggages with exact arguments
  • communicate_info is empty, so no string check applies
db ✓
communicate ✓
USER_STOP

The agent looked up the user, found the right reservation among several, confirmed the changes and payment method, then made all three writes: passenger swap, cabin upgrade, and bags. The final database hashes match the gold state and the run ends on USER_STOP, so reward is 1. This is what the eval is designed to measure: multi-step tool use under policy constraints, done correctly.

reward = 0

Failure: cancels a reservation the policy forbids

Task 0 agent: openai/gpt-4o-mini

Cancel reservation EHGLP3. The booking is more than 24 hours old, basic economy, no travel insurance, so policy says no cancellation.

  • Database must end unmodified (the gold solution makes no changes)
  • NL assertion (not machine-checked): agent should refuse the cancellation
db ✗
communicate ✓
USER_STOP

The agent pulled the reservation, saw a basic economy fare booked more than 24 hours ago with no insurance, and cancelled it anyway. It even promised a $208 refund. The checker compares the final database against the gold database (unmodified), the hashes differ, reward is 0. The same model refused this exact cancellation in a different epoch; sampling variance flips the outcome.

reward = 0

Failure: right database, wrong number to the user

Task 18 agent: openai/gpt-4o-mini

Downgrade all five business-class reservations to economy, then report the total amount saved. The correct total is $23,553.

  • Database must match the gold state: 5 update_reservation_flights calls with exact cabin, flights, and payment IDs
  • communicate_info: the string "23553" must appear in an agent message
db ✓
communicate ✗
USER_STOP

The agent executed all five downgrades correctly, and the database check passed. Then it computed the savings from a partial list of fares and told the user $14,965 instead of $23,553. The communicate check greps every agent message for the literal string "23553", finds nothing, and zeroes the whole task. One wrong arithmetic answer erased five correct database writes.

reward = 0

Failure: invents tool arguments, then loops on the same rejected call

Task 11 agent: meta-llama/llama-3.1-8b-instruct

Remove passenger Sophia from reservation GV1N64. Policy forbids changing the passenger count; the correct move is to downgrade both passengers to basic economy and refund $5,244.

  • Database must match the gold state: one update_reservation_flights call downgrading to basic_economy
  • communicate_info: the string "5244" must appear in an agent message
db ✗
communicate ✗
USER_STOP

The agent’s very first move is a write with fully invented arguments: a reservation ID the user never gave and a passenger "John Doe" born 1990-01-01 who exists nowhere in the data. After finding the real reservation it sends update_reservation_passengers with a one-passenger array against a two-passenger booking, gets "number of passengers does not match", and retries the identical call four more times, at one point hallucinating "Sophia Smith, 1992-01-01" as the second passenger. It never consults the policy (removing a passenger is forbidden; the correct move is a cabin downgrade), quotes a fabricated $200 refund, and both checks fail. The arguments are well-formed JSON that the tool schema accepts; they are just wrong about the world, which is why small models can look fine on schema-level tool-call metrics and still fail tasks like this.

reward = 0

Failure: retries the same broken tool call until the run dies

Task 14 agent: openai/gpt-4o-mini

Rebook the cheapest business round trip and split payment across gift cards, one certificate, and a credit card. The correct split puts $44 on the card.

  • Database must match the gold state (cancel + rebook with the exact payment split)
  • communicate_info: gift card and certificate sums, and the final card charge
db ✗
communicate ✗
MAX_STEPS

The agent built a booking where the payment amounts did not sum to the ticket price. The tool rejected it with the same error every time, and the agent retried the identical call dozens of times until the run hit its 200-step ceiling. Any run that terminates on MAX_STEPS scores 0 before the database is even compared. This transcript is 202 messages long; the excerpt below is the loop.

reward = 1

Pass without doing anything: refusal tasks reward inaction

Task 0 agent: openai/gpt-4o-mini

Same task as the policy-break failure: the user wants to cancel EHGLP3, and policy says no.

  • Gold action list is empty, so the database check passes as long as nothing changed
  • communicate_info is empty, so the communicate check passes vacuously
  • The NL assertion ("agent should refuse") is metadata only and never machine-checked
db ✓
communicate ✓
USER_STOP

This run earned a legitimate pass: the agent checked the reservation, cited the 24-hour rule, and refused. But look at what the checker actually verified: an untouched database and an empty communicate list. An agent that stonewalled every request, or transferred to a human immediately, would score identically. Refusal tasks measure "did nothing break", so they inflate scores for overly cautious models.

Methodology

Scores aggregate successful runs from the newest 90 days, weighted by task count, with a minimum of 45 graded tasks per model-provider pair. A model's headline score uses its default routing (not pinned to a provider) when one exists; otherwise it falls back to the median provider. The standard deviation is measured across runs for that representative result. Cost, time, and token figures are per-task averages from the same runs. Best value is the cheapest Pareto-optimal model within 5 points of the top score.

These are the same measurements that power Exacto routing. See the docs for how routing works, or browse all models to try one.

API access

These scores are available through OpenRouter's public benchmarks API, so you can retrieve the same model-level results programmatically.

GET https://openrouter.ai/api/v1/benchmarks?source=openrouter
Authorization: Bearer <API key>

Use task_type=agentic to filter to tau_bench_verified_airline. Each item represents one model and includes accuracy, accuracy_stddev, avg_cost_per_task, total_tasks, and last_run_timestamp. See the benchmarks API docs.

Download and cite

A snapshot of this leaderboard is published in the benchmark-leaderboard open dataset (filter on the benchmark column), regenerated daily with no API key required. The current snapshot was generated on Sep 29, 2026, 1:45 AM UTC. It is licensed under CC BY 4.0, so you can reuse and republish it with attribution to OpenRouter.

  • Download CSV (text/csv)
  • Download JSON (application/json)
  • Snapshot manifest (SHA-256 checksums of the files above)
  • Latest manifest (mutable alias, updated as snapshots are published)

Column definitions, methodology and redaction rules are in the dataset README.

Cite this leaderboard

OpenRouter (2026). τ²-Bench: Airline results on OpenRouter, 2026-09-29 snapshot. https://openrouter.ai/benchmarks/tau2-bench-airline. Licensed under CC BY 4.0.

Frequently asked questions

It tests whether a model can do an airline support agent’s job: follow a policy manual, talk to a simulated customer, and call the right tools to search flights, change bookings, and issue refunds. The model needs no domain knowledge, and is scored entirely on how well it executes tool calls across multi-step task trajectories.

A task is graded pass or fail. The agent has to finish the conversation within its step budget, leave the booking system in the state the reference solution produces, and tell the customer what it did; missing any one of those scores zero, and partial credit is not awarded.

Runs execute against the same provider endpoints that serve OpenRouter traffic, so a score reflects both the model and the provider running it. These accuracy scores are one of the signals Exacto routing uses to steer tool-calling traffic away from endpoints that fall behind their peers.

The customer side of every conversation is played by another model, and OpenRouter keeps that simulator fixed so agent scores stay comparable across runs. The simulator has failure modes of its own, so absolute scores shift whenever it changes, which is why τ²-Bench numbers from different sources rarely line up exactly.

Scores aggregate repeated runs of the same task set, weighted by how many tasks each run graded. A model’s headline score is a single representative result rather than its best-performing provider.

Yes. The public benchmarks API returns the same model-level results from GET https://openrouter.ai/api/v1/benchmarks?source=openrouter with an API key. Filter with task_type=agentic to reach tau_bench_verified_airline.