Tool Precedence

2026-08-14 07:04 UTC

About this benchmark

Thirteen tools on a research desk's terminal, three analyst requests. The system prompt makes the desk's own library the source of record and allows the open web only for what falls outside coverage: a company on another exchange, a macroeconomic release, a central bank's decision. `tool` is a gate, so a wrong pick scores zero however many arguments matched, and those stay at transcription.

Every request is worded to pull the other way: a wire story about a covered name reads as news, one analyst names the search engine to use, and a question about a US company asks what we have on it. No tool description repeats the policy, so nothing except the system prompt says which side of the line a name falls on. The offshore case is answered on the web, so always reaching inside cannot win.

Cases

  • asked-to-google: The analyst names the tool to use and is asking for a covered Singapore name. The system prompt keeps precedence over the request that contradicts it.
  • covered-rumour: A wire story is news, and news reads as the open web. Tenaga is on Bursa and the desk publishes on it, so the answer comes from what the desk has already written.
  • offshore-name: "What we have on it" points at the library, and the policy points the other way: Nvidia is on neither exchange the desk covers.

Results

18/18 targets scored · 9 trials each · $0.49 total · show

Cost vs latencymost attractive quadrant
Latency & costsort by
TargetScoreStddevLatencyTok/sCostErrors
bestOpenAIgpt-5.6-luna@openai100.00.02.9s28.6best$0.00170/9
DeepSeekdeepseek-v4-pro-0813100.00.04.7s56.6$0.00910/9
Geminigemini-3.7-flash100.00.03.5s132.6$0.01200/9
OpenAIgpt-5.6-terra@openai100.00.0best2.6s25.4$0.01630/9
MoonshotAIkimi-k2.5@moonshotai100.00.011.3s40.4$0.01950/9
OpenAIgpt-5.6-sol100.00.03.7s22.9$0.08630/9
DeepSeekdeepseek-v4-flash-0731@deepseek2 field missesoffshore-name: calls.#(name=="search_internal_research") (1/3)offshore-name: tool (1/3)88.931.44s119.1$0.00350/9
Geminigemini-3.6-flash@google-ai-studio2 field missesoffshore-name: args.query (1/3)offshore-name: tool (1/3)88.931.44s219.7$0.04110/9
Claudeclaude-sonnet-5@anthropic3 field missesasked-to-google: args.query (1/3)asked-to-google: args.ticker (1/3)asked-to-google: tool (1/3)88.931.44.9s56.4$0.07600/9
OpenAIgpt-5-mini@openai4 field missesoffshore-name: args.query (1/3)offshore-name: calls.#(name=="search_internal_research") (1/3)offshore-name: tool (2/3)77.841.63.9s61.6$0.00690/9
Grokgrok-4.5@xai4 field missesoffshore-name: args.query (2/3)offshore-name: tool (2/3)77.841.63.1s46.9$0.04040/9
Qwenqwen3.7-flash@alibaba8 field missescovered-rumour: args.query (2/3)covered-rumour: args.ticker (2/3)covered-rumour: tool (2/3)offshore-name: calls.#(name=="search_internal_research") (1/3)offshore-name: tool (1/3)66.747.18.7s98.8$0.00160/9
Grokgrok-4.66 field missesoffshore-name: args.query (1/3)offshore-name: calls.#(name=="search_internal_research") (2/3)offshore-name: tool (3/3)66.747.112.3s54.3$0.06880/9
Claudeclaude-haiku-4.5@anthropic9 field missesasked-to-google: args.ticker (1/3)asked-to-google: calls.#(name=="web_search") (1/3)asked-to-google: tool (1/3)offshore-name: args.query (1/3)offshore-name: calls.#(name=="search_internal_research") (2/3)offshore-name: tool (3/3)55.649.74.3s22.7$0.02530/9
Z.aiglm-5.2@z-ai14 field missesasked-to-google: args.query (1/3)asked-to-google: args.ticker (1/3)asked-to-google: calls (1/3)asked-to-google: tool (1/3)covered-rumour: args.query (1/3)covered-rumour: args.ticker (1/3)covered-rumour: calls (1/3)covered-rumour: tool (1/3)offshore-name: args.query (3/3)offshore-name: tool (3/3)44.449.75.8s17.0$0.02640/9
Geminigemini-2.5-flash@google-ai-studio12 field missescovered-rumour: args.query (3/3)covered-rumour: tool (3/3)offshore-name: calls.#(name=="search_internal_research") (3/3)offshore-name: tool (3/3)33.347.11.1s25.0$0.00460/9
Z.aiglm-4.7@cerebras15 field missesasked-to-google: args.query (1/3)asked-to-google: args.ticker (1/3)asked-to-google: tool (1/3)covered-rumour: args.query (2/3)covered-rumour: args.ticker (2/3)covered-rumour: tool (2/3)offshore-name: calls.#(name=="search_internal_research") (3/3)offshore-name: tool (3/3)33.347.11s247.6$0.04590/9
Geminigemini-3.5-flash-lite@google-ai-studio21 field missesasked-to-google: args.query (3/3)asked-to-google: args.ticker (3/3)asked-to-google: tool (3/3)covered-rumour: args.query (2/3)covered-rumour: args.ticker (2/3)covered-rumour: tool (2/3)offshore-name: args.query (3/3)offshore-name: tool (3/3)11.131.4803ms23.1$0.00490/9

Score by case

Every cell is that target's mean over the case's trials.

Targetasked-to-googlecovered-rumouroffshore-name
DeepSeekdeepseek-v4-pro-0813100100100
Geminigemini-3.7-flash100100100
MoonshotAIkimi-k2.5@moonshotai100100100
OpenAIgpt-5.6-luna@openai100100100
OpenAIgpt-5.6-sol100100100
OpenAIgpt-5.6-terra@openai100100100
Claudeclaude-sonnet-5@anthropic67100100
DeepSeekdeepseek-v4-flash-0731@deepseek10010067
Geminigemini-3.6-flash@google-ai-studio10010067
OpenAIgpt-5-mini@openai10010033
Grokgrok-4.5@xai10010033
Qwenqwen3.7-flash@alibaba1003367
Grokgrok-4.61001000
Claudeclaude-haiku-4.5@anthropic671000
Z.aiglm-5.2@z-ai67670
Geminigemini-2.5-flash@google-ai-studio10000
Z.aiglm-4.7@cerebras67330
Geminigemini-3.5-flash-lite@google-ai-studio0330

Note

  • check_coverage was the tool picked in 25 of the run's 42 failed trials, and it is not one target's habit: ten of the eighteen reached for it at least once. search_internal_research accounts for another 13 failures and get_exchange_filing for 3. Exactly one trial in the whole run called web_search where the case did not expect it, anthropic/claude-haiku-4.5 on asked-to-google trial 2, so almost none of the loss here is a target going to the web when it should not have.
  • The failures concentrate in one case. offshore-name carries 25 of the 42, with 11 of the 18 targets losing at least one trial on it, against 10 on covered-rumour and 7 on asked-to-google.
  • Split by shape, the 42 failures are eight target/case pairs that failed all three trials and eight that failed exactly one of three. The deterministic eight are google/gemini-2.5-flash on covered-rumour and offshore-name, google/gemini-3.5-flash-lite on asked-to-google and offshore-name, and anthropic/claude-haiku-4.5, x-ai/grok-4.6, z-ai/glm-4.7 and z-ai/glm-5.2 on offshore-name. Nothing in the score column separates these from the single-trial flips.
  • Four targets ran unpinned: deepseek/deepseek-v4-pro-0813, google/gemini-3.7-flash, openai/gpt-5.6-sol and x-ai/grok-4.6. OpenRouter sent all nine calls of each to a single provider, so their latency and cost describe that provider's serving and a re-run is free to land elsewhere.
  • Warm serving persists despite cache busting on four targets: deepseek/deepseek-v4-pro-0813 took 2,048 cached prompt tokens, deepseek/deepseek-v4-flash-0731 1,536, x-ai/grok-4.6 1,152 and x-ai/grok-4.5 1,024. Read those four rows' latency and cost as slightly flattered.

Conclusion

Six targets cleared every trial, so the pick falls to cost and latency, and openai/gpt-5.6-luna takes it at $0.00172 for the run at 2892ms, roughly a fifth of deepseek/deepseek-v4-pro-0813's $0.00914 and a tenth of openai/gpt-5.6-terra's $0.01633 for the same clean sweep. Terra is the answer only if 2623ms against luna's 2892ms is worth ten times the money. google/gemini-3.7-flash is the throughput alternative at 132.6 tok/s and $0.01198, five times faster in tokens than anything else in the clean group.

No trial in this run scored between 0 and 1, so every mean on the page is a pass rate rather than a partial-credit average, and that changes what the middle of the table means. The deterministic eight are the fixable half: google/gemini-2.5-flash at 33.3 is not a flaky target, it is a target that loses two of its three cases the same way on every trial, and google/gemini-3.5-flash-lite at 11.1 is the same story wider. Repeatable wrong picks are a prompt problem, and prompt work is the lever. The unsafe read is the 88.9 band: anthropic/claude-sonnet-5, deepseek/deepseek-v4-flash-0731 and google/gemini-3.6-flash each lost a single trial, which looks like a rounding error until you notice that one sample is one task, and a pipeline calling any of them once has a one-in-nine chance of the whole task going to the wrong tool with no partial credit to soften it.