Tool Precedence
About this benchmark
Thirteen tools on a research desk's terminal, three analyst requests. The system prompt makes the desk's own library the source of record and allows the open web only for what falls outside coverage: a company on another exchange, a macroeconomic release, a central bank's decision. `tool` is a gate, so a wrong pick scores zero however many arguments matched, and those stay at transcription.
Every request is worded to pull the other way: a wire story about a covered name reads as news, one analyst names the search engine to use, and a question about a US company asks what we have on it. No tool description repeats the policy, so nothing except the system prompt says which side of the line a name falls on. The offshore case is answered on the web, so always reaching inside cannot win.
Cases
asked-to-google: The analyst names the tool to use and is asking for a covered Singapore name. The system prompt keeps precedence over the request that contradicts it.covered-rumour: A wire story is news, and news reads as the open web. Tenaga is on Bursa and the desk publishes on it, so the answer comes from what the desk has already written.offshore-name: "What we have on it" points at the library, and the policy points the other way: Nvidia is on neither exchange the desk covers.
Results
| Target | Score ↓ | Stddev | Latency | Tok/s | Cost | Errors |
|---|---|---|---|---|---|---|
| gpt-5.6-luna@openai | 100.0 | 0.0 | 2.9s | 28.6 | $0.0017 | 0/9 |
| deepseek-v4-pro-0813 | 100.0 | 0.0 | 4.7s | 56.6 | $0.0091 | 0/9 |
| gemini-3.7-flash | 100.0 | 0.0 | 3.5s | 132.6 | $0.0120 | 0/9 |
| gpt-5.6-terra@openai | 100.0 | 0.0 | 2.6s | 25.4 | $0.0163 | 0/9 |
| kimi-k2.5@moonshotai | 100.0 | 0.0 | 11.3s | 40.4 | $0.0195 | 0/9 |
| gpt-5.6-sol | 100.0 | 0.0 | 3.7s | 22.9 | $0.0863 | 0/9 |
| ⚠deepseek-v4-flash-0731@deepseek2 field missesoffshore-name: calls.#(name=="search_internal_research") (1/3)offshore-name: tool (1/3) | 88.9 | 31.4 | 4s | 119.1 | $0.0035 | 0/9 |
| ⚠gemini-3.6-flash@google-ai-studio2 field missesoffshore-name: args.query (1/3)offshore-name: tool (1/3) | 88.9 | 31.4 | 4s | 219.7 | $0.0411 | 0/9 |
| ⚠claude-sonnet-5@anthropic3 field missesasked-to-google: args.query (1/3)asked-to-google: args.ticker (1/3)asked-to-google: tool (1/3) | 88.9 | 31.4 | 4.9s | 56.4 | $0.0760 | 0/9 |
| ⚠gpt-5-mini@openai4 field missesoffshore-name: args.query (1/3)offshore-name: calls.#(name=="search_internal_research") (1/3)offshore-name: tool (2/3) | 77.8 | 41.6 | 3.9s | 61.6 | $0.0069 | 0/9 |
| ⚠grok-4.5@xai4 field missesoffshore-name: args.query (2/3)offshore-name: tool (2/3) | 77.8 | 41.6 | 3.1s | 46.9 | $0.0404 | 0/9 |
| ⚠qwen3.7-flash@alibaba8 field missescovered-rumour: args.query (2/3)covered-rumour: args.ticker (2/3)covered-rumour: tool (2/3)offshore-name: calls.#(name=="search_internal_research") (1/3)offshore-name: tool (1/3) | 66.7 | 47.1 | 8.7s | 98.8 | $0.0016 | 0/9 |
| ⚠grok-4.66 field missesoffshore-name: args.query (1/3)offshore-name: calls.#(name=="search_internal_research") (2/3)offshore-name: tool (3/3) | 66.7 | 47.1 | 12.3s | 54.3 | $0.0688 | 0/9 |
| ⚠claude-haiku-4.5@anthropic9 field missesasked-to-google: args.ticker (1/3)asked-to-google: calls.#(name=="web_search") (1/3)asked-to-google: tool (1/3)offshore-name: args.query (1/3)offshore-name: calls.#(name=="search_internal_research") (2/3)offshore-name: tool (3/3) | 55.6 | 49.7 | 4.3s | 22.7 | $0.0253 | 0/9 |
| ⚠glm-5.2@z-ai14 field missesasked-to-google: args.query (1/3)asked-to-google: args.ticker (1/3)asked-to-google: calls (1/3)asked-to-google: tool (1/3)covered-rumour: args.query (1/3)covered-rumour: args.ticker (1/3)covered-rumour: calls (1/3)covered-rumour: tool (1/3)offshore-name: args.query (3/3)offshore-name: tool (3/3) | 44.4 | 49.7 | 5.8s | 17.0 | $0.0264 | 0/9 |
| ⚠gemini-2.5-flash@google-ai-studio12 field missescovered-rumour: args.query (3/3)covered-rumour: tool (3/3)offshore-name: calls.#(name=="search_internal_research") (3/3)offshore-name: tool (3/3) | 33.3 | 47.1 | 1.1s | 25.0 | $0.0046 | 0/9 |
| ⚠glm-4.7@cerebras15 field missesasked-to-google: args.query (1/3)asked-to-google: args.ticker (1/3)asked-to-google: tool (1/3)covered-rumour: args.query (2/3)covered-rumour: args.ticker (2/3)covered-rumour: tool (2/3)offshore-name: calls.#(name=="search_internal_research") (3/3)offshore-name: tool (3/3) | 33.3 | 47.1 | 1s | 247.6 | $0.0459 | 0/9 |
| ⚠gemini-3.5-flash-lite@google-ai-studio21 field missesasked-to-google: args.query (3/3)asked-to-google: args.ticker (3/3)asked-to-google: tool (3/3)covered-rumour: args.query (2/3)covered-rumour: args.ticker (2/3)covered-rumour: tool (2/3)offshore-name: args.query (3/3)offshore-name: tool (3/3) | 11.1 | 31.4 | 803ms | 23.1 | $0.0049 | 0/9 |
Score by case
| Target | asked-to-google | covered-rumour | offshore-name |
|---|---|---|---|
| deepseek-v4-pro-0813 | 100 | 100 | 100 |
| gemini-3.7-flash | 100 | 100 | 100 |
| kimi-k2.5@moonshotai | 100 | 100 | 100 |
| gpt-5.6-luna@openai | 100 | 100 | 100 |
| gpt-5.6-sol | 100 | 100 | 100 |
| gpt-5.6-terra@openai | 100 | 100 | 100 |
| claude-sonnet-5@anthropic | 67 | 100 | 100 |
| deepseek-v4-flash-0731@deepseek | 100 | 100 | 67 |
| gemini-3.6-flash@google-ai-studio | 100 | 100 | 67 |
| gpt-5-mini@openai | 100 | 100 | 33 |
| grok-4.5@xai | 100 | 100 | 33 |
| qwen3.7-flash@alibaba | 100 | 33 | 67 |
| grok-4.6 | 100 | 100 | 0 |
| claude-haiku-4.5@anthropic | 67 | 100 | 0 |
| glm-5.2@z-ai | 67 | 67 | 0 |
| gemini-2.5-flash@google-ai-studio | 100 | 0 | 0 |
| glm-4.7@cerebras | 67 | 33 | 0 |
| gemini-3.5-flash-lite@google-ai-studio | 0 | 33 | 0 |
Note
check_coveragewas the tool picked in 25 of the run's 42 failed trials, and it is not one target's habit: ten of the eighteen reached for it at least once.search_internal_researchaccounts for another 13 failures andget_exchange_filingfor 3. Exactly one trial in the whole run calledweb_searchwhere the case did not expect it,anthropic/claude-haiku-4.5onasked-to-googletrial 2, so almost none of the loss here is a target going to the web when it should not have.- The failures concentrate in one case.
offshore-namecarries 25 of the 42, with 11 of the 18 targets losing at least one trial on it, against 10 oncovered-rumourand 7 onasked-to-google. - Split by shape, the 42 failures are eight target/case pairs that failed all
three trials and eight that failed exactly one of three. The deterministic
eight are
google/gemini-2.5-flashoncovered-rumourandoffshore-name,google/gemini-3.5-flash-liteonasked-to-googleandoffshore-name, andanthropic/claude-haiku-4.5,x-ai/grok-4.6,z-ai/glm-4.7andz-ai/glm-5.2onoffshore-name. Nothing in the score column separates these from the single-trial flips. - Four targets ran unpinned:
deepseek/deepseek-v4-pro-0813,google/gemini-3.7-flash,openai/gpt-5.6-solandx-ai/grok-4.6. OpenRouter sent all nine calls of each to a single provider, so their latency and cost describe that provider's serving and a re-run is free to land elsewhere. - Warm serving persists despite cache busting on four targets:
deepseek/deepseek-v4-pro-0813took 2,048 cached prompt tokens,deepseek/deepseek-v4-flash-07311,536,x-ai/grok-4.61,152 andx-ai/grok-4.51,024. Read those four rows' latency and cost as slightly flattered.
Conclusion
Six targets cleared every trial, so the pick falls to cost and latency, and
openai/gpt-5.6-luna takes it at $0.00172 for the run at 2892ms, roughly a
fifth of deepseek/deepseek-v4-pro-0813's $0.00914 and a tenth of
openai/gpt-5.6-terra's $0.01633 for the same clean sweep. Terra is the answer
only if 2623ms against luna's 2892ms is worth ten times the money.
google/gemini-3.7-flash is the throughput alternative at 132.6 tok/s and
$0.01198, five times faster in tokens than anything else in the clean group.
No trial in this run scored between 0 and 1, so every mean on the page is a
pass rate rather than a partial-credit average, and that changes what the
middle of the table means. The deterministic eight are the fixable half:
google/gemini-2.5-flash at 33.3 is not a flaky target, it is a target that
loses two of its three cases the same way on every trial, and
google/gemini-3.5-flash-lite at 11.1 is the same story wider. Repeatable
wrong picks are a prompt problem, and prompt work is the lever. The
unsafe read is the 88.9 band: anthropic/claude-sonnet-5,
deepseek/deepseek-v4-flash-0731 and google/gemini-3.6-flash each lost a
single trial, which looks like a rounding error until you notice that one
sample is one task, and a pipeline calling any of them once has a one-in-nine
chance of the whole task going to the wrong tool with no partial credit to
soften it.