Guideline Lint

2026-09-26 06:06 UTC

About this benchmark

Staff write a page of bullets that Bumi Niaga Bank's support bot reads on every chat. A linter checks it one bullet at a time, so each case is a single bullet and its section heading, not the page. The model names the one writing rule it breaks, or `none`: addressed to staff, too long, only prohibits, restates the base prompt, or holds a credential or personal data. Exact label scores one.

Five controls break nothing but look as if they do: a bullet naming a team, two related steps in one bullet, a "never" with an action attached, a specific apology, and a public hotline. A repeated bullet needs the whole page to spot, so that rule is left out. The grade ignores confidence, so a hedged right answer scores as a sure one.

Cases

  • chained-steps: Seven loosely related steps chained into one bullet.
  • customer-nric: A named customer and her NRIC number; no password to spot.
  • named-team: Control: names a team, but the instruction is to the bot.
  • never-then-do: Control: opens with "never" but gives the action to take.
  • only-avoid: Three prohibitions and nothing for the bot to do instead.
  • portal-password: A back-office login and password inside a bullet every chat sees.
  • public-hotline: Control: a public business number is not personal data.
  • related-pair: Control: two steps, but closely related actions may share a bullet.
  • restates-base: Tone and no-fabrication, which the base prompt already covers.
  • specific-apology: Control: about tone, but specific to delayed transfers and tied to a tool.
  • staff-voice: Addresses support agents about the bot instead of instructing the bot.

Results

12/12 targets scored · 55 trials each · $0.44 total · show

Cost vs latencymost attractive quadrant
Latency & costsort by
TargetScore ↓StddevLatencyTok/sCostErrors
bestjev-1.13@typesafe100.00.0best332ms200.5best$0.00170/55
OpenAIgpt-6-luna@openai100.00.02.2s21.5$0.00360/55
OpenAIgpt-5.6-luna@openai100.00.01.9s7.4$0.00560/55
Z.aiglm-5.3-flash100.00.02.1s82.9$0.00800/55
DeepSeekdeepseek-v4.1-flash100.00.02.3s78.3$0.01880/55
Claudeclaude-haiku-4.5@anthropic100.00.01.3s7.6$0.03580/55
Geminigemini-3.8-flash@google-ai-studio100.00.02.5s83.2$0.05900/55
mimo-v2.6-pro-ultraspeed100.00.08.3s7.1$0.12220/55
MoonshotAIkimi-k2.6100.00.018s40.9$0.13640/55
⚠Geminigemini-3.5-flash-lite@google-ai-studio98.213.41.1s10.2$0.00810/55
⚠Qwenqwen3.8-flash@alibabaerrorPOST "https://openrouter.ai/api/v1/chat/completions": 429 Too Many Requests {"message":"Provider returned error","code":429,"metadata":{"raw":"qwen/qwen3.8-flash is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"Alibaba","is_byok":false,"limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing"}}97.515.612s16.2$0.006315/55
⚠Z.aiglm-5.3-flashx@z-ai0.00.03.7s85.3$0.02960/55

Score by case

Every cell is that target's mean over the case's trials.

Targetchained-stepscustomer-nricnamed-teamnever-then-doonly-avoidportal-passwordpublic-hotlinerelated-pairrestates-basespecific-apologystaff-voice
Claudeclaude-haiku-4.5@anthropic100100100100100100100100100100100
DeepSeekdeepseek-v4.1-flash100100100100100100100100100100100
Geminigemini-3.8-flash@google-ai-studio100100100100100100100100100100100
MoonshotAIkimi-k2.6100100100100100100100100100100100
OpenAIgpt-5.6-luna@openai100100100100100100100100100100100
OpenAIgpt-6-luna@openai100100100100100100100100100100100
jev-1.13@typesafe100100100100100100100100100100100
mimo-v2.6-pro-ultraspeed100100100100100100100100100100100
Z.aiglm-5.3-flash100100100100100100100100100100100
Geminigemini-3.5-flash-lite@google-ai-studio10010010080100100100100100100100
Qwenqwen3.8-flash@alibaba10010010010010010010010010080100
Z.aiglm-5.3-flashx@z-ai00000000000

Note

  • z-ai/glm-5.3-flashx@z-ai scores zero, but its labels are right: the first line of all 55 trials names the expected label. None of them is the requested {"label": ...} object. Fourteen are a bare word; the other 41 put the label on its own line, mostly bolded, and follow it with an explanation. The row measures output shape, not judgement. It ran pinned to @z-ai only; this run did not try free routing.
  • qwen/qwen3.8-flash@alibaba errored on 15 of 55 trials, every one a 429 rate limit from the pinned Alibaba provider. Its 97.5 is the mean of the 40 trials that returned. Its one wrong answer is specific-apology on trial 5, labelled restates_base instead of none.
  • google/gemini-3.5-flash-lite@google-ai-studio missed one trial: never-then-do on trial 5, labelled only_avoid instead of none. The other four trials of that case were right.
  • typesafe/jev-1.13@typesafe picked the right label on all 55 trials, but its probabilities show two close calls. On chained-steps it gives too_long 0.68 to 0.74, with none taking 0.25 to 0.31. On named-team it gives none 0.61 to 0.75, with staff_voice taking 0.23 to 0.37. Every other case comes back at 0.86 or higher.

Conclusion

Nine targets swept every trial, so cost and latency decide it, and the classifier wins both. typesafe/jev-1.13 averages 332ms and costs $0.0017 for 55 trials. The fastest perfect chat model, anthropic/claude-haiku-4.5, takes 1,302ms for $0.0358. The cheapest, openai/gpt-6-luna, costs $0.0036 but takes 2,159ms. Pick Jev. Among the chat models, haiku-4.5 is the choice when latency matters and gpt-6-luna when cost does.

Jev's label is right every time, but a linter that only flags above a probability cutoff would act on the probability, not the label. At a 0.7 cutoff, chained-steps goes unflagged on one trial in five (0.68), so a real violation slips through. Set the cutoff below 0.68 or treat 0.6 to 0.75 as needing review. The two single misses, from gemini-3.5-flash-lite and qwen3.8-flash, are one-off false flags on controls, not a field that fails every time. GLM flashx is a shape problem and Qwen's row is mostly rate limiting; neither mean describes judgement.