Guideline Lint
About this benchmark
Staff write a page of bullets that Bumi Niaga Bank's support bot reads on every chat. A linter checks it one bullet at a time, so each case is a single bullet and its section heading, not the page. The model names the one writing rule it breaks, or `none`: addressed to staff, too long, only prohibits, restates the base prompt, or holds a credential or personal data. Exact label scores one.
Five controls break nothing but look as if they do: a bullet naming a team, two related steps in one bullet, a "never" with an action attached, a specific apology, and a public hotline. A repeated bullet needs the whole page to spot, so that rule is left out. The grade ignores confidence, so a hedged right answer scores as a sure one.
Cases
chained-steps: Seven loosely related steps chained into one bullet.customer-nric: A named customer and her NRIC number; no password to spot.named-team: Control: names a team, but the instruction is to the bot.never-then-do: Control: opens with "never" but gives the action to take.only-avoid: Three prohibitions and nothing for the bot to do instead.portal-password: A back-office login and password inside a bullet every chat sees.public-hotline: Control: a public business number is not personal data.related-pair: Control: two steps, but closely related actions may share a bullet.restates-base: Tone and no-fabrication, which the base prompt already covers.specific-apology: Control: about tone, but specific to delayed transfers and tied to a tool.staff-voice: Addresses support agents about the bot instead of instructing the bot.
Results
| Target | Score ↓ | Stddev | Latency | Tok/s | Cost | Errors |
|---|---|---|---|---|---|---|
| jev-1.13@typesafe | 100.0 | 0.0 | 332ms | 200.5 | $0.0017 | 0/55 |
| gpt-6-luna@openai | 100.0 | 0.0 | 2.2s | 21.5 | $0.0036 | 0/55 |
| gpt-5.6-luna@openai | 100.0 | 0.0 | 1.9s | 7.4 | $0.0056 | 0/55 |
| glm-5.3-flash | 100.0 | 0.0 | 2.1s | 82.9 | $0.0080 | 0/55 |
| deepseek-v4.1-flash | 100.0 | 0.0 | 2.3s | 78.3 | $0.0188 | 0/55 |
| claude-haiku-4.5@anthropic | 100.0 | 0.0 | 1.3s | 7.6 | $0.0358 | 0/55 |
| gemini-3.8-flash@google-ai-studio | 100.0 | 0.0 | 2.5s | 83.2 | $0.0590 | 0/55 |
| mimo-v2.6-pro-ultraspeed | 100.0 | 0.0 | 8.3s | 7.1 | $0.1222 | 0/55 |
| kimi-k2.6 | 100.0 | 0.0 | 18s | 40.9 | $0.1364 | 0/55 |
| ⚠gemini-3.5-flash-lite@google-ai-studio | 98.2 | 13.4 | 1.1s | 10.2 | $0.0081 | 0/55 |
| ⚠qwen3.8-flash@alibabaerrorPOST "https://openrouter.ai/api/v1/chat/completions": 429 Too Many Requests {"message":"Provider returned error","code":429,"metadata":{"raw":"qwen/qwen3.8-flash is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"Alibaba","is_byok":false,"limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing"}} | 97.5 | 15.6 | 12s | 16.2 | $0.0063 | 15/55 |
| ⚠glm-5.3-flashx@z-ai | 0.0 | 0.0 | 3.7s | 85.3 | $0.0296 | 0/55 |
Score by case
| Target | chained-steps | customer-nric | named-team | never-then-do | only-avoid | portal-password | public-hotline | related-pair | restates-base | specific-apology | staff-voice |
|---|---|---|---|---|---|---|---|---|---|---|---|
| claude-haiku-4.5@anthropic | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| deepseek-v4.1-flash | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| gemini-3.8-flash@google-ai-studio | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| kimi-k2.6 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| gpt-5.6-luna@openai | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| gpt-6-luna@openai | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| jev-1.13@typesafe | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| mimo-v2.6-pro-ultraspeed | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| glm-5.3-flash | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| gemini-3.5-flash-lite@google-ai-studio | 100 | 100 | 100 | 80 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| qwen3.8-flash@alibaba | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 80 | 100 |
| glm-5.3-flashx@z-ai | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
Note
z-ai/glm-5.3-flashx@z-aiscores zero, but its labels are right: the first line of all 55 trials names the expected label. None of them is the requested{"label": ...}object. Fourteen are a bare word; the other 41 put the label on its own line, mostly bolded, and follow it with an explanation. The row measures output shape, not judgement. It ran pinned to@z-aionly; this run did not try free routing.qwen/qwen3.8-flash@alibabaerrored on 15 of 55 trials, every one a 429 rate limit from the pinned Alibaba provider. Its 97.5 is the mean of the 40 trials that returned. Its one wrong answer isspecific-apologyon trial 5, labelledrestates_baseinstead ofnone.google/gemini-3.5-flash-lite@google-ai-studiomissed one trial:never-then-doon trial 5, labelledonly_avoidinstead ofnone. The other four trials of that case were right.typesafe/jev-1.13@typesafepicked the right label on all 55 trials, but its probabilities show two close calls. Onchained-stepsit givestoo_long0.68 to 0.74, withnonetaking 0.25 to 0.31. Onnamed-teamit givesnone0.61 to 0.75, withstaff_voicetaking 0.23 to 0.37. Every other case comes back at 0.86 or higher.
Conclusion
Nine targets swept every trial, so cost and latency decide it, and the
classifier wins both. typesafe/jev-1.13 averages 332ms and costs $0.0017
for 55 trials. The fastest perfect chat model, anthropic/claude-haiku-4.5,
takes 1,302ms for $0.0358. The cheapest, openai/gpt-6-luna, costs $0.0036
but takes 2,159ms. Pick Jev. Among the chat models, haiku-4.5 is the choice
when latency matters and gpt-6-luna when cost does.
Jev's label is right every time, but a linter that only flags above a
probability cutoff would act on the probability, not the label. At a 0.7
cutoff, chained-steps goes unflagged on one trial in five (0.68), so a
real violation slips through. Set the cutoff below 0.68 or treat 0.6 to 0.75
as needing review. The two single misses, from gemini-3.5-flash-lite and
qwen3.8-flash, are one-off false flags on controls, not a field that fails
every time. GLM flashx is a shape problem and Qwen's row is mostly rate
limiting; neither mean describes judgement.