Plate Verify

2026-10-08 04:56 UTC

About this benchmark

Each case is a 1483x796 carpark barrier frame and one candidate registration number. The model answers yes if the candidate matches the plate, ignoring spaces and case, or no otherwise. One label per trial, so the score is a plain success rate.

Four candidates are right and four are near misses an OCR pass makes: B read as 8, a stuttered extra digit, two digits swapped, and a two-line plate read bottom first. Five frames are night captures, and two right claims look wrong: a two-letter prefix and a spaced reading. Rubber-stamping every claim scores half.

Cases

  • bkh-line-order: BKH over 3187 on a two-line plate in fog, claimed with the lines read bottom first.
  • bll-b-as-8: BLL 7493 on a clear wet-floor frame, claimed with the leading B read as an 8.
  • bnx-spaced: BNX 6781 at night, claimed with the space the plate shows. Spacing is not a mismatch.
  • bry-extra-digit: BRY 590, a three-digit plate, claimed with a stuttered extra 0.
  • dfd-streak: DFD 3011 at night behind a vertical light streak, the trailing 11 in a narrow font. A correct claim on a hard frame.
  • vhd-transposed: VHD 4609 beside two lit headlamps, claimed with the last two digits swapped.
  • vn-two-letter: VN 6410 at night. A two-letter prefix that reads as a missing character.
  • wqh-daylight: WQH 4748 in daylight, the clean control.

Results

11/11 targets scored · 40 trials each · $0.52 total · show

Cost vs latencymost attractive quadrant
Latency & costsort by
TargetScore ↓StddevLatencyTok/sCostErrors
bestClaudeclaude-haiku-5.5@anthropic100.00.0best2.9s22.4best$0.00910/40
Geminigemini-3.8-flash@google-ai-studio100.00.06.6s31.0$0.06750/40
OpenAIgpt-6.1-sol@openai100.00.03.6s6.7$0.16790/40
⚠pplx-decider-v1.1-27b@perplexity97.515.61.1s0.9$0.00110/40
⚠DeepSeekdeepseek-v4.1-flash95.021.811.1s20.7$0.01780/40
⚠OpenAIgpt-6-sol@openai90.030.04.4s20.7$0.19450/40
⚠OpenAIgpt-6-luna@openai77.541.84s18.6$0.00940/40
⚠OpenAIgpt-6-luna-decisions@openai62.548.4675ms0.0$0.00560/40
⚠clef-flash60.049.01.3s0.0$0.01380/40
⚠Qwenqwen3.7-flash@alibaba57.549.42.7s9.4$0.00160/40
⚠clef50.050.02.2s0.0$0.02750/40

Score by case

Every cell is that target's mean over the case's trials.

Targetbkh-line-orderbll-b-as-8bnx-spacedbry-extra-digitdfd-streakvhd-transposedvn-two-letterwqh-daylight
Claudeclaude-haiku-5.5@anthropic100100100100100100100100
Geminigemini-3.8-flash@google-ai-studio100100100100100100100100
OpenAIgpt-6.1-sol@openai100100100100100100100100
pplx-decider-v1.1-27b@perplexity80100100100100100100100
DeepSeekdeepseek-v4.1-flash10010010010060100100100
OpenAIgpt-6-sol@openai8010010010010010040100
OpenAIgpt-6-luna@openai201001001001001000100
OpenAIgpt-6-luna-decisions@openai801001002001000100
clef-flash10010080100010000
Qwenqwen3.7-flash@alibaba8001008040060100
clef1001000100010000

Note

  • This run needed code that is not committed yet: image input on the classify loader and cloudflare/clef routed to the decisions endpoint. git_sha is 83fc84e-dirty, so no commit reproduces these numbers until that change lands.
  • cloudflare/clef answered no on all 40 trials, and its yes probability stayed between 0.13 and 0.21 on every case, right claim or wrong. cloudflare/clef-flash stayed between 0.33 and 0.55 and said yes only on bnx-spaced, 4 of 5 trials. Neither separates a right reading from a wrong one; their 50.0 and 60.0 are a constant answer, not partial skill. Both reported 16,384 prompt tokens on every trial.
  • deepseek/deepseek-v4.1-flash was routed to three providers. The 4 trials BaseTen served used 167 to 169 prompt tokens, so the image never reached the model and it answered from the text alone; DigitalOcean and Together used 868 to 876. Its 95.0 mixes both.
  • qwen/qwen3.7-flash said yes on all 5 trials of bll-b-as-8 and of vhd-transposed: it passed a wrong reading as correct 12 times in 20 trials of wrong claims.
  • openai/gpt-6-luna and openai/gpt-6-luna-decisions both rejected the correct vn-two-letter claim on all 5 trials; Luna Decisions did so at a yes probability of 0.06 to 0.08.

Conclusion

Three targets score 100. anthropic/claude-haiku-5.5 is the pick: $0.00023 a task at 2,945ms, against $0.0017 and 6,598ms for google/gemini-3.8-flash and $0.0042 and 3,558ms for openai/gpt-6.1-sol. Among decision models, perplexity/pplx-decider-v1.1-27b scores 97.5 at $0.000027 a task and 1,056ms. Its one miss was bkh-line-order at a yes probability of 0.53; sending anything between 0.3 and 0.7 to review would have caught it and flagged 3 of 40 trials, all on that case.

The Luna misses on vn-two-letter are fixable: the same case fails on every trial, so the error is predictable. Qwen 3.7 Flash is unsafe for this job, since its failures are wrong readings passed as correct, which is what a verifier is there to stop. Both Clef models are unusable here: a flat probability means the mean reflects the yes/no split of the corpus, not what the model saw.