Plate Verify
About this benchmark
Each case is a 1483x796 carpark barrier frame and one candidate registration number. The model answers yes if the candidate matches the plate, ignoring spaces and case, or no otherwise. One label per trial, so the score is a plain success rate.
Four candidates are right and four are near misses an OCR pass makes: B read as 8, a stuttered extra digit, two digits swapped, and a two-line plate read bottom first. Five frames are night captures, and two right claims look wrong: a two-letter prefix and a spaced reading. Rubber-stamping every claim scores half.
Cases
bkh-line-order: BKH over 3187 on a two-line plate in fog, claimed with the lines read bottom first.bll-b-as-8: BLL 7493 on a clear wet-floor frame, claimed with the leading B read as an 8.bnx-spaced: BNX 6781 at night, claimed with the space the plate shows. Spacing is not a mismatch.bry-extra-digit: BRY 590, a three-digit plate, claimed with a stuttered extra 0.dfd-streak: DFD 3011 at night behind a vertical light streak, the trailing 11 in a narrow font. A correct claim on a hard frame.vhd-transposed: VHD 4609 beside two lit headlamps, claimed with the last two digits swapped.vn-two-letter: VN 6410 at night. A two-letter prefix that reads as a missing character.wqh-daylight: WQH 4748 in daylight, the clean control.
Results
| Target | Score ↓ | Stddev | Latency | Tok/s | Cost | Errors |
|---|---|---|---|---|---|---|
| claude-haiku-5.5@anthropic | 100.0 | 0.0 | 2.9s | 22.4 | $0.0091 | 0/40 |
| gemini-3.8-flash@google-ai-studio | 100.0 | 0.0 | 6.6s | 31.0 | $0.0675 | 0/40 |
| gpt-6.1-sol@openai | 100.0 | 0.0 | 3.6s | 6.7 | $0.1679 | 0/40 |
| ⚠pplx-decider-v1.1-27b@perplexity | 97.5 | 15.6 | 1.1s | 0.9 | $0.0011 | 0/40 |
| ⚠deepseek-v4.1-flash | 95.0 | 21.8 | 11.1s | 20.7 | $0.0178 | 0/40 |
| ⚠gpt-6-sol@openai | 90.0 | 30.0 | 4.4s | 20.7 | $0.1945 | 0/40 |
| ⚠gpt-6-luna@openai | 77.5 | 41.8 | 4s | 18.6 | $0.0094 | 0/40 |
| ⚠gpt-6-luna-decisions@openai | 62.5 | 48.4 | 675ms | 0.0 | $0.0056 | 0/40 |
| ⚠clef-flash | 60.0 | 49.0 | 1.3s | 0.0 | $0.0138 | 0/40 |
| ⚠qwen3.7-flash@alibaba | 57.5 | 49.4 | 2.7s | 9.4 | $0.0016 | 0/40 |
| ⚠clef | 50.0 | 50.0 | 2.2s | 0.0 | $0.0275 | 0/40 |
Score by case
| Target | bkh-line-order | bll-b-as-8 | bnx-spaced | bry-extra-digit | dfd-streak | vhd-transposed | vn-two-letter | wqh-daylight |
|---|---|---|---|---|---|---|---|---|
| claude-haiku-5.5@anthropic | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| gemini-3.8-flash@google-ai-studio | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| gpt-6.1-sol@openai | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| pplx-decider-v1.1-27b@perplexity | 80 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| deepseek-v4.1-flash | 100 | 100 | 100 | 100 | 60 | 100 | 100 | 100 |
| gpt-6-sol@openai | 80 | 100 | 100 | 100 | 100 | 100 | 40 | 100 |
| gpt-6-luna@openai | 20 | 100 | 100 | 100 | 100 | 100 | 0 | 100 |
| gpt-6-luna-decisions@openai | 80 | 100 | 100 | 20 | 0 | 100 | 0 | 100 |
| clef-flash | 100 | 100 | 80 | 100 | 0 | 100 | 0 | 0 |
| qwen3.7-flash@alibaba | 80 | 0 | 100 | 80 | 40 | 0 | 60 | 100 |
| clef | 100 | 100 | 0 | 100 | 0 | 100 | 0 | 0 |
Note
- This run needed code that is not committed yet: image input on the classify
loader and
cloudflare/clefrouted to the decisions endpoint.git_shais83fc84e-dirty, so no commit reproduces these numbers until that change lands. cloudflare/clefanswerednoon all 40 trials, and itsyesprobability stayed between 0.13 and 0.21 on every case, right claim or wrong.cloudflare/clef-flashstayed between 0.33 and 0.55 and saidyesonly onbnx-spaced, 4 of 5 trials. Neither separates a right reading from a wrong one; their 50.0 and 60.0 are a constant answer, not partial skill. Both reported 16,384 prompt tokens on every trial.deepseek/deepseek-v4.1-flashwas routed to three providers. The 4 trials BaseTen served used 167 to 169 prompt tokens, so the image never reached the model and it answered from the text alone; DigitalOcean and Together used 868 to 876. Its 95.0 mixes both.qwen/qwen3.7-flashsaidyeson all 5 trials ofbll-b-as-8and ofvhd-transposed: it passed a wrong reading as correct 12 times in 20 trials of wrong claims.openai/gpt-6-lunaandopenai/gpt-6-luna-decisionsboth rejected the correctvn-two-letterclaim on all 5 trials; Luna Decisions did so at ayesprobability of 0.06 to 0.08.
Conclusion
Three targets score 100. anthropic/claude-haiku-5.5 is the pick: $0.00023 a
task at 2,945ms, against $0.0017 and 6,598ms for google/gemini-3.8-flash and
$0.0042 and 3,558ms for openai/gpt-6.1-sol. Among decision models,
perplexity/pplx-decider-v1.1-27b scores 97.5 at $0.000027 a task and
1,056ms. Its one miss was bkh-line-order at a yes probability of 0.53;
sending anything between 0.3 and 0.7 to review would have caught it and
flagged 3 of 40 trials, all on that case.
The Luna misses on vn-two-letter are fixable: the same case fails on every
trial, so the error is predictable. Qwen 3.7 Flash is unsafe for this job,
since its failures are wrong readings passed as correct, which is what a
verifier is there to stop. Both Clef models are unusable here: a flat
probability means the mean reflects the yes/no split of the corpus, not what
the model saw.