The tasks are dull on purpose.
Read a receipt. Pull the fields off a support email. Lift the header facts out of a research PDF. This is the work a business runs on all day, and nothing here is designed to defeat anyone.
$ bytemark --help
Frontier evals measure how well the most expensive models handle the hardest problems on earth. Almost no production work is hard. It is a receipt, a support ticket, a PDF nobody wants to read. ByteMark checks which models get that job exactly right, then ranks the ones that did on seconds and cents.
$ bytemark -eval plate-recognition -trials 20
9/14 models scored a flat 100% on every field, every trial. sorted by cost. last, the worst score of the same run.
same answer. 39.5× the price, 9.6× the wait.
Read a receipt. Pull the fields off a support email. Lift the header facts out of a research PDF. This is the work a business runs on all day, and nothing here is designed to defeat anyone.
9 of 14 models scored a flat 100% on every field, every trial. A leaderboard sorted by score is a 9-way tie, and tells you nothing about which one to ship.
Those tied models are 39.5× apart on the bill and 9.6× apart on the wait. The crown goes to the cheapest model that got everything right. Being frontier earns nothing here.
Extraction
Receipt ParseReceipt fields from a one-page PDFpdfgpt-5.6-luna9/14 clean at 100%$0.00043.9sReport ExtractEquity research report details from multi-page PDFspdfgpt-5.6-luna6/13 clean at 100%$0.00284.8sTicket ExtractSupport ticket fields from plain-text customer messagestextgemini-3.5-flash-lite7/10 clean at 100%$0.00061.2sPlate RecognitionOne registration number read off a carpark camera frameimageqwen3.7-flash11/16 clean at 100%$0.00013.3sTool Use
Tool CallWhether a banking assistant reaches for the right tool when a near-duplicate sits beside itsingle-turngpt-5.6-luna14/16 clean at 100%$0.00012.6sTool HaystackWhether the right tool is found in a list of fifty where nine others look like itsingle-turnqwen3.7-flash14/14 clean at 100%$0.00033.4sTool PrecedenceWhether a stated order of preference beats the tool the request sounds likesingle-turngpt-5.6-luna6/18 clean at 100%$0.00022.9sTool LoopWhether a model chains calls, threads the identifiers between them, and stopsmulti-turnqwen3.7-flash11/15 clean at 100%$0.00057.8sAdversarial
Prompt LeakWhether a support bot's system prompt survives five extraction attemptssystem promptqwen3.7-flash14/15 clean at 0%$0.00016.8sPersona HoldWhether a support bot stays in character under bait and hostilityuser promptqwen3.7-flash7/19 clean at 100%$0.000210.9sPrompt InjectionInstructions hidden in the document a model was asked to readdocumentgpt-5.6-luna8/21 clean at 100%$0.00012.5sPlanned