Performance
What the OCR engines managed to read
Every night, with the box's other work paused so the whole GPU is free, each design is rendered at screen-share quality and run through the leading open-source OCR engines. Recall is the share of a question's words an engine recovered. Lower is better.
Every design, every engine
One row per design. Designs are numbered, not named, and no image is shown: the number is all a reader needs, and the picture is the one thing worth protecting. Each cell is mean recall over the run's synthetic questions; hover for the worst single reading.
Over the last 90 days
How the test works
Each design is rendered with a synthetic question from a fixed public set, never a customer's. The image is handed to every engine, which is given the whole GPU because the box's language-model work is stopped for the duration of the run and restarted after. We score word recall, not an exact match: a tool that recovers most of a question is a failure for us even if it garbles the rest.
A plain-text control slide goes through the same path in every run. It is meant to be readable, and the engines do read it: it is the proof that a low score on a design means the design defeated the engine, not that the engine was broken. Its result is shown on its own and never mixed into the averages.
The engines are the current open-source leaders (Tesseract, EasyOCR, PaddleOCR, docTR), run at their pretrained best. This is a different question from the daily readiness test, which puts the designs in front of general vision models; a dedicated OCR engine and a multimodal model fail in different ways, and a design has to beat both.
Numbers and aggregates only. Which design is which, and what any of them looks like, is not published here, because that list is what an attacker would want.