Veritas Charts

Benchmarks

Veritas, measured the hard way.

Every Veritas score here comes from a graded run: real engines, public test sets and hidden tests, marked by code rather than by a model. The frontier models' scores are their published numbers, each with its source. Where nobody has published a number, the chart says so.

Against the frontier

These are the two public tests where Veritas can stand next to Claude, GPT and Gemini on exactly the same questions. The whiskers on the Veritas bars are 95% intervals: with 30 questions, a few points either way is noise.

Veritas, measured here Closed frontier models (published) Not measured yet Early: under a quarter measured

Frontier tests Veritas can't run yet

The labs lead with these now. They are here so the comparison isn't cherry-picked, with the reason each one isn't measured for Veritas.

Veritas Hard

Four suites built for what goes wrong in daily use: code that must pass hidden tests, instructions with every constraint checked, answers buried in long files, and pages checked for bugs and the AI-template look.

Each cell is the share of tasks fully passed. The small line underneath gives partial credit, meaning tests or constraints met. Hover a cell for details.

What it costs in time

Median seconds per task, from the request to the finished answer. Veritas checks its own work before it answers, and that takes time.

How this was measured

Rules

These hold for every number on this page

    The suites

    GPQA questions are never published, at the authors' request

    Sources