Benchmarks
Veritas, measured the hard way.
Every Veritas score here comes from a graded run: real engines, public test sets and hidden tests, marked by code rather than by a model. The frontier models' scores are their published numbers, each with its source. Where nobody has published a number, the chart says so.
Against the frontier
These are the two public tests where Veritas can stand next to Claude, GPT and Gemini on exactly the same questions. The whiskers on the Veritas bars are 95% intervals: with 30 questions, a few points either way is noise.
Frontier tests Veritas can't run yet
The labs lead with these now. They are here so the comparison isn't cherry-picked, with the reason each one isn't measured for Veritas.
Veritas Hard
Four suites built for what goes wrong in daily use: code that must pass hidden tests, instructions with every constraint checked, answers buried in long files, and pages checked for bugs and the AI-template look.
What it costs in time
Median seconds per task, from the request to the finished answer. Veritas checks its own work before it answers, and that takes time.