Veritas Charts

Veritas models, side by side

Which Veritas for which job.

Each Veritas model ran the same tasks, graded by code rather than by opinion. Here is which one to reach for, how far apart they really are, and how long each one makes you wait.

Small samples: most of these picks are close calls. When two models are too close to call, the quicker one is recommended.

Pick by task

The model to reach for first, for each kind of work. Under each pick are all four models: the share of tasks they got fully right and their typical time.

Head to head

Every model on every kind of task. The circled score is the pick. Hover a cell for the details.

Compared with Veritas-Pro

Veritas-Pro is the one everyone has. Each model is compared with it on the same tasks, so the gap is like for like. The whiskers show the likely range: when they cross zero, the difference could be luck.

Share right, compared with Veritas-Pro

Percentage points, on the tasks both models ran

Time taken, compared with Veritas-Pro

Median time per task, as a multiple of Veritas-Pro's

How long each one takes

Median seconds per task, from the request to the finished answer. The careful models check their own work before they answer, and that takes time.

Which plan has it

Prompts a day on the website for each model. Veritas-Experimental now comes with Pro.

How this was measured

Rules

These hold for every number on this page

    The tasks

    What each kind of work was tested with