Veritas models, side by side
Which Veritas for which job.
Each Veritas model ran the same tasks, graded by code rather than by opinion. Here is which one to reach for, how far apart they really are, and how long each one makes you wait.
Small samples: most of these picks are close calls. When two models are too close to call, the quicker one is recommended.
Pick by task
The model to reach for first, for each kind of work. Under each pick are all four models: the share of tasks they got fully right and their typical time.
Head to head
Every model on every kind of task. The circled score is the pick. Hover a cell for the details.
Compared with Veritas-Pro
Veritas-Pro is the one everyone has. Each model is compared with it on the same tasks, so the gap is like for like. The whiskers show the likely range: when they cross zero, the difference could be luck.
Share right, compared with Veritas-Pro
Time taken, compared with Veritas-Pro
How long each one takes
Median seconds per task, from the request to the finished answer. The careful models check their own work before they answer, and that takes time.
Which plan has it
Prompts a day on the website for each model. Veritas-Experimental now comes with Pro.