Model observatory
Model observatory
Live benchmarks of the leading AI models — and the part most comparisons skip: how they wire together into one system, each model placed in the role its numbers earn.
A snapshot of the Artificial Analysis leaderboard taken on the date above — real figures, refreshed by hand. Connect a data key to update them live.
01 · The index
Every model, on the same axes
Higher intelligence costs more and runs slower — almost always. That trade-off is the whole game. Sort the table, or read the frontier on the chart: up and to the left is smarter and cheaper.
| Model | |||||
|---|---|---|---|---|---|
01 Claude Fable 5 Anthropic | 60 | — | — | 7.70 | 1M |
02 Claude Opus 4.8 Anthropic | 56 | 65 | 32 | 3.85 | 1M |
03 GPT-5.5 OpenAI | 55 | 68 | 122 | 4.35 | 922k |
04 GLM-5.2 Z AI | 51 | 116 | 1.4 | 0.90 | 1M |
05 Gemini 3.5 Flash Google | 50 | 165 | 18 | 1.31 | 1M |
06 Claude Sonnet 4.6 Anthropic | 47 | 52 | 101 | 2.31 | 1M |
07 Gemini 3.1 Pro Google | 46 | 132 | 25 | 1.74 | 1M |
08 Qwen3.7 Max Alibaba | 46 | 198 | 2.5 | 1.43 | 1M |
09 GPT-5.3 Codex OpenAI | 44 | 90 | 84 | 1.87 | 400k |
10 MiniMax-M3 MiniMax | 44 | 84 | 3.6 | 0.22 | 1M |
11 DeepSeek V4 Pro DeepSeek | 44 | 78 | 1.7 | 0.18 | 1M |
02 · The orchestration
One system, seven roles, the right model in each
A real task isn't one prompt to one model. It's a pipeline — route, plan, reason, retrieve, code, write, review. Viviar assigns each step to the model whose benchmarks win it, and the indices above are the receipts. Hover a role to see the pick and the reason.
Breaks the goal into steps and assigns the sub-agents. The most demanding seat — it needs raw intelligence.
Selected for: Intelligence — 60
This is how AI becomes infrastructure, not a toy.
The same discipline behind this dashboard — the right execution layer for the right verified workflow — is how we build operational systems.