Did your model quietly change?
On a weekly schedule we run the same 30 graded tasks against each model. Graders are deterministic: exact answers, formats and arithmetic, never another model’s opinion.
Leaderboard
| Model | Provider | Score | 14-day average | Change | Last run |
|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | 100.0 | 100.0 | +0.0 | 20 min ago |
| Claude Sonnet 5.5 | Anthropic | 91.7 | 91.7 | +0.0 | 80 min ago |
| Claude Haiku 5.5 | Anthropic | 90.0 | 90.0 | +0.0 | 20 min ago |
| Claude Opus 5.5 | Anthropic | 88.3 | 88.3 | +0.0 | 79 min ago |
Hear about it first
Drift alerts are part of Pro: one message when a model you depend on shifts.