Skip to content
Quality drift

Did your model quietly change?

On a weekly schedule we run the same 30 graded tasks against each model. Graders are deterministic: exact answers, formats and arithmetic, never another model’s opinion.

Leaderboard

ModelProviderScore14-day averageChangeLast run
Claude Fable 5.1 Anthropic 100.0 100.0 +0.0 20 min ago
Claude Sonnet 5.5 Anthropic 91.7 91.7 +0.0 80 min ago
Claude Haiku 5.5 Anthropic 90.0 90.0 +0.0 20 min ago
Claude Opus 5.5 Anthropic 88.3 88.3 +0.0 79 min ago
Score out of 100 · change against the average of the earlier runs in the last 14 days

Hear about it first

Drift alerts are part of Pro: one message when a model you depend on shifts.

See plans