Anthropic API latency and performance
DegradedHow fast Anthropic answers, from real calls: time to first token and total response time for every model we probe. This is the speed of a Super Intelligence service as a caller sees it.
p95 first token · 1 hour502 ms
Fastest model · p95 first token · 24 hours502 msClaude Haiku 4.5
Performance by model, last 24 hours
p50 is the typical call; p95 is the slow one in twenty.
| Model | Time to first token · p50 | p95 | Total response time · p50 | p95 |
|---|---|---|---|---|
| Claude Haiku 4.5 | 462 ms | 502 ms | 498 ms | 538 ms |
Time to first token is the wait for the first word of a short streamed answer. Total response time is until that answer is complete. Both are measured from our probe, so your own latency also depends on where you call from and how long your answers are.
Last 24 hours
Time to first token
p50
p95
800 ms400 ms0 ms
81 min agoNow
Last 7 days
Time to first token
p50
p95
800 ms400 ms0 ms
2 hours agoNow
Last 30 days
Time to first token
p50
p95
800 ms400 ms0 ms
2 hours agoNow