Skip to content

Anthropic API latency and performance

Degraded

How fast Anthropic answers, from real calls: time to first token and total response time for every model we probe. This is the speed of a Super Intelligence service as a caller sees it.

p95 first token · 1 hour502 ms
Fastest model · p95 first token · 24 hours502 msClaude Haiku 4.5

Performance by model, last 24 hours

p50 is the typical call; p95 is the slow one in twenty.

Model Time to first token · p50p95 Total response time · p50p95
Claude Haiku 4.5 462 ms 502 ms 498 ms 538 ms

Time to first token is the wait for the first word of a short streamed answer. Total response time is until that answer is complete. Both are measured from our probe, so your own latency also depends on where you call from and how long your answers are.

Last 24 hours

Time to first token p50 p95
800 ms400 ms0 ms
81 min agoNow

Last 7 days

Time to first token p50 p95
800 ms400 ms0 ms
2 hours agoNow

Last 30 days

Time to first token p50 p95
800 ms400 ms0 ms
2 hours agoNow

More on Anthropic