Skip to content

Methodology: how we measure Super Intelligence uptime

Exactly what we send to each service, how often, what counts as an outage, and what we do not claim.

We measure Super Intelligence services: model APIs, gateways, inference platforms, apps and tools, and media and infrastructure services.

Three signals, always labelled

  • Measured: our own real API calls. This is the only signal that drives the headline status when we have it.
  • Official: what the provider’s own status page says. It is shown beside our measurement, and it drives the headline only for services we do not call ourselves.
  • User reports: people pressing “I’m having a problem”. A spike raises a flag; it never declares an outage on its own. A spike is at least 10 reports in an hour and 3 times the 7-day average for that hour.

What a probe is

  • Liveness: a request to the provider’s models-list endpoint. It uses no tokens and answers one question: is the API reachable?
  • Inference: a streaming request with the prompt Reply with exactly: OK, at most 5 output tokens (16 or 256 for models that reason before they answer). We time the first content token (time to first token), read until the reply is complete, and pass the probe only if the reply contains OK.
  • Website check: a request for a public page of a service we cannot call, such as an app. It shows that the site answers, not that the product works, and every page labels it as a website check.
  • Every probe has a 30 s timeout and is never retried. A retry would hide exactly the failures we are here to see.
  • At most 2 requests are in flight to any one host, and probes are spread across each minute so they never arrive in a burst.

How often

This instance runs the lean profile: API reachability every 60 s, a small model every 5 min, a flagship model every 15 min, official status pages every 2 min and website checks every 2 min. When an API check or a model call fails, an official page reports an incident, or user reports spike, we call that service’s models every 60 s until 30 min pass without a new trigger. At 80% of the daily budget the intervals between model calls double.

What the statuses mean

  • Operational: at least 98% of calls in the last 5 min succeeded, and p95 first-token time is at most 2 times the 7-day baseline.
  • Degraded: 90% to 98% success, or p95 above 2 times the baseline, or more than 2% wrong replies.
  • Partial outage: 50% to 90% success, or only some models or regions failing.
  • Major outage: under 50% success for 3 consecutive windows. Under 50% for fewer windows shows Partial outage, and a service is in Major outage only when every model we call is.

A window needs at least 2 failed calls before its success rate counts: one failed call alone never changes a status. A window only counts as a new bad window when it contains a result we have not counted before; one failed call is never counted three times. With liveness data alone we can say Operational or Major outage, never Degraded: judging latency needs real inference. The latency rule needs a baseline first: 30 calls over 6 hours of history. Until then a model is judged on success alone.

Uptime is the share of our checks that succeeded. “Errors” is the rest.

Incidents and lead time

An incident opens after 2 consecutive bad windows and resolves after 10 min of operational checks. For a service we do not call, the incident opens from the provider’s status page and says so. If the provider’s status page acknowledges a problem while our incident is open, we record when, and the difference is the lead time. If the provider backdates its incident to before our detection, we claim no lead.

What we do not count

  • A 401, 402, 403 or 429 answer on our own key is our problem, not the provider’s. Those results are left out of success rates, and a model we cannot call for that reason shows no figure at all.
  • If our own scheduler stops, every page says “Monitoring degraded” instead of blaming providers.
  • A service or model we have no data for is not shown. It appears with its first result. A service we hold no API key for can still appear through its official page and user reports.
  • A check that could not be judged is left out of success rates: the provider refused our request as malformed, or a model that reasons before it answers used the whole output limit thinking. Neither says the provider is down.

Limits you should know

  • We measure from Mumbai. Your region may differ.
  • A very short answer says nothing about long generations, tool use or rate limits on your account.
  • Hourly and daily latency figures are built from per-minute percentiles: the p95 of an hour is the 95th percentile of its minutes’ p95 values, weighted by the number of calls. That is good for a chart, not a substitute for raw data.
  • Quality drift uses deterministic graders on a fixed set of 30 tasks. It detects change, not “intelligence”.
  • Providers whose terms forbid published benchmarks are excluded from comparisons and rankings.

Data marked “Sample data” is generated, not measured, or is measured data shown beside generated history. It exists so that a local or demo instance looks complete.