Methods

How we make sure the numbers mean something.

Principles, every rule with its version and status, the limits of our data, and a log of every method change, including the ones that made us look worse.

Principles

Rules before results

A study's rule is written and versioned before its numbers are read. Any change creates a new version; old versions keep their results.

Controls must pass

Statistical tests include a negative control that must show nothing and a planted positive control that must be found. If either fails, the result is not interpreted.

Honest denominators

Rates come with their counts and 95% intervals. Missing values show as a dash, never as zero. Too-small samples get no verdict.

Negative results stay public

Retired rules, failed ideas and misses are listed with the same prominence as successes.

Reproducible

Fixed random seeds, open CSV tables under CC BY 4.0, and a JSON copy of every number shown.

Sources you can check

Every official incident links to the company's or provider's own page.

Rule registry

RuleStudyDefinitionStatusSince
cs-v1Upstream signalsCompanies sharing a provider declare incidents within 60 min; chance below 1 in 1,000 (Poisson-binomial on own rates, hour-of-week adjusted).Active2026-10-10
sf-v1Shared fateClose pairs of incident declarations within shared-provider groups vs. whole-week rotation null; negative and injection controls; Holm per provider.Active2026-10-10
sf-radar-v1Shared fate (own probes)Same test on five-minute reachability sweeps; verdict only with at least 10 pairs observed or expected.Supporting2026-10-10
val-v1Radar accuracyDOWN episodes (two failed sweeps) matched to official incidents ±30 min; Wilson intervals; bootstrap for lead time.Active2026-10-10
conc-v1ConcentrationHHI per layer over organisations (registrable domains); blast radius over hosting, CDN and DNS.Active2026-10-10
ew-v2Early warningPer-organisation latency degradation in shared groups, confirmed by re-probe and two consecutive sweeps.Supporting2026-10-09
ew-v1Early warning (retired)Retired after 0 of 10 warnings were confirmed; the record stays public.Retired2026-10-09
outage-v1Outage trading rule (retired)Retired after a pre-registered event study found no tradable stock reaction to provider outages.Retired2026-10-09

Data sources

SourceWhat we takeRefresh
182 company status pages (Statuspage / incident.io API)Incident title, impact, declared / started / resolved times, linkEvery 10 minutes
AWS Health Dashboard public historySignificant AWS events by service and regionEvery 10 minutes
Google Cloud status (incidents.json)Significant Google Cloud incidentsEvery 10 minutes
Cloudflare, Akamai, Vercel status pagesProvider incidents with impactEvery 10 minutes
ShadowGraph radarHTTP reachability and latency of ~900 services from one cloud locationEvery 5 minutes
Public DNS, IP ownership (ASN), published cloud IP rangesHosting provider, cloud region, CDN, DNS and email provider per serviceWeekly re-mapping

Known limitations

  • Self-reported ground truth. Status pages are written by the companies. Some under-report, some over-report, and declaration times lag the real start.
  • Selection. Companies with machine-readable status pages are mostly developer and SaaS businesses; banks, retailers and media are under-represented.
  • Inferred infrastructure. A CDN hides the origin, so hosting is observable for fewer organisations; multi-cloud set-ups are reduced to the most common provider.
  • One vantage point. All probes run from one cloud location. Regional outages elsewhere can be missed and our own network can fail; sweeps with more than 50% failures are excluded.
  • Short probe history. Raw probes are kept three days; sweeps, hourly aggregates and outage episodes are kept permanently since 8 October 2026.
  • Correlation, not mechanism. A lift says companies on a shared layer declare incidents together more often than chance. It does not prove which component failed.

Changelog

  1. cs-v1 introduced. Upstream signals become a second detection channel. Before launch two attribution rules were tightened: an incident that names another provider no longer counts for this one, and a provider confirmation must start within two hours of the signal (a long-running minor incident elsewhere had counted as confirmation).
  2. sf-v1 revised before publication. The first draft rotated incident histories by arbitrary offsets. Its negative control failed (random groups showed ×2.1) because incidents cluster in working hours. The null now rotates by whole weeks; the negative control passes.
  3. val-v1 introduced. First validation of radar outage detection against official status pages.
  4. ew-v1 retired, ew-v2 introduced. 0 of 10 early warnings were confirmed; the same three hosts per region gained latency together, a measurement-origin artefact.
  5. Outage trading rule retired. A pre-registered event study of 26 outages found no tradable stock reaction (day-0 −0.34%, t −0.55).