Eight sections of checks to run before your second region carries production traffic, each scored 1 to 5 against evidence you can actually produce: a dashboard, a signed target, a dated test result. Topology and dependencies carry the heaviest weights, so the total tells you whether to launch, to rehearse, or to go back and settle the data story.
The trap most infrastructure groups walked into: stand up a second set of pods behind a global load balancer, draw the database as one box because the diagram was made before anyone asked, switch replication on without ever measuring lag, then test failover by draining traffic with the writes still pointed home and record the result as a successful drill.
What the teams whose failover actually works do: settle the topology first, marking every datastore active-active or active-passive with the RPO each choice buys, list every write path that assumes a single primary, then lose a region on purpose with production traffic on it and rehearse the failback, which is the harder half nobody schedules.
Eight sections, and they do not count equally. Data topology carries x4 and dependencies carry x4, because that is where these programmes actually die. Cost and compliance carry x2. The weights sum to 24, so the weighted total runs from 24 to 120, and a comfortable score on the easy sections cannot bury a 5 on the two that decide whether failover works.
Every line is scored 1 to 5 against something you can put on a screen: a measured p99 replication lag, a signed RPO, a dependency inventory, a dated rehearsal report. A confident answer with no artifact behind it is a 5. Nothing above a 3 goes live, and a 4 is not a caveat for the launch review, it is the reason the launch moves.
Each line carries the specific failure it is hunting. A health check that would stay green with the primary down. A vault or KMS key that is single-region, which quietly makes everything single-region. Monitoring that runs only in the region that failed. The red flag is what stops a room arguing about the wording and sends somebody to go and look.
Per service tier, in seconds or minutes, with the price of each target stated out loud. Synchronous replication across regions adds latency to every write, permanently, and that is a trade the business makes rather than engineering. Separate region loss from data corruption, because replication copies corruption faithfully. Start the RTO clock at impact, not at declaration.
Postgres, the cache, object storage, search, the queue. Each one marked active-active or active-passive in writing, with a conflict rule chosen rather than an inherited last-write-wins. Then list every write path that assumes a single primary: auto-increment IDs, uniqueness constraints, counters, advisory locks, cron jobs that assume they run once. The list is the deliverable.
Identity, secrets and KMS, DNS and the registrar, CI and the artifact registry, payments, feature flags, Terraform state, the job scheduler, and the observability stack meant to explain the outage. Owned by the infrastructure lead, one named owner per line. Almost every failed failover we review dies here rather than in the application code.
With production traffic, run by someone who did not write the runbook, and timed. Record the measured RTO and RPO, then check them against the targets from step one. Rehearse the failback as well, because reconciling divergent writes on the way home is harder than the outbound trip and it is where the tidy drill stops being tidy.
No, the standby is the risk. A cold region drifts, and drift is found by the incident it ruins rather than by anyone reading a diagram. Uptime Institute puts 57% of major outages above $100,000, so the question was never whether you use the region. It is whether it works the one time you do
No, and that is the point. Nothing above a 3 goes live. A 4 in cost or compliance can sometimes be carried with a date and an owner against it. A 4 in topology or dependencies is structural, carries a x4 weight, and will reappear as a data problem in the middle of the incident.
Yes, start there. Most systems called active-active are active-passive with a global load balancer in front and a bill priced like the former. The test takes a
minute: state the conflict rule for two concurrent writes to the same row. If nobody in the room can, the label is aspirational.
One afternoon, then four weeks. Infrastructure and data leads can score all eight sections in a session, because scoring exposes what is not evidenced rather than requiring fresh analysis. The four weeks are for producing the evidence behind the two highest weighted rows: a measured lag number, a signed RPO, a dependency inventory, a rehearsal report.
No, it audits it. A DR plan states what should happen. This scores whether each part of it has been measured, priced, agreed and rehearsed, which is a different question entirely. Teams scoring well on rehearsal while scoring badly on topology are practising a failover their data cannot survive, and the drill goes green anyway.
VPs of infrastructure with a second region already funded. It assumes you have compute in two places, replication switched on, and no signed answer yet to how much committed data you can afford to lose, or what you are willing to make slower for every user on a normal day to protect it.
Drop your details and we'll send 57% Of Major Outages Cost Over $100,000. Most Second Regions Are Decoration. straight to your inbox - no spam, unsubscribe anytime.
Book a 30-minute resilience review
Download the checklist