Logiciel Contact Us
Success Stories Tech News Contact Us
framework

57% Of Major Outages Cost Over $100,000. Most Second Regions Are Decoration.

Eight sections of checks to run before your second region carries production traffic, each scored 1 to 5 against evidence you can actually produce: a dashboard, a signed target, a dated test result. Topology and dependencies carry the heaviest weights, so the total tells you whether to launch, to rehearse, or to go back and settle the data story.

In depth

The Second Region Goes Up In A Week. The Data Question Is The Actual Project.

01

The trap most infrastructure groups walked into: stand up a second set of pods behind a global load balancer, draw the database as one box because the diagram was made before anyone asked, switch replication on without ever measuring lag, then test failover by draining traffic with the writes still pointed home and record the result as a successful drill.

02

What the teams whose failover actually works do: settle the topology first, marking every datastore active-active or active-passive with the RPO each choice buys, list every write path that assumes a single primary, then lose a region on purpose with production traffic on it and rehearse the failback, which is the harder half nobody schedules.

The detail

What Makes This A Scored Method And Not A List.

Zone · 01

The Weights Are The Argument

Eight sections, and they do not count equally. Data topology carries x4 and dependencies carry x4, because that is where these programmes actually die. Cost and compliance carry x2. The weights sum to 24, so the weighted total runs from 24 to 120, and a comfortable score on the easy sections cannot bury a 5 on the two that decide whether failover works.

Zone · 02

Scored Against Evidence Only

Every line is scored 1 to 5 against something you can put on a screen: a measured p99 replication lag, a signed RPO, a dependency inventory, a dated rehearsal report. A confident answer with no artifact behind it is a 5. Nothing above a 3 goes live, and a 4 is not a caveat for the launch review, it is the reason the launch moves.

Zone · 03

Every Check Names Its Red Flag

Each line carries the specific failure it is hunting. A health check that would stay green with the primary down. A vault or KMS key that is single-region, which quietly makes everything single-region. Monitoring that runs only in the region that failed. The red flag is what stops a room arguing about the wording and sends somebody to go and look.

By the numbers

The figures that make it a board-level conversation.

57%
of organisations say their most recent major outage cost more than $100,000, and for the second year running one in five put it above $1 million
41%
of enterprises put a single hour of downtime at $1 million to over $5 million, which is the figure a signed RPO is supposed to be argued against
2 in 3
reported outages trace to third-party service providers
Inside the report

What you'll take away.

01

get the RPO and RTO signed by the business

Per service tier, in seconds or minutes, with the price of each target stated out loud. Synchronous replication across regions adds latency to every write, permanently, and that is a trade the business makes rather than engineering. Separate region loss from data corruption, because replication copies corruption faithfully. Start the RTO clock at impact, not at declaration.

02

settle the topology datastore by datastore

Postgres, the cache, object storage, search, the queue. Each one marked active-active or active-passive in writing, with a conflict rule chosen rather than an inherited last-write-wins. Then list every write path that assumes a single primary: auto-increment IDs, uniqueness constraints, counters, advisory locks, cron jobs that assume they run once. The list is the deliverable.

03

inventory the dependencies outside the application

Identity, secrets and KMS, DNS and the registrar, CI and the artifact registry, payments, feature flags, Terraform state, the job scheduler, and the observability stack meant to explain the outage. Owned by the infrastructure lead, one named owner per line. Almost every failed failover we review dies here rather than in the application code.

04

lose a region on purpose, then fail back

With production traffic, run by someone who did not write the runbook, and timed. Record the measured RTO and RPO, then check them against the targets from step one. Rehearse the failback as well, because reconciling divergent writes on the way home is harder than the outbound trip and it is where the tidy drill stops being tidy.

Questions

Frequently asked.

Is this not overkill for a standby region we may never use?

No, the standby is the risk. A cold region drifts, and drift is found by the incident it ruins rather than by anyone reading a diagram. Uptime Institute puts 57% of major outages above $100,000, so the question was never whether you use the region. It is whether it works the one time you do

Can we launch with a section still scoring 4?

No, and that is the point. Nothing above a 3 goes live. A 4 in cost or compliance can sometimes be carried with a date and an owner against it. A 4 in topology or dependencies is structural, carries a x4 weight, and will reappear as a data problem in the middle of the incident.

We already run active-active, so do we need to score section B?

Yes, start there. Most systems called active-active are active-passive with a global load balancer in front and a bill priced like the former. The test takes a
minute: state the conflict rule for two concurrent writes to the same row. If nobody in the room can, the label is aspirational.

How long does running the checklist actually take?

One afternoon, then four weeks. Infrastructure and data leads can score all eight sections in a session, because scoring exposes what is not evidenced rather than requiring fresh analysis. The four weeks are for producing the evidence behind the two highest weighted rows: a measured lag number, a signed RPO, a dependency inventory, a rehearsal report.

Does this replace the disaster recovery plan we already have?

No, it audits it. A DR plan states what should happen. This scores whether each part of it has been measured, priced, agreed and rehearsed, which is a different question entirely. Teams scoring well on rehearsal while scoring badly on topology are practising a failover their data cannot survive, and the drill goes green anyway.

Who is this checklist for?

VPs of infrastructure with a second region already funded. It assumes you have compute in two places, replication switched on, and no signed answer yet to how much committed data you can afford to lose, or what you are willing to make slower for every user on a normal day to protect it.

Get the framework

Have it emailed to you.

Drop your details and we'll send 57% Of Major Outages Cost Over $100,000. Most Second Regions Are Decoration. straight to your inbox - no spam, unsubscribe anytime.

Download framework
Next step

Put this into practice.

Book a 30-minute resilience review

Download the checklist