"Active-active multi-region design: the honest trade-offs"
· 12 minute read
Every few years a team decides it needs active-active multi-region, usually after an outage, with the requirement phrased as "we can never go down." Active-active is legitimate for a narrow set of workloads — and an expensive, operationally brutal mistake for most others. The honest conversation: what each tier costs, what it buys, and why the hard part is always the data, never the traffic.
The spectrum: what each tier costs and buys
Multi-region is not binary. It is a spectrum, and each step up roughly doubles the operational burden:
Backup and restore. Back up to a second region; on failure, restore and redeploy. Cost: near zero. RPO: hours. RTO: hours to a day. Fine for internal tools where a day of downtime is an inconvenience, not existential.
Pilot light. Minimal core idles in the second region: replication running, critical services at near-zero scale. On failure, scale up and cut traffic over. Cost: small. RPO: minutes to an hour. RTO: tens of minutes to hours. The trap: a pilot light is only as good as its last test — untested ones fail exactly when needed.
Warm standby. Near-full capacity running in the second region, serving little or no traffic. Cost: ~2x infra. RPO: seconds to minutes. RTO: minutes. Most serious teams should stop and think hard here, because the next step is a cliff.
Active-passive. Second region fully live, serving reads or a subset of users; writes anchored in primary. Cost: 2x+ infra plus replication. RPO: seconds. RTO: minutes. Failover: shift traffic, promote the passive data plane to accept writes.
Active-active. All regions serve all traffic, writes included, concurrently. Cost: 2x+ infra plus conflict resolution plus the permanent operational tax below. RPO/RTO: near zero — a region loss is a capacity reduction, not an outage.
The key question at each tier: what RPO/RTO does the business actually need, in dollars per hour — not what sounds good in planning. Most teams discover "warm standby would have been fine" only after two years of paying for active-active.
Why active-active is a data problem, not a traffic problem
Routing traffic to multiple regions is solved — DNS, anycast, global load balancers are commodities. The unsolved problem is concurrent writes to the same logical data across 100ms+ of inter-region latency.
Physics sets the terms: ~70ms for light to cross the Atlantic one way, longer in practice. Strong cross-region consistency pays that latency on every write — or you accept divergent writes that must be reconciled. No third option. Every active-active design answers one question: when two regions accept conflicting writes to the same record, what happens? If your design doc doesn't answer that per entity type, you have a diagram, not a design.
This is also why "active-active for reads" is barely a decision: read replicas across regions are standard practice with well-understood lag semantics. The cliff is write-write concurrency. Teams that say "we need active-active" usually need multi-region reads plus fast failover of writes — which is active-passive, at half the complexity.
Conflict resolution strategies
Last-writer-wins (LWW) is the default everyone drifts into: latest timestamp wins, the other write is silently discarded. Simple, convergent, and wrong wherever the losing write mattered — concurrent balance updates (money vanishes), two admins editing one config (a change disappears unnoticed), any counter. Use LWW only where silent overwrite is genuinely harmless: a display name, a cached preference. That set is smaller than teams assume. And never use wall-clock timestamps for "latest" across regions — clock skew elects the wrong winner; use hybrid logical clocks.
CRDTs merge deterministically: counters, sets, maps, text. Right for genuinely mergeable data — collaborative editing, distributed counters, add-only sets. Limits: not every operation has a CRDT (a bank transfer isn't a counter), metadata grows (tombstones accumulate), and "deterministic merge" is not "the merge the business wanted." Adopt per data structure with domain confirmation of the merge semantics.
Application-level merge resolves conflicts with domain logic: union the carts, keep both addresses and flag for review. Most correct, most expensive — every entity needs its merge function tested against every conflict shape. Reserve for entities where conflicts are both likely and business-meaningful.
Partitioning by entity or region avoids conflicts by construction: each entity's writes go only to its home region; cross-region access is read-only or routed home. Most working "active-active" systems are actually partitioned this way — single-homed writes, global reads. It trades write locality for some users for eliminating conflict resolution entirely. For the large middle of workloads: partition the writes, replicate the reads.
Decision rule: partition first. For what can't be partitioned, use CRDTs where the merge semantics fit the domain, application merge where conflicts are meaningful, and LWW only where silent overwrite is genuinely harmless. Document the choice per entity type — a table, in the design doc, reviewed.
The CAP reality for your specific workload
CAP is not a slogan; it is a per-operation question. During a network partition between regions, each write operation chooses: block until the partition heals (consistency, sacrificing availability) or accept the write locally and reconcile later (availability, sacrificing consistency). "We chose availability" is meaningless until you say for which operations, because the answer differs within one system: accepting a like during a partition is fine; accepting a funds transfer during a partition is a reconciliation nightmare.
Do the exercise per critical path. List your top ten writes. For each: if regions can't talk, accept or reject? What's the reconciliation procedure, who owns the queue, what's its SLA? Most teams find 8 of 10 can be partition-tolerant — and the remaining 2 (payments, inventory, anything with external financial effect) must block or route to one region. That finding is your multi-region data architecture.
Global vs regional data classification
Before any replication topology, classify every dataset:
- Global, single-writer: reference data, configuration, product catalog. Replicate everywhere; writes go to one place.
- Global, multi-writer: user-generated content, collaboration data. Needs conflict resolution — the expensive category. Minimize what lands here.
- Regional, partitioned: user profiles, orders, tenant data. Written in the home region, readable everywhere. The bulk of most systems.
- Regional, ephemeral: sessions, rate-limit counters, caches. Doesn't need cross-region replication at all — rebuild locally.
Classification typically shrinks the "needs active-active" surface to a fraction of the data. The common failure is defaulting everything to global multi-writer and discovering the conflict-resolution bill later. Classify first; choose replication per class, not one topology for everything.
DNS and traffic steering (slower than you think)
DNS failover is slower and less reliable than its reputation: TTLs are widely ignored, real-world propagation takes minutes at best, and mobile clients notoriously cache beyond TTL. If your RTO needs DNS to propagate in 60 seconds, test that against your actual clients.
Prefer anycast or global load balancer health checks for fast failover: traffic shifts on health-check failure, in seconds, independent of client DNS behavior. Use DNS for the coarse decision (which region set) and health-checked balancing for the fast one (which region). Keep a manual traffic-shift runbook — during a real outage, the automation is sometimes a casualty too.
Steering policy matters as much as mechanism. Route users to the nearest healthy region, but respect partitioning: a user homed in region A shouldn't take writes in region B just because it's closer. Geo-routing without partition awareness recreates the conflicts partitioning was meant to avoid.
Replication lag budgets and what breaks when you exceed them
Every cross-region read of replicated data is a read at some lag. Define the lag budget per dataset — "user profile reads may lag writes by up to 5 seconds" — and then decide what breaks when reality exceeds it. Because it will: replication falls behind during traffic spikes, network degradation, and failovers.
Design for lag explicitly. After a write in region A, a user routed to region B: do they see their write? If yes, you need sticky routing, version-gated reads (wait for replica version >= write version, with timeout), or home-region reads. "Hope the lag is small" is not a strategy. For derived data, decide: computed per region (divergent during lag — fine for dashboards, not billing) or computed once and replicated (consistent but lagged).
Monitor lag as a first-class SLI per replication stream, alerting well below budget — it's the multi-region equivalent of queue depth. And load-test replication under failure: kill a consumer, measure lag growth, and observe behavior at 10x budgeted lag. That behavior is your real consistency story.
Failover testing: game days and forced region evacuation
An untested failover is a hope, not a plan. Multi-region systems need regular, realistic failure testing:
Game days simulate region loss with the team watching: traffic shifts, replication reverses, on-call runs the runbook. Start announced, in business hours; go unannounced only when the announced ones are boring. Every game day produces action items that actually get done — one that finds the same gaps twice is theater.
Forced region evacuation is stronger medicine: regularly run production with a region drained to prove the system works degraded. If evacuation causes an outage, your active-active isn't — a region loss should be a capacity reduction, not an outage. Schedule evacuations like deploys: routine, automated, reversible.
Test the data plane, not just traffic. Shifting traffic is the easy half. The hard half: does replication reverse cleanly, do conflict paths engage, do partitioned writes re-home, and does the evacuated region rejoin without corruption? Rejoin is where most multi-region data loss happens — two divergent datasets meeting again. Test it explicitly with verification (checksums, reconciliation reports), not just "traffic looks fine."
The operational cost: on-call, deploys, schema changes
This is the section vendors skip. Active-active roughly doubles infrastructure cost and more than doubles operational cost:
On-call across regions means 3 AM incidents for someone, always — follow-the-sun or accept one team carrying nights. Debugging spans regions, and every dashboard, runbook, and alert needs a region dimension. "Is this regional or global?" becomes every incident's first question.
Deploys are region-aware: rolling per region with health verification, haltable mid-rollout, with compatibility windows where old and new coexist across regions. Every deploy is a multi-region deploy. Feature flags help enormously — per-region flag-gated rollouts decouple code deployment from behavior change.
Schema changes are the sharpest edge: a column added in region A must be readable by region B's old code until B deploys — every change backward compatible, expand-then-contract as the only safe pattern. Replicating a schema change old code can't parse halts replication, which becomes a lag incident, then a consistency incident. The discipline that was "good practice" single-region becomes load-bearing multi-region. Budget the time; teams underestimate it by half.
Honest accounting: single-region ops cost, doubled, plus 50% cross-region coordination tax (replication monitoring, reconciliation, multi-region deploys, game days). If that doesn't fit the budget, you have an active-active aspiration that will degrade into untested warm standby.
When active-passive is the right answer (most of the time)
Active-passive with fast failover covers the realistic requirement for the large majority of systems: survive a regional outage with minutes of downtime and seconds of data loss. It gives you the headline benefit — "we survive region loss" — without conflict resolution, without multi-writer reconciliation, without the permanent operational tax.
Choose active-passive when: writes tolerate brief unavailability during failover; your conflict-resolution answer is "we don't have one"; the team is too small for 24/7 multi-region on-call; or the business can't quantify downtime cost but knows the engineering cost. That last one describes most companies most of the time.
Active-active earns its keep when: downtime is measured in millions per minute with contractual penalties; users are globally distributed and need write locality (not just read locality) in multiple regions; or the workload is naturally partitionable with genuinely concurrent multi-region writes (collaboration, messaging, gaming). If none of those apply, active-passive plus aggressive RTO optimization is the better engineering call — and it's honest to say so.
Decision framework
| Question | Points toward active-active | Points toward active-passive |
|---|---|---|
| Cost of 5 min write downtime | Contractual / existential | Annoying but survivable |
| Write locality needs | Users write from everywhere, latency-sensitive | Writes tolerate home-region routing |
| Conflict resolution | Answered per entity, tested | "We'll figure it out" |
| Team size / on-call | Follow-the-sun staffed | One team, one timezone |
| Data classification | Mostly partitionable | Large global multi-writer surface |
| Schema change discipline | Expand-contract, enforced | Ad-hoc migrations |
| Failover testing appetite | Quarterly game days minimum | "We'll test before we need it" |
If more than two rows point right, you don't want active-active — you want active-passive with a fast, tested failover. Revisit annually; the answer changes as the business grows.
Anti-patterns
Active-active for vanity. "Our competitors are multi-region" is not an architecture requirement. Build to your RPO/RTO numbers, not to a marketing checkbox.
Shared global database with synchronous replication across oceans. This is active-active in name and a latency catastrophe in practice: every write pays intercontinental round trips, and a network partition halts writes everywhere. If your "multi-region" database needs synchronous cross-ocean replication to be correct, you have built the world's most expensive single-region database.
Failover that has never been tested. The pilot light that hasn't lit in a year, the runbook last updated by someone who left, the DNS change "we'll make when it happens." Untested failover fails; this is as close to a law as distributed systems gets.
Treating regions as identical. Regions differ: instance types, quotas, network performance, feature availability, even clock behavior. A deploy that works in one region can fail in another for region-specific reasons. Test in every region you operate, not just the primary.
Forgetting the humans. Multi-region incidents need clear ownership: who declares a region dead, who authorizes traffic shifts, who talks to the cloud provider. "The automation handles it" works until the automation is what's broken. The runbook names names.
Closing
Active-active is a specific, expensive answer to a specific question: which writes must survive region loss with zero downtime, and what happens when they conflict? Answer per entity — conflict resolution designed and tested, lag budgeted and monitored, failovers rehearsed until boring, ops cost honestly accounted. If that program doesn't fit, build active-passive with a fast, tested failover and spend the savings on what actually causes your outages. The best multi-region architecture is the simplest one that meets your real RPO/RTO.