All notes

Platform architecture

Progressive delivery for platform changes

· 4 minute read

Platform changes are high-blast-radius by definition: a bad rollout of a shared build pipeline, service mesh configuration, or base image affects every team at once. Yet platform teams often deploy changes with less progressive-delivery discipline than the product teams they serve. The fix is to treat the platform as the most important production system in the organization — because it is.

Ring the rollout by blast radius

Order consumers by how much damage a bad change can do and how quickly they can detect it. A common ring structure: the platform team's own services first, then a volunteer early-adopter cohort, then a representative slice of general workloads, then the regulated or revenue-critical systems last.

Each ring needs explicit entry and exit criteria. Entry: the change works in the previous ring with defined success metrics. Exit: no blocking issues within the soak period, or automatic rollback. Time-based promotion without evidence is just slow big-bang deployment.

The early-adopter cohort is worth cultivating deliberately. A group of teams that opts into platform changes early, in exchange for influence over the roadmap, gives you realistic feedback before general rollout. Treat them as partners: communicate what is changing, what to watch for, and how to report problems.

Make every change reversible

Progressive delivery only works if rollback is fast, practiced, and unembarrassing. For platform changes this means versioning everything: pipeline definitions, base images, admission policies, mesh configuration, and feature flags on new platform behavior. A change that cannot be rolled back in minutes is a change that should not roll forward.

Practice rollbacks on non-incidents. A team that has never rolled back a platform change will hesitate during a real incident, and hesitation at platform scale is expensive. Game days for platform rollbacks build the muscle memory and expose the gaps in runbooks.

Measure the platform's own delivery health

Platform teams should hold themselves to the same delivery metrics they recommend: change failure rate for platform changes, time to detect a bad platform rollout, time to roll back, and the number of teams affected by platform incidents. Publish these internally.

Also measure adoption friction directly. When a platform change requires action from product teams — a pipeline migration, a base image update, a new required check — track completion rate over time and support load generated. A technically excellent change that half the organization cannot adopt is not excellent.

Communicate like a product launch

Internal platform changes fail socially more often than technically. Announce changes through the channels teams actually read, with enough lead time for planning. Explain what changes, who is affected, what action is required, and where to get help. After rollout, report what happened: what worked, what broke, what you learned.

Maintain a platform changelog with the same care as a customer-facing one. Teams plan their quarters around platform capabilities; surprise changes destroy the trust that makes the platform model work.

Separate experimentation from commitment

Not every platform idea deserves a rollout. Use feature flags and opt-in pilots to test new capabilities with willing teams before committing the whole organization. Kill experiments that do not demonstrate value — a platform cluttered with half-adopted capabilities is worse than a smaller platform with full adoption.

The goal is a platform that changes safely and continuously, earning the trust that lets it change again. Progressive delivery is how platform teams ship at the speed their product teams need without becoming the organization's single point of failure.

Know when progressive delivery is the wrong tool

Progressive delivery is not free. Rings, soak periods, and rollback automation cost engineering time, and for some changes the cost exceeds the risk. A purely additive platform capability behind a feature flag — a new optional pipeline template, an additional base image variant — can often ship directly with good monitoring. The discipline is in distinguishing these from changes that alter existing behavior.

The test is reversibility and blast radius. If the change only adds a new path that nobody uses until they opt in, ship it and watch adoption. If it modifies a path every team already depends on, ring it. When in doubt, ask what the rollback looks like: if you cannot describe it in one sentence, the change needs progressive delivery regardless of how small it seems.

Also recognize that progressive delivery does not substitute for pre-production validation. Rings catch integration surprises and scale effects; they are a poor substitute for testing. A change that fails in the first ring because nobody tried it in staging is not a progressive-delivery success story — it is a testing failure that the rings happened to contain. Keep the testing bar high and let the rings handle what testing cannot predict.

Done well, progressive delivery becomes invisible — just how the platform team ships. That invisibility is the goal: product teams should experience platform change as uneventful, because the drama happened in an early ring where it belonged.

Continue reading