All notes

Mobile

"Mobile release trains: CI/CD when the app store is your deploy target"

· 13 minute read

Web teams deploy when they feel like it. Mobile teams deploy when Apple and Google let them. That asymmetry rewrites every assumption CI/CD makes: there is no rollback, no instant revert, no "fix it in prod in five minutes." A bad binary can sit in a user's pocket for days because they haven't updated, and the 1.0-star reviews it earns are permanent. Mobile release engineering is the discipline of being right before the artifact leaves your control — because once it does, you don't control it anymore.

The fundamental asymmetry

You cannot take back a release. On the web, a deploy is reversible: route traffic away, roll back, move on. On mobile, shipping 2.14.0 means every user who installs it keeps it until they decide to update. Your worst bug will live on devices for weeks alongside your fix. This makes release correctness a property you build into the pipeline, not something you verify after deploy. The blast radius of a mistake is bounded only by how fast users adopt the fix, and adoption is not under your control.

Review queues are part of your lead time. Apple's review is a human process with variable latency — typically fast, occasionally days, sometimes a rejection that resets the clock. Google Play's staged rollouts are mechanical but still slow to propagate. A "we'll fix it Tuesday" plan that assumes web-speed iteration collapses the first time review takes 48 hours. Your release plan must treat store review as infrastructure with its own SLO — one you don't set.

Verdict: mobile CI/CD is release assurance engineering. The question is never "can we ship fast" — it is "can we prove this binary is safe to put in pockets we don't control."

The release train: cadence is a contract

The release train is the foundational discipline: releases cut on a fixed cadence — weekly, biweekly — regardless of whether any given feature is "ready." The train leaves on schedule; what isn't on it catches the next one.

Cadence removes the shipping decision from every feature. Without a train, every release becomes a negotiation: is this feature done enough? With a train, the question inverts: the branch is cut at 2 PM Tuesday, and whatever passed the gates is on it. Features that miss the cut wait. This single change eliminates most release anxiety, because shipping stops being an event and becomes a routine.

Train cuts need a release captain, not a hero. Every train has one accountable person who runs the cut, watches the gates, makes the go/no-go call, and owns the rollout. This rotates. The captain's authority is bounded and real: they can pull a feature off the train, but they cannot stop the train for one. When captains rotate, everyone learns release engineering; when one person owns it forever, you have a bus factor of one in your most critical pipeline.

The cut is a freeze, and the freeze is short. Between branch cut and store submission there is a hardening window — regression tests, manual sanity passes, performance checks. Keep it to hours or a single day. A long freeze means the trunk diverges from the release branch, merges become painful, and engineers context-switch between "the release" and "real work" for a week. If your hardening window keeps growing, your test suite is weak, not your discipline.

  Trunk (main)  ──────────────────────────────────────────────►
                     │
      cut 2PM Tue    │ branch release/2.14 ──► hardening ──► submit
                     │      │                                  │
      Feature A ─────┼──────┘ merged before cut: ON THE TRAIN  │
      Feature B ─────┼────────── merged after cut: NEXT TRAIN  │
                     │                                         ▼
                gates: unit ▸ integration ▸ beta ▸ signing ▸ store

Rules of thumb: biweekly trains for most consumer apps; weekly when your flag discipline is strong enough that trains are nearly automatic; monthly is usually a symptom of manual testing bottlenecks. The cadence you can hold without fire drills is the right one.

Versioning: the contract between you and the device

Mobile versioning has two layers, and confusing them causes real incidents. The marketing version (2.14.0) is what users see; the build number is what the OS uses to decide "is this an update." Both need pipeline discipline.

Build numbers must be monotonic and pipeline-generated. Never hand-assigned, never "I'll bump it." The CI system owns the build number — a counter from the pipeline run or a commit-derived sequence — and it increments on every store-bound build. A duplicate build number uploaded to App Store Connect gets rejected; a non-monotonic one confuses update checks. Make the pipeline generate it from a single source of truth and fail loudly if it ever goes backward.

Semver applies to the app surface, not the binary. Use semantic versioning for the marketing version the way your API consumers think: major for breaking changes (forced-upgrade thresholds, dropped OS versions), minor for features, patch for fixes. But remember the app doesn't get upgraded by a dependency manager — the user's device does it. Version numbers on mobile are communication: to users, to support teams ("what version are you on?"), and to your own analytics.

Every build is traceable to a commit. Tag the commit, stamp the build with the short SHA, and record it in the release manifest. When a crash report arrives against build 1847, the first question — "what code is this?" — must have a one-command answer. This sounds basic; it breaks the first time a build is re-triggered from a slightly different branch tip.

# release-manifest.json — generated by CI for every store build
{
  "marketing_version": "2.14.0",
  "build_number": 1847,
  "commit": "a3f9c1e2",
  "branch": "release/2.14",
  "train": "2026-10-13",
  "captain": "on-call rotation, not a name",
  "feature_flags_snapshot": { "checkout_v2": "off", "dark_mode": "100%" },
  "api_compatibility": { "min_server": "api-v47", "tested_against": "api-v52" },
  "signing_identity": "Apple Distribution: <redacted thumbprint>",
  "stores": ["app-store", "play"]
}

Verdict: the release manifest is the source of truth for "what shipped." If it's generated by hand, you don't have one.

Signing and provisioning: the infrastructure nobody wants to own

Code signing is where mobile CI/CD goes to die. Certificates expire, provisioning profiles mismatch, keystores live on someone's laptop, and the failure message is always inscrutable. Treat signing like any other infrastructure: codified, automated, monitored.

Signing secrets live in a secret manager, not in the repo. Distribution certificates, provisioning profiles, Play service account keys, keystore files and their passwords — all of it goes into your vault (AWS Secrets Manager, GCP Secret Manager, whatever you already use) with rotation runbooks. The build job fetches them at build time and discards them after. A signing credential checked into git is a compromise waiting to be discovered by the same scanners you run on your own code.

Certificate expiry is a monitored event. Apple distribution certificates expire yearly; profiles expire too. Put their expiry dates in your monitoring with alerts at 60, 30, and 14 days. The failure mode is always the same: nobody notices until the release train can't sign, and the train doesn't wait. This is a five-minute automation that prevents a guaranteed incident.

Rule of thumb: if more than one human knows the signing password by heart, your signing infrastructure is already broken — it's just a secret that's distributed rather than managed.

Beta channels: your last real feedback loop

Between CI and the store there should be a beta channel — TestFlight on iOS, internal/closed testing tracks on Play — and it should be used every train, not just for "big" releases.

Dogfood every train internally. The beta build of every release candidate goes to the whole company, and support, QA, and product open it on their own devices. Most teams underuse this: the beta exists, but nobody is required to look. Make the beta gate explicit — the release captain doesn't submit to review until the beta build has a clean bill from dogfooding, or at least a defined window (24–48 hours) with no new crash clusters.

Crash and performance telemetry gates the rollout. Instrument crash-free session rates and key performance metrics (cold start, ANR rate, frame drops) per build, and set gate thresholds: if the beta build's crash rate exceeds the current production baseline by a defined margin, the train doesn't proceed. Pick thresholds in advance and write them down — "we'll know a bad build when we see it" is how bad builds ship.

Verdict: the beta channel is the closest thing mobile has to a staging environment. Treat it as a gate, not a formality, or it adds zero value.

The two-store strategy: review and rollout

iOS: submit early, use phased release. App Store review is the uncertain part — submit as early as the build is stable, then use phased release (1% → 5% → 15% → 50% → 100% over 7 days) to control rollout. You can pause a phased release mid-rollout; you cannot pause a full release. Every production iOS release should go through the phased schedule. The pause button is your rollback substitute, and it only works if you haven't already reached 100%.

Android: staged rollouts with holdback. Play's staged rollout works the same way — 1%, 5%, 20%, 50%, 100% — and you should also always hold a small control group at the old version for telemetry comparison. Play gives you one more weapon: you can halt a rollout entirely and supersede it with a new release. Use staged rollouts for every release, and define the promotion criteria in advance: crash rate delta, key business metric movement, support ticket volume.

Keep the stores in sync deliberately. The ideal is same-day submission to both stores, but review timing differs. Decide your policy: release Android as soon as its rollout completes, or wait for iOS approval and go together? Cross-platform features with server-side dependencies make this decision for you — if 2.14 on either platform requires API v47, neither ships until both are approved. Record the policy per release; "we'll figure it out" produces one platform carrying a feature the other can't support.

Rule of thumb: 100% day-one releases are a declaration that your beta process is decorative. If you're confident enough to skip phasing, your beta channel should be empty — and it isn't.

Feature flags: decoupling deploy from release

On the web, feature flags are a convenience. On mobile, they are load-bearing infrastructure — the only mechanism that decouples "the binary shipped" from "the feature is on."

Every feature merges behind a flag, off by default. The train ships binaries containing inactive code. Activation happens server-side, per user cohort, after the binary is in the wild. This is what makes trains safe: the release candidate can contain five half-built features, and the train still ships, because flags are off and the flag states are what get validated in beta.

Flags need the same lifecycle discipline as code. A flag is born (defined, default off), rolled out (cohorts, percentages), and — critically — retired. Retired means the code path is removed, not just "flag set to 100% forever." Stale flags accumulate: dead code paths, config surface that nobody understands, and the day a flag service has an outage and everything falls back to defaults, you discover which defaults are still correct. Schedule flag retirement as part of the release that no longer needs it.

Kill switches are flags with a different SLA. Every release should carry kill switches for its riskiest new behavior — not gradual rollouts, but hard off-switches that a human can flip in seconds when the crash dashboard lights up. The difference between a flag and a kill switch is operational: kill switches are documented in the runbook, visible on the release dashboard, and tested (flip it in beta, confirm the behavior dies). An untested kill switch is a rumor.

// Feature flag evaluation — client-side, with server-pushed config
interface ReleaseConfig {
  version: string;               // config schema version
  flags: Record<string, FlagState>;
  killSwitches: Record<string, boolean>;
}

function evaluateFlag(name: string, user: UserContext): boolean {
  // Kill switch always wins — checked first, no caching across flips
  if (releaseConfig.killSwitches[name] === true) return false;
  const flag = releaseConfig.flags[name];
  if (!flag) return flagDefaults[name] ?? false;  // fail closed
  return flag.enabled && inCohort(user.id, flag.rolloutPercent);
}

Fail closed, always. If the flag config can't be fetched, the feature is off. This is the mobile equivalent of a circuit breaker: a flag-service outage must degrade to the safest state, not the most exciting one. Test the outage path — airplane mode at launch, flag service down — because it will happen.

Verdict: on mobile, the release is the flag flip, not the store submission. The binary is just the envelope.

Hotfixes: the train has an express lane

Production incidents on mobile need a hotfix path that doesn't demolish the train schedule — but "hotfix" must mean something precise, or every urgent request becomes one.

Define what qualifies. A hotfix is a production crash affecting a large cohort, a data-loss bug, or a security issue — shipped as the smallest possible diff against the current production build, not against trunk. Everything else rides the next train. Write the criteria down and make the release captain the gatekeeper. Without written criteria, the hotfix lane becomes the lane every stakeholder lobbies for, and your release process degrades into ad-hoc deploys — exactly the chaos the train was built to prevent.

Hotfix branches from the release tag, not from main. The hotfix is cut from the exact commit that produced the production build, takes the minimal fix plus flag changes, runs the full gate suite, and goes through an accelerated but complete store submission — including review. Apple offers expedited review for critical fixes; use it, but don't plan on it being instant. The hotfix then merges forward into main so the fix isn't lost in the next train.

Rule of thumb: if you're hotfixing more than once a quarter, the problem is your beta gates, not your luck. Fix the gates.

API compatibility: the server moves under your feet

Mobile clients talk to backends that deploy independently, and old app versions talk to new backends for months. The contract discipline that web teams take for granted — client and server deploy together — doesn't exist.

The backend never breaks an in-the-wild client. Additive changes only: new fields, new endpoints, new enum values that old clients ignore. Breaking changes ride behind API versioning, and old versions stay live until the oldest supported app version is below a negligible adoption threshold. This is a backend design constraint driven entirely by mobile realities, and backend teams that don't internalize it will break your users.

Define a minimum supported version and enforce it. Every release knows the oldest server version it supports (in the manifest) and every server knows the oldest client it tolerates. Below that floor, the app shows an upgrade prompt — not a crash, not silent breakage, a clear "please update to continue" screen. Forced upgrades are a user-hostile tool; reserve them for genuine incompatibility (security, data model breaks), and make the upgrade path one tap to the store.

Verdict: on mobile, backward compatibility isn't a nice-to-have — it's the load-bearing wall. Every "small" breaking change is a cohort of users with a broken app and no way to fix it except updating.

Anti-patterns

Manual store submission. Someone uploading the binary from their laptop "just this once" breaks traceability, signing discipline, and the manifest. The store submission is a CI job or it doesn't happen.

Testing only the latest OS. Your users run three major OS versions. Test the release candidate on the oldest supported OS — that's where the WebView quirks, permission behaviors, and layout bugs live.

Flags without owners. Every flag has an owner and a retirement date. Flags without owners become permanent, and permanent flags become the config sprawl that fails in the worst way during an outage.

Treating review rejection as a surprise. Rejections happen — guideline changes, metadata issues, false positives. Build the appeal/retry path into the schedule: submit early enough that one rejection cycle doesn't miss the train.

Closing

Mobile release engineering is the art of shipping with confidence into an environment you don't control. The train gives you cadence; the gates give you evidence; flags give you a release mechanism independent of the store; phased rollouts give you the closest thing to a rollback that exists. None of it is glamorous. All of it is the difference between a release process that runs itself and one that runs you.

Continue reading