The anatomy of a release cadence, from cut to rollback

Every release process is four verbs — cut, verify, ship, roll back — and an owner for each. The shapes real teams run, what each demands upstream, and where they break at 10, 40, and 200 engineers.

Andrei Gaspar

Editor, How They Ship · · 11 min read

Share
Amber circle, cream card and teal pill on charcoal

Ask an engineer how their team ships and you will usually get a tool name: "GitHub Actions", "Argo", "we have a deploy bot". Ask them to draw it — the path a commit takes from merge to a user's browser, with a name next to every arrow — and the room goes quiet. That silence is the most useful thing this piece can give you, so we will come back to it at the end.

A release process, whatever the tooling, is four verbs. Cut: decide which commits are in this release. Verify: collect the evidence that it is safe. Ship: move it in front of users. Roll back: undo it when the evidence was wrong. Each verb has an owner, a duration, and a failure mode, and the shape of a team's cadence is nothing more than the answers to those twelve questions.

The claim of this piece is that the fourth verb decides the other three. How fast a team can undo a release sets how much verification it needs before shipping, which sets how big a batch it can afford to cut, which sets the cadence. Teams that want to ship more often almost never get there by shipping faster. They get there by making undo cheap and then letting everything upstream relax.

Four verbs, twelve questions

Cut answers: which commits? The honest options are the head of a branch at a scheduled time (a train), every individual merge (continuous), a tag someone chooses (a batch), or a branch someone stabilizes for days (a mobile release). The question behind it is whether a human decides the contents, and if so, whether that human can say no.

Verify answers: what evidence gates the ship? CI is the floor. Above it sit some combination of a staging environment, a canary slice in production, automated canary analysis against a baseline, manual QA, and a human approval. The question behind it is what the last five releases actually caught at each stage, which almost nobody has written down.

Ship answers: how does it reach users? All at once, by percentage, by region, by cohort (employees first), or through a third party's gate — an app store, a partner's change window. The question behind it is who can press the button and whether they need to be the author.

Roll back answers: how do we undo it, and how long does that take from the moment someone decides? Redeploying the previous artifact, flipping a flag, a forward fix, or, for a shipped binary, waiting. The question behind it is whether the undo has been done under real conditions in the last quarter, or only in a runbook.

For each verb the three follow-ups are the same: who does it, how long does it take, and can a person who did not write the change do it at three in the morning. Any answer of "it depends who's around" is a hole in the cadence.

The shapes real teams actually run

Only a handful of honest shapes exist. Most teams are running one of them, sometimes without having chosen it.

ShapeCutVerifyShipRoll backWho owns ship
Continuous deploy on mergeEvery mergeCI, then canary in productionAutomatic, progressiveRedeploy previous, or flag offNobody; the pipeline
Release train, no managerWhatever is on trunk at departureCI, short soakScheduled, progressiveRevert and wait for next train, or flagA rotation, not a role
Daily batchA human picks the momentCI, often stagingOnce a day, usually all at onceRedeploy previousWhoever is "on deploy"
Mobile's forced trainA branch cut and stabilizedCI, manual QA, store reviewStore phased rolloutCannot undo the binary; flags and a hotfix buildA release owner per version

Continuous deploy on merge is the shape Instagram's engineering write-up on continuous deployment describes arriving at. The useful thing about that account is that it does not skip the intermediate states: the team lived in each one and abandoned it for a reason. The defining property is that the merge is the cut. There is no second decision. Verification therefore has to be automatic, because there is no human between merge and users to look at the evidence.

A release train without a release manager leaves on schedule with whatever is merged. Missed it? Next one. The rule sounds administrative and is the most consequential one on this list, because it makes holding the train impossible, which makes "just one more fix" impossible, which pushes every engineer toward changes that are safe to ship half-done — which is to say, behind a flag.

The daily batch is where most teams between ten and forty engineers actually live. Someone decides it is time, the pipeline builds what is there, someone watches a dashboard for a while. Its honest cost is that the batch is the incident size: when the deploy is bad, the search space is a day of merges from several people.

Mobile's forced train is the one shape nobody chooses. The binary goes through a store, review takes time you do not control, and once a version is on a device you cannot pull it back — you can only ship another one or turn a feature off from the server. Mobile teams run the most disciplined flag practice in the building because the flag is the only rollback they have.

Facebook's "Rapid release at massive scale" post describes moving the web tier from a cut branch pushed on a schedule to a near-continuous push from master, with exposure widening in tiers — employees, then a small slice of production, then everyone. Slack's "Deploys at Slack" post describes a rotation of deploy commanders who own a deploy window and watch a staged rollout. Neither is a template; both confirm that the tiered ship and the rotating human are the two pieces that survive at scale while everything else gets automated.

What each shape demands upstream

The shape you pick is a bill you send to the rest of the engineering process. Nobody adopts continuous deploy and keeps two-week feature branches; the shape does not permit it.

Trunk-based development is the entry fee for anything faster than a daily batch. If integration happens on merge, integration has to be small and frequent, which means branches that live for hours, not weeks. Teams that run a train on top of long-lived branches discover that the train's contents are decided by whose merge landed first, and the merge conflicts become the release manager.

Flag discipline follows directly. If a change can be shipped before it is finished, it needs a switch, and the switch needs an owner and an expiry. Pete Hodgson's feature toggles article on martinfowler.com is still the reference for the taxonomy — release, ops, experiment, and permission toggles — and for the point that these have different lifetimes. A team with one kind of flag for all four purposes has a cleanup problem coming.

Small changes are the third demand. Not as a virtue, as arithmetic: a bad continuous deploy has one commit in it, so the rollback decision is trivial and the fix is obvious. A bad daily batch has a day of commits, and the first twenty minutes of the incident are spent finding which one. The size of the change is the size of the search.

Fast CI is the constraint people underestimate. Suppose a team merges thirty times a day into a pipeline whose CI takes twenty-five minutes. If merges are serialized through a queue that verifies each one against the current head, the queue clears at most about two an hour and falls behind by mid-morning. Either CI gets faster, merges get batched, or the queue speculates — testing several candidate merges in parallel on the assumption that most will pass, which is what Uber's "Keeping Master Green at Scale" paper describes for its monorepo submit queue. For a smaller team the point is simpler: the queue's throughput is CI time divided into the hour, and that number gets checked against the merge rate before adopting continuous deploy, not after.

Decoupled migrations are the demand that bites hardest during rollback. If a deploy carries a schema change the previous version of the code cannot run against, "redeploy previous" is no longer an undo. The expand/contract pattern — add the column, deploy code that writes both, backfill, deploy code that reads the new one, then drop — exists so that every step is reversible on its own. It costs several deploys instead of one; that is the bill.

Sponsored: Gitdailies — Install today. Move faster daily.

Where it strains: ten, forty, two hundred

Release processes do not fail at random. They fail at specific team sizes for specific reasons, and the replacements follow a recognizable order.

At ten engineers everyone deploys, the deploy is a script, and rollback is running the script with the previous tag. There is no release manager because everyone knows what is in the release — they wrote it, or sat next to the person who did. The strain is tacit knowledge: the process lives in one or two heads, and the day one of them is on vacation, deploys stop. The first artifact worth producing at this size is not a tool; it is a page a new hire could follow to deploy and to roll back, tested by making the new hire do it.

At forty engineers several teams are merging into one deploy. Team A's bad merge now blocks team B's fix, and someone has to decide whether to hold, revert, or push through. That someone is the release manager, and the role appears whether or not anyone creates it: somebody starts being the person who "knows what's in the release", and within a quarter that is their job. Staging appears at the same size for the same reason — it is the first place all the teams' changes meet. The strain arrives about a year later. The release manager is a human queue with single-threaded capacity, staging drifts from production in data and shape and scale, and the freeze — the manager's tool for controlling risk — starts accumulating the biggest batch of the year for the first deploy after it.

At two hundred engineers the release manager disappears, replaced by machinery that does the same job deterministically: a merge queue decides contents, automated canary analysis decides verification, a rotation supplies a human who watches but does not choose. Staging is usually demoted or deleted around here. It did not stop being useful in principle; keeping it representative of production became a project nobody wanted to own, and a canary slice of real traffic answers the same question with real data. The deploy itself splits: the monolith's single pipeline becomes per-service pipelines, so team A's failure no longer blocks team B at all.

The freeze fails most visibly on this path, and for a mechanical reason rather than a cultural one. A freeze stops changes. It does not stop work. Merges pile up on trunk, or worse, on branches, and when the freeze lifts the team ships more change in one deploy than it has all quarter, on the first day everyone is back, with a rollback that has to unwind weeks of migrations. The replacement is not "no freeze" but a fast undo and a progressive rollout, at which point a bad change over a holiday costs a few minutes of a small slice of traffic and one page, and the freeze has nothing left to buy.

Rollback is the hidden design constraint

Take a bad deploy. Its cost to users is roughly the fraction of traffic exposed, multiplied by the time from ship to undo. That time is three intervals: how long until someone or something notices (detection), how long until they decide to roll back (decision), and how long the undo mechanically takes (execution).

Suppose detection is five minutes from a good alert, decision is two minutes because the rule is "roll back first, investigate after", and execution is three minutes because the previous artifact is still warm and the pipeline can redeploy it without a rebuild. Exposure is ten minutes at whatever slice the rollout had reached — say five percent. A team with those numbers can afford to ship on merge with nothing more than CI in front of production, because the worst case is small and short.

Now suppose execution is ninety minutes, because the deploy rebuilds from source, the migration is not backward compatible, and the person who knows the runbook is not on call. Exposure is over an hour and a half at whatever slice the rollout had reached, and the rollout was probably one hundred percent because the team never invested in progressive delivery. That team needs staging, manual QA, and a sign-off before every ship, and it is right to have them. Its cadence is a consequence of its undo, not its ambition.

The things that make execution slow are consistent across teams: schema migrations coupled to code, stateful services that need a warm-up, caches holding the new shape of data, clients caching the old API, and rebuilds instead of re-promotions of an artifact. Each is fixable, and each fix relaxes a constraint upstream. The DORA research's four metrics — deployment frequency, lead time, change failure rate, time to restore — are widely quoted; read together, we take them to say that restore time is the lever and the other three follow.

Draw your own

This is the exercise, and it takes thirty minutes.

Get the team in a room with a whiteboard, or the remote equivalent. Do not discuss. Each person, independently, draws the release process: a commit on the left, a user on the right, every stage in between as a box, and every arrow labeled with who or what moves the change along and how long it typically takes. Include the rollback path. Include the pager.

Then compare.

If nobody can draw it, that is the finding. The process exists — code reaches production somehow — but as habit and tacit knowledge, and it will fail the day the habit-holder is unavailable. The next step is writing down what actually happens today, however embarrassing, and having someone who did not write it follow it.

If everyone draws something different, that is the second finding, and the more common one. The differences cluster at the same places: who decides the cut, what verification is required versus merely available, and how rollback works for a deploy that included a migration. Those differences are your incident backlog in advance.

If everyone draws the same thing, check it against the twelve questions:

  1. Who decides the cut, and can they say no?
  2. How long from merge to the cut, at the median and at the worst in the last month?
  3. What did each verification stage actually catch in the last five releases?
  4. How long does verification take, and who waits on it?
  5. Who can trigger ship? Does it have to be the author?
  6. What fraction of users see the change in the first ten minutes?
  7. How is a bad release detected, and how fast, measured not assumed?
  8. Who decides to roll back, and what is the rule?
  9. How long does the undo take once decided?
  10. When was it last done for real, and by whom?
  11. Which change in the last quarter could not have been rolled back, and why?
  12. What happens to all of the above at three in the morning on a public holiday?

Any answer of the form "it depends" is a box on the drawing with no owner.

What to do on Monday

Three measurements, all cheap, all more useful than a new deploy tool.

First, count deploys per week and lead time from merge to production for the last month. Pull it from the pipeline's history; a script against the CI API is an afternoon's work. You will learn which shape you are actually running, which is often not the one people describe.

Second, run a rollback drill. Pick a service, deploy a trivial change, and have the on-call engineer — not the pipeline's author — roll it back, timed, following only what is written down. The gap between the runbook and reality is the most valuable number you will collect this quarter.

Third, list the last five things each verification stage caught. If staging's list is empty or is entirely things CI could have caught with one more test, you have found something to stop doing, and the evidence to justify it.

The teams whose processes are worth copying did not start with a better shape. They started with a faster undo, and then noticed that most of the ceremony upstream had nothing left to protect.

Andrei Gaspar

Editor, How They Ship

This blog exists thanks to the support of our sponsors:

GitdailiesQA.techAppSignalSuperlinked
Become a sponsor

Comments

Loading comments…

Subscribe

Five useful things about shipping, every Friday.

No launch announcements, no growth hacks. Just what we read, tested and changed our minds about this week.

We’ll email you a confirmation first. Your address is used only to send you this newsletter, and every issue has an unsubscribe link. Privacy policy

Keep reading