What good teams stop doing, and how they knew it was safe
Freezes, release managers, staging, manual QA gates, sign-off meetings: each bought something once. What replaced each, the evidence needed before deleting it, and the bill that arrives after.
Andrei Gaspar
Every control in a release process was installed by someone who had just been burned. The freeze followed a holiday outage. The staging environment followed a change that worked on a laptop. The sign-off meeting followed a deploy nobody knew was happening. None of them were stupid at the time. Any account of their removal that starts from "these were always waste" was written by someone who was not there for the outage.
The teams that ship well have removed most of these controls. They did not do it by declaring them wasteful. They did it by working out what each control was buying, building a mechanism that bought the same thing more cheaply, collecting the evidence that the mechanism worked, and only then deleting the control. When the order is reversed — delete first, build the replacement after the next incident — the team learns the wrong lesson and reinstalls the control with a stricter version of itself.
This piece walks the five controls teams most often stop: freezes, release managers, staging environments, manual QA gates, and sign-off meetings. For each: what it bought, what replaced it, what evidence should exist before it goes, and what it cost the teams that stopped. Then the bill that arrives after adoption day, and the handful of things that stay a human step on purpose.
The shape of every safe deletion
A control is a way of reducing the cost of a bad change. That cost has two factors: how many users are exposed, and for how long. Every control on the list works by adding verification before the ship, which reduces the probability of a bad change reaching anyone. Every replacement on the list works instead by reducing exposure — fewer users, less time — which reduces the cost of a bad change that does reach someone. Detection speed, undo speed, progressive rollout, flags, canaries, and production verification are all exposure reducers.
Which means the evidence for deleting a pre-ship control is always an exposure number: how fast a bad change is detected, how fast it is undone, and how many users saw it. If a team cannot state those three numbers from measurement, it has not earned the deletion, however tired it is of the control.
| Control | What it bought | What replaces it | Evidence before deleting |
|---|---|---|---|
| Freeze | No change during a fragile period | Fast undo, progressive rollout, on-call coverage | Measured rollback time; change failure rate over the last quarter |
| Release manager | A human who knows the contents and can say no | Merge queue, generated changelog, a rotation | Contents of any release listable by script; anyone on rotation can ship and undo |
| Staging | Seeing everything together before users | Canary slice, preview environments, contract tests | The log of what staging caught, and what else would have caught it |
| Manual QA gate | A human exercising the product | Automated critical-path tests, exploratory testing off the release path, dogfooding | The QA catch log classified by what else would have caught it |
| Sign-off meeting | Shared awareness and an accountability point | The deploy channel, rollout dashboards, generated release notes | The list of decisions the meeting changed in the last quarter |
Freezes
A freeze buys the absence of change during a period when the team is thin, the stakes are high, or both. It is the simplest control there is: nothing goes out, nothing breaks.
Its cost is that it does not stop work. Merges accumulate, or worse, branches grow, and the first deploy after the freeze carries more change than any other deploy in the year, on the day people are still catching up from leave, with a rollback that has to unwind every migration the freeze bottled up. There is also the run-up: the days before a freeze are dangerous because everyone is pushing to land before the door shuts, and the changes that land are the ones that were not quite ready.
What replaces it is a fast undo plus a progressive rollout. If a bad change reaches five percent of traffic, is detected in five minutes and undone in three, its cost over a holiday is a few minutes of a small slice and one page. There is nothing left for the freeze to buy. Some teams keep a weaker form — a reduced-risk window where only flagged changes ship and nobody deploys a migration — and that is an honest middle state.
The evidence before deleting: a rollback time measured in a drill within the last quarter, by the on-call engineer rather than the pipeline's author; a change failure rate low enough that a bad change over the window is unlikely; and an on-call rotation that actually covers the window, because a fast undo is worthless without someone to trigger it. The bill: someone is on call over the holidays, with the flag discipline and rollback tooling that make that a page rather than a weekend.
Release managers
A release manager buys a person who knows what is in the release and has the standing to say no. At around forty engineers this role appears whether anyone creates it or not, because that is the size at which several teams start merging into one deploy and someone has to arbitrate when team A's bad change blocks team B's fix.
Its cost is that it is a human queue with single-threaded capacity and unwritten state. What the release manager knows lives in their head, their decisions are not reproducible, and they are on vacation sometimes. They also become the freeze's owner and the sign-off meeting's chair, so the other controls on this list consolidate around them.
What replaces the role is three mechanisms. A merge queue makes the contents of trunk deterministic: what is in the release is whatever the queue admitted, and a script can list it. Generated release notes — from merged pull request titles or commit trailers — replace the manager's knowledge with something anyone can read. And a rotation supplies the human who watches: Slack's "Deploys at Slack" post describes a deploy commander who owns a deploy window and watches the rollout, which is a rotation rather than a role. The commander watches; the machinery decides.
The evidence before deleting: the contents of any release can be produced by a script; any engineer on the rotation can trigger a deploy and a rollback following what is written down; and a bad change from one team can be reverted without blocking another team's ship. The bill: the rotation itself, which is time from every engineer on it, and the tooling that makes the rotation's job a watch rather than a judgment.
Sponsored:
Staging environments
Staging buys the ability to see everything together before users do. It is the first place all the teams' changes meet, and looking at it feels safer than not looking.
Its costs are drift, contention, and false confidence. Staging's data is not production's data, its scale is not production's scale, and its configuration diverges the week after it was last rebuilt. There is usually one of it, so teams queue for it, which makes it a batching mechanism. And the confidence it gives is calibrated to the wrong distribution: staging catches the bugs that reproduce on small, clean data, and the bugs that hurt in production are the ones that need production's data and load to appear.
What replaces it is a set of things. A canary — a small slice of production traffic on the new version, compared against the old — answers the same question with real data. The Google SRE workbook's chapter on canarying releases is the reference, including why the comparison needs a control running the old version beside it. Preview environments per pull request cover the visual checks UI teams used staging for. Contract tests cover the seams between services. Flags allow a dark launch: the code ships, the feature is off, and the people who built it exercise it in production. Cindy Sridharan's writing on testing in production lays out the broader case, with the caveat that testing in production is a discipline with prerequisites, not a license to skip testing.
The evidence before deleting: a log of what staging actually caught in the last quarter, each entry classified by what else would have caught it. Most teams that do this find the list is short and mostly things a missing unit test or a canary would have found. If it contains things nothing else would have caught — a third-party sandbox integration, a rehearsal of a large migration against a production-shaped snapshot — those are the reasons to keep staging for that class of change while deleting it from the path for everything else.
The bill: the canary's arithmetic. A slice small enough to be safe may be too small to see a small regression — a five percent slice watched for five minutes will not surface an error rate that moved by a fraction of a percent, because the traffic is not there. Teams that delete staging and keep a canary must size the canary against the regression they want to detect, which means knowing their traffic and their error baseline. That is a measurement most staging-era teams have never had to make.
Manual QA gates
A manual QA gate buys a human exercising the product in ways a script would not think to. That is a real thing, and the teams that removed the gate did not remove the humans.
Its cost is cycle time and batching. A gate that takes a day to clear adds a day to every change and encourages bundling, for the same reason slow review does. It has a subtler cost too: a QA gate lets engineers treat correctness as someone else's stage, and the tests they would otherwise have written do not get written.
What replaces the gate is a move, not a deletion. The critical paths — the ones whose breakage is an incident — get automated end-to-end tests in CI. Exploratory testing continues, but off the release path: it runs continuously against production or a preview, files what it finds, and does not block a ship. Employee-first rollout gives the exploratory testers real data and the rest of the team a stake in noticing. The gate becomes a stream.
The evidence before deleting: the QA catch log, classified. Which of last quarter's catches were on a critical path that now has an automated test? Which were exploratory findings the same people would have found post-ship at lower cost? Which would have been caught by nothing else? The third category is the argument for keeping a gate on that class of change. The bill: end-to-end tests are expensive to keep honest. A flaky test either gets fixed promptly or gets ignored, and an ignored flaky test is a gate that is permanently open. Teams that stop manual QA acquire a test-maintenance budget, and the ones that do not fund it end up with neither control.
Sign-off meetings
A sign-off meeting buys shared awareness — everyone knows what is going out — and an accountability point where someone says yes. It usually appears after a deploy that surprised someone senior.
Its cost is that it is a batch, held at a time, attended by people whose attention is the team's scarcest resource, producing decisions that are almost always "yes". A meeting scheduled weekly makes the release weekly regardless of what the pipeline could do.
What replaces it is asynchronous visibility: a deploy channel where every ship and rollback posts automatically with the changelog attached, a rollout dashboard anyone can open, generated release notes. The accountability point moves from a person in a room to a rule in the pipeline: a change that touches a flagged path needs a named approver, and everything else ships on the rotation's watch.
The evidence before deleting: the list of decisions the meeting actually changed in the last quarter. If the answer is "it held one release for a day, once", the meeting costs more attention than it buys. The bill: the automation that posts to the channel, and the cultural work of getting the people who used the meeting to check the channel instead.
The tax that comes due after adoption day
Every replacement above has an adoption story that ends on the day it was switched on. The account worth reading is the one that continues into the following year, because that is when the bill arrives.
The merge queue has a throughput ceiling that is decided by CI time. A queue that verifies each merge against the current head serially can clear at most sixty divided by the CI minutes per hour: twenty-minute CI, three merges an hour, twenty-four in a working day. A team merging fifty times a day discovers this in the first week and has three exits. Faster CI, which is a project. Batching, which reintroduces the bisect problem the queue was supposed to remove. Or speculation — testing several candidate merges in parallel on the assumption that most pass — which is what Uber's "Keeping Master Green at Scale" paper describes for its monorepo, with a probabilistic model deciding what to speculate on. Speculation is machinery with its own failure modes, and flaky tests become queue-killers: a test that fails one time in fifty is a nuisance in CI and a stall in a queue that must be green to advance.
The monorepo, if adopted alongside, brings its own tax: build and test times that scale with the repository unless the build system understands the dependency graph, which means adopting one that does, which means a team to own it. Potvin and Levenberg's "Why Google Stores Billions of Lines of Code in a Single Repository" in Communications of the ACM is honest about this — the monorepo's benefits are real and its tooling investment is not optional.
Flags accumulate. A flag that shipped a feature and was never removed is a branch in the code that nobody tests, and a codebase with hundreds of them has a configuration space nobody can reason about. Pete Hodgson's feature toggles article makes the point that different kinds of toggle have different lifetimes, and the ones meant to be temporary need an expiry and an owner. Teams that replace freezes and staging with flags acquire a flag-cleanup practice or a flag-cleanup problem.
Canaries need the arithmetic above and a baseline to compare against — which means the metrics that define "bad" must exist, be reliable, and be watched by a person or an analysis job. A canary nobody watches is a five-percent deploy with extra steps. Preview environments cost money per pull request and need production-like data to be worth anything, which is a data-masking problem.
None of these bills are reasons not to adopt. They are the reason an adoption account that stops at the switch-on is not evidence of anything.
What stays a human step, on purpose
The teams that have removed the most controls tend to be the clearest about which ones they keep, and the principle is consistent: automate the reversible, keep a human on the irreversible, and make the human step cheap.
The irreversible list is short. A migration that destroys data — dropping a column, deleting rows — gets a named approver and runs at a time someone chose. Anything that leaves the company's control once done: an email to every customer, a price change, a key rotation, a change to a contract-bound API that a partner integrates against with notice periods. The first deploy of a new service, before its rollback has ever been exercised. And the decision to abort a rollout on ambiguous signals, such as a canary that is slightly worse on one metric and slightly better on another, which is a judgment, and the rotation's watcher is there to make it.
The rule that makes this work is that the human step is a button, not a meeting. An approver clicks; they do not attend. The deploy pipeline pauses at the step and posts to a channel; it does not wait for Thursday. A human step that costs an hour of calendar time will be routed around, and then the control is gone without anyone having decided to delete it.
What to do on Monday
Pick one control from the list — the one the team complains about most. Write down, in a sentence, what it buys. Then write down the incident that installed it, if anyone remembers.
Then answer the evidence question from the table. For a freeze: when was rollback last timed, and what was the number? For staging: what did it catch last quarter? For the meeting: what did it change? The answers are usually a day's work to collect and are the entire argument, either way.
If the evidence says the replacement is not in place, the Monday task is building it, and the control stays until it is. If the evidence says the replacement has been doing the job for a quarter, delete the control, announce it, and note the date — so that when the next outage comes and someone proposes reinstalling it, the team can check whether the outage was one the control would have caught, or one the replacement missed.
The teams worth copying did not stop doing things because the things were annoying. They stopped when they could show, with numbers, that the thing had nothing left to buy.
Andrei Gaspar
Editor, How They Ship


Comments
Loading comments…