Code review latency is the number that predicts everything
Hours from PR opened to first real review comment sets batch size, deploy frequency, and how many changes sit half-done. How review runs at teams that ship well, and how to get your own number.
Andrei Gaspar
If you can only know one number about a team's engineering process, make it the hours from a pull request being opened to its first substantive review comment. Not time to merge, which is confounded by how big the change was and how many rounds it took. Not review count, which is a vanity metric. Time to first review.
The reason is what happens while the author waits. An engineer who opens a change and hears nothing for a day does not sit still. They start the next thing, on top of the first, and now the second change depends on an unreviewed first. When the review does arrive and asks for a rework, both changes move. The author learns the lesson every team learns from slow review: open fewer, bigger changes, so the wait is amortized. Bigger changes take longer to review, which lengthens the wait, which teaches the lesson again.
That loop is why review latency predicts everything downstream: pull request size, deploy batch size, deploy frequency, the number of changes in flight, and how a team feels about its own throughput.
Why first-review time, not total review time
Total time in review measures the whole conversation, and the whole conversation is legitimately long for a change that deserves it. A five-hundred-line refactor of the billing path should take three rounds. What matters is whether the author waited a day for the first round to begin.
Time to first review isolates the part that is under the reviewer's control and is felt by the author. It is also almost purely a matter of habit. Google's public engineering practices document on review speed states the expectation plainly: respond shortly after a review arrives if you are not in the middle of focused work, and one business day is the maximum. Google's published study of its review practice — "Modern Code Review: A Case Study at Google", Sadowski and colleagues, 2018 — describes a culture built around small changes and fast turnaround. You do not need Google's tooling to adopt the habit. You need a team agreement about what "shortly" means and a number that shows whether it is being kept.
A working target for a team of ten to forty: a median first review under four working hours, a ninetieth percentile under one working day. Those are not research findings; they are the shape of the distribution at teams whose review process does not show up in their retros as a complaint. The ninetieth percentile matters more than the median: a team can have a two-hour median and still have a fifth of its changes wait three days, and the authors of those changes are the ones bundling.
Review is a queue, and Little's law says what that costs
Little's law is the one piece of queueing theory worth knowing by heart: the average number of items in a system equals the arrival rate multiplied by the average time each item spends there. L = λW. It holds for any stable system regardless of how arrivals are distributed, which makes it safe to apply to review.
A worked illustration: suppose a team of eight engineers opens twenty pull requests a week, and the average time from open to merge is two working days, which is 0.4 of a five-day week. Then on average 20 × 0.4 = 8 pull requests are open at any moment. That is one per engineer, which sounds fine until you remember that each one is a context the author is holding in their head while working on something else.
Now suppose the same team gets its average time in review down to half a day, 0.1 of a week. Open pull requests at any moment: 20 × 0.1 = 2. Six engineers out of eight have nothing waiting. The change in flight is the change they are working on.
Two caveats. First, Little's law describes; it does not explain. It says nothing about how you get W down — more reviewer attention, smaller changes, fewer rounds — only what the system looks like when you do. Second, smaller changes raise λ. If the team splits its work into forty pull requests a week instead of twenty, and gets each through in a quarter of a day, then L = 40 × 0.05 = 2: the same steady state with half the batch size per change. The queue length does not care whether you got there by speed or by size; the authors do, because a small change that waits is cheaper to hold than a large one.
The law also makes capacity drops visible. If two reviewers are out for a week and review time doubles, open changes double with it and do not come back down until the backlog drains. A team that measures its queue sees the spike and can respond. A team that does not experiences the same week as "everything felt slow" and blames the sprint.
Sponsored:
Ownership: CODEOWNERS and the ways it fails
Routing is the part of review that tooling handles well. GitHub's CODEOWNERS file (GitLab has an equivalent) is the standard mechanism: a path pattern maps to the people or teams whose approval is required. It answers the question every open pull request otherwise asks — who is supposed to look at this — and its failure modes are consistent enough to list.
The hot path with two owners. One directory that every feature touches, owned by the two people who built it. Every change in the company queues on two calendars. The per-path latency for that directory is the team's real bottleneck, and it will not appear in the overall median because most changes do not touch it. Measure review latency per owner pattern, not just per repository.
The team owner nobody owns. Ownership assigned to a team alias distributes the notification and diffuses the responsibility. The fix that works is a rotation inside the team — a reviewer of the day, whose job for that day is the queue — so that the alias resolves to a person.
Ownership rot. Files move, people leave, patterns stop matching. An entry that points at a departed engineer either blocks merges or, depending on configuration, silently requires nobody. A monthly check that every owner is current and every pattern matches a file is cheap.
The cross-cutting change. A rename that touches twelve directories needs twelve approvals. This is the mechanism working as designed and producing a bad outcome. Teams handle it with a fallback owner group that can approve repository-wide, or by agreeing that mechanical changes — generated code, formatting, dependency bumps — go through a different path.
CODEOWNERS is a routing table, not a review process. It makes sure a change lands in front of someone with context. It does nothing about how long that person takes.
Stacked diffs, and the problem they exist to solve
A stack is a sequence of small changes, each depending on the one below, each reviewed on its own. The practice grew up at Facebook around Phabricator, where the unit of review was a single diff and a chain of them was natural; the same model has since been re-created on top of Git by tools such as Graphite, ghstack, and Sapling, with varying amounts of friction.
The problem a stack solves is precise. A reviewer can give a fifteen-hundred-line change two kinds of review: a careful one, which takes half a day and delays every other change in the queue, or the one they actually give it, which is a skim with a comment about a variable name. Split into five three-hundred-line changes, each gets a real review in twenty minutes, and the first can merge while the fifth is still being written.
The cost is tooling and habit. Git does not natively understand a stack, so a rework to the second change means rebasing the third, fourth, and fifth, and a pull-request model that assumes one branch per change fights the idea at every step. Teams that adopt stacking either pay for a tool that manages the rebases or accept that a few engineers become the stack experts and everyone else avoids it.
Whether a team needs stacks is decided by its pull-request size distribution, not by fashion. If the median change is already under two hundred lines and the ninetieth percentile under five hundred, stacking is machinery for a problem the team does not have. If the tail is regularly over a thousand lines and those are the changes with multi-day first-review times, the tail is the argument.
Ship, Show, Ask: a dial for trust
Rouan Wilsenach's Ship / Show / Ask, published on martinfowler.com, is the review pattern most worth adopting on a Monday, because it needs no tooling at all. The author puts each change on one of three settings.
Ship: merge straight to trunk, no review. For changes the author is confident in and the team has agreed are low risk: a typo, a dependency bump with green tests, a change in a corner of the codebase the author owns outright.
Show: open the pull request and merge it immediately, then let people review after the fact. The change ships; the conversation still happens; nobody waits. For changes where the author wants a second pair of eyes but does not need one to proceed, and for changes worth showing to the team as a pattern.
Ask: open the pull request and wait. For changes the author is unsure about, changes in areas they do not know, and changes that would be expensive to undo.
The pattern's value for latency is that it takes an entire class of changes out of the queue. In the illustration above, if a third of the team's changes are Ship or Show, λ for the Ask queue drops by a third, and reviewer attention concentrates on the changes that actually wanted it. Its value for culture is that it makes trust explicit and adjustable. A new engineer's dial sits at Ask for a while; a senior engineer's sits mostly at Show; the team can talk about moving it. A team where the dial is welded to Ask for everyone has decided that no one is trusted, and its review latency will say so.
The prerequisite is a fast undo. Ship and Show are only safe when a wrong change can be reverted in minutes and detected before it reaches many users. A team with a ninety-minute rollback and no progressive rollout should fix the rollback first.
What gets measured, and what got gamed
Once a team publishes a review metric, the metric starts to change the behavior it measures, and not always in the intended direction. The games are predictable.
The rubber stamp. If time to first review is the number, the cheapest way to move it is to approve within a minute of opening. Detect it by plotting approval latency against lines changed; the cluster of fast approvals on large changes is the problem. Counter it by measuring time to first comment as well as time to first approval, and by treating a zero-comment approval on a change over some size as a review that did not happen.
The size game. If pull-request size is the number, changes get split into pieces that pass the threshold and mean nothing on their own — a commit that adds an unused function, a commit that adds its only caller. Or the reverse: if review count is the number, changes get bundled to reduce it. The tell is a change that cannot be understood without its siblings, which is precisely the property a stack is supposed to have solved.
The LGTM escape. A reviewer who has been asked to keep latency down and has too many changes in the queue learns that "LGTM" is faster than reading. The comments-per-review ratio drifts to zero, the change failure rate drifts up a quarter later, and nobody connects them.
The honest response is to measure a small basket rather than a single number, to look at distributions rather than averages, and to look at the metrics in a retro rather than on a leaderboard. Time to first review, time to merge, lines changed, comments per review, and — the one that keeps the others honest — the fraction of merged changes later reverted or hot-fixed. A team that improves the first four and worsens the fifth has learned to game its own dashboard.
Batch size and deploy frequency
Review latency sets a floor under lead time. If the median first review takes a day, no change reaches production in less than a day plus everything else. Engineers then make a rational choice: if every change costs a day of waiting regardless of size, make each change worth the wait. Changes get bigger. Bigger changes need more review rounds, take longer to verify, and fail more often when deployed, because the search space when something goes wrong is the whole change.
The DORA research has for years reported the same association across its survey population: teams with short lead times deploy more often and have lower change failure rates, not higher. The mechanism is batch size, and review latency is where batch size is decided. A team that wants to deploy more often and has not looked at its review queue is working on the wrong end of the pipeline.
The connection runs the other way too. A team that ships on merge with a progressive rollout has made every change individually visible in production, which makes small changes cheap to ship and large ones expensive to debug. The deploy cadence pulls review toward small changes; small changes pull review latency down; the loop reverses.
Onboarding reviewers without burning them
The teams that keep latency low do not have a review specialist. Everyone reviews, which means every new engineer becomes a reviewer early, and how that is done decides whether they become a good one or a burned-out one.
A workable pattern: from the first week, the new engineer is a second reviewer alongside someone experienced, on small changes in the area they are learning. Their job is to read and ask questions, not to approve. After a few weeks they become the first reviewer on small changes, with the experienced reviewer watching. The dial moves toward independent review as their comments start catching things.
Three protections matter. The new reviewer must be explicitly allowed to say "I do not know this area well enough" and hand it on, or the queue fills with reviews they are afraid to decline and afraid to approve. Their review load needs a cap — a rotation with a maximum number of open reviews per person, not a free-for-all where the most conscientious person absorbs the most. And they must not be made the second owner of the hot path just because the first owner needs relief; that is how a team creates the next bottleneck while solving the current one. The reviewer-of-the-day rotation earns its keep here: a bounded day with the queue as the only responsibility, instead of an unbounded trickle competing with their own work.
The numbers a team should know, and how to get them cheaply
None of this needs a product. The data is in the version control host, and a script can pull it in an afternoon.
| Number | What it tells you | Where to get it |
|---|---|---|
| Time to first review (median, p90) | The wait authors feel | PR opened time to first review or comment timestamp |
| Time to merge (median, p90) | Total lead time through review | PR opened to merged |
| Lines changed per PR (median, p90) | Batch size and whether the tail needs stacking | Additions plus deletions |
| Comments per approval | Whether review is happening or being stamped | Review comment count divided by approvals |
| Per-path latency | Where the hot path is | Group the above by owner pattern |
| Open PRs at any moment | The queue, for Little's law | Count of open, non-draft PRs sampled daily |
| Reverts per hundred merges | Whether speed is costing correctness | Merged PRs whose title or body reverts a prior one |
With the GitHub CLI, a single command gets the raw material for most of the table:
gh pr list --state merged --limit 500 \
--json number,createdAt,mergedAt,additions,deletions,reviews,comments,filesA short script computes the intervals, takes medians and ninetieth percentiles, and groups by the paths in files against the owner patterns. Run it weekly and post the table in the team channel. The first week the numbers will surprise someone; the fourth week they will have moved without anyone being told to move them, because a visible queue gets attended to.
Two cautions. Use working hours, not wall-clock, or every change opened on a Friday afternoon looks like a disaster. And use percentiles, never means; a single change that waited two weeks drags an average past any target and says nothing about the other forty-nine.
What to do on Monday
Pull the table above for the last ninety days. Find the ninetieth-percentile time to first review and look at the changes in that tail — the specific ones, with names. Ask what they have in common. It will be one of three things: a path with too few owners, a size nobody wanted to read, or a reviewer whose queue was full because they are also the person everyone asks.
Then pick the smallest intervention that addresses it: a rotation, a fallback owner group, an agreement that changes over some size get split before they are opened, or a team conversation about which changes belong on Show instead of Ask. Rerun the table in a month.
The teams that ship well did not get there by tolerating slow review and compensating downstream. They treated the review queue as the first stage of the deploy pipeline, gave it a number, and kept the number where authors could feel it.
Andrei Gaspar
Editor, How They Ship



Comments
Loading comments…