Measuring the Tea Break
I am not a fan of pull requests (PRs). I've written about them, lost bloody battles over them (Hi, Amazon), and helped many software teams move past them.
I advocate for continuous integration (CI), merging code into mainline frequently — typically several times a day. This is technically possible with PRs, but very uncommon. Martin Fowler covers CI in depth in Continuous Integration, as do Humble & Farley in chapter 3 of Continuous Delivery.
Many people see "PR" and think "Code Review." The two are not synonymous. There are better ways to review code, such as pairing and continuous review after a merge, without interfering with CI. As Fowler puts it, "Even when done well, Pre-Integration Reviews always [introduce] some latency into the integration process, encouraging a lower integration frequency."
As AI revolutionizes the software industry, I'm seeing companies turn to measuring PRs: their length, their duration, their frequency. Recently, DX published an that nicely captures this trend: "Is there a relationship between cycle time and PR throughput?"
The research is competent, and the framing is exactly backward.
The report
Across 500+ organizations, DX compared how long PRs sat open against how much each developer shipped. The longer PRs sat, the less each developer shipped, and the effect grew stronger the more a team shipped. Among the teams shipping least, the relationship all but disappeared.
The data looks sound, and the analysis appears correct. But has anyone (besides me) asked, "Why the hell are we measuring PRs?" Why aren't we looking at something that measures value to the customer, like the DORA metrics?
Lead time to change (LTTC) comes to mind. PRs worsen LTTC. Why are we measuring something that slows down delivery?
The tea break
Let's step back from PRs for a moment, and imagine measuring something else, like the runners in the London Marathon. Some runners use caffeine on a race day. Others load up on carbohydrates (sugars) for fuel. Imagine that some trainers started encouraging their runners to stop for tea.
Caffeine? Check. Carbs? Check. Brilliant! Nobody's ledger has a line for the four minutes sitting still.
So if runners are stopping for tea and I measure how quickly the waitress serves it, I'd find that elite tea-drinking runners are sensitive to slow service in tea rooms. I'd have statistics mapped to quantiles, with lots of numbers. It would be an impressive report: "We studied 500 marathon participants and found the elite runners most sensitive to tea service speed..."
Of course, this is nuts. Stopping for tea slows progress toward finishing the marathon. No trainer would encourage a tea stop because she knows it would hurt her client's goal of a fast race. There are better ways to carb-load and caffeinate.
You see, the marathon trainers have figured it out: "tea time hurts race time." Software leaders, by and large, have not. The DX report tells us that teams using PRs have parked changes that "often spend significant amounts of time waiting, not being actively coded or reviewed." We should listen. As leaders, don't make it worse by measuring and lauding a thing that slows your team down. Eliminate the PR.
DORA for the win
Over a decade of research, DORA found that four measures predict software delivery performance: deployment frequency, lead time for changes, change failure rate, and time to restore service. Further, they found that speed and stability rise together rather than trading off.
And DORA identifies trunk-based development (TBD), my preferred approach to CI, as a capability that predicts delivery performance: In Accelerate (2018), Nicole Forsgren, Jez Humble, and Gene Kim — the researchers behind the DORA program — report that TBD was one of the practices separating high performers: "Teams that did well had fewer than three active branches at any time, their branches had very short lifetimes (less than a day) before being merged into trunk ..."
Inserting a PR into the merge process slows down that integration; it adds a constraint to the flow. PRs, as typically used, lengthen your lead time for changes. Adopting TBD shortens it.
The DORA metrics are surprisingly durable. They were effective before AI. They are effective with AI. This is true because they don't care about AI; they reveal how well your software teams deliver, irrespective of how you generated the code.
All the things that lead you to good DORA numbers, like CI/CD, fast pipelines, and robust automated tests, won't tolerate stopping the flow of work for something like a PR. My highest-performing teams have been the ones with good DORA numbers and, not coincidentally, they use CI (no PRs). You need a DORA-strong system to absorb what AI produces.
In the future, I will make a further anti-PR case that humans should no longer be reading code, but that is another article. Stay tuned.
Happy building!
Further Reading
Why Pull Requests Are A BAD IDEA — Dave Farley on how the review gate caps your integration frequency
Non-Blocking Continuous Code Reviews, a Case Study — Thierry De Pauw on a team that reviewed everything, and blocked on nothing
On The Evilness of Feature Branching — Thierry De Pauw on why open-source workflows don't transplant into colocated teams
RefinementCodeReview — Martin Fowler on review as continuous refinement rather than a gate
Ship / Show / Ask — Rouan Wilsenach's graduated alternative, if you're not ready to drop the gate entirely
Async Code Reviews Are Killing Your Company's Throughput — Dragan Stepanovic puts numbers on the queueing cost
References
Seth Orell, Code Review Musings, Ownership Matters, August 9, 2023
Martin Fowler, Continuous Integration, martinfowler.com, updated 2024
Jez Humble & Dave Farley, Continuous Delivery, Addison-Wesley, 2010
Martin Fowler, Patterns for Managing Source Code Branches, martinfowler.com, May 28, 2020
Brian Houck & Grace Fu, , DX, August 19, 2026
Nanci Guest et al., International Society of Sports Nutrition position stand: caffeine and exercise performance, Journal of the International Society of Sports Nutrition 18:1, January 2, 2021
D. Travis Thomas, Kelly Anne Erdman & Louise Burke, Nutrition and Athletic Performance, Medicine & Science in Sports & Exercise 48(3), March 2016
Nicole Forsgren, Jez Humble & Gene Kim, Accelerate: The Science of Lean Software and DevOps, IT Revolution, 2018
Nathen Harvey & Derek DeBellis, State of AI-assisted Software Development, DORA / Google Cloud, September 23, 2025
Originally published at Ownership Matters.