Teams often assume that shipping faster means shipping less safely. Research on software delivery found the opposite: teams that deploy often, in small changes, with automated checks, tend to have lower failure rates and recover faster.
This guide covers the four DORA delivery metrics, how to design quality gates that give fast feedback, why a written definition of done matters, and a worked example in which a team's change failure rate falls from 15.0% to 8.3% as its deployment frequency rises.
Before You Start
Why Quality Gates and Delivery Metrics Matter
Speed and Stability Are Not a Trade-Off
Research summarized in Accelerate found that teams delivering frequently also tend to have lower failure rates and faster recovery.
Gates Give Fast Feedback
Automated checks that run on every change catch problems while the author still remembers the code.
Manual Approvals Slow Without Protecting
A review board that meets weekly adds waiting. Automated gates and small changes reduce risk more effectively.
Metrics Show Whether Changes Work
Four delivery measures show whether practices are improving speed and stability together.
The Four Delivery Metrics
The DORA (DevOps Research and Assessment) research program identified four measures of software delivery performance, described in Accelerate and in the annual State of DevOps reports:
| Metric | Definition | Type |
|---|---|---|
| Deployment frequency | How often the team deploys to production | Throughput |
| Lead time for changes | Time from a change being committed to running in production | Throughput |
| Change failure rate | Share of deployments that cause a failure needing remediation | Stability |
| Time to restore service | How long it takes to recover from a failure in production (recent DORA reports call this failed deployment recovery time) | Stability |
DORA groups teams into performance tiers, and the definitions and thresholds have been refined over time. Check the current report before comparing yourself with a benchmark, and compare mainly with your own trend.
Designing Quality Gates
- Fast first. Put quick checks, such as build, unit tests, and linting, early, and slower ones later.
- Clear ownership and thresholds. Decide what fails a gate (for example, a critical vulnerability) and what only warns.
- Trustworthy checks. Flaky tests train people to ignore failures. Fix or remove them.
- A written definition of done. Agree what "done" means for a change and for a release, including tests, documentation, and monitoring. See the Definition of Done and Quality Gate Checklist.
- Beware targets on proxies. A code coverage target can be met by tests that assert nothing. Use coverage to find untested areas, not as a goal.
Worked Example: Smaller Batches and Automated Gates
A team measured its delivery over 30 days, then introduced automated integration tests, a security scan, and canary releases, and shipped smaller changes. The figures are illustrative.
| Metric | Before | After |
|---|---|---|
| Deployments in 30 days | 40 | 60 |
| Failed deployments | 6 | 5 |
| Change failure rate | 6 / 40 = 15.0% | 5 / 60 = 8.3% |
| Median lead time for changes | 3 days | 1 day |
| Median time to restore | 2 hours | 35 minutes |
Deployment frequency rose by half, lead time fell by two thirds, and the change failure rate fell from 15.0% to 8.3%, even though the team shipped more often. Restore time improved partly because canary releases limited each failure's impact and rollback was automatic. The team does not claim that the gates alone produced the change: batch size fell at the same time, and one month of data is a small sample. It will keep tracking the four measures and also watch a balancing measure, engineer satisfaction, to make sure the process is not just adding toil.
Record the four measures monthly in the Definition of Done and Quality Gate Checklist workbook, which calculates change failure rate and charts the trend.
Where Gates Go in the Pipeline
Gates are placed so that the cheapest, fastest checks run first and the slower, more expensive ones run later on changes that have already passed the early ones. The order gives fast feedback to developers.
Fast first. Compile, unit tests, and linting give feedback within minutes. If these take half an hour, developers will work around them. Move slow checks to later stages or run them in parallel.
Gate on what matters. Each gate should prevent a kind of failure you have seen or fear, such as broken builds, vulnerable dependencies, or failed migrations. A gate with no history of catching anything deserves a review.
Progressive delivery. Canary releases, feature flags, and staged rollouts limit the blast radius. Pair them with monitoring and an automatic or quick rollback.
Treat flaky tests as defects. Tests that fail at random teach people to rerun and ignore failures. Quarantine, fix, or remove them.
Using Delivery Metrics Without Gaming Them
Delivery metrics such as deployment frequency, lead time, change failure rate, and time to restore are useful for teams that want to improve. They become harmful when they are turned into targets for individuals or used to compare teams without context.
- Measure teams and systems, not people. Individual targets on these measures invite cheating and discourage collaboration.
- Use the four together. Speed measures without stability measures can push teams to ship faster and break more. Stability without speed can freeze delivery.
- Define each measure carefully, for example what counts as a deployment and as a failure, and keep the definition constant.
- Look at trends and ranges, not single points, and compare a team with its own past first.
- Pair numbers with conversation. Ask the team what is slowing them and what would help.
Goodhart's law states that when a measure becomes a target it ceases to be a good measure. If code coverage is a target, tests can be written to touch lines without checking anything. Use coverage as a diagnostic (where are we blind?) and not as a goal.
Keep the pipeline itself healthy. Watch build times, flaky test rates, and the time changes wait in queues. A slow or unreliable pipeline is a flow problem that also deserves improvement.
Pitfalls. Comparing teams by raw deployment counts; counting failures inconsistently; gate bloat that slows delivery without reducing risk; and skipping the retrospective on failed changes. See the Flow Metrics Guide, the Incident Management Guide, and the Definition of Done Checklist. Figures in the examples are illustrative.
Self-Assessment Questions
- Do automated checks run on every change and give results in minutes?
- Do we trust our tests, and do we fix flaky ones?
- Do we have a written definition of done?
- Do we track deployment frequency, lead time, change failure rate, and time to restore?
- Do we compare ourselves with our own trend, not just with benchmarks?
Common Mistakes
Gates That Are Too Slow
A gate that takes hours discourages small changes. Keep feedback fast.
Manual Approval as the Main Control
Approval boards add delay and do not reliably catch defects. Automate the checks and reduce batch size.
Gaming the Metrics
Deploying trivial changes inflates frequency. Look at the four measures together, and at outcomes.
Ignoring Restore Time
Failure will happen. Practice recovery, and make rollback quick.
Quality Gates and Delivery Metrics: Frequently Asked Questions
What are the four DORA metrics?
They are deployment frequency (how often you deploy to production), lead time for changes (time from commit to production), change failure rate (the share of deployments causing a failure that needs remediation), and time to restore service after a production failure. The first two measure throughput and the last two measure stability.
What is a quality gate?
A quality gate is an automated or manual checkpoint in a delivery pipeline that a change must pass before moving on, such as passing unit tests, a security scan below a severity threshold, or integration tests. Good gates give fast, clear feedback and block changes that fail.
Should we set a code coverage target?
Use coverage to discover untested areas, but be careful about setting a numeric target as a goal. When coverage becomes a target, teams can meet it with tests that execute code without checking behavior. Pair it with other measures such as escaped defects and change failure rate.
Sources and Further Reading
- Nicole Forsgren, Jez Humble, and Gene Kim, Accelerate: The Science of Lean Software and DevOps.
- DORA (DevOps Research and Assessment), annual State of DevOps reports.
- Jez Humble and David Farley, Continuous Delivery.
- ISO/IEC 25010, systems and software quality models.