Quality gates are checkpoints where code must meet specific criteria before proceeding to the next stage. This page synthesizes practices from Google, Microsoft, and SRE disciplines [1][2][3].
Quality Gates by Stage
Stage 1: Pre-submit (Before Merge)
Gate
Google
Microsoft
Code review
Mandatory, no solo commits
Mandatory peer review
Certification
Readability in language
—
Static analysis
Tricorder
PREfast
Tests
All affected tests pass
All tests pass
Coverage
Varies by Test Certified level
80% for new features
Stage 2: Post-submit (CI/Staging)
Gate
Google
Microsoft
Build status
TAP - main must stay green
Daily builds
Integration tests
Dependency-based selection
Cross-component
Milestone criteria
—
M2: 65%, M3: 70% coverage
Stage 3: Release/Deployment
Gate
Google/SRE
Microsoft
Error budget
Must be positive to deploy
—
Canary
1% rollout first
—
PRR
Production Readiness Review
—
Approval
Automated (budget-based)
Ship Room committee
CI/CD Pipeline
Pipeline Stages
Code → Build → Unit Test → Integration Test → Staging → E2E Test → Canary → Production
Key Practices
Practice
Description
Benefit
Trunk-based development
All developers commit to main
No merge conflicts
Pre-submit testing
Tests run before merge
Catch bugs early
Post-submit testing
Full test suite on main
Catch integration bugs
Canary deployment
1% rollout first
Limit blast radius
Automated rollback
Revert if metrics degrade
Fast recovery
Scale at Google (TAP)
Google’s Test Automation Platform handles massive scale [4]:
Metric
Value
Daily test runs
150 million
Unique test targets
5.5 million
Commit rate
1 per second
Test failure rate
Only 1.15% ever failed
Scale at Facebook
Facebook’s continuous deployment research shows [5]:
Metric
Value
Updates per dev/week
3.5 (~daily)
Critical issues vs deployments
Nearly constant
Feedback time
10 minutes from smoke tests
“The number of critical issues arising from deployments was almost constant regardless of the number of deployments.” [5]
“If feedback takes more than a few minutes, add more machines. CPU hours are cheaper than developer time.”
Facebook’s Approach
10-minute feedback from integration smoke tests
Thousands of compute nodes for test parallelization
3.5 deploys per developer per week sustained
Rollback Strategies
Roll Back vs Roll Forward
Strategy
When to Use
Speed
Roll back
Known-good previous version exists
Fast (seconds-minutes)
Roll forward
Fix is simple, rollback is risky
Varies
Automated Rollback Triggers
Trigger
Action
Error rate spike
Automatic revert
Latency regression
Stop rollout
Crash rate increase
Immediate rollback
Error budget exhausted
Feature freeze
Rollback Requirements
For effective rollback capability:
Immutable deployments - Previous version still available
Database compatibility - Schema supports both versions
Feature flags - Can disable without redeploy
Monitoring - Know when to trigger
Comparison: Continuous vs Milestone
Aspect
Continuous (Google/Facebook)
Milestone (Microsoft)
Release frequency
Multiple times/day
Weeks to months
Quality gate
Error budget
Bug bar
Ship decision
Automated
Committee
Rollback
Instant (flags, canary)
Patch release
Risk model
Small changes, fast feedback
Staged stabilization
Test strategy
Continuous, selective
Milestone exit criteria
Both models achieve high quality through different mechanisms. The choice depends on:
Product type (web service vs packaged software)
User expectations (continuous updates vs stability)
Regulatory requirements (audit trails, approvals)
Key Metrics Summary
Metric
Target
Source
Canary population
1%
Google, Netflix
Pre-submit test time
< 30 min
Google
Feedback loop
< 10 min
Facebook
Deploys per dev/week
3.5
Facebook
Test failure rate
~1%
Google TAP
References
T. Winters, T. Manshreck, and H. Wright, Software Engineering at Google: Lessons Learned from Programming Over Time. O’Reilly Media, 2020.
A. Page, K. Johnston, and B. Rollison, How We Test Software at Microsoft. Microsoft Press, 2008.
B. Beyer, C. Jones, J. Petoff, and N. R. Murphy, Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media, 2016.
A. Memon et al., “Taming Google-Scale Continuous Testing,” in ICSE-SEIP 2017, IEEE, 2017, pp. 233–242. doi: 10.1109/ICSE-SEIP.2017.16.
T. Savor, M. Douglas, M. Gentili, L. Williams, K. Beck, and M. Stumm, “Continuous Deployment at Facebook and OANDA,” in ICSE Companion 2016, ACM, 2016, pp. 21–30. doi: 10.1145/2889160.2889223.
Disclaimer: AI is used for text summarization, polishing and explaining. Authors have verified all facts and claims. In case of an error, feel free to file an issue.