Selected Materials
A curated selection of books and papers from the SQR and QAM courses. Each entry is chosen for its foundational or canonical contribution — not for volume. One annotation line explains what it contributes and why it is here.
Textbooks
Naik, K. & Tripathy, P. — Software Testing and Quality Assurance: Theory and Practice (2008, Wiley) The primary SQR course textbook; covers the full testing lifecycle from black-box methods through reliability engineering in a single coherent framework.
[1] — Introduction to Software Testing (Ammann & Offutt, 2016, Cambridge UP) Rigorous graph-model treatment of coverage criteria and test generation; the theoretical foundation behind the adequacy-criteria and domain-testing lectures.
Industry Books
[2] — Software Engineering at Google (Winters, Manshreck & Wright, 2020, O’Reilly) How Google operationalises the 70/20/10 test pyramid, mandatory code review, and error budgets; the industry benchmark for engineering quality at scale.
[3] — Site Reliability Engineering (Beyer et al., 2016, O’Reilly) Introduces SLOs, error budgets, and toil reduction as engineering disciplines; the definitive source for reliability as a practice, not just a goal.
[4] — How We Test Software at Microsoft (Page, Johnston & Rollison, 2008, Microsoft Press) Microsoft’s milestone quality gates, SDET role, and ship-room culture; the counterpart to the Google book, showing a different model that reaches similar conclusions.
Whittaker, J., Arbon, J. & Carollo, J. — How Google Tests Software (2012, Addison-Wesley) Explains Google’s testing org structure, the “testing on the toilet” programme, and continuous integration; complements SWE at Google with a practitioner’s perspective.
[5] — The Art of Computer Systems Performance Analysis (Jain, 1991, Wiley) The reference textbook for systematic performance measurement and modelling; covers experimental design, workload characterisation, and queuing models in one volume.
[6] — Building Maintainable Software (Visser et al., 2016, O’Reilly) SIG’s ten guidelines with industry-calibrated thresholds; translates abstract maintainability into actionable, measurable code-level rules.
Key Papers by Topic
Defining Quality (L1)
[7] — Five quality views (transcendental, user, manufacturing, product, value). The conceptual lens every quality model discussion starts from; still the most-cited framing paper fifty years later.
[8] — First metric-based quality factor–criterion–metric hierarchy. The direct ancestor of ISO 25010; shows how abstract quality goals decompose into measurable attributes.
Metrics and Measurement (L2)
[9] — Goal/Question/Metric paradigm. The “define the question before choosing the metric” discipline that prevents vanity metrics and governs measurement programmes.
[10] — Cyclomatic complexity as a structural coverage criterion. Foundational for basis-path testing and code-risk classification; over 4000 citations.
Verification Overview (L3)
[11] — TDD at Microsoft and IBM: 40–90% fewer defects with 15–35% longer initial development. The most-cited empirical trade-off data for TDD adoption decisions.
[12] — HP inspection ROI: 10:1 return ratio. The canonical industry-scale business case for shifting defect removal earlier in the lifecycle.
Coverage Criteria (L4)
[13] — When test-suite size is controlled, coverage–fault-detection correlation drops to near zero. The ICSE landmark result that grounds all critical discussion of coverage as a quality proxy.
[14] — MC/DC criterion applicability in avionics (DO-178B). Explains why safety-critical domains require decision coverage beyond simple branch coverage.
Random Testing and Fuzzing (L6)
[15] — SAGE whitebox fuzzer found one-third of all Windows 7 security bugs before release. The industrial-scale proof that greybox fuzzing outperforms black-box random testing.
Code Inspection and Review (A5)
[16] — Original IBM formal inspection process; 90%+ defect detection before testing. The paper that established structured review as an engineering discipline.
[17] — Cisco study (200 programmers, 3.2M LOC): 8–12× cheaper per defect than testing; 200–400 LOC/hour is the optimal review rate. The most-cited empirical guide for review calibration.
[18] — Modern code review at Google: median 24 lines, 1 reviewer, completed in under 4 hours. Documents how asynchronous lightweight review works at scale without formal inspection overhead.
Static Analysis (A6)
[19] — Static analysis of Linux/OS X by modelling developer beliefs found 1000+ bugs. The “bugs as deviant behaviour” framing that underpins modern inter-procedural checkers.
[20] — Google Tricorder at scale: the 10%-false-positive rule and developer-facing presentation that drove adoption. Shows that usability, not precision, determines whether static analysis is actually used.
Combinatorial Testing (A7)
[21] — NIST study: 93%/98%/100% fault detection at 2/3/4-way interaction coverage; no failure required more than 6-way. The empirical foundation for pairwise testing adoption.
[22] — AETG algorithm: pairwise test-suite size grows logarithmically vs. exponentially for exhaustive testing. The economic argument that makes combinatorial testing practical.
[23] — Introduction to Combinatorial Testing (Kuhn, Kacker & Lei, CRC Press). The primary reference textbook for the combinatorial testing lecture.
Software Reliability (L8)
[24] — Operational profiles and reliability growth models. The framework for measuring, predicting, and managing software reliability as a quantifiable engineering property.
[25] — Dependability taxonomy: fault → error → failure chain; unified vocabulary for reliability and security. The standard reference for any precise reliability discussion.
[26] — Empirical proof that independently developed modules fail on the same inputs. The decisive evidence against naive N-version programming as a fault-tolerance strategy.
Performance and Queuing Theory (L9, A9)
[27] — Little’s Law L = λW. The single equation from which all queuing results follow; essential for any capacity-planning discussion.
[28] — “The Tail at Scale”: p99/p999 latency dominates user experience and how to tame it with hedged requests. Reframed how the industry thinks about latency SLOs.
[29] — Universal Scalability Law extends Amdahl’s Law with a coherence penalty. The practical formula for predicting scalability limits from measured throughput data.
Security (L10)
[30] — Eight protection principles (least privilege, fail-safe defaults, complete mediation, …). Fifty years old and still the design checklist every secure system starts from.
[31] — Definitive threat modelling reference: STRIDE, attack trees, the four-question framework. The book that operationalised threat modelling for software teams.
[32] — “Build security in” via touchpoints: 50% of security issues require architectural fixes that testing cannot catch. The argument for shifting security left.
Maintainability (A10)
[33] — Lehman’s Laws: a system that is not actively adapted will degrade in quality over time. The theoretical basis for continuous refactoring investment.
[34] — “Software Aging”: programs must be rejuvenated or retired; passive maintenance is not sufficient. Gives practitioners a vocabulary for the technical-debt conversation.
[35] — Mozilla and Linux DSM study: highly-tangled architectures accumulate 2–4× more defects in changed modules. Empirical link between coupling and defect density.
Cost of Quality (A1)
[36] — NIST study: inadequate software testing costs the US economy $22–60 billion annually. The macroeconomic data behind the business case for quality investment.
[37] — The modern CoQ model: optimal quality spend shifts toward 100% conformance as processes improve. Challenges the traditional view that quality and cost necessarily trade off.
- P. Ammann and J. Offutt, Introduction to Software Testing, 2nd ed. Cambridge University Press, 2016.
- T. Winters, T. Manshreck, and H. Wright, Software Engineering at Google: Lessons Learned from Programming Over Time. O’Reilly Media, 2020.
- B. Beyer, C. Jones, J. Petoff, and N. R. Murphy, Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media, 2016.
- A. Page, K. Johnston, and B. Rollison, How We Test Software at Microsoft. Microsoft Press, 2008.
- R. Jain, The Art of Computer Systems Performance Analysis. Wiley, 1991.
- J. Visser, S. Rigal, R. van der Leij, P. van Eck, and G. Wijnholds, Building Maintainable Software: Ten Guidelines for Future-Proof Code. O’Reilly Media, 2016.
- D. A. Garvin, “What Does ‘Product Quality’ Really Mean?,” Sloan Management Review, p. 25, 1984.
- J. A. McCall, P. K. Richards, and G. F. Walters, “Factors in Software Quality. Volume I. Concepts and Definitions of Software Quality,” GENERAL ELECTRIC CO SUNNYVALE CA, Nov. 1977. Accessed: January 19, 2022. [Online]. Available at: https://apps.dtic.mil/sti/citations/ADA049014
- V. R. Basili, “Applying the Goal/Question/Metric paradigm in the experience factory,” Software quality assurance and measurement: A worldwide perspective, vol. 7, no. 4, pp. 21–44, 1993.
- T. J. McCabe, “A complexity measure,” IEEE Transactions on software Engineering, no. 4, pp. 308–320, 1976.
- N. Nagappan, E. M. Maximilien, T. Bhat, and L. Williams, “Realizing quality improvement through test driven development: results and experiences of four industrial teams,” Empirical Software Engineering, vol. 13, no. 3, pp. 289–302, 2008, doi: 10.1007/s10664-008-9062-z.
- R. B. Grady and T. Van Slack, “Key lessons in achieving widespread inspection use,” IEEE Software, vol. 11, no. 4, pp. 46–57, 1994, doi: 10.1109/52.300082.
- L. Inozemtseva and R. Holmes, “Coverage is Not Strongly Correlated with Test Suite Effectiveness,” in International Conference on Software Engineering (ICSE), ACM, 2014, pp. 435–445. doi: 10.1145/2568225.2568271.
- J. J. Chilenski and S. P. Miller, “Applicability of Modified Condition/Decision Coverage to Software Testing,” Software Engineering Journal, vol. 9, no. 5, pp. 193–200, 1994, doi: 10.1049/sej.1994.0025.
- E. Bounimova, P. Godefroid, and D. Molnar, “Billions and Billions of Constraints: Whitebox Fuzz Testing in Production,” in Proceedings of the 35th International Conference on Software Engineering (ICSE), 2013, pp. 122–131.
- M. E. Fagan, “Design and Code Inspections to Reduce Errors in Program Development,” IBM Systems Journal, vol. 15, no. 3, pp. 182–211, 1976, doi: 10.1147/sj.153.0182.
- J. Cohen, S. Teleki, and E. Brown, Best Kept Secrets of Peer Code Review. SmartBear Software, 2006.
- C. Sadowski, E. Söderberg, L. Church, M. Sipko, and A. Bacchelli, “Modern Code Review: A Case Study at Google,” in ICSE-SEIP 2018, ACM, 2018, pp. 181–190. doi: 10.1145/3183519.3183525.
- D. Engler, D. Y. Chen, S. Hallem, A. Chou, and B. Chelf, “Bugs as Deviant Behavior: A General Approach to Inferring Errors in Systems Code,” in Proceedings of the 18th ACM Symposium on Operating Systems Principles (SOSP), 2001, pp. 57–72. doi: 10.1145/502034.502041.
- C. Sadowski, J. van Gogh, C. Jaspan, E. Söderberg, and C. Winter, “Tricorder: Building a Program Analysis Ecosystem,” in Proceedings of the 37th International Conference on Software Engineering (ICSE), 2015, pp. 598–608. doi: 10.1109/ICSE.2015.76.
- D. R. Kuhn, D. R. Wallace, and A. M. Gallo, “Software Fault Interactions and Implications for Software Testing,” IEEE Transactions on Software Engineering, vol. 30, no. 6, pp. 418–421, June 2004, doi: 10.1109/TSE.2004.24.
- D. M. Cohen, S. R. Dalal, M. L. Fredman, and G. C. Patton, “The AETG System: An Approach to Testing Based on Combinatorial Design,” IEEE Transactions on Software Engineering, vol. 23, no. 7, pp. 437–444, 1997, doi: 10.1109/32.605761.
- D. R. Kuhn, R. N. Kacker, and Y. Lei, Introduction to Combinatorial Testing. CRC Press, 2013.
- J. D. Musa, Software Reliability Engineering: More Reliable Software Faster and Cheaper, 2nd ed. McGraw-Hill, 2004.
- A. Avizienis, J.-C. Laprie, B. Randell, and C. Landwehr, “Basic Concepts and Taxonomy of Dependable and Secure Computing,” IEEE Transactions on Dependable and Secure Computing, vol. 1, no. 1, pp. 11–33, 2004, doi: 10.1109/TDSC.2004.2.
- J. C. Knight and N. G. Leveson, “An Experimental Evaluation of the Assumption of Independence in Multiversion Programming,” IEEE Transactions on Software Engineering, vol. SE-12, no. 1, pp. 96–109, 1986, doi: 10.1109/TSE.1986.6312924.
- J. D. C. Little, “A Proof for the Queuing Formula: L = λW,” Operations Research, vol. 9, no. 3, pp. 383–387, 1961, doi: 10.1287/opre.9.3.383.
- J. Dean and L. A. Barroso, “The Tail at Scale,” Communications of the ACM, vol. 56, no. 2, pp. 74–80, 2013, doi: 10.1145/2408776.2408794.
- N. J. Gunther, Guerrilla Capacity Planning: A Tactical Approach to Planning for Highly Scalable Applications and Services. Springer, 2007.
- J. H. Saltzer and M. D. Schroeder, “The Protection of Information in Computer Systems,” Proceedings of the IEEE, vol. 63, no. 9, pp. 1278–1308, 1975, doi: 10.1109/PROC.1975.9939.
- A. Shostack, Threat Modeling: Designing for Security. Wiley, 2014.
- G. McGraw, Software Security: Building Security In. Addison-Wesley, 2006.
- M. M. Lehman, “Programs, Life Cycles, and Laws of Software Evolution,” Proceedings of the IEEE, vol. 68, no. 9, pp. 1060–1076, 1980, doi: 10.1109/PROC.1980.11805.
- D. L. Parnas, “Software Aging,” in Proceedings of the 16th International Conference on Software Engineering (ICSE), IEEE Computer Society Press, 1994, pp. 279–287. doi: 10.1109/ICSE.1994.296790.
- A. MacCormack, J. Rusnak, and C. Y. Baldwin, “Exploring the Structure of Complex Software Designs: An Empirical Study of Open Source and Proprietary Code,” Management Science, vol. 52, no. 7, pp. 1015–1030, 2006, doi: 10.1287/mnsc.1060.0552.
- G. Tassey, “The Economic Impacts of Inadequate Infrastructure for Software Testing,” National Institute of Standards and Technology (NIST), Planning Report 02-3, 2002.
- S. T. Knox, “Modeling the Cost of Software Quality,” 4, 1993.