Selected Materials

A curated selection of books and papers from the SQR and QAM courses. Each entry is chosen for its foundational or canonical contribution — not for volume. One annotation line explains what it contributes and why it is here.


Textbooks

Naik, K. & Tripathy, P.Software Testing and Quality Assurance: Theory and Practice (2008, Wiley) The primary SQR course textbook; covers the full testing lifecycle from black-box methods through reliability engineering in a single coherent framework.

[1]Introduction to Software Testing (Ammann & Offutt, 2016, Cambridge UP) Rigorous graph-model treatment of coverage criteria and test generation; the theoretical foundation behind the adequacy-criteria and domain-testing lectures.


Industry Books

[2]Software Engineering at Google (Winters, Manshreck & Wright, 2020, O’Reilly) How Google operationalises the 70/20/10 test pyramid, mandatory code review, and error budgets; the industry benchmark for engineering quality at scale.

[3]Site Reliability Engineering (Beyer et al., 2016, O’Reilly) Introduces SLOs, error budgets, and toil reduction as engineering disciplines; the definitive source for reliability as a practice, not just a goal.

[4]How We Test Software at Microsoft (Page, Johnston & Rollison, 2008, Microsoft Press) Microsoft’s milestone quality gates, SDET role, and ship-room culture; the counterpart to the Google book, showing a different model that reaches similar conclusions.

Whittaker, J., Arbon, J. & Carollo, J.How Google Tests Software (2012, Addison-Wesley) Explains Google’s testing org structure, the “testing on the toilet” programme, and continuous integration; complements SWE at Google with a practitioner’s perspective.

[5]The Art of Computer Systems Performance Analysis (Jain, 1991, Wiley) The reference textbook for systematic performance measurement and modelling; covers experimental design, workload characterisation, and queuing models in one volume.

[6]Building Maintainable Software (Visser et al., 2016, O’Reilly) SIG’s ten guidelines with industry-calibrated thresholds; translates abstract maintainability into actionable, measurable code-level rules.


Key Papers by Topic

Defining Quality (L1)

[7] — Five quality views (transcendental, user, manufacturing, product, value). The conceptual lens every quality model discussion starts from; still the most-cited framing paper fifty years later.

[8] — First metric-based quality factor–criterion–metric hierarchy. The direct ancestor of ISO 25010; shows how abstract quality goals decompose into measurable attributes.

Metrics and Measurement (L2)

[9] — Goal/Question/Metric paradigm. The “define the question before choosing the metric” discipline that prevents vanity metrics and governs measurement programmes.

[10] — Cyclomatic complexity as a structural coverage criterion. Foundational for basis-path testing and code-risk classification; over 4000 citations.

Verification Overview (L3)

[11] — TDD at Microsoft and IBM: 40–90% fewer defects with 15–35% longer initial development. The most-cited empirical trade-off data for TDD adoption decisions.

[12] — HP inspection ROI: 10:1 return ratio. The canonical industry-scale business case for shifting defect removal earlier in the lifecycle.

Coverage Criteria (L4)

[13] — When test-suite size is controlled, coverage–fault-detection correlation drops to near zero. The ICSE landmark result that grounds all critical discussion of coverage as a quality proxy.

[14] — MC/DC criterion applicability in avionics (DO-178B). Explains why safety-critical domains require decision coverage beyond simple branch coverage.

Random Testing and Fuzzing (L6)

[15] — SAGE whitebox fuzzer found one-third of all Windows 7 security bugs before release. The industrial-scale proof that greybox fuzzing outperforms black-box random testing.

Code Inspection and Review (A5)

[16] — Original IBM formal inspection process; 90%+ defect detection before testing. The paper that established structured review as an engineering discipline.

[17] — Cisco study (200 programmers, 3.2M LOC): 8–12× cheaper per defect than testing; 200–400 LOC/hour is the optimal review rate. The most-cited empirical guide for review calibration.

[18] — Modern code review at Google: median 24 lines, 1 reviewer, completed in under 4 hours. Documents how asynchronous lightweight review works at scale without formal inspection overhead.

Static Analysis (A6)

[19] — Static analysis of Linux/OS X by modelling developer beliefs found 1000+ bugs. The “bugs as deviant behaviour” framing that underpins modern inter-procedural checkers.

[20] — Google Tricorder at scale: the 10%-false-positive rule and developer-facing presentation that drove adoption. Shows that usability, not precision, determines whether static analysis is actually used.

Combinatorial Testing (A7)

[21] — NIST study: 93%/98%/100% fault detection at 2/3/4-way interaction coverage; no failure required more than 6-way. The empirical foundation for pairwise testing adoption.

[22] — AETG algorithm: pairwise test-suite size grows logarithmically vs. exponentially for exhaustive testing. The economic argument that makes combinatorial testing practical.

[23]Introduction to Combinatorial Testing (Kuhn, Kacker & Lei, CRC Press). The primary reference textbook for the combinatorial testing lecture.

Software Reliability (L8)

[24] — Operational profiles and reliability growth models. The framework for measuring, predicting, and managing software reliability as a quantifiable engineering property.

[25] — Dependability taxonomy: fault → error → failure chain; unified vocabulary for reliability and security. The standard reference for any precise reliability discussion.

[26] — Empirical proof that independently developed modules fail on the same inputs. The decisive evidence against naive N-version programming as a fault-tolerance strategy.

Performance and Queuing Theory (L9, A9)

[27] — Little’s Law L = λW. The single equation from which all queuing results follow; essential for any capacity-planning discussion.

[28] — “The Tail at Scale”: p99/p999 latency dominates user experience and how to tame it with hedged requests. Reframed how the industry thinks about latency SLOs.

[29] — Universal Scalability Law extends Amdahl’s Law with a coherence penalty. The practical formula for predicting scalability limits from measured throughput data.

Security (L10)

[30] — Eight protection principles (least privilege, fail-safe defaults, complete mediation, …). Fifty years old and still the design checklist every secure system starts from.

[31] — Definitive threat modelling reference: STRIDE, attack trees, the four-question framework. The book that operationalised threat modelling for software teams.

[32] — “Build security in” via touchpoints: 50% of security issues require architectural fixes that testing cannot catch. The argument for shifting security left.

Maintainability (A10)

[33] — Lehman’s Laws: a system that is not actively adapted will degrade in quality over time. The theoretical basis for continuous refactoring investment.

[34] — “Software Aging”: programs must be rejuvenated or retired; passive maintenance is not sufficient. Gives practitioners a vocabulary for the technical-debt conversation.

[35] — Mozilla and Linux DSM study: highly-tangled architectures accumulate 2–4× more defects in changed modules. Empirical link between coupling and defect density.

Cost of Quality (A1)

[36] — NIST study: inadequate software testing costs the US economy $22–60 billion annually. The macroeconomic data behind the business case for quality investment.

[37] — The modern CoQ model: optimal quality spend shifts toward 100% conformance as processes improve. Challenges the traditional view that quality and cost necessarily trade off.


  1. P. Ammann and J. Offutt, Introduction to Software Testing, 2nd ed. Cambridge University Press, 2016.
  2. T. Winters, T. Manshreck, and H. Wright, Software Engineering at Google: Lessons Learned from Programming Over Time. O’Reilly Media, 2020.
  3. B. Beyer, C. Jones, J. Petoff, and N. R. Murphy, Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media, 2016.
  4. A. Page, K. Johnston, and B. Rollison, How We Test Software at Microsoft. Microsoft Press, 2008.
  5. R. Jain, The Art of Computer Systems Performance Analysis. Wiley, 1991.
  6. J. Visser, S. Rigal, R. van der Leij, P. van Eck, and G. Wijnholds, Building Maintainable Software: Ten Guidelines for Future-Proof Code. O’Reilly Media, 2016.
  7. D. A. Garvin, “What Does ‘Product Quality’ Really Mean?,” Sloan Management Review, p. 25, 1984.
  8. J. A. McCall, P. K. Richards, and G. F. Walters, “Factors in Software Quality. Volume I. Concepts and Definitions of Software Quality,” GENERAL ELECTRIC CO SUNNYVALE CA, Nov. 1977. Accessed: January 19, 2022. [Online]. Available at: https://apps.dtic.mil/sti/citations/ADA049014
  9. V. R. Basili, “Applying the Goal/Question/Metric paradigm in the experience factory,” Software quality assurance and measurement: A worldwide perspective, vol. 7, no. 4, pp. 21–44, 1993.
  10. T. J. McCabe, “A complexity measure,” IEEE Transactions on software Engineering, no. 4, pp. 308–320, 1976.
  11. N. Nagappan, E. M. Maximilien, T. Bhat, and L. Williams, “Realizing quality improvement through test driven development: results and experiences of four industrial teams,” Empirical Software Engineering, vol. 13, no. 3, pp. 289–302, 2008, doi: 10.1007/s10664-008-9062-z.
  12. R. B. Grady and T. Van Slack, “Key lessons in achieving widespread inspection use,” IEEE Software, vol. 11, no. 4, pp. 46–57, 1994, doi: 10.1109/52.300082.
  13. L. Inozemtseva and R. Holmes, “Coverage is Not Strongly Correlated with Test Suite Effectiveness,” in International Conference on Software Engineering (ICSE), ACM, 2014, pp. 435–445. doi: 10.1145/2568225.2568271.
  14. J. J. Chilenski and S. P. Miller, “Applicability of Modified Condition/Decision Coverage to Software Testing,” Software Engineering Journal, vol. 9, no. 5, pp. 193–200, 1994, doi: 10.1049/sej.1994.0025.
  15. E. Bounimova, P. Godefroid, and D. Molnar, “Billions and Billions of Constraints: Whitebox Fuzz Testing in Production,” in Proceedings of the 35th International Conference on Software Engineering (ICSE), 2013, pp. 122–131.
  16. M. E. Fagan, “Design and Code Inspections to Reduce Errors in Program Development,” IBM Systems Journal, vol. 15, no. 3, pp. 182–211, 1976, doi: 10.1147/sj.153.0182.
  17. J. Cohen, S. Teleki, and E. Brown, Best Kept Secrets of Peer Code Review. SmartBear Software, 2006.
  18. C. Sadowski, E. Söderberg, L. Church, M. Sipko, and A. Bacchelli, “Modern Code Review: A Case Study at Google,” in ICSE-SEIP 2018, ACM, 2018, pp. 181–190. doi: 10.1145/3183519.3183525.
  19. D. Engler, D. Y. Chen, S. Hallem, A. Chou, and B. Chelf, “Bugs as Deviant Behavior: A General Approach to Inferring Errors in Systems Code,” in Proceedings of the 18th ACM Symposium on Operating Systems Principles (SOSP), 2001, pp. 57–72. doi: 10.1145/502034.502041.
  20. C. Sadowski, J. van Gogh, C. Jaspan, E. Söderberg, and C. Winter, “Tricorder: Building a Program Analysis Ecosystem,” in Proceedings of the 37th International Conference on Software Engineering (ICSE), 2015, pp. 598–608. doi: 10.1109/ICSE.2015.76.
  21. D. R. Kuhn, D. R. Wallace, and A. M. Gallo, “Software Fault Interactions and Implications for Software Testing,” IEEE Transactions on Software Engineering, vol. 30, no. 6, pp. 418–421, June 2004, doi: 10.1109/TSE.2004.24.
  22. D. M. Cohen, S. R. Dalal, M. L. Fredman, and G. C. Patton, “The AETG System: An Approach to Testing Based on Combinatorial Design,” IEEE Transactions on Software Engineering, vol. 23, no. 7, pp. 437–444, 1997, doi: 10.1109/32.605761.
  23. D. R. Kuhn, R. N. Kacker, and Y. Lei, Introduction to Combinatorial Testing. CRC Press, 2013.
  24. J. D. Musa, Software Reliability Engineering: More Reliable Software Faster and Cheaper, 2nd ed. McGraw-Hill, 2004.
  25. A. Avizienis, J.-C. Laprie, B. Randell, and C. Landwehr, “Basic Concepts and Taxonomy of Dependable and Secure Computing,” IEEE Transactions on Dependable and Secure Computing, vol. 1, no. 1, pp. 11–33, 2004, doi: 10.1109/TDSC.2004.2.
  26. J. C. Knight and N. G. Leveson, “An Experimental Evaluation of the Assumption of Independence in Multiversion Programming,” IEEE Transactions on Software Engineering, vol. SE-12, no. 1, pp. 96–109, 1986, doi: 10.1109/TSE.1986.6312924.
  27. J. D. C. Little, “A Proof for the Queuing Formula: L = λW,” Operations Research, vol. 9, no. 3, pp. 383–387, 1961, doi: 10.1287/opre.9.3.383.
  28. J. Dean and L. A. Barroso, “The Tail at Scale,” Communications of the ACM, vol. 56, no. 2, pp. 74–80, 2013, doi: 10.1145/2408776.2408794.
  29. N. J. Gunther, Guerrilla Capacity Planning: A Tactical Approach to Planning for Highly Scalable Applications and Services. Springer, 2007.
  30. J. H. Saltzer and M. D. Schroeder, “The Protection of Information in Computer Systems,” Proceedings of the IEEE, vol. 63, no. 9, pp. 1278–1308, 1975, doi: 10.1109/PROC.1975.9939.
  31. A. Shostack, Threat Modeling: Designing for Security. Wiley, 2014.
  32. G. McGraw, Software Security: Building Security In. Addison-Wesley, 2006.
  33. M. M. Lehman, “Programs, Life Cycles, and Laws of Software Evolution,” Proceedings of the IEEE, vol. 68, no. 9, pp. 1060–1076, 1980, doi: 10.1109/PROC.1980.11805.
  34. D. L. Parnas, “Software Aging,” in Proceedings of the 16th International Conference on Software Engineering (ICSE), IEEE Computer Society Press, 1994, pp. 279–287. doi: 10.1109/ICSE.1994.296790.
  35. A. MacCormack, J. Rusnak, and C. Y. Baldwin, “Exploring the Structure of Complex Software Designs: An Empirical Study of Open Source and Proprietary Code,” Management Science, vol. 52, no. 7, pp. 1015–1030, 2006, doi: 10.1287/mnsc.1060.0552.
  36. G. Tassey, “The Economic Impacts of Inadequate Infrastructure for Software Testing,” National Institute of Standards and Technology (NIST), Planning Report 02-3, 2002.
  37. S. T. Knox, “Modeling the Cost of Software Quality,” 4, 1993.

This site uses Just the Docs, a documentation theme for Jekyll.