Empirical Effectiveness of Static Analysis

Static analysis effectiveness depends on false positive rates, cost relative to alternatives, and tool precision. This page synthesizes empirical evidence from industry studies to answer a practical question: how well does static analysis actually work?


The 10% False Positive Rule

The single most important metric for static analysis adoption is not accuracy – it is developer tolerance for false positives. Two large-scale industrial studies converge on the same threshold.

Google enforces a strict false positive policy for all Tricorder analyzers [1]:

Threshold Action
< 10% effective false positive rate Analyzer remains active
> 10% Analyzer placed on probation
> 25% Analyzer disabled immediately

Platform-wide, Tricorder achieves approximately 5% “not useful” rate. Critically, Google measures developer perception – the rate at which developers click “not useful” on findings – rather than theoretical accuracy metrics.

eBay reached the same threshold through a different path [2]:

Phase False Positive Rate Developer Response
Initial deployment (default config) 50% Declared tool “useless”
After customization (75% of checks disabled) 10% Acceptance and adoption

Key insight: The threshold for developer tolerance is remarkably consistent across organizations – around 10%. Above this, developers simply ignore all warnings.


Cost-Effectiveness Benchmarks

Static analysis is not a replacement for testing or inspection – it is a cheaper complement that catches different fault types.

Metric Value Source
Development vs. Testing cost 10x cheaper to find defects during development [2]
ASA vs. Inspection cost 60-72% of manual inspection cost [3]
ASA defect yield 23-37% [3]
Testing defect yield 63-98% [3]
Combined (ASA + Testing) Higher than either alone [3]

Interpretation: ASA does not replace testing – it complements it. ASA catches assignment and checking faults that testing misses. Testing catches functional and algorithmic faults that ASA misses. The combination yields better results than either technique alone.


Tool Precision in Practice

A comparative study of six Java static analysis tools [4] reveals wide variation in precision and coverage:

Tool Precision Coverage Notes
SonarQube 18% Broadest rule set Many rules, many false positives
Better Code Hub 29% Design quality SIG-based metrics
Coverity Scan 37% Deep analysis More precise, fewer rules
LGTM (CodeQL) Security focus Taint analysis strength
FindSecBugs Security focus Java-specific
CheckStyle 86% Style only High precision because issues are trivial

Critical finding: Less than 0.4% agreement between tools at the line level. Different tools detect fundamentally different defect types. No single tool is sufficient.

The implication is clear: organizations serious about static analysis must deploy multiple tools covering different defect categories, rather than relying on a single tool for comprehensive coverage.


Scale at Google

Google’s static analysis infrastructure demonstrates that SA can operate at extreme scale [1]:

Metric Value
Findings per day ~93,000 across ~31,000 changelists
Code reviews per day ~50,000 – SA findings integrated into review
“Not useful” rate ~5% platform-wide
Unique Linter users 18,000+ (separate from Tricorder)

Key architectural decisions that enable this scale:

  • Analyzers contributed by domain experts across the company, not a central SA team
  • Suggested fixes significantly increase the fix rate – developers are more likely to act when the tool provides a one-click resolution
  • Project-level configuration, not per-user – ensures team consistency
  • Code review integration – findings appear as review comments, not standalone reports

Symbolic Execution Effectiveness

Symbolic execution bridges static and dynamic analysis. KLEE results on the COREUTILS suite [5] demonstrate what automated analysis can achieve:

Metric Value
Programs tested 452 automatically
Line coverage 90%+ (median 94.7%)
Coverage improvement Beat 15 years of manually-maintained developer tests by 16.8%
Bugs found 56 serious bugs, 3 undetected for 15+ years
Notable discoveries Bugs in tools considered “thoroughly tested” (sort, cut, comm)

These results demonstrate that automated techniques can surpass manual test suite construction for coverage, even on mature, well-tested software.


Complementary Strengths of V&V Techniques

No single verification technique catches all fault types. The Nortel Networks study [3] quantified the complementary strengths:

Fault Type ASA Testing Inspection
Assignment faults Strong Weak Medium
Checking faults Strong Medium Medium
Functional faults Weak Strong Medium
Algorithmic faults Weak Strong Strong
Interface faults Medium Medium Strong
Design faults Weak Weak Strong

Best practice: Combine all three techniques. Each catches fault types the others miss. ASA excels at mechanical errors (wrong assignment, missing check), testing excels at behavioral errors (wrong output, wrong algorithm), and inspection excels at structural errors (bad interface, poor design).


Summary

Dimension Key Finding
False positive threshold 10% – consistent across Google and eBay
Cost advantage 10x cheaper than finding defects in testing
ASA vs. inspection cost 60-72% of manual inspection cost
Tool agreement < 0.4% at line level – no single tool suffices
Symbolic execution 90%+ coverage automatically, beats manual tests
Best strategy Combine ASA + testing + inspection

Further Exploration


References

  1. C. Sadowski, J. van Gogh, C. Jaspan, E. Söderberg, and C. Winter, “Tricorder: Building a Program Analysis Ecosystem,” in Proceedings of the 37th International Conference on Software Engineering (ICSE), 2015, pp. 598–608. doi: 10.1109/ICSE.2015.76.
  2. C. Jaspan, I.-C. Chen, and A. Sharma, “Understanding the Value of Program Analysis Tools,” in Companion to the 22nd ACM SIGPLAN Conference on Object-Oriented Programming Systems and Applications (OOPSLA), 2007, pp. 963–970. doi: 10.1145/1297846.1297964.
  3. J. Zheng, L. Williams, N. Nagappan, W. Snipes, J. P. Hudepohl, and M. A. Vouk, “On the Value of Static Analysis for Fault Detection in Software,” IEEE Transactions on Software Engineering, vol. 32, no. 4, pp. 240–253, 2006, doi: 10.1109/TSE.2006.38.
  4. V. Lenarduzzi, N. Saarimäki, and D. Taibi, “A Critical Comparison on Six Static Analysis Tools: Detection, Agreement, and Precision,” Journal of Systems and Software, vol. 198, p. 111575, 2023, doi: 10.1016/j.jss.2022.111575.
  5. C. Cadar, D. Dunbar, and D. Engler, “KLEE: Unassisted and Automatic Generation of High-Coverage Tests for Complex Systems Programs,” in Proceedings of the 8th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 2008, pp. 209–224.

Disclaimer: AI is used for text summarization, polishing and explaining. Authors have verified all facts and claims. In case of an error, feel free to file an issue.


This site uses Just the Docs, a documentation theme for Jekyll.