Empirical Evidence for Exploratory Testing
The effectiveness of exploratory testing has been studied through multiple controlled experiments and industrial case studies. The evidence converges on three robust findings: ET matches scripted testing in defect detection, is dramatically more efficient, and produces fewer false positives.
Reading the Evidence: A Quick Statistics Glossary
The experiments below use standard statistical measures. Here is what they mean:
| Term | Meaning | Rule of Thumb |
|---|---|---|
| p-value | The probability that the observed difference happened by chance. A small p-value means the result is unlikely to be a coincidence. | p < 0.05 = statistically significant; p < 0.001 = highly significant; “n.s.” = not significant |
| Effect size (d) | How large the difference is, independent of sample size. A significant p-value tells you the difference is real; effect size tells you if it matters. | d = 0.2 small, d = 0.5 medium, d = 0.8 large, d > 1.0 very large |
| n.s. | “Not significant” — the difference could easily be due to chance | Does not mean “no difference” — it means we cannot be confident one exists |
Example: If ET finds 7.04 defects and scripted finds 6.37 defects with p=0.088 (n.s.), it means: the raw numbers favor ET, but the difference is small enough that random variation could explain it. We cannot conclude ET is better — but we can conclude it is not worse.
Controlled Experiments
Itkonen & Mantyla (2007, 2014) — The Landmark Study
The most influential ET experiment and its replication [1], [2]:
| Metric | ET | TCT | p-value | Source |
|---|---|---|---|---|
| Defects found | 7.04 | 6.37 | 0.088 (n.s.) | [1] |
| Total effort | 1.5h | 8.5h (7h design + 1.5h exec) | — | [1] |
| False reports | 1.03 | 2.08 | 0.000 | [1] |
| Defects (replication) | 5.47 | 6.06 | 0.093 (n.s.) | [2] |
| Total effort (replication) | 4.58h | 19.47h | d=-1.33 (large) | [2] |
| False reports (replication) | 1.84 | 2.90 | 0.006 | [2] |
n=79 students (2007), n=51 students (2014). TCT = Test-Case-Based Testing.
ET was 4-6x more efficient than scripted testing when total effort (including test case design time) is accounted for, with a large effect size (d=1.33).
Afzal et al. (2015) — Equal-Time Experiment
Under a strict 90-minute equal-time constraint for all activities [3]:
| Metric | ET | TCT | p-value | Effect Size |
|---|---|---|---|---|
| Defects found | 8.34 | 1.83 | 1.16e-10 | d=2.065 (huge) |
n=70 (46 students + 24 practitioners). Largest effect size in any ET study.
When time is truly equal (including test case design), ET found 4.6x more defects. This is the strongest evidence that ET’s efficiency advantage translates directly into effectiveness under time pressure.
No significant difference between students and practitioners (p=0.07), challenging the assumption that ET requires extensive experience [3].
Shah et al. (2014) — Three-Way Comparison
Comparison of ET, Hybrid Testing (HT), and Scripted Testing (ST) with experienced practitioners [4]:
| Metric | ET | Hybrid | Scripted | ANOVA |
|---|---|---|---|---|
| Defects found | 13.16 | 10.67 | 7.0 | F=20.53, p<0.001 |
| Functionality coverage | 6.67 | 8.83 | 7.83 | — |
n=6 experienced testers (10+ years).
ET found the most defects but achieved the lowest systematic coverage. The hybrid approach (combining requirement-based test cases for coverage with test missions for exploratory freedom) provided the best balance [4].
Prakash & Gopalakrishnan (2011) — Efficiency Comparison
Both groups found similar bug counts, but [5]:
- ET found bugs earlier and across more diverse quality criteria
- Test case writing effort: 12 hours (scripted) vs. 2 hours (checklists for ET) — a 6x difference
Summary of Experimental Evidence
{
"$schema": "https://vega.github.io/schema/vega-lite/v5.json",
"title": "ET vs. Scripted Testing: Defect Detection Across Studies",
"width": 450,
"height": 250,
"data": {
"values": [
{"study": "Itkonen 2007", "method": "ET", "defects": 7.04},
{"study": "Itkonen 2007", "method": "Scripted", "defects": 6.37},
{"study": "Itkonen 2014", "method": "ET", "defects": 5.47},
{"study": "Itkonen 2014", "method": "Scripted", "defects": 6.06},
{"study": "Afzal 2015", "method": "ET", "defects": 8.34},
{"study": "Afzal 2015", "method": "Scripted", "defects": 1.83},
{"study": "Shah 2014", "method": "ET", "defects": 13.16},
{"study": "Shah 2014", "method": "Scripted", "defects": 7.0}
]
},
"mark": "bar",
"encoding": {
"x": {"field": "study", "type": "nominal", "title": "Study", "axis": {"labelAngle": 0}},
"y": {"field": "defects", "type": "quantitative", "title": "Mean Defects Found"},
"xOffset": {"field": "method"},
"color": {
"field": "method",
"type": "nominal",
"scale": {"domain": ["ET", "Scripted"], "range": ["#2e7d32", "#d32f2f"]},
"title": "Method"
}
}
}
Chart reconstructed from published means. See [1], [2], [3], [4] for original data.
| Finding | Consistency |
|---|---|
| ET >= scripted in defect detection | 4/5 experiments |
| ET dramatically more efficient (4-6x) | 5/5 experiments |
| Scripted produces more false positives | 2/2 experiments measuring this |
| ET achieves lower systematic coverage | 2/2 experiments measuring this |
The Efficiency Story
The most consistent finding across all studies is ET’s dramatic efficiency advantage when total effort (including test case design) is counted:
{
"$schema": "https://vega.github.io/schema/vega-lite/v5.json",
"title": "Total Effort: ET vs. Scripted Testing (hours)",
"width": 450,
"height": 250,
"data": {
"values": [
{"study": "Itkonen 2007", "method": "ET", "hours": 1.5},
{"study": "Itkonen 2007", "method": "Scripted (design + execution)", "hours": 8.5},
{"study": "Itkonen 2014", "method": "ET", "hours": 4.58},
{"study": "Itkonen 2014", "method": "Scripted (design + execution)", "hours": 19.47},
{"study": "Prakash 2011", "method": "ET", "hours": 2.0},
{"study": "Prakash 2011", "method": "Scripted (design + execution)", "hours": 12.0}
]
},
"mark": "bar",
"encoding": {
"y": {"field": "study", "type": "nominal", "title": "Study", "sort": null},
"x": {"field": "hours", "type": "quantitative", "title": "Total Hours"},
"yOffset": {"field": "method"},
"color": {
"field": "method",
"type": "nominal",
"scale": {"domain": ["ET", "Scripted (design + execution)"], "range": ["#2e7d32", "#d32f2f"]},
"title": "Method"
}
}
}
Chart reconstructed from published means. Scripted testing effort includes test case design time. See [1], [2], [5] for original data.
Industrial Evidence
Defect Detection Rates
Studies of ET in real organizations show high defect detection rates [6]:
| Context | ET Method | Defect Rate |
|---|---|---|
| Industrial sessions | Session-based ET | 4.8-8.7 defects/hour |
| Benchmarks | Usage-based testing | <3 defects/hour |
| Benchmarks | Functional testing | 2.47 defects/hour |
Defect Types
ET excels at finding certain defect types that scripted testing misses [1]:
| Defect Type | ET | Scripted | ET Advantage |
|---|---|---|---|
| Usability defects | 19 | 5 | 380% more |
| GUI defects | 70 | 49 | 143% more |
The Role of Tester Knowledge
ET effectiveness depends heavily on the tester’s knowledge and cognitive engagement [7]:
| Factor | Evidence |
|---|---|
| Effort does not equal effectiveness | ET effort does not correlate with defects found (r=0.025, p=0.86) [7] |
| Experience matters | More experienced practitioners significantly more likely to use ET (Fisher’s exact p=0.016) [8] |
| Domain knowledge critical | None of 7 industrial ET practitioners had formal testing training — effectiveness relied on domain knowledge [6] |
| Personality effects | Extroverts more effective at ET (75% vs. 51.5% control) [9] |
| But novices can learn | No significant difference between students and practitioners in one study (p=0.07) [3] |
In ET, it is the quality of the tester’s thinking, not the quantity of effort, that determines outcomes.
Teaching ET
“It is definitely easier to start learning testing when it is fully scripted. If we do freestyle then it would be difficult” — Focus group participant [10]
The evidence suggests ET can be taught through structured progression:
- Start novices with low exploration (detailed charters, steps provided)
- Gradually increase exploration as domain knowledge grows
- Teach heuristics and tours as a vocabulary for exploration
- Use session debriefings as a coaching mechanism
ET Complements Scripted Testing
The evidence consistently shows ET and scripted testing are complementary, not competing approaches:
flowchart LR
subgraph et["Exploratory Testing"]
direction TB
E1["Usability defects"]
E2["GUI defects"]
E3["Edge cases"]
E4["Unknown unknowns"]
end
subgraph st["Scripted Testing"]
direction TB
S1["Specification conformance"]
S2["Regression detection"]
S3["Systematic coverage"]
S4["Traceability to requirements"]
end
subgraph overlap["Both Find"]
direction TB
O1["Functional defects"]
O2["Logic errors"]
end
et --- overlap --- st
style et fill:#c8e6c9,stroke:#388e3c
style st fill:#bbdefb,stroke:#1976d2
style overlap fill:#fff9c4,stroke:#f9a825
| Dimension | ET Strength | Scripted Strength |
|---|---|---|
| Defect types | Usability, GUI, edge cases [1] | Conformance, specification-based |
| Efficiency | Higher (no design overhead) | — |
| Coverage | — | Higher systematic coverage [4] |
| Traceability | — | Full traceability to requirements |
| Adaptability | High (responds to discoveries) | Low (follows script) |
“ET is not a replacement for existing test-case based approaches, but a complementary testing approach suitable for certain situations” [6]
References
- J. Itkonen, M. V. Mäntylä, and C. Lassenius, “Defect Detection Efficiency: Test Case Based vs. Exploratory Testing,” in International Symposium on Empirical Software Engineering and Measurement (ESEM), 2007, pp. 61–70. doi: 10.1109/ESEM.2007.56.
- J. Itkonen and M. V. Mäntylä, “Are test cases needed? Replicated comparison between exploratory and test-case-based software testing,” Empirical Software Engineering, vol. 19, no. 2, pp. 303–342, 2014, doi: 10.1007/s10664-013-9266-8.
- W. Afzal, A. N. Ghazi, J. Itkonen, R. Torkar, A. Andrews, and K. Bhatti, “An experiment on the effectiveness and efficiency of exploratory testing,” Empirical Software Engineering, vol. 20, no. 3, pp. 844–878, 2015, doi: 10.1007/s10664-014-9301-4.
- S. M. A. Shah, U. Alvi, Ç. Gencel, and K. Petersen, “Comparing a Hybrid Testing Process with Scripted and Exploratory Testing: An Experimental Study with Practitioners,” in International Conference on Product-Focused Software Process Improvement (PROFES), 2014. doi: 10.1007/978-3-319-06862-6_13.
- V. Prakash and S. Gopalakrishnan, “Testing efficiency exploited: Scripted versus exploratory testing,” in International Conference on Emerging Trends in Electrical and Computer Technology (ICECTECH), 2011. doi: 10.1109/ICECTECH.2011.5941824.
- J. Itkonen and K. Rautiainen, “Exploratory testing: a multiple case study,” in International Symposium on Empirical Software Engineering (ISESE), 2005. doi: 10.1109/ISESE.2005.1541817.
- J. Itkonen, M. V. Mäntylä, and C. Lassenius, “The Role of the Tester’s Knowledge in Exploratory Software Testing,” IEEE Transactions on Software Engineering, vol. 39, no. 5, pp. 707–724, 2013, doi: 10.1109/tse.2012.55.
- D. Pfahl, H. Yin, M. V. Mäntylä, and J. Münch, “How is exploratory testing used? A state-of-the-practice survey,” in ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM), 2014, pp. 1–10. doi: 10.1145/2652524.2652531.
- L. Shoaib, A. Nadeem, and A. Akbar, “An empirical evaluation of the influence of human personality on exploratory software testing,” in IEEE International Multitopic Conference (INMIC), 2009, pp. 1–6. doi: 10.1109/inmic.2009.5383088.
- A. N. Ghazi, K. Petersen, E. Bjarnason, and P. Runeson, “Exploratory Testing: One Size Doesn’t Fit All,” 2017.
Disclaimer: AI is used for text summarization, polishing and explaining. Authors have verified all facts and claims. In case of an error, feel free to file an issue.