Article detail · 2008
Implications of ceiling effects in defect predictors
- Year
- 2008
- Type
- conference-paper
Abstract
OpenAlex · English
Context: There are many methods that input static code features and output a predictor for faulty code modules. These data mining methods have hit a "performance ceiling"; i.e., some inherent upper bound on the amount of information offered by, say, static code features when identifying modules which contain faults. Objective: We seek an explanation for this ceiling effect. Perhaps static code features have "limited information content"; i.e. their information can be quickly and completely discovered by even simple learners. Method:An initial literature review documents the ceiling effect in other work. Next, using three sub-sampling techniques (under-, over-, and micro-sampling), we look for the lower useful bound on the number of training instances. Results: Using micro-sampling, we find that as few as 50 instances yield as much information as larger training sets. Conclusions: We have found much evidence for the limited information hypothesis. Further progress in learning defect predictors may not come from better algorithms. Rather, we need to be improving the information content of the training data, perhaps with case-based reasoning methods.
Topics
Citations
OpenAlex cited_by_count. Not a WoS or Scopus citation count; those sources have no separate column here.
186 citations
OpenAlex cited_by_count (cache / database)
25 publications in the local catalog that cite this work (OpenAlex reference match; not the full global list).
- On the relative value of cross-company and within-company data for defect prediction 2009
- Empirical evaluation of the effects of mixed project data on learning defect predictors 2012
- Validation of network measures as indicators of defective modules in software systems 2009
- Practical considerations in deploying statistical methods for defect prediction: A case study within the Turkish telecommunications industry 2010
- Practical considerations in deploying AI for defect prediction 2009
- Deriving thresholds of software metrics to predict faults on open source software Replicated case studies 2016
- Influence of confirmation biases of developers on software quality: an empirical study 2012
- Empirical Evaluation of Mixed-Project Defect Prediction Models 2011
- Defect prediction using social network analysis on issue repositories 2011
- Merits of using repository metrics in defect prediction for open source projects 2009