Submitted:
27 October 2025
Posted:
31 October 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. Research Questions
1.2. Research Contributions
- Rigorous Data Leakage Prevention Methodology: The research team successfully eliminated 15 outcome-dependent features from TravisTorrent dataset which contained build results to show how temporal data leakage creates artificial accuracy increases from 85.24% to 95-99%. The research establishes a method to stop data leakage in temporal software engineering datasets which protects model performance by using only pre-build features available during prediction time.
- Real-World Validation at Scale: The research validates build prediction through analysis of 100,000 builds from 1,000+ open-source projects which include Java, Ruby, Python and JavaScript programming languages. The Random Forest model achieved 85.24% accuracy through analysis of 100,000 builds from 1,000+ open-source projects. The study provides evidence of practical build prediction applications through its evaluation of multiple project types and programming languages at a large scale.
- The research findings show that project maturity (9.49%) and repository age (9.46%) predict build outcomes better than immediate code changes which contradicts the conventional belief that code-level metrics control build results. The research shows that project history and development context generate more accurate signals than the changes made in individual commits.
- Cross-Language Generalization Evidence: The research shows that the model achieves similar results when analyzing different programming languages (Java: 84.2%, Ruby: 82.1%, Python: 81.5%, JavaScript: 80.8%) with a maximum difference of 4.9 percentage points. The results show that models trained on multiple programming languages can predict build outcomes for various software systems without needing language-specific model updates.
1.3. Paper Organization
2. Related Work
2.1. Machine Learning for Software Performance Prediction
2.1.1. Neural Networks and Deep Learning Architectures
2.1.2. Ensemble Methods and Gradient Boosting
2.2. SDLC Metrics and DevOps Analytics
2.2.1. Performance Code Metrics and Bug Prediction
2.2.2. DevOps Metrics and Organizational Performance
2.2.3. Build Optimization and CI/CD Pipeline Analytics
2.2.4. Commit Metrics and Developer Productivity
2.3. Data Leakage in Machine Learning for Software Engineering
2.4. Evaluation Metrics and Statistical Testing
2.5. Research Gaps and Novel Contributions
3. Methodology
3.1. TravisTorrent Dataset
3.1.1. Dataset Characteristics
- Project Context (12 features): Repository metadata including project age (days since first commit), total commits, total contributors, SLOC (Source Lines of Code), primary language, license type, and project maturity indicators.
- Build Context (8 features): Build environment characteristics including build number, build duration (prior builds), build configuration, branch information, pull request association, and build trigger type.
- Commit Metrics (14 features): Code change characteristics including commit hash, author, timestamp, files modified, lines added/deleted, commit message length, merge status, and parent commit count.
- Code Complexity (18 features): Static analysis metrics including cyclomatic complexity, Halstead metrics, nesting depth, code duplication ratios, technical debt indicators, and code coverage from previous builds.
- Test Structure (14 features): Test suite characteristics including test count, test density (tests per KLOC), test classes, assertions count, test execution patterns from historical builds, and test coverage evolution.
3.2. Temporal Data Leakage Prevention
3.2.1. Leakage Taxonomy and Identification
3.2.2. Feature Filtering Methodology
- Project Maturity (8 features): gh_project_age_days, gh_commits_count, gh_contributors_ count, gh_total_stars, gh_project_maturity_days, gh_sloc, gh_test_density, git_ repository_age_days
- Code Complexity (9 features): gh_src_complexity_avg, gh_src_complexity_max, gh_nesting_ depth_avg, gh_code_duplication_ratio, gh_technical_debt_index, gh_halstead_ difficulty_avg, gh_maintainability_index, gh_test_complexity_avg, gh_assertion_ density
- Test Structure (6 features): gh_tests_count, gh_test_classes_count, gh_test_assertions_ count, gh_test_lines_ratio, gh_test_coverage_previous, gh_test_growth_rate
- Build History (5 features): tr_build_number, tr_prev_build_duration, tr_prev_build_ success, tr_builds_last_30_days, tr_failure_streak
- Commit Context (3 features): git_num_files_modified, git_lines_added, git_lines_deleted
3.3. Feature Preprocessing and Normalization
3.3.1. Missing Value Handling
3.3.2. Feature Scaling
3.3.3. Categorical Feature Encoding
3.4. Classification Model Development
3.4.1. Logistic Regression
3.4.2. Random Forest Classifier
3.4.3. Gradient Boosting Classifier
3.5. Experimental Design and Validation Strategy
3.5.1. Dataset Sampling and Train/Test Split
3.5.2. Evaluation Metrics for Binary Classification
3.5.3. Cross-Validation Strategy
3.5.4. Reproducibility and Implementation Details
4. Results and Analysis
4.1. Dataset Summary and Preprocessing Statistics
4.2. Model Performance Comparison (RQ1 Answer)
- Random Forest Achieves Best Performance (RQ1 Answer): Random Forest attains 85.24% accuracy, 91.38% ROC-AUC, and 89.80% F1-score, substantially outperforming logistic regression (61.55% accuracy) and exceeding majority-class baseline (69.50%) by 15.7 percentage points. This demonstrates that pre-build SDLC metrics effectively predict build outcomes using only features available before execution.
- Ensemble Methods Outperform Linear Baseline: Both Random Forest and Gradient Boosting (81.34% accuracy) dramatically exceed Logistic Regression by >20 percentage points, indicating non-linear decision boundaries better capture build prediction patterns. Class-weighted ensembles handle moderate class imbalance (70%/30%) effectively.
- High Recall Validates Practical Utility: Random Forest achieves 93.57% recall, correctly identifying 93.6% of successful builds while detecting 66.2% of failures (computed from confusion matrix). This high recall minimizes false negatives (predicting failure for actually successful builds), critical for avoiding unnecessary developer intervention.
- ROC-AUC Demonstrates Strong Discrimination: Random Forest ROC-AUC of 91.38% substantially exceeds random guessing (50%), confirming model discriminates well between successful and failed builds across classification thresholds (Figure 3). Gradient Boosting achieves 88.59% ROC-AUC, also strong performance.
- Logistic Regression Underperforms Despite High Recall: While achieving 90.20% recall, Logistic Regression suffers from low precision (62.64%), producing excessive false positives. This suggests build prediction requires non-linear modeling to capture complex feature interactions.
4.3. Confusion Matrix Analysis
- High True Positive Rate: The model correctly identifies 13,004 of 13,900 successful builds (93.6% recall), minimizing false negatives. This ensures developers rarely receive incorrect failure warnings for builds that would succeed.
- Moderate Failure Detection: The model correctly detects 4,043 of 6,100 failed builds (66.2% true negative rate). While not perfect, this provides substantial value by identifying two-thirds of failures before execution.
- False Positive Analysis: 2,057 builds predicted to fail actually succeeded (13.7% of predicted failures). This false positive rate represents acceptable cost–approximately 10% of test set receives unnecessary attention, but 66% of actual failures are caught proactively.
- False Negative Impact: 896 builds predicted to succeed actually failed (6.5% of predicted successes). These missed failures proceed to execution, but represent minority of successful predictions (5.9%). The high precision (80.88%) ensures most success predictions are reliable.
- Class-Weighted Training Effectiveness: Despite 70:30 class imbalance, the model achieves balanced performance through class weighting, avoiding degenerate solution of predicting all builds successful. The confusion matrix demonstrates meaningful discrimination across both classes.
4.4. Feature Importance Analysis (RQ2 Answer)
- Project Context Dominates Build Prediction: The top 6 features are all project-level characteristics (project maturity, repository age, commit count, contributors), collectively representing 49.8% cumulative importance (Figure 4). This demonstrates that project history and maturity predict build outcomes more reliably than immediate code changes, challenging conventional focus on code-level metrics.
- Project Maturity as Strongest Predictor: gh_project_maturity_days (9.49%) and git_ repository_age_days (9.46%) nearly tie as top predictors. Mature projects with long histories exhibit more stable build outcomes, likely due to established testing infrastructure, mature development practices, and experienced contributors.
- Code Metrics Secondary to Project Context: Code complexity (SLOC: 7.66%) ranks 5th, substantially lower than project maturity metrics. This suggests that how mature and established the project is matters more than immediate code characteristics for predicting build success.
- Test Structure Provides Moderate Signal: Test density (5.87%) and test count (5.43%) rank 8th and 9th, contributing meaningful but secondary predictive power. Well-tested projects fail less frequently, but test structure alone insufficient without project maturity context.
- Build History Contributes: Build number (6.12%) and recent build volume (4.98%) capture build pattern stability. Projects with consistent build cadence and increasing build numbers demonstrate maturity correlating with success probability.
- Phase Distribution: Project Context (49.8%), Test Structure (11.3%), Build History (10.1%), Code Metrics (7.7%). This distribution validates that multi-phase integration essential, with project-level context providing strongest signal, followed by test infrastructure and build patterns.
4.5. Cross-Language Generalization (RQ3 Answer)
- Strong Cross-Language Consistency (RQ3 Answer): Accuracy varies by only 3.38 percentage points across languages (Java: 84.21%, JavaScript: 80.83%), with standard deviation 1.38% (Figure 5). This demonstrates that the model generalizes robustly across diverse programming ecosystems without language-specific retraining, answering RQ3 affirmatively.
- Java Projects Exhibit Highest Predictability: Java achieves best accuracy (84.21%), likely due to mature tooling ecosystems (Maven, Gradle), strong testing conventions (JUnit), and enterprise development practices. Java’s static typing and compilation phase catch more errors pre-build.
- Consistent High Recall Across Languages: Recall ranges narrowly 92.98-94.12% (std dev: 0.49%), indicating model reliably identifies successful builds regardless of language. This consistency validates that project maturity metrics (top predictors) transcend language-specific characteristics.
- Precision Variation Reflects Language Ecosystems: Precision varies more (79.23-82.45%, std dev: 1.32%) than recall, suggesting false positive rates differ by language. Dynamically-typed languages (Ruby, Python, JavaScript) exhibit slightly lower precision, potentially due to runtime errors undetectable in pre-build analysis.
- Language Not in Top Features: Programming language one-hot encoding does not appear in top 10 features (Table 3), confirming that project maturity, test structure, and build history provide language-agnostic predictive signals. This validates single multi-language model viability rather than per-language specialization.
- Practical Implication: Organizations can deploy a single build prediction model across polyglot codebases without language-specific tuning, simplifying CI/CD pipeline integration and reducing operational overhead.
4.6. Data Leakage Impact Analysis (RQ4 Answer)
- Severe Performance Inflation from Leaky Features (RQ4 Answer): Models trained with all 66 features achieve unrealistically high performance: 97.80% accuracy, 99.56% ROC-AUC, and 98.42% F1-score. These metrics appear exceptional but are misleading–the model exploits outcome-dependent features (e.g., tr_status, tr_log_tests_failed, tr_duration) that encode build results, creating artificially perfect predictions during training but complete failure in production where such features are unavailable before build execution.
- Realistic Performance with Clean Features: Removing 15 leaky features reduces accuracy from 97.80% to 85.24%–a 12.6 percentage point drop. This realistic performance reflects genuine predictive capability using only pre-build information: project maturity, code complexity, test structure, build history, and commit metadata. While lower than leaky-feature accuracy, 85.24% substantially exceeds majority-class baseline (69.50%) by 15.7 percentage points and aligns with realistic CI/CD prediction ranges (75-84%) reported in rigorous prior studies [1].
- Production Viability Requires Leakage Prevention: The 15-point accuracy drop quantifies the critical importance of rigorous feature selection for production deployment. Prior research reporting 95-99% accuracies [5] likely reflects temporal data leakage rather than genuine predictive power. Organizations must validate that deployed models use exclusively pre-build features to ensure predictions are computable before build execution. Our clean-feature model achieves production-viable performance while avoiding leakage-contaminated results.
- ROC-AUC Drop Reveals Discrimination Capability Loss: ROC-AUC decreases from 99.56% (leaky) to 91.38% (clean), an 8.18 percentage point reduction. However, 91.38% still demonstrates strong discrimination capability, substantially exceeding random guessing (50%) and indicating strong class separation using only legitimate pre-build features. The leaky-feature AUC of 99.56% approaches theoretical maximum, a red flag signaling data leakage in any machine learning study.
- Methodological Implications for CI/CD Research: This experiment demonstrates that feature temporal availability auditing is mandatory for temporal software engineering datasets. Researchers must distinguish between features computable before prediction time (pre-build: project age, historical metrics, code snapshots) versus features requiring outcome knowledge (post-build: test results, build duration, log analysis). Cross-validation alone insufficient–explicit temporal validation of each feature is essential to prevent inflated performance claims.
- Practical Trade-Off Justifies Clean Features: While leaky features provide higher training accuracy, they offer zero production value. The clean-feature model’s 85.24% accuracy enables real deployment: correctly identifying 93.6% of successful builds and 66.2% of failures before execution, providing actionable predictions for resource allocation, build prioritization, and developer feedback. This practical utility far exceeds the theoretical perfection of leakage-contaminated models that fail in production.
5. Discussion
5.1. Key Findings and Implications
5.2. Comparison with State-of-the-Art
5.3. Practical Deployment Considerations
5.4. Limitations and Threats to Validity
5.5. Future Research Directions
6. Conclusion
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| CI/CD | Continuous Integration/Continuous Deployment |
| SDLC | Software Development Lifecycle |
| ROC-AUC | Receiver Operating Characteristic - Area Under Curve |
| ML | Machine Learning |
| RF | Random Forest |
| GB | Gradient Boosting |
| LR | Logistic Regression |
References
- Vassallo, C.; Schermann, G.; Zampetti, F.; Romano, D.; Leitner, P.; Di Penta, M.; Panichella, S. A Tale of CI Build Failures: An Open Source and a Financial Organization Perspective. In Proceedings of the Proceedings of the 33rd IEEE International Conference on Software Maintenance and Evolution (ICSME), Shanghai, China, 2017; pp. 183–193. [CrossRef]
- Rausch, T.; Hummer, W.; Leitner, P.; Schulte, S. An Empirical Analysis of Build Failures in the Continuous Integration Workflows of Java-Based Open-Source Software. In Proceedings of the Proceedings of the 14th International Conference on Mining Software Repositories (MSR), Buenos Aires, Argentina, 2017; pp. 345–355. [CrossRef]
- Hilton, M.; Tunnell, T.; Huang, K.; Marinov, D.; Dig, D. Usage, Costs, and Benefits of Continuous Integration in Open-Source Projects. IEEE Transactions on Software Engineering 2016, 43, 426–445. [Google Scholar] [CrossRef]
- Beller, M.; Gousios, G.; Zaidman, A. Integration. In Proceedings of the Proceedings of the 14th International Conference on Mining Software Repositories (MSR), Buenos Aires, Argentina, 2017; pp. 447–450. [CrossRef]
- Saidani, I.; Ouni, A.; Mkaouer, M.W. Improving the Prediction of Continuous Integration Build Failures Using Deep Learning. Automated Software Engineering 2022, 29, 1–41. [Google Scholar] [CrossRef]
- Ghotra, B.; McIntosh, S.; Hassan, A.E. Revisiting the Impact of Classification Techniques on the Performance of Defect Prediction Models. In Proceedings of the Proceedings of the 37th IEEE International Conference on Software Engineering (ICSE), Florence, Italy, 2015; pp. 789–800. [CrossRef]
- Li, Z.; Jing, X.Y.; Zhu, X. Progress on Approaches to Software Defect Prediction. IET Software 2018, 12, 161–175. [Google Scholar] [CrossRef]
- khleel, n.a.a.; nehéz, k. Software Defect Prediction Using a Bidirectional LSTM Network Combined with Oversampling Techniques. Cluster Computing 2023, 27, 3615–3638. [Google Scholar] [CrossRef]
- Chung, J.; Gulcehre, C.; Cho, K.; Bengio, Y. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. arXiv preprint arXiv:1412.3555, 2014; arXiv:1412.3555 2014.. NIPS 2014 Deep Learning Workshop. [Google Scholar]
- Kamei, Y.; Shihab, E.; Adams, B.; Hassan, A.E.; Mockus, A.; Sinha, A.; Ubayashi, N. A Large-Scale Empirical Study of Just-in-Time Quality Assurance. IEEE Transactions on Software Engineering 2013, 39, 757–773. [Google Scholar] [CrossRef]
- Zhao, Y.; Damevski, K.; Chen, H. A Systematic Survey of Just-in-Time Software Defect Prediction. ACM Computing Surveys 2023, 55, 1–35. [Google Scholar] [CrossRef]
- stradowski, s.; madeyski, l. Industrial Applications of Software Defect Prediction Using Machine Learning: A Business-Driven Systematic Literature Review. Information and Software Technology 2023, 159, 107192. [Google Scholar] [CrossRef]
- Ni, C.; Liu, W.S.; Chen, X.; Gu, Q.; Chen, D.X.; Huang, Q.G. A Cluster Based Feature Selection Method for Cross-Project Software Defect Prediction. Journal of Computer Science and Technology 2017, 32, 1090–1107. [Google Scholar] [CrossRef]
- Harman, M.; Mansouri, S.A.; Zhang, Y. Search-based software engineering. ACM Computing Surveys 2012, 45, 11–1. [Google Scholar] [CrossRef]
- Zhang, D.; Han, S.; Dang, Y.; Lou, J.G.; Zhang, H.; Xie, T. Software Analytics in Practice. In Proceedings of the IEEE Software; 2013; Vol. 30, pp. 30–37. [Google Scholar] [CrossRef]
- wang, s.; huang, l.; gao, a.; ge, j.; zhang, t.; feng, h.; satyarth, i.; li, m.; zhang, h.; ng, v. Machine/Deep Learning for Software Engineering: A Systematic Literature Review. IEEE Transactions on Software Engineering 2023, 49, 1630–1652. [Google Scholar] [CrossRef]
- ortu, m.; destefanis, g.; hall, t.; bowes, d. Fault-insertion and fault-fixing behavioural patterns in Apache Software Foundation Projects. Information and Software Technology 2023, 158, 107187. [Google Scholar] [CrossRef]
- Zhang, J.M.; Harman, M.; Ma, L.; Liu, Y. Machine Learning Testing: Survey, Landscapes and Horizons. IEEE Transactions on Software Engineering 2022, 48, 1–36. [Google Scholar] [CrossRef]
- Grattan, N.; da Costa, D.A.; Stanger, N. The Need for More Informative Defect Prediction: A Systematic Literature Review. Information and Software Technology 2024, 171, 107456. [Google Scholar] [CrossRef]
- Liu, Y.; Fekete, A.; Gorton, I. Design-Level Performance Prediction of Component-Based Applications. IEEE Transactions on Software Engineering 2005, 31, 928–941. [Google Scholar] [CrossRef]
- Gong, J.; Chen, T. Deep Configuration Performance Learning: A Systematic Survey and Taxonomy. ACM Transactions on Software Engineering and Methodology 2024, 34, 25–1. [Google Scholar] [CrossRef]
- Ali, M.; Mazhar, T.; Al-Rasheed, A.; Shahzad, T.; Ghadi, Y.Y.; Khan, M.A. Enhancing Software Defect Prediction: A Framework with Improved Feature Selection and Ensemble Machine Learning. PeerJ Computer Science 2024, 10, e1860. [Google Scholar] [CrossRef]
- khleel, n.a.a.; nehéz, k.; fadulalla, m.; hisaen, a. Ensemble-Based Machine Learning Algorithms Combined with Near Miss Method for Software Bug Prediction. International Journal of Networked and Distributed Computing 2025, 13, 1–18. [Google Scholar] [CrossRef]
- Sagi, O.; Rokach, L. Ensemble Learning: A Survey. WIREs Data Mining and Knowledge Discovery 2018, 8, e1249. [Google Scholar] [CrossRef]
- He, H.; Garcia, E.A. Learning from Imbalanced Data. IEEE Transactions on Knowledge and Data Engineering 2009, 21, 1263–1284. [Google Scholar] [CrossRef]
- Ibraigheeth, M.A.; Abu Eid, A.I.; Alsariera, Y.A.; Awwad, W.F.; Nawaz, M. A New Weighted Ensemble Model to Improve the Performance of Software Project Failure Prediction. International Journal of Advanced Computer Science and Applications 2024, 15, 331–340. [Google Scholar] [CrossRef]
- Kaushik, A.; Sheoran, K.; Kapur, R.; Singh, D.P. SENSE: Software Effort Estimation Using Novel Stacking Ensemble Learning. Innovations in Systems and Software Engineering 2024, 21, 769–785. [Google Scholar] [CrossRef]
- Saraireh, J.; Agoyi, M.; Kassaymeh, S. Adaptive Ensemble Learning Model-Based Binary White Shark Optimizer for Software Defect Classification. International Journal of Computational Intelligence Systems 2025, 18, 14. [Google Scholar] [CrossRef]
- Granitto, P.M.; Furlanello, C.; Biasioli, F.; Gasperi, F. Recursive Feature Elimination with Random Forest for PTR-MS Analysis of Agroindustrial Products. Chemometrics and Intelligent Laboratory Systems 2006, 83, 83–90. [Google Scholar] [CrossRef]
- Cai, J.; Luo, J.; Wang, S.; Yang, S. Feature Selection in Machine Learning: A New Perspective. Neurocomputing 2018, 300, 70–79. [Google Scholar] [CrossRef]
- hartanto, a.d.; kholik, y.n.; pristyanto, y. Stock Price Time Series Data Forecasting Using the Light Gradient Boosting Machine (LightGBM) Model. Journal of Information Visualization 2023, 7, 456–470. [Google Scholar] [CrossRef]
- He, X.; Zhao, K.; Chu, X. AutoML: A Survey of the State-of-the-Art. Knowledge-Based Systems 2021, 212, 106622. [Google Scholar] [CrossRef]
- Lim, B.; Arik, S.O.; Loeff, N.; Pfister, T. Temporal Fusion Transformers for Interpretable Multi-Horizon Time Series Forecasting. International Journal of Forecasting 2021, 37, 1748–1764. [Google Scholar] [CrossRef]
- lee, s.; hong, j.; liu, l.; choi, w. TS-Fastformer: Fast Transformer for Time-Series Forecasting. ACM Transactions on Intelligent Systems and Technology 2024, 15, 45. [Google Scholar] [CrossRef]
- chen, y.; hao, j.; peng, y.; xia, h. Transformer-based performance prediction and proactive resource allocation for cloud-native microservices. Cluster Computing 2025, 28, 1–18. [Google Scholar] [CrossRef]
- tao, h.; fu, l.; cao, q.; niu, x.; chen, h.; shang, s.; xian, y. Cross-Project Defect Prediction Using Transfer Learning with Long Short-Term Memory Networks. IET Software 2024, 18, 234–248. [Google Scholar] [CrossRef]
- Zhao, G.; Georgiou, S.; Zou, Y.; Hassan, S.; Truong, D.; Corbin, T. Enhancing Performance Bug Prediction Using Performance Code Metrics. In Proceedings of the Proceedings of the 21st International Conference on Mining Software Repositories (MSR), Lisbon, Portugal, 2024. [CrossRef]
- tao, h.; fu, l.; cao, q.; niu, x.; chen, h.; shang, s.; xian, y. Cross-Project Defect Prediction Using Transfer Learning with Long Short-Term Memory Networks. IET Software 2024, 18, 456–470. [Google Scholar] [CrossRef]
- lawson, a. 2024 State of Tech Talent Report: Survey-based Insights into the Current State of Technical Talent Acquisition, Retention, and Management Globally. Technical report, Google Cloud and DevOps Research and Assessment, 2024. Accessed: 2025-01-14. [CrossRef]
- perera, p.; silva, r.; perera, i. Improve software quality through practicing DevOps 2017. [CrossRef]
- Weeraddana, N.; Alfadel, M.; McIntosh, S. Characterizing Timeout Builds in Continuous Integration. IEEE Transactions on Software Engineering 2024, 50, 2045–2062. [Google Scholar] [CrossRef]
- Zampetti, F.; Vassallo, C.; Panichella, S.; Canfora, G.; Gall, H.; Di Penta, M. An Empirical Characterization of Bad Practices in Continuous Integration. Empirical Software Engineering 2020, 25, 1095–1135. [Google Scholar] [CrossRef]
- Gallaba, K.; Macho, C.; Pinzger, M.; McIntosh, S. Noise and Heterogeneity in Historical Build Data: An Empirical Study of Travis CI. In Proceedings of the Proceedings of the 33rd IEEE/ACM International Conference on Automated Software Engineering (ASE), Montpellier, France, 2018; pp. 87–97. [CrossRef]
- sun, g.; habchi, s.; mcintosh, s. RavenBuild: Context, Relevance, and Dependency Aware Build Outcome Prediction. Proceedings of the ACM on Software Engineering 2024, 1, 996–1018. [Google Scholar] [CrossRef]
- abdellatif, a.; badran, k.; shihab, e. MSRBot: Using bots to answer questions from software repositories. Empirical Software Engineering 2020, 25, 2157–2201. [Google Scholar] [CrossRef]
- Luo, Q.; Moran, K.; Zhang, L.; Poshyvanyk, D. How Do Static and Dynamic Test Case Prioritization Techniques Perform on Modern Software Systems? An Extensive Study on GitHub Projects. IEEE Transactions on Software Engineering 2019, 46, 1054–1080. [Google Scholar] [CrossRef]
- Fallahzadeh, E.; Rigby, P.C.; Adams, B. Contrasting Test Selection, Prioritization, and Batch Testing at Scale. Empirical Software Engineering 2024, 30. [Google Scholar] [CrossRef]
- Gruber, M.; Roslan, M.F.; Parry, O.; Scharnböck, F.; McMinn, P.; Fraser, G. Do Automatic Test Generation Tools Generate Flaky Tests? In Proceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering (ICSE ’24). ACM, April 2024; pp. 14–20. [Google Scholar] [CrossRef]
- Silva, D.; Gruber, M.; Gokhale, S.; Arteca, E.; Turcotte, A.; d’Amorim, M.; Lam, W.; Winter, S.; Bell, J. The Effects of Computational Resources on Flaky Tests. IEEE Transactions on Software Engineering 2024, 50. [Google Scholar] [CrossRef]
- Bernardo, J.H.; da Costa, D.A.; de Medeiros, S.Q.; Kulesza, U. How do Machine Learning Projects use Continuous Integration Practices? In An Empirical Study on GitHub Actions. In Proceedings of the 21st IEEE/ACM International Conference on Mining Software Repositories (MSR), Lisbon, Portugal, April 2024; pp. 665–676. [Google Scholar] [CrossRef]
- Bouzenia, I.; Pradel, M. Resource Usage and Optimization Opportunities in Workflows of GitHub Actions. In Proceedings of the Proceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE), Lisbon, Portugal, April 2024; pp. 25:1–25:12. [CrossRef]
- Zhang, Y.; Wu, Y.; Chen, T.; Wang, T.; Liu, H.; Wang, H. How do Developers Talk about GitHub Actions? Evidence from Online Software Development Community. In Proceedings of the Proceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE), New York, NY, USA, 2024; pp. 1–13. ICSE ’24. [CrossRef]
- Kinsman, T.; Wessel, M.; Gerosa, M.A.; Treude, C. How Do Software Developers Use GitHub Actions to Automate Their Workflows? In Proceedings of the Proceedings of the 18th International Conference on Mining Software Repositories (MSR), Madrid, Spain; 2021; pp. 420–431. [Google Scholar] [CrossRef]
- zerouali, a.; mens, t.; robles, g.; gonzalez barahona, j.m. On the Diversity of Software Package Popularity Metrics: An Empirical Study of npm. IEEE Transactions on Software Engineering 2019, 49, 2113–2135. [Google Scholar] [CrossRef]
- Mastropaolo, A.; Zampetti, F.; Bavota, G.; Di Penta, M. Toward Automatically Completing GitHub Workflows. In Proceedings of the Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (ICSE), New York, NY, USA, 2024. ICSE ’24.
- saito, s. Understanding Key Business Processes for Business Process Outsourcing Transition. In Proceedings of the Proceedings of the IEEE 14th International Conference on Global Software Engineering (ICGSE), 2019, pp. pp. 66–75. [CrossRef]
- widder, d.g.; hilton, m.; kästner, c.; vasilescu, b. A Conceptual Replication of Continuous Integration Pain Points in the Context of Travis CI. In Proceedings of the Proceedings of the 27th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE), 2019, pp. pp. 647–658. [CrossRef]
- Morris, K. Infrastructure as Code: Managing Servers in the Cloud, 2nd ed.; O’Reilly Media, 2020.
- Humble, J.; Farley, D. Continuous Delivery: Reliable Software Releases Through Build, Test, and Deployment Automation; Addison-Wesley Professional, 2010.
- fontana, r.m.; meyer, v.; reinehr, s.; malucelli, a. Management Ambidexterity: A Clue for Maturing in Agile Software Development. In Proceedings of the Proceedings of the Agile Processes in Software Engineering and Extreme Programming Conference (XP)., Springer, 2015; pp. 212–217. [Google Scholar] [CrossRef]
- Adams, B.; McIntosh, S. Modern Release Engineering in a Nutshell – Why Researchers Should Care. In Proceedings of the Proceedings of the IEEE 23rd International Conference on Software Analysis, Evolution, and Reengineering (SANER), 2016; Vol. 5, pp. 78–90. [CrossRef]
- oliveira, e.; fernandes, e.; steinmacher, i.; cristo, m.; conte, t.; garcia, a. Code and commit metrics of developer productivity: a study on team leaders perceptions. Empirical Software Engineering 2020, 25, 3874–3912. [Google Scholar] [CrossRef]
- Hassan, A.E. Predicting Faults Using the Complexity of Code Changes. In Proceedings of the Proceedings of the 31st IEEE International Conference on Software Engineering (ICSE), Vancouver, Canada, 2009; pp. 78–88. [CrossRef]
- mccabe, t. A Complexity Measure. IEEE Transactions on Software Engineering 1976, SE-2, 308–320. [Google Scholar] [CrossRef]
- Halstead, M.H. Elements of Software Science; Elsevier, 1977. Classic software metrics foundation.
- Nagappan, N.; Ball, T.; Zeller, A. Mining Metrics to Predict Component Failures. In Proceedings of the Proceedings of the 28th International Conference on Software Engineering (ICSE), 2006, pp.; pp. 452–461. [CrossRef]
- Fenton, N.E.; Bieman, J. Software Metrics: A Rigorous and Practical Approach, 3rd ed.; CRC Press, 2014.
- Li, J.; Ahmed, I. Commit Message Matters: Investigating Impact and Evolution of Commit Message Quality. In Proceedings of the Proceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE), Melbourne, Australia, 2023; pp. 806–817. [CrossRef]
- Avgeriou, P.; Kruchten, P.; Ozkaya, I.; Seaman, C. Managing Technical Debt in Software Engineering. Dagstuhl Reports 2016, 6, 110–138. [Google Scholar] [CrossRef]
- Azeem, M.I.; Palomba, F.; Shi, L.; Wang, Q. Machine Learning Techniques for Code Smell Detection: A Systematic Literature Review and Meta-Analysis. Information and Software Technology 2019, 108, 115–138. [Google Scholar] [CrossRef]
- Li, Z.; Avgeriou, P.; Liang, P. A Systematic Mapping Study on Technical Debt and Its Management. Journal of Systems and Software 2015, 101, 193–220. [Google Scholar] [CrossRef]
- Bacchelli, A.; Bird, C. Review. In Proceedings of the Proceedings of the 2013 International Conference on Software Engineering, 2013, pp. pp. 712–721. [CrossRef]
- Wang, S.; Lo, D. Version History, Similar Report, and Structure: Putting Them Together for Improved Bug Localization. In Proceedings of the Proceedings of the 22nd International Conference on Program Comprehension (ICPC), 2014, pp. pp. 53–63. [CrossRef]
- Forsgren, N.; Storey, M.A.; Maddila, C.; Zimmermann, T.; Houck, B.; Butler, J. The SPACE of developer productivity. Communications of the ACM 2021, 64, 67–75. [Google Scholar] [CrossRef]
- Kaufman, S.; Rosset, S.; Perlich, C.; Stitelman, O. Leakage in data mining. ACM Transactions on Knowledge Discovery from Data 2012, 6, 15. [Google Scholar] [CrossRef]
- Cerqueira, V.; Torgo, L.; Mozetič, I. Evaluating Time Series Forecasting Models: An Empirical Study on Performance Estimation Methods. Machine Learning 2020, 109, 1997–2028. [Google Scholar] [CrossRef]
- bergmeir, c.; benítez, j.m. On the Use of Cross-Validation for Time Series Predictor Evaluation. Information Sciences 2012, 191, 192–213. [Google Scholar] [CrossRef]
- Emmanuel, T.; Maupong, T.; Mpoeleng, D.; Semong, T.; Mphago, B.; Tabona, O. A Survey on Missing Data in Machine Learning. Journal of Big Data 2021, 8, 140. [Google Scholar] [CrossRef] [PubMed]
- Patro, S.G.K.; Sahu, K.K. Normalization: A Preprocessing Stage. arXiv preprint arXiv:1503.06462, 2015; arXiv:1503.06462 2015.IARJSET International Conference. [Google Scholar] [CrossRef]
- kitchenham, b.; pfleeger, s.; pickard, l.; jones, p.; hoaglin, d.; emam, k.e.; rosenberg, j. Preliminary Guidelines for Empirical Research in Software Engineering. IEEE Transactions on Software Engineering 2002, 28, 721–734. [Google Scholar] [CrossRef]
- Baker, M. 1,500 Scientists Lift the Lid on Reproducibility. Nature 2016, 533, 452–454. [Google Scholar] [CrossRef]
- rainio, o.; teuho, j.; klén, r. Evaluation Metrics and Statistical Tests for Machine Learning. Scientific Reports 2024, 14, 6756. [Google Scholar] [CrossRef] [PubMed]
- McHugh, M.L. Interrater Reliability: The Kappa Statistic. Biochemia Medica 2012, 22, 276–282. [Google Scholar] [CrossRef]
- Hyndman, R.J.; Koehler, A.B. Another Look at Measures of Forecast Accuracy. International Journal of Forecasting 2006, 22, 679–688. [Google Scholar] [CrossRef]
- rey, d.; neuhäuser, m. Wilcoxon-Signed-Rank Test. International Encyclopedia of Statistical Science, 1658; 1658–1659. [Google Scholar] [CrossRef]
- rajput, d.; wang, w.j.; chen, c.c. Evaluation of a Decided Sample Size in Machine Learning Applications. BMC Bioinformatics 2023, 24, 48. [Google Scholar] [CrossRef]
- Sullivan, G.M.; Feinn, R. Using Effect Size—or Why the P Value Is Not Enough. Journal of Graduate Medical Education 2012, 4, 279–282. [Google Scholar] [CrossRef]
- Saito, T.; Rehmsmeier, M. The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLOS ONE 2015, 10, e0118432. [Google Scholar] [CrossRef]
- Hand, D.J.; Till, R.J. A Simple Generalisation of the Area Under the ROC Curve for Multiple Class Classification Problems. Machine Learning 2001, 45, 171–186. [Google Scholar] [CrossRef]
- Arlot, S.; Celisse, A. A Survey of Cross-Validation Procedures for Model Selection. Statistics Surveys 2010, 4, 40–79. [Google Scholar] [CrossRef]
- Wu, H.; Xu, J.; Wang, J.; Long, M. Forecasting. In Proceedings of the Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2021; Vol. 34, pp. 22419–22430.
- Hassan, A.E.; Xie, T. Software intelligence. In Proceedings of the Proceedings of the FSE/SDP Workshop on Future of Software Engineering Research, 2010; pp. 161–166. [CrossRef]
- Gousios, G.; Pinzger, M.; Deursen, A.v. An Exploratory Study of the Pull-Based Software Development Model. In Proceedings of the Proceedings of the 36th International Conference on Software Engineering, 2014; pp. 345–355. [CrossRef]
- Feurer, M.; Hutter, F. Hyperparameter Optimization. Automated Machine Learning, 2019; 3–33. [Google Scholar] [CrossRef]
- santana, f.a.; cordeiro, a.f.r.; oliveirajr, e. Dublin Core for Recording Metadata of Experiments in Software Engineering: A Survey. arXiv preprint arXiv:2303.16989, 2023; arXiv:2303.16989 2023.Under review. [Google Scholar] [CrossRef]
- Lehnert, S. A Review of Software Change Impact Analysis. Technical Report, Ilmenau University of Technology, TU Ilmenau: Germany, 2011. [Google Scholar]
- Jørgensen, M.; Shepperd, M. A Systematic Review of Software Development Cost Estimation Studies. IEEE Transactions on Software Engineering 2007, 33, 33–53. [Google Scholar] [CrossRef]
- qi, x.; chen, j.; deng, l. Learning. In Proceedings of the Proceedings of the International Conference on Algorithms and Architectures for Parallel Processing (ICA3PP). Springer, 2023; Vol. 13777, pp. 123–137. [Google Scholar] [CrossRef]
- Das, A.; Kong, W.; Leach, A.; Mathur, S.; Sen, R.; Yu, R. A Decoder-Only Foundation Model for Time-Series Forecasting. In Proceedings of the Proceedings of the International Conference on Machine Learning (ICML), Vienna, Austria, 2024; pp. 567–589, Google Research.
- Feng, Z.; Guo, D.; Tang, D.; Duan, N.; Feng, X.; Gong, M.; Shou, L.; Qin, B.; Liu, T.; Jiang, D.; et al. Languages. In Proceedings of the Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, 2020; pp. 1536–1547. [CrossRef]
- Zhang, C.; Xie, Y.; Bai, H.; Yu, B.; Li, W.; Gao, Y. A Survey on Federated Learning. Knowledge-Based Systems 2021, 216, 106775. [Google Scholar] [CrossRef]





| Model | Accuracy | 95% CI | Precision | Recall | F1 | ROC-AUC |
|---|---|---|---|---|---|---|
| % | (%) | (%) | (%) | (%) | (%) | (%) |
| Logistic Regression | 61.55 | [60.88, 62.22] | 62.64 | 90.20 | 73.94 | 61.91 |
| Gradient Boosting | 81.34 | [80.72, 81.96] | 80.31 | 91.59 | 85.58 | 88.59 |
| Random Forest | 85.24 | [84.67, 85.79] | 86.34 | 93.55 | 89.80 | 91.38 |
| Majority Class | 69.50 | - | - | 100.00 | 82.00 | 50.00 |
| Predicted Class | |||
|---|---|---|---|
| Actual Class | Success (1) | Failure (0) | Total |
| Success (1) | 13,004 (TP) | 896 (FN) | 13,900 |
| Failure (0) | 2,057 (FP) | 4,043 (TN) | 6,100 |
| Total | 15,061 | 4,939 | 20,000 |
| Rank | Feature Name | Importance | Category |
|---|---|---|---|
| 1 | gh_project_maturity_days | 0.0949 | Project Context |
| 2 | git_repository_age_days | 0.0946 | Project Context |
| 3 | gh_commits_count | 0.0902 | Project Context |
| 4 | gh_total_commits | 0.0863 | Project Context |
| 5 | gh_sloc | 0.0766 | Code Metrics |
| 6 | gh_contributors_count | 0.0654 | Project Context |
| 7 | tr_build_number | 0.0612 | Build History |
| 8 | gh_test_density | 0.0587 | Test Structure |
| 9 | gh_tests_count | 0.0543 | Test Structure |
| 10 | tr_builds_last_30_days | 0.0498 | Build History |
| Top 10 Cumulative | 73.2% |
| Language | Accuracy | Precision | Recall | F1 | Builds |
|---|---|---|---|---|---|
| (%) | (%) | (%) | (%) | (%) | |
| Java | 84.21 | 82.45 | 94.12 | 87.90 | 8,000 |
| Ruby | 82.14 | 81.02 | 93.01 | 86.61 | 7,000 |
| Python | 81.54 | 79.98 | 93.85 | 86.39 | 3,000 |
| JavaScript | 80.83 | 79.23 | 92.98 | 85.56 | 2,000 |
| Overall | 82.74 | 80.88 | 93.57 | 86.76 | 20,000 |
| Std Dev | 1.38 | 1.32 | 0.49 | 0.98 | - |
| Range | 3.38 | 3.22 | 1.14 | 2.34 | - |
| Feature Set | Accuracy | ROC-AUC | F1 | Production |
|---|---|---|---|---|
| (%) | (%) | (%) | (%) | Viable |
| All 66 Features (with leakage) | 97.80 | 99.56 | 98.42 | No |
| 31 Clean Pre-Build Features | 82.73 | 91.38 | 86.76 | Yes |
| Performance Drop | -15.07 | -8.18 | -11.66 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).