Submitted:
10 June 2026
Posted:
11 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. Research Questions
- RQ1. How much of the reported performance of flow-based IDS models is attributable to the choice of train-test split rather than to the model?
- RQ2. Does a model trained on one flow-based dataset transfer to another, and if not, can a small target calibration buffer recover useful performance?
- RQ3. Can the observed transfer behaviour be explained by measurable distributional drift in the shared feature space?
1.2. Contributions
- We show, on CSE-CIC-IDS2018, that the train-test protocol is the dominant variable in reported performance: a random split and a class-stratified temporal split agree (macro-F1 ≈ 0.79–0.82 across three model families), while a class-blind chronological split collapses to macro-F1 ≈ 0.02. The gap is statistically significant (95% bootstrap CI [0.75, 0.80], p < 0.001) and model-agnostic.
- We attribute the collapse to unseen-attack generalization failure rather than to gradual concept drift, using a stratified-temporal control that holds class coverage fixed, and we verify the harness is leakage-free with a label-shuffle baseline.
- We characterize how recalibration and discrimination come apart under cross-dataset transfer across two distinct CIC-DDoS2019 regimes: a deployment-realistic, time-ordered target buffer repairs confidence reliably (ECE ≤ 0.11 in both regimes at the largest (10%) buffer) but restores discrimination unevenly, weakly in one regime (MCC ≈ 0.26) and substantially in the other (MCC ≈ 0.71), and the calibration metrics do not reveal which. Repairing confidence therefore does not guarantee discrimination, so calibration alone is an insufficient health signal for a drifted detector.
- We provide a distributional explanation: 36 of 45 common-core features exceed the severe-drift threshold of PSI ≥ 0.25 (KS up to 0.79), concentrated in timing, packet-length, and rate features consistent with a DDoS-flood signature.
- We document a fully specified, deterministic pipeline (seeds, bootstrap confidence intervals, and a leakage control) and provide a reproducibility manifest (Appendix A) listing the exact releases, class counts, label mapping, feature set, and statistic settings; the complete code and configuration are available from the author on request.
2. Background and Related Work
2.1. Concept Drift and Temporal Evaluation Bias
2.2. Benchmark Criticism And Cross-Dataset Generalization
2.3. Calibration in Intrusion Detection
3. Datasets
4. A Drift Taxonomy for Flow-Based IDS
5. Methodology
5.1. Data Preparation and Harmonisation
5.2. Drift Statistics
5.3. Models and Evaluation Protocols
6. Results
6.1. The Evaluation Protocol Is the Hidden Variable


6.2. Cross-Dataset Transfer Collapses Zero-Shot; Recalibration Repairs Confidence but Restores Discrimination only Unevenly
6.3. A Distributional Explanation for the Transfer Collapse
7. Discussion
8. Recommendations and Future Work
8.1. Recommendations for IDS Evaluation
- Report at least one time-aware split (chronological or stratified-temporal) in addition to any random split, and report the gap.
- When claiming generalization, report at least one cross-dataset or cross-regime transfer result.
- Retain timestamps, day identifiers, and scenario identifiers as metadata even when they are excluded from model inputs.
- State the exact dataset release, cleaning script, label mapping, and feature-alignment procedure.
- Report calibration (ECE, Brier) and discrimination (MCC, AUPRC) metrics, plus false-positive burden, never accuracy or weighted-F1 alone, especially under class imbalance.
- Accompany headline gaps with confidence intervals and a leakage control.
8.2. Future Work
9. Conclusion
Data Availability
Use of AI Tools
Conflicts of Interest
Appendix A. Reproducibility Manifest
References
- Canadian Institute for Cybersecurity, “CSE-CIC-IDS2018 on AWS,” University of New Brunswick. [Online]. Available: https://www.unb.ca/cic/datasets/ids-2018.html.
- I. Sharafaldin, A. H. Lashkari, and A. A. Ghorbani, “Toward generating a new intrusion detection dataset and intrusion traffic characterization,” in Proc. 4th Int. Conf. Information Systems Security and Privacy (ICISSP), 2018, pp. 108–116.
- Canadian Institute for Cybersecurity, “DDoS Evaluation Dataset (CIC-DDoS2019),” University of New Brunswick. [Online]. Available: https://www.unb.ca/cic/datasets/ddos-2019.html.
- M. Cantone, C. Marrocco, and A. Bria, “Machine learning in network intrusion detection: A cross-dataset generalization study,” IEEE Access, vol. 12, pp. 144489–144508, 2024. [CrossRef]
- L. Liu, G. Engelen, T. Lynar, D. Essam, and W. Joosen, “Error prevalence in NIDS datasets: A case study on CIC-IDS-2017 and CSE-CIC-IDS-2018,” in Proc. IEEE Conf. Communications and Network Security (CNS), 2022, pp. 254–262. [CrossRef]
- A. Raskovalov, N. Gabdullin, and V. Dolmatov, “Investigation and rectification of NIDS datasets and standardized feature set derivation for network attack detection with graph neural networks,” arXiv:2212.13994, 2022.
- C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proc. 34th Int. Conf. Machine Learning (ICML), 2017, pp. 1321–1330.
- J. C. Platt, “Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods,” in Advances in Large Margin Classifiers, A. J. Smola et al., Eds. Cambridge, MA: MIT Press, 1999, pp. 61–74.
- J. Gama, P. Medas, G. Castillo, and P. Rodrigues, “Learning with drift detection,” in Advances in Artificial Intelligence - SBIA 2004, Lecture Notes in Computer Science, vol. 3171, 2004, pp. 286–295.
- F. Pendlebury, F. Pierazzi, R. Jordaney, J. Kinder, and L. Cavallaro, “TESSERACT: Eliminating experimental bias in malware classification across space and time,” in Proc. 28th USENIX Security Symposium, 2019, pp. 729–746.
- D. Arp, E. Quiring, F. Pendlebury, A. Warnecke, F. Pierazzi, C. Wressnegger, L. Cavallaro, and K. Rieck, “Dos and don’ts of machine learning in computer security,” in Proc. 31st USENIX Security Symposium, 2022, pp. 3971–3988.
- J. G. Moreno-Torres, T. Raeder, R. Alaiz-Rodriguez, N. V. Chawla, and F. Herrera, “A unifying view on dataset shift in classification,” Pattern Recognition, vol. 45, no. 1, pp. 521–530, 2012. [CrossRef]
- M. Sarhan, S. Layeghy, N. Moustafa, and M. Portmann, “NetFlow datasets for machine learning-based network intrusion detection systems,” in Big Data Technologies and Applications (BDTA), Lecture Notes of ICST, 2020, pp. 117–135. [CrossRef]
- M. P. Naeini, G. F. Cooper, and M. Hauskrecht, “Obtaining well calibrated probabilities using Bayesian binning,” in Proc. AAAI Conf. Artificial Intelligence, vol. 29, no. 1, 2015, pp. 2901–2907.
- A. H. Lashkari, G. Draper-Gil, M. S. I. Mamun, and A. A. Ghorbani, “Characterization of Tor traffic using time based features,” in Proc. 3rd Int. Conf. Information Systems Security and Privacy (ICISSP), 2017, pp. 253–262.
- N. Siddiqi, Credit Risk Scorecards: Developing and Implementing Intelligent Credit Scoring. Hoboken, NJ, USA: Wiley, 2006.
- NIST/SEMATECH, “Kolmogorov-Smirnov goodness-of-fit test,” in e-Handbook of Statistical Methods. [Online]. Available: https://www.itl.nist.gov/div898/handbook/eda/section3/eda35g.htm.




| Dataset | Collection context | Attack coverage | Drift relevance |
|---|---|---|---|
| CSE-CIC-IDS2018 | Enterprise/cloud AWS testbed; traffic organized by scenario days | Brute force, botnet, DoS, DDoS, web attacks, infiltration | Enterprise baseline; strong temporal-leakage risk if day/scenario structure is ignored |
| CIC-DDoS2019 | DDoS-focused CIC testbed; two separate capture days | Reflection/amplification DDoS families; Portmap appears only in the testing-day regime | Attack-scope and label-distribution drift; severe class imbalance; novel-family transfer target |
| Drift layer | Meaning | Where it appears here |
|---|---|---|
| Feature drift | The distribution of measured variables changes | 36/45 common-core features shift between datasets (Section 6.3) |
| Label-distribution drift | The frequency of benign vs. attack classes changes | Within the shared {benign, ddos} space, benign prevalence falls from ~38% (CSE-CIC-IDS2018) to ~3% (CIC-DDoS2019) |
| Predictive concept drift | The feature→label relationship changes | Zero-shot cross-dataset transfer collapses to MCC ≤ 0 (Section 6.2) |
| Temporal/context drift | The order or context of traffic affects evaluation | Random vs. chronological splitting changes macro-F1 by 0.78 (Section 6.1) |
| Model | Protocol | Macro-F1 | Weighted-F1 | MCC | ECE |
|---|---|---|---|---|---|
| RF | Random | 0.807 | 0.921 | 0.907 | 0.002 |
| RF | Temporal (stratified) | 0.787 | 0.907 | 0.891 | 0.018 |
| RF | Chronological (class-blind) | 0.021 | 0.051 | 0.026 | 0.734 |
| LightGBM | Random | 0.824 | 0.921 | 0.909 | 0.003 |
| LightGBM | Temporal (stratified) | 0.811 | 0.900 | 0.884 | 0.026 |
| LightGBM | Chronological (class-blind) | 0.025 | 0.057 | 0.057 | 0.826 |
| MLP | Random | 0.786 | 0.901 | 0.885 | 0.002 |
| MLP | Temporal (stratified) | 0.796 | 0.898 | 0.880 | 0.016 |
| MLP | Chronological (class-blind) | 0.021 | 0.050 | 0.010 | 0.829 |
| Regime | Setting | Macro-F1 | Weighted-F1 | MCC | AUPRC | FP-rate | Brier | ECE |
| 1 (12 Jan) | Zero-shot | 0.027 | 0.002 | −0.043 | 0.551 | 0.006 | 1.776 | 0.928 |
| 1 | Buffer 1% | 0.207 | 0.356 | 0.065 | 0.540 | 0.037 | 0.650 | 0.407 |
| 1 | Buffer 5% | 0.547 | 0.893 | 0.262 | 0.540 | 0.147 | 0.151 | 0.140 |
| 1 | Buffer 10% | 0.542 | 0.888 | 0.249 | 0.539 | 0.186 | 0.130 | 0.108 |
| 2 (11 Mar) | Zero-shot | 0.046 | 0.005 | −0.078 | 0.874 | 0.008 | 1.729 | 0.906 |
| 2 | Buffer 1% | 0.853 | 0.973 | 0.712 | 0.870 | 0.177 | 0.157 | 0.224 |
| 2 | Buffer 5% | 0.847 | 0.972 | 0.694 | 0.870 | 0.268 | 0.043 | 0.039 |
| 2 | Buffer 10% | 0.846 | 0.971 | 0.693 | 0.870 | 0.268 | 0.042 | 0.023 |
| Feature | PSI | KS |
|---|---|---|
| fwd_pkt_len_min | 2.51 | 0.792 |
| fwd_pkt_len_mean | 2.63 | 0.787 |
| fwd_seg_size_mean | 2.63 | 0.787 |
| pkt_len_mean | 4.23 | 0.763 |
| flow_iat_mean | 3.89 | 0.758 |
| flow_pkts/s | 3.53 | 0.750 |
| fwd_iat_min | 6.35 | 0.747 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).