Submitted:
22 August 2026
Posted:
24 August 2026
You are already at the latest version
Abstract
Network intrusion detection requires accurate classification of high-volume network traffic while maintaining reliable evaluation and practical analyst support. This study presents a supervised network anomaly-classification workflow using the BigFlow-NIDS V2 benchmark, integrating leakage-aware data separation, SMOTE-Tomek balancing, FlowTransformer classification, confidence and novelty scoring, bounded Large Language Model (LLM) assistance, and analyst triage. The quantitative experiment uses a controlled binary subset of Benign and Scanning traffic, comprising 3,972 usable flows and 795 independent test observations. On the independent test set, FlowTransformer achieved 0.9107 accuracy, 0.9119 macro-precision, 0.9105 macro-recall, 0.9106 macro-F1, 0.9105 balanced accuracy, MCC of 0.8225, Cohen's kappa of 0.8213, and macro PR-AUC of 0.959. A RandomForest baseline achieved 0.9308 accuracy. Paired McNemar, Wilcoxon, and t-tests indicated a statistically significant difference, with the baseline performing better on this test set. The LLM was restricted to bounded advisory analysis, while analyst triage remained separate from model outputs. The results provide a reproducible binary evaluation and a basis for extending the workflow to the full multiclass benchmark.
Keywords:
network intrusion detection
; BigFlow-NIDS V2
; FlowTransformer
; random forest
; statistical validation
; large language model
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.