Submitted:
10 September 2026
Posted:
14 September 2026
You are already at the latest version
Abstract
Machine learning-based network intrusion detection, the explanation of detector decisions, and alert triage have each developed as a largely separate line of research. As a result, a detector’s aggregate performance score still tells a security analyst little about how far to trust an individual alert: it does not reveal how reliable the detector is for each attack type, whether the explanation of a decision is faithful, or on what calibrated basis some alerts should be automated and others escalated. We propose a confusion-aware triage framework that addresses these three concerns within a single, unified approach. From a trained detector’s predictions, the framework estimates the reliability of each attack type, validates the explanations of its decisions, and routes each alert, by that reliability, into automatic handling, analyst assistance, or escalation, presenting each escalated alert as a compact decision card. We evaluate the framework on UNSW-NB15 and CIC-IDS2017 with six detector families over five seeds. Per-attack-type reliability is estimable, stable across training runs, and transfers to unseen traffic, whereas a detector’s confidence overstates its reliability on the rarest attacks; the explanations are faithful for seven of the eight routed attack types; and routing by reliability escalates 27.9% of alerts while catching 75.7% of the detector’s errors, and abstains where the detector is already reliable. Routing by reliability matches routing by confidence on error coverage while remaining calibrated, per-attack-type, and auditable. The framework is proposed as an analyst-facing decision aid, and its evaluation with security analysts is left to future work.
Keywords:
network intrusion detection
; explainable AI
; per-attack-type reliability
; alert triage
; explanation faithfulness
; selective classification
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.