Preprint
Article

This version is not peer-reviewed.

Auditing Explanation Faithfulness in Privacy-Preserving Federated Intrusion Detection

Submitted:

20 August 2026

Posted:

20 August 2026

You are already at the latest version

Abstract
Security operations centres (SOCs) face two simultaneous mandates: keep the network traffic they analyse private, and make their AI-based detectors explainable. These goals are usually assumed to be in tension, because in image and language models differential privacy (DP) is reported to degrade post-hoc explanations. We present the first systematic audit of explanation faithfulness under formally-accounted (ε,δ)-DP in federated, tabular deep intrusion detection (IDS), and we add a previously-unstudied axis: the interaction between Byzantine-robust aggregation and explanation quality. Across two datasets (ToN_IoT, CICIDS2017), two architectures (an MLP and a compact tabular transformer), three attribution methods (Integrated Gradients, GradientSHAP, attention rollout), privacy budgets ε ∈ {8, 4, 1} (ε ∈ {4, 1} for the transformer), and heterogeneity levels, and using paired Wilcoxon tests with Cliff’s δ effect sizes, we find — contrary to our own anchor hypothesis and to the image/NLP literature — that client-side DP-SGD does not degrade explanation faithfulness. Measured by comprehensiveness and the AOPC MoRF−LeRF separation it significantly increases self-faithfulness (large δ, all seeds), an effect that attenuates at the strongest budget and survives a confidence-ceiling (f0) normalization. A pre-registered capacity-matched control attributes this faithfulness gain largely to DP’s utility cost — non-DP models throttled to the same accuracy by three independent mechanisms reproduce it — leaving only a small, clipping-driven residual in the centralized setting. DP’s real cost falls on utility and, most sharply, on rare-attack detection (mitm F1 0.65→0.05). Per-class utility and faithfulness decouple: well-supported classes hold their F1 while their attribution gap widens. A strong model-poisoning attack collapses FedAvg in both utility (0.83→0.20) and faithfulness (gap 0.58→0.21), whereas median, trimmed-mean and multi-Krum preserve both — though median/Krum cost some faithfulness when no attack is present, making trimmed-mean the best all-round trade-off. Among attribution methods, attention rollout is the most fragile. Practically: for explainable federated IDS, gradient/Shapley attributions remain trustworthy under DP, and trimmed-mean aggregation buys robustness at little explanation cost.
Keywords: 
;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.