Preprint
Article

This version is not peer-reviewed.

MedBackdoorGuard: Detecting and Mitigating Backdoor Attacks in Deep Learning Models for MRI-Based Alzheimer's Disease and Mild Cognitive Impairment Diagnosis

Submitted:

20 September 2026

Posted:

20 September 2026

You are already at the latest version

Abstract
Deep learning classifiers trained on structural magnetic resonance imaging (MRI) are increasingly proposed for early detection of Alzheimer’s disease (AD) and mild cognitive impairment (MCI), yet the pipelines that produce these models depend on shared datasets, pre-trained weights, and federated updates that an adversary can poison. A backdoored diagnostic model behaves correctly on ordinary scans but collapses to an attacker-chosen label whenever a hidden trigger is present, which turns a prognostic instrument for cognitive aging into a silent source of misdiagnosis. This article proposes MedBackdoorGuard, a defense framework that detects and mitigates backdoor attacks in MRI-based AD and MCI classifiers without prior knowledge of the trigger. The framework couples spectral latent screening of training data with latent trigger isolation in activation space, then immunizes the model through trigger-sensitive neuron pruning and fine-tuning under an anatomical attention consistency constraint that anchors the decision on medial temporal structures. A runtime inference guard adds superimposition entropy and anatomically informed reconstruction to purify suspicious scans at deployment. Across ADNI, OASIS-1, and a public four-class dementia MRI benchmark under patch, blended, warping, and frequency triggers, MedBackdoorGuard reduces the mean attack success rate from 97.4% to 3.0% while holding clean accuracy within 0.6 percentage points of the undefended model, reaches a poisoned-sample detection AUROC of 97.4%, and identifies 59 of 60 backdoored models. Compared with the strongest defense baseline, attack success rate falls by 15.1 percentage points and clean accuracy rises by 1.7 percentage points. The framework restores MCI recall on triggered scans from 3.9% to 88.3%, which is the outcome that matters for early detection in cognitive aging.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.