Preprint
Article

This version is not peer-reviewed.

Conditional Latent Diffusion for Synthetic Brain MRI in Alzheimer’s Disease: A Preprocessing-Focused Pipeline

Submitted:

06 August 2026

Posted:

07 August 2026

You are already at the latest version

Abstract
Alzheimer’s disease (AD) detection from structural magnetic resonance imaging (MRI) using deep learning depends on large, labelled datasets. However, many cohorts contain only a few hundred subjects, for which standard augmentation adds very limited diversity. Latent diffusion models (LDMs) offer an alternative. The current study presents a reproducible two-stage pipeline that generates 2D coronal brain MRI conditioned on diagnosis and trained on 295 subjects from the Alzheimer’s Disease Neuroimaging Initiative (ADNI). In the first stage, a variational autoencoder with an adversarial objective compressed 256 × 256 slices into a 32 × 32 × 8 latent space. Then, a class-conditional LDM with classifier-free guidance generated AD and cognitively normal (CN) images. A baseline and a redesigned pipeline adding MNI152 registration, anatomically guided slice selection, and increased latent capacity were compared on the same subjects. The redesigned pipeline reached a Kernel Inception Distance of 0.030 ± 0.002 and a bias-corrected Fréchet Inception Distance (FID∞) of 42.14. A classifier trained only on synthetic data achieved an area under the curve (AUC) of 0.754 (95% CI, 0.60 to 0.91), compared with 0.810 for real data. No instance memorisation was detected across 880 samples, and a control analysis exposed a 32.7-percentage-point inflation in the standard memorisation metric under unequal reference sets. The redesigned preprocessing pipeline improves synthesis quality at a small-cohort scale and offers an approach that may be transferable to other privacy-restricted medical imaging domains.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Alzheimer’s disease is the most common cause of dementia, affecting over 55 million people worldwide, with projections reaching 139 million by 2050 [1]. The disease follows a progressive neurodegenerative course in which neurofibrillary tangle pathology begins in the entorhinal cortex, spreads to the hippocampus, and eventually reaches neocortical association areas [2]. Structural brain magnetic resonance imaging (MRI) can detect the atrophy years before clinical diagnosis, making hippocampal volume loss, ventricular enlargement, and cortical thinning as established imaging biomarkers for early detection.
Deep learning has been widely applied to the automated detection of Alzheimer’s disease from structural MRI. A systematic review of 116 studies reported a mean Convolutional Neural Network (CNN) accuracy of 78.5% for predicting conversion from mild cognitive impairment (MCI) to Alzheimer’s disease, with 107 of the 116 studies relying on data from the ADNI database [3]. Anatomically informed feature design can exceed this average, as Herzog and Magoulas [4] separated cognitively normal subjects from Alzheimer’s disease cases at 93.0% accuracy on ADNI data using machine learning applied to features of brain asymmetry. These methods require large, labelled datasets to train effectively, yet many clinical research groups have access to cohorts of only 50 to 500 subjects. Class imbalance compounds the problem, with cognitively normal participants typically outnumbering those with Alzheimer’s disease. Multi-site acquisition introduces further confounds, as scanner-dependent intensity distributions across ADNI’s 54 sites are classifiable by manufacturer with approximately 99% accuracy [5]. Privacy regulations under frameworks such as the UK GDPR additionally constrain the sharing of medical imaging data between institutions, even when de-identified.
Generative models offer a fundamentally different approach to data scarcity. Rather than augmenting existing images through geometric transformations, which cannot introduce new anatomical variation, generative models learn the underlying data distribution and synthesise new samples from it [6]. Generative adversarial networks produce sharp outputs but suffer from mode collapse and training instability that worsens as dataset size decreases [7]. Denoising diffusion probabilistic models avoid these instabilities through iterative noise prediction [8], and latent diffusion models (LDMs) brought this approach to practical resolution by performing the diffusion process in a compressed latent space rather than at pixel resolution, reducing computational cost by 16- to 64-fold [9].
Several studies have applied LDMs to brain MRI. Pinaya et al. [10] trained a VQ-VAE autoencoder paired with a conditional diffusion model on 31,740 UK Biobank T1-weighted volumes at 160 × 224 × 160 voxels. Ventricular volume conditioning yielded a correlation of r = 0.972 between specified and measured volumes. The UK Biobank cohort is predominantly healthy, though, and the study did not condition on disease status. Müller-Franzes et al. [11] introduced Medfusion, a latent DDPM with eightfold spatial compression, and demonstrated generalisation across fundoscopy, histopathology, and chest radiography, achieving FID values of 11.63 to 30.03 and consistently outperforming GAN baselines in both precision and recall. Khader et al. [12] extended this to 3D volumes using VQ-GAN with DDPM across chest CT, brain MRI, and knee MRI. In a radiologist assessment, 50 of 50 generated ADNI brain-MRIs were rated at least largely realistic. Dhinagar et al. [13] trained a conditional LDM on 4,098 ADNI T1-weighted scans from 1,188 subjects and showed a 3-percentage-point improvement in AD classification AUC when augmenting with synthetic data, using classifier-free guidance with 15% label dropout and a guidance scale of 2. Counterfactual heat maps from that study highlighted ventricular and temporal regions consistent with known atrophy patterns. Across these studies, two characteristics recur. All operate at cohort scales of 998 to 31,740 subjects, and the preprocessing methodology is minimally documented. Dhinagar et al. describe their preprocessing (N4 bias correction, skull stripping, linear registration, and min-max scaling) in a single paragraph. Pinaya et al. devote a brief subsection to their UK Biobank pipeline. Müller-Franzes et al. provide no brain MRI preprocessing details at all.
This gap matters because preprocessing decisions directly determine what the generative model receives as input. Without spatial registration, anatomical structures occupy different voxel positions across subjects, forcing the diffusion model to average over misaligned features and producing blurred outputs. Without principled slice selection, coronal sections may sample diagnostically uninformative regions while under-representing the hippocampus and entorhinal cortex, where Alzheimer’s pathology is most pronounced and where classification evidence is strongest. Mendoza-Léon et al. [14] classified AD at 90.0% accuracy from single coronal slices, and Chen et al. [15] improved diagnosis using a slice-level attention mechanism that emphasises informative slices and discards redundant ones. Despite this, no published work documents preprocessing at this level of detail for brain MRI synthesis or reports a same-cohort comparison between a baseline and an anatomically informed pipeline at a small-cohort scale. Variable-density anatomical slice selection, which densely samples disease-relevant regions while sparsely covering reference anatomy, has not been combined with generative modelling in the synthesis literature.
This study addresses these gaps through a two-stage conditional LDM pipeline for generating synthetic 2D coronal brain MRI conditioned on Alzheimer’s disease diagnosis, trained on 295 subjects from ADNI. The pipeline comprises a VAEGAN autoencoder [16] that compresses 256 × 256 images into a 32 × 32 × 8 latent space, followed by a conditional LDM with classifier-free guidance. Two pipeline configurations were applied to the same 295 subjects. The baseline (v1) omitted spatial registration, used fixed-index slice extraction, encoded 224 × 224 images into a four-channel latent, and set the LPIPS (Learned Perceptual Image Patch Similarity) [17] loss weight to 0.001, effectively disabling perceptual optimisation. The enhanced version (v2) incorporated MNI152 registration, variable-density anatomical slicing informed by Braak staging neuropathology, an eight-channel latent at 256 × 256, and a literature-informed LPIPS weight of 0.5. Because these changes were applied together, the comparison shows the combined effect of the redesigned pipeline, and the augmentation ablation then isolates training-side effects, indicating that the improvement does not stem from insufficient data diversity.
Thus, the work makes four contributions. First, the six-stage preprocessing pipeline with MNI152 registration and anatomically guided slice selection is documented in detail not found in existing brain MRI synthesis literature, and the v1-to-v2 redesign on the same 295 subjects shows that these design choices, taken together, substantially improve synthesis quality when cohort size is limited. Second, the pipeline achieves a TSTR AUC of 0.754, retaining 93% of real-data discriminative signal from a cohort approximately one fourteenth the size of that used by the nearest comparable study [13], establishing clinical utility at small scale. Third, a 32.7- percentage-point set-size confound in the Scardace et al. [18] nearest-neighbour memorisation metric is identified and quantified, with a corrected protocol using equal-sized reference sets and real-image control baselines which found zero instance memorisation across 880 generated samples. Fourth, a multi-tier evaluation protocol combining bias-corrected FID∞ [19], KID as primary distributional metric [20], subject-level TSTR with the Yagis et al. [21] correction, and memorisation severity grading provides a practical framework for honest assessment at a small-cohort scale.
The rest of this paper is organised as follows. Section 2 describes the dataset, preprocessing pipeline, model architecture, training protocol, and evaluation framework. Section 3 presents the experimental results and compares model performance. Section 4 discusses the findings in the context of existing literature and identifies limitations. Section 5 concludes with a summary of contributions and directions for future work.

2. Materials and Methods

2.1. Dataset

Data were obtained from the Alzheimer’s Disease Neuroimaging Initiative [22], a multi-site longitudinal study launched in 2003 to develop clinical, imaging, genetic, and biochemical biomarkers for the early detection and tracking of Alzheimer’s disease. The raw dataset comprised 6,169 preprocessing rows across 603 subjects, with each subject potentially represented by multiple preprocessing variants, acquisition visits, and sequence types. To eliminate data leakage from repeated measurements, a strict one-row-per-subject selection protocol was applied. Selection followed a priority hierarchy favouring the most complete preprocessing available for each subject. Priority 1 combined N3 bias field correction with B1 inhomogeneity correction and gradient warping distortion correction. Where Priority 1 was unavailable, the protocol fell back to progressively less complete correction combinations (N3 + gradient warping, then B1 + gradient warping, then gradient warping alone). Baseline visits were preferred over follow-up scans, and MPR sequences over MPR-R when multiple acquisitions existed for the same subject. The resulting cohort comprised 295 unique subjects, of whom 118 had an Alzheimer’s diagnosis (AD) (40%), and 177 were cognitively normal (CN) controls (60%), with 97.6% receiving Priority 1 preprocessing.
The cohort was partitioned into training, validation, and test sets at the subject level using a 70/15/15 ratio. Subject-level partitioning is essential because Yagis et al. [21] demonstrated that slice-level splitting inflates classification accuracy by approximately 29% on ADNI data, as adjacent slices from the same subject share anatomy and effectively constitute data leakage. Scanner manufacturer alone is classifiable with approximately 99% accuracy across ADNI’s 54 acquisition sites [5], so stratification was performed jointly on diagnosis and acquisition site using MultilabelStratifiedShuffleSplit from the iterstrat library with a fixed random seed of 42. The split proceeded in two stages, first separating 15% for the test set, then splitting the remaining subjects into training and validation at an 82.4/17.6 ratio. A chi-square test confirmed no significant class imbalance across the resulting train, validation, and test partitions (chi-squared = 3.053, p = 0.217; Table 1).

2.2. Preprocessing Pipeline

The preprocessing pipeline consists of six sequential stages, each producing verified outputs that serve as inputs to the next (Figure 1). A checkpoint gate follows every stage, which requires all outputs to meet predefined criteria before processing continues. This design prevents error propagation, ensuring that a quality failure at any stage is caught before it corrupts downstream results. All 295 subjects passed every stage with a 0% rejection rate.
Brain extraction (Stage 2) removes non-brain tissue using HD-BET [23], a deep learning skull-stripping tool trained on multicentric clinical data that outperforms six established algorithms in a comparative evaluation. Running in accurate mode with test-time augmentation, HD-BET achieved a 100% success rate across all 295 volumes with an average runtime of approximately 4 seconds per subject. All outputs were verified for RAS orientation and 1 mm isotropic voxel spacing.
Spatial registration (Stage 3) aligns all volumes to the MNI152 standard-space template [24] using a twelve-degree-of-freedom affine transformation implemented with ANTs and SimpleITK. Registration is the most consequential preprocessing step for generative modelling. Without spatial alignment, anatomical structures occupy different voxel positions across subjects, forcing the generative model to average over misaligned features and producing blurred outputs. With registration, the diffusion model can focus its capacity on learning anatomical variation rather than spatial location. All 295 volumes were registered to a common shape of 182 × 218 × 182 voxels at 1 mm isotropic resolution, with a mean normalised cross-correlation of 0.563 (range 0.414 to 0.648). Ventricle alignment across subjects was confirmed through visual inspection. Because affine registration normalises gross brain volume by design, the AD atrophy signal is preserved in fine structural features including hippocampal shape, ventricular enlargement, and cortical thinning.
Quality control (Stage 4) applies automated outlier detection across six metrics computed for every volume. These include brain volume, mean and standard deviation of intensity, mask contiguity, shape consistency, and NaN/Inf detection. Seven subjects exceeded the three-standard-deviation threshold on at least one metric. All seven were accepted after a visual review of three-plane montages confirmed the absence of genuine artefacts. The 0% rejection rate reflects the standardised acquisition protocol maintained across ADNI sites.
Anatomically guided slice extraction (Stage 5) produces twenty coronal slices per subject at specific MNI Y-coordinates ranging from +58 mm (frontal pole) to −100 mm (occipital pole), following a variable-density sampling design informed by Braak- staging neuropathology [2,25]. Seven slices are densely sampled through the hippocampal formation (Y = −4 to −32 mm at 4 to 8 mm spacing), targeting the structures that show earliest atrophy in Braak stages III and IV, and that carry the highest discriminative value for AD classification [14]. Four additional slices capture the posterior cingulate cortex and precuneus (Y = −45 to −62 mm), regions showing early metabolic abnormalities in the Alzheimer’s continuum. The remaining nine slices provide sparse frontal and occipital reference anatomy at 10 to 16 mm intervals. Prior to slicing, each volume is normalised to [0,1] using per-volume percentile clipping at the 0.5th and 99.5th percentiles, which is the evidence-based standard for generative brain MRI preprocessing [26]. The 99.5th percentile values ranged from 47.3 to 2,271.5 across the 295 subjects, a nearly 50-fold variation in raw intensity that the normalisation successfully harmonises to a common range. Slices are extracted at 256 × 256 resolution using bicubic (order 3) interpolation and saved as float32 arrays. All twenty slice positions were verified to contain between 42% and 75% brain tissue, confirming that no slice captures predominantly empty space.
Data splitting (Stage 6) partitions the cohort into training, validation, and test sets using the subject-level stratified protocol described in Section 2.1, producing 5,900 slices in total across three non-overlapping sets with manifest files tracking each subject through the complete pipeline.

2.3. Model Architecture

The pipeline follows the two-stage latent diffusion paradigm introduced by Rombach et al. [9] and applied to medical imaging by Müller-Franzes et al. [11] as Medfusion. Stage 1 compresses input images into a compact latent representation. Stage 2 generates new latent codes conditioned on diagnostic class, which are then decoded back to pixel space (Figure 2). This separation is motivated by computational and quality considerations. Running diffusion at 256 × 256 pixel resolution would be impractical for small-dataset research on accessible hardware, whereas a 32 × 32 latent resolution reduces dimensionality by approximately 64-fold. The architecture also introduces a fundamental constraint. Autoencoder reconstruction fidelity sets an absolute ceiling on generation quality, making autoencoder design and training the first-order concern.

2.3.1. Stage 1 (VAEGAN Autoencoder)

The autoencoder is implemented using MONAI’s AutoencoderKL [27] with a four-level encoder-decoder architecture. The encoder applies successive downsampling along the channel dimensions (64, 128, 256, 512), each stage containing two residual blocks, reducing the input from 256 × 256 × 1 to a 32 × 32 × 8 latent representation (downsampling factor f = 8). Self-attention is applied only at the deepest level, while non-local attention is enabled in both encoder and decoder for global context aggregation. The eight-channel latent space follows the Medfusion finding that eight channels preserve fine medical structures more faithfully than the four channels used in standard natural-image autoencoders [11]. The downsampling factor of f = 8 was selected based on Rombach et al. [9], who demonstrated that f = 4 and f = 8 achieve the strongest trade-off between compression and perceptual fidelity.
The loss function combines five components. L1 pixel reconstruction loss (weight 1.0) provides the primary reconstruction signal. Learned perceptual loss using LPIPS with VGG-16 features (weight 0.5) penalises high-level structural differences invisible to pixel-wise metrics [17]. Multi-scale structural similarity loss (weight 0.1) preserves luminance, contrast, and structural patterns across spatial scales. KL divergence regularisation (β = 1 × 10⁻⁶) provides a deliberately minimal constraint on latent space structure, contributing only 0.000025 to the total loss at the observed raw KL of 25 and producing a near-deterministic posterior with a standard deviation of 0.0116 [28]. A PatchGAN adversarial loss with hinge criterion (weight 0.01) sharpens reconstructions once introduced in Phase 2. The LPIPS weight of 0.5 follows the Medfusion loss configuration and represents one of the changes from the baseline pipeline, which used 0.001 and effectively disabled perceptual optimisation.
Training follows a two-phase schedule. Phase 1 (epochs 0 to 39) trains with reconstruction losses only, allowing the autoencoder to learn reasonable image reconstructions before introducing adversarial pressure. Phase 2 (epoch 40 onward) adds a three-layer PatchGAN discriminator with spectral normalisation and R1 gradient penalty (γ = 10), stabilisation techniques essential for small-dataset adversarial training [7]. The optimiser was AdamW with a learning rate of 1 × 10⁻⁴, weight decay of 1 × 10⁻⁴, and gradient clipping at 1.0. Early stopping with patience of 30 epochs on validation SSIM selected epoch 70 as the best checkpoint. Total training time was 3.8 hours on an A100-80GB GPU.
The trained latent space has a mean of −3.16, a standard deviation of 5.40, and a range of [−48.7, 34.4]. The scale factor 0.1852 (= 1/5.40) normalises latents to approximately unit variance before the diffusion stage, serving as a critical constant that links the two pipeline stages.

2.3.2. Stage 2 (Conditional Latent Diffusion Model)

The diffusion model uses MONAI’s DiffusionModelUNet with channel dimensions of (128, 256, 256), deliberately compact compared to Stable Diffusion’s (320, 640, 1280). This sizing reflects both the modest latent resolution of 32 × 32 and the small training set of 4,080 slices. Dar et al. [29] demonstrated that smaller generative architectures memorise less patient data, making a compact design preferable for medical imaging. The network comprises three levels with self-attention at levels two and three (operating at 16 × 16 and 8 × 8 resolution), totalling approximately 39.1 million parameters. Input and output channels are set to eight, matching the autoencoder’s latent dimensionality.
Class-conditional generation is implemented via cross-attention. A learned embedding maps three tokens (CN at index 0, AD at index 1, and an unconditional token at index 2) to a 64-dimensional space. At each attention level, UNet features attend to the projected class embedding through query, key, and value projections. The unconditional token is used during training for classifier-free guidance [30]. With probability p = 0.15, the class label is replaced by the unconditional token, enabling the model to learn both conditional and unconditional score estimates. At inference, two forward passes per timestep produce conditioned and unconditioned predictions, combined as ε = ε_uncond + w(ε_cond − ε_uncond). The guidance scale w = 2.0 was determined by an empirical sweep over {1.0, 2.0, 3.0, 4.0, 5.0, 7.5}. This is lower than the literature recommendation of 3 to 5 but was empirically optimal for binary conditioning on this small dataset, as higher values reduced diversity without improving fidelity.
The noise schedule uses a scaled linear beta schedule with β_start = 0.0015 and β_end = 0.0195 over T = 1,000 timesteps. This wider range compared to Stable Diffusion (0.00085 to 0.012) compensates for the higher variance of the latent space (std ≈ 5.4 versus approximately 0.7 for natural image autoencoders), ensuring that the forward process reaches sufficient noise levels to enable generation from pure noise. The training objective is the simplified noise-prediction loss of Ho et al. [8], in which the model predicts the noise added at each timestep, conditioned on the noisy latent, the timestep, and the class label. Inference uses DDIM with 50 deterministic steps, matching the sampler and step count adopted by Pinaya et al. [10], who reduced the reverse process from 1,000 DDPM steps to 50 DDIM steps for brain MRI generation.
Training used AdamW with a learning rate of 2.5 × 10⁻⁵, weight decay of 0.01, cosine decay after a 500-step warmup, and gradient clipping at 1.0. Batch size was 16 with a WeightedRandomSampler to oversample the minority AD class to approximately 50/50 per batch. Exponential moving average with a decay of 0.999 was applied to both UNet and class embedding weights for inference. Early stopping with patience of 25 selected epoch 95 as the best checkpoint from 121 total epochs, with a validation loss of 0.0559. The training and validation loss gap stayed below 0.003 across all 121 epochs, with no evidence of overfitting. Total training time was 82.3 minutes on an L4 GPU using 4.7 GB of 22.5 GB available VRAM.
The complete generation pipeline proceeds in four steps. A latent tensor is sampled from a standard normal distribution with shape 32 × 32 × 8. DDIM denoising runs for 50 steps with class conditioning to produce a denoised latent. The latent is then divided by the scale factor (0.1852) to restore the autoencoder’s native scale. Finally, the VAEGAN decoder maps the latent space back to a 256 × 256 synthetic MRI slice, with output values clipped to [0,1]. A critical implementation detail governs this connection. The DDIM scheduler in MONAI 1.4.0 defaults to clip_sample = True, which clips denoised latents to [−1, 1] at every sampling step. Because the autoencoder’s latent space ranges from −48.7 to 34.4, this default destroys the latent distribution and produces noise rather than brain anatomy. Setting clip_sample = False resolves this completely, but it is not apparent from training metrics alone, since the training and sampling code paths use different scheduler configurations.

2.4. Computational Environment

Preprocessing was performed locally on Windows 11 (Intel Core i9-11900H, Nvidia RTX 3060 notebook, 40 GB RAM) with WSL2 Ubuntu, Python 3.12, using HD-BET and ANTs/SimpleITK . Model training and evaluation were performed on Google Colab Pro. The VAEGAN required an A100-80GB GPU (3.8 hours, approximately 40 GB VRAM). All subsequent training and evaluation were run on the L4 GPU (82.3 minutes for LDM training, 4.7 GB VRAM). Software versions were MONAI 1.4.0, PyTorch 2.0+, and NumPy 1.26.4.

2.5. Evaluation Framework

Evaluation follows a five-tier protocol designed for small-cohort synthesis where standard metrics have known limitations.
Tier 1 assesses autoencoder reconstruction quality using SSIM, PSNR, and LPIPS [17], establishing the quality ceiling for generation. Because any loss of detail at this stage propagates irreversibly to generated images, reconstruction metrics must exceed their targets before proceeding to diffusion model evaluation.
Tier 2 measures distributional similarity between real and generated image sets. KID [20] is the primary metric because it is an unbiased polynomial MMD² estimator that avoids the O(1/N) upward bias that affects FID at small sample sizes [19]. FID is reported secondarily using the FID∞ extrapolation method, in which FID is computed at six sample sizes (N = 200, 350, 500, 650, 800, 880) and extrapolated to N approaching infinity via linear regression on 1/N, yielding a bias-corrected estimate. Precision and recall [31] provide complementary information, separately quantifying fidelity (the fraction of generated images falling within the real-data manifold) and diversity (the fraction of the real distribution covered by generated images). All distributional metrics use Inception v3 features. DINOv2 features were evaluated but showed domain mismatch effects on greyscale brain MRI, with DINOv2-based FID reaching 477 versus Inception-based FID of 46 on identical images, a tenfold discrepancy attributable to DINOv2’s natural-image training distribution rather than to genuine quality differences. DINOv2 is retained only for memorisation detection (Tier 4), where the Scardace et al. [18] framework was originally validated. Whether an image is flagged depends only on whether it lies nearer to the training set than to the held-out set, a comparison of the ordering of two distances rather than their absolute scale, and is therefore far more robust to the feature-space distortion that inflates DINOv2-based FID than a distributional distance would be; because the real-image control is computed in the same space, any systematic component of that distortion shifts the generated and control distributions together and is absorbed in the gap between them.
Tier 3 assesses memorisation using the Scardace et al. [18] nearest-neighbour ratio metric, with two critical additions that are absent from the original methodology. Training images are subsampled to N = 880 to match the holdout set size, eliminating a set-size confound that inflates apparent memorisation when reference sets are unequal. The identical metric is also run on real test images as a leave-one-out control, establishing a null baseline flagging rate. Severity is graded into three categories based on the nearest-neighbour ratio. Values below 0.5 indicate instance memorisation (near-copies of the training data), values between 0.5 and 1.0 indicate distributional proximity (closer to the training data than to the holdout set, which is expected for well-trained models), and values at or above 1.0 indicate distributional generalisation.
Tier 4 evaluates clinical utility through the Train-on-Synthetic-Test-on-Real (TSTR) paradigm. A ResNet-18 classifier [32] is trained exclusively on synthetic images and evaluated on real test data, with a corresponding Train-on-Real-Test-on-Real (TRTR) experiment providing the upper bound. Subject-level aggregation is performed by averaging the predicted probabilities across the twenty slices per subject; a correction is required because Yagis et al. [21] demonstrated that slice-level evaluation inflates AUC by approximately 29% on ADNI data. Ninety-five percent confidence intervals for the subject-level AUCs were obtained from the test-set class counts (18 AD, 26 CN) using the method of Hanley and McNeil [33,34].
Tier 5 provides an augmentation ablation. The LDM is retrained with enhanced online augmentation comprising random rotation (±3°) and Gaussian intensity jitter (σ = 0.03), in addition to the horizontal flip (p = 0.5) already present in the base training. This ablation distinguishes preprocessing-driven improvements from augmentation-driven improvements and tests whether the base model’s capacity was already well matched to the dataset.

2.6. Baseline Comparison

Two pipeline configurations were applied to the same 295 ADNI subjects, and they differed in five respects. The baseline pipeline (v1) omitted spatial registration, extracted slices at a fixed index rather than at anatomically guided coordinates, encoded 224 × 224 images into a four-channel latent, and set the LPIPS perceptual loss weight to 0.001. The enhanced pipeline (v2) added MNI152 registration and variable-density anatomical slicing, encoded 256 × 256 images into an eight-channel latent, and used an LPIPS weight of 0.5. The v1 pipeline produced images described during development as visually unrealistic, with a raw FID of 42.87. Because the two configurations differ in preprocessing, loss weighting, and representational capacity at once, the comparison identifies the combined effect of this pipeline redesign rather than the effect of any single factor. The controlled comparison in this study is instead the augmentation ablation (Section 2.5, Section 3.6), which varies training-time augmentation alone against a fixed architecture and cohort.

3. Results

3.1. Reconstruction Quality

The VAEGAN autoencoder achieved its best validation performance at epoch 70, exceeding all predefined targets by substantial margins. PSNR reached 38.44 dB against a target of 32 dB, exceeding the target by 6.44 dB. SSIM was 0.9861 against a target of 0.95. LPIPS was 0.0166 against a target of 0.08, a 4.8-fold improvement over the threshold. Visual inspection of reconstructions at epoch 70 confirmed that anatomical structures, including sulci, ventricles, grey-white matter boundaries, and the cortical ribbon, were faithfully preserved. AD-specific features such as ventricular enlargement and temporal horn prominence were accurately reproduced, as were CN-specific features including compact ventricles and preserved cortical thickness (Figure 3).
Training dynamics validated the two-phase schedule. Phase 1 reached a PSNR of approximately 35.6 dB before the PatchGAN discriminator was introduced at epoch 40. The discriminator loss spiked to 253.9 at Phase 2 onset, as expected when a discriminator encounters reconstructions already close to real images, then converged to equilibrium at 1.0 by epoch 50. Quality degraded beyond the best epoch, with PSNR dropping from 38.44 dB at epoch 70 to 34.3 dB by epoch 90, validating the early stopping strategy. These reconstruction metrics set the quality ceiling for the subsequent generation stage.

3.2. Qualitative Appearance

The synthetic images capture the class-defining gross morphology (Figure 4): enlarged lateral ventricles and widened cortical sulci in the Alzheimer’s disease cases, and compact ventricles with a fuller cortical mantle in the cognitively normal cases. Fine cortical detail is resolved less sharply than in the real slices, and occasional generations show incomplete or asymmetric structure. This separation between well-formed class-level anatomy and imperfect fine detail frames the quantitative results that follow: the distributional metrics in Section 3.3 measure the residual gap, and the precision of 0.269 reported in Table 2 quantifies the fraction of samples lying outside the real-image manifold.

3.3. Distributional Metrics

The primary distributional metric, Inception v3 KID, measured 0.030 ± 0.002 overall for the enhanced pipeline, within the acceptable range of 0.01 to 0.05 for medical image synthesis (Table 2). Per-class KID revealed that AD images achieved 0.023 compared to 0.037 for CN. This asymmetry is consistent with ventricular enlargement being a spatially large, coherent signal that the diffusion model captures readily, whereas CN brains exhibit subtler inter-subject variation that is harder to match in distributional terms.
The raw FID at N = 880 was 46.23, but this value overestimates the true distributional distance due to FID’s O(1/N) sample-size bias [19]. FID∞ extrapolation, computed at six sample sizes and regressed against 1/N, yielded a bias-corrected estimate of 42.14 with R² = 0.988 (Figure 5). The near-perfect linear fit validates the 1/N bias model and confirms that the raw FID overestimated the true value by approximately 4.1 points, a 9.7% relative bias. The corrected FID∞ of 42.14 is close to the baseline pipeline’s raw FID of 42.87, but the two values are not directly comparable: only the enhanced estimate is sample-size corrected, and the pipelines also differ in input resolution and latent dimensionality, so FID at this scale does not reliably rank the two configurations. The evidence for the v1-to-v2 improvement is instead the qualitative shift from visually unrealistic to anatomically plausible output, together with the enhanced pipeline’s absolute KID, precision, recall, and TSTR results. Per-class raw FIDs of 56.03 for AD and 60.45 for CN were both worse than the overall value, consistent with bias amplification when 880 samples are split into smaller class-specific subsets.
Inception precision was 0.269 and recall was 0.414. The pattern of precision falling below recall is typical of models operating through a spatial bottleneck. The 32 × 32 latent space captures general anatomical structure effectively, producing good coverage of the real distribution (recall 0.414, meaning 41% of the real distribution is represented), but smooths fine details that would place generated images precisely within the real-data manifold (precision 0.269, meaning 27% of generated images match real-image statistics closely). This suggests that the primary quality limitation is autoencoder compression rather than diffusion model capacity.

3.4. Memorisation Analysis

None of the 880 generated images exhibited instance memorisation. The decisive finding is the number of severe cases.. Zero images produced a nearest-neighbour ratio below 0.5 in DINOv2 feature space, the threshold below which generated images would resemble near-copies of specific training examples. True instance memorisation typically produces ratios of 0.05 to 0.30. The worst-case generated image had a ratio of 0.807, far above this range.
The uncorrected flagging rate, computed with the original unequal reference sets (4,080 training images versus 880 test images), was 90.3%. A control experiment applying the same metric to real test images under equal-sized reference conditions (880 by 880) revealed a baseline flagging rate of 38.5%. This control establishes that the Scardace et al. [18] metric produces substantial artefactual flagging from set-size imbalance alone. After correcting by equalising reference set sizes, the generated flagging rate dropped to 57.6%, a gap of 19.1 percentage points above the control baseline (Figure 6). The difference between that naive rate and the corrected rate is 32.7 percentage points; a controlled comparison that holds the generated and held-out sets fixed and varies only the training reference, from 4,080 down to 880, attributes 29.6 of those points to set size alone. The corrected estimate is robust to the subsample: across 1,000 independent training subsamples, the flagging rate stayed within a 95% interval of 52 to 66 percentage points, the generated-minus-control gap remained positive in every subsample (95% interval 11 to 26 percentage points), and no generated image was flagged as instance memorisation in any subsample.
The distribution of generated ratios was unimodal and centred at a mean of 0.993, with a range of [0.807, 1.275]. The absence of a bimodal cluster rules out the presence of a subpopulation of memorised copies hidden within the overall distribution. Per-image LPIPS distance to the nearest training neighbour was 0.266, and per-image SSIM to the nearest training neighbour was 0.629. Both values confirm that generated images are perceptually distinct from their closest real counterparts. For comparison, actual autoencoder reconstructions achieve an SSIM of 0.986 and LPIPS of 0.017 against the same originals, placing the generated-to-nearest-neighbour distances firmly in the “different images sharing anatomical characteristics” range rather than the “near-copies” range. The severity classification places all 880 samples outside the instance-memorisation band (Table 3).

3.5. Clinical Utility

A ResNet-18 classifier trained exclusively on synthetic data and evaluated on the 44 real test subjects achieved a subject-level AUC of 0.754 (95% CI 0.60 to 0.91), with each subject’s score computed as the mean predicted probability across its twenty slices. The corresponding Train-on-Real-Test-on-Real upper bound was 0.810 (95% CI 0.67 to 0.95), yielding a gap of 5.6 percentage points. The synthetic data therefore retains approximately 93% of the discriminative signal present in real data. Both AUCs exceed the 0.5 chance level, but the confidence intervals are wide and overlap, so the 5.6-percentage-point gap is not statistically distinguishable at this test-set size (n = 44).
Validation AUC during TSTR training reached 0.841, close to the 0.864 obtained under TRTR, so the synthetic training distribution supported classifier learning comparable to that from real data. Slice-level AUC values of 0.720 for TSTR and 0.750 for TRTR are reported for completeness, though subject-level evaluation is the valid metric following the Yagis et al. [21] correction for slice-level inflation. The per-class KID results (AD 0.023 versus CN 0.037) provide complementary evidence that the model learned disease-specific features, as the lower AD KID indicates that the prominent ventricular enlargement characteristic of Alzheimer’s disease is captured more faithfully than the subtler inter-subject variation present in the CN group.

3.6. Augmentation Ablation

To distinguish preprocessing-driven quality from augmentation-driven quality, the LDM was retrained from scratch with enhanced online augmentation comprising random rotation (±3°) and Gaussian intensity jitter (σ = 0.03), in addition to the horizontal flip already present in the base training. All other hyperparameters remained identical. The retrained model converged to an essentially identical validation loss of 0.0558, compared to 0.0559 for the base model, and reached its best epoch earlier (epoch 69 versus 95).
Every distributional quality metric degraded (Table 4). KID increased from 0.030 ± 0.002 to 0.056 ± 0.001. FID∞ rose from 42.14 to 67.85. Precision collapsed from 0.269 to 0.060, a 4.5-fold deterioration. Recall fell from 0.414 to 0.211.
Clinical utility was unaffected. TSTR AUC was 0.748 versus 0.754 for the base model, with the TSTR gap remaining at 5.6 percentage points in both conditions. Memorisation was similarly unchanged, moving from 57.6% to 59.4% , confirming that augmentation did not reduce memorisation because there was no memorisation problem to address.

4. Discussion

The comparison between the baseline and enhanced pipelines supports the central finding of this study, subject to an important limit on what it can isolate. Both configurations were trained on the same 295 ADNI subjects, but five elements changed together between them. Three concerns the pipeline and its training objective: the enhanced version added MNI152 registration and anatomically guided variable-density slice extraction, and it raised the perceptual loss weight from 0.001 to 0.5. Two concerns representational capacity: the latent grew from four to eight channels, and the input resolution increased from 224 × 224 to 256 × 256. The baseline produced images with limited anatomical realism and a raw FID of 42.87, whereas the enhanced pipeline reached a KID of 0.030 ± 0.002, an FID∞ of 42.14, and a TSTR AUC of 0.754. Because these factors moved simultaneously, the contrast measures the combined effect of a pipeline redesign rather than the effect of preprocessing in isolation. Two considerations bound this preprocessing-versus-capacity confound. FID at this sample size does not reliably rank the two configurations (Section 3.3), so the case for the redesign rests on the qualitative shift from unrealistic to anatomically plausible output and on the enhanced pipeline’s absolute distributional and downstream metrics, not on a head-to-head score. The augmentation ablation (Section 3.6) is the controlled experiment in this study, holding the enhanced architecture and the 295-subject cohort fixed while varying only training-time augmentation. Registration remains the most plausible single contributor on mechanistic grounds, because without spatial alignment, anatomical structures occupy different voxel positions across subjects, forcing the diffusion model to average over misaligned features and blurring its outputs; separating its contribution from the concurrent increase in capacity would require a full factorial design.
The augmentation ablation strengthens this interpretation. Retraining the LDM with enhanced augmentation (±3° rotation and Gaussian intensity jitter with σ = 0.03) degraded every distributional metric while leaving the TSTR gap unchanged at 5.6 percentage points. The mechanism behind this dissociation is that the score-matching training loss is augmentation-invariant (validation loss converged to 0.0558 versus 0.0559 for the base model), meaning the noise prediction task is equally learnable from slightly rotated or jittered latents. The distributional degradation arises because the model generates from a wider augmented distribution at inference, pushing samples outside the tight statistical manifold defined by real brain MRI. Clinical utility is preserved because the classification boundary depends on gross anatomical features (ventricular size, hippocampal volume, cortical thickness) that are robust to mild spatial and intensity perturbations. This confirms that the base model’s capacity was already well matched to the 4,080-slice training set, consistent with Karras et al. [7], and that the quality gains between v1 and v2 stem from the redesigned pipeline rather than from insufficient data diversity.
The TSTR result requires careful interpretation. A 93% retention of discriminative signal from 295 subjects is a strong outcome, though estimated from only 44 test subjects with correspondingly wide confidence intervals, and TSTR measures clinical utility rather than anatomical realism. A model could theoretically score well on TSTR by preserving gross ventricular differences between AD and CN while generating anatomically implausible cortical detail, because the classifier focuses on the most discriminative features rather than overall image quality. The combination of TSTR with KID, precision and recall, visual inspection, and memorisation analysis provides substantially stronger evidence than any individual metric, consistent with the multi-metric evaluation framework recommended by Deo et al. [35]. The autoencoder quality ceiling (SSIM 0.9861) is not fully exploited by the diffusion model, as the precision of 0.269 suggests that fine anatomical details are partially smoothed during generation. This gap indicates room for improvement in the generative stage, likely through additional training data rather than architectural changes.
Direct comparison with published brain MRI synthesis studies is complicated by differences in cohort size, evaluation methodology, feature extractors, and preprocessing. Müller-Franzes et al. [11] reported Medfusion FID values between 11.63 and 30.03 across three non-brain modalities with datasets of 19,958 to 223,414 images. The FID∞ of 42.14 achieved here on 880 test images from 295 subjects is higher than expected, given the substantially smaller training set, but still falls within a reasonable range for the data scale. Medfusion’s precision ranged from 0.66 to 0.70 across those datasets, compared to 0.269 in this project, reflecting the more constrained generative capacity available from 4,080 training slices and the 32 × 32 spatial bottleneck. Dhinagar et al. [13] demonstrated a 3-percentage-point AUC improvement when augmenting AD classification with LDM-generated synthetic data from 4,098 ADNI scans, approximately 14 times the cohort available here. The standalone TSTR AUC of 0.754 compares favourably, retaining 93% of real-data discriminative signal without any real-data augmentation. Dar et al. [29] reported an average memorisation rate of 37.2% across medical imaging datasets using a contrastive copy detection framework. This project found zero instances of memorisation using a nearest-neighbour ratio metric with set-size controls, though the two studies employ different detection methodologies and definitions of memorisation, making direct numerical comparison inappropriate. The comparison spans cohorts from 295 to 31,740 subjects under differing architectures and protocols, and only this study and Dhinagar et al. [13] condition generation on diagnosis (Table 5).
The memorisation analysis produced a finding with implications beyond this specific pipeline. The Scardace et al. [18] nearest-neighbour ratio metric, applied without correction using unequal reference sets (4,080 versus 880), flagged 90.3% of generated images. The control experiment on real test images under equal-sized reference conditions revealed a baseline flagging rate of 38.5%, indicating a 32.7 percentage point inflation from set-size imbalance. After correction, the generated flagging rate dropped to 57.6%, a gap of 19.1 percentage points above the control. Zero images met the threshold for instance memorisation (ratio below 0.5), and the minimum observed ratio of 0.807 falls far above the 0.05 to 0.30 range characteristic of genuine pixel-level copying. The metric cannot distinguish distributional proximity from instance memorisation without severity analysis, and the recommended three-part reporting protocol (equal-sized reference sets, real-image control baseline, and severity grading rather than headline flagging percentage) is straightforward to implement and should be standard practice for small-cohort evaluation.
Six limitations constrain the interpretation of these results. The two-dimensional slice-based generation approach produces twenty independent coronal slices per subject without enforcing volumetric consistency, meaning a generated subject may exhibit inconsistent anatomy across adjacent slices and limiting clinical application to per-slice analysis. The cohort of 295 subjects operates at the lower bound for diffusion model training, and Bonnaire et al. [36] showed that the memorisation onset timescale scales linearly with dataset size, placing this cohort in an elevated-risk regime. The empirical evidence (zero severe cases) suggests the risk did not materialise with the chosen architecture, but this cannot be assumed for larger or differently structured models. The conditioning is limited to a binary AD versus CN distinction, whereas real diagnostic utility would require finer-grained conditioning on clinical variables such as CDR score, hippocampal volume, or disease duration. The ADNI cohort over-represents white, educated North American participants, and results may not generalise to other populations or acquisition protocols. Inception v3 features, used for KID and FID computation, were trained on real images and may not capture all diagnostically relevant features of brain MRI. There was no domain-specific feature extractor, such as RadImageNet [37], which was evaluated. No radiologist reader study was conducted, meaning all visual assessments were performed by the researcher rather than a clinical expert, limiting the strength of anatomical realism claims. The v1-to-v2 comparison changed five elements simultaneously: registration, slice selection, and LPIPS weight, along with an increase in latent channels from 4 to 8 and in input resolution from 224 × 224 to 256 × 256. The design therefore cannot separate the preprocessing changes from the concurrent increase in representational capacity. The augmentation ablation provides partial evidence that training-side interventions were not the driver, but a full factorial design varying each factor individually would be needed to attribute the improvement to any single change.
Future work should pursue five directions. Extending to three-dimensional volumetric generation using 3D latent diffusion models [10,12] would enforce inter-slice consistency, enabling volumetric clinical assessment. Replacing the DDPM formulation with flow matching [38] would reduce inference time by five- to tenfold while maintaining equivalent generation quality, representing the highest-value architectural improvement for practical deployment. Finer-grained conditioning on continuous clinical variables would produce a severity spectrum rather than a binary distinction, and the existing cross-attention mechanism can accommodate this with minimal architectural modifications. Applying the preprocessing pipeline to non-ADNI datasets such as OASIS and AIBL would test the generalisability of both the methodology and the preprocessing-matters finding across different acquisition protocols and demographics. Adopting FID∞ extrapolation as standard practice and conducting formal radiologist reader studies would strengthen evaluation methodology for small-cohort synthesis more broadly.

5. Conclusions

This study developed and evaluated a two-stage conditional latent diffusion model pipeline for generating synthetic 2D coronal brain MRI conditioned on Alzheimer’s disease diagnosis, trained on data from 295 subjects of the ADNI database. The central experiment applied two pipeline configurations to the same subjects, and the augmentation ablation held the architecture and cohort fixed to isolate training-side effects. Together, these show that the quality gains follow from the redesigned, anatomically informed pipeline rather than from increased data diversity, a comparison and ablation pairing that is absent from the existing synthesis literature.
The enhanced pipeline, incorporating MNI152 registration, variable-density anatomical slice selection informed by Braak-staging neuropathology, and literature-informed loss configuration, achieved a KID of 0.030 ± 0.002 and an FID∞ of 42.14, producing synthetic images that capture class-characteristic gross morphology from a cohort approximately one fourteenth the size used by the nearest comparable study. A TSTR evaluation demonstrated that a classifier trained exclusively on synthetic data retains 93% of real-data discriminative signal (AUC 0.754 versus 0.810), establishing clinical utility at a small-cohort scale. An augmentation ablation confirmed that additional data augmentation degraded distributional quality (KID from 0.030 to 0.056, precision from 0.269 to 0.060) while leaving clinical utility unchanged (TSTR gap 5.6 percentage points in both conditions), providing evidence that the quality gains follow from the redesigned pipeline rather than from training-side interventions.
The memorisation analysis found zero instances of memorisation across 880 generated samples and identified a 32.7-percentage-point set-size confound in the nearest-neighbour ratio metric. The corrected three-part reporting protocol (equal-sized reference sets, real-image control baseline, and severity grading) provides a practical contribution to evaluation methodology that extends beyond this specific pipeline.
These results demonstrate that rigorous, anatomically informed preprocessing can partially compensate for small cohort size in brain MRI synthesis, and that preprocessing methodology deserves the same documentation and experimental attention as model architecture. Extending this work to three-dimensional volumetric generation, replacing the DDPM formulation with flow matching for faster inference, and validating the preprocessing-matters finding across non-ADNI cohorts represent the most consequential next steps toward clinical deployment of synthetic brain MRI for Alzheimer’s disease research.

Author Contributions

Conceptualization, S.F. and N.H.; methodology, S.F.; software, S.F.; validation, S.F. and N.H.; formal analysis, S.F.; investigation, S.F.; data curation, S.F.; writing—original draft preparation, S.F.; writing—review and editing, S.F. and N.H.; visualization, S.F.; supervision, N.H.; project administration, S.F. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

This study used de-identified data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI), collected under Institutional Review Board approval at each participating ADNI site. Ethical approval for this secondary analysis was granted by Northumbria University.

Data Availability Statement

The imaging data analysed in this study were obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI; https://adni.loni.usc.edu) and are governed by the ADNI Data Use Agreement. They cannot be redistributed by the authors and are available to qualified researchers on application through ADNI. The analysis code and trained model weights are available at https://github.com/soheilfallah/brain-mri-ad-synthesis and will be released publicly upon acceptance. The repository contains no ADNI participant data or participant-derived outputs, consistent with the ADNI Data Use Agreement.

Acknowledgments

The authors thank the Alzheimer’s Disease Neuroimaging Initiative (ADNI) investigators and all ADNI study participants. Data collection and sharing for the Alzheimer’s Disease Neuroimaging Initiative (ADNI) is funded by the National Institute on Aging (National Institutes of Health Grant U19AG024904). The grantee organization is the Northern California Institute for Research and Education. In the past, ADNI has also received funding from the National Institute of Biomedical Imaging and Bioengineering, the Canadian Institutes of Health Research, and private sector contributions through the Foundation for the National Institutes of Health (FNIH) including generous contributions from the following: AbbVie, Alzheimer’s Association; Alzheimer’s Drug Discovery Foundation; Araclon Biotech; BioClinica, Inc.; Biogen; Bristol-Myers Squibb Company; CereSpir, Inc.; Cogstate; Eisai Inc.; Elan Pharmaceuticals, Inc.; Eli Lilly and Company; EuroImmun; F. Hoffmann-La Roche Ltd and its affiliated company Genentech, Inc.; Fujirebio; GE Healthcare; IXICO Ltd.; Janssen Alzheimer Immunotherapy Research & Development, LLC.; Johnson & Johnson Pharmaceutical Research & Development LLC.; Lumosity; Lundbeck; Merck & Co., Inc.; Meso Scale Diagnostics, LLC.; NeuroRx Research; Neurotrack Technologies; Novartis Pharmaceuticals Corporation; Pfizer Inc.; Piramal Imaging; Servier; Takeda Pharmaceutical Company; and Transition Therapeutics. A complete listing of ADNI investigators is available at http://adni.loni.usc.edu/wp-content/uploads/how_to_apply/ADNI_Acknowledgement_List.pdf. During the preparation of this manuscript, the author used Claude (Anthropic, claude.ai) as a coding assistant for Python/PyTorch/MONAI implementation and debugging of pipeline components. No ADNI participant-level data, derived participant-level data, or ADNI-identifiable content was transmitted to or through the tool. The authors have reviewed and edited all tool output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AD Alzheimer’s Disease
ADNI Alzheimer’s Disease Neuroimaging Initiative
ANTs Advanced Normalization Tools
AUC Area Under the Receiver Operating Characteristic Curve
CDR Clinical Dementia Rating
CI Confidence Interval
CN Cognitively Normal
DDIM Denoising Diffusion Implicit Models
DDPM Denoising Diffusion Probabilistic Models
DINOv2 Self-Distillation with No Labels, version 2
FID Fréchet Inception Distance
GAN Generative Adversarial Network
HD-BET High-Definition Brain Extraction Tool
KID Kernel Inception Distance
KL Kullback–Leibler
LDM Latent Diffusion Model
LPIPS Learned Perceptual Image Patch Similarity
MNI Montreal Neurological Institute
MRI Magnetic Resonance Imaging
PSNR Peak Signal-to-Noise Ratio
SSIM Structural Similarity Index Measure
TSTR Train on Synthetic, Test on Real
TRTR Train on Real, Test on Real
VAE Variational Autoencoder
VAEGAN Variational Autoencoder with Generative Adversarial Network

References

  1. Knopman, D.S.; Amieva, H.; Petersen, R.C.; Chételat, G.; Holtzman, D.M.; Hyman, B.T.; Nixon, R.A.; Jones, D.T. Alzheimer Disease. Nat. Rev. Dis. Prim. 2021, 7, 33. [Google Scholar] [CrossRef] [PubMed]
  2. DeTure, M.A.; Dickson, D.W. The Neuropathological Diagnosis of Alzheimer’s Disease. Mol. Neurodegener. 2019, 14, 32. [Google Scholar] [CrossRef] [PubMed]
  3. Grueso, S.; Viejo-Sobera, R. Machine Learning Methods for Predicting Progression from Mild Cognitive Impairment to Alzheimer’s Disease Dementia: A Systematic Review. Alz Res. Ther. 2021, 13, 162. [Google Scholar] [CrossRef] [PubMed]
  4. Herzog, N.J.; Magoulas, G.D. Brain Asymmetry Detection and Machine Learning Classification for Diagnosis of Early Dementia. Sensors 2021, 21, 778. [Google Scholar] [CrossRef] [PubMed]
  5. Kushol, R.; Parnianpour, P.; Wilman, A.H.; Kalra, S.; Yang, Y.-H. Effects of MRI Scanner Manufacturers in Classification Tasks with Deep Learning Models. Sci. Rep. 2023, 13, 16791. [Google Scholar] [CrossRef] [PubMed]
  6. Kazerouni, A.; Aghdam, E.K.; Heidari, M.; Azad, R.; Fayyaz, M.; Hacihaliloglu, I.; Merhof, D. Diffusion Models in Medical Imaging: A Comprehensive Survey. Med. Image Anal. 2023, 88, 102846. [Google Scholar] [CrossRef] [PubMed]
  7. Karras, T.; Aittala, M.; Hellsten, J.; Laine, S.; Lehtinen, J.; Aila, T. Training Generative Adversarial Networks with Limited Data. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc., 2020; Vol. 33, pp. 12104–12114. [Google Scholar]
  8. Ho, J.; Jain, A.; Abbeel, P. Denoising Diffusion Probabilistic Models. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc., 2020; pp. 6840–6851. [Google Scholar]
  9. Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New Orleans, LA, USA, June 2022; pp. 10674–10685. [Google Scholar]
  10. Pinaya, W.H.L.; Tudosiu, P.-D.; Dafflon, J.; Da Costa, P.F.; Fernandez, V.; Nachev, P.; Ourselin, S.; Cardoso, M.J. Brain Imaging Generation with Latent Diffusion Models. In Proceedings of the Deep Generative Models: Second MICCAI Workshop, DGM4MICCAI 2022, Held in Conjunction with MICCAI 2022, Singapore, September 22, 2022, Proceedings, September 22 2022; Springer-Verlag: Berlin, Heidelberg; pp. 117–126. [Google Scholar]
  11. Müller-Franzes, G.; Niehues, J.M.; Khader, F.; Arasteh, S.T.; Haarburger, C.; Kuhl, C.; Wang, T.; Han, T.; Nolte, T.; Nebelung, S.; et al. A Multimodal Comparison of Latent Denoising Diffusion Probabilistic Models and Generative Adversarial Networks for Medical Image Synthesis. Sci. Rep. 2023, 13, 12098. [Google Scholar] [CrossRef] [PubMed]
  12. Khader, F.; Müller-Franzes, G.; Tayebi Arasteh, S.; Han, T.; Haarburger, C.; Schulze-Hagen, M.; Schad, P.; Engelhardt, S.; Baeßler, B.; Foersch, S.; et al. Denoising Diffusion Probabilistic Models for 3D Medical Image Generation. Sci. Rep. 2023, 13, 7303. [Google Scholar] [CrossRef] [PubMed]
  13. Dhinagar, N.J.; Thomopoulos, S.I.; Laltoo, E.; Thompson, P.M. Counterfactual MRI Generation with Denoising Diffusion Models for Interpretable Alzheimer’s Disease Effect Detection. In Proceedings of the 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC); IEEE: Orlando, FL, USA, 15 July 2024; pp. 1–6. [Google Scholar]
  14. Mendoza-Léon, R.; Puentes, J.; Uriza, L.F.; Hernández Hoyos, M. Single-Slice Alzheimer’s Disease Classification and Disease Regional Analysis with Supervised Switching Autoencoders. Comput. Biol. Med. 2020, 116, 103527. [Google Scholar] [CrossRef] [PubMed]
  15. Chen, L.; Qiao, H.; Zhu, F. Alzheimer’s Disease Diagnosis With Brain Structural MRI Using Multiview-Slice Attention and 3D Convolution Neural Network. Front. Aging Neurosci. 2022, 14, 871706. [Google Scholar] [CrossRef] [PubMed]
  16. Larsen, A.B.L.; Sønderby, S.K.; Larochelle, H.; Winther, O. Autoencoding beyond Pixels Using a Learned Similarity Metric. In Proceedings of the 33rd International Conference on Machine Learning; PMLR, June 11 2016; pp. 1558–1566. [Google Scholar]
  17. Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2018; pp. 586–595. [Google Scholar]
  18. Scardace, A.; Puglisi, L.; Guarnera, F.; Battiato, S.; Ravì, D. A Novel Metric for Detecting Memorization in Generative Models for Brain MRI Synthesis. In Proceedings of the 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), March 2026; pp. 3868–3877. [Google Scholar]
  19. Chong, M.J.; Forsyth, D. Effectively Unbiased FID and Inception Score and Where to Find Them. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Seattle, WA, USA, June 2020; pp. 6069–6078. [Google Scholar]
  20. Bińkowski, M.; Sutherland, D.J.; Arbel, M.; Gretton, A. Demystifying MMD GANs. In Proceedings of the International Conference on Learning Representations (ICLR), 2018. [Google Scholar]
  21. Yagis, E.; Atnafu, S.W.; García Seco De Herrera, A.; Marzi, C.; Scheda, R.; Giannelli, M.; Tessa, C.; Citi, L.; Diciotti, S. Effect of Data Leakage in Brain MRI Classification Using 2D Convolutional Neural Networks. Sci. Rep. 2021, 11, 22544. [Google Scholar] [CrossRef] [PubMed]
  22. Alzheimer’s Disease Neuroimaging Initiative. Available online: https://adni-lde.loni.usc.edu/ (accessed on 16 July 2026).
  23. Isensee, F.; Schell, M.; Pflueger, I.; Brugnara, G.; Bonekamp, D.; Neuberger, U.; Wick, A.; Schlemmer, H.; Heiland, S.; Wick, W.; et al. Automated Brain Extraction of Multisequence MRI Using Artificial Neural Networks. Hum. Brain Mapp. 2019, 40, 4952–4964. [Google Scholar] [CrossRef] [PubMed]
  24. Fonov, V.; Evans, A.C.; Botteron, K.; Almli, C.R.; McKinstry, R.C.; Collins, D.L. Unbiased Average Age-Appropriate Atlases for Pediatric Studies. NeuroImage 2011, 54, 313–327. [Google Scholar] [CrossRef]
  25. Braak, H.; Braak, E. Neuropathological Stageing of Alzheimer-Related Changes. Acta Neuropathol. 1991, 82, 239–259. [Google Scholar] [CrossRef] [PubMed]
  26. Peng, W.; Adeli, E.; Bosschieter, T.; Park, S.H.; Zhao, Q.; Pohl, K.M. Generating Realistic Brain MRIs via a Conditional Diffusion Probabilistic Model. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2023; Lecture Notes in Computer Science; Greenspan, H., Madabhushi, A., Mousavi, P., Salcudean, S., Duncan, J., Syeda-Mahmood, T., Taylor, R., Eds.; Springer Nature Switzerland: Cham, 2023; Vol. 14227, pp. 14–24. ISBN 978-3-031-43992-6. [Google Scholar]
  27. Cardoso, M.J.; Li, W.; Brown, R.; Ma, N.; Kerfoot, E.; Wang, Y.; Murrey, B.; Myronenko, A.; Zhao, C.; Yang, D.; et al. MONAI: An Open-Source Framework for Deep Learning in Healthcare. arXiv 2022. [Google Scholar] [CrossRef]
  28. Kouzelis, T.; Kakogeorgiou, I.; Gidaris, S.; Komodakis, N. EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling. In Proceedings of the 42nd International Conference on Machine Learning, 2025; PMLR: Vancouver, Canada; Vol. 267. [Google Scholar]
  29. Dar, S.U.H.; Seyfarth, M.; Ayx, I.; Papavassiliu, T.; Schoenberg, S.O.; Siepmann, R.M.; Laqua, F.C.; Kahmann, J.; Frey, N.; Baeßler, B.; et al. Unconditional Latent Diffusion Models Memorize Patient Imaging Data. Nat. Biomed. Eng. 2025. [Google Scholar] [CrossRef] [PubMed]
  30. Ho, J.; Salimans, T. Classifier-Free Diffusion Guidance. arXiv 2022. [Google Scholar] [CrossRef]
  31. Kynkäänniemi, T.; Karras, T.; Laine, S.; Lehtinen, J.; Aila, T. Improved Precision and Recall Metric for Assessing Generative Models. In Proceedings of the Advances in Neural Information Processing Systems; Wallach, H., Larochelle, H., Beygelzimer, A., Alché-Buc, F. d’, Fox, E., Garnett, R., Eds.; Curran Associates, Inc., 2019; Vol. 32. [Google Scholar]
  32. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Las Vegas, NV, USA, June 2016; pp. 770–778. [Google Scholar]
  33. Hanley, J.A.; McNeil, B.J. The Meaning and Use of the Area under a Receiver Operating Characteristic (ROC) Curve. Radiology 1982, 143, 29–36. [Google Scholar] [CrossRef] [PubMed]
  34. Hajian-Tilaki, K. Receiver Operating Characteristic (ROC) Curve Analysis for Medical Diagnostic Test Evaluation. Casp. J. Intern Med. 2013, 4, 627–635. [Google Scholar]
  35. Deo, Y.; Jia, Y.; Lassila, T.; Smith, W.A.P.; Lawton, T.; Kang, S.; Frangi, A.F.; Habli, I. Metrics That Matter: Evaluating Image Quality Metrics for Medical Image Generation. arXiv 2025. [Google Scholar] [CrossRef]
  36. Bonnaire, T.; Urfin, R.; Biroli, G.; Mézard, M. Why Diffusion Models Don’t Memorize: The Role of Implicit Dynamical Regularization in Training. arXiv 2025. [Google Scholar] [CrossRef]
  37. Mei, X.; Liu, Z.; Robson, P.M.; Marinelli, B.; Huang, M.; Doshi, A.; Jacobi, A.; Cao, C.; Link, K.E.; Yang, T.; et al. RadImageNet: An Open Radiologic Deep Learning Research Dataset for Effective Transfer Learning. Radiol. Artif. Intell. 2022, 4, e210315. [Google Scholar] [CrossRef]
  38. Lipman, Y.; Chen, R.T.Q.; Ben-Hamu, H.; Nickel, M.; Le, M. Flow Matching for Generative Modeling. arXiv 2023. [Google Scholar] [CrossRef]
Figure 1. Six-stage preprocessing pipeline converting raw ADNI acquisitions into model-ready coronal slices. Grey boxes denote processing stages, each annotated with its input, process, and output; green diamonds (CP1–CP6) denote checkpoint gates that every subject must satisfy before the next stage proceeds. The pipeline reduces 6,169 preprocessing rows from 603 subjects to one volume per subject (295 subjects; 118 AD, 177 CN), applies HD-BET brain extraction and affine registration to MNI152 space (12 degrees of freedom; 182 × 218 × 182), extracts 20 anatomically guided coronal slices per subject, and produces a subject-level stratified 70/15/15 split of 5,900 slices (training, 4,080; validation, 940; test, 880). Variant selection (Stage 1) reduces the 6,169 available preprocessing rows to one volume per subject using the priority hierarchy described in Section 2.1. The resulting 295 T1-weighted volumes share a consistent preprocessing provenance, with 97.6% receiving the highest-priority combination of corrections.
Figure 1. Six-stage preprocessing pipeline converting raw ADNI acquisitions into model-ready coronal slices. Grey boxes denote processing stages, each annotated with its input, process, and output; green diamonds (CP1–CP6) denote checkpoint gates that every subject must satisfy before the next stage proceeds. The pipeline reduces 6,169 preprocessing rows from 603 subjects to one volume per subject (295 subjects; 118 AD, 177 CN), applies HD-BET brain extraction and affine registration to MNI152 space (12 degrees of freedom; 182 × 218 × 182), extracts 20 anatomically guided coronal slices per subject, and produces a subject-level stratified 70/15/15 split of 5,900 slices (training, 4,080; validation, 940; test, 880). Variant selection (Stage 1) reduces the 6,169 available preprocessing rows to one volume per subject using the priority hierarchy described in Section 2.1. The resulting 295 T1-weighted volumes share a consistent preprocessing provenance, with 97.6% receiving the highest-priority combination of corrections.
Preprints 227206 g001
Figure 2. Two-stage conditional latent diffusion pipeline. In Stage 1 (perceptual compression, trained first), an encoder E maps an input image x to a compact latent representation z of size 32 × 32 × 8. In Stage 2 (latent diffusion, trained second), a forward process adds noise to the latent over T steps to reach z<sub>T</sub>, and a UNet learns the reverse process, denoising over T steps to recover while conditioning on the diagnostic class label through cross-attention. A decoder D reconstructs the denoised latent into the output image . The dashed box marks the shared latent space linking the two stages.
Figure 2. Two-stage conditional latent diffusion pipeline. In Stage 1 (perceptual compression, trained first), an encoder E maps an input image x to a compact latent representation z of size 32 × 32 × 8. In Stage 2 (latent diffusion, trained second), a forward process adds noise to the latent over T steps to reach z<sub>T</sub>, and a UNet learns the reverse process, denoising over T steps to recover while conditioning on the diagnostic class label through cross-attention. A decoder D reconstructs the denoised latent into the output image . The dashed box marks the shared latent space linking the two stages.
Preprints 227206 g002
Figure 3. Autoencoder reconstruction quality on held-out test slices. Each column shows one example mid-brain slice at the hippocampal level, giving two Alzheimer’s disease (AD) cases and two cognitively normal (CN) cases. Rows show the original (top), reconstruction (middle), and absolute per-pixel difference (bottom), with difference maps windowed to [0, 0.10] in normalised-intensity units. Each column header reports the peak signal-to-noise ratio (PSNR) of that slice. Over the full test set (n = 880 slices), the autoencoder reached a mean PSNR of 38.66 dB and a mean structural similarity (SSIM) of 0.989, and reconstruction error stayed low for both classes with no visible class-dependent bias.
Figure 3. Autoencoder reconstruction quality on held-out test slices. Each column shows one example mid-brain slice at the hippocampal level, giving two Alzheimer’s disease (AD) cases and two cognitively normal (CN) cases. Rows show the original (top), reconstruction (middle), and absolute per-pixel difference (bottom), with difference maps windowed to [0, 0.10] in normalised-intensity units. Each column header reports the peak signal-to-noise ratio (PSNR) of that slice. Over the full test set (n = 880 slices), the autoencoder reached a mean PSNR of 38.66 dB and a mean structural similarity (SSIM) of 0.989, and reconstruction error stayed low for both classes with no visible class-dependent bias.
Preprints 227206 g003
Figure 4. Real and synthetic coronal brain MRI slices by diagnostic class. (a) Alzheimer’s disease (AD); (b) cognitively normal (CN). In each panel, the top row shows real held-out test slices, and the bottom row shows synthetic slices from the enhanced pipeline, selected as median-typical generations by brain-area fraction rather than for visual quality. Synthetic images reproduce class-characteristic gross morphology, with enlarged ventricles and widened sulci in AD and compact ventricles in CN, while fine cortical detail is resolved less sharply than in the real slices.
Figure 4. Real and synthetic coronal brain MRI slices by diagnostic class. (a) Alzheimer’s disease (AD); (b) cognitively normal (CN). In each panel, the top row shows real held-out test slices, and the bottom row shows synthetic slices from the enhanced pipeline, selected as median-typical generations by brain-area fraction rather than for visual quality. Synthetic images reproduce class-characteristic gross morphology, with enlarged ventricles and widened sulci in AD and compact ventricles in CN, while fine cortical detail is resolved less sharply than in the real slices.
Preprints 227206 g004
Figure 5. Bias-corrected FID via sample-size extrapolation. Fréchet Inception Distance (FID) was computed at six sample sizes (N = 200, 350, 500, 650, 800, 880) and regressed against 1/N following the extrapolation method of Chong and Forsyth [19]; the intercept at 1/N → 0 gives the bias-free estimate FID∞. Blue circles show the measured FID at each N, the grey line is the linear fit (R² = 0.988), and the orange diamond marks FID∞ = 42.14. The raw FID at the largest sample (N = 880) was 46.23, so the extrapolation removes about 4.1 points of upward small-sample bias.
Figure 5. Bias-corrected FID via sample-size extrapolation. Fréchet Inception Distance (FID) was computed at six sample sizes (N = 200, 350, 500, 650, 800, 880) and regressed against 1/N following the extrapolation method of Chong and Forsyth [19]; the intercept at 1/N → 0 gives the bias-free estimate FID∞. Blue circles show the measured FID at each N, the grey line is the linear fit (R² = 0.988), and the orange diamond marks FID∞ = 42.14. The raw FID at the largest sample (N = 880) was 46.23, so the extrapolation removes about 4.1 points of upward small-sample bias.
Preprints 227206 g005
Figure 6. Nearest-neighbour distance-ratio distributions for the memorisation control analysis. For each image, the ratio is the distance to its nearest training neighbour divided by the distance to its nearest held-out test neighbour in the DINOv2 feature space [18]; a ratio below 1 (dashed line) indicates an image lying closer to the training set than to held-out data. The grey distribution is a size-matched real-image control (n = 880) that establishes the expected baseline, and the blue outline is the generated set (n = 880). The generated distribution is shifted toward lower ratios: 57.6% fall below 1, compared with 38.5% for the control, a gap of 19.1 percentage points. No sample fell below 0.5, and no instance-level match was found, indicating distributional proximity to the training data without verbatim copying.
Figure 6. Nearest-neighbour distance-ratio distributions for the memorisation control analysis. For each image, the ratio is the distance to its nearest training neighbour divided by the distance to its nearest held-out test neighbour in the DINOv2 feature space [18]; a ratio below 1 (dashed line) indicates an image lying closer to the training set than to held-out data. The grey distribution is a size-matched real-image control (n = 880) that establishes the expected baseline, and the blue outline is the generated set (n = 880). The generated distribution is shifted toward lower ratios: 57.6% fall below 1, compared with 38.5% for the control, a gap of 19.1 percentage points. No sample fell below 0.5, and no instance-level match was found, indicating distributional proximity to the training data without verbatim copying.
Preprints 227206 g006
Table 1. Cohort demographics and data split.
Table 1. Cohort demographics and data split.
Split Subjects AD CN %AD Slices
Train 204 76 128 37.3 4,080
Val 47 24 23 51.1 940
Test 44 18 26 40.9 880
Total 295 118 177 40 5,900
Table 2. Distributional generation-quality metrics for the enhanced pipeline, computed on 880 generated images against the held-out test set using Inception v3 features. KID and FID are lower-is-better; precision and recall are higher-is-better. KID is reported as the mean ± standard deviation over 100 random subsets. FID∞ is the bias-corrected value obtained by extrapolating FID to infinite sample size [19] (Figure 5).
Table 2. Distributional generation-quality metrics for the enhanced pipeline, computed on 880 generated images against the held-out test set using Inception v3 features. KID and FID are lower-is-better; precision and recall are higher-is-better. KID is reported as the mean ± standard deviation over 100 random subsets. FID∞ is the bias-corrected value obtained by extrapolating FID to infinite sample size [19] (Figure 5).
Metric Value
Kernel Inception Distance (KID), overall 0.030 ± 0.002
KID, AD 0.023
KID, CN 0.037
Fréchet Inception Distance (FID), raw (N = 880) 46.23
FID∞ (bias-corrected) 42.14
Precision 0.269
Recall 0.414
Table 3. Three-way memorisation severity classification of all 880 generated samples, graded by the nearest-neighbour distance ratio (distance to nearest training image divided by distance to nearest held-out test image) in DINOv2 feature space. A ratio below 0.5 denotes instance memorisation (near-copies of training data), a ratio from 0.5 to below 1.0 denotes distributional proximity (closer to training than to held-out data, expected for a well-trained model), and a ratio of 1.0 or above denotes distributional generalisation.
Table 3. Three-way memorisation severity classification of all 880 generated samples, graded by the nearest-neighbour distance ratio (distance to nearest training image divided by distance to nearest held-out test image) in DINOv2 feature space. A ratio below 0.5 denotes instance memorisation (near-copies of training data), a ratio from 0.5 to below 1.0 denotes distributional proximity (closer to training than to held-out data, expected for a well-trained model), and a ratio of 1.0 or above denotes distributional generalisation.
Category Ratio range Count Percentage
Instance memorisation < 0.5 0 / 880 0.0
Distributional proximity 0.5 to < 1.0 507 / 880 57.6
Distributional generalisation ≥ 1.0 373 / 880 42.4
Table 4. Effect of additional data augmentation on the enhanced pipeline. Both configurations use the identical two-stage architecture and the same 295 ADNI subjects; the augmented variant adds ±3° rotation and Gaussian intensity jitter (σ = 0.03) during latent-diffusion training. KID, FID∞, and memorisation flagging are lower-is-better; precision, recall, TSTR AUC, and TRTR AUC are higher-is-better.
Table 4. Effect of additional data augmentation on the enhanced pipeline. Both configurations use the identical two-stage architecture and the same 295 ADNI subjects; the augmented variant adds ±3° rotation and Gaussian intensity jitter (σ = 0.03) during latent-diffusion training. KID, FID∞, and memorisation flagging are lower-is-better; precision, recall, TSTR AUC, and TRTR AUC are higher-is-better.
Metric Enhanced (v2) Enhanced + augmentation
KID (Inception) 0.030 ± 0.002 0.056 ± 0.001
FID∞ 42.14 67.85
Precision 0.269 0.060
Recall 0.414 0.211
TSTR AUC (subject-level) 0.754 0.748
TRTR AUC (subject-level) 0.810 0.803
Memorisation flagged (ratio < 1), % 57.6 59.4
Instance memorisation (ratio < 0.5) 0 / 880 0 / 880
Table 5. Contextualisation of this study against representative diffusion-based medical image synthesis work. The studies differ substantially in cohort size, image dimensionality, resolution, feature extractor, and evaluation protocol; reported values are shown as published and are not directly comparable across rows, so the table situates this study’s scale and scope rather than ranking methods. Among diffusion models applied to brain MRI, only Dhinagar et al. [13] and this study condition generation on diagnosis, and this study operates at the smallest cohort scale.
Table 5. Contextualisation of this study against representative diffusion-based medical image synthesis work. The studies differ substantially in cohort size, image dimensionality, resolution, feature extractor, and evaluation protocol; reported values are shown as published and are not directly comparable across rows, so the table situates this study’s scale and scope rather than ranking methods. Among diffusion models applied to brain MRI, only Dhinagar et al. [13] and this study condition generation on diagnosis, and this study operates at the smallest cohort scale.
Study Cohort Image data Architecture Diagnosis-conditional Reported performance
Pinaya et al. [10] 31,740 subjects (UK Biobank) 3D T1-w, 160×224×160 LDM (VQ-VAE + diffusion) No Ventricular-volume conditioning r = 0.972
Khader et al. [12] 998 scans (ADNI subset) 3D, 64×64×64 VQ-GAN + DDPM No 50/50 ADNI images rated realistic (radiologist)
Müller-Franzes et al. [11] 19,958–223,414 images (non-brain) 2D, 256×256 Medfusion (latent DDPM) No FID 11.63–30.03; precision 0.66–0.70
Dhinagar et al. [13] 1,188 subjects / 4,098 scans (ADNI) 3D T1-w Conditional LDM/DDPM Yes (AD/CN) +3 pp downstream AD-classification AUC
This study 295 subjects / 4,080 slices (ADNI) 2D coronal, 256×256 Class-conditional LDM (Medfusion-style)
Yes (AD/CN) FID∞ 42.14; KID 0.030; TSTR AUC 0.754; 0 instance memorisation
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings