Submitted:
24 September 2026
Posted:
24 September 2026
You are already at the latest version
Abstract
Federated learning (FL) has emerged as a promising novel paradigm for enabling multi-institutional collaborations and expediting data access for training and benchmarking Healthcare AI models, without the need for data centralization. During FL, each collaborator performs local computations on their own data and sends model updates to a central server, which collects and aggregates all model updates into a global consensus model. Several FL aggregation (FLAg) strategies have been developed, each focusing on improving the effectiveness of FL, while addressing distinct technical and data-related challenges – from handling varying resource environments, to patient privacy, and to data imbalance across collaborators. In this article, we provide a systematic survey and taxonomy of FLAg strategies designed for healthcare, paired with their relation to known FL challenges, in our attempt to guide readers navigate and interpret the current scattered body of relevant literature. Our analysis reveals that significant progress has been made in improving model utility and privacy guarantees, while mitigating stragglers, system heterogeneity, fairness and biases, uncertainty, security, and data imbalance. We also touch upon current research gaps between current methodological advances and real-world deployment in healthcare, including limitations in i) evaluation beyond model utility, ii) clinical risk assessment, and iii) governance requirements for transparent and accountable FL.
Keywords:
federated learning
; aggregation
; healthcare
1. Introduction
Deep Learning (DL) has brought significant advantages in medical image analysis, including automated feature extraction, effective handling of large-scale image data, and improved model utility in both classification [1] and segmentation [2] workloads. However, models trained on single institutional data collections often lack generalization, resulting in significant utility variations across multi-institutional data distributions.
Multi-institutional studies in the medical domain [3,4] are facilitated through Centralized Data Sharing (CDS) in which DL models are trained on data pooled in one center (Figure 1a). CDS is hindered by privacy concerns (among others) as medical data must comply with confidentiality and security regulations, i.e., Personal Health information under the General Data Protection Regulation (GDPR) [5] and the Health Insurance Portability and Accountability Act (HIPAA) [6]. Federated Learning (FL) [7,8] enables multi-institutional collaborations while circumventing data sharing (Figure 1b). A central server initializes a global model parameters and broadcasts them to all collaborating institutions. During a training cycle (i.e., a federated round) each institution performs local training on its own data, sends the model updates to the central server, the server collects the updated models from all institutions and aggregates them in a global consensus model. It then broadcasts the updated global model back to all the collaborators. This process is repeated iteratively until model convergence is reached.
FL was made popular in 2017 with the Federated Averaging (FedAvg) algorithm firstly used for auto-complete typing functionality of phones [9] and then facilitating the very first FL study in healthcare [10]. During FedAvg the server aggregates collaborator updates by averaging their model parameters. In their evaluation, both the convex and non-convex problems were shown to converge toward the global minimum as the number of local training epochs and communication rounds increased.
This strategy works remarkably well for independent and identical datasets (IID), as they have a similar objective for all the clients. In medical applications, however, data tends to be non-IID as they are collected from different healthcare institutions that exhibit variations in terms of disease frequency, patient ethnicity, geographic location, gender, and socioeconomic status. This leads to biased models that even further worsen healthcare disparities [11,12]. Thus, fair aggregation strategies are required to ensure the global model neither favors some institutions over others, nor a patient population group dominates the global consensus model.
FL comes with other technical challenges (such as heterogeneous computational resources across collaborators) that can create communication bottlenecks, increase synchronization delays, under constrained network conditions hinder convergence, and potentially lead to model divergence. The most straightforward solution is a wait-free FL framework that allows asynchronous training [13]. Since the primary focus of FL is to incorporate knowledge from institutions with varying levels of computational capability, FLAg strategies should account for the computational load of clients without compromising model fairness, as highlighted in [9,14,15].
Medical datasets are often highly imbalanced and noisy due to variations in annotation quality or protocol, and image resolution. DL systems in healthcare are expected to provide reliable and well-calibrated predictions, rather than solely achieving high predictive utility [11]. Data heterogeneity and communication constraints, can cause clients to complete local training at different rates. In synchronous FL, the server must wait for slower clients, commonly referred to as stragglers, before performing model aggregation, resulting in increased training latency and reduced system efficiency. Therefore, FLAg strategies designed for healthcare applications should not only ensure fair contributions from participating hospitals, but also effectively mitigate the impact of stragglers [9,13,14].
Existing literature [16] has divided FL methods into horizontal, vertical, and federated transfer learning, describing their key differences and features. Kairouz et al. [17] provided a comprehensive theory of effective and efficient implementations of FL, while Rauniyar et al. [18] created a taxonomy-based review of FL applications in medical fields, with a focus on cancer detection and future model development. Pati et al [19] and Raza et al [20] gave an overview of FL applications in healthcare, mainly focusing on identifying threats and their mitigation through varying privacy preserving mechanisms. Other studies focused on FL problems when targeting a particular medical application, such as COVID-19 [21,22,23], cancer diagnostics [24,25,26] or rare diseases [27].
Other recent surveys have discussed various challenges to overcome FL including non-IID, stragglers, privacy, security, data uncertainty, asynchronous training, and a very limited focus given to aggregation strategies available in the current practice. In this review, we provide a comprehensive and structured overview of FLAg strategies and how they are addressing the challenges faced in FL Healthcare ecosystem that relate to either to data or technical issues (Figure 2).
2. Aggregation Strategies
FL emerged as one of the key methods for decentralized model training in privacy sensitive domains such as healthcare. The effectiveness of FL is highly influenced by the aggregation approach, i.e., how the central server combines the client parameters. A wide range of FLAg strategies have been proposed in various domains since 2016. Each of these FLAg strategies was developed to address various challenges such as the presence of system heterogeneity between different institutions, the non-IID nature of the data, and the variation in the client reliability due to resource constraints. For those new to the field, these strategies can be challenging to grasp because they are scattered and inconsistently presented across the literature. There is a need for a clear and structured overview that organizes these approaches into a coherent road map, helping readers navigate and interpret the existing body of work more effectively. This section aims to categorize FLAg strategies according to their underlying strategy (Figure 2), followed by a comparative discussion of their strengths and limitations in real-world scenarios. A summarized version of all families can be found in (Suppl.Table S1).
2.1. Parameter Aggregation
This is one of the earliest strategies, where a global model is constructed by aggregating parameters from each local client model after every few iterations of local training (Figure 2).
The historically first aggregation strategy (FedAvg [9]) falls under this category. Clients train locally before sending updates to the server, requiring fewer communication rounds while being relatively easy to implement. A shortcoming of this method are large update sizes, which is prone to model divergence under non-IID data. FedProx [14] sought to fix this by adding a proximal term to weight drift, forcing the global consensus model to maintain the impact of variable local updates. HeteroFL [28] performs partial weight aggregation that enables clients with heterogeneous resources to train models with different capacities while still contributing to a single global consensus model. Similarly, FlexiFed [29] addresses model architecture heterogeneity by aligning parameters across different client architectures with a reference map before aggregation, instead of blindly averaging the parameters.
Parameter aggregation comes with its own issues, including the risk of information leakage through model inversion [30], computational cost, and slower convergence when it comes large DL models [9,14], and the potential to wash out local patterns, especially in the case of non-IID data, leading to utility issues [12].
2.2. Gradient Aggregation
To reduce communication overhead and have fine-grained control over global optimization, an alternative is for the server to aggregate local gradients from clients rather than model weights (Figure 2).
FedSGD [9] and DPSGD-FL [31] were the earlier established gradient based aggregation strategies for FL which demonstrates that decentralized training can achieve convergence behavior comparable to centralized stochastic gradient descent (CSGD). FedGS [32] proposes a balanced framework that sends reduced information through gradient sparsification with a higher degree of freedom for additional control over communication and computation trade-off. FedNova [33] addresses the issue of objective inconsistency caused by the unfairly scaled gradients resulting from different numbers of local update steps across clients. It normalizes the local gradient updates to ensure a more balanced contribution from clients during aggregation.
As clients share the direction of optimization, gradient aggregation reduces communication overhead across distributed clinical sites. However, this approach may lead to oscillatory optimization with instability in updates. Under heterogeneous medical environments, conflicting gradient directions may degrade global model utility compared to conventional weight aggregation.
2.3. Cluster-Based Aggregation
Cluster-based aggregation was proposed to overcome the client drift inherent in heterogeneous datasets. In this approach, clients are organized into clusters based on criteria such as data distribution, computational capacity, or client availability, and aggregation is performed per cluster (Figure 2). In most cases, each cluster maintains its own global model, resulting in multiple cluster-level models rather than a single global model representing all the clients.
Along these lines, ClusterFL [34] groups clients based on the similarity of their updates and performs aggregation within each cluster. By dropping slow-converging or weakly correlated clients within each cluster, this method reduces communication overhead and accelerates convergence. However, this method also introduces cluster-wise dropout for slow clients and correlation-based client selection, which can exclude clinically valuable but resource-constrained participants, partially undermining the purpose of FL. FedCluster [35] extended this idea to grouping clients based on time zones and availability, employing cluster cycling to alternate training between clients. This reduces client drift and improves convergence under non-IID data settings.
The aforementioned strategies are based on hard association assumptions to form clusters of IID clients, but comes with the shortcoming that FL cannot effectively exploit similarities between different clusters. To overcome this, a soft clustering way is proposed. Fedsoft [36] allows every local dataset to follow a mixture of multiple source distributions. However, this approach does not explicitly consider similarities between clusters, potentially limiting its ability to leverage shared structure across clients. FedSim [37] further encourages collaboration among similar clients, while minimizing the variance from dissimilar clients through clustering.
Methods such as IFCA [38] which use dynamic clustering can cause clients to shift between clusters, potentially destabilizing training. Moreover, the generation of multiple global models challenges the traditional notion of generalization, as most clustering techniques prioritize cluster-level utility over a single unified global consensus model.
By leveraging clients with similar data distributions, cluster-based aggregation handles heterogeneous datasets, including differences between populations or institutions. Additionally, it provides a form of partial personalization, allowing separate cluster-level models rather than a single global model, which better accommodates local variability.
2.4. Knowledge Distillation Aggregation
Knowledge distillation aggregation allow clients to maintain different local models while contributing to an FL consortium. In this framework, clients send their predictions, logits, or a smaller model to the server instead of the local model parameters (Figure 2). Unlike conventional FL, where the global model is obtained by directly averaging client model parameters,the server uses the clients predictions to train a global student model. This allows clients who have different computational capacities and data constraints to train heterogeneous models and make them adaptable to the shared FL global ecosystem.
Knowledge distillation has evolved through several research directions. FedMD [39] was among the first methods to demonstrate that clients with heterogeneous model architectures can collaborate by exchanging prediction logits instead of model parameters. Building on this idea, FedDF [40] introduced server-side ensemble distillation, where the global model is trained by aggregating the predictions of local models rather than averaging their parameters. FedBE [41] further enhanced this framework by incorporating Bayesian inference to generate multiple candidate global models before compressing them into a single distilled model, accounting for uncertainty during aggregation. Subsequent methods extended these ideas to address specific challenges, including data-free distillation (FedFTG [42]), improved initialization and privacy with auxiliary data (FedAUX [43]), and ensemble-style model fusion (FedFusion [44]), among others (Suppl.Table S1).
Overall, knowledge distillation offers greater flexibility than parameter aggregation by allowing clients to train heterogeneous models tailored to their local data modalities (e.g., medical imaging, electronic health records, or biosignals) and computational resources, making it well suited to heterogeneous healthcare environments. However, because aggregation relies on model outputs rather than model parameters or raw data, the distilled global model may lose information contained in local representations, particularly for rare or complex patterns. Additionally, although sharing logits reduces the exposure of model parameters, prediction outputs can still reveal sensitive information through inference attacks, and biases present in local models may propagate into the distilled global model.
2.5. Representation Aggregation
Representation-based aggregation addresses non-IID data by aggregating feature-level representations learned by client models instead of their complete model parameters (Figure 2). By aligning latent feature representations across clients, these strategies improve collaboration under heterogeneous data distributions, such as class imbalance or varying patient populations.
Representation-based aggregation has evolved to improve feature alignment under heterogeneous data distributions. FedPer [45] was among the first methods to separate shared feature extraction from task-specific prediction by aggregating only the feature extractor while keeping classifiers local, enabling personalization under non-IID settings. Building on representation learning, MOON [46] introduced contrastive learning to align local and global feature representations, reducing client drift and improving training stability. FedProto [47] further generalized this idea by exchanging class prototypes instead of model parameters, allowing clients to communicate compact feature-level summaries while remaining architecture-independent. Subsequent methods extended these ideas to improve representation learning and personalization, including explicit decoupling of shared representations and local heads (FedRep [48]) and adaptive feature-aware aggregation for highly heterogeneous data (FedMix [49]), among others (Suppl.Table S1).
However, the quality of the aggregated model depends on the learned feature representations, which may become biased toward dominant client populations, potentially skewing the global model toward particular distributions [50]. Careful design of weighting and representation alignment strategies is critical to ensure equitable and effective aggregation.
2.6. Ensemble-Based Aggregation
Ensemble-based aggregation addresses client heterogeneity by combining the predictions of multiple local models rather than directly averaging their parameters (Figure 2). By operating in the prediction space, these strategies mitigate utility degradation caused by direct parameter averaging under strong statistical heterogeneity, such as that encountered by FedAvg [9].
Ensemble-based aggregation has evolved through several research directions. FedBE [41] introduced Bayesian model averaging in function space, treating client models as posterior samples to explicitly account for uncertainty during aggregation. FedAUX [43] further improved ensemble quality by assigning confidence-aware weights to client predictions using auxiliary data, allowing more reliable models to contribute more strongly to the aggregated output. More recently, ensemble techniques have been extended to graph-structured data, where ensemble Graph Neural Networks (GNNs) [51] combine node embeddings or predictions from multiple GNNs to improve robustness and generalization. Other ensemble-based strategies, including distillation-driven approaches such as FedDF, FedFusion, and FedSEAL, further compress ensemble knowledge into a single deployable global model (Suppl.Table S1).
Despite providing robustness and flexibility, ensemble-based aggregation necessitates the maintenance of multiple models, which increases computational complexity and resource requirements.
2.7. Attention-Based Aggregation
Another way to handle non-IID data is through an attention-based aggregation strategy. This approach introduces an adaptive mechanism for combining client updates by assigning importance to model parameters, representations or feature embeddings for each client during aggregation in the server (Figure 2). Instead of averaging all the parameters of the model, the server uses attention mechanisms to dynamically weigh client contributions based on their relevance and informativeness. This approach is particularly effective in addressing challenges posed by the non-IID nature of data distributions or multi-modality modeling, where each client trained data from different modalities.
Attention-based aggregation has evolved through several research directions. The Attentive Aggregation method [52] was among the first to introduce attention mechanisms for federated aggregation by iteratively assigning adaptive weights to client updates based on their similarity to the global model. Building on this idea, FedAMP [53] enabled personalized aggregation by allowing clients to receive greater contributions from similar peers, thereby reducing negative transfer under heterogeneous data distributions. More recently, HEALNET [54] extended attention mechanisms to multimodal healthcare by learning modality-specific parameter spaces within a unified global model. Subsequent methods further adapted attention for specialized tasks, including hierarchical client clustering for traffic prediction (FedDA [55]) and attention-guided client selection (FedABC [56]), among others (Suppl.Table S1).
The main advantage of attention-based aggregation is that it prioritizes more relevant and informative clients, allowing adaptive personalization [57]. However, this type of bias amplification directly conflicts with fairness requirements in healthcare AI. The attention weights are not interpretable even though they are mathematically defined. These weights do not always provide clear or clinically meaningful explanations for decision-making. This lack of traceability can hinder trust and adoption in real-world applications.
2.8. Meta-Learning Aggregation
The meta-learning global model is built based on client model learning. Instead of learning one static global model from local parameters or local SGD from clients, the server will learn the initialization rules and adaptation to each client’s data. Meta-learning focuses on learning algorithm utility for arbitrary tasks across devices (Figure 2).
FedMeta [58] demonstrates this by facilitating collaborative training across distributed clients with a meta-learner, which can quickly adapt to a client’s local data using only a few gradient updates. Building on this idea, pFedMe [59] formulates FL as a bi-level optimization problem, learning a global initialization that can rapidly adapt to individual client needs. Similarly, Per-FedAvg [60] extends the traditional FedAvg framework by incorporating meta-learning, allowing clients to achieve personalization through a few local gradient updates. Moving beyond shared initialization, a meta-learning inspired personalized FL approach, pFedHN [61] introduces a model generator approach, enabling stronger and more flexible personalization with relatively small communication overhead.
Meta-learning introduces a shift from simply learning a single global model to learning how to learn across diverse clients [7]. This becomes especially important in healthcare settings, where data are inherently heterogeneous. By capturing shared knowledge while allowing rapid adaptation to local distributions, meta-learning enables models to perform better on rare diseases and personalization [62], which are often under-represented in global training. However, meta-learning also increases computational complexity to the system as it requires additional optimization steps to learn adaptable model parameters. These strategies are sensitive to data distribution shifts, which can negatively impact utility if local datasets are non-IID [63]. A critical concern in clinical settings is explainability, and interpretable meta-learning models tend to function as complex black-box systems, making it difficult for clinicians to understand or trust their predictions, thereby limiting their practical adoption in healthcare environments.
2.9. Optimization-Based Aggregation
In optimized-based aggregation, the server treats the aggregation process as a constrained optimization problem instead of naive parameter averaging. These strategies actively adjust how updates are combined to improve convergence, stability, and fairness, especially under non-IID data (Figure 2).
Optimization-based aggregation has evolved through several research directions. Oort [64] improves training efficiency by prioritizing clients that provide high-quality and timely updates, accelerating convergence compared with uniform client sampling. FedDyn [65] further stabilizes optimization by introducing dynamic regularization to reduce divergence caused by heterogeneous client objectives. In parallel, FedAdam and FedYogi [66] adapt server-side optimization by incorporating momentum and adaptive learning rates, enabling more robust convergence under heterogeneous data distributions. Subsequent methods further refined optimization through drift correction and update alignment, including FedDC [67] and correlation-aware optimization in FedCorr [68], among others (Suppl.Table S1).
Optimization-based aggregation combines local updates into a coherent and high-performing global model with stability and faster convergence. However, rapid convergence makes the system sensitive to hyper parameters, with even minor misconfiguration in learning rates or momentum terms significantly affecting model utility. In healthcare, particularly in the presence of rare disease cases, adaptive aggregation mechanisms may carelessly control under-represented updates if they contribute less to the global objective.
2.10. Sparsification Aggregation
In sparsification-based aggregation clients send masked parameters or a subset of model parameters to the server, with the goal of easing the communication bottleneck. Masking strategies of this nature include top-k gradient selection, magnitude pruning, random sparsification, or structured pruning masking, enforcing only the important parameter subsets to reach the server (Figure 2).
SparseFed [69] was among the first methods to reduce communication costs by applying global top-k update sparsification together with device-level gradient clipping, while also improving robustness against model poisoning attacks. Building on this idea, FedSPA [70] incorporated sparsification directly into the aggregation process by enforcing a shared sparse parameter structure across clients. Lossless Gradient Sparsification [71] further improved communication efficiency by representing gradients in an alternative space that increases their compressibility without sacrificing model utility. Subsequent strategies extended these ideas through adaptive parameter selection and sparsification strategies, including FedSparse [72] and Top-k/Rand-k gradient selection methods [73], among others (Suppl.Table S1).
Sparsification is compatible for large-scale applications, allowing efficient network bandwidth usage in distributed learning [72]. However, as this method depends on the selection strategy, fairness in aggregation of under-represented data distribution could be missed. It also complicates debugging, making it difficult to trace individual contributions for clinical validation because of irreversible sparsification.
2.11. Peer-to-Peer Aggregation
Peer-to-peer aggregation eliminates the need for a central server, allowing clients to converge through consensus rules without strict central averaging. Generally, peer-to-peer aggregator and client have a common predefined communication protocol that they abide (Figure 2).
The segmented gossip approach [74] was among the first decentralized aggregation strategies, combining model partitioning with gossip-based communication to improve bandwidth utilization and convergence without a central server. Building on FL optimization, SecureD-FL [75] employs the Alternating Direction Method of Multipliers (ADMM) to enable peers to collaboratively compute the global model through distributed consensus. BAFFLE [76] further extended this paradigm by leveraging blockchain and smart contracts to coordinate aggregation, while improving computational efficiency through parameter-space decomposition. Subsequent strategies expanded peer-to-peer aggregation to broader FL environments, including neighborhood-based serverless aggregation [77] and other blockchain-enabled frameworks such as Block-chain-based Federated Learning framework with Committee consensus BFLC [78], among others (Suppl.Table S1).
Peer-to-peer aggregation improves robustness by eliminating the single point of failure. However, it comes with notable challenges. For example, gossip or ring-based topologies can become inefficient as the number of clients grows [74]. The lack of a central authority makes it harder to detect and mitigate malicious updates, which increases security threats. The absence of centralized tracking complicates the contribution, fairness, and critical auditing concerns in sensitive healthcare data [79].
2.12. Confidential Aggregation
Confidential aggregation addresses privacy risks posed by the aggregator itself. The standard threat model assumes an honest-but-curious adversary: the aggregator executes the aggregation protocol faithfully, but may attempt to extract sensitive information from the individual model updates submitted by clients. Confidential aggregation strategies prevent such leakage by ensuring that the server can perform the aggregation without ever seeing individual client updates in plain text. (Figure 2)
Homomorphic encryption (HE) [80] allows the aggregator to perform computations directly on encrypted model updates, ensuring that the server never gains access to raw gradient or weight values, while still producing correct aggregated results. This provides strong cryptographic privacy guaranties at the cost of significant computational and communication overhead, driven by the explosion in ciphertext size and computational complexity, which can be prohibitive in large-scale healthcare deployments with numerous participants.
Secure aggregation protocols, most notably SAFELearn [81], enable the server to compute the sum of client updates without learning individual contributions. Through pairwise masking and secret sharing techniques, each client’s update is encrypted in such a way that the masks cancel out only after all updates are combined, ensuring that only the final aggregate is revealed to the server. While computationally more efficient than HE, secure aggregation remains vulnerable to inference attacks on the final aggregate itself and requires careful handling of client dropouts.
Trusted Execution Environments (TEEs) provide a hardware-based approach to confidential aggregation by processing client updates within an isolated, attested enclave, such as Intel SGX or AMD SEV. The aggregator’s own operating system, and its system administrators, cannot inspect computations within the enclave providing a practical alternative to pure cryptographic strategies with lower computational overhead. Recent work has demonstrated the feasibility of TEE-based FL in healthcare settings [82], though limitations in secure memory capacity and vulnerabilities to side-channel attacks remain active concerns.
Secure Multi-Party Computation (SMPC) extends the principle of confidential aggregation by distributing the aggregation computation across multiple non-colluding servers, such that no single party can access individual client updates in plain text. Secure Multi-Party Computation such SMPC protocols[83] and Arithmetic, Boolean, and Yao (ABY) [84,85] enable parties to jointly compute aggregate model parameters while maintaining privacy guarantees even if some servers are compromised. SMPC-based FL mechanisms [86] have been increasingly adopted for applications requiring strong privacy guarantees across institutional boundaries, at the cost of increased communication rounds and protocol complexity.
Confidential aggregation strategies form the backbone of privacy-preserving FL in healthcare by providing cryptographic security guarantees that client updates remain confidential during the aggregation process. However, the stronger the cryptographic guarantee, the higher the computational and communication overhead, which raise a critical tension in resource-constrained healthcare deployments. Furthermore, confidentiality during aggregation does not eliminate the risk of information leakage from the aggregated model itself, underscoring the need for complementary privacy-preserving techniques.
2.13. Byzantine-Robust Aggregation
Byzantine robustness refers to a system’s ability to maintain correct operation and reach consensus even when some nodes are compromised, malfunctioning, or acting maliciously. In FL aggregation, it tackles a general security threat within the FL network, where the adversaries could be any of the protocol participants. For example, a malicious client (or a coalition of clients) may submit intentionally corrupted or crafted updates to degrade the global model’s accuracy or to inject targeted backdoors. Unlike the honest-but-curious adversary in confidential aggregation, byzantine adversaries do not follow the protocol and may send arbitrary values. The goal of Byzantine-robust aggregation is to ensure that the final global model remains accurate and uncorrupted. This distinction is critical in healthcare, where a poisoned model could lead to misdiagnosis or incorrect clinical decisions. (Figure 2)
Statistically robust aggregation strategies defend against Byzantine attacks by replacing or modifying the standard averaging operation with aggregation functions that are inherently resistant to extreme values. Solutions like Multi-Krum [87] select a single client update that is most similar to its neighbors, effectively discarding outliers, while its generalized variant selects and averages a trusted subset of updates. Beyond statistical filtering, trust-based aggregation strategies maintain a reputation or trust score for each client and weight their contributions accordingly. FLTrust [88] introduces a trusted root model that validates client updates before aggregation, ensuring that only updates consistent with trustworthy behavior influence the global consensus model. Similarly, FORTA [89] enhances the Krum algorithm by incorporating discrete Fourier transform guidance to better distinguish between benign and malicious updates under heterogeneous data conditions.
An emerging direction in Byzantine-robust aggregation is the utilization of computational verifiability, which enables clients or auditors to cryptographically confirm that aggregation was performed correctly on legitimate updates. Techniques leveraging zero-knowledge proofs [90], TEEs [91], and authenticated data structures [92] allow participants to verify that the server has not tampered with the aggregation process and that all included updates originated from registered clients [93]. Verifiable aggregation addresses a critical trust gap in cross-institutional healthcare FL where no single authority is universally trusted, though these strategies currently incur significant computational and communication overhead that limits practical deployment.
Collectively, Byzantine-robust aggregation strategies provide a multi-layered defense against adversarial manipulation of the global model, with methodologies from statistical outlier rejection to trust-based client weighting to cryptographically verifiable computation. A key concern in healthcare is that aggressive filtering of anomalous updates may inadvertently suppress legitimate contributions from institutions with underrepresented patient populations, creating a tension between security and fairness. Verifiable aggregation addresses the trust gap in cross-institutional collaboration, where no single authority is universally trusted, though practical adoption remains constrained by computational overhead.
3. Challenges in Healthcare FL
Deploying FL in real-world healthcare settings comes with significant challenges, span across technical (Figure 3) and data-related challenges (Figure 3).
FL practitioners must carefully select the right FLAg strategy, as it directly affects the utility, fairness, and reliability of the model. Simple FLAg strategies can fail under complex conditions, leading to biased predictions or unstable training. In healthcare, where decisions can impact patient outcomes, careful consideration of both system constraints and data characteristics is essential to ensure robust, trustworthy, and clinically-relevant models.
This section discusses key technical and data-related challenges encountered when implementing FL in healthcare and highlights the FLAg strategies designed to address them.
3.1. Technical Challenges
3.1.1. Stragglers
Stragglers are clients whose local training or communication takes significantly longer than other clients per FL round. Their delayed model updates can slow down or stall the aggregation process, particularly in synchronous FL algorithms, where the server typically waits for updates from a predefined set of clients before proceeding to the next round.
Although a natural response for stragglers is to ease the synchronization requirement through asynchronous strategies, this strategy do not universally solve the problem and generally more effective in large-scale FL settings with hundreds of participating clients (e.g., at scale of 100 or more) to achieve utility and training efficiency comparable to centralized learning [94]. Straggler mitigation in cases like healthcare requires approaches that explicitly accommodates heterogeneous computational capabilities rather than relying solely on asynchronous participation. Existing aggregation strategies address this challenge through several complementary mechanisms. Sparsification based aggregation strategies such as [95] allow flexibility to clients to train and contribute sub-models suited for their computational capabilities and eliminate the likelihood that resource limited clients become persistent stragglers during aggregation. In contrast, rather than preventing the occurrence of stragglers, optimization-based aggregation strategies accommodate differences in local training progress and correct the resulting objective inconsistency during aggregation. For example, FedNova [33] uses normalized averaging allows fast clients perform more local updates per communication round and stragglers to train minimal rounds and attempt to remove objective inconsistencies and ensures fair contribution from all clients during aggregation. Clustering-based aggregation addresses the problem by grouping clients based on their computational capabilities and aggregate them without requiring every participant to progress at the pace of the slowest client. Peer-to-peer aggregation further relaxes dependence on a centralized synchronization barrier. Instead of eliminating or grouping the clients, clients exchange and update models with peers according to their availability. Faster participants can therefore continue exchanging updates without stalling for stragglers, effectively distributing the synchronization burden across the network.
Collectively, these strategies illustrate a shift from strict synchronization toward adaptive and selective aggregation strategies to mitigate the negative effects of stragglers and the heterogeneity caused by them in FL systems.
3.1.2. Asynchronous Training
Asynchronous training allows the global model to update continuously as client updates arrive, rather than waiting for all participating sites to complete local training.
While this improves system responsiveness and avoids delays caused by stragglers, it introduces challenges such as slower or unstable convergence and scalability issues due to inconsistent update timing and heterogeneous data distributions.
Traditional asynchronous aggregation suffers from communication bottleneck when the consortium is large, resulting in suboptimal convergence and while the parameter traffic to the server is congested. Theoretical analysis of Gradient-based Asynchronous Decentralized-PSGD [94] shows that the clients to use stale weights to compute gradients shows consistent convergence rate similar to D-PSGD and SGD, suggesting that the issue is not with asynchronism itself but the way it is managed. Methods like FedBuff [96] attempt a buffered asynchronous aggregation where the server aggregates clients updates in a secure buffer before performing the update. This provides a good guaranty of convergence and achieves better scalability, especially in large healthcare collaborative settings with many participating nodes. Similarly, hierarchical FL Async-HFL[97] utilizes hierarchical asynchronous aggregation to aggregate in multi-layer manner (tier 1 -cloud layer, tier 2 - gateway layer , tier 3 - edge layer), which collaboratively optimize model convergence and stability. Despite these advances, fully asynchronous strategies still struggle with the dual challenges of staleness and data heterogeneity.
Strategies such as knowledge distillation based or meta-learning based can focus on issue caused by staleness and data heterogeneity. By solely relying on sharing knowledge predictions or initializations, these approaches indirectly makes sure to reduce the impact of negative effects of data heterogeneity and staleness.
Rather than eliminating the challenges of asynchronism, these approaches progressively constrain and manage it. The main focus of these strategies is turning a source of instability into a tunable design.
3.1.3. Varying Resource Environments
Large healthcare institutions may be wealthy in terms of resources, such as powerful GPUs, stable network connectivity, and dedicated AI teams. On the other hand, smaller clinics or regional hospitals might operate on limited CPU resources, intermittent connectivity, and heavy clinical workloads. Apart from that, data heterogeneity may further be amplified due to differences in imaging devices and acquisition protocols. This resource variation immediately breaks the assumption that all clients can train the same model under the same conditions, bringing model heterogeneity. Heterogeneous models vary in model size and types across clients, which creates stale updates and inconsistency in training with each client.
Many aggregation strategies have attemped to handle model heterogeneity. For instance, clustering-based strategies such as ClusterFL [34] and FedCluster [35] take the first step by preserving model diversity, allowing multiple client-specific models to coexist and contribute collectively, rather than collapsing them into a single averaged solution. Knowledge distillation approaches like FedMD [39], FedDF [44], and FedKD [98] shift the focus from parameter sharing to knowledge sharing, enabling clients with different model architectures and capacities to exchange soft predictions, thereby decoupling learning from strict structural uniformity. These approaches alone do not resolve the instability introduced by inconsistent local training, which is where optimization-based strategies such as FedProx [14], FedNova [33], and SCAFFOLD [99] play a critical role by correcting objective mismatches, mitigating client drift, and stabilizing convergence despite uneven update frequencies.
Attention-based aggregation strategies like FedAtt [57] and FedAMP [53] introduce a more selective mechanism, dynamically weighting client contributions based on their relevance or reliability, ensuring that faster or higher-quality updates are emphasized without entirely discarding slower participants. Together, these strategies illustrate a shift from rigid averaging toward adaptive, multi-faceted aggregation, where diversity is not suppressed but structured—balancing flexibility, stability, and fairness in the face of real-world healthcare heterogeneity. Peer-to-peer aggregation like Fedlesscan [100] includes a clustering-based FL strategy designed specifically for peer-to-peer environments to make it staleness-aware, mitigate slow model updates and avoid wasted contribution of clients.
These approaches shift aggregation from a rigid averaging process to a more flexible, heterogeneity-aware mechanism that sustain learning in healthcare environments.
3.1.4. Quantization Aware
The transmission of full-precision model updates becomes substantial communication overhead in FL, particularly when large models are trained across client with resource heterogeneity. Quantization-aware strategies address this by allowing clients to send compressed updates while aiming to retain sufficient information for effective global aggregation.
One of the most commonly discussed problems in compression methods, as discussed in [101], [102] and [103] is Quantization noise. This is the noise generated when high precision numerical value is represented in low precision numerical value. In FL, the quantization noise is not uniform across clients, where differences in hardware and implementation can introduce variability in how updates are quantized, making aggregation noisier and potentially less reliable. This issue becomes more in non-IID healthcare data settings, where client updates are already misaligned. Adding quantization noise can further distort their contribution, amplifying heterogeneity rather than mitigating it.
Ultimately, the effect of quantization depends on type of information being communication. Strategies such as Gradient-based and parameter based approaches provides a effective framework for applying quantization to the updates exchanged between clients and the server. However, quantization may introduce approximation errors into the transmitted updates, leading to the development of methods that adjust the quantization process according to the characteristics of the local updates. For instance, FedWSQ [102] employs distribution-aware non-uniform quantization to reduce quantization error by considering the statistical distribution of local model updates, while weight standardization is incorporated to improve learning stability under heterogeneous data. Adaptive quantization is considered in AdaGQ [104], where the quantization resolution is dynamically adjusted according to gradient characteristics across training rounds and heterogeneous communication conditions across clients.
Collectively, these approaches demonstrate that the effects of quantization can be controlled through distribution-aware, and adaptive quantization mechanisms rather than by applying a uniform fixed-precision representation to all client updates. This balance is more important in heterogeneous healthcare environments, where variations in local data distributions and system resources may result in different update and communication characteristics across participating clients.
3.1.5. Security
In FL for healthcare, security in aggregation refers to ensuring both that the model training process and the output are robust against adversarial manipulation by adversaries. Unlike privacy threats that target data confidentiality, security threats aim to compromise model integrity. The standard threat model assumes a Byzantine adversary: a malicious client (or a coalition of clients) that submits arbitrarily corrupted updates to degrade model accuracy or inject a targeted backdoor. Or a malicious server that manipulates the aggregation process. In healthcare, a poisoned model could lead to systematic misdiagnosis with direct clinical consequences.
Model poisoning attacks allow malicious clients to send manipulated updates that degrade model utility or introduce backdoor, particularly dangerous in healthcare, where incorrect predictions can have clinical consequences. Inference attacks such as gradient inversion or membership inference can exploit shared updates to reconstruct sensitive patient data, even without direct access to it. Unreliable or heterogeneous clients may behave unpredictably, either due to faults or adversarial intent, making it difficult to distinguish between benign noise and malicious contribution. Secure communication overhead can become significant, especially when dealing with bandwidth-constrained hospital networks, creating a tension between security and efficiency.
FL must ensure model integrity in the presence of malicious or faulty clients, which is addressed through robust aggregation strategies. Algorithms such as Multi-Krum [105] select updates that are closest to the majority, filtering out anomalous or adversarial contributions, while trimmed mean and coordinate-wise median reduce the influence of extreme gradient values by aggregating only central tendencies. Bulyan [106], a hybrid method, further strengthens the robustness by combining outlier filtering with robust averaging to defend against strong Byzantine attacks.
To reduce reliance on centralized servers and improve communication resilience, FL aggregation approaches such as gossip learning and decentralized SGD (D-PSGD) [94] allow clients to exchange model updates directly in a peer-to-peer network, eliminating single points of failure while still enabling global learning. Finally, meta-learning-based approaches such as FedMeta [58], Per-FedAvg [60], and pFedMe [59] shift the objective from learning a single global model to learning a shared initialization that can rapidly adapt to each client’s local distribution, enabling strong personalization in heterogeneous healthcare environments.
Together, these aggregation strategies form a multi-layered framework where privacy, robustness, efficiency, and personalization are jointly addressed rather than treated in isolation.
3.2. Data Related Challenges
3.2.1. Uncertainty
Uncertainty in healthcare FL can broadly arise from aleatoric and epistemic sources. Aleatoric uncertainty reflects the noise present in clinical observations, such as measurement noise and ambiguity in medical data. In contrast, epistemic uncertainty arises from limitations in the model’s knowledge, which may result from insufficient, sparse, or unrepresentative training data across participating clients
The main challenge with uncertainty in healthcare is that it is very difficult to distinguish outlier data or data anomalies from non-IID data in healthcare applications with rare-disease conditions and unbalanced data. To achieve robust aggregation, deviations in data and model updates are mostly identified and handled to maintain stability. However, in heterogeneous settings such deviations may arise from rare cases, rather than anomalous or unreliable data. This may result in instability and ambiguity in model aggregation and bias amplification.
In order to handle client level uncertainty, oort [64] provides a client selection framework that filters the participating clients before aggregation. Instead of aggregating all updates, Oort allows the selection of high-utility and reliable clients, effectively reducing the impact of noisy, slow, or low-quality data sources and thereby stabilizing the aggregated global model under heterogeneous healthcare settings.
Model level uncertainty is addressed via probability-based ensemble aggregation, FedBE [41] incorporates uncertainty in decision-making by placing a distribution over model weights and marginalizing these models to form a whole predictive distribution. In this case, local models are treated as samples from a posterior distribution rather than deterministic parameter estimates. The server aggregates these model updates through Bayesian ensemble, capturing model-level uncertainty and improving robustness against variability across hospitals, patient demographics, and data collection protocols. Another approach to mitigate model level uncertainty is TM-FAL [107], which uses temporal uncertainty strategy instead of relying on a single model prediction to avoid inconsistent anomalies and suitable for long-tailed datasets.
These methods collectively indicate that uncertainty in FL arises at multiple stages of the learning pipeline, including client participation, model estimation, and aggregation. Existing methods tries to handle specific aspects of uncertainty in highly heterogeneous healthcare settings and still remain as an open challenge.
3.2.2. Privacy
Privacy in FL refers to the protection of sensitive patient information such as medical images, electronic health records, and genomic data from being exposed during training, even though raw data never leaves the hospital. Institutions share model updates instead of sharing raw data. However, these updates indirectly may leak private information, making privacy a key concern rather than something that being guaranteed.
One of the major threats is model inversion, where an intruder may exploit access to model parameters or gradients to reconstruct sensitive patient information, such as medical images or underlying clinical features, effectively reversing the learning process to approximate original health records. Similarly, membership inference is when the attacker attempts to determine whether a specific patient’s data was part of the training dataset. In healthcare, this can expose whether an individual was treated for a particular condition at a hospital, leading to serious privacy violations. Attribute inference on the other hand, is when sensitive hidden attributes such as disease status, genetic predisposition, or demographic factors are inferred from seemingly innocuous model outputs or intermediate representations, even when those attributes were never explicitly included in the training data.
Secure aggregation methods, such as SecAgg [108], homomorphic encryption [109], and secret sharing, where individual client updates are cryptographically masked so that the server can only observe the aggregated result, all mitigate reconstruction of patient-level information. Differential privacy [110], [70] provides a complementary statistical defense by adding calibrated noise to model updates, offering formal privacy guarantees bounded by a privacy budget that quantifies the trade-off between confidentiality and model accuracy [17].
Overall, privacy-preserved FL aggregation strategies provide important safeguards against information leakage, while enabling collaborative model training across healthcare institutions. However, achieving strong privacy guarantees seems to come at the cost of model utility, communication efficiency, and computational overhead, highlighting the need for methods that can effectively balance privacy protection and learning utility.
3.2.3. Data Imbalance
Data imbalance is common since each institution contributes differently in the size and composition of their local datasets. Large tertiary hospitals may contribute data from broad patient populations, whereas community hospitals or specialized clinics may contain fewer samples concentrated around particular diseases or demographic groups. Consequently, healthcare FL may exhibit imbalance at multiple levels, including unequal dataset sizes across institutions, skewed class distributions within individual clients, and long-tailed representation of rare diseases or patient populations.
One major issue is biased global model learning on long tailed datasets, where clients with larger or more diverse datasets dominate the aggregation process, causing the global model to be skewed toward majority populations while under performing on under-represented groups or smaller healthcare institutions. This leads to poor generalization across clients, especially in healthcare settings where rare diseases, minority populations, or small clinics may be systematically ignored.
Optimization based aggregation addresses the main statistical challenge associated with data imbalance is global objective drift. For example, FedProx [14], controls client drift during local optimization and stabilizes contributions from clients with heterogeneous data distributions, thereby improving convergence stability even when data distributions are uneven. In contrast, FedWeightedAvg [102] adjusts the contribution of individual clients during aggregation based on factors such as dataset size or estimated importance. Beyond differences in client-level contributions, imbalance may also occur within local datasets through unequal class representation, TM-FAL [107] addresses this class imbalance through class-aware, pseudo-label-guided active sample selection by promoting a more balanced representation of classes throughout training.
Meta learning strategies such as FedMeta[58], Per-FedAvg[60] and pFedMe[59] address the imbalance by learning a shared initialization that can quickly adapt to each client’s local data distribution, making the global model less sensitive to skewed or uneven datasets.
Attention-based strategies such as FedAtt[57] and FedAMP[53] dynamically assign higher weights to more informative or under-represented client updates. This allows the aggregation process to selectively emphasize useful contributions, while reducing the dominance of large or redundant datasets.
Ensemble-based aggregation such as FedSDC [111] addresses class imbalance by introducing shuffled multi-head classifier and incorporate ensemble strategy on top of a shared feature extractor for decision stability.
Overall, existing approaches such as clustering, meta-learning, attention-based aggregation, and ensemble strategies have demonstrated their effectiveness in addressing the consequence of data imbalance in FL. However, maintaining unbiased model utility across highly heterogeneous and long-tailed healthcare datasets remains a significant challenge, particularly when client data distributions differ substantially in size, quality, and class representation.
3.2.4. Fairness and Bias
Fairness in healthcare FL concerns avoiding systematic disparities in model utility across participating institutions and patient populations, while ensuring that the aggregation process does not systematically disadvantage particular clients or underrepresented groups.
Maintaining fairness in healthcare FL is a challenging task as it involves data from multiple clients of different geographical regions and demographic groups. Such difference among these data sources creates biases in model utility through either over or under-representation.
In order to improve personalization, clustered based aggregation approaches such as clusterFL [34] and Flexifed [29] group institutions with similar characteristics, reducing bias introduced by aggregating fundamentally different populations.
Representation based aggregations address heterogeneity by aligning the information learned across clients. For example, MOON [46], uses contrastive learning to constrain local representations relative to global and previous local representations, whereas FedProto [47] exchanges and aggregates class-level prototypes to promote consistency in representation spaces across clients. By encouraging semantically related information to remain aligned across heterogeneous sites, representation-based approaches can improve personalization when local feature distributions differ substantially [27].
Personalization strategies, including Ditto [112] and pFedMe [59], are especially valuable in healthcare as they allow each institution to adapt the global model to its local patient population, accounting for variations in disease prevalence, equipment, and clinical protocols. Similarly, approaches based on adversarial training such as [113], propose individual fairness on the global model as well as on client-side through robust optimization where a model is trained on both clean and adversarially perturbed inputs to improve resistance against worst-case perturbations.
Collectively these types of aggregation address different sources of unfairness in FL healthcare systems. Achieving fairness requires more than balanced participation during aggregation. It requires aggregation and optimization strategies that account for both institutional heterogeneity and population-level differences throughout the learning process. Robust approaches further supports this objective by improving model resilience to variations across clients and patient populations diversity.
4. Discussion
In this article, we provided a systematic survey and a taxonomy of FLAg strategies designed for healthcare, while taking into account their relation to distinct technical and data-related FL challenges. Significant progress is recognized in FLAg strategies to improve model utility and privacy guarantees, while mitigating these challenges (Figure 3), e.g., stragglers, system heterogeneity, fairness and biases, uncertainty, security, and data imbalance. With this focused content, this article intends to serve as a guide to readers navigating and interpreting the current scattered body of relevant literature, towards developing and deploying scalable, equitable, and clinically reliable FLAg strategies.
The following trade-off patterns are recognized across current FLAg strategies:
- Robustness vs Fairness: FLAg strategies designed to resist anomalous or outlier updates also tend to down-weight updates that deviate from the majority. In healthcare, such outliers are not necessarily malicious, but they may reflect long-tailed data distributions, including rare disease populations or minority cohorts. Thus, improving robustness should be done with care as it is often a byproduct of suppressing underrepresented features, compromising healthcare equity.
- Personalization vs Generalization: Personalized FL approaches adapt models to local hospital data distributions or particular individual as a target utility. In contrast, generalization targets single global model that works consistently throughout. The key challenge here is that one must completely choose one side. In practice, it is not alright and many approaches try to balance both by using Shared global backbone and local fine-tuning or clustering similar clients.
- Stability vs Convergence rate: Strategies that stabilize training through fewer local epochs with smaller updates, regularization, and conservative aggregation can reduce high-frequency fluctuations in model utility during training and improve convergence stability across heterogeneous clients. However, these strategies may slow down convergence rate, increasing training time and computational costs. In clinical settings, unstable models are unreliable. The goal here is not to choose stability over speed, it is to avoid naive stability that kills learning efficiency.
- Computational expensiveness vs Communication overhead: Reducing communication to fewer rounds or sending compressed updates can reduce communication overhead. However, performing more local epochs between federated rounds and the additional processing required for compression and decompression may increase the local computation burden on participating hospitals. In contrast, reducing local computation leads to more frequent communication among the clients. Hospitals have uneven infrastructure with some able to handle complex computing, whereas others cannot handle it. Hence, design choices made with only one collaborating institution in mind could penalize the other(s).
- Traceability vs trustworthiness: Improving aggregation traceability by documenting can improve transparency and auditability. However, it can expose vulnerabilities such as enabling adversarial manipulation or privacy leakage. In healthcare, trust is not just about visibility, but about guarantees under adversarial and sensitive conditions.
These trade-offs arise from the interaction of statistical heterogeneity across institutions, such as differences in patient populations and disease distributions; system-level constraints, including computational capacity and communication bandwidth; and strict ethical and regulatory requirements related to privacy, fairness, and explainability. As a result, FLAg strategies should not be simply optimized for model utility but must balance the aforementioned criteria.
Frameworks such as NVIDIA FLARE [114], Flower [115], FATE [116], OpenFL [117], and IBM FL [118] play a critical role in enabling real-world deployment. Each platform is based on several critical features such as model compatibility, machine learning library support, expandability, privacy provisions, and commercial usage [119]. These frameworks provide end-to-end infrastructure for orchestrating distributed training across multiple institutions, handling tasks such as client-server coordination, secure data exchange, model versioning, and fault tolerance. They also support integration with hospital IT systems, allowing models to be trained directly on local clinical data while complying with privacy regulations like HIPAA. In practice, they simplify deployment by offering scalable architectures, APIs for customization, and support for heterogeneous environments where different hospitals may use different hardware or data formats. As a result, these frameworks act as the operational backbone that bridges theoretical FL methods with practical, large-scale, and regulation-compliant healthcare applications.
Benchmark datasets and organized computational challenges have played an important role in advancing FL research in healthcare. In medical imaging, the Federated Tumor Segmentation (FeTS) challenges provide a notable example, using multi-institutional multi-parametric brain MRI data to facilitate the development and evaluation of FLAg strategies for segmentation workloads [120,121]. These challenges have enabled the systematic evaluation of FLAg strategies across data originating from different healthcare institutions. However, much of the current literature still relies on curated benchmark datasets that are often preprocessed and standardized. While such settings support reproducible evaluation, they may not fully capture the noise, missing values, inconsistencies, and other complexities encountered in real-world clinical data, highlighting the need to further evaluate FLAg strategies under realistic clinical conditions.
Several gaps are identified in FL research between current methodological advances and real-world deployment in healthcare, including limitations in i) evaluation beyond model utility, ii) clinical risk assessment, and iii) governance requirements for transparent and accountable FL. Most FL studies focus on algorithmic improvements in isolation, assuming ideal infrastructure, stable connectivity, and homogeneous client capabilities. In reality, healthcare institutions differ widely in computational resources, data storage systems, and network reliability. The current literature also still continues to discuss the model utility centric as an evaluation metric, which is not sufficient for healthcare applications. There should be more emphasis on fairness, uncertainty, and clinical risk assessment for the model as a standard evaluation framework. Furthermore, there is a minimal understanding of regulatory and governance requirements for FL deployment in healthcare. Although privacy-preserving techniques are widely discussed, there is still a limited work to make sure there is transparency, auditability, and compliance with evolving healthcare regulations. Also, accountability, model validation, and cross-institutional trust remain largely not addressed, highlighting the need for frameworks that align technical design with legal and ethical standards.
Addressing these challenges requires a shift from purely algorithmic improvements toward system-aware designs that integrate adaptive aggregation, efficient communication protocols, and robust data standardization strategies tailored for clinical environments.
5. Conclusion
This review article offers a taxonomy of current FLAg strategies and how they strategically handle technical and data-related challenges during their deployment. It further also discusses the trade-off that these FLAg strategies bring. However, to have a more robust approach to clinical variability in real-world deployment, methodological innovation proven with system-level integration is necessary as a future direction. We hope that this survey will inspire fellow researchers to devote their efforts on investigating FL in the healthcare domain, targeting the deployment of scalable, equitable, and clinically reliable FLAg strategies.
Acknowledgments
This work was partially supported by the Informatics Technology for Cancer Research (ITCR) program of the National Cancer Institute (NCI) of the National Institutes of Health (NIH) under award number U24CA279629. The content is solely the authors’ responsibility and does not represent the official views of the NIH.
Supplementary Materials
The following supporting information can be downloaded at the website of this paper posted on Preprints.org, Table S1. Comparison of the reviewed FLAg strategies, according to their key idea, strengths, and limitations.
References
- Gulshan, V.; Peng, L.; Coram, M.; Stumpe, M.C.; Wu, D.; Narayanaswamy, A.; Venugopalan, S.; Widner, K.; Madams, T.; Cuadros, J.; et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. jama 2016, 316, 2402–2410. [Google Scholar] [CrossRef] [PubMed]
- Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Proceedings of the International Conference on Medical image computing and computer-assisted intervention; Springer, 2015; pp. 234–241. [Google Scholar]
- Bakas, S.; Akbari, H.; Sotiras, A.; Bilello, M.; Rozycki, M.; Kirby, J.S.; Freymann, J.B.; Farahani, K.; Davatzikos, C. Advancing the cancer genome atlas glioma MRI collections with expert segmentation labels and radiomic features. Sci. Data 2017, 4, 170117. [Google Scholar] [CrossRef] [PubMed]
- Davatzikos, C.; Barnholtz-Sloan, J.S.; Bakas, S.; Colen, R.; Mahajan, A.; Quintero, C.B.; Capellades Font, J.; Puig, J.; Jain, R.; Sloan, A.E.; et al. AI-based prognostic imaging biomarkers for precision neuro-oncology: the ReSPOND consortium. Neuro-oncology 2020, 22, 886–888. [Google Scholar] [CrossRef] [PubMed]
- EU, G. General data protection regulation. Off. J. Eur. Union 2016. [Google Scholar] [CrossRef]
- Act, A.; et al. Health insurance portability and accountability act of 1996. Public Law 1996, 104, 1–16. [Google Scholar]
- Bakas, S.; Li, X.; Shah, P.; Roth, H.R. Federated learning in healthcare: From research to real-world deployment. Annu. Rev. Biomed. Eng. 2026, 28. [Google Scholar] [CrossRef] [PubMed]
- Rieke, N.; Hancox, J.; Li, W.; Milletari, F.; Roth, H.R.; Albarqouni, S.; Bakas, S.; Galtier, M.N.; Landman, B.A.; Maier-Hein, K.; et al. The future of digital health with federated learning. npj Digit. Med. 2020, 3, 119. [Google Scholar] [CrossRef] [PubMed]
- McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; y Arcas, B.A. Communication-efficient learning of deep networks from decentralized data. In Proceedings of the Artificial intelligence and statistics; Pmlr, 2017; pp. 1273–1282. [Google Scholar]
- Sheller, M.J.; Reina, G.A.; Edwards, B.; Martin, J.; Bakas, S. Multi-institutional deep learning modeling without sharing patient data: A feasibility study on brain tumor segmentation. In Proceedings of the International MICCAI Brainlesion Workshop; Springer, 2018; pp. 92–104. [Google Scholar]
- Zhang, J.; Lin, M.; Wang, Y. Federated Learning with Non-IID Data: Bridging the Gap between Global and Local Models. Proc. arXiv 2019, arXiv:1905.02236. [Google Scholar]
- Zhao, Y.; Li, M.; Lai, L.; Suda, N.; Civin, D.; Chandra, V. Federated learning with non-iid data. arXiv 2018, arXiv:1806.00582. [Google Scholar]
- Xie, C.; Koyejo, S.; Gupta, I. Asynchronous federated optimization. arXiv 2019, arXiv:1903.03934. [Google Scholar]
- Li, T.; Sahu, A.K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; Smith, V. Federated optimization in heterogeneous networks. Proc. Mach. Learn. Syst. 2020, 2, 429–450. [Google Scholar]
- Li, X.; Sahu, A.K.; Talwalkar, A.; Smith, V. Fair Resource Allocation in Federated Learning. In Proceedings of the Proceedings of AAAI, 2020. [Google Scholar]
- Yang, Q.; Liu, Y.; Chen, T.; Tong, Y. Federated machine learning: Concept and applications. ACM Trans. Intell. Syst. Technol. (TIST) 2019, 10, 1–19. [Google Scholar] [CrossRef]
- Kairouz, P.; McMahan, H.B. Advances and Open Problems in Federated Learning. Found. Trends Mach. Learn. 2021, 14, 1–210. [Google Scholar] [CrossRef]
- Rauniyar, A.; Hagos, D.H.; Jha, D.; Håkegård, J.E.; Bagci, U.; Rawat, D.B.; Vlassov, V. Federated learning for medical applications: A taxonomy, current trends, challenges, and future research directions. IEEE Internet Things J. 2023, 11, 7374–7398. [Google Scholar] [CrossRef]
- Pati, S.; Kumar, S.; Varma, A.; Edwards, B.; Lu, C.; Qu, L.; Wang, J.J.; Lakshminarayanan, A.; Wang, S.h.; Sheller, M.J.; et al. Privacy preservation for federated learning in health care. Patterns 2024, 5. [Google Scholar] [CrossRef] [PubMed]
- Dasaradharami Reddy, K.; Gadekallu, T.R. A comprehensive survey on federated learning techniques for healthcare informatics. Comput. Intell. Neurosci. 2023, 2023, 8393990. [Google Scholar] [CrossRef] [PubMed]
- Naz, S.; Phan, K.T.; Chen, Y.P.P. A comprehensive review of federated learning for COVID-19 detection. Int. J. Intell. Syst. 2022, 37, 2371–2392. [Google Scholar] [CrossRef] [PubMed]
- Nazir, S.; Kaleem, M. Federated learning for medical image analysis with deep neural networks. Diagnostics 2023, 13, 1532. [Google Scholar] [CrossRef] [PubMed]
- Majeed, A.; Zhang, X.; Hwang, S.O. Applications and challenges of federated learning paradigm in the big data era with special emphasis on COVID-19. Big Data Cogn. Comput. 2022, 6, 127. [Google Scholar] [CrossRef]
- Naeem, A.; Anees, T.; Naqvi, R.A.; Loh, W.K. A comprehensive analysis of recent deep and federated-learning-based methodologies for brain tumor diagnosis. J. Pers. Med. 2022, 12, 275. [Google Scholar] [CrossRef] [PubMed]
- Chowdhury, A.; Kassem, H.; Padoy, N.; Umeton, R.; Karargyris, A. A review of medical federated learning: Applications in oncology and cancer research. In Proceedings of the International MICCAI Brainlesion Workshop; Springer, 2021; pp. 3–24. [Google Scholar]
- Rehman, M.H.U.; Hugo Lopez Pinaya, W.; Nachev, P.; Teo, J.T.; Ourselin, S.; Cardoso, M.J. Federated learning for medical imaging radiology. Br. J. Radiol. 2023, 96, 20220890. [Google Scholar] [CrossRef] [PubMed]
- Wang, J.; Ma, F. Federated learning for rare disease detection: a survey. Rare Dis. Orphan Drugs J. 2023, 2, N–A. [Google Scholar] [CrossRef]
- Diao, E.; Ding, J.; Tarokh, V. Heterofl: Computation and communication efficient federated learning for heterogeneous clients. arXiv 2020, arXiv:2010.01264. [Google Scholar]
- Wang, K.; He, Q.; Chen, F.; Chen, C.; Huang, F.; Jin, H.; Yang, Y. FlexiFed: Personalized Federated Learning for Edge Clients with Heterogeneous Model Architectures. In Proceedings of the Proceedings of the ACM Web Conference 2023 (WWW ’23); Association for Computing Machinery, 2023; pp. 2979–2990. [Google Scholar] [CrossRef]
- Fredrikson, M.; Jha, S.; Ristenpart, T. Model inversion attacks that exploit confidence information and basic countermeasures. In Proceedings of the Proceedings of the 22nd ACM SIGSAC conference on computer and communications security; 2015; pp. 1322–1333. [Google Scholar]
- Lian, X.; Zhang, C.; Zhang, H.; Hsieh, C.J.; Zhang, W.; Liu, J. Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent. Adv. Neural Inf. Process. Syst. 2017, 30. [Google Scholar]
- Han, P.; Wang, S.; Leung, K.K. Adaptive gradient sparsification for efficient federated learning: An online learning approach. In Proceedings of the 2020 IEEE 40th international conference on distributed computing systems (ICDCS); IEEE, 2020; pp. 300–310. [Google Scholar]
- Wang, J.; Liu, Q.; Liang, H.; Joshi, G.; Poor, H.V. Tackling the objective inconsistency problem in heterogeneous federated optimization. Adv. Neural Inf. Process. Syst. 2020, 33, 7611–7623. [Google Scholar]
- Ouyang, X.; Xie, Z.; Zhou, J.; Xing, G.; Huang, J. Clusterfl: A clustering-based federated learning system for human activity recognition. ACM Trans. Sens. Netw. 2022, 19, 1–32. [Google Scholar] [CrossRef]
- Chen, C.; Chen, Z.; Zhou, Y.; Kailkhura, B. Fedcluster: Boosting the convergence of federated learning via cluster-cycling. In Proceedings of the 2020 IEEE international conference on big data (Big Data); IEEE, 2020; pp. 5017–5026. [Google Scholar]
- Ruan, Y.; Joe-Wong, C. Fedsoft: Soft clustered federated learning with proximal local updating. Proc. Proc. AAAI Conf. Artif. Intell. 2022, 36, 8124–8131. [Google Scholar] [CrossRef]
- Palihawadana, C.; Wiratunga, N.; Wijekoon, A.; Kalutarage, H. Fedsim: Similarity guided model aggregation for federated learning. Neurocomputing 2022, 483, 432–445. [Google Scholar] [CrossRef]
- Ghosh, A.; Chung, J.; Yin, D.; Ramchandran, K. An efficient framework for clustered federated learning. Adv. Neural Inf. Process. Syst. 2020, 33, 19586–19597. [Google Scholar]
- Li, D.; Wang, J. Fedmd: Heterogenous federated learning via model distillation. arXiv 2019, arXiv:1910.03581. [Google Scholar]
- Lin, T.; Kong, L.; Stich, S.U.; Jaggi, M. Ensemble distillation for robust model fusion in federated learning. Adv. Neural Inf. Process. Syst. 2020, 33, 2351–2363. [Google Scholar]
- Chen, H.Y.; Chao, W.L. Fedbe: Making bayesian model ensemble applicable to federated learning. arXiv 2020, arXiv:2009.01974. [Google Scholar]
- Zhang, L.; Shen, L.; Ding, L.; Tao, D.; Duan, L.Y. Fine-tuning global model via data-free knowledge distillation for non-iid federated learning. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2022; pp. 10164–10173. [Google Scholar]
- Sattler, F.; Korjakow, T.; Rischke, R.; Samek, W. Fedaux: Leveraging unlabeled auxiliary data in federated learning. IEEE Trans. Neural Netw. Learn. Syst. 2021, 34, 5531–5543. [Google Scholar] [CrossRef] [PubMed]
- Li, D.; Xie, W.; Li, Y.; Fang, L. FedFusion: Manifold-driven federated learning for multi-satellite and multi-modality fusion. IEEE Trans. Geosci. Remote Sens. 2023, 62, 1–13. [Google Scholar] [CrossRef]
- Arivazhagan, M.G.; Aggarwal, V.; Singh, A.K.; Choudhary, S. Federated learning with personalization layers. arXiv 2019, arXiv:1912.00818. [Google Scholar]
- Li, Q.; He, B.; Song, D. Model-contrastive federated learning. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2021; pp. 10708–10717. [Google Scholar]
- Tan, Y.; Long, G.; Liu, L.; Zhou, T.; Lu, Q.; Jiang, J.; Zhang, C. Fedproto: Federated prototype learning across heterogeneous clients. Proc. Proc. AAAI Conf. Artif. Intell. 2022, 36, 8432–8440. [Google Scholar] [CrossRef]
- Collins, L.; Hassani, H.; Mokhtari, A.; Shakkottai, S. Exploiting shared representations for personalized federated learning. In Proceedings of the International conference on machine learning; PMLR, 2021; pp. 2089–2099. [Google Scholar]
- Wicaksana, J.; Yan, Z.; Zhang, D.; Huang, X.; Wu, H.; Yang, X.; Cheng, K.T. Fedmix: Mixed supervised federated learning for medical image segmentation. IEEE Trans. Med. Imaging 2022, 42, 1955–1968. [Google Scholar] [CrossRef] [PubMed]
- Zhao, Z.; Pang, S.; Wang, L.; Qiao, S.; Zhao, Y.; Wang, S.; Lyu, Z. MedFedProto: a semi-supervised classification framework for medical images based on federated prototypical learning. Expert Syst. With Appl. 2025, 130332. [Google Scholar]
- Pfeifer, B.; Chereda, H.; Martin, R.; Saranti, A.; Clemens, S.; Hauschild, A.C.; Beißbarth, T.; Holzinger, A.; Heider, D. Ensemble-GNN: federated ensemble learning with graph neural networks for disease module discovery and classification. Bioinformatics 2023, 39, btad703. [Google Scholar] [CrossRef] [PubMed]
- Ji, S.; Pan, S.; Long, G.; Li, X.; Jiang, J.; Huang, Z. Learning private neural language modeling with attentive aggregation. In Proceedings of the 2019 International joint conference on neural networks (IJCNN); IEEE, 2019; pp. 1–8. [Google Scholar]
- Huang, Y.; Chu, L.; Zhou, Z.; Wang, L.; Liu, J.; Pei, J.; Zhang, Y. Personalized cross-silo federated learning on non-iid data. Proc. Proc. AAAI Conf. Artif. Intell. 2021, 35, 7865–7873. [Google Scholar] [CrossRef]
- Hemker, K.; Simidjievski, N.; Jamnik, M. Healnet: Multimodal fusion for heterogeneous biomedical data. Adv. Neural Inf. Process. Syst. 2024, 37, 64479–64498. [Google Scholar] [CrossRef]
- Zhang, C.; Dang, S.; Shihada, B.; Alouini, M.S. Dual attention-based federated learning for wireless traffic prediction. In Proceedings of the IEEE INFOCOM 2021-IEEE conference on computer communications; IEEE, 2021; pp. 1–10. [Google Scholar]
- Ye, W.; An, X.; Wang, J.; Yan, X.; Carle, G. FedABC: Attention-Based Client Selection for Federated Learning with Long-Term View. In Proceedings of the ICC 2025-IEEE International Conference on Communications; IEEE, 2025; pp. 801–806. [Google Scholar]
- Ji, S.; Pan, S.; Long, G.; Li, X.; Jiang, J.; Huang, Z. Learning private neural language modeling with attentive aggregation. In Proceedings of the 2019 International joint conference on neural networks (IJCNN); IEEE, 2019; pp. 1–8. [Google Scholar]
- Chen, Y.; Sun, X.; Jin, Y. Federated Meta-Learning with Fast Convergence and Efficient Communication. arXiv 2018, arXiv:1802.07876. [Google Scholar]
- T Dinh, C.; Tran, N.; Nguyen, J. Personalized federated learning with moreau envelopes. Adv. Neural Inf. Process. Syst. 2020, 33, 21394–21405. [Google Scholar]
- Fallah, A.; Mokhtari, A.; Ozdaglar, A. Personalized Federated Learning: A Meta-Learning Approach. arXiv 2020, arXiv:2002.07948. [Google Scholar]
- Shamsian, A.; Navon, A.; Fetaya, E.; Chechik, G. Personalized federated learning using hypernetworks. In Proceedings of the International conference on machine learning; PMLR, 2021; pp. 9489–9502. [Google Scholar]
- Gao, J.; Li, Y. FedMetaMed: Federated meta-learning for personalized medication in distributed healthcare systems. In Proceedings of the 2024 IEEE international conference on bioinformatics and biomedicine (BIBM); IEEE, 2024; pp. 6384–6391. [Google Scholar]
- Zhang, T.; Zhang, S.; Chen, Z.; Bengio, Y.; Liu, D. PMFL: Partial Meta-Federated Learning for heterogeneous tasks and its applications on real-world medical records. In Proceedings of the 2022 IEEE International Conference on Big Data (Big Data); IEEE, 2022; pp. 4453–4462. [Google Scholar]
- Lai, F.; Zhu, X.; Madhyastha, H.V.; Chowdhury, M. Oort: Efficient federated learning via guided participant selection. In Proceedings of the 15th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 21), 2021; pp. 19–35. [Google Scholar]
- Acar, D.A.E.; Zhao, Y.; Navarro, R.M.; Mattina, M.; Whatmough, P.N.; Saligrama, V. Federated learning based on dynamic regularization. arXiv 2021, arXiv:2111.04263. [Google Scholar]
- Reddi, S.; Charles, Z.; Zaheer, M.; Garrett, Z.; Rush, K.; Konečnỳ, J.; Kumar, S.; McMahan, H.B. Adaptive federated optimization. arXiv 2020, arXiv:2003.00295. [Google Scholar]
- Gao, L.; Fu, H.; Li, L.; Chen, Y.; Xu, M.; Xu, C.Z. Feddc: Federated learning with non-iid data via local drift decoupling and correction. In Proceedings of the 2022 IEEE/CVF conference on computer vision and pattern recognition (CVPR); IEEE, 2022; pp. 10102–10111. [Google Scholar]
- Soni, A.; Mishra, R. Fair-select: a federated learning approach to ensure fairness in selection of participants. Multimed. Tools Appl. 2025, 84, 32193–32207. [Google Scholar] [CrossRef]
- Panda, A.; Mahloujifar, S.; Bhagoji, A.N.; Chakraborty, S.; Mittal, P. Sparsefed: Mitigating model poisoning attacks in federated learning with sparsification. In Proceedings of the International Conference on Artificial Intelligence and Statistics; PMLR, 2022; pp. 7587–7624. [Google Scholar]
- Hu, R.; Gong, Y.; Guo, Y. Federated learning with sparsification-amplified privacy and adaptive optimization. arXiv 2020, arXiv:2008.01558. [Google Scholar]
- Kim, D.Y.; Han, D.J.; Seo, J.; Moon, J. Achieving Lossless Gradient Sparsification via Mapping to Alternative Space in Federated Learning. In Proceedings of the ICML, 2024; pp. 23867–23900. [Google Scholar]
- Li, J.; Zhang, Y.; Li, Y.; Gong, X.; Wang, W. FedSparse: A communication-efficient federated learning framework based on sparse updates. Electronics 2024, 13, 5042. [Google Scholar] [CrossRef]
- Lu, S.; Li, R.; Liu, W.; Guan, C.; Yang, X. Top-k sparsification with secure aggregation for privacy-preserving federated learning. Comput. Secur. 2023, 124, 102993. [Google Scholar] [CrossRef]
- Hu, C.; Jiang, J.; Wang, Z. Decentralized federated learning: A segmented gossip approach. arXiv 2019, arXiv:1908.07782. [Google Scholar]
- Jeon, B.; Ferdous, S.; Rahman, M.R.; Walid, A. Privacy-preserving decentralized aggregation for federated learning. In Proceedings of the IEEE INFOCOM 2021-IEEE conference on computer communications workshops (INFOCOM WKSHPS); IEEE, 2021; pp. 1–6. [Google Scholar]
- Ramanan, P.; Nakayama, K. Baffle: Blockchain based aggregator free federated learning. In Proceedings of the 2020 IEEE international conference on blockchain (Blockchain); IEEE, 2020; pp. 72–81. [Google Scholar]
- Catak, F.O.; Kuzlu, M.; Dalveren, Y.; Ozdemir, G. Serverless federated learning: Decentralized spectrum sensing in heterogeneous networks. Phys. Commun. 2025, 70, 102634. [Google Scholar] [CrossRef]
- Li, Y.; Chen, C.; Liu, N.; Huang, H.; Zheng, Z.; Yan, Q. A blockchain-based decentralized federated learning framework with committee consensus. IEEE Netw. 2020, 35, 234–241. [Google Scholar] [CrossRef]
- Rückel, T.; Sedlmeir, J.; Hofmann, P. Fairness, integrity, and privacy in a scalable blockchain-based federated learning system. Comput. Netw. 2022, 202, 108621. [Google Scholar] [CrossRef]
- Zhang, C.; Li, S.; Xia, J.; Wang, W.; Yan, F.; Liu, Y. BatchCrypt:Efficient homomorphic encryption for Cross-Silo federated learning. In Proceedings of the 2020 USENIX annual technical conference (USENIX ATC 20), 2020; pp. 493–506. [Google Scholar]
- Fereidooni, H.; Marchal, S.; Miettinen, M.; Mirhoseini, A.; Möllering, H.; Nguyen, T.D.; Rieger, P.; Sadeghi, A.R.; Schneider, T.; Yalame, H.; et al. SAFELearn: Secure aggregation for private federated learning. In Proceedings of the 2021 IEEE security and privacy workshops (SPW); IEEE, 2021; pp. 56–62. [Google Scholar]
- Mo, F.; Haddadi, H.; Katevas, K.; Marin, E.; Perino, D.; Kourtellis, N. PPFL: Privacy-preserving federated learning with trusted execution environments. In Proceedings of the Proceedings of the 19th annual international conference on mobile systems, applications, and services, 2021; pp. 94–108. [Google Scholar]
- Damgård, I.; Pastro, V.; Smart, N.; Zakarias, S. Multiparty computation from somewhat homomorphic encryption. In Proceedings of the Annual cryptology conference; Springer, 2012; pp. 643–662. [Google Scholar]
- Demmler, D.; Schneider, T.; Zohner, M. ABY-A framework for efficient mixed-protocol secure two-party computation. In Proceedings of the Ndss, 2015. [Google Scholar]
- Patra, A.; Schneider, T.; Suresh, A.; Yalame, H. {ABY2. 0}: Improved {mixed-protocol} secure {two-party} computation. In Proceedings of the 30th USENIX Security Symposium (USENIX Security 21), 2021; pp. 2165–2182. [Google Scholar]
- Liu, F.; Zheng, Z.; Shi, Y.; Tong, Y.; Zhang, Y. A survey on federated learning: a perspective from multi-party computation. Front. Comput. Sci. 2024, 18, 181336. [Google Scholar] [CrossRef]
- Colosimo, F.; De Rango, F. Median-krum: A joint distance-statistical based byzantine-robust algorithm in federated learning. In Proceedings of the Proceedings of the Int’l ACM Symposium on Mobility Management and Wireless Access, 2023; pp. 61–68. [Google Scholar]
- Sánchez, P.M.S.; Celdrán, A.H.; Xie, N.; Bovet, G.; Pérez, G.M.; Stiller, B. Federatedtrust: A solution for trustworthy federated learning. Future Gener. Comput. Syst. 2024, 152, 83–98. [Google Scholar] [CrossRef]
- Shahul, U.; Harshan, J. FORTA: Byzantine-Resilient FL Aggregation via DFT-Guided Krum. In Proceedings of the 2025 IEEE Information Theory Workshop (ITW); IEEE, 2025; pp. 788–793. [Google Scholar]
- Wang, Z.; Dong, N.; Sun, J.; Knottenbelt, W.; Guo, Y. Zero-Knowledge Proof-Based Gradient Aggregation for Federated Learning. IEEE Trans. Big Data 2024, 11, 447–460. [Google Scholar] [CrossRef]
- Hahn, C.; Kim, H.; Kim, M.; Hur, J. Versa: Verifiable secure aggregation for cross-device federated learning. IEEE Trans. Dependable Secur. Comput. 2021, 20, 36–52. [Google Scholar] [CrossRef]
- Zhou, H.; Yang, G.; Huang, Y.; Dai, H.; Xiang, Y. Privacy-preserving and verifiable federated learning framework for edge computing. IEEE Trans. Inf. Forensics Secur. 2022, 18, 565–580. [Google Scholar] [CrossRef]
- Guo, X.; Liu, Z.; Li, J.; Gao, J.; Hou, B.; Dong, C.; Baker, T. Communication-efficient and fast verifiable aggregation for federated learning. IEEE Trans. Inf. Forensics Secur. 2020, 16, 1736–1751. [Google Scholar] [CrossRef]
- Lian, X.; Zhang, W.; Zhang, C.; Liu, J. Asynchronous decentralized parallel stochastic gradient descent. In Proceedings of the International conference on machine learning; PMLR, 2018; pp. 3043–3052. [Google Scholar]
- Horváth, S.; Laskaridis, S.; Almeida, M.; Leontiadis, I.; Venieris, S.I.; Lane, N.D. FjORD: Fair and Accurate Federated Learning under Heterogeneous Targets with Ordered Dropout. Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) 2021, 34, 12876–12889. [Google Scholar]
- Nguyen, J.; Malik, K.; Zhan, H.; Yousefpour, A.; Rabbat, M.; Malek, M.; Huba, D. Federated learning with buffered asynchronous aggregation. In Proceedings of the International conference on artificial intelligence and statistics; PMLR, 2022; pp. 3581–3607. [Google Scholar]
- Yu, X.; Cherkasova, L.; Vardhan, H.; Zhao, Q.; Ekaireb, E.; Zhang, X.; Mazumdar, A.; Rosing, T. Async-HFL: Efficient and robust asynchronous federated learning in hierarchical IoT networks. In Proceedings of the Proceedings of the 8th ACM/IEEE Conference on Internet of Things Design and Implementation, 2023; pp. 236–248. [Google Scholar]
- Wu, C.; Wu, F.; Lyu, L.; Huang, Y.; Xie, X. Communication-efficient federated learning via knowledge distillation. Nat. Commun. 2022, 13, 2032. [Google Scholar] [CrossRef] [PubMed]
- Karimireddy, S.P.; Kale, S.; Mohri, M.; Reddi, S.; Stich, S.; Suresh, A.T. Scaffold: Stochastic controlled averaging for federated learning. In Proceedings of the International conference on machine learning; PMLR, 2020; pp. 5132–5143. [Google Scholar]
- Elzohairy, M.; Chadha, M.; Jindal, A.; Grafberger, A.; Gu, J.; Gerndt, M.; Abboud, O. Fedlesscan: Mitigating stragglers in serverless federated learning. In Proceedings of the 2022 IEEE international conference on big data (Big Data); IEEE, 2022; pp. 1230–1237. [Google Scholar]
- Alistarh, D.; Grubic, D.; Li, J.; Tomioka, R.; Vojnovic, M. QSGD: Communication-efficient SGD via gradient quantization and encoding. Adv. Neural Inf. Process. Syst. 2017, 30. [Google Scholar]
- Kim, S.W.; Kim, S.; Kim, J.; Ji, S.; Lee, S.H. Fedwsq: Efficient federated learning with weight standardization and distribution-aware non-uniform quantization. In Proceedings of the 2025 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE, 2025; pp. 4616–4625. [Google Scholar]
- Reisizadeh, A.; Mokhtari, A.; Hassani, H.; Jadbabaie, A.; Pedarsani, R. Fedpaq: A communication-efficient federated learning method with periodic averaging and quantization. In Proceedings of the International conference on artificial intelligence and statistics; PMLR, 2020; pp. 2021–2031. [Google Scholar]
- Liu, H.; He, F.; Cao, G. Communication-efficient federated learning for heterogeneous edge devices based on adaptive gradient quantization. In Proceedings of the IEEE INFOCOM 2023-IEEE Conference on Computer Communications; IEEE, 2023; pp. 1–10. [Google Scholar]
- Bareilles, G.; Bouaziz, W.; Fageot, J.; El-Mhamdi, E.M. Byzantine Machine Learning: MultiKrum and an optimal notion of robustness. arXiv 2026, arXiv:2602.03899. [Google Scholar]
- Liu, Y.; Chen, C.; Lyu, L.; Wu, F.; Wu, S.; Chen, G. Byzantine-robust learning on heterogeneous data via gradient splitting. In Proceedings of the International Conference on Machine Learning; PMLR, 2023; pp. 21404–21425. [Google Scholar]
- Yan, Y.; Feng, C.M.; Li, Y.; Xie, J.; Chen, J.; Elhoseiny, M.; Hu, M.; Wu, K.; Zhu, L. Temporal model-based federated active medical image classification. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer, 2025; pp. 604–614. [Google Scholar]
- Ngo, K.H.; Östman, J.; Durisi, G.; Graell i Amat, A. Secure aggregation is not private against membership inference attacks. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases; Springer, 2024; pp. 180–198. [Google Scholar]
- Firdaus, M.; Larasati, H.T.; Hyune-Rhee, K. Blockchain-based federated learning with homomorphic encryption for privacy-preserving healthcare data sharing. Internet Things 2025, 31, 101579. [Google Scholar] [CrossRef]
- Pentyala, S.; Neophytou, N.; Nascimento, A.; De Cock, M.; Farnadi, G. Privfairfl: Privacy-preserving group fairness in federated learning. arXiv 2022, arXiv:2205.11584. [Google Scholar]
- Gao, W.; Lan, L.; Liu, Y.; Wang, R.; Fan, X. Shuffle-Diversity Collaborative Federated Learning for Imbalanced Medical Image Analysis. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer, 2025; pp. 572–582. [Google Scholar]
- Li, T.; Hu, S.; Beirami, A.; Smith, V. Ditto: Fair and robust federated learning through personalization. In Proceedings of the International conference on machine learning; PMLR, 2021; pp. 6357–6368. [Google Scholar]
- Li, J.; Zhu, T.; Ren, W.; Raymond, K.K. Improve individual fairness in federated learning via adversarial training. Comput. Secur. 2023, 132, 103336. [Google Scholar] [CrossRef]
- Roth, H.R.; Cheng, Y.; Wen, Y.; Yang, I.; Xu, Z.; Hsieh, Y.T.; Kersten, K.; Harouni, A.; Zhao, C.; Lu, K.; et al. Nvidia flare: Federated learning from simulation to real-world. arXiv 2022, arXiv:2210.13291. [Google Scholar]
- Beutel, D.J.; Topal, T.; Mathur, A.; Qiu, X.; Fernandez-Marques, J.; Gao, Y.; Sani, L.; Li, K.H.; Parcollet, T.; de GusmÃĢo, P.P.B.; et al. Flower: A friendly federated learning research framework. arXiv 2020, arXiv:2007.14390. [Google Scholar]
- Liu, Y.; Fan, T.; Chen, T.; Xu, Q.; Yang, Q. Fate: An industrial grade platform for collaborative learning with data protection. J. Mach. Learn. Res. 2021, 22, 1–6. [Google Scholar]
- Foley, P.; Sheller, M.J.; Edwards, B.; Pati, S.; Riviera, W.; Sharma, M.; Narayana Moorthy, P.; Wang, S.h.; Martin, J.; Mirhaji, P.; et al. OpenFL: the open federated learning library. Phys. Med. Biol. 2022, 67, 214001. [Google Scholar] [CrossRef] [PubMed]
- Ludwig, H.; Baracaldo, N.; Thomas, G.; Zhou, Y.; Anwar, A.; Rajamoni, S.; Ong, Y.; Radhakrishnan, J.; Verma, A.; Sinn, M.; et al. Ibm federated learning: an enterprise framework white paper v0. 1. arXiv 2020, arXiv:2007.10987. [Google Scholar]
- Li, X.; Xu, Z.; Fu, H. (Eds.) Federated Learning for Medical Imaging: Principles, Algorithms, and Applications; The MICCAI Society Book Series; Academic Press: Cambridge, MA, USA, 2025. [Google Scholar] [CrossRef]
- Zenk, M.; Baid, U.; Pati, S.; Linardos, A.; Edwards, B.; Sheller, M.; Foley, P.; Aristizabal, A.; Zimmerer, D.; Gruzdev, A.; et al. Towards fair decentralized benchmarking of healthcare AI algorithms with the Federated Tumor Segmentation (FeTS) challenge. Nat. Commun. 2025, 16, 6274. [Google Scholar] [CrossRef] [PubMed]
- Linardos, A.; Pati, S.; Baid, U.; Edwards, B.; Foley, P.; Ta, K.; Chung, V.; Sheller, M.; Khan, M.I.; Jafaritadi, M.; et al. The MICCAI Federated Tumor Segmentation (FeTS) Challenge 2024: Efficient and Robust Aggregation Methods for Federated Learning. arXiv 2025, arXiv:2512.06206. [Google Scholar]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. Imagenet classification with deep convolutional neural networks. Adv. Neural Inf. Process. Syst. 2012, 25. [Google Scholar]
- Simonyan, K.; Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv 2014, arXiv:1409.1556. [Google Scholar]
- Shafiq, M.; Gu, Z. Deep residual learning for image recognition: A survey. Appl. Sci. 2022, 12, 8972. [Google Scholar] [CrossRef]
- Bakas, S.; Reyes, M.; Jakab, A.; Bauer, S.; Rempfler, M.; Crimi, A.; Shinohara, R.T.; Berger, C.; Ha, S.M.; Rozycki, M.; et al. Identifying the best machine learning algorithms for brain tumor segmentation, progression assessment, and overall survival prediction in the BRATS challenge. arXiv 2018, arXiv:1811.02629. [Google Scholar]
- Rajpurkar, P.; Irvin, J.; Zhu, K.; Yang, B.; Mehta, H.; Duan, T.; Ding, D.; Bagul, A.; Langlotz, C.; Shpanskaya, K.; et al. Chexnet: Radiologist-level pneumonia detection on chest x-rays with deep learning. arXiv 2017, arXiv:1711.05225. [Google Scholar]
- Li, Q.; Wen, Z.; Wu, Z.; Hu, S.; Wang, N.; Li, Y.; Liu, X.; He, B. A survey on federated learning systems: Vision, hype and reality for data privacy and protection. IEEE Trans. Knowl. Data Eng. 2021, 35, 3347–3366. [Google Scholar] [CrossRef]
- Alahmari, S.; Alghamdi, I. A comprehensive survey on energy-efficient and privacy-preserving federated learning for edge intelligence and IoT. Results Eng. 2025, 107849. [Google Scholar] [CrossRef]
- Lim, W.Y.B.; Luong, N.C.; Hoang, D.T.; Jiao, Y.; Liang, Y.C.; Yang, Q.; Niyato, D.; Miao, C. Federated learning in mobile edge networks: A comprehensive survey. IEEE Commun. Surv. Tutor. 2020, 22, 2031–2063. [Google Scholar] [CrossRef]
- Lyu, L.; Yu, H.; Yang, Q. Threats to federated learning: A survey. arXiv 2020, arXiv:2003.02133. [Google Scholar]
- Sheller, M.J.; Edwards, B.; Reina, G.A.; Martin, J.; Pati, S.; Kotrotsou, A.; Milchenko, M.; Xu, W.; Marcus, D.; Colen, R.R.; et al. Federated learning in medicine: facilitating multi-institutional collaborations without sharing patient data. Sci. Rep. 2020, 10, 12598. [Google Scholar] [CrossRef] [PubMed]
- Mohri, M.; Sivek, G.; Suresh, A.T. Agnostic federated learning. In Proceedings of the International conference on machine learning; PMLR, 2019; pp. 4615–4625. [Google Scholar]
- Bonawitz, K.; Ivanov, V.; Kreuter, B.; Marcedone, A.; et al. Practical Secure Aggregation for Federated Learning on User-Held Data. In Proceedings of the Proceedings of NeurIPS, 2017. [Google Scholar]
- Yu, H.; Jin, R.; Yang, S. On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization. In Proceedings of the Proceedings of the 36th International Conference on Machine Learning (ICML) PMLR; Chaudhuri, K., Salakhutdinov, R., Eds.; Proceedings of Machine Learning Research, 2019; Vol. 97, pp. 7184–7193. [Google Scholar]
- Bian, J.; Xu, X.; Zhao, Y.; Li, J.; Zhang, Z.L. FedSeal: Semi-Supervised Federated Learning with Self-Ensemble Learning and Negative Learning. Proc. Proc. AAAI Conf. Artif. Intell. 2023, 37, 7720–7728. [Google Scholar]
- Li, Y.; Gui, J.; Deng, Z.; Meng, F.; Wu, Y. Fedqs: Optimizing gradient and model aggregation for semi-asynchronous federated learning. Adv. Neural Inf. Process. Syst. 2026, 38, 86502–86550. [Google Scholar]
- Bach, F.; Moulines, E. Non-strongly-convex smooth stochastic approximation with convergence rate O (1/n). Adv. Neural Inf. Process. Syst. 2013, 26. [Google Scholar]
- Lee, G.; Jeong, M.; Shin, Y.; Bae, S.; Yun, S.Y. Preservation of the global knowledge by not-true distillation in federated learning. Adv. Neural Inf. Process. Syst. 2022, 35, 38461–38474. [Google Scholar] [CrossRef]
- Wang, H.; Yurochkin, M.; Sun, Y.; Papailiopoulos, D.; Khazaeni, Y. Federated learning with matched averaging. arXiv 2020, arXiv:2002.06440. [Google Scholar]
- Siddika, F.; Hossen, M.A.; Munoz, J.P.; Roosta, T.G.; Sharma, A.; Jannesari, A. FedReFT: Federated Representation Fine-Tuning with All-But-Me Aggregation. Proc. Find. Assoc. Comput. Linguist. EACL 2026, 2026, 4341–4362. [Google Scholar] [CrossRef]
- Gupta, A.; Misra, S.; Pathak, N.; Das, D. FedCare: Federated learning for resource-constrained healthcare devices in IoMT system. IEEE Trans. Comput. Soc. Syst. 2023, 10, 1587–1596. [Google Scholar] [CrossRef]
- Ahmed, S.T.; Vinoth Kumar, V.; Mahesh, T.; Narasimha Prasad, L.; Velmurugan, A.; Muthukumaran, V.; Niveditha, V. FedOPT: federated learning-based heterogeneous resource recommendation and optimization for edge computing. Soft Comput. 2024, 1–12. [Google Scholar] [CrossRef]
- Gupta, K.; Fournarakis, M.; Reisser, M.; Louizos, C.; Nagel, M. Quantization robust federated learning for efficient inference on heterogeneous devices. arXiv 2022, arXiv:2206.10844. [Google Scholar]
- Cui, H.; Qu, Z.; Ye, B.; Tang, B.; Zhuang, T.; Wang, X.; Zeng, Y. Lightweight Adaptive Quantization Algorithms for Federated Learning with Heterogeneous Clients. IEEE Transactions on Mobile Computing 2025. [Google Scholar] [CrossRef]
- Chen, J.; Ma, B.; Cui, H.; Xia, Y. FedEvi: Improving federated medical image segmentation via evidential weight aggregation. In Proceedings of the International conference on medical image computing and computer-assisted intervention; Springer, 2024; pp. 361–372. [Google Scholar]
Figure 1.
Comparison of centralized learning and federated learning (FL) workflows. (a) In centralized learning, hospitals transfer raw patient data to a central server, where a single global model is trained and then deployed back to each site. (b) In FL, patient data never leaves the acquiring institution/hospital. Each site trains the model locally and shares only model updates, which the server iteratively aggregates into an improved global consensus model that is broadcast back to each collaborating institution.
Figure 1.
Comparison of centralized learning and federated learning (FL) workflows. (a) In centralized learning, hospitals transfer raw patient data to a central server, where a single global model is trained and then deployed back to each site. (b) In FL, patient data never leaves the acquiring institution/hospital. Each site trains the model locally and shares only model updates, which the server iteratively aggregates into an improved global consensus model that is broadcast back to each collaborating institution.

Figure 2.
Fan chart taxonomy of FL aggregation (FLAg) strategies in healthcare. Each colored ring represents an aggregation family, and each radial column corresponds to a technical or data-related challenge: from stragglers and asynchronous training to fairness and data imbalance. A color-filled cell indicates that the aggregation family has been reported to address that challenge, giving a compact map of the literature covered in this review.
Figure 2.
Fan chart taxonomy of FL aggregation (FLAg) strategies in healthcare. Each colored ring represents an aggregation family, and each radial column corresponds to a technical or data-related challenge: from stragglers and asynchronous training to fairness and data imbalance. A color-filled cell indicates that the aggregation family has been reported to address that challenge, giving a compact map of the literature covered in this review.

Figure 3.
Technical (a-e) and data-related (f-i) challenges in healthcare FL and the aggregation families that mitigate them. (a) Stragglers: Slow clients stall synchronous training rounds. Optimization strategies normalize updates so fast and slow sites contribute fairly, clustering groups clients by capability, and representative and peer to peer strategies keep contributions useful without waiting for the slowest site.(b) Asynchronous: Clients train and communicate at different rates. Parameter and Gradient based strategies directly aggregate information, while knowledge distillation, meta-learning and representation-based strategies enable knowledge sharing without much effect on the stalesness (c) Varying Resource Environment: Institutions differ in compute power, network quality, and devices. Parameter, clustering, ensemble, and attention strategies balance uneven contributions, knowledge distillation lets different model designs share predictions, and peer to peer exchange avoids stalling on weak sites. (d) Quantization Aware: Compressed updates save bandwidth but add noise. Gradient based strategies send low bit updates, and sparsification strategies transmit only the most informative parts of the model. (e) Security: Malicious clients or servers threaten model integrity. Byzantine strategies filter out harmful updates, confidential strategies hide individual updates from the server, and meta learning limits what is shared. (f) Uncertainty: Medical data is noisy and unreliable across sites. Parameter strategies select reliable clients, ensemble strategies capture model level uncertainty, and attention strategies reduce the weight of unreliable contributions. (g) Privacy: Shared updates can leak patient information. Confidential strategies hide individual updates, knowledge distillation shares predictions instead of parameters, and sparsification, representative, ensemble, meta learning, and peer to peer strategies reduce exposure by sharing less. (h) Data Imbalance: Clients differ in dataset size and class distribution. Parameter and ensemble strategies balance client contributions, clustering and representative strategies group and align similar sites, attention favors under represented clients, and meta learning adapts to each local dataset. (i) Fairness: The global model may favor some patient groups over others. Parameter strategies weight contributions fairly, clustering groups similar institutions, and representative strategies align features across sites.
Figure 3.
Technical (a-e) and data-related (f-i) challenges in healthcare FL and the aggregation families that mitigate them. (a) Stragglers: Slow clients stall synchronous training rounds. Optimization strategies normalize updates so fast and slow sites contribute fairly, clustering groups clients by capability, and representative and peer to peer strategies keep contributions useful without waiting for the slowest site.(b) Asynchronous: Clients train and communicate at different rates. Parameter and Gradient based strategies directly aggregate information, while knowledge distillation, meta-learning and representation-based strategies enable knowledge sharing without much effect on the stalesness (c) Varying Resource Environment: Institutions differ in compute power, network quality, and devices. Parameter, clustering, ensemble, and attention strategies balance uneven contributions, knowledge distillation lets different model designs share predictions, and peer to peer exchange avoids stalling on weak sites. (d) Quantization Aware: Compressed updates save bandwidth but add noise. Gradient based strategies send low bit updates, and sparsification strategies transmit only the most informative parts of the model. (e) Security: Malicious clients or servers threaten model integrity. Byzantine strategies filter out harmful updates, confidential strategies hide individual updates from the server, and meta learning limits what is shared. (f) Uncertainty: Medical data is noisy and unreliable across sites. Parameter strategies select reliable clients, ensemble strategies capture model level uncertainty, and attention strategies reduce the weight of unreliable contributions. (g) Privacy: Shared updates can leak patient information. Confidential strategies hide individual updates, knowledge distillation shares predictions instead of parameters, and sparsification, representative, ensemble, meta learning, and peer to peer strategies reduce exposure by sharing less. (h) Data Imbalance: Clients differ in dataset size and class distribution. Parameter and ensemble strategies balance client contributions, clustering and representative strategies group and align similar sites, attention favors under represented clients, and meta learning adapts to each local dataset. (i) Fairness: The global model may favor some patient groups over others. Parameter strategies weight contributions fairly, clustering groups similar institutions, and representative strategies align features across sites.

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.