Submitted:
17 September 2026
Posted:
18 September 2026
You are already at the latest version
Abstract
Earth observation (EO) is entering a petabyte-scale era in which multi-sensor satellite archives capture environmental processes across broad spatial, spectral, and temporal dimensions. EO data-processing workflows, however, still rely heavily on task-specific supervised models that require large amounts of labeled data and may generalize poorly when applied to new regions, time periods, sensors, or environmental conditions. This review provides a structured synthesis of the rapidly evolving literature on EO foundation models (EO-FMs), which seek to address these limitations through large-scale pretraining and efficient downstream adaptation. We organize the field of EO-FMs through a data-objective-architecture-adaptation-evidence lens that links training data and harmonization, pretraining objectives, backbone design, adaptation strategies, and empirical evidence across downstream applications. We examine representative model families including Prithvi-EO-2.0, Copernicus-FM, TerraFM, DOFA, SatCLIP, EarthPT, and Aurora, and compare their capabilities across tasks, regions, time periods, and sensing modalities. Across the literature, EO-FMs show strong potential for improving label efficiency, supporting multimodal learning, and enabling the reuse of pretrained representations across applications such as land-cover mapping, hazard response, ecosystem monitoring, urban climate analysis, and Earth-system forecasting. At the same time, empirical evidence on downstream performance, label efficiency, and generalization remains heterogeneous across datasets, sensing modalities, metrics, and benchmark protocols, which limits direct model ranking and underscores the need for task-aware interpretation. We argue that the next phase of EO-FM research should move beyond proof-of-concept gains toward more rigorous and capability-aware benchmarking, explicit assessment of predictive uncertainty and calibration, transparent preprocessing and adaptation procedures, and geographically and temporally robust evaluation. Meeting these requirements would strengthen the reliability, comparability, and scientific utility of EO-FMs for scalable and decision-relevant Earth observation analysis.
Keywords:
earth observation foundation models
; remote sensing
; self-supervised learning
; transfer learning
; multimodal learning
; geospatial artificial intelligence
; hazard monitoring
; urban climate
; benchmarking
1. Introduction
Earth observation (EO) has become fundamental to environmental monitoring, disaster response, ecosystem assessment, agricultural intelligence, and climate-risk analysis because modern satellite systems provide dense and repeated observations across space, time, and diverse sensing modalities. EO archives increasingly encompass optical, synthetic aperture radar (SAR), thermal, hyperspectral, and other geophysical observations, providing complementary information about Earth-surface and atmospheric processes. At the same time, multi-mission harmonization has improved the consistency and temporal density of long-term satellite records. For example, the Harmonized Landsat and Sentinel-2 (HLS) Version 2.0 surface reflectance dataset provides harmonized Landsat and Sentinel-2 observations at 30 m spatial resolution and substantially increases observation frequency relative to a single sensor (Ju et al., 2025).
However, the growing abundance of EO data has not automatically produced analytical models that generalize reliably across different deployment conditions. Many remote-sensing workflows may still rely on task-specific supervised models trained for a particular application, sensor, or geographic setting using labeled data that can be costly to produce and may vary in availability and consistency across regions and institutions. Domain shift occurs when the statistical characteristics of the data encountered during model deployment differ from those represented in the training data; in EO, such shifts can arise from changes in geography, seasonality, acquisition geometry, atmospheric conditions, land-cover composition, spatial resolution, or sensor characteristics (Ma et al., 2024). Under these conditions, models that perform strongly within one benchmark or study region may show reduced performance when applied elsewhere. The combination of costly labeling requirements and limited generalization has therefore motivated growing interest in EO foundation models (EO-FMs), for which large-scale pretraining offers a pathway toward more reusable representations and reduced dependence on task-specific labels (Lacoste et al., 2023; Xiong et al., 2024).
The broader foundation-model paradigm emerged from natural language processing and computer vision, where transformer architectures and large-scale self-supervised or weakly supervised pretraining enabled learned representations to be reused across multiple downstream tasks (Vaswani et al., 2017; He et al., 2022; Radford et al., 2021). EO, however, is not simply another image domain. EO observations are often multispectral or multimodal and are strongly conditioned by sensor physics, spatial resolution, revisit cycles, acquisition geometry, geolocation, and temporal context. These characteristics complicate the direct application of generic vision foundation models to EO and motivate sensor-aware input representations, harmonization procedures, spatiotemporal encoding, and evaluation protocols tailored to Earth-monitoring problems (Lacoste et al., 2023; Wang et al., 2025; Xiong et al., 2024).
Recent studies illustrate the diversity of design strategies emerging within EO-FMs. Prithvi-EO-2.0 was pretrained on 4.2 million global HLS time-series samples, and its temporal-location variants incorporate acquisition-date and geolocation metadata through temporal and location embeddings to support multi-temporal EO applications (Szwarcman et al., 2026). Copernicus-FM advances toward a unified multi-mission Copernicus framework with metadata-aware encoding across Sentinel-based inputs (Wang et al., 2025). TerraFM is designed for joint SAR-optical representation learning from large-scale Sentinel-1 and Sentinel-2 imagery, whereas Dynamic One-For-All (DOFA) uses wavelength-conditioned dynamic parameterization to support flexible processing across multiple sensors and spectral configurations (Danish et al., 2026; Xiong et al., 2024). Other approaches broaden the EO-FM paradigm further: SatCLIP learns general-purpose geographic representations by aligning satellite imagery with location information (Klemmer et al., 2025), EarthPT treats EO as a sequence-modeling problem through autoregressive forecasting (Smith et al., 2023), and Aurora extends foundation-model approaches to Earth-system prediction using more than one million hours of geophysical data (Bodnar et al., 2025).
Despite this progress, the field remains uneven in both maturity and comparability. EO-FMs are often evaluated using different datasets, sensing modalities, spatial resolutions, downstream tasks, adaptation procedures, and performance metrics, making broad claims of model superiority difficult to justify. GEO-Bench was introduced partly to improve comparability through curated classification and segmentation tasks and standardized evaluation procedures for Earth-monitoring models (Lacoste et al., 2023). Even so, EO-FMs differ substantially in their intended capabilities: some emphasize multi-temporal optical representation learning, others multimodal SAR-optical learning, cross-sensor flexibility, geographic representation, or forecasting. Consequently, model suitability should be interpreted relative to the required sensing modalities, temporal structure, downstream task, adaptation strategy, and evaluation setting rather than through a single headline performance score.
Against this background, this review examines EO-FMs from five interconnected perspectives: training data and data preparation, pretraining objectives, architectural design, downstream adaptation and generalization, and empirical evidence. Specifically, the review aims to: (i) clarify the data and design constraints that distinguish EO-FMs from conventional task-specific remote-sensing models; (ii) explain how heterogeneous EO observations are encoded into reusable representations; (iii) compare representative EO-FM families across optical, multimodal, geographic-representation, and forecasting-oriented settings; and (iv) identify persistent gaps in benchmarking, generalization, predictive uncertainty and calibration, interpretability, reproducibility, and operational deployment. We argue that EO-FMs should be understood not as a single model class, but as an emerging family of pretrained geospatial representation systems whose scientific and operational value depends on the interaction among pretraining data, model design, downstream adaptation, and rigorous empirical evaluation.
2. Review Methodology
This review followed a structured literature review approach (Snyder, 2019), using predefined search, screening and synthesis procedures to identify and compare studies on Earth observation foundation models (EO-FMs). Because EO-FMs constitute a rapidly evolving research area, the search focused primarily on literature published between January 2020 and August 2026. A limited number of earlier foundational studies on transformer architectures, self-supervised learning, masked autoencoders, and vision-language learning were also retained when necessary to establish the conceptual and methodological background of EO-FMs. The review aimed not only to summarize individual studies but also to compare how EO-FMs differ in their training data, pretraining objectives, architectural design, downstream adaptation strategies, and empirical evaluation across EO applications.
2.1. Search Strategy
Relevant literature was identified through searches of Google Scholar, Scopus, Web of Science, IEEE Xplore, arXiv, and publisher websites. The search combined terms related to Earth observation, remote sensing, foundation models, self-supervised learning, and downstream adaptation. Representative search terms included “earth observation foundation model,” “remote sensing foundation model,” “geospatial foundation model,” “satellite foundation model,” “multimodal remote sensing foundation model,” “self-supervised remote sensing transformer,” “remote sensing masked autoencoder,” and “EO foundation model benchmarking.” These terms were also combined, where appropriate, with application-oriented terms such as “land-cover mapping,” “flood mapping,” “wildfire mapping,” “urban climate,” “ecosystem monitoring,” “image retrieval,” and “forecasting.”
The literature search was last updated in August 2026. The database and source searches identified 312 records before duplicate removal. Backward and forward citation tracking of relevant model, benchmark, and review studies identified 107 additional records not captured by the initial keyword searches. Thus, 419 records were identified in total. After duplicate removal and screening, 27 studies were retained in the final review corpus. The complete set of included studies is provided in Supplementary Table S1.
2.2. Inclusion and Exclusion Criteria
Studies were included if they focused on EO-FMs or closely related large-scale pretrained representation-learning systems for remote sensing, provided sufficient methodological detail for comparison, and contributed empirical or methodological evidence relevant to downstream adaptation, representation reuse across applications, or generalization across regions, time periods, or sensing modalities. Studies were excluded if they focused exclusively on conventional task-specific supervised models, discussed general-purpose AI foundation models without a clear EO connection, lacked sufficient empirical or methodological information, or did not contribute directly to the analytical scope of the review. The detailed eligibility criteria are summarized in Table 1.
2.3. Screening and Study Selection
The literature selection process was conducted in three stages. First, duplicate and clearly overlapping records were removed. Second, titles and abstracts were screened to assess their relevance to EO foundation models and related pretrained geospatial representation-learning systems. Third, the full texts of potentially eligible studies were assessed against the inclusion and exclusion criteria. During full-text screening, particular attention was given to studies addressing one or more of the following dimensions: training data and data preparation, pretraining objectives, architectural design, downstream adaptation and generalization, and empirical evidence. Following screening, 27 studies were retained for the final synthesis. The final corpus was organized into two analytical groups. Core model studies formed the basis of the principal comparisons of EO-FM design, adaptation, and performance, whereas contextual studies supported broader discussion of benchmarking, interpretability, reproducibility, uncertainty, governance, and operational deployment.
2.4. Data Extraction and Synthesis
For each included study, information was extracted on model name, publication year, sensing modality, pretraining dataset, pretraining objective, backbone architecture, downstream adaptation strategy, application domain, evaluation setting, reported performance metrics, and stated limitations. Where reported, information on label efficiency, cross-region or cross-sensor generalization, temporal generalization, computational requirements, uncertainty or calibration, and model accessibility was also recorded. The extracted information was synthesized across five interconnected dimensions: training data and data preparation, pretraining objectives, architectural design, downstream adaptation and generalization, and empirical evidence. This structure enabled comparison of how EO-FMs ingest and represent heterogeneous EO observations, incorporate spatial and temporal context, adapt pretrained representations to downstream applications, and are evaluated across different experimental settings.
2.5. Reporting Approach
Because EO-FM studies differ substantially in datasets, sensing modalities, downstream tasks, adaptation procedures, performance metrics, and evaluation protocols, numerical results were not treated as directly comparable across all studies. Instead, the review adopts a task-aware comparative evidence synthesis to assess the reported strengths, limitations, and suitability of representative EO-FMs for applications such as multispectral classification, SAR–optical fusion, hazard mapping, and image retrieval. Direct numerical comparisons are emphasized only where models were evaluated under sufficiently comparable tasks, datasets, and protocols. This approach avoids presenting a universal model ranking when reported performance differences may instead reflect variation in benchmark design, modality coverage, preprocessing, adaptation strategy, or evaluation procedure.
3. Background and Analytical Framework
Recent surveys have documented the rapid expansion of foundation-model approaches across remote-sensing applications, architectures, and learning paradigms (Huo et al., 2025). Figure 1 summarizes the analytical framework used in this review to compare Earth observation foundation models (EO-FMs). Rather than prescribing a universal architecture for EO-FMs, the framework organizes the literature around five interconnected dimensions: training data and data preparation, pretraining objectives, architectural design, downstream adaptation and generalization, and empirical evidence. EO-FMs build on large-scale pretraining to learn reusable geospatial representations from broad and heterogeneous EO archives. Unlike conventional remote-sensing models, which are commonly developed for a specific task, sensor, or study region, EO-FMs are intended to support multiple downstream applications through reuse and adaptation of pretrained representations. Their scientific and operational value therefore depends not only on model scale, but also on the representativeness of the pretraining data, treatment of sensor and temporal information, downstream adaptation strategy, and robustness under geographic, temporal, and sensor shifts (Lacoste et al., 2023; Szwarcman et al., 2026; Wang et al., 2025; Xiong et al., 2024).
3.1. From Task-Specific GeoAI to EO Foundation Models
Historically, deep learning in remote sensing has largely been organized around task-specific applications, such as optical land-cover classification, SAR-based flood segmentation, crop-type mapping, and burn-scar detection. Such models can perform well within the domain for which they are trained but often require substantial labeled data and additional training when applied to new applications or deployment conditions.
The foundation-model paradigm offers a different workflow by shifting emphasis from independently training a model for each application toward broad pretraining followed by downstream adaptation. Pretraining is often self-supervised or weakly supervised, allowing models to exploit large EO archives without requiring labels for every observation. However, EO foundation models are not necessarily label-free. Many downstream applications, including land-cover classification and semantic segmentation, still require labeled examples to train a task-specific prediction head or to fine-tune the pretrained model. The relevant question is therefore not whether EO-FMs eliminate labeled data, but whether pretraining enables representations to be reused across applications with lower task-specific labeling requirements or improved performance compared with training separate models from scratch (Lacoste et al., 2023; Szwarcman et al., 2026; Xiong et al., 2024).
3.2. Core Concepts and Working Definitions
Because the EO-FM literature is still evolving, clear working definitions are necessary to support consistent comparison across studies. In this review, an EO foundation model refers to a pretrained model, commonly but not exclusively transformer-based, that learns reusable representations from broad EO datasets and is intended for adaptation to multiple downstream applications. Pretraining may use self-supervised, weakly supervised, supervised, or combined learning signals depending on the model design. This review distinguishes representation reuse, downstream adaptation, and generalization because these concepts describe different properties of an EO-FM. Representation reuse concerns whether a pretrained backbone or embedding can support multiple applications; downstream adaptation concerns how the pretrained model is specialized for a particular task; and generalization concerns whether the resulting model maintains performance under changed deployment conditions. Table 2 summarizes the working definitions used throughout the review.
3.3. Why EO Data Are Not “Just Another Vision Dataset”
EO data differ fundamentally from conventional natural-image datasets. First, EO observations commonly extend beyond RGB imagery to multispectral and hyperspectral measurements and may include modalities such as synthetic aperture radar (SAR), thermal imagery, elevation data, and atmospheric observations. Second, EO observations are geolocated and time-stamped, so location, seasonality, revisit intervals, and acquisition geometry often carry information that is directly relevant to interpretation. Third, observations from different sensors are not automatically comparable because platforms differ in spectral response functions, spatial resolution, viewing geometry, calibration, and preprocessing history.
In this review, harmonization refers to procedures that reduce such inconsistencies so that observations from different sensors or acquisition periods can be compared or jointly analyzed. These procedures can include geometric co-registration, radiometric adjustment, spatial-resolution alignment, temporal matching, and consistent metadata or preprocessing conventions. Harmonization is therefore distinct from multimodal fusion, which combines complementary information from different data sources or sensing modalities within a model. These characteristics complicate direct reuse of generic computer-vision models and motivate EO-specific approaches to sensor-aware encoding, harmonization, multimodal fusion, and benchmark design (Ju et al., 2025; Lacoste et al., 2023; Xiong et al., 2024; Wang et al., 2025).
3.4. Data Modalities, Harmonization and Multimodal Fusion
In EO-FMs, the composition and preparation of the pretraining corpus are first-order design choices because they determine which spatial, spectral, temporal, and cross-sensor relationships a model can learn. Harmonized optical products are particularly useful for multi-mission pretraining because they reduce systematic differences among observations while increasing temporal coverage. The Harmonized Landsat and Sentinel-2 (HLS) Version 2.0 dataset, for example, provides globally harmonized Landsat and Sentinel-2 surface-reflectance observations at 30 m resolution (Ju et al., 2025). Prithvi-EO-2.0 builds on HLS through pretraining with 4.2 million global time-series samples and its temporal-locational variants incorporate temporal and location embeddings. SSL4EO-L demonstrates large-scale Landsat-based self-supervised pretraining using approximately 5 million image patches (Szwarcman et al., 2026; Stewart et al., 2023).
Multimodal fusion addresses a different problem. Rather than making observations from different sensors directly comparable, fusion seeks to exploit complementary information contained in different sensing modalities. Copernicus-FM learns across multiple Copernicus data streams using sensor and metadata information, whereas TerraFM combines Sentinel-1 SAR and Sentinel-2 optical observations through modality-specific representations and cross-attention mechanisms (Wang et al., 2025; Danish et al., 2026). Thus, harmonization and fusion should be treated as related but distinct components of EO-FM data design: harmonization improves comparability, whereas fusion integrates complementary information.
3.5. Pretraining Objectives
EO-FMs adapt several families of pretraining objectives from computer vision and sequence modeling to the spatial, spectral, temporal, and multimodal characteristics of EO data. One major family is masked reconstruction, in which portions of the input are intentionally removed and the model learns to reconstruct the missing information. Masked autoencoding in vision commonly operates by masking spatial image patches and reconstructing missing information (He et al., 2022). EO-specific extensions can additionally exploit the temporal and multispectral structure of satellite imagery. SatMAE, for example, independently masks image patches across time and uses distinct spectral positional encodings for grouped multispectral bands (Cong et al., 2022). Cross-modal masked learning extends this idea to multiple input sources; MultiMAE, for example, uses masking and reconstruction to learn predictive relationships across modalities (Bachmann et al., 2022).
Masked reconstruction can provide rich spatial, spectral, and temporal representations and may support downstream applications involving incomplete or heterogeneous observations. However, successful reconstruction of artificial masks during pretraining should not automatically be interpreted as an ability to perform cloud removal or reconstruct cloud-obscured observations. Such capabilities require task-specific training or direct evaluation.
A second family is contrastive alignment, in which representations of related observations are encouraged to become similar while unrelated observations remain distinguishable. In EO, such relationships may involve imagery, geographic coordinates, text, acquisition times, or complementary sensing modalities. SatCLIP, for example, aligns satellite imagery with geographic coordinates to learn reusable location representations (Klemmer et al., 2025).
A third family is temporal prediction and forecasting, in which models learn temporal dynamics by predicting future observations or states. EarthPT, for example, uses autoregressive prediction of reflectance sequences as a learning objective (Smith et al., 2023).
The relationship between pretraining objectives and downstream applications is not one-to-one. Nevertheless, several practical associations can be identified. Masked reconstruction often produces spatial and spectral representations useful for classification and segmentation; contrastive and multimodal alignment are particularly relevant to retrieval, matching, and geographically or semantically aligned applications; and temporal-prediction objectives are naturally suited to forecasting and time-series analysis. The same pretrained objective can nevertheless support multiple downstream applications after appropriate adaptation, so these associations should be interpreted as tendencies rather than fixed rules.
3.6. Downstream Adaptation, Representation Reuse, and Generalization
Rather than treating transfer as a single concept, this review distinguishes representation reuse, downstream adaptation, and generalization. Representation reuse refers to applying the same pretrained backbone or embeddings to different downstream applications, such as classification, segmentation, regression, or change detection. The ability of one pretrained backbone to support several applications is therefore described here as reuse across downstream tasks rather than as evidence of generalization.
Downstream adaptation refers to the mechanism used to specialize a pretrained model for a target application. Common approaches include linear probing, full fine-tuning, and parameter-efficient methods such as adapters, learnable prompts where supported, or low-rank parameter updates. Linear probing typically freezes the pretrained encoder and trains only a lightweight prediction layer, whereas full fine-tuning updates all or most model parameters. Parameter-efficient adaptation modifies only a relatively small subset of parameters, potentially reducing computational and storage requirements. Generalization, by contrast, concerns whether the adapted model remains reliable when evaluated under changed conditions, such as new geographic regions, seasons or years, sensors, spectral configurations, or environmental settings. These different forms of generalization should be assessed explicitly rather than inferred solely from performance across multiple tasks.
Different EO-FMs provide different mechanisms and capabilities for representation reuse and adaptation. Prithvi-EO-2.0 emphasizes multi-temporal optical EO representations, Copernicus-FM supports heterogeneous sensor and metadata inputs, DOFA uses wavelength-conditioned mechanisms to accommodate different spectral configurations, and SatCLIP provides geographic embeddings that can be incorporated into spatially structured prediction tasks (Szwarcman et al., 2026; Wang et al., 2025; Xiong et al., 2024; Klemmer et al., 2025). Consequently, the usefulness of an EO-FM should be assessed relative to the intended downstream application, available data, and deployment conditions rather than treated as a universal model property.
4. Functional Architecture and Downstream Adaptation of EO Foundation Models
Earth observation foundation models (EO-FMs) operate within a broader workflow that extends from EO data preparation to downstream prediction. It is important, however, to distinguish operations performed before data enter the model from those that constitute the EO-FM core. Radiometric processing, geometric co-registration, cloud and quality masking, resampling, and related data-preparation procedures are generally external to the model, whereas input representation, pretrained representation learning, and optional multimodal or contextual integration occur within the EO-FM. Downstream adaptation and task-specific output generation then specialize the pretrained representation for a target application. Figure 2 summarizes these functional stages and distinguishes external data preparation, the EO-FM core, downstream adaptation, evaluation checkpoints, generalization evidence, and application contexts.
4.1. End-to-End EO-FM Workflow
A typical EO-FM workflow can be represented by seven functional components: EO inputs, external data preparation, model-internal input representation, pretrained representation learning, optional fusion or contextual integration, downstream adaptation with a task-specific interface, and downstream tasks or outputs. Within this workflow, EO inputs and data preparation remain external to the foundation model; input representation, pretrained representation learning, and optional fusion or contextual integration constitute the EO-FM core; and downstream adaptation specializes the pretrained representation for a target application. Depending on the sensor and application, external data preparation may include radiometric correction or normalization, geometric co-registration, cloud and quality masking, spatial resampling, temporal compositing, tiling, and harmonization where appropriate. These operations prepare observations for model ingestion but should not be interpreted as components of the foundation-model architecture itself.
Prepared observations are subsequently converted into model-compatible representations using operations such as patch embedding, band-aware encoding, wavelength-conditioned encoding, modality-specific embedding, or temporal and geospatial tokenization. These representations are processed by a pretrained backbone to produce reusable latent features. In multimodal or metadata-rich EO-FMs, optional fusion or contextual modules may additionally integrate information from different sensing modalities, acquisition times, geographic locations, or other metadata. The presence of an encoder does not imply that every EO-FM requires a corresponding general-purpose decoder. A decoder is used when required by the pretraining objective or downstream task. Masked-autoencoding approaches, for example, commonly employ a reconstruction decoder during pretraining, whereas classification may require only a prediction head and semantic segmentation may use a task-specific decoder. The task interface is therefore determined by the target application rather than being a universal component of every EO-FM.
4.2. Sensor-Aware Input Representation
Sensor-aware input representation is particularly important for EO-FMs intended to process heterogeneous sensors or sensing modalities. EO observations are not standardized in the same manner as conventional RGB image datasets: optical, SAR, thermal, hyperspectral, and atmospheric measurements differ in spectral response, spatial resolution, acquisition geometry, noise characteristics, and physical interpretation. Applying identical input mappings to these heterogeneous observations can therefore discard information that is meaningful for downstream analysis. Different EO-FMs address this challenge in different ways. DOFA uses wavelength-conditioned dynamic parameterization to accommodate varying spectral configurations and support representation learning across multiple sensors (Xiong et al., 2024). Copernicus-FM incorporates dynamic hypernetworks and metadata-aware encoding to process heterogeneous Copernicus observations (Wang et al., 2025). TerraFM instead uses modality-specific representations and cross-attention to integrate complementary Sentinel-1 SAR and Sentinel-2 optical information at the feature level (Danish et al., 2026). Sensor-aware representation should not, however, be equated with universal cross-sensor harmonization. When observations from multiple optical missions are intended to represent comparable physical quantities, radiometric, geometric, spatial, or temporal harmonization may be appropriate. In contrast, multimodal systems may deliberately preserve sensor-specific characteristics and integrate them later through feature-level fusion or cross-attention. The appropriate strategy therefore depends on whether the objective is to make observations directly comparable or to preserve and combine complementary sensor-specific information.
4.3. Spatiotemporal Context and Metadata Encoding
EO observations are georeferenced and time-dependent, making geographic location, acquisition time, seasonality, and temporal recurrence potentially informative components of model input. Several EO-FMs therefore encode such contextual information explicitly rather than treating it only as ancillary metadata. The temporal and location variants of Prithvi-EO-2.0 incorporate acquisition-time and geolocation information through temporal and location embeddings within a multi-temporal backbone pretrained on 4.2 million global HLS time-series samples (Szwarcman et al., 2026). SatCLIP follows a different strategy by aligning satellite imagery with geographic coordinates to learn reusable geographic representations (Klemmer et al., 2025).
Such contextual information can improve representation learning when downstream processes depend on seasonal variation, repeated observations, or geographic structure. However, stronger contextual encoding does not automatically imply stronger generalization. A model may learn stable geographic or seasonal shortcuts that perform well within familiar regions or periods while weakening robustness under underrepresented locations or atypical environmental conditions. Evaluation should therefore examine whether spatiotemporal encoding contributes to performance under geographic and temporal shifts rather than infer generalization from within-domain accuracy alone.
4.4. Backbone Architectures and Multimodal Fusion
Transformer-based architectures are widely used in EO-FMs because they support flexible tokenization, long-range dependency modeling, and attention-based interaction among spatial, spectral, temporal, and multimodal inputs. Masked autoencoding has also become an important representation-learning strategy, in which information is deliberately removed from the input and reconstructed from the remaining context (He et al., 2022). For EO data, masking and representation learning can extend beyond spatial image patches to temporal and multispectral structures (Cong et al., 2022), while multimodal masked learning can operate across different input modalities (Bachmann et al., 2022). Prithvi-EO-2.0 uses an MAE-style geospatial backbone for multi-temporal optical EO data (Szwarcman et al., 2026). TerraFM combines representation learning with adaptive SAR–optical integration through modality-specific processing and cross-attention (Danish et al., 2026), whereas
DOFA and Copernicus-FM introduce mechanisms designed to accommodate heterogeneous sensor configurations (Xiong et al., 2024; Wang et al., 2025). Multimodal fusion should also be distinguished from harmonization. Harmonization reduces inconsistencies among observations so that they can be compared more consistently, whereas fusion combines complementary information within a joint representation or prediction framework. Architectural mechanisms such as cross-attention, modality-specific token streams, and metadata-aware integration therefore address a different problem from external sensor harmonization. Architectural flexibility alone should not be interpreted as evidence of broad generalization. The usefulness of a backbone ultimately depends on the relationship between its pretraining data and the target application, the sensing characteristics represented during pretraining, the downstream adaptation strategy, and its performance under the geographic, temporal, or sensor conditions relevant to deployment.
4.5. Downstream Adaptation and Task Interfaces
The practical usefulness of an EO-FM depends not only on its pretrained backbone but also on how its representations are adapted to a specific downstream application. Common adaptation strategies include linear probing, full fine-tuning, and parameter-efficient fine-tuning.
In linear probing, the pretrained backbone remains frozen and only a lightweight prediction layer is trained for the downstream task. This approach provides a relatively direct indication of the information already contained in the pretrained representation while requiring limited additional computation. In full fine-tuning, most or all pretrained parameters are updated using downstream data. This provides greater task-specific flexibility but requires additional computation and can increase the risk of overfitting when labeled data are limited. Parameter-efficient fine-tuning modifies only a limited subset of model parameters through approaches such as adapters, low-rank updates such as LoRA, or learnable prompt tokens where supported. These methods can reduce computational and storage requirements while retaining most of the pretrained backbone. Learnable prompt tuning should be distinguished from natural-language prompting. In vision-oriented EO-FMs, prompts may refer to trainable vectors or tokens used to adapt the representation, whereas text-based instruction or prompt tuning requires an explicit language or vision-language interface and is therefore not applicable to all EO-FMs.
The downstream task interface also varies with the application. Classification may require a lightweight prediction head, semantic segmentation generally requires dense prediction components or a task-specific decoder, retrieval may operate directly on learned embeddings, and forecasting models require output structures capable of representing future observations or states. Model comparisons should therefore report both the pretrained backbone and the adaptation procedure because downstream performance can be influenced substantially by how the representation is specialized.
4.6. Reproducibility and Evaluation Requirements
Reproducible evaluation of EO-FMs requires documentation of the complete experimental configuration rather than the pretrained model weights alone. Relevant information includes preprocessing procedures and product versions, sensor and band configurations, spatial resolution, tiling strategy, cloud and quality masking, train-validation-test splitting procedures, amount of downstream training data, adaptation strategy, layers or parameters updated during fine-tuning, optimization settings, and evaluation metrics. Where predictive uncertainty is evaluated, studies should additionally specify how uncertainty is estimated and whether the resulting predictions are calibrated.
Spatial and temporal leakage are general machine-learning concerns rather than problems unique to EO-FMs. They are nevertheless particularly important in EO because observations frequently exhibit strong spatial and temporal autocorrelation. Overlapping tiles, geographically adjacent samples, repeated observations of the same locations, or closely related temporal observations distributed across training and test sets can yield optimistic performance estimates when the intended application involves new regions or time periods (Roberts et al., 2017).
Evaluation should also distinguish performance attributable to the pretrained representation from improvements introduced through downstream adaptation. Comparisons become difficult when models are assessed using different fine-tuning regimes, different quantities of labeled data, or substantially different hyperparameter-optimization procedures. Explicit reporting of these choices helps determine whether observed gains arise primarily from the pretrained representation, the downstream adaptation procedure, the data-processing workflow, or interactions among these components.
Accordingly, EO-FM performance should be interpreted relative to the downstream task, sensing modality, preprocessing and splitting procedures, adaptation strategy, amount of labeled data, and intended deployment setting. This provides a more informative basis for assessing representation reuse and generalization than treating downstream performance as an intrinsic and universally applicable property of the pretrained model.
Table 3.
Core functional components of the EO-FM workflow.
| Tier | Description | Function |
| Input and encoding layer | Ingests heterogeneous EO inputs, including multispectral optical, SAR, thermal, atmospheric, and metadata streams after preprocessing and tiling. | Converts sensor-specific observations into model compatible tokens or embedding using patch embeddings, band-aware encoders, wavelength-conditioned modules, and metadata tokens. |
| Representation learning layer | Pretrained backbone, typically transformer-based and often coupled with masked, contrastive, or forecasting objectives. | Encodes spatial, spectral, temporal, and cross-modal context into reusable latent representations. |
| Fusion and context layer | Optional multimodal or metadata-aware components such as cross-attention, dynamic hypernetworks, and geospatial/temporal embeddings. | Integrates complementary modalities and acquisition context while preserving or combining sensor-specific information. |
| Adaptation and task interface | Lightweight heads or parameter-efficient modules for classification, segmentation, regression, retrieval, or forecasting. | Converts pretrained representations into downstream outputs and can reduce task-specific labeling requirements in some settings. |
| Evaluation and reproducibility layer | Benchmark protocols, calibration checks, and transparent preprocessing documentation. | Supports transparent comparison of downstream performance, adaptation, and generalization claims and assessment of scientific or operational suitability. |
5. Model Landscape and Taxonomy
The current EO-FM landscape is diverse in both model design and intended use. Some models are developed primarily for multispectral optical representation learning, others for multimodal SAR–optical integration or cross-sensor flexibility, and others for temporal forecasting, geographic representation, or embedding-based downstream analysis. These model families differ in their pretraining data, sensing modalities, learning strategies, and downstream capabilities. Figure 3 organizes representative EO-FMs according to their dominant functional characteristics and illustrates how different model families support different scientific and operational needs.
5.1. Taxonomy Dimensions
For review purposes, the taxonomy in Figure 3 compares representative EO-FMs along four descriptive dimensions: (i) primary data emphasis, (ii) typical learning strategy, (iii) representative models, and (iv) characteristic downstream strengths. Primary data emphasis distinguishes models centered on multispectral optical imagery from those incorporating SAR, multiple sensors, temporal sequences, geographic coordinates, or embedding products. Typical learning strategy refers to the dominant representation-learning approach, including masked autoencoding, multimodal fusion, sensor-aware encoding, autoregressive forecasting, and contrastive image–location alignment. Representative models illustrate each family without implying rigid boundaries, because some models combine characteristics from more than one category. Characteristic strengths summarize the downstream capabilities most closely associated with each model family in the literature reviewed here. These dimensions provide a consistent basis for comparing models with substantially different sensing assumptions and intended uses.
5.2. Optical and Multispectral Vision Foundation Models
A first major group consists of optical and multispectral EO-FMs. This family is relatively mature because large optical archives, including harmonized Landsat and Sentinel observations, provide extensive spatial and temporal coverage for pretraining. Prithvi-EO-2.0 is a representative example: it was pretrained on 4.2 million global HLS time-series samples at 30 m resolution, and its temporal-location variants incorporate temporal and location information into the representation-learning framework (Szwarcman et al., 2026). Its reported evaluations span land-cover mapping, crop mapping, hazard response, and ecosystem-related applications, illustrating how multi-temporal optical representations can be reused across different downstream tasks. Overall, optical and multispectral EO-FMs are particularly relevant when downstream applications rely primarily on optical observations and when pretrained representations can reduce the need for extensive task-specific training. Their effectiveness nevertheless depends on the representativeness of the pretraining archive and the correspondence between the pretraining data and downstream sensing conditions.
5.3. Multimodal and Cross-Sensor Foundation Models
A second family emphasizes multimodal and cross-sensor EO learning, particularly the integration of observations from different sensing systems. This direction is important because different modalities provide complementary information. For example, optical imagery provides rich spectral information, whereas SAR provides structural information and can remain useful under cloud cover. TerraFM uses globally distributed Sentinel-1 and Sentinel-2 observations together with modality-specific representations and adaptive cross-attention to support classification and segmentation (Danish et al., 2026). DOFA uses wavelength-conditioned dynamic parameterization to accommodate multiple spectral configurations and sensor inputs (Xiong et al., 2024). Copernicus-FM extends heterogeneous-input learning across multiple Copernicus observation streams through metadata-aware encoding and a large multimission training corpus that includes both surface and atmospheric observations (Wang et al., 2025). Clay represents another multisensor approach, using sensor-aware inputs and metadata to generate reusable EO representations across different supported sensing configurations. Its evidence base is currently weighted more heavily toward openly released model resources, documentation, and deployment-oriented use than toward the standardized comparative evaluations available for several benchmark-focused EO-FMs (Clay Foundation, 2024).
Together, these models illustrate different strategies for representing heterogeneous EO data. TerraFM emphasizes SAR–optical fusion, DOFA emphasizes spectral and sensor adaptability, Copernicus-FM emphasizes multimission and metadata-aware representation learning, and Clay emphasizes reusable multisensor representations and accessible downstream use. The reviewed evidence demonstrates the feasibility of these approaches, but it does not establish that multimodal or multisensor models are uniformly superior to modality-specific EO-FMs. Their relative usefulness depends on the target application, available sensors, and the way complementary information is represented and integrated.
5.4. Time-Series and Forecasting-Oriented EO Foundation Models
A third family treats EO primarily as a sequence-modeling problem rather than as static image understanding. These models emphasize temporal dynamics, recurrence, and future-state prediction. EarthPT is a representative example: it uses a 700-million-parameter autoregressive transformer to forecast surface-reflectance time series and is explicitly designed around EO temporal structure (Smith et al., 2023). Aurora extends this logic into the broader Earth-system domain by using more than one million hours of heterogeneous geophysical data for forecasting applications including weather, air quality, ocean waves, and tropical-cyclone tracks (Bodnar et al., 2025). Aurora is therefore better viewed as an adjacent Earth-system foundation model rather than a conventional satellite-image EO-FM. These models demonstrate that the foundation-model paradigm extends beyond classification and segmentation backbones. Their characteristic strengths lie primarily in temporal representation and forecasting, so their downstream capabilities and evaluation criteria differ substantially from those of static EO image encoders.
5.5. Location-Centered and Embedding-Oriented Geospatial Models
A fourth family focuses on geographic representation, location-conditioned prediction, and embedding-based geospatial analysis rather than dense pixel-level mapping. SatCLIP is a representative example, using contrastive pretraining between satellite imagery and geographic coordinates to learn reusable location embeddings (Klemmer et al., 2025). This approach is particularly relevant to applications in which geographic location itself provides useful contextual information, including ecological, environmental, and socio-spatial prediction. The importance of this family is both conceptual and technical: EO foundation models do not necessarily need to operate primarily as segmentation backbones or multimodal image encoders. Some instead provide reusable geographic representations that can be incorporated into downstream analytical workflows. Their usefulness should therefore be assessed relative to location-conditioned prediction and related geographic tasks rather than directly compared with dense-mapping models using a single performance metric.
5.6. Global Embedding Products and 'Foundation Model as a Dataset'
An emerging deployment pattern treats the foundation model not only as a pretrained network but also as a provider of analysis-ready embeddings. AlphaEarth Foundations is an important example of this approach. Its embedding-field framework integrates spatial, temporal, and measurement context across multiple sources, and its representations are released through annual satellite embedding products that can support downstream mapping and monitoring workflows (Brown et al., 2025). This approach can lower computational barriers for users who do not need to execute the underlying backbone locally. Clay also partly reflects this deployment-oriented direction through its emphasis on accessible pretrained representations and downstream usability. At the same time, embedding products increase the importance of clear versioning, provenance, preprocessing documentation, and reproducibility, because changes to the underlying model or data-processing workflow can affect subsequent analyses.
5.7. Comparative Synthesis of the Current Landscape
The taxonomy above indicates that the EO-FM landscape is developing along several complementary pathways: optical and multispectral representation models, multimodal and cross-sensor models, time-series and forecasting-oriented models, location-centered and embedding-oriented geospatial models. These families differ in the observations they use, the representations they learn, and the downstream capabilities they expose.
These differences also mean that results from one model family should not automatically be interpreted as evidence of broader superiority across other downstream applications. Classification and segmentation models, geographic embedding models, and forecasting-oriented systems address fundamentally different prediction settings. Table 4 therefore summarizes representative models according to their functional family, modality emphasis, downstream capability, and available evidence.
The broader implication is that the taxonomy represents complementary model-design pathways rather than a hierarchy of EO-FMs. Model selection should be guided by the correspondence among the model’s input data, representation-learning strategy, sensing and temporal characteristics, and the requirements of the intended downstream application.
6. Applications, Empirical Evidence, and Research Gaps
Earth observation foundation models (EO-FMs) are increasingly evaluated not only by benchmark performance but also by their usefulness across downstream applications under limited labels, heterogeneous sensing conditions, and distribution shift. However, the empirical evidence remains heterogeneous because studies differ in datasets, sensing modalities, spatial and temporal resolutions, adaptation procedures, data splits, and performance metrics. Reported performance should therefore be interpreted relative to the target application and experimental setting rather than as evidence of universal model superiority. Figure 4 synthesizes three related aspects of the current EO-FM literature: major application domains, representative empirical evidence, and the research gaps and limitations that affect reliable downstream use. The figure also highlights four priorities for strengthening future evaluation: standardized and reproducible assessment, geographically, temporally, and sensor-disjoint testing, uncertainty and failure analysis, and transparent reporting of downstream adaptation.
6.1. Comparative Performance Evidence Across Representative EO-FMs
Comparative evidence should be interpreted relative to task type, sensing modality, adaptation procedure, and evaluation design. Models developed for multi-temporal optical mapping, multimission observations, geographic representation, or Earth-system forecasting address substantially different problems and are therefore not directly interchangeable. Table 5 presents selected reported results as a task-specific comparative evidence summary rather than a universal leaderboard. Strong performance on one benchmark family does not necessarily imply better performance under a different task, sensor configuration, or deployment setting (Lacoste et al., 2023; Simumba et al., 2026).
These results illustrate capability–application correspondence rather than an overall ranking. Each model is evaluated within a different combination of task, sensing modality, and experimental setting.
6.2. Land-Cover and Land-Use Mapping
Land-cover classification and semantic segmentation are among the most frequently evaluated EO-FM applications because labeled datasets are relatively available and these tasks align well with pretrained multispectral representations. Prithvi-EO-2.0 reports multi-temporal results across crop and land-use applications (Szwarcman et al., 2026), TerraFM reports strong classification performance on EuroSAT and BigEarthNet variants (Danish et al., 2026), and DOFA supports classification and segmentation across different sensor configurations (Xiong et al., 2024). These findings demonstrate the usefulness of pretrained EO representations for downstream adaptation, but strong within-benchmark performance should not automatically be interpreted as geographic or sensor generalization. Evaluation on geographically disjoint regions, different acquisition periods, or previously unseen sensor conditions provides stronger evidence of robustness beyond the development setting (Lacoste et al., 2023; Simumba et al., 2026).
6.3. Change Detection and Disturbance Monitoring
Change detection and disturbance monitoring can benefit from EO-FMs because many applications require comparison across repeated observations rather than interpretation of a single image. Multi-temporal representations may help distinguish persistent environmental change from normal seasonal variation, which is relevant to vegetation disturbance, wildfire impacts, flood progression, and land-use change.
Apparent change can also arise from phenology, illumination or viewing differences, atmospheric conditions, sensor artifacts, or inconsistent preprocessing and labels. These factors complicate interpretation because differences between observations do not necessarily represent true changes in land-surface conditions. Models with explicit temporal structure, such as Prithvi-EO-2.0 and EarthPT, are therefore relevant to disturbance-oriented analysis (Szwarcman et al., 2026; Smith et al., 2023). Robust evidence nevertheless requires temporally—and where appropriate geographically—disjoint evaluation rather than reliance on average benchmark performance alone (Lacoste et al., 2023; Simumba et al., 2026).
6.4. Hazards and Disaster Response
Hazard applications are particularly relevant to EO-FMs because event-specific labels are often limited when rapid mapping is required. Published evaluations of Prithvi-EO-2.0 provide direct evidence for flood and wildfire mapping, while Copernicus-FM includes flood-related evaluation within Copernicus-Bench; TerraFM provides complementary evidence from multisensor classification and segmentation benchmarks relevant to environmental mapping (Szwarcman et al., 2026; Wang et al., 2025; Danish et al., 2026). Operational performance nevertheless remains sensitive to event type, sensor availability, acquisition timing, spatial resolution, class imbalance, and geographic or temporal shift. Benchmark and application results therefore indicate potential but do not establish general readiness across disaster settings. Recent work has also examined Sentinel-1 and Sentinel-2 foundation-model representations for building-damage retrieval in disaster-response applications (Dietrich et al., 2026). Section 7 examines representative flood and wildfire workflows in greater detail, focusing on sensing configuration, preprocessing, downstream adaptation, and evaluation across published hazard applications.
6.5. Urban Climate Analytics
Urban climate remains an emerging EO-FM application area, including urban heat-island analysis, exposure mapping, and built-environment characterization. Geospatial foundation-model research suggests potential for reusable representations across urban settings (Janowicz et al., 2025), while early direct work has explored foundation-model technology for urban heat-island applications (Bhamjee et al., 2024). Generalization to cities not represented during downstream training remains particularly challenging because urban form, climate, land-cover composition, and socioeconomic context vary substantially among cities. Transparent uncertainty assessment and reporting are also important where model outputs may inform threshold-sensitive heat or exposure assessments (Kermarrec et al., 2025). The current evidence base, however, remains too limited to treat urban climate as a mature EO-FM application domain.
6.6. Agriculture, Ecosystems, and Biodiversity
Agricultural and ecosystem applications are particularly relevant to EO-FMs because many underlying processes are temporally dynamic. Crop classes may appear spectrally similar on a single date but differ substantially in their phenological trajectories across the growing season. Repeated observations can capture changes associated with emergence, canopy development, peak growth, senescence, and harvest. Prithvi-EO-2.0 processes multi-temporal HLS observations, and its temporal-location variants explicitly incorporate temporal information into the representation-learning framework, making its design well aligned with crop classification and segmentation tasks that depend on seasonal development (Szwarcman et al., 2026). Different model families represent temporal and spatial context in different ways. EarthPT learns temporal structure through autoregressive prediction of surface-reflectance sequences and is therefore oriented toward temporal dynamics and forecasting (Smith et al., 2023), whereas SatCLIP provides geographic embeddings that can support ecological and socio-environmental prediction where location is informative (Klemmer et al., 2025).
These differences illustrate why model–application alignment matters: multi-temporal optical models may be appropriate for phenology-sensitive crop mapping, forecasting-oriented models for vegetation trajectories, and location-centered representations for geographically structured ecological prediction. Nevertheless, temporal sampling, cloud contamination, field-label quality, spatial resolution, and mismatch between field measurements and EO pixels can limit downstream performance. Geographic split design and scale correspondence therefore remain important considerations in agricultural and ecological applications (Simumba et al., 2026).
6.7. Cross-Cutting Research Gaps
Several recurring limitations restrict the interpretation and comparison of EO-FM performance. First, benchmark comparability remains incomplete. Models are frequently evaluated on different application-specific datasets, while even studies using similar datasets may differ in preprocessing, data splits, downstream adaptation, and hyperparameter selection. Relevant preprocessing differences include cloud and shadow masking, radiometric normalization, co-registration, spatial resampling, band selection, temporal compositing, and tiling or overlap rules. Such choices can alter the information presented to a model and affect reported performance. The broader issue is therefore not simply differences in sensing modality, but the lack of common evaluation settings and consistent adaptation and reporting protocols.
Second, generalization under distribution shift remains uneven. Relevant shifts include application to a new geographic region, a different season or year, an unseen or differently configured sensor, or rare environmental conditions such as extreme floods, wildfires, or droughts. Performance under these conditions provides stronger evidence of generalization than evaluation on data closely matched to the development setting.
Third, many evaluations continue to emphasize accuracy-oriented metrics while providing less information about predictive uncertainty, calibration, and failure analysis. Calibration concerns whether predicted confidence corresponds to observed correctness, while failure analysis examines the classes, regions, sensors, events, or environmental conditions under which errors occur. These issues are particularly important for applications such as flood delineation, wildfire mapping, environmental exposure assessment, and emergency-response support.
Finally, evaluations should distinguish improvements attributable to the pretrained representation from those introduced by downstream adaptation. The amount of labeled data, fine-tuning strategy, updated layers, augmentation procedures, and hyperparameter optimization can all influence reported performance. Explicit reporting of these choices is therefore necessary when attributing downstream gains to an EO-FM.
Table 6 indicates that EO-FMs offer reusable representations, reduced task-specific labeling requirements in some settings, and enhanced capacity to model temporal or multimodal EO information. These benefits nevertheless depend on appropriate preprocessing, clearly documented downstream adaptation, and spatially, temporally, or sensor-disjoint evaluation when generalization beyond the development setting is claimed.
7. Representative Hazard Applications and Model Adaptation Workflows
Hazard applications provide a useful testing ground for Earth observation foundation models (EO-FMs) because event-specific labels are often limited and environmental conditions can vary substantially across locations, acquisition periods, and sensors. Floods and wildfires rarely occur under identical conditions, making downstream adaptation and generalization particularly important. EO-FMs are therefore relevant not because they eliminate the need for task-specific adaptation, but because pretrained representations may reduce labeling requirements and provide a stronger starting point for rapid mapping and post-event assessment. Figure 5 illustrates two representative workflows -(A) flood extent mapping and (B) wildfire mapping—and shows how sensor inputs, preprocessing, pretrained representations, downstream adaptation, and evaluation interact in practice.
7.1. Flood Extent Mapping
Flood extent mapping is particularly relevant to EO-FMs because rapid delineation is often required under cloud-covered conditions. Synthetic aperture radar (SAR) is valuable in this setting because microwave observations are largely unaffected by cloud cover and can distinguish open water from many surrounding land surfaces. Multimodal input is not a prerequisite for EO-FM-based flood mapping: a pretrained model designed for SAR observations alone can be adapted for flood segmentation, whereas multimodal models can additionally incorporate optical imagery when suitable cloud-free observations are available.
Copernicus-FM provides a direct SAR-based example because its multimission Sentinel framework includes Sentinel-1 observations and reports 77.7 mIoU on the Flood-S1 task in Copernicus-Bench (Wang et al., 2025). Prithvi-EO-2.0 provides a complementary optical example. Although evaluated using the Sen1Floods11 flood benchmark, the Prithvi experiment uses Sentinel-2 imagery and the same six optical bands represented during pretraining. Its 600M-TL model reports 90.3 mIoU, 97.7 mF1, and 83.1 IoU for the water class (Szwarcman et al., 2026). These results demonstrate the applicability of pretrained EO representations to flood segmentation, but they arise from different sensing modalities, model architectures, and evaluation settings and should therefore not be interpreted as directly comparable rankings.
A flood-mapping workflow begins with sensor-appropriate preprocessing. For Sentinel-1, relevant steps may include radiometric calibration, terrain correction, speckle treatment where appropriate, spatial alignment, and tiling. Sentinel-2 imagery can be incorporated when cloud-free observations are available, but optical supplementation is optional rather than required. Metadata such as acquisition time, orbit geometry or incidence angle, and sensor identity may also be incorporated when supported by the model architecture. Such information can help account for systematic acquisition differences that might otherwise be confused with changes in surface conditions, although the contribution of individual metadata variables should be demonstrated empirically.
The pretrained representation can then be adapted to flood segmentation. When labeled data are very limited, the backbone may remain frozen while a lightweight segmentation head is trained. Parameter-efficient fine-tuning can update a limited subset of model parameters, whereas full fine-tuning updates the backbone together with the downstream head. The appropriate adaptation strategy therefore depends on label availability, computational resources, and the degree of mismatch between the target flood conditions and the data represented during pretraining. Claims of geographic or event-level generalization should be supported by evaluation on floods, regions, or acquisition periods that were not represented during downstream training. Such evaluation is particularly important because flood boundaries can remain sensitive to surface roughness, land cover, SAR acquisition characteristics, and local hydrologic conditions.
7.2. Wildfire Mapping
Wildfire applications include burn-scar delineation, burn-severity or burn-intensity mapping, active-fire support, and post-fire recovery monitoring. These tasks are challenging because wildfire signatures vary with vegetation type, burn intensity, topography, season, acquisition timing, and post-fire surface conditions. Prithvi-EO-2.0 provides direct empirical evidence for this application family. Its 600M model reports 90.5 mIoU, 98.1 mF1, and 83.2 IoU for the burn-scar class in wildfire-scar mapping and has also been evaluated for burn-intensity mapping (Szwarcman et al., 2026). These results demonstrate the potential of multi-temporal pretrained representations for wildfire applications, although their interpretation remains specific to the datasets and downstream adaptation procedures used in the reported experiments.
A wildfire-mapping workflow commonly uses pre-event and post-event optical observations. Sentinel-2 imagery alone can be used, and harmonization is not inherently required. Harmonized products become relevant when observations from different optical missions, such as Landsat and Sentinel-2, are combined to increase temporal coverage or construct consistent pre- and post-event image sequences. SAR observations may also provide complementary information where optical availability is limited by clouds or other acquisition constraints.
The pretrained backbone can then be adapted for burn-scar segmentation, burn-intensity mapping, or severity-related prediction. As with flood mapping, adaptation can range from a frozen backbone with a lightweight task head to parameter-efficient or full fine-tuning. Multi-temporal EO-FMs may be particularly useful because pre-event observations establish seasonal baseline conditions against which post-fire disturbance can be evaluated, potentially helping distinguish fire-related change from normal vegetation phenology. The suitability of a particular adaptation strategy should therefore be evaluated empirically rather than inferred solely from the scale of the pretrained model. Event-disjoint testing across different fires, vegetation conditions, regions, and acquisition periods provides stronger evidence that the adapted representation remains useful beyond the events used during downstream training.
7.3. Model Adaptation, Reproducibility, and Evaluation for Hazard Applications
Hazard applications also illustrate why evaluation should approximate the intended deployment scenario. This does not require testing a model during an actual emergency. Rather, the evaluation should reproduce the main information constraint that would occur when the model is applied to a new event: the test flood or wildfire should not already be represented in the downstream training data. For example, if the intended use is mapping a previously unseen flood, randomly assigning neighboring tiles from the same flood event to both training and test sets can produce overly optimistic performance estimates. A more informative design would hold out an entire flood event, geographic region, or acquisition period. Similarly, wildfire evaluation can hold out complete fire events rather than randomly mixing patches from the same burn area across training and testing. Spatially disjoint splits, event-level holdouts, and temporally separated evaluation therefore provide stronger evidence of generalization to new hazard conditions (Roberts et al., 2017).
Evaluation should also clearly report the downstream adaptation procedure, because results obtained with a frozen encoder and lightweight head are not directly comparable with those obtained through parameter-efficient or full fine-tuning. The amount of labeled data and the extent of parameter updating should therefore be reported when attributing downstream performance gains to the pretrained representation. Hazard evaluation should also consider predictive uncertainty, calibration, and failure analysis where these quantities are available. Predictive uncertainty reflects the degree of uncertainty associated with a model output, whereas calibration assesses whether reported confidence is consistent with observed correctness. Failure analysis can further identify conditions under which errors concentrate, such as vegetated floodplains, complex urban surfaces, low-severity burns, rare classes, or unfamiliar geographic settings. These diagnostics are particularly important when model outputs are intended to support risk-sensitive or emergency-response decisions, where reliability and applicability should be evaluated alongside predictive performance (Tiggeloven et al., 2025).
8. Discussion: Toward Trustworthy EO Foundation Models for Science and Policy
The literature reviewed in this paper suggests that Earth observation foundation models (EO-FMs) represent an important development in geospatial artificial intelligence, but their significance extends beyond increasing model size. Their broader contribution lies in learning reusable representations from large EO archives that can support downstream adaptation across diverse applications. Across representative model families, the evidence indicates potential advantages in label efficiency, representation reuse, multimodal and temporal learning, and geographic scalability. However, these benefits depend on the compatibility between model inputs and target data, the downstream adaptation strategy, and the evaluation design (Lacoste et al., 2023; Szwarcman et al., 2026; Wang et al., 2025). EO-FMs should therefore not be interpreted as universally superior replacements for conventional remote-sensing workflows. Their scientific and operational value depends on how effectively pretraining data, sensor-aware representation learning, temporal structure, downstream adaptation, and evaluation are aligned with the intended application. Figure 6 summarizes six considerations that influence whether EO-FMs can progress from high-performing research models to reliable tools for environmental science and policy: transparency, evaluation fit, predictive uncertainty, scientific validity, governance, and sustainability.
8.1. From Benchmark Performance to Application-Relevant Utility
One of the clearest conclusions from this review is that benchmark performance alone is insufficient to establish practical utility. Before benchmark scores are compared, a more fundamental question is whether a model is compatible with the target sensing data and application. A model pretrained for optical inputs cannot directly process raw SAR observations unless its architecture and input representation are designed or adapted to support SAR. A SAR-based application should first consider models that support SAR inputs or multimodal architectures capable of incorporating SAR observations.
Model comparison should therefore begin with modality and task compatibility, followed by evaluation of performance within the relevant application setting. Prithvi-EO-2.0, TerraFM, Copernicus-FM, and DOFA illustrate this distinction because they differ substantially in supported sensing configurations, representation strategies, and downstream capabilities (Szwarcman et al., 2026; Danish et al., 2026; Wang et al., 2025; Xiong et al., 2024). Benchmark frameworks remain important once appropriate model candidates have been identified. GEO-Bench improves comparability through standardized classification and segmentation tasks, while GEO-Bench-2 extends this logic toward capability-oriented evaluation and demonstrates why a single aggregate score cannot adequately summarize model suitability across heterogeneous EO applications (Lacoste et al., 2023; Simumba et al., 2026). Model selection should therefore consider not only benchmark performance but also sensing modality, task requirements, data availability, adaptation requirements, and the consequences of prediction errors.
8.2. Why No Single EO-FM Dominates Across All Settings
The diversity of EO-FMs reflects genuine differences in how Earth observation information is represented rather than simply an immature stage of model development. Optical and multispectral models, multimodal systems, sensor-flexible architectures, geographic embedding models, and Earth-system forecasting models address different aspects of the EO problem. Prithvi-EO-2.0 emphasizes multi-temporal optical representation learning, Copernicus-FM combines multimission observations with metadata-aware learning, TerraFM emphasizes multimodal SAR–optical representation learning, DOFA provides flexibility across different spectral and sensor configurations, SatCLIP learns geographic representations from location information, and Aurora focuses on forecasting Earth-system processes (Szwarcman et al., 2026; Wang et al., 2025; Danish et al., 2026; Xiong et al., 2024; Klemmer et al., 2025; Bodnar et al., 2025). Accordingly, asking which EO-FM is “best” in absolute terms is generally less informative than asking which model is best aligned with the target data, task, and intended output. The current model landscape is therefore better understood as a set of complementary pretrained representation systems rather than a hierarchy expected to converge toward a single universally dominant architecture.
8.3. Trustworthiness: Transparency, Calibration, and Governance
If EO-FMs are to support environmental decision-making, trustworthiness must become an explicit component of model evaluation. EO-FM outputs may contribute to flood delineation, wildfire assessment, land-use monitoring, environmental exposure estimation, or climate-related decision support. In these settings, average predictive performance alone does not establish reliability. Transparency begins with clear documentation of the data-processing steps applied before observations are provided to the model. Relevant preprocessing choices can include cloud and shadow masking, radiometric normalization, geometric co-registration, spatial resampling, band selection, temporal compositing, and tiling. These operations can alter the information available to the model and therefore influence downstream performance. Preprocessing procedures, software or dataset versions, and model versions should consequently be documented sufficiently to support reproducibility.
Predictive confidence also requires careful interpretation. Calibration refers to the correspondence between predicted confidence and observed correctness. For example, among predictions assigned approximately 90% confidence, a well-calibrated classifier would be expected to be correct approximately 90% of the time. Calibration is therefore different from accuracy: a model can achieve high overall accuracy while still producing poorly calibrated confidence estimates. Reporting calibration, predictive uncertainty, and failure patterns is particularly important where model outputs contribute to threshold-sensitive or risk-sensitive decisions. Trustworthiness also depends on transparent model release and maintenance practices. When pretrained models, embeddings, or derived datasets are updated, changes in model weights, preprocessing pipelines, or source data can affect longitudinal comparability. Clear versioning, intended-use statements, documentation of known limitations, and transparent update practices therefore become important elements of responsible EO-FM deployment (Janowicz et al., 2025; “Towards Responsible Geospatial Foundation Models,” 2025).
8.4. Scientific Validity and Interpretability
Strong downstream performance does not necessarily demonstrate that an EO-FM has learned scientifically meaningful environmental relationships. EO applications often aim not only to predict outcomes but also to support understanding of land-surface and Earth-system processes. A model may achieve high predictive performance by exploiting geographic, seasonal, or acquisition-related regularities without necessarily learning the physical processes of interest. This issue is particularly relevant for models that incorporate geographic coordinates, temporal metadata, or other contextual information. Such information can improve prediction, but it can also act as a shortcut when strong correlations between location and target variables are present (Klemmer et al., 2025; Janowicz et al., 2025). Interpretability should therefore be considered part of scientific validation rather than only a model-transparency exercise.
Useful approaches include attribution analysis, sensitivity testing, ablation of geographic or temporal context, and comparison with established environmental relationships. These analyses can help determine whether performance improvements reflect meaningful representation of environmental processes or reliance on contextual correlations. Strong downstream performance should thus be viewed as necessary, but not sufficient, evidence of scientific reliability.
8.5. Model Adaptation Across Environments, Sensors, and Applications
A central practical challenge for EO-FMs is adapting pretrained representations to target settings that differ from those emphasized during pretraining. These differences may involve geographic environment, acquisition period, sensor configuration, spectral characteristics, spatial resolution, or downstream task. Distribution shift is one manifestation of this broader problem, but the more practically relevant question is how EO-FMs should be adapted when the target setting changes. Different levels of adaptation may be appropriate. A frozen backbone with a lightweight task head may be sufficient when pretrained features already align well with the target problem. Parameter-efficient approaches, such as adapters or low-rank updates, provide an intermediate option when some specialization is required but computational resources or labels are limited. Full fine-tuning may be appropriate when stronger adjustment to the target application is necessary.
The appropriate strategy therefore depends on the correspondence between the pretrained representation and the downstream data. Generalization claims should subsequently be tested using geographic, temporal, sensor, or event-level holdouts appropriate to the intended application. In operational settings, a realistic workflow may therefore combine pretrained representations, targeted downstream adaptation, local validation, and explicit uncertainty assessment rather than relying on completely frozen or zero-shot models.
8.6. Sustainability and Access
The increasing scale of EO-FMs introduces additional questions concerning computational sustainability and accessibility. Larger pretraining datasets and models can expand representational capacity but also require substantial computational resources, storage, engineering expertise, and energy. These requirements may concentrate EO-FM development within well-resourced institutions and make other researchers or operational agencies dependent on externally maintained models or embedding products. Precomputed global embeddings can lower barriers to downstream use, but they also increase the importance of version control, documentation, and reproducibility because users may not control the underlying training or preprocessing pipeline (Google DeepMind, 2025; Google Earth Engine, 2025; Janowicz et al., 2025). A sustainable EO-FM ecosystem therefore requires not only increasingly capable models but also efficient model variants, reproducible preprocessing workflows, accessible benchmarks, and transparent release practices. Computational efficiency and accessibility should be considered alongside predictive performance when assessing the broader scientific utility of EO-FMs.
8.7. A Forward-Looking Research Agenda
Several priorities emerge from this review. First, future EO-FM development should improve sensor-aware representation learning while making the treatment of modality differences explicit, particularly as thermal, hyperspectral, SAR, atmospheric, and other EO observations are incorporated.
Second, greater attention is needed to downstream adaptation across sensors, geographic environments, and applications, including systematic comparison of frozen representations, parameter-efficient adaptation, and full fine-tuning.
Third, evaluation should move beyond average benchmark performance toward application-relevant testing that includes geographically, temporally, or sensor-disjoint data when generalization is claimed. Predictive uncertainty, calibration, failure analysis, and performance under rare or difficult conditions should increasingly complement conventional accuracy metrics.
Fourth, interpretability research should examine whether learned representations correspond to known environmental processes rather than relying only on downstream predictive performance.
Finally, continued progress will require more open and reproducible model and benchmark infrastructures, transparent documentation of preprocessing and model updates, and computationally efficient alternatives that broaden access beyond a small number of large institutions. Recent perspective work similarly suggests that the long-term value of EO-FMs will depend not only on model scale, but also on scientific validity, transparency, sustainability, multisensor integration, and reliable evaluation (Janowicz et al., 2025; Zhu et al., 2026).
9. Conclusions
Earth observation foundation models are reshaping geospatial analysis by enabling Large EO Archives to support reusable representations across a growing range of downstream applications. This review shows, however, that EO-FMs should not be treated as a single model class or ranked through a universal performance criterion. Optical, multimodal, cross sensor, geographic-embedding and forecasting oriented models address different sensing and analytical requirements. Consequently, their usefulness depends on the alignment among pretraining data, model architecture, downstream adaptation, target sensing modalities and application context.
The review evidence demonstrates substantial potential for label-efficient adaptation, multi-modal and temporal representations learning, representation reuse. But evidence of generalization and operational reliability remains uneven. Progress should therefore be judged not only by model scale or benchmark accuracy but by reproduceable evaluation under relevant geographic, temporal and sensor conditions, transparent processing and adaptations, and explicit assessment of predictive uncertainty, calibration and failure modes. Ultimately the scientific value of EO-FMs will depend on their ability to provide reusable, generalizable and trustworthy geospatial representations that remain reliable beyond the conditions under which they were developed.
Supplementary Materials
The following supporting information can be downloaded at the website of this paper posted on Preprints.org.
Author Contributions
Conceptualization, M.M.M., S.W.M., P.F., Y.Z. and A.N.V.; methodology, M.M.M., S.W.M. and Y.Z.; writing—original draft preparation, M.M.M.; writing—review and editing, M.M.M., S.W.M., P.F., Y.Z., B.H., A.N.V. and Y.L.; supervision, S.W.M. and P.F. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Data Availability Statement
No new primary data were generated in this study. The literature corpus included in the review is reported in Supplementary Table S1.
Acknowledgments
During the preparation of this manuscript, the authors used OpenAI ChatGPT for language refinement, editorial restructuring, and assistance in developing two graphical layouts. The authors reviewed and edited all outputs and take full responsibility for the content of this publication.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Bachmann, R.; Mizrahi, D.; Atanov, A.; Zamir, A. MultiMAE: Multi-modal multi-task masked autoencoders. In Computer Vision – ECCV 2022; Springer, 2022; pp. 348–367. [Google Scholar] [CrossRef]
- Bhamjee, M.; Debary, H.; Gaffoor, Z.; Govindasamy, T.; Mahlasi, C.; Fiaz, M.; Vos, E.; Klein, L.; Makhanya, S.; Watson, C.; Kuehnert, J. Detection and characterization of urban heat islands with machine learning. IGARSS 2024 – 2024 IEEE International Geoscience and Remote Sensing Symposium; 2024; pp. 1693–1699. [Google Scholar] [CrossRef]
- Bodnar, C.; Bruinsma, W. P.; Lucic, A.; Stanley, M.; Allen, A.; Brandstetter, J.; et al. A foundation model for the Earth system. Nature 2025, 641(8065), 1180–1187. [Google Scholar] [CrossRef] [PubMed]
- Brown, C. F.; Kazmierski, M. R.; Pasquarella, V. J.; Rucklidge, W. J.; Samsikova, M.; Zhang, C.; Shelhamer, E.; Lahera, E.; Wiles, O.; Ilyushchenko, S.; Gorelick, N.; Zhang, L. L.; Alj, S.; Schechter, E.; Askay, S.; Guinan, O.; Moore, R.; Boukouvalas, A.; Kohli, P. AlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label data [Preprint]. arXiv. 2025. Available online: https://arxiv.org/abs/2507.22291.
- Clay Foundation. Pretrained model release v1.5. Clay Foundation Model. 19 November 2024. Available online: https://clay-foundation.github.io/model/release-notes/specification.html.
- Cong, Y.; Khanna, S.; Meng, C.; Liu, P.; Rozi, E.; He, Y.; Burke, M.; Lobell, D. B.; Ermon, S. SatMAE: Pre-training transformers for temporal and multi-spectral satellite imagery. Advances in Neural Information Processing Systems 35 2022, 197–211. [Google Scholar] [CrossRef]
- Danish, M. S.; Munir, M. A.; Shah, S. R. A.; Khan, M. H.; Anwer, R. M.; Laaksonen, J.; Khan, F. S.; Khan, S. TerraFM: A scalable foundation model for unified multisensor Earth observation. International Conference on Learning Representations (ICLR 2026); 2026. [Google Scholar]
- Dietrich, O.; Alfredsson, M.; Arens, E.; Metzger, N.; Peters, T.; Scheibenreif, L.; Wegner, J. D.; Schindler, K. The potential of Copernicus satellites for disaster response: Retrieving building damage from Sentinel-1 and Sentinel-2. ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences XI-2-2026 2026, 831–838. [Google Scholar] [CrossRef]
- Google DeepMind. AlphaEarth Foundations helps map our planet in unprecedented detail. 30 July 2025. Available online: https://deepmind.google/blog/alphaearth-foundations-helps-map-our-planet-in-unprecedented-detail/.
- Google Earth Engine. Satellite Embedding V1 (annual). Earth Engine Data Catalog. 2025. Available online: https://developers.google.com/earth-engine/datasets/catalog/GOOGLE_SATELLITE_EMBEDDING_V1_ANNUAL.
- He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2022; pp. 16000–16009. [Google Scholar] [CrossRef]
- Huo, C.; Chen, K.; Zhang, S.; Wang, Z.; Yan, H.; Shen, J.; Hong, Y.; Qi, G.; Fang, H.; Wang, Z. When remote sensing meets foundation model: A survey and beyond. Remote Sensing 2025, 17(2), 179. [Google Scholar] [CrossRef]
- Janowicz, K.; Mai, G.; Huang, W.; Zhu, R.; Lao, N.; Cai, L. GeoFM: How will geo-foundation models reshape spatial data science and GeoAI? International Journal of Geographical Information Science 2025, 39(9), 1849–1865. [Google Scholar] [CrossRef]
- Ju, J.; Zhou, Q.; Freitag, B.; Roy, D. P.; Zhang, H. K.; Sridhar, M.; Mandel, J.; Arab, S.; Schmidt, G. L.; Crawford, C.; Gascon, F.; Strobl, P. A.; Masek, J. G.; Neigh, C. S. R. The Harmonized Landsat and Sentinel-2 version 2.0 surface reflectance dataset. Remote Sensing of Environment 324 2025, 114723. [Google Scholar] [CrossRef]
- Kermarrec, G.; Montillet, J.-P.; Li, D. Uncertainty in urban climate modeling: Bridging the gap between science and policy. PLOS Climate 2025, 4(10), e0000743. [Google Scholar] [CrossRef]
- Klemmer, K.; Rolf, E.; Robinson, C.; Mackey, L.; Rußwurm, M. SatCLIP: Global, general-purpose location embeddings with satellite imagery. Proceedings of the AAAI Conference on Artificial Intelligence 2025, 39(4), 4347–4355. [Google Scholar] [CrossRef]
- Lacoste, A.; Lehmann, N.; Rodriguez, P.; Sherwin, E. D.; Kerner, H.; Lütjens, B.; Irvin, J. A.; Dao, D.; Alemohammad, H.; Drouin, A.; Gunturkun, M.; Huang, G.; Vazquez, D.; Newman, D.; Bengio, Y.; Ermon, S.; Zhu, X. X. GEO-Bench: Toward foundation models for Earth monitoring. In Advances in Neural Information Processing Systems 36; Datasets and Benchmarks Track, 2023. [Google Scholar] [CrossRef]
- Ma, Y.; Chen, S.; Ermon, S.; Lobell, D. B. Transfer learning in environmental remote sensing. Remote Sensing of Environment 301 2024, 113924. [Google Scholar] [CrossRef]
- Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; et al. Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning Proceedings of Machine Learning Research 2021, Vol. 139, 8748–8763. Available online: https://proceedings.mlr.press/v139/radford21a.html.
- Roberts, D. R.; Bahn, V.; Ciuti, S.; Boyce, M. S.; Elith, J.; Guillera-Arroita, G.; et al. Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography 2017, 40(8), 913–929. [Google Scholar] [CrossRef]
- Simumba, N.; Lehmann, N.; Fraccaro, P.; Alemohammad, H.; De Mel, G.; Khan, S.; Maskey, M.; Longépé, N.; Zhu, X. X.; Kerner, H.; Bernabe Moreno, J.; Lacoste, A. GEO-Bench-2: From performance to capability, rethinking evaluation in geospatial AI. Transactions on Machine Learning Research 2026. [Google Scholar] [CrossRef]
- Smith, M. J.; Fleming, L.; Geach, J. E. EarthPT: A time series foundation model for Earth observation [Preprint]. arXiv. 2023. Available online: https://arxiv.org/abs/2309.07207.
- Snyder, H. Literature review as a research methodology: An overview and guidelines. Journal of Business Research 104 2019, 333–339. [Google Scholar] [CrossRef]
- Stewart, A. J.; Lehmann, N.; Corley, I. A.; Wang, Y.; Chang, Y.-C.; Ait Ali Braham, N.; Sehgal, S.; Robinson, C.; Banerjee, A. SSL4EO-L: Datasets and foundation models for Landsat imagery. Advances in Neural Information Processing Systems 36 2023, 59787–59807. [Google Scholar] [CrossRef]
- Szwarcman, D.; Roy, S.; Fraccaro, P.; Gíslason, Þ. E.; Blumenstiel, B.; Ghosal, R.; de Oliveira, P. H.; de Sousa Almeida, J. L.; Sedona, R.; Kang, Y.; Chakraborty, S.; Wang, S.; Gomes, C.; Kumar, A.; Gaur, V.; Truong, M.; Godwin, D.; Khallaghi, S.; Lee, H.; Hsu, C.-Y.; Asanjan, A. A.; Mujeci, B.; Shidham, D.; Balogun, R. O.; Kolluru, V.; Keenan, T.; Arévalo, P.; Li, W.; Alemohammad, H.; Olofsson, P.; Mayer, T.; Hain, C.; Kennedy, R.; Zadrozny, B.; Bell, D.; Cavallaro, G.; Watson, C.; Maskey, M.; Ramachandran, R.; Moreno, J. B. Prithvi-EO-2.0: A versatile multitemporal foundation model for Earth observation applications. IEEE Transactions on Geoscience and Remote Sensing 64 2026, 4400120. [Google Scholar] [CrossRef]
- Tiggeloven, T.; Pfeiffer, S.; Matanó, A.; van den Homberg, M.; Thalheimer, L.; Reichstein, M.; Torresan, S. The role of artificial intelligence for early warning systems: Status, applicability, guardrails, and ways forward. iScience 2025, 28(11), 113689. [Google Scholar] [CrossRef] [PubMed]
- Towards responsible geospatial foundation models. Nature Machine Intelligence 7 2025, 1189. [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Advances in Neural Information Processing Systems 2017, Vol. 30, 5998–6008. [Google Scholar]
- Wang, Y.; Xiong, Z.; Liu, C.; Stewart, A. J.; Dujardin, T.; Bountos, N. I.; Zavras, A.; Gerken, F.; Papoutsis, I.; Leal-Taixé, L.; Zhu, X. X. Towards a unified Copernicus foundation model for Earth vision. In Proceedings of the IEEE/CVF International Conference on Computer Vision; 2025; pp. 9888–9899. Available online: https://openaccess.thecvf.com/content/ICCV2025/html/Wang_Towards_a_Unified_Copernicus_Foundation_Model_for_Earth_Vision_ICCV_2025_paper.html.
- Xiong, Z.; Wang, Y.; Zhang, F.; Stewart, A. J.; Hanna, J.; Borth, D.; Papoutsis, I.; Le Saux, B.; Camps-Valls, G.; Zhu, X. X. Neural plasticity-inspired multimodal foundation model for Earth observation. arXiv. 2024. Available online: https://arxiv.org/abs/2403.15356.
- Zhu, X. X.; Xiong, Z.; Wang, Y.; Stewart, A. J.; Heidler, K.; Wang, Y.; Yuan, Z.; Dujardin, T.; Xu, Q.; Shi, Y. On the foundations of Earth foundation models. Communications Earth & Environment 7 2026, 103. [Google Scholar] [CrossRef]
Figure 1.
Analytical framework used in this review to compare EO foundation models across training data and data preparation, pretraining objectives, architectural design, downstream adaptation and generalization, and empirical evidence.
Figure 1.
Analytical framework used in this review to compare EO foundation models across training data and data preparation, pretraining objectives, architectural design, downstream adaptation and generalization, and empirical evidence.

Figure 2.
End-to-end EO-FM workflow distinguishing external EO inputs and external data preparation from EO-FM core comprising input representation, pretrained representation learning, optional multimodal and contextual integration followed by downstream adaptation, and task-specific outputs. Evaluation checkpoints and evidence of adaptation and generalization are considered across the workflow.
Figure 2.
End-to-end EO-FM workflow distinguishing external EO inputs and external data preparation from EO-FM core comprising input representation, pretrained representation learning, optional multimodal and contextual integration followed by downstream adaptation, and task-specific outputs. Evaluation checkpoints and evidence of adaptation and generalization are considered across the workflow.

Figure 3.
Four-panel functional taxonomy of representative Earth observation foundation-model families: (1) optical/multispectral, (2) multimodal/cross-sensor, (3) time-series/forecasting, and (4) location-centered/embedding-oriented models. Each panel summarizes the primary data emphasis, typical learning strategy, representative models, and characteristic downstream strengths.
Figure 3.
Four-panel functional taxonomy of representative Earth observation foundation-model families: (1) optical/multispectral, (2) multimodal/cross-sensor, (3) time-series/forecasting, and (4) location-centered/embedding-oriented models. Each panel summarizes the primary data emphasis, typical learning strategy, representative models, and characteristic downstream strengths.

Figure 4.
Applications, representative empirical evidence, and cross-cutting research gaps in Earth observation foundation models (EO-FMs).
Figure 4.
Applications, representative empirical evidence, and cross-cutting research gaps in Earth observation foundation models (EO-FMs).

Figure 5.
Representative EO-FM adaptation workflows for (A) flood extent mapping and (B) wildfire mapping, showing sensor inputs, preprocessing, downstream adaptation, evaluation on unseen events or regions, and application outputs with predictive uncertainty and calibration information where available.
Figure 5.
Representative EO-FM adaptation workflows for (A) flood extent mapping and (B) wildfire mapping, showing sensor inputs, preprocessing, downstream adaptation, evaluation on unseen events or regions, and application outputs with predictive uncertainty and calibration information where available.

Figure 6.
Key conditions for trustworthy EO foundation models in scientific and policy applications.
Figure 6.
Key conditions for trustworthy EO foundation models in scientific and policy applications.

Table 1.
Eligibility criteria used to select literature for the review.
| Criterion type | Inclusion criteria | Exclusion criteria |
| Topical scope | Studies focused on EO foundation models, remote-sensing foundation models, geospatial foundation models, or closely related pretrained representation-learning systems for EO data | Studies focused only on general computer vision, NLP, or AI foundation models without clear EO relevance |
| Model scope | Large pretrained EO backbones, self-supervised EO models, multimodal EO-FMs, geographic/retrieval representation models, or forecasting-oriented EO or Earth-System foundation models directly relevant to the review. | Conventional task-specific supervised models without reusable pretrained representations |
| Methodological content | Studies reporting one or more of: pretraining corpus, learning objective, architecture, downstream adaptation strategy, generalization setting, or benchmark design | Insufficient methodological information to support comparison |
| Evidence base | Core model studies providing empirical evaluation, benchmark results, downstream performance, label-efficiency evidence, case studies, or evaluation across regions, time periods, sensors, or modalities; contextual studies providing substantive methodological, benchmarking, interpretability, reproducibility, uncertainty, or governance evidence | Studies providing neither empirical evidence relevant to EO applications nor substantive methodological or analytical contributions to the review |
| Document type | Peer-reviewed articles, conference papers, technical reports, benchmark papers, and high-relevance preprints. | Editorials, opinion pieces, news items, blog-only posts, promotional webpages, or slide decks without substantive technical information |
| Language and access | Full-text studies available in English | Non-English studies or records without accessible full text |
| Time window | Primarily January 2020–August 2026, with selected earlier foundational studies retained for conceptual or methodological background | Studies outside the primary review period that were not required for conceptual or methodological grounding |
| Relevance to review dimensions | training data and data preparation, pretraining objectives, architectural design, downstream adaptation and generalization, and empirical evaluation. | Studies not sufficiently relevant to these analytical dimensions |
Table 2.
Core concepts and working definitions used in this review.
| Term | Working definition |
| EO foundation model (EO-FM) | A model pretrained on broad EO data to learn reusable representations that can support multiple downstream applications through task-specific adaptation. |
| Pretraining corpus | The large EO dataset or collection used during pretraining, potentially spanning sensors, locations, time periods, resolutions, and modalities. |
| Pretraining objective | The learning signal used to learn representations, including masked reconstruction, contrastive alignment, temporal prediction, supervised objectives, or combinations of these approaches. |
| Harmonization | Procedures that make observations from different sensors, missions, or acquisition periods more comparable by reducing geometric, radiometric, spatial, temporal, or metadata inconsistencies. |
| Multimodal fusion | The combination of complementary information from different sensing modalities or data sources, such as SAR and optical imagery, within a joint representation or prediction framework. |
| Representation reuse | Use of the same pretrained backbone or embeddings across different downstream applications, such as classification, segmentation, regression, or change detection. |
| Downstream adaptation | Methods used to specialize a pretrained model for a particular downstream task, including linear probing, full fine-tuning, adapters, learnable prompts, or other parameter-efficient approaches. |
| Generalization | The ability of an adapted model to maintain performance when evaluated under conditions not adequately represented during training, such as new regions, time periods, sensors, or environmental conditions. |
Table 4.
Representative EO foundation models and their dominant taxonomy characteristics.
| Model | Dominant family | Main modality logic | Primary downstream capability | Typical evidence base | Source |
| Prithvi-EO-2.0 | Optical / multispectral, multi-temporal | Harmonized Landsat–Sentinel optical time series | Segmentation, classification, and multi-temporal downstream applications | GEO-Bench and application-specific evaluations | Szwarcman et al. (2026) |
| Clay | Multisensor/ embedding-oriented | Multisensor EO inputs with wavelength, spatial, temporal, and location metadata | Reusable embeddings for downstream analysis | Documentation and deployment-oriented resources | Clay Foundation (2024) |
| TerraFM | Multimodal SAR–optical | Sentinel-1 SAR + Sentinel-2 optical with adaptive fusion | Classification and segmentation | GEO-Bench and Copernicus-Bench | Danish et al. (2026) |
| DOFA | Cross-sensor / spectral-flexible | Wavelength-conditioned dynamic parameterization enabling a shared Transformer to accommodate different sensors and spectral configurations | Classification and segmentation across heterogeneous sensors and spectral configurations | Multi-sensor evaluation across 12 EO tasks, including sensors not represented during pretraining | Xiong et al. (2024) |
| Copernicus-FM | Multimission / metadata-aware | Sentinel-based surface and atmospheric observations with sensor- and metadata-aware encoding | Multimission downstream tasks across heterogeneous Copernicus observations | Copernicus-Bench and ICCV evaluation | Wang et al. (2025) |
| EarthPT | Time-series / forecasting | Autoregressive EO surface-reflectance forecasting | Forecasting and temporal representation | Forecasting experiments and downstream embedding evaluation | Smith et al. (2023) |
| Aurora | Earth-system forecasting | Large-scale heterogeneous geophysical data integration | Weather, air-quality, ocean-wave, and tropical-cyclone forecasting | Peer-reviewed Nature evaluation and operationally relevant forecasting experiments | Bodnar et al. (2025) |
| SatCLIP | Location-centered / geographic representation | Contrastive alignment of satellite imagery and geographic coordinates | Geographic embeddings and location-conditioned prediction | AAAI paper and evaluation across downstream geospatial prediction tasks | Klemmer et al. (2025) |
| AlphaEarth Foundations | Foundation model as a dataset / embedding-oriented | Embedding-field integration across multiple EO sources and contextual information | Precomputed annual embeddings for mapping and monitoring | Preprint and released embedding-product evaluation | Brown et al. (2025) |
Source note. Table 4 is the authors’ synthesis based on Prithvi-EO-2.0 (Szwarcman et al., 2026), Clay Foundation documentation (Clay Foundation, 2024), TerraFM (Danish et al., 2026), DOFA (Xiong et al., 2024), Copernicus-FM (Wang et al., 2025), EarthPT (Smith et al., 2023), Aurora (Bodnar et al., 2025), SatCLIP (Klemmer et al., 2025), and AlphaEarth Foundations (Brown et al., 2025).
Table 5.
Task-aware comparative evidence from representative EO foundation models.
| Model | Representative evaluation setting | Selected reported performance | Interpretation for model use | Source |
| Prithvi-EO-2.0 | GEO-Bench and application studies | The 600M-TL model reported an 8% aggregate improvement over Prithvi-EO-1.0 across GEO-Bench tasks; 90.3 mIoU on Sen1Floods11; 90.5 mIoU and 98.1 mF1 on wildfire-scar mapping | Multi-temporal optical applications, including land-cover, crop, and hazard mapping | Szwarcman et al. (2026) |
| Copernicus-FM | Copernicus-Bench multimission evaluation | 87.2 OA on EuroSAT-S1; 97.9 OA on EuroSAT-S2; 77.9 mAP on BigEarthNet-S1; 79.0 mAP on BigEarthNet-S2; 77.7 mIoU on Flood-S1; 66.7 mIoU on Cloud-S2; RMSE 2.8 on AQ-NO2-S5P | Multimission Copernicus applications, including metadata-aware processing | Wang et al. (2025) |
| TerraFM | GEO-Bench and Copernicus-Bench | 87.8 OA on EuroSAT-S1; 99.1 OA on EuroSAT-S2; 84.4 mAP on BigEarthNet-S2; 67.9 mIoU on Cloud-S2; 55.4 mIoU on DFC2020-S1. | SAR–optical classification and segmentation | Danish et al. (2026) |
| DOFA | Cross-sensor classification and segmentation | 91.9 frozen OA and 96.1 fine-tuned OA for DOFA ViT-Large on RESISC-45, with additional evaluation across EO classification and segmentation tasks. | Applications requiring sensor and spectral flexibility. | Xiong et al. (2024) |
| SatCLIP | Geographic and location-conditioned tasks | Evaluated across multiple location-dependent downstream tasks using general-purpose location embeddings. | Geographic representation and location-conditioned prediction. | Klemmer et al. (2025) |
| EarthPT | EO time-series forecasting | Typical NDVI prediction error of approximately 0.05 over a five-month test-set horizon. | Temporal representation and forecasting. | Smith et al. (2023) |
| Aurora | Earth-system forecasting | Reported improvements for air quality, ocean waves, tropical-cyclone tracks, and weather forecasting. | Broader Earth-system forecasting rather than conventional EO image mapping. | Bodnar et al. (2025) |
Note. OA = overall accuracy; mAP = mean average precision; mIoU = mean intersection over union; RMSE = root mean square error.
Table 6.
Common EO-FM strengths and limitations in practice.
| Strength or capability | Why it matters for EO | Limitation to consider |
| Label-efficient downstream adaptation | Can reduce dependence on expensive task-specific labels | Benefits may decrease when the downstream domain differs substantially from pretraining |
| Cross-sensor adaptability | Can accommodate different spectral configurations or sensing modalities | Performance may degrade for unseen sensors or poorly represented acquisition conditions |
| Spatiotemporal representation | Can capture seasonality, recurrence, and environmental dynamics | Temporal leakage or geographic shortcuts can inflate apparent performance |
| Reusable pretrained representations | Can reduce repeated model development and task-specific supervision | Undocumented model or preprocessing changes can reduce reproducibility |
| Standardized benchmarking | Improves consistency of model comparison | Existing benchmarks may not represent rare events or genuinely out-of-distribution deployment |
| Multimodal learning | Can combine complementary information from multiple sensing sources | Benefits depend on modality availability, task relevance, and representation strategy |
Source note.Table 6 is the authors’ synthesis based on benchmark and perspective literature on EO/Geo foundation models, especially GEO-Bench, GEO-Bench-2, Copernicus-FM, and the GeoFM perspective paper (Janowicz et al., 2025; Lacoste et al., 2023; Simumba et al., 2026; Wang et al., 2025).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.