Submitted:
07 August 2026
Posted:
12 August 2026
You are already at the latest version
Abstract
Egocentric vision has become a central paradigm for studying human behaviour, interaction, and intent from a first-person perspective, supporting a growing range of tasks such as action recognition, temporal segmentation, and action anticipation. Progress in this area has been driven largely by the availability of public datasets, which differ substantially in sensing configurations, annotation strategies, application domains, and temporal structure. Despite the rapid expansion of available resources, the egocentric dataset landscape remains fragmented, with limited cross-dataset interoperability and uneven coverage across domains, modalities, and task formulations. This survey presents a comprehensive, dataset-centric analysis of the current egocentric and action-related video ecosystem. We curate and systematically analyse 75 publicly documented datasets spanning more than a decade of research, covering both egocentric and selected exocentric benchmarks that are widely used for action-centric and anticipatory modelling. To organise this landscape, we introduce a unified taxonomy structured around five complementary design dimensions: perspective, domain, sensing modality, annotation structure, and task support. This taxonomy enables consistent comparison across heterogeneous datasets and provides a principled framework for analysing how design choices influence supported tasks and evaluation practices. Beyond cataloguing datasets, we examine temporal trends, distributional patterns, and co-occurrence statistics across the curated corpus, revealing recurring biases toward specific domains, limited multimodal coverage, and a widespread reliance on post hoc constructions for action anticipation. We further discuss annotation heterogeneity, scale–diversity trade-offs, and the ethical and legal constraints that shape egocentric data collection and release. By consolidating dispersed resources into a coherent taxonomy and identifying systematic gaps in current benchmarks, this survey aims to support informed dataset selection, facilitate cross-dataset analysis, and guide future efforts in egocentric dataset design and benchmarking.
Keywords:
egocentric vision
; action anticipation
; datasets
; taxonomy
; first-person video
1. Introduction
Public datasets have long been recognised as the driving force behind the solid foundations of modern computer vision infrastructure. Without these open datasets, research and innovation would lose momentum. More specifically, in the field of egocentric vision, rapid advancements in both hardware and software have significantly expanded the range and quality of available data sources. The increasing use of RGB cameras, depth sensors, gaze trackers, inertial units, and audio devices has enabled the creation of large-scale multimodal datasets with considerably richer representations than before. Therefore, the egocentric vision community has developed several benchmark datasets that differ in scale, sensing configurations, annotation detail, and target tasks. These datasets capture daily activities, social interactions, hand-object manipulations, and both short and long-term procedural behaviours, offering diverse resources for first-person perception research.
Egocentric vision, by definition, focuses on capturing the world from the perspective of an acting individual, allowing detailed analysis of hand-object interactions, gaze behaviour, goal-directed manipulation, and extended activity sequences. This unique viewpoint has supported a rapidly expanding range of applications, including cognitive and assistive robotics, immersive AR systems, activity recognition, action anticipation, and more broadly, human-computer interaction. However, despite these developments, progress in egocentric vision remains fundamentally tied to the availability, diversity, and openness of the underlying datasets.
Prior work in computer vision has repeatedly demonstrated that public, well-documented datasets are not merely convenient resources but core scientific infrastructure. As highlighted in [1], datasets identify key research areas, including the topics the community investigates, the way algorithms are benchmarked, and the evaluation standards that become widely adopted. Similarly, [2] underscores that the historical evolution of computer vision has been inseparable from the development of increasingly larger and more diverse benchmarks. According to [3,4], the pivotal importance of datasets is essential for ensuring the reproducibility of experiments to unveil significant insights and patterns obtained directly from the data, underscoring the significance of reproducibility in scientific research. These assertions are foundational to the integrity, consistency, and long-term sustainable advancement of computer vision research.
In the domain of egocentric vision, there is a notable emphasis on data dependency. Studies like [5] and [6] argue that progress in first-person action recognition, action anticipation, and behavior analysis has been predominantly facilitated by the introduction of large-scale datasets. Recent works, such as [7], state that the field is still struggling with a persistent inadequacy of specialized, high-quality egocentric datasets, particularly those covering various settings, multi-sensor data, and extended temporal labels. Regarding action anticipation and human forecasting tasks, [8] similarly note that the available benchmarks exhibit inconsistencies across domains and lack consistency in their annotation protocols and prediction horizons. Despite the progress enabled by recent benchmarks, the current egocentric dataset landscape remains fragmented and uneven. Existing datasets differ substantially in sensing configurations, annotation schemas, action taxonomies, and domain coverage, leading to limited interoperability and significant barriers to cross-dataset comparison. Most of the widely used benchmarks tend to focus on a restricted set of environments, most commonly kitchens, daily living spaces, driving scenes, or constrained industrial workflows, while offering sparse temporal labeling and limited variation in sensing setups. Another constraint arising from this restricted domain coverage is its impact on domain generalization and transfer learning. Specifically, prior studies in the field of computer vision indicate that models trained on a narrow set of visual categories often fail to generalize well to unfamiliar conditions. According to the authors in [9], the constrained scope of domains in numerous benchmarks results in assessments that lack the capacity to adequately represent the variability present in real-world scenarios. This limitation is further illustrated in [10] where it was observed that the occurrence of inadequate generalization frequently originates from a restricted range of domains represented in training datasets. Complementary findings from transfer learning surveys [11,12] emphasise that mismatches between pre-training datasets and downstream tasks remain a key bottleneck, especially when pre-training corpora do not capture the types of environments or behaviours encountered in egocentric vision. These observations reinforce the idea that insufficient domain and annotation variability is a fundamental constraint that directly affects the reliability and robustness of trained models, rather than just a question of dataset design.
Moreover, existing surveys typically organise datasets around specific tasks such as egocentric action recognition, video summarisation, future prediction, or action anticipation, rather than providing a unified, cross-domain perspective on the full dataset ecosystem. As a result, the community still lacks a comprehensive and up-to-date resource that aims to catalogue all the egocentric and action-related datasets, analyze their characteristics, and organize them within a coherent taxonomy.
Furthermore, the creation and dissemination of egocentric datasets are subject to ethical and legal restrictions that are less restrictive in many other areas of computer vision. Recordings captured from a first-person perspective typically encompass intimate environments, including bystanders, as well as biometric markers such as facial features, vocal characteristics, unique physical gestures, and patterns of eye contact. Within legal frameworks like the EU General Data Protection Regulation (GDPR), biometric data used for uniquely identifying a person is treated as a special category of personal data and are therefore subject to stricter conditions on collection, processing, and sharing [13]. For instance, recent large-scale egocentric benchmarks such as [14] and [15] explicitly document their reliance on institutional ethics review, informed consent procedures, and, where necessary, de-identification of personally identifiable information before public release. While these measures are essential for protecting participants, they also increase the cost and complexity of dataset collection and can limit how openly data may be distributed or extended.
Given these scientific, practical, and regulatory constraints, a more structured understanding of the egocentric dataset landscape is still lacking. Existing surveys provide valuable summaries of specific tasks and sub-domains, however, they typically discuss datasets only in the context of individual problems or benchmark settings. Notable examples include surveys on egocentric future prediction [5], action recognition [6], broader egocentric trends [7], and human action anticipation [8]. While this task-centred perspective is useful for comparing methods within a particular problem, it gives only a partial account of how datasets differ across domains, sensing modalities, annotation practices, and temporal structure. As the number and diversity of available egocentric datasets continue to grow, a unified taxonomy becomes essential for organising this landscape, revealing systematic gaps, and supporting more informed decisions about dataset selection, learning approaches, and evaluation protocols.
Building on these observations, this study provides a comprehensive, dataset-centric analysis of egocentric and action-oriented video resources. We curate an extensive collection of 75 publicly accessible datasets spanning more than a decade of research and covering diverse environments, activities, and sensing configurations. A central aim of this work is to introduce a unified taxonomy that organises datasets according to perspective, domain, modality, and annotation design, offering a structured framework for understanding how existing resources relate to one another. Beyond cataloguing datasets, we analyse common patterns and discrepancies across current resources, identify systematic gaps in domain coverage and multimodal support, and summarise the ethical and regulatory factors that influence dataset collection and public dissemination. Collectively, these contributions establish a consolidated foundation for researchers working with egocentric video data and for future efforts on dataset design and benchmarking.
The remainder of this paper is structured as follows. Section 2 defines the background and scope of the survey, introducing egocentric vision and action-centric video understanding, and formalising the tasks, modalities, and dataset characteristics considered. Section 3 presents a unified taxonomy for egocentric and action-related datasets, organising existing resources along five core design dimensions and motivating the criteria used for systematic categorisation. Section 4 provides a comprehensive overview of the curated dataset corpus, analysing temporal trends, perspective, domain and modality distributions, and highlighting structural patterns in current dataset design. Section 5 discusses the broader implications of the survey findings, examining persistent limitations in the dataset landscape, ethical and accessibility considerations, and outlining directions for future dataset development and benchmarking.
2. Background and Scope of the Survey
In this section, we outline the conceptual foundations and scope of this work. We begin by defining egocentric vision and positioning it within the broader landscape of action-centric video understanding. Particular attention is given to the sensing configurations, viewpoint constraints, and behavioural cues that differentiate first-person footage from conventional third-person video. We then summarise the modalities, annotation types, and temporal structures encountered across the datasets considered in this survey, highlighting the methodological diversity that motivates a principled taxonomy. Finally, we formalise the inclusion criteria and scope of our analysis to ensure that the subsequent classifications and comparisons are grounded in transparent and replicable assumptions.
2.1. Egocentric Vision: Definition and Characteristics
Egocentric vision, also known as first-person vision, refers to visual information captured by a camera mounted on the head or body of an individual, providing a perspective that matches the camera-wearer’s line of sight. Unlike third-person or exocentric video, which observes actions from an external point of view, egocentric footage reflects the perceptual and motor experience of the individual, capturing what they see and how they interact with objects and their surroundings. This perspective introduces distinctive visual and temporal characteristics, including motion induced by head and body movement, frequent visibility of the individual’s hands and manipulated objects, and a fragmented, dynamically changing view of the environment driven by the camera-wearer’s gaze direction.
The literature reviewed in this work consistently highlights these features. Núñez-Marcos et al. [6], for example, describe egocentric video as exhibiting substantial non-linear and unpredictable motion arising from the wearer’s movements, alongside limited access to broader scene context. They identify ego-motion, hand-object interaction, and a constantly shifting field of view tied to the individual’s focus as key challenges for visual recognition systems. Early foundational work in egocentric action understanding [16] similarly notes that wearable-camera footage is dominated by frequent ego-motion and that alignment with the wearer’s attention exposes behaviourally rich cues such as hand pose, head movement, and gaze signals that are largely absent from conventional third-person video datasets.
In addition, egocentric vision imposes unique assumptions on the structure of activities. Because the camera wearer is both the observer and the actor, egocentric video encodes the intentional structure of interaction: tasks are organised around hand-object contact, object state changes, and gaze-driven attention patterns. These interaction-centric signals have become central to modern egocentric benchmarks such as EPIC-KITCHENS [15,17], EGTEA Gaze+[16,18,19] and Ego4D [14], which explicitly annotate hands, objects, object states, and gaze. Their design reflects the need to capture the full sensorimotor context underlying everyday activities, enabling downstream tasks such as action recognition, hand-object interaction modelling, and action anticipation. Overall, the defining characteristics of egocentric video wearer-driven motion, embodied viewpoints, interaction-centric scene structure, and attention-linked visual framing form the conceptual basis of the dataset taxonomy developed in this survey. These properties shape both the modalities typically captured (e.g., RGB, depth, IMU, gaze, audio) and the annotation paradigms required (e.g., hand masks, object bounding boxes, action segments, narrations). Understanding these characteristics is essential for interpreting the methodological diversity of the datasets analysed in this work and provides the foundation for the analyses conducted in subsequent sections.
2.2. Egocentric Video Understanding and Action-Centric Tasks
Egocentric video understanding focuses on analysing human behaviour from the perspective of an acting individual, where actions are observed as part of an ongoing stream of goal-directed interaction rather than as isolated motion patterns. Unlike third-person video analysis, which often emphasises full-body kinematics and scene-level context, egocentric understanding centres on object manipulation, hand activity, visual attention, and task progression. As a result, the definition and formulation of action-centric tasks in egocentric vision differ both conceptually and practically from their exocentric counterparts.
Early work in egocentric vision primarily addressed action recognition, where the objective is to identify the activity being performed based on visual evidence accumulated over time. In first-person video, however, actions are frequently short, overlapping, and visually ambiguous, with their semantic meaning determined by object interactions rather than body motion alone. This has led to task formulations that explicitly incorporate hand–object interactions, verb–noun action decompositions, and object state changes, as seen in influential benchmarks such as EPIC-KITCHENS and EGTEA Gaze+. These formulations reflect the observation that, in egocentric settings, actions are more naturally defined through what is being manipulated and how, rather than through global pose or movement. Beyond recognition, egocentric video has motivated tasks that operate over longer temporal extents, including temporal segmentation, activity parsing, and procedural understanding. Such tasks aim to decompose extended activity streams into meaningful action units or phases, capturing the structure of multi-step behaviours such as cooking, assembly, or daily routines. Prior surveys note that these tasks are particularly challenging in first-person video due to weak visual boundaries between actions, frequent interleaving of activities, and variations in execution order [6,7]. Consequently, egocentric datasets often support task formulations that emphasise temporal continuity and contextual reasoning rather than discrete clip-level classification.
A defining characteristic of egocentric video understanding is the growing emphasis on future-oriented tasks. Action anticipation, future action forecasting, and goal inference aim to predict what the camera wearer will do next based on partial observations. These tasks rely on cues such as object affordances, hand trajectories, gaze patterns, and task context, which frequently precede observable action execution. Surveys on action anticipation highlight that first-person data is particularly well suited for such predictive tasks, as it captures early indicators of intent that are unavailable or attenuated in third-person views [5,8]. At the same time, these tasks introduce intrinsic uncertainty, since multiple future actions may be consistent with the same observed state. In addition to action-centric tasks, egocentric datasets increasingly support interaction-focused and multimodal problem formulations. These include hand–object interaction modelling, object state transition detection, gaze and attention prediction, and joint vision–language reasoning based on narrations or spoken instructions. Such tasks reflect the embodied and sensor-rich nature of egocentric data and blur the boundary between perception and higher-level reasoning. As noted in recent surveys, the diversity of supported tasks across egocentric benchmarks is both a strength and a source of fragmentation, as datasets are often designed with specific task assumptions that limit direct comparability [6,7].
In general, egocentric video understanding encompasses a broad spectrum of action-centric tasks that differ in temporal scope, semantic abstraction, and predictive intent. These task formulations are tightly coupled to dataset design choices, including annotation strategies and temporal organisation, and they motivate the need for a structured framework to analyse how datasets support different forms of inference. This perspective provides the conceptual foundation for the taxonomy introduced in Section 3, where task support is formalised as a distinct dimension of dataset design.
2.3. Modalities and Annotation Paradigms in Egocentric Datasets
Egocentric datasets differ from conventional video benchmarks not only in viewpoint, but also in the breadth and role of sensory information they capture. Because first-person video reflects the perceptual experience of an acting individual, egocentric data collection often extends beyond RGB video to include complementary sensory signals that convey motion, attention, and interaction. As a result, both the choice of sensing modalities and the design of annotation paradigms become central elements of dataset construction rather than secondary implementation details.
RGB video constitutes the core modality in nearly all egocentric datasets, providing access to scene appearance, manipulated objects, and hand activity. However, first-person RGB streams are frequently affected by rapid ego-motion, occlusions, and viewpoint changes, which can obscure fine-grained interaction cues. To mitigate these limitations, many datasets incorporate additional modalities such as depth, inertial measurements, audio, or gaze signals. Depth and three-dimensional information enable spatial reasoning about object geometry and hand–object alignment, while inertial sensors provide direct measurements of head or body motion that are difficult to infer reliably from video alone. Audio recordings capture environmental sounds, object interactions, and speech, offering complementary cues for action boundaries and context, particularly in kitchen, industrial, and social settings [7,14].
Gaze and eye-tracking data represent a distinctive modality in egocentric vision, as they provide direct access to the camera wearer’s attentional focus. Prior work has shown that gaze often precedes physical interaction and serves as an early indicator of intent, making it especially valuable for predictive and anticipatory tasks [16,18]. Datasets that include gaze supervision therefore support forms of reasoning that are difficult to realise using visual appearance alone, but they also introduce additional complexity in data collection and synchronisation.
Beyond sensing, annotation paradigms in egocentric datasets exhibit substantial diversity in both form and granularity. Unlike many third-person benchmarks that rely primarily on clip-level labels, egocentric datasets frequently combine multiple annotation types to capture different aspects of behaviour. These may include frame-level labels such as hand presence or object visibility, temporal annotations defining action segments, and higher-level semantic labels describing tasks or procedures. Several datasets further annotate object states, interaction events, or narrations aligned with the video stream, reflecting the fact that actions in egocentric settings are often defined by changes in object configuration rather than by motion patterns alone [15,17]. The coexistence of multiple annotation paradigms within a single dataset reflects the inherently multi-layered nature of egocentric behaviour, but it also creates challenges for dataset interoperability and method comparison. Differences in label definitions, temporal alignment, and semantic abstraction complicate the transfer of models across benchmarks and make it difficult to compare results obtained under different annotation assumptions. Surveys of egocentric vision consistently identify this heterogeneity as a major source of fragmentation in the field, particularly as datasets expand to support increasingly diverse tasks and modalities [6,7].
The combination of rich multimodal sensing and heterogeneous annotation paradigms distinguishes egocentric datasets from traditional video benchmarks. These characteristics enable more expressive modelling of interaction, intent, and context, but they also introduce design trade-offs that directly influence which learning problems can be addressed. Understanding how modalities and annotations are selected and combined at the dataset level is therefore essential for interpreting experimental results and motivates the structured treatment of these aspects in the taxonomy introduced in the next section.
2.4. Temporal Structures and Dynamics in Egocentric Data
A defining property of egocentric video is the temporal organisation of human behaviour. Unlike third-person datasets, where actions are often short, visually distinct, and body-motion centric, egocentric datasets capture extended, multi-step activities driven by object manipulation, task goals, and interaction sequences. As a result, temporal structure becomes a central modelling axis, shaping both the annotation design and the types of tasks that datasets support. More specifically, temporal structure in egocentric datasets is not only a characteristic of the recorded behaviour, but also a deliberate design choice that shapes how data can be used. Decisions such as whether actions are annotated at the frame or segment level, whether temporal boundaries are sharply defined or loosely specified, and whether future events are explicitly labelled determine which forms of inference a dataset can support. For example, datasets with densely annotated temporal boundaries facilitate fine-grained segmentation and early action detection, whereas datasets that encode extended procedural structure enable reasoning over long-term dependencies and task progression. As a result, temporal organisation acts as an implicit constraint on task formulation and evaluation, even when datasets capture similar activities.
Many egocentric datasets provide frame-level labels, such as hand presence, object visibility, or action primitives. These annotations support fine-grained temporal reasoning but often depend on dense supervisory effort. More commonly, datasets adopt segment-level annotations, in which actions are defined over temporally contiguous intervals. These segments may vary widely in duration from sub-second atomic manipulations to multi-minute procedural steps, reflecting the inherently hierarchical nature of first-person behaviour. A second temporal characteristic is the presence of overlapping or interleaved actions. Egocentric activities frequently involve parallel processes, such as preparing an object while reaching for another, making the temporal boundaries between actions less distinct than in curated third-person benchmarks. This motivates annotation schemes that capture verb–noun pairs, object state changes, and interactions, as seen in modern datasets such as EPIC-KITCHENS and Ego4D. Egocentric datasets also exhibit extended temporal dependencies. Daily activities, assembly tasks, and long procedures unfold over hundreds or thousands of frames, with meaningful relationships between early and later stages of the sequence. These long-range dependencies underpin tasks such as action anticipation, next-object prediction, and future forecasting, which rely on modelling behavioural intent and procedural progression rather than isolated motion cues.
Finally, temporal dynamics in egocentric video are strongly influenced by ego-motion, gaze shifts, and task-driven visual framing. Rapid head movements may produce momentary instability, while gaze-aligned framing emphasises objects and regions of interest at different points in time. Together, these factors shape the temporal structure of egocentric datasets and motivate the need for specialised temporal models, annotation conventions, and evaluation protocols. Understanding these dynamics is crucial for interpreting the diverse dataset designs surveyed in this work and provides the conceptual foundation for the unified taxonomy that is later introduced.
2.5. Objectives and Scope
This survey provides a systematic and comprehensive overview of datasets relevant to egocentric vision and action anticipation. To ensure methodological transparency and reproducibility, we follow a structured, PRISMA-inspired [20] selection procedure that emphasises systematic identification, screening, and eligibility checking. Our final corpus comprises 75 distinct datasets, each documented with a complete metadata record capturing modality, task, perspective, domain, and annotation type. The initial dataset pool was formed through an extensive Google Scholar search, followed by a detailed examination of the resulting literature, ultimately identifying 77 action-anticipation papers and 12 egocentric-vision studies.
To compile the dataset corpus, we enumerated all datasets referenced across these works and collected their corresponding publications and documentation. For each dataset, we constructed a unified metadata entry describing its sensing modalities, annotation types, supported tasks, domain category, and whether it adopts an egocentric or exocentric perspective. During consolidation, multiple references to the same dataset were merged, naming inconsistencies were standardised, and ambiguous cases were resolved by consulting the primary sources. All metadata entries were manually verified for internal coherence and completeness. A dataset was included in the final list only if its metadata profile could be reliably reconstructed from publicly available documentation. This requirement ensures that all datasets incorporated into our taxonomy can be analysed uniformly across key dimensions such as domain, modality, annotation structure, and anticipation relevance, an essential prerequisite for reproducible cross-dataset comparison and principled categorisation in Section 3.
Furthermore, to define the final dataset corpus, we applied a set of principled inclusion criteria to ensure analytical consistency. Each dataset was required to contain human-centered video depicting actions, interactions, or procedures, and to provide meaningful supervisory signals such as action labels, temporal segments, object annotations, or interaction markers. Datasets were included only when accompanied by sufficient publicly available documentation, such as an associated publication, technical report, or official project description, allowing reliable metadata reconstruction and traceable source verification. In addition, a dataset was retained only if it was relevant to egocentric vision or action anticipation, either through an explicit first-person viewpoint or through its established use as a benchmark for forecasting or action-centric tasks. Finally, to support the structured comparisons developed in subsequent sections, we retained only datasets whose domain, modality, annotation structure, and task definitions could be consistently identified and aligned with the taxonomy introduced in the following sections. This alignment guarantees coherence across the survey and enables reproducible, dataset-level analyses.
3. Taxonomy of Egocentric and Action-Related Datasets
This section presents a unified taxonomy that consistently organises datasets related to egocentric vision and action anticipation. The proposed taxonomy offers a systematic framework for comparing diverse datasets based on standardised criteria, despite variations in sensing setups, annotation methods, task objectives, and application domains. Rather than grouping datasets by individual benchmarks or problem formulations, the taxonomy serves as an abstraction of common design principles occur across the dataset landscape.
3.1. Design Principles and Scope of the Taxonomy
The proposed taxonomy is organised around five complementary dimensions: perspective, domain, modality, annotation structure, and task support. These dimensions were selected to capture the core dataset design choices that most strongly influence model development, training strategies, and evaluation protocols in egocentric vision and action anticipation. Together, they provide a compact yet expressive representation of the dataset landscape without imposing assumptions tied to specific benchmarks, dataset families, or learning paradigms.
A central design principle of the taxonomy is orthogonality. Each dimension captures a distinct aspect of dataset construction and can be analysed independently of the others. For instance, datasets originating from the same application domain may differ substantially in sensing modalities or annotation granularity, while datasets sharing similar modalities may support different tasks or temporal horizons. By separating these dimensions, the taxonomy avoids conflating unrelated design factors and enables flexible, cross-cutting comparisons across heterogeneous datasets. Another guiding principle is dataset-agnosticism. The taxonomy is not tailored to a particular benchmark suite or evaluation protocol. Instead, it is designed to accommodate a broad range of datasets relevant to action-centric video understanding, including both egocentric and exocentric resources that are commonly used for recognition, temporal segmentation, and anticipation tasks. This allows the taxonomy to unify dataset ecosystems that are often treated separately in the literature, while preserving their essential distinctions. The taxonomy is also designed to be operational. Each dataset considered in this survey can be mapped unambiguously to a specific configuration within the taxonomy based on publicly available documentation. This ensures that the taxonomy is not merely conceptual, but can support reproducible dataset analysis, structured comparison, and principled discussion of dataset trends and gaps.
3.2. Taxonomy Dimensions
3.2.1. Perspective Dimension
The perspective dimension constitutes the most fundamental axis of the proposed taxonomy, as it distinguishes datasets according to the viewpoint from which actions and interactions are observed. As illustrated in Figure 1, this dimension separates datasets into egocentric and exocentric categories, reflecting whether visual data are captured from the first-person perspective of an acting individual or from an external observer’s viewpoint.
This survey primarily focuses on egocentric datasets, which are recorded using wearable or body-mounted sensors and reflect the visual stream experienced by the actor during task execution. This perspective is characterised by strong ego-motion, frequent hand and object visibility, gaze-aligned framing, and an interaction-centric scene structure. As a result, egocentric datasets are particularly well suited for modeling fine-grained human-object interactions, procedural activities, and anticipatory behaviours, where understanding intent and future actions requires access to the actor’s perceptual context. Exocentric datasets, in contrast, capture actions from a third-person viewpoint, typically using static or mobile cameras placed in the environment. These datasets often provide a more stable global view of the scene and clearer body pose information, and have therefore been widely used in action recognition, temporal segmentation, and forecasting benchmarks. While exocentric datasets do not capture the embodied perceptual experience of the actor, several are included in this survey due to their historical importance and continued use in action-centric and anticipatory modelling. Significantly, the proposed taxonomy treats perspective as an independent design dimension rather than a strict inclusion criterion. Although egocentric datasets form the core of our analysis, selected exocentric datasets are retained when they play a significant role in benchmarking or methodological development. This design choice enables a unified view of the dataset ecosystem and supports comparative analysis across perspectives, highlighting how viewpoint assumptions influence visual cues, annotation strategies, and task formulations. By explicitly modelling perspective as a top-level taxonomy dimension, the proposed framework clarifies how viewpoint assumptions shape dataset design and downstream modelling choices, and provides a foundation for analysing cross-perspective transfer, hybrid learning setups, and emerging efforts to bridge egocentric and exocentric representations.
3.2.2. Domain Dimension
Domain dimension categorizes datasets according to the real-world environment, activity context, and application setting in which the data were captured. It characterizes where actions occur and what types of activities are performed, providing essential contextual information for interpreting visual appearance, interaction patterns, temporal structure, and task semantics. Prior studies on dataset design and evaluation have repeatedly highlighted domain as a critical factor shaping dataset bias, model behaviour, and downstream generalisation [1,2]. As illustrated in Figure 2, the proposed taxonomy organises egocentric datasets into a diverse set of domains that reflect the breadth of scenarios explored in contemporary first-person video research.
A large proportion of egocentric datasets are situated in domains related to daily activities, kitchen and cooking scenarios. This observation is consistent with existing surveys of egocentric vision and action anticipation, which report a strong concentration of benchmarks focused on routine household tasks and object-centric procedures [5,6,7]. These environments typically involve routine object manipulation, sequential procedures, and goal-directed behaviour, making them particularly suitable for studying fine-grained hand–object interactions, action segmentation, and action anticipation. Their prevalence reflects both their relevance to human-centred applications and the practical advantages they offer for controlled data collection, dense annotation, and repeatable experimental design.
Beyond everyday activities, the taxonomy explicitly includes domains such as sports and exercise, entertainment, and outdoor activities, which introduce greater motion dynamics, environmental variability, and interaction diversity. Sports-oriented datasets often capture rapid, skill-intensive movements and dynamic body–environment interactions, while outdoor scenarios introduce challenges related to lighting variation, clutter, and scene unpredictability. Such domains are frequently cited as important prototype environments for evaluating the robustness of egocentric models under less constrained and more realistic conditions [7,8].
In addition, further distinction is being made in domains involving structured and safety-critical settings, including industrial and assembly, healthcare and medical, and education and training. Datasets in these categories typically emphasise procedural correctness, long-horizon task execution, and fine-grained action sequencing, often under strict constraints or high-stakes conditions. Prior work has shown that these domains pose distinct challenges for temporal modelling and supervision design, particularly in anticipation and forecasting tasks [5,8]. Such settings are especially relevant for applications in assistive robotics, skill assessment, and decision support, and frequently require precise temporal annotations and multimodal sensing. Social and multi-agent contexts are captured through the social interaction domain, which encompasses activities involving communication, collaboration, and interpersonal dynamics. These scenarios introduce additional complexity due to the presence of multiple actors, conversational cues, and social intent, often motivating the inclusion of audio, language, or gaze annotations, as discussed in prior egocentric and action-centric surveys [6,7]. Finally, the taxonomy accounts for synthetic or simulated datasets and cross-domain or mixed-activity datasets, which either rely on virtual environments for scalable data generation or intentionally combine multiple domains to support generalisation and transfer learning studies.
In particular, the proposed taxonomy treats domain as an independent design dimension rather than a proxy for task, modality, or annotation strategy. Datasets drawn from the same domain may differ substantially in sensing configuration, supervision granularity, and supported tasks, while similar tasks may be studied across multiple domains. This observation aligns with findings from domain generalisation and transfer learning literature, which emphasise that limited domain diversity is a key factor underlying poor cross-dataset generalisation [9,10,11,12]. By modelling domain explicitly and independently, the taxonomy enables systematic analysis of domain coverage, exposes biases toward frequently studied environments, and facilitates the identification of underexplored or emerging application settings within the egocentric dataset landscape. This domain-level organisation provides a principled foundation for analysing generalisation, robustness, and transfer across environments in later sections of the survey.
3.2.3. Modality Dimension
The modality dimension characterises datasets according to the types of sensory signals available during data acquisition and processing. It defines what information is observed beyond raw visual appearance and directly determines the perceptual cues accessible to learning models. In egocentric vision, the selection of appropriate modalities fundamentally influences the design decision, as first-person viewpoints naturally support rich multimodal sensing that reflects the embodied experience of the actor. Figure 3 provides an overview of the sensing modalities represented across contemporary egocentric datasets.
The most prevalent modality remains RGB video, which serves as the foundational signal in nearly all datasets. RGB streams capture scene appearance, object context, and hand–object interactions, forming the basis for tasks such as action recognition, segmentation, and anticipation. However, egocentric RGB data is often affected by rapid ego-motion, motion blur, and frequent occlusions, motivating the inclusion of complementary modalities that stabilise or enrich perception. To address motion-related challenges, many datasets incorporate motion and kinematic signals, including inertial measurement units (IMU), optical flow, and ego-motion estimates. These modalities provide explicit information about head, body, or hand movement and are especially valuable in domains involving dynamic activities such as sports, outdoor tasks, or physically intensive procedures. Motion signals often improve robustness under visually ambiguous conditions and support fine-grained temporal reasoning. Depth and 3D information constitute another important modality class. Depth sensors, stereo cameras, and reconstructed 3D geometry enable reasoning about spatial relationships, object affordances, and interaction geometry from a first-person perspective. These signals are particularly prominent in datasets targeting hand pose estimation, human–object interaction modelling, and procedural understanding, where accurate spatial alignment between hands, tools, and objects is essential.
Egocentric datasets also increasingly include audio as a complementary modality. First-person audio captures environmental sounds, object interactions, and speech, providing cues that are often weak or absent in visual streams alone. Audio is especially informative in kitchen, industrial, and social interaction domains, where acoustic events correlate strongly with action boundaries and object state changes. Its inclusion has enabled multimodal learning setups that fuse auditory and visual information for improved action understanding and anticipation. Another distinctive modality in egocentric vision is gaze and eye-tracking data. Gaze signals provide direct access to the actor’s attentional focus, revealing intent and task progression before observable physical actions occur. Datasets incorporating gaze enable early anticipation, intent prediction, and fine-grained analysis of human attention mechanisms, offering supervision signals that are largely unavailable in exocentric settings. Finally, a growing subset of datasets integrates language-based modalities, such as narrations, spoken instructions, procedural text, or aligned transcripts. These signals bridge perception and semantics, supporting tasks like instruction following, video–language grounding, and multimodal reasoning. Language annotations are particularly relevant for long-horizon procedural activities, where visual cues alone may be insufficient to disambiguate task structure or goals.
Derived and post-processed modalities. As illustrated in Figure 3, the taxonomy explicitly distinguishes between raw sensor signals and derived modalities. Modalities represented with dashed outlines correspond to signals that are not directly captured by sensors, but are obtained through post-processing or model-based estimation applied to raw data. Examples include optical flow computed from RGB frames, hand or body pose estimated via vision-based models, 3D scene reconstructions, gaze vectors inferred from eye-tracking pipelines, and SLAM (Simultaneous Localization and Mapping) or visual-inertial odometry trajectories. While derived modalities often provide higher-level, task-relevant representations, their availability depends on algorithmic choices, annotation pipelines, and post-hoc processing decisions rather than on the sensing hardware alone. Notably, the proposed taxonomy treats modality as an independent design dimension rather than a proxy for task complexity or dataset scale. Datasets sharing the same modalities may target different tasks, while similar tasks may be supported by vastly different sensing configurations. By explicitly modelling modality as a separate axis, the taxonomy enables systematic analysis of multimodal coverage, exposes biases toward visually dominant datasets, and highlights opportunities for richer sensor fusion and representation learning in future dataset design. This modality-level organisation provides a principled foundation for understanding how sensory choices and post-processing pipelines shape model capabilities, generalisation behaviour, and robustness across domains and tasks.
3.2.4. Annotation Structure Dimension
The annotation structure dimension describes how supervisory information is represented within a dataset and plays a decisive role in determining which learning settings and evaluation protocols can be meaningfully supported. In egocentric vision, annotation design is particularly influential due to the complexity of first-person interactions, the uncertainty of temporal boundaries, and the need to encode both perceptual detail and semantic intent. Instead of treating annotations as a single, uniform attribute, the proposed taxonomy decomposes annotation structure into a set of complementary axes that capture variation in temporal resolution, semantic level, and interaction representation.
One major source of variation concerns temporal granularity. Egocentric datasets may provide annotations at the frame level, the segment level, or through multi-level temporal descriptions. Frame-level supervision includes signals such as hand visibility, gaze fixation, object presence, or pixel-wise masks, and enables precise temporal analysis of interaction dynamics. However, such annotations are costly to produce and are therefore often confined to short clips or limited subsets of data. Segment-level annotations, which associate actions or activities with contiguous time intervals, constitute the dominant approach in large-scale egocentric benchmarks. These segments may correspond to atomic manipulations, verb–noun action units, or higher-level procedural steps, and typically exhibit substantial variation in duration. Some datasets further introduce hierarchical temporal annotations, decomposing extended activities into sub-actions or phases in order to reflect the inherently multi-scale organisation of egocentric behaviour.
A second axis relates to semantic abstraction. Low-level annotations capture perceptual entities such as hands, objects, object states, poses, or motion primitives, whereas higher-level annotations encode semantic constructs including actions, goals, tasks, or procedures. Many egocentric datasets combine multiple levels of abstraction, for example by annotating object interactions alongside the semantic actions they realise. This layered supervision facilitates the learning of representations that connect perception with meaning, but it also raises challenges related to label consistency, ontology design, and alignment across datasets. In practice, differences in action vocabularies, verb–noun formulations, and task definitions substantially limit interoperability between benchmarks.
Interaction representation constitutes a third defining component of annotation structure. Unlike third-person datasets, egocentric benchmarks frequently include interaction-focused annotations, such as hand–object contact, object state transitions, or gaze-to-target associations. These annotations reflect the embodied nature of first-person vision, where actions are characterised primarily by manipulation and attention rather than by full-body motion. Datasets that provide such interaction cues are especially valuable for studying human–object interaction, skill execution, and anticipatory reasoning, as they reveal causal and intentional signals that precede observable outcomes. At the same time, the availability and form of interaction annotations vary considerably across datasets, ranging from explicit relational labels to interaction information that is only implicitly encoded within action segments.
Finally, the annotation structure dimension accounts for whether a dataset provides anticipation-aligned supervision. Although many egocentric datasets include temporal annotations, only a limited subset explicitly supports anticipatory evaluation by defining observation windows, prediction horizons, or labels associated with future actions. In most cases, anticipation tasks are constructed by truncating sequences from recognition-oriented datasets and reinterpreting existing annotations. While this practice has accelerated methodological development, it also introduces ambiguity with respect to temporal alignment, action onset definition, and evaluation consistency. By making annotation structure explicit, the proposed taxonomy distinguishes datasets that are deliberately designed for anticipation from those that merely allow anticipatory use through post-hoc adaptation. Figure 4 provides a structured visual summary of the annotation structure dimension introduced above. The taxonomy organises egocentric annotations into five complementary categories according to the nature of the supervisory signal they encode. Frame-level annotations capture instantaneous perceptual states, such as object presence, hand occupancy, or gaze fixation, and support fine-grained temporal analysis. Temporal annotations represent actions and behaviours over extended intervals, including action segments, verb–noun sequences, narrations, and explicitly defined anticipation targets. Interaction-level annotations focus on relational and contact-based information, such as hand–object contact points, object state transitions, and manipulation affordances, reflecting the interaction-centric nature of egocentric vision. Spatial and trajectory-based annotations encode longer-term geometric structure, including 3D human and object trajectories, camera motion estimates, and global scene alignment, enabling reasoning over space, motion, and embodiment. Finally, audio, text, and language annotations capture complementary semantic and contextual signals, such as speech transcripts, narration timestamps, and dense video captions, which are increasingly used for multimodal learning and intent modelling. Together, these categories illustrate how egocentric datasets distribute supervision across perceptual, temporal, relational, spatial, and semantic dimensions, without implying that all datasets provide all annotation types.
3.2.5. Task Support Dimension
Task support dimension describes the classes of learning and inference problems that a dataset is designed to enable. In egocentric vision, task formulation has a direct impact on annotation strategies, evaluation protocols, and the interpretation of reported results. Instead of the assumption that the datasets are uniformly applicable across action-centric problems, the proposed taxonomy distinguishes datasets according to the tasks they are explicitly constructed to support, independently of the sensing modalities or annotation formats they employ.
A fundamental distinction within the task taxonomy is between retrospective and prospective tasks. Retrospective tasks operate on complete or partially observed sequences to infer past or ongoing activities and include action recognition, temporal segmentation, detection, and summarisation. These tasks account for the majority of existing egocentric benchmarks and have driven much of the progress in first-person video understanding. Prospective tasks, by contrast, require reasoning about future events, actions, or intentions from incomplete observations. Action anticipation, future action forecasting, next-object prediction, and goal inference belong to this category and impose substantially different modelling assumptions and evaluation requirements. Within the class of prospective tasks, action anticipation occupies a distinctive role. Whereas recognition aims to classify an action once sufficient visual evidence has accumulated, anticipation requires predicting actions before their execution is completed or even initiated. This temporal asymmetry introduces inherent uncertainty, since multiple future actions may be consistent with the same observed context. As a result, anticipation benchmarks require datasets with clearly specified observation windows, prediction horizons, and action onset definitions in order to enable consistent and reproducible evaluation. When such temporal constraints are absent, early recognition is often indistinguishable from genuine predictive inference.
In many cases, datasets that are commonly used for anticipation were originally developed for recognition or segmentation. Anticipation benchmarks are therefore often constructed by truncating input sequences and reinterpreting existing action annotations. Although this practice has facilitated rapid experimentation and comparison, it also blurs the conceptual boundary between retrospective and prospective inference and introduces ambiguity in temporal alignment and evaluation consistency. The task support dimension explicitly distinguishes datasets that natively support anticipation through deliberate temporal design from those that allow anticipatory use only through post hoc adaptation. Beyond recognition and anticipation, egocentric datasets support a wide range of additional tasks, including hand object interaction modelling, object state change detection, skill assessment, procedural understanding, gaze and attention modelling, and multimodal reasoning. Some datasets are intentionally designed to support multiple tasks, while others focus on a narrow problem formulation. Task support is therefore determined not simply by the presence of labels, but by the joint specification of annotation structure, temporal organisation, and evaluation protocols.
Figure 5 presents a structured overview of the task landscape in egocentric vision. Tasks are grouped according to their semantic scope and temporal orientation, ranging from low-level perceptual tasks such as object recognition and tracking, through action and activity understanding, to prospective tasks including anticipation and forecasting. The taxonomy further highlights complementary task families related to gaze and attention, hand object interaction and manipulation, three-dimensional mapping and tracking, vision language reasoning, and summarisation. This organisation clarifies which datasets are appropriate for specific research objectives and exposes systematic gaps in task coverage across existing benchmarks.
4. Dataset Overview
4.1. Dataset Corpus Summary
This survey curates a corpus of 75 publicly documented datasets depicted in Table A1 that are relevant to egocentric vision and action-centric video understanding, covering a period of more than a decade. The collection reflects the progressive evolution of first-person data resources, from early small-scale benchmarks focusing on isolated activities to recent large-scale datasets designed for long-horizon, multimodal, and anticipatory reasoning. Egocentric datasets prevail in the corpus, showcasing the increased focus on first-person perception, interaction-centric modelling, and intent-aware reasoning in current studies on computer vision research. A smaller but non-negligible subset of exocentric datasets is retained due to their historical importance and continued use as benchmarks for action recognition, temporal segmentation, and forecasting. In addition, several datasets adopt mixed or multi-perspective configurations, enabling comparative analysis across viewpoints and supporting emerging research on cross-perspective transfer and joint ego-exo learning.
From a sensing perspective, the corpus is characterised by a strong reliance on RGB video, which remains the foundational modality across nearly all datasets. However, a substantial fraction of recent datasets augment visual streams with additional signals such as audio, inertial measurements, gaze tracking, depth, or hand pose, reflecting a clear trend toward richer multimodal representations. While unimodal datasets remain prevalent, particularly among earlier benchmarks, the increasing availability of multimodal data has expanded the range of tasks that can be meaningfully supported, including anticipation, interaction modelling, and procedural reasoning. Specifically, the dataset corpus covers a wide range of domains, particularly focusing on daily activities and kitchen environments in various applications. This highlights the significance of these environments for applications focused on human users and the benefits they provide for structured data gathering and labelling. Concurrently, the corpus includes datasets drawn from sports, industrial and assembly tasks, healthcare, education, social interaction, and outdoor activities, enabling analysis of how domain characteristics influence sensing choices, annotation strategies, and supported tasks.
In general, the curated dataset corpus provides broad coverage across perspective, modality, domain, annotation structure, and task support, forming a representative snapshot of the current egocentric and action-related dataset landscape. This diversity establishes the empirical foundation for the analyses that follow. In the remainder of this section, we examine how datasets are distributed across time, perspective, domain, and modality, and how scale and temporal structure vary across benchmarks.
4.2. Temporal Evolution of Datasets (2008–2025)
A look at dataset release timelines shows a clear change both in volume and in viewpoint over time. Early benchmarks are relatively few and are mostly built around exocentric views, which aligns with the earlier emphasis on third-person action recognition and fixed camera setups. From the mid-2010s onward, dataset releases begin to increase more rapidly, alongside the wider availability of wearable sensors and a growing interest in first-person data. In the most recent years, egocentric datasets make up most new releases, pointing to a longer-term move toward studying interaction, intention, and behaviour from the actor’s perspective. The continued rise in the total number of available datasets suggests that the field is moving into a more stable phase. New datasets are often released alongside existing benchmarks instead of supplanting them, typically adding coverage for additional viewpoints, domains, or sensing setups. As a result, recent work tends to expand what is already available rather than redefining the benchmark space.
Figure 6.
Temporal evolution of egocentric and exocentric datasets between 2008 and 2025. Stacked bars indicate the number of datasets introduced per year, grouped by perspective (egocentric, exocentric, and multi-view).Years with no new dataset releases are explicitly included to preserve the continuity of the timeline.
Figure 6.
Temporal evolution of egocentric and exocentric datasets between 2008 and 2025. Stacked bars indicate the number of datasets introduced per year, grouped by perspective (egocentric, exocentric, and multi-view).Years with no new dataset releases are explicitly included to preserve the continuity of the timeline.

Figure 7.
Cumulative perspective trends over time. The curves depict the cumulative number of datasets introduced between 2008 and 2025 for each perspective category (egocentric, exocentric, and multi-view). The trajectories highlight a gradual early dominance of exocentric datasets, followed by a pronounced acceleration in egocentric dataset growth after the mid-2010s, indicating a sustained shift toward first-person data collection in action-centric and anticipatory research.
Figure 7.
Cumulative perspective trends over time. The curves depict the cumulative number of datasets introduced between 2008 and 2025 for each perspective category (egocentric, exocentric, and multi-view). The trajectories highlight a gradual early dominance of exocentric datasets, followed by a pronounced acceleration in egocentric dataset growth after the mid-2010s, indicating a sustained shift toward first-person data collection in action-centric and anticipatory research.

4.3. Distribution by Perspective, Domain, and Modality
Dataset design in egocentric and action-centric video benchmarks can be described along three core dimensions: viewpoint perspective, sensing modality, and application domain. Considering these dimensions jointly reveals recurring patterns in how datasets are constructed, as well as uneven coverage of behaviours and interaction contexts. Figure 8, Figure 9 and Figure 10 summarise these distributions, presenting both individual dimension statistics and their combined structure.
Viewed together, the distributions indicate that egocentric datasets constitute the majority of available benchmarks, while exocentric and multi-view datasets remain concentrated in domains where stable scene observation is required. With respect to sensing configuration, RGB video is almost universally present, whereas datasets offering additional sensory signals are comparatively limited. The joint distribution of perspective and domain further suggests that certain environments are consistently associated with specific viewpoints, pointing to recurring design choices rather than uniform sampling of real-world activity settings.
4.4. Summary and Transition
This section examined the curated dataset corpus at a broad level, focusing on how egocentric and action-related datasets vary over time, viewpoint, application domain, and sensing modality. Looking across release years, the data show a steady rise in new datasets over the last decade, alongside a noticeable move toward egocentric capture in more recent work. This shift mirrors increased interest in first-person perception and interaction-driven modelling, and is closely tied to the growing practicality of wearable sensors for large-scale data collection.
Several recurring design patterns emerge from the distributional analysis. Egocentric datasets now dominate the benchmark landscape, while exocentric and multi-view datasets remain in use for settings where a stable scene layout or global motion information is important. Most datasets rely on RGB video, with only a smaller fraction incorporating additional signals such as audio, gaze, depth, or inertial measurements. The limited use of these modalities is largely driven by the practical effort required to capture, synchronise, and annotate them.
From a domain perspective, datasets are concentrated around everyday activities, with kitchen-based scenarios appearing far more often than other settings. Other environments, including sports, industrial work, healthcare, social interaction, and outdoor scenarios, receive comparatively less attention. These patterns are shaped by practical constraints, perceived application relevance, and established research trajectories, and they influence which behaviours and interactions are most frequently studied. Examining perspective and domain jointly further reveals that certain environments are repeatedly paired with specific viewpoints, suggesting common conventions in dataset construction.
Quantitative analysis was intentionally restricted to dataset attributes that could be reconstructed in a consistent manner across sources. Measures of dataset scale were handled with particular caution, as statistics such as video count, total duration, and annotation density are reported inconsistently across benchmarks. Focusing on directly comparable attributes helps avoid conclusions based on aggregate statistics that are defined differently across datasets. As a whole, the analysis outlines how current datasets are typically designed and where coverage remains limited. These observations set the stage for a closer discussion of annotation practices, accessibility, and structural limitations in existing benchmarks. The following sections build on this overview to examine where current datasets remain insufficient and how future dataset efforts might address these shortcomings.
5. Conclusion and Future Directions
Despite the substantial progress enabled by recent egocentric datasets, the analysis conducted in this survey reveals a number of persistent limitations and structural gaps in the current dataset landscape. These limitations are not isolated artefacts of individual benchmarks, but rather recurring patterns that emerge across perspective, domain, modality, annotation structure, and task support. Identifying these gaps is essential for interpreting reported results, understanding the scope of current benchmarks, and guiding future dataset design.
A first and most evident limitation concerns domain concentration. As shown in the dataset distribution analysis, a large fraction of existing egocentric datasets are situated in a small number of environments, most notably kitchen and daily living settings. While these domains are well suited for studying object-centric manipulation and procedural activities, their dominance results in a skewed representation of real-world behaviour. Domains involving outdoor environments, social interaction, safety-critical scenarios, or high-variability conditions remain comparatively underrepresented. This imbalance has direct implications for generalisation: models trained and evaluated primarily on routine, highly structured activities may fail to transfer to more diverse or less constrained settings. Consequently, performance gains reported on dominant benchmarks may overestimate real-world robustness, particularly for tasks such as action anticipation that rely on contextual and environmental cues.
A second gap relates to the limited native support for anticipation-oriented evaluation. Although many datasets are frequently used for action anticipation and future prediction, only a small subset has been explicitly designed with anticipatory supervision in mind. In most cases, anticipation benchmarks are constructed post hoc by truncating recognition-oriented datasets and reinterpreting existing action annotations. This practice introduces ambiguity regarding observation windows, prediction horizons, and the precise definition of action onset. As a result, evaluation protocols vary widely across studies, making it difficult to compare results or assess progress consistently. The lack of datasets that natively encode anticipation-specific temporal structure remains a fundamental limitation for principled evaluation of predictive models.
A further limitation concerns modality imbalance and sparse multimodal coverage. While RGB video is nearly ubiquitous across egocentric datasets, additional modalities such as audio, gaze, inertial measurements, and depth are available only in a minority of benchmarks. Even when present, these modalities are often confined to specific domains or collected under limited conditions. This uneven distribution restricts the development and evaluation of genuinely multimodal models, as methods trained on one dataset may not be transferable to others with different sensing configurations. Moreover, the scarcity of datasets combining multiple complementary modalities constrains systematic study of sensor fusion strategies and limits understanding of how non-visual cues contribute to anticipation and interaction modelling.
Annotation heterogeneity represents another persistent challenge. Egocentric datasets differ substantially in how actions, interactions, and temporal boundaries are defined. Variations in annotation granularity, semantic abstraction, and label vocabularies complicate cross-dataset comparison and hinder reuse of trained models. Even when datasets target similar activities, differences in verb–noun formulations, temporal segmentation criteria, and interaction definitions often prevent direct alignment. This lack of standardisation does not merely reflect stylistic differences, but has concrete methodological consequences: models trained under one annotation regime may perform poorly or unpredictably when transferred to another, and reported improvements may be specific to a particular labeling convention rather than indicative of general progress.
Finally, the dataset landscape exhibits a recurring scale–diversity trade-off. Large-scale datasets typically focus on a narrow set of domains and tasks, enabling extensive training but limited environmental diversity. Conversely, datasets that span multiple domains or include richer multimodal signals are often smaller in scale, restricting their suitability for data-intensive learning approaches. This trade-off constrains the development of models that are both statistically robust and broadly generalisable. In practice, researchers are frequently forced to prioritise either scale or diversity, with few datasets offering both simultaneously.
Taken together, these limitations highlight that current egocentric datasets, while powerful, provide an incomplete and uneven foundation for action-centric and anticipatory modelling. The gaps identified above are not shortcomings of individual dataset efforts, but rather systemic characteristics of the field’s evolution to date. Addressing these issues will require coordinated efforts toward broader domain coverage, clearer anticipation-oriented annotation protocols, richer multimodal sensing, and improved alignment across annotation schemes. These observations motivate the guidelines for future dataset development discussed in the following subsection and provide context for interpreting results reported across existing benchmarks.
Author Contributions
Conceptualization, A.M. and E.S.; methodology, A.M.; data curation, A.M.; writing—original draft preparation, A.M.; writing—review and editing, A.M. and E.S.; supervision, E.S.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
Not applicable.
Conflicts of Interest
The authors declare no conflicts of interest.
Appendix A Dataset Inventory
Table A1.
Unified overview of action-centric video datasets. Datasets are summarized according to capture perspective, sensing modalities, annotation structure, scale, and application context. Abbreviations: Persp.=Perspective; E=Egocentric; X=Exocentric; M=Multi-view; RGB=Color video; D=Depth; A=Audio; IMU=Inertial Measurement Unit; Gz=Gaze; Hd=Head pose; Act.=Action labels; Obj.=Object annotations; Tmp.=Temporal boundaries; Cls.=Number of classes; Avg.=Average duration (seconds); Tot.=Total duration (hours); Vids.=Number of videos.
Table A1.
Unified overview of action-centric video datasets. Datasets are summarized according to capture perspective, sensing modalities, annotation structure, scale, and application context. Abbreviations: Persp.=Perspective; E=Egocentric; X=Exocentric; M=Multi-view; RGB=Color video; D=Depth; A=Audio; IMU=Inertial Measurement Unit; Gz=Gaze; Hd=Head pose; Act.=Action labels; Obj.=Object annotations; Tmp.=Temporal boundaries; Cls.=Number of classes; Avg.=Average duration (seconds); Tot.=Total duration (hours); Vids.=Number of videos.
| Dataset | Year | Persp. | Modalities | Annotations | Cls. | Avg. (s) | Tot. (h) | Vids. | Context | |||||||
| RGB | D | A | IMU | Gz | Hd | Act. | Obj. | Tmp. | ||||||||
| 50 Salads [21] | 2013 | X | ✓ | ✓ | ✓ | ✓ | ✓ | 10 | N/A | 4.5 | 50 | Kitchen, food preparation | ||||
| ActivityNet [22] | 2015 | X | ✓ | ✓ | ✓ | 203 | N/A | 849 | 27801 | Daily human activities | ||||||
| ADL [23] | 2022 | X | ✓ | ✓ | ✓ | ✓ | 31 | N/A | N/A | N/A | Smart home, indoor, daily living | |||||
| AEA [24] | 2024 | E | ✓ | ✓ | ✓ | N/A | N/A | 7.3 | 143 | Indoor daily livings, smart home activities | ||||||
| Assembly101 [25] | 2022 | M | ✓ | ✓ | ✓ | ✓ | ✓ | 202(C) 1380(FG) |
426 | 513 | 4321 | Assembly / Disassembly | ||||
| AssemblyHands [26] | 2023 | M | ✓ | ✓ | ✓ | ✓ | 6 | N/A | N/A | N/A | Procedural assembly, hand–object interactions | |||||
| AVA [27] | 2018 | X | ✓ | ✓ | ✓ | 80 | N/A | 107.5 | 430 | Movies, daily human activities | ||||||
| BDDA [28] | 2024 | X | ✓ | N/A | N/A | N/A | N/A | Autonomous driving, adverse weather and lighting | ||||||||
| BDD100K [29] | 2020 | X | ✓ | ✓ | ✓ | ✓ | 10 | 40 | 1111 | 100000 | Autonomous driving | |||||
| BEOID [30] | 2014 | E | ✓ | ✓ | ✓ | 75 | 16.3 | N/A | 58 | Daily object interactions(kitchen, office, gym, household) | ||||||
| BON [31] | 2022 | E | ✓ | ✓ | ✓ | 18 | 5 | 3.7 | 2639 | Office/workplace activities | ||||||
| Breakfast [32] | 2014 | X | ✓ | ✓ | ✓ | 10(A) 48(C) |
N/A | 77 | 520 | Kitchen, cooking activities | ||||||
| CAD-120 [33] | 2012 | X | ✓ | ✓ | ✓ | ✓ | ✓ | 10 | N/A | N/A | 120 | Indoor daily activities | ||||
| Charades [34] | 2016 | X | ✓ | ✓ | ✓ | ✓ | 157(A) 46(Obj) |
30.1 | 82.3 | 9848 | Indoor daily activities, household environments | |||||
| Charades-Ego [35] | 2018 | M | ✓ | ✓ | ✓ | 157 | 31.2 | 69.3 | 8000 | Indoor daily activities, household environments | ||||||
| CMU-MMAC [36] | 2008 | M | ✓ | ✓ | ✓ | ✓ | ✓ | 31(A) 5(recipes) |
900 | 6.25 | 25 | Kitchen, cooking activities | ||||
| COIN [37] | 2019 | X | ✓ | ✓ | ✓ | 180 | 141.6 | 476.6 | 11827 | Instructional/daily activities | ||||||
| CrossTask [38] | 2019 | X | ✓ | ✓ | ✓ | 83 | N/A | 51 | 4700 | Instructional, how-to activities | ||||||
| EasyCom [39] | 2021 | E | ✓ | ✓ | N/A | N/A | 5.3 | 15 | Conversational interactions | |||||||
| EDINA [40] | 2022 | E | ✓ | ✓ | N/A | N/A | 16 | N/A | Indoor daily activities, scene understanding | |||||||
| EDUB-Seg [41] | 2017 | E | ✓ | ✓ | N/A | N/A | N/A | N/A | Daily life, photo streams (wearable camera) | |||||||
| EGTEA Gaze + [18] | 2018 | E | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 106 | 4.2 | 28 | 86 | Kitchen/food preparation with first-person video, audio, binocular gaze tracking, and sparse hand masks | |||
| Ego4D [14] | 2022 | E | ✓ | ✓ | ✓ | ✓ | N/A | N/A | 3025 | 3670 | Daily life activities | |||||
| EgoBody3M [42] | 2024 | E | N/A | N/A | 31.8 | 2688 | Indoor environments/VR body motion | |||||||||
| EgoCVR [43] | 2024 | E | ✓ | ✓ | ✓ | ✓ | N/A | 7.9 | N/A | 2295 | Daily-life activities (derived from Ego4D and FHO) | |||||
| EgoExo4D [44] | 2024 | M | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | N/A | N/A | 1286 | 5035 | Skilled human activities (e.g., cooking, bike repair, health care, music, etc. | ||
| EgoExo-Fitness [45] | 2024 | M | ✓ | ✓ | ✓ | 12 | N/A | 32 | 1276 | Fitness/exercise activities | ||||||
| EgoExoLearn [46] | 2024 | M | ✓ | ✓ | ✓ | 19 | N/A | N/A | 1600 | Procedural activities | ||||||
| Ego-HOIBench [47] | 2025 | E | ✓ | ✓ | ✓ | 18 | N/A | N/A | N/A | Daily human–object interactions | ||||||
| EgoGesture [48] | 2018 | E | ✓ | ✓ | ✓ | ✓ | 83 | N/A | N/A | 2081 | Hand gesture interaction | |||||
| EgoK360 [49] | 2020 | E | ✓ | ✓ | ✓ | 45 | N/A | N/A | 127 | Daily human activities captured with 360° cameras | ||||||
| EgoMe [50] | 2025 | M | ✓ | ✓ | ✓ | ✓ | ✓ | 184 | 18.25 | 82.46 | 15804 | Real-world daily activities | ||||
| EgoObjects [51] | 2023 | E | ✓ | ✓ | 368 | N/A | N/A | N/A | Indoor everyday object-centric scenes captured from wearable devices across households and offices | |||||||
| EgoOops [52] | 2024 | E | ✓ | ✓ | ✓ | N/A | N/A | N/A | N/A | Procedural daily activities with mistake actions, aligned to instructional texts | ||||||
| EgoPet [53] | 2024 | E | ✓ | N/A | N/A | 84 | 819 | Animal perception and interaction | ||||||||
| EgoProceL [54] | 2022 | E | ✓ | ✓ | ✓ | 16 | 769.2 | 62 | 329 | Procedural daily-life activities | ||||||
| EgoTracks [55] | 2023 | E | ✓ | ✓ | ✓ | N/A | 367.9 | 602.9 | 5708 | Long-term object tracking in daily-life activities (sourced from Ego4D) | ||||||
| EgoVid5M [56] | 2024 | E | ✓ | ✓ | ✓ | N/A | N/A | N/A | 5M | Large-scale video–action dataset derived from Ego4D for video generation | ||||||
| EPIC-Fields [57] | 2023 | E | ✓ | ✓ | ✓ | ✓ | 90K | N/A | 99 | 671 | Kitchen/cooking activities with 3D camera pose and geometry augmentation | |||||
| EPIC-KITCHENS-100 [15] | 2022 | E | ✓ | ✓ | ✓ | ✓ | ✓ | 97(V), 300(N), 4053(A) | N/A | 100 | 700 | Daily unscripted kitchen activities | ||||
| EPIC-KITCHENS-55 [17] | 2018 | E | ✓ | ✓ | ✓ | ✓ | ✓ | 125(V), 331(N) | N/A | 55 | 432 | Daily kitchen activities | ||||
| EPIC-Sounds [58] | 2025 | E | ✓ | ✓ | ✓ | ✓ | 44 audio classes | 514 | 100 | 700 | Daily-life kitchen activities/sounds (derived from EPIC-KITCHENS-100 audio) | |||||
| EPIC-Tent [59] | 2019 | E | ✓ | ✓ | ✓ | 12(A), 38(FG) | 816 | 5.4 | 24 | Outdoor camping scenario, recording of non-rigid object manipulation during tent assembly | ||||||
| First-Person-Social-Interactions [16] | 2012 | E | ✓ | ✓ | 6 | N/A | N/A | 8 | Day-long real-world social events (theme parks, group outings), captured with head-mounted cameras, focusing on social interaction patterns based on face attention and first-person motion | |||||||
| FPHA [60] | 2018 | E | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 45 | N/A | 58.6 | 1175 | Daily hand–object manipulation actions captured with an RGB-D shoulder-mounted camera, featuring accurate 3D hand pose annotations obtained via magnetic sensors, designed for studying first-person hand action recognition, hand pose estimation, and hand–object interaction. | |||
| FPV-O [61] | 2018 | E | ✓ | ✓ | ✓ | 20 | N/A | 3.02 | 12 | Office activities recorded with a chest-mounted GoPro camera, covering person-to-person interactions, person-to-object interactions, and locomotion, annotated with temporal action segments in a realistic office environment. | ||||||
| GTEA [19] | 2011 | E | ✓ | ✓ | ✓ | 7 | N/A | N/A | 28 | Kitchen/food preparation | ||||||
| GTEA Gaze [16] | 2012 | E | ✓ | ✓ | ✓ | ✓ | 7 | N/A | N/A | 17 | Kitchen/food preparation activities captured from a first-person perspective with synchronized eye-gaze tracking for studying gaze-guided action recognition | |||||
| GTEA Gaze + [62] | 2015 | E | ✓ | ✓ | ✓ | ✓ | 44 | N/A | 9 | 86 | Kitchen/food preparation activities captured in real-world environments with head-mounted cameras and binocular eye-gaze tracking | |||||
| H2O [63] | 2021 | E | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 36 | N/A | N/A | N/A | Indoor hand–object manipulation | |||
| HD-EPIC [64] | 2025 | E | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | N/A | 15.9 | 41.3 | 156 | Unscripted kitchen activities in real home environments | ||
| HOI4D [65] | 2022 | E | ✓ | ✓ | ✓ | ✓ | 16 | N/A | N/A | 4K | Indoor human–object interaction with category-level object manipulation | |||||
| HoloAssist [66] | 2023 | E | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 414(C,A), 39(V), 90(N) | N/A | 166 | 2221 | Real-world assistive, interactive task completion |
| HowToDiv [67] | 2025 | E | ✓ | ✓ | N/A | N/A | 24 | 507 | Step-by-step procedural task assistance instructional videos transformed into expert–novice dialogues | |||||||
| IndustReal [68] | 2024 | E | ✓ | ✓ | ✓ | 12 | N/A | N/A | 200 | Industrial-like procedural tasks with explicit modeling of correct and erroneous execution steps | ||||||
| JAAD [69] | 2016 | E | ✓ | ✓ | N/A | 5-15 | N/A | 346 | Urban driving scenarios focusing on joint attention between drivers and pedestrians | |||||||
| KrishnaCam [70] | 2016 | E | ✓ | N/A | N/A | 70.2 | 460 | Longitudinal single-person video stream for scene understanding and motion prediction | ||||||||
| MECCANO [71] | 2023 | E | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 61(A), 12(V), 20(Obj) | 21 | 7 | 20 | Industrial-like human–object interaction during assembly tasks | ||
| MPII Cooking [72] | 2012 | X | ✓ | ✓ | 65 | N/A | >8 | 44 | Continuous kitchen recordings for fine-grained cooking activity classification and detection | |||||||
| MPII Cooking 2 [73] | 2015 | X | ✓ | ✓ | ✓ | ✓ | 67(FC, A), 155(Obj), 59(A) | N/A | >27 | 273 | Fine-grained and composite cooking activity understanding | |||||
| MultiTHUMOS [74] | 2018 | X | ✓ | ✓ | ✓ | 65 | N/A | 30 | 413 | Dense multilabel action annotations over long, unconstrained internet videos extending the THUMOS dataset | ||||||
| N-EPIC-Kitchens [75] | 2021 | E | ✓ | ✓ | ✓ | 8 | N/A | N/A | N/A | Event-based action recognition dataset extending EPIC-Kitchens with simulated neuromorphic event streams | ||||||
| Nymeria [76] | 2024 | E | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | N/A | ~900 | 300 | 1200 | Large-scale multimodal dataset capturing daily human motion with synchronized vision, audio, IMUs, gaze, full-body motion capture, and hierarchical language annotations | |||
| PH-Ego [77] | 2023 | E | ✓ | ✓ | ✓ | 9 | ~390 | N/A | N/A | Personal-health–oriented action anticipation dataset capturing scripted kitchen activities such as hygiene and food preparation for studying domain shift and untrimmed anticipation | ||||||
| PIE [78] | 2019 | X | ✓ | ✓ | ✓ | ✓ | 6 | ~600 | >6 | N/A | On-board driving dataset for pedestrian intention estimation and trajectory prediction in urban traffic scenes | |||||
| THUMOS14 [79] | 2017 | X | ✓ | ✓ | ✓ | 101 | N/A | ~254 | N/A | Large-scale benchmark for action recognition and temporal action detection in untrimmed real-world videos collected from YouTube | ||||||
| TV-I [80] | 2010 | X | ✓ | ✓ | ✓ | 4 | N/A | N/A | 300 | Realistic human–human interaction recognition dataset composed of clips extracted from TV shows for interaction retrieval in unconstrained scenes | ||||||
| TVSeries [81] | 2016 | X | ✓ | ✓ | 30 | N/A | 30 | 27 | Realistic long-form TV show footage for online action detection with dense frame-level temporal annotations across multiple concurrent actions | |||||||
| Tasty-Videos [82] | 2019 | X | ✓ | ✓ | ✓ | N/A | 54 | 37.6 | 2511 | Short-form instructional cooking videos with temporally aligned procedural steps for zero-shot action anticipation and recipe understanding | ||||||
| UnrealEgo [83] | 2022 | E | ✓ | ✓ | ✓ | 24 | N/A | N/A | N/A | Large-scale synthetic dataset generated in Unreal Engine for training and evaluating action recognition models under controlled yet photorealistic environments | ||||||
| UnrealEgo2 [84] | 2024 | E | ✓ | ✓ | N/A | N/A | 14 | N/A | Large-scale synthetic dataset for 3D human pose estimation with diverse full-body motions rendered from virtual eyeglasses-based cameras | |||||||
| VISOR [85] | 2022 | E | ✓ | ✓ | ✓ | ✓ | ✓ | 257 | 720 | 36 | 179 | Pixel-level segmentation benchmark built on EPIC-KITCHENS, focusing on hands and active objects with long-term temporal consistency for video object segmentation, interaction understanding, and long-term reasoning | ||||
| VOST [86] | 2023 | E | ✓ | ✓ | ✓ | 155(Obj), 51(Trans. types) | 21.2 | 4.2 | 713 | Video Object Segmentation benchmark focusing on objects undergoing complex transformations (e.g., cutting, breaking, peeling), with dense instance-level masks over relatively long clips | ||||||
| WEAR [87] | 2024 | E | ✓ | ✓ | ✓ | ✓ | 18 | N/A | 44 | 347 | Outdoor sports dataset combining video and wearable IMU data for long-duration activity recognition in real-world athletic scenarios | |||||
| YouCook2 [88] | 2018 | X | ✓ | ✓ | N/A | 316 | 175.6 | 2K | Large-scale instructional cooking video dataset with temporally segmented procedural steps and natural language descriptions for learning and understanding complex procedures | |||||||
References
- Paullada, A.; Raji, I.D.; Bender, E.M.; Denton, E.; Hanna, A. Data and its (dis) contents: A survey of dataset development and use in machine learning research. Patterns 2021, 2. [CrossRef]
- Salari, A.; Djavadifar, A.; Liu, X.; Najjaran, H. Object recognition datasets and challenges: A review. Neurocomputing 2022, 495, 129–152. [CrossRef]
- Hernandez, J.A.; Colom, M. Reproducible research policies and software/data management in scientific computing journals: a survey, discussion, and perspectives. Frontiers in Computer Science 2025, 6, 1491823. [CrossRef]
- Intelligence, N.M. How to be responsible in AI publication. Nature Machine Intelligence 2021, 3. [CrossRef]
- Rodin, I.; Furnari, A.; Mavroeidis, D.; Farinella, G.M. Predicting the future from first person (egocentric) vision: A survey. Computer Vision and Image Understanding 2021, 211, 103252. [CrossRef]
- Núñez-Marcos, A.; Azkune, G.; Arganda-Carreras, I. Egocentric vision-based action recognition: A survey. Neurocomputing 2022, 472, 175–197. [CrossRef]
- Li, X.; Qiu, H.; Wang, L.; Zhang, H.; Qi, C.; Han, L.; Xiong, H.; Li, H. Challenges and Trends in Egocentric Vision: A Survey. arXiv preprint arXiv:2503.15275 2025.
- Lai, B.; Toyer, S.; Nagarajan, T.; Girdhar, R.; Zha, S.; Rehg, J.M.; Kitani, K.; Grauman, K.; Desai, R.; Liu, M. Human action anticipation: A survey. arXiv preprint arXiv:2410.14045 2024.
- Zhang, X.; He, Y.; Xu, R.; Yu, H.; Shen, Z.; Cui, P. Nico++: Towards better benchmarking for domain generalization. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 16036–16047.
- Wang, Y.; Wu, Y.; Zhang, H. Lost domain generalization is a natural consequence of lack of training domains. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2024, Vol. 38, pp. 15689–15697. [CrossRef]
- Panda, A.; Panigrahi, D.; Mitra, S.; Mittal, S.; Rahimi, S. Transfer Learning Applied to Computer Vision Problems: Survey on Current Progress, Limitations, and Opportunities. arXiv preprint arXiv:2409.07736 2024.
- Zhao, Z.; Alzubaidi, L.; Zhang, J.; Duan, Y.; Gu, Y. A comparison review of transfer learning and self-supervised learning: Definitions, applications, advantages and limitations. Expert Systems with Applications 2024, 242, 122807. [CrossRef]
- Regulation (EU) 2016/679 ... (GDPR), 2016. OJ L 119, 04.05.2016; https://eur-lex.europa.eu/eli/reg/2016/679/oj.
- Grauman, K.; Westbury, A.; Byrne, E.; Chavis, Z.; Furnari, A.; Girdhar, R.; Hamburger, J.; Jiang, H.; Liu, M.; Liu, X.; et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 18995–19012.
- Damen, D.; Doughty, H.; Farinella, G.M.; Furnari, A.; Kazakos, E.; Ma, J.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; et al. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International Journal of Computer Vision 2022, 130, 33–55. [CrossRef]
- Fathi, A.; Li, Y.; Rehg, J.M. Learning to recognize daily actions using gaze. In Proceedings of the European Conference on Computer Vision. Springer, 2012, pp. 314–327.
- Damen, D.; Doughty, H.; Farinella, G.M.; Fidler, S.; Furnari, A.; Kazakos, E.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; et al. Scaling egocentric vision: The EPIC-Kitchens dataset. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 720–736.
- Li, Y.; Liu, M.; Rehg, J.M. In the Eye of the Beholder: Joint learning of gaze and actions in first-person video. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 619–635.
- Fathi, A.; Ren, X.; Rehg, J.M. Learning to recognize objects in egocentric activities. In Proceedings of the CVPR. IEEE, 2011, pp. 3281–3288.
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. bmj 2021, 372. [CrossRef]
- Stein, S.; McKenna, S.J. Combining embedded accelerometers with computer vision for recognizing food preparation activities. In Proceedings of the Proceedings of the 2013 ACM International Joint Conference on Pervasive and Ubiquitous Computing, 2013, pp. 729–738.
- Caba Heilbron, F.; Escorcia, V.; Ghanem, B.; Carlos Niebles, J. ActivityNet: A large-scale video benchmark for human activity understanding. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 961–970.
- Pan, J.; Yuan, Z.; Zhang, X.; Chen, D. YouHome system and dataset: Making your home know you better. In Proceedings of the Proceedings of the IEEE International Symposium on Smart Electronic Systems (iSES). IEEE, 2022, pp. 414–420.
- Lv, Z.; Charron, N.; Moulon, P.; Gamino, A.; Peng, C.; Sweeney, C.; Miller, E.; Tang, H.; Meissner, J.; Dong, J.; et al. Aria Everyday Activities Dataset. arXiv preprint arXiv:2402.13349 2024.
- Sener, F.; Chatterjee, D.; Shelepov, D.; He, K.; Singhania, D.; Wang, R.; Yao, A. Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 21096–21106.
- Ohkawa, T.; He, K.; Sener, F.; Hodan, T.; Tran, L.; Keskin, C. AssemblyHands: Towards egocentric activity understanding via 3D hand pose estimation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 12999–13008.
- Gu, C.; Sun, C.; Ross, D.A.; Vondrick, C.; Pantofaru, C.; Li, Y.; Vijayanarasimhan, S.; Toderici, G.; Ricco, S.; Sukthankar, R.; et al. AVA: A video dataset of spatio-temporally localized atomic visual actions. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 6047–6056.
- Assion, F.; Gressner, F.; Augustine, N.; Klemenc, J.; Hammam, A.; Krattinger, A.; Trittenbach, H.; Philippsen, A.; Riemer, S. A-BDD: Leveraging data augmentations for safe autonomous driving in adverse weather and lighting. arXiv preprint arXiv:2408.06071 2024.
- Yu, F.; Chen, H.; Wang, X.; Xian, W.; Chen, Y.; Liu, F.; Madhavan, V.; Darrell, T. BDD100K: A diverse driving dataset for heterogeneous multitask learning. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 2636–2645.
- Damen, D.; Haines, O.; Leelasawassuk, T.; Calway, A.; Mayol-Cuevas, W. Multi-user egocentric online system for unsupervised assistance on object usage. In Proceedings of the European Conference on Computer Vision. Springer, 2014, pp. 481–492.
- Tadesse, G.A.; Bent, O.; Weldemariam, K.; Istiak, M.A.; Hasan, T.; Cavallaro, A. BON: An extended public domain dataset for human activity recognition. arXiv preprint arXiv:2209.05077 2022.
- Kuehne, H.; Arslan, A.; Serre, T. The language of actions: Recovering the syntax and semantics of goal-directed human activities. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 780–787.
- Koppula, H.S.; Gupta, R.; Saxena, A. Human Activity Learning using Object Affordances from RGB-D Videos. arXiv preprint arXiv:1208.XXXX 2012.
- Sigurdsson, G.A.; Varol, G.; Wang, X.; Farhadi, A.; Laptev, I.; Gupta, A. Hollywood in Homes: Crowdsourcing data collection for activity understanding. In Proceedings of the European Conference on Computer Vision. Springer, 2016, pp. 510–526.
- Sigurdsson, G.A.; Gupta, A.; Schmid, C.; Farhadi, A.; Alahari, K. Actor and observer: Joint modeling of first and third-person videos. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 7396–7404.
- De la Torre Frade, F.; Hodgins, J.K.; Bargteil, A.W.; Martin Artal, X.; Macey, J.C.; Collado I Castells, A.; Beltran, J. Guide to the Carnegie Mellon University Multimodal Activity (CMU-MMAC) Database. Technical Report CMU-RI-TR-08-22, Pittsburgh, PA, 2008.
- Tang, Y.; Ding, D.; Rao, Y.; Zheng, Y.; Zhang, D.; Zhao, L.; Lu, J.; Zhou, J. COIN: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 1207–1216.
- Zhukov, D.; Alayrac, J.B.; Cinbis, R.G.; Fouhey, D.; Laptev, I.; Sivic, J. Cross-task weakly supervised learning from instructional videos. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 3537–3545.
- Donley, J.; Tourbabin, V.; Lee, J.S.; Broyles, M.; Jiang, H.; Shen, J.; Pantic, M.; Ithapu, V.K.; Mehra, R. EasyCom: An augmented reality dataset to support algorithms for easy communication in noisy environments. arXiv preprint arXiv:2107.04174 2021.
- Do, T.; Vuong, K.; Park, H.S. Egocentric Scene Understanding via Multimodal Spatial Rectifier. In Proceedings of the CVPR, 2022.
- Dimiccoli, M.; Bolanos, M.; Talavera, E.; Aghaei, M.; Nikolov, S.G.; Radeva, P. SR-clustering: Semantic regularized clustering for egocentric photo streams segmentation. Computer Vision and Image Understanding 2017, 155, 55–69. [CrossRef]
- Zhao, A.; Tang, C.; Wang, L.; Li, Y.; Dave, M.; Tao, L.; Twigg, C.D.; Wang, R.Y. EgoBody3M: Egocentric body tracking on a VR headset using a diverse dataset. In Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 375–392.
- Hummel, T.; Karthik, S.; Georgescu, M.I.; Akata, Z. EgoCVR: An egocentric benchmark for fine-grained composed video retrieval. In Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 1–17.
- Grauman, K.; Westbury, A.; Torresani, L.; Kitani, K.; Malik, J.; Afouras, T.; Kumar, A.; Baiyya, V.; Bansal, S.; Boote, B.; et al. Ego-Exo4D: Understanding skilled human activity from first- and third-person perspectives. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 19383–19400.
- Li, Y.M.; Huang, W.J.; Wang, A.L.; Zeng, L.A.; Meng, J.K.; Zheng, W.S. EgoExo-Fitness: Towards egocentric and exocentric full-body action understanding. In Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 363–382.
- Huang, Y.; Chen, G.; Xu, J.; Zhang, M.; Yang, L.; Pei, B.; Zhang, H.; Dong, L.; Wang, Y.; Wang, L.; et al. EgoExoLearn: A dataset for bridging asynchronous ego- and exo-centric view of procedural activities in real world. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 22072–22086.
- Deng, K.; Wang, Y.; Chau, L.P. Egocentric Human-Object Interaction Detection: A New Benchmark and Method. arXiv preprint arXiv:2506.14189 2025.
- Zhang, Y.; Cao, C.; Cheng, J.; Lu, H. EgoGesture: A new dataset and benchmark for egocentric hand gesture recognition. IEEE Transactions on Multimedia 2018, 20, 1038–1050. [CrossRef]
- Bhandari, K.; DeLaGarza, M.A.; Zong, Z.; Latapie, H.; Yan, Y. EgoK360: A 360° egocentric kinetic human activity video dataset. In Proceedings of the IEEE International Conference on Image Processing (ICIP). IEEE, 2020, pp. 266–270.
- Qiu, H.; Shi, Z.; Wang, L.; Xiong, H.; Li, X.; Li, H. EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World. arXiv preprint arXiv:2501.19061 2025.
- Zhu, C.; Xiao, F.; Alvarado, A.; Babaei, Y.; Hu, J.; El-Mohri, H.; Culatana, S.; Sumbaly, R.; Yan, Z. EgoObjects: A large-scale egocentric dataset for fine-grained object understanding. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 20110–20120.
- Haneji, Y.; Nishimura, T.; Kameko, H.; Shirai, K.; Yoshida, T.; Kajimura, K.; Yamamoto, K.; Cui, T.; Nishimoto, T.; Mori, S. EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos with Procedural Texts, 2024, [arXiv:cs.CV/2410.05343].
- Bar, A.; Bakhtiar, A.; Tran, D.; Loquercio, A.; Rajasegaran, J.; LeCun, Y.; Globerson, A.; Darrell, T. EgoPet: Egomotion and interaction data from an animal’s perspective. In Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 377–394.
- Bansal, S.; Arora, C.; Jawahar, C. My View is the Best View: Procedure Learning from Egocentric Videos. In Proceedings of the European Conference on Computer Vision (ECCV), 2022.
- Tang, H.; Liang, K.J.; Grauman, K.; Feiszli, M.; Wang, W. EgoTracks: A long-term egocentric visual object tracking dataset. Advances in Neural Information Processing Systems 2023, 36, 75716–75739. [CrossRef]
- Wang, X.; Zhao, K.; Liu, F.; Wang, J.; Zhao, G.; Bao, X.; Zhu, Z.; Zhang, Y.; Wang, X. EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation. arXiv preprint arXiv:2411.08380 2024.
- Tschernezki, V.; Darkhalil, A.; Zhu, Z.; Fouhey, D.; Laina, I.; Larlus, D.; Damen, D.; Vedaldi, A. Epic Fields: Marrying 3D geometry and video understanding. Advances in Neural Information Processing Systems 2023, 36, 26485–26500. [CrossRef]
- Huh, J.; Chalk, J.; Kazakos, E.; Damen, D.; Zisserman, A. EPIC-SOUNDS: A Large-Scale Dataset of Actions that Sound. IEEE Transactions on Pattern Analysis and Machine Intelligence 2025.
- Jang, Y.; Sullivan, B.; Ludwig, C.; Gilchrist, I.; Damen, D.; Mayol-Cuevas, W. EPIC-Tent: An egocentric video dataset for camping tent assembly. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, 2019, pp. 0–0.
- Garcia-Hernando, G.; Yuan, S.; Baek, S.; Kim, T.K. First-person hand action benchmark with RGB-D videos and 3D hand pose annotations. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 409–419.
- Abebe, G.; Catala, A.; Cavallaro, A. A first-person vision dataset of office activities. In Proceedings of the IAPR Workshop on Multimodal Pattern Recognition of Social Signals in Human-Computer Interaction. Springer, 2018, pp. 27–37.
- Li, Y.; Ye, Z.; Rehg, J.M. Delving into egocentric actions. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 287–295.
- Kwon, T.; Tekin, B.; Stühmer, J.; Bogo, F.; Pollefeys, M. H2O: Two hands manipulating objects for first-person interaction recognition. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 10138–10148.
- Perrett, T.; Darkhalil, A.; Sinha, S.; Emara, O.; Pollard, S.; Parida, K.K.; Liu, K.; Gatti, P.; Bansal, S.; Flanagan, K.; et al. HD-EPIC: A highly-detailed egocentric video dataset. In Proceedings of the Proceedings of the CVPR, 2025, pp. 23901–23913.
- Liu, Y.; Liu, Y.; Jiang, C.; Lyu, K.; Wan, W.; Shen, H.; Liang, B.; Fu, Z.; Wang, H.; Yi, L. HOI4D: A 4D egocentric dataset for category-level human-object interaction. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 21013–21022.
- Wang, X.; Kwon, T.; Rad, M.; Pan, B.; Chakraborty, I.; Andrist, S.; Bohus, D.; Feniello, A.; Tekin, B.; Frujeri, F.V.; et al. HoloAssist: An egocentric human interaction dataset for interactive AI assistants in the real world. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 20270–20281.
- Aggarwal, L.; Bahirwani, V.; Li, L.; Colaco, A. Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark. arXiv preprint arXiv:2508.11192 2025.
- Schoonbeek, T.J.; Houben, T.; Onvlee, H.; Van der Sommen, F.; et al. IndustReal: A dataset for procedure step recognition handling execution errors in egocentric videos in an industrial-like setting. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 4365–4374.
- Kotseruba, I.; Rasouli, A.; Tsotsos, J.K. Joint Attention in Autonomous Driving (JAAD). arXiv preprint arXiv:1609.04741 2016.
- Singh, K.K.; Fatahalian, K.; Efros, A.A. KrishnaCam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks. In Proceedings of the Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 2016, pp. 1–9.
- Ragusa, F.; Furnari, A.; Farinella, G.M. MECCANO: A multimodal egocentric dataset for human behavior understanding in the industrial-like domain. Computer Vision and Image Understanding 2023, 235, 103764. [CrossRef]
- Rohrbach, M.; Amin, S.; Andriluka, M.; Schiele, B. A database for fine grained activity detection of cooking activities. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2012, pp. 1194–1201.
- Rohrbach, M.; Rohrbach, A.; Regneri, M.; Amin, S.; Andriluka, M.; Pinkal, M.; Schiele, B. Recognizing Fine-Grained and Composite Activities Using Hand-Centric Features and Script Data. International Journal of Computer Vision 2015, pp. 1–28.
- Yeung, S.; Russakovsky, O.; Jin, N.; Andriluka, M.; Mori, G.; Fei-Fei, L. Every moment counts: Dense detailed labeling of actions in complex videos. International Journal of Computer Vision 2018, 126, 375–389.
- Plizzari, C.; Planamente, M.; Goletto, G.; Cannici, M.; Gusso, E.; Matteucci, M.; Caputo, B. E2(GO) MOTION: Motion Augmented Event Stream for Egocentric Action Recognition. arXiv preprint arXiv:2112.03596 2021.
- Ma, L.; Ye, Y.; Hong, F.; Guzov, V.; Jiang, Y.; Postyeni, R.; Pesqueira, L.; Gamino, A.; Baiyya, V.; Kim, H.J.; et al. Nymeria: A massive collection of multimodal egocentric daily motion in the wild. In Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 445–465.
- Rodin, I.; Furnari, A.; Mavroeidis, D.; Farinella, G.M. Egocentric action anticipation for personal health. In Proceedings of the ICASSP. IEEE, 2023, pp. 1–5.
- Rasouli, A.; Kotseruba, I.; Kunic, T.; Tsotsos, J.K. PIE: A Large-Scale Dataset and Models for Pedestrian Intention Estimation and Trajectory Prediction. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision, 2019.
- Idrees, H.; Zamir, A.R.; Jiang, Y.G.; Gorban, A.; Laptev, I.; Sukthankar, R.; Shah, M. The THUMOS challenge on action recognition for videos in the wild. Computer Vision and Image Understanding 2017, 155, 1–23.
- Patron-Perez, A.; Marszalek, M.; Zisserman, A.P.; Reid, I.D. High five: Recognising human interactions in TV shows 2010.
- De Geest, R.; Gavves, E.; Ghodrati, A.; Li, Z.; Snoek, C.; Tuytelaars, T. Online action detection. In Proceedings of the European Conference on Computer Vision. Springer, 2016, pp. 269–284.
- Sener, F.; Yao, A. Zero-Shot Anticipation for Instructional Activities. In Proceedings of the ICCV, 2019.
- Akada, H.; Wang, J.; Shimada, S.; Takahashi, M.; Theobalt, C.; Golyanik, V. UnrealEgo: A new dataset for robust egocentric 3D human motion capture. In Proceedings of the European Conference on Computer Vision. Springer, 2022, pp. 1–17.
- Akada, H.; Wang, J.; Golyanik, V.; Theobalt, C. 3D human pose perception from egocentric stereo videos. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 767–776.
- Darkhalil, A.; Shan, D.; Zhu, B.; Ma, J.; Kar, A.; Higgins, R.; Fidler, S.; Fouhey, D.; Damen, D. EPIC-Kitchens VISOR benchmark: Video segmentations and object relations. Advances in Neural Information Processing Systems 2022, 35, 13745–13758.
- Tokmakov, P.; Li, J.; Gaidon, A. Breaking the “Object” in Video Object Segmentation. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023.
- Bock, M.; Kuehne, H.; Van Laerhoven, K.; Moeller, M. WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 2024, 8. https://doi.org/10.1145/3699776.
- Zhou, L.; Xu, C.; Corso, J. Towards automatic learning of procedures from web instructional videos. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2018, Vol. 32.
Figure 1.
Perspective taxonomy illustrating the distinction between egocentric and exocentric viewpoints.
Figure 1.
Perspective taxonomy illustrating the distinction between egocentric and exocentric viewpoints.

Figure 2.
Taxonomy of egocentric domains spanning daily living, sports, industrial activities, healthcare, social interaction, and other real-world settings.
Figure 2.
Taxonomy of egocentric domains spanning daily living, sports, industrial activities, healthcare, social interaction, and other real-world settings.

Figure 3.
Multimodal taxonomy of egocentric sensing, organised into visual, motion, audio, language, and 3D/environmental modalities. Solid-bordered elements denote raw sensor signals directly captured during data acquisition, while dashed-bordered elements indicate derived modalities obtained through post-processing or model-based estimation (e.g., optical flow, pose estimation, etc.).
Figure 3.
Multimodal taxonomy of egocentric sensing, organised into visual, motion, audio, language, and 3D/environmental modalities. Solid-bordered elements denote raw sensor signals directly captured during data acquisition, while dashed-bordered elements indicate derived modalities obtained through post-processing or model-based estimation (e.g., optical flow, pose estimation, etc.).

Figure 4.
Structured taxonomy of annotation types in egocentric datasets. The figure groups supervisory signals according to their primary function: frame-level perceptual annotations, temporal action and anticipation labels, interaction-centric annotations capturing hand–object and relational events, spatial and trajectory-based annotations encoding 3D structure and motion, and audio–text annotations providing linguistic and acoustic context. The taxonomy is illustrative rather than exhaustive and highlights the diversity of annotation practices across egocentric benchmarks.
Figure 4.
Structured taxonomy of annotation types in egocentric datasets. The figure groups supervisory signals according to their primary function: frame-level perceptual annotations, temporal action and anticipation labels, interaction-centric annotations capturing hand–object and relational events, spatial and trajectory-based annotations encoding 3D structure and motion, and audio–text annotations providing linguistic and acoustic context. The taxonomy is illustrative rather than exhaustive and highlights the diversity of annotation practices across egocentric benchmarks.

Figure 5.
Taxonomy of tasks in egocentric video understanding. The figure organises tasks according to their semantic scope and temporal orientation, ranging from low-level perceptual tasks to action and activity understanding, and further to prospective tasks such as anticipation and forecasting. Additional task families related to gaze and attention, hand–object interaction and manipulation, three-dimensional mapping and tracking, vision–language reasoning, and summarisation are shown to illustrate the breadth of task formulations supported by egocentric datasets. The taxonomy highlights the diversity of task support across benchmarks and clarifies which dataset designs are appropriate for different research objectives.
Figure 5.
Taxonomy of tasks in egocentric video understanding. The figure organises tasks according to their semantic scope and temporal orientation, ranging from low-level perceptual tasks to action and activity understanding, and further to prospective tasks such as anticipation and forecasting. Additional task families related to gaze and attention, hand–object interaction and manipulation, three-dimensional mapping and tracking, vision–language reasoning, and summarisation are shown to illustrate the breadth of task formulations supported by egocentric datasets. The taxonomy highlights the diversity of task support across benchmarks and clarifies which dataset designs are appropriate for different research objectives.

Figure 8.
Distribution of datasets by perspective across the curated corpus. The figure provides a static snapshot of the proportion of egocentric, exocentric, and multi-view datasets, independent of their year of release.
Figure 8.
Distribution of datasets by perspective across the curated corpus. The figure provides a static snapshot of the proportion of egocentric, exocentric, and multi-view datasets, independent of their year of release.

Figure 9.
Distribution of sensing modalities across the curated dataset corpus. The figure reports the number of datasets that include each modality. Counts are non-exclusive, as individual datasets may provide multiple modalities (e.g., RGB video combined with audio, gaze, or inertial measurements).
Figure 9.
Distribution of sensing modalities across the curated dataset corpus. The figure reports the number of datasets that include each modality. Counts are non-exclusive, as individual datasets may provide multiple modalities (e.g., RGB video combined with audio, gaze, or inertial measurements).

Figure 10.
Joint distribution of dataset perspective and application domain. Stacked bars show the number of datasets per domain, grouped by perspective, highlighting domain-specific preferences for egocentric, exocentric, and multi-view data collection.
Figure 10.
Joint distribution of dataset perspective and application domain. Stacked bars show the number of datasets per domain, grouped by perspective, highlighting domain-specific preferences for egocentric, exocentric, and multi-view data collection.

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.