Preprint
Review

This version is not peer-reviewed.

EgoDatasets: A Dataset-Centric Survey and Taxonomy for Egocentric Vision and Action Anticipation

Submitted:

07 August 2026

Posted:

12 August 2026

You are already at the latest version

Abstract
Egocentric vision has become a central paradigm for studying human behaviour, interaction, and intent from a first-person perspective, supporting a growing range of tasks such as action recognition, temporal segmentation, and action anticipation. Progress in this area has been driven largely by the availability of public datasets, which differ substantially in sensing configurations, annotation strategies, application domains, and temporal structure. Despite the rapid expansion of available resources, the egocentric dataset landscape remains fragmented, with limited cross-dataset interoperability and uneven coverage across domains, modalities, and task formulations. This survey presents a comprehensive, dataset-centric analysis of the current egocentric and action-related video ecosystem. We curate and systematically analyse 75 publicly documented datasets spanning more than a decade of research, covering both egocentric and selected exocentric benchmarks that are widely used for action-centric and anticipatory modelling. To organise this landscape, we introduce a unified taxonomy structured around five complementary design dimensions: perspective, domain, sensing modality, annotation structure, and task support. This taxonomy enables consistent comparison across heterogeneous datasets and provides a principled framework for analysing how design choices influence supported tasks and evaluation practices. Beyond cataloguing datasets, we examine temporal trends, distributional patterns, and co-occurrence statistics across the curated corpus, revealing recurring biases toward specific domains, limited multimodal coverage, and a widespread reliance on post hoc constructions for action anticipation. We further discuss annotation heterogeneity, scale–diversity trade-offs, and the ethical and legal constraints that shape egocentric data collection and release. By consolidating dispersed resources into a coherent taxonomy and identifying systematic gaps in current benchmarks, this survey aims to support informed dataset selection, facilitate cross-dataset analysis, and guide future efforts in egocentric dataset design and benchmarking.
Keywords: 
;  ;  ;  ;  

1. Introduction

Public datasets have long been recognised as the driving force behind the solid foundations of modern computer vision infrastructure. Without these open datasets, research and innovation would lose momentum. More specifically, in the field of egocentric vision, rapid advancements in both hardware and software have significantly expanded the range and quality of available data sources. The increasing use of RGB cameras, depth sensors, gaze trackers, inertial units, and audio devices has enabled the creation of large-scale multimodal datasets with considerably richer representations than before. Therefore, the egocentric vision community has developed several benchmark datasets that differ in scale, sensing configurations, annotation detail, and target tasks. These datasets capture daily activities, social interactions, hand-object manipulations, and both short and long-term procedural behaviours, offering diverse resources for first-person perception research.
Egocentric vision, by definition, focuses on capturing the world from the perspective of an acting individual, allowing detailed analysis of hand-object interactions, gaze behaviour, goal-directed manipulation, and extended activity sequences. This unique viewpoint has supported a rapidly expanding range of applications, including cognitive and assistive robotics, immersive AR systems, activity recognition, action anticipation, and more broadly, human-computer interaction. However, despite these developments, progress in egocentric vision remains fundamentally tied to the availability, diversity, and openness of the underlying datasets.
Prior work in computer vision has repeatedly demonstrated that public, well-documented datasets are not merely convenient resources but core scientific infrastructure. As highlighted in [1], datasets identify key research areas, including the topics the community investigates, the way algorithms are benchmarked, and the evaluation standards that become widely adopted. Similarly, [2] underscores that the historical evolution of computer vision has been inseparable from the development of increasingly larger and more diverse benchmarks. According to [3,4], the pivotal importance of datasets is essential for ensuring the reproducibility of experiments to unveil significant insights and patterns obtained directly from the data, underscoring the significance of reproducibility in scientific research. These assertions are foundational to the integrity, consistency, and long-term sustainable advancement of computer vision research.
In the domain of egocentric vision, there is a notable emphasis on data dependency. Studies like [5] and [6] argue that progress in first-person action recognition, action anticipation, and behavior analysis has been predominantly facilitated by the introduction of large-scale datasets. Recent works, such as [7], state that the field is still struggling with a persistent inadequacy of specialized, high-quality egocentric datasets, particularly those covering various settings, multi-sensor data, and extended temporal labels. Regarding action anticipation and human forecasting tasks, [8] similarly note that the available benchmarks exhibit inconsistencies across domains and lack consistency in their annotation protocols and prediction horizons. Despite the progress enabled by recent benchmarks, the current egocentric dataset landscape remains fragmented and uneven. Existing datasets differ substantially in sensing configurations, annotation schemas, action taxonomies, and domain coverage, leading to limited interoperability and significant barriers to cross-dataset comparison. Most of the widely used benchmarks tend to focus on a restricted set of environments, most commonly kitchens, daily living spaces, driving scenes, or constrained industrial workflows, while offering sparse temporal labeling and limited variation in sensing setups. Another constraint arising from this restricted domain coverage is its impact on domain generalization and transfer learning. Specifically, prior studies in the field of computer vision indicate that models trained on a narrow set of visual categories often fail to generalize well to unfamiliar conditions. According to the authors in [9], the constrained scope of domains in numerous benchmarks results in assessments that lack the capacity to adequately represent the variability present in real-world scenarios. This limitation is further illustrated in [10] where it was observed that the occurrence of inadequate generalization frequently originates from a restricted range of domains represented in training datasets. Complementary findings from transfer learning surveys [11,12] emphasise that mismatches between pre-training datasets and downstream tasks remain a key bottleneck, especially when pre-training corpora do not capture the types of environments or behaviours encountered in egocentric vision. These observations reinforce the idea that insufficient domain and annotation variability is a fundamental constraint that directly affects the reliability and robustness of trained models, rather than just a question of dataset design.
Moreover, existing surveys typically organise datasets around specific tasks such as egocentric action recognition, video summarisation, future prediction, or action anticipation, rather than providing a unified, cross-domain perspective on the full dataset ecosystem. As a result, the community still lacks a comprehensive and up-to-date resource that aims to catalogue all the egocentric and action-related datasets, analyze their characteristics, and organize them within a coherent taxonomy.
Furthermore, the creation and dissemination of egocentric datasets are subject to ethical and legal restrictions that are less restrictive in many other areas of computer vision. Recordings captured from a first-person perspective typically encompass intimate environments, including bystanders, as well as biometric markers such as facial features, vocal characteristics, unique physical gestures, and patterns of eye contact. Within legal frameworks like the EU General Data Protection Regulation (GDPR), biometric data used for uniquely identifying a person is treated as a special category of personal data and are therefore subject to stricter conditions on collection, processing, and sharing [13]. For instance, recent large-scale egocentric benchmarks such as [14] and [15] explicitly document their reliance on institutional ethics review, informed consent procedures, and, where necessary, de-identification of personally identifiable information before public release. While these measures are essential for protecting participants, they also increase the cost and complexity of dataset collection and can limit how openly data may be distributed or extended.
Given these scientific, practical, and regulatory constraints, a more structured understanding of the egocentric dataset landscape is still lacking. Existing surveys provide valuable summaries of specific tasks and sub-domains, however, they typically discuss datasets only in the context of individual problems or benchmark settings. Notable examples include surveys on egocentric future prediction [5], action recognition [6], broader egocentric trends [7], and human action anticipation [8]. While this task-centred perspective is useful for comparing methods within a particular problem, it gives only a partial account of how datasets differ across domains, sensing modalities, annotation practices, and temporal structure. As the number and diversity of available egocentric datasets continue to grow, a unified taxonomy becomes essential for organising this landscape, revealing systematic gaps, and supporting more informed decisions about dataset selection, learning approaches, and evaluation protocols.
Building on these observations, this study provides a comprehensive, dataset-centric analysis of egocentric and action-oriented video resources. We curate an extensive collection of 75 publicly accessible datasets spanning more than a decade of research and covering diverse environments, activities, and sensing configurations. A central aim of this work is to introduce a unified taxonomy that organises datasets according to perspective, domain, modality, and annotation design, offering a structured framework for understanding how existing resources relate to one another. Beyond cataloguing datasets, we analyse common patterns and discrepancies across current resources, identify systematic gaps in domain coverage and multimodal support, and summarise the ethical and regulatory factors that influence dataset collection and public dissemination. Collectively, these contributions establish a consolidated foundation for researchers working with egocentric video data and for future efforts on dataset design and benchmarking.
The remainder of this paper is structured as follows. Section 2 defines the background and scope of the survey, introducing egocentric vision and action-centric video understanding, and formalising the tasks, modalities, and dataset characteristics considered. Section 3 presents a unified taxonomy for egocentric and action-related datasets, organising existing resources along five core design dimensions and motivating the criteria used for systematic categorisation. Section 4 provides a comprehensive overview of the curated dataset corpus, analysing temporal trends, perspective, domain and modality distributions, and highlighting structural patterns in current dataset design. Section 5 discusses the broader implications of the survey findings, examining persistent limitations in the dataset landscape, ethical and accessibility considerations, and outlining directions for future dataset development and benchmarking.

2. Background and Scope of the Survey

In this section, we outline the conceptual foundations and scope of this work. We begin by defining egocentric vision and positioning it within the broader landscape of action-centric video understanding. Particular attention is given to the sensing configurations, viewpoint constraints, and behavioural cues that differentiate first-person footage from conventional third-person video. We then summarise the modalities, annotation types, and temporal structures encountered across the datasets considered in this survey, highlighting the methodological diversity that motivates a principled taxonomy. Finally, we formalise the inclusion criteria and scope of our analysis to ensure that the subsequent classifications and comparisons are grounded in transparent and replicable assumptions.

2.1. Egocentric Vision: Definition and Characteristics

Egocentric vision, also known as first-person vision, refers to visual information captured by a camera mounted on the head or body of an individual, providing a perspective that matches the camera-wearer’s line of sight. Unlike third-person or exocentric video, which observes actions from an external point of view, egocentric footage reflects the perceptual and motor experience of the individual, capturing what they see and how they interact with objects and their surroundings. This perspective introduces distinctive visual and temporal characteristics, including motion induced by head and body movement, frequent visibility of the individual’s hands and manipulated objects, and a fragmented, dynamically changing view of the environment driven by the camera-wearer’s gaze direction.
The literature reviewed in this work consistently highlights these features. Núñez-Marcos et al. [6], for example, describe egocentric video as exhibiting substantial non-linear and unpredictable motion arising from the wearer’s movements, alongside limited access to broader scene context. They identify ego-motion, hand-object interaction, and a constantly shifting field of view tied to the individual’s focus as key challenges for visual recognition systems. Early foundational work in egocentric action understanding [16] similarly notes that wearable-camera footage is dominated by frequent ego-motion and that alignment with the wearer’s attention exposes behaviourally rich cues such as hand pose, head movement, and gaze signals that are largely absent from conventional third-person video datasets.
In addition, egocentric vision imposes unique assumptions on the structure of activities. Because the camera wearer is both the observer and the actor, egocentric video encodes the intentional structure of interaction: tasks are organised around hand-object contact, object state changes, and gaze-driven attention patterns. These interaction-centric signals have become central to modern egocentric benchmarks such as EPIC-KITCHENS [15,17], EGTEA Gaze+[16,18,19] and Ego4D [14], which explicitly annotate hands, objects, object states, and gaze. Their design reflects the need to capture the full sensorimotor context underlying everyday activities, enabling downstream tasks such as action recognition, hand-object interaction modelling, and action anticipation. Overall, the defining characteristics of egocentric video wearer-driven motion, embodied viewpoints, interaction-centric scene structure, and attention-linked visual framing form the conceptual basis of the dataset taxonomy developed in this survey. These properties shape both the modalities typically captured (e.g., RGB, depth, IMU, gaze, audio) and the annotation paradigms required (e.g., hand masks, object bounding boxes, action segments, narrations). Understanding these characteristics is essential for interpreting the methodological diversity of the datasets analysed in this work and provides the foundation for the analyses conducted in subsequent sections.

2.2. Egocentric Video Understanding and Action-Centric Tasks

Egocentric video understanding focuses on analysing human behaviour from the perspective of an acting individual, where actions are observed as part of an ongoing stream of goal-directed interaction rather than as isolated motion patterns. Unlike third-person video analysis, which often emphasises full-body kinematics and scene-level context, egocentric understanding centres on object manipulation, hand activity, visual attention, and task progression. As a result, the definition and formulation of action-centric tasks in egocentric vision differ both conceptually and practically from their exocentric counterparts.
Early work in egocentric vision primarily addressed action recognition, where the objective is to identify the activity being performed based on visual evidence accumulated over time. In first-person video, however, actions are frequently short, overlapping, and visually ambiguous, with their semantic meaning determined by object interactions rather than body motion alone. This has led to task formulations that explicitly incorporate hand–object interactions, verb–noun action decompositions, and object state changes, as seen in influential benchmarks such as EPIC-KITCHENS and EGTEA Gaze+. These formulations reflect the observation that, in egocentric settings, actions are more naturally defined through what is being manipulated and how, rather than through global pose or movement. Beyond recognition, egocentric video has motivated tasks that operate over longer temporal extents, including temporal segmentation, activity parsing, and procedural understanding. Such tasks aim to decompose extended activity streams into meaningful action units or phases, capturing the structure of multi-step behaviours such as cooking, assembly, or daily routines. Prior surveys note that these tasks are particularly challenging in first-person video due to weak visual boundaries between actions, frequent interleaving of activities, and variations in execution order [6,7]. Consequently, egocentric datasets often support task formulations that emphasise temporal continuity and contextual reasoning rather than discrete clip-level classification.
A defining characteristic of egocentric video understanding is the growing emphasis on future-oriented tasks. Action anticipation, future action forecasting, and goal inference aim to predict what the camera wearer will do next based on partial observations. These tasks rely on cues such as object affordances, hand trajectories, gaze patterns, and task context, which frequently precede observable action execution. Surveys on action anticipation highlight that first-person data is particularly well suited for such predictive tasks, as it captures early indicators of intent that are unavailable or attenuated in third-person views [5,8]. At the same time, these tasks introduce intrinsic uncertainty, since multiple future actions may be consistent with the same observed state. In addition to action-centric tasks, egocentric datasets increasingly support interaction-focused and multimodal problem formulations. These include hand–object interaction modelling, object state transition detection, gaze and attention prediction, and joint vision–language reasoning based on narrations or spoken instructions. Such tasks reflect the embodied and sensor-rich nature of egocentric data and blur the boundary between perception and higher-level reasoning. As noted in recent surveys, the diversity of supported tasks across egocentric benchmarks is both a strength and a source of fragmentation, as datasets are often designed with specific task assumptions that limit direct comparability [6,7].
In general, egocentric video understanding encompasses a broad spectrum of action-centric tasks that differ in temporal scope, semantic abstraction, and predictive intent. These task formulations are tightly coupled to dataset design choices, including annotation strategies and temporal organisation, and they motivate the need for a structured framework to analyse how datasets support different forms of inference. This perspective provides the conceptual foundation for the taxonomy introduced in Section 3, where task support is formalised as a distinct dimension of dataset design.

2.3. Modalities and Annotation Paradigms in Egocentric Datasets

Egocentric datasets differ from conventional video benchmarks not only in viewpoint, but also in the breadth and role of sensory information they capture. Because first-person video reflects the perceptual experience of an acting individual, egocentric data collection often extends beyond RGB video to include complementary sensory signals that convey motion, attention, and interaction. As a result, both the choice of sensing modalities and the design of annotation paradigms become central elements of dataset construction rather than secondary implementation details.
RGB video constitutes the core modality in nearly all egocentric datasets, providing access to scene appearance, manipulated objects, and hand activity. However, first-person RGB streams are frequently affected by rapid ego-motion, occlusions, and viewpoint changes, which can obscure fine-grained interaction cues. To mitigate these limitations, many datasets incorporate additional modalities such as depth, inertial measurements, audio, or gaze signals. Depth and three-dimensional information enable spatial reasoning about object geometry and hand–object alignment, while inertial sensors provide direct measurements of head or body motion that are difficult to infer reliably from video alone. Audio recordings capture environmental sounds, object interactions, and speech, offering complementary cues for action boundaries and context, particularly in kitchen, industrial, and social settings [7,14].
Gaze and eye-tracking data represent a distinctive modality in egocentric vision, as they provide direct access to the camera wearer’s attentional focus. Prior work has shown that gaze often precedes physical interaction and serves as an early indicator of intent, making it especially valuable for predictive and anticipatory tasks [16,18]. Datasets that include gaze supervision therefore support forms of reasoning that are difficult to realise using visual appearance alone, but they also introduce additional complexity in data collection and synchronisation.
Beyond sensing, annotation paradigms in egocentric datasets exhibit substantial diversity in both form and granularity. Unlike many third-person benchmarks that rely primarily on clip-level labels, egocentric datasets frequently combine multiple annotation types to capture different aspects of behaviour. These may include frame-level labels such as hand presence or object visibility, temporal annotations defining action segments, and higher-level semantic labels describing tasks or procedures. Several datasets further annotate object states, interaction events, or narrations aligned with the video stream, reflecting the fact that actions in egocentric settings are often defined by changes in object configuration rather than by motion patterns alone [15,17]. The coexistence of multiple annotation paradigms within a single dataset reflects the inherently multi-layered nature of egocentric behaviour, but it also creates challenges for dataset interoperability and method comparison. Differences in label definitions, temporal alignment, and semantic abstraction complicate the transfer of models across benchmarks and make it difficult to compare results obtained under different annotation assumptions. Surveys of egocentric vision consistently identify this heterogeneity as a major source of fragmentation in the field, particularly as datasets expand to support increasingly diverse tasks and modalities [6,7].
The combination of rich multimodal sensing and heterogeneous annotation paradigms distinguishes egocentric datasets from traditional video benchmarks. These characteristics enable more expressive modelling of interaction, intent, and context, but they also introduce design trade-offs that directly influence which learning problems can be addressed. Understanding how modalities and annotations are selected and combined at the dataset level is therefore essential for interpreting experimental results and motivates the structured treatment of these aspects in the taxonomy introduced in the next section.

2.4. Temporal Structures and Dynamics in Egocentric Data

A defining property of egocentric video is the temporal organisation of human behaviour. Unlike third-person datasets, where actions are often short, visually distinct, and body-motion centric, egocentric datasets capture extended, multi-step activities driven by object manipulation, task goals, and interaction sequences. As a result, temporal structure becomes a central modelling axis, shaping both the annotation design and the types of tasks that datasets support. More specifically, temporal structure in egocentric datasets is not only a characteristic of the recorded behaviour, but also a deliberate design choice that shapes how data can be used. Decisions such as whether actions are annotated at the frame or segment level, whether temporal boundaries are sharply defined or loosely specified, and whether future events are explicitly labelled determine which forms of inference a dataset can support. For example, datasets with densely annotated temporal boundaries facilitate fine-grained segmentation and early action detection, whereas datasets that encode extended procedural structure enable reasoning over long-term dependencies and task progression. As a result, temporal organisation acts as an implicit constraint on task formulation and evaluation, even when datasets capture similar activities.
Many egocentric datasets provide frame-level labels, such as hand presence, object visibility, or action primitives. These annotations support fine-grained temporal reasoning but often depend on dense supervisory effort. More commonly, datasets adopt segment-level annotations, in which actions are defined over temporally contiguous intervals. These segments may vary widely in duration from sub-second atomic manipulations to multi-minute procedural steps, reflecting the inherently hierarchical nature of first-person behaviour. A second temporal characteristic is the presence of overlapping or interleaved actions. Egocentric activities frequently involve parallel processes, such as preparing an object while reaching for another, making the temporal boundaries between actions less distinct than in curated third-person benchmarks. This motivates annotation schemes that capture verb–noun pairs, object state changes, and interactions, as seen in modern datasets such as EPIC-KITCHENS and Ego4D. Egocentric datasets also exhibit extended temporal dependencies. Daily activities, assembly tasks, and long procedures unfold over hundreds or thousands of frames, with meaningful relationships between early and later stages of the sequence. These long-range dependencies underpin tasks such as action anticipation, next-object prediction, and future forecasting, which rely on modelling behavioural intent and procedural progression rather than isolated motion cues.
Finally, temporal dynamics in egocentric video are strongly influenced by ego-motion, gaze shifts, and task-driven visual framing. Rapid head movements may produce momentary instability, while gaze-aligned framing emphasises objects and regions of interest at different points in time. Together, these factors shape the temporal structure of egocentric datasets and motivate the need for specialised temporal models, annotation conventions, and evaluation protocols. Understanding these dynamics is crucial for interpreting the diverse dataset designs surveyed in this work and provides the conceptual foundation for the unified taxonomy that is later introduced.

2.5. Objectives and Scope

This survey provides a systematic and comprehensive overview of datasets relevant to egocentric vision and action anticipation. To ensure methodological transparency and reproducibility, we follow a structured, PRISMA-inspired [20] selection procedure that emphasises systematic identification, screening, and eligibility checking. Our final corpus comprises 75 distinct datasets, each documented with a complete metadata record capturing modality, task, perspective, domain, and annotation type. The initial dataset pool was formed through an extensive Google Scholar search, followed by a detailed examination of the resulting literature, ultimately identifying 77 action-anticipation papers and 12 egocentric-vision studies.
To compile the dataset corpus, we enumerated all datasets referenced across these works and collected their corresponding publications and documentation. For each dataset, we constructed a unified metadata entry describing its sensing modalities, annotation types, supported tasks, domain category, and whether it adopts an egocentric or exocentric perspective. During consolidation, multiple references to the same dataset were merged, naming inconsistencies were standardised, and ambiguous cases were resolved by consulting the primary sources. All metadata entries were manually verified for internal coherence and completeness. A dataset was included in the final list only if its metadata profile could be reliably reconstructed from publicly available documentation. This requirement ensures that all datasets incorporated into our taxonomy can be analysed uniformly across key dimensions such as domain, modality, annotation structure, and anticipation relevance, an essential prerequisite for reproducible cross-dataset comparison and principled categorisation in Section 3.
Furthermore, to define the final dataset corpus, we applied a set of principled inclusion criteria to ensure analytical consistency. Each dataset was required to contain human-centered video depicting actions, interactions, or procedures, and to provide meaningful supervisory signals such as action labels, temporal segments, object annotations, or interaction markers. Datasets were included only when accompanied by sufficient publicly available documentation, such as an associated publication, technical report, or official project description, allowing reliable metadata reconstruction and traceable source verification. In addition, a dataset was retained only if it was relevant to egocentric vision or action anticipation, either through an explicit first-person viewpoint or through its established use as a benchmark for forecasting or action-centric tasks. Finally, to support the structured comparisons developed in subsequent sections, we retained only datasets whose domain, modality, annotation structure, and task definitions could be consistently identified and aligned with the taxonomy introduced in the following sections. This alignment guarantees coherence across the survey and enables reproducible, dataset-level analyses.

4. Dataset Overview

4.1. Dataset Corpus Summary

This survey curates a corpus of 75 publicly documented datasets depicted in Table A1 that are relevant to egocentric vision and action-centric video understanding, covering a period of more than a decade. The collection reflects the progressive evolution of first-person data resources, from early small-scale benchmarks focusing on isolated activities to recent large-scale datasets designed for long-horizon, multimodal, and anticipatory reasoning. Egocentric datasets prevail in the corpus, showcasing the increased focus on first-person perception, interaction-centric modelling, and intent-aware reasoning in current studies on computer vision research. A smaller but non-negligible subset of exocentric datasets is retained due to their historical importance and continued use as benchmarks for action recognition, temporal segmentation, and forecasting. In addition, several datasets adopt mixed or multi-perspective configurations, enabling comparative analysis across viewpoints and supporting emerging research on cross-perspective transfer and joint ego-exo learning.
From a sensing perspective, the corpus is characterised by a strong reliance on RGB video, which remains the foundational modality across nearly all datasets. However, a substantial fraction of recent datasets augment visual streams with additional signals such as audio, inertial measurements, gaze tracking, depth, or hand pose, reflecting a clear trend toward richer multimodal representations. While unimodal datasets remain prevalent, particularly among earlier benchmarks, the increasing availability of multimodal data has expanded the range of tasks that can be meaningfully supported, including anticipation, interaction modelling, and procedural reasoning. Specifically, the dataset corpus covers a wide range of domains, particularly focusing on daily activities and kitchen environments in various applications. This highlights the significance of these environments for applications focused on human users and the benefits they provide for structured data gathering and labelling. Concurrently, the corpus includes datasets drawn from sports, industrial and assembly tasks, healthcare, education, social interaction, and outdoor activities, enabling analysis of how domain characteristics influence sensing choices, annotation strategies, and supported tasks.
In general, the curated dataset corpus provides broad coverage across perspective, modality, domain, annotation structure, and task support, forming a representative snapshot of the current egocentric and action-related dataset landscape. This diversity establishes the empirical foundation for the analyses that follow. In the remainder of this section, we examine how datasets are distributed across time, perspective, domain, and modality, and how scale and temporal structure vary across benchmarks.

4.2. Temporal Evolution of Datasets (2008–2025)

A look at dataset release timelines shows a clear change both in volume and in viewpoint over time. Early benchmarks are relatively few and are mostly built around exocentric views, which aligns with the earlier emphasis on third-person action recognition and fixed camera setups. From the mid-2010s onward, dataset releases begin to increase more rapidly, alongside the wider availability of wearable sensors and a growing interest in first-person data. In the most recent years, egocentric datasets make up most new releases, pointing to a longer-term move toward studying interaction, intention, and behaviour from the actor’s perspective. The continued rise in the total number of available datasets suggests that the field is moving into a more stable phase. New datasets are often released alongside existing benchmarks instead of supplanting them, typically adding coverage for additional viewpoints, domains, or sensing setups. As a result, recent work tends to expand what is already available rather than redefining the benchmark space.
Figure 6. Temporal evolution of egocentric and exocentric datasets between 2008 and 2025. Stacked bars indicate the number of datasets introduced per year, grouped by perspective (egocentric, exocentric, and multi-view).Years with no new dataset releases are explicitly included to preserve the continuity of the timeline.
Figure 6. Temporal evolution of egocentric and exocentric datasets between 2008 and 2025. Stacked bars indicate the number of datasets introduced per year, grouped by perspective (egocentric, exocentric, and multi-view).Years with no new dataset releases are explicitly included to preserve the continuity of the timeline.
Preprints 227403 g006
Figure 7. Cumulative perspective trends over time. The curves depict the cumulative number of datasets introduced between 2008 and 2025 for each perspective category (egocentric, exocentric, and multi-view). The trajectories highlight a gradual early dominance of exocentric datasets, followed by a pronounced acceleration in egocentric dataset growth after the mid-2010s, indicating a sustained shift toward first-person data collection in action-centric and anticipatory research.
Figure 7. Cumulative perspective trends over time. The curves depict the cumulative number of datasets introduced between 2008 and 2025 for each perspective category (egocentric, exocentric, and multi-view). The trajectories highlight a gradual early dominance of exocentric datasets, followed by a pronounced acceleration in egocentric dataset growth after the mid-2010s, indicating a sustained shift toward first-person data collection in action-centric and anticipatory research.
Preprints 227403 g007

4.3. Distribution by Perspective, Domain, and Modality

Dataset design in egocentric and action-centric video benchmarks can be described along three core dimensions: viewpoint perspective, sensing modality, and application domain. Considering these dimensions jointly reveals recurring patterns in how datasets are constructed, as well as uneven coverage of behaviours and interaction contexts. Figure 8, Figure 9 and Figure 10 summarise these distributions, presenting both individual dimension statistics and their combined structure.
Viewed together, the distributions indicate that egocentric datasets constitute the majority of available benchmarks, while exocentric and multi-view datasets remain concentrated in domains where stable scene observation is required. With respect to sensing configuration, RGB video is almost universally present, whereas datasets offering additional sensory signals are comparatively limited. The joint distribution of perspective and domain further suggests that certain environments are consistently associated with specific viewpoints, pointing to recurring design choices rather than uniform sampling of real-world activity settings.

4.4. Summary and Transition

This section examined the curated dataset corpus at a broad level, focusing on how egocentric and action-related datasets vary over time, viewpoint, application domain, and sensing modality. Looking across release years, the data show a steady rise in new datasets over the last decade, alongside a noticeable move toward egocentric capture in more recent work. This shift mirrors increased interest in first-person perception and interaction-driven modelling, and is closely tied to the growing practicality of wearable sensors for large-scale data collection.
Several recurring design patterns emerge from the distributional analysis. Egocentric datasets now dominate the benchmark landscape, while exocentric and multi-view datasets remain in use for settings where a stable scene layout or global motion information is important. Most datasets rely on RGB video, with only a smaller fraction incorporating additional signals such as audio, gaze, depth, or inertial measurements. The limited use of these modalities is largely driven by the practical effort required to capture, synchronise, and annotate them.
From a domain perspective, datasets are concentrated around everyday activities, with kitchen-based scenarios appearing far more often than other settings. Other environments, including sports, industrial work, healthcare, social interaction, and outdoor scenarios, receive comparatively less attention. These patterns are shaped by practical constraints, perceived application relevance, and established research trajectories, and they influence which behaviours and interactions are most frequently studied. Examining perspective and domain jointly further reveals that certain environments are repeatedly paired with specific viewpoints, suggesting common conventions in dataset construction.
Quantitative analysis was intentionally restricted to dataset attributes that could be reconstructed in a consistent manner across sources. Measures of dataset scale were handled with particular caution, as statistics such as video count, total duration, and annotation density are reported inconsistently across benchmarks. Focusing on directly comparable attributes helps avoid conclusions based on aggregate statistics that are defined differently across datasets. As a whole, the analysis outlines how current datasets are typically designed and where coverage remains limited. These observations set the stage for a closer discussion of annotation practices, accessibility, and structural limitations in existing benchmarks. The following sections build on this overview to examine where current datasets remain insufficient and how future dataset efforts might address these shortcomings.

5. Conclusion and Future Directions

Despite the substantial progress enabled by recent egocentric datasets, the analysis conducted in this survey reveals a number of persistent limitations and structural gaps in the current dataset landscape. These limitations are not isolated artefacts of individual benchmarks, but rather recurring patterns that emerge across perspective, domain, modality, annotation structure, and task support. Identifying these gaps is essential for interpreting reported results, understanding the scope of current benchmarks, and guiding future dataset design.
A first and most evident limitation concerns domain concentration. As shown in the dataset distribution analysis, a large fraction of existing egocentric datasets are situated in a small number of environments, most notably kitchen and daily living settings. While these domains are well suited for studying object-centric manipulation and procedural activities, their dominance results in a skewed representation of real-world behaviour. Domains involving outdoor environments, social interaction, safety-critical scenarios, or high-variability conditions remain comparatively underrepresented. This imbalance has direct implications for generalisation: models trained and evaluated primarily on routine, highly structured activities may fail to transfer to more diverse or less constrained settings. Consequently, performance gains reported on dominant benchmarks may overestimate real-world robustness, particularly for tasks such as action anticipation that rely on contextual and environmental cues.
A second gap relates to the limited native support for anticipation-oriented evaluation. Although many datasets are frequently used for action anticipation and future prediction, only a small subset has been explicitly designed with anticipatory supervision in mind. In most cases, anticipation benchmarks are constructed post hoc by truncating recognition-oriented datasets and reinterpreting existing action annotations. This practice introduces ambiguity regarding observation windows, prediction horizons, and the precise definition of action onset. As a result, evaluation protocols vary widely across studies, making it difficult to compare results or assess progress consistently. The lack of datasets that natively encode anticipation-specific temporal structure remains a fundamental limitation for principled evaluation of predictive models.
A further limitation concerns modality imbalance and sparse multimodal coverage. While RGB video is nearly ubiquitous across egocentric datasets, additional modalities such as audio, gaze, inertial measurements, and depth are available only in a minority of benchmarks. Even when present, these modalities are often confined to specific domains or collected under limited conditions. This uneven distribution restricts the development and evaluation of genuinely multimodal models, as methods trained on one dataset may not be transferable to others with different sensing configurations. Moreover, the scarcity of datasets combining multiple complementary modalities constrains systematic study of sensor fusion strategies and limits understanding of how non-visual cues contribute to anticipation and interaction modelling.
Annotation heterogeneity represents another persistent challenge. Egocentric datasets differ substantially in how actions, interactions, and temporal boundaries are defined. Variations in annotation granularity, semantic abstraction, and label vocabularies complicate cross-dataset comparison and hinder reuse of trained models. Even when datasets target similar activities, differences in verb–noun formulations, temporal segmentation criteria, and interaction definitions often prevent direct alignment. This lack of standardisation does not merely reflect stylistic differences, but has concrete methodological consequences: models trained under one annotation regime may perform poorly or unpredictably when transferred to another, and reported improvements may be specific to a particular labeling convention rather than indicative of general progress.
Finally, the dataset landscape exhibits a recurring scale–diversity trade-off. Large-scale datasets typically focus on a narrow set of domains and tasks, enabling extensive training but limited environmental diversity. Conversely, datasets that span multiple domains or include richer multimodal signals are often smaller in scale, restricting their suitability for data-intensive learning approaches. This trade-off constrains the development of models that are both statistically robust and broadly generalisable. In practice, researchers are frequently forced to prioritise either scale or diversity, with few datasets offering both simultaneously.
Taken together, these limitations highlight that current egocentric datasets, while powerful, provide an incomplete and uneven foundation for action-centric and anticipatory modelling. The gaps identified above are not shortcomings of individual dataset efforts, but rather systemic characteristics of the field’s evolution to date. Addressing these issues will require coordinated efforts toward broader domain coverage, clearer anticipation-oriented annotation protocols, richer multimodal sensing, and improved alignment across annotation schemes. These observations motivate the guidelines for future dataset development discussed in the following subsection and provide context for interpreting results reported across existing benchmarks.

Author Contributions

Conceptualization, A.M. and E.S.; methodology, A.M.; data curation, A.M.; writing—original draft preparation, A.M.; writing—review and editing, A.M. and E.S.; supervision, E.S.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

Not applicable.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A Dataset Inventory

Table A1. Unified overview of action-centric video datasets. Datasets are summarized according to capture perspective, sensing modalities, annotation structure, scale, and application context. Abbreviations: Persp.=Perspective; E=Egocentric; X=Exocentric; M=Multi-view; RGB=Color video; D=Depth; A=Audio; IMU=Inertial Measurement Unit; Gz=Gaze; Hd=Head pose; Act.=Action labels; Obj.=Object annotations; Tmp.=Temporal boundaries; Cls.=Number of classes; Avg.=Average duration (seconds); Tot.=Total duration (hours); Vids.=Number of videos.
Table A1. Unified overview of action-centric video datasets. Datasets are summarized according to capture perspective, sensing modalities, annotation structure, scale, and application context. Abbreviations: Persp.=Perspective; E=Egocentric; X=Exocentric; M=Multi-view; RGB=Color video; D=Depth; A=Audio; IMU=Inertial Measurement Unit; Gz=Gaze; Hd=Head pose; Act.=Action labels; Obj.=Object annotations; Tmp.=Temporal boundaries; Cls.=Number of classes; Avg.=Average duration (seconds); Tot.=Total duration (hours); Vids.=Number of videos.
Dataset Year Persp. Modalities Annotations Cls. Avg. (s) Tot. (h) Vids. Context
RGB D A IMU Gz Hd Act. Obj. Tmp.
50 Salads [21] 2013 X 10 N/A 4.5 50 Kitchen, food preparation
ActivityNet [22] 2015 X 203 N/A 849 27801 Daily human activities
ADL [23] 2022 X 31 N/A N/A N/A Smart home, indoor, daily living
AEA [24] 2024 E N/A N/A 7.3 143 Indoor daily livings, smart home activities
Assembly101 [25] 2022 M 202(C)
1380(FG)
426 513 4321 Assembly / Disassembly
AssemblyHands [26] 2023 M 6 N/A N/A N/A Procedural assembly, hand–object interactions
AVA [27] 2018 X 80 N/A 107.5 430 Movies, daily human activities
BDDA [28] 2024 X N/A N/A N/A N/A Autonomous driving, adverse weather and lighting
BDD100K [29] 2020 X 10 40 1111 100000 Autonomous driving
BEOID [30] 2014 E 75 16.3 N/A 58 Daily object interactions(kitchen, office, gym, household)
BON [31] 2022 E 18 5 3.7 2639 Office/workplace activities
Breakfast [32] 2014 X 10(A)
48(C)
N/A 77 520 Kitchen, cooking activities
CAD-120 [33] 2012 X 10 N/A N/A 120 Indoor daily activities
Charades [34] 2016 X 157(A)
46(Obj)
30.1 82.3 9848 Indoor daily activities, household environments
Charades-Ego [35] 2018 M 157 31.2 69.3 8000 Indoor daily activities, household environments
CMU-MMAC [36] 2008 M 31(A)
5(recipes)
900 6.25 25 Kitchen, cooking activities
COIN [37] 2019 X 180 141.6 476.6 11827 Instructional/daily activities
CrossTask [38] 2019 X 83 N/A 51 4700 Instructional, how-to activities
EasyCom [39] 2021 E N/A N/A 5.3 15 Conversational interactions
EDINA [40] 2022 E N/A N/A 16 N/A Indoor daily activities, scene understanding
EDUB-Seg [41] 2017 E N/A N/A N/A N/A Daily life, photo streams (wearable camera)
EGTEA Gaze + [18] 2018 E 106 4.2 28 86 Kitchen/food preparation with first-person video, audio, binocular gaze tracking, and sparse hand masks
Ego4D [14] 2022 E N/A N/A 3025 3670 Daily life activities
EgoBody3M [42] 2024 E N/A N/A 31.8 2688 Indoor environments/VR body motion
EgoCVR [43] 2024 E N/A 7.9 N/A 2295 Daily-life activities (derived from Ego4D and FHO)
EgoExo4D [44] 2024 M N/A N/A 1286 5035 Skilled human activities (e.g., cooking, bike repair, health care, music, etc.
EgoExo-Fitness [45] 2024 M 12 N/A 32 1276 Fitness/exercise activities
EgoExoLearn [46] 2024 M 19 N/A N/A 1600 Procedural activities
Ego-HOIBench [47] 2025 E 18 N/A N/A N/A Daily human–object interactions
EgoGesture [48] 2018 E 83 N/A N/A 2081 Hand gesture interaction
EgoK360 [49] 2020 E 45 N/A N/A 127 Daily human activities captured with 360° cameras
EgoMe [50] 2025 M 184 18.25 82.46 15804 Real-world daily activities
EgoObjects [51] 2023 E 368 N/A N/A N/A Indoor everyday object-centric scenes captured from wearable devices across households and offices
EgoOops [52] 2024 E N/A N/A N/A N/A Procedural daily activities with mistake actions, aligned to instructional texts
EgoPet [53] 2024 E N/A N/A 84 819 Animal perception and interaction
EgoProceL [54] 2022 E 16 769.2 62 329 Procedural daily-life activities
EgoTracks [55] 2023 E N/A 367.9 602.9 5708 Long-term object tracking in daily-life activities (sourced from Ego4D)
EgoVid5M [56] 2024 E N/A N/A N/A 5M Large-scale video–action dataset derived from Ego4D for video generation
EPIC-Fields [57] 2023 E 90K N/A 99 671 Kitchen/cooking activities with 3D camera pose and geometry augmentation
EPIC-KITCHENS-100 [15] 2022 E 97(V), 300(N), 4053(A) N/A 100 700 Daily unscripted kitchen activities
EPIC-KITCHENS-55 [17] 2018 E 125(V), 331(N) N/A 55 432 Daily kitchen activities
EPIC-Sounds [58] 2025 E 44 audio classes 514 100 700 Daily-life kitchen activities/sounds (derived from EPIC-KITCHENS-100 audio)
EPIC-Tent [59] 2019 E 12(A), 38(FG) 816 5.4 24 Outdoor camping scenario, recording of non-rigid object manipulation during tent assembly
First-Person-Social-Interactions [16] 2012 E 6 N/A N/A 8 Day-long real-world social events (theme parks, group outings), captured with head-mounted cameras, focusing on social interaction patterns based on face attention and first-person motion
FPHA [60] 2018 E 45 N/A 58.6 1175 Daily hand–object manipulation actions captured with an RGB-D shoulder-mounted camera, featuring accurate 3D hand pose annotations obtained via magnetic sensors, designed for studying first-person hand action recognition, hand pose estimation, and hand–object interaction.
FPV-O [61] 2018 E 20 N/A 3.02 12 Office activities recorded with a chest-mounted GoPro camera, covering person-to-person interactions, person-to-object interactions, and locomotion, annotated with temporal action segments in a realistic office environment.
GTEA [19] 2011 E 7 N/A N/A 28 Kitchen/food preparation
GTEA Gaze [16] 2012 E 7 N/A N/A 17 Kitchen/food preparation activities captured from a first-person perspective with synchronized eye-gaze tracking for studying gaze-guided action recognition
GTEA Gaze + [62] 2015 E 44 N/A 9 86 Kitchen/food preparation activities captured in real-world environments with head-mounted cameras and binocular eye-gaze tracking
H2O [63] 2021 E 36 N/A N/A N/A Indoor hand–object manipulation
HD-EPIC [64] 2025 E N/A 15.9 41.3 156 Unscripted kitchen activities in real home environments
HOI4D [65] 2022 E 16 N/A N/A 4K Indoor human–object interaction with category-level object manipulation
HoloAssist [66] 2023 E 414(C,A), 39(V), 90(N) N/A 166 2221 Real-world assistive, interactive task completion
HowToDiv [67] 2025 E N/A N/A 24 507 Step-by-step procedural task assistance instructional videos transformed into expert–novice dialogues
IndustReal [68] 2024 E 12 N/A N/A 200 Industrial-like procedural tasks with explicit modeling of correct and erroneous execution steps
JAAD [69] 2016 E N/A 5-15 N/A 346 Urban driving scenarios focusing on joint attention between drivers and pedestrians
KrishnaCam [70] 2016 E N/A N/A 70.2 460 Longitudinal single-person video stream for scene understanding and motion prediction
MECCANO [71] 2023 E 61(A), 12(V), 20(Obj) 21 7 20 Industrial-like human–object interaction during assembly tasks
MPII Cooking [72] 2012 X 65 N/A >8 44 Continuous kitchen recordings for fine-grained cooking activity classification and detection
MPII Cooking 2 [73] 2015 X 67(FC, A), 155(Obj), 59(A) N/A >27 273 Fine-grained and composite cooking activity understanding
MultiTHUMOS [74] 2018 X 65 N/A 30 413 Dense multilabel action annotations over long, unconstrained internet videos extending the THUMOS dataset
N-EPIC-Kitchens [75] 2021 E 8 N/A N/A N/A Event-based action recognition dataset extending EPIC-Kitchens with simulated neuromorphic event streams
Nymeria [76] 2024 E N/A ~900 300 1200 Large-scale multimodal dataset capturing daily human motion with synchronized vision, audio, IMUs, gaze, full-body motion capture, and hierarchical language annotations
PH-Ego [77] 2023 E 9 ~390 N/A N/A Personal-health–oriented action anticipation dataset capturing scripted kitchen activities such as hygiene and food preparation for studying domain shift and untrimmed anticipation
PIE [78] 2019 X 6 ~600 >6 N/A On-board driving dataset for pedestrian intention estimation and trajectory prediction in urban traffic scenes
THUMOS14 [79] 2017 X 101 N/A ~254 N/A Large-scale benchmark for action recognition and temporal action detection in untrimmed real-world videos collected from YouTube
TV-I [80] 2010 X 4 N/A N/A 300 Realistic human–human interaction recognition dataset composed of clips extracted from TV shows for interaction retrieval in unconstrained scenes
TVSeries [81] 2016 X 30 N/A 30 27 Realistic long-form TV show footage for online action detection with dense frame-level temporal annotations across multiple concurrent actions
Tasty-Videos [82] 2019 X N/A 54 37.6 2511 Short-form instructional cooking videos with temporally aligned procedural steps for zero-shot action anticipation and recipe understanding
UnrealEgo [83] 2022 E 24 N/A N/A N/A Large-scale synthetic dataset generated in Unreal Engine for training and evaluating action recognition models under controlled yet photorealistic environments
UnrealEgo2 [84] 2024 E N/A N/A 14 N/A Large-scale synthetic dataset for 3D human pose estimation with diverse full-body motions rendered from virtual eyeglasses-based cameras
VISOR [85] 2022 E 257 720 36 179 Pixel-level segmentation benchmark built on EPIC-KITCHENS, focusing on hands and active objects with long-term temporal consistency for video object segmentation, interaction understanding, and long-term reasoning
VOST [86] 2023 E 155(Obj), 51(Trans. types) 21.2 4.2 713 Video Object Segmentation benchmark focusing on objects undergoing complex transformations (e.g., cutting, breaking, peeling), with dense instance-level masks over relatively long clips
WEAR [87] 2024 E 18 N/A 44 347 Outdoor sports dataset combining video and wearable IMU data for long-duration activity recognition in real-world athletic scenarios
YouCook2 [88] 2018 X N/A 316 175.6 2K Large-scale instructional cooking video dataset with temporally segmented procedural steps and natural language descriptions for learning and understanding complex procedures

References

  1. Paullada, A.; Raji, I.D.; Bender, E.M.; Denton, E.; Hanna, A. Data and its (dis) contents: A survey of dataset development and use in machine learning research. Patterns 2021, 2. [CrossRef]
  2. Salari, A.; Djavadifar, A.; Liu, X.; Najjaran, H. Object recognition datasets and challenges: A review. Neurocomputing 2022, 495, 129–152. [CrossRef]
  3. Hernandez, J.A.; Colom, M. Reproducible research policies and software/data management in scientific computing journals: a survey, discussion, and perspectives. Frontiers in Computer Science 2025, 6, 1491823. [CrossRef]
  4. Intelligence, N.M. How to be responsible in AI publication. Nature Machine Intelligence 2021, 3. [CrossRef]
  5. Rodin, I.; Furnari, A.; Mavroeidis, D.; Farinella, G.M. Predicting the future from first person (egocentric) vision: A survey. Computer Vision and Image Understanding 2021, 211, 103252. [CrossRef]
  6. Núñez-Marcos, A.; Azkune, G.; Arganda-Carreras, I. Egocentric vision-based action recognition: A survey. Neurocomputing 2022, 472, 175–197. [CrossRef]
  7. Li, X.; Qiu, H.; Wang, L.; Zhang, H.; Qi, C.; Han, L.; Xiong, H.; Li, H. Challenges and Trends in Egocentric Vision: A Survey. arXiv preprint arXiv:2503.15275 2025.
  8. Lai, B.; Toyer, S.; Nagarajan, T.; Girdhar, R.; Zha, S.; Rehg, J.M.; Kitani, K.; Grauman, K.; Desai, R.; Liu, M. Human action anticipation: A survey. arXiv preprint arXiv:2410.14045 2024.
  9. Zhang, X.; He, Y.; Xu, R.; Yu, H.; Shen, Z.; Cui, P. Nico++: Towards better benchmarking for domain generalization. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 16036–16047.
  10. Wang, Y.; Wu, Y.; Zhang, H. Lost domain generalization is a natural consequence of lack of training domains. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2024, Vol. 38, pp. 15689–15697. [CrossRef]
  11. Panda, A.; Panigrahi, D.; Mitra, S.; Mittal, S.; Rahimi, S. Transfer Learning Applied to Computer Vision Problems: Survey on Current Progress, Limitations, and Opportunities. arXiv preprint arXiv:2409.07736 2024.
  12. Zhao, Z.; Alzubaidi, L.; Zhang, J.; Duan, Y.; Gu, Y. A comparison review of transfer learning and self-supervised learning: Definitions, applications, advantages and limitations. Expert Systems with Applications 2024, 242, 122807. [CrossRef]
  13. Regulation (EU) 2016/679 ... (GDPR), 2016. OJ L 119, 04.05.2016; https://eur-lex.europa.eu/eli/reg/2016/679/oj.
  14. Grauman, K.; Westbury, A.; Byrne, E.; Chavis, Z.; Furnari, A.; Girdhar, R.; Hamburger, J.; Jiang, H.; Liu, M.; Liu, X.; et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 18995–19012.
  15. Damen, D.; Doughty, H.; Farinella, G.M.; Furnari, A.; Kazakos, E.; Ma, J.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; et al. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International Journal of Computer Vision 2022, 130, 33–55. [CrossRef]
  16. Fathi, A.; Li, Y.; Rehg, J.M. Learning to recognize daily actions using gaze. In Proceedings of the European Conference on Computer Vision. Springer, 2012, pp. 314–327.
  17. Damen, D.; Doughty, H.; Farinella, G.M.; Fidler, S.; Furnari, A.; Kazakos, E.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; et al. Scaling egocentric vision: The EPIC-Kitchens dataset. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 720–736.
  18. Li, Y.; Liu, M.; Rehg, J.M. In the Eye of the Beholder: Joint learning of gaze and actions in first-person video. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 619–635.
  19. Fathi, A.; Ren, X.; Rehg, J.M. Learning to recognize objects in egocentric activities. In Proceedings of the CVPR. IEEE, 2011, pp. 3281–3288.
  20. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. bmj 2021, 372. [CrossRef]
  21. Stein, S.; McKenna, S.J. Combining embedded accelerometers with computer vision for recognizing food preparation activities. In Proceedings of the Proceedings of the 2013 ACM International Joint Conference on Pervasive and Ubiquitous Computing, 2013, pp. 729–738.
  22. Caba Heilbron, F.; Escorcia, V.; Ghanem, B.; Carlos Niebles, J. ActivityNet: A large-scale video benchmark for human activity understanding. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 961–970.
  23. Pan, J.; Yuan, Z.; Zhang, X.; Chen, D. YouHome system and dataset: Making your home know you better. In Proceedings of the Proceedings of the IEEE International Symposium on Smart Electronic Systems (iSES). IEEE, 2022, pp. 414–420.
  24. Lv, Z.; Charron, N.; Moulon, P.; Gamino, A.; Peng, C.; Sweeney, C.; Miller, E.; Tang, H.; Meissner, J.; Dong, J.; et al. Aria Everyday Activities Dataset. arXiv preprint arXiv:2402.13349 2024.
  25. Sener, F.; Chatterjee, D.; Shelepov, D.; He, K.; Singhania, D.; Wang, R.; Yao, A. Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 21096–21106.
  26. Ohkawa, T.; He, K.; Sener, F.; Hodan, T.; Tran, L.; Keskin, C. AssemblyHands: Towards egocentric activity understanding via 3D hand pose estimation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 12999–13008.
  27. Gu, C.; Sun, C.; Ross, D.A.; Vondrick, C.; Pantofaru, C.; Li, Y.; Vijayanarasimhan, S.; Toderici, G.; Ricco, S.; Sukthankar, R.; et al. AVA: A video dataset of spatio-temporally localized atomic visual actions. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 6047–6056.
  28. Assion, F.; Gressner, F.; Augustine, N.; Klemenc, J.; Hammam, A.; Krattinger, A.; Trittenbach, H.; Philippsen, A.; Riemer, S. A-BDD: Leveraging data augmentations for safe autonomous driving in adverse weather and lighting. arXiv preprint arXiv:2408.06071 2024.
  29. Yu, F.; Chen, H.; Wang, X.; Xian, W.; Chen, Y.; Liu, F.; Madhavan, V.; Darrell, T. BDD100K: A diverse driving dataset for heterogeneous multitask learning. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 2636–2645.
  30. Damen, D.; Haines, O.; Leelasawassuk, T.; Calway, A.; Mayol-Cuevas, W. Multi-user egocentric online system for unsupervised assistance on object usage. In Proceedings of the European Conference on Computer Vision. Springer, 2014, pp. 481–492.
  31. Tadesse, G.A.; Bent, O.; Weldemariam, K.; Istiak, M.A.; Hasan, T.; Cavallaro, A. BON: An extended public domain dataset for human activity recognition. arXiv preprint arXiv:2209.05077 2022.
  32. Kuehne, H.; Arslan, A.; Serre, T. The language of actions: Recovering the syntax and semantics of goal-directed human activities. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 780–787.
  33. Koppula, H.S.; Gupta, R.; Saxena, A. Human Activity Learning using Object Affordances from RGB-D Videos. arXiv preprint arXiv:1208.XXXX 2012.
  34. Sigurdsson, G.A.; Varol, G.; Wang, X.; Farhadi, A.; Laptev, I.; Gupta, A. Hollywood in Homes: Crowdsourcing data collection for activity understanding. In Proceedings of the European Conference on Computer Vision. Springer, 2016, pp. 510–526.
  35. Sigurdsson, G.A.; Gupta, A.; Schmid, C.; Farhadi, A.; Alahari, K. Actor and observer: Joint modeling of first and third-person videos. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 7396–7404.
  36. De la Torre Frade, F.; Hodgins, J.K.; Bargteil, A.W.; Martin Artal, X.; Macey, J.C.; Collado I Castells, A.; Beltran, J. Guide to the Carnegie Mellon University Multimodal Activity (CMU-MMAC) Database. Technical Report CMU-RI-TR-08-22, Pittsburgh, PA, 2008.
  37. Tang, Y.; Ding, D.; Rao, Y.; Zheng, Y.; Zhang, D.; Zhao, L.; Lu, J.; Zhou, J. COIN: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 1207–1216.
  38. Zhukov, D.; Alayrac, J.B.; Cinbis, R.G.; Fouhey, D.; Laptev, I.; Sivic, J. Cross-task weakly supervised learning from instructional videos. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 3537–3545.
  39. Donley, J.; Tourbabin, V.; Lee, J.S.; Broyles, M.; Jiang, H.; Shen, J.; Pantic, M.; Ithapu, V.K.; Mehra, R. EasyCom: An augmented reality dataset to support algorithms for easy communication in noisy environments. arXiv preprint arXiv:2107.04174 2021.
  40. Do, T.; Vuong, K.; Park, H.S. Egocentric Scene Understanding via Multimodal Spatial Rectifier. In Proceedings of the CVPR, 2022.
  41. Dimiccoli, M.; Bolanos, M.; Talavera, E.; Aghaei, M.; Nikolov, S.G.; Radeva, P. SR-clustering: Semantic regularized clustering for egocentric photo streams segmentation. Computer Vision and Image Understanding 2017, 155, 55–69. [CrossRef]
  42. Zhao, A.; Tang, C.; Wang, L.; Li, Y.; Dave, M.; Tao, L.; Twigg, C.D.; Wang, R.Y. EgoBody3M: Egocentric body tracking on a VR headset using a diverse dataset. In Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 375–392.
  43. Hummel, T.; Karthik, S.; Georgescu, M.I.; Akata, Z. EgoCVR: An egocentric benchmark for fine-grained composed video retrieval. In Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 1–17.
  44. Grauman, K.; Westbury, A.; Torresani, L.; Kitani, K.; Malik, J.; Afouras, T.; Kumar, A.; Baiyya, V.; Bansal, S.; Boote, B.; et al. Ego-Exo4D: Understanding skilled human activity from first- and third-person perspectives. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 19383–19400.
  45. Li, Y.M.; Huang, W.J.; Wang, A.L.; Zeng, L.A.; Meng, J.K.; Zheng, W.S. EgoExo-Fitness: Towards egocentric and exocentric full-body action understanding. In Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 363–382.
  46. Huang, Y.; Chen, G.; Xu, J.; Zhang, M.; Yang, L.; Pei, B.; Zhang, H.; Dong, L.; Wang, Y.; Wang, L.; et al. EgoExoLearn: A dataset for bridging asynchronous ego- and exo-centric view of procedural activities in real world. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 22072–22086.
  47. Deng, K.; Wang, Y.; Chau, L.P. Egocentric Human-Object Interaction Detection: A New Benchmark and Method. arXiv preprint arXiv:2506.14189 2025.
  48. Zhang, Y.; Cao, C.; Cheng, J.; Lu, H. EgoGesture: A new dataset and benchmark for egocentric hand gesture recognition. IEEE Transactions on Multimedia 2018, 20, 1038–1050. [CrossRef]
  49. Bhandari, K.; DeLaGarza, M.A.; Zong, Z.; Latapie, H.; Yan, Y. EgoK360: A 360° egocentric kinetic human activity video dataset. In Proceedings of the IEEE International Conference on Image Processing (ICIP). IEEE, 2020, pp. 266–270.
  50. Qiu, H.; Shi, Z.; Wang, L.; Xiong, H.; Li, X.; Li, H. EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World. arXiv preprint arXiv:2501.19061 2025.
  51. Zhu, C.; Xiao, F.; Alvarado, A.; Babaei, Y.; Hu, J.; El-Mohri, H.; Culatana, S.; Sumbaly, R.; Yan, Z. EgoObjects: A large-scale egocentric dataset for fine-grained object understanding. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 20110–20120.
  52. Haneji, Y.; Nishimura, T.; Kameko, H.; Shirai, K.; Yoshida, T.; Kajimura, K.; Yamamoto, K.; Cui, T.; Nishimoto, T.; Mori, S. EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos with Procedural Texts, 2024, [arXiv:cs.CV/2410.05343].
  53. Bar, A.; Bakhtiar, A.; Tran, D.; Loquercio, A.; Rajasegaran, J.; LeCun, Y.; Globerson, A.; Darrell, T. EgoPet: Egomotion and interaction data from an animal’s perspective. In Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 377–394.
  54. Bansal, S.; Arora, C.; Jawahar, C. My View is the Best View: Procedure Learning from Egocentric Videos. In Proceedings of the European Conference on Computer Vision (ECCV), 2022.
  55. Tang, H.; Liang, K.J.; Grauman, K.; Feiszli, M.; Wang, W. EgoTracks: A long-term egocentric visual object tracking dataset. Advances in Neural Information Processing Systems 2023, 36, 75716–75739. [CrossRef]
  56. Wang, X.; Zhao, K.; Liu, F.; Wang, J.; Zhao, G.; Bao, X.; Zhu, Z.; Zhang, Y.; Wang, X. EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation. arXiv preprint arXiv:2411.08380 2024.
  57. Tschernezki, V.; Darkhalil, A.; Zhu, Z.; Fouhey, D.; Laina, I.; Larlus, D.; Damen, D.; Vedaldi, A. Epic Fields: Marrying 3D geometry and video understanding. Advances in Neural Information Processing Systems 2023, 36, 26485–26500. [CrossRef]
  58. Huh, J.; Chalk, J.; Kazakos, E.; Damen, D.; Zisserman, A. EPIC-SOUNDS: A Large-Scale Dataset of Actions that Sound. IEEE Transactions on Pattern Analysis and Machine Intelligence 2025.
  59. Jang, Y.; Sullivan, B.; Ludwig, C.; Gilchrist, I.; Damen, D.; Mayol-Cuevas, W. EPIC-Tent: An egocentric video dataset for camping tent assembly. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, 2019, pp. 0–0.
  60. Garcia-Hernando, G.; Yuan, S.; Baek, S.; Kim, T.K. First-person hand action benchmark with RGB-D videos and 3D hand pose annotations. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 409–419.
  61. Abebe, G.; Catala, A.; Cavallaro, A. A first-person vision dataset of office activities. In Proceedings of the IAPR Workshop on Multimodal Pattern Recognition of Social Signals in Human-Computer Interaction. Springer, 2018, pp. 27–37.
  62. Li, Y.; Ye, Z.; Rehg, J.M. Delving into egocentric actions. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 287–295.
  63. Kwon, T.; Tekin, B.; Stühmer, J.; Bogo, F.; Pollefeys, M. H2O: Two hands manipulating objects for first-person interaction recognition. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 10138–10148.
  64. Perrett, T.; Darkhalil, A.; Sinha, S.; Emara, O.; Pollard, S.; Parida, K.K.; Liu, K.; Gatti, P.; Bansal, S.; Flanagan, K.; et al. HD-EPIC: A highly-detailed egocentric video dataset. In Proceedings of the Proceedings of the CVPR, 2025, pp. 23901–23913.
  65. Liu, Y.; Liu, Y.; Jiang, C.; Lyu, K.; Wan, W.; Shen, H.; Liang, B.; Fu, Z.; Wang, H.; Yi, L. HOI4D: A 4D egocentric dataset for category-level human-object interaction. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 21013–21022.
  66. Wang, X.; Kwon, T.; Rad, M.; Pan, B.; Chakraborty, I.; Andrist, S.; Bohus, D.; Feniello, A.; Tekin, B.; Frujeri, F.V.; et al. HoloAssist: An egocentric human interaction dataset for interactive AI assistants in the real world. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 20270–20281.
  67. Aggarwal, L.; Bahirwani, V.; Li, L.; Colaco, A. Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark. arXiv preprint arXiv:2508.11192 2025.
  68. Schoonbeek, T.J.; Houben, T.; Onvlee, H.; Van der Sommen, F.; et al. IndustReal: A dataset for procedure step recognition handling execution errors in egocentric videos in an industrial-like setting. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 4365–4374.
  69. Kotseruba, I.; Rasouli, A.; Tsotsos, J.K. Joint Attention in Autonomous Driving (JAAD). arXiv preprint arXiv:1609.04741 2016.
  70. Singh, K.K.; Fatahalian, K.; Efros, A.A. KrishnaCam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks. In Proceedings of the Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 2016, pp. 1–9.
  71. Ragusa, F.; Furnari, A.; Farinella, G.M. MECCANO: A multimodal egocentric dataset for human behavior understanding in the industrial-like domain. Computer Vision and Image Understanding 2023, 235, 103764. [CrossRef]
  72. Rohrbach, M.; Amin, S.; Andriluka, M.; Schiele, B. A database for fine grained activity detection of cooking activities. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2012, pp. 1194–1201.
  73. Rohrbach, M.; Rohrbach, A.; Regneri, M.; Amin, S.; Andriluka, M.; Pinkal, M.; Schiele, B. Recognizing Fine-Grained and Composite Activities Using Hand-Centric Features and Script Data. International Journal of Computer Vision 2015, pp. 1–28.
  74. Yeung, S.; Russakovsky, O.; Jin, N.; Andriluka, M.; Mori, G.; Fei-Fei, L. Every moment counts: Dense detailed labeling of actions in complex videos. International Journal of Computer Vision 2018, 126, 375–389.
  75. Plizzari, C.; Planamente, M.; Goletto, G.; Cannici, M.; Gusso, E.; Matteucci, M.; Caputo, B. E2(GO) MOTION: Motion Augmented Event Stream for Egocentric Action Recognition. arXiv preprint arXiv:2112.03596 2021.
  76. Ma, L.; Ye, Y.; Hong, F.; Guzov, V.; Jiang, Y.; Postyeni, R.; Pesqueira, L.; Gamino, A.; Baiyya, V.; Kim, H.J.; et al. Nymeria: A massive collection of multimodal egocentric daily motion in the wild. In Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 445–465.
  77. Rodin, I.; Furnari, A.; Mavroeidis, D.; Farinella, G.M. Egocentric action anticipation for personal health. In Proceedings of the ICASSP. IEEE, 2023, pp. 1–5.
  78. Rasouli, A.; Kotseruba, I.; Kunic, T.; Tsotsos, J.K. PIE: A Large-Scale Dataset and Models for Pedestrian Intention Estimation and Trajectory Prediction. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision, 2019.
  79. Idrees, H.; Zamir, A.R.; Jiang, Y.G.; Gorban, A.; Laptev, I.; Sukthankar, R.; Shah, M. The THUMOS challenge on action recognition for videos in the wild. Computer Vision and Image Understanding 2017, 155, 1–23.
  80. Patron-Perez, A.; Marszalek, M.; Zisserman, A.P.; Reid, I.D. High five: Recognising human interactions in TV shows 2010.
  81. De Geest, R.; Gavves, E.; Ghodrati, A.; Li, Z.; Snoek, C.; Tuytelaars, T. Online action detection. In Proceedings of the European Conference on Computer Vision. Springer, 2016, pp. 269–284.
  82. Sener, F.; Yao, A. Zero-Shot Anticipation for Instructional Activities. In Proceedings of the ICCV, 2019.
  83. Akada, H.; Wang, J.; Shimada, S.; Takahashi, M.; Theobalt, C.; Golyanik, V. UnrealEgo: A new dataset for robust egocentric 3D human motion capture. In Proceedings of the European Conference on Computer Vision. Springer, 2022, pp. 1–17.
  84. Akada, H.; Wang, J.; Golyanik, V.; Theobalt, C. 3D human pose perception from egocentric stereo videos. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 767–776.
  85. Darkhalil, A.; Shan, D.; Zhu, B.; Ma, J.; Kar, A.; Higgins, R.; Fidler, S.; Fouhey, D.; Damen, D. EPIC-Kitchens VISOR benchmark: Video segmentations and object relations. Advances in Neural Information Processing Systems 2022, 35, 13745–13758.
  86. Tokmakov, P.; Li, J.; Gaidon, A. Breaking the “Object” in Video Object Segmentation. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023.
  87. Bock, M.; Kuehne, H.; Van Laerhoven, K.; Moeller, M. WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 2024, 8. https://doi.org/10.1145/3699776.
  88. Zhou, L.; Xu, C.; Corso, J. Towards automatic learning of procedures from web instructional videos. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2018, Vol. 32.
Figure 1. Perspective taxonomy illustrating the distinction between egocentric and exocentric viewpoints.
Figure 1. Perspective taxonomy illustrating the distinction between egocentric and exocentric viewpoints.
Preprints 227403 g001
Figure 2. Taxonomy of egocentric domains spanning daily living, sports, industrial activities, healthcare, social interaction, and other real-world settings.
Figure 2. Taxonomy of egocentric domains spanning daily living, sports, industrial activities, healthcare, social interaction, and other real-world settings.
Preprints 227403 g002
Figure 3. Multimodal taxonomy of egocentric sensing, organised into visual, motion, audio, language, and 3D/environmental modalities. Solid-bordered elements denote raw sensor signals directly captured during data acquisition, while dashed-bordered elements indicate derived modalities obtained through post-processing or model-based estimation (e.g., optical flow, pose estimation, etc.).
Figure 3. Multimodal taxonomy of egocentric sensing, organised into visual, motion, audio, language, and 3D/environmental modalities. Solid-bordered elements denote raw sensor signals directly captured during data acquisition, while dashed-bordered elements indicate derived modalities obtained through post-processing or model-based estimation (e.g., optical flow, pose estimation, etc.).
Preprints 227403 g003
Figure 4. Structured taxonomy of annotation types in egocentric datasets. The figure groups supervisory signals according to their primary function: frame-level perceptual annotations, temporal action and anticipation labels, interaction-centric annotations capturing hand–object and relational events, spatial and trajectory-based annotations encoding 3D structure and motion, and audio–text annotations providing linguistic and acoustic context. The taxonomy is illustrative rather than exhaustive and highlights the diversity of annotation practices across egocentric benchmarks.
Figure 4. Structured taxonomy of annotation types in egocentric datasets. The figure groups supervisory signals according to their primary function: frame-level perceptual annotations, temporal action and anticipation labels, interaction-centric annotations capturing hand–object and relational events, spatial and trajectory-based annotations encoding 3D structure and motion, and audio–text annotations providing linguistic and acoustic context. The taxonomy is illustrative rather than exhaustive and highlights the diversity of annotation practices across egocentric benchmarks.
Preprints 227403 g004
Figure 5. Taxonomy of tasks in egocentric video understanding. The figure organises tasks according to their semantic scope and temporal orientation, ranging from low-level perceptual tasks to action and activity understanding, and further to prospective tasks such as anticipation and forecasting. Additional task families related to gaze and attention, hand–object interaction and manipulation, three-dimensional mapping and tracking, vision–language reasoning, and summarisation are shown to illustrate the breadth of task formulations supported by egocentric datasets. The taxonomy highlights the diversity of task support across benchmarks and clarifies which dataset designs are appropriate for different research objectives.
Figure 5. Taxonomy of tasks in egocentric video understanding. The figure organises tasks according to their semantic scope and temporal orientation, ranging from low-level perceptual tasks to action and activity understanding, and further to prospective tasks such as anticipation and forecasting. Additional task families related to gaze and attention, hand–object interaction and manipulation, three-dimensional mapping and tracking, vision–language reasoning, and summarisation are shown to illustrate the breadth of task formulations supported by egocentric datasets. The taxonomy highlights the diversity of task support across benchmarks and clarifies which dataset designs are appropriate for different research objectives.
Preprints 227403 g005
Figure 8. Distribution of datasets by perspective across the curated corpus. The figure provides a static snapshot of the proportion of egocentric, exocentric, and multi-view datasets, independent of their year of release.
Figure 8. Distribution of datasets by perspective across the curated corpus. The figure provides a static snapshot of the proportion of egocentric, exocentric, and multi-view datasets, independent of their year of release.
Preprints 227403 g008
Figure 9. Distribution of sensing modalities across the curated dataset corpus. The figure reports the number of datasets that include each modality. Counts are non-exclusive, as individual datasets may provide multiple modalities (e.g., RGB video combined with audio, gaze, or inertial measurements).
Figure 9. Distribution of sensing modalities across the curated dataset corpus. The figure reports the number of datasets that include each modality. Counts are non-exclusive, as individual datasets may provide multiple modalities (e.g., RGB video combined with audio, gaze, or inertial measurements).
Preprints 227403 g009
Figure 10. Joint distribution of dataset perspective and application domain. Stacked bars show the number of datasets per domain, grouped by perspective, highlighting domain-specific preferences for egocentric, exocentric, and multi-view data collection.
Figure 10. Joint distribution of dataset perspective and application domain. Stacked bars show the number of datasets per domain, grouped by perspective, highlighting domain-specific preferences for egocentric, exocentric, and multi-view data collection.
Preprints 227403 g010
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.