Submitted:
09 August 2026
Posted:
11 August 2026
You are already at the latest version
Abstract
Recent advances in artificial intelligence (AI) are reshaping smart glasses from egocentric capture and display devices into platforms for wearable intelligence. Smart glasses increasingly serve as wearable AI systems that connect first-person observation with real-time assistance under strict form-factor constraints. We frame this transition through the lens of AI smart glasses and define them as a system-level concept in which egocentric sensing, resource-aware computing, intelligent reasoning, multimodal interaction, and real-world application constraints are co-designed for personalized assistance in the physical world. To systematically study this perspective, we organize the survey around four connected dimensions. First, we examine the hardware foundation that bounds sensing, computation, feedback delivery, and sustained deployment. Second, we study wearable intelligence, where egocentric signals are transformed into perceptual, contextual, and agentic capabilities. Third, we discuss interaction design, through which users request, receive, correct, and regulate assistance during ongoing activity. Fourth, we analyze application scenarios across healthcare, accessibility, situated learning, daily life assistance, cultural tourism, and industrial support, showing how domain requirements reshape system design and evaluation. We further identify five cross-cutting research challenges for future AI smart glasses: next-generation hardware, trustworthy egocentric intelligence, lifelong personalized memory, proactive intelligence, and embodied foundation models. By centering smart glasses as wearable-intelligence platforms, this survey provides a unified framework for organizing technologies, applications, and open challenges in this emerging area.
Keywords:
AI smart glasses
; egocentric sensing
; multimodal agents
; wearable intelligence
; wearable interaction
; personalized intelligent assistant
1. Introduction
Digital computing increasingly supports information access, interpersonal communication, work coordination, and skill acquisition. However, much of this support still depends on devices that must be held, watched, or explicitly operated, interrupting physical activity and limiting access to the user’s immediate context. Wearable technologies address this limitation by placing sensing, computation, and feedback on the body, enabling digital assistance that is more continuous, contextual, and compatible with mobile or hands-busy situations [1,2]. Among wearable devices, smart glasses are distinctive because their sensors and feedback channels are aligned with the user’s visual field and auditory space. This egocentric form factor allows them to observe the world from the wearer’s perspective and provide hands-free assistance without requiring a full shift of attention to a separate device [3,4].
Recent advances in artificial intelligence (AI), including large multimodal models [5,6,7], retrieval-augmented generation [8,9], and AI agents [10,11,12], are changing the role of smart glasses in wearable computing. Rather than serving mainly as wearable sensing and display devices, smart glasses are becoming platforms for wearable intelligence that combine the glasses form factor with AI-driven perception, reasoning, and action. Work such as EgoLife [13], SUPERGLASSES [14], and VisionClaw [15] illustrates this emerging direction across daily memory, question answering, and agentic service scenarios. In these examples, smart glasses convert egocentric observations into task-relevant states, link those states to external or personal knowledge, infer what the wearer may need, and deliver assistance through external tools, connected services, or situated feedback. In this survey, we use AI smart glasses to denote this emerging class and define it as a system-level concept in which egocentric sensing, resource-aware computing, contextual and agentic intelligence, multimodal interaction, and real-world application constraints are co-designed to support personalized intelligent assistance in the physical world.
Realizing this vision requires coordination across device capabilities, model inference, user interaction, and deployment constraints. Each dimension is shaped by the wearable form factor. Richer sensing may improve perceptual coverage but increases power consumption, device complexity, and privacy concerns for wearers and bystanders; larger models and cloud-enhanced processing may improve contextual and agentic reasoning but add latency, energy cost, and connectivity dependence; and more proactive or expressive assistance may reduce user effort but can also interrupt attention, expose sensitive context, or act on uncertain intent. Taken together, these trade-offs show that AI smart glasses must be designed as an integrated system: they need to transform egocentric sensing into contextual interpretation and situated support while remaining controllable, deployable, and usable in daily wearable contexts.
To systematically characterize AI smart glasses, this survey focuses on smart glasses and related egocentric wearable devices that contribute to the sense–understand–act process that supports situated assistance under wearable constraints. As illustrated in Figure 1, we organize the main body into four connected analytical parts, followed by a cross-cutting discussion of open problems and future directions. This organization follows the path by which wearable intelligence becomes operational in smart glasses: a physical and computational substrate supports perceptual, contextual, and agentic capabilities; interaction design exposes those capabilities to users; application scenarios examine how they are shaped by domain-specific requirements; and future directions synthesize the unresolved research agenda across these components.
First, the Hardware Foundation (Section 2) studies the physical and computational substrate on which AI smart glasses operate. Sensing arrays define what the device can observe from the wearer’s perspective, while computing platforms determine how local and cloud-assisted inference are coordinated under wearable constraints. Interaction units and supporting infrastructure then determine how assistance is delivered, controlled, and sustained during daily use. This part establishes a basic design principle: the intelligence of smart glasses is bounded by the physical constraints of wearable deployment.
Building on this substrate, Wearable Intelligence (Section 3) examines how AI smart glasses convert egocentric signals into assistance. We organize this process into three levels: Perception Intelligence extracts spatial, temporal, conceptual, and social structure from egocentric observations; Contextual Intelligence relates these observations to the personal, environmental, temporal, and knowledge state; Agentic Intelligence moves from understanding to action through intent reasoning, planning, and execution. These levels form a closed loop from observation to response, rather than a collection of isolated recognition models.
Next, Interaction Design (Section 4) studies how wearers request, receive, and correct assistance through AI smart glasses during ongoing activity. Since the interaction between wearers and smart glasses is mobile, situated, and attention-constrained, we organize it around intent capture, reference grounding, feedback design, proactive interaction, and inclusive usability. Together, these aspects show how model capability is exposed through interfaces that fit attention, activity, ability, and social context.
Fourth, Application Scenarios (Section 5) examine how the framework is instantiated across healthcare, accessibility, situated learning, daily life assistance, cultural tourism, and industrial support, where assistance is shaped by different risks, responsibilities, and user needs. After these four analytical parts, Open Problems and Future Directions (Section 6) identifies five cross-cutting research challenges for future AI smart glasses: next-generation hardware, trustworthy egocentric intelligence, lifelong personalized memory, proactive intelligence, and embodied foundation models. Addressing these challenges is necessary for personalized assistance that is adaptive, dependable, and trustworthy in real-world wearable contexts.
This organization distinguishes our survey from adjacent reviews, which provide important foundations but cover only part of the system-level perspective on AI smart glasses. Egocentric-vision surveys organize egocentric datasets, perception tasks, and recognition challenges [16]; smart-glasses interaction surveys focus on wearable input and output interfaces [4]; and application-oriented reviews examine healthcare or industrial support [17,18], or domain-specific acceptance [19]. In contrast, this survey centers on AI smart glasses as wearable-intelligence platforms for personalized assistance, emphasizing the coupling among hardware constraints, intelligence capabilities, interaction design, and real-world deployment.
Contributions
This survey makes the following contributions:
- Conceptual framing: We frame AI smart glasses as integrated wearable-intelligence platforms for personalized intelligent assistance, moving beyond views of smart glasses as sensing or display devices.
- Analytical framework: We organize AI smart glasses into four connected parts: hardware foundation, wearable intelligence, interaction design, and application scenarios, and synthesize representative work across sensing, reasoning, interaction, and deployment.
- Applications and future agenda: We connect the framework to real-world application scenarios and outline cross-cutting future directions toward personalized assistance that is adaptive, dependable, and trustworthy in wearable contexts.
2. Hardware Foundation
As the physical foundation of AI smart glasses, advances in wearable hardware have continuously expanded the capabilities of smart glasses over the past decade. As illustrated in Figure 2, this evolution reflects a gradual transition from simple wearable devices to AI assistants. Early attempts, represented by Google Glass, pursued an all-in-one vision by integrating smartphone-like functionalities into a lightweight wearable form factor. However, limited hardware technologies and immature application ecosystems prevented this design from achieving widespread adoption. Therefore, the industry shifted toward hardware specialization, where camera-centric products focused on egocentric capture, audio smart glasses popularized open-ear listening and voice assistants, and display smart glasses advanced optical displays and spatial mapping. These specialized product lines progressively matured the underlying sensing, interaction, and display technologies that later became the hardware foundation of AI smart glasses.
The emergence of multimodal foundation models has fundamentally reshaped this trajectory by integrating these previously specialized hardware capabilities into unified intelligent systems. Unlike traditional smart glasses that simply integrated cameras or headphones, AI smart glasses are evolving into embodied intelligent assistants capable of continuously observing the physical world, contextually reasoning over multimodal information, and proactively assisting users in everyday activities. Such capabilities place unprecedented demands on wearable hardware, which must deliver real-time intelligence under stringent constraints on size, weight, and power consumption.
From the perspective of an intelligent assistant, supporting the complete intelligence pipeline requires four fundamental hardware capabilities, as illustrated in Figure 3. Sensing Arrays(§Section 2.1) continuously observe both the surrounding environment and the wearer through egocentric visual, motion, and spatial signals, providing the sensory foundation for wearable intelligence. Computing Platforms(§Section 2.2) support contextual reasoning over heterogeneous observations by coordinating on-device processors with cloud-based AI models. Interaction Units(§Section 2.3) establish communication channels for capturing user input and presenting system feedback, enabling wearable assistance under mobile and attention-constrained conditions. Supporting Infrastructure(§Section 2.4) provides reliable energy and storage, allowing smart glasses to operate reliably throughout prolonged daily use. Rather than treating cameras, chips, controls, speakers, displays, batteries, and storage as isolated hardware components, we organize them as a progressive hardware capability stack that enables AI smart glasses to function as embodied intelligent assistants. Table 1 summarizes representative AI smart-glasses products and their hardware specifications along this capability stack.
2.1. Sensing Arrays
Sensing arrays serve as a gateway through which AI smart glasses perceive the physical world. Unlike traditional wearable devices that collect isolated physiological or motion signals, AI smart glasses continuously observe both the environment and the wearer from an egocentric perspective to support downstream reasoning and response. By capturing visual scenes, user motions, and surrounding spatial structures, sensing arrays establish the hardware basis for transforming raw sensory inputs into wearable intelligence.
Visual sensing provides the primary means of acquiring environmental information for smart glasses. Most commercial devices are equipped with ultra-wide cameras ranging from 12 to 16 megapixels, providing sufficient visual fidelity for object recognition, scene observation, and activity recording [20]. Continuous egocentric imaging enables AI assistants to observe what users see and establish a shared visual context for subsequent reasoning. Recent products further improve imaging quality through High Dynamic Range (HDR) imaging and low-light enhancement, preserving image fidelity under challenging illumination conditions [21]. Beyond environmental observation, some visual sensors support face recognition and virtual contact card generation, enabling social interaction assistance [22].
Complementing visual perception, Inertial Measurement Units (IMUs) enable smart glasses to detect the wearer’s movements. This dynamic information allows the system to stabilize captured imagery and virtual content while providing important cues about user behaviors and attention [23]. Motion sensing also supports low-latency interaction by detecting head gestures and body movements, thereby complementing interaction modalities [24]. As AI assistants become increasingly proactive, continuous motion awareness provides essential behavioral context for understanding user intentions and activities throughout daily life.
Spatial sensing further extends environment observation to a 3D scene. By integrating depth sensors, distance sensors, and magnetometers, modern smart-glasses estimate distances and calibrate object orientations within a scene, thereby constructing rich spatial signals [25]. The fusion of diverse signals enables AI assistants to understand the wearer’s surroundings, providing the foundational infrastructure for real-time spatial localization and navigation assistance.
2.2. Computing Platforms
Building on the multimodal observations captured by sensing arrays, computing platforms provide the processing foundation that enables AI smart glasses to interpret their surroundings and support user assistance. In practice, this processing foundation can be implemented through different computing modes, ranging from on-device computing for low-latency, privacy-sensitive inference to cloud-enhanced computing for scalable model capacity and service integration.
Modern smart glasses employ highly integrated on-device computing to enable real-time intelligence inference under stringent resource constraints. The System-on-Chip (SoC) integrates Central Processing Units (CPUs), Graphics Processing Units (GPUs), Neural Processing Units (NPUs), and Image Signal Processors (ISPs) into a unified platform, enabling efficient execution of sensor processing and neural inference workloads. Representative platforms such as the Qualcomm Snapdragon AR1 Gen 1 further optimize this integration for visual understanding, speech recognition, and multimodal feature extraction on edge devices [26]. The tightly coupled design reduces inference latency and improves energy efficiency, making it the dominant computing solution in recent commercial smart glasses.
To further trade off computational capability with energy efficiency, some systems employ hierarchical on-device computing architectures. Instead of executing all workloads on a single processor, these architectures separate always-on sensing from computationally intensive reasoning. An ultra-low-power microcontroller unit (MCU) continuously monitors wake words, inertial measurements, and other lightweight sensing signals, while the primary processor remains dormant until user intent or environmental events are detected. This event-driven scheduling reduces unnecessary computation and extends operating time without sacrificing responsiveness. Commercial products such as Xiaomi AI Glasses [27] and RayNeo V4 [25] have adopted dual-chip architectures that combine low-power control processors with high-performance AI chips, demonstrating that hierarchical computing resource management has become an increasingly practical strategy for continuous wearable intelligence.
Despite the development of on-device chips, the computational demands of multimodal foundation models still exceed the capabilities of current wearable on-device processors. Consequently, cloud-enhanced computing has become an indispensable component for smart glasses. AI assistants typically perform lightweight inference locally, while computationally intensive reasoning is delegated to cloud-hosted foundation models. For instance, Rokid AI Glasses support dynamic switching among multiple AI models, including DeepSeek, Qwen, and GLM AI, allowing users to select specialized reasoning capabilities for different tasks [28]. This collaborative architecture not only overcomes hardware computational limitations but also enables smart glasses to leverage rapidly evolving large multimodal models without frequent hardware upgrades.
2.3. Interaction Units
Interaction units provide the physical input and output channels through which AI smart glasses connect computational results with the wearer. While sensing arrays observe the physical world and computing platforms convert observations into task-relevant outputs, interaction units determine how user input can be captured and how system feedback can be presented under mobile and attention-constrained conditions. By combining acoustic, visual, and physical channels, they allow AI assistance to remain continuously accessible while reducing the cognitive and operational burden of device access.
Audio interaction has become the primary communication channel for AI smart glasses. Microphone arrays capture user speech for voice commands and conversation content, while open-ear speakers deliver AI-generated responses and multimedia notifications. Recent audio devices improve communication quality through beamforming, adaptive volume adjustment, far-field speech enhancement, and acoustic leakage suppression, allowing AI assistants to support natural conversations across diverse acoustic environments [29]. As large multimodal models reshape human-AI interaction, speech is increasingly evolving into the most intuitive interface between humans and wearable AI.
Visual feedback complements audio interaction by presenting information-rich content directly on the lenses. Rather than requiring users to shift their attention toward handheld devices, smart glasses overlay digital information onto the visual field, enabling glanceable access to navigation cues, translation subtitles, and contextual notifications. Commercial systems such as Meta Ray-Ban Display and INMO GO series employ Micro-LED displays and diffractive waveguide optics to achieve lightweight optical designs while maintaining sufficient brightness and readability for outdoor use [22,30]. Compared with transient voice responses, visual feedback allows AI assistants to communicate complex information continuously without interrupting ongoing activities, making it particularly valuable for navigation and multilingual communication assistance.
Physical interaction provides reliable and low-latency control when voice communication is unavailable or undesirable. A capacitive touchpad integrated into the temples supports intuitive gestures, such as tapping and swiping, for assistant activation and media control. Meanwhile, mechanical buttons provide reliable shortcuts for photography, recording, translation, and other latency-sensitive functions. These lightweight input mechanisms further complement audio interaction by enabling users to communicate explicit intentions in noisy environments or scenarios requiring rapid operation.
2.4. Supporting Infrastructure
Supporting infrastructure forms the foundational layer that enables AI smart glasses to function as persistent intelligent assistants in real-world environments. Without a reliable energy supply and data storage, the capabilities of observation, reasoning, and response cannot be sustained. Consequently, supporting infrastructure determines both the operating duration of AI services and the reliable storage of personal and contextual data under the stringent size, weight, and power constraints of wearable devices.
Among these supporting components, energy supply is the most immediate constraint. Despite integrating cameras, microphones, processors, and displays, most commercial smart glasses are powered by batteries with capacities of only 200-300mAh, representing less than one twentieth of the battery capacity of a typical smartphone. The limited energy budgets require every stage of sensing, computation, and communication to be carefully optimized for power efficiency. Rather than simply increasing battery capacity, recent devices improve service continuity through system-level energy management, including low-power computing architectures [27], event-driven task scheduling [25], and portable charging cases that enable additional fast-charging cycles [28]. They enable users to replenish energy opportunistically during daily routines, extending the effective operating duration of AI assistance from several hours to an entire day.
Personal data storage also plays a crucial role in enabling personalized intelligence throughout long-term human-AI interactions. Unlike conventional storage systems that primarily preserve application resources and multimedia files, the storage of AI smart glasses is a repository for user-specific knowledge, including local personalized AI models, personal preferences, interaction histories, and contextual records accumulated during daily use. By retaining and updating these personalized memories, AI assistants can build a deeper understanding of users and deliver context-aware, personalized support over time. Meanwhile, personal memory also strengthens privacy preservation. Sensitive information, such as facial features, voiceprints, and frequently accessed personal data, can be processed and retained locally without transmission to cloud servers, reducing privacy risks while maintaining continuous personalized services. As foundation models become increasingly capable of long-term memory and personalized reasoning, future smart glasses are expected to rely more heavily on on-device storage to preserve persistent user profiles and contextual knowledge across extended interactions. Consequently, storage systems are evolving from passive data repositories into hardware substrates that enable continuous personalization and long-term intelligent assistance.
Looking beyond current AI smart glasses, next-generation systems are expected to support more persistent, contextual, and proactive forms of wearable intelligence. The capabilities summarized in Figure 2, including predictive cognition, real-time perceptual augmentation, emotion-aware communication, persistent personal memory, and cross-sensory translation, illustrate how future assistance may extend from explicit responses to continuous situated support. From a hardware perspective, these capabilities intensify requirements across the entire capability stack: multimodal sensing must capture richer egocentric, behavioral, spatial, and social signals; heterogeneous computing must coordinate local inference and cloud-enhanced reasoning under latency and energy constraints; interaction units must deliver timely and controllable feedback without overwhelming attention; and supporting infrastructure must sustain long-term operation and personal data management. Therefore, the hardware foundation of AI smart glasses should be understood not only as a collection of components, but as the enabling substrate that bounds how wearable intelligence can be realized in daily environments.
3. Wearable Intelligence
my-box=[ rectangle, draw=DarkBlue, rounded corners, text opacity=1, minimum height=1.5em, minimum width=5em, inner sep=2pt, align=center, fill opacity=.5, ] leaf=[ my-box, fill=yellow!32, text=black, align=center, font=, inner xsep=5pt, inner ysep=4pt, text width=38em, ] leaf2=[ leaf, fill=purple!20, ] leaf3=[ leaf, fill=LightBlue, ]
Unlike traditional AI systems that primarily perceive the world from third-person observations, smart glasses require a continuous egocentric perspective to provide actionable assistance within the constraints of wearable devices. This capability is referred to as Wearable Intelligence, which describes how AI smart glasses progressively convert egocentric observations into intelligent assistance.
In this survey, we organize wearable intelligence into three hierarchical levels, as illustrated in Figure 4. Perception Intelligence(§Section 3.1) extracts directly observable information from egocentric multimodal streams, including spatial structures, temporal dynamics, semantic concepts, and social interactions. Building upon these basic perceptual representations, Contextual Intelligence(§Section 3.2) enriches observations with the user’s personal history, surrounding environments, temporal continuity, and external knowledge, enabling context-aware understanding. Finally, Agentic Intelligence(§Section 3.3) transforms contextual understanding into goal-directed assistance through intent reasoning, task planning, and action execution. The three levels form a progressive intelligence hierarchy that moves from observation to understanding and ultimately to assistance, illustrating how AI smart glasses evolve from passive sensing devices into proactive wearable assistants.
3.1. Perception Intelligence
Perception intelligence enables smart glasses to extract information directly from egocentric observations. As the entry point to the intelligence hierarchy, perception intelligence transforms raw sensory signals into structured representations that support subsequent contextual understanding and agentic decision-making. In this survey, we organize perception intelligence into four interconnected capabilities, as illustrated in Figure 5. Spatial perception (§Section 3.1.1) captures geometric structures and spatial relationships within the physical world. Temporal perception (§Section 3.1.2) characterizes dynamic changes and future evolution from continuous egocentric observations. Conceptual perception (§Section 3.1.3) recognizes meaningful entities and behaviors from egocentric experiences. Finally, social perception (§Section 3.1.4) understands interpersonal attention and social interactions in human-centered environments. Together, these capabilities provide complementary perspectives on the observable world, enabling smart glasses to answer where things are, how the world evolves, what things mean, and who people interact with, providing the perceptual foundation for AI smart glasses.
3.1.1. Spatial Perception
Spatial perception in smart glasses refers to the ability to understand the geometric structures of objects and their spatial relationships from an egocentric perspective. Starting with wearer-centric sensing and localization, smart glasses further extend spatial perception of the surrounding environment through egocentric observations, enabling the system to understand the spatial characteristics of nearby objects and to model the spatial relationships and physical interactions between the user and the physical world. Based on these characteristics, this capability can be categorized into two closely related perspectives. The first is wearer-centric spatial perception, which estimates the user’s own body states and poses. The second is environment-centric spatial perception, which enables smart glasses to perceive the spatial distribution of nearby objects and the user’s position in dynamic environments.
Wearer-centric spatial perception provides the geometric foundation for understanding user behaviors in smart glasses. Early studies primarily explored single-modality solutions that relied on user visual inputs for pose estimation [62]. EgoEgo proposed a generic pose estimation framework based on head motion, demonstrating the feasibility of recovering full-body motion from egocentric observations [63]. However, image-based approaches often suffer from instability in real-world smart glasses scenarios due to severe self-occlusion, rapid ego-motion, and continuous viewpoint transitions.
To improve robustness under continuous wearable sensing conditions, recent research increasingly shifts toward multimodal spatial perception by integrating smart glasses with body-worn inertial measurement units (IMUs), enabling complementary fusion between visual observations and motion signals [31,64]. Recent progress has also been significantly accelerated by the emergence of large-scale multimodal egocentric datasets. Representative datasets such as Ego-Exo4D [32], Nymeria [65], and SimXR [66] provide synchronized egocentric videos, eye tracking, IMU signals, and full-body motion annotations captured using Project Aria glasses. EMHI further demonstrates that cross-modal fusion between stereo vision and wearable IMU signals can effectively mitigate self-occlusion and sensor drift during smart glasses deployment [67]. Building upon these multimodal resources, recent systems increasingly explore closed-loop embodied perception frameworks. Ego4o supports arbitrary combinations of sparse IMUs, egocentric visual streams, and motion descriptions, enabling mutual enhancement between spatial perception and embodied understanding [68].
Environment-centric spatial perception enables smart glasses to focus on geometric properties of nearby objects and the user’s position during continuous daily interactions. One important direction is object localization, where smart glasses identify and localize target objects within the current scene. This capability is particularly important for everyday wearable assistance, such as helping users quickly locate misplaced items during daily activities. Early object localization methods mainly relied on feature matching and region proposal mechanisms, which often struggled under severe viewpoint variations and cluttered environments of smart glasses. CoCoFormer enhances localization accuracy by constructing viewpoint transformations in contrastive learning [69]. PRVQL further introduces a feature knowledge base to alleviate background interference caused by cluttered backgrounds [70]. More recently, RELOCATE shifts toward pretrained representation-driven spatial reasoning to improve the robustness of smart glasses under real-world scenarios [71]. Beyond 2D localization, EAGLE unifies 2D and 3D localization under egocentric visual settings, enabling more consistent spatial grounding for smart glasses [33].
Another important direction is place localization, where smart glasses estimate the user’s location by matching observations against reference images in a database. Reliable place localization is essential for AI smart glasses that aim to provide continuous navigation assistance and environment-aware services across indoor and outdoor scenarios. R2Former combine global scene representations with relational reasoning mechanisms, enabling smart glasses to maintain stable environmental awareness during continuous daily activities [72]. SemVPR further improves the efficiency of wearable place localization, facilitating more practical real-time deployment for continuous navigation and environment-aware assistance in smart glasses [73].
3.1.2. Temporal Perception
Beyond 3D spatial structure, the real world also evolves along the temporal dimension, which encodes information about the past, present, and potential future states of objects and environments. For AI smart glasses, temporal perception refers to the capability of continuously understanding dynamic changes and performing predictive reasoning from egocentric streams over time. Such capability enables systems to provide real-time assistance and proactive warning during daily activities. Recent research on temporal perception for AI smart glasses can be broadly categorized into two closely related directions: present object tracking and future trajectory prediction in egocentric vision.
Egocentric object tracking is the foundation of understanding dynamic changes. By continuously tracking objects in egocentric streams, smart glasses can identify which objects are moving and maintain the temporal consistency of user-scene interactions over time. Compared with third-person tracking tasks, object tracking in smart glasses is substantially more challenging due to head motion, rapid locomotion, and hand occlusions. To systematically evaluate these challenges, TREK-150 establishes a dedicated benchmark to measure temporal tracking capability under egocentric settings [74]. EgoTracks further extends the evaluation toward longer-duration egocentric videos and reveals the unique difficulties of long-term egocentric tracking [34]. These benchmark studies collectively demonstrate that temporal perception in smart glasses is not merely a viewpoint adaptation problem, but rather a complex combination of intermittent observations, drastic appearance changes, and persistent interaction-induced occlusions.
As smart glasses systems continue to evolve, object tracking in 3D environments can not only help users locate objects of interest but also provide location information for path planning and navigation. For example, IT3DEgo leverages camera poses and depth information to project 2D instances into 3D space, enabling persistent object tracking under real-world egocentric scenarios and providing more accurate spatial grounding for smart glasses applications [75].
Future trajectory prediction requires smart glasses to reason about future motion evolution from historical observations, representing a more advanced form of temporal perception. In crowded streets or complex manipulation scenarios, smart glasses need to anticipate pedestrian movements or object motion trends, allowing proactive assistance and hazard warnings. EgoNav perceives surrounding environments and generates multiple plausible future paths, addressing the uncertainty during wearable navigation planning [35]. For finer-grained human-object interaction scenarios, HOIMotion further extends trajectory prediction toward body parts interacting with manipulated objects, leveraging historical body poses and object states to predict human motion [76]. However, trajectory prediction in smart glasses faces unique noise problems, where the historical observations are not entirely reliable. EgoTraj-Bench addresses this issue by constructing a real-world benchmark, which reveals a substantial gap between idealized offline evaluation and real-world smart glasses deployment [36].
Forecasting the users’ hand motion plays a vital role in proactive interaction and intent understanding in smart glasses. Hand trajectories often directly reflect the next interaction goals. Ego4D introduces an egocentric hand trajectory prediction benchmark, requiring systems to continuously predict future 2D hand coordinates in image space [3]. To further extend this capability into real-world interaction understanding, EgoPAT3D learns future action reasoning by continuously forecasting hand interaction targets in 3D environments [77]. While trajectories may indicate targets, uncertain hand motions may imply hesitation or decision-making processes. Consequently, hand trajectory prediction in smart glasses is not merely a coordinate regression problem but also a critical window into future user intentions and interaction demands. USST achieves global temporal perception from 3D hand trajectories, aiming to infer user planning behaviors from future motion evolution [78]. MADiff further incorporates hand-scene interaction relationships together with motion signals of smart glasses to guide future action reasoning [79].
3.1.3. Conceptual Perception
Conceptual perception in smart glasses refers to the capability of interpreting the meaning of entities and behaviors from egocentric observations. Rather than focusing on geometry or dynamics, conceptual perception focuses on understanding what these observations represent. This capability enables AI smart glasses to recognize meaningful objects, interpret user behaviors, and anticipate future activities, providing the semantic foundation for higher-level contextual intelligence. According to different targets, conceptual perception in smart glasses can be broadly divided into daily entity understanding and human behavior understanding.
Daily entity understanding serves as the primary channel through which smart glasses associate visual objects with concepts. As one of the most fundamental semantic capabilities, it answers the question of what category this object belongs to, which is essential for wearable assistance scenarios such as object search, environment-aware interaction, and daily activity support. Large-scale egocentric datasets accelerate the development of conceptual perception in smart glasses systems. Ego4D provides a comprehensive benchmark for object understanding under real-world egocentric settings [3]. Early recognition methods mainly relied on closed-set learning paradigms that required predefined object categories[80,81]. However, it is difficult to generalize to diverse wearable environments, where users continuously encounter previously unseen objects during daily activities.
In recent years, vision-language pretraining has gained significant attention for its ability to leverage massive amounts of internet data [82,83,84,85]. EgoVLP and EgoVLPv2 have explored this paradigm in egocentric video understanding [37,38]. It enables systems to align egocentric observations with flexible descriptions rather than a fixed set of predefined categories, which improves the object recognition capability of AI smart glasses under open-world scenarios.
Understanding human behavior marks the beginning of user activity perception in smart glasses. Compared with object recognition, behavior understanding introduces stronger temporal and interaction dynamics into conceptual perception by answering what the user is doing. This capability directly determines whether smart glasses can provide context-aware, proactive assistance throughout daily activities. Early studies mainly adapted general action recognition models [86,87,88] to egocentric datasets such as Ego4D [3], EPIC-Kitchens-100 [89], and Assembly101 [90], enabling smart glasses systems to recognize user activities. These approaches were also generally limited to predefined action vocabularies, making them difficult to generalize to the diversity of real-world user behaviors and interaction patterns.
Smart glasses require open-world behavioral and state understanding beyond fixed labels, since human activities and affective cues continuously evolve across environments, tasks, and available sensing channels [91,92,93,94]. The strong zero-shot generalization capabilities of multimodal models have opened new opportunities for action understanding across diverse scenarios [95,96,97]. ActionCLIP demonstrates that semantic knowledge learned from multimodal models can be transferred to egocentric action recognition, allowing smart glasses to identify unseen actions without requiring additional task-specific annotations [98]. Beyond recognizing novel action categories, OAP explicitly models the relationship between user actions and manipulated objects, improving the semantic generalization of wearable systems across diverse interaction scenarios [99].
Recent research extends understanding of behavior from recognizing current actions to anticipating future activities and interaction intentions [100,101,102]. Human behaviors are rarely isolated semantic units, but instead form temporally correlated event sequences. By semantically predicting what users are likely to do next, smart glasses can establish the foundation for task planning and anticipatory interaction. AntGPT leverages LLMs to perform goal-oriented reasoning over long-term action representation, enabling prediction of future user behaviors from egocentric activity streams [39]. VLog further expands the semantic vocabulary of novel events encountered during inference, providing a new perspective for open-world behavior anticipation in smart glasses systems [103].
3.1.4. Social Perception
Social perception refers to the ability to understand interpersonal signals from egocentric observations. As humans are social beings, many important cues in daily life are not conveyed through physical objects or the environment, but through interactions among individuals. Consequently, social perception captures the interpersonal dimension of the observable world, enabling smart glasses to understand human-centered social environments. This capability is particularly important for AI smart glasses that aim to provide effective social assistance and communication support in real-world settings. According to the social dimension, social perception in smart glasses can be broadly divided into two related directions: self-oriented social perception and other-oriented social perception.
Estimating the wearer’s attention serves as the foundation of self-oriented social perception. By inferring the user’s gaze direction, smart glasses can identify which objects or individuals currently attract the user’s attention, providing important clues about the user’s intentions and social engagement. Accurately estimating user attention remains challenging in practical applications. Therefore, some studies exploit multiple supplementary cues beyond gaze signals alone. MCN [40] demonstrates that hand actions and gaze behaviors exhibit strong correlations, showing that modeling action distributions can significantly improve attention prediction performance. MVGT [104] further captures both eye movements and attended targets by leveraging multi-view observations [105,106], improving gaze estimation accuracy under dynamic interaction scenarios. More recently, the integration of eye-tracking sensors into modern smart glasses enables direct estimation of gaze directions and visual attention distributions, providing increasingly reliable measurements of user attention in natural environments [107,108].
The social attention estimation plays a critical role in conversational assistance, where smart glasses need to determine the intended speaker among multiple participants. By grounding user attention to specific partners, wearable systems can selectively amplify relevant speech while suppressing background noise and competing voices, thereby improving communication quality in crowded social environments [109,110]. SAAL pioneers this direction by jointly leveraging egocentric videos and multichannel audio signals to infer socially attended speakers in multi-party conversations [41]. Given the dynamic nature of social interactions, recent studies further extend social attention estimation toward conversational focus understanding. For example, AV-CONV models the evolving interaction patterns among multiple active speakers, enabling smart glasses to continuously track shifts in conversational attention over time [42].
Other-oriented social perception treats each individual as an independent social agent, analyzing their behaviors and conversations. By modeling these interpersonal relationships, smart glasses are able to perceive the social structure of ongoing interactions and maintain awareness of multiple participants in social environments. To facilitate the study of social modeling, recent research has introduced a variety of benchmarks spanning dialogue understanding [43,111], social video question answering [112,113,114], and social behavior recognition [115]. These tasks evaluate whether systems can infer latent social relationships and interaction structures from continuous multimodal observations, significantly expanding the ability of smart glasses to perceive social signals in real-world settings. Recent advances in multimodal large language models (MLLMs) further improve the integration of visual and auditory social cues for smart glasses [116]. By jointly modeling visual observations and speech content, MLLM-based systems can more reliably identify conversational turns and interaction targets in social environments. Recent studies demonstrate that aligning attention across multiple participants significantly improves the robustness of social perception, enabling smart glasses to better understand social relationships in multi-person environments [117].
Perception intelligence establishes the foundation of AI smart glasses by transforming continuous egocentric observations into structured representations of the human-centered world. As the first layer of the intelligence hierarchy, it enables smart glasses to answer the fundamental question of what is directly observable. Despite substantial progress, current perception models are still often developed as isolated task-specific solutions, making it difficult to build unified representations that run continuously under wearable resource constraints. A central challenge is to maintain coherent, multimodal, and temporally continuous representations of the user’s surroundings while balancing accuracy, robustness, latency, and energy cost. Evaluation should therefore move beyond task-specific accuracy and consider perceptual continuity, multimodal consistency, computational efficiency, and robustness in real-world deployments.
3.2. Contextual Intelligence
Contextual intelligence enables smart glasses to interpret egocentric observations as user-specific, situated, temporally continuous, and knowledge-grounded evidence. It extends perception intelligence by asking not only what is visible or happening, but what the observation means for this wearer, in this environment, at this time, and relative to relevant external or personal knowledge. As illustrated in Figure 6, we organize this capability into four complementary forms: personalized context (§Section 3.2.1), which adapts interpretation to user histories and preferences; environmental context (§Section 3.2.2), which abstracts the surrounding scene into affordances and situational meanings; temporal context (§Section 3.2.3), which links the current moment to prior episodes, routines, and future needs; and knowledge context (§Section 3.2.4), which connects egocentric observations to external retrieval and personal memory.
3.2.1. Personalized Context
Personalized context incorporates user-specific information and preferences into contextual reasoning, enabling the same observation to be interpreted differently for individuals. Unlike perception intelligence, which focuses on universally observable signals, personalized context understanding requires the system to reason about the user’s unique habits, preferences, routines, and historical experiences. Such capability is essential for AI smart glasses to evolve from generic perception systems into personalized assistants that can provide user-specific support in daily life. According to their functional emphasis, existing research can be broadly categorized into three closely related directions: personalized memory, personalized question answering, and personalized recommendation.
Personalized memory organizes a user’s experiences, preferences, and behaviors into structured personal knowledge. By continuously accumulating and learning from this knowledge, smart glasses can move from generic understanding to user-specific meaning inferences. Yo’LLaVA directly injects personal visual information as additional contextual signals, enabling user-specific concept understanding and personalized visual reasoning [44]. MemoryBank introduces storage and retrieval modules for long-term personality modeling [118], while Generative Agents represent personal experiences in natural language and retrieve relevant memories to support behavioral prediction and planning [119]. These studies demonstrate the importance of memory as a contextual substrate for personalized intelligence in smart glasses. However, accumulating historical data often leads to memory redundancy and retrieval inefficiency. Consequently, recent research has shifted from memory storage toward memory organization and abstraction. AMEGO proposes an active memory framework that enables smart glasses to organize user interactions into structured representations in real time [50]. EgoSelf further distills user long-term behaviors into profiles for future interaction prediction [45]. These methods indicate a transition from passive memory recording toward active personal knowledge construction for smart glasses.
Personalized question answering transforms personal memories into an interactive retrieval interface with natural language. Rather than merely retrieving visually similar content, personalized question answering requires smart glasses to understand which experiences are relevant to a user’s query and why they matter. The Ego4D episodic memory benchmark first formalizes this problem by introducing natural language queries that require systems to identify temporal segments containing the requested information [3]. EgoLifeQA extends the task toward long-context question answering grounded in daily life experiences, covering scenarios such as recalling past events, monitoring health habits, and providing personalized assistance [13]. These benchmarks highlight the unique role of smart glasses as lifelong memory companions rather than conventional information retrieval systems [120,121]. Recent research increasingly emphasizes user-centric understanding with retrieval. MyEgo formulates personalized egocentric question answering, where systems model the wearer as the central actor and distinguish personal belongings and user activities [122]. Meanwhile, it is gradually evolving from visual matching toward intent-aware semantic reasoning. EgoRetriever introduces reflective chain-of-thought reasoning to infer user intentions, enabling smart glasses to actively understand what information users are seeking [123]. WAG further organizes multimodal wearable sensor data into personalized knowledge graphs to support answer generation [124]. Together, these studies demonstrate a broader shift from memory retrieval toward personalized contextual reasoning grounded in long-term user experiences.
Personalized recommendations [125,126,127] extend contextual understanding to proactively deliver information and suggestions. By combining current observations with long-term preference records, smart glasses can infer user interests and provide tailored assistance in real time. Typical examples include recommending menu items aligned with dietary preferences, highlighting news topics of personal interest, or suggesting actions based on historical behavioral patterns. The integration of large multimodal models into recommendation systems has substantially enhanced the ability of personalized assistants to generate user-specific suggestions [128,129,130,131]. Rather than relying solely on static preference profiles, RecICL demonstrates that in-context learning enables dynamic recommendation adaptation under continuously changing user preferences [132]. CAP-LLM further combines long-term user interests with real-world news content to help users efficiently discover personalized information [133]. Recent studies have also explored evaluation benchmarks for AI smart glasses. PerRecBench proposes a measure of the system’s inference ability of user preferences while reducing biases introduced by content quality and external ratings [134]. VoiceBench [135] and Mobile-Bench [136] further establish evaluation protocols for recommendation under different deployment scenarios. More recently, RecBench+ introduces realistic interactive environments and provides optimization strategies for improving recommendation performance [137].
3.2.2. Environmental Context
Environmental context interprets the user’s surrounding scene as a situated condition rather than a collection of visible objects. For AI smart glasses, this capability differs from environment-centric spatial perception and conceptual perception. Spatial perception estimates where objects and places are; conceptual perception identifies what objects and actions are present; environmental context infers what the surrounding situation means for the wearer. A kitchen counter with ingredients, a crowded transit platform, a store shelf, or a meeting room may contain recognizable entities, but the useful assistance depends on latent meanings such as activity setting, task affordance, safety risk, social appropriateness, and whether an interruption would be acceptable.
One line of work models the environment through functional zones and affordances. EGO-TOPO [46] learns topological maps from egocentric video, linking activity-centric zones with the likely actions they support. This is important for smart glasses because a place is not only a geometric container; it also encodes what the wearer can plausibly do next. A counter, doorway, workstation, or shelf becomes contextually meaningful when the system can associate it with preparation, navigation, repair, search, or comparison. Such affordance representations help the assistant move from recognizing objects to interpreting environmental opportunity.
A second line of work emphasizes a persistent environmental state. Smart glasses observe the world through a narrow, unstable egocentric field of view, so the current frame may omit relevant parts of the surrounding place. EgoEnv [47] addresses this limitation by learning human-centric environment representations that are predictive of the camera wearer’s potentially unseen local surroundings. The Aria Everyday Activities dataset [48] provides complementary smart-glasses evidence by aligning egocentric video, gaze, speech transcripts, trajectories, and 3D scene context in everyday environments. Together, these works suggest that environmental context should be maintained as a compact state of the surrounding situation, rather than recomputed from isolated frames.
Environmental context also requires relational and interaction-aware structure. Egocentric Action Scene Graphs [138] extend verb-noun action labels into temporally evolving graphs of actions, interacted objects, and their relationships, providing a richer substrate for long-form egocentric understanding. Grounding 3D Scene Affordance From Egocentric Interactions [139] further connects egocentric interaction videos with affordance regions in 3D scenes through the Ego-SAG framework and VSAD dataset. For smart glasses, these representations are useful because assistance often depends on where interaction is possible, which objects are functionally related, and how the user’s current manipulation changes the meaning of the surrounding scene.
The final step is to connect environmental meaning with the user’s implicit need. Visual Intention Grounding for Egocentric Assistants [140] shows that egocentric assistants must identify objects relevant to an intention, ignore unrelated contextual objects, and reason about uncommon object functionality. WearVQA [141] highlights the deployment difficulty: environmental reasoning must remain reliable under wearable image degradations such as blur, occlusion, poor lighting, unusual viewpoints, and clutter. AiGet [49] then demonstrates a smart-glasses pattern in which gaze, environmental context, and user profiles trigger low-disruption knowledge discovery during everyday moments. ContextAgent [142] and LifeEval [143] broaden this direction by evaluating whether sensory context can support proactive decisions and daily-life assistance.
3.2.3. Temporal Context
Temporal context situates the current egocentric observation within a longer continuity of life events. For AI smart glasses, this capability differs from temporal perception: perception estimates motion, tracks objects, or forecasts trajectories, whereas temporal context interprets what the current moment means relative to prior episodes, ongoing activities, recurring routines, and plausible future intentions. This distinction is important because wearable assistance often depends less on a single frame than on whether the user has already attempted a task, whether an event is recurring, whether the current interruption is timely, and how much historical evidence is needed before acting.
The first role of temporal context is to convert continuous wearable streams into activity episodes. The Aria Everyday Activities dataset [48] illustrates why this is a smart-glasses problem rather than a generic video problem: it records everyday scenarios with Project Aria glasses and synchronizes egocentric video, eye gaze, speech transcripts, user trajectories, and 3D scene context across shared locations. ActSonic [144] reaches the same goal through a different sensing route, using miniature speakers and microphones on eyeglasses to recognize everyday activities from inaudible acoustic reflections. Together, these systems show that temporal context is built from multimodal evidence distributed across time. Vision provides scene and object continuity, gaze indicates attention, speech gives task cues, trajectories indicate where the user is in an activity, and acoustic or inertial signals can preserve activity awareness when cameras are costly, occluded, or socially inappropriate.
The second role is episodic localization: the system must recover when a relevant event occurred and how it connects to the current query. Episodic Memory Question Answering [145] formalizes this requirement for egocentric augmented-reality assistants by asking a model to build a spatio-temporal scene memory and localize answers within a tour of a home environment. SpotEM [146] then addresses the efficiency problem that arises when this memory spans hours, using query-conditioned clip selection and low-cost semantic indexing so that an assistant does not exhaustively process every clip before answering. For smart glasses, this efficiency is not a secondary optimization. A temporal-context layer must support questions about prior moments while respecting battery, heat, storage, and latency limits.
The third role is to maintain a compact long-horizon state. AMEGO [50] constructs active memories from very long egocentric videos by retaining key locations and object interactions, and evaluates sequencing, concurrency, and temporal grounding queries over the Active Memories Benchmark. Online episodic memory extends this direction to streaming settings: ESOM [51] processes each frame once, discovers and tracks objects, and stores spatio-temporal object information in a compact memory for later visual queries. These works are especially relevant to glasses because the device cannot assume offline access to a complete video archive. Temporal context must be updated incrementally, decide what to keep, preserve enough ordering information to support later reasoning, and expose a compact state that downstream agents can use without replaying a full daily record.
The fourth role is to abstract evolving life context from repeated episodes. A temporal memory is useful only when it can distinguish an isolated event from a habit, a completed task from an unfinished one, and a stable routine from a recent change. EgoMemReason [52] highlights this gap by evaluating week-long egocentric video understanding through entity memory, event memory, and behavior memory, where evidence may be sparse and separated by long temporal intervals. For smart glasses, such behavior-level memory is the bridge from retrospective recall to prospective assistance: if the system knows that an object is usually used after a particular activity, that a route deviation is unusual, or that a repeated interaction has previously required help, it can reason about likely next needs before the user issues a fully specified command.
3.2.4. Knowledge Context
Knowledge context gives smart glasses access to information beyond the current sensor frame. For AI smart glasses, this capability has two complementary forms: retrieval from external knowledge sources and memory over the user’s own longitudinal experience. The former helps the device answer situated questions with evidence from documents, web pages, maps, manuals, or search results; the latter helps it reuse prior observations, locations, routines, and preferences. In both cases, knowledge is useful only after it is grounded in egocentric perception and delivered through low-friction wearable interaction.
Retrieval-augmented generation (RAG) is the main technical route for external knowledge grounding. Classical work in dense retrieval and RAG shows how a language model can retrieve evidence from a non-parametric corpus and generate answers conditioned on that evidence [147,148,149]. More recent RAG variants improve scale, adaptivity, hierarchical retrieval, graph-structured evidence, and long-context reading [150,151,152,153,154]. In smart glasses, however, RAG is not a standalone text pipeline. It must first convert gaze, speech, pointing, OCR, object detections, and dialogue history into a retrievable query. GazePointAR [54] illustrates this prerequisite by using gaze, pointing, and conversation history to resolve pronouns and object references in wearable AR before answering context-dependent questions.
Recent smart-glasses systems integrate RAG more directly into the assistant loop. QA-Dragon introduces query-aware domain and search routing to coordinate text and image retrieval agents, supporting multimodal, multi-turn, and multi-hop reasoning for knowledge-intensive visual question answering [155]. SUPERGLASSES [14] frames smart-glasses visual question answering as an agentic external-knowledge problem: its SUPERLENS pipeline decomposes the user query, detects relevant objects, searches multimodal web evidence, and reasons over retrieved results. CRAG-MM [156] extends this direction to multimodal, multi-turn RAG, emphasizing that wearable questions often require evidence across image, text, and dialogue turns rather than a single retrieved passage. WearVox [60] adds the audio side of the problem by evaluating search-grounded question answering and tool use from egocentric multichannel recordings collected with wearable devices. These works show that RAG for smart glasses is best understood as perceptually grounded retrieval: the system must decide what the user is referring to, retrieve evidence from the right source, and compress the answer into a form that does not overload attention.
External retrieval is also becoming proactive. AiGet [49] uses gaze, environmental context, and user profiles to surface knowledge during everyday moments rather than waiting for a fully specified query. Egocentric Co-Pilot [59] and VisionClaw [15] further connect egocentric perception with web-native tools and agentic execution, allowing retrieved knowledge to support navigation, document understanding, shopping, scheduling, and device control. This shift is important for AI smart glasses because the device is worn in situations where the user may not know what to ask, may not want to speak a long prompt, or may need help before a task failure occurs. The design problem is therefore not only retrieval accuracy, but also trigger timing, user interruption, and evidence transparency.
Memory provides the second form of knowledge context. General memory-augmented agents show how observations, reflections, and interaction histories can be stored and reused across sessions [118,119,157,158,159], while recent systems emphasize structured, evolving, and scalable memory for long-running agents [160,161]. Smart glasses make this problem concrete because they can continuously observe the user’s environment. MemX [162] is an early eyewear system that captures attention-relevant moments as compact visual memory, showing how sensing and energy constraints shape what should be stored. EgoLife [13] later introduces EgoRAG for long-context question answering over daily egocentric recordings, turning life logs into retrievable personal evidence for an egocentric assistant.
Wearable memory is not only a storage layer; it changes interaction. SpeechLess [163] binds prior interactions to spatial, temporal, activity, and referent context so that users can issue micro-utterances or under-specified requests in repeated situations. AI4Service [164] similarly uses a memory unit to personalize proactive service assistance with AI smart glasses. These systems suggest a memory pipeline for AI smart glasses: capture salient events, summarize and index them, retrieve relevant memories when a user asks or when the system detects an opportunity, and update or forget memories as context changes.
Taken together, personalized, environmental, temporal, and knowledge contexts show that AI smart glasses need a contextual substrate rather than isolated recognition modules. This substrate must decide which user history, scene affordance, past episode, or external evidence should condition assistance, and it must expose that state to interaction and agentic modules without overwhelming the wearer. The central challenge is selective contextualization: the system must preserve enough evidence to support continuity and proactive assistance, while giving users control over what is recorded, abstracted, retrieved, forgotten, or used to trigger intervention.
3.3. Agentic Intelligence
Agentic intelligence turns smart glasses from passive perception and question-answering devices into goal-directed assistants. As illustrated in Figure 7, we organize this capability into three connected stages: intent reasoning (§Section 3.3.1), which infers what the user wants and whether assistance is appropriate; task planning (§Section 3.3.2), which converts the inferred goal into executable steps under egocentric and wearable constraints; and action execution (§Section 3.3.3), which carries out those steps through tools, web services, connected devices, or situated guidance. This organization follows the agent loop from understanding to planning to acting, while keeping the smart-glasses form factor central.
3.3.1. Intent Reasoning
Intent reasoning asks what the user is trying to accomplish and whether assistance is appropriate. For smart glasses, intent is often implicit, brief, and situated in the user’s field of view rather than expressed as a complete instruction. The system must combine speech, gaze, gesture, dialogue history, activity state, environmental cues, and personal context to distinguish an explicit command from a passing observation, a curiosity cue, or a moment where interruption would be unwelcome. This differs from intent recognition in desktop assistants because the relevant evidence may be distributed across what the wearer sees, what the wearer is manipulating, where attention is directed, and what has happened moments or days earlier.
General language-agent work provides a useful abstraction but not a complete solution. ReAct [57] shows that reasoning and acting can be interleaved so that an agent updates its beliefs from observations while pursuing a goal. Embodied-agent work further argues that agents operating in physical environments need world models and user models rather than only text-level plans [165]. In smart glasses, this means that intent reasoning should estimate not only the semantic content of an utterance, but also the referent, urgency, confidence, social appropriateness, and permission boundary of the next action. A request such as a short deictic question, a glance at a product, or repeated hesitation near an appliance may require different levels of inference and different thresholds for proactive response.
Recent wearable systems show how intent reasoning begins with grounded reference. GazePointAR [54] resolves ambiguous pronouns by combining gaze, pointing, and dialogue context, making the user’s physical target interpretable before the assistant answers about an object. Gazeify Then Voiceify [166] extends this idea to displayless smart glasses by using gaze to select a physical object and voice to confirm or describe it, which is especially relevant when the device cannot rely on visual overlays for disambiguation. Persistent Assistant [167] studies intent grounding through gaze and gesture for repeated everyday decisions, emphasizing that efficient target specification and multimodal feedback can reduce physical demand in persistent assistant use. These systems suggest a design pattern in which smart glasses first stabilize the object, event, or person under discussion, then pass that grounded intent to retrieval, planning, or execution modules.
A second line of work broadens intent reasoning from explicit commands to proactive assistance. AiGet [49] uses gaze, environmental context, and user profiles to infer moments where background knowledge may be useful during everyday activities. ContextAgent [142] and WAGIBench [55] evaluate whether agents can infer service needs or high-level user goals from wearable sensory streams, personal context, and digital context. AI4Service [164] makes this proactive setting smart-glasses-specific by detecting service opportunities from context and personal memory. These systems shift the question from "what did the user ask?" to "is there a useful, timely, and acceptable action now?" and position intent reasoning as the trigger for downstream assistance.
Intent also affects system resource allocation. Intention-aware semantic communication for AI smart glasses [56] treats inferred task intent as a variable that determines which visual semantics should be preserved before offloading to a server-side model. This connects user modeling with systems design: if the agent knows whether the downstream goal is navigation, text reading, product lookup, or safety monitoring, it can compress, transmit, or discard different parts of the egocentric stream. Intent reasoning, therefore, acts as the entry point of the agentic loop: it grounds what the user means, estimates whether assistance should occur, and passes a confidence-aware goal representation to downstream planning.
3.3.2. Task Planning
Task planning converts inferred intent into executable steps. General agent frameworks provide useful abstractions: ReAct [57] interleaves reasoning traces with actions, Toolformer [168] studies when and how a language model should call external tools, and HuggingGPT [169] decomposes a user request into subtasks, selects specialized models, executes them, and summarizes the result. Hierarchical reasoning [170] provides another fundamental capability for task planning because complex goals must often be decomposed into subgoals, ordered operations, and tool-level actions. For smart glasses, this planning loop must additionally coordinate perception, retrieval, memory, interaction, and device constraints. A plan may require selecting a visual crop, reading text, querying the web, retrieving personal memory, asking a clarification question, generating step-by-step guidance, or deciding that the task is unsafe to automate.
The distinctive planning problem is that the agent’s state is incomplete and unstable. A desktop agent usually plans from a relatively persistent interface state, whereas a glasses agent plans from moving egocentric video, noisy audio, partial hand-object views, and attention signals that may change as the wearer turns their head. Planning, therefore, begins with state construction: the system must decide which frame, object, utterance, past event, or external document should become part of the working state. Ego4D [3] provides broad egocentric task substrates such as episodic memory, forecasting, and hand-object interaction, showing why planning in egocentric settings depends on temporal evidence and future action prediction rather than isolated image recognition. MM-Ego [171] further contributes to long egocentric video understanding and memory pointer prompting, which is useful when an agent must plan from an extended visual history rather than a single frame.
Egocentric planning benchmarks make this problem more explicit. EgoPlan-Bench [58] evaluates whether multimodal models can plan human-level actions from egocentric observations, while EgoPlan-Bench2 [172] expands this evaluation to broader real-world daily and work scenarios. EgoToM [173] evaluates goal, belief, and future-action reasoning from egocentric videos, highlighting that planning requires modeling what the user or nearby people are likely to know and do next. LifeEval [143] moves toward assistive daily-life evaluation, where the agent must combine multimodal perception, contextual reasoning, and task-oriented assistance across everyday scenarios. Together, these benchmarks suggest that smart-glasses planning should be evaluated at the level of goal completion and recovery, not only at the level of next-token explanations or single-frame answers.
Smart-glasses systems then connect planning ability to deployed assistance. SUPERGLASSES [14] decomposes visual questions into object identification, multimodal web search, and evidence-based reasoning. Egocentric Co-Pilot [59] compresses long-horizon egocentric context and coordinates perception, reasoning, and web tools. Vinci [174] presents a real-time embodied smart assistant based on egocentric vision-language modeling, illustrating how planning must operate within short interaction cycles rather than offline video analysis. These systems show a recurring architecture: construct a compact egocentric state, decompose the user’s goal, select tools or models, monitor intermediate outputs, and return either an action or a clarification request.
3.3.3. Action Execution
Action execution closes the agent loop by carrying out the planned step through digital tools, connected devices, or user-facing guidance. General agent research treats execution as a problem of reliable tool selection, command generation, and environment interaction [175,176,177,178]. WebGPT [179] demonstrates browser-assisted question answering through search and navigation commands; Toolformer [168] and Gorilla [180] study how language models can select and call external APIs; Mind2Web [181] and WebArena [182] evaluate web agents that must choose interface actions over realistic websites. These works are useful foundations, but smart glasses add a harder grounding problem: the action target may be a physical object, a spoken request, a bystander-sensitive scene, or a remembered context rather than a clean browser state.
Execution in smart glasses, therefore, begins before the tool call. The agent must verify that the perceived target and planned operation are sufficiently grounded. WearVQA [141] illustrates the perception side of this boundary by evaluating visual question answering under authentic wearable image conditions, where motion blur, occlusion, unusual viewpoints, and low-quality frames can degrade downstream decisions. WearVox [60] captures the audio side by evaluating tool calling, speech translation, and search-grounded question answering under egocentric multichannel audio conditions. These benchmarks show that execution reliability depends on the quality of the sensory channel that triggered the action. A mistaken object reference, side conversation, or noisy transcription can lead the agent to call the right tool for the wrong situation.
Recent smart-glasses agents connect this grounded state to concrete services. SUPERGLASSES [14] executes an external-knowledge pipeline by coupling object detection, query decomposition, multimodal web search, and reasoning over retrieved evidence. Ego2Web [61] adds a web-agent view: it pairs egocentric video with online tasks, requiring an agent to understand the user’s surroundings and complete a related web workflow. Egocentric Co-Pilot [59] similarly emphasizes web-mediated action, where egocentric perception is connected to online services and assistive workflows. VisionClaw [15] presents an always-on smart-glasses agent that combines live egocentric perception with OpenClaw agents to support shopping, document note-taking, meeting assistance, event creation, and device control. These systems show that execution for AI smart glasses is not just API calls; it is a controlled handoff from situated perception to external action.
Execution also includes the delivery of user-facing guidance when the planned step cannot be completed entirely through digital tools or connected devices. In smart glasses, such guidance must remain grounded in the perceived object, location, or activity state, so that instructions, warnings, or confirmations refer to the correct situation. The detailed design of feedback channels and repair cues is treated as an interaction-design problem, but agentic execution must still produce outputs that are concise, grounded, and safe to act on.
Agentic intelligence integrates intent reasoning, task planning, and action execution into a closed loop for AI smart glasses: intent reasoning grounds the user’s goal in egocentric perception and personal context, planning converts that goal into perceptual, retrieval, dialogue, and tool-use operations, and execution transfers the plan into services, devices, or situated feedback. The main tension is autonomy under wearable constraints: proactive and direct execution can reduce effort, but they also increase the risk of distraction, privacy exposure, inappropriate interruption, and unsafe action when intent or perception is uncertain. This tension requires calibrated intent estimates, uncertainty-aware planning, user-controllable proactivity, low-latency feedback, and auditable traces of what the agent perceived, retrieved, planned, and did, with evaluation focused on integrated task success, recovery, attention cost, privacy risk, and user trust in continuous real-world use.
4. Interaction Design
Interaction design is the user-facing layer through which AI smart glasses become usable as situated intelligent assistants in daily life. Unlike phones or desktop assistants, smart glasses operate during mobile, social, and attention-constrained activities. Users may issue short utterances, glance at objects, touch the frame, perform subtle hand gestures, or expect assistance without opening an explicit application. Therefore, interaction design for AI smart glasses is not simply the addition of input and output modalities. It is the control layer that captures intent, grounds references, selects feedback, regulates proactive intervention, and keeps assistance usable in public settings.
As shown in Figure 8, we organize interaction design into five connected functions. Intent Capture(§Section 4.1) detects engagement through speech, gaze, touch, hand gestures, and other cues. Reference Grounding(§Section 4.2) resolves intended referents through situated evidence and memory. Feedback Design(§Section 4.3) combines channel selection with repair-oriented feedback for confirmation, correction, and trust calibration. Proactive Interaction(§Section 4.4) couples trigger modeling with intervention decisions. Inclusive Usability(§Section 4.5) treats user diversity, accessibility, privacy, transparency, and control as cross-cutting constraints. Earlier surveys mainly characterized smart-glasses interaction as micro-interactions and device control [4]; recent AI-mediated systems instead treat interaction as a closed loop among multimodal sensing, contextual reasoning, user correction, and situated assistance.
4.1. Intent Capture
Intent capture concerns how smart glasses infer that the wearer wants to engage the system and what interaction form that intention should take. Speech is a natural starting point because it supports hands-free control, but it is vulnerable to noise, socially exposed in public, and ambiguous when nearby conversations occur. WearVox [60] captures this difficulty by evaluating wearable voice-assistant tasks such as search-grounded question answering, tool calling, side-talk rejection, and speech translation. EchoSpeech [183] complements audible speech with eyewear-mounted acoustic sensing for silent speech recognition, reducing dependence on vocalized commands in public or noisy environments. These works show that speech-based intent capture must address not only recognition accuracy, but also addressability and social exposure.
Gaze provides a second intent channel because smart glasses are naturally aligned with the wearer’s visual attention. GazeTrak [184] explores acoustic gaze estimation on a glasses frame, while EyeGesener [185] targets intentional eye gestures and filters daily eye movements to reduce false activation. Together, these systems show how gaze and eye motion can support low-friction intent capture, while also raising calibration and reliability challenges.
Touch and hand gestures provide more deliberate forms of control and confirmation. FingerGlass [186] uses a frame-mounted fingerprint sensor to support finger-aware touch-based interaction, while Helios [187] studies low-power event-based gesture recognition for always-on smart eyewear. GlassMessaging [188] further shows that everyday smart-glasses interaction often combines voice and manual input during mobile activities. Taken together, these systems suggest that speech, gaze, touch, hand gestures, and other input channels should be coordinated adaptively rather than treated as independent interaction choices.
4.2. Reference Grounding
Reference grounding connects a captured intention to the place, object, event, or prior experience that the wearer means. This is a core smart-glasses problem because wearable commands are often short and deictic. Grounding spans both the current scene and prior experience: situated grounding resolves references in the immediate environment, whereas memory grounding links an utterance or gesture to earlier interaction history. Without such grounding, accurate speech recognition still produces an incomplete query because the system cannot identify what should be passed to perception, reasoning, or agentic planning.
Situated grounding addresses references to visible objects and ongoing activities. GazePointAR [54] demonstrates this requirement by combining gaze, pointing, voice, and conversation history to resolve ambiguous pronouns in wearable augmented reality. Its significance lies not only in multimodal fusion but also in showing that reference resolution is a prerequisite for assistance in physical space. Gazeify Then Voiceify [166] extends this principle to displayless smart glasses, where gaze selects a physical object, and voice confirms or describes it without visual overlays. This design is important for audio-first glasses, where the system cannot rely on displaying a highlighted target and must instead support spoken confirmation and repair.
Memory grounding expands reference resolution beyond the current frame. SpeechLess [163] uses personalized spatial memory to support micro-utterances in everyday augmented reality, reducing the need for long public commands when repeated places, objects, and activities are already known. Persistent Assistant [167] similarly treats grounding as an ongoing process supported by embodied input and multimodal feedback. These systems suggest that reference grounding in AI smart glasses should align situated evidence with longer-term memory, rather than treating each utterance as a one-shot mapping from language to command.
4.3. Feedback Design
Feedback design determines how smart glasses return information, request clarification, and expose uncertainty without overloading the wearer. For this form factor, the central question is how to select an appropriate feedback channel while also supporting repair-oriented feedback for confirmation, correction, and trust calibration. Smart glasses must provide useful feedback without pulling excessive attention away from the surrounding world.
Feedback channel selection is constrained by the wearable form factor. Visual overlays can support spatial guidance but may occlude the scene; audio can preserve displayless use but may interfere with conversations or environmental awareness; haptics and subtle notifications can reduce visual load but carry limited information. Accessibility-oriented navigation illustrates this trade-off: visual and audio wayfinding guidance on smart glasses for people with low vision affects not only navigation performance, but also attention, confidence, and the coordination between assistance and residual vision [189]. GlassMessaging [188] reaches a related conclusion in everyday communication, where head-mounted messaging can improve access during multitasking, but must balance speed, accuracy, visual demand, and the appropriateness of voice input.
Repair-oriented feedback becomes necessary when AI assistance is grounded, uncertain, or awaiting user confirmation. Persistent Assistant [167] treats multimodal feedback as part of the intent-grounding loop rather than a final display stage. Gazeify Then Voiceify [166] shows the same principle in displayless interaction, where gaze-grounded selection is converted into voice-mediated confirmation. Feedback for AI smart glasses should therefore operate as a lightweight repair channel between perception and user control. It should provide enough evidence for correction and trust calibration, while avoiding explanations that compete with ongoing activity.
4.4. Proactive Interaction
Proactive interaction occurs when smart glasses initiate, compress, or reshape assistance from context rather than waiting for a fully specified command. Its design depends on trigger modeling, which detects when assistance may be relevant, and intervention decision, which determines whether and how the system should act. Proactivity can reduce interaction effort, but poorly timed prompts can distract the wearer, expose private context, or make the device socially intrusive.
Trigger modeling uses situated and personal context to identify moments where assistance may be useful. AiGet [49] exemplifies this direction by using gaze, environmental context, and user profiles to surface knowledge during everyday moments. SpeechLess [163] approaches the same problem through spatial memory, allowing short utterances to trigger contextually appropriate assistance. In both cases, context reduces interaction effort only when the system can infer relevance from the wearer’s activity, surroundings, and prior experience.
Intervention decision governs whether a detected opportunity should become an action, a clarification request, or no interruption at all. AI4Service [164] uses context and memory to provide proactive service assistance with AI smart glasses, while Persistent Assistant [167] studies persistent everyday assistance through intent grounding and feedback. In this section, these systems are best understood as interaction designs for timing, interruption management, confirmation, and user control, rather than as planning architectures. Future systems should condition proactive behavior on confidence, urgency, privacy sensitivity, and social context, with the goal of deciding when to remain silent, when to ask for confirmation, and when immediate feedback is justified.
4.5. Inclusive Usability
Inclusive usability evaluates whether interaction remains practical across users, abilities, environments, and social settings. Smart glasses are worn on the face, near the eyes and ears, and often in the presence of bystanders, so interaction choices have consequences beyond task efficiency. The main constraints are user diversity and accessibility, which determine whether input and feedback remain usable across abilities and contexts, together with privacy, transparency, and control, which determine whether sensing and AI assistance remain understandable and interruptible.
User diversity and accessibility require interaction techniques to be evaluated against real user capabilities rather than aggregate performance alone. Low-vision wayfinding demonstrates that smart-glasses guidance should account for sensory ability, navigation context, attention, and confidence [189]. Interaction design work for AI smart glasses and older workers similarly emphasizes embodied cognition, cognitive load, physical comfort, and socio-emotional needs, showing that efficient assistance can still fail when age-related usability requirements are ignored [190]. Input systems such as EchoSpeech [183] and FingerGlass [186] can reduce dependence on public speech or large gestures, but they also introduce calibration, learnability, and fatigue constraints.
Privacy, transparency, and control cut across all modalities. Vision-language model interactions on smart glasses can expose bystanders, sensitive documents, or private scenes, especially when image capture and cloud processing are hidden behind a simple voice query [191]. This makes consent, legibility, and user control central to interaction design. Interfaces should help wearers and nearby people understand when sensing is active, what information is used, and how assistance can be paused, corrected, deleted, or bypassed.
Taken together, intent capture, reference grounding, feedback design, proactive interaction, and inclusive usability define the interaction layer of AI smart glasses. Table 2 summarizes representative systems across this interaction loop. Their shared challenge is not maximizing modality richness, but coordinating input, grounding, feedback, intervention, and user control under wearable constraints. Future evaluation should therefore move beyond recognition accuracy or task completion to assess whether assistance is addressable, grounded, repairable, well-timed, accessible, privacy-aware, and socially acceptable during continuous real-world use.
5. Application Scenarios
Application scenarios test whether smart glasses can turn egocentric sensing into timely, situated assistance rather than merely recording or displaying information. Across healthcare, accessibility, situated learning, daily life assistance, cultural tourism, and industrial support, the same intelligence stack must be adapted to different risks: clinical accountability, assistive agency, learning transfer, personalization, cultural fidelity, and workplace safety. Figure 9 illustrates a representative workflow for applying AI smart glasses across these domains. The following subsections examine each domain in turn, drawing on representative systems, use scenarios, and deployment concerns to show how the wearable-intelligence capabilities in §Section 3 must be coordinated with the interaction principles in §Section 4 under domain-specific constraints.
5.1. Healthcare
Healthcare is a demanding domain for AI smart glasses because clinical assistance must be delivered while practitioners keep their hands free, protect sensitive information, and make time-critical judgments. Compared with phones or wall-mounted displays, smart glasses preserve the clinician’s egocentric view while making patient information, procedural guidance, remote expertise, and alerts available at the point of care. A recent systematic review shows that AI-powered smart glasses are being explored across care delivery, monitoring, telemedicine, and personalized intervention, but clinical use remains constrained by privacy, validation, standardization, battery life, engagement, and medical ethics [17]. Healthcare, therefore, makes the survey’s broader argument concrete: sensing, contextual intelligence, interaction, and deployment constraints must be designed as one system.
Emergency care illustrates this integration. At an incident scene, an emergency medical technician must treat the patient while forming a rapid account of symptoms, measurements, bystander reports, and triage requirements. EMSGlass [192] models this setting as a multimodal smart-glasses system: EMSNet combines scene evidence, text, and vital signs to infer the incident state, while EMSServe supports low-latency field deployment. The workflow relies on conceptual perception to recognize visible clinical and environmental cues, contextual intelligence to relate those cues to the patient’s medical state, and agentic intelligence to translate the inference into next-step guidance. It also depends on the interaction loop in §Section 4: feedback must be concise, interruptible, and urgency-aware, so that questions, confirmations, risk cues, and protocol reminders support treatment rather than divert attention from the patient.
Other healthcare work broadens the setting without changing the central requirement for situated assistance. In nursing education, smart glasses make clinical procedures visible from the learner’s point of view, supporting demonstration, remote supervision, and reflection [193]. Adoption studies further show that usefulness is not only technical: acceptance depends on how the device changes contact between professionals and patients [194]. For older adults and long-term care, scoping reviews suggest that smart glasses may support independence and daily functioning, although evidence remains small and fragmented [195,196]. These cases highlight temporal context and memory because useful assistance often depends on routines, medication histories, repeated confusion, or deviations from expected behavior.
Healthcare shows both opportunity and caution. AI smart glasses can provide hands-free, context-aware support when clinicians, caregivers, or patients cannot shift attention to another device, but errors in perception, timing, command interpretation, or data handling can directly affect care quality and trust. Clinical deployment, therefore, requires validated perception and reasoning, auditable interaction histories, conservative proactivity, and feedback designs that preserve professional control while protecting patient privacy under realistic workflow pressure.
5.2. Accessibility
Accessibility is one of the clearest areas where smart glasses can become assistive infrastructure rather than convenience devices. For users whose visual, language, motor, or cognitive abilities make handheld interfaces difficult, the value lies in situated access: the device observes the same environment as the wearer and returns guidance without requiring a phone or a shift of attention. Early low-vision wayfinding work shows that visual and audio guidance on smart glasses changes how users allocate attention and coordinate residual vision with device feedback, not only whether they complete a route [189]. AI smart glasses extend this role from route display to situated interpretation, question answering, and support for independent participation in everyday tasks and social settings.
A representative workflow appears in LLM-based smart-glasses assistance for visually impaired users. A wearer entering an unfamiliar transit station may ask whether the platform entrance is nearby, relying on the glasses to interpret the egocentric view and produce a concise answer. Lee et al. [197] illustrates this connection between smart-glasses sensing and language-model assistance. In such a scenario, spatial perception grounds the surrounding layout, conceptual perception reads visible objects and text, and contextual intelligence filters these cues through the user’s mobility goal. Agentic intelligence then determines whether the system should answer directly, ask for clarification, or recommend a safer step. The response must remain brief, confidence-aware, and deliverable through audio or sparse visual cues so that guidance supports orientation instead of competing with environmental awareness.
Accessibility work beyond visual impairment shows why model capability is insufficient by itself. Co-design with people with aphasia found that smart glasses may support everyday access and communication, but the same study also showed how hands-free gestures, visual clutter, heavy form factors, and socially conspicuous interfaces can become barriers [198]. CollabLens, studied with blind and low-vision participants and sighted peers, shows that smart glasses can support inclusive collaboration by giving BLV users more flexible access to visual information while changing how sighted collaborators provide help [199]. Accessibility, therefore, includes both individual task completion and participation in shared social settings.
Assistive smart glasses should be judged by whether they increase independence and agency without adding new burdens. Reliable perception matters, but it must be paired with interaction designs that remain private, socially legible, repairable, and adaptable across users. The main risks are overconfident guidance, excessive interruption, inaccessible feedback, and assumptions that one modality or assistance style works for everyone.
5.3. Situated Learning
Situated learning is grounded in the context where knowledge is used, making it a natural fit for smart glasses: the device shares the learner’s egocentric view and can attach an explanation to the object or action currently being studied. Reviews of augmented reality smart glasses in industrial assembly show this value for procedural work, while also showing that deployment depends on usable authoring, robust tracking, readable feedback, and integration with production settings [18]. For AI smart glasses, the shift is from static overlays to situated instruction that recognizes attention, tracks task state, retrieves relevant knowledge, and decides whether guidance would help or distract.
Maintenance training for a less experienced operator illustrates the pattern. The worker approaches an unfamiliar machine, looks at the target component, and asks for the next repair step while keeping both hands near the equipment. MARMA [200], although developed as a mobile augmented-reality maintenance assistant rather than a glasses-only system, provides a useful procedural pattern for aligning recognized equipment with step-by-step guidance. On AI smart glasses, the assistant would keep instructions synchronized with the recognized component, task progress, and relevant safety knowledge. Voice can request clarification, gaze or head pose can ground references such as “this valve,” and short feedback can confirm the next action or warn about a skipped safety step without blocking the field of view.
The same design logic extends to classrooms, laboratories, and production support. Mixed-reality learning can attach abstract concepts to visible procedures and spatial demonstrations [201], while smart-classroom research emphasizes that AI support must remain aligned with pedagogy, teacher control, and institutional readiness [202,203]. In workplace settings, smart glasses are most useful when guidance and expert assistance are embedded in the work process rather than treated as separate displays [204]. Adoption and evaluation studies sharpen the same point: task fit and social setting shape acceptance [19], and faster task completion can still coincide with higher mental workload [205]. For heterogeneous users, interaction design must also account for embodied cognition, comfort, and socio-emotional needs [190].
AI smart glasses are most valuable for situated learning when they operate as contextual tutors rather than passive displays. Their effectiveness depends on whether guidance is synchronized with task state, learner attention, and pedagogical goals, while keeping errors, cognitive burden, interruption cost, and knowledge-base maintenance under control.
5.4. Daily Life Assistance
Daily life assistance is the broadest and least bounded application area for AI smart glasses. Unlike healthcare, accessibility, or situated learning, it usually begins from small opportunistic needs that arise while the user is already shopping, commuting, reading, meeting, or moving through familiar places. Smart glasses are well matched to this setting because they are aligned with the user’s gaze, speech, movement, and surrounding objects. The research challenge is not to support one predefined task, but to build a persistent assistant that can judge relevance, timing, and the appropriate use of personal or external knowledge.
Proactive knowledge discovery during shopping or commuting illustrates this role. AiGet [49] uses smart glasses to turn everyday moments into contextual knowledge opportunities: a wearer may glance at a product, poster, or unfamiliar place without issuing a full command, and the system surfaces a short suggestion only when the expected value outweighs the interruption cost. In this workflow, perception grounds the visible object or text, personal and external knowledge determine what might matter, and intent reasoning estimates whether the moment is appropriate for intervention. The user may refine, ignore, or expand the suggestion through short voice prompts, while gaze and dialogue history ground ambiguous references such as “this one” or “that sign.”
Other work expands daily life assistance from momentary suggestions to longitudinal memory and agentic services. EgoLife [13] combines long-duration daily-life recordings with EgoGPT and EgoRAG, showing why future glasses need retrieval over personal episodes rather than isolated frame understanding. SUPERGLASSES [14] emphasizes external-knowledge visual question answering, where the assistant must connect the object in view with multimodal evidence. VisionClaw [15] illustrates the action side by linking live perception to downstream tools for daily service tasks. Together, these systems move daily life assistance beyond recognition toward a loop of memory, retrieval, planning, and execution.
Daily life assistance highlights the promise and fragility of always-available wearable intelligence. The value lies in timely, personalized support across small everyday moments, but trust can erode quickly when suggestions are mistimed, memory is used inappropriately, or inferred context is difficult to inspect and correct. Useful systems must therefore keep assistance brief, socially acceptable, privacy-aware, and controllable across repeated daily use.
5.5. Cultural Tourism
In cultural tourism, smart glasses can turn a visit into a situated interpretation rather than a sequence of disconnected labels or guidebook entries. Early systems used overlays to connect exhibits, landmarks, routes, and reconstructions with the visitor’s field of view, reducing dependence on handheld guides. Adoption studies suggest that visitors accept such systems when the content is useful, the device remains comfortable, and the experience fits the social setting of tourism [206,207]. The central challenge is to move beyond predefined content: the glasses must infer what the visitor is attending to, retrieve reliable cultural knowledge, and deliver interpretation at the right level of detail.
A visitor exploring an unfamiliar heritage site illustrates this workflow. TouristicAR [208] delivered context-aware tourism content through smart glasses at UNESCO World Heritage sites in Malaysia, connecting physical landmarks with digital information. Litvak and Kuflik [209] similarly used visitor location and viewing orientation to identify points of interest and present corresponding explanations. In an AI smart-glasses version, the system would localize the visitor, recognize the attended landmark or artifact, retrieve the relevant cultural record, and decide whether the moment calls for a short explanation, a historical narrative, a comparison, or a route suggestion. Follow-up questions, language switching, and continued movement can then be grounded through gaze, dialogue history, and spatial context.
Cultural-tourism applications differ less in their display technology than in the kind of interpretation they require. A museum guide may need to compare artifacts or adapt the tour to a visitor’s language and interests, whereas a heritage-site assistant may need to explain damaged structures through archival evidence or reconstruction. Interactive exhibitions and scenic areas further shift the task from explanation to participation, route choice, accessibility support, or nearby services. Across these settings, the assistant must keep perception, localization, cultural knowledge, visitor intent, narrative generation, and navigation aligned within a continuous visit.
Cultural tourism requires AI smart glasses to balance situated interpretation with factual reliability and social fit. The assistant must adapt to visitor attention, movement, language, and interests, while preserving historical authenticity, distinguishing verified knowledge from generated explanation, and remaining robust to localization error, lighting variation, network instability, and crowded public settings.
5.6. Industrial Support
Industrial smart glasses are valuable when workers need information without taking their hands, eyes, or attention away from physical tasks. Early systems used augmented reality to place equipment status and operating guidance in the field of view. In livestock farming, for example, Caria et al. [210] used GlassUp F4 smart glasses to provide animal and production information during ongoing work, with later evaluation under real farming conditions [211]. The central challenge is to move from predefined retrieval and visualization toward work-context understanding, abnormal-state detection, and timely guidance.
A representative inspection workflow can be abstracted from these industrial settings. The glasses first identify the work object and relate it to nearby equipment, then use worksite evidence and operational knowledge to infer what the worker needs next. As the task unfolds, the system may highlight the relevant part, present the next instruction, warn about a hazardous condition, or connect the worker with a remote expert. Voice, gaze, or gesture can confirm results and report anomalies, while the system updates its interpretation and records the completed operation.
Existing industrial systems emphasize different stages of this workflow. In aquaculture, Xi et al. [212] used a smart headset to collect prawn-related observations during feeding-tray inspection and convert them into evidence about pond conditions. In shipbuilding, Kunkera et al. [213] connected augmented reality with ship models, digital twins, and production workflows so that workers could localize components and update assembly status in context. In mining and construction, Baek and Choi [214] used smart glasses and Bluetooth beacons to warn workers when equipment proximity became dangerous. These cases show that industrial support is not a single visualization task; the system must adapt its perception, knowledge source, and feedback timing to the risk structure of the workplace.
Industrial deployment leaves little room for fragile assistance. Because incorrect recognition, delayed warnings, or inappropriate procedural guidance may affect worker safety and production reliability, AI smart glasses must remain usable under obstruction, drift, lighting variation, noise, network instability, battery limits, comfort constraints, and legacy-system requirements. The key test is whether workers can rely on the system for timely guidance while retaining the ability to inspect, override, or correct its decisions.
6. Open Problems and Future Directions
Building on the preceding discussion of hardware foundation, wearable intelligence, interaction design, and application scenarios, this section discusses several promising future research directions for AI smart glasses, as illustrated in Figure 10.
6.1. Next-Generation Hardware
The future development of AI smart glasses relies not only on stronger perception and reasoning algorithms, but also on continuous advances in underlying hardware. Existing devices remain constrained by display quality, computing performance, storage capacity, battery life, thermal efficiency, and wearing comfort, making it difficult to support persistent, real-time, multimodal, and personalized intelligent services. Future research should explore lightweight full-color and high-resolution display modules, wider fields of view, and lower-power optical solutions for presenting navigation cues, environmental annotations, and spatially grounded interactions. Meanwhile, more powerful yet energy-efficient processors, neural processing units, and vision accelerators are needed to support on-device multimodal model inference. Larger memory and local storage capacities would further enable the preservation of personalized memories, interaction histories, multimodal contexts, and lightweight models. In addition, future devices should improve battery capacity, thermal management, camera and microphone quality, wireless connectivity, and overall form-factor design, while leveraging edge-cloud or smartphone-assisted computing to reduce the energy consumption and weight of the glasses themselves. Ultimately, next-generation smart glasses should achieve a systematic balance among computing capability, display quality, battery life, comfort, and cost, thereby enabling truly all-day wearable intelligent assistants.
6.2. Trustworthy Egocentric Intelligence
Existing research mainly focuses on enhancing the perception, reasoning, and interaction capabilities of AI smart glasses, while their trustworthiness remains largely underexplored. Trustworthy AI smart glasses should be fair, explainable, safe, robust, and privacy-preserving. Fairness requires smart glasses to avoid biased perception and decision-making across different users, scenarios, and social groups. For example, when assisting users in education, shopping, or healthcare-related scenarios, the system should not rely on biased assumptions about users’ gender, age, occupation, or cultural background. Explainability requires AI smart glasses to justify their observations, recommendations, and actions, helping users understand why certain objects are recognized, why certain suggestions are provided, and whether the system can be trusted in high-stakes scenarios. Furthermore, future research should also investigate safety and robustness under noisy egocentric inputs, dynamic environments, and adversarial instructions, as well as privacy protection for both users and bystanders captured by always-on sensors.
6.3. Lifelong Personalized Memory
AI smart glasses naturally have the potential to become long-term personal assistants, since they can continuously perceive users’ environments, behaviors, routines, and preferences. However, most existing systems still rely on short-term interaction contexts, making it difficult to provide persistent and personalized assistance over time. Lifelong personalized memory aims to enable AI smart glasses to remember useful user-specific information, such as daily habits, preferred interaction styles, frequently used objects, and long-term interests. The key challenge is deciding what should be stored, updated, retrieved, forgotten, or compressed. Future research should develop adaptive and controllable memory mechanisms that can support long-term personalization while allowing users to inspect, edit, and delete their memories.
6.4. Proactive Intelligence
Most current AI systems are reactive, responding only after users issue explicit queries or commands. In contrast, AI smart glasses provide a unique opportunity for proactive intelligence, where the system can infer user intent from real-time egocentric context and offer assistance before the user explicitly asks. For instance, when a user gazes at a product, the glasses may provide price comparison or personalized suggestions; when a user enters an unfamiliar environment, they may offer timely navigation cues; when a user attends a meeting, they may surface relevant notes or action items. However, excessive proactivity may become intrusive or distracting. Future research should therefore study when to intervene, what to provide, and how to balance proactive assistance with user control in AI smart glasses.
6.5. Embodied Foundation Models
Although existing foundation models have achieved strong performance in language, vision, and multimodal reasoning, they are not specifically designed for the embodied setting of AI smart glasses. Smart glasses require models to interpret the world from egocentric observation, reason over temporal dynamics, and interact with users in continuous perception-action loops. Embodied foundation models should go beyond static image-text understanding by modeling spatial relations, human activities, object affordances, user intentions, and real-world task progress. For example, when a user looks at an object, the system should not only recognize what it is, but also infer how it can be used and whether it is relevant to the current task. Future research should explore efficient embodied foundation models that support real-time perception, grounded reasoning, and lightweight deployment for AI smart glasses.
7. Conclusions
AI smart glasses are becoming an important platform for wearable intelligence because they combine egocentric perception, hands-free interaction, continuous availability, and intelligent reasoning during daily activities. However, research on AI smart glasses remains dispersed across device design, model capabilities, interaction techniques, and domain-specific deployments, making a systematic overview necessary. To bridge this gap, this survey reviewed AI smart glasses through a system-level taxonomy spanning hardware foundation, wearable intelligence, interaction design, and application scenarios, providing researchers with a structured understanding of how wearable assistance is built, exposed to users, and adapted to real-world contexts. Since AI smart glasses are still in the early stages of development, we also discussed current limitations and future research directions to make such systems more deployable, trustworthy, personalized, and proactive. We hope this survey helps clarify the design space of AI smart glasses and supports future research on wearable intelligence in real-world settings.
References
- Starner, T.E. Wearable computers: No longer science fiction. IEEE Pervasive Comput. 2002, 1, 86–88. [Google Scholar] [CrossRef]
- Patel, S.; Park, H.; Bonato, P.; Chan, L.; Rodgers, M.M. A review of wearable sensors and systems with application in rehabilitation. J. Neuroeng. Rehabil. 2012, 9, 21–21. [Google Scholar] [CrossRef] [PubMed]
- Grauman, K.; Westbury, A.; Byrne, E.; Chavis, Z.; Furnari, A.; Girdhar, R.; Hamburger, J.; Jiang, H.; Liu, M.; Liu, X.; et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp. 18995–19012. [Google Scholar]
- Lee, L.H.; Hui, P. Interaction Methods for Smart Glasses: A Survey. IEEE Access 2017, 6, 28712–28732. [Google Scholar] [CrossRef]
- Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual instruction tuning. Adv. Neural Inf. Process. Syst. 2023, 36, 34892–34916. [Google Scholar] [CrossRef]
- Yuan, Y.; Li, W.; Liu, J.; Tang, D.; Luo, X.; Qin, C.; Zhang, L.; Zhu, J. Osprey: Pixel understanding with visual instruction tuning. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 28202–28211. [Google Scholar]
- Yuan, X.; Zhou, L.; Sun, Z.; Zhou, Z.; Lan, J. Instruction-guided multi-granularity segmentation and captioning with large multimodal model. Proc. Proc. AAAI Conf. Artif. Intell. 2025, Vol. 39, 9725–9733. [Google Scholar] [CrossRef]
- Fan, W.; Ding, Y.; Ning, L.; Wang, S.; Li, H.; Yin, D.; Chua, T.S.; Li, Q. A survey on rag meeting llms: Towards retrieval-augmented large language models. In Proceedings of the Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining, 2024; pp. 6491–6501. [Google Scholar]
- Wang, S.; Liu, C.; Ding, Y.; Lin, S.; Ng, S.K.; Xin, X.; Fan, W. Mixture-of-Experts Knowledge Graph Retrieval-Augmented Generation for Multi-Agent LLM-based Recommendation. arXiv 2026, arXiv:2605.28175. [Google Scholar]
- Ning, L.; Liang, Z.; Jiang, Z.; Qu, H.; Ding, Y.; Fan, W.; Wei, X.y.; Lin, S.; Liu, H.; Yu, P.S.; et al. A survey of webagents: Towards next-generation ai agents for web automation with large foundation models. In Proceedings of the Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, 2025; pp. 6140–6150. [Google Scholar]
- Jiang, Z.; Chen, Y.; Pan, Y.; Hu, Z.; Fan, W.; Li, Q.; Wang, H.; Wang, J.; Ou, W. A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing. arXiv 2026, arXiv:2608.04625. [Google Scholar]
- Wu, P.; Chen, P.Q.; Li, X.; Fan, W.; Li, Q. Datamart-Agent: LLM-Driven Game-Theoretic Agent for Data Marketplace Modeling. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2026, 2026; pp. 32509–32531. [Google Scholar]
- Yang, J.; Liu, S.; Guo, H.; Dong, Y.; Zhang, X.; Zhang, S.; Wang, P.; Zhou, Z.; Xie, B.; Wang, Z.; et al. Egolife: Towards egocentric life assistant. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 28885–28900. [Google Scholar]
- Jiang, Z.; Yuan, X.; Qu, H.; Lin, S.; Liu, K.; Fan, W.; Qing, L. SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 2165–2175. [Google Scholar]
- Liu, X.; Lee, D.; Gonzalez, E.J.; Gonzalez-Franco, M.; Suzuki, R. VisionClaw: Always-On AI Agents through Smart Glasses. arXiv 2026, arXiv:2604.03486. [Google Scholar]
- Li, X.; Qiu, H.; Wang, L.; Zhang, H.; Qi, C.; Han, L.; Xiong, H.; Li, H. Challenges and trends in egocentric vision: A survey. Mach. Intell. Res. 2026, 23, 1–33. [Google Scholar] [CrossRef]
- Wang, B.; Zheng, Y.; Han, X.; Kong, L.; Xiao, G.; Xiao, Z.; Chen, S. A systematic literature review on integrating AI-powered smart glasses into digital health management for proactive healthcare solutions. npj Digit. Med. 2025, 8, 410. [Google Scholar] [CrossRef] [PubMed]
- Danielsson, O.; Holm, M.; Syberfeldt, A. Augmented reality smart glasses in industrial assembly: Current status and future challenges. J. Ind. Inf. Integr. 2020, 20, 100175. [Google Scholar] [CrossRef]
- Koutromanos, G.; Kazakou, G. Augmented reality smart glasses use and acceptance: A literature review. Comput. Educ. X Real. 2023, 2, 100028. [Google Scholar] [CrossRef]
- Meta-Platforms. Ray-Ban Meta smart glasses. 2025. Available online: https://www.meta.com/ai-glasses/ray-ban-meta.
- Huawei. Huawei AI glasses. 2026. Available online: https://consumer.huawei.com/cn/audio/ai-glasses.
- INMO. INMO Go3. 2025. Available online: https://inmolens.com/go3.
- XREAL. XREAL One Pro. 2024. Available online: https://www.xreal.com/cn/one-pro.
- Rokid. Rokid Max Pro. 2023. Available online: https://arstudio.rokid.com/profile.
- RayNeo. RayNeo X3 Pro. 2025. Available online: https://rayneo.cn/x3pro.html.
- Snapdragon. Snapdragon AR1 Gen 1. 2023. Available online: https://www.qualcomm.com/xr-vr-ar/products/ar-series/snapdragon-ar1-gen-1-platform.
- Xiaomi. Xiaomi AI glasses. 2025. Available online: https://www.mi.com/prod/xiaomi-ai-glasses.
- Rokid. Rokid AI Glasses. 2024. Available online: https://glasses.rokid.com/.
- RayNeo. RayNeo V4. 2026. Available online: https://rayneo.cn/v4.html.
- INMO. INMO Go2. 2024. Available online: https://inmolens.com/go2.
- Yi, X.; Zhou, Y.; Habermann, M.; Golyanik, V.; Pan, S.; Theobalt, C.; Xu, F. Egolocate: Real-time motion capture, localization, and mapping with sparse body-mounted sensors. ACM Trans. Graph. (TOG) 2023, 42, 1–17. [Google Scholar] [CrossRef]
- Grauman, K.; Westbury, A.; Torresani, L.; Kitani, K.; Malik, J.; Afouras, T.; Ashutosh, K.; Baiyya, V.; Bansal, S.; Boote, B.; et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 19383–19400. [Google Scholar]
- Cao, Y.; Liu, Y.; Wang, G.; Liu, Z.; Wang, K.; Zhang, X.; Yu, J.; Tu, X. EAGLE: Episodic Appearance-and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision. Proc. Proc. AAAI Conf. Artif. Intell. 2026, Vol. 40, 2634–2642. [Google Scholar] [CrossRef]
- Tang, H.; Liang, K.J.; Grauman, K.; Feiszli, M.; Wang, W. EgoTracks: A long-term egocentric visual object tracking dataset. Adv. Neural Inf. Process. Syst. 2023, 36, 75716–75739. [Google Scholar] [CrossRef]
- Wang, W.; Liu, C.K.; Kennedy, M., III. EgoNav: Egocentric scene-aware human trajectory prediction. arXiv 2024, arXiv:2403.19026. [Google Scholar]
- Liu, J.; Zhou, J.; Ye, K.; Lin, K.Y.; Wang, A.; Liang, J. EgoTraj-Bench: Towards Robust Trajectory Prediction Under Ego-view Noisy Observations. arXiv 2025, arXiv:2510.00405. [Google Scholar]
- Lin, K.Q.; Wang, J.; Soldan, M.; Wray, M.; Yan, R.; Xu, E.Z.; Gao, D.; Tu, R.C.; Zhao, W.; Kong, W.; et al. Egocentric video-language pretraining. Adv. Neural Inf. Process. Syst. 2022, 35, 7575–7586. [Google Scholar] [CrossRef]
- Pramanick, S.; Song, Y.; Nag, S.; Lin, K.Q.; Shah, H.; Shou, M.Z.; Chellappa, R.; Zhang, P. Egovlpv2: Egocentric video-language pre-training with fusion in the backbone. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023; pp. 5285–5297. [Google Scholar]
- Zhao, Q.; Wang, S.; Zhang, C.; Fu, C.; Do, M.Q.; Agarwal, N.; Lee, K.; Sun, C. Antgpt: Can large language models help long-term action anticipation from videos? Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 56677–56697. [Google Scholar]
- Huang, Y.; Cai, M.; Li, Z.; Lu, F.; Sato, Y. Mutual context network for jointly estimating egocentric gaze and action. IEEE Trans. Image Process. 2020, 29, 7795–7806. [Google Scholar] [CrossRef]
- Ryan, F.; Jiang, H.; Shukla, A.; Rehg, J.M.; Ithapu, V.K. Egocentric auditory attention localization in conversations. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp. 14663–14674. [Google Scholar]
- Jia, W.; Liu, M.; Jiang, H.; Ananthabhotla, I.; Rehg, J.M.; Ithapu, V.K.; Gao, R. The audio-visual conversational graph: From an egocentric-exocentric perspective. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 26396–26405. [Google Scholar]
- Lee, S.; Lai, B.; Ryan, F.; Boote, B.; Rehg, J.M. Modeling multimodal social interactions: new challenges and baselines with densely aligned representations. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 14585–14595. [Google Scholar]
- Nguyen, T.; Liu, H.; Li, Y.; Cai, M.; Ojha, U.; Lee, Y.J. Yo’llava: Your personalized language and vision assistant. Adv. Neural Inf. Process. Syst. 2024, 37, 40913–40951. [Google Scholar] [CrossRef]
- Wang, Y.; Xu, Y.; Li, X.; Hong, J.; Wang, Y.; Chen, C.W.; Zhu, W. EgoSelf: From Memory to Personalized Egocentric Assistant. arXiv 2026, arXiv:2604.19564. [Google Scholar]
- Nagarajan, T.; Li, Y.; Feichtenhofer, C.; Grauman, K. Ego-topo: Environment affordances from egocentric video. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020; pp. 163–172. [Google Scholar]
- Nagarajan, T.; Ramakrishnan, S.K.; Desai, R.; Hillis, J.; Grauman, K. EgoEnv: Human-centric environment representations from egocentric video. Adv. Neural Inf. Process. Syst. 2023, 36, 60130–60143. [Google Scholar] [CrossRef]
- Lv, Z.; Charron, N.; Moulon, P.; Gamino, A.; Peng, C.; Sweeney, C.; Miller, E.; Tang, H.; Meissner, J.; Dong, J.; et al. Aria Everyday Activities Dataset. arXiv 2024, arXiv:2402.13349. [Google Scholar]
- Cai, R.; Janaka, N.; Kim, H.; Chen, Y.; Zhao, S.; Huang, Y.; Hsu, D. Aiget: Transforming everyday moments into hidden knowledge discovery with ai assistance on smart glasses. In Proceedings of the Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025; pp. 1–26. [Google Scholar]
- Goletto, G.; Nagarajan, T.; Averta, G.; Damen, D. Amego: Active memory from long egocentric videos. In Proceedings of the European Conference on Computer Vision, 2024; Springer; pp. 92–110. [Google Scholar]
- Manigrasso, Z.; Dunnhofer, M.; Furnari, A.; Nottebaum, M.; Finocchiaro, A.; Marana, D.; Forte, R.; Farinella, G.M.; Micheloni, C. Online episodic memory visual query localization with egocentric streaming object memory. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2026; pp. 3951–3960. [Google Scholar]
- Wang, Z.; Zhang, Y.; Yu, S.; Zhang, C.; Zhao, Z.; Yoon, J.; Lee, H.; Bertasius, G.; Bansal, M. EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding. arXiv 2026, arXiv:2605.09874. [Google Scholar]
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.t.; Rocktäschel, T.; et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Adv. Neural Inf. Process. Syst. 2020, 33, 9459–9474. [Google Scholar]
- Lee, J.; Wang, J.; Brown, E.; Chu, L.; Rodriguez, S.S.; Froehlich, J.E. GazePointAR: A context-aware multimodal voice assistant for pronoun disambiguation in wearable augmented reality. In Proceedings of the Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 2024; pp. 1–20. [Google Scholar]
- Veerabadran, V.; Xiao, F.; Kamra, N.; Matias, P.; Chen, J.; Drooff, C.; Roads, B.; Williams, R.J.; Henderson, E.; Zhao, X.; et al. Benchmarking egocentric multimodal goal inference for assistive wearable agents. Adv. Neural Inf. Process. Syst. 2026, 38. [Google Scholar]
- Jiang, P.; Liu, F.; Guo, J.; Wen, C.K.; Jin, S.; Zhang, J. Intention-Aware Semantic Agent Communications for AI Glasses. arXiv 2026, arXiv:2604.23691. [Google Scholar]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. React: Synergizing reasoning and acting in language models. arXiv 2022, arXiv:2210.03629. [Google Scholar]
- Chen, Y.; Ge, Y.; Ge, Y.; Ding, M.; Li, B.; Wang, R.; Xu, R.; Shan, Y.; Liu, X. Egoplan-bench: Benchmarking multimodal large language models for human-level planning. Int. J. Comput. Vis. 2026, 134, 118. [Google Scholar] [CrossRef]
- Yang, S.; Huang, Y.; Cai, W.; Sun, S.; Fang, F.; He, Y.; Xie, Y.; Deng, J.; Zhang, H.; Song, J.; et al. Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI. Proc. Proc. ACM Web Conf. 2026, 2026, 8862–8873. [Google Scholar] [CrossRef]
- Lin, Z.; Xu, Y.; Sun, K.; Zheng, J.; Huang, Y.; Appini, S.T.; Narang, K.; Tao, R.; Jain, I.K.; Arora, S.; et al. WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables. arXiv 2025, arXiv:2601.02391. [Google Scholar]
- Yu, S.; Shu, L.; Yang, A.; Fu, Y.; Sunkara, S.; Wang, M.; Chen, J.; Bansal, M.; Gong, B. Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 25633–25643. [Google Scholar]
- Jiang, H.; Ithapu, V.K. Egocentric pose estimation from human vision span. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE, 2021; pp. 10986–10994. [Google Scholar]
- Li, J.; Liu, K.; Wu, J. Ego-body pose estimation via ego-head pose estimation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp. 17142–17151. [Google Scholar]
- Guzov, V.; Jiang, Y.; Hong, F.; Pons-Moll, G.; Newcombe, R.; Liu, C.K.; Ye, Y.; Ma, L. Hmd 2: Environment-aware motion generation from single egocentric head-mounted device. In Proceedings of the 2025 International Conference on 3D Vision (3DV); IEEE, 2025; pp. 1394–1405. [Google Scholar]
- Ma, L.; Ye, Y.; Hong, F.; Guzov, V.; Jiang, Y.; Postyeni, R.; Pesqueira, L.; Gamino, A.; Baiyya, V.; Kim, H.J.; et al. Nymeria: A massive collection of multimodal egocentric daily motion in the wild. In Proceedings of the European Conference on Computer Vision, 2024; Springer; pp. 445–465. [Google Scholar]
- Luo, Z.; Cao, J.; Khirodkar, R.; Winkler, A.; Kitani, K.; Xu, W. Real-time simulated avatar from head-mounted sensors. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 571–581. [Google Scholar]
- Fan, Z.; Dai, P.; Su, Z.; Gao, X.; Lv, Z.; Zhang, J.; Du, T.; Wang, G.; Zhang, Y. Emhi: A multimodal egocentric human motion dataset with hmd and body-worn imus. Proc. Proc. AAAI Conf. Artif. Intell. 2025, Vol. 39, 2879–2887. [Google Scholar] [CrossRef]
- Wang, J.; Dabral, R.; Luvizon, D.; Cao, Z.; Liu, L.; Beeler, T.; Theobalt, C. Ego4o: Egocentric human motion capture and understanding from multi-modal input. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 22668–22679. [Google Scholar]
- Xu, M.; Li, Y.; Fu, C.Y.; Ghanem, B.; Xiang, T.; Pérez-Rúa, J.M. Where is my wallet? modeling object proposal sets for egocentric visual query localization. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp. 2593–2603. [Google Scholar]
- Fan, B.; Feng, Y.; Tian, Y.; Liang, J.C.; Lin, Y.; Huang, Y.; Fan, H. Prvql: Progressive knowledge-guided refinement for robust egocentric visual query localization. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 5156–5165. [Google Scholar]
- Khosla, S.; Schwing, A.; Hoiem, D.; et al. Relocate: A simple training-free baseline for visual query localization using region-based representations. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 3697–3706. [Google Scholar]
- Zhu, S.; Yang, L.; Chen, C.; Shah, M.; Shen, X.; Wang, H. R2former: Unified retrieval and reranking transformer for place recognition. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 19370–19380. [Google Scholar]
- Zhang, S.; Mao, H.; Chen, Q.; Kim, Y. Efficient Visual Place Recognition Through Multimodal Semantic Knowledge Integration. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 5601–5610. [Google Scholar]
- Dunnhofer, M.; Furnari, A.; Farinella, G.M.; Micheloni, C. Visual object tracking in first person vision. Int. J. Comput. Vis. 2023, 131, 259–283. [Google Scholar] [CrossRef] [PubMed]
- Zhao, Y.; Ma, H.; Kong, S.; Fowlkes, C. Instance tracking in 3d scenes from egocentric videos. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 21933–21944. [Google Scholar]
- Hu, Z.; Yin, Z.; Haeufle, D.; Schmitt, S.; Bulling, A. Hoimotion: Forecasting human motion during human-object interactions using egocentric 3d object bounding boxes. IEEE Trans. Vis. Comput. Graph. 2024, 30, 7375–7385. [Google Scholar] [CrossRef] [PubMed]
- Li, Y.; Cao, Z.; Liang, A.; Liang, B.; Chen, L.; Zhao, H.; Feng, C. Egocentric prediction of action target in 3d. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2022; pp. 20971–20980. [Google Scholar]
- Bao, W.; Chen, L.; Zeng, L.; Li, Z.; Xu, Y.; Yuan, J.; Kong, Y. Uncertainty-aware state space transformer for egocentric 3d hand trajectory forecasting. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 13702–13711. [Google Scholar]
- Ma, J.; Chen, X.; Bao, W.; Xu, J.; Wang, H. Madiff: Motion-aware mamba diffusion models for hand trajectory prediction on egocentric videos. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. [Google Scholar]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. ImageNet classification with deep convolutional neural networks. Commun. ACM 2017, 60, 84–90. [Google Scholar] [CrossRef]
- Liu, X.; Peng, H.; Zheng, N.; Yang, Y.; Hu, H.; Yuan, Y. Efficientvit: Memory efficient vision transformer with cascaded group attention. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 14420–14430. [Google Scholar]
- Lin, J.; Yin, H.; Ping, W.; Molchanov, P.; Shoeybi, M.; Han, S. Vila: On pre-training for visual language models. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024; pp. 26689–26699. [Google Scholar]
- Wang, Y.; Wang, Z.; Xu, B.; Du, Y.; Lin, K.; Xiao, Z.; Yue, Z.; Ju, J.; Zhang, L.; Yang, D.; et al. Time-r1: Post-training large vision language model for temporal video grounding. Adv. Neural Inf. Process. Syst. 2026, 38, 83330–83364. [Google Scholar]
- Liu, J.; Wang, Y.; Ma, H.; Wu, X.; Ma, X.; Wei, X.; Jiao, J.; Wu, E.; Hu, J. Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input: J. Liu et al. Int. J. Comput. Vis. 2026, 134, 114. [Google Scholar]
- Xu, B.; Mei, Y.; Zheng, S.; Jin, Q.; et al. Egodtm: Towards 3d-aware egocentric video-language pretraining. Adv. Neural Inf. Process. Syst. 2026, 38, 122549–122576. [Google Scholar]
- Feichtenhofer, C.; Fan, H.; Malik, J.; He, K. Slowfast networks for video recognition. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2019; pp. 6202–6211. [Google Scholar]
- Wu, C.Y.; Li, Y.; Mangalam, K.; Fan, H.; Xiong, B.; Malik, J.; Feichtenhofer, C. Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition. In Proceedings of the Proceedings of the ieee/cvf conference on computer vision and pattern recognition, 2022; pp. 13587–13597. [Google Scholar]
- Patrick, M.; Campbell, D.; Asano, Y.; Misra, I.; Metze, F.; Feichtenhofer, C.; Vedaldi, A.; Henriques, J.F. Keeping your eye on the ball: Trajectory attention in video transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12493–12506. [Google Scholar]
- Damen, D.; Doughty, H.; Farinella, G.M.; Furnari, A.; Kazakos, E.; Ma, J.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; et al. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. Int. J. Comput. Vis. 2022, 130, 33–55. [Google Scholar] [CrossRef]
- Sener, F.; Chatterjee, D.; Shelepov, D.; He, K.; Singhania, D.; Wang, R.; Yao, A. Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp. 21096–21106. [Google Scholar]
- Wang, M.; Xing, J.; Jiang, B.; Chen, J.; Mei, J.; Zuo, X.; Dai, G.; Wang, J.; Liu, Y. A multimodal, multi-task adapting framework for video action recognition. Proc. Proc. AAAI Conf. Artif. Intell. 2024, Vol. 38, 5517–5525. [Google Scholar] [CrossRef]
- Quan, Z.; Chen, J.; Deguchi, D.; Sun, J.; Zhang, C.; Li, Y.; Murase, H. Semantic matters: A constrained approach for zero-shot video action recognition. Pattern Recognit. 2025, 162, 111402. [Google Scholar] [CrossRef]
- He, W.J.; Zhu, X.; Zhang, Z. Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition. Proc. Proc. AAAI Conf. Artif. Intell. 2026, Vol. 40, 17463–17471. [Google Scholar] [CrossRef]
- Huang, Z.; He, W.J.; Hu, B.; Zhang, Z. Grading-Inspired Complementary Enhancing for Multimodal Sentiment Analysis. Inf. Fusion 2026, 104174. [Google Scholar] [CrossRef]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International conference on machine learning. PmLR, 2021; pp. 8748–8763. [Google Scholar]
- Kim, M.J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.P.; Sanketi, P.R.; Vuong, Q.; et al. OpenVLA: An Open-Source Vision-Language-Action Model. In Proceedings of the Conference on Robot Learning. PMLR, 2025; pp. 2679–2713. [Google Scholar]
- Qu, H.; Cai, Y.; Liu, J. Llms are good action recognizers. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024; pp. 18395–18406. [Google Scholar]
- Wang, M.; Xing, J.; Mei, J.; Liu, Y.; Jiang, Y. Actionclip: Adapting language-image pretrained models for video action recognition. IEEE Trans. Neural Netw. Learn. Syst. 2023, 36, 625–637. [Google Scholar] [CrossRef] [PubMed]
- Chatterjee, D.; Sener, F.; Ma, S.; Yao, A. Opening the vocabulary of egocentric actions. Adv. Neural Inf. Process. Syst. 2023, 36, 33174–33187. [Google Scholar] [CrossRef]
- Girase, H.; Agarwal, N.; Choi, C.; Mangalam, K. Latency matters: Real-time action forecasting transformer. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 18759–18769. [Google Scholar]
- Mittal, H.; Agarwal, N.; Lo, S.Y.; Lee, K. Can’t make an omelette without breaking some eggs: Plausible action anticipation using large video-language models. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 18580–18590. [Google Scholar]
- Guo, H.; Agarwal, N.; Lo, S.Y.; Lee, K.; Ji, Q. Uncertainty-aware action decoupling transformer for action anticipation. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024; pp. 18644–18654. [Google Scholar]
- Lin, K.Q.; Shou, M.Z. Vlog: Video-language models by generative retrieval of narration vocabulary. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025; pp. 3218–3228. [Google Scholar]
- Miao, Q.; Golani, V.R.; Xu, J.; Dutta, P.P.; Hoai, M.; Samaras, D. Multi-view Gaze Target Estimation. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 5371–5381. [Google Scholar]
- He, W.J.; Zhang, Z.; Zhu, X. Dual-correlation-guided anchor learning for scalable incomplete multi-view clustering. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 17336–17349. [Google Scholar] [CrossRef] [PubMed]
- He, W.J.; Liu, C.; Luo, Z.; Zhang, Z. Prototype-tailored cross-view alignment for partially-aligned incomplete multi-view clustering. Pattern Recognit. 2026, 114233. [Google Scholar] [CrossRef]
- Meyer, J.; Zimmer, A.; Vilches, S. Ambient Light Robust Eye-Tracking for Smart Glasses Using Laser Feedback Interferometry Sensors with Elongated Laser Beams. Proc. ACM Hum.-Comput. Interact. 2025, 9, 1–17. [Google Scholar] [CrossRef]
- Ma, R.; Morimoto, Y.; Ho, J.S.; Shiu, S.; Zhu, J. mmET: mmWave Radar-Based Eye Tracking on Smart Glasses. In Proceedings of the Proceedings of the 23rd ACM Conference on Embedded Networked Sensor Systems, 2025; pp. 30–42. [Google Scholar]
- Jiang, H.; Murdock, C.; Ithapu, V.K. Egocentric deep multi-channel audio-visual active speaker localization. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp. 10544–10552. [Google Scholar]
- Xue, Z.; Song, Y.; Grauman, K.; Torresani, L. Egocentric video task translation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp. 2310–2320. [Google Scholar]
- Chang, K.K.; Cramer, M.H.; Ho, A.; Nguyen, T.T.; Yuan, Y.; Bamman, D. Multimodal conversation structure understanding. Proceedings of the Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics 2026, Volume 1, 7437–7458. [Google Scholar] [CrossRef]
- Lei, J.; Yu, L.; Bansal, M.; Berg, T. Tvqa: Localized, compositional video question answering. In Proceedings of the Proceedings of the 2018 conference on empirical methods in natural language processing, 2018; pp. 1369–1379. [Google Scholar]
- Mathur, L.; Qian, M.; Liang, P.P.; Morency, L.P. Social genome: Grounded social reasoning abilities of multimodal models. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025; pp. 24879–24902. [Google Scholar]
- Hyun, L.; Sung-Bin, K.; Han, S.; Yu, Y.; Oh, T.H. Smile: Multimodal dataset for understanding laughter in video with language models. Proc. Find. Assoc. Comput. Linguist. NAACL 2024, 2024, 1149–1167. [Google Scholar] [CrossRef]
- Cao, X.; Virupaksha, P.; Jia, W.; Lai, B.; Ryan, F.; Lee, S.; Rehg, J.M. Socialgesture: Delving into multi-person gesture understanding. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 19509–19519. [Google Scholar]
- Zhang, H.; Xu, H.; Zhu, Y.; Wang, P.; Zhu, H.; Zhou, J.; Zhang, J.; et al. Can large language models help multimodal language analysis? mmla: A comprehensive benchmark. Adv. Neural Inf. Process. Syst. 2026, 38. [Google Scholar]
- Ouyang, L.; Huang, Y.; Zhang, M.; Kang, C.; Furuta, R.; Sato, Y. Multi-speaker attention alignment for multimodal social interaction. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 24608–24619. [Google Scholar]
- Zhong, W.; Guo, L.; Gao, Q.; Ye, H.; Wang, Y. Memorybank: Enhancing large language models with long-term memory. Proc. Proc. AAAI Conf. Artif. Intell. 2024, Vol. 38, 19724–19731. [Google Scholar] [CrossRef]
- Park, J.S.; O’Brien, J.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the Proceedings of the 36th annual acm symposium on user interface software and technology, 2023; pp. 1–22. [Google Scholar]
- Karpukhin, V.; Oguz, B.; Min, S.; Lewis, P.; Wu, L.; Edunov, S.; Chen, D.; Yih, W.t. Dense passage retrieval for open-domain question answering. In Proceedings of the Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), 2020; pp. 6769–6781. [Google Scholar]
- Yuan, X.; Zhang, Z.; Wang, X.; Wu, L. Semantic-aware adversarial training for reliable deep hashing retrieval. IEEE Trans. Inf. Forensics Secur. 2023, 18, 4681–4694. [Google Scholar] [CrossRef]
- Xiao, J.; Zhang, S.; Zhu, P.; Yao, A. Ego-Grounding for Personalized Question-Answering in Egocentric Videos. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 40537–40547. [Google Scholar]
- Tang, Y.; Zhang, J.; Qin, X.; Yu, J.; Qiu, M.; Gou, G.; Xiong, G.; Lin, Q.; Rajmohan, S.; Zhang, D.; et al. Memory-Augmented Personalized Retrieval for Long-Context Egocentric Video, 2026. [CrossRef]
- Lu, Z.; Abbasian, M.; Rahmani, A.M. Query-Conditioned Graph Retrieval for Contextualized LLM Reasoning in Personalized Wearable Data. arXiv 2026, arXiv:2605.18763. [Google Scholar]
- Jiang, Z.; Chen, Y.; Wang, S.; Qu, H.; Jindong, Z.; Fan, W.; Qing, L.; Liang, D.; Wang, J. Atomic Intent Reasoning: Bringing LLM Semantics to Industrial Cross-Domain Recommendations. arXiv 2026, arXiv:2606.10357. [Google Scholar]
- Qu, H.; Lin, S.; Ding, Y.; Wang, Y.; Fan, W. Diffusion generative recommendation with continuous tokens. In Proceedings of the Proceedings of the ACM Web Conference 2026, 2026; pp. 7259–7270. [Google Scholar]
- Wu, Y.; Liu, C.; Fan, W.; Zhang, R. Beyond Static Diffusion: Explicitly Modeling Temporal Patterns in Sequential Recommendation. In Proceedings of the Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2026; pp. 1983–1993. [Google Scholar]
- Huang, J.; Wang, S.; Ning, L.B.; Fan, W.; Li, Q. ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning. Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics 2026, Volume 1, 21040–21055. [Google Scholar] [CrossRef]
- Qu, H.; Fan, W.; Zhao, Z.; Li, Q. Tokenrec: Learning to tokenize id for llm-based generative recommendations. IEEE Transactions on Knowledge and Data Engineering, 2025. [Google Scholar]
- Wang, S.; Fan, W.; Feng, Y.; Shanru, L.; Ma, X.; Wang, S.; Yin, D. Knowledge graph retrieval-augmented generation for llm-based recommendation. Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 2025, Volume 1, 27152–27168. [Google Scholar] [CrossRef]
- Ning, L.; Fan, W.; Li, Q. Exploring backdoor attack and defense for llm-empowered recommendations. IEEE Transactions on Knowledge and Data Engineering, 2026. [Google Scholar]
- Bao, K.; Yan, M.; Zhang, Y.; Zhang, J.; Wang, W.; Feng, F.; He, X. Customizing in-context learning for dynamic interest adaption in llm-based recommendation. Proc. Find. Assoc. Comput. Linguist. ACL 2025, 2025, 14278–14291. [Google Scholar] [CrossRef]
- Wilson, R.; Graham, C.; Carter, C.; Yang, Z.; Gu, R. CAP-LLM: Context-Augmented Personalized Large Language Models for News Headline Generation. arXiv 2025, arXiv:2508.03935. [Google Scholar]
- Tan, Z.; Zeng, Z.; Zeng, Q.; Wu, Z.; Liu, Z.; Mo, F.; Jiang, M. Can Large Language Models Understand Preferences in Personalized Recommendation? arXiv 2025, arXiv:2501.13391. [Google Scholar]
- Chen, Y.; Yue, X.; Zhang, C.; Gao, X.; Tan, R.T.; Li, H. Voicebench: Benchmarking llm-based voice assistants. Trans. Assoc. Comput. Linguist. 2026, 14, 378–398. [Google Scholar] [CrossRef]
- Deng, S.; Xu, W.; Sun, H.; Liu, W.; Tan, T.; Liujianfeng, L.; Li, A.; Luan, J.; Wang, B.; Yan, R.; et al. Mobile-bench: An evaluation benchmark for llm-based mobile agents. Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics 2024, Volume 1, 8813–8831. [Google Scholar] [CrossRef]
- Huang, J.; Wang, S.; Ning, L.; Fan, W.; Wang, S.; Yin, D.; Li, Q. Towards next-generation recommender systems: A benchmark for personalized recommendation assistant with llms. In Proceedings of the Proceedings of the Nineteenth ACM International Conference on Web Search and Data Mining, 2026; pp. 217–226. [Google Scholar]
- Rodin, I.; Furnari, A.; Min, K.; Tripathi, S.; Farinella, G.M. Action scene graphs for long-form understanding of egocentric videos. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 18622–18632. [Google Scholar]
- Liu, C.; Zhai, W.; Yang, Y.; Luo, H.; Liang, S.; Cao, Y.; Zha, Z.J. Grounding 3d scene affordance from egocentric interactions. arXiv 2024, arXiv:2409.19650. [Google Scholar]
- Sun, P.; Xiao, J.; Tse, T.H.E.; Li, Y.; Akula, A.; Yao, A. Visual intention grounding for egocentric assistants. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 2512–2522. [Google Scholar]
- Chang, E.; Huang, Z.; Liao, Y.; Bhavsar, S.; Param, A.; Stark, T.; Ahmadyan, A.; Yang, X.; Wang, J.; Abdullah, A.; et al. WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios. Adv. Neural Inf. Process. Syst. 2026, 38. [Google Scholar]
- Yang, B.; Xu, L.; Zeng, L.; Liu, K.; Jiang, S.; Lu, W.; Chen, H.; Jiang, X.; Xing, G.; Yan, Z. Contextagent: Context-aware proactive llm agents with open-world sensory perceptions. Adv. Neural Inf. Process. Syst. 2026, 38, 167509–167543. [Google Scholar]
- Gao, H.; Zhang, K.; Wang, S.; Chen, M.; Cao, Q.; Wang, X.; Zhu, Y.; Min, X.; Sun, W.; Zhu, D.; et al. LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 32892–32902. [Google Scholar]
- Mahmud, S.; Parikh, V.; Liang, Q.; Li, K.; Zhang, R.; Ajit, A.; Gunda, V.; Agarwal, D.; Guimbretière, F.; Zhang, C. ActSonic: recognizing everyday activities from inaudible acoustic wave around the body. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 2024, 8, 1–32. [Google Scholar] [CrossRef]
- Datta, S.; Dharur, S.; Cartillier, V.; Desai, R.; Khanna, M.; Batra, D.; Parikh, D. Episodic memory question answering. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp. 19119–19128. [Google Scholar]
- Ramakrishnan, S.K.; Al-Halah, Z.; Grauman, K. Spotem: Efficient video search for episodic memory. In Proceedings of the International Conference on Machine Learning. PMLR, 2023; pp. 28618–28636. [Google Scholar]
- Guu, K.; Lee, K.; Tung, Z.; Pasupat, P.; Chang, M. Retrieval augmented language model pre-training. In Proceedings of the International conference on machine learning. PMLR, 2020; pp. 3929–3938. [Google Scholar]
- Izacard, G.; Grave, E. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the Proceedings of the 16th conference of the european chapter of the association for computational linguistics: main volume, 2021; pp. 874–880. [Google Scholar]
- Liu, C.; Ning, L.; Ding, Y.; Fan, W. Inference Cost Attacks for Retrieval-Augmented Large Language Models. In Proceedings of the Proceedings of the ACM Web Conference 2026, 2026; pp. 7564–7575. [Google Scholar]
- Borgeaud, S.; Mensch, A.; Hoffmann, J.; Cai, T.; Rutherford, E.; Millican, K.; Van Den Driessche, G.B.; Lespiau, J.B.; Damoc, B.; Clark, A.; et al. Improving language models by retrieving from trillions of tokens. In Proceedings of the International conference on machine learning. PMLR, 2022; pp. 2206–2240. [Google Scholar]
- Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; Hajishirzi, H. Self-rag: Learning to retrieve, generate, and critique through self-reflection. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 9112–9141. [Google Scholar]
- Edge, D.; Trinh, H.; Cheng, N.; Bradley, J.; Chao, A.; Mody, A.; Truitt, S.; Metropolitansky, D.; Ness, R.O.; Larson, J. From local to global: A graph rag approach to query-focused summarization. arXiv 2024, arXiv:2404.16130. [Google Scholar]
- Yuan, X.; Ning, L.; Ye, Q.; Fan, W.; Li, Q. mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA. In Proceedings of the Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2026; pp. 2274–2285. [Google Scholar]
- Jiang, Z.; Ma, X.; Chen, W. Longrag: Enhancing retrieval-augmented generation with long-context llms. arXiv 2024, arXiv:2406.15319. [Google Scholar]
- Jiang, Z.; Wu, P.; Yuan, X.; Fan, W.; Li, Q. QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering. arXiv 2025, arXiv:2508.05197. [Google Scholar]
- Wang, J.; Yang, X.; Sun, K.; Suresh, P.; Sharma, S.; Czyzewski, A.; Andersen, D.; Appini, S.; Banerjee, A.; Choudhary, S.; et al. CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark. arXiv 2025, arXiv:2510.26160. [Google Scholar]
- Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: Language agents with verbal reinforcement learning. Adv. Neural Inf. Process. Syst. 2023, 36, 8634–8652. [Google Scholar] [CrossRef]
- Wang, W.; Dong, L.; Cheng, H.; Liu, X.; Yan, X.; Gao, J.; Wei, F. Augmenting language models with long-term memory. Adv. Neural Inf. Process. Syst. 2023, 36, 74530–74543. [Google Scholar] [CrossRef]
- Packer, C.; Wooders, S.; Lin, K.; Fang, V.; Patil, S.G.; Stoica, I.; Gonzalez, J.E. MemGPT: Towards LLMs as Operating Systems. arXiv 2023, arXiv:2310.08560. [Google Scholar]
- Xu, W.; Liang, Z.; Mei, K.; Gao, H.; Tan, J.; Zhang, Y. A-mem: Agentic memory for llm agents. Adv. Neural Inf. Process. Syst. 2026, 38, 17577–17604. [Google Scholar]
- Chhikara, P.; Khant, D.; Aryan, S.; Singh, T.; Yadav, D. Mem0: Building production-ready ai agents with scalable long-term memory. arXiv 2025, arXiv:2504.19413. [Google Scholar]
- Chang, Y.; Zhao, Y.; Dong, M.; Wang, Y.; Lu, Y.; Lv, Q.; Dick, R.P.; Lu, T.; Gu, N.; Shang, L. MemX: An attention-aware smart eyewear system for personalized moment auto-capture. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 2021, 5, 1–23. [Google Scholar]
- Kim, Y.; Jadeja, D.; Pradhan, D.; Yang, Y.; Kaufman, A.E. SpeechLess: Micro-utterance with Personalized Spatial Memory-aware Assistant in Everyday Augmented Reality. In Proceedings of the 2026 IEEE Conference on Virtual Reality and 3D User Interfaces (VR); IEEE, 2026; pp. 217–227. [Google Scholar]
- Wen, Z.; Wang, Y.; Liao, C.; Yang, B.; Li, J.; Liu, W.; He, H.; Feng, B.; Liu, X.; Lyu, Y.; et al. Ai for service: Proactive assistance with ai glasses. arXiv 2025, arXiv:2510.14359. [Google Scholar]
- Fung, P.; Bachrach, Y.; Celikyilmaz, A.; Chaudhuri, K.; Chen, D.; Chung, W.; Dupoux, E.; Gong, H.; Jégou, H.; Lazaric, A.; et al. Embodied ai agents: Modeling the world. arXiv 2025, arXiv:2506.22355. [Google Scholar]
- Zhang, Z.; Yu, M.; Wang, T.; Todi, K.; Fernandes, A.S.; Liu, Y.; Xia, H.; Grossman, T.; Jonker, T.R. Gazeify Then Voiceify: Physical Object Referencing Through Gaze and Voice Interaction with Displayless Smart Glasses. In Proceedings of the Proceedings of the 31st International Conference on Intelligent User Interfaces, 2026; pp. 1821–1841. [Google Scholar]
- Cho, H.; Fashimpaur, J.; Sendhilnathan, N.; Browder, J.; Lindlbauer, D.; Jonker, T.R.; Todi, K. Persistent assistant: Seamless everyday AI interactions via intent grounding and multimodal feedback. In Proceedings of the Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025; pp. 1–19. [Google Scholar]
- Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Hambro, E.; Zettlemoyer, L.; Cancedda, N.; Scialom, T. Toolformer: Language models can teach themselves to use tools. Adv. Neural Inf. Process. Syst. 2023, 36, 68539–68551. [Google Scholar] [CrossRef]
- Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; Zhuang, Y. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face. Adv. Neural Inf. Process. Syst. 2023, 36, 38154–38180. [Google Scholar] [CrossRef]
- Jiang, Z.; Wu, P.; Liang, Z.; Chen, P.Q.; Yuan, X.; Jia, Y.; Tu, J.; Li, C.; Ng, P.H.; Li, Q. Hibench: Benchmarking llms capability on hierarchical structure reasoning. Proceedings of the Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining 2025, V. 2, 5505–5515. [Google Scholar] [CrossRef]
- Ye, H.; Zhang, H.; Daxberger, E.; Chen, L.; Lin, Z.; Li, Y.; Zhang, B.; You, H.; Xu, D.; Gan, Z.; et al. MMEgo: Towards building egocentric multimodal LLMs for video QA. Proc. Int. Conf. Learn. Represent. 2025, Vol. 2025, 71705–71723. [Google Scholar]
- Qiu, L.; Chen, Y.; Ge, Y.; Ge, Y.; Shan, Y.; Liu, X. Egoplan-bench2: A benchmark for multimodal large language model planning in real-world scenarios. Int. J. Comput. Vis. 2026, 134, 222. [Google Scholar] [CrossRef]
- Li, Y.; Veerabadran, V.; Iuzzolino, M.L.; Roads, B.D.; Celikyilmaz, A.; Ridgeway, K. EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos. ArXiv 2025, abs/2503.22152. [Google Scholar]
- Huang, Y.; Xu, J.; Pei, B.; Yang, L.; Zhang, M.; He, Y.; Chen, G.; Chen, X.; Wang, Y.; Nie, Z.; et al. Vinci: A real-time smart assistant based on egocentric vision-language model for portable devices. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 2025, 9, 1–33. [Google Scholar] [CrossRef]
- Chen, Y.; Chen, J.; Meng, R.; Yin, J.; Li, N.; Fan, C.; Wang, C.; Pfister, T.; Yoon, J. Tumix: Multi-agent test-time scaling with tool-use mixture. Proc. Int. Conf. Learn. Represent. 2026, Vol. 2026, 55848–55878. [Google Scholar]
- Wu, P.; Li, X. Dynamic Action Space Reinforcement Learning for Optimal Trading Execution. In Proceedings of the Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems, 2026; pp. 1883–1891. [Google Scholar]
- Wu, P.; Li, X. Timing Optimization in Dynamic Discrete Action Space Lifelong Reinforcement Learning. In Proceedings of the Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems, 2026; pp. 3310–3312. [Google Scholar]
- Ning, L.B.; Zhu, Y.; Huang, H.; Wang, X.; Chang, Y.; Li, Q.; Fan, W. When Efficiency Becomes a Vulnerability: Computational Cost Attacks on WebAgents. Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics 2026, Volume 1, 38315–38335. [Google Scholar] [CrossRef]
- Nakano, R.; Hilton, J.; Balaji, S.; Wu, J.; Long, O.; Kim, C.; Hesse, C.; Jain, S.; Kosaraju, V.; Saunders, W.; et al. WebGPT: Browser-assisted question-answering with human feedback. ArXiv 2021, abs/2112.09332. [Google Scholar]
- Patil, S.G.; Zhang, T.; Wang, X.; Gonzalez, J.E. Gorilla: Large language model connected with massive apis. Adv. Neural Inf. Process. Syst. 2024, 37, 126544–126565. [Google Scholar] [CrossRef]
- Deng, X.; Gu, Y.; Zheng, B.; Chen, S.; Stevens, S.; Wang, B.; Sun, H.; Su, Y. Mind2web: Towards a generalist agent for the web. Adv. Neural Inf. Process. Syst. 2023, 36, 28091–28114. [Google Scholar] [CrossRef]
- Zhou, S.; Xu, F.F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Ou, T.; Bisk, Y.; Fried, D.; et al. Webarena: A realistic web environment for building autonomous agents. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 15585–15606. [Google Scholar]
- Zhang, R.; Li, K.; Hao, Y.; Wang, Y.; Lai, Z.; Guimbretière, F.; Zhang, C. EchoSpeech: continuous silent speech recognition on minimally-obtrusive eyewear powered by acoustic sensing. In Proceedings of the Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, 2023; pp. 1–18. [Google Scholar]
- Li, K.; Zhang, R.; Chen, B.; Chen, S.; Yin, S.; Mahmud, S.; Liang, Q.; Guimbretière, F.; Zhang, C. Gazetrak: Exploring acoustic-based eye tracking on a glass frame. In Proceedings of the Proceedings of the 30th Annual International Conference on Mobile Computing and Networking, 2024; pp. 497–512. [Google Scholar]
- Sun, T.; Zhao, Y.; Xie, W.; Li, J.; Ma, Y.; Zhang, J. EyeGesener: Eye gesture listener for smart glasses interaction using acoustic sensing. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 2024, 8, 1–28. [Google Scholar] [CrossRef]
- Xu, Z.; Pei, H.; Feng, J.; Zhou, J. FingerGlass: Enhancing Smart Glasses Interaction via Fingerprint Sensing. In Proceedings of the Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025; pp. 1–18. [Google Scholar]
- Bhattacharyya, P.; Mitton, J.; Page, R.; Morgan, O.; Menzies, B.; Homewood, G.; Jacobs, K.; Baesso, P.; Trickett, D.; Mair, C.; et al. Helios: An extremely low power event-based gesture recognition for always-on smart eyewear. In Proceedings of the European Conference on Computer Vision, 2024; Springer; pp. 168–184. [Google Scholar]
- Janaka, N.; Gao, J.; Zhu, L.; Zhao, S.; Lyu, L.; Xu, P.; Nabokow, M.; Wang, S.; Ong, Y. GlassMessaging: Towards Ubiquitous Messaging Using OHMDs. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 2023, 7, 1–32. [Google Scholar]
- Zhao, Y.; Kupferstein, E.; Rojnirun, H.; Findlater, L.; Azenkot, S. The effectiveness of visual and audio wayfinding guidance on smartglasses for people with low vision. In Proceedings of the Proceedings of the 2020 CHI conference on human factors in computing systems, 2020; pp. 1–14. [Google Scholar]
- Guo, Y.; Li, D. Interaction Design Strategies of AI Smart Glasses for Older Workers: An Embodied Cognition Perspective and Usability Evaluation. Appl. Sci. 2026, 16, 2768. [Google Scholar] [CrossRef]
- Zhang, Z.; Bao, C.; Pan, X.; Chang, C.M.; Igarashi, T.; Zhang, G. Through the lens of privacy: Exploring privacy protection in vision-language model interactions on smart glasses. In Proceedings of the Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 2025; pp. 1–8. [Google Scholar]
- Jin, L.; Gunawardena, P.; Haroon, A.; Wang, R.; Lee, S.; Stoleru, R.; Middleton, M.; Huo, Z.; Kim, J.; Moats, J. A Smart-Glasses for Emergency Medical Services via Multimodal Multitask Learning. arXiv 2025, arXiv:2511.13078. [Google Scholar]
- Romare, C.; Skär, L. The use of smart glasses in nursing education: A scoping review. Nurse Educ. Pract. 2023, 73, 103824. [Google Scholar] [CrossRef] [PubMed]
- Zuidhof, N.; Peters, O.; Verbeek, P.P.; Ben Allouch, S. Social acceptance of smart glasses in health care: model evaluation study of anticipated adoption and social interaction. JMIR Form. Res. 2025, 9, e49610. [Google Scholar] [CrossRef] [PubMed]
- Burch, B.F.; Gu, J.; Betz, G.; Huang, C.M.; McPherson, R.; Ryan, A.S.; Resnick, B. Smart Glasses for Older Adults With Cognitive Impairment: A Scoping Review. J. Am. Med. Dir. Assoc. 2025, 26, 105831. [Google Scholar] [CrossRef] [PubMed]
- Delgado-Morales, C.; García-Iglesias, J.J.; Duarte-Hueros, A. AI-Enabled Smart Glasses for Active Aging: Scoping Review. JMIR Aging 2026, 9, e81157. [Google Scholar] [CrossRef] [PubMed]
- Lee, Z.Y.; Sun, Y.C.; Wong, Y.N.; Hsu, C.H. Exploring LLM-based Assistants with Smart Glasses for the Visually Impaired. In Proceedings of the Proceedings of the Tenth ACM/IEEE Symposium on Edge Computing, 2025; pp. 1–8. [Google Scholar]
- Curtis, H.; Neate, T. Making smartglasses accessible: perspectives and prototypes from co-design with people with aphasia. Sci. Rep. 2025, 15, 38309. [Google Scholar] [CrossRef] [PubMed]
- Ding, J.; Zhang, Y.; Zhu, X.T.; Yang, K.; Wei, Y.; Wang, S.; Liu, Y.; Jiao, Y. Reshaping Inclusive Interpersonal Dynamics through Smart Glasses in Mixed-Vision Social Activities. In Proceedings of the Proceedings of the 2026 Designing Interactive Systems Conference, 2026; pp. 4384–4399. [Google Scholar]
- Konstantinidis, F.K.; Kansizoglou, I.; Santavas, N.; Mouroutsos, S.G.; Gasteratos, A. Marma: A mobile augmented reality maintenance assistant for fast-track repair procedures in the context of industry 4.0. Machines 2020, 8, 88. [Google Scholar] [CrossRef]
- Tang, Y.M.; Au, K.M.; Lau, H.C.; Ho, G.T.; Wu, C.H. Evaluating the effectiveness of learning design with mixed reality (MR) in higher education. Virtual Real. 2020, 24, 797–807. [Google Scholar] [CrossRef]
- Dimitriadou, E.; Lanitis, A. A critical evaluation, challenges, and future perspectives of using artificial intelligence and emerging technologies in smart classrooms. Smart Learn. Environ. 2023, 10, 12. [Google Scholar] [CrossRef] [PubMed]
- Jia, Y.; Wu, X.; Hao, L.; QinglinZhang, Q.; Hu, Y.; Zhao, S.; Fan, W. Uni-retrieval: A multi-style retrieval framework for stem’s education. Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 2025, Volume 1, 10182–10197. [Google Scholar] [CrossRef]
- Danielsson, O.; Holm, M.; Syberfeldt, A. Augmented reality smart glasses for operators in production: Survey of relevant categories for supporting operators. Procedia CIRP 2020, 93, 1298–1303. [Google Scholar] [CrossRef]
- Alessa, F.M.; Alhaag, M.H.; Al-Harkan, I.M.; Ramadan, M.Z.; Alqahtani, F.M. A neurophysiological evaluation of cognitive load during augmented reality interactions in various industrial maintenance and assembly tasks. Sensors 2023, 23, 7698. [Google Scholar] [CrossRef] [PubMed]
- Tom Dieck, M.C.; Jung, T.; Han, D.I. Mapping requirements for the wearable smart glasses augmented reality museum application. J. Hosp. Tour. Technol. 2016, 7, 230–253. [Google Scholar] [CrossRef]
- Han, D.I.D.; Tom Dieck, M.C.; Jung, T. Augmented Reality Smart Glasses (ARSG) visitor adoption in cultural tourism. Leis. Stud. 2019, 38, 618–633. [Google Scholar] [CrossRef]
- Obeidy, W.K.; Arshad, H.; Huang, J.Y. TouristicAR: A Smart Glass Augmented Reality Application for UNESCO World Heritage Sites in Malaysia. J. Telecommun. Electron. Comput. Eng. (JTEC) 2018, 10, 101–108. [Google Scholar]
- Litvak, E.; Kuflik, T. Enhancing cultural heritage outdoor experience with augmented-reality smart glasses. Personal. Ubiquitous Comput. 2020, 24, 873–886. [Google Scholar] [CrossRef]
- Caria, M.; Sara, G.; Todde, G.; Polese, M.; Pazzona, A. Exploring smart glasses for augmented reality: A valuable and integrative tool in precision livestock farming. Animals 2019, 9, 903. [Google Scholar] [CrossRef] [PubMed]
- Caria, M.; Todde, G.; Sara, G.; Piras, M.; Pazzona, A. Performance and usability of smartglasses for augmented reality in precision livestock farming operations. Appl. Sci. 2020, 10, 2318. [Google Scholar] [CrossRef]
- Xi, M.; Rahman, A.; Nguyen, C.; Arnold, S.; McCulloch, J. Smart headset, computer vision and machine learning for efficient prawn farm management. Aquac. Eng. 2023, 102, 102339. [Google Scholar] [CrossRef]
- Kunkera, Z.; Željković, I.; Mimica, R.; Ljubenkov, B.; Opetuk, T. Development of augmented reality technology implementation in a shipbuilding project realization process. J. Mar. Sci. Eng. 2024, 12, 550. [Google Scholar] [CrossRef]
- Baek, J.; Choi, Y. Smart glasses-based personnel proximity warning system for improving pedestrian safety in construction and mining sites. Int. J. Environ. Res. Public Health 2020, 17, 1422. [Google Scholar] [CrossRef] [PubMed]
Figure 1.
Overview of this survey. AI smart glasses are organized as integrated wearable-intelligence platforms that connect hardware foundation, wearable intelligence, interaction design, application scenarios, and cross-cutting research challenges.
Figure 1.
Overview of this survey. AI smart glasses are organized as integrated wearable-intelligence platforms that connect hardware foundation, wearable intelligence, interaction design, application scenarios, and cross-cutting research challenges.

Figure 2.
Evolution of smart glasses. Smart glasses have progressed from an early exploratory phase to a stage of hardware specialization and, more recently, to the convergence of wearable hardware and multimodal foundation models. Looking ahead, they are expected to evolve into proactive, embodied intelligent assistants that augment human cognition, perception, and interaction.
Figure 2.
Evolution of smart glasses. Smart glasses have progressed from an early exploratory phase to a stage of hardware specialization and, more recently, to the convergence of wearable hardware and multimodal foundation models. Looking ahead, they are expected to evolve into proactive, embodied intelligent assistants that augment human cognition, perception, and interaction.

Figure 3.
Overview of AI smart glasses hardware capabilities. By synthesizing the hardware components into a capability stack, AI smart glasses are able to effectively observe, reason, and respond to users in a unified workflow.
Figure 3.
Overview of AI smart glasses hardware capabilities. By synthesizing the hardware components into a capability stack, AI smart glasses are able to effectively observe, reason, and respond to users in a unified workflow.

Figure 4.
Hierarchical taxonomy of wearable intelligence for AI smart glasses, consisting of perception intelligence, contextual intelligence, and agentic intelligence.
Figure 4.
Hierarchical taxonomy of wearable intelligence for AI smart glasses, consisting of perception intelligence, contextual intelligence, and agentic intelligence.

Figure 5.
Overview of perception intelligence in AI smart glasses. Continuous egocentric sensing is organized into spatial, temporal, conceptual, and social perception, which respectively ground geometric relationships, dynamic evolution, semantic entities and behaviors, and interpersonal signals for downstream contextual and agentic intelligence.
Figure 5.
Overview of perception intelligence in AI smart glasses. Continuous egocentric sensing is organized into spatial, temporal, conceptual, and social perception, which respectively ground geometric relationships, dynamic evolution, semantic entities and behaviors, and interpersonal signals for downstream contextual and agentic intelligence.

Figure 6.
Overview of contextual intelligence in AI smart glasses. Egocentric observations are interpreted through personalized, environmental, temporal, and knowledge context, allowing the system to connect what is perceived with the wearer’s history, surrounding situation, life-event continuity, and external or personal knowledge.
Figure 6.
Overview of contextual intelligence in AI smart glasses. Egocentric observations are interpreted through personalized, environmental, temporal, and knowledge context, allowing the system to connect what is perceived with the wearer’s history, surrounding situation, life-event continuity, and external or personal knowledge.

Figure 7.
Overview of agentic intelligence in AI smart glasses. Situated observations and contextual evidence are transformed into an intent–plan–act loop: intent reasoning grounds the user’s goal and assistance boundary, task planning decomposes the goal under wearable constraints, and action execution delivers tool use, connected-device control, or situated feedback.
Figure 7.
Overview of agentic intelligence in AI smart glasses. Situated observations and contextual evidence are transformed into an intent–plan–act loop: intent reasoning grounds the user’s goal and assistance boundary, task planning decomposes the goal under wearable constraints, and action execution delivers tool use, connected-device control, or situated feedback.

Figure 8.
Overview of interaction design in AI smart glasses. The interaction loop captures intent from speech, gaze, touch, hand gestures, and other cues; grounds references through situated and memory evidence; selects feedback channels and repair cues; governs proactive intervention; and constrains all stages by accessibility, privacy, transparency, and user control.
Figure 8.
Overview of interaction design in AI smart glasses. The interaction loop captures intent from speech, gaze, touch, hand gestures, and other cues; grounds references through situated and memory evidence; selects feedback channels and repair cues; governs proactive intervention; and constrains all stages by accessibility, privacy, transparency, and user control.

Figure 9.
Representative workflow of AI smart glasses across healthcare, accessibility, situated learning, daily life assistance, cultural tourism, and industrial support.
Figure 9.
Representative workflow of AI smart glasses across healthcare, accessibility, situated learning, daily life assistance, cultural tourism, and industrial support.

Figure 10.
Overview of open problems and future research directions for AI smart glasses. The five directions highlight advances in next-generation hardware, trustworthy egocentric intelligence, lifelong personalized memory, proactive intelligence, and embodied foundation models, which jointly support efficient, reliable, personalized, and context-aware wearable intelligence.
Figure 10.
Overview of open problems and future research directions for AI smart glasses. The five directions highlight advances in next-generation hardware, trustworthy egocentric intelligence, lifelong personalized memory, proactive intelligence, and embodied foundation models, which jointly support efficient, reliable, personalized, and context-aware wearable intelligence.

Table 1.
Hardware specifications of representative AI smart glasses. The dash (-) indicates that the information is not disclosed in public materials; Snapdragon AR1 and AR1 both refer to the Qualcomm Snapdragon AR1 Gen1 chipset; mic and spk denote microphone and speaker respectively.
Table 1.
Hardware specifications of representative AI smart glasses. The dash (-) indicates that the information is not disclosed in public materials; Snapdragon AR1 and AR1 both refer to the Qualcomm Snapdragon AR1 Gen1 chipset; mic and spk denote microphone and speaker respectively.
| Product | Release Time |
Observation | Reasoning | Responsiveness | Sustainability | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Camera | Motion | Spatial | Local Inference | Cloud AI | Audio | Display | Touch | Battery | Storage | ||
| Resolution | Stabilization | Awareness | Control | ||||||||
| Ray-Ban Meta | 2023.09 | 12 MP | – | Snapdragon AR1 | Meta AI | 5-mic, 2-spk | × | 160 mAh | 32 GB | ||
| Rokid AI Glasses | 2024.11 | 12 MP | – | Snapdragon AR1 | Multi-LLM | 4-mic, 2-spk | Binocular | 210 mAh | 32 GB | ||
| INMO Air3 | 2024.11 | 16 MP | Space Computing Chip | GLM AI | 4-mic, 2-spk | Binocular | 660 mAh | 128 GB | |||
| RayNeo X3 Pro | 2025.05 | 12 MP | Snapdragon AR1 | Qwen | 3-mic, 4-spk | Binocular | 245 mAh | 32 GB | |||
| Xiaomi AI Glasses | 2025.06 | 12 MP | – | AR1 + BES2700 | MiMo | 5-mic, 2-spk | × | 263 mAh | 32 GB | ||
| Meta Ray-Ban Display | 2025.09 | 12 MP | – | Snapdragon AR1 | Meta AI | 6-mic, 2-spk | Monocular | 248 mAh | 32 GB | ||
| INMO GO3 | 2025.11 | 8 MP | – | Unisoc W337 | GLM AI | 4-mic, 2-spk | Binocular | 270 mAh | 64 GB | ||
| Qwen G1 | 2026.03 | 12 MP | – | AR1 + BES2800 | Qwen | 6-mic, 2-spk | × | 272 mAh | 64 GB | ||
| Huawei AI Glasses | 2026.04 | 12 MP | – | Self-developed Chip | PanguLM | 3-mic, 2-spk | × | 252 mAh | 64 GB | ||
| RayNeo V4 | 2026.05 | 9 MP | – | AR1 + BES2800BP | Qwen | 4-mic, 2-spk | × | 250 mAh | 64 GB | ||
Table 2.
Representative systems organized by interaction-design dimensions for AI smart glasses.
| Interaction dimension |
Design focus | Representative systems | Main constraints |
|---|---|---|---|
| Intent capture | Speech input | WearVox [60]; EchoSpeech [183] | Noise, addressability, exposure |
| Gaze input | GazeTrak [184]; EyeGesener [185] | Calibration, false activation | |
| Touch and gesture input | FingerGlass [186]; Helios [187] | Power, learnability, fatigue | |
| Reference grounding | Situated grounding | GazePointAR [54]; Gazeify Then Voiceify [166] | Deixis, confirmation |
| Memory grounding | SpeechLess [163]; Persistent Assistant [167] | Personalization, repair | |
| Feedback design | Channel selection | Zhao et al. [189]; GlassMessaging [188] | Attention load, accessibility |
| Repair-oriented feedback | Persistent Assistant [167]; Gazeify Then Voiceify [166] | Confirmation, trust calibration | |
| Proactive interaction | Trigger modeling | AiGet [49]; SpeechLess [163] | Timing, context reliability |
| Intervention decision | AI4Service [164]; Persistent Assistant [167] | Interruption, privacy, control | |
| Inclusive usability | User diversity and accessibility | Zhao et al. [189]; Guo et al. [190] | Cognitive load, comfort |
| Privacy, transparency, and control | Lens Privacy [191] | Consent, legibility, social acceptability |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.