Submitted:
29 June 2026
Posted:
02 July 2026
You are already at the latest version
Abstract
As Large Language Models (LLMs) become the cognitive cores of embodied agents, traditional text-centric safety alignments prove critically insufficient; semantic hallucinations and adversarial jailbreaks bypass digital filters to precipitate catastrophic kinetic hazards. This survey presents the first NLP-centric examination of safety and alignment for language-conditioned embodied agents. We introduce a kinetically-grounded threat taxonomy, demonstrating how robotic embodiments physically amplify linguistic vulnerabilities across white-box, black-box, and cross-modal attacks. To mitigate these threats, we propose a defense-in-depth architecture synergizing internal representation engineering with dual semantic-physical grounding. Finally, we establish a unified cognitive error taxonomy and collate diagnostic benchmarks targeting logic consistency and abstention, outlining a critical roadmap toward physical-consequence-aware alignment.
Keywords:
large language models (LLMs)
; embodied agents
; physical safety
; AI alignment
; adversarial attacks
; vision-language-action (VLA) models
; semantic-to physical amplification
; defense-in-depth
; representation engineering
; dual grounding
; affordance reasoning
; jailbreaking
; cognitive fault taxonomy
; robotics safety
; physical-consequence-aware alignment
1. Introduction
The integration of Large Language Models (LLMs) and Vision-Language Models (VLMs) into embodied agents marks a profound paradigm shift: language is no longer physically inert Ahn et al. (2022). By translating open-vocabulary semantics directly into kinetic actuation, foundation models expose a critical epistemic gap between fluent text generation and rigid physical constraints Soh and Lim (2026). This transition fractures the core assumption of traditional NLP safety. While conversational jailbreaks or hallucinations merely cause digital confusion, embodied vulnerabilities undergo immediate semantic-to-physical amplification, precipitating irreversible real-world hazards Li et al. (2026a). Furthermore, current alignment paradigms (e.g., RLHF Ouyang et al. (2022), DPO Rafailov et al. (2023)) prioritize toxic text suppression, often defaulting to over-refusal. In dynamic physical settings, such text-centric "harmlessness" introduces the fatal risk of inaction—an agent politely refusing an instruction while freezing on a busy highway remains a catastrophic threat. Consequently, semantic safety must evolve from lexical filtering to rigorous kinetic alignment.
This survey introduces a core thesis to bridge this gap: the physical consequences of language-level attacks shatter the assumption of physical compliance, as semantic hallucinations translate instantly into kinetic hazards. A singular adversarial prompt—such as coercing a planner to bypass spatial constraints—undergoes physical amplification, yielding divergent kinetic outcomes. It may cause a manageable path deviation in a wheeled robot, trigger severe fall fragmentation in a bipedal humanoid, or result in high-energy collisions for a fixed-base manipulator. Securing embodied agents therefore demands a shift from text-centric filtering to a physically-grounded safety paradigm, one that accounts for the specific kinetic failure modes of each embodiment.
We present a comprehensive, NLP-centric examination of safety and alignment for language-driven embodied agents, structured around a unified taxonomy (Figure 2). By synthesizing adversarial red-teaming, representation engineering, and 3D physical grounding, our contributions are threefold:
- 1.
- A Semantic-to-Physical Threat Surface: We map how cognitive layer exploits bypass digital boundaries to trigger embodiment-specific physical vulnerabilities, demonstrating that the same linguistic attack yields different catastrophic outcomes depending on the robot’s form-factor.
- 2.
- An Integrated Defense-in-Depth Architecture: As illustrated in Figure 4, we construct a layered defense that synergizes intrinsic white-box interventions with dual semantic-physical grounding to robustly filter semantic hallucinations.
- 3.
- A Cognitive Error Evaluation Framework: Moving beyond mechanical routing metrics, we propose a unified error taxonomy and collate NLP-centric benchmarks (e.g., FoMER Dissanayake et al. (2025), SAFEL Son et al. (2025)) to rigorously audit logic consistency and abstention recall.
As previewed in Figure 1 and detailed in our taxonomy (Figure 2), the remainder of this paper is structured as follows: §2 deconstructs language grounding vulnerabilities. §3 analyzes morphology-amplified attacks. §4 surveys defense architectures. §5 establishes our unified error taxonomy and collates diagnostic benchmarks. §6 outlines future trajectories for physical-consequence-aware alignment.
Figure 1.
Teaser: Safety and Alignment of Language-Conditioned Embodied Agents. Panel ① maps cognitive-layer exploits to task paradigms and attacks (§2, §3); Panel ② illustrates semantic-to-physical amplification and the defense-in-depth architecture (§4); Panel ③ shows how identical language attacks produce divergent physical outcomes across robot embodiments (§3.2). The bottom bar previews the fault taxonomy and diagnostic benchmarks (§5.1–§5.2).
Figure 1.
Teaser: Safety and Alignment of Language-Conditioned Embodied Agents. Panel ① maps cognitive-layer exploits to task paradigms and attacks (§2, §3); Panel ② illustrates semantic-to-physical amplification and the defense-in-depth architecture (§4); Panel ③ shows how identical language attacks produce divergent physical outcomes across robot embodiments (§3.2). The bottom bar previews the fault taxonomy and diagnostic benchmarks (§5.1–§5.2).

Figure 2.
Comprehensive Taxonomy of Vulnerabilities, Attacks, and Defenses in Language-Conditioned Embodied AI.
Figure 2.
Comprehensive Taxonomy of Vulnerabilities, Attacks, and Defenses in Language-Conditioned Embodied AI.

2. Language Task Paradigms and Grounding Challenges
As language grounding remains a primary bottleneck in embodied AI Lisondra et al. (2026), we categorize task paradigms by interactivity to analyze how linguistic ambiguity manifests as specific physical failure modes.
2.1. Language-Driven Task Paradigms
Rather than treating language grounding merely as a functional bottleneck Lisondra et al. (2026), we categorize task paradigms by their interactivity to expose a structural misalignment: how discrete linguistic ambiguities deterministically collapse into continuous physical failures.
- Single-Turn Tasks. A solitary command forces autonomy without corrective feedback. Goal-oriented tasks (ALFRED Shridhar et al. (2020), REVERIE Qi et al. (2020)) demand planning under uncertainty, compelling agents to resolve latent planning burdens by hallucinating missing sub-steps. Yet, this grounding difficulty is non-monotonic: Huang et al. (2026) reveals a “Complexity Paradox” where coarser instructions paradoxically improve success rates by triggering shallow, vision-dominant policies. This exposes a critical vulnerability—language is often not causally grounding behavior, but acting as a superficial prior that the agent easily bypasses. Route-oriented tasks (Room-to-Room Chang et al. (2017)) enforce strict sequential dependencies, where a single misgrounded landmark shatters the execution chain.
- Multi-Turn Tasks. Progressive instruction delivery introduces an interaction axis. Interactive tasks (JustAsk Chi et al. (2020), RobotSlang Banerjee et al. (2021)) mitigate ambiguity by permitting clarification, shifting the cognitive burden to the ask-or-act dilemma: querying incurs overhead, but acting on uncertainty risks catastrophic grounding errors. However, overall task success often masks brittle interaction logic; thus, Zorzi et al. (2026) forcefully advocates for decoupling interactive reasoning from navigation metrics, arguing that true collaborative intelligence requires calibrated, uncertainty-driven communication rather than brute-force trial and error. Conversely, Unidirectional execution tasks (CEREALBAR Suhr et al. (2019)) strictly forbid queries, precipitating cumulative misalignment.
2.2. Vulnerabilities in Language Grounding
In Vision-Language-Action (VLA) systems, NLP errors transcend text to become kinetic hazards Li et al. (2026a). This semantic-to-physical transmission exposes a fatal tension between low-dimensional linguistic abstraction and rigid physical constraints, driven by three structural vulnerabilities.
- Referential Ambiguity and Hallucination. Beyond mere scene-task inconsistencies—which inflate object hallucinations by 40× Chakraborty et al. (2025)—the deeper crisis lies in action hallucinations Soh and Lim (2026). Rather than mere category confusion, this exposes a fundamental topological barrier: mapping continuous latent priors to disconnected physical free spaces unavoidably creates “seams” of invalid actions (e.g., grasping through solids). Consequently, generative VLAs inherently synthesize affordance violations that bypass digital safeguards.
- Spatial Relation Deficits. Language fundamentally bottlenecks 6-DoF geometric reality. State-of-the-art models systematically fail on metric-semantic queries Padhan et al. (2026), producing plans that are semantically fluent yet geometrically disastrous due to unresolved frame-of-reference ambiguities. Alarmingly, static benchmarks mask this spatial blindness: VLAs exhibit semantic feature collapse Xu et al. (2026), exploiting lexical-kinematic shortcuts that instantly shatter under causal layout shifts, exposing apparent spatial comprehension as a fragile statistical illusion.
- Cumulative Instruction Deviation. In long-horizon tasks, early semantic misalignments do not merely persist; they metastasize. Formalized as cascading failures Zeng et al. (2026) and identified as the definitive VLA bottleneck Zhang et al. (2026b), these compounding errors expose severe flaws in autoregressive physical control. Confronted with physical dissonance, models regress to behavioral inertia Xu et al. (2026) and “over-reasoning” Lim et al. (2026)—consuming tokens to rationalize failed states rather than adapting. This epistemic drift traps the agent in a self-justifying narrative, where fluent text generation masquerades as physical planning, irreversibly decoupled from kinetic reality.
3. Adversarial Language Attacks with Morphology-Aware Consequences
In standard NLP, adversarial attacks terminate at toxic text generation. In embodied AI, semantic jailbreaks bypass digital boundaries to become irreversible kinetic hazards—a paradigm-shifting threat in Vision-Language-Action (VLA) deployment Li et al. (2026a,b). This semantic-to-physical transmission exploits a profound output-action mismatch Zhang et al. (2025b): models may maintain linguistic safety while simultaneously leaking hazardous physical commands. We deconstruct these attack vectors, tracing how each subverts the VLA’s cognitive alignment to dictate asymmetric physical consequences.
3.1. The Cognitive Layer as an Attack Entry Point
Adversaries target embodied planners along a spectrum of access and modality, each exploiting a distinct cognitive vulnerability.
- White-Box Reasoning Layer Attacks. When adversaries breach the inference pipeline, they bypass semantic reasoning to surgically manipulate the planner’s mathematical substrate. Logits-based hijacking (e.g., VulMine Li et al. (2025)) reveals that safety alignments merely suppress, rather than erase, harmful priors, making them easily amplified. Furthermore, this cognitive subversion transcends language: FreezeVLA Wang et al. (2025b) demonstrates that adversarial visual noise can deterministically trap agents in an “action-freezing” state. This severs the digital mind from physical execution, weaponizing persistent inaction to bypass standard safety monitors entirely.
- Black-Box Adversarial Prompts. Without weight access, attackers weaponize the VLA’s foundational mandate for compliance. Static obfuscations like nested narratives (DeepInception Li et al. (2023a)) exploit the model’s “self-losing under authority” to bypass filters. Meanwhile, automated red-teaming (RoboPAIR Robey et al. (2025)) and voice-based exploits (BadRobot Zhang et al. (2025b)) expose a fatal conceptual deception: because VLAs rely on statistical token matching rather than true ethical reasoning, they seamlessly rationalize consequentially identical but lexically distinct malicious acts. The very autoregressive fluency that drives brilliant physical planning thus enables systematic kinetic sabotage.
- Cross-Modality Semantic Injection. Multimodal architectures enable adversaries to deploy perceptual Trojan horses, embedding malicious payloads within visual Gong et al. (2025); Ma et al. (2024) or acoustic Shen et al. (2024) streams. These perturbations trigger severe cross-modal representation drift, bypassing linguistic safeguards to directly corrupt physical affordances. As revealed by the AgentSafe benchmark Ying et al. (2025b), agents often successfully perceive environmental hazards visually but critically fail to translate this awareness into safe planning. This systemic perception-execution disconnect highlights the profound inadequacy of text-centric defensive paradigms for robust embodied control.
Figure 4.
The Defense-in-Depth Architecture for Language-Conditioned Embodied AI. Adversarial inputs are systematically intercepted across four modalities. Pre-processing and post-guardrails (§4.3) provide peripheral shielding, while intrinsic latent space alignment (§4.1) and robust dual semantic-physical grounding (§4.2) form the core cognitive defenses. Errors bypassing early layers are increasingly filtered by physical and kinematic realities.
Figure 4.
The Defense-in-Depth Architecture for Language-Conditioned Embodied AI. Adversarial inputs are systematically intercepted across four modalities. Pre-processing and post-guardrails (§4.3) provide peripheral shielding, while intrinsic latent space alignment (§4.1) and robust dual semantic-physical grounding (§4.2) form the core cognitive defenses. Errors bypassing early layers are increasingly filtered by physical and kinematic realities.

3.2. Physical Amplification of Embodied Attacks
Classical NLP alignment falsely assumes embodiment-agnostic consequences. In Embodied AI, a singular adversarial payload undergoes physical amplification: its kinetic manifestation is strictly governed by the agent’s morphology, converting identical semantic errors into asymmetric, irreversible hazards:
- Industrial Kinematic Collisions: In rigid manipulators, semantic jailbreaks bypass high-level filters to disregard designated safety zones Li et al. (2026a). This actively weaponizes the robotic embodiment, translating malicious language tokens into aggressive kinetic strikes against humans Zhang et al. (2025b).
4. Defense: Alignment, Grounding, and Purification
To prevent semantic manipulation from manifesting as irreversible physical harm, embodied agents demand multi-layered safeguards resolving the inherent safety-latency paradox. As depicted in Figure 4, we formalize a decoupled defense-in-depth architecture Li et al. (2026a,b) intercepting adversarial inputs across modalities. This unified pipeline comprises intrinsic latent alignment (§4.1), dual semantic-physical grounding (§4.2), and peripheral purification encompassing input sanitization and execution guardrails (§4.3).
4.1. Intrinsic Alignment and Inference-Time Guardrails
Traditional alignment (RLHF Ouyang et al. (2022), DPO Rafailov et al. (2023)) exposes a fatal cross-domain safety misalignment in robotics: penalizing toxic text fundamentally ignores kinetic consequences. Mere textual refusal is dangerously insufficient; cognitive freezing mid-task transforms into physical peril. True embodied alignment must map unsafe directives to safe fallback physical actions. The SAFEL benchmark Son et al. (2025) exposes this crisis, proving text-aligned LLMs completely fail at transition modeling, routinely hallucinating fatal execution paths for contextually subtle, non-toxic hazards.
Bypassing the catastrophic forgetting of full fine-tuning, the frontier shifts from behavioral pruning to inference-time representation intervention. While RAIN Li et al. (2023c) achieves heuristic self-correction without weight updates, CEE Yang et al. (2025a) and LatentGuard Shu et al. (2025) prove that steering activations via SLERP-based subspace rotation and disentangled latent supervision selectively enforces safety boundaries. This reveals a profound insight: embodied safety is fundamentally a representation geometry problem. Planners already encapsulate safety priors; robust defense merely requires geometric extraction at inference, proving true alignment is a structural property of the latent space, rather than an aftermarket behavioral patch.
4.2. Dual Semantic-Physical Grounding
To prevent physical hallucinations, linguistic commands demand dual grounding—a necessity formally verified by SENTINEL’s Zhan et al. (2025) multi-level temporal logic (LTL/CTL) evaluations.
- 3D Scene Representations as Spatial Anchors. Raw 2D perception is epistemically fragile. Enforcing geometric consensus via 3D Gaussian Splatting Matsuki et al. (2024); Qin et al. (2024) mitigates semantic point ambiguity and inherently filters view-dependent visual trojans. Yet, grounding is not an absolute panacea: StealthAttack Ke et al. (2025) reveals that density-guided point injection easily poisons these explicit representations. This exposes a recursive epistemic vulnerability: while we rely on 3D semantics to anchor physical reality, the spatial anchor itself can be seamlessly and adversarially rewritten.
- Affordance Reasoning as a Kinematic Filter. Perfect geometry does not guarantee kinetic feasibility. Affordance reasoning bridges this gap by mapping abstract semantics into structured geometric and positional flows Su et al. (2025). Crucially, the ADAPT benchmark Chen et al. (2026a) exposes execution vulnerabilities through resource-level affordance failures. By testing agents against absent physical prerequisites (e.g., unavailable ovens), it proves that true physical alignment demands proactive environmental verification rather than relying on ungrounded semantic priors. By proactively filtering kinematically invalid commands Ju et al. (2025); Qian et al. (2024), affordance serves as a strict mechanical safeguard, recalibrating the LLM’s illusion of control against genuine physical boundaries.
4.3. Prompt and Input Purification
When latent alignment fails, sanitization must intercept adversarial payloads before kinetic execution.
4.3.1. Preprocessing and Neutralization
Because autoregressive fluency lets models rationalize obfuscated sabotage Yuan et al. (2024b), defenses must decouple intent from syntax. Rather than naive filtering, ASF Khachaturov and Mullins (2025) excises out-of-distribution suffixes, while DR-Smoothing Lin et al. (2026) pairs local disruption with global rectification to restore inputs to safe in-distribution manifolds. Structurally, unimodal sanitization ignores embodied threats. Since visual noise induces cross-modal representation drift Yan et al. (2025), ECSO Gou et al. (2025) bypasses visual Trojans via image-to-text transformation to reactivate linguistic safety priors, while AdaShield Wang et al. (2024) prepends adaptive shield prompts to enforce holistic safety verification across both modalities.
4.3.2. Post-Processing Execution Guardrails
As a final barrier, execution guardrails intercept latent hazards before kinetic actuation. Instead of heuristic filtering, SmoothLLM Robey et al. (2023) applies randomized input smoothing, while LSD Phute et al. (2023) leverages zero-shot self-examination for harm detection. In embodied contexts, RoboGuard Ravichandran et al. (2026) formally synthesizes open-vocabulary rules into Signal Temporal Logic (STL), and RoboSafe Wang et al. (2025a) generates executable safety predicates via bidirectional memory reasoning. This exposes a necessary architectural paradox: the cognitive planner is inherently treated as an untrusted adversary, enforcing a strict epistemic decoupling where high-level semantic fluency is subordinated to formal downstream containment.
4.4. Cross-Modal Consistency Verification
Beyond semantic constraints, kinetic safety demands strict consensus across independent sensor streams. While adversarial payloads can spoof visual perception via optimized visual perturbations Wang et al. (2025b), they mathematically struggle to simultaneously corrupt orthogonal metric modalities (e.g., LiDAR or proprioception). To operationalize this asymmetry, embodied architectures deploy decoupled dual-loop mechanisms Li et al. (2026a). When the slow, high-level VLA planner hallucinates a semantically safe but physically disastrous trajectory, the high-frequency “fast reflex” loop utilizes Control Barrier Functions (CBFs) to trigger immediate kinematic overrides. This strict sensorimotor decoupling ensures language-driven statistical illusions never override rigid geometric realities.
5. Evaluation: Unified Taxonomy and Diagnostic Benchmarks
5.1. A Unified Error Taxonomy for Embodied Language Agents
Transcending reductive binary success metrics that mask compensatory over-reasoning Lim et al. (2026); Xu et al. (2026) and evade text-only safety protocols Ying et al. (2025b), we decompose embodied vulnerabilities into three epistemic domains, tracing the structural degradation from latent linguistic intent to physical execution.
- Cognitive and Semantic Faults (“Brain” Failures). Originating pre-execution, these errors reflect severe implicit intent misresolution Lim et al. (2026). RoboInspector Ying et al. (2025a) formalizes this structural collapse, proving that instruction granularity deficits and cognitive limits drive models to generate Disorder and Nonsense plans. This exposes a profound epistemic flaw: LLMs merely approximate physical syntax through compensatory over-reasoning, blindly generating ungrounded scripts that expose their semantic comprehension as a statistical illusion.
- Kinematic and Execution Faults (“Body” Failures). Logically fluent plans predictably shatter against physical friction. Formalized within a goal-conditioned POMDP framework Garrabé et al. (2025), Badpose and Infeasible Ying et al. (2025a) execution errors expose disembodied planners ignoring continuous kinematics. These breakdowns reveal a dangerous paradigm: the cognitive layer assumes absolute physical compliance, fatally treating the robotic embodiment as a deterministic text-renderer.
- Contextual and Integration Faults (“Nervous System” Failures). These sensorimotor disconnects between latent plans and dynamic environments trap agents in cascading failures. Moving beyond static simulators, AgentSafe Ying et al. (2025b) and HazardArena Chen et al. (2026b) expose a profound alignment crisis: agents successfully perceiving hazards still execute fatal actions under semantic shifts. This cross-modal decoupling reveals that while models statistically master what to do, they remain dangerously oblivious to when they must abstain.
5.2. Diagnostic Dimensions and Systemic Benchmark Gaps
Table 1 maps benchmarks across seven diagnostic dimensions within our System 2-Immune-Somatic framework (metric details in Appendices A1 and A2). Yet, a topological analysis reveals a profound systemic flaw: an overwhelming epistemic bias toward static semantic reasoning at the expense of dynamic kinematic integration, masking severe vulnerabilities in continuous physical control.
- The Bias Toward Static “Brain” Evaluation. While foundational functions are densely covered, they operate in a vacuum. REI-Bench Jiang et al. (2025) and ECBench Dang et al. (2025) effectively expose coreferential vagueness and robot-centric spatial blindness. Concurrently, hallucination diagnostics (MIRAGE Zhang et al. (2025c), AbstainEQA Wu et al. (2025)) map metacognitive boundaries. However, this static paradigm treats agents as passive, disembodied conversationalists. By severing semantic reasoning from active execution, these metrics fail to capture how linguistically “safe” plans predictably collapse into physical violations under mechanical friction.
- The Sparsity of Cognitive Consistency. Furthermore, evaluations for cognitive consistency remain alarmingly sparse. Over extended horizons, autoregressive planners undergo severe epistemic drift, fatally treating rigid physical states as mutable text. Although FoMER Dissanayake et al. (2025) scrutinizes reasoning trail correctness, PRISM Lim et al. (2026) explicitly isolates memory retention failures, and RoboCerebra Han et al. (2026) evaluates System-2 performance in dynamically evolving environments, dynamic red-teaming is virtually absent. Consequently, we cannot audit how early semantic misalignments metastasize into cascading kinetic failures, masking the fundamental inability to maintain a causal world model.
- The Nascent Frontier: End-to-End Physical Consequences. The most glaring revelation is the severe underdevelopment of diagnostics for Semantic-Physical Sync and adversarial Physical Feasibility, exposing the historical neglect of the “Somatic System”. While pioneering frameworks like RoboBench Luo et al. (2025) have finally begun to measure embodied feasibility beyond symbolic matching, and HazardArena Chen et al. (2026b) introduces controlled semantic risk scenarios, they represent exceptions rather than the rule. The vast majority of existing metrics still validate the linguistic plausibility of generated plans but ignore true physical execution limits. By failing to broadly test whether language-driven trajectories trigger cross-modal conflicts or synchronize with high-frequency proprioceptive feedback, the field dangerously equates statistical text generation with genuine sensorimotor intelligence.
6. Conclusion
Integrating foundation models into robotic systems highlights a critical limitation in current AI alignment: linguistic harmlessness does not guarantee physical safety. Standard textual safeguards fail to protect against semantic ambiguities or spatial misunderstandings that can result in real-world kinetic hazards. By reframing embodied alignment to focus on physical consequences, this survey demonstrates that robust defense requires tightly coupling the model’s cognitive constraints with mechanical realities. This is achieved by combining targeted internal representation engineering with dynamic semantic-physical grounding, and by evaluating system reliability through structured cognitive fault taxonomies rather than relying solely on binary task success rates.
Limitations
This survey focuses on language-conditioned embodied agents, excluding purely deterministic or non-semantic reinforcement learning systems. Consequently, our threat taxonomy emphasizes cognitive reasoning failures over low-level mechanical faults.
Additionally, embodied AI safety faces a pronounced Sim2Real gap. Most reviewed benchmarks and defenses are validated in simulations, which cannot fully replicate physical hardware variables like kinetic or sensor noise. Finally, given the rapid evolution of foundation models, the attack vectors detailed herein represent a contemporary snapshot; emergent architectures may introduce novel vulnerabilities beyond our current taxonomy.
Potential Risks and Ethical Considerations
Detailing advanced adversarial attacks (e.g., latent space hijacking, prompt jailbreaking) introduces a dual-use risk, as malicious actors could weaponize these exploits to induce physical harm. We mitigate this by pairing our threat analysis with defense-in-depth strategies, emphasizing that exposing these flaws is essential for developing resilient guardrails.
Furthermore, deploying language-driven agents in high-stakes environments introduces profound accountability challenges. If an agent causes physical damage due to a semantic hallucination, attributing liability among model providers, manufacturers, and users remains unresolved. We urge the community to prioritize physically grounded safety evaluations prior to real-world deployment.
Table A1.
Detailed benchmark metrics and implementation methods for cognitive, safety, hallucination, and abstention dimensions (Part I & II of Table 1).
Table A1.
Detailed benchmark metrics and implementation methods for cognitive, safety, hallucination, and abstention dimensions (Part I & II of Table 1).
| Benchmark | Metrics Category | Specific Metric Name | Implementation Method |
|---|---|---|---|
| Part I: Cognitive & Semantic Reasoning | |||
| REI-Bench Jiang et al. (2025) | Instruction Comprehension | Coreferential vagueness resolution | Simulation-based task success rate rollout in AI2-THOR |
| RoboBench Luo et al. (2025) | Instruction Comprehension | Explicit/Implicit intention parsing | MLLM-as-world-simulator for embodied feasibility validation |
| Perception Reasoning | Robotic/Object/Scene-centric cognition | Multiple-choice visual question answering accuracy | |
| ECBench Dang et al. (2025) | Perception Reasoning | Robot-centric & Scene-based cognition | VideoQA evaluated via ECEval (GPT-4o multi-level scoring) |
| EmbodiedBench Yang et al. (2025b) | Instruction Comprehension | Intent misinterpretation | Simulator-based holistic task success rate |
| Perception Reasoning | Spatial misalignment | Object-centric spatial relation checks in simulation | |
| Cognitive Consistency | Logic decay | Rule-based temporal consistency verifier over action sequence | |
| ALFRED Shridhar et al. (2020) | Instruction Comprehension | Goal misalignment, unseen template failures | Simulator strict state checking (Goal-Condition Success) |
| Cognitive Consistency | Subgoal coherence, plan stability | Path-weighted task success (PLW) in simulation | |
| TEACh Min et al. (2022) | Instruction Comprehension | Dialogue history dependence (Teacher-forcing bias) | Action prediction accuracy with vs. without conversational history |
| AgentBench Liu et al. (2024) | Instruction Comprehension | Interactive task success rate | Rule-based verification of environment states (e.g., OS, Database) |
| Cognitive Consistency | Turn Limit Exceeded (TLE) rate | Interaction round counting before task termination | |
| FoMER Dissanayake et al. (2025) | Cognitive Consistency | Reasoning alignment, Missing step detection | LLM-based and human evaluation against ground-truth reasoning chains |
| RoboCerebra Han et al. (2026) | Cognitive Consistency | Memory exploration success, Decision accuracy | Simulator-based long-horizon hierarchical subtask execution |
| CALVIN Mees et al. (2022) | Cognitive Consistency | Sequential task success rate | Simulated continuous control roll-out evaluating sequence completion |
| VIMA-Bench Jiang et al. (2023) | Instruction Comprehension | Multimodal prompt generalization | Simulation-based task success across L1-L4 splits |
| Perception Reasoning | Novel visual concept grounding | Zero-shot execution success on unseen objects | |
| Ego4D Grauman et al. (2022) | Perception Reasoning | Episodic memory retrieval | Spatio-temporal localization error for visual queries |
| Cognitive Consistency | Long-horizon action forecasting | Edit distance and mAP for future action anticipation | |
| Part II: Safety, Hallucination & Abstention | |||
| MIRAGE Zhang et al. (2025c) | Hallucination Detection | Utility Score (US), Hallucination Rate (HR) | Contextual snapshot evaluation via LLM-as-a-Judge (o4-mini) with risk-aware prompts |
| SAFE Gu et al. (2026) | Failure Detection | Failure detection accuracy | Learning to detect anomalies/failures directly from observation history and instructions |
| POPE Li et al. (2023b) | Hallucination Detection | F1 Score on object hallucination | Polling-based Yes/No visual QA for binary object presence classification |
| MMHal-Bench Sun et al. (2024) | Hallucination Detection | Overall Hallucination Score & Rate | GPT-4-based open-ended response scoring across 8 hallucination sub-categories |
| AbstainEQA Wu et al. (2025) | Abstention & Uncertainty | Abstention rate on ambiguous queries | Benchmarking agent abstention behavior on unanswerable or noisy visual questions |
| SAFEL Son et al. (2025) | Abstention & Uncertainty | Refusal Recall, Transition Modeling | Evaluating safe fallback and physical execution refusal on contextually subtle hazards |
| SAFEL Son et al. (2025) | Abstention & Uncertainty | Refusal Recall, Transition Modeling | Evaluating safe fallback on contextually subtle hazards |
| KnowNo Ren et al. (2023) | Abstention & Uncertainty | Statistical help-asking rate | Conformal prediction sets providing formal guarantees for task success |
| HAZARD Zhou et al. (2024) | Abstention & Uncertainty | System damage rate, Rescue success | Dynamic simulation (fire/flood/wind) testing temporal hazard awareness |
| SENTINEL Zhan et al. (2025) | Kinetic Guardrails | State invariance & reachability | Multi-level formal verification (LTL/CTL) over plans and trajectories |
| R-Judge Yuan et al. (2024a) | Safety Monitoring | Safety judgment F1, Risk identification | Auditing LLMs on human-annotated multi-turn interaction records |
| Safety-Gym Ray et al. (2019) | Kinetic Guardrails | Constraint violation cost | CMDP-based evaluation enforcing strict spatial hazard boundaries |
| Physical Feasibility | Collision-free locomotion | Continuous control simulation penalizing unsafe state transitions | |
Table A2.
Detailed benchmark metrics and implementation methods for kinetic actuation and sensorimotor sync dimensions (Part III of Table 1).
Table A2.
Detailed benchmark metrics and implementation methods for kinetic actuation and sensorimotor sync dimensions (Part III of Table 1).
| Benchmark | Metrics Category | Specific Metric Name | Implementation Method |
|---|---|---|---|
| Part III: Kinetic Actuation & Sensorimotor Sync (The “Somatic System”) | |||
| SENTINEL Zhan et al. (2025) | Physical Feasibility | State invariance & reachability | Multi-level formal verification (LTL/CTL) over continuous trajectories |
| ADAPT Chen et al. (2026a) | Physical Feasibility | Unspecified affordance constraints | Commonsense planning evaluation across domain-adapted VLMs |
| POMDP Garrabé et al. (2025) | Physical Feasibility | Subgoal mischaracterization | Expected outcome mismatch evaluation via goal-conditioned POMDP |
| RLBench James et al. (2020) | Physical Feasibility | Kinematic-aware task success | Continuous control roll-out evaluating strict joint-space constraints |
| LIBERO Liu et al. (2023a) | Physical Feasibility | Lifelong knowledge transfer | Continuous control evaluation across sequentially non-stationary task suites |
| ManiSkill Gu et al. (2023) | Physical Feasibility | Generalizable physical skill success | End-to-end execution rollout with fully physical grasp constraints |
| RoboTHOR Deitke et al. (2020) | Perception Reasoning | Epistemic mapping in clutter | Object-goal navigation under severe visual occlusion |
| Physical Feasibility | Embodied kinematic efficiency | Success weighted by Path Length (SPL) in complex topologies | |
| Semantic-Physical Sync | Cross-domain sensorimotor gap | Direct performance variance auditing across physical and digital twins | |
| VLA-Arena Zhang et al. (2025a) | Physical Feasibility | Dynamic obstacle & state preservation | Executing strict kinetic constraint checking across extrapolative tasks |
| VLSA Hu et al. (2025) | Kinetic Guardrails | Constraint layer violation rate | Enforcing runtime safe exploration via plug-and-play constraint modules |
| HazardArena Chen et al. (2026b) | Semantic-Physical Sync | Safety intention vs. feasibility | Evaluating action trajectories across physical and psychosocial hazards |
| AI2-THOR Kolve et al. (2017) | Perception Reasoning | Causal state transitions | Simulating actionable properties (e.g., thermal, slicing) under interaction |
| Semantic-Physical Sync | Continuous extended state tracking | Verifying state-action coherence across prolonged manipulation chains | |
| Habitat Szot et al. (2021) | Perception Reasoning | Kinematic scene rearrangement | Success weighted by Path Length (SPL) in interactive tasks |
| Semantic-Physical Sync | High-frequency dynamic control | Continuous rigid-body physics simulation roll-outs | |
| BEHAVIOR Srivastava et al. (2022) | Instruction Comprehension | Initial-to-goal logic state grounding | BDDL-based predicate evaluation (e.g., sliced, cooked) |
| Physical Feasibility | Long-horizon execution viability | Simulation-based evaluation across 1,000+ daily activities | |
| Semantic-Physical Sync | Continuous physical state transitions | Tracking extended properties (e.g., thermal, wetness) | |
| SAPIEN Xiang et al. (2020) | Physical Feasibility | Articulated part-level manipulation | Evaluating physically realistic joint limits and actuation |
| Semantic-Physical Sync | Causal kinematic interactions | Part-state transitions within strict physical constraints | |
| SayCan Ahn et al. (2022) | Instruction Comprehension | Affordance-grounded probability | Multiplying LLM intent likelihood with learned value functions |
| Physical Feasibility | Zero-shot physical execution | End-to-end real-world robotic task completion rate | |
| Mobile ALOHA Fu et al. (2024) | Physical Feasibility | Non-holonomic bimanual coordination | Real-world task success via whole-body imitation learning |
| Semantic-Physical Sync | Dynamic sensorimotor execution | End-to-end trajectory roll-outs in unstructured environments | |
| Open X-Embodiment O’Neill et al. (2024) | Physical Feasibility | Morphological generalization | Positive transfer evaluation across distinct kinematic structures |
| Cognitive Consistency | Cross-embodiment policy adaptation | Zero-shot execution success across diverse robot platforms | |
References
- Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, and 1 others. 2022. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691. [CrossRef]
- Shurjo Banerjee, Jesse Thomason, and Jason Corso. 2021. The robotslang benchmark: Dialog-guided robot localization and navigation. In Conference on Robot Learning, pages 1384–1393. PMLR.
- Trishna Chakraborty, Udita Ghosh, Xiaopan Zhang, Fahim Faisal Niloy, Yue Dong, Jiachen Li, Amit K Roy-Chowdhury, and Chengyu Song. 2025. Heal: An empirical study on hallucinations in embodied agents driven by large language models. arXiv preprint arXiv:2506.15065. [CrossRef]
- Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. 2017. Matterport3d: Learning from rgb-d data in indoor environments. arXiv preprint arXiv:1709.06158. [CrossRef]
- Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. 2025. Jailbreaking black box large language models in twenty queries. In 2025 IEEE Conference on Secure and Trustworthy Machine Learning, pages 23–42, Copenhagen, Denmark. IEEE. [CrossRef]
- Pei-An Chen, Yong-Ching Liang, Jia-Fong Yeh, Hung-Ting Su, Yi-Ting Chen, Min Sun, and Winston Hsu. 2026a. Adapt: Benchmarking commonsense planning under unspecified affordance constraints. arXiv preprint arXiv:2604.14902. [CrossRef]
- Zixing Chen, Yifeng Gao, Li Wang, Yunhan Zhao, Yi Liu, Jiayu Li, Xiang Zheng, Zuxuan Wu, Cong Wang, Xingjun Ma, and 1 others. 2026b. Hazardarena: Evaluating semantic safety in vision-language-action models. arXiv preprint arXiv:2604.12447. [CrossRef]
- Ta-Chung Chi, Minmin Shen, Mihail Eric, Seokhwan Kim, and Dilek Hakkani-Tur. 2020. Just ask: An interactive learning framework for vision and language navigation. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 2459–2466.
- Xuanming Cui, Alejandro Aparcedo, Young Kyun Jang, and Ser-Nam Lim. 2024. On the robustness of large multimodal models against image adversarial attacks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24625–24634, Seattle, WA, USA. IEEE. [CrossRef]
- Ronghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin, Boqiang Zhang, Long Li, Liuyi Wang, Qinyang Zeng, Xin Li, and Lidong Bing. 2025. Ecbench: Can multi-modal foundation models understand the egocentric world? a holistic embodied cognition benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24593–24602.
- Matt Deitke, Winson Han, Alvaro Herrasti, Aniruddha Kembhavi, Eric Kolve, Roozbeh Mottaghi, Jordi Salvador, Dustin Schwenk, Eli VanderBilt, Matthew Wallingford, and 1 others. 2020. Robothor: An open simulation-to-real embodied ai platform. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3164–3174.
- Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing. 2023. Multilingual jailbreak challenges in large language models. Preprint, arXiv:2310.06474. [CrossRef]
- Dinura Dissanayake, Ahmed Heakl, Omkar Thawakar, Noor Ahsan, Ritesh Thawkar, Ketan More, Jean Lahoud, Rao Anwer, Hisham Cholakkal, Ivan Laptev, and 1 others. 2025. How good are foundation models in step-by-step embodied reasoning? arXiv preprint arXiv:2509.15293. [CrossRef]
- Zipeng Fu, Tony Z Zhao, and Chelsea Finn. 2024. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation. arXiv preprint arXiv:2401.02117. [CrossRef]
- Émiland Garrabé, Pierre Teixeira, Mahdi Khoramshahi, and Stéphane Doncieux. 2025. Enhancing robustness in language-driven robotics: A modular approach to failure reduction. In 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 16717–16724. IEEE. [CrossRef]
- Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang. 2025. Figstep: Jailbreaking large vision-language models via typographic visual prompts. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 23951–23959. [CrossRef]
- Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Hang Xu, Zhenguo Li, Dit-Yan Yeung, James T. Kwok, and Yu Zhang. 2025. Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation. In Computer Vision – ECCV 2024, pages 388–404, Cham. Springer Nature Switzerland. [CrossRef]
- Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, and 1 others. 2022. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 18995–19012.
- Jiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, and 1 others. 2023. Maniskill2: A unified benchmark for generalizable manipulation skills. arXiv preprint arXiv:2302.04659. [CrossRef]
- Qiao Gu, Yuanliang Ju, Shengxiang Sun, Igor Gilitschenski, Haruki Nishimura, Masha Itkina, and Florian Shkurti. 2026. Safe: Multitask failure detection for vision-language-action models. Advances in Neural Information Processing Systems, 38:40041–40076.
- Songhao Han, Boxiang Qiu, Yue Liao, Siyuan Huang, Chen Gao, Shuicheng Yan, and Si Liu. 2026. Robocerebra: A large-scale benchmark for long-horizon robotic manipulation evaluation. Advances in Neural Information Processing Systems, 38.
- Songqiao Hu, Zeyi Liu, Shuang Liu, Jun Cen, Zihan Meng, and Xiao He. 2025. Vlsa: Vision-language-action models with plug-and-play safety constraint layer. arXiv preprint arXiv:2512.11891. [CrossRef]
- Chenguang Huang, Oier Mees, Andy Zeng, and Wolfram Burgard. 2023. Audio visual language maps for robot navigation. In International Symposium on Experimental Robotics, pages 105–117. Springer.
- Sukai Huang, Chenyuan Zhang, Fucai Ke, Zhixi Cai, Gholamreza Haffari, Lizhen Qu, and Hamid Rezatofighi. 2026. Mini-behavior-gran: Revealing u-shaped effects of instruction granularity on language-guided embodied agents. arXiv preprint arXiv:2604.17019. [CrossRef]
- Wensi Huang, Shaohao Zhu, Meng Wei, Jinming Xu, Xihui Liu, Hanqing Wang, Tai Wang, Feng Zhao, and Jiangmiao Pang. 2025. Vl-ln bench: Towards long-horizon goal-oriented navigation with active dialogs. arXiv preprint arXiv:2512.22342. [CrossRef]
- Hakan Inan, Khanh Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. 2023. Llama guard: Llm-based input-output safeguard for human-ai conversations. Preprint, arXiv:2312.06674. [CrossRef]
- Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023. Baseline defenses for adversarial attacks against aligned language models. arXiv preprint arXiv:2309.00614. [CrossRef]
- Stephen James, Zicong Ma, David Rovick Arrojo, and Andrew J Davison. 2020. Rlbench: The robot learning benchmark & learning environment. IEEE Robotics and Automation Letters, 5(2):3019–3026. [CrossRef]
- Chenxi Jiang, Chuhao Zhou, and Jianfei Yang. 2025. Rei-bench: Can embodied agents understand vague human instructions in task planning? arXiv preprint arXiv:2505.10872. [CrossRef]
- Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan. 2023. Vima: Robot manipulation with multimodal prompts.
- Yuanchen Ju, Kaizhe Hu, Guowei Zhang, Gu Zhang, Mingrun Jiang, and Huazhe Xu. 2025. Robo-abc: Affordance generalization beyond categories via semantic correspondence for robot manipulation. In European Conference on Computer Vision, pages 222–239. Springer. [CrossRef]
- Bo-Hsu Ke, You-Zhe Xie, Yu-Lun Liu, and Wei-Chen Chiu. 2025. Stealthattack: Robust 3d gaussian splatting poisoning via density-guided illusions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 27400–27411.
- David Khachaturov and Robert Mullins. 2025. Adversarial suffix filtering: a defense pipeline for llms. arXiv preprint arXiv:2505.09602. [CrossRef]
- Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, and 1 others. 2017. Ai2-thor: An interactive 3d environment for visual ai. arXiv preprint arXiv:1712.05474. [CrossRef]
- Manling Li, Shiyu Zhao, Qineng Wang, Kangrui Wang, Yu Zhou, Sanjana Srivastava, Cem Gokmen, Tony Lee, Li Erran Li, Ruohan Zhang, Weiyu Liu, Percy Liang, Li Fei-Fei, Jiayuan Mao, and Jiajun Wu. 2024. Embodied agent interface: Benchmarking llms for embodied decision making. In Advances in Neural Information Processing Systems, volume 37, pages 100428–100534. Curran Associates, Inc. [CrossRef]
- Qi Li, Bo Yin, Weiqi Huang, Ruhao Liu, Bojun Zou, Runpeng Yu, Jingwen Ye, Weihao Yu, and Xinchao Wang. 2026a. Vision-language-action safety: Threats, challenges, evaluations, and mechanisms. arXiv preprint arXiv:2604.23775. [CrossRef]
- Xiao Li, Xiang Zheng, Yifeng Gao, Xinyu Xia, Yixu Wang, Xin Wang, Ye Sun, Yunhan Zhao, Ming Wen, Jiayu Li, and 1 others. 2026b. Safety in embodied ai: A survey of risks, attacks, and defenses. arXiv preprint arXiv:2605.02900. [CrossRef]
- Xuan Li, Zhanke Zhou, Jianing Zhu, Jiangchao Yao, Tongliang Liu, and Bo Han. 2023a. Deepinception: Hypnotize large language model to be jailbreaker. Preprint, arXiv:2311.03191. [CrossRef]
- Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. 2023b. Evaluating object hallucination in large vision-language models. In The 2023 Conference on Empirical Methods in Natural Language Processing. [CrossRef]
- Yuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang, and Hongyang Zhang. 2023c. Rain: Your language models can align themselves without finetuning. arXiv preprint arXiv:2309.07124. [CrossRef]
- Yuxi Li, Yi Liu, Yuekang Li, Ling Shi, Gelei Deng, Shengquan Chen, and Kailong Wang. 2025. Lockpicking llms: A logit-based jailbreak using token-level manipulation. Preprint, arXiv:2405.13068. [CrossRef]
- Yunn Kang Lim, Pengzhan Sun, Ziyi Bai, Xun Xu, Angela Yao, Xulei Yang, and Shijie Li. 2026. Prism: Planning and reasoning with intent in simulated embodied environments. arXiv preprint arXiv:2605.11534. [CrossRef]
- Jieru Lin, Zhiwei Yu, and Börje F Karlsson. 2025. Switch: Benchmarking modeling and handling of tangible interfaces in long-horizon embodied scenarios. arXiv preprint arXiv:2511.17649. [CrossRef]
- Zheng Lin, Zhenxing Niu, Haoxuan Ji, and Haichang Gao. 2026. Guaranteed jailbreaking defense via disrupt-and-rectify smoothing. arXiv preprint arXiv:2605.10582. [CrossRef]
- Matthew Lisondra, Beno Benhabib, and Goldie Nejat. 2026. Embodied ai with foundation models for mobile service robots: A systematic review. Robotics, 15(3):55.
- Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. 2023a. Libero: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36:44776–44791. [CrossRef]
- Huihan Liu, Alice Chen, Yuke Zhu, Adith Swaminathan, Andrey Kolobov, and Ching-An Cheng. 2023b. Interactive robot learning from verbal correction. arXiv preprint arXiv:2310.17555. [CrossRef]
- Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, and 1 others. 2024. Agentbench: Evaluating llms as agents. In International Conference on Learning Representations, volume 2024, pages 52989–53046.
- Yulin Luo, Chun-Kai Fan, Menghang Dong, Jiayu Shi, Mengdi Zhao, Bo-Wen Zhang, Cheng Chi, Jiaming Liu, Gaole Dai, Rongyu Zhang, and 1 others. 2025. Robobench: A comprehensive evaluation benchmark for multimodal large language models as embodied brain. arXiv preprint arXiv:2510.17801. [CrossRef]
- Siyuan Ma, Weidi Luo, Yu Wang, Xiaogeng Liu, Muhao Chen, Bo Li, and Chaowei Xiao. 2024. Visual-roleplay: Universal jailbreak attack on multimodal large language models via role-playing image character. Preprint, arXiv:2405.20773. [CrossRef]
- Hidenobu Matsuki, Riku Murai, Paul HJ Kelly, and Andrew J Davison. 2024. Gaussian splatting slam. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18039–18048.
- Oier Mees, Lukas Hermann, Erick Rosete-Beas, and Wolfram Burgard. 2022. Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. IEEE Robotics and Automation Letters, 7(3):7327–7334. [CrossRef]
- So Yeon Min, Hao Zhu, Ruslan Salakhutdinov, and Yonatan Bisk. 2022. Don’t copy the teacher: Data and model challenges in embodied dialogue. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9361–9368. [CrossRef]
- Piotr Mirowski, Andras Banki-Horvath, Keith Anderson, Denis Teplyashin, Karl Moritz Hermann, Mateusz Malinowski, Matthew Koichi Grimes, Karen Simonyan, Koray Kavukcuoglu, Andrew Zisserman, and 1 others. 2019. The streetlearn environment and dataset. arXiv preprint arXiv:1903.01292. [CrossRef]
- Khanh Nguyen, Debadeepta Dey, Chris Brockett, and Bill Dolan. 2019. Vision-based navigation with language-based assistance via imitation learning with indirect intervention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12527–12537.
- Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, Arash Vahdat, and Anima Anandkumar. 2022. Diffusion models for adversarial purification. arXiv preprint arXiv:2205.07460. [CrossRef]
- Mahdi Nikdan, Soroush Tabesh, Elvir Crnčević, and Dan Alistarh. 2024. Rosa: Accurate parameter-efficient fine-tuning via robust adaptation. arXiv preprint arXiv:2401.04679. [CrossRef]
- Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744.
- Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew Wang, Anikait Singh, and 260 others. 2024. Open x-embodiment: Robotic learning datasets and rt-x models : Open x-embodiment collaboration0. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903. [CrossRef]
- Paul Pacaud, Ricardo Garcia, Shizhe Chen, and Cordelia Schmid. 2025. Guardian: Detecting robotic planning and execution errors with vision-language models. arXiv preprint arXiv:2512.01946. [CrossRef]
- Swagat Padhan, Lakshya Jain, Bhavya Minesh Shah, Omkar Patil, Thao Nguyen, and Nakul Gopalan. 2026. Meanings and measurements: Multi-agent probabilistic grounding for vision-language navigation. arXiv preprint arXiv:2603.19166. [CrossRef]
- Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gokhan Tur, and Dilek Hakkani-Tur. 2022. Teach: Task-driven embodied agents that chat. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 2017–2025. [CrossRef]
- Mansi Phute, Alec Helbling, Matthew Hull, ShengYun Peng, Sebastian Szyller, Cory Cornelius, and Duen Horng Chau. 2023. Llm self defense: By self examination, llms know they are being tricked. arXiv preprint arXiv:2308.07308. [CrossRef]
- Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. 2020. Reverie: Remote embodied visual referring expression in real indoor environments. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9982–9991.
- Shengyi Qian, Weifeng Chen, Min Bai, Xiong Zhou, Zhuowen Tu, and Li Erran Li. 2024. Affordancellm: Grounding affordance from vision language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7587–7597.
- Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister. 2024. Langsplat: 3d language gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20051–20060.
- Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. CoRR, abs/2305.18290. [CrossRef]
- Zachary Ravichandran, Alexander Robey, Vijay Kumar, George J Pappas, and Hamed Hassani. 2026. Safety guardrails for llm-enabled robots. IEEE Robotics and Automation Letters. [CrossRef]
- Alex Ray, Joshua Achiam, and Dario Amodei. 2019. Benchmarking safe exploration in deep reinforcement learning. arXiv preprint arXiv:1910.01708, 7(1):2. [CrossRef]
- Allen Z Ren, Anushri Dixit, Alexandra Bodrova, Sumeet Singh, Stephen Tu, Noah Brown, Peng Xu, Leila Takayama, Fei Xia, Jake Varley, and 1 others. 2023. Robots that ask for help: Uncertainty alignment for large language model planners. arXiv preprint arXiv:2307.01928. [CrossRef]
- Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018. Semantically equivalent adversarial rules for debugging nlp models. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (volume 1: long papers), pages 856–865. [CrossRef]
- Alexander Robey, Zachary Ravichandran, Vijay Kumar, Hamed Hassani, and George J Pappas. 2025. Jailbreaking llm-controlled robots. In 2025 IEEE International Conference on Robotics and Automation (ICRA), pages 11948–11956. IEEE. [CrossRef]
- Alexander Robey, Eric Wong, Hamed Hassani, and George J Pappas. 2023. Smoothllm: Defending large language models against jailbreaking attacks. arXiv preprint arXiv:2310.03684. [CrossRef]
- Xinyue Shen, Yixin Wu, Michael Backes, and Yang Zhang. 2024. Voice jailbreak attacks against gpt-4o. Preprint, arXiv:2405.19103. [CrossRef]
- Mohit Shridhar, Lucas Manuelli, and Dieter Fox. 2023. Perceiver-actor: A multi-task transformer for robotic manipulation. In Proceedings of The 6th Conference on Robot Learning, volume 205, pages 785–799. PMLR.
- Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2020. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10740–10749.
- Huizhen Shu, Xuying Li, and Zhuo Li. 2025. Latentguard: Controllable latent steering for robust refusal of attacks and reliable response generation. arXiv preprint arXiv:2509.19839. [CrossRef]
- Bruno Siciliano, Oussama Khatib, and Torsten Kröger. 2008. Springer handbook of robotics, volume 200. Springer. [CrossRef]
- Harold Soh and Eugene Lim. 2026. Action hallucination in generative visual-language-action models. arXiv preprint arXiv:2602.06339. [CrossRef]
- Yejin Son, Minseo Kim, Sungwoong Kim, Seungju Han, Jian Kim, Dongju Jang, Youngjae Yu, and Chan Young Park. 2025. Subtle risks, critical failures: A framework for diagnosing physical safety of llms for embodied decision making. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 25703–25744. [CrossRef]
- Sanjana Srivastava, Chengshu Li, Michael Lingelbach, Roberto Martín-Martín, Fei Xia, Kent Elliott Vainio, Zheng Lian, Cem Gokmen, Shyamal Buch, Karen Liu, and 1 others. 2022. Behavior: Benchmark for everyday household activities in virtual, interactive, and ecological environments. In Conference on robot learning, pages 477–490. PMLR.
- Chenyu Su, Weiwei Shang, Chen Qian, Fei Zhang, and Shuang Cong. 2025. Resemact: Advancing fine-grained robotic manipulation via semantic structuring and affordance refinement. arXiv preprint arXiv:2507.18262. [CrossRef]
- Hung-Ting Su, Ting-Jun Wang, Jia-Fong Yeh, Min Sun, and Winston H Hsu. 2026. Vln-nf: Feasibility-aware vision-and-language navigation with false-premise instructions. arXiv preprint arXiv:2604.10533. [CrossRef]
- Alane Suhr, Claudia Yan, Jack Schluger, Stanley Yu, Hadi Khader, Marwa Mouallem, Iris Zhang, and Yoav Artzi. 2019. Executing instructions in situated collaborative interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2119–2130, Hong Kong, China. Association for Computational Linguistics. [CrossRef]
- Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liangyan Gui, Yu-Xiong Wang, Yiming Yang, and 1 others. 2024. Aligning large multimodal models with factually augmented rlhf. In Findings of the Association for Computational Linguistics: ACL 2024, pages 13088–13110. [CrossRef]
- Andrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Singh Chaplot, Oleksandr Maksymets, and 1 others. 2021. Habitat 2.0: Training home assistants to rearrange their habitat. Advances in neural information processing systems, 34:251–266.
- Yuchuang Tong, Haotian Liu, and Zhengtao Zhang. 2024. Advancements in humanoid robots: A comprehensive review and future prospects. IEEE/CAA Journal of Automatica Sinica, 11(2):301–328. [CrossRef]
- Hanqing Wang, Shaoyang Wang, Yiming Zhong, Zemin Yang, Jiamin Wang, Zhiqing Cui, Jiahao Yuan, Yifan Han, Mingyu Liu, and Yuexin Ma. 2026a. Affordance-r1: Reinforcement learning for generalizable affordance reasoning in multimodal large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 9738–9746. [CrossRef]
- Le Wang, Zonghao Ying, Xiao Yang, Quanchen Zou, Zhenfei Yin, Tianlin Li, Jian Yang, Yaodong Yang, Aishan Liu, and Xianglong Liu. 2025a. Robosafe: Safeguarding embodied agents via executable safety logic. arXiv preprint arXiv:2512.21220. [CrossRef]
- Xianhao Wang, Xiaojian Ma, Haozhe Hu, Rongpeng Su, Yutian Cheng, Zhou Ziheng, Hangxin Liu, Lei Liu, Bin Li, and Qing Li. 2026b. Chain of interaction benchmark (coin): When reasoning meets embodied interaction. arXiv preprint arXiv:2604.16886. [CrossRef]
- Xin Wang, Jie Li, Zejia Weng, Yixu Wang, Yifeng Gao, Tianyu Pang, Chao Du, Yan Teng, Yingchun Wang, Zuxuan Wu, and 1 others. 2025b. Freezevla: Action-freezing attacks against vision-language-action models. arXiv preprint arXiv:2509.19870. [CrossRef]
- Yu Wang, Xiaogeng Liu, Yu Li, Muhao Chen, and Chaowei Xiao. 2024. Adashield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting. In European Conference on Computer Vision, pages 77–94. Springer. [CrossRef]
- Yu Wang, Xiaogeng Liu, Yu Li, Muhao Chen, and Chaowei Xiao. 2025c. Adashield : Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting. In Computer Vision – ECCV 2024, pages 77–94, Cham. Springer Nature Switzerland. [CrossRef]
- Jimmy Wu, Rika Antonova, Adam Kan, Marion Lepert, Andy Zeng, Shuran Song, Jeannette Bohg, Szymon Rusinkiewicz, and Thomas Funkhouser. 2023. Tidybot: Personalized robot assistance with large language models. Autonomous Robots, 47(8):1087–1102. [CrossRef]
- Tao Wu, Chuhao Zhou, Guangyu Zhao, Haozhi Cao, Yewen Pu, and Jianfei Yang. 2025. When robots should say" i don’t know": Benchmarking abstention in embodied question answering. arXiv preprint arXiv:2512.04597. [CrossRef]
- Fanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia, Hao Zhu, Fangchen Liu, Minghua Liu, Hanxiao Jiang, Yifu Yuan, He Wang, and 1 others. 2020. Sapien: A simulated part-based interactive environment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11097–11107.
- Jiannan Xiang, Tianhua Tao, Yi Gu, Tianmin Shu, Zirui Wang, Zichao Yang, and Zhiting Hu. 2024. Language models meet world models: Embodied experiences enhance language models. Advances in neural information processing systems, 36.
- Haiweng Xu, Sipeng Zheng, Hao Luo, Wanpeng Zhang, Ziheng Xi, and Zongqing Lu. 2026. Unmasking the illusion of embodied reasoning in vision-language-action models. arXiv preprint arXiv:2604.18000. [CrossRef]
- Yuping Yan, Yuhan Xie, Yixin Zhang, Lingjuan Lyu, Handing Wang, and Yaochu Jin. 2025. When alignment fails: Multimodal adversarial attacks on vision-language-action models. arXiv preprint arXiv:2511.16203. [CrossRef]
- Jirui Yang, Zheyu Lin, Shuhan Yang, Zhihui Lu, and Xin Du. 2025a. Concept enhancement engineering: A lightweight and efficient robust defense against jailbreak attacks in embodied ai. arXiv e-prints, pages arXiv–2504. [CrossRef]
- Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, Marziyeh Movahedi, Manling Li, and 1 others. 2025b. Embodiedbench: Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents. arXiv preprint arXiv:2502.09560. [CrossRef]
- Sheng Yin, Xianghe Pang, Yuanzhuo Ding, Menglan Chen, Yutong Bi, Yichen Xiong, Wenhao Huang, Zhen Xiang, Jing Shao, and Siheng Chen. 2024. Safeagentbench: A benchmark for safe task planning of embodied llm agents. arXiv preprint arXiv:2412.13178. [CrossRef]
- Chenduo Ying, Linkang Du, Peng Cheng, and Yuanchao Shu. 2025a. Roboinspector: Unveiling the unreliability of policy code for llm-enabled robotic manipulation. arXiv preprint arXiv:2508.21378. [CrossRef]
- Zonghao Ying, Le Wang, Yisong Xiao, Jiakai Wang, Yuqing Ma, Jinyang Guo, Zhenfei Yin, Mingchuan Zhang, Aishan Liu, and Xianglong Liu. 2025b. Agentsafe: Benchmarking the safety of embodied agents on hazardous instructions. arXiv preprint arXiv:2506.14697. [CrossRef]
- Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, and 1 others. 2024a. R-judge: Benchmarking safety risk awareness for llm agents. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 1467–1490. [CrossRef]
- Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2024b. GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher. In The Twelfth International Conference on Learning Representations.
- Xiaoxue Zang, Ashwini Pokle, Marynel Vázquez, Kevin Chen, Juan Carlos Niebles, Alvaro Soto, and Silvio Savarese. 2018. Translating navigation instructions in natural language to a high-level plan for behavioral robot navigation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2657–2666, Brussels, Belgium. Association for Computational Linguistics. [CrossRef]
- Xiyin Zeng, Yuyu Sun, Haoyang Li, Shouqiang Liu, and Hao Wang. 2026. Recapa: Hierarchical predictive correction to mitigate cascading failures. arXiv preprint arXiv:2604.21232. [CrossRef]
- Simon Sinong Zhan, Yao Liu, Philip Wang, Zinan Wang, Qineng Wang, Zhian Ruan, Xiangyu Shi, Xinyu Cao, Frank Yang, Kangrui Wang, and 1 others. 2025. Sentinel: A multi-level formal framework for safety evaluation of llm-based embodied agents. arXiv preprint arXiv:2510.12985. [CrossRef]
- Borong Zhang, Jiahao Li, Jiachen Shen, Yishuai Cai, Yuhao Zhang, Yuanpei Chen, Juntao Dai, Jiaming Ji, and Yaodong Yang. 2025a. Vla-arena: An open-source framework for benchmarking vision-language-action models. arXiv preprint arXiv:2512.22539. [CrossRef]
- Hangtao Zhang, Chenyu Zhu, Xianlong Wang, Ziqi Zhou, Changgan Yin, Minghui Li, Lulu Xue, Yichen Wang, Shengshan Hu, Aishan Liu, and 1 others. 2025b. Badrobot: Jailbreaking embodied llm agents in the physical world. In The Thirteenth International Conference on Learning Representations.
- Tao Zhang, Kaixian Qu, Zhibin Li, Jiajun Wu, Marco Hutter, Manling Li, and Fan Shi. 2026a. Using large language models for embodied planning introduces systematic safety risks. arXiv preprint arXiv:2604.18463. [CrossRef]
- Weichen Zhang, Yiyou Sun, Pohao Huang, Jiayue Pu, Heyue Lin, and Dawn Song. 2025c. Mirage-bench: Llm agent is hallucinating and where to find them. arXiv preprint arXiv:2507.21017. [CrossRef]
- Zhexin Zhang, Junxiao Yang, Pei Ke, and Minlie Huang. 2023. Defending large language models against jailbreaking attacks through goal prioritization. Preprint, arXiv:2311.09096. [CrossRef]
- Zhilong Zhang, Wenyu Luo, Haonan Wang, Yifei Sheng, Yidi Wang, Hanyuan Guo, Haoxiang Ren, Xinghao Du, Yuhan Che, Tongtong Cao, and 1 others. 2026b. Anticipation-vla: Solving long-horizon embodied tasks via anticipation-based subgoal generation. arXiv preprint arXiv:2605.01772. [CrossRef]
- Qinhong Zhou, Sunli Chen, Yisong Wang, Haozhe Xu, Weihua Du, Hongxin Zhang, Yilun Du, Joshua B Tenenbaum, and Chuang Gan. 2024. Hazard challenge: Embodied decision making in dynamically changing environments. arXiv preprint arXiv:2401.12975. [CrossRef]
- Edoardo Zorzi, Francesco Taioli, Yiming Wang, Marco Cristani, Alessandro Farinelli, Alberto Castellini, and Loris Bazzani. 2026. Benchmarking interaction, beyond policy: a reproducible benchmark for collaborative instance object navigation. arXiv preprint arXiv:2604.00265. [CrossRef]
Table 1.
Taxonomy of epistemic and physical alignment diagnostics across embodied benchmarks. Env. denotes the evaluation manifold: S (Simulation), R (Real-world), or S/R (Both). Adv. indicates explicit adversarial red-teaming to expose behavioral inertia. ✓ denotes targeted metric isolation for structural vulnerabilities; × signifies a diagnostic blind spot.
Table 1.
Taxonomy of epistemic and physical alignment diagnostics across embodied benchmarks. Env. denotes the evaluation manifold: S (Simulation), R (Real-world), or S/R (Both). Adv. indicates explicit adversarial red-teaming to expose behavioral inertia. ✓ denotes targeted metric isolation for structural vulnerabilities; × signifies a diagnostic blind spot.
| Benchmark | Env. | Adv. | Instr. Comp. | Percep. Reas. | Cog. Consis. | Halluc. Detect. | Abstention | Physical Feas. | Sem.-Phys. Sync |
|---|---|---|---|---|---|---|---|---|---|
| Part I: Epistemic Foundations & Semantic Reasoning (System 2 “Brain”) | |||||||||
| REI-Bench Jiang et al. (2025) | S | × | ✓ | × | × | × | × | × | × |
| RoboBench Luo et al. (2025) | S/R | × | ✓ | ✓ | ✓ | ✓ | × | × | ✓ |
| ECBench Dang et al. (2025) | R | × | × | ✓ | × | ✓ | × | × | × |
| EmbodiedBench Yang et al. (2025b) | S | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| ALFRED Shridhar et al. (2020) | S | × | ✓ | ✓ | ✓ | × | × | × | × |
| TEACh Padmakumar et al. (2022) | S | × | ✓ | ✓ | ✓ | × | × | × | × |
| AgentBench Liu et al. (2024) | S | × | ✓ | × | ✓ | × | × | × | × |
| FoMER Dissanayake et al. (2025) | S | × | × | × | ✓ | ✓ | ✓ | ✓ | ✓ |
| RoboCerebra Han et al. (2026) | S | × | × | × | ✓ | × | × | ✓ | × |
| CALVIN Mees et al. (2022) | S | × | ✓ | × | ✓ | × | × | ✓ | × |
| VIMA-Bench Jiang et al. (2023) | S | × | ✓ | ✓ | × | × | × | ✓ | × |
| Ego4D Grauman et al. (2022) | R | × | × | ✓ | ✓ | × | × | × | × |
| Part II: Metacognitive Boundaries & Kinetic Guardrails (The “Immune System”) | |||||||||
| MIRAGE Zhang et al. (2025c) | S | × | ✓ | ✓ | ✓ | ✓ | × | × | × |
| SAFE Gu et al. (2026) | S | × | × | × | × | ✓ | × | × | × |
| POPE Li et al. (2023b) | R | ✓ | × | ✓ | × | ✓ | × | × | × |
| MMHal-Bench Sun et al. (2024) | R | ✓ | × | ✓ | × | ✓ | × | × | × |
| AbstainEQA Wu et al. (2025) | S | ✓ | ✓ | ✓ | × | × | ✓ | × | × |
| AgentSafe Ying et al. (2025b) | S | × | ✓ | ✓ | × | × | ✓ | × | × |
| SAFEL Son et al. (2025) | S | × | ✓ | × | ✓ | × | ✓ | × | × |
| SafeAgentBench Yin et al. (2024) | S | × | ✓ | × | ✓ | × | ✓ | × | × |
| HAZARD Zhou et al. (2024) | S | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ |
| R-Judge Yuan et al. (2024a) | S | × | ✓ | × | ✓ | ✓ | ✓ | × | × |
| Safety-Gym Ray et al. (2019) | S | × | × | × | × | × | ✓ | ✓ | × |
| Part III: Kinetic Actuation & Sensorimotor Sync (The “Somatic System”) | |||||||||
| SENTINEL Zhan et al. (2025) | S | × | ✓ | × | × | × | ✓ | ✓ | × |
| ADAPT Chen et al. (2026a) | S | × | ✓ | × | × | × | × | ✓ | ✓ |
| LIBERO Liu et al. (2023a) | S | × | ✓ | ✓ | × | × | × | ✓ | × |
| RLBench James et al. (2020) | S | × | ✓ | ✓ | × | × | × | ✓ | × |
| LIBERO Liu et al. (2023a) | S | × | ✓ | ✓ | × | × | × | ✓ | × |
| ManiSkill Gu et al. (2023) | S | × | × | ✓ | × | × | × | ✓ | × |
| RoboTHOR Deitke et al. (2020) | S/R | × | × | ✓ | × | × | × | ✓ | ✓ |
| VLA-Arena Zhang et al. (2025a) | S | × | ✓ | ✓ | × | × | × | ✓ | × |
| SafeLIBERO Hu et al. (2025) | S | × | ✓ | ✓ | × | × | ✓ | ✓ | × |
| HazardArena Chen et al. (2026b) | S | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ |
| AI2-THOR Kolve et al. (2017) | S | × | × | ✓ | × | × | × | ✓ | ✓ |
| Habitat Szot et al. (2021) | S | × | × | ✓ | × | × | × | ✓ | ✓ |
| BEHAVIOR Srivastava et al. (2022) | S | × | ✓ | ✓ | ✓ | × | × | ✓ | ✓ |
| SAPIEN Xiang et al. (2020) | S | × | × | × | × | × | × | ✓ | ✓ |
| SayCan Ahn et al. (2022) | R | × | ✓ | ✓ | × | × | × | ✓ | ✓ |
| Mobile ALOHA Fu et al. (2024) | R | × | × | × | × | × | × | ✓ | ✓ |
| Open X-Embodiment O’Neill et al. (2024) | S/R | × | ✓ | ✓ | × | × | × | ✓ | × |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.