Submitted:
04 October 2026
Posted:
07 October 2026
You are already at the latest version
Abstract
Embodied AI aims to develop agents that perceive, reason, and act in physical environments. Beyond recognizing objects and scenes, agents must identify possible interactions, determine their relevance to the current task and embodiment, and translate them into executable behavior. Affordance learning supports this process by grounding perception in action possibilities and connecting it with task reasoning and control. Existing affordance surveys, however, provide only partial coverage of the field, often focusing on isolated aspects of perception or robotic interaction and lacking a comprehensive view of the full perception–reasoning–action pipeline in embodied intelligence. To bridge this gap, this survey presents a unified framework organized around three complementary roles: affordance perception identifies possible interactions and their locations; affordance reasoning selects opportu-nities relevant to the task, context, and agent state; and affordance-guided action uses affordance representations for planning and policy learning. Within this framework, we develop fine-grained taxonomies, compare representative methods, and trace the field’s progression from passive visual recognition toward action-aware embodied learning. We further organize representative datasets and evaluation protocols and identify open challenges and future directions. Overall, this survey clarifies how affordance learning connects perception, reasoning, and action to support generalizable and physically grounded embodied intelligence.
Keywords:
affordance learning
; embodied AI
; computer vision
; robot learning
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.