Computer Science and Mathematics

Sort by

Article
Computer Science and Mathematics
Robotics

Hong Su

Abstract: Autonomous robots require not only reliable task execution but also the ability to identify what new local knowledge should be learned, when explicit thinking is necessary, and when learned knowledge can be reused automatically. However, existing methods often bind learned knowledge to predefined tasks or fixed reasoning procedures, making it difficult to autonomously discover reusable local concerns and reduce repeated reasoning after adaptation. This paper proposes a thinking-triggered aspect learning framework for autonomous knowledge acquisition and transfer. Routine behavior is handled by automatic models, while novelty, uncertainty, or periodic heartbeat inspection triggers thinking to localize unresolved concerns. The robot then discovers task-independent aspects containing relevant factors, applicability conditions, handling principles, and local automatic models. Validated aspects are reused across tasks and domains, while aspect-specific calibration and checkpoint decay progressively convert explicit thinking into automatic processing; environmental changes can reactivate thinking and initiate new aspect acquisition. Experiments show aspect and factor F1 scores of 0.975 and 0.965, perfect zero-shot cross-task and cross-domain transfer in all successfully validated source runs, and a reduction in thinking rate from 0.508 to 0.100 after familiarization; the proposed method also reduces LLM calls by 76.4% compared with Always Think while achieving 0.976 post-validation detection and handling success.

Article
Computer Science and Mathematics
Robotics

Hong Su

Abstract: Autonomous robots should be able to initiate meaningful activities even when no explicit task is currently assigned. However, most existing systems remain task-centered or rely on predefined internal mechanisms, limiting their ability to preserve and acquire cognitive influences that remain relevant after current reasoning has ended. This paper proposes a self-initiated activity-triggering framework in which robot activities can arise from external perception, autonomous internal processes, or prolonged inactivity. Internal processes maintain independent states and provide non-commanding messages whose effects are interpreted contextually by the main thinking process. An independent process manager further learns acquired internal processes from previous thinking experience. Process formation is gated by future influence rather than recurrence alone and supports both recurrent delayed needs and rare but severe future vigilance needs. Learned activation signatures determine what future evidence should reactivate acquired processes. Experiments show that the proposed external triggering mechanism achieves 96.00% strict accuracy while reducing LLM tokens by 76.1% relative to centralized reasoning. The learned process manager achieves 98.00% acquisition F1, while activation signatures improve strict balanced future correctness from 62.67% to 82.33%. Compared with repeated history-based LLM reasoning, process internalization reduces LLM calls by 91.80% and token consumption by 94.07%.

Article
Computer Science and Mathematics
Robotics

Giulio Leone

,

Daniela D’Auria

Abstract: Diagnostic uncertainty in neurological rehabilitation motivates robotic systems that can select informative sensing actions adaptively rather than rely on fixed assessment protocols or static clinical records. This work introduces the Embodied Evidence Acquisition and Reasoning Loop (EARL), a typed architecture in which five role-specialized critics propose and assess sequential sensing actions, a probabilistic simulation sandbox estimates their expected information value, and a deterministic safety governor retains exclusive execution authority. Observed, derived, and simulated evidence remain provenance-distinct in a replayable, hash-linked ledger. EARL was evaluated retrospectively on 64 subjects from the PhysioNet Gait in Neurodegenerative Disease Database using repeated subject-level cross-validation, 11 predeclared conditions, and 3520 replay-verified runs. The primary endpoint was area under cumulative posterior-entropy reduction. EARL achieved 7.689 (95% CI 7.535–7.840), exceeding fixed-order and random selection by 0.459 and 0.763, respectively; the EARL-minus-EIG difference was -0.321. EARL nevertheless achieved higher macro accuracy (55.8% versus 49.5%) and a lower Brier score (0.781 versus 0.813) than pure expected-information-gain selection, demonstrating a trade-off among uncertainty reduction, discrimination, calibration, and safety-constrained decision making. All 10 deterministic safety-conformance scenarios produced their expected outcomes. These results establish reproducible software behavior for bounded sequential retrospective sensing; they do not establish clinical diagnostic performance, treatment benefit, physical-robot safety, or patient efficacy.

Article
Computer Science and Mathematics
Robotics

Yanghua Li

,

Bin Liu

,

Cheng Zhu

Abstract: Multi-rotor unmanned aerial vehicles (UAVs) used in post-disaster reconnaissance and target search face two challenges: limited prior information and the disruption of task chains when targets are discovered dynamically. Conventional coverage path planning (CPP) algorithms rely on fixed sweeping directions and static task allocation, which leads to high turning energy consumption, imbalanced workloads among vehicles, and an inability to respond to unexpected target discoveries. To address these issues, this paper proposes the Dynamic Adaptive Reallocation for Terrain Sweeping (DARTS) framework.First, a direction-adaptive sweeping mechanism is designed that generates boustrophedon (back-and-forth) sweep paths along horizontal, vertical, diagonal, and PCA-derived principal directions, and automatically selects the fewest-turn coverage pattern for each sub-region. Second, a two-level load-balanced allocation model is constructed: in the offline stage, a constrained 0-1 integer-programming model solved by a branch-and-bound constraint-programming search is employed to simultaneously optimize the load variance and regional clustering under hard battery-capacity constraints; in the online stage, when a UAV discovers a target it switches from a covering to a tracking role, and the released remaining sub-regions of the affected UAV are treated as regional task blocks that are migrated by a greedy heuristic with minimal overhead on the basis of a composite score combining the real-time load and spatial distance.Simulation results show that, compared with the conventional Divide Areas Algorithm for Optimal Multi-Robot Coverage Path Planning (DARP) with greedy allocation, DARTS reduces the coefficient of variation (CV) of multi-UAV path lengths from 0.111 to 0.082 (a relative reduction of 26.5%) and, at the per-region scan-path level, the average number of turns by up to 54.5% relative to the worst fixed sweeping pattern. Moreover, when a target-discovering UAV switches from covering to tracking the target, the event-triggered reallocation raises the target discovery rate from 68.3% to 100% and the coverage completion rate from 61.8% to 83.1%, at a transfer overhead of only 13.1% of the released path length, while the initial CP allocation solves in under 0.1 s at the scale of the main experiments. These results confirm the robustness and computational efficiency of the proposed algorithm in dynamic scenarios.

Article
Computer Science and Mathematics
Robotics

Hong Su

Abstract: Learning from prior experience is essential for autonomous robots, but directly reusing historical actions is often insufficient when the environment changes and previously successful behaviors are no longer applicable. This paper proposes an experience-to-thought learning framework that learns the underlying thinking activities behind historical robot materials rather than only their surface behaviors. An LLM analyzes accumulated experiences to extract diverse thoughts, including prediction, calculation, comparison, risk evaluation, causal analysis, reflection, planning, and verification. A model-agnostic temporal thought-learning mechanism then associates these thoughts with the evolving states and contextual conditions under which they are useful, allowing the same thought to guide different coordinated action sequences in different situations. The framework further supports open-ended growth of the thinking repertoire by detecting when existing thinking is insufficient, discovering candidate thoughts, and consolidating them into reusable new thinking activities. Experiments show that temporal thought learning substantially improves thought selection under temporally ambiguous situations and transfers more effectively than direct action-plan learning when the required behavior changes. The framework also successfully discovers, consolidates, and reuses a new thinking activity in previously unseen situations.

Review
Computer Science and Mathematics
Robotics

Yuqing Zhu

,

Alexander Alexandrovich Boryaev

Abstract: Decentralized swarms of unmanned aerial vehicles represent a revolutionary technology with various applications, but so far there are no systemic analytical approaches that combine technical, tactical, and ethical aspects. This paper presents a systematic review of 47 publications from IEEE Xplore, Scopus, and Web of Science (2018–2026), developing original analytical models for swarm deployment. A taxonomy of three architectural types (centralized, hierarchical, decentralized) with quantitative resilience assessment shows decentralized swarms maintain effectiveness with up to 30% agent loss, versus 5% for centralized. An analytical J/S model evaluates communication channel vulnerability to electronic warfare, while a grid weight model wi(t+1)=wi(t)+αdi(t)−βλwi(t) with threshold wth=10 enables decentralized targeting without a single point of failure. Comparative analysis of MANET and DTN protocols reveals latency, reliability, and power trade offs, justifying a hybrid scheme for combat conditions. Economic analysis demonstrates cost exchange ratios reaching 1:1000, and computing infrastructure assessment shows a 100 UAV swarm requires ≈2 PetaFLOPS with 70% ground and 30% onboard distribution using NPUs. Findings provide practical recommendations for swarm architecture selection, hybrid communication protocols, specialized autopilot modes, automated ground infrastructure, and mandatory human in the loop confirmation for critical targeting decisions. Future research priorities include experimental validation, integration with large language models, and cyber security assurance.

Article
Computer Science and Mathematics
Robotics

Hong Su

Abstract: Autonomous robots must continually improve their behavior policies to operate reliably in complex, partially observable, and changing environments. However, most robot-learning methods assume that suitable training data are externally provided or passively accumulated, without enabling robots to reason about missing materials or how they should be constructed. This paper proposes TG-MAC, a thought-guided material acquisition and construction framework for autonomous robot learning. TG-MAC identifies environment-dependent training-material gaps and addresses them through delayed-feedback retrospective association, change-aware multimodal spatiotemporal evidence expansion, goal-relevant factor selection, and thought-guided active acquisition and supplementation. These mechanisms can be dynamically selected and combined according to material gain, reliability, cost, and risk, while supplemented materials are validated before policy updating. Controlled simulations over 300 scenarios show that TG-MAC achieves 98.0% delayed-outcome association accuracy, compared with 87.0% without thought-guided triggering; obtains a factor-selection F1 of 71.2% and policy success of 84.7% in complex multimodal environments; and reaches 80.7% policy success using only 2.12 real interactions per round, compared with 73.3% and 4.34 interactions for active acquisition alone.

Article
Computer Science and Mathematics
Robotics

Hong Su

Abstract: Long-horizon autonomous robots must retain and reuse important information from sensing, learning, and reasoning so that past events can continue to influence future decisions and actions. However, conventional memory retrieval and fixed-interval reminders either fail to activate information when it is indirectly relevant or repeatedly process it after its importance has decreased. This paper proposes a persistent influence model for autonomous robots. Important sensor states, newly learned knowledge and skills, externally supplied information, and conclusions autonomously judged important by a thinking module are represented as persistent influence items. These items can be independently reconsidered at adaptive intervals or attached to the prompts of related future tasks. During each processing cycle, the thinking module may request additional sensing, perform analysis, apply learned knowledge, execute a skill, or take protective action. An independent regulation module gradually updates influence strength, processing interval, and state according to elapsed time, new evidence, and action outcomes, preventing high-consequence information from being removed prematurely. Experiments show that the proposed method reduced cumulative temperature excess by 72.6% in a long-horizon latent-risk scenario, achieved an 83.33% task-success rate, and increased hazard prevention for thinking-identified issues to 41.18%, while reducing unnecessary inspections by 79.7% compared with fixed-interval persistence.

Article
Computer Science and Mathematics
Robotics

Malte Herrmann

,

Dominykas Strazdas

,

Ayoub Al-Hamadi

Abstract: With Industry 5.0, human-robot collaboration has moved to the center of attention. This introduces new challenges, where workspaces become highly dynamic leading to safety concerns for robots and especially for humans. This makes it important to have a realistic and accurate digital representation of a production cell and its components in the environments for movement planning and remote supervision. This paper presents a digital twin of an Industry 5.0 smart production cell, synchronizing objects bidirectional between a real and virtual workspace with minimal effort using only a single RGB-D camera. A fine-tuned YOLO-based detector identifies tools and items in the scene, estimates their spatial position, and spawns them in Unity relative to the robot via coordinate transformation. Experiments with different scanning velocities demonstrate a mean planar spawn deviation of 5.82mm (standard deviation 2.24mm) at 0.1m/s at 0.7m height, while maintaining a constant depth bias of −0.54mm. Once spawned, objects can be manipulated freely via drag-and-drop within the simulation. Upon confirmation, a motion-planning module calculates trajectories to execute these changes physically. Across 95 trials and 1805 object placements, the system achieves 100% success within working bounds, successfully executing complex tasks such as repositioning objects and stacking them into pyramid structures. The presented system provides a framework to see, spawn, and synchronize industrial workspaces, enabling rapid setup and safe remote supervision of smart production cells in high dynamic Industry 5.0 environments.

Article
Computer Science and Mathematics
Robotics

Emin Bayramov

,

Zoltán Istenes

Abstract: Predicting where a traffic agent will go and how it will get there are fundamentally different questions, yet most learning-based forecasters conflate them. When a model entangles agent intent (e.g., turning left) with driving style (e.g., cautious versus aggressive execution), spurious correlations arise that degrade predictions under distribution shift. This paper proposes SC-Mamba, a forecasting architecture built on selective state-space modeling and causal representation learning that explicitly disentangles these two factors. The model factorizes its latent space into independent intent and style variables, enforcing their separation through counterfactual do-interventions applied during training. A differentiable Signal Temporal Logic (STL) safety layer additionally constrains predicted trajectories to satisfy formal driving rules. Evaluated on the Argoverse 2 Motion Forecasting Dataset, SC-Mamba achieves 0.73 m minADE6, 1.22 m minFDE6, and 14.7% Miss Rate using only 1.35M parameters at 6.16 ms per scene—competitive with Transformer baselines at a fraction of computational cost. Critically, it is demonstrated that the learned intent latents correspond to identifiable maneuver categories (left turn, right turn, straight), style latents correlate with measurable execution dynamics (speed, curvature), and counterfactual manipulation of either factor independently controls trajectory shape without retraining. These results establish that linear-complexity SSM backbones can support structured causal reasoning for motion prediction while maintaining real-time performance.

Article
Computer Science and Mathematics
Robotics

Alireza Rezaee

Abstract: During indoor navigation, a mobile robot often needs to cross doors while relying on multiple sensors that may be faulty or temporarily unreliable. This paper presents a fault-tolerant framework for mobile-robot door-crossing behaviour based on a dynamic Bayesian network (DBN). The framework combines behaviour learning, sensor-fault detection, fault isolation, structural adaptation of the Bayesian network, and corrective control. Faults are detected by evaluating sensor probabilities and modifying the DBN structure when abnormal sensor readings are identified. Five fault types are considered: constant, drift, shock, crosstalk, and spike faults. Experiments in static and dynamic environments show that the proposed DBN-based structures substantially reduce behavioural failure compared with a conventional Bayesian-network controller. The best configuration, combining fault detection with the third manoeuvre strategy, achieves a failure rate of 7.2% in the tested dynamic environment.

Article
Computer Science and Mathematics
Robotics

Molly Watson

,

Zach Carter

,

Yeganeh Madadi

Abstract: Simultaneous localization and mapping (SLAM) is a foundational capability for autonomous navigation in unknown environments. Its performance is strongly coupled to the type, quality, and reliability of available sensor data, limiting the portability of navigation systems across heterogeneous mobile robot platforms. This paper presents a cross-platform adaptive navigation framework that decouples localization providers from platform-specific sensing configurations. A sensor abstraction layer normalizes heterogeneous and low-fidelity sensor inputs into a unified representation, enabling structured operational modes constructed according to available sensing modalities, computational constraints, and environmental characteristics. A learning-based performance prediction module is further designed to estimate impending SLAM degradation and support proactive mode switching. Due to middleware constraints within the Pepper NAOqi stack, this predictive component was not deployed during experimental evaluation and remains part of the proposed architecture for future validation. Experimental results on real indoor navigation tasks demonstrate improved robustness and portability compared to fixed SLAM configurations without manual retuning.

Article
Computer Science and Mathematics
Robotics

Alireza Shojaei

Abstract: A widely held assumption in cross-embodiment robot learning is that morphologically similar robots transfer behavior more easily, so a similarity measure, a learned transferability predictor, or a sufficiently diverse pretraining set should predict or improve transfer. We test this assumption under an enforced acceptance gate, requiring every claim to beat its strongest trivial baseline by a margin whose bootstrap 95% confidence interval excludes that baseline on independent units, and we refute it four ways. A morphology-distance predictor fails to beat a target-only prior, with Spearman rho = 0.283 versus 0.579 over 42 independent pairs. A transferability oracle trained on full morphology features, with rho = 0.762 over 29 robots, fails to beat a one-bit arm/not-arm indicator, which reaches 0.834; it predicts robot class, not morphology. On the 812-pair suite there is no transfer law; the mean gain is +0.8 percentage points, within-target variation across sources is under-dispersed relative to evaluation noise, with variance ratio 0.53, and every candidate pairwise predictor has a confidence interval spanning zero. The apparent benefit of pretraining diversity vanishes once total data volume is held fixed, with Delta = -0.026 and CI [-0.092, +0.034]; breadth never beats depth at any tested budget. A same-body control shows the assay detects transfer when present, with Delta = +0.067 and p = 0.0006, and a morphology-equivariant graph policy attains zero-shot transfer that distance still fails to grade. The outcome level is set by the target's own trainability and raw data volume; short of exact body identity, no measured relation between bodies moves it.

Article
Computer Science and Mathematics
Robotics

Alireza Shojaei

Abstract: Every system that reached zero-shot cross-embodiment manipulation in the first half of 2026 made the same move, deleting body information from the interface between task reasoning and motor control, whether through body-agnostic handheld data, masked end-effectors, language-coded actions, or contact-intent latents. None of these systems tests that the deletion is what causes transfer, characterizes what the interface still retains, or asks whether the interface must be symbolic. This paper supplies all three on a scene-controlled manipulation substrate where appearance confounds cannot operate. A causal interface ladder over five source and five held-out arms shows that a body-blind end-effector interface transfers zero-shot while leaking body channels back into it collapses transfer once the leak passes a threshold, a gap of $0.157$ that every held-out arm reproduces, and that injecting body identity is actively harmful. At matched body-blindness and identical upstream information, a structured symbolic coding of the interface beats a language-token coding by $0.109$ with the margin compounding over task depth, while a low-capacity continuous latent falls below the task's precision floor. On a released vision-language-action model with scene controlled by robot-swap rendering, most apparent body recoverability is scene appearance, yet a modest scene-invariant residue exceeds a raw-pixel control in all three folds, and an in-model test finds the decoded action body-light. Recoverability is not reliance, at the interface and inside the released model alike, which is the mechanism the zero-shot wave depends on and the boundary it must respect.

Article
Computer Science and Mathematics
Robotics

Alireza Shojaei

Abstract: Cross-embodiment transfer, the reuse of learned behavior across robots of different morphologies, is a central goal of generalist robot learning, and the quantitative claims made about it fail in a small set of specific, recurring ways. A transfer claim can ride a robot-class prior that a single bit reproduces, tighten a confidence interval by treating seeds as independent observations, credit pretraining diversity for what is data volume, or certify an invariant representation with a linear probe that a nonlinear probe falsifies. This paper defines an enforced measurement standard of eight checks, each backed by a runnable tool, that catches these failure modes before a claim ships. The centerpiece is an acceptance gate that recomputes a claim's metric, runs its strongest trivial baseline, bootstraps the difference over independent units, and emits pass or fail. The standard is demonstrated through eleven documented failure-mode case studies drawn from a real cross-embodiment research program, each stated as the tempting claim, the diagnostic that exposes it, the corrected analysis, and the check that catches it, and each traced to a released artifact. The eleven cases span every check, from a correlation of 0.98 that proves geometrically trivial to an identity probe on a released vision-language-action model that a raw eight-by-eight-pixel control exposes as appearance-confounded. The work is positioned within the 2025-2026 movement toward statistical rigor in robot-policy evaluation, to which it adds the confound set specific to cross-embodiment transfer, an enforced gate rather than a checklist, and a worked record of the checks correcting real claims.

Article
Computer Science and Mathematics
Robotics

Alireza Shojaei

Abstract: A central ambition of cross-embodiment robot learning is a single representation of the task that is invariant to the body, a shared task-state that means the same thing on a four-legged robot as on an eight-legged one, or on a Panda arm as on a UR10e. The dominant approach learns such a representation as a continuous latent, trained for control-sufficiency and scrubbed of body identity by an adversary. We prove this is impossible exactly when behavior is body-coupled. The central object is a sufficiency-invariance bound, which states that any continuous task-state \(z\) that is \(\varepsilon\)-sufficient to predict a body's realized outcome \(y\) leaks the body identity \(m\) at a floor set by how body-coupled the behavior is, \(I(z;m) \ge I(y;m) - \kappa(\varepsilon)\). The lower bound is constructive, obtained by composing the sufficiency decoder with a classifier of \(m\) from \(y\) to build a probe of \(m\) from \(z\), and its assumption-free content is the measured accuracy of that probe. Lifting it to the closed form requires a margin condition, \(\kappa(\varepsilon) \le \inf_t [\delta(t) + \varepsilon/t^2]\). A rate-distortion complement shows when a coarse task-state escapes the floor, namely when bodies reach the same coarse symbol through different fine realizations. We validate the bound on two non-commensurable substrates. On locomotion the floor is $0.90$ under a strong probe against a linear reading of $0.37$, every control-sufficient continuous latent leaks topology above $0.98$, and a coarse state reaches near chance while retaining most task signal. On manipulation the floor is $0.86$ against a linear $0.34$, the leak exceeds $0.80$, and the coarse state again sheds the body. The morphology-invariant interface between deliberation and control must therefore be coarse.

Article
Computer Science and Mathematics
Robotics

Alireza Shojaei

Abstract: Companion work shows that neither morphological similarity nor data diversity governs policy transfer across robot bodies, and that any continuous task representation sufficient for control re-encodes the body it came from. This paper supplies the constructive counterpart, a task factorization whose cross-body interface is a coarse symbolic progress state. On a depth K sequential-reach task over six simulated arms, we compare three behavior-cloned agents that differ only in how the observation is factored. A reactive policy that must infer the active subgoal from perception solves single reaches, with success 0.628, but collapses at depths 2 to 4, with success 0.069, 0.000, and 0.003. A deliberative agent whose plan supplies the active subgoal holds between 0.631 and 0.558 and matches an oracle policy hand-fed the progress state at every depth. Under an enforced acceptance gate requiring every claim to beat its strongest baseline with a bootstrap confidence interval excluding it on independent units, the cross-morphology contrast pooled over depths 2 to 4 is 0.590 versus 0.018, with gap CI [+0.34, +0.81] over six arms, and a pre-registered depth 1 control gap of -0.002. A recurrent policy tuned per morphology to convergence reaches depth 1 parity, with success 0.606, yet still collapses at deeper tasks, with success 0.006, 0.000, and 0.000, so memory does not substitute for the symbolic state under behavior cloning. The factorization also matches the oracle-fed monolith's best success with 2.5 to 3 times less demonstration data, beats a language-token coding of the same information, with gap +0.109 and exact p = 0.031, runs zero-shot on 13 unseen arms, and composes independently trained skills with zero composite demonstrations.

Article
Computer Science and Mathematics
Robotics

Emin Bayramov

,

Zoltán Istenes

Abstract: Motion forecasting models in the autonomous driving domain achieve high accuracy but cannot explain their predictions, creating a barrier to safety certification. This paper presents CogSig-Mamba, a model that produces causally validated temporal explanations alongside trajectory predictions. Inspired by hippocampal memory, the model follows a five-stage process: (1) synaptic tagging, where a top-k sparse gate selects which observation windows drove the prediction; (2) evidence encoding, which consolidates window content into memory representations; (3) reverse replay, which confirms causal faithfulness of tags through removal experiments; (4) spatial context, integrating road geometry for grounded predictions; and (5) constructive retrieval, which explains each predicted behavior via mode-specific attention to produce a complete Cognitive Signature. Evaluated on the Argoverse 2 dataset, CogSig-Mamba achieves minADE6 = 0.908 m and minFDE6 = 1.949 m with only 1.9M parameters. Removing tagged windows shifts predictions by 4.6 m on average, while removing untagged windows produces negligible impact (0.46 m), confirming causal faithfulness across all 24,988 validation scenarios. To the best of the authors’ knowledge, this is the first motion forecaster with verified temporal credit assignment, supporting the audit trails required by ISO 21448 for safety-critical deployment.

Article
Computer Science and Mathematics
Robotics

Kailin Lyu

,

Kangyi Wu

,

Pengna Li

,

Wenxuan Song

,

Di Wu

,

Jianwei He

,

Junting Chen

,

Ning Yang

,

Zebin Han

,

Kaiwen Luo

+47 authors

Abstract: Vision-and-Language Navigation (VLN) requires embodied agents to ground natural language instructions in visual perception and make navigation decisions in complex 3D environments, making it a central problem in embodied artificial intelligence. Since the introduction of the Room-to-Room (R2R) benchmark, VLN has made substantial progress. In recent years, as research settings have gradually expanded from closed and single indoor benchmark scenarios to open-world environments, the field has undergone a profound paradigm shift from passive instruction following on fixed benchmarks to autonomous cognitive navigation in open-world settings. However, existing surveys mainly organize prior work according to technical taxonomies, lacking a systematic characterization of this paradigm evolution. To address this gap, this survey proposes an evolution-centered unified analytical framework that reviews contemporary VLN research across four progressive layers: perception, cognition, learning, and generalization. It reveals the intrinsic connections and evolutionary logic among different technical lines, identifies key open challenges at each dimension, and outlines future research directions. This survey aims to provide VLN researchers with a clear panoramic view of capability evolution, while offering the broader embodied intelligence community a systematic roadmap from closed-benchmark evaluation toward trustworthy open-world deployment.

Article
Computer Science and Mathematics
Robotics

Abraham Goodman

,

Theodoros Theodoridis

,

Saleem Ameen

,

Guowu Wei

Abstract: The recognition of surgical instruments using Artificial Intelligence (AI) in Minimally Invasive Surgery (MIS) offers significant opportunities for data-driven improvements in surgical training and patient safety, with surgical instrument recognition being a critical component. MIS remains challenging due to complex intraoperative conditions that limit conventional real-time object detection AI algorithms. This paper optimises the state-of-the-art YOLOv8 object detection architecture for surgical instrument recognition and takes a novel approach to deal with imbalance of the dataset. A large-scale dataset of 25 surgical videos, consolidated from CholecTrack20, CholecT50, and Cholec80, underwent custom cleaning and strategic partitioning to address severe class imbalance. A systematic tournament identified YOLOv8l as the best-performing variant of the YOLOv8 versions, achieving Mean Average Precision (mAP)@0.5 of 64.4% on the test set. Despite hyperparameter tuning, the attempt led to overfitting, while the data-balancing strategy, despite a slight reduction in overall mAP@0.5 to 60.4%, the approach significantly improved per-class accuracy, notably doubling performance for the rarely used instruments such as Scissors and Clippers. This study establishes a new performance baseline for surgical instrument recognition using a carefully configured YOLOv8l model, underscoring that data imbalance, rather than architecture, is the primary limitation. Future progress for surgical instrument recognition will hinge on data-centric strategies for robust and clinically reliable models.

of 13