Submitted:
09 July 2026
Posted:
10 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- RQ1 – Methodological Landscape: Which Learning from Demonstration (LfD) architectures, trajectory learning algorithms, and mathematical representations are most commonly used in robotic deburring and polishing tasks?
- RQ2 – System and Sensory Configuration: How are robotic systems, control architectures, and sensory modalities (e.g., force sensing, vision, multimodal perception) designed and integrated to support LfD-based surface finishing operations?
- RQ3 – Evaluation and Performance Metrics: What performance evaluation criteria and validation metrics are used to assess the effectiveness of LfD-based robotic deburring and polishing methods?
- RQ4 – Industrial Deployment Challenges: What are the key technical challenges, limitations, and research gaps that hinder the transition of LfD-based robotic surface finishing methods from laboratory settings to industrial applications?
2. Methodology
2.1. Protocol and Registration
2.2. Eligibility Criteria
- Thematic Relevance: The study must explicitly investigate robotic surface finishing operations, specifically deburring or polishing.
- Methodological Approach: The proposed robotic control or programming framework must incorporate an LfD-based method (e.g., kinesthetic teaching, teleoperated demonstration, or haptic imitation learning) where motion or force policies are extracted from human demonstrations.
- Technical Depth: The study must provide experimental validation or detailed theoretical formulations regarding trajectory generation, force regulation, control architectures, or sensor configurations.
- Publication Type and Language: Only peer-reviewed academic journal articles or conference proceedings published in the English language were included.
- Out-of-Scope Operations: Studies addressing generic robotic manipulation tasks (e.g., pick-and-place, assembly, peg-in-hole insertion, or trajectory tracking in free space) without a specific surface finishing context.
- Non-Demonstration Control: Research utilizing traditional robotic programming (e.g., offline CAD/CAM programming, manual teach-pendant jogging) or model-based adaptive controllers that do not learn or adapt based on human demonstration data.
- Insufficient Quality or Length: Extended abstracts, editorial prefaces, technical reports, white papers, book reviews, or unpublished master’s/doctoral theses.
- Accessibility Barriers: Articles for which the full-text version was unavailable or lacked sufficient data to extract parameters necessary for answering the research questions.
2.3. Search Strategy and Information Sources
- Block A (Learning Paradigm): `"Learning from Demonstration"`, `"LfD"`, `"Imitation Learning"`, `"Programming by Demonstration"`, `"DMP"`, `"Dynamic Movement Primitives"`, `"ProMP"`, `"Probabilistic Movement Primitives"`, `"Skill learning"`, `"Task learning"`.
- Block B (Application Context): `"Deburring"`, `"Polishing"`, `"Surface finishing"`, `"Surface-finishing"`, `"Finishing"`, `"Robotic finish*"`.
2.4. Screening and Selection Process
2.5. Methodological Quality Assessment
- QA1 — Experimental Validation: Does the study validate the proposed LfD framework on a physical robotic manipulator performing a surface finishing task?
- QA2 — Reproducibility and Statistical Reporting: Are the experiments repeated over multiple trials, and are quantitative statistical measures (e.g., mean and standard deviation) reported?
- QA3 — Baseline Comparison: Is the proposed method compared against a baseline (e.g., standard model-based control, human expert performance, or alternative LfD algorithms)?
- QA4 — Generalizability: Is the learned skill tested on varying workpiece geometries, orientations, or materials?
- QA5 — Quantitative Performance Metrics: Does the study report quantitative evaluation metrics for both kinematic trajectory accuracy and dynamic force tracking?
- QA6 — Industrial Relevance: Is the task validated using an industrial workpiece, or are real-world industrial deployment constraints (e.g., tool wear, cycle time, worker safety) discussed?
- High Quality (H): Cumulative score of 9 to 12.
- Moderate Quality (M): Cumulative score of 5 to 8.
- Low Quality (L): Cumulative score of 0 to 4.
2.6. Data Extraction and Synthesis
2.7. Effect Measures and Synthesis Rationale
3. Results
3.1. Study Selection and PRISMA Flow Diagram
3.2. Quantitative Study Characteristics
3.3. Bibliometric Analysis
3.3.1. Temporal Distribution
3.3.2. Geographical Distribution
3.3.3. Quantitative Methodological and Sensory Distributions
3.4. Methodological Quality Assessment Results
| Source | QA1 | QA2 | QA3 | QA4 | QA5 | QA6 | Total | Rating |
| Min et al. [2] | 2 | 1 | 1 | 1 | 2 | 2 | 9 | High |
| Li et al. [3] | 2 | 2 | 2 | 1 | 2 | 1 | 10 | High |
| Si et al. [4] | 2 | 1 | 1 | 1 | 1 | 1 | 7 | Moderate |
| Wang et al. [5] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Wang et al. [6] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Acikgoz et al. [7] | 2 | 1 | 1 | 1 | 1 | 2 | 8 | Moderate |
| Zhai et al. [8] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Fischer et al. [9] | 2 | 2 | 2 | 2 | 1 | 1 | 10 | High |
| Kulak et al. [10] | 2 | 1 | 1 | 1 | 1 | 1 | 7 | Moderate |
| Zhang et al. [11] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Wu et al. [12] | 2 | 1 | 1 | 1 | 2 | 2 | 9 | High |
| Wu et al. [13] | 2 | 1 | 1 | 2 | 2 | 1 | 9 | High |
| Haninger et al. [14] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Möhl et al. [15] | 2 | 1 | 2 | 2 | 1 | 2 | 10 | High |
| Parvizi et al. [16] | 2 | 1 | 1 | 1 | 1 | 2 | 8 | Moderate |
| Wang et al. [17] | 2 | 1 | 2 | 2 | 2 | 1 | 10 | High |
| Xu et al. [18] | 2 | 2 | 2 | 1 | 2 | 1 | 10 | High |
| Shen et al. [19] | 2 | 1 | 1 | 2 | 2 | 2 | 10 | High |
| Hamdan et al. [20] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Liao et al. [21] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Li et al. [22] | 2 | 1 | 1 | 2 | 1 | 1 | 8 | Moderate |
| Ke et al. [23] | 2 | 1 | 2 | 2 | 1 | 2 | 10 | High |
| Nemec et al. [24] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Duarte et al. [25] | 2 | 1 | 1 | 1 | 1 | 1 | 7 | Moderate |
| Fully Met (score=2) | 24 | 2 | 6 | 6 | 14 | 6 | - | - |
| % Fully Met | 100% | 8.3% | 25.0% | 25.0% | 58.3% | 25.0% | - | - |
3.5. RQ1: Methodological Landscape
3.5.1. Dynamic Movement Primitives (DMP) and Variants
- Force-Controlled Dynamic Coupling DMPs (FDC-DMP): Introduced by Shen et al. [19], this framework adds a coupling term driven by interaction forces directly into the DMP acceleration equation, enabling the robot to dynamically modify its path (e.g., avoiding obstacles during bus body polishing) without altering the global target.
- B-Spline DMPs (BDMPs): Wang et al. [17] proposed replacing the standard Gaussian basis functions in DMPs with B-splines. This modification significantly improves trajectory modeling accuracy with a smaller number of basis functions. The forcing term is parameterized as , where represents the B-spline basis functions. When optimized using Policy Improvement with Path Integrals () reinforcement learning, BDMPs demonstrate high generalization capability for polishing trajectories and force profiles under unseen workpiece positions.
- Riemannian DMPs: Liao et al. [21] extended DMPs to Riemannian manifolds (e.g., Cartesian space and 2-D sphere manifolds) to simultaneously model human motion, 3-D endpoint stiffness, and contact forces from a one-shot demonstration, solved via Quadratic Programming (QP). Trajectories on the manifold are generated by mapping the states to the tangent space , maintaining geometrical properties of robot orientations.
- Neural Network-Augmented DMPs: Wang et al. [5] integrated a Phase-Modulated Diagonal Recurrent Neural Network (PMDRNN) with DMPs to adaptively predict trajectory offsets based on real-time force tracking deviations, mitigating environmental uncertainties.
3.5.2. Probabilistic and Statistical Models
- Gaussian Mixture Models and Regression (GMM-GMR): Used to model the joint distribution of time, space, and force parameters. Wu et al. [13] utilized GMMs to encode human polishing dynamics, combining GMR with a variable impedance controller to regulate contact compliance. Zhai et al. [8] integrated GMM-GMR with a vector-valued Gaussian Process (GP) to online-modulate robotic trajectories when subjected to human external forces.
- Probabilistic Movement Primitives (ProMPs): Unlike DMPs, ProMPs capture the statistical variance of demonstrations. ProMPs represent a trajectory as a linear combination of basis functions: , where captures the statistical variance across multiple demonstrations. Wang et al. [6] developed Arc-Length ProMPs (AL-ProMP), mapping the probability distribution of contact forces to spatial coordinates (arc-length ) rather than time t, formulating the trajectory as . This formulation prevents trajectory distortions during non-linear speed scaling.
- Fourier Movement Primitives (FMP): Grounded in signal processing, Kulak et al. [10] proposed FMPs using Fourier series as basis functions: . FMPs approximate periodic, multi-frequency signals (e.g., circular polishing patterns) from unaligned demonstrations without requiring phase alignment or frequency extraction.
3.5.3. Deep Learning and Generative AI Architectures
- Diffusion Policies: Ke et al. [23] and Li et al. [3] utilized diffusion models to generate continuous, expert-like motion-force trajectories. In the DP-RRL framework [3], the Diffusion Policy generates a trajectory distribution by iteratively denoising a random sequence using a noise predictor conditioned on observation O. The residual RL agent then predicts a displacement to correct the reference force based on contact dynamics.
- Neural Ordinary Differential Equations (Hyper-NODEs): Xu et al. [18] developed a Hyper-NODE architecture to generate continuous position and orientation (quaternion) trajectories. The system dynamics are modeled as:where represents the continuous hidden state, and are the weights generated by a hypernetwork. Paired with Control Barrier Functions (CBF), the system guarantees obstacle avoidance in local non-polishing areas (LNP-areas) while estimating admittance control parameters.
- Neural Network Morphing: Möhl et al. [15] designed a keypoint-driven non-linear morphing network to transfer demonstrated trajectories between 3D point clouds of geometrically similar objects without CAD models.
3.5.4. Autonomous Dynamical Systems (DS)
- Stable Limit Cycles: Duarte et al. [25] represented periodic human polishing movements (e.g., circular or elliptical motions) using a time-invariant DS with a stable limit cycle attractor. The system is formulated as a second-order nonlinear dynamical system of the form , where the trajectories are forced to converge asymptotically to a closed orbit . This formulation guarantees that the robot converges back to the demonstrated polishing pattern even after being physically displaced.
- DS-based Imitation Learning: Si et al. [4] proposed a dynamically stable energy field framework to provide virtual haptic guidance during teleoperated human demonstrations, reducing operator physical workload.
3.5.5. Direct Adaptive Control and Parameter Estimation
3.6. RQ2: System and Sensory Configuration
3.6.1. Sensory Modalities and Multimodal Perception
- Force and Torque Sensing: As the primary driver of closed-loop execution, force feedback is integrated in 87.5% of the studies (either as the sole sensor or in multimodal setups). While end-effector 6-DOF F/T sensors are standard, Hamdan et al. [20] proposed a dual-force sensor configuration. One sensor measures the human operator’s guiding force (), while the second measures the tool-workpiece interaction force (). By calculating the environmental reaction force:the system isolates the environmental dynamics (stiffness, friction) from human guidance inputs.
- Vision and Spatial Perception: To handle geometrically complex surfaces, depth sensors (e.g., overhead or wrist-mounted RGB-D cameras) are integrated. These sensors capture raw 3D point clouds, which are processed via PointNet++ or keypoint-based neural networks to reconstruct surface meshes [23] or guide trajectory morphing [15].
- Kinematic and Biometric Tracking: Teleoperation and kinesthetic demonstration interfaces utilize haptic devices (e.g., Geomagic Touch) or wearable inertial measurement units (IMUs). More advanced setups incorporate surface electromyography (sEMG) sensors on the human arm to capture synergistic muscle activity, translating muscle co-contraction directly into robot joint stiffness parameters.
- Instrumented Tools: To facilitate platform-independent demonstrations, Fischer et al. [9] designed custom instrumented tools and mechanical alignment plates to record high-quality contact forces and orientations directly on the workpiece.
3.6.2. Control Architectures
- Variable Impedance and Admittance Control: These strategies model the robot-workpiece interface as a mass-spring-damper system. The low-level dynamic behavior is governed by the admittance control law:where , , and are the desired mass, damping, and stiffness matrices, is the reference trajectory, x is the actual position, is the external interaction force, and is the target reference force. Variable impedance controllers (Wu et al., 2023, 2025) modulate stiffness and damping online. For instance, stiffness is lowered when transitioning onto hard, brittle materials (e.g., iron) to prevent impact chatter, and increased on soft materials (e.g., wood) to ensure uniform material removal.
- Iterative Learning Control (ILC): To compensate for repetitive tracking errors, ILC is integrated with impedance control. Zhang et al. [11] combined GMM trajectory models with Dynamic Time Warping ILC (DTW-ILC), iteratively updating the robot’s reference path based on stiffness estimation to achieve fast force convergence over multiple polishing passes.
3.6.3. Robotic Systems and Hardware Integration
3.7. RQ3: Evaluation and Performance Metrics
3.7.1. Kinematic and Trajectory Accuracy Metrics
- Dynamic Time Warping (DTW) Distance: Quantifies the spatiotemporal similarity between the demonstrated human trajectory and the robot’s executed path, especially when feed rates vary.
- Root Mean Square Error (RMSE): Calculates the spatial deviation (in millimeters) between the executed end-effector path and the demonstrated trajectory over N samples:
- Pearson Correlation Coefficient (r): Measures the shape similarity of the trajectories:with values closer to 1.0 indicating high imitation fidelity.
- Relative Smoothness ( for position, for orientation): Evaluates the jerk of the generated trajectory to ensure smooth robotic motion.
3.7.2. Force Tracking and Dynamic Interaction Metrics
- Force Root Mean Square Error (Force RMSE): The primary metric to evaluate force tracking performance, measuring the deviation between the executed contact force and the demonstrated reference force profile :
- Maximum Impact Force (): Evaluates system compliance and safety during initial tool contact or material transitions.
- Mean Force Deviation (): Quantifies the stability of the normal force during continuous polishing.
3.7.3. Surface Quality and Process-Specific Metrics
- Surface Roughness (Ra, Rq, Rz): Measured using contact profilometers or white-light interferometers. A reduction in average roughness () verifies successful surface smoothing.
- Material Removal Rate (MRR): Replicating the expert’s material removal strategy is evaluated based on Preston’s equation:where is the thickness of the removed material, is Preston’s coefficient (depending on tool and material properties), P is the contact pressure (directly proportional to normal force ), and v is the relative tool speed. Studies like Kim et al. (2023) and Min et al. [2] estimate these parameters online to adapt forces dynamically on curved surfaces.
- Remaining Stain Ratio (RSR): In cleaning applications, image segmentation is used to calculate the percentage of stains remaining on the surface post-execution.
3.8. RQ4: Industrial Deployment Challenges
3.8.1. The Sim-to-Real Gap and Complex Contact Dynamics
3.8.2. Demonstration Quality and Hardware Constraints
- Kinematic Interference: The physical weight and joint limits of the robot arm restrict the operator’s natural movement, leading to distorted demonstrations.
- Sensor Noise: High-speed spindle rotation and pneumatic tool vibrations generate significant mechanical noise, degrading force/torque sensor readings during the teaching phase.
- Cognitive Overload: Controlling the robot’s spatial path, tool orientation, and contact force simultaneously in real time places a high cognitive demand on the human expert.
3.8.3. Generalization to Complex Geometries and LNP-Areas
- Local Non-Polishing (LNP) Areas: Industrial components often feature functional geometry (e.g., threaded holes, slots, and ribs) that must remain untouched. Autonomously detecting and avoiding these LNP-areas while maintaining a constant normal force on the surrounding freeform surface remains an open control problem.
- Geometric Generalization: Trajectory generalization models (e.g., DMPs) often distort orientations when scaling trajectories to highly curved, non-planar workpieces, risking collision or uneven polishing.
3.8.4. Multimodal Perception and Computational Complexity
3.8.5. Need for Robust Human-in-the-Loop (HITL) Systems
4. Discussion
4.1. Methodological and Academic Perspectives: Temporal Evolution (2016–2026)
| Metric Category | Metric | Performance Baseline Range | Primary Algorithmic Family | Representative Studies |
| Force Tracking Accuracy | Force RMSE (Root Mean Square Error) | DMP + Adaptive Variable Impedance Control | Wang et al. [5], Liao et al. [21] | |
| Surface Quality | Average Surface Roughness () | (Initial: ) | GMM-GMR & ILC-Based Polishing | Min et al. [2], Zhang et al. [11] |
| Spatial Trajectory Accuracy | Trajectory RMSE & Pearson Correlation (r) | RMSE < , | Deep Generative AI (Diffusion Policy / Hyper-NODE) | Li et al. [3], Ke et al. [23], Xu et al. [18] |
4.2. Task-Specific Synthesis: Polishing vs. Deburring
4.3. Economic and Operational Implications for SMEs
4.4. Comparative Synthesis of the Included Studies
4.5. Identification of Research Gaps in the LfD Literature
4.6. How the Current Study Addresses These Research Gaps
- Addressing the Task Imbalance: By exposing the severe deficit in deburring research (only 12.5% of the corpus), this study provides a clear technical analysis of why deburring is exceptionally challenging (high-frequency transient contact, mechanical chattering) and catalogued the haptic and PSO-based control strategies used by the few successful deburring studies. This guides future researchers toward the control and haptic configurations necessary to tackle deburring.
- Providing a Mechatronic and Sensory Roadmap: This study maps the exact mechatronic configurations required to support LfD (such as the dual-force sensor configuration of Hamdan et al. to isolate human forces from environmental reaction forces). By documenting the frequency of force-only (70.8%) vs. multimodal (16.7%) setups, it provides a design guideline for building hardware platforms capable of perceiving both geometric and tactile environments.
- Synthesizing Quantitative Performance Benchmarks: To assist researchers in bridging the sim-to-real gap, this work extracted and synthesized concrete, quantitative baseline performance metrics reported in physical experiments (force tracking RMSE of , trajectory RMSE under , and finished surface roughness of ). These benchmarks provide a standard against which simulated policies can be verified and validated.
- Structuring Algorithmic Trade-offs: By categorizing the LfD algorithms into five methodological families (Table 6) and comparing their characteristics (Table 10), this study maps which algorithms are best suited for specific challenges. For instance, it highlights that while DMPs excel at spatial scaling, probabilistic models (like GMM-GMR) are better suited for online trajectory deformation during human intervention (HITL), and generative AI is best for visual point-cloud mapping, allowing practitioners to select the optimal control architecture.
5. Limitations of this Study
- Database Coverage and Search Strategy: The literature search was restricted to Scopus, Web of Science, IEEE Xplore, and Google Scholar. While these are the primary repositories for engineering and robotics research, some publications indexed in regional or specialized databases may have been omitted. Additionally, the keyword-based search queries centered around terms like "Learning from Demonstration" and "DMP" might have missed relevant studies that utilize alternative terminologies, such as "skill transfer" or "human-robot co-manipulation," despite sharing the same underlying architecture.
- Language Bias: In accordance with the screening protocol, only articles published in the English language were included. Given that 62.5% of the included literature originates from China, and that countries like Japan and Germany possess strong academic and industrial backgrounds in robotic manufacturing, excluding non-English publications likely introduced a language bias. Highly innovative papers published in Chinese, Japanese, or German may have been overlooked.
- Exclusion of Grey Literature: To ensure scientific rigor, only peer-reviewed journal articles and conference proceedings were included, while patents, technical white papers, and corporate reports were excluded. Because real-world industrial implementations of LfD are frequently protected as commercial trade secrets, the exclusion of grey literature may have limited our ability to map the exact degree of current commercial adoption.
- Selection Bias and Single Screener Limitation: Since the screening, eligibility selection, and data extraction processes were conducted by a single primary researcher rather than by two independent reviewers as recommended by the PRISMA 2020 [1] guidelines [1], there is an inherent risk of selection bias. Although borderline or ambiguous cases were re-evaluated multiple times to ensure coding consistency and adherence to the eligibility protocol, the lack of a second independent auditor represents a methodological limitation.
6. Future Directions
6.1. Multimodal Perception and High-Dimensional Datasets
6.2. Bridging the Sim-to-Real Gap and Safe Exploration
6.3. Interactive Human-in-the-Loop (HITL) Systems
7. Conclusions
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 2021, 372. [CrossRef]
- Min, K.; Ni, F.; Chen, Z.; Liu, H. A Force Control Method Integrating Human Skills for Complex Surface Finishing. Machines 2024, 12, 756. [CrossRef]
- Li, Y.; Lyu, Q.; Yang, J.; Salam, Y.; Wang, W. A Hybrid Framework Using Diffusion Policy and Residual RL for Force-Sensitive Robotic Manipulation. IEEE Robotics and Automation Letters 2025, 10, 10266–10273. [CrossRef]
- Si, W.; Jin, Z.; Lu, Z.; Wang, N.; Yang, C. A Stable Guidance Method for Teleoperation-based Robot Learning from Demonstration. In Proceedings of the IEEE International Conference on Automation Science and Engineering (CASE). IEEE, 2024, pp. 2376–2381. [CrossRef]
- Wang, Y.; et al. Adaptive Tuning of Robotic Polishing Skills based on Force Feedback Model. In Proceedings of the IEEE International Conference on Robotics and Biomimetics (ROBIO). IEEE, 2023, pp. 1–7. [CrossRef]
- Wang, Y.; Chen, C.; Peng, F.; Zheng, Z.; Gao, Z.; Yan, R.; Tang, X. AL-ProMP: Force-relevant skills learning and generalization method for robotic polishing. Robotics and Computer-Integrated Manufacturing 2023, 82, 102538. [CrossRef]
- Acikgoz, K.; Parvizi, P.; Donder, A.; Ugurlu, M.C.; Konukseven, E.I. Dynamic movement primitives and force feedback: Teleoperation in precision grinding process. In Proceedings of the International Conference on Electrical and Electronics Engineering (ELECO). IEEE, 2017, pp. 722–726. [CrossRef]
- Zhai, X.; Ou, Y.; Xu, Z.; Jiang, L.; Zhou, X.; Wu, H. Effective learning and online modulation for robotic variable impedance skills. In Proceedings of the IEEE International Conference on Robotics and Biomimetics (ROBIO). IEEE, 2022, pp. 1–6. [CrossRef]
- Fischer, A.; Unger, C.; Kugi, A.; Hartl-Nesic, C. Few-Shot Learning of a Force-Based Industrial Cleaning Process using an Instrumented Tool. IFAC-PapersOnLine 2025, 59, 103–108. [CrossRef]
- Kulak, T.; Silvério, J.; Calinon, S. Fourier movement primitives: an approach for learning rhythmic robot skills from demonstrations. In Proceedings of the Robotics: Science and Systems (RSS), 2020. [CrossRef]
- Zhang, R.; Xia, J.; Ma, J.; Huang, D.; Zhang, X.; Li, Y. Human-robot interactive skill learning and correction for polishing based on dynamic time warping iterative learning control. IEEE Transactions on Control Systems Technology 2024, 32, 2310–2320. [CrossRef]
- Wu, H.; Zhai, X.; Zheng, H.; Liao, Z.; Xu, Z.; Zhou, X. Learning stability-guaranteed skill and adaptive control strategies from demonstrations for heterogeneous component robotic machining. Journal of Manufacturing Processes 2025, 151, 506–520. [CrossRef]
- Wu, H.; Zhai, X.; Wu, X.; Gu, S.; Liao, Z.; Xu, Z.; Zhou, X. Learning Stable Nonlinear Dynamics and Interactive Force-Aware Variable Impedance Control for Robotic Contact Tasks. Procedia Computer Science 2023, 226, 127–133. [CrossRef]
- Haninger, K.; Hegeler, C.; Peternel, L. Model predictive impedance control with Gaussian processes for human and environment interaction. Robotics and Autonomous Systems 2023, 165, 104431. [CrossRef]
- Möhl, P.; Pratheepkumar, A.; Ikeda, M.; Pichler, A. Morphing based transfer of demonstrated surface finishing trajectories to point clouds of similar objects. Procedia Computer Science 2025, 253, 1002–1011. [CrossRef]
- Parvizi, P.; Ugurlu, M.C.; Acikgoz, K.; Konukseven, E.I. Parametrization of robotic deburring process with motor skills from motion primitives of human skill model. In Proceedings of the International Conference on Methods and Models in Automation and Robotics (MMAR). IEEE, 2017, pp. 373–378. [CrossRef]
- Wang, Y.; Chen, C.; Hong, Y.; Zheng, Z.; Gao, Z.; Peng, F.; Tang, X. PI2-BDMPs in combination with contact force model: A robotic polishing skill learning and generalization approach. IEEE/ASME Transactions on Mechatronics 2025. [CrossRef]
- Xu, X.; Qian, K.; Liu, A.; Yue, Z.; Huang, W. Polishing via ODEs: Adaptive admittance control for robot polishing based on Neural ODEs. Journal of Manufacturing Processes 2025, 155, 428–442. [CrossRef]
- Shen, N.; Mao, J.; Li, J.; Mao, Z. Research on trajectory learning and modification method based on improved dynamic movement primitives. Robotics and Computer-Integrated Manufacturing 2024, 89, 102748. [CrossRef]
- Hamdan, S.; Aydin, Y.; Oztop, E.; Basdogan, C. Robotic Learning of Haptic Skills from Expert Demonstration for Contact-Rich Manufacturing Tasks. In Proceedings of the IEEE International Conference on Automation Science and Engineering (CASE). IEEE, 2024, pp. 2334–2341. [CrossRef]
- Liao, Z.; Tassi, F.; Gong, C.; Leonori, M.; Zhao, F.; Jiang, G.; Ajoudani, A. Simultaneously learning of motion, stiffness, and force from human demonstration based on Riemannian DMP and QP optimization. IEEE Transactions on Automation Science and Engineering 2024, 22, 7773–7785. [CrossRef]
- Li, X.; Yang, C.; Feng, Y. The Generalization of Robot Skills Based on Dynamic Movement Primitives. IFAC-PapersOnLine 2020, 53, 265–270. [CrossRef]
- Ke, S.; Zhang, J.; Zhao, H.; Guo, Y.; Wei, Z.; Pan, J.; Ding, H. Visual-Guided Diffusion Policy and Mesh-DMP Integration for Robotic Freeform Surface Polishing. In Proceedings of the International Conference on Intelligent Robotics and Applications. Springer, Singapore, 2025, pp. 77–90. [CrossRef]
- Nemec, B.; Yasuda, K.; Mullennix, N.; Likar, N.; Ude, A. Learning by demonstration and adaptation of finishing operations using virtual mechanism approach. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2018, pp. 7219–7225. [CrossRef]
- Duarte, N.F.; Santos-Victor, J. Robot Imitation of Polishing Motions by Observing Humans: From Human Non-Verbal Cues to Stable Limit Cycles. In Proceedings of the IEEE International Conference on Development and Learning (ICDL). IEEE, 2024, pp. 1–7. [CrossRef]

| Criteria Category | Inclusion Criteria | Exclusion Criteria |
| Application Domain | Robotic deburring and polishing operations | Generic assembly, pick-and-place, or free-space trajectory tracking |
| Methodology | Learning from Demonstration (LfD, PbD, imitation learning, DMP, GMM, etc.) | Standard CNC programming, offline CAD/CAM paths, non-learning adaptive control |
| Evidence Type | Experimental validation or detailed simulation on finishing tasks | Abstract concepts without task validation, technical sheets, patents |
| Publication Medium | Peer-reviewed journal articles and conference papers | Dissertations, white papers, book chapters, editorials, abstracts |
| Language | English language only | Non-English language publications |
| Database | Search Query |
|---|---|
| Scopus | TITLE-ABS-KEY ( ( "Learning from Demonstration" OR "LfD" OR "Imitation Learning" OR "Programming by Demonstration" OR "DMP" OR "Dynamic Movement Primitives" OR "ProMP" OR "Probabilistic Movement Primitives" OR "Skill learning" OR "Task learning" ) AND ( "Deburring" OR "Polishing" OR "Surface finishing" OR "Surface-finishing" OR "Finishing" OR "Robotic finish*" ) ) AND PUBYEAR > 2015 AND PUBYEAR < 2027 AND ( LIMIT-TO ( DOCTYPE , "ar" ) OR LIMIT-TO ( DOCTYPE , "cp" ) ) AND ( LIMIT-TO ( LANGUAGE , "English" ) ) |
| Web of Science | TS=(( "Learning from Demonstration" OR "LfD" OR "Imitation Learning" OR "Programming by Demonstration" OR "DMP" OR "Dynamic Movement Primitives" OR "ProMP" OR "Probabilistic Movement Primitives" OR "Skill learning" OR "Task learning") AND ("Deburring" OR "Polishing" OR "Surface finishing" OR "Surface-finishing" OR "Finishing" OR "Robotic finish*")) AND PY=(2016-2026) AND DT=(ARTICLE OR PROCEEDINGS PAPER) AND LA=(ENGLISH) |
| IEEE Xplore | ("Learning from Demonstration" OR "LfD" OR "Imitation Learning" OR "Programming by Demonstration" OR "DMP" OR "Dynamic Movement Primitives" OR "ProMP" OR "Probabilistic Movement Primitives" OR "Skill learning" OR "Task learning") AND ("Deburring" OR "Polishing" OR "Surface finishing" OR "Surface-finishing" OR "Finishing" OR "Robotic finish*") |
| Google Scholar | ("Learning from Demonstration" OR "Imitation Learning" OR "Programming by Demonstration" OR "Dynamic Movement Primitives" OR "Probabilistic Movement Primitives") AND (Deburring OR Polishing OR "Surface finishing" OR "Surface-finishing" OR Finishing OR "Robotic finish*") |
| No | Authors | Year | Robot Type | Application Area | Algorithm Used |
| 1 | Min et al. | 2024 | Franka Emika Panda | Violin surface polishing | Computed-torque impedance control |
| 2 | Li et al. | 2025 | 7-DOF arm | Basin cleaning | Diffusion Policy + Residual RL |
| 3 | Si et al. | 2024 | - | Polishing, Ultrasound scanning | DS-based imitation learning |
| 4 | Wang et al. | 2023a | - | Polishing | PMDRNN + DMPs |
| 5 | Wang et al. | 2023b | - | Disc polishing | AL-ProMP |
| 6 | Acikgoz et al. | 2017 | Deburring machine | Grinding / Deburring | DMPs |
| 7 | Zhai et al. | 2022 | Franka Emika Panda | Button pressing / Polishing | GMM-GMR, Var. Impedance |
| 8 | Fischer et al. | 2025 | - | Surface cleaning | ProMPs (Few-Shot) |
| 9 | Kulak et al. | 2020 | Franka Emika Panda | Polishing, 8-shape drawing | Fourier Movement Primitives (FMP) |
| 10 | Zhang et al. | 2024 | - | Polishing | DTW-ILC + GMM |
| 11 | Wu et al. | 2025 | - | Machining heterogeneous components | PC-GMM-DS, Var. Impedance |
| 12 | Wu et al. | 2023 | Franka Emika Panda | Polishing, Grinding | GMM-GMR, Var. Impedance |
| 13 | Haninger et al. | 2023 | - | Collaborative polishing / Assembly | MPC + Gaussian Processes |
| 14 | Möhl et al. | 2025 | - | Trajectory morphing for finishing | Neural network morphing |
| 15 | Parvizi et al. | 2017 | - | Deburring | Modified DMPs (sDMP) |
| 16 | Wang et | 2025 | - | Polishing | PI2-BDMPs |
| 17 | Xu et al. | 2025 | - | Polishing (Rust removal) | Neural ODEs (Hyper-NODEs) |
| 18 | Shen et al. | 2024 | - | Bus body polishing | FDC-DMP |
| 19 | Hamdan et al. | 2024 | - | Polishing | MLP-based force learning |
| 20 | Liao et al. | 2024 | Franka Emika Panda | Polishing / Button pressing | Riemannian DMP + QP |
| 21 | Li et al. | 2020 | - | Desktop finishing | DMPs + Vision |
| 22 | Ke et al. | 2025 | - | Freeform polishing | Vision-Diffusion + Mesh-DMP |
| 23 | Nemec et al. | 2018 | - | Grinding, Polishing | Virtual mechanism + ILC |
| 24 | Duarte et al. | 2024 | - | Polishing | Dynamical system (Limit cycle) |
| Publication Year | Frequency | Percentage (%) | Cum. Percentage (%) |
| 2017 | 2 | 8.3% | 8.3% |
| 2018 | 1 | 4.2% | 12.5% |
| 2019 | 0 | 0.0% | 12.5% |
| 2020 | 2 | 8.3% | 20.8% |
| 2021 | 0 | 0.0% | 20.8% |
| 2022 | 1 | 4.2% | 25.0% |
| 2023 | 4 | 16.7% | 41.7% |
| 2024 | 7 | 29.2% | 70.8% |
| 2025 | 7 | 29.2% | 100.0% |
| Total | 24 | 100.0% | 100.0% |
| Country / Region | Frequency | Percentage (%) | Key Institutions |
| China | 15 | 62.5% | Huazhong University of Science and Technology, Harbin Institute of Technology |
| Turkey | 3 | 12.5% | Middle East Technical University, Koç University |
| Austria | 2 | 8.3% | PROFACTOR GmbH, Johannes Kepler University Linz |
| Germany | 1 | 4.2% | Fraunhofer Institute for Manufacturing Engineering and Automation |
| Slovenia | 1 | 4.2% | Jožef Stefan Institute |
| Switzerland | 1 | 4.2% | Idiap Research Institute / EPFL |
| Portugal | 1 | 4.2% | Instituto Superior Técnico, University of Lisbon |
| Total | 24 | 100.0% | - |
| Algorithmic Cluster | Frequency | Percentage (%) | Representative Methods |
| Dynamic Movement Primitives (DMP) & Variants | 9 | 37.5% | FDC-DMP, B-Spline DMP, Riemannian DMP |
| Probabilistic & Statistical Models | 8 | 33.3% | GMM-GMR, AL-ProMP, Fourier MP, Gaussian Processes |
| Deep Learning & Generative AI | 4 | 16.7% | Diffusion Policy, Residual RL, Hyper-NODEs, MLPs |
| Autonomous Dynamical Systems (DS) | 2 | 8.3% | Stable Limit Cycles, DS-based Imitation |
| Direct Impedance Control & Parameter Estimation | 1 | 4.2% | Computed-Torque Impedance control |
| Total | 24 | 100.0% | - |
| Sensory Modality | Frequency | Percentage (%) | Key Hardware Elements |
| Force / Torque Sensing Only | 17 | 70.8% | 6-DOF F/T sensors, Joint torque sensors |
| Vision Only (RGB / RGB-D) | 3 | 12.5% | Depth cameras, Point clouds |
| Multimodal (Force + Vision / Haptic) | 4 | 16.7% | RGB-D + F/T sensor, Haptic interface + Force feedback |
| Total | 24 | 100.0% | - |
| No | Study (Year) | LfD Category | Core Commonalities with Corpus | Unique Differences & Key Contributions |
|---|---|---|---|---|
| 1 | Min et al. [2] | Direct Impedance | Focuses on polishing; utilizes force control and collaborative robot platforms. | Decouples human demonstrations into separate "motion skills" (discrete pose sequences) and "force skills," validating on complex violin surfaces. |
| 2 | Li et al. [3] | Deep Learning | Focuses on cleaning/polishing; utilizes force control and collaborative robot platforms. | Combines a generative Diffusion Policy (for motion-force generation) with a Residual RL agent (for online force updates) using point clouds. |
| 3 | Si et al. [4] | Dynamical Systems | Focuses on polishing; utilizes force control and haptic teleoperation interfaces. | Introduces a dynamically stable energy field-based virtual haptic guidance force that decays iteratively to reduce operator workload. |
| 4 | Wang et al. [5] | DMP | Focuses on polishing; utilizes force control and collaborative robot platforms. | Integrates a Phase-Modulated Diagonal Recurrent Neural Network (PMDRNN) to adaptively predict trajectory offsets based on force errors. |
| 5 | Wang et al. [6] | Probabilistic | Focuses on polishing; utilizes force control and collaborative robot platforms. | Proposes Arc-Length ProMPs (AL-ProMP) to decouple force scaling and speed scaling in the spatial coordinate (arc-length) domain. |
| 6 | Acikgoz et al. [7] | DMP | Focuses on deburring; utilizes haptic teleoperation interfaces. | Developed for deburring; human guides a 1-DOF haptic knob, and a high-speed piezoelectric actuator executes micro-adjustments on the workpiece. |
| 7 | Zhai et al. [8] | Probabilistic | Focuses on polishing; utilizes force control and collaborative robot platforms. | Couples GMM-GMR with a vector-valued Gaussian Process to enable online trajectory deformation under human physical intervention. |
| 8 | Fischer et al. [9] | Probabilistic | Focuses on cleaning; utilizes haptic interfaces and movement primitives. | Developed a location-invariant few-shot cleaning framework using an instrumented manual tool to capture expert data independent of the robot platform. |
| 9 | Kulak et al. [10] | Probabilistic | Focuses on polishing; utilizes collaborative robots and joint torque sensing. | Uses Fourier series basis functions (FMP) to learn periodic tasks from unaligned demonstrations without temporal or phase registration. |
| 10 | Zhang et al. [11] | Probabilistic | Focuses on polishing; utilizes force control and collaborative robot platforms. | Combines GMM with Dynamic Time Warping Iterative Learning Control (DTW-ILC) to estimate environment stiffness and update reference paths. |
| 11 | Wu et al. [12] | Probabilistic | Focuses on polishing; utilizes force control and variable impedance architectures. | Tailored for Heterogeneous Material Components (HMCs); uses PC-GMM-DS and SMoGP to regulate rapid transitions across wood-iron splicing. |
| 12 | Wu et al. [13] | Probabilistic | Focuses on polishing; utilizes force control and variable impedance architectures. | Learns a non-parametric, globally stable GMM for rhythmic motions, optimizing variable impedance via GMR to minimize control torque. |
| 13 | Haninger et al. [14] | Probabilistic | Focuses on co-manipulation polishing; utilizes force control and collaborative robot platforms. | Captures task uncertainty using Gaussian Processes (GPs) and solves trajectory and impedance planning online using a non-linear MPC. |
| 14 | Möhl et al. [15] | Deep Learning | Focuses on trajectory transfer; utilizes depth cameras and vision data. | Direct trajectory transfer between scan point clouds of similar objects using keypoint-driven neural network morphing, without CAD models. |
| 15 | Parvizi et al. [16] | DMP | Focuses on deburring; utilizes haptic teleoperation interfaces. | Uses Particle Swarm Optimization (PSO) to parameterize sDMPs to capture expert force responses under sharp corner and circular geometries. |
| 16 | Wang et al. [17] | DMP | Focuses on polishing; utilizes force control and collaborative robot platforms. | Introduces B-spline DMPs (BDMPs) requiring fewer basis functions, optimized via Policy Improvement with Path Integrals () for generalization. |
| 17 | Xu et al. [18] | Deep Learning | Focuses on polishing; utilizes admittance control and collaborative robot platforms. | Uses Hyper-NODEs to generate smooth position-quaternion trajectories, combined with CLF/CBF for obstacle avoidance in LNP-areas. |
| 18 | Shen et al. [19] | DMP | Focuses on polishing; utilizes force control and collaborative robot platforms. | Introduces force-controlled dynamic coupling terms (FDC-DMP) using virtual coupling forces to dynamically alter local paths. |
| 19 | Hamdan et al. [20] | Deep Learning | Focuses on polishing; utilizes force control and admittance control. | Employs a dual-force sensor configuration to isolate human guide forces () from environmental reaction forces (), training an MLP. |
| 20 | Liao et al. [21] | DMP | Focuses on polishing; utilizes force control and collaborative robot platforms. | Uses Riemannian DMPs and QP optimization to simultaneously learn motion, 3-D endpoint stiffness, and applied forces from a one-shot demonstration. |
| 21 | Li et al. [22] | DMP | Focuses on finishing; utilizes collaborative robot platforms and movement primitives. | Pairs DMPs with machine vision object detection to automatically recognize workpiece locations and generalize trajectory paths. |
| 22 | Ke et al. [23] | DMP | Focuses on polishing; utilizes force control and depth cameras. | Combines a Diffusion Policy to generate continuous spatial actions from RGB-D images and embeds them on freeform meshes using Mesh-DMP. |
| 23 | Nemec et al. [24] | DMP | Focuses on polishing; utilizes force control and haptic teleoperation interfaces. | Models the tool as a "virtual mechanism" (augmented kinematic chain) for redundancy resolution, refining trajectories via Iterative Learning Control. |
| 24 | Duarte et al. [25] | Dynamical Systems | Focuses on polishing; utilizes collaborative robot platforms and human data. | Models all circular/ellipse polishing motions as a time-invariant dynamical system with a stable limit cycle attractor, mapping non-verbal human cues. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).