Submitted:
31 October 2025
Posted:
03 November 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
2.1. Model-Based Reinforcement Learning
- Model Learning: estimation of the transition and reward models from interaction data, often through system identification, regression, or neural network approximation.
- Planning: utilization of the learned model to simulate trajectories, enabling the agent to evaluate policies without direct environment interaction.
- Policy Optimization: improvement of the policy or value function based on real and simulated experience.
2.2. Q-Learning
| Algorithm 1 Q-learning Algorithm |
|
3. Approach
3.1. Mass–Spring–Damper Dynamics with Hardening Nonlinearity
3.2. Piecewise Linear Model with Membership Functions (PLM)
3.3. Nonlinear Auto-Regressive Exogenous (NLARX)
4. Methodology


5. Performance Analysis
6. Discussion
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Keesman, K.J. System Identification: An Introduction; Springer: Berlin, Heidelberg, 2011. [Google Scholar]
- Ljung, L. System Identification. In Signal Analysis and Prediction; Birkhäuser Boston: Boston, MA, 1998; pp. 163–173. [Google Scholar]
- Landau, I.D.; Zito, G. Digital Control Systems: Design, Identification and Implementation; Springer: London, 2006. [Google Scholar]
- Chen, C.W.; et al. Integrated System Identification and State Estimation for Control of Flexible Space Structures. Journal of Guidance, Control, and Dynamics 1992, 15, 88–95. [Google Scholar] [CrossRef]
- Nelles, O. Nonlinear System Identification. Measurement Science and Technology 2002, 13, 646. [Google Scholar] [CrossRef]
- Majji, M.; Juang, J.N.; Junkins, J.L. Observer/Kalman-Filter Time-Varying System Identification. Journal of Guidance, Control, and Dynamics 2010, 33, 887–900. [Google Scholar] [CrossRef]
- Kuo, C.H.; Schoen, M.P.; Chinvorarat, S.; Huang, J.K. Closed-Loop System Identification by Residual Whitening. Journal of Guidance, Control, and Dynamics 2000, 23, 406–411. [Google Scholar] [CrossRef]
- Pintelon, R.; Schoukens, J. System Identification: A Frequency Domain Approach; Wiley: Hoboken, NJ, 2012. [Google Scholar]
- Lee, H.; Huang, J.K.; Hsiao, M.H. Frequency Domain Closed-Loop System Identification with Known Feedback Dynamics. In Proceedings of the AIAA Guidance, Navigation, and Control Conference, 1995.
- Huang, J.K.; Lee, H.C.; Schoen, M.P.; Hsiao, M.H. State-Space System Identification from Closed-Loop Frequency Response Data. Journal of Guidance, Control, and Dynamics 1996, 19, 1378–1380. [Google Scholar] [CrossRef]
- Chinvorarat, S.; Lu, B.; Huang, J.K.; Schoen, M.P. Setpoint Tracking Predictive Control by System Identification Approach. In Proceedings of the Proceedings of the 1999 American Control Conference. IEEE, 1999, Vol. 1, pp. 331–335.
- Chiuso, A.; Pillonetto, G. System Identification: A Machine Learning Perspective. Annual Review of Control, Robotics, and Autonomous Systems 2019, 2, 281–304. [Google Scholar] [CrossRef]
- Cui, M.; Khodayar, M.; Chen, C.; Wang, X.; Zhang, Y.; Khodayar, M.E. Deep Learning-Based Time-Varying Parameter Identification for System-Wide Load Modeling. IEEE Transactions on Smart Grid 2019, 10, 6102–6114. [Google Scholar] [CrossRef]
- Brunke, L.; Greeff, M.; Hall, A.W.; Yuan, Z.; Zhou, S.; Panerati, J.; Schoellig, A.P. Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning. Annual Review of Control, Robotics, and Autonomous Systems 2022, 5, 411–444. [Google Scholar] [CrossRef]
- Jaman, G.G.; Monson, A.; Chowdhury, K.R.; Schoen, M.; Walters, T. System Identification and Machine Learning Model Construction for Reinforcement Learning Control Strategies Applied to LENS System. In Proceedings of the 2022 Intermountain Engineering, Technology and Computing (IETC), 2022, pp. 1–6. [CrossRef]
- Farheen, N.; Jaman, G.G.; Schoen, M.P. Model-Based Reinforcement Learning with System Identification and Fuzzy Reward. In Proceedings of the 2024 Intermountain Engineering, Technology and Computing (IETC), 2024, pp. 80–85. [CrossRef]
- Ross, S.; Bagnell, J.A. Agnostic System Identification for Model-Based Reinforcement Learning. arXiv preprint arXiv:1203.1007 2012.
- Martinsen, A.B.; Lekkas, A.M.; Gros, S. Combining System Identification with Reinforcement Learning-Based MPC. In Proceedings of the IFAC-PapersOnLine, 2020, Vol. 53, pp. 8130–8135.
- Shuprajhaa, T.; Sujit, S.K.; Srinivasan, K. Reinforcement Learning-Based Adaptive PID Controller Design for Control of Linear/Nonlinear Unstable Processes. Applied Soft Computing 2022, 128, 109450. [Google Scholar] [CrossRef]
- Hafner, R.; Riedmiller, M. Reinforcement Learning in Feedback Control: Challenges and Benchmarks from Technical Process Control. Machine Learning 2011, 84, 137–169. [Google Scholar] [CrossRef]
- Sutton, R.S. Dyna, an integrated architecture for learning, planning, and reacting. ACM SIGART Bulletin 1991, 2, 160–163. [Google Scholar] [CrossRef]
- Moerland, T.M.; Broekens, J.; Jonker, C.M. Model-based Reinforcement Learning: A Survey. Foundations and Trends in Machine Learning 2023, 16, 1–118. [Google Scholar] [CrossRef]
- Watkins, C.J.; Dayan, P. Q-learning. Machine Learning 1992, 8, 279–292. [Google Scholar] [CrossRef]







| Model | MAE | STD | SSE [%] | Overshoot [%] | Settling Time [s] | Training Samples |
|---|---|---|---|---|---|---|
| NLARX | 0.31 | 0.09 | 30 | 12 | 7 | 12,000 |
| PLM | 0.03 | 0.10 | 3 | 10 | 7 | 60,000 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).