Submitted:
08 September 2026
Posted:
09 September 2026
You are already at the latest version
Abstract
Marine cable winch systems used in offshore operations are exposed to substantial tension fluctuations produced by vessel heave motion and irregular ocean-wave disturbances. These fluctuations can compromise operational stability and may contribute to undesirable slack–re-tension cycles, mechanical loading, fatigue, and reduced reliability. Conventional proportional-integral (PI) controllers remain attractive because of their simplicity and dependable steady-state behavior, but fixed-gain feedback control may be less effective when operating conditions and disturbance patterns vary. This study investigates a reinforcement-learning approach based on Proximal Policy Optimization (PPO) combined with Long Short-Term Memory (LSTM) recurrent memory for adaptive marine winch tension control. The controller is trained entirely in a Unity ML-Agents simulation environment containing a nonlinear drum-winch plant and a stochastic JONSWAP irregular-wave disturbance model. The agent uses a nine-dimensional observation vector comprising normalized tension, setpoint, tracking error, drum angular velocity, rope length, motor torque, dominant wave-phase sine and cosine components, and tension rate of change. A continuous normalized motor-torque command is generated at 60 ms decision intervals. Training employs a graduated curriculum over an 8-20 kN setpoint range, with an initial fixed 8 kN phase followed by randomized setpoints. The final reported model is evaluated against a manually tuned PI controller using randomized JONSWAP realizations. Across the combined operating range, the PPO-LSTM controller achieves a mean tension error of 1.724 kN compared with 1.780 kN for the PI controller, corresponding to a 3.1% improvement that is not statistically significant at the 0.05 level. More importantly, the reinforcement-learning controller reduces error standard deviation by 17.3% overall, with reductions of 26.3% and 39.2% in the 12-15 kN and 15-20 kN ranges, respectively. At the 95th percentile of errors, the overall reduction is 6.4%. The results indicate that the main benefit of the recurrent RL architecture is improved consistency and robustness rather than a large reduction in average tracking error. The study also identifies curriculum bias at low setpoints, increased slack incidence in the 8-12 kN range, catastrophic forgetting during targeted high-setpoint fine-tuning, and the simulation-to-reality gap as important limitations. The expanded discussion therefore emphasizes both the potential and the engineering constraints of PPO-LSTM control for marine winch applications.

Keywords:
reinforcement learning
; proximal policy optimization
; LSTM
; marine winch
; tension control
; JONSWAP
; irregular waves
; active heave compensation
; Unity ML-Agents
; simulation
1. Introduction
Marine winches constitute a fundamental part of offshore handling and lifting systems. They are used whenever cables, loads, subsea devices, or other suspended equipment must be deployed, recovered, positioned, or maintained from a vessel that is itself moving in a dynamic marine environment. The controlled movement of the winch drum determines the amount of cable deployed and therefore strongly influences the mechanical state of the suspended system. In a stationary industrial setting, tension regulation can often be treated as a relatively well-defined feedback-control problem. In offshore applications, however, the supporting vessel is continually affected by environmental excitation, and the resulting motion is transmitted through the cable to the load. The controller must therefore regulate tension while simultaneously compensating for a disturbance that is dynamic, irregular, and only partially predictable.
Wave-induced heave motion is particularly important because it acts directly along the vertical direction in which many offshore lifting operations are performed. A vessel moving upward or downward changes the effective geometry and velocity of the cable system. The resulting relative motion between the winch, rope, and suspended load can generate significant changes in tension even when the commanded winch operation is unchanged. Under irregular waves, the tension signal may contain multiple frequency components and may exhibit transient peaks, oscillations, and periods of reduced tension. Such behavior is undesirable because the mechanical system must repeatedly absorb changing loads.
The consequences of poor tension regulation extend beyond tracking error alone. Large tension peaks may increase structural loading and accelerate fatigue or mechanical wear, whereas insufficient tension can permit the rope to become slack. Slack conditions are particularly problematic because the subsequent restoration of tension can create impact-like transients. The source manuscript cites prior work showing alternating tension and slack conditions under irregular excitation and notes that repeated impacts following slack events can contribute to rope failure. These considerations establish a strong motivation for active heave compensation and robust control strategies capable of maintaining stable tension over a range of environmental and operating conditions.
Conventional PI and PID controllers remain widely attractive for this class of engineering problems. Their principal advantages are interpretability, modest computational requirements, straightforward implementation, and the possibility of tuning the controller using established control-engineering procedures. A PI controller can provide accurate tracking around a nominal operating point by responding to the instantaneous error and its accumulated history. Nevertheless, a fixed-gain PI controller is fundamentally reactive. If the sea state, payload condition, rope behavior, or other plant characteristics differ from those assumed during tuning, the controller may no longer provide the same quality of response. Feed-forward and model-based active heave compensation can reduce this limitation, but their performance depends on the quality and timing of disturbance measurements or predictions and on the adequacy of the underlying model.
Reinforcement learning provides a different perspective. Rather than explicitly prescribing the control law, an RL agent learns a policy through interaction with an environment. At each decision step, the agent observes a representation of the system state, selects an action, and receives a reward describing the quality of that action. Repeated interaction allows the policy to improve. In model-free RL, the learned policy does not require a complete analytical description of the plant dynamics, which is attractive for nonlinear systems in which accurate modeling can be difficult.
The present work focuses on PPO because it is a widely used policy-gradient method for continuous control. PPO constrains policy updates using a clipped surrogate objective, limiting excessively large changes to the policy during training. The approach is combined with LSTM recurrent memory because wave disturbances are sequential processes. A single instantaneous measurement may not contain sufficient information to characterize the current phase and recent evolution of a disturbance. A recurrent network can maintain an internal state and thereby exploit temporal dependencies in the observation sequence.
The paper uses a JONSWAP spectrum to represent irregular wave excitation. The JONSWAP formulation is appropriate for fetch-limited sea states and permits the generation of stochastic wave realizations by combining frequency components with randomized phases. In the simulation, the dominant peak frequency is randomized between episodes. The PPO-LSTM agent also receives sine and cosine representations of the dominant wave phase, giving it an implicit representation of wave-state information while avoiding a discontinuity associated with directly representing an angular phase variable.
The study is designed around three central contributions. First, it establishes a Unity ML-Agents simulation environment that couples a nonlinear drum-winch model with a JONSWAP disturbance generator. Second, it trains a PPO-LSTM controller using a graduated curriculum over an 8–20 kN setpoint range. Third, it compares the learned controller with a manually tuned PI controller over randomized episodes and evaluates not only average error but also error dispersion, high-percentile error, and slack incidence. The source manuscript describes this as the first application, within its stated scope, of PPO with LSTM recurrent memory to marine winch tension control in a JONSWAP simulation environment.
1.1. Background on Marine Winch Tension Control
The control objective in a marine winch can be expressed conceptually as maintaining the measured rope tension near a desired reference despite disturbances introduced by vessel motion and other dynamic effects. Let denote the measured tension and the desired tension. The instantaneous tracking error is e_T = T_setpoint − T. A controller must choose a motor torque that causes the drum to modify rope motion in a direction that reduces this error. The problem becomes more difficult when the rope is coupled to an external moving boundary, because the same drum action can have different effects depending on the current vessel motion, rope length, and dynamic state.
Active heave compensation is therefore closely related to the broader problem of tension control. A compensation system can use measurements or predictions of vessel motion to counteract disturbances before they propagate into the load. However, the quality of such compensation depends on sensors, estimation, communication, actuation, and model accuracy. The RL formulation examined here instead attempts to learn a policy that incorporates dynamic information directly into its decision process.
1.2. Motivation for Recurrent Reinforcement Learning
The key motivation for LSTM memory is temporal dependence. If wave-induced tension variations were independent from one sampling instant to the next, historical information would have limited value. In realistic irregular-wave signals, however, consecutive observations are correlated. The current phase of a dominant component and its recent evolution can provide information about the likely near-term behavior of disturbance. The source implementation explicitly supplies and and also uses recurrent hidden state. This combination allows the learned policy to condition its torque command on both instantaneous measurements and a compact representation of recent history.
This distinction is important when interpreting the expected advantage of the proposed controller. The principal hypothesis is not necessarily that RL will eliminate mean error. Instead, the hypothesis is that temporal memory and learned adaptation may produce a smoother response, reduce variability and limiting the magnitude of difficult excursions. The reported results support this interpretation: mean-error improvement is modest, whereas standard-deviation reduction is considerably larger.
1.3. Research Questions and Scope
The study can be understood through four practical questions. First, can PPO-LSTM achieve average tension-tracking performance comparable with a tuned PI controller? Second, can recurrent memory reduce the variability of tracking error under randomized JONSWAP excitation? Third, does the learned policy retain useful performance across different tension setpoints rather than specializing in one operating point? Fourth, what limitations become visible when the controller is trained through a curriculum or fine-tuned for a narrower operating range? These questions define the scope of simulation study.
The work remains a simulation study. No physical winch test bench or hardware-in-the-loop validation is reported in the source manuscript. Consequently, the numerical results should be interpreted as evidence about the behavior of the proposed policy within the stated simulation model rather than as proof of direct readiness for deployment on offshore hardware.
1.4. Structure of the Paper
· Section 1 (Introduction) presents the motivation and background of reinforcement learning (RL)-based tension control for marine winches operating under irregular wave disturbances. The section introduces the challenges associated with wave-induced vessel motion, rope-tension fluctuations, and the limitations of conventional fixed-gain PI control. It also presents the objective of research, the proposed PPO-LSTM control strategy, the use of JONSWAP irregular-wave disturbances, and the main contributions of the study.
· Section 2 (Marine Winch System and Dynamic Model) describes the simulated marine winch plant and its main physical components, including the torque-controlled drum winch, motor actuator, drum rotational dynamics, rope, and seabed-contact or submerged-payload condition. The section formulates the rope-tension dynamics using a first-order differential equation incorporating environmental stiffness, rope velocity, and tension relaxation. The relationships between drum angular velocity, vessel motion, rope velocity, motor torque, and tension are also presented, together with actuator saturation and physical operating constraints.
· Section 3 (JONSWAP Wave Disturbance and Simulation Environment) presents the modelling of irregular marine-wave excitation using the JONSWAP spectrum. The section describes the generation of stochastic wave realizations, wave-frequency and phase variations, and the coupling between wave-induced vessel motion and rope dynamics. The Unity-based simulation environment and its integration with the reinforcement learning framework are also described, including the observation, action, episode, and fault-reset mechanisms used during training and evaluation.
· Section 4 (PPO-LSTM Reinforcement Learning Control Methodology) presents the proposed reinforcement learning controller based on Proximal Policy Optimization (PPO) with Long Short-Term Memory (LSTM) recurrent dynamics. The section defines the observation vector, including tension, tension reference, tracking error, drum angular velocity, rope length, motor torque, wave-phase information, and tension rate. The continuous torque-control action, LSTM temporal representation, reward formulation, actuator constraints, and PPO training procedure are presented. The curriculum-learning strategy and randomized tension-reference range of are also detailed.
· Section 5 (Simulation Results and Comparative Evaluation) presents the performance evaluation of the proposed PPO-LSTM controller against the conventional PI controller under randomized JONSWAP irregular-wave disturbances. The section evaluates the controllers using mean absolute tension error, standard deviation of tension error, -percentile error, and slack-event rate. The results are analyzed over the complete operating range and within the , , and sub-ranges. Attention is given to tracking consistency, high-tension operation, statistical significance, and the effects of the training strategy.
· Section 6 (Discussion) analyses the principal findings and their implications for reinforcement learning-based marine winch control. The section discusses the comparable mean tracking accuracy of PPO-LSTM and PI control, the reduction in tension-error variability achieved by the recurrent RL controller, and the improved consistency observed at higher tension levels. It further examines the limitations associated with curriculum bias, low-tension performance, catastrophic forgetting during high-setpoint fine-tuning, and slack events. The section also discusses the implications for controller robustness, safety, generalization, and potential future deployment.
· Section 7 (Conclusions) summarizes the main findings of the study and evaluates the effectiveness of PPO-LSTM tension control under JONSWAP irregular-wave disturbances. The section highlights the principal performance improvements relative to PI control, particularly the reduction in tension-error variability, while acknowledging the absence of a statistically significant improvement in mean tracking error. Finally, the section outlines future research directions, including improved setpoint randomization, domain randomization, incorporation of wave-forecast information, continual-learning methods to mitigate catastrophic forgetting, and experimental validation using a physical marine-winch test platform.
2. Materials and Methods
2.1. Simulation Environment
The simulation environment was implemented in Unity 2022.3 LTS using the Unity ML-Agents Toolkit as is shown in Figure 1. The training configuration uses the ML-Agents Python package together with PyTorch as the deep-learning backend. Unity provides a configurable environment in which the winch plant, disturbance generation, observations, actions, episode logic, and fault-reset mechanisms can be executed repeatedly. This arrangement is particularly useful for RL because a large number of interactions can be generated without exposing physical equipment to unsafe exploratory actions.
Each training or evaluation episode has a maximum duration of 2000 decision steps. With a decision period of , this corresponds to approximately 120 seconds of simulated time. At the beginning of an episode, a new JONSWAP wave realization is generated and the tension setpoint is selected from the configured operating range. Consequently, the controller is exposed to different disturbance realizations and, during randomized training, different target tensions.
The simulation architecture can be divided into four functional blocks. The first block is the disturbance generator, which produces a stochastic wave realization and associated vessel-heave input. The second is the winch plant, which converts motor torque and relative rope motion into tension dynamics. The third is the observation interface, which extracts and normalizes plant and wave-state variables. The fourth is the PPO-LSTM policy, which converts the observation sequence into a continuous torque command. This command is scaled to the physical motor-torque limits before being applied to the simulated plant.
2.2. Winch Plant Model
The simulated plant represents a torque-controlled drum winch coupled through a rope to a seabed-contact or submerged-payload condition. The rope tension is represented by a first-order differential equation:
Here, represents environmental stiffness, is the effective rope velocity, and represents a tension-relaxation coefficient. The source model uses and . Rope velocity is derived from drum angular velocity together with wave-induced vessel motion.
This formulation captures three important effects at a reduced-order level. First, changes in relative rope velocity produce changes in tension through the environmental stiffness term. Second, tension is not assumed to persist indefinitely because the relaxation term causes a gradual decay. Third, the drum’s rotational state and external vessel motion jointly determines the effective rope velocity. The model therefore contains dynamic coupling between actuator input, drum motion, environmental excitation, and measured tension.
The motor is represented as a first-order torque actuator with saturation limits, while drum inertia is included through a rotational dynamics equation. Actuator saturation is important because it prevents the simulated policy from generating unrealistically large torque commands. The environment also contains a fault-reset mechanism that reinitializes an episode if tension exceeds physical limits. This mechanism prevents unstable early exploratory behavior from causing persistent numerical divergence during training.
2.3. JONSWAP Wave Disturbance Model
Ocean wave disturbances were modeled using the JONSWAP (Joint North Sea Wave Observation Project) directional wave spectrum [13], which is standard for North Sea fetch-limited sea states. The JONSWAP spectrum, is given by:
The peak enhancement factor is . The spectral-width parameter is for frequencies at or below the peak and above the peak. The normalization constant is selected according to the simulation formulation, as is shown in Figure 2.
The evaluation uses a significant wave-force level of , representing the moderate sea-state condition specified in the source. The spectrum is discretized into component frequencies between . Each component receives an independent random phase offset uniformly distributed over . The peak frequency is randomized between at the episode level. This procedure creates different time-domain wave realizations while preserving the overall statistical structure of the selected spectrum.
The use of multiple frequency components is important because a single sinusoidal disturbance would not adequately represent the irregularity that motivates the study. By varying the phase realization and peak frequency between episodes, the simulation exposes the controller to different temporal patterns. The agent additionally receives the sine and cosine of the dominant wave phase. This representation avoids the numerical discontinuity that would occur if phase were provided directly at the boundary.
for three representative peak frequencies with peak enhancement factor . (b) Sample time-domain heave displacement realization frequency components used as disturbance input to the winch plant model.
2.4. PPO-LSTM Agent Architecture
The controller uses PPO with LSTM recurrent memory in both actor and critic networks. The architecture contains three fully connected hidden layers of 256 units each, followed by an LSTM layer with hidden units. The LSTM sequence length is steps. Because the decision interval is , the sequence length corresponds to approximately 15.36 seconds of simulated history when interpreted continuously. This provides the recurrent controller with a substantial temporal window for learning relationships between wave evolution and tension response.
PPO is based on repeated policy improvement using trajectories generated by the current policy. The clipped objective restricts the ratio between the updated and previous policy probabilities to a specified interval. In the present training configuration, the clipping parameter is during the first phase and during the second phase. The lower value in the second phase reduces the permitted policy change as training moves from a specialized initial behavior toward the broader randomized operating range.
Generalized Advantage Estimation is used with lambda = 0.95, while the discount factor is gamma = 0.995. These settings determine how future rewards and multi-step temporal information contribute to the policy update. The entropy coefficient is reduced from 0.005 in Phase 1 to 0.003 in Phase 2. The learning rate is also reduced from 3.0 x 10^-4 to 1.0 x 10^-4. Batch size and buffer size are both maintained at 2048 and 20480, respectively, while observation normalization is enabled in both phases. Table 1 summarizes the key training hyperparameters.
2.5. Observation and Action Space
The agent receives nine scalar observations at each decision step. The first is the normalized current tension, followed by the normalized setpoint and normalized tracking error. The fourth observation is drum angular velocity, which provides information about actuator and mechanical motion. The fifth is deployed rope length, reflecting the changing configuration of the winch system. The sixth is the current motor torque. The seventh and eighth are the sine and cosine of the dominant wave phase. The ninth is the normalized tension rate of change. Therefore, the agent receives a nine-dimensional continuous observation vector at each decision step, as summarized in Table 2. The single continuous action output represents a normalized torque command in the range which is scaled to the motor torque limits before application to the plant.
2.6. Reward Function
The reward is designed to prioritize tension tracking while explicitly discouraging near-slack conditions.
where: , , and . The first term imposes a continuous penalty proportional to absolute tracking error. The second term is inactive when tension is safely above the slack threshold and becomes increasingly negative as tension approaches zero.
This reward design encodes an engineering preference that cannot be expressed adequately by mean error alone. Two controllers may produce similar average tracking error while differing substantially in the frequency and depth of low-tension events. By assigning an additional penalty to low tension, the RL agent is encouraged to avoid trajectories that approach rope slack. The evaluation nevertheless uses a separate slack metric, defined as the proportion of episodes in which minimum tension falls below , allowing the safety-related outcome to be assessed independently of the training reward.
2.7. Curriculum Training Strategy
Training was conducted in two phases using a graduated curriculum approach, as is shown in Figure 3. In Phase 1 (8 million steps), the setpoint was fixed at to allow the agent to develop basic tension-tracking behavior in low-tension conditions. In Phase 2 (5 million additional steps, initialized from Phase 1 weights), the setpoint was randomized uniformly from at each episode reset, with reduced learning rate and clip epsilon to prevent catastrophic forgetting of the Phase 1 policy. The Phase 2 model was used for all reported evaluations.
An additional Phase 3 fine-tune targeting the subrange (4 million steps, ) was also investigated but resulted in catastrophic forgetting at lower setpoints and was not adopted as the final model. This finding is discussed in Section 4.
covers the full setpoint range. Phase 2 begins at steps with reduced learning rate (). Smoothed curve: 20-point moving average.
2.8. PI Controller Reference
A discrete-time PI controller is implemented as the reference baseline. It uses the same simulated plant and the same JONSWAP disturbance conditions as the RL agent, ensuring that the comparison is performed within the same environmental framework. The PI gains are manually tuned to minimize mean tension error across the full operating range.
The PI controller receives only the current tension and the desired setpoint. It does not receive the explicit wave-phase information available to the RL controller. This configuration represents a conventional feedback baseline without active-heave feed-forward compensation. The comparison should therefore be interpreted as a comparison between a tuned fixed-gain feedback controller and a recurrent learning-based controller with access to additional state information.
2.9. Evaluation Protocol and Statistical Analysis
Both the PI controller and RL agent were evaluated to be over 300 independent episodes ( MaxSteps episodes; MaxSteps episodes) with setpoints randomized uniformly from and independent JONSWAP realizations at each episode (see Figure 4).
Four principal performance metrics are considered. MeanErr measures average absolute tension-tracking error. StdErr describes the dispersion of the tracking error and is used as an indicator of control consistency. P95 represents the 95th percentile of error and therefore characterizes a difficult upper-tail portion of the observed error distribution. Slack rate is the proportion of episodes whose minimum recorded tension is below . The results are also divided into low (), medium (), and high () setpoint ranges.
Statistical significance is assessed with the two-sided Mann–Whitney U test. This non-parametric test is appropriate for comparing two independent samples without assuming normally distributed error data. A significance threshold of is applied. The rank-biserial correlation coefficient is reported as an effect-size measure. This combination allows the study to distinguish between numerical improvements and improvements that meet the selected statistical significance criterion.
3. Results
3.1. Overall Performance Comparison
Across the combined operating range, the RL agent achieves a mean tension error of , while the PI controller achieves . The corresponding improvement is . However, the Mann–Whitney U test gives , so the overall mean-error difference does not reach the 0.05 significance threshold. The appropriate interpretation is therefore that the RL controller demonstrates comparable average performance rather than a statistically established improvement in mean error, as is shown in Table 3.
* Statistically significant at (two-sided Mann-Whitney U test).
3.2. Low Setpoint Range
In the range, the PI controller performs better in mean tracking error. The PI mean error is compared with for the RL controller, corresponding to a 4.3% regression for RL. The difference is statistically significant at . This is the clearest operating region in which the learned controller underperforms the baseline.
It can attribute this result to curriculum bias. The initial training phase uses a fixed setpoint for 8 million steps. Although the second phase introduces randomized setpoints, the learned policy begins its broader training from a strongly specialized initial condition. The observed regression suggests that this initial specialization may persist and affect generalization over the remainder of the low-tension interval.
The slack data provides an additional indication of the low-setpoint limitation. The RL slack rate is (), compared with () for PI. Mean minimum tension is also lower for RL in this range, 3.89 kN versus . These results suggest that the low-setpoint weakness is not restricted to average tracking error; it is also visible in the occurrence of low-tension events.
3.3. Medium Setpoint Range:
In the range, the two controllers exhibit essentially equivalent mean tracking performance. The PI mean error is and the RL mean error is , an improvement of only 0.6%. The p-value is 0.763, indicating no statistical evidence of a difference in mean error.
Despite the near-identical means, the RL controller has substantially lower error variability. The standard deviation decreases from for PI to for RL, representing a reduction. The 95th-percentile error also decreases from 2.032 to 1.946 kN, an improvement of . Thus, the principal benefit in this operating region is not better central tendency but tighter concentration of the error distribution.
3.4. High Setpoint Range: 15–20 kN
At higher tension setpoints, the RL controller provides its clearest mean-performance advantage. Mean error decreases from for PI to for RL, corresponding to a 2.7% improvement. The p-value of 0.065 is close to but above the selected threshold, so the difference is not conventionally statistically significant under the stated test criterion.
The variance result is much stronger. Error standard deviation falls from 0.283 kN for PI to for RL, a 39.2% reduction. The P95 error decreases from 2.566 to 2.433 kN, representing a 5.2% improvement. Both results indicate that the RL controller produces a more consistent response at the higher tension levels that are particularly relevant to demanding offshore lifting conditions.
Neither controller records a slack event in the N range. Thus, the advantage of the RL controller at high setpoints is not primarily a reduction in slack occurrence; rather, it is reflected in reduced tracking-error variability and improved upper-tail performance.
3.5. Variance Reduction and Extreme-Condition Robustness
The strongest aggregate finding is the reduction in error variability. Across the complete N range, standard deviation decreases from for PI to for RL, corresponding to a reduction. In the medium and high ranges, the reductions are and 39.2%, respectively. The low range shows a smaller reduction in standard deviation despite the worse mean error, as is shown in Table 4.
The overall P95 reduction is 6.4%. Because P95 focuses on the upper tail of the error distribution, the result indicates that the RL controller is less likely to produce very large tracking deviations under the most challenging realizations represented in the evaluation set. The improvement is particularly visible in the and ranges.
3.6. Slack Prevention
Table 5 summarizes the slack prevention performance, defined as episodes in which the minimum recorded tension fell below . At the range, both controllers achieved zero slack events. At , both controllers had equivalent slack rates ). At , the RL agent exhibited a slightly higher slack rate (), consistent with the overall regression at this sub-range. The overall slack rates were 2.1% (RL) vs. ), a modest difference that may warrant attention in safety-critical deployments.
4. Discussion
4.1. Interpretation of the Main Result
The results support a specific interpretation of the proposed PPO-LSTM controller. The principal advantage is not a dramatic improvement in average tracking accuracy. The overall mean-error reduction is only 3.1% and is not statistically significant. Instead, the most consistent advantage is a reduction in the dispersion of the error distribution. This distinction is important because a controller that produces a slightly better average error but substantially fewer large deviations may be attractive in an offshore mechanical system.
The high-setpoint results reinforce this interpretation. In the range, mean error improves by 2.7%, but standard deviation falls by 39.2%. The P95 error also decreases by 5.2%. The difference between these metrics indicates that the learned controller is changing the shape of the error distribution more strongly than its center. In engineering terms, the controller appears to produce more repeatable behavior even when its average tracking accuracy remains close to the baseline.
4.2. Possible Role of LSTM Temporal Memory
The recurrent architecture provides a plausible explanation for the observed reduction in variability. The agent does not operate solely on a single instantaneous measurement. Its LSTM maintains an internal state that summarizes information from previous observations. Because the wave disturbance is temporally correlated, recent observations can contain useful information about the phase and trend of future tension changes. The explicit sine and cosine phase inputs further expose a representation of the dominant wave state.
The evidence in the source manuscript does not isolate the contribution of LSTM from the contribution of PPO, wave-phase observations, or other architectural choices. Therefore, the improvement should not be interpreted as a proven causal effect of LSTM memory alone. A dedicated ablation study comparing PPO without recurrence, PPO-LSTM without wave-phase inputs, and PPO-LSTM with the complete observation vector would be needed to quantify the individual contribution of each element.
4.3. Comparison with Fixed-Gain PI Control
The PI controller has an important advantage: its behavior is relatively transparent and can be understood through its feedback structure and tuned gains. It does not require a large training process, a recurrent neural network, or a simulation environment. In the low setpoint region it also achieves better mean error and lower slack incidence. These characteristics should not be overlooked when evaluating an RL controller for safety-critical applications.
The RL controller, however, has a different information structure. It receives wave-phase information and multiple plant-state variables and can form a nonlinear mapping from the observation history to torque. This flexibility can be beneficial when the relationship between disturbance, plant state, and required control action is complex. The observed reduction in error variability suggests that the policy is using information unavailable to the PI baseline in a useful way.
A fair engineering comparison should therefore distinguish between controller complexity and control performance. The present results show that PPO-LSTM can reach PI-comparable mean performance while providing lower variability. They do not demonstrate that RL is universally superior to PI. Instead, they show that a recurrent learning-based policy can be competitive within the defined simulation environment and operating range.
4.4. Curriculum Bias at Low Setpoints
The low-range regression is one of the most informative findings in the study. Training begins at a single setpoint of for 8 million steps. This strategy makes the initial learning problem easier, but it also creates a strong distributional bias. When the policy is later exposed to the 8–20 kN range, it must adapt from a policy that has already been shaped around one operating point.
The statistical result at , where RL is 4.3% worse than PI with p = 0.020, indicates that the bias is not merely theoretical. It has a measurable effect on generalization. The most direct proposed remedy is to randomize setpoints from the beginning of training rather than maintaining a single initial target. Another possibility would be to use a shorter initial specialization phase or a curriculum in which the setpoint range expands progressively without remaining fixed for such a large fraction of total training.
4.5. Catastrophic Forgetting During Fine-Tuning
The Phase 3 experiment demonstrates a second learning-related limitation. The controller was fine-tuned for 4 million steps on the range using a learning rate of . Although the objective was to improve a region in which the RL policy was already competitive, the specialization caused substantial degradation at . This is a classic form of catastrophic forgetting: adaptation to a new distribution modifies parameters in a way that damages previously acquired behavior.
The finding has direct relevance to practical deployment. If a controller is retrained after new operating conditions are encountered, improving one operating region should not compromise established performance elsewhere. The source manuscript proposes elastic weight consolidation (EWC) and experience replay from earlier sub-ranges as possible remedies. Both approaches are consistent with the broader requirement that future RL training should preserve knowledge acquired across the operating envelope.
4.6. Simulation-to-Reality Gap
The most important deployment limitation is the difference between the simulated plant and a physical winch. The source model includes nonlinear dynamics, tension relaxation, actuator saturation, drum inertia, and wave-induced disturbance, but it does not capture every source of physical uncertainty. The manuscript specifically identifies rope-elasticity variation, sheave friction, and sensor noise as examples of omitted or simplified effects.
A policy can therefore learn behaviors that are highly effective in the simulated environment but less effective when exposed to unmodeled physical phenomena. This is the well-known simulation-to-reality gap. Domain randomization is one proposed strategy: rather than training on a single fixed model, parameters such as stiffness, friction, actuator characteristics, and measurement noise can be randomized within plausible ranges. The policy then learns behavior that is less dependent on one exact simulation configuration.
Physical validation is ultimately necessary before any safety-critical deployment. The source manuscript proposes validation on a scaled winch test bench. Such experiments would allow the controller to be evaluated under measured sensor noise, actuator delay, mechanical backlash, rope compliance, and other effects that are difficult to represent completely in a reduced-order simulation.
4.7. Safety Considerations
The increased slack rate in the low operating range deserves explicit consideration. Across all episodes, RL produces 7 slack events out of 330 compared with 3 out of 300 for PI. Although the absolute event counts are small, the difference matters because slack can have consequences beyond the instantaneous control metric. A controller intended for offshore lifting should therefore not be evaluated solely on mean error or even standard deviation. Minimum tension, peak tension, actuator saturation, rate of tension change, and recovery behavior following severe disturbances should also be monitored.
A practical RL controller would also require supervisory safety mechanisms independent of the learned policy. The simulation already contains a fault-reset mechanism when tension exceeds physical limits. Physical implementation could similarly employ hard torque limits, tension limits, emergency shutdown logic, and rule-based fallback control. Such mechanisms are especially important because neural policies can produce unexpected actions outside the training distribution.
4.8. Reproducibility and Benchmarking
The use of Unity ML-Agents provides a reproducible computational framework for repeated simulation-based experiments. The source manuscript specifies the principal plant parameters, wave-generation settings, observation variables, reward function, training phases, and major PPO-LSTM hyperparameters. These details establish a useful baseline for subsequent research.
Future benchmarking should extend this baseline through controlled ablations and repeated random seeds. For example, the statistical robustness of the reported differences could be evaluated across multiple independent training runs rather than a single final policy. Additional comparisons could include PI controllers with feed-forward compensation, model-predictive control, non-recurrent PPO, alternative recurrent architectures, and robust control methods. Such comparisons would clarify whether the observed performance is specific to the proposed architecture or represents a broader advantage of learning-based control.
4.9. Limitations of the Present Study
Several limitations should be considered when interpreting the results. First, the investigation is entirely simulation-based. Second, the JONSWAP evaluation uses a fixed significant wave-force level of 4.0 kN, even though peak frequency and phase realizations are randomized. Third, the PI baseline is a fixed-gain controller without wave feed-forward information, so the comparison does not represent every possible conventional AHC implementation. Fourth, the RL and PI evaluation sample counts differ, with 330 RL episodes and 300 PI episodes.
Fifth, the study does not report a formal ablation analysis that isolates the contribution of LSTM memory, phase information, reward shaping, or curriculum design. Sixth, the results are reported for the specified range and therefore should not be extrapolated beyond that range without additional testing. Finally, the simulation does not include all physical uncertainties that may affect hardware performance. These limitations do not invalidate the reported simulation findings, but they define the boundary within which those findings should be interpreted.
5. Conclusions
This expanded study presents a PPO-LSTM reinforcement-learning controller for marine drum-winch tension regulation under irregular JONSWAP wave disturbances. The controller is trained in a Unity ML-Agents environment containing a nonlinear winch model, stochastic wave excitation, actuator limits, a nine-dimensional observation vector, and a reward function that combines tracking accuracy with a penalty against near-slack conditions. A graduated curriculum is used to develop the policy from an initial 8 kN operating point toward randomized setpoints over 8–20 kN.
The main result is that the PPO-LSTM controller achieves mean performance comparable to the manually tuned PI baseline. Over the complete range, mean tension error is 1.724 kN for RL and 1.780 kN for PI, a 3.1% improvement that is not statistically significant at alpha = 0.05. Therefore, the study does not establish a statistically significant improvement in average tracking accuracy.
The stronger result concerns consistency. Overall error standard deviation decreases from 0.439 to 0.363 kN, corresponding to a 17.3% reduction. The reduction reaches 26.3% in the 12–15 kN range and 39.2% in the 15–20 kN range. The 95th-percentile error is reduced by 6.4% overall, with the largest improvements occurring in the medium and high setpoint ranges. These findings indicate that the RL policy can produce a more consistent tension response under the randomized wave conditions used for evaluation.
The study also demonstrates important limitations. At 8–12 kN, RL performs 4.3% worse than PI in mean error, with p = 0.020, and has a higher slack rate. The most plausible explanation identified in the source manuscript is bias introduced by the initial fixed-8 kN curriculum. Furthermore, targeted fine-tuning on 15–20 kN causes catastrophic forgetting at lower setpoints, illustrating the difficulty of improving one region of the operating envelope without degrading another.
The results support PPO-LSTM as a promising candidate for further investigation rather than as a completed replacement for conventional marine winch controllers. Its most notable advantage is enhanced consistency, particularly at higher tension setpoints, while its principal challenges include low-setpoint generalization, safe learning, and transfer from simulation to physical hardware. Future research should investigate randomized setpoint training from the beginning, domain randomization, explicit wave forecasts or Motion Reference Unit measurements, continual-learning methods such as elastic weight consolidation, and physical validation on a scaled winch test bench.
Overall, the work demonstrates the feasibility of using recurrent reinforcement learning to address marine winch tension regulation under irregular-wave excitation. The Unity ML-Agents environment and the reported PPO-LSTM benchmark provide a foundation for systematic comparison of learning-based and conventional control methods. A successful transition toward practical offshore deployment will require stronger validation, broader disturbance distributions, safety supervision, and careful treatment of uncertainty. Within the limits of the present simulation study, however, the results indicate that recurrent RL can match conventional PI control in average performance while offering a meaningful reduction in tension-error variability.
Author Contributions
Conceptualization, all authors; methodology: Cosmin Munteniță, Adrian Filipescu, Răzvan Șolea, Adrian Șerbencu; software: Cosmin Munteniță, Adrian Șerbencu; validation: Adrian Filipescu; formal analysis: Cosmin Munteniță, Adrian Filipescu, Răzvan Șolea; writing—original draft preparation: Cosmin Munteniță and Adrian Filipescu; writing—review and editing: Cosmin Munteniță and Adrian Filipescu; project administration: Cosmin Munteniță and Adrian Filipescu; funding acquisition, Cosmin Munteniță and Adrian Filipescu. All authors have read and agreed to the published version of the manuscript.
Funding
This article (APC) will be supported by “Dunarea de Jos” University of Galati.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
Data availability is not applicable to this article as the study did not report any data.
Conflicts of Interest
The authors declare no conflicts of interest.:
Abbreviations
The following abbreviations are used in this manuscript:
| AHC | Active Heave Compensation |
| EWC | Elastic Weight Consolidation |
| GAE | Generalized Advantage Estimation |
| JONSWAP | Joint North Sea Wave Observation Project |
| LSTM | Long Short-Term Memory |
| MPC | Model Predictive Control |
| MRU | Motion Reference Unit |
| PI | Proportional-Integral |
| PPO | Proximal Policy Optimization |
| RL | Reinforcement Learning |
| ROV | Remotely Operated Vehicle |
References
- Xie, T.; Huang, L.; Xu, J.; Guo, Y.; Ou, Y. Dynamic Analysis of Active Heave Compensation System for Marine Winch under the Impact of Irregular Waves. J. Mar. Sci. Eng. 2023, 11, 240. [CrossRef]
- Du, R.; Wang, N.; Rao, H. Modeling and Adaptive Boundary Robust Control of Active Heave Compensation Systems. J. Mar. Sci. Eng. 2023, 11, 484. [CrossRef]
- Kang, Y.; Fang, H.; Wang, C.; Liu, S. Research on Heave Compensation Systems and Control Methods for Deep-Sea Mining. J. Mar. Sci. Eng. 2025, 13, 652. [CrossRef]
- Chen, H.; Xie, J.; Han, J.; Shi, W.; Charpentier, J.-F.; Benbouzid, M. Position Control of Heave Compensation for Offshore Cranes Based on a Particle Swarm Optimized Model Predictive Trajectory Path Controller. J. Mar. Sci. Eng. 2022, 10, 1427. [CrossRef]
- Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; MIT Press: Cambridge, MA, USA, 2018.
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347.
- Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735-1780. [CrossRef]
- Park, H.; Chakir, S.; Kim, Y.; Lee, D.-H. A Robust Payload Control System Design for Offshore Cranes: Experimental Study. Electronics 2021, 10, 462. [CrossRef]
- Bae, J.-H.; Cha, J.-H.; Ha, S. Heave reduction of payload through crane control based on deep reinforcement learning using dual offshore cranes. Journal of Computational Design and Engineering 2023, 10(1), 414–424. [CrossRef]
- Juliani, A.; Berges, V.-P.; Teng, E.; Cohen, A.; Harper, J.; Elion, C.; Goy, C.; Gao, Y.; Henry, H.; Mattar, M.; Lange, D. Unity: A General Platform for Intelligent Agents. arXiv 2020, arXiv:1809.02627.
- Tobin, J.; Fong, R.; Ray, A.; Schneider, J.; Zaremba, W.; Abbeel, P. Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World. In Proceedings of the 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada, 24-28 September 2017; pp. 23-30.
- Lei, J.; Wang, Q.; Li, H.; Su, R.; Wang, L.; Liu, C. Coordinated Control of Automatic Drilling Feed and Heave Compensation for Offshore Hydraulic Hoisting Systems: A Co-Simulation Study. J. Mar. Sci. Eng. 2026, 14, 1184. [CrossRef]
- Hasselmann, K.; Barnett, T.P.; Bouws, E.; Carlson, H.; Cartwright, D.E.; Enke, K.; Ewing, J.A.; Gienapp, H.; Hasselmann, D.E.; Kruseman, P.; et al. Measurements of Wind-Wave Growth and Swell Decay During the Joint North Sea Wave Project (JONSWAP). Dtsch. Hydrogr. Z. Suppl. 1973, 8, 1-95.
- Mann, H.B.; Whitney, D.R. On a Test of Whether One of Two Random Variables Is Stochastically Larger than the Other. Ann. Math. Stat. 1947, 18, 50-60. [CrossRef]
- Kirkpatrick, J.; Pascanu, R.; Rabinowitz, N.; Veness, J.; Desjardins, G.; Rusu, A.A.; Milan, K.; Quan, J.; Ramalho, T.; Grabska-Barwinska, A.; et al. Overcoming Catastrophic Forgetting in Neural Networks. Proc. Natl. Acad. Sci. USA 2017, 114, 3521-3526. [CrossRef]
Figure 1.
System architecture of the Unity ML-Agents simulation environment for marine winch tension control. The PPO-LSTM agent receives a nine-dimensional observation vector and outputs a continuous torque command. The JONSWAP wave model provides stochastic heave disturbances to the winch plant model.
Figure 1.
System architecture of the Unity ML-Agents simulation environment for marine winch tension control. The PPO-LSTM agent receives a nine-dimensional observation vector and outputs a continuous torque command. The JONSWAP wave model provides stochastic heave disturbances to the winch plant model.

Figure 2.
JONSWAP wave disturbance model. (a) Directional wave energy spectrum

Figure 3.
PPO-LSTM training curve showing cumulative episode reward versus training steps. Phase 1,

Figure 4.
Performance comparison between PI controller and RL (PPO-LSTM, v15_ft) agent across setpoint sub-ranges. (a) Mean tracking error ± standard deviation; percentages indicate improvement of RL over PI (positive = RL better). (b) Relative reduction in standard deviation and 95th-percentile tracking error achieved by the RL agent versus PI controller.
Figure 4.
Performance comparison between PI controller and RL (PPO-LSTM, v15_ft) agent across setpoint sub-ranges. (a) Mean tracking error ± standard deviation; percentages indicate improvement of RL over PI (positive = RL better). (b) Relative reduction in standard deviation and 95th-percentile tracking error achieved by the RL agent versus PI controller.

Table 1.
PPO-LSTM training hyperparameters.
| Parameter | Phase 1 (8M steps) | Phase 2 (13M steps) |
| Learning rate | 3.0 × 10−4 | 1.0 × 10−4 |
| LR schedule | Linear decay | Linear decay |
| Clip epsilon (ε) | 0.20 | 0.15 |
| Entropy coefficient (β) | 0.005 | 0.003 |
| GAE lambda (λ) | 0.95 | 0.95 |
| Epochs per update | 3 | 2 |
| Batch size | 2048 | 2048 |
| Buffer size | 20480 | 20480 |
| Discount factor (γ) | 0.995 | 0.995 |
| Hidden units | 256 | 256 |
| Num. hidden layers | 3 | 3 |
| LSTM memory size | 256 | 256 |
| LSTM sequence length | 256 | 256 |
| Observation normalization | Enabled | Enabled |
Table 2.
Agent observation vector.
| Index | Observation | Description |
| o0 | Tension (normalized) | Current rope tension, normalized to setpoint range |
| o1 | Setpoint (normalized) | Target tension setpoint |
| o2 | Error (normalized) | Tension tracking error (setpoint − tension) |
| o3 | Angular velocity | Drum angular velocity ) |
| o4 | Rope length | Deployed rope length (normalized) |
| o5 | Motor torque | Current motor torque command (normalized) |
| o6 | Sine of dominant wave phase at peak frequency | |
| o7 | Cosine of dominant wave phase at peak frequency | |
| o8 | Tension rate of change, normalized by |
Table 3.
Mean tension error comparison: PI controller vs. RL v15_ft agent. : rank-biserial correlation (positive = RL better); Significance: * p < 0.05; n.s. = not significant.
Table 3.
Mean tension error comparison: PI controller vs. RL v15_ft agent. : rank-biserial correlation (positive = RL better); Significance: * p < 0.05; n.s. = not significant.
| Setpoint range | PI mean ± std (kN) | RL mean ± std (kN) | RL improvement | p-value |
| 1.295 ± 0.159 | 1.351 ± 0.144 | -4.3% | 0.020* | |
| 1.736 ± 0.198 | 1.726 ± 0.146 | +0.6% | 0.763 | |
| 2.180 ± 0.283 | 2.121 ± 0.172 | +2.7% | 0.065 | |
| 1.780 ± 0.439 | 1.724 ± 0.363 | +3.1% | 0.134 |
Table 4.
Error standard deviation and 95th percentile comparison. Positive values in P95 Improvement indicate RL advantage (lower worst-case error).
Table 4.
Error standard deviation and 95th percentile comparison. Positive values in P95 Improvement indicate RL advantage (lower worst-case error).
| Range | Std Reduction | RL P95 (kN) | P95 Improvement | |||
| 0.159 | 0.144 | −9.4% | 1.516 | 1.521 | ≈0.0% | |
| 0.198 | 0.146 | −26.3% | 2.032 | 1.946 | +4.2% | |
| 0.283 | 0.172 | −39.2% | 2.566 | 2.433 | +5.2% | |
| 0.439 | 0.363 | −17.3% | 2.465 | 2.308 | +6.4% |
Table 5.
Slack events ) and mean minimum tension.
| Range | PI slack rate | RL slack rate | PI T_min mean (kN) | RL T_min mean (kN) |
| 2.1% (2/95) | 4.8% (6/124) | 4.16 | 3.89 | |
| 1.2% (1/81) | 1.1% (1/90) | 5.37 | 5.56 | |
| 0.0% (0/124) | 0.0% (0/116) | 7.92 | 7.55 | |
| 1.0% (3/300) | 2.1% (7/330) | 6.04 | 5.63 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.