3. Rated-Condition Results and Discussion
3.1. Output-Voltage Regulation
Figure 4.
Output-voltage responses of the three control strategies.
Figure 4.
Output-voltage responses of the three control strategies.
All three strategies regulate the output voltage without overshoot. The PI+TPS baseline reaches a steady-state average voltage of 59.386298 V, corresponding to a voltage error of 0.613702 V. Direct D1-D2-D3 compensation yields 59.386370 V and an error of 0.613630 V. The difference from the baseline is only 72 μV in average voltage, showing that the learned direct actions provide essentially no additional voltage-regulation benefit in the investigated case.
The PI+TPS+RL-φ strategy reaches a steady-state average voltage of 59.566354 V and reduces the voltage error to 0.433646 V, a 29.34% reduction relative to the baseline. Its steady-state RMSE is also reduced from 0.618845 V to 0.440578 V. The improvement is therefore not limited to the final sample; it is maintained over the full steady-state evaluation window. This confirms a persistent improvement rather than a momentary numerical advantage.
The startup trajectories should be interpreted together with the current response. The φ-level strategy enters the 2% tolerance band without overshoot through a more gradual voltage build-up, so the shorter settling time does not imply a more aggressive initial command.
3.2. Current Stress and Dynamic Performance
The steady-state average current-stress index of PI+TPS is 4.548739. The RL-D123 value is 4.549014, a change of only +0.006%, which should be interpreted as practically unchanged rather than as a meaningful deterioration. The RL-φ value is 4.531973, corresponding to a modest 0.369% reduction. Although the steady-state reduction is small, it is achieved together with a clear improvement in voltage accuracy.
Table 4.
Steady-state performance under the rated condition.
Table 4.
Steady-state performance under the rated condition.
| Current-stress change |
Voltage-error change |
Average current stress |
Voltage error / V |
Average Vo / V |
Control strategy |
| — |
— |
4.548739 |
0.613702 |
59.386298 |
PI+TPS |
| +0.006% |
−0.012% |
4.549014 |
0.613630 |
59.386370 |
PI+TPS+RL-D123 |
| −0.369% |
−29.34% |
4.531973 |
0.433646 |
59.566354 |
PI+TPS+RL-φ |
Figure 5.
Current-stress responses of the three control strategies.
Figure 5.
Current-stress responses of the three control strategies.
The dynamic indices in
Table 5 provide additional information. The 2% settling time of RL-φ is 5.593 ms, compared with 70.986 ms for PI+TPS. Its full-transient peak inductor current is 36.8378 A, approximately 47.24% lower than the baseline value of 69.8237 A. This substantial peak reduction is consistent with the slower, more gradual initial voltage build-up visible in the waveform. Thus, the φ-level strategy does not simply make the response more aggressive; rather, it achieves faster entry into the 2% band while limiting the largest startup-current excursion.
Peak and RMS current describe different operating phenomena: the full-transient peak reflects startup stress, whereas the steady-state RMS value characterizes the conduction burden after regulation has been established. A lower startup peak can therefore coexist with nearly unchanged steady-state RMS current.
3.3. Steady-State Inductor Current
The three steady-state inductor-current waveforms have similar high-frequency shapes and nearly identical RMS values of approximately 6.132 A. Their steady-state current peaks are 14.278669 A, 14.264871 A, and 14.242303 A for PI+TPS, RL-D123, and RL-φ, respectively. The differences are small, which is consistent with the modest change in the average current-stress index. The main current-related advantage of RL-φ is therefore the reduction of the startup peak rather than a major alteration of the steady-state switching-current waveform.
Figure 6.
Steady-state leakage-inductor-current waveforms over selected switching periods.
Figure 6.
Steady-state leakage-inductor-current waveforms over selected switching periods.
3.4. TPS Variables
To examine whether the learned corrections alter the modulation hierarchy,
Figure 7 compares D1, D2, and D3 for all three strategies. The traces are separated by method so that startup transitions and steady-state ripple can be inspected without curve overlap.
The direct-compensation agent converges to small nearly constant offsets, with steady-state averages of approximately +5×10⁻⁴, −5×10⁻⁴, and +5×10⁻⁴ for ΔD1, ΔD2, and ΔD3. Consequently, the final TPS variables remain very close to those of the baseline. This observation explains why the voltage and current-stress performance of RL-D123 is almost unchanged.
For RL-φ, the final D1, D2, and D3 trajectories show small bounded periodic variations after the initial transient. These variations should not be described as instability: the variables remain confined around their steady operating values, the output voltage remains regulated, and the quantitative smoothness measures remain comparable to those of the baseline. The periodic component reflects fine discrete-time adjustment of the phase-shift command and the resulting TPS remapping.
3.5. Reinforcement Learning Compensation Outputs
Figure 8 isolates the learned compensation signals themselves. This comparison helps distinguish a genuinely state-dependent corrective policy from a nearly constant bias that produces only a small shift in the final operating point.
The RL-D123 compensation terms are of the order of 10⁻⁴ to 10⁻³ and remain close to constant values during the tested interval. The learned policy therefore behaves mainly as a small static bias correction rather than as a state-dependent dynamic compensation law. This is consistent with the negligible performance difference from PI+TPS.
The phase-shift correction remains close to 10⁻⁴ and exhibits a small bounded oscillation in steady state. Its steady-state average is 9.794274×10⁻⁵, and its maximum steady-state action jump is only 1.492277×10⁻⁷. The plotted vertical scale magnifies these variations; in absolute terms, they are small. The agent therefore provides fine phase-shift regulation without introducing a large disturbance into the baseline loop.
3.6. Control Smoothness
The CSI values are all in the same order of magnitude, and the differences are small. Accordingly, the results do not support a claim of a dramatic smoothness improvement. They do show that the RL compensation layers do not destroy the inherent smoothness of the TPS control variables. RL-D123 is almost indistinguishable from the baseline, whereas RL-φ gives the lowest CSI, the lowest maximum jump, and slightly smaller standard deviations. The reduction in CSI relative to PI+TPS is approximately 0.276%. This result supports the more limited conclusion that the small periodic variations observed in the RL-φ waveforms remain bounded and do not degrade overall modulation smoothness.
Table 6.
Steady-state control-smoothness comparison.
Table 6.
Steady-state control-smoothness comparison.
| Std(D3) |
Std(D2) |
Std(D1) |
Maximum control jump |
CSI |
Control strategy |
| 1.5510×10⁻² |
1.0340×10⁻² |
1.5510×10⁻² |
6.6612×10⁻³ |
9.1683×10⁻⁶ |
PI+TPS |
| 1.5511×10⁻² |
1.0340×10⁻² |
1.5511×10⁻² |
6.6521×10⁻³ |
9.1689×10⁻⁶ |
PI+TPS+RL-D123 |
| 1.5387×10⁻² |
1.0258×10⁻² |
1.5387×10⁻² |
6.6303×10⁻³ |
9.1430×10⁻⁶ |
PI+TPS+RL-φ |
3.7. Overall Discussion
The comparison shows that action-space dimension alone is not a measure of control capability. The three-dimensional RL-D123 agent has more direct authority, but the learned actions are small and provide no meaningful improvement in the tested condition. A plausible interpretation is that direct additive modification of strongly coupled TPS variables makes it difficult for the agent to discover a useful policy without additional structure, constraints, or richer operating-condition excitation [
5].
The one-dimensional RL-φ structure is more aligned with the hierarchy of the original controller. The agent modifies a high-level command, while the deterministic TPS mapping enforces the relationship among D1, D2, and D3. This reduces the learning dimension and embeds prior physical structure in the control architecture. Under the rated condition considered here, this arrangement improves steady-state voltage accuracy and startup peak current while keeping the final modulation variables bounded [
6,
7,
10].
The rated-condition comparison isolates the influence of the action insertion point, whereas the variable-input-voltage study evaluates cross-condition robustness after retraining. The two agents were trained independently and their optimizer settings were tuned separately; therefore, the results demonstrate a structural trend rather than a universal optimum. Multiple random seeds, identical training budgets, load steps, reference changes, wide-power-range operation, and experimental implementation remain necessary for a complete statistical and hardware-level validation.