Preprint
Article

This version is not peer-reviewed.

Control System Design of a Pitch‐Decoupled VTOL UAV Using Reinforcement Learning

Submitted:

10 August 2026

Posted:

11 August 2026

You are already at the latest version

Abstract
The paper presents the design, development, and simulation of an intelligent control system for the pitch-decoupled vertical takeoff and landing (VTOL) unmanned aerial vehicle (UAV). In contrast to conventional tilt-rotor VTOL UAVs, the mechanical structure of the pitch-decoupled VTOL UAV allows passive transition from vertical to horizontal flight modes and vice versa without the use of additional servo actuators. A detailed kinematic scheme and the dynamic equations of motion of the pitch-decoupled VTOL UAV equipped with a two-degree-of-freedom robotic arm are developed using the Denavit-Hartenberg parameters and the Euler-Lagrange equations. The aerodynamic stability of the UAV is investigated, and the aerodynamic coefficients used in the complete dynamics model are derived. On this basis, a multivariable control system is developed. To address the well-known transition-phase issues of VTOL UAVs, a neural controller is designed using the reinforcement learning actor-critic method. It is shown that the proposed neural controller provides a smoother and more accurate transition phase and subsequent horizontal flight of the pitch-decoupled VTOL UAV compared with conventional and cascaded PID controllers.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Unmanned aerial vehicles (UAVs), including vertical takeoff and landing (VTOL) UAVs are currently widely used in both civilian and military applications. The demand for UAVs continues to grow because of their extensive range of applications in many important areas, including rescue operations, film production, package delivery, aerial photography, agriculture, reconnaissance, and military operations. Such wide-ranging applications stimulate the development of new UAV models that are more practical, versatile, easy to control, and energy-efficient.
The design and manufacturing of contemporary VTOL UAVs encounter various challenges, including control-system accuracy, energy consumption related to structural and mechanical design, and versatility of application. VTOL UAVs represent one of the most rapidly developing and scientifically significant classes of UAVs. The design of their control systems often requires fundamentally new approaches and methodologies, owing to their structural characteristics and diverse flight regimes.
Generally, there are three main types of multirotor VTOL UAVs, as well as several non-standard configurations [1,2,3]. The simplest configuration, typically used in small-sized vehicles, is a hybrid model that combines a multirotor UAV and a fixed-wing UAV (Figure 1(a)). The other two principal types of multirotor VTOL UAVs, conventionally referred to as tilt-rotor and tilt-wing VTOL UAVs, are widely used in practical applications (Figure 1(b),(c)) [4]. The main distinction between tilt-rotor and tilt-wing VTOL UAVs and the hybrid model in Figure 1(a) is the requirement for additional actuators (servomotors), which rotate either the rotors relative to the wings or the entire wings with fixed rotors relative to the UAV body. These rotations are necessary when switching between vertical takeoff and landing modes and the horizontal flight mode [5].
Irrespective of the multirotor VTOL UAV type, there are three distinct flight phases or modes:
  • Multirotor mode (vertical takeoff and/or landing),
  • Transition mode (switching between vertical and horizontal flight),
  • Fixed-wing mode (horizontal flight).
Essentially, the control of takeoff, landing, and hovering modes can be achieved using conventional control algorithms developed for standard multirotor UAVs (quadcopters, octocopters, etc.) and fixed-wing aircraft. However, the transition mode is unique to VTOL UAVs and presents numerous challenges because of the highly nonlinear and complex nature of the vehicle dynamics during this mode. In addition, each VTOL UAV configuration shown in Figure 1 requires individual consideration because of its specific structural characteristics.
During the transition phase, strong mechanical and aerodynamic forces act on the UAV body due to variations in the thrust-vector angle. This may result in thrust loss, attitude instability, and increased complexity of the overall control problem.
In general, traditional linear controllers, such as PID, LQR, and pole-placement controllers, do not perform satisfactorily during the transition phase. Consequently, many researchers investigate various nonlinear control approaches and hybrid strategies, including gain-scheduling, sliding mode control, backstepping, and others. In some studies, the transition-control problem is often formulated as an optimal-control problem aimed at deriving optimal control laws for the transition phase. For example, in [6], the transition control of a ducted-fan VTOL UAV is formulated as a constrained Markov decision process, and constrained policy optimization is used to learn safe control policies that regulate state transitions while satisfying safety constraints. However, the solutions proposed in [6] are rather sophisticated and require significant computational resources and complex numerical algorithms [7,8].
One of the major drawbacks of the tilt-rotor and tilt-wing VTOL UAVs in Figure 1(b) and 1(c), compared with the hybrid configuration in Figure 1(a), is the necessity for additional actuators that provide transitions between takeoff/landing and horizontal flight modes.
To overcome this disadvantage, Apkarian [9,10,11] proposed a new type of tilt-rotor VTOL UAV called the pitch-decoupled UAV, which was patented in the United States [9]. The overall view of the pitch-decoupled VTOL UAV equipped with a two-degree-of-freedom robotic arm is shown in Figure 2. The main idea behind J. Apkarian’s patent is illustrated in Figure 3.
Similar to standard hybrid VTOL UAV in Figure 1(a), the fuselage of the pitch-decoupled UAV consists of two parts: the quadcopter part and the fixed-wing aircraft part (hereafter referred to as the quad and wing parts). The rotors are passively coupled to the aircraft frame about the pitch axis through a parallelogram hinge mechanism. During takeoff and landing, the propellers are positioned vertically; therefore, the UAV operates similarly to a conventional quadcopter (Figure 2 (a)). During longitudinal (horizontal) flight, the propellers rotate by 90 degrees relative to the UAV frame and function similarly to those of a standard tilt-rotor VTOL UAV (Figure 2 (b)). A unique feature of this design is that the front and rear propellers are kinematically linked, and their rotation relative to the UAV body is achieved not by additional servomotors, as in the configuration shown in Figure 1(b), but by differential thrust generated by the front and rear rotor pairs. To improve control of the pitch axis, the UAV is additionally equipped with a fifth rotor located near the tail. The VTOL UAV also includes ailerons and an elevator to provide additional flight control during fixed-wing operation.
The main structural and kinematic features of the pitch-decoupled VTOL UAV, compared with a conventional tilt-rotor UAV, are considered in [10]. Based on a simplified mathematical model, it is shown that the vertical component of the rotor thrust depends only on the angle γ , which is defined as the difference between aircraft pitch angle and the relative tilt angle of the rotor hinge mechanism. The major challenges associated with controlling this type of VTOL UAV using cascaded PID controllers are discussed in the tutorial [11].
The novel pitch-decoupled VTOL UAV concept has attracted considerable attention from researchers in this field. In [12], a single PID-based controller is proposed that seamlessly manages the transition from vertical to longitudinal flight. A more sophisticated nonlinear dynamic model of the pitch-decoupled VTOL UAV, formulated as a constrained system of three rigid bodies using the Lagrangian approach, is presented in [13]. The model developed in [13] also includes an aerodynamic subsystem that generates forces and moments acting on the wing and fuselage of the aircraft. A detailed mathematical model of the pitch-decoupled VTOL UAV equipped with a robotic arm is developed in [14], and a PID-based control system for hovering mode is investigated using a Simulink model.
This paper is devoted to the development of a neural controller for a pitch-decoupled VTOL UAV equipped with a robotic arm, which provides smooth transitions between vertical takeoff and landing modes and horizontal flight mode.
The paper is organized as follows. Section 2 presents the kinematic analysis of the pitch-decoupled VTOL UAV equipped with a two-degree-of-freedom (DOF) robotic arm using the Denavit-Hartenberg formalism. In Section 3, the dynamic equations of the UAV’s angular rotations are derived using the Euler-Lagrange equations. Section 4 presents the translational dynamics of the UAV. Section 5 is devoted to the design of the attitude control system based on the developed mathematical model and conventional PID controllers. Section 6 presents the design of a matrix neural controller based on the reinforcement learning actor-critic method.

2. Kinematics of the Pitch-Decoupled UAV

In this section, a detailed kinematic scheme and the kinematic equations of the pitch -decoupled VTOL UAV equipped with a two-DOF robot manipulator are developed based on the parameters, which can be regarded as certain modifications of the well-known Denavit-Hartenberg (DH) parameters used in robotics [15]. Such an approach allows both the relative angular positions and linear displacements to be represented in a united form for an arbitrary number of reference frames (coordinate systems) using homogeneous transformation matrices [16,17]. The generalized coordinates and reference frames chosen for the derivation of kinematic equations are shown in Figure 4, where the quad part of the VTOL UAV is modeled as a flat plate with four fixed rotors.
Let { O I } denote a right-hand inertial reference frame with the origin at O I and the axes denoted X I ,
Y I , and Z I . , and let { O i } ( i = 0 , 1 , 2... , 6 ) denote orthogonal reference frames with origins at O i and axes X i ,
Y i , and Z i .
In Figure 4, { O 0 } is a reference frame with the origin O 0 at the center of mass (CM) of the VTOL UAV with the robotic arm. The axes of { O 0 } are aligned with the axes of the inertial frame { O I } . The orientation of the VTOL UAV with respect to { O 0 } is determined by three successive rotations of the intermediary frames { O 1 } , { O 2 } , and { O 3 } through the yaw, roll, and pitch angles. These angles determine the first three generalized coordinates q 1 (yaw), q 2 (roll) and q 3 (pitch) of the UAV. The corresponding homogeneous transformation matrix can be written in the form:
T 3 ( q 1 , q 2 , q 3 ) = R 0 ( q 1 , q 2 , q 3 )       0 3 x 1           0 1 x 3                       1 (1)
In the matrix T 3 ( q 1 , q 2 , q 3 ) (1), 0 3 x 1 and 0 1 x 3 are zero vectors of the indicated sizes, and R 0 ( q 1 , q 2 , q 3 ) is the following orthogonal rotation matrix [18]:
R 0 ( q 1 , q 2 , q 3 ) = c q 1 c q 3 s q 1 s q 2 s q 3         c q 2 s q 1           c q 1 s q 3 + c q 3 s q 1 s q 2   c q 3 s q 1 + c q 1 s q 2 s q 3               c q 2 c q 1           s q 1 s q 3 c q 1 c q 3 s q 2         c q 1 s q 3                                       s q 2                                 c q 2 c q 3 , (2)
where c and s are shorthand forms for cosine and sine, respectively. Note that the matrix R 0 ( q 1 , q 2 , q 3 ) (2) determines the orientation of the UAV’s body-fixed frame { O 3 } with respect to the inertial frame { O I } .
The next two generalized coordinates, q 4 and q 5 , describe rotations of the robotics arm with respect to the body-fixed frame { O 3 } . Finally, the sixth generalized coordinate, q 6 , is the pitch angle of the quad part of the UAV relative to the wing part [10]. Each of these three generalized coordinates q 4 ,   q 5 , and q 6 has an associated reference frame and a corresponding homogeneous transformation matrix A i q i . To develop the configuration kinematic equations for the pitch-decoupled UAV, we introduce the following DH-like parameters and basic transformations: a i , the linear displacement of the reference frame { O i } relative to { O i 1 } along the axis X i ; d i , the linear displacement of the reference frame { O i } relative to { O i 1 } , along the axis Z i ; α i , the rotation (twist) of { O i } about the axis X i ; q i ,   the rotation of { O i } about the axis Z i . Each basic transformation has an associated homogeneous transformation matrix. The product of these matrices yields the resulting homogeneous transformation matrix
A i q i =   cos q i   sin q i , 0 a i cos α i sin q i cos α i cos q i   sin α i 0 sin α i sin q i cos q i sin α i   cos α i d i 0 0 0 1   (3)
I = 4, 5, 6.
The introduced parameters a i ,   d i ,   α i ,   q i , and the matrix A i q i (3) differ from the standard DH parameters [15] but are more convenient for describing the kinematics of the UAV in Figure 4. Based on the kinematic model shown in Figure 4, we have
A 4 q 4 = c q 4     s q 4         0           0 s q 4         c q 4         0         0 0                 0             1         d 4 0                 0             0         1   , A 5 q 5 =     c q 5       s q 5         0           0     0               0               1           0 s q 5     c q 5         0         d 5       0               0               0         1 , (4)
A 6 q 6 = c q 6         s q 6           0           a 6   0                   0                 1             0 s q 6     c q 6             0             0 0                     0                 0             1     (5)
Besides the presented reference frames { O 0 } { O 6 } , an additional frame { O b } is introduced. This frame is not associated with any generalized coordinate but is required to describe the position and orientation of the tail rotor and of the quad part of the UAV. The corresponding homogeneous transformation matrix has a simple form
A b =   1   0 0 0 0 1   0 0 0 0 1 d b 0 0 0 1   (6)
By multiplying the introduced matrices A i q i in accordance with the kinematic model in Figure 4, we obtain the position and orientation of each reference frame relative to the initial reference frame { O 0 } . Performing the multiplications yields the following homogeneous transformation matrices:
T 4 = T 3 A 4 ,       T 5 = T 4 A 5 = T 3 A 4 A 5 ,       T b = T 3 A b ,       T 6 = T b A 6 . (7)
To calculate the torques generated by the rotors, it is necessary to determine the rotor displacements with respect to the CM O 0 of the UAV. Note that the positions r 1 6 ,   r 2 6 ,  
r 3 6 ,   r 4 6 of the four rotors on the quad part in the reference frame { O 6 } are determined as follows:
r 1 6 = 0.7071 L R           0 0.7071 L R           1   ;       r 2 6 = 0.7071 L R           0 0.7071 L R           1   ;     r 3 6 = 0.7071 L R           0 0.7071 L R           1   ;     r 4 6 = 0.7071 L R           0 0.7071 L R           1   , (8)
where L R is the distance from the rotors to the origin O 6 , and the fourth, a unity, coordinate is added to conform with the homogeneous matrix notations.
Similarly, the displacement vector of the fifth (tail) rotor in the reference frame { O b } can be defined as
r 5 O b = L 5 0 0 1 T .   ( 9 )
Multiplication of the rotor displacement vectors by the corresponding transformation matrices yields
r 1 0 = T 6 r 1 6 ,       r 2 0 = T 6 r 2 6 ,       r 3 0 = T 6 r 3 6 ,       r 4 0 = T 6 r 4 6 ,       r 5 0 = T b r 5 0 b , (10)
the first three coordinates of which determine the rotor displacements (rotors positions vectors) in the reference frame { O o }, which is attached to the CM of the system. In what follows, we denote the indicated three-dimensional vectors by r 1 R ,   r 2 R ,   r 3 R ,   r 4 R ,   r 5 R . Furthermore, let us denote γ 6 a three-dimensional vector composed of the first three entries of the second column of the matrix T 6 , and Z b , a vector formed in the same manner from the third column of T b :
γ 6 = T 6 ( 1 , 2 ) T 6 ( 2 , 2 ) T 6 ( 3 , 2 ) ,           Z b = T b ( 1 , 3 ) T b ( 2 , 3 ) T b ( 3 , 3 ) . (11)
Note that the vector γ 6 determines the direction of the thrusts F 1 ,   F 2 ,
F 3 ,   F 4 , generated by the first four rotors, in the frame { O 6 } , and Z b determines the direction of the thrust F 5 in the frame { O b } .
Then, the total torque M ¯ generated by the five rotors can be expressed as
M ¯ = M b + k φ 5 F 5 Z b   +   i = 1 4 M i + 1 i 1 k ψ F i γ 6   . (12)
In the equation (12),
M i = r i R × γ 6 F i ,         i = 1 , 2 , 3 , 4 (13)
M b = r 5 R × Z b F 5 ,   (14)
where × is the symbol of cross product, are the torques acting on the UAV and generated by all five rotors, k ψ is a constructive parameter determining the reaction torques acting on the airframe by the rotors of the quad part, and k φ 5 is the same parameter of the tail (fifth) rotor [3].
Note that in the vertical takeoff and landing flight modes, the generalized coordinate q 6 (the passive pitch angle γ of the quad part) is equal to zero and the total vertical thrust F Σ of the quad part rotors is
F Σ = F 1 + F 2 + F 3 + F 4 . (15)
This thrust is directed along the body-fixed axis Z 3 . In the horizontal flight mode, the passive pitch angle is equal to 90 , and the total thrust F Σ (15) is directed along the axis X 3 .
It should also be noted that, besides the rotor thrusts, the generalized coordinates q 1 ,    
  q 2 ,      
q 3 , which determine the orientation of the UAV, are affected by tilts of the elevator and the ailerons. This point is discussed further in the derivation of the dynamics equations of the pitch-decoupled UAV.
The block diagram shown in Figure 5 illustrates schematically the kinematic links of the pitch-decoupled UAV.
Here, the arrows directed from left to right represent the torques or forces generated by the rotors, actuators of the robotic arm ( F 4 ,       F 5 ), the ailerons ( F 6 ), and the elevator plus the fifth (tail) rotor ( F 7 ) ;   τ r and τ y are the torques generated on the quad frame about the roll and yaw axes, respectively, and R γ is the rotation matrix
R γ = cos ( γ )         sin ( γ ) sin ( γ )                 cos ( γ ) = cos ( q 6 )         sin ( q 6 ) sin ( q 6 )                   cos ( q 6 ) (16)

3. Rotational Dynamics of the Pitch-Decoupled UAV

The dynamics equations of angular rotations of the pitch-decoupled VTOL UAV are derived using the Euler-Lagrange equations, which for a mechanical system with six degrees of freedom have the following form [19,20]:
d d t L q ˙ i L q i = F i i = 1 , 2 , 3 , 4 , 5 , 6 , (17)
where
L = K P (18)
is the Lagrangian, defined as the difference between the total kinetic K and potential P energies; q i are the generalized coordinates, and F i are generalized forces arising from all external and dissipative forces and torques, including those arising from potential energy and from the elevator and ailerons deflections [21].
Let ω denote the vector of angular velocity of the UAV’s wing part (of the fuselage), to which the body-fixed frame { O 3 } is attached. Then the kinetic energy K 0 of the wing part (or, in other words, of the UAV’s fuselage) is equal to
K 0 = 1 2 ω T I 0 ω , (19)
where
J 0 =       I 0 x x             I 0 x y             I 0 x z   I 0 x y                   I 0 y y             I 0 y z   I 0 x z             I 0 y z                   I 0 z z (20)
is the inertia tensor of the fuselage.
The vector ω in (19) can be parameterized by generalized coordinates q 1 , q 2 , and q 3 (by yaw, roll, and pitch angles) as [22]
ω = ω q 0 , d q 0 d t = S q 0 d q 0 d t , (21)
where q 0 = q 1 , q 2 , q 3 T , and
S ( q 0 ) = S ( q 1 , q 2 , q 3 ) =   cos ( q 2 ) sin ( q 3 )             cos ( q 3 )               0           sin ( q 2 )                                       0                         1 cos ( q 2 ) cos ( q 3 )                   sin ( q 3 )               0 . (22)
Then, the kinetic energy K 0 (19) can be rewritten in the form
K 0 = 1 2 d q 0 T d t S T q 0 I 0 S q 0 d q 0 d t . (23)
Note that the generalized coordinates q 4 ,
q 5 , and q 6 in the kinematic model shown in Figure 4 are associated with three rigid bodies. Each of these bodies (two links of the robotic arm and the quad part) has one degree of freedom relative to the bodies, to which they are attached. This allows the kinetic energies K 4 , K 5 , and K 6 of these bodies to be represented in a unified form, using the homogeneous transformation matrices, as
K 4 = 1 2 j = 1 4 k = 1 4 t r T 4 q j J 4 T 4 T q k d q j d t d q k d t , (24)
K 5 = 1 2 j = 1 5 k = 1 5 t r T 5 q j J 5 T 5 T q k d q j d t d q k d t , (25)
K 6 = 1 2 j = 1 3 k = 1 3 t r T 6 q j J 6 T 6 T q k d q j d t d q k d t + j = 1 3 t r T 6 q j J 6 T 6 T q 6 d q j d t d q 6 d t +                                                           1 2 t r T 6 T q 6 J 6 T 6 q 6 d q 6 d t 2 . (26)
where
J i = I i x x + I i y y + I i z z / 2                               I i x y                                                             I i x z                                   m i x ¯ i                 I i x y                                           I i x x I i y y + I i z z / 2                             I i y z                                     m i y ¯ i                 I i x z                                                               I i y z                                 I i x x + I i y y I i z z / 2           m i z ¯ i               m i x ¯ i                                                             m i y ¯ i                                                     m i z ¯ i                                 m i       (27)
i = 4 , 5 , 6 , are the pseudo inertia matrices, and m i x ¯ i , m i y ¯ i ,   m i z ¯ i are first inertia moments with respect to the corresponding rotation axis [20,21].
It should be noted that the transformation matrix T 6 in (26) does not depend on the generalized coordinates q 4 and q 5 .
To make the following expressions more compact, let us introduce special notations for the traces of transformation matrices:
a j k ( 4 ) ( q ) = t r T 4 q j J 4 T 4 T q k ,                   j , k = 1 , 2 , 3 , 4 , (28)
a j k ( 5 ) ( q ) = t r T 5 q j J 5 T 5 T q k ,                 j , k = 1 , 2 , 3 , 4 , 5 , (29)
a j k ( 6 ) ( q ) = t r T 6 q j J 6 T 6 T q k ,                             j , k = 1 , 2 , 3 , 6. (30)
Then the total kinetic energy of the UAV’s rotational motion can be represented in the form
      K =   K 0   + K 4   + K 5   + K 6   = 1 2 d q 0 T d t S T q 0 I 0 S q 0 d q 0 d t + 1 2 j = 1 4 k = 1 4 a j k ( 4 ) ( q ) d q j d t d q k d t   +   1 2 j = 1 5 k = 1 5 a j k ( 5 ) ( q ) d q j d t d q k d t   + 1 2 j = 1 3 k = 1 3 a j k ( 6 ) ( q ) d q j d t d q k d t + k = 1 3 a k 6 ( 6 ) d q k d t d q 6 d t + 1 2 a 66 ( 6 ) ( q ) d q 6 d t 2 . (31)
The potential energy of the pitch-decoupled UAV is given by the equation
P = m 0 g ¯   r ¯ 0 i = 4 , 5 , 6 m i g ¯     T i r ¯ i i , (32)
where g ¯ = g x     g y     g z     0 is the gravitational vector, m 0 ,       m i   ( i = 4 , 5 , 6 ) are the masses of the fuselage and the corresponding rigid bodies, r ¯ i i are the positions of the CM of the i th body in the associated coordinate frame [21].
Substitution of the equations (31) and (32) into (17) yields the dynamics equations of the pitch-decoupled UAV. For the vector q 0 of the Euler angles q 1 ,
q 2 , and q 3 we have:
S T q 0 I 0 S q 0 d 2 q 0 d t 2 + S T q 0 I 0 d S q 0 d t d q 0 d t + d S T q 0 d t I 0 S q 0 d q 0 d t  
      q 0 d q 0 T d t S T q 0 I 0 S q 0 d q 0 d t     +     Ξ 123 = S T q 0 M + S T ( q o )   0 M A i M E l M 4 M 5   0 , (33)
where
q 0 d q 0 T d t S T q 0 I 0 S q 0 d q 0 d t =   d q 0 T d t [ q 1 S q 0 ] T I 0 S q 0 d q 0 d t d q 0 T d t [ q 2 S q 0 ] T I 0 S q 0 d q 0 d t d q 0 T d t [ q 3 S q 0 ] T I 0 S q 0 d q 0 d t , (34)
q i denotes the partial derivative with respect to q i ; M is the vector of torques (12); M A i and M E l are the torques produced by the ailerons and elevator tilts; M 4 is an internal torque, generated by the first actuator of the robotic arm; Ξ 123 is a three-dimensional vector with the components
Ξ 123 ( p ) = r = 4 5 k = 1 r a k p ( r ) d 2 q k d t 2   +   k = 1 r m = 1 r a k p ( r ) q m d q k d t d q m d t +   2 k = 1 r m = 1 r a k m ( r ) q p d q k d t d q m d t + k = 1 3 a k p ( 6 ) d 2 q k d t 2   + 2 k = 1 3 m = 1 3 a k m ( 6 ) q p d q k d t d q m d t + 2 a p p ( 6 ) d 2 q 6 d t 2 + 2 k = 1 3 a k 6 ( 6 ) q p d q k d t d q 6 d t + m = 1 3 k = 1 3 a k m ( 6 ) q p d q m d t d q k d t + 4 k = 1 3 a k 6 ( 6 ) q p d q k d t d q 6 d t   +   2 a 66 ( 6 ) q p d q 6 d t 2 + k = 4 , 5 , 6 m k g ¯     T k q p r ¯ k k (35)
p = 1 , 2 , 3.
Equation (33) can be written in a more compact and transparent form [22]. Note that the time derivative of the angular velocity vector ω in (21) is
d ω d t = S q 0 d 2 q 0 d t 2 + d S q 0 d t d q 0 d t (36)
Then the equation (33) can be written as
I 0 d ω d t + S T q 0 d S T q 0 d t I 0 S q 0 d q 0 d t   q 0 d q 0 T d t S T q 0 I 0 S q 0 d q 0 d t   +   (37)
          S T q 0 Ξ 123 ( q 0 , q R A ) = M +   0 M A i M E l S T q 0 M 4 M 5   0 , where q R A = q 4 ,   q 5 T , M A i and   M E l denote the moments generated by ailerons and elevator tilts, the superscript T denotes the combination of matrix transposition and inversion, and the transposed inverse of S q 0 is
S T ( q 0 ) =   sin ( q 3 ) cos ( q 2 )                   cos ( q 3 )                   sin ( q 3 ) tan ( q 2 )             0                                     0                                             1         cos ( q 3 ) cos ( q 2 )                     sin ( q 3 )                     cos ( q 3 ) tan ( q 2 ) (38)
Comparing this expression with the Newton-Euler’s equations of rotational motion of a rigid body in the body-fixed frame [22]
I 0 d ω d t + ω × I 0 ω = τ (39)
implies
S T q 0 d S T q 0 d t I 0 S q 0 d q 0 d t       q 0 d q 0 T d t S T q 0 I 0 S q 0 d q 0 d t   = ω × I 0 ω , (40)
and the equation (37) can be written in the following compact form:
I 0 d ω d t + ω × I 0 ω   +   S T q 0   Ξ 123 ( q ) = M +   0 M A i M E l S T q 0 M 4   M 5   0 . (41)
This equation, in which the last term on the left-hand side describes dynamic cross-coupling of the fuselage and the robotic arm, is more convenient for analysis and design of the attitude control system in the takeoff and landing modes.
The dynamics equations for the generalized coordinates q 4 ,
q 5 , and q 6 are
k = 1 4 a k 4 ( 4 ) d 2 q k d t 2 +   2 k = 1 4 m = 1 4 a k 4 ( 4 ) q m d q k d t d q m d t +   k = 1 5 a k 4 ( 5 ) d 2 q k d t 2 + 2 k = 1 5 m = 1 5 a k 4 ( 5 ) q m d q k d t d q m d t +   1 2 m = 1 4 k = 1 4 a k m ( 4 ) q 4 d q m d t d q k d t +   1 2 m = 1 4 k = 1 4 a k 4 ( 4 ) q m d q m d t d q k d t +   1 2 m = 1 5 k = 1 5 a k m ( 5 ) q 4 d q m d t d q k d t +     1 2 m = 1 5 k = 1 5 a k 4 ( 5 ) q m d q m d t d q k d t +   m 4 g ¯ T 4   q 4 r ¯ 4 4 + m 5 g ¯ T 5   q 4 r ¯ 5 5 = M 4 (42)
k = 1 5 a k 5 ( 5 ) d 2 q k d t 2   +   k = 1 5 m = 1 5 a k 5 ( 5 ) q m d q k d t d q m d t +   2 k = 1 5 m = 1 5 a 5 k ( 5 ) q m d q k d t d q m d t +   m 5 g ¯ T 5   q 5 r ¯ 5 5 = M 5 , (43)
where M 5 is internal torque produced by the second actuator of the robotic arm, and
k = 1 3 a k 6 ( 6 ) d 2 q k d t 2   +   2 k = 1 3 m = 1 3 a k 6 ( 6 ) q m d q k d t d q m d t +   2 a 66 ( 6 ) d 2 q 6 d t 2 +   2 k = 1 3 a k 6 ( 6 ) q 6 d q k d t d q 6 d t +   m = 1 3 k = 1 3 a k m ( 6 ) q 6 d q m d t d q k d t + 4 m = 1 3 a 66 ( 6 ) q m d q m d t d q 6 d t +   2 a 66 ( 6 ) q 6 d q 6 d t 2 +   m 6 g ¯ T 6   q 6 r ¯ 6 6 =   cos ( 45 ) ( F 1 + F 3 )       ( F 2 + F 4 )     (44)
Equations (33), (42)-(44) can be written in the matrix form
M ( q ) d 2 q d t 2 + C q , d q d t d q d t + G ( q ) = τ (45)
The entries M i j of the symmetric inertia matrix M ( q ) in (45) are
Mij = (STI0S)ij+ aij(4) + aij(5) + aij(6), i,j ≤ 3,
M i 4 = a i 4 ( 4 ) + a i 4 ( 5 ) ,     M i 5 = a i 5 ( 5 ) ,     M i 6 = a i 6 ( 6 ) ,               i = 1 , 2 , 3 ,
M 44 = a 44 ( 4 ) + a 44 ( 5 ) ,     M 45 = a 45 ( 5 ) ,     M 46 = 0 ,
M 55 = a 55 ( 5 ) ,       M 56 =   0 ,       M 66 = a 66 ( 6 ) .  
The entries of the Coriolis/centrifugal terms (Christoffel symbols) of the matrix C q , d q / d t in (45) are determined in terms of the coefficients M i j (46) as
C i j q , d q d t = 1 2 k = 1 6 M i j q k   + M i k q j + M j k q i   d q k d t                                     i , j = 1 , 2 , 3 , 4 , 5 , 6. (47)
The components G i ( q ) of the gravity vector G ( q ) in (45) are
G i = k = 4 , 5 , 6 m k g ¯     T k ( q ) q i r ¯ k k ,                                   i = 1 , 2 , 3 , (48)
G 4 = m 4 g ¯ T 4   q 4 r ¯ 4 4 + m 5 g ¯ T 5   q 4 r ¯ 6 ,     G 5 = m 5 g ¯ T 5   q 5 r ¯ 5 5 ,     G 6 = m 6 g ¯ T 6   q 6 r ¯ 6 6 .      
Finally, the components of the vector τ in (45) are
τ 1 τ 2 τ 3 = S T ( q o ) M ¯     + S T ( q o )   0 M A i M E l M 4 M 5   0 , (49)
τ 4 =   M 4   ,       τ 5 = M 5 ,         τ 6 =   cos ( 45 ) ( F 1 + F 3 )       ( F 2 + F 4 ) . (50)
The derived differential equations of the rotational motion (33)-(50) were used in the development of the comprehensive (expanded) dynamic model of the pitch-decoupled UAV in the Simulink package environment.

4. Translational Dynamics of the UAV

The main forces acting on the pitch-decoupled UAV during the horizontal (longitudinal) flight are schematically shown in Figure 5, where F 1 ,   F 2 ,   F 3 ,   F 4 ,   F 5   are the thrusts generated be the five rotors, F Σ is the total thrust of the quad part’s four rotors, Lift and Drag are aerodynamic forces.
If all forces are represented with respect to the inertial reference frame, then, according to Newton’s second law, the simplified differential equations describing the translational motion of the UAV are given by the equation
m Σ d 2 η d t 2 = m Σ g Z I + F A + R 0 ( q 0 ) sin ( q 6 ) 0 cos ( q 6 ) F Σ , (51)
where vector η = η x , η y , η z T represents the linear position of the UAV in the inertial reference frame, M Σ is the total mass of the UAV, F Σ is the total thrust of the quad part rotors (15), R 0 ( q 0 ) is the orthogonal rotation matrix (2), g is the gravitational constant,
F A denotes the aerodynamic forces acting on the UAV in the wind frame. For simplicity, sideslip forces are neglected, and the aerodynamic forces are assumed to be transformed to the inertial reference frame.
To calculate the aerodynamic characteristics of the pitch-decoupled UAV, it is necessary to determine its lift force F L , drag force F D , and the aerodynamic pitching moment M p   about the pitch axis, which are given by the following expressions [23,24]:
F L = 1 2   ρ V 2 S C L α ,   F D = 1 2   ρ V 2 S C D α ,   M p = 1 2 c S C m α ,   ( 52 )
where α is an angle of attack, and the lift C L , drag   C D , and pitching moment C M coefficients are equal to
C L = 2 L ρ V 2 S ,     C D = 2 D ρ V 2 S ,   C M = 2 M ρ V 2 c S .   ( 53 )
In the expressions (52), ρ is the air density, S is the wing area, V is the flight velocity relative to the air, c is the mean aerodynamic chord [25]. Besides, the coefficients L ,
D , and M in (53) are actual aerodynamic lift and drag forces, and the pitching moment.
The study of the UAV’s aerodynamic characteristics, including the lift and drag coefficients, as well as the stability properties, was performed using the “XFLR5” aerodynamics simulator [26,27,28], considering external forces applied along the longitudinal and lateral axes of the UAV. Figure 6 illustrates the lift, drag and pitching-moment coefficients obtained for different aerodynamic airflow simulation runs. Figure 7 presents the results of analyzing the aerodynamic stability under external forces. The graphs in Figure 7 indicate that the UAV is aerodynamically stable in both longitudinal and lateral directions. The results of aerodynamic simulations were used in the design of the neural controller described in Section 6.

5. Control System of the Pitch-Decoupled UAV

In this Section, the control of the pitched-decoupled UAV in different flight modes is considered in detail. Unlike the transient modes common to tilt-rotors UAVs, in which the UAV switches between the vertical and horizontal flights by tilting the rotors from 0 to 90 or vice versa, special attention in Subsection 5.2 is devoted to a flight regime, in which the pitch angle γ of the quad part is constant but not equal to the limiting (marginal) values 0 or 90 . Geometrically, this corresponds to a “sloped” (climb or descent) flight of the pitch-decoupled UAV. Schematically, the sloped flight mode is illustrated in Figure 8. In this mode, the pitch angle of the plane is zero, and the given constant (but not 0 or 90 ) tilt of the quad part is maintained by a special feedback loop, the error signal of which is calculated as a difference of the given tilt and the signal of the potentiometer measuring the relative tilt of the quad part.
For any given “intermediate” tilt of the quad part, there are three different situations. If the vertical component F z of the total thrust F Σ is equal to the gravitational constant g = 9.81 (Figure 8a), the UAV is in the horizontal flight with tilted rotors. If the vertical component is larger ( F z > g ), the UAV is in the climb flight (Figure 8b). Finally, if F z < g , the UAV is in the descent flight. Note that in all these situations, the zero pitch angle of the UAV is provided by the control system of the tail rotor.

5.1. Control System in Takeoff, Landing, and Hovering Modes

We specifically discuss the flight of the pitch-decoupled UAV in the standard takeoff, landing, and hover modes because in these modes the pitch-decoupled UAV is controlled like a conventional quadcopter. Moreover, the robotic arm is usually used when the UAV is in hover mode. In the present configuration, in these modes, the quad pitch angle γ (also referred to as Q p i t c h angle) is zero and the quad part is rigidly fixed to the UAV’s fuselage.
To proceed, we shall make some assumptions and simplifications needed for the design of the corresponding controllers in the mentioned modes. First, we assume that the mass and inertia matrix I 6 of the quad part are accounted for in the mass m 0 and in the inertia matrix I 0 of the fuselage, where the matrix I 0 is assumed to be diagonal. This simplifies the differential equations (32), (42), (43), in which all the terms containing index 6 must be eliminated. In addition, the equation (44) must be discarded.
Generally, in controlling the flight of quadcopters, the altitude z and rotation angles (roll ϕ , pitch θ , and yaw ψ ) are usually chosen as four control variables [29].
Below, we shall assume that the angles and angular velocities of the UAV are so small that nonlinear terms in the dynamic equations of rotational motions can be neglected, the cosines of all angles are approximately equal to one, and sines of the angles are zero. At this stage, we assume that there are no dynamic cross-couplings between the fuselage of the UAV and the robotic arm.
On these assumptions, the dynamics equations of multirotor UAVs have the following linearized form [3,29]:
d 2 η z d t 2 = 1 m Σ F Σ g , (54)
I 0 d 2 q 0 d t 2 = τ , I 4 d 2 q 4 d t 2 = M 4 , I 5 d 2 q 5 d t 2 = M 5 , (55)
where m Σ is the mass of the UAV; g – the gravitational constant; I 0 is the diagonal tensor of inertia with the components I x ,   I y ,   I z on the principal diagonal; vector τ =
[ τ x , τ y , τ z ] T combines the principal non-conservative forces and moments applied to the UAV airframe by the aerodynamics of the four quad part rotors (assuming no external disturbances); I 4 and I 5 are inertia moments of the robotic arm links.
Denoting by F ¯ the four-dimensional vector of thrusts F i ( F ¯ = [ F 1 , F 2 , F 3 , F 4 ] T ), the mapping of F ¯ to the vector F Σ , τ T can be written in matrix form
F Σ   τ = D M F ¯ . (56)
In (56), the 4 × 4 full-rank numerical matrix D M for the quad part rotors allocation in Figure 4 is equal to
D M =         1                             1                           1                           1     2 2 L               2 2 L           2 2 L         2 2 L 2 2 L               2 2 L               2 2 L         2 2 L   k ψ                     k ψ                   k ψ                     k ψ , (57)
where L is the distance of the rotors from the point Q 6 in Figure 4. Structurally, the matrix D M (59) describes the kinematic cross-coupling between the separate channels of the UAVs in hover mode.
The matrix block diagram of the UAV’s control system in hover mode and of the two-dimensional control system of the robotic arm is shown in Figure 9.
In Figure 9: I is an 3 × 3 identity matrix; I 2 × 2 is an 2 × 2 identity matrix q R A is a two-dimensional vector with components q 4 and q 5 ; P is a 2 × 3 projection matrix
P = 1       0         0 0       1         0 ; (58)
w M ( s ) is the transfer function of brushless DC motors (rotors) of the quad part; d i a g { w i R ( s ) } is a 4 × 4 diagonal matrix of the transfer functions w i R ( s ) ;
K Re g ( s )   = K D d i a g { w i R ( s ) } (59)
is the transfer matrix of the matrix regulator with the numerical matrix K D equal to
K D = D M 1 =   1 4               2 4 L               2 4 L                   1 4 k ψ   1 4               2 4 L                     2 4 L                     1 4 k ψ   1 4         2 4 L                     2 4 L                 1 4 k ψ   1 4             2 4 L                 2 4 L                     1 4 k ψ . (60)
The superscripts Ref in Figure 9 denote the reference (input) signals.
The transfer function w R A ( s ) in Figure 9 represents the identical transfer functions of separate channels of the robotic arm and has the form
w R A ( s ) = w R C ( s ) w R M ( s ) , (61)
where w R M ( s ) is the transfer function of the link actuators (DC motors), and w R C ( s ) is the transfer function of the controllers.
Assuming for simplicity that the transfer functions w i R ( i = 1 , 2 , 3 , 4 ) in (59) are identical, that is w i R ( s ) = w R ( s ) , the transfer matrix W O ( s ) of the open-loop UAV’s control system in hover mode takes the form
W O ( s ) = w M ( s ) w R ( s ) s 2 M Σ 1 D M K D , (62)
where the matrix M Σ has the following block-diagonal form:
M Σ = m Σ               0 1 × 3 0 3 × 1             I 0 . (63)
The dynamics of the DC motors in (61) and (62) are approximated by the first-order transfer functions
w M ( s ) = K M T M s + 1 , w R M ( s ) = K R M T R M s + 1 , (64)
Note that if the matrix K D in the regulator K Re g ( s ) (59) is chosen as in (60), then the transfer matrix of the open-loop system W O ( s ) reduces to the diagonal form
W O ( s ) = w M ( s ) w R ( s ) s 2 M Σ 1 =     1 m Σ             0 1 × 3 0 3 × 1             I 0 1 w M ( s ) w R ( s ) s 2 (65)
This equation implies that the attitude control channels of the UAV in hover mode are decoupled (since matrix I 0 is assumed diagonal) and the design of controllers in these channels can be performed as for common single-input single-output (SISO) control systems [30]. The PD controllers with the following transfer function:
w R ( s ) = 0.237   +   1.32 s 0.0026 s + 1 (66)
were used as controllers w R ( s ) in (65). This transfer function was obtained by applying the graphical user interface (GUI) pidTuner in MATLAB [30]. Analogously, the PD controllers of the form:
w R C ( s ) = 87   +   158 s 0.0011 s + 1 (67)
were used as the controllers of the robotic arm.
The step responses of the attitude control system of the pitched-decoupled UAV with the robotic arm in hover mode are shown in Figure 10. The control of the UAV altitude is considered in the next section.

5.2. PID Control of the Pitch-Decoupled UAV in Transition and Sloped Modes

In this section, we concentrate mainly on the dynamics of the pitched-decoupled UAV in transition mode using standard PID controllers. A control architecture is proposed, in which, to improve the controllability of the UAV, additional velocity-control loops for the three translational axes are introduced. In this system, the control of the UAV in the transition mode essentially reduces to the velocity control of the UAV’s fuselage. The selection of parameters of PID controllers was performed by automatic design tools available in Control System Toolbox in MATLAB [30]. Dynamics analysis of the control system was performed in the Simulink environment, based on the developed nonlinear mathematical model of the UAV, in which, also, the aerodynamic forces (lift and drag) were accounted for. The matrix block diagram of control system of the pitched-decoupled UAV in the transition mode is shown in Figure 11. It should be noted that, virtually, the robotic arm is not (or, practically, cannot be) used in the transition mode. For that reason, the control system of the robotic arm is not shown in the block diagram in Figure 11. Moreover, the aerodynamic (lift and drag) forces, the influence of which is of a complex nonlinear nature, are also not presented in the block diagram. In Figure 11: v Ref and v are the inputs and outputs of the velocity control system; w v ( s ) is the diagonal matrix of the following PID controllers:
w y a w ( s ) = 3 + 1 s + 0.03 s , w p i t c h ( s ) = 1 + 0.05 s + 0.1 s , w r o l l ( s ) = 3 + 0.5 s + 0.03 s ; (68)
Γ is the matrix of the form
Γ = 1 m Σ R 0 ( q 0 ) sin ( q 6 ) 0 cos ( q 6 ) ; (69)
and
w Q ( s ) = 2 + 0.5 s + 0.02 s (70)
is the PID controller of the passive pitch angle γ = q 6 ; I Q is the inertia moment of the quad part about the rotation axis.
Let us consider the transition phase of the pitch-decoupled UAV with the PD controllers w Q s (70) for two scenarios. Figure 12 shows the results of 3D simulation in case when the UAV switches from hover mode to longitudinal flight by tilting the quad part by γ = 90 . The dynamic model of the UAV was developed based on nonlinear equations ()-() and the results of aerodynamic simulation of the UAV presented in Section 4. In this scenario, the UAV starts the flight at t = 0 and gains the prescribed altitude of 85 m. At t = 50   s , the UAV switches to longitudinal flight. As can be seen from Figure 12, during the transition phase, the UAV’s altitude drops to 16 m (at t = 54   s ). Afterwards, due to aerodynamic characteristics of the UAV, it starts gaining velocity and altitude. The graphs in Figure 13 show the Q p i t c h and altitude response graphs.
In the second scenario, the UAV is in hover mode at the prescribed altitude 5 m, and at t = 8 s, the passive pitch angle changes from γ = 0 to γ = 30 . Figure 14 shows the changes in the Q p i t c h angle and the flight altitude. Figure 15 presents the horizontal velocity, as well as the aerodynamic lift and drag forces of the UAV. Figure 16 shows the total thrust generated by the rotors, which decreases after completing the transition phase due to the increase in the aerodynamic lift force.
The parameters of the PID controllers (68) in the velocity control system are determined based on the linearized model of the UAV and for the fixed value of the angle γ . In practical flights, the required angle γ is calculated in advance and introduced manually by the operator. The required parameters of the velocity PID controllers may change substantially for different angles γ . The most common strategies to overcome that difficulty are based on using adaptive, gain-scheduling, and intelligent control techniques, or their combinations [31,32,33,34], as well as other strategies [29,30].

6. Actor-Critic Method and Neural Controller

One of the most effective approaches for designing intelligent controllers is reinforcement learning because it allows one to develop an optimal control policy through interaction with the environment, particularly for systems that are highly nonlinear, uncertain, or difficult to model accurately [35,36,37]. Among reinforcement learning algorithms, the actor-critic family has become one of the most widely used in robotics, UAVs, autonomous vehicles, and process control because it combines the strengths of policy-based and value-based methods [38,39,40].
This section describes a neural controller designed and trained using the reinforcement-learning actor–critic method. The objective is to develop a controller capable of ensuring a smooth transition of a pitch-decoupled UAV from quadcopter mode to airplane mode and back in various working conditions, while minimizing tracking errors.
Briefly, the training procedure can be described as follows. At each time step t , the main control agent selects an action, A t , receives a reward R t from the environment, and transitions from state S t to state S t + 1 . The ultimate objective of reinforcement learning is to maximize the expected discounted return over time
G t = R t + 1 + γ 2 R t + 2   + k = 0 γ k R t + k + 1 max   , (71)
where γ is the discount factor.
If an agent follows policy π at time t , then π ( a , s ) denotes the probability of selecting action A t = a while the system is in state S t = s . Maximizing the expected return yields the optimal policy for the agent.
The action-value function associated with policy π is denoted by q π and is given by the following expression:
q π s , a = E π G T     S t = s ,   A t = a   ] = E π k = 0 γ k R t + k + 1   |   S t = s ,   A t = a   ( 72 )
The central concept of reinforcement learning is represented by the Bellman equation,
q * s , a = E R t + 1 + γ * max q * s ' , a ' .   73
This equation defines the optimal action-value function that the agent must learn iteratively. In this study, the Deep Deterministic Policy Gradient (DDPG) algorithm was employed [41,42]. DDPG uses an actor–critic architecture in which the critic serves as a value-function approximator, while the actor is trained using feedback from the critic by selecting actions that maximize the estimated critic value. The actor–critic structure used for the pitch-decoupled UAV training is illustrated in Figure 17.
The actor network is denoted by μ s θ μ ) and parameterized by θ μ , whereas the critic network is parameterized by θ q , and its loss function has the following form:
L θ q = E q s t , a t | θ q y t 2 ,   ( 74 )
where q s t , a t | θ q is the output of the critic network, and the target value is defined as:
y t = ( R s t , a t + γ q s t + 1 , μ s t + 1 θ q ) ) .   ( 75 )
The overall policy gradient is given by the following expression:
θ μ = E s t a q s , a θ q | s = s t , a = μ s t θ μ μ s θ μ | s = s t   ( 76 )
Here, a   q s , a   θ q ) and θ μ   μ s θ μ are the gradients of the critic and actor neural networks, respectively. Consequently, the actor is trained based on the action values returned by the critic network. It is important to note that the target value y t employs the same actor-critic networks with the same parameters θ μ a n d θ q , which may cause instability during training. To address this issue, target copies of the actor and critic networks, q ' s , a | θ q ' and μ ' s | θ μ ' , are introduced. Their parameters θ q ' and θ μ ' are updated more slowly than those of the primary networks using a soft update factor w 1 :
θ q ' = w θ q + 1 w θ q '   ( 77 )
θ μ ' = w θ μ + 1 w θ μ '   ( 78 )
As described previously, the neural controller replaces the conventional velocity PID controllers to achieve smooth transition behavior and accurate velocity tracking. The resulting velocity control system architecture, incorporating the neural controller, is shown in Figure 18. Here, the neural controller continuously governs the passive pitch angle γ = q 6 depending on the current state of the control system, required flight regime and flying conditions.
The state vector is defined as follows:
S 10 × 1 =   v x v z v x r v z r v x ˙ v z ˙ A t 1 2 × 1 T ε x ε z   T ,   ( 79 )
where
ε x = v x r v x ,   ε z = v z r v z ,   ( 80 ) are the velocity tracking errors.
The reward function is defined as follows:
R = p 1 t T p 2 ε x 2 p 3 ε z 2 p 4 i A i t 1   2   ,   ( 81 )
where p i ,   i = 1 4 are the weighting coefficients of the reward components, and the term p 4 i A i t 1   2 represents the actuator effort penalty. The training episode is terminated if any of the following conditions is violated:
45 < ε x < 45 7 < ε z < 7 0 < v x r < 25 2 < v z r < 2 .   ( 82 )
The pitch-decoupled UAV model and the training procedure were implemented in the Simulink environment using the Reinforcement Learning Designer Toolbox for agent construction [43]. Figure 19 presents the overall training process, during which the agent progressively improves its control policy by maximizing the expected return over successive episodes.
To evaluate the behavior of the control system incorporating the neural controller, two transition-phase scenarios were investigated. In each scenario, the pitch-decoupled UAV takes off and reaches a specified altitude Z . Then, at t = 5   s , the transition phase is initiated during which the UAV accelerates toward a prescribed horizontal velocity v x . Both scenarios were simulated in the Simulink environment using different values of the reference parameters Z 1 , Z 2 and v x 1 , v x 2 :
Scenario 1: Z 1 =   5   m ,     v x 1 = 25   m / s .
Scenario 2: Z 1 =   10   m ,     v x 1 = 10   m / s .
Figure 20 and Figure 21 present the simulation results for Scenario 1, and Figure 22 and Figure 23 for Scenario 2, where γ denotes the passive pitch angle and Z is the altitude. In both cases, the transition phase proceeds smoothly, without overshoot or steady-state error, confirming the effectiveness of the proposed neural controller.
Figure 21 and Figure 23 show the horizontal and vertical velocity responses (blue lines) of the pitch-decoupled UAV together with their reference signals (black dot lines) for Scenarios 1 and 2, respectively. The results demonstrate that the neural controller ensures smooth transition and stable flight at the prescribed horizontal velocities while maintaining minimal tracking error. The transition phase begins when the reference horizontal velocity v x changes from zero to the prescribed constant value. The primary function of the neural controller is not to directly initiate the transition, but rather to track the reference velocities with minimal error by adjusting the γ angle and rotors’ total thrust.
The required γ ( Q p i t c h ) angle, which is calculated by the neural controller, is shown in Figure 20 and Figure 22. In Scenario 2, the transition phase begins before the pitch-decoupled UAV reaches its reference altitude. Nevertheless, the neural controller successfully manages the transition without producing overshoot or steady-state errors.
The simplified, without the altitude and attitude control subsystems, block diagram of the pitch-decoupled UAV’s control system with the velocity neural controller is shown in Figure 24.
Here, W η ( s ) is a diagonal position controller with, usually, standard SISO PID controllers on the main diagonal. In the block diagram, there are two multivariable feedback loops, and one passively controlled (by difference of the paired thrusts) SISO loop, which governs the transitions between different flight modes during entire flight envelope. The altitude and attitude control subsystems, which operate during the hover and standard vertical takeoff and landing modes, as well as control system of the robotic arm, are discussed in Subsection 5.1.

7. Discussion

The article addresses an important challenge in pitch-decoupled VTOL UAV operation: achieving reliable and smooth transitions between hovering and fixed-wing (horizontal-flight) modes. Its principal contribution is the application of an actor–critic reinforcement-learning velocity controller to a pitch-decoupled UAV architecture that differs substantially from conventional tilt-rotor systems. In addition, the study considers a pitch-decoupled UAV equipped with a two-DOF robotic arm and derives a complete mathematical model of the system using modified Denavit–Hartenberg parameters. The proposed approach can accommodate UAVs equipped with robotic manipulators of various configurations and with an arbitrary number of degrees of freedom.
The present study focuses on investigating the main characteristics of the pitch-decoupled UAV control architecture and on designing a neural controller. Therefore, control of the pitch-decoupled UAV in conventional fixed-wing flight—for example, through the elevator, ailerons, or their combination—is not considered. In addition, the SISO tail-rotor control system is assumed to operate perfectly; that is, it is assumed to maintain the fuselage pitch angle at zero during flight. A comprehensive investigation of the pitch-decoupled UAV’s dynamics that accounts for these factors is beyond the scope of this paper.
Another distinctive feature of the study is that the UAV control architecture is formulated within the multivariable control-system framework. This formulation enables clear representation of the kinematic and dynamic cross-couplings among the subsystems and blocks that constitute the overall control system.
Compared with closely related studies on pitch-decoupled UAV control [10,11,12,13], the principal distinguishing features of the present work are the proposed DDPG-based transition controller and the inclusion of a two-DOF robotic arm in the multibody model.
Nevertheless, the proposed neural controller has several limitations, some of which are inherent to reinforcement-learning-based control. At this stage, robustness, stability, and safety have not been evaluated under variations in mass, robotic-arm configuration and payload, aerodynamic coefficients, wind disturbances, sensor noise, or partial actuator failures. These limitations, together with the need for experimental validation and practical deployment of the neural controller, define important directions for future research.
A further promising direction is the development of a hybrid control architecture for the pitch-decoupled UAV that integrates a learning-based controller with a classical model-based controller. Such architectures may provide an effective means of improving stability, robustness, and practical reliability of the UAV’s control system [33,34].

8. Conclusions

This study develops a modeling and control framework for a pitch-decoupled VTOL UAV with a two-DOF robotic arm. The platform’s passive coupling mechanism offers an attractive alternative to conventional tilt-rotor systems by using differential thrust rather than dedicated rotor-tilting actuators. The article combines kinematic and dynamic modeling, aerodynamic analysis, conventional PID-based control, and a DDPG actor–critic neural controller for transition flight. The simulation results demonstrate that the proposed neural controller provides better performance and tracking accuracy during transition modes than the conventional PID controller by continuously governing the passive pitch angle of the pitch-decoupled VTOL UAV.

Author Contributions

Conceptualization, J.A., N.N., O.G. and H.D.; methodology, J.A., N.N., O.G., Z.K. and H.D.; software, N.N., O.G., A.K., G.K. and K.B.; validation, K.B., V.M., G.K. and Z.K.; formal analysis, K.B., V.M., G.K. H.D. and A.K.; investigation, J.A., N.N., O.G.; data curation, K.B., V.M., Z.K. and G.K..; writing—original draft preparation N.N., O.G and A.K.; writing—review and editing, O.G., N.N and J.A.; visualization, K.B., V.M., G.K. and A.K.; supervision, O.G. and N.N..; project administration, O.G. and N.N. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
VTOL Vertical Takeoff and Landing
UAV Unmanned Aerial Vehicle
CM Center of Mass
PID Proportional-Integral-Derivative

References

  1. Hassanalian, M.; Abdelkefi, A. Classifications, applications, and design challenges of drones: A review. Prog. Aerosp. Sci. 2017, 91. [Google Scholar] [CrossRef]
  2. Huang, Y.; Thomson, S.; Hoffmann, C.; Lan, Y.; Fritz, B. Development and prospect of unmanned aerial vehicle technologies for agricultural production management. Int. J. Agric. Biol. Eng. 2013, 6, 1–10. [Google Scholar] [CrossRef]
  3. Mahony, R.; Kumar, V.; Corke, P. Multirotor Aerial Vehicles: Modeling, Estimation, and Control of Quadrotor. IEEE Robot. Autom. Mag. 2012, 19, 20–32. [Google Scholar] [CrossRef]
  4. Saeed, A.S.; Younes, A.B.; Islam, S.; Dias, J.; Seneviratne, L.; Cai, G. A review on the platform design, dynamic modeling and control of hybrid UAVs. Proc. 2015 Int. Conf. Unmanned Aircr. Syst. (ICUAS) 2015, 2015, 806–815. [Google Scholar] [CrossRef]
  5. Flores, G.; Lozano, R. Transition flight control of the quad-tilting rotor convertible MAV. Proc. 2013 Int. Conf. Unmanned Aircr. Syst. (ICUAS) 2013, 2013, 789–794. [Google Scholar] [CrossRef]
  6. Fu, Y.; Zhao, W.; Liu, L. Safe Reinforcement Learning for Transition Control of Ducted-Fan UAVs. Drones 2023, 7, 332. [Google Scholar] [CrossRef]
  7. Escareño, J.; Stone, R.H.; Sanchez, A.; Lozano, R. Modeling and control strategy for the transition of a convertible tail-sitter UAV. Proc. 2007 Eur. Control Conf. (ECC) 2007, 2007, 3385–3390. [Google Scholar] [CrossRef]
  8. Escareño, J.; Salazar, S.; Lozano, R. Modelling and Control of a Convertible VTOL Aircraft. In Proceedings of the 45th IEEE Conference on Decision and Control, 2006; pp. 69–74. [Google Scholar]
  9. Apkarian, J. Hybrid Multicopter and Fixed Wing Aerial Vehicle. 2018. Available online: https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/9873508.
  10. Apkarian, J. Pitch-decoupled VTOL/FW aircraft: First flights. Proceedings of the 2017 Workshop on Research, Education and Development of Unmanned Aerial Systems (RED-UAS) 2017, 2017, 258–263. [Google Scholar] [CrossRef]
  11. Apkarian, J. Attitude Control of Pitch-Decoupled VTOL Fixed Wing Tiltrotor. 2018. [Google Scholar] [CrossRef]
  12. Patience, C. A.; Nahon, M. Control of a Passively-Coupled Hybrid Aircraft. 2020 International Conference on Unmanned Aircraft Systems (ICUAS), Athens, Greece, 2020; pp. 1620–1627. [Google Scholar] [CrossRef]
  13. Chiappinelli, R.; Cohen, M.; Doff-Sotta, M.; Nahon, M.; Forbes, J. R.; Apkarian, J. Modeling and Control of a Passively-Coupled Tilt-Rotor Vertical Takeoff and Landing Aircraft. 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada, 2019; pp. 4141–4147. [Google Scholar] [CrossRef]
  14. Singh, K. WORKING OF ROBOTIC ARM IN INDUSTRIES. 2021, 38. [Google Scholar] [PubMed]
  15. Craig, J.J. Introduction to Robotics: Mechanics and Control; Prentice Hall, 2005; 408 pp. [Google Scholar]
  16. Niku, S.B. Introduction to Robotics: Analysis, Control, Applications; Whilley, 2019. [Google Scholar]
  17. Mason, M.T. Mechanics of Robotic Manipulation; The MIT Press, 2001. [Google Scholar]
  18. Henderson, D.M. Euler angles, quaternions, and transformation matrices for space shuttle analysis // Nasa Shuttle Program, Mission Planning and AnalysisDivision; July 1977; p. 11. [Google Scholar]
  19. Siciliano, B.; Khatib, O. (Eds.) Springer Handbook of Robotics; Springer, 2008; p. 1611 pp. [Google Scholar]
  20. Paul, R.P. Robot Manipulators: Mathematics, Programming and Control; The MIT Press, 1981; p. 279 pp. [Google Scholar]
  21. Jazar, R.N. Theory of Applied Robotics: Kinematics, Dynamics, and Control; Springer, 2010; p. 905 pp. [Google Scholar]
  22. Bernstein, D.S.; Goel, A.; Kouba, O. Deriving Euler’s Equation for Rigid-Body Rotation via Lagrangian Dynamics with Generalized Coordinates. Mathematics 2023, 11, 2727. [Google Scholar] [CrossRef]
  23. Durham, W. Aircraft Flight Dynamics and Control; Wiley, August 2013; p. 306. [Google Scholar]
  24. Cory, R.; Russ, T. Experiments in Fixed-Wing UAV Perching. AIAA Guidance, Navigation and Control Conference and Exhibit AIAA Guidance, 2008; p. 12. [Google Scholar]
  25. Anderson, J.D. Fundamentals of Aerodynamics, 6th ed.; McGraw-Hill Education: New York, NY, USA, 2017; ISBN 978-1259129919. [Google Scholar]
  26. Sanchez, F. Aerodynamics: Airfoil Analysis Using XFLR5 // Universidad Carlos III de Madrid 2017. 11. [PubMed]
  27. Sharma, S. Analysis of a Tiltrotor Vertical Take-off and Landing Unmanned Aerial Vehicle: CFD Approach. IOP Conf. Ser.: Mater. Sci. Eng 1116.
  28. Sanjiv, P.; Shailendra, R.; Saugat, G.; Kshitiz, S.; Sudip, B. Aerodynamic and StabilityAnalysis of Blended Wing Body Aircraft. Int. J. Mech. Eng. Appl. 2016, no 4, 143–151. [Google Scholar] [CrossRef]
  29. Maaruf, M.; Mahmoud, M.S.; Ma’arif, A. A Survey of Control Methods for Quadrotor UAV. Int. J. Robot. Control Syst. 2022, Vol. 2(No. 4), 652–665. Available online: https://pubs2.ascee.org/index.php/ijrcs. [CrossRef]
  30. Chong, Kiam A.; Yun, G. Li. PID Control System Analysis, Design, and Technology // Control Systems Technology. IEEE Transactions 2005, vol. 13, 559–576. [Google Scholar] [CrossRef]
  31. Theys, B.; Vos, G.D.; Schutter, J.D. A control approach for transitioning VTOL UAVs with continuously varying transition angle and controlled by differential thrust. Proc. 2016 Int. Conf. Unmanned Aircr. Syst. (ICUAS) 2016, 2016, 118–125. [Google Scholar] [CrossRef]
  32. Hwangbo, J.; Sa, I.; Siegwart, R.; Hutter, M. Control of a Quadrotor With Reinforcement Learning. IEEE Robot. Autom. Lett. 2017, PP, 1–1. [Google Scholar] [CrossRef]
  33. Zulu, A.; John, S. A Review of Control Algorithms for Autonomous Quadrotors. (widely cited review). 2020, arXiv:1602.02622. [Google Scholar] [CrossRef]
  34. Van An, Vo; Mien, T. L.; Binh, N. V. A Comprehensive Survey of UAV Control Algorithms: Integrating Classical Methods with Artificial Intelligence for Enhanced Trajectory Tracking Includes an ANN–PID hybrid controller. AIMS Electron. Electr. Eng. 2026, 10(1), 26–53. Available online: http://aimspress.com. [CrossRef]
  35. Puterman, M. L. Markov decision processes: discrete stochastic dynamic programming; John Wiley & Sons: Hoboken, New Jersey, 2014; p. 684. [Google Scholar]
  36. Muddasar, N.; Syed, R.; Antonio, C. A Gentle Introduction to Reinforcement Learning and its Application in Different Fields. IEEE Access 2020, vol. 8, 25. [Google Scholar] [CrossRef]
  37. Hafner, R.; Riedmiller, M. Reinforcement learning in feedback control. Mach. Learn. 2011, vol. 84(1- 2), 137–169. [Google Scholar] [CrossRef]
  38. Degris, T.; White, M.; Sutton, R. S. Linear off-policy actor-critic. In 29th International Conference on Machine Learning. 2011. (ICML’12).; Omnipress: Madison, WI, USA; pp. 179–186.
  39. Zhang, Y.; Yu, Z. Path Following Control for UAV Using Deep Reinforcement Learning Approach. Guid. Navig. Control. 2021, 01(No. 01), 18. [Google Scholar] [CrossRef]
  40. Hu, Z.; Gau, X.; Wa, K.; Zhai, Y.; Wang, Q. Relevant experience learning: A deep reinforcement learning method for UAV autonomous motion planning in complex unknown environments. Chin. J. Aeronaut. 2021, Volume 34(Issue 12), 187–204. [Google Scholar] [CrossRef]
  41. Silver, D.; Lever, G.; Heess, N.; Degris, T.; Wierstra; D. Riedmiller, M. Deterministic Policy Gradient Algorithms. Proceedings of the 31st International Conference on Machine Learninng Proceedings of Machine Learning Research 2014, vol. 32(no. 1), 387–395. [Google Scholar]
  42. Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.M.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous control with deep reinforcement learning; abs/1509.02971; CoRR, 2015; p. 14. [Google Scholar]
  43. Abdallah, M.; Lashin, M.; El-Awamry, A.; Gabr, W.I. Reinforcement Learning-Based Optimal Path Planning for Mobile Robot with Obstacles Avoidance. In Proceedings of the 2024 12th International Japan-Africa Conference on Electronics, Communications, and Computations (JAC-ECC), 2024; pp. 1–6. [Google Scholar] [CrossRef]
Figure 1. Different types of VTOL UAVs: (a) hybrid, (b) tilt-rotor, (c) tilt-wing.
Figure 1. Different types of VTOL UAVs: (a) hybrid, (b) tilt-rotor, (c) tilt-wing.
Preprints 227767 g001
Figure 2. Pitch-decoupled VTOL UAV with a robotic arm.
Figure 2. Pitch-decoupled VTOL UAV with a robotic arm.
Preprints 227767 g002
Figure 3. Mechanical design of a pitch-decoupled VTOL UAV: (a) takeoff and landing modes, (b) horizontal-flight mode.
Figure 3. Mechanical design of a pitch-decoupled VTOL UAV: (a) takeoff and landing modes, (b) horizontal-flight mode.
Preprints 227767 g003
Figure 4. Kinematic model of the pitch-decoupled VTOL UAV with an attached robotic arm.
Figure 4. Kinematic model of the pitch-decoupled VTOL UAV with an attached robotic arm.
Preprints 227767 g004
Figure 5. Schematic block diagram of kinematic links of the pitch-decoupled UAV.
Figure 5. Schematic block diagram of kinematic links of the pitch-decoupled UAV.
Preprints 227767 g005
Figure 5. Forces acting on the pitch-decoupled UAV during horizontal flight ( q 6 = 90 ) .
Figure 5. Forces acting on the pitch-decoupled UAV during horizontal flight ( q 6 = 90 ) .
Preprints 227767 g006
Figure 6. Aerodynamic coefficients of the pitch-decoupled UAV obtained for different simulation runs; (a) C L over α , (b) C D versus α ; (d) C M versus α ; (d) C L / C D versus α .
Figure 6. Aerodynamic coefficients of the pitch-decoupled UAV obtained for different simulation runs; (a) C L over α , (b) C D versus α ; (d) C M versus α ; (d) C L / C D versus α .
Preprints 227767 g007
Figure 7. Aerodynamic stability analysis of the UAV: (a) lateral axis, (b) longitudinal axis.
Figure 7. Aerodynamic stability analysis of the UAV: (a) lateral axis, (b) longitudinal axis.
Preprints 227767 g008
Figure 8. Illustration of the “sloped” mode: (a) horizontal flight, (b) climb flight.
Figure 8. Illustration of the “sloped” mode: (a) horizontal flight, (b) climb flight.
Preprints 227767 g009
Figure 9. Matrix block diagram of the control system of the pitch-decoupled UAV with the robotic arm in hover mode.
Figure 9. Matrix block diagram of the control system of the pitch-decoupled UAV with the robotic arm in hover mode.
Preprints 227767 g010
Figure 10. Step responses of the pitch-decoupled UAV with the robotic arm: (a) roll, pitch, and yaw axes; (b) rotation angles of the robot’s links.
Figure 10. Step responses of the pitch-decoupled UAV with the robotic arm: (a) roll, pitch, and yaw axes; (b) rotation angles of the robot’s links.
Preprints 227767 g011
Figure 11. Matrix block diagram of the control system of the pitch-decoupled UAV in the transition modes.
Figure 11. Matrix block diagram of the control system of the pitch-decoupled UAV in the transition modes.
Preprints 227767 g012
Figure 12. Transition phase simulation ( γ = 90 ): 3D simulation.
Figure 12. Transition phase simulation ( γ = 90 ): 3D simulation.
Preprints 227767 g013
Figure 13. Transition phase simulation ( γ = 90 ): Q p i t c h and altitude response graphs.
Figure 13. Transition phase simulation ( γ = 90 ): Q p i t c h and altitude response graphs.
Preprints 227767 g014
Figure 14. Transition phase simulation ( γ = 30 ): Q p i t c h and altitude response graphs.
Figure 14. Transition phase simulation ( γ = 30 ): Q p i t c h and altitude response graphs.
Preprints 227767 g015
Figure 15. Transition phase simulation ( γ = 30 ): horizontal velocity and aerodynamic forces.
Figure 15. Transition phase simulation ( γ = 30 ): horizontal velocity and aerodynamic forces.
Preprints 227767 g016
Figure 16. Transition phase simulation ( γ = 30 ): total thrust change.
Figure 16. Transition phase simulation ( γ = 30 ): total thrust change.
Preprints 227767 g017
Figure 17. Actor-critic diagram implemented for the pitched-decoupled UAV training.
Figure 17. Actor-critic diagram implemented for the pitched-decoupled UAV training.
Preprints 227767 g018
Figure 18. Block diagram of the VTOL UAV control system with the neural controller.
Figure 18. Block diagram of the VTOL UAV control system with the neural controller.
Preprints 227767 g019
Figure 19. Training process over episodes.
Figure 19. Training process over episodes.
Preprints 227767 g020
Figure 20. Transition-phase simulation for Scenario 1: Q p i t c h and altitude response graphs.
Figure 20. Transition-phase simulation for Scenario 1: Q p i t c h and altitude response graphs.
Preprints 227767 g021
Figure 21. Transition phase simulation for Scenario 1: horizontal and vertical velocities.
Figure 21. Transition phase simulation for Scenario 1: horizontal and vertical velocities.
Preprints 227767 g022
Figure 22. Transition phase simulation for Scenario 2: Q p i t c h and altitude response graphs.
Figure 22. Transition phase simulation for Scenario 2: Q p i t c h and altitude response graphs.
Preprints 227767 g023
Figure 23. Transition phase simulation for Scenario 2: horizontal and vertical velocities.
Figure 23. Transition phase simulation for Scenario 2: horizontal and vertical velocities.
Preprints 227767 g024
Figure 24. Simplified matrix block diagram of the control system of the pitch-decoupled UAV.
Figure 24. Simplified matrix block diagram of the control system of the pitch-decoupled UAV.
Preprints 227767 g025
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings