4. Solution of the Optimal Control Problem for a Single Semi-Oscillation
In the previous section, it was demonstrated how to resolve the initial problem (
2) by first solving an auxiliary problem
and find the dependency of the optimal time
T on the terminal value
C.
Here the condition denotes the monotonicity of the trajectory , which corresponds to one semi-oscillation.
First, the question of controllability will be examined, and the range of values for
C for which problem (
5) has a solution will be defined.
The following notations will be introduced
The largest value
can be attained with the control
because with such control, acceleration is maximized when
and deceleration is minimized when
.
Similarly, the smallest value
can be reached analogously with the control
Solving the differential equation with the boundary conditions from system (
5) and with control (
6) or (
7), it is obtained
where
To apply PMP [
1] introduce the notation
and rewrite (
5) in the form of a system of first-order differential equations
Now let the terminal value
C satisfy condition (
8), which ensures the controllability of the system.
Write the Pontryagin function
and denote its upper boundary
If
,
, and
constitute a solution to the optimal control problem (
10), then the following three conditions are satisfied:
I) There exist continuous functions
and
, which never simultaneously become zero and are solutions to the adjoint system.
II) For any
, the maximum condition is satisfied
III) For any
, a specific inequality is occured
From condition (
12) for the maximum of the function
H, the optimal control is obtained in the form
Let us show that the case of singular control in formula (
13), specifically when
over a non-zero length interval of time is impossible, assuming the opposite. This means considering the existence of a time interval during which
. In such an interval, determining the value of optimal control from the maximum condition would not be feasible.
Given the continuity of the functions and , it is possible either for over some interval or for over a certain time period.
If
, then
must also be identically zero. However, this conclusion, derived from the second equation of the adjoint system (
11), implies that
, contradicting the maximum principle’s condition I).
In the scenario where
, it follows that
. Such a case is deemed impossible, as the controlled system cannot stay in a zero state under any control value, given that the term of the system’s differential equation (
5), which includes the control, would also equate to zero.
This reasoning leads to the formulation of a statement:
Lemma 1. Optimal control is limited to only two values, 1 and , dictated by the sign of the product . Considering the case where this product equals zero as non-existent is justified by the fact that the control value at a single point or a finite number of points lacks any impact on the trajectory of the controlled system.
Now, consider condition III. It represents the greatest interest at values and .
At
, the condition is expressed as
At
, the condition becomes
Given the boundary conditions that
, and considering the control value
is always positive, with
and
, the following additional conditions are derived from (
14) and (
15)
Now, exploring the potential form of optimal control and the number of switches. It is already known that the value of optimal control is determined by the sign of the product .
The trajectory , due to its monotonic nature, crosses zero only once. This moment in time is denoted as .
Thus, control may only change its value at the point and at points where the sign of the adjoint variable changes. If at point , both and change their signs simultaneously, then the control value remains unchanged.
Firstly, consider an interval of time where control
. Then, the general solution
of the differential equation from system (4) and
from the adjoint system (
5) will take a specific form
where constants
,
,
,
must be determined from the boundary conditions on the interval of constant control. The value of the adjoint variable
is not of interest, as it does not enter into formula (
13).
Now, consider an interval of time during which control
. Similarly, it is obtained that
where constants
,
,
,
are also to be determined from the boundary conditions.
It’s now proposed that the adjoint variable
turns to zero at most twice within the interval
, either in
. For instance, let
, where
. Then, within the interval
, the control value does not change, and this leads to a contradiction with formulas (
17), (
18) because the distance between zeros of the function
(for example
for formulas (
17)) exceeds the maximum length of an interval of constancy of sign and monotonicity of the function
(for example
or
).
Thus, it is proven that
Lemma 2. In problem 5, optimal control can have no more than one switch in each of the intervals and
The function
has a continuous derivative (as the right-hand side of the second equation of the adjoint system (
11) is continuous) and turns to zero no more than twice within the interval
. Moreover, these zeros cannot both lie within the same subinterval
or
. This leads to 10 different cases (
Figure 2) of sign changes for the function
over the interval
. Dashed gray lines on the graph indicate scenarios that contradict the PMP, while solid red lines indicate cases with no contradiction with PMP found. A detailed analysis of these cases is provided.
If
, then
for
, leading to cases 1) and 2). In case 1), a constant control equal to 1 is maintained throughout the entire time interval. Case 2) is not possible, as
and does not satisfy condition (
16).
If
turns to zero twice within the interval
, there exist
and
such that
and
, leading to cases 3) and 4). These cases contradict condition (
16) since
and
have the same sign.
If turns to zero once at a point and does not equal zero within the interval , cases 5) and 6) are obtained. Case 5) is impossible because .
If turns to zero once at a point and does not equal zero within the interval , cases 7) and 8) emerge. Case 8) is not feasible, as .
Finally, if does not turn to zero within the interval , cases 9) and 10) are considered. Case 9) is possible if . Case 10) is possible if .
After analyzing cases 1)-10), it is determined that the following statement holds
Lemma 3. Optimal control (bang-bang) in the problem (5) can be one of the five types represented in Figure 3.
It is noted that all types of control satisfying the maximum principle (illustrated in
Figure 3) differ in the length of the segment where the control value equals
, and its placement respectively to the point
.
Introducing the parameter
, the values of
and
T can be distinctly determined from the equation and three boundary conditions (excluding the condition
) of problem (
5) by substituting the corresponding control. This results in the determination of the end time
and the terminal value
as functions of the unknown parameter
s.
For control type 3 (illustrated in
Figure 3),
corresponds, and for control type 1, the smallest value
. For control type 5 the largest value
is obtained as the longest possible duration of motion under constant control
, that is
, with the moments of time
and
T derived from formula (
18) and the conditions
,
,
,
, aiming to minimize
. Similarly, from formula (
18), the smallest value of
is obtained. Controls of type 2 and 4 correspond to intermediate values of
s within intervals
and
.
Knowing the switching moment of control and having an analytical solution (formulas
17-
18), the end time
T and the terminal trajectory value
can be explicitly calculated as functions of the parameter
s.
Let us consider
, then for
,
and from formula (
17) and the initial condition
it’s found
,
leading to
and
.
Subsequently, for
,
and from formula (
18) and the continuity of
at
, similarly,
.
Finally, for
,
and from formula (
17) and the continuity of
at
, it’s found
where
.
From formula (
19) and the condition
, the end moment of time
Simplifying expressions (
19) - (
20), ultimately, for
, one obtains
Conducting analogous calculations for the case
, one obtains
Noting that formulas (
21)-(
22) parametrically define a certain curve
depicting the dependency of the end time on the terminal value
C when utilizing controls that satisfy the maximum principle. The parametric formulation of the function allows for the calculation of the first two derivatives of
as functions of the variable
C. Thus, the following properties of the function
are established
Lemma 4.
Uniquely determine the function , defined for .
The function is continuous for .
The function is differentiable for . At the endpoints of the interval, the derivative equals infinity, while at the point corresponding to the parameter , the derivative equals zero. Let be denoted.
The function decreases on the interval and increases on the interval .
The second derivative of the function is negative on the intervals . This condition signifies that the function is concave down for .
Remark 1. Note that the constancy of the sign of the second derivative was established by calculations via symbolic mathematics by Wolfram.
Investigating the properties of the function , it was found that each permissible terminal value C corresponds to a unique control that satisfies the PMP. Therefore the statement is following
Lemma 5. The function , defined by formulas (21)-(22), determines the optimal time in problem (5).
Considering an example with given parameters
,
, it’s calculated that
,
. From (
21), it’s found that
.
Figure 4 would illustrate the graph of the function
, demonstrating how the optimal time varies with different terminal values
C within the specified range.
It has been demonstrated that each value of
s unequivocally corresponds to a specific optimal control and an optimal trajectory, leading to a particular terminal point
. Different optimal trajectories, corresponding to various types of controls, are presented in
Figure 5. Controls of types 1 and 5 correspond to trajectories reaching the extreme points of the reachability set. Control of type 2 corresponds to the upper branch of the
curve (left branch on
Figure 4). Control of type 3, which has no switches, corresponds to the trajectory with the minimum possible time. Control of type 4 corresponds to the lower branch of the
curve (right branch on
Figure 4).
Trajectories are constructed for the given parameter values on the
Figure 4, but the general character of the drawing does not change with different parameter values.