Equality path constraints and the augmented Hamiltonian
Consider the Bolza problem
u(⋅)minJ=Φ(x(tf))+∫t0tfL(x,u,t)dt, subject to
x˙=f(x,u,t) and the mixed equality path constraint
c(x,u,t)=0. (a) Introduce a time-varying multiplier
μ(t) for the path constraint and construct the
augmented Hamiltonian
H=L+λTf+μTc. (b) Derive the state equation, costate equation, and control-stationarity
condition in terms of H.
(c) Show that variation with respect to
μ(t) recovers the path constraint.
(d) Specialize the result to a control-only constraint
c(u,t)=0. (e) Specialize the result to a state-only constraint
c(x,t)=0 and identify the additional term in the costate equation.
Inequality path constraints and complementarity
Suppose a minimization problem contains the scalar inequality path
constraint
c(x,u,t)≤0. Using
H=H+μc, derive the pointwise conditions
c≤0,μ≥0,μc=0. Then:
(a) derive the necessary conditions on an inactive arc;
(b) derive the necessary conditions on an active arc;
(c) explain why the active-set structure is generally unknown in advance;
(d) for the control bound
u(t)−umax≤0, determine the multiplier behavior when u(t)<umax and when
u(t)=umax;
(e) state how the inequality would be imposed in a direct-collocation
nonlinear program.
Order of a state inequality constraint
Consider the double integrator
x˙1=x2,x˙2=u, with the state path constraint
c(x)=x1−xmax≤0. (a) Differentiate c repeatedly along the dynamics until the control
appears explicitly.
(b) Show that the constraint has order two.
(c) On a constrained arc, derive the tangency conditions
x1=xmax,x2=0. (d) Determine the control required to remain on the constrained arc.
(e) State the consistency conditions that must hold at entry into the
constrained arc.
(f) Explain why enforcing x1≤xmax only at collocation nodes
does not guarantee continuous-time feasibility.
(g) Formulate a dense-grid verification test for the reconstructed state
trajectory.
Principle of optimality and the HJB equation
Define the optimal cost-to-go
J∗(x,t)=u(⋅)min[Φ(x(tf))+∫ttfL(x(τ),u(τ),τ)dτ]. Starting from Bellman’s principle of optimality over the short interval
[t,t+Δt]:
(a) write the dynamic-programming recursion;
(b) expand
J∗(x(t+Δt),t+Δt) to first order;
(c) derive the Hamilton–Jacobi–Bellman equation
0=Jt∗+umin[L+∇xJ∗Tf]; (d) derive the terminal condition for J∗;
(e) show that, along the optimal trajectory,
dtdJ∗=−L(x∗,u∗,t); (f) identify the relationship between the value-function gradient and the
costate.
Scalar HJB equation and Riccati feedback
Consider
x˙=−2x+u with
J=21x2(tf)+21∫ttf(x2+u2)dτ. Assume
J∗(x,t)=21P(t)x2. (a) Derive Jt∗ and Jx∗.
(b) Substitute the assumed value function into the HJB equation.
(c) Minimize the HJB expression with respect to u.
(d) Derive the optimal feedback law.
(e) Show that P(t) satisfies
P˙=P2+4P−1. (f) Show that the terminal condition is
P(tf)=1. (g) Derive the optimal closed-loop state equation.
(h) Show that
λ(t)=P(t)x∗(t). Matrix HJB formulation of finite-horizon LQR
Consider
x˙=A(t)x+B(t)u, with performance index
J=21xT(tf)Sfx(tf)+21∫ttf[xTQx+uTRu]dτ, where Q(t)=QT(t)⪰0 and
R(t)=RT(t)≻0.
Assume
J∗(x,t)=21xTS(t)x. (a) Derive
Jt∗=21xTS˙x. (b) Derive the minimizing control.
(c) Derive the Riccati differential equation
−S˙=ATS+SA−SBR−1BTS+Q. (d) Derive the terminal condition
S(tf)=Sf. (e) Derive the feedback gain and the closed-loop state equation.
(f) Show that
λ=Sx∗. (g) Explain why the Riccati equation is integrated backward while the
state equation is integrated forward.
Bang–bang and singular control
Consider
u(⋅)minJ=∫01x2(t)u(t)dt subject to
x˙1=x2,x˙2=−x2+u, 0≤u(t)≤2, and
x1(0)=0,x2(0)=1,x1(1)=1,x2(1)=1. (a) Construct the Hamiltonian.
(b) Derive the costate equations.
(c) Define the switching function
ϕ(t)=x2(t)+λ2(t). (d) Derive the bang–bang minimization rule for ϕ>0 and
ϕ<0.
(e) Impose
ϕ=ϕ˙=0 on a candidate singular arc.
(f) Derive the singular control.
(g) Verify that the singular state and control satisfy all endpoint
conditions.
(h) Determine the order of the singular arc and state the corresponding
generalized Legendre–Clebsch check.
Analytical direct shooting for the minimum-energy double integrator
Consider
u(⋅)minJ=21∫0Tu2(t)dt subject to
x˙1=x2,x˙2=u, with fixed endpoint data
x1(0)=a0,x2(0)=v0,x1(T)=af,x2(T)=vf. (a) Derive the Hamiltonian and costate equations.
(b) Show that λ1 is constant and λ2 is affine in
time.
(c) Derive the optimal control as an affine function of time.
(d) Integrate the state equations analytically.
(e) Reduce the terminal conditions to a 2×2 linear system for
the two unknown initial costates.
(f) Solve this system explicitly.
(g) Explain why the original infinite-dimensional problem has been reduced
to a finite-dimensional shooting problem.
General neighboring-optimal-control derivation
Let
(x∗,u∗,λ∗)
be a nominal nonlinear optimal trajectory. Define
δx=x−x∗,δu=u−u∗,δλ=λ−λ∗. Assume the neighboring dynamics and second-order cost are
δx˙=F(t)δx+G(t)δu, δ2J=21δxT(tf)Pfδx(tf)+21∫t0tf[δxTQδx+2δxTMδu+δuTRδu]dt. (a) Construct the neighboring Hamiltonian.
(b) Derive the neighboring costate equation.
(c) Derive the stationarity condition.
(d) Show that
δu=−R−1(MTδx+GTδλ). (e) Assume
δλ=P(t)δx and derive the generalized Riccati equation.
(f) Show that
P(tf)=Pf. (g) Derive the neighboring feedback gain
K(t)=R−1(MT+GTP). (h) Show that the standard time-varying LQR Riccati equation is recovered
when M=0.
Neighboring feedback for a nonlinear scalar system
Consider the nonlinear dynamics
x˙=x+x2+u and the nominal trajectory
x∗(t)=0,u∗(t)=0. The neighboring cost is
δ2J=21pfδx2(tf)+21∫0tf(qδx2+2mδxδu+rδu2)dt, with r>0.
(a) Linearize the dynamics about the nominal trajectory and identify
F and G.
(b) Derive the scalar neighboring Riccati equation.
(c) Derive the feedback gain
K(t)=rm+GP(t). (d) Write the neighboring correction
δu(t)=−K(t)δx(t). (e) Write the complete implemented control
u=u∗+δu.
(f) Show that the correction is zero on the nominal trajectory.
(g) Describe a numerical test using several initial perturbation
magnitudes.
(h) State the residuals and qualitative behaviors that should be checked
to determine when the neighboring approximation is no longer reliable.