7 · control theory
inputs, nonholonomic constraints, optimality, and reduction
Control theory enlarges the geometric framework of classical mechanics by admitting external inputs and non-integrable constraints. The same Lie-theoretic and symplectic tools developed earlier now decide controllability, characterize optimal trajectories, and reduce the dynamics of constrained multibody and continuum systems to intrinsic equations on reduced spaces.
controllability and the caratheodory-chow theorem
A control-affine system on a manifold $M$ is given by $$ \dot x=f_0(x)+\sum_{i=1}^m u_i f_i(x),\qquad u\in U\subset\mathbb{R}^m. $$ The accessible set from a point $x_0$ is the set of all points reachable in finite time by admissible controls. The system is accessible if the accessible set has non-empty interior and controllable if every point can be reached from every other point.
The Lie algebra $\operatorname{Lie}\{f_0,f_1,\dots,f_m\}$ generated by the drift and control vector fields evaluates to a distribution $\Delta_{\mathrm{Lie}}(x)\subset T_x M$. The Caratheodory-Chow theorem asserts that if $\Delta_{\mathrm{Lie}}(x)=T_x M$ at every $x$ (the Lie-algebra rank condition), then the system is locally accessible; when the manifold is connected and the system is symmetric (or the drift can be neutralized), local accessibility upgrades to global controllability. The theorem converts an analytic question about reachable sets into a purely algebraic computation with Lie brackets.
definition (control-affine system). smooth vector fields $f_0,f_1,\dots,f_m$ on $M$ and a control set $U\subset\mathbb{R}^m$ define the dynamics $\dot x=f_0(x)+\sum u_i f_i(x)$ for measurable $u:[0,T]\to U$.
definition (reachable / accessible set). a point $y$ is reachable from $x_0$ in time $T$ if there exists an admissible control such that the corresponding trajectory satisfies $x(0)=x_0$ and $x(T)=y$. the accessible set $\mathcal{A}(x_0)$ is the union over $T\ge 0$ of all such points. the system is locally accessible at $x_0$ if $\mathcal{A}(x_0)$ contains a nonempty open set of every neighborhood of $x_0$ (in the relative topology of the orbit); it is controllable on a connected $M$ if $\mathcal{A}(x)=M$ for every $x$.
definition (lie algebra rank condition, LARC). let $\mathcal{L}=\operatorname{Lie}\{f_0,\dots,f_m\}$ be the smallest Lie algebra of vector fields containing the $f_i$. the system satisfies LARC at $x$ if $\{X(x):X\in\mathcal{L}\}=T_x M$.
theorem (orbit theorem / rashevsky-chow for driftless systems). consider the driftless system $\dot x=\sum_{i=1}^m u_i f_i(x)$ with $U=\mathbb{R}^m$ (or any open neighborhood of $0$). if LARC holds at every point of a connected manifold $M$, then the system is controllable: any two points can be joined by a piecewise smooth trajectory of the control system.
proof. the orbit $\mathcal{O}(x_0)$ of $x_0$ under the group of diffeomorphisms generated by the flows of $\pm f_1,\dots,\pm f_m$ (and hence by their brackets, via the commutator formula of chapter 3) is an immersed submanifold whose tangent space at each point is exactly $\Delta_{\mathrm{Lie}}(x)$ (Stefan-Sussmann / orbit theorem). LARC forces $\dim\mathcal{O}(x_0)=\dim M$. the orbit is therefore open. it is also closed: if $y_k\in\mathcal{O}(x_0)$ converges to $y$, then near $y$ the same full-rank distribution implies that a neighborhood of $y$ lies in a single orbit, so $y\in\mathcal{O}(x_0)$. a nonempty open-and-closed subset of a connected manifold is all of $M$, so $\mathcal{O}(x_0)=M$. every point of the orbit is reachable by a finite composition of integral curves of $\pm f_i$, which are admissible control arcs with bang controls $u=\pm e_i$. $\square$
theorem (local accessibility under LARC, with drift). if LARC holds at $x_0$ for the control-affine family $\{f_0,\dots,f_m\}$, then the system is locally accessible at $x_0$.
proof. the accessible set from $x_0$ for small times contains the points obtained by concatenating short arcs of the fields $f_0\pm\varepsilon f_i$. the orbit theorem applied to the family of all brackets involving $f_0$ and the $f_i$ shows that the tangent cone to the reachable set spans $T_{x_0}M$. a cone with nonempty interior in $T_{x_0}M$ implies, by the flow-box and continuous dependence, that the reachable set has nonempty interior in $M$ near $x_0$. $\square$
remark (global controllability with drift). LARC alone does not imply global controllability when a drift is present (e.g. $\dot x=1$ on $\mathbb{R}$ with no control). symmetry (invariance under $u\mapsto -u$ after feedback) or the ability to cancel $f_0$ upgrades local accessibility to global controllability on connected manifolds.
singular trajectories, regularization and boundary conditions
Extremals that fail to satisfy the strict Legendre-Clebsch condition are called singular. Their control is determined by the vanishing of the switching function and all its time derivatives along the trajectory. Regularization embeds the singular locus into a higher-dimensional smooth manifold on which the extremal becomes an ordinary integral curve of a desingularized vector field.
For linear boundary-value problems the self-adjointness of a differential operator is equivalent to the Lagrangian character of the boundary conditions inside a symplectic vector space of dimension twice the order of the equation. The set of all Lagrangian subspaces is the Lagrangian Grassmannian; a continuous path of self-adjoint boundary conditions corresponds to a curve in this Grassmannian, and the Maslov index of the path counts the conjugate points. Thus symplectic linear algebra furnishes a complete topological classification of self-adjoint realizations.
definition (switching function, single input). along an extremal $(x,\lambda)$ of a control-affine problem with Pontryagin Hamiltonian $H=\langle\lambda,f_0\rangle+u\langle\lambda,f_1\rangle-L$, the switching function is $\sigma(t)=\langle\lambda(t),f_1(x(t))\rangle$ (or its analog if $L$ depends on $u$). on intervals where $\sigma$ changes sign one has bang-bang control; if $\sigma\equiv 0$ on a time interval the control is singular there.
definition (legendre-clebsch condition). for a smooth Hamiltonian maximized in the control, the strict Legendre-Clebsch condition is $\partial^2 H/\partial u^2<0$ along the extremal (maximality of a regular interior control). failure of the inequality forces the singular case and a higher-order test.
definition (lagrangian subspace). a subspace $L\subset(V,\omega)$ of a symplectic vector space of dimension $2n$ is Lagrangian if $L=L^\omega$ (equivalently $\dim L=n$ and $\omega|_{L}=0$).
theorem (self-adjoint realizations are lagrangian). let $A$ be a formally self-adjoint ordinary differential operator of order $n$ on an interval, with domain determined by linear boundary conditions $\gamma(y)=0$ at the endpoints, where $\gamma$ takes values in $\mathbb{C}^{2n}$ after the standard trace map of jets of order $n-1$ at the two ends. identify the trace space with a symplectic vector space $(V,\omega)$ of dimension $2n$ whose form is the boundary pairing coming from Green's formula. then the boundary conditions define a self-adjoint realization if and only if the graph of admissible traces is a Lagrangian subspace of $V$.
proof. Green's formula for a formally self-adjoint operator reads $$ \langle Ay,z\rangle_{L^2}-\langle y,Az\rangle_{L^2} =\omega\bigl(\gamma(y),\gamma(z)\bigr) $$ for the canonical boundary symplectic form $\omega$. the maximal operator acts on all sufficiently regular $y$ with free traces; a restriction of the domain by $\gamma(y)\in L$ for a fixed subspace $L\subset V$ makes the boundary form vanish identically on the domain if and only if $\omega|_{L}=0$. maximality of a self-adjoint extension among symmetric extensions is equivalent to $\dim L=n$, i.e. $L$ Lagrangian (otherwise one can enlarge $L$ while keeping $\omega|_L=0$). $\square$
remark (maslov). a continuous family of self-adjoint boundary conditions is a path in the Lagrangian Grassmannian $\Lambda(n)$. conjugate points of a linear Hamiltonian boundary-value problem are the instants at which the evolving Lagrangian plane intersects a reference plane nontrivially; their algebraic count is the Maslov index, already met as metaplectic phase in chapter 6.
optimal control and the pontryagin maximum principle
An optimal-control problem seeks to minimize a cost $$ \int_0^T L(x,u)\,dt+K(x(T)) $$ subject to the controlled dynamics and endpoint constraints. The Pontryagin maximum principle is obtained by regarding the problem as a constrained variational problem on the space of state-input curves and applying the Lagrange-multiplier rule in the Banach space of trajectories. The resulting necessary conditions assert the existence of an adjoint covector $\lambda(t)$ such that the Pontryagin Hamiltonian $$ H(x,\lambda,u)=\langle\lambda,f(x,u)\rangle-L(x,u) $$ is maximized pointwise with respect to $u$, and the pair $(x,\lambda)$ satisfies Hamilton's equations on $T^*M$. When the maximum is attained in the interior of the control set one recovers the classical Euler-Lagrange equations; when it is attained on the boundary the maximization condition supplies the optimal feedback.
definition (bolza problem, free final time or fixed). minimize $J[u]=\int_0^T L(x(t),u(t))\,dt+K(x(T))$ subject to $\dot x=f(x,u)$, $x(0)=x_0$, and possibly $x(T)\in\mathcal{T}$ for a target manifold $\mathcal{T}$.
theorem (pontryagin maximum principle, free endpoint, fixed $T$, compact $U$). let $u_*$ be an optimal control with trajectory $x_*$, and assume $f,L$ are $C^1$ in $(x,u)$ and $U$ is compact. then there exists an absolutely continuous covector $\lambda:[0,T]\to T^*M$ along $x_*$, never zero, such that $$ \dot x_*=\partial_\lambda H,\qquad \dot\lambda=-\partial_x H $$ evaluated at $(x_*,\lambda,u_*)$ (Hamilton equations on $T^*M$), the maximality condition $$ H\bigl(x_*(t),\lambda(t),u_*(t)\bigr) =\max_{v\in U}H\bigl(x_*(t),\lambda(t),v\bigr) $$ holds for almost every $t$, and the transversality condition $\lambda(T)=-dK(x_*(T))$ holds when the terminal cost is free of further constraints.
proof (multiplier form of first variation). work in a single chart for notational simplicity (globalize by partition of unity on time intervals). admissible variations of the control $u_*+\varepsilon v$ produce state variations $\delta x$ solving the linearized dynamics $$ \delta\dot x =D_x f(x_*,u_*)\delta x+D_u f(x_*,u_*)v,\qquad \delta x(0)=0. $$ the first variation of the cost is $$ \delta J =\int_0^T\bigl( L_x\cdot\delta x+L_u\cdot v \bigr)\,dt +K_x\cdot\delta x(T). $$ introduce an adjoint curve $\lambda$ solving the final-value problem $$ \dot\lambda=-\lambda\cdot D_x f+L_x,\qquad \lambda(T)=-K_x(x_*(T)) $$ (row covector notation). integration by parts of $\lambda\cdot\delta\dot x$ against the variational equation cancels the $\delta x$ bulk terms and leaves $$ \delta J =\int_0^T\bigl( L_u-\lambda\cdot D_u f \bigr)\cdot v\,dt. $$ for $u_*$ to be a local minimum against all needle or compactly supported variations $v$ with values keeping the control in $U$, the integrand must not allow decrease, which forces the pointwise maximization of $H=\lambda\cdot f-L$ over $U$ almost everywhere (standard needle variation argument: if some $v_0\in U$ gave a strictly larger $H$ on a set of positive measure, a needle variation toward $v_0$ would decrease $J$). rearranging the adjoint ODE with the same $H$ gives Hamilton's equations. if $\lambda\equiv 0$ then maximality and the ODE force a degenerate abnormal problem; for free endpoints with $K$ present and normal multipliers one normalizes nonzero $\lambda$. $\square$
remark (interior maximum). if for each $t$ the maximum of $H$ over $U$ is attained at an interior point where $\partial_u H=0$ and $\partial_{uu}H<0$, the maximality condition reduces to a feedback $u=u(x,\lambda)$ and, eliminating $u$, one recovers Euler-Lagrange dynamics for the reduced Lagrangian when $f$ is a pure control (chapter 1).
nonholonomic constraints and dirac structures
Nonholonomic constraints (rolling without slipping, skating, etc.) define a non-integrable distribution $D\subset TM$. The appropriate geometric object that unifies the constraint distribution with the symplectic structure of the unconstrained phase space is a Dirac structure: a maximal isotropic subbundle of $TM\oplus T^*M$ with respect to the neutral pairing $$ \langle(X,\alpha),(Y,\beta)\rangle=\alpha(Y)+\beta(X). $$ A Dirac structure simultaneously generalizes graphs of symplectic forms, graphs of Poisson tensors, and graphs of closed two-forms.
The Lagrange-Poincare-Dirac reduction procedure quotients the Dirac structure by a symmetry group that preserves both the Lagrangian and the constraint distribution. The reduced object is again a Dirac structure on a smaller space, and the dynamics become the implicit Euler-Poincare equations $$ \Bigl( \frac{d}{dt}\frac{\delta\ell}{\delta\xi} -\operatorname{ad}^*_\xi\frac{\delta\ell}{\delta\xi}, \frac{\delta\ell}{\delta\xi} \Bigr)\in\mathfrak{d}, $$ where $\mathfrak{d}$ is the reduced Dirac structure. These equations govern the controlled motion of multibody robotic systems with nonholonomic wheels or contact constraints and automatically preserve the energy balance and the constraint forces.
definition (dirac structure). a subbundle $L\subset TM\oplus T^*M$ is a Dirac structure if it is maximally isotropic for the pairing $\langle(X,\alpha),(Y,\beta)\rangle=\alpha(Y)+\beta(X)$, i.e. $L=L^\perp$ with respect to that pairing (fiberwise). smoothness and integrability conditions (Courant involutivity) are often imposed for integrable Dirac structures.
theorem (graph of a symplectic form is dirac). if $(M,\omega)$ is symplectic, then $$ L_\omega =\bigl\{(X,\iota_X\omega):X\in TM\bigr\} $$ is a Dirac structure. similarly the graph of a Poisson bivector defines a Dirac structure.
proof. the pairing of two elements $(X,\iota_X\omega)$ and $(Y,\iota_Y\omega)$ is $$ \iota_X\omega(Y)+\iota_Y\omega(X) =\omega(X,Y)+\omega(Y,X)=0, $$ so $L_\omega$ is isotropic. $\dim L_\omega=\dim M=\frac12\dim(TM\oplus T^*M)$, so maximality follows. $\square$
definition (nonholonomic affine constraint). a distribution $D\subset TM$ of constant rank, not necessarily integrable, restricts admissible velocities by $\dot q\in D_q$. for a mechanical Lagrangian $L$ on $TQ$, the nonholonomic equations are the Lagrange-d'Alembert principle: $\delta\int L=\mathbf{0}$ for variations with $\delta q\in D$ and $\dot q\in D$.
remark (dirac encoding). the pair consisting of the constraint force annihilator and the Legendre-transformed symplectic relation can be assembled as a single Dirac structure whose implicit dynamics reproduce Lagrange-d'Alembert. reduction by a symmetry group free on the constraint distribution produces the reduced Dirac structure $\mathfrak{d}$ and the displayed Euler-Poincare-Dirac equation, which collapses to ordinary Euler-Poincare when $D=TM$.
infinite-dimensional multisymplectic reduction
When the configuration space is itself a space of maps (fluids, elasticity, plasma), the symplectic form is replaced by a multisymplectic form, a closed non-degenerate form of higher degree on the first-jet bundle of the field. The symmetry group is typically a group of diffeomorphisms or gauge transformations. Multisymplectic reduction by this group yields the Euler-Poincare equations on the dual of the Lie algebra of vector fields (or of the gauge algebra).
Special solutions of the reduced equations include peakons (weak solutions of the Camassa-Holm equation obtained by reduction of the EPDiff equation) and vortex filaments (reductions of the incompressible Euler equations supported on curves). The same geometric mechanism that produces the rigid-body Euler equations in finite dimensions therefore produces the fundamental models of continuum mechanics once the appropriate infinite-dimensional symmetry group is reduced.
definition (EPDiff). the Euler-Poincare equation on the diffeomorphism group of a Riemannian manifold for a kinetic-energy Lagrangian on vector fields is called EPDiff. in one space dimension with $H^1$ metric it reduces to the Camassa-Holm equation $m_t+um_x+2u_x m=0$ with $m=u-u_{xx}$.
remark. peakons $u=\sum_i p_i e^{-|x-q_i|}$ arise as weak solutions formally tracking singular momentum supported at points; vortex filaments similarly concentrate vorticity on curves. both are geometric singular solutions of reduced Euler-Poincare dynamics rather than ad hoc PDE constructions.
summary
Collectively, these constructions extend the symplectic and Lie-theoretic framework of classical mechanics to systems with inputs, non-integrable constraints, and infinite-dimensional configuration spaces. Controllability is decided by Lie brackets, optimality by maximization of a Hamiltonian, and reduction by Dirac or multisymplectic quotient, furnishing a unified geometric language for the analysis and design of controlled mechanical systems.
exercises
exercise 1 (heisenberg / nonholonomic planner). on $\mathbb{R}^3$ with controls $f_1=\partial_x+y\partial_z$ and $f_2=\partial_y-x\partial_z$ (and no drift), compute $[f_1,f_2]$ and verify LARC everywhere. conclude controllability by Chow.
exercise 2 (linear time-invariant controllability). for $\dot x=Ax+Bu$ on $\mathbb{R}^n$, show that LARC at $0$ is equivalent to $\mathrm{rank}[B,AB,\dots,A^{n-1}B]=n$ (Kalman rank condition). (Hint: the Lie algebra is spanned by the constant fields $A^k b_j$ after identifying vector fields with their values.)
exercise 3 (bang maximality). for $\dot x=u$ on $\mathbb{R}$ with $u\in[-1,1]$ and cost $\int_0^T x^2\,dt$, write the Pontryagin Hamiltonian and show that an optimal control is bang-bang or singular only where the switching function vanishes identically on an interval. argue that a singular arc would require $x\equiv 0$.
exercise 4 (dirac graph). verify directly that $L_\omega$ for $\omega=dp\wedge dq$ on $\mathbb{R}^2$ is maximal isotropic for the neutral pairing on $T\mathbb{R}^2\oplus T^*\mathbb{R}^2$.
exercise 5 (bang-bang intuition for the double integrator). on $\ddot q=u$ with $|u|\le 1$, write the PMP Hamiltonian and show that an optimal control (minimizing time to the origin) must satisfy $u=-\operatorname{sign}(\lambda_p)$ whenever the switching function $\sigma=\lambda_p$ is nonzero. explain what a singular arc would require of $\sigma$ and why generic time-optimal arcs for this system are bang-bang.