Model Based Control (MBC) is a class of control-system design paradigms in which an explicit mathematical model of the plant’s dynamics — encoding kinematics, inertia tensors, contact forces, aerodynamics, or learned neural representations — is embedded within the controller to predict, plan, and…
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:hasPart rb:ModelPredictiveControl))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:hasPart rb:IterativeLQR))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:hasPart rb:DifferentialDynamicProgramming))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:hasPart rb:WholeBodyController))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:hasPart rb:SystemIdentification))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:hasPart rb:CostFunction))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:hasPart rb:PredictionHorizon))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:hasPart rb:ConstraintSet))
## Dependency Relationships
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:requires rb:DynamicModel))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:requires rb:StateEstimator))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:requires rb:OptimisationSolver))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:requires rb:JacobianComputation))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:dependsOn rb:LagrangianMechanics))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:dependsOn rb:ConvexOptimisation))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:dependsOn rb:ContactMechanics))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:dependsOn rb:NumericalIntegration))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:dependsOn rb:StateEstimation))
## Capability Relationships
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:enables rb:RealTimeMotionPlanning))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:enables rb:ConstraintSatisfaction))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:enables rb:SampleEfficientLearning))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:enables rb:RobustLocomotion))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:enables rb:CompliantManipulation))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:supports rb:LeggedLocomotion))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:supports rb:AerialRobotics))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:supports rb:SurgicalRobotics))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:supports rb:AutonomousDriving))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:supports rb:ProcessControl))
## Implementation Relationships
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:implements rb:RecedinghorizonControl))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:implements rb:QuadraticProgramming))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:implements rb:SequentialQuadraticProgramming))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:implements rb:InteriorPointMethods))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:implements rb:DifferentialDynamicProgramming))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:implements rb:KoopmanOperatorMethods))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:uses rb:MuJoCo))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:uses rb:PinocchioLibrary))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:uses rb:DrakeToolbox))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:uses rb:CasADi))
## Reduction Relationships
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:reduces rb:PhysicalTrialRequirements))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:reduces rb:ConstraintViolationRate))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:reduces rb:SampleComplexity))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:reduces rb:EnergyConsumption))
SubClassOf(rb:ModelBasedControl
ObjectSomeValuesFrom(rb:reduces rb:TrajectoryTrackingError))
## Data Properties
DataPropertyAssertion(rb:hasIdentifier rb:ModelBasedControl "RB-9018"^^xsd:string)
DataPropertyAssertion(rb:authorityScore rb:ModelBasedControl "0.87"^^xsd:decimal)
DataPropertyAssertion(rb:typicalMPCHorizon rb:ModelBasedControl "20"^^xsd:integer)
DataPropertyAssertion(rb:iLQRSolveTimeMs rb:ModelBasedControl "10"^^xsd:decimal)
DataPropertyAssertion(rb:sampleEfficiencyGainFactor rb:ModelBasedControl "50"^^xsd:integer)
## Annotations
AnnotationAssertion(rdfs:label rb:ModelBasedControl "Model Based Control"@en)
AnnotationAssertion(rdfs:comment rb:ModelBasedControl "Control paradigm that embeds explicit mathematical models of plant dynamics within the controller to predict, plan, and optimise trajectories under physical and operational constraints, encompassing MPC, iLQR, DDP, Whole-Body Control, contact-implicit planning, robust/tube MPC, and Koopman-operator methods."@en)
AnnotationAssertion(dcterms:identifier rb:ModelBasedControl "RB-9018"^^xsd:string)
AnnotationAssertion(dcterms:subject rb:ModelBasedControl "Robotics, Control Theory, Model Predictive Control, Trajectory Optimisation, Legged Locomotion"@en)
)
Property Characteristics
AsymmetricObjectProperty(rb:requires) AsymmetricObjectProperty(rb:enables) AsymmetricObjectProperty(rb:implements) AsymmetricObjectProperty(rb:reduces) TransitiveObjectProperty(rb:dependsOn)
About Model Based Control
- Model Based Control (MBC) represents the dominant paradigm for high-performance robot and process control where physical understanding of the plant can be codified mathematically. In contrast to pure reactive or model-free approaches — which map sensor signals to actuator commands without an internal predictive model — MBC maintains an explicit representation of how the system will evolve under a given sequence of control inputs, enabling anticipatory planning, constraint enforcement, and principled handling of uncertainty.
- The intellectual heritage of MBC spans optimal control theory (Bellman 1957 dynamic programming, Pontryagin 1962 maximum principle), state-space methods (Kalman 1960), and industrial process control (Richalet et al. 1978 MAC; Cutler and Ramaker 1980 DMC), which first applied receding-horizon ideas to constrained multivariable chemical plants. These historical threads converge in modern robotics MBC: whole-body torque control draws on recursive Newton-Euler dynamics (Featherstone 2008), legged locomotion planners exploit Linear Inverted Pendulum model reductions (Kajita 2001), and manipulation controllers use rigid-body dynamics computed by Pinocchio or Drake at kHz rates.
- The 2020s have seen a dramatic convergence of classical MBC with deep learning, driven by three forces: (1) the availability of differentiable physics engines (MuJoCo, Brax, PyBullet) enabling gradient-back-propagation through physical simulations; (2) learned dynamics models that can capture unmodelled phenomena (friction hysteresis, cable stretching, aerodynamic wake effects) beyond classical parametric models; and (3) sim-to-real transfer techniques that allow controllers trained entirely in simulation to transfer to physical robots with minimal fine-tuning.
Core Algorithmic Families
Model Predictive Control (MPC)
MPC is a receding-horizon optimal control scheme solving the finite-horizon problem: minimise J = Σ_{k=0}^{N-1} [xᵀQx + uᵀRu] + x_NᵀP x_N subject to x_{k+1} = Ax_k + Bu_k (linear model), x_k ∈ X, u_k ∈ U at each sample time, applying only the first optimal input and repeating. For nonlinear systems, the dynamics are linearised at each step (linearisation-based NMPC) or handled directly by interior-point NLP solvers (full NMPC). Key solver technologies:
-
OSQP (Stellato et al. 2020): operator splitting QP solver achieving <1 ms solve times for linear MPC problems with 100s of variables.
-
HPIPM (Frison & Diehl 2020): high-performance interior-point solver for structured QPs arising from MPC; achieves 0.1–1 ms solve times on ARM Cortex-A72 class CPUs.
-
acados (Verschueren et al. 2021): open-source framework for real-time NMPC using SQP with Gauss-Newton Hessian approximations; deployed in commercial UAVs, Formula Student autonomous cars, and legged robots.
-
FORCES Pro (Embotech): commercial embedded QP/NLP solver achieving sub-ms NMPC on automotive-grade ECUs (sub-10W power). Linear MPC is now routinely deployed in automotive (Tesla lane-keeping, Waymo path following), energy systems (power grid frequency regulation), and aerospace (SpaceX Falcon 9 landing MPC using convex lossless relaxation — Açıkmeşe & Ploen 2007). For robotics, convex MPC formulations linearising centroidal dynamics enable 500 Hz–1 kHz gait controllers for Boston Dynamics’ Spot (Di Carlo et al. 2018) and MIT Cheetah 3 (Bledt et al. 2018).
Iterative LQR (iLQR) and Differential Dynamic Programming (DDP)
DDP (Mayne 1966; Jacobson & Mayne 1970) minimises a trajectory cost J = Σ l(x_t, u_t) + l_f(x_T) by iterating: Backward pass: Compute value function V(x_t) = min_{u_t} [l(x_t,u_t) + V(x_{t+1}(x_t,u_t))] via second-order Taylor expansion around the current nominal trajectory (x̄, ū). This yields quadratic local models Q_xx, Q_uu, Q_xu and linear feedback gains k_t, K_t. Forward pass: Simulate the improved control u_t = ū_t + k_t + K_t(x_t - x̄_t) to obtain a new nominal trajectory with reduced cost; add line search for convergence. iLQR (Todorov & Li 2004; Tassa et al. 2012) drops second-order dynamics terms (f_xx, f_xu, f_uu ≈ 0), reducing backward-pass complexity while maintaining super-linear convergence in practice. iLQR/DDP are the algorithmic backbone of:
-
MuJoCo MPC (MJPC) (Howell et al. 2022/2024): Google DeepMind’s open-source real-time predictive control system that runs iLQR and CEM (cross-entropy method) asynchronously in separate threads, continuously re-planning at >100 Hz on a standard workstation CPU. MJPC demonstrated real-time whole-body dexterous manipulation, cartwheel locomotion, and in-hand manipulation of complex objects.
-
ALTRO (Howell et al. 2019): AL-iLQR for constrained trajectory optimisation, using augmented Lagrangian to handle inequality and equality constraints not amenable to standard DDP.
-
GPU-accelerated DDP: Park et al. (2021) and Plancher et al. (2021) demonstrated 100× speedups by parallelising DDP rollouts across GPU thread blocks, enabling ensemble planning for robustness.
Whole-Body Control (WBC) and Centroidal Dynamics
Whole-body control formulates robot motion as a hierarchical QP: minimise Σᵢ wᵢ ‖Jᵢ q̈ + J̇ᵢ q̇ - ẍᵢ_d‖² + w_τ ‖τ‖² subject to M(q)q̈ + C(q,q̇)q̇ + g(q) = τ + Jcᵀ λ (rigid-body dynamics) Jc q̈ + J̇c q̇ = 0 (rigid contact: zero acceleration at contact) λ ∈ FC (friction cone: Coulomb friction, normal force ≥ 0) τ_min ≤ τ ≤ τ_max (actuator limits) This resolved motion control framework (Kanoun et al. 2011; Escande et al. 2014) runs at 1 kHz on embedded RT-Linux or real-time Xenomai platforms. Key implementations:
-
mc_rtc (CNRS/AIST): real-time multi-contact control framework used on HRP-4, HRP-5, and Atlas.
-
OCS2 (Farshidian et al. 2020, ETH Zurich): receding-horizon whole-body MPC framework coupling a centroidal MPC with a WBC; deployed on ANYmal-C for stair climbing and object pushing.
-
Crocoddyl (Mastalli et al. 2020, LAAS-CNRS/Edinburgh): efficient DDP library for robot dynamics with analytical derivatives via Pinocchio; 10 ms solve times for 100-node humanoid trajectory optimisation.
Contact-Implicit Trajectory Optimisation
Contact-implicit MO removes the need to pre-specify contact mode sequences by treating contact forces and contact indicator variables as decision variables within the NLP. Complementarity constraints λ ≥ 0, φ(q) ≥ 0, λ·φ = 0 (non-penetration, non-negative normal force, complementarity) are typically relaxed via:
-
Smoothed complementarity (Stewart & Trinkle 1996; Posa et al. 2016): replaces λ·φ = 0 with a sigmoidal approximation, allowing gradient-based solvers (SNOPT, IPOPT) to handle contact implicitly.
-
Contact-implicit MPC (Manchester et al. 2020; Drnach & Zhao 2021): runs smoothed contact-implicit NLP at 10-50 Hz, enabling real-time gait transitions without mode enumeration; demonstrated on Cassie bipedal robot.
-
Lossless Convexification (Açıkmeşe 2007): converts non-convex thrust constraints in rocket landing to convex QP by variable transformation; adopted by SpaceX for Falcon 9/Starship guidance.
Robust, Tube, and Stochastic MPC
Real-world deployments face bounded parameter uncertainty ‖Δf‖ ≤ ε and unbounded (Gaussian) disturbance noise w. Robust MPC variants:
-
Tube MPC (Mayne et al. 2005): computes a nominal trajectory and an invariant “tube” around it; the tracking controller maintains the actual state within the tube regardless of disturbances up to the specified bound. Provides formal recursive feasibility and constraint-satisfaction guarantees.
-
Min-max MPC: solves the worst-case optimisation problem over all realisations of uncertainty; conservative but formally safe; tractable only for ellipsoidal uncertainty sets.
-
Stochastic MPC (Mesbah 2016): propagates probability distributions through the prediction horizon using polynomial chaos expansion or scenario trees; satisfies constraints with specified probability (chance constraints Pr[x ∈ X] ≥ 1-ε).
-
GP-MPC (Hewing et al. 2020): Gaussian processes provide posterior uncertainty estimates that are propagated through the MPC horizon via moment matching; deployed on a Formula Student race car at ETH Zurich achieving 97% of optimal lap time with probabilistic constraint satisfaction.
Koopman Operator MPC
Koopman’s theorem guarantees the existence of an infinite-dimensional linear operator K acting on the space of observables such that K Φ(x) = Φ(f(x)), effectively lifting nonlinear dynamics into a (globally) linear space. In practice, Extended Dynamic Mode Decomposition (EDMD; Williams et al. 2015) identifies a finite-rank approximation K̂ from data by regressing lifted state vectors z_{t+1} against z_t, where z = [x; Φ(x)] includes polynomial, radial basis, or deep-network features. Once K̂ is identified:
-
Linear MPC with standard QP solvers applies directly to z_t;
-
Real-time 100 Hz Koopman MPC has been demonstrated on soft-robot arms (Bruder et al. 2021), human lower-limb exoskeletons, and quadrotor attitude control;
-
Deep Koopman Networks (Lusch et al. 2018; Azencot et al. 2020) learn Φ jointly with K̂ end-to-end, capturing high-dimensional nonlinear phenomena such as fluid wake dynamics and muscle-tendon interactions.
Components and Architecture
-
Model-based control systems decompose into four architectural layers, each with well-defined interfaces and characteristic computational budgets:
Layer 1: Perception and State Estimation
The perception layer reconstructs the full system state x = [q; q̇; p_CoM; L] from partial sensor observations. Key components:
-
IMU integration and fusion: Strapdown integration of gyroscope/accelerometer readings provides high-rate (1–4 kHz) body velocity estimates; extended Kalman filter (EKF) or Unscented KF fuses IMU with kinematics to estimate base velocity. The Invariant EKF (InEKF; Hartley et al. 2019) exploits Lie group symmetry to achieve globally consistent error propagation on SO(3) × ℝ³, reducing linearisation errors on fast-rotating platforms.
-
Visual-inertial odometry (VIO): ORBSLAM3 (Campos et al. 2021), VINS-Mono (Qin et al. 2018), and Kimera-VIO (Rosinol et al. 2020) fuse stereo cameras or depth sensors with IMU to provide drift-corrected 6-DoF pose estimates. Relevant for legged robots operating in GPS-denied environments; typical position accuracy 1–3 cm over 100 m traversed.
-
Contact state detection: Leg force-torque sensors (FTS) or motor current observers classify contact (in-contact / airborne) from normal force estimates; contact Jacobians J_c are updated accordingly in the WBC QP. False contact detection is a leading cause of locomotion failure — research groups (ETH, Oxford) use learned contact detection from IMU vibration patterns.
-
Proprioceptive forward kinematics: Joint encoder readings q provide precise link positions; the Recursive Newton-Euler Algorithm (RNEA) and forward kinematic Jacobians map joint-space state to Cartesian-space tool and CoM positions. Model-based observers (Mistry et al. 2010) estimate contact forces from joint torque sensors without explicit FTS. The state estimator outputs (x̂, P_x) — posterior mean and covariance — consumed by the MPC. Estimator frequency determines MPC feedback rate; 500 Hz is typical for legged robots.
Layer 2: Model and Prediction
The dynamic model f(x, u) predicts x_{t+1} from current state and control input. Three model archetypes co-exist:
-
Physics-based rigid-body models: Pinocchio 2.6+ (Carpentier et al. 2019), Drake (Tedrake et al.), and RBDL compute the equations of motion M(q)q̈ + C(q,q̇)q̇ + g(q) = τ + J_cᵀ λ via Composite Rigid Body Algorithm (CRBA, O(n²)) or Articulated Body Algorithm (ABA, O(n)). Analytical derivatives (via Pinocchio’s Casadi/Autodiff interface or Drake’s gradient computation) enable gradient-based trajectory optimisation with <1 ms evaluation times for 36-DoF humanoid.
-
Data-driven models: Neural ODE (Chen et al. 2018) parameterises ẋ = f_θ(x, u) with a neural network and integrates with ODE solvers (RK4, Dormand-Prince); trained from physical interaction data. GP dynamics (Deisenroth & Rasmussen 2011 — PILCO) provide calibrated posterior uncertainty. Neural network ensembles (Chua et al. 2018 — PETS) capture epistemic uncertainty via disagreement, halting planning when predicted variance exceeds a threshold.
-
Hybrid residual models: The rigid-body model handles known kinematics and inertia; a neural residual δf_θ(x, u) corrects for unmodelled effects (foot compliance, motor flexibility, aerodynamic drag, cable dynamics). This architecture (Lee et al. 2020, Learning-Based WBC) preserves the interpretability and physical constraints of the rigid-body model while adapting to real-world deviations. The prediction model is evaluated N times (N=20–50 typical) at each MPC step to unroll the prediction horizon. GPU parallelism (Isaac Gym, MJX) allows 1,000–16,384 parallel rollouts simultaneously for sampling-based MPC variants.
Layer 3: Planning and Optimisation
The planning layer solves the finite-horizon OCP at the control frequency (10–100 Hz for NMPC, 100–1000 Hz for QP-based linear MPC):
-
Linear MPC (QP): For convex centroidal dynamics or linearised models, the OCP reduces to a quadratic program min 0.5 zᵀHz + fᵀz s.t. Az = b, Gz ≤ h, where z stacks state and control decision variables. OSQP/HPIPM solve this in <1 ms on embedded CPUs.
-
Nonlinear MPC (NLP): Full NMPC with nonlinear dynamics solved via IPOPT (interior-point, BFGS Hessian approximation), SNOPT (SQP, sparse Schur complement), or acados (RTI-SQP — single SQP step per control cycle). Solve times: 10–100 ms for 20-node horizon, 40-DoF system.
-
DDP/iLQR: Unconstrained trajectory optimisation via backward-forward passes; O(N·n³) per iteration for n state dimensions, N horizon steps. Warm-started from previous solution; typically 5–20 iterations to convergence.
-
Sampling-based MPC (MPPI, CEM): Generate K=1,000–16,384 random control perturbations, roll out in parallel on GPU, compute importance-weighted cost average (MPPI) or select elite trajectories (CEM). No gradient required — applicable to non-differentiable dynamics. MJPC uses CEM + iLQR asynchronously. Frameworks providing autodiff interfaces: CasADi (symbolic, C-code generation), JAX (functional autodiff, GPU-native), PyTorch (dynamic computation graph, CUDA), and Pinocchio’s analytical derivatives (fastest, requires model specification in Pinocchio format).
Layer 4: Execution and Feedback
The execution layer converts planned torques/forces to hardware commands at rates of 1–4 kHz:
-
Torque control: Desired joint torques τ_d from the WBC QP are sent to actuator amplifiers via CAN bus, EtherCAT, or real-time UDP. Series elastic actuators (SEA) and quasi-direct-drive (QDD) motors require additional inner-loop torque controllers compensating spring deflection or motor reluctance.
-
Impedance control: A Cartesian impedance law F = K_d (x_d - x) + D_d (ẋ_d - ẋ) provides passive compliance during contact transients, preventing high-frequency torque spikes when the contact model is inaccurate. Variable impedance (Hogan 1985; Buchli et al. 2011 learning variable impedance) adapts stiffness to task phase.
-
Disturbance observers (DOB): Q-filter-based DOBs (Ohnishi 1987; Liu & Goldenberg 1996) estimate lumped disturbance d = τ_ext + Δf(x, u) from measured joint velocities and command torques, injecting feedforward compensation τ_ff = -d_hat at kHz rates. Effective for managing model-reality gaps without re-identification.
-
Real-time operating systems: RT-Linux (PREEMPT_RT patch), Xenomai 3, and QNX Neutrino provide deterministic scheduling with <100 μs jitter for the control loop. ROS 2 real-time executors and OROCOS RTT manage inter-layer communication with lock-free message passing.
-
-
The interfaces between layers are standardised by frameworks such as ROS 2 real-time middleware, OROCOS RTT, and ETH’s signal_logger/raisim ecosystem, enabling modular substitution of components (e.g., swapping a GP dynamics model for a neural ODE without altering the MPC solver).
-
System identification (SysID) spans layers 2 and 4: offline SysID (Ljung 1999 prediction error methods; subspace methods N4SID) estimates static model parameters θ (inertia, friction) from excitation experiments before deployment; online SysID (recursive least squares, EKF parameter estimation) adapts θ in real time to handle payload changes, wear, and environmental variation without requiring offline re-identification experiments.
Use Cases and Major Application Families
Legged Locomotion
Quadruped and bipedal locomotion represents the most demanding MBC deployment: contact topology changes at 10–50 Hz (gait cycle), terrain geometry is uncertain, and falls are catastrophic.
Boston Dynamics Spot and Atlas (2018–2026): The operational MBC stack combines four tightly coupled layers:
-
Footstep planner (5–10 Hz): generates desired foot placements using Linear Inverted Pendulum MPC (LIPM) or centroidal model, considering terrain height maps from lidar/stereo.
-
Convex centroidal MPC (40–500 Hz): optimises CoM trajectory and ground reaction force (GRF) profiles over a 400 ms horizon using a 12-state single rigid body (SRB) model; OSQP solves the QP in <2 ms.
-
Whole-body controller QP (1 kHz): maps desired GRFs and task-space accelerations to joint torques via constrained optimisation; handles contact constraints, friction cone membership, and actuator limits simultaneously.
-
Torque controller (4 kHz): PD feedback + model-based feedforward on each joint; joint-level compliance via SEA spring state estimation. Di Carlo et al. (2018) demonstrated this stack on MIT Cheetah 3, achieving 2.45 m/s trotting and robust recovery from push disturbances (150 N lateral impulse). Boston Dynamics’ Spot commercial product uses the same architecture with terrain-adaptive footstep planning.
ANYmal (ETH Anybotics): OCS2 (Farshidian et al.) framework coupling centroidal MPC (40 Hz, 20-node horizon) with WBC (400 Hz). Demonstrated: stair climbing (30 cm steps at 0.3 m/s), outdoor navigation on rubble and mud, payload-carrying (10 kg backpack), and collaborative inspection of industrial facilities (Skanska/ABB deployments in 2024). The Learning-based WBC (Lee et al. 2020) adds a neural residual to the rigid-body model, improving torque tracking by 23% on hardware.
IIT HyQ and HyQ-Real: Fahmi et al. (2020) deployed terrain-adaptive passive WBC with real-time contact point updates from an elevation map (ETH Anymal-style). HyQ demonstrated reactive stepping over unexpected 15 cm obstacles at 150 ms reaction time; the terrain-adaptive MPC updates contact normal directions from depth sensor measurements at 20 Hz.
MIT Humanoid and Unitree H1/G1 (2024–2026): MIT’s 2024 humanoid demonstrator uses MJPC (iLQR + CEM asynchronous) for real-time whole-body manipulation — catching thrown objects at 3 m/s, using tools, and opening doors without task-specific reward engineering. Unitree H1’s open-source SDK exposes a WBC interface; the community has ported Crocoddyl and OCS2 controllers targeting 40 Hz NMPC.
Aerial Robotics (UAV/MAV)
MPC is the de facto standard for multirotor trajectory tracking, replacing earlier PID cascade controllers:
-
ETH Zurich RPG/RSL (Scaramuzza, D’Andrea): demonstrated quadrotors catching balls mid-flight (Mueller et al. 2019, 4 ms reaction window), flying through gaps at 10 m/s (Foehn et al. 2021), and collaborative cable-suspended payload transport. Full NMPC with 10-step horizon at 100 Hz; acados SQP solver, <10 ms solve time on Odroid XU4.
-
AlphaPilot (2021): Foehn et al. demonstrated MPC-guided autonomous drone racing on the AlphaPilot championship circuit, achieving superhuman lap times (0.3 s faster than champion human pilot). The controller uses polynomial minimum-snap trajectory planning with a 3D point-mass MPC tracking layer.
-
Fixed-wing / VTOL: Tube MPC with aerodynamic envelope constraints (stall speed, maximum load factor, bank angle) applied in aerospace research at DLR, ONERA, and BAE Systems. The constraint set enforces structural limits and prevents departure from controlled flight during agile manoeuvres.
-
Swarm coordination: Convex distributed MPC (DMPC) enables collision-free trajectory optimisation across swarms of 10–100 UAVs, decomposing the multi-agent problem into per-agent QPs with inter-agent collision constraints (Augugliaro et al. 2012, ETH Flying Machine Arena).
Automotive and Autonomous Driving
The global automotive MPC deployment spans embedded ECUs to cloud-based planning:
-
Tesla Autopilot: Kinematic bicycle model MPC for lane-keeping and lane-changing at 50–130 km/h; QP-based with 3 s prediction horizon; runs on Tesla’s custom HW3/HW4 inference chip.
-
Waymo Driver: Full NMPC over 5–8 s horizon with probabilistic occupancy grid constraints; chance constraints Pr[collision] < 10⁻⁶ per scenario. Scenario-based stochastic MPC handles pedestrian trajectory uncertainty via Monte Carlo sampling.
-
Adaptive cruise control (ACC): Continental ARS/MFC, Bosch SCC, and Aptiv ACC modules use linear MPC with 2–4 s horizon on Aurix TC3xx or Renesas R-Car SoCs; OSQP solves the QP in <0.5 ms.
-
Autonomous racing: MPC controllers at the Indy Autonomous Challenge (2021–2023) achieved 270 km/h lap speeds with friction-circle MPC enforcing tyre dynamics constraints (Raji et al. 2022). Formula Student autonomous cars (ETH AMZ, TU Munich) use NMPC with GP tyre models at 100 Hz.
-
Parking and low-speed manoeuvring: Nonlinear MPC with kinematic bicycle model enforces geometric constraints (kerb avoidance, lane boundaries, turning radius limits) for automated parking in sub-1 m precision; deployed in Mercedes EQS, BMW 7 Series, and Volvo XC90 automated parking systems.
Industrial Process Control
MPC originated in chemical and petroleum process industries and remains the dominant advanced control technology there:
-
Refinery operations: RMPCT (Honeywell), DMCplus (AspenTech), and SMOC (Shell) regulate distillation columns (50–200 tray), fluid catalytic crackers (FCC), and hydrotreaters with 50–500 controlled/manipulated variables, dead times of 5–60 minutes, and strong multi-variable interactions. Global installed base: >10,000 advanced MPC applications in refineries, delivering 2–5% throughput uplift and 10–30% energy savings per implementation (DMCplus benchmark studies, 2018).
-
Pharmaceutical batch control: FDA Process Analytical Technology (PAT) guidelines encourage model-based monitoring; MPC regulates crystallisation reactors, granulation fluidised beds, and lyophilisation chambers to maintain product quality attributes (particle size distribution, moisture content) within specification.
-
Building HVAC: MPC-based building climate control (Siemens Navigator, Schneider Electric EcoStruxure) optimises chiller, AHU, and pump scheduling over 24-hour horizons incorporating weather forecasts and occupancy predictions, achieving 15–30% energy savings versus rule-based controls in commercial buildings.
-
Power systems: Grid frequency regulation, wind turbine control (Cortez-Perez 2020 GP-MPC for fatigue-load minimisation), and battery energy storage system (BESS) dispatch via MPC are active deployment areas, with the National Grid ESO trialling MPC-based inertia emulation at multiple UK grid-scale batteries (2025).
Surgical and Medical Robotics
Model-based control is central to surgical robot safety and performance:
-
da Vinci (Intuitive Surgical): Inverse-kinematics model enforces remote centre of motion (RCM) — the invariant point through which laparoscopic tools pass the abdominal wall. Constrained motion scaling (master-slave with 3:1–10:1 motion attenuation) uses model-based Jacobian computation updated at 1 kHz. The da Vinci Xi incorporates active force feedback estimation from joint motor currents.
-
Tissue-interaction MPC (UCL/Imperial): Research at UCL Hamlyn Centre and Imperial Personal Robotics Lab implements viscoelastic tissue contact models within real-time force-control MPC. The controller minimises tissue deformation while tracking surgeon-defined tool trajectories, predicting interaction forces 100–200 ms ahead using Hunt-Crossley contact mechanics models. Critical for preventing perforation during microsurgery and endonasal skull-base procedures.
-
Rehabilitation exoskeletons: Lower-limb exoskeletons (ETH Hocoma Lokomat, Ekso Bionics EksoGT) use model-based impedance control with gait-phase-dependent impedance parameters, updated by a high-level MPC that adapts assistance level based on measured user interaction forces and EMG signals.
-
Radiotherapy positioning: 6-DoF patient positioning stages use model-based feedforward (computed torque) plus disturbance-observer feedback to achieve <0.5 mm positioning accuracy during beam delivery, compensating for respiratory motion via pre-identified motion models.
Space and Spacecraft Control
-
Rocket landing (SpaceX Falcon 9/Starship): The powered descent guidance problem (Açıkmeşe & Ploen 2007) uses lossless convexification to transform the non-convex fuel-optimal control problem into a second-order cone program (SOCP) solvable in real time. The G-FOLD algorithm (Acikmese et al. 2013) demonstrated this on Masten Space XOMBIE in 2014; adapted for Falcon 9 booster recovery (2015–present) with 1 m landing accuracy.
-
Attitude control: MPC-based attitude and orbit control systems (AOCS) on small satellites enforce slew rate limits, reaction wheel saturation constraints, and solar exclusion zones simultaneously — infeasible for classical decoupled PID. ESA’s OPS-SAT demonstrator (2020) ran experimental MPC AOCS in orbit.
-
Proximity operations: Spacecraft rendezvous and docking MPC (NASA Orion, ESA ATV) enforces safety exclusion zones, approach corridor constraints, and fuel consumption limits over 10–30 minute horizons, with the Clohessy-Wiltshire linear model enabling convex QP formulations.
Academic Context
-
Model-based control has deep roots in the mathematical control theory community, with four intellectual lineages contributing the modern synthesis:
Optimal Control Theory (1950s–1970s)
-
Richard Bellman (RAND Corporation / USC): Dynamic Programming (1957) and the Principle of Optimality established the backward-sweep value function framework that underpins all DDP and MPC theory. Bellman’s curse of dimensionality — exponential grid scaling with state dimension — motivated the low-dimensional model approximations and receding-horizon strategies that define practical MBC.
-
Lev Pontryagin (Steklov Institute, Moscow): Maximum Principle (1962) provided necessary conditions for optimality in continuous-time optimal control without requiring Bellman’s full value function, enabling trajectory optimisation via adjoint (co-state) equations. Modern DDP backward passes solve the discrete-time Pontryagin conditions via second-order value expansion.
-
Rudolf Kalman (Bell Labs / Stanford): LQR / LQG framework (1960) formalised the linear-quadratic optimal control problem and its connection to state estimation (Kalman filter), establishing the separation principle and enabling practical implementation on analog computers. The Riccati equation solution P is the infinite-horizon analogue of MPC’s terminal cost matrix.
-
Early aerospace validation: Apollo Guidance Computer (MIT IL, 1969) ran offline-optimised powered-descent guidance; Saturn V attitude control used gain-scheduled LQR; Space Shuttle used discrete-time LQG with 25 Hz update rate.
Industrial MPC (1970s–1990s)
-
Jacques Richalet (ADERSA, France): Model Algorithmic Control / MAC (1978) — first commercially deployed MPC, using impulse-response step models for multivariable constrained control of chemical plants. Applied to Elf Aquitaine refinery distillation column with 20 inputs, 10 outputs, 1-hour dead time.
-
Charles Cutler and Brian Ramaker (Shell): Dynamic Matrix Control / DMC (1980 ACC) — state-space MPC formulation with quadratic cost, constraint handling via LP, applied to FCC units at Shell Deer Park refinery; over 1,000 DMC applications installed by 1995.
-
Carlos Garcia, David Prett, and Manfred Morari (Caltech/Shell): QDMC (1989) added rigorous stability analysis, multiple-model uncertainty, and robust constraint satisfaction. Morari’s group (ETH Zurich, 1990–2015) produced the definitive theoretical stability and optimality framework (Mayne et al. 2000).
-
David Mayne (Imperial College London, later UC Davis): Unified MPC stability proofs (Mayne, Rawlings, Rao, Scokaert 2000) covering both infinite-horizon and finite-horizon with terminal sets; tube MPC for robust constraint satisfaction (Mayne, Seron, Raković 2005). Mayne is widely considered the pre-eminent theorist of constrained MPC.
Robot Control (1980s–2000s)
-
Oussama Khatib (Stanford AI Lab): Operational-space formulation (1987) computed task-space dynamics from rigid-body equations, enabling intuitive specification of robot tasks in end-effector space while computing joint torques that satisfy full Newton-Euler dynamics. Foundational for WBC.
-
Bruno Siciliano (Naples Federico II), Luigi Villani, Giuseppe Oriolo: Textbook formalisation of robot dynamics, trajectory planning, computed torque control, and impedance control (Siciliano et al. 2010) — the standard European robotics curriculum text.
-
Roy Featherstone (ANU, then IIT): Spatial algebra formulation and recursive Newton-Euler algorithm (RNEA, 1987; Rigid Body Dynamics Algorithms 2008) enabling O(n) computation of robot dynamics — the computational foundation of all real-time MBC for manipulators and legged robots.
-
Emanuel Todorov (Salk Institute, then UW/Google): iLQR and DDP for robotics (Todorov & Li 2004; Tassa, Erez, Todorov 2012); MuJoCo physics engine (2012, open-sourced 2022 by DeepMind). Todorov’s work catalysed the modern fusion of trajectory optimisation with deep RL.
-
Nicolas Mansard (LAAS-CNRS, Toulouse): Multi-contact DDP, Crocoddyl library, and the MEMMO (Memory of Motion) project; key figure in humanoid whole-body MPC. Collaborated extensively with Edinburgh Robotarium (Sethu Vijayakumar).
Model-Based Reinforcement Learning (2010s–2020s)
-
Marc Deisenroth (UCL, now Imperial): PILCO (2011) — Gaussian process dynamics model with analytic policy gradients; first algorithm to learn cartpole in seconds of real interaction. PILCO’s probabilistic uncertainty propagation through the prediction horizon influenced subsequent GP-MPC and safe-MBC research.
-
Sergey Levine (UC Berkeley RAIL): Guided policy search, MBPO, PETS (2015–2019) — systematic study of learned dynamics models within MPC planning loops. Levine’s 2020–2024 survey series on MBRL established the taxonomic framework (pure MPC / Dyna / latent world model) used by most subsequent papers.
-
Danijar Hafner (Google DeepMind): DreamerV1 (2019) → DreamerV2 (2020) → DreamerV3 (2023): progressive demonstration that recurrent world models trained from pixels can support planning and imagination-based policy learning achieving superhuman performance on 150+ tasks including Atari, DMControl, Minecraft, and robotic manipulation.
-
Key academic venues: IEEE ICRA, IEEE/RSJ IROS, Robotics: Science and Systems (RSS), IFAC World Congress, IEEE CDC, Learning for Dynamics and Control (L4DC), NeurIPS, ICML, CoRL (Conference on Robot Learning), Journal of Field Robotics, IJRR, Automatica, IEEE Transactions on Robotics.
-
Current Landscape (2026)
-
The year 2025–2026 represents a period of deep integration between classical MBC and deep learning, with six dominant trends shaping research and commercial deployment:
Trend 1: Learned World Models for MPC
-
DeepMind’s DreamerV3 (Hafner et al. 2023) — recurrent state-space model (RSSM) with categorical latent states, trained from pixels via reconstruction loss — achieves superhuman performance on 150+ tasks including Atari, DMControl Suite, Minecraft, and humanoid locomotion, using model-based imagination rollouts for both actor-critic training and online planning.
-
TD-MPC2 (Hansen et al. 2023) scales latent-state MPC to 80+ continuous control tasks, achieving 10× sample efficiency over model-free TD3/SAC; uses a shared latent dynamics model for multi-task planning and a model-predictive path integral for action selection.
-
IRIS (Micheli et al. 2022, EPFL): transformer-based discrete world model for Atari, demonstrating that attention mechanisms can capture long-range temporal dependencies in environment dynamics better than recurrent models for sparse-reward tasks.
-
Genie 2 (Google DeepMind 2024): generative interactive environment model trained on internet video, producing photo-realistic playable worlds from a single image prompt; first step toward foundation environment models supporting MPC planning in imagined scenarios without prior robot data.
Trend 2: Foundation Dynamics Models
-
UniSim (Yang et al. 2023, Stanford/Berkeley): pre-trains a neural simulator on diverse robot interaction data (grasping, pushing, pouring), enabling zero-shot transfer to new robot morphologies and task configurations via a conditioning interface. Reduces physical trials for manipulation policy development from 500 to <50.
-
RoboDreamer (2024, CMU/MIT): diffusion-based world model for contact-rich manipulation planning; models high-dimensional visual observations with probabilistic dynamics, enabling MPC in visual space for tasks like folding cloth and pouring liquid.
-
Transformer dynamics models (2024–2025): multiple groups (ETH, Berkeley, DeepMind) training transformer-based dynamics models on large robot datasets (Open X-Embodiment, DROID) achieving cross-robot generalisation of dynamic predictions — early-stage “GPT-for-robot-physics” paradigm.
Trend 3: MuJoCo MPC (MJPC) Ecosystem (2022–2026)
-
Google DeepMind’s open-source MJPC release (Howell et al. 2022, updated 2023–2025) created a common benchmark for real-time MBC research with an asynchronous iLQR + CEM architecture that continuously re-plans at >100 Hz on a standard CPU workstation.
-
MuJoCo 3.0 (2023) with MJX (MuJoCo on JAX/GPU) enables automatic differentiation through physics simulation and GPU-accelerated rollouts — critical for both MBRL training and GPU-parallel MPPI planning.
-
Downstream MJPC work: hand manipulation (Allshire et al. 2024), locomotion benchmarking on 8 quadruped/biped platforms, integration with ROS 2 for hardware deployment on Unitree and Franka robots, and a Python API enabling researchers to define custom cost functions in JAX.
Trend 4: Hardware-Efficient and Sampling-Based MPC
-
NVIDIA Isaac Gym / Isaac Lab (2021–2024) and MJX (MuJoCo on JAX, 2023): 16,384 parallel trajectories on a single A100 GPU, enabling gradient-free MPPI and CEM MPC with 10 ms planning latency at 100 Hz.
-
MPPI deployments (2022–2025): aggressive off-road driving at 10 m/s on rough terrain (Georgia Tech AUTORALLY project), quadrotor aerobatics at 20 m/s (ETH), legged locomotion over debris fields (MIT), and robot arm manipulation under unknown friction — all without requiring differentiable dynamics models.
-
FPGA/ASIC MPC accelerators: Embotech and academic groups (ETH, Imperial) developing hardware-accelerated QP chips (Xilinx Zynq UltraScale+, custom ASIC) targeting <100 μs QP solve times for automotive MPC at <5W power consumption.
Trend 5: Koopman MPC Maturation and Commercial Adoption
-
Multiple 2024 IJRNC benchmarks: Koopman MPC achieves within 5% of full NMPC performance on chemical reactors, robot arms, and building HVAC systems, solving as a linear QP 100× faster than equivalent NLP.
-
Commercial adoption: Mitsubishi Electric deploying Koopman HVAC MPC in Japanese commercial buildings (2025 pilot); Johnson Controls integrating Koopman-based chiller optimisation in OpenBlue platform; ABB applying Deep Koopman MPC to industrial motor drive control.
-
Deep Koopman Networks (Azencot et al. 2020; Shi et al. 2022): end-to-end learning of observable functions Φ jointly with operator K̂, handling high-dimensional inputs (camera images, full robot state vectors) — enabling Koopman MPC from raw sensor streams without manual feature engineering.
Trend 6: Contact-Rich Manipulation and Dexterous Hands
-
2025 manipulation breakthroughs: MIT EVS and CMU TCDP demonstrated real-time 10 Hz contact-implicit MPC for plug insertion, bolt driving, and snap-fit assembly using smoothed complementarity NLPs solved on GPU clusters.
-
Dexterous hand MPC: Multi-fingered dexterous manipulation (Shadow Hand, Allegro Hand) using contact-implicit MPC on 24-DoF systems; demonstrated in-hand reorientation of arbitrary objects without pre-specified grasp sequences (2025, MIT CSAIL + ETH RSL).
-
Sim-to-real gap for contact: Physics-based randomisation (friction coefficient ± 50%, compliance ± 30%) in MuJoCo training combined with online GP residual correction enables <5% real-to-sim performance gap for contact-rich tasks — down from 30–50% in 2020.
-
UK Context
-
The UK has a distinguished research tradition in model-based control, spanning theoretical foundations (Mayne at Imperial, Maciejowski at Cambridge), field robotics (Oxford, Edinburgh), rehabilitation engineering (Imperial, UCL), and industrial deployment (BAE Systems, Rolls-Royce, AMRC):
Oxford Robotics Institute (ORI)
-
Leadership: Professor Ingmar Posner (perception for autonomy), Professor Maurice Fallon (state estimation and legged robotics), Dr. Ioannis Havoutis (locomotion MPC)
-
Research focus: Model-based SLAM and terrain-adaptive locomotion MPC for field robots. ORI’s legged robotics group uses terrain-aware convex MPC that updates contact point predictions from 3D lidar point clouds at 20 Hz, enabling reactive stepping on uneven outdoor terrain.
-
Projects: ONR-funded legged robot autonomy (2022–2026); Oxford Mars rover testbed for planetary exploration MPC; GOALS project (Geometry-aware Online Adaptive Locomotion using SLAM) — integrates SLAM elevation maps directly into the MPC prediction model.
-
UK Space Agency collaboration: Lunar rover MPC for ESA Prospect mission (planned 2026 Moon South Pole landing), requiring model-based control under 1–3 s communication latency and uncertain terrain geometry.
Edinburgh Robotarium / Edinburgh Centre for Robotics (ECR)
-
Leadership: Professor Sethu Vijayakumar (whole-body control, learning for robotics), Dr. Carlos Mastalli (DDP/Crocoddyl), Dr. Wolfgang Merkt
-
Research focus: Whole-body MPC for manipulation and locomotion; GP-based learning for robot dynamics; human-robot interaction under model uncertainty.
-
Crocoddyl co-development: Edinburgh co-developed Crocoddyl with LAAS-CNRS (Nicolas Mansard); the library achieves 10 ms solve time for 100-node humanoid trajectory optimisation via Pinocchio analytical derivatives and Cholesky factorisation.
-
MEMMO project (EU H2020): Memory of Motion — MPC warm-started from a database of pre-computed DDP trajectories reduces online solve time by 60% for humanoid motion planning; demonstrated on Talos and Pepper humanoids.
-
ANYmal deployment: Edinburgh operates 2 ANYmal-C robots instrumented with 3D lidar and RGB-D cameras; OCS2-based WBC-MPC stack runs for autonomous inspection and human-robot collaboration tasks.
Imperial College London
-
Personal Robotics Lab (Professor Petar Kormushev): model-based RL for manipulation and rehabilitation robotics; MPC-based gait optimisation for lower-limb prosthetic devices; online SysID of soft tissue mechanical properties (Hunt-Crossley contact model) integrated within MPC force-control loop for simulated laparoscopy.
-
Hamlyn Centre for Robotic Surgery (Professor Guang-Zhong Yang, Professor Philip Pratt): model-based control for surgical robots, including RCM constraint enforcement, force-controlled tissue interaction, and MPC for anastomosis task assistance.
-
Control and Power Group (Professor Eric Kerrigan): embedded MPC theory for resource-constrained hardware; sparse QP reformulations and code-generation tools (FPGA-based MPC accelerators); contributions to robust and stochastic MPC for energy systems.
-
David Mayne legacy: Mayne (Emeritus) formulated tube MPC at Imperial; the Control and Power Group continues robust MPC research extending Mayne’s framework to nonlinear systems.
Manchester Control Systems Group
-
Leadership: Professor Guido Herrmann, Professor Zhengtao Ding (adaptive control), Dr. Timothy Rogers (stochastic MPC)
-
Robust and adaptive MPC: tube MPC variants for uncertain nonlinear systems; disturbance-observer-based MPC; adaptive control for systems with parametric uncertainty (unknown payload, changing friction).
-
Power systems MPC (EPSRC 2023–2028): robust MPC for grid frequency regulation under renewable intermittency; integrating battery energy storage system (BESS) dispatch with wind/solar generation forecasts; collaboration with National Grid ESO.
-
Collaborative manufacturing: model-based adaptive control for UR10e cobots on shared workspaces with human operators; impedance MPC adapting stiffness in real time based on measured human proximity.
Cambridge Machine Intelligence Lab
-
GP-MPC theory (Professor Richard Turner, Dr. Mark van der Wilk): sparse GP approximations (inducing point methods) enabling GP-MPC on problems with >1000 training points; probabilistic safety certificates for stochastic MPC via GP posterior confidence bounds.
-
Probabilistic numerics (Turner’s group): treats numerical integration of ODEs as Bayesian inference, providing uncertainty estimates on the numerical solution itself — directly improving accuracy of GP predictions in MPC rollouts.
-
Secondmind (formerly PROWLER.io, Cambridge spin-out): commercialised GP-MPC for industrial optimisation (supply chain, energy scheduling) before acquisition; alumni now at Google DeepMind and Wayve contributing to learned MPC for autonomous driving.
Northern English Industry
-
BAE Systems (Samlesbury and Warton, Lancashire): robust MPC and gain-scheduled LQR for Typhoon Eurofighter flight control (active since 2003); leading Project Tempest / UK GCAP sixth-generation combat air, where model-based control is a critical technology for supercruise and high-AoA manoeuvring.
-
Rolls-Royce (Derby): model-based turbofan engine control using gain-scheduled MPC with adaptive loop handling degradation-driven model shifts (compressor blade erosion, combustor fouling). Rolls-Royce Research Partnerships with Loughborough University on model-based health monitoring and prognostic MPC.
-
AMRC (Advanced Manufacturing Research Centre, Sheffield, Catapult member): model-based adaptive control for CNC machining centres (real-time cutting force models within MPC feed-rate optimisation achieving 15% cycle time reduction, 30% tool wear reduction) and robotic welding cells (adaptive MPC compensating for thermal distortion).
-
Dyson (Malmesbury, Wiltshire): model-based control for robotic vacuum cleaners (brush motor impedance control) and surgical robotics division (constrained MPC for tool-tissue interaction force limiting).
-
Model-Based vs Model-Free RL: Comparative Analysis
-
The relationship between MBC and model-free reinforcement learning is one of the most actively debated axes in robot learning research (2020–2026). The trade-offs are well characterised:
Sample Efficiency
-
Model-based methods (PETS, MBPO, DreamerV3, TD-MPC2) consistently achieve 10×–1000× better sample efficiency than model-free methods (PPO, SAC, TD3) on continuous control benchmarks:
- PILCO learns cartpole in ~20 physical trials; model-free REINFORCE requires 5,000+
- MBPO achieves HalfCheetah 10K reward with 5K steps; SAC requires 300K
- TD-MPC2 achieves humanoid locomotion in 100K steps; DDPG requires 2M+
-
Sample efficiency advantage is most pronounced for low-dimensional systems (<50 state dimensions) where accurate dynamics models can be learned; the gap narrows for high-dimensional visual control where model learning itself requires data.
Asymptotic Performance
-
Model-free RL often achieves higher asymptotic performance given sufficient data:
- PPO + domain randomisation achieves 99% success on dexterous manipulation tasks in simulation; MBPO plateaus at 85–90% due to model bias
- Model bias — systematic prediction errors in regions not visited during training — causes compounding errors over long prediction horizons, degrading plan quality
- Model-free methods bypass this by directly optimising the policy without long-horizon rollouts
-
Exception: well-calibrated physics-based models (Pinocchio, Drake) achieve near-perfect asymptotic performance because model error is negligible (< 1% for rigid-body systems without significant unmodelled effects)
Constraint Satisfaction
-
MBC provides hard constraint guarantees that model-free RL cannot: a QP-based WBC with friction cone constraints is provably guaranteed to produce contact forces within the feasible set, regardless of how improbable that region was during training; model-free RL can only satisfy constraints softly via reward shaping.
-
For safety-critical applications (surgical robotics, aerospace, nuclear): MBC with formal constraint satisfaction is mandatory; model-free RL with soft constraints is insufficiently reliable.
Interpretability and Debugging
-
MBC cost functions have direct physical meaning: a term ‖x_CoM - x_ref‖² directly penalises CoM tracking error; an engineer can inspect which constraint was binding and why the controller chose a particular trajectory.
-
Model-free policies are opaque: a PPO policy network output depends on 1M+ parameters with no interpretable decomposition of the control decision.
Practical Hybridisation (Dyna and Beyond)
-
Most competitive 2024–2026 systems use hybrid architectures that combine MBC and model-free RL:
- Dyna-style (Sutton 1991): model-free value function + model-based synthetic rollouts for policy improvement. MBPO generates 95% synthetic / 5% real rollouts.
- Model-predictive RL: DreamerV3 trains an actor-critic using imagined rollouts from a world model; the actor is model-free but benefits from model-generated training data.
- TD-MPC / TDMPC2: latent MPC layer extracts structured planning signal; model-free Q-function critic corrects for model bias near the planning horizon boundary.
- Model-based pre-training + model-free fine-tuning (Nagabandi et al. 2018): MPC with neural dynamics model initialises a policy; SAC fine-tuning corrects residual errors in 10–50 physical episodes.
-
Future Directions (2026–2030)
2026–2027: Hardware-Accelerated Real-Time NMPC
-
Humanoid-scale NMPC: 40+ DoF whole-body NMPC at 100 Hz requires sub-10 ms NLP solve times — presently achievable only on multi-core x86 workstations, not embedded ARM/RISC-V SoCs. Emerging approaches:
- Neural value function approximators replacing backward-pass DDP sweeps (VFMPC)
- GPU-parallelised SQP with warm-starting from previous solution
- FPGA/ASIC QP chips (Embotech, academic groups at ETH/Imperial): <100 μs QP solve at <5W power
-
Unified MPC frameworks: acados 2.0 and Crocoddyl 2.0 (anticipated 2026) target unified Python-native interfaces for physics models, cost functions, and constraints, with backend selection (CPU SQP / GPU MPPI / FPGA QP) via a single API.
2027–2028: Foundation Dynamics Models
-
Pre-trained transformer dynamics models generalising across robot morphologies, enabling MPC without per-robot system identification:
-
Cross-embodiment dynamics: one model predicts dynamics for quadruped, biped, manipulator, and aerial platforms by conditioning on a morphology descriptor vector
-
Zero-shot contact prediction: pre-trained contact dynamics models from the Open X-Embodiment dataset generalise to novel object geometries and material properties with <10 physical trials
-
Expected timeline: manipulation generalisation by 2027, locomotion by 2029
2027–2028: Formal Safety with Learned Models
-
-
Control Barrier Functions (CBFs) + neural dynamics: CBF conditions h(x) ≥ 0 ∀t are verified via neural network Lyapunov certificates, providing forward-invariance guarantees even under neural dynamics model uncertainty.
-
Hamilton-Jacobi reachability with GP dynamics: computes backward reachable tube under GP posterior dynamics, certifying safe operating envelopes with probabilistic guarantees (Berkenkamp et al. 2017 approach scaled to 12–40 DoF).
-
Expected TRL-5 (technology validated in relevant environment) for medical robotics and aerospace applications by 2027.
2028–2030: Koopman MPC for Biological and Clinical Systems
-
Artificial pancreas: EU clinical trials of Koopman MPC closed-loop insulin delivery (planned 2027); targets HbA1c reduction of 0.3–0.5% vs PID-based control in Type 1 diabetes.
-
Cardiac electrophysiology: model-based pacing control for atrial fibrillation using Koopman-linearised electromechanical heart model; research phase at Imperial BioMedEng and King’s College London.
-
Neural prosthetics: Koopman MPC for functional electrical stimulation (FES) coordinating paralysed limb muscles; building on UCL and Imperial BCI group work.
2026–2030: Carbon-Aware and Net-Zero MPC
-
Building-HVAC MPC integrating National Grid ESO real-time marginal carbon intensity data: 20–40% operational carbon reduction beyond energy-optimal MPC, achieved by pre-cooling during low-carbon periods and shifting loads.
-
Manchester and Edinburgh scaling to district-level heating networks (DHN) under EPSRC Net Zero project funding 2025–2028; MPC horizon of 24–72 hours integrating building thermal models, weather forecasts, and grid carbon intensity signals.
-
Industrial process MPC explicitly optimising carbon footprint alongside economic cost: first academic formulations (2024) demonstrating 15–25% scope 1 emission reductions on ethylene cracker test cases.
Risks, Limitations, and Failure Modes
- Model-based control carries distinctive failure modes that must be actively managed in deployed systems:
- Model mismatch / reality gap: If the dynamics model f(x,u) deviates significantly from the true plant (e.g., unmodelled joint flexibility, slipping contacts, payload changes), MPC plans become suboptimal or infeasible. Residual learning, online SysID, and Gaussian-process model augmentation are standard mitigations; tube MPC provides formal robustness bounds up to ‖Δf‖ ≤ ε.
- Computational latency: NMPC solve times of 10–100 ms introduce feedback delays that can destabilise fast dynamics (contact transients at 1–10 ms timescales). Solutions: RTI-SQP single-iteration approximation; pre-committed first-step application; computation-aware MPC that accounts for solve-time uncertainty in the prediction model.
- Horizon truncation and terminal cost design: MPC optimises over a finite horizon N steps; stability and performance depend critically on the terminal cost V_f(x_N) accurately approximating the infinite-horizon value function. Incorrect terminal cost design is a leading source of MPC instability — evidenced by feasibility loss or oscillatory behaviour beyond the horizon.
- Local optimality of DDP/iLQR: DDP backward passes converge to locally optimal solutions; highly non-convex landscapes (multi-modal contact dynamics, narrow passages) may trap the optimizer in poor local minima. Multiple random restarts, warm-starting from prior solutions, and hybrid sampling+DDP approaches (MJPC CEM+iLQR) are standard mitigations.
- Contact model singularities: Rigid-body contact mechanics produce impulsive forces at contact events (coefficient of restitution) creating discontinuities in the dynamics. MPC handling discrete contact mode transitions requires either enumeration (exponential in horizon) or relaxation (complementarity smoothing, which introduces artificial dynamics).
- Constraint infeasibility: Overconstrained WBC QPs (too many task-space objectives relative to actuator DoF) may become infeasible under strict equality constraints; hierarchical task resolution (task space prioritisation via null-space projection) or soft constraint relaxation (penalised slack variables) is required.
- Sensor noise amplification: Model inversion in computed torque control amplifies sensor noise — high-gain feedforward based on noisy q̈ estimates produces chattering torques. Filtering (Butterworth, Kalman) is required; filtered derivatives introduce additional phase lag.
Tooling Ecosystem and Software Libraries
-
The MBC software ecosystem has matured considerably between 2015 and 2026, converging on a small number of dominant open-source libraries:
Physics Simulation Engines
-
MuJoCo 3.x (DeepMind, open-source 2022): fast rigid-body simulation with analytic gradients, implicit integration for stiff systems, and GPU backend (MJX, 2023); standard simulation environment for MBC research and RL training. 100,000+ GitHub stars; adopted by OpenAI, Berkeley, ETH, Oxford, Edinburgh.
-
Drake (MIT/TRI, open-source): comprehensive robot modelling and control framework with symbolic computation, optimization interfaces (OSQP, SNOPT, IPOPT wrappers), and Lyapunov analysis tools; used by Toyota Research Institute and Google for manipulation and autonomous driving research.
-
PyBullet / Bullet3: lighter-weight simulator with Python interface; widely used for RL training; less accurate for contact dynamics than MuJoCo but faster for simple scenes.
-
Isaac Gym / Isaac Lab (NVIDIA, 2021–2024): GPU-native simulation using PhysX, enabling 16,384+ parallel environments for RL and sampling-based MPC; dominant for massively parallel policy training.
-
Brax (Google, JAX-native 2021): differentiable physics engine in JAX; fully GPU/TPU accelerated; 1,000× faster than PyBullet for parallel rollouts; used for evolutionary algorithms and gradient-based MPC.
Robot Dynamics Libraries
-
Pinocchio 2.x / 3.x (LAAS-CNRS / Edinburgh, open-source): rigorous rigid-body dynamics (RNEA, CRBA, ABA) with analytical derivatives via CasADi and code generation; used within Crocoddyl, mc_rtc, and OCS2. C++ core with Python bindings; 15 ms evaluation for 36-DoF humanoid.
-
RBDL (Martin Felis, open-source): lighter-weight C++ dynamics library; popular for embedded deployment and teaching.
MPC and Trajectory Optimisation Frameworks
-
acados (SYSCOP Freiburg / KU Leuven): production-grade NMPC framework with real-time iterations (RTI-SQP), HPIPM backend, C code generation for embedded deployment; Python/MATLAB interfaces; deployed in commercial UAVs and Formula Student race cars.
-
Crocoddyl (LAAS-CNRS / Edinburgh): efficient DDP library with Pinocchio derivatives; supports multi-contact dynamics, actuation constraints, and abstract cost functions; primary framework for humanoid whole-body MPC in European academic labs.
-
OCS2 (ETH RSL): whole-body MPC for legged robots; ROS interface; supports SLQ (Sequential Linear Quadratic) and DDP backends; deployed on ANYmal-C and ANYmal-D.
-
CasADi (KU Leuven, open-source): symbolic framework for algorithmic differentiation and NLP formulation; widely used as an optimisation modelling language within MPC frameworks; interfaces to IPOPT, SNOPT, OSQP, and HPIPM.
-
FORCES Pro (Embotech): commercial embedded solver for QP/NLP from model structures; generates C code optimised for target hardware (automotive ECUs, ARM Cortex-A); <1 ms QP solve on Renesas R-Car H3.
State Estimation and Perception Tools
-
ORBSLAM3 (Campos et al. 2021): visual, visual-inertial, and RGB-D SLAM; real-time 6-DoF pose estimation at 25–30 Hz; integrates with robot state estimators for MPC state feedback.
-
Kimera-VIO / Kimera-Semantics (MIT SPARK Lab): visual-inertial odometry + 3D semantic map; provides state for terrain-aware MPC in field robots.
-
contact_estimation packages (ETH RSL, Oxford ORI): contact state classifiers from motor currents and IMU signals; required inputs to WBC QP contact Jacobian computation.
-
Research and Literature
- Bellman, R. (1957). Dynamic Programming. Princeton University Press. — Foundational DP framework underlying all receding-horizon methods.
- Pontryagin, L.S. et al. (1962). The Mathematical Theory of Optimal Processes. Interscience. — Maximum principle; necessary conditions for optimality in continuous-time OCP.
- Kalman, R.E. (1960). “A new approach to linear filtering and prediction problems.” Journal of Basic Engineering, 82(1), 35–45. — LQR/LQG foundations enabling model-based state feedback.
- Richalet, J., Rault, A., Testud, J.L., & Papon, J. (1978). “Model predictive heuristic control: Applications to industrial processes.” Automatica, 14(5), 413–428. — First MPC industrial deployment (MAC algorithm).
- Mayne, D.Q., Rawlings, J.B., Rao, C.V., & Scokaert, P.O.M. (2000). “Constrained model predictive control: Stability and optimality.” Automatica, 36(6), 789–814. — Rigorous stability proof for constrained MPC; 10,000+ citations.
- Rawlings, J.B., Mayne, D.Q., & Diehl, M. (2017). Model Predictive Control: Theory, Computation, and Design (2nd ed.). Nob Hill Publishing. — Comprehensive graduate textbook; standard reference.
- Mayne, D.Q., Seron, M.M., & Raković, S.V. (2005). “Robust model predictive control of constrained linear systems with bounded disturbances.” Automatica, 41(2), 219–224. — Tube MPC foundational paper.
- Tassa, Y., Erez, T., & Todorov, E. (2012). “Synthesis and stabilization of complex behaviors through online trajectory optimisation.” IEEE/RSJ IROS, 4906–4913. — iLQR/DDP real-time motion synthesis in MuJoCo; seminal robotics DDP paper.
- Tassa, Y., Mansard, N., & Todorov, E. (2014). “Control-limited differential dynamic programming.” IEEE ICRA, 1168–1175. — DDP with control constraints via backward-pass QP at each node.
- Todorov, E. & Li, W. (2004). “A generalized iterative LQG method for locally optimal feedback control of constrained nonlinear stochastic systems.” ACC, 300–306. — iLQR derivation and Gaussian noise extension.
- Di Carlo, J., Wensing, P.M., Katz, B., Bledt, G., & Kim, S. (2018). “Dynamic locomotion in the MIT cheetah 3 through convex model-predictive control.” IEEE/RSJ IROS, 1–9. — Convex MPC for quadruped at 500 Hz; basis for Boston Dynamics Spot controller.
- Bellicoso, C.D. et al. (2019). “Advances in real-world applications for legged robots.” Journal of Field Robotics, 36(7), 1311–1326. — ANYmal WBC-MPC stack in field deployments.
- Fahmi, S., Mastalli, C., Focchi, M., & Semini, C. (2020). “Passive whole-body control for quadruped robots: Experimental validation over challenging terrain.” IEEE Robotics & Automation Letters, 4(3), 2553–2560. — IIT HyQ terrain-adaptive MPC.
- Howell, T.A., Jackson, B.E., & Manchester, Z. (2019). “ALTRO: A fast solver for constrained trajectory optimization.” IEEE/RSJ IROS, 7674–7679. — AL-iLQR for general inequality constraints.
- Howell, T.A. et al. (2022). “Predictive sampling: Real-time behaviour synthesis with MuJoCo.” arXiv:2212.09051. — MJPC open-source release; CEM + iLQR asynchronous architecture.
- Mastalli, C. et al. (2020). “Crocoddyl: An efficient and versatile framework for multi-contact optimal control.” IEEE ICRA, 2536–2542. — Crocoddyl DDP library with Pinocchio analytical derivatives.
- Farshidian, F., Jelavic, E., Satapathy, A., Giftthaler, M., & Buchli, J. (2017). “Real-time motion planning of legged robots: A model predictive control approach.” IEEE-RAS Humanoids, 577–584. — OCS2 framework basis.
- Nagabandi, A., Kahn, G., Fearing, R.S., & Levine, S. (2018). “Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning.” IEEE ICRA, 7559–7566. — Neural dynamics model within MPC loop; MBRL survey anchor.
- Chua, K., Calandra, R., McAllister, R., & Levine, S. (2018). “Deep reinforcement learning in a handful of trials using probabilistic dynamics models.” NeurIPS, 4754–4765. — PETS: probabilistic ensemble + trajectory sampling; 100× sample efficiency vs model-free.
- Janner, M., Fu, J., Zhang, M., & Levine, S. (2019). “When to trust your model: Model-based policy optimization.” NeurIPS, 12519–12530. — MBPO: synthetic rollouts from learned model augment model-free SAC.
- Hafner, D., Lillicrap, T., Norouzi, M., & Ba, J. (2023). “Mastering diverse domains through world models.” arXiv:2301.04104. — DreamerV3: world-model MPC achieving superhuman performance on 150+ tasks.
- Hansen, N., Wang, X., & Su, H. (2023). “TD-MPC2: Scalable, robust world models for continuous control.” arXiv:2310.16828. — TD-MPC2: latent-space MPC competitive with model-free at 10× sample efficiency.
- Williams, G. et al. (2018). “Information theoretic MPC for model-based reinforcement learning.” IEEE ICRA, 1714–1721. — MPPI: sampling-based MPC using importance-weighted path integrals.
- Stellato, B. et al. (2020). “OSQP: An operator splitting solver for quadratic programs.” Mathematical Programming Computation, 12, 637–672. — OSQP: industry-standard embedded QP solver for linear MPC.
- Verschueren, R. et al. (2021). “acados — a modular open-source framework for fast embedded optimal control.” Mathematical Programming Computation. — acados: NMPC framework with real-time SQP; deployed in commercial UAVs and race cars.
- Hewing, L., Wabersich, K.P., Menner, M., & Zeilinger, M.N. (2020). “Learning-based model predictive control: Toward safe learning in control.” Annual Review of Control, Robotics, and Autonomous Systems, 3, 269–296. — Comprehensive survey of GP-MPC and safe learning for control.
- Bruder, D., Fu, X., Gillespie, R.B., Remy, C.D., & Vasudevan, R. (2021). “Data-driven control of soft robots using Koopman operator theory.” IEEE Transactions on Robotics, 37(3), 948–961. — Koopman MPC for soft robots; linear QP replacing nonlinear OCP.
- Siciliano, B., Sciavicco, L., Villani, L., & Oriolo, G. (2010). Robotics: Modelling, Planning and Control. Springer. — Standard robotics dynamics and model-based control textbook.
- Tedrake, R. (2024). Underactuated Robotics: Algorithms for Walking, Running, Swimming, Flying, and Manipulation. MIT Press (online edition). — Open graduate textbook; DDP, LQR, MPC for underactuated systems.
Metadata
- Term ID: RB-9018
- Domain: robotics
- Ontological domain confirmed:
robotics— correct; no domain correction required - IRI: http://narrativegoldmine.com/robotics#ModelBasedControl
- Quality Score: 0.52
- Authority Score: 0.87
- Version: 2.1.0
- Phase 6 Enrichment: Sonnet 4.6 worker; enrichment date 2026-05-17
- Research basis: Canonical textbooks (Rawlings, Siciliano, Tedrake), IEEE ICRA/IROS proceedings 2012–2024, NeurIPS/ICLR 2018–2023, arXiv preprints 2022–2025
- UK Context: Oxford Robotics Institute, Edinburgh Robotarium/ECR, Imperial Personal Robotics Lab, Manchester Control Systems, Cambridge MIL, BAE Systems, Rolls-Royce, AMRC Sheffield
- Domain corrections: None required —
roboticsdomain is correct for Model Based Control - Wikilinks: 71 cross-concept links covering control theory, robotics, optimisation, learning, and tools
- OWL axioms: 42 axioms across 5 families (Compositional 8, Dependency 9, Capability 10, Implementation 10, Reduction 5)
- Provenance references: 27 canonical references spanning 1957–2024
- Key algorithms covered: MPC, iLQR, DDP, WBC, contact-implicit MPC, tube MPC, GP-MPC, Koopman MPC, MPPI, CEM, MBPO, DreamerV3, TD-MPC2
- Key software: MuJoCo/MJPC, Drake, Pinocchio, Crocoddyl, OCS2, acados, OSQP, HPIPM, CasADi, Isaac Gym/MJX
Provenance
- Rawlings, J.B., Mayne, D.Q., & Diehl, M. (2017). Model Predictive Control: Theory, Computation, and Design (2nd ed.). Nob Hill Publishing.
- Siciliano, B., Sciavicco, L., Villani, L., & Oriolo, G. (2010). Robotics: Modelling, Planning and Control. Springer.
- Tedrake, R. (2024). Underactuated Robotics (online edition). MIT.
- Mayne, D.Q., Rawlings, J.B., Rao, C.V., & Scokaert, P.O.M. (2000). Constrained MPC: Stability and optimality. Automatica, 36(6), 789–814.
- Tassa, Y., Erez, T., & Todorov, E. (2012). Synthesis and stabilization of complex behaviors through online trajectory optimisation. IEEE/RSJ IROS.
- Tassa, Y., Mansard, N., & Todorov, E. (2014). Control-limited DDP. IEEE ICRA.
- Di Carlo, J. et al. (2018). Dynamic locomotion in the MIT cheetah 3 through convex MPC. IEEE/RSJ IROS.
- Bellicoso, C.D. et al. (2019). Advances in real-world applications for legged robots. Journal of Field Robotics.
- Fahmi, S. et al. (2020). Passive whole-body control for quadruped robots. IEEE RA-L.
- Howell, T.A. et al. (2022). Predictive sampling: Real-time behaviour synthesis with MuJoCo. arXiv:2212.09051.
- Mastalli, C. et al. (2020). Crocoddyl: An efficient and versatile framework for multi-contact optimal control. IEEE ICRA.
- Nagabandi, A. et al. (2018). Neural network dynamics for model-based deep RL. IEEE ICRA.
- Chua, K. et al. (2018). PETS: Deep RL in a handful of trials. NeurIPS.
- Janner, M. et al. (2019). MBPO: When to trust your model. NeurIPS.
- Hafner, D. et al. (2023). DreamerV3: Mastering diverse domains through world models. arXiv:2301.04104.
- Hansen, N. et al. (2023). TD-MPC2: Scalable, robust world models. arXiv:2310.16828.
- Williams, G. et al. (2018). MPPI: Information theoretic MPC. IEEE ICRA.
- Stellato, B. et al. (2020). OSQP. Mathematical Programming Computation, 12, 637–672.
- Verschueren, R. et al. (2021). acados. Mathematical Programming Computation.
- Hewing, L. et al. (2020). Learning-based MPC: Safe learning in control. Annual Review of Control, Robotics, AS, 3, 269–296.
- Bruder, D. et al. (2021). Koopman operator theory for soft robots. IEEE Transactions on Robotics, 37(3), 948–961.
- Mayne, D.Q. et al. (2005). Robust MPC with bounded disturbances. Automatica, 41(2), 219–224.
- Howell, T.A. et al. (2019). ALTRO: Fast solver for constrained trajectory optimisation. IEEE/RSJ IROS.
- Farshidian, F. et al. (2017). Real-time motion planning of legged robots: MPC approach. IEEE-RAS Humanoids.
- Mesbah, A. (2016). Stochastic model predictive control. IEEE Control Systems Magazine, 36(6), 30–44.
- Richalet, J. et al. (1978). Model predictive heuristic control: MAC. Automatica, 14(5), 413–428.
- Bellman, R. (1957). Dynamic Programming. Princeton University Press.