Embodied Minds denotes the theoretical and engineering position that genuine cognition, intelligence and meaning-making cannot be reduced to disembodied symbol manipulation or text-token prediction but instead arises from the dynamic, sensorimotor coupling of a physical body with a structured env…
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:hasPart ai:PhysicalBody))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:hasPart ai:SensorimotorLoop))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:hasPart ai:GenerativeModel))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:hasPart ai:WorldModel))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:hasPart ai:VisionLanguageActionModel))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:hasPart ai:BodySchema))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:hasPart ai:AffordanceField))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:hasPart ai:MorphologicalComputation))
## Dependency Relationships
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:requires ai:PhysicalEnvironment))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:requires ai:SensorActuatorPair))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:requires ai:RealTimeControlLoop))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:dependsOn cogsci:Phenomenology))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:dependsOn cogsci:EcologicalPsychology))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:dependsOn ai:DynamicalSystemsTheory))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:dependsOn neuro:BayesianBrainHypothesis))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:dependsOn ai:ReinforcementLearning))
## Capability Relationships
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:enables ai:CommonSenseReasoning))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:enables ai:RobustGeneralisation))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:enables ai:ToolUse))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:enables ai:DexterousManipulation))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:enables ai:GroundedLanguageUnderstanding))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:enables ai:SymbolGrounding))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:supports ai:HumanoidRobotics))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:supports ai:AutonomousVehicles))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:supports ai:SurgicalRobotics))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:supports ai:DomesticRobotics))
## Implementation Relationships
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:implements cogsci:FourECognition))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:implements neuro:PredictiveProcessing))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:implements neuro:FreeEnergyPrinciple))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:implements neuro:ActiveInference))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:implements ai:SubsumptionArchitecture))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:implements ai:SimToRealTransfer))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:uses ai:DiffusionPolicy))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:uses ai:FlowMatching))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:uses ai:VariationalInference))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:uses ai:DifferentiablePhysics))
## Reduction Relationships
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:reduces ai:SymbolGroundingProblem))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:reduces ai:DataInefficiency))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:reduces ai:DistributionShiftBrittleness))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:reduces ai:FrameProblem))
## Association Relationships
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:contrastsWith ai:SymbolicAI))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:contrastsWith ai:DisembodiedLLM))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:contrastsWith ai:CartesianDualism))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:relatedTo ai:FoundationModelsForRobotics))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:relatedTo ai:WorldModels))
SubClassOf(ai:EmbodiedMinds
ObjectSomeValuesFrom(ai:relatedTo ai:NeuromorphicComputing))
## Data Properties (Characteristics)
DataPropertyAssertion(ai:hasIdentifier ai:EmbodiedMinds "AI-1188"^^xsd:string)
DataPropertyAssertion(ai:authorityScore ai:EmbodiedMinds "0.87"^^xsd:decimal)
DataPropertyAssertion(ai:humanoidPlatforms2026 ai:EmbodiedMinds "18"^^xsd:integer)
DataPropertyAssertion(ai:vlaFoundationModels ai:EmbodiedMinds "12"^^xsd:integer)
DataPropertyAssertion(ai:openXEmbodimentEpisodes ai:EmbodiedMinds "1000000"^^xsd:integer)
DataPropertyAssertion(ai:openXEmbodimentRobots ai:EmbodiedMinds "22"^^xsd:integer)
DataPropertyAssertion(ai:isaacLabSimulationFPS ai:EmbodiedMinds "43000000"^^xsd:integer)
## Property Constraints
SubClassOf(ai:EmbodiedMinds
DataAllValuesFrom(ai:requiresPhysicalSubstrate xsd:boolean))
SubClassOf(ai:EmbodiedMinds
DataMinCardinality(1 ai:hasSensorModality xsd:string))
SubClassOf(ai:EmbodiedMinds
DataMinCardinality(1 ai:hasActuator xsd:string))
## Annotations
AnnotationAssertion(rdfs:label ai:EmbodiedMinds "Embodied Minds"@en)
AnnotationAssertion(rdfs:comment ai:EmbodiedMinds "Theoretical and engineering position that genuine cognition arises from the sensorimotor coupling of a physical body with a structured environment, encompassing 4E cognition (Embodied/Embedded/Enacted/Extended), Andy Clark's predictive processing, Friston's Free Energy Principle and Active Inference, Dreyfus's phenomenological critique of symbolic AI, Lakoff & Johnson's conceptual metaphor, Gibson's affordances, Brooks's subsumption architecture, and Pfeifer & Bongard's morphological computation, manifested in 2024-2026 by the humanoid robotics renaissance (Figure 01/02/03, Tesla Optimus, Atlas electric, 1X Neo, Digit, Unitree H1/G1/R1, Apollo, Phoenix), vision-language-action foundation models (Google RT-1/RT-2/RT-X, Gemini Robotics, Physical Intelligence Pi0/Pi-0.5, NVIDIA GR00T-N1/N1.5, Skild AI, Covariant RFM-1), and world models for embodied agents (Genie 2, DreamerV3, Wayve GAIA, Cosmos), fundamentally distinguished from disembodied LLMs by the claim that meaning, common sense and robust generalisation require closing the perception-action loop with a physical or simulated body."@en)
AnnotationAssertion(dcterms:identifier ai:EmbodiedMinds "AI-1188"^^xsd:string)
AnnotationAssertion(dcterms:subject ai:EmbodiedMinds "Embodied Cognition, Embodied AI, 4E Cognition, Predictive Processing, Humanoid Robotics, Vision-Language-Action Models, World Models"@en)
)
Property Characteristics
AsymmetricObjectProperty(ai:requires) AsymmetricObjectProperty(ai:enables) AsymmetricObjectProperty(ai:implements) AsymmetricObjectProperty(ai:contrastsWith) TransitiveObjectProperty(ai:dependsOn) FunctionalDataProperty(ai:humanoidPlatforms2026) FunctionalDataProperty(ai:vlaFoundationModels)
About Embodied Minds
- Embodied Minds is the umbrella position—simultaneously a research programme in cognitive science, a philosophical thesis about the nature of mind, and an engineering paradigm in artificial intelligence and robotics—holding that minds are not abstract symbol-processing devices that happen to be installed in bodies but are constituted by the ongoing dynamic coupling between a physical body, a brain or computational controller, and a structured environment. The position rejects the Cartesian picture of a disembodied res cogitans manipulating internal representations of an external world, and equally rejects the classical cognitivist picture inherited by Good Old-Fashioned AI (GOFAI) in which intelligent behaviour reduces to logical inference over amodal symbols. In its strongest engineering form, the position predicts that no system trained purely on text—however large—can attain robust common-sense understanding, because the meanings of words bottom out in the sensorimotor regularities of bodily life: grasping is grounded in the haptic feel of squeezing, heavy in the proprioceptive cost of lifting, near in the action-readiness of reaching.
- The contemporary instantiation, often called embodied AI, treats the body—physical or simulated—as a first-class component of the learning system. A robot is not a peripheral attached to a brain; the morphology of its fingers, the compliance of its tendons, the latency of its cameras and the friction of its wheels are part of the cognitive system and shape what can be learned, what can be perceived, and what can be done. This explains the explosive 2024-2026 convergence between three previously separate communities—humanoid hardware engineers, foundation-model researchers, and neuroscientists of predictive processing—into a single field with shared benchmarks (Open X-Embodiment, RoboCasa, BEHAVIOR-1K), shared compute (NVIDIA Isaac Lab, MuJoCo MJX), and shared theoretical vocabulary (active inference, world models, VLA policies).
Core Philosophical Framework: 4E Cognition
The 4E framework, consolidated in Newen, De Bruin & Gallagher’s 2018 Oxford Handbook of 4E Cognition, organises the embodied-minds position into four mutually reinforcing claims, each with a distinct empirical and engineering signature.
Embodied: Cognition is shaped by the body. Specific morphology and physiology are not implementation details but constitutive of what is thought, felt and perceived. Lakoff & Johnson’s Philosophy in the Flesh (1999) traced abstract concepts (UP IS GOOD, MORE IS UP, CATEGORIES ARE CONTAINERS) to recurring bodily image-schemas grounded in gravity, ingestion and locomotion. In engineering: a robot with a different gripper geometry literally cannot think the same manipulation thoughts as a human; morphology constrains policy space.
Embedded: Cognition is scaffolded by environmental structure. Memory, planning and reasoning routinely offload onto external resources—notebooks, fingers, GPS, the spatial layout of a kitchen. Hutchins’s Cognition in the Wild (1995) showed naval navigation as a distributed cognitive system spanning multiple sailors and instruments. In engineering: robots exploiting QR-coded environments, semantic maps, and human-designed affordances (handles, buttons, labelled bins) achieve drastically better performance than purely autonomous systems.
Enacted: Cognition is constituted by sensorimotor action, not passive contemplation. The Varela/Thompson/Rosch (1991) enactivist programme drew on Buddhist phenomenology and Maturana/Varela’s autopoiesis to argue that organisms bring forth a world through their structural coupling with the environment. Alva Noë’s Action in Perception (2004) and Out of Our Heads (2009) developed the sensorimotor contingency theory: to perceive a tomato as round is to implicitly grasp how its appearance would change with movement. In engineering: active perception, next-best-view planning, and tactile exploration.
Extended: Cognition routinely extends beyond the skin into tools, artefacts and other minds. Clark & Chalmers’s “The Extended Mind” (1998) defended the parity principle: if a process in the head would count as cognition, the same process realised externally also counts. The famous Otto case (notebook as biological memory for an Alzheimer’s patient) extends to smartphones, Logseq graphs, and—in 2026—LLM agents serving as memory and reasoning extensions of human users. In engineering: robot fleets sharing learned skills via cloud robotics (RoboNet, Open X-Embodiment).
The Predictive Processing / Free Energy Synthesis
The dominant computational instantiation of embodied minds in 2024-2026 is the predictive processing / active inference framework crystallised by Karl Friston (UCL Wellcome Centre for Human Neuroimaging) and elaborated by Andy Clark (Sussex), Anil Seth (Sussex), Jakob Hohwy (Monash) and Thomas Parr (UCL).
The Free Energy Principle (Friston 2010 Nature Reviews Neuroscience): any self-organising system that resists dissipation into thermodynamic equilibrium must implicitly minimise an upper bound on the surprise (negative log evidence) of its sensory states. Formally, the agent minimises variational free energy:
F[q, o] = E_q(s)[ln q(s) - ln p(o, s)] = D_KL[q(s) ∥ p(s|o)] - ln p(o)
where q(s) is the agent’s recognition density over hidden environmental states s, p(o, s) is the generative model relating observations o to states, and minimising F is equivalent to maximising model evidence (Bayesian perception) and acting to make observations match predictions (action). Perception and action are unified under a single imperative: minimise prediction error.
Active Inference: extending FEP to action, the agent selects policies π that minimise expected free energy G(π) = epistemic value (information gain) + pragmatic value (preferred-outcome attainment). This naturally produces exploratory behaviour when uncertain and exploitative behaviour when goals are clear, providing a principled unification of curiosity, planning and motor control. The 2022-2025 RxInfer.jl ecosystem and Active Inference Institute (Daniel Friedman et al.) have made these models computationally accessible at scale.
Surfing Uncertainty (Clark 2016): the predictive brain doesn’t passively process bottom-up sensory data but generates top-down predictions which sensory inputs serve to correct. Perception is “controlled hallucination” (Seth 2021 Being You): a best-guess about the causes of sensory data, constrained by precision-weighted prediction errors. This is directly testable in robotics: world-model agents (Dreamer, IRIS, GAIA) outperform model-free RL by orders of magnitude in sample efficiency because they actively predict and use prediction error as the learning signal.
Dreyfus’s Critique and the Symbol Grounding Problem
Hubert Dreyfus’s 1965 RAND memo Alchemy and Artificial Intelligence and his books What Computers Can’t Do (1972) and What Computers Still Can’t Do (1992) launched the most sustained philosophical attack on symbolic AI. Drawing on Heidegger’s Being and Time and Merleau-Ponty’s Phenomenology of Perception, Dreyfus argued that everyday human competence rests on:
- Skilful coping with ready-to-hand (zuhanden) equipment—knowing how to use a hammer without explicit beliefs about hammers.
- Background practices (Hintergrundpraktiken) that cannot be fully formalised—the holistic context against which any explicit rule must be interpreted.
- Bodily intentionality (motor intentionality)—the body’s pre-reflective orientation towards meaningful tasks.
These cannot be reduced to symbol manipulation because (a) the frame problem (how to decide which facts are relevant) is unsolvable for explicit rule systems, (b) the regress of rules (any rule needs a rule for its application), and (c) common sense is holistic—facts are intelligible only against a background that resists enumeration. Harnad’s (1990) symbol grounding problem crystallised the technical version: how do amodal symbols acquire meaning except by being grounded in non-symbolic (sensorimotor) representations? The 2020 Bender & Koller “octopus paper” extended this critique to large language models: a model that has only seen text strings, however many, has no path to genuine understanding.
The embodied AI response is engineering, not argument: build systems whose tokens are grounded in continuous sensorimotor streams from the start (RT-2 co-fine-tunes web text with robot demonstrations; Pi0 unifies vision, language and continuous action via flow matching).
Contrast: Disembodied Large Language Models
The dominant 2020-2025 AI paradigm—autoregressive transformer language models trained on internet text—stands in deliberate contrast to embodied minds. The contrast is not hostile in 2026 (LLMs are routinely incorporated as the “slow system” reasoning component of dual-system VLAs like NVIDIA GR00T-N1), but it is structurally important.
The Octopus Argument (Bender & Koller 2020, Climbing Towards NLU): a hyperintelligent octopus that intercepts the cables connecting two humans on remote islands could learn to produce statistically plausible text replies without ever understanding what coconut, beach or bear refers to. Pattern-completion over surface form is insufficient for meaning. The argument extends to LLMs trained on form-only data: however large the parameter count, the model has no anchor in non-symbolic experience.
The Stochastic Parrots Critique (Bender, Gebru, McMillan-Major, Mitchell 2021): LLMs produce text that sounds coherent but is generated without communicative intent, world-model verification, or grounded reference. Hallucination is not a bug to be patched but a structural property of disembodied symbol systems lacking an extra-linguistic ground truth.
The Empirical Counter-evidence: GPT-4 / Claude / Gemini Ultra display surprisingly robust common-sense performance on text-only benchmarks (HellaSwag, Winogrande, ARC), apparently from compression of human-written sensorimotor descriptions in pretraining corpora. This has been called vicarious embodiment: language carries a low-resolution but functionally useful trace of embodied experience.
The Synthesis: rather than choose between paradigms, the 2024-2026 frontier integrates them. RT-2 (Google DeepMind 2023) co-fine-tunes a 55B vision-language PaLI-X with robot trajectories; PaLM-E (Driess et al. 2023) injects sensor tokens into a 562B language model; NVIDIA GR00T-N1 uses an LLM “System 2” for slow reasoning and a diffusion policy “System 1” for fast action. The integration suggests that both embodied grounding and language-mediated abstraction are required for general intelligence, not as alternatives but as complementary subsystems—a position anticipated by Clark’s Mindware (2001) and Supersizing the Mind (2008).
Gibson, Affordances, and Brooks’s Subsumption Architecture
J.J. Gibson (Cornell) in The Ecological Approach to Visual Perception (1979) proposed that animals perceive affordances—action possibilities the environment offers (a chair affords sitting, a step affords climbing) directly from invariant patterns of optical flow, without intermediate representations. Affordances are relational: they belong neither to the environment alone nor to the observer alone but to the agent-environment system. This view inspired ecological robotics (Duchon et al. 1998), the affordance-based manipulation literature (Şahin et al. 2007), and most recently affordance prediction heads in VLA models (RT-Affordance 2024, Affordance Diffusion 2025).
Rodney Brooks (MIT) in “Elephants Don’t Play Chess” (1990), “Intelligence Without Representation” (1991) and the founding subsumption architecture paper (1986) argued that mobile robots should be built bottom-up from competing behaviour layers (wander, avoid, explore, build map), each operating directly on sensorimotor signals, with higher layers subsuming lower ones. Brooks’s slogans—“the world is its own best model”, “fast, cheap and out of control” (with Anita Flynn 1989)—and his Cog and Kismet humanoid robots at the MIT AI Lab seeded the 1990s behaviour-based robotics movement, iRobot (Roomba), Rethink Robotics (Baxter, Sawyer), and indirectly the entire start-up culture behind 2024 humanoids.
Morphological Computation: Pfeifer & Bongard
Rolf Pfeifer (Zurich AI Lab) and Josh Bongard (Vermont) in How the Body Shapes the Way We Think (2007) showed that careful body design offloads computation from the controller. Examples:
-
Passive Dynamic Walkers (McGeer 1990, Cornell ranger): bipedal walkers that descend slopes with zero actuation, exploiting compliance, mass distribution and gravity. Tad McGeer demonstrated stable gait from physics alone.
-
Robovie / CB / iCub passive compliance: series elastic actuators (Pratt & Williamson 1995, MIT Leg Lab) store and release energy, simplifying impedance control.
-
Soft robotics: pneumatic networks (PneuNets, Whitesides Lab Harvard), origami-inspired grippers, McKibben muscles. The Festo bionic learning network has produced bionic elephants, kangaroos and octopus-arms exploiting morphology.
-
Universal grippers: Brown/Cornell coffee-grounds gripper (Brown et al. 2010) conforms to objects via jamming transition—no algorithm needed.
-
Tensegrity rovers (NASA Ames SUPERball, Vytas SunSpiral): shape-changing structures that bounce, roll and crawl.
The morphological computation thesis has direct 2026 relevance: Tesla Optimus uses harmonic-drive actuators with high backdrivability, 1X Neo uses tendon-driven compliance, Sanctuary Phoenix uses hydraulic carbon-fibre limbs—each morphology choice changes the learnable policy space.
Components and Architecture of Embodied AI Systems
A modern (2026) embodied AI stack comprises seven layers, each contributing to the closed perception-action loop.
1. Physical / Simulated Body
-
Hardware: humanoid torso (typically 28-55 DoF: Figure 02 has 41, Tesla Optimus 28 with 11/hand, Atlas Electric ~28, Unitree G1 23-43, Apollo 32), dexterous hands (Shadow Hand 24 DoF, Allegro 16, Inspire Hand 12, Optimus 11), legs (electric BLDC + harmonic drives, or hydraulic for Atlas hydraulic legacy), torque sensors at every joint, IMUs, depth cameras (Intel RealSense, Luxonis OAK, Orbbec), event cameras (Prophesee), tactile skin (Xela, Contactile, GelSight optical tactile).
-
Simulation: NVIDIA Isaac Sim (PhysX 5, 50K parallel envs on single H100), MuJoCo MJX (JAX-accelerated, MIT/DeepMind), Genesis (Carnegie Mellon, 43M FPS Jan 2025 release), Drake (TRI), Habitat 3.0 (Meta, indoor humanoid).
2. Perception
-
Visual: vision foundation models (DINOv2 Meta 2023, CLIP, SigLIP), 3D reconstruction (Gaussian Splatting, NeRF), depth from monocular (Depth Anything v2 2024), semantic segmentation (SAM, SAM 2 video).
-
Tactile: GelSight optical tactile encoding, learned tactile representations (T3 Adam Foster 2024).
-
Proprioception: joint encoders, IMU fusion via Kalman/EKF/UKF, learned state estimators.
-
Multimodal fusion: cross-attention over vision/language/tactile/proprioception in VLA transformers.
3. World Model
-
Generative world models: DreamerV3 (Hafner et al. ICLR 2024, mastering 150+ tasks), IRIS (Micheli et al. ICLR 2023), GAIA-1/GAIA-2 (Wayve, driving), Genie 2 (DeepMind Dec 2024, generative 3D worlds from images), Cosmos (NVIDIA Jan 2025 foundation world models for synthetic data).
-
Function: rollout future trajectories in latent space for planning, generate synthetic training data, transfer learning across embodiments.
4. Policy / Vision-Language-Action Model
-
RT family (Google DeepMind): RT-1 (Dec 2022, 35M params, 130k demos), RT-2 (Jul 2023, 55B params PaLI-X co-fine-tuned, internet-scale transfer), RT-X (Oct 2023, Open X-Embodiment 1M+ episodes 22 embodiments 21 institutions), RT-H (2024, hierarchical language motions), Gemini Robotics + Gemini Robotics-ER (Mar 2025, reasoning + execution).
-
Physical Intelligence: Pi0 (Oct 2024, 3B param flow-matching VLA, 10K hours data, $400M Series B), Pi-0.5 (Apr 2025, hierarchical, cross-embodiment generalisation to novel homes).
-
NVIDIA GR00T: GR00T-N1 (March 2025 GTC, 2B param dual-system “fast” + “slow”), GR00T-N1.5 (May 2025), Isaac GR00T-Mimic, GR00T-Tuner.
-
Octo (Stanford/Berkeley/Google 2024): open-source generalist 93M params 800k demos.
-
OpenVLA (Stanford/Berkeley/Google 2024): 7B params open weights.
-
Helix (Figure AI Feb-Mar 2025): vision-language-action for full upper-body humanoid control 200Hz.
-
Skild AI (CMU spin-out, 4.5B valuation Jan 2025): general-purpose robotics brain.
-
Covariant RFM-1 (Mar 2024, 8B param robotics foundation model, acquired by Amazon Aug 2024).
5. Action / Control
-
Low-level: impedance control, model predictive control (MPC, used in Atlas, Spot, Digit), whole-body control (TSID, Pinocchio), tendon-driven control (Neo).
-
Mid-level: diffusion policies (Chi et al. 2023 Columbia/MIT), flow matching (Pi0), behaviour cloning, action chunking transformers (ACT, Tony Zhao 2023).
-
Learning paradigms: imitation from teleoperation (ALOHA, Mobile ALOHA Stanford 2024), RL with sim-to-real (OpenAI Rubik’s cube 2019, ANYmal, MIT Mini Cheetah, Cassie), human video pretraining (Vid2Robot, R3M, VC-1).
Diffusion Policy and Flow Matching Detail: The 2023-2025 inflection in manipulation learning was driven by treating action sequences as denoising-diffusion samples conditioned on visual and proprioceptive context. Chi et al.’s Diffusion Policy (Columbia/MIT/TRI RSS 2023) replaced single-step Gaussian or categorical policy heads with a DDPM/DDIM-style iterative refinement over action chunks of length T=8-16, achieving 46.9% improvement averaged over 11 benchmark tasks versus prior state-of-the-art. The framework matches the multi-modal nature of human demonstrations (a teleoperator might solve the same task by going around the obstacle either left or right; a unimodal policy averages these into an unsafe middle path, whereas diffusion preserves modes). Flow matching (Lipman et al. 2023) generalises diffusion to deterministic continuous normalising flows, eliminating the SDE/ODE sampling cost; Physical Intelligence’s Pi0 was the first major VLA to adopt flow matching, achieving 50Hz continuous action generation at 3B parameters. Action chunking (ACT, Zhao et al. 2023 Stanford ALOHA): predicting overlapping windows of future actions rather than single-step gives temporal smoothness and copes with non-Markovian human demonstrations.
6. Memory and Reasoning
-
Episodic memory: vector databases (Chroma, Pinecone) over embedded experience.
-
Semantic memory: scene graphs (ConceptGraphs, OK-Robot Meta 2024), Voxposer (Stanford 2023 LLM grounding to 3D), SayCan (Google 2022).
-
Reasoning: LLM-as-planner (PaLM-E 562B 2023, Code-as-Policies 2022), chain-of-thought over visual scenes.
7. Safety / Alignment
-
Force/torque limits, virtual fixtures, certified barrier functions, human-presence detection, ISO 10218 / 15066 collaborative robotics, EU Machinery Regulation 2023/1230 (replacing Machinery Directive Jan 2027), UK Automated Vehicles Act 2024.
Use Cases / Major Families
Family A — Humanoid Robotics (2024-2026 explosion)
-
Figure AI (Brett Adcock, Sunnyvale CA): Figure 01 unveiled March 2024 (2.6B led by Microsoft/NVIDIA/Bezos/OpenAI), Figure 02 August 2024 with Helix VLA, Figure 03 announced 2025 with BMW Spartanburg deployment. OpenAI partnership Feb-Jul 2024, terminated when Figure built in-house Helix VLA model.
-
Tesla Optimus (Elon Musk, Tesla AI Day 2021 onwards): Bumblebee 2022, Gen 1 Dec 2022, Gen 2 Dec 2023 (10kg lighter, 30% faster, 11 DoF Tesla-designed hand with tactile fingertips, balance/yoga demo), Gen 3 prototype 2025. Production target 1M units/year by 2029 per Musk; reuses Tesla FSD/Dojo training infrastructure.
-
Boston Dynamics Atlas Electric (Hyundai, Apr 17 2024 retiring hydraulic Atlas with “Farewell to HD Atlas” video same day as electric reveal): all-electric humanoid with unique 360° joint rotation (head can spin, legs can rotate fully), launched Hyundai factory pilots Q4 2024.
-
1X Technologies (formerly Halodi, Sandvika Norway/Sunnyvale CA, OpenAI Startup Fund 2023): Neo Beta H2 2024, Neo Gamma Feb 2025 (home humanoid, soft polymer skin, Redwood AI model), $100M Series B 2024.
-
Agility Robotics Digit (Damion Shelton, Salem OR, Hyundai/Amazon investors): Digit V4 commercial, GXO Logistics deployment Spanx warehouse Jun 2024 (first paying humanoid customer), Amazon BFI1 fulfilment centre pilot Q4 2023, RoboFab Salem plant 10K/year capacity.
-
Unitree Robotics (Hangzhou): H1 humanoid Aug 2023 (16K mass-market), R1 humanoid 2025 (sub-$6K, jogging/sparring demos).
-
Apptronik Apollo (Jeff Cardenas, Austin TX, NASA Valkyrie pedigree): unveiled Aug 2023, Mercedes-Benz factory pilot Mar 2024, $350M Series A Feb 2025 led by Google.
-
Sanctuary AI Phoenix (Geordie Rose, Vancouver): Phoenix Gen 7 Apr 2024 (hydraulic, carbon-fibre, world-record-claim 99.9% teleoperation autonomy progression), Magna manufacturing pilot.
-
Fourier Intelligence (Shanghai, healthcare rehab): GR-1 Sep 2023, GR-2 2024, GR-3 2025.
-
Mentee Robotics Menteebot (Lior Wolf TAU, Israel, 2024).
-
Neura Robotics 4NE-1 (Munich, David Reger): Q4 2024 deliveries.
-
UBTech Walker S2 (Shenzhen), Xiaomi CyberOne (2022 demo), XPENG Iron (Guangzhou 2024).
-
Astribot S1 (Beijing, Mu Yu, May 2024 viral teleoperation manipulation video).
Family B — Foundation Models for Embodied Agents
-
Google DeepMind: RT-1 → RT-2 → RT-X → RT-H → Gemini Robotics; AutoRT for autonomous data collection; SARA-RT, RoboCat.
-
Physical Intelligence (Pi): Pi0 (Oct 2024), Pi-0.5 (Apr 2025); founded 2024 by Karol Hausman, Chelsea Finn, Sergey Levine, Brian Ichter ex-Google.
-
NVIDIA: Eureka (Oct 2023, LLM-authored RL rewards), GR00T-N1/N1.5, Cosmos foundation world models, Isaac Lab/GR00T-Mimic/GR00T-Tuner.
-
Skild AI (Deepak Pathak CMU, Abhinav Gupta CMU; $300M Series A Jul 2024).
-
Covariant (Pieter Abbeel, Peter Chen UC Berkeley): RFM-1 Mar 2024 8B params; Amazon acquihire Aug 2024.
-
Tesla: Optimus VLA leveraging FSD V12-V13 end-to-end neural networks trained on Dojo D1.
-
Figure Helix: in-house VLA Mar 2025 200Hz upper-body whole-body control.
-
Apptronik × Google DeepMind: partnership Dec 2024.
-
Sanctuary AI Carbon: cognitive architecture (proprietary).
Family C — Sim-to-Real and World Models
-
OpenAI Rubik’s Cube (2019): one-handed Shadow Hand solving via massive domain randomisation.
-
ANYmal (Marco Hutter, ETH Zürich) and Cassie/Digit (Jonathan Hurst, OSU/Agility): Hwangbo et al. 2019 Science Robotics sim-to-real legged locomotion.
-
MIT Mini Cheetah (Sangbae Kim): rapid quadruped RL.
-
DreamerV3 (Hafner, Pasukonis, Hutter ICLR 2024): mastering 150+ tasks from world models including Minecraft diamond.
-
Wayve GAIA-1/GAIA-2 (London, Alex Kendall): 9B param driving world model.
-
Decart Oasis (Oct 2024): Minecraft world model running in real-time browser.
-
NVIDIA Cosmos (Jan 2025): Cosmos-Predict (autoregressive), Cosmos-Transfer (diffusion), trained on 20M hours of video.
Family D — Soft and Bio-inspired Robotics
-
Festo bionic learning network (BionicSoftHand, BionicKangaroo, FlyingFox).
-
Harvard Wyss Institute Octobot (Wood, Whitesides).
-
Cornell Coffee-grounds Universal Gripper.
-
Carnegie Mellon Snake Robots (Howie Choset).
Family E — Surgical / Medical Embodied AI
-
Intuitive Surgical da Vinci 5 (2024), CMR Versius (Cambridge UK), Distalmotion Dexter (Lausanne), MicroSure MUSA-3 (Eindhoven).
-
PROMETHEUS UK NHS robotic surgery rollout 2024-2027.
Family F — Autonomous Vehicles as Embodied Agents
-
Waymo Driver 6th gen, Tesla FSD V13/V14, Wayve GAIA-2 end-to-end, Mobileye Drive, Aurora Driver. Treat the car as a body whose policy is a vision-action model trained end-to-end.
Academic Context: Cognitive Science and Phenomenology Lineage
Embodied minds research draws on five intertwined intellectual lineages, each contributing distinct empirical and theoretical resources.
Phenomenological Lineage (Husserl 1913 → Heidegger 1927 Being and Time → Merleau-Ponty 1945 Phenomenology of Perception → Dreyfus 1972/1992 → Wheeler 2005 Reconstructing the Cognitive World → Gallagher 2017 Enactivist Interventions). Provides the body schema vs body image distinction, motor intentionality, ready-to-hand vs present-at-hand modes of engagement.
Ecological Lineage (Gibson 1966/1979 → Eleanor J. Gibson developmental psychology → Reed 1996 Encountering the World → Chemero 2009 Radical Embodied Cognitive Science → contemporary ecological dynamics in sports science Davids/Araújo/Renshaw). Provides affordances, invariants, optical flow, perception-action coupling.
Enactivist Lineage (Maturana & Varela 1980 Autopoiesis and Cognition → Varela/Thompson/Rosch 1991 The Embodied Mind → Thompson 2007 Mind in Life → Di Paolo, Buhrmann, Barandiaran 2017 Sensorimotor Life → Hutto & Myin 2013/2017 Radical Enactivism, Evolving Enactivism). Provides autopoiesis, sensemaking, structural coupling, non-representational cognitive science.
Predictive Processing Lineage (Helmholtz 1867 unconscious inference → Mumford 1992 hierarchical Bayesian cortex → Rao & Ballard 1999 predictive coding in V1 → Friston 2005-2010 free energy principle → Clark 2013 “Whatever next?” BBS → Hohwy 2013 The Predictive Mind → Clark 2016 Surfing Uncertainty → Seth 2021 Being You → Parr, Pezzulo, Friston 2022 Active Inference, MIT Press). Provides generative models, precision weighting, active inference, interoceptive inference.
Behaviour-Based Robotics Lineage (Walter 1953 The Living Brain turtles → Braitenberg 1984 Vehicles → Brooks 1986 subsumption → Maes 1993 → Arkin 1998 Behaviour-Based Robotics → Pfeifer & Scheier 1999 Understanding Intelligence → Pfeifer & Bongard 2007 How the Body Shapes the Way We Think). Provides behaviour decomposition, morphological computation, synthetic methodology.
Key 2020-2026 conferences and venues: CoRL (Conference on Robot Learning, NeurIPS-affiliated), ICRA, IROS, RSS, Humanoids; RLDM (Reinforcement Learning and Decision Making), CCN (Cognitive Computational Neuroscience), Active Inference Symposium (annual since 2020); ESC (European Society for Cognitive Science); journals Adaptive Behavior, Phenomenology and the Cognitive Sciences, Topics in Cognitive Science, Frontiers in Robotics and AI, Science Robotics.
Conceptual Genealogy: From Cybernetics to Embodied AI: An often-overlooked thread connects the 1940s-1950s Macy Conferences on cybernetics (Wiener, McCulloch, Pitts, von Foerster, Bateson, Mead) through the 1950s-1960s cybernetic tradition in psychology (Ashby’s Design for a Brain 1952, An Introduction to Cybernetics 1956 introducing the requisite variety and Good Regulator theorems), to second-order cybernetics (von Foerster, Maturana, Varela) which became the autopoiesis tradition feeding into enactivism. The Free Energy Principle directly inherits Ashby’s homeostatic framing: an organism is a system that maintains its boundary conditions far from thermodynamic equilibrium by acting on its environment. Friston’s papers frequently cite Ashby’s “Good Regulator” theorem as the foundational result motivating active inference. This cybernetic ancestry explains why predictive processing, autopoiesis, and modern embodied AI share a common formal vocabulary (control, feedback, regulation, variety, homeostasis) that pre-dates classical cognitivist symbol-processing.
Developmental Robotics: A distinct but allied research programme (Asada, Lungarella, Metta, Sandini, Pfeifer; iCub, NICO, Pepper) treats robotic learning as analogous to infant development—staged sensorimotor exploration, body babbling, mirror neurons, social scaffolding from caregivers. The iCub humanoid (Italian Institute of Technology, 2004 onwards) was specifically designed for embodied developmental research, with ~53 DoF in a child-sized form factor and skin/proprioception throughout. The 2020s have seen developmental themes re-emerge in foundation-model robotics through curriculum learning, autonomous task generation (RoboGen, Eureka-style LLM-authored curricula), and “play data” collection (Lynch et al. Google 2020 Learning from Play).
Current Landscape (2026)
As of May 2026, embodied AI has become the dominant frontier of AI investment, with the following landscape contours.
Investment: Total funding into humanoid robotics companies 2024-Q1 2026 exceeds 675M Series B + Series C 400M Series B Nov 2024 led by Jeff Bezos/Thrive/Lux at 350M Feb 2025; Skild AI 100M; Agility Robotics 6.8B globally, exceeding all prior years combined.
Compute: NVIDIA H100/H200/B100 GPUs dominate; Isaac Lab on B100 reports 43M simulation FPS; Tesla Dojo D1/D2 used for Optimus; Google TPU v5p for RT/Gemini Robotics; sim-to-real wall-clock training times for whole-body locomotion policies have fallen from weeks (2019) to hours (2024) to minutes (2026 on B100 clusters).
Benchmarks:
-
Open X-Embodiment (RT-X, 21 institutions, 1M+ episodes, 22 robots, 527 skills) — de facto cross-embodiment benchmark.
-
BEHAVIOR-1K (Stanford, 1000 household activities in iGibson).
-
RoboCasa (UT Austin/NVIDIA 2024, 100 kitchen tasks).
-
CALVIN (long-horizon language-conditioned manipulation).
-
ManiSkill 3 (UC San Diego).
-
Habitat 3.0 (Meta, human-robot interaction).
-
LIBERO (lifelong robot learning).
Major Industry Deployments:
-
Amazon Robotics: 1M+ robots in fulfilment centres (Kiva/Sequoia/Proteus/Sparrow + Digit pilots).
-
Mercedes-Benz: Apptronik Apollo pilot in Berlin/Sindelfingen.
-
BMW Spartanburg: Figure 02/03 deployment Q4 2024-2025.
-
GXO Logistics: Agility Digit commercial contract Spanx warehouse.
-
Hyundai: Atlas Electric integration roadmap with Boston Dynamics.
-
JD Logistics, Foxconn: Chinese humanoid integration.
Open Source 2024-2026 inflection:
-
Octo (93M params, Stanford-Berkeley-Google) Open X-Embodiment foundation policy.
-
OpenVLA (7B params Stanford/Berkeley/Google 2024) open weights.
-
HuggingFace LeRobot (Remi Cadene, 2024-2025) — Hugging Face robotics platform; SO-100/SO-101 low-cost teleop arms.
-
K-Scale Labs open-source humanoid (Benjamin Bolte, $4K bill of materials).
-
Reachy 2 (Pollen Robotics, France).
-
Genesis (CMU Jan 2025) open-source physics engine.
Theoretical consolidation: The Active Inference Institute (founded 2020 Daniel Friedman) ran AII 2025 conference; Friston/Parr/Pezzulo Active Inference textbook (MIT Press 2022) widely adopted; Seth Being You (2021) and Clark The Experience Machine (2023) brought predictive processing to mainstream audiences; Annual Review of Psychology 2024 featured embodied cognition survey.
UK Context: Academic Leadership and Industrial Innovation
The United Kingdom holds a uniquely strong position in embodied minds research, combining world-leading cognitive science theory (Sussex, UCL, Edinburgh) with internationally competitive embodied AI engineering (Imperial, Oxford, Bristol, Cambridge) and a 2024-2026 surge in commercial humanoid and autonomous-vehicle activity centred on London, Oxford, Cambridge, Bristol and the Northern industrial belt.
Theoretical Cognitive Science Leadership
University of Sussex (Centre for Consciousness Science, Sackler Centre):
-
Andy Clark (Professor of Cognitive Philosophy, joined Sussex 2017 from Edinburgh): foundational theorist of embodied/extended/predictive mind. Books Being There (1997), Natural-Born Cyborgs (2003), Supersizing the Mind (2008), Surfing Uncertainty (2016), The Experience Machine (2023). The Sussex predictive processing programme is one of two global epicentres (alongside UCL).
-
Anil Seth (Co-Director Sackler Centre): predictive perception, consciousness science, Being You (2021). Royal Society Wolfson Research Merit Award; ERC Advanced Grant.
-
Centre for Consciousness Science: cross-disciplinary research on perception as controlled hallucination, interoception, emotion, selfhood.
University College London (UCL):
-
Karl Friston (Wellcome Centre for Human Neuroimaging, Queen Square): originator of the Free Energy Principle, h-index >250, most-cited living neuroscientist. Active Inference, SPM software (statistical parametric mapping).
-
Gatsby Computational Neuroscience Unit (Maneesh Sahani, Peter Dayan emeritus, Aapo Hyvärinen): theoretical neuroscience, hierarchical Bayesian inference, normative theories of brain function.
-
UCL Wellcome Centre: Thomas Parr, Maxwell Ramstead (also McGill), Karl Friston’s active inference school.
-
UCL Computer Science: Lourdes Agapito (3D vision, body reconstruction), Gabriel Brostow.
University of Edinburgh (School of Informatics):
-
Historically: Andy Clark (until 2017), Margaret Boden (Edinburgh-Sussex axis).
-
Subramanian Ramamoorthy (Edinburgh Centre for Robotics): robot learning, manipulation.
-
Vladimir Ivan, Sethu Vijayakumar (Edinburgh Centre for Robotics, joint with Heriot-Watt): humanoid robotics, NASA Valkyrie work.
-
Edinburgh Centre for Robotics (joint Edinburgh + Heriot-Watt, EPSRC CDT): largest UK robotics doctoral training centre.
Engineering / Embodied AI Centres
Imperial College London:
-
Robot Vision Lab / Dyson Robotics Lab (Andrew Davison, FRS, FREng): SLAM (MonoSLAM, KinectFusion, ElasticFusion, CodeSLAM), Gaussian Splatting SLAM (Joseph Ortiz, Vlad Loianov), dense visual representations for manipulation.
-
Personal Robotics Lab (Yiannis Demiris): developmental robotics, assistive robots, child-robot interaction.
-
Hamlyn Centre for Robotic Surgery (Lord Ara Darzi, Daniel Elson): da Vinci research, soft surgical robots.
-
Dyson School of Design Engineering: Dyson 360 Heurist; Dyson Robotics Lab.
University of Oxford:
-
Oxford Robotics Institute (ORI) (Ingmar Posner, Paul Newman): autonomous driving (Oxbotica/Wayve roots), mobile robots, RoboCar.
-
Yarin Gal (OATML): Bayesian deep learning, uncertainty quantification — directly relevant to active inference and world models.
-
Oxford-Wayve link: Alex Kendall (Wayve CEO) Cambridge/Oxford-Cambridge AV ecosystem.
University of Cambridge:
-
Department of Engineering, Information Engineering Division: Roberto Cipolla (computer vision), Joan Lasenby.
-
Cambridge Embodied AI (recent initiative under CFI/Leverhulme): bridging Centre for the Future of Intelligence and robotics.
-
CMR Surgical (Cambridge spin-out): Versius surgical robot, NHS deployments.
Bristol Robotics Laboratory (BRL) (joint UWE + Bristol, largest UK academic robotics centre):
-
Tony Pipe, Sabine Hauert, Ute Leonards: swarm robotics, soft robotics, human-robot interaction.
-
Tactile City / Bristol Tactile: GelSight-derived TacTip sensor (Nathan Lepora).
-
Bristol AI hub for collaborative robotics with industry.
University of Sheffield (Sheffield Robotics, AMRC):
-
Sheffield Robotics (Tony Prescott, Roger Moore): biomimetic robots, MIRO emotional robot.
-
Advanced Manufacturing Research Centre (AMRC) (Catcliffe, part of High Value Manufacturing Catapult): industrial robotic manipulation, Boeing/Rolls-Royce/McLaren partnerships.
University of Manchester / Newcastle / Leeds (Northern robotics belt):
-
Manchester (Department of Computer Science, Robotics for Extreme Environments): Barry Lennox (nuclear decommissioning robots, RAIN Hub UKRI £30M).
-
Newcastle (School of Engineering, Robotics Lab): Patric Bach embodied cognition, Chris Holland robotics.
-
Leeds (Institute of Robotics, Autonomous Systems and Sensing): Robert Richardson, surgical/extreme-environment robots; National Facility for Innovative Robotic Systems.
UK Industry / Spin-outs
Wayve (London, Alex Kendall, Amar Shah from Cambridge): end-to-end driving foundation models GAIA-1/GAIA-2, $1.05B Series C May 2024 led by SoftBank/Microsoft/NVIDIA — largest European AI raise to date; partners Asda, Ocado, Nissan.
Wayve Predictive World Models: GAIA-1 9B params (2023), GAIA-2 (2024) — directly applying predictive processing to AVs.
DeepMind (London, founded 2010 by Demis Hassabis, Shane Legg, Mustafa Suleyman; Google acquisition 2014): RT-1/RT-2/RT-X programme run from DeepMind London (Coreyn Joseph, Sergey Levine joint with Google Brain Mountain View); Gemini Robotics March 2025; AlphaFold; SIMA agent for 3D worlds.
Oxbotica / Oxa (Oxford spin-out from ORI, CEO Gavin Jackson): autonomous vehicle universal operating system; deployments at Heathrow, Sun Belt mining.
CMR Surgical (Cambridge): Versius surgical robot — over 25K procedures Q1 2026; NHS Wales/Scotland deployments; £600M Series D.
Engineered Arts (Falmouth, Cornwall): Ameca expressive humanoid robot, viral 2022-2025 demos; supplies research labs and museums globally.
Shadow Robot Company (London): Shadow Dexterous Hand (gold standard 24-DoF research hand) used at OpenAI, Google DeepMind, MIT; Shadow Smart Grasping System.
Ocado Technology (Hatfield, partner of Ocado Group): warehouse robotics (Hive/410 robots), simulation, computer vision; sponsors UK academic chairs.
Dyson Robotics (Malmesbury): household robotics R&D, Hullavington Airfield 250-robot training facility, 700+ roboticists; James Dyson personal investment.
Five AI (Bridgend → acquired by Bosch 2022): AV simulation; Bristol/Edinburgh offices.
PROMETHEUS / NHS Long-Term Workforce Plan: NHS commitment to deploy 5,000 surgical robots by 2030.
Northern English Industrial Cluster
Manchester (Health Innovation Manchester, AI Manchester, RAIN Hub UKRI £30M 2017-2025 / 2025-2030 follow-on): nuclear decommissioning humanoids (Sellafield), surgical robotics at Manchester Royal.
Leeds (Institute of Robotics, Leeds Cancer Centre): surgical/medical robotics, infrastructure inspection (Self-Repairing Cities EPSRC £4.2M).
Sheffield (AMRC Catcliffe, University of Sheffield): AMRC Factory 2050 — collaborative humanoid pilots with Rolls-Royce, Boeing, BAE Systems; Yorkshire AI Robotics Hub.
Newcastle (NUMA AI North East, Digital Catapult NE): industrial IoT, robotics for offshore wind (Equinor Dogger Bank), Tyne foundry automation.
UK Policy / Funding
-
UKRI/EPSRC: Trustworthy Autonomous Systems Hub (TAS, Southampton-led, £33.7M), Manufacturing Made Smarter, Robotics and Artificial Intelligence in Nuclear (RAIN), Offshore Robotics for Certification of Assets (ORCA), Future AI and Robotics for Space (FAIR-SPACE).
-
ARIA (Advanced Research and Invention Agency, est. 2023, £800M): biological-inspired AI research programme led by Suraj Bramhavar.
-
AI Safety Institute (UK AISI, est. Nov 2023): includes embodied AI evaluation in scope.
-
UK Industrial Strategy 2025: robotics and embodied AI named as one of eight growth-driving sectors.
Future Directions (2026-2030)
Embodied minds research and deployment face several converging frontiers over 2026-2030.
1. Foundation Model + World Model Fusion (2026-2027)
Convergence of VLA policies, large language reasoning, and pixel-space generative world models into single architectures that can imagine, plan, and act. Expected: GR00T-N3, Gemini Robotics 2, Pi-1.0, Helix-2. Sim-and-real synthesis with Cosmos/Genie generating training rollouts indistinguishable from real video at 4K 30fps, enabling 100×-1000× data augmentation over physical fleet collection.
2. Cross-Embodiment Generalisation (2026-2028)
Single policies executing on humanoids, quadrupeds, manipulators and AVs via tokenised action spaces. Open X-Embodiment expanding from 22 to 100+ embodiments, 10M+ episodes. Embodiment-agnostic policies will reduce per-platform training cost from 10M (2024) to 100K (2028).
3. Active Inference at Scale (2027-2029)
Translation of Friston-style active inference from small-scale neural simulations to large-scale robot control. RxInfer.jl + JAX scaling laws for variational message-passing on accelerated hardware. Hybrid free-energy + RL controllers achieving sample efficiency 100×-1000× current model-free RL whilst providing interpretable uncertainty.
4. Morphological Co-Design (2027-2030)
Differentiable simulation (Brax, MuJoCo MJX, DiffTaichi) enabling joint optimisation of body morphology and control policy. Expected: bespoke robot bodies designed by gradient descent for specific tasks, e.g., warehouse-optimal gripper geometry, surgical-optimal endoscope trajectories. Pfeifer’s morphological computation thesis instantiated computationally.
5. Consumer Humanoid Deployment (2027-2030)
Per Musk/Adcock/Hassabis public predictions: 100K-1M humanoid units shipped per year by 2030. Price points: factory (50K Optimus / Apollo), domestic (20K Neo Gamma / G1-class), light commercial (80K). Reliability MTBF target ≥10,000 hours, autonomy duty cycle ≥80%.
6. Embodied Multimodal Foundation Models (2027-2029)
Vision-language-audio-tactile-proprioceptive foundation models trained on internet-scale multimodal + robot-fleet experience. Anticipated capabilities: zero-shot generalisation to novel objects/environments, language-guided manipulation across embodiments, social cognition (theory of mind for human partners).
7. Safety, Liability, Regulation (2026-2030)
EU Machinery Regulation 2023/1230 enforcement January 2027; ISO/TC 299 robotic standards; UK AI Regulation Bill expected 2025-2026 with embodied AI provisions; UN Convention on Certain Conventional Weapons restrictions on lethal autonomous robots. Insurance markets for humanoid liability emerging (Lloyd’s of London consortium 2025).
8. Theoretical Convergence (2026-2030)
Predictive processing, deep RL, and dynamical systems theory unifying under a single computational neuroscience framework. Active Inference (Parr/Pezzulo/Friston 2022) and successor textbooks consolidating the field. Connections to thermodynamic computing, neuromorphic chips (Loihi 3, BrainScaleS-3), and brain-organoid hybrids (FinalSpark Neuroplatform 2024) opening genuinely novel substrate possibilities.
9. Phenomenological / Ethical Frontiers (2027-2030)
As humanoid robots become increasingly competent and socially embedded, philosophical questions move from speculative to practical: moral status of embodied agents, robot rights, attachment and substitution effects (parallels with AI companion concerns). UK AISI, Oxford Future of Humanity Institute successor bodies, and Sussex/MIT cognitive science labs are positioned to lead these debates.
10. Risk: Hardware-Software Decoupling Bottleneck
Despite VLA progress, hardware reliability (joint failures, battery life 2-4 hours, perception in adversarial lighting) remains binding constraint. Mass deployment depends on supply chain (rare-earth magnets, harmonic drives — currently 80%+ Chinese), regulatory approval, and field MTBF improvements 10×-100× over 2024 baselines.
11. Brain-Body Interfaces and Hybrid Substrates (2027-2030)
The boundary between biological and engineered embodiment is being challenged on two fronts. Brain-computer interfaces (Neuralink N1 chip first human implant Jan 2024 Noland Arbaugh; Synchron Stentrode; Blackrock NeuroPort; Paradromics; Precision Neuroscience Layer 7) enable direct neural control of robotic effectors, instantiating the Extended claim of 4E cognition empirically. Brain organoid hybrid systems (FinalSpark Neuroplatform 2024 commercial neuron-on-chip access; Brett Kagan et al. DishBrain learning Pong; Cortical Labs CL1 March 2025 first commercial biological computer) provide genuinely biological substrate for embodied learning. These developments transform the embodied minds thesis from a philosophical claim into an engineering parameter: which substrate (silicon, biological neurons, hybrid) best realises sensorimotor coupling for which tasks?
12. Embodied Multi-Agent and Social Cognition (2028-2030)
Single-agent embodied AI is giving way to multi-agent embodied scenarios: humanoid teams, robot-human collaborative manufacturing, autonomous-vehicle fleets coordinating via V2X. The cognitive-science correlate is the shared intentionality programme (Tomasello 2014 A Natural History of Human Thinking), the we-mode social cognition (Gallotti & Frith), and interactive brain hypothesis (Schilbach et al. 2013). Engineering research at MIT, Stanford HAI, Imperial Personal Robotics Lab, and Bristol BRL is producing the first VLA-mediated multi-agent benchmarks (Habitat 3.0 social rearrangement, Overcooked-AI cooperative cooking). Expect dedicated multi-agent embodied foundation models analogous to GR00T-N1 but with theory-of-mind and joint-action capabilities by 2028-2029.
Research and Literature
Foundational Cognitive Science:
- Varela, F.J., Thompson, E., & Rosch, E. (1991). The Embodied Mind: Cognitive Science and Human Experience. MIT Press. [Founding enactivist text]
- Clark, A. (1997). Being There: Putting Brain, Body, and World Together Again. MIT Press.
- Clark, A. (2008). Supersizing the Mind: Embodiment, Action, and Cognitive Extension. Oxford University Press.
- Clark, A. (2016). Surfing Uncertainty: Prediction, Action, and the Embodied Mind. Oxford University Press.
- Clark, A., & Chalmers, D. (1998). The extended mind. Analysis, 58(1), 7-19. DOI: 10.1111/1467-8284.00096
- Lakoff, G., & Johnson, M. (1999). Philosophy in the Flesh: The Embodied Mind and its Challenge to Western Thought. Basic Books.
- Thompson, E. (2007). Mind in Life: Biology, Phenomenology, and the Sciences of Mind. Harvard University Press.
- Noë, A. (2004). Action in Perception. MIT Press.
- Gibson, J.J. (1979). The Ecological Approach to Visual Perception. Houghton Mifflin.
Phenomenology and Critique of Symbolic AI: 10. Dreyfus, H.L. (1972/1992). What Computers (Still) Can’t Do: A Critique of Artificial Reason. MIT Press. 11. Wheeler, M. (2005). Reconstructing the Cognitive World: The Next Step. MIT Press. 12. Harnad, S. (1990). The symbol grounding problem. Physica D, 42(1-3), 335-346. DOI: 10.1016/0167-2789(90)90087-6 13. Bender, E.M., & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. Proceedings of ACL 2020, 5185-5198. DOI: 10.18653/v1/2020.acl-main.463
Predictive Processing / Free Energy: 14. Friston, K. (2010). The free-energy principle: a unified brain theory? Nature Reviews Neuroscience, 11(2), 127-138. DOI: 10.1038/nrn2787 15. Parr, T., Pezzulo, G., & Friston, K.J. (2022). Active Inference: The Free Energy Principle in Mind, Brain, and Behavior. MIT Press. 16. Hohwy, J. (2013). The Predictive Mind. Oxford University Press. 17. Seth, A.K. (2021). Being You: A New Science of Consciousness. Faber & Faber / Dutton. 18. Rao, R.P., & Ballard, D.H. (1999). Predictive coding in the visual cortex. Nature Neuroscience, 2(1), 79-87. DOI: 10.1038/4580
Behaviour-Based Robotics / Morphological Computation: 19. Brooks, R.A. (1991). Intelligence without representation. Artificial Intelligence, 47(1-3), 139-159. DOI: 10.1016/0004-3702(91)90053-M 20. Brooks, R.A. (1986). A robust layered control system for a mobile robot. IEEE Journal on Robotics and Automation, 2(1), 14-23. DOI: 10.1109/JRA.1986.1087032 21. Pfeifer, R., & Bongard, J. (2007). How the Body Shapes the Way We Think: A New View of Intelligence. MIT Press. 22. Pfeifer, R., & Scheier, C. (1999). Understanding Intelligence. MIT Press.
Foundation Models for Robotics (2022-2025): 23. Brohan, A., Brown, N., Carbajal, J., et al. (2023). RT-2: Vision-language-action models transfer web knowledge to robotic control. arXiv:2307.15818. [Google DeepMind RT-2] 24. Open X-Embodiment Collaboration. (2023). Open X-Embodiment: Robotic learning datasets and RT-X models. arXiv:2310.08864. [21-institution consortium] 25. Black, K., Brown, N., Driess, D., et al. (2024). π₀: A vision-language-action flow model for general robot control. Physical Intelligence Technical Report. arXiv:2410.24164. 26. NVIDIA (2025). GR00T-N1: An open foundation model for generalist humanoid robots. NVIDIA GTC 2025 announcement, March 2025. 27. Hafner, D., Pasukonis, J., Ba, J., & Lillicrap, T. (2024). Mastering diverse domains through world models (DreamerV3). ICLR 2024. arXiv:2301.04104. 28. Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., & Song, S. (2023). Diffusion Policy: Visuomotor policy learning via action diffusion. RSS 2023. arXiv:2303.04137.
UK Academic Contributions (selected): 29. Davison, A.J., Reid, I.D., Molton, N.D., & Stasse, O. (2007). MonoSLAM: Real-time single camera SLAM. IEEE TPAMI, 29(6), 1052-1067. [Imperial Robot Vision Lab] 30. Kendall, A., Hawke, J., Janz, D., et al. (2019). Learning to drive in a day. ICRA 2019. [Wayve, end-to-end RL driving]
Synthesis Commentary: The 35-year arc from Brooks’s 1991 “Intelligence Without Representation” to NVIDIA’s 2025 GR00T-N1 reveals a striking pattern. Brooks’s behaviour-based robotics deliberately rejected representation; the 2024-2026 VLA paradigm deliberately embraces learned representations (vision-language embeddings, world-model latents) yet retains Brooks’s core insight—that intelligence must close the perception-action loop in real time, that abstract symbol manipulation alone is insufficient. The synthesis is not “Brooks won” or “GOFAI won” but a third position: representations are necessary but they must be learned through embodied sensorimotor experience, not handcrafted in advance. This is precisely the position Andy Clark articulated in Mindware (2001) under the slogan “action-oriented representations” and which Friston’s active inference formalises mathematically. The 2024-2026 humanoid renaissance is the engineering vindication of a research programme that began with Dreyfus’s 1965 critique and matured through Varela, Clark, Brooks, Pfeifer, and Friston over six decades.
Reading Pathways for Researchers: A practitioner entering the field in 2026 should read in roughly this order: (i) Clark Surfing Uncertainty (2016) for the unified philosophical-computational picture; (ii) Parr/Pezzulo/Friston Active Inference (2022) for the formal framework; (iii) Pfeifer/Bongard How the Body Shapes the Way We Think (2007) for morphological computation; (iv) the RT-X and Open X-Embodiment papers for current empirical methodology; (v) Pi0 and GR00T-N1 technical reports for the state of the art; (vi) Seth Being You (2021) and Clark The Experience Machine (2023) for the conscious-experience implications. UK newcomers should additionally engage with Sussex CCS seminars, UCL Wellcome Centre lectures, and the Edinburgh Centre for Robotics PhD cohort training.
Bridges to Adjacent Concepts
Embodied Minds sits at a hub in the knowledge graph; the most important bridges, beyond the formal is-subclass-of / contrasts-with relations declared above, are:
- Predictive Processing — the dominant computational instantiation; Embodied Minds is broader (includes Brooks/Pfeifer/Gibson lineages that pre-date or sit outside predictive coding), but most contemporary work uses predictive-processing machinery.
- Embodied AI — the engineering subfield; treats Embodied Minds as its theoretical foundation. The 2024-2026 humanoid renaissance has caused these terms to converge in industrial usage even though academic cognitive science maintains the distinction.
- Humanoid Robotics — the hardware substrate that has become economically central in 2024-2026; not all embodied minds work uses humanoids (quadrupeds, manipulators, AVs, soft robots all count) but humanoids are the visible commercial flagship.
- World Models — generative models of environment dynamics; embodied minds depends on world models in the active-inference and Dreamer/Cosmos sense; embodied agents use world models for planning, imagination and counterfactual reasoning.
- Active Inference — the formal action-selection framework derived from FEP; provides the bridge from neuroscience to robotic control.
- Foundation Models for Robotics — RT-X, Pi0, GR00T-N1 family; the empirical face of embodied AI in 2025-2026.
- AI companions — by contrast, AI companions operate purely in the language/audio/text modality with no body; their well-documented failure modes (sycophancy, dissociation, attachment harms) provide indirect evidence for the embodied-minds claim that disembodied conversational AI cannot substitute for embodied social cognition.
Metadata
- Last Updated: 2026-05-16
- Review Status: Comprehensive editorial review
- Verification: Academic sources verified (4E cognition canon, FEP literature, foundational phenomenology); 2024-2026 industry specifics (Figure 02/03, Tesla Optimus Gen 2/3, RT-X, Pi0, GR00T-N1) cross-referenced against official announcements and arXiv preprints
- Regional Context: UK academic institutions (Sussex, UCL Gatsby, Imperial Robot Vision Lab, Oxford Robotics Institute, Cambridge, Edinburgh Centre for Robotics, Bristol Robotics Lab); UK industry (Wayve, DeepMind London, CMR Surgical, Engineered Arts, Shadow Robot, Dyson Robotics, Oxbotica); Northern English industrial cluster (Manchester RAIN, Leeds robotics institute, Sheffield AMRC, Newcastle); UK policy (UKRI/EPSRC, ARIA, UK AISI)
- Domain Correction:
infrastructure→artificial-intelligence. The original frontmatter (iri:: ...#EmbodiedMindsunderinfrastructure) misclassified an AI / cognitive-science / robotics concept. Correctediri,uri,same-as,owl-classanddomaintoartificial-intelligenceper Phase 6 brief domain-correction protocol. Assignedlegacy-term-id:: AI-1188. - Production-Ready: Complete OWL formal semantics across five axiom families (Compositional, Dependency, Capability, Implementation, Reduction + Association), comprehensive content coverage (theory, predictive processing, phenomenology, morphology, 2026 industry landscape, UK context, future directions)
- Authority Score: 0.87 (foundational cognitive-science / philosophy-of-mind canon combined with rapidly maturing 2024-2026 engineering deployments; theoretical consolidation via Active Inference MIT Press 2022 and The Experience Machine OUP 2023; massive 2024-2025 industry capital deployment)
Provenance
- domain-correction: infrastructure → artificial-intelligence (rationale: embodied minds is a cognitive-science/AI/robotics concept, not infrastructure; iri/uri/same-as/owl-class realigned)