Task Planning is the computational sub-field of artificial intelligence and robotics concerned with automatically synthesising finite sequences of discrete actions — called plans — that transform a given initial world state into a state satisfying a specified goal condition, subject to action pre…

Classical symbolic planning represents world states as sets of logical ground atoms, actions as STRIPS operators with ADD/DELETE lists, and plans as sequences satisfying preconditions at each step. The landmark Planning Domain Definition Language (PDDL; McDermott et al. 1998 AIPS planning competition, revised PDDL 2.1 Fox & Long 2003 temporal, PDDL 3.0 Gerevini & Long 2005 trajectory constraints, PDDL+ 2006 hybrid mixed-discrete/continuous, PDDL 3.1 Kovacs 2011) standardised representation across research groups enabling systematic solver comparison. The Fast Downward planning system (Helmert 2006 JAIR) introduced the SAS⁺ multi-valued variable framework, causal graph analysis, and pattern database heuristics now embedded in near-universal use; its successor Scorpion (Seipp & Röger 2018) added cost-optimal planning with operator-counting heuristics. The competing LAMA planner (Richter & Westphal 2010) won the IPC 2008 satisficing track using additive and landmark heuristics in an anytime weighted-A* framework, establishing the current IPC (International Planning Competition) benchmark ecosystem running biennially since 1998 with 30+ participating systems by IPC 2023.

Hierarchical Task Network (HTN) planning decomposes abstract high-level tasks into subtask networks via decomposition methods, enabling rich domain knowledge encoding and polynomial tractability for many practical problems. SHOP (Nau et al. 1999) and its successor SHOP2 (Nau et al. 2003 JAIR, cited 2,400+) introduced an ordered-task-network formulation enabling fast forward-chaining depth-first search with strong soundness and completeness guarantees for totally-ordered networks; the current SHOP3 (Nau et al., University of Maryland, published 2021) extends SHOP2 with PDDL 3.1 compatibility, external calls to learning modules, and hybrid HTN/PDDL interleaving. Practical HTN systems include HDDL (Höller et al. 2020 AAAI) standardising hierarchical domain description and the pandaPI/PANDA planner family achieving competitive IPC-HTN 2020 results.

Behaviour Trees (BTs) represent reactive task execution structures as directed acyclic trees of control flow nodes (Sequence, Fallback/Selector, Parallel, Decorator) and leaf action/condition nodes, combining the modularity of finite state machines with execution reactivity and online tree policy switching. Originally developed for game AI (Halo 2 AI director circa 2004), BTs were formalised for robotics by Colledanchise & Ögren (2016 IROS, textbook 2018 IEEE Press) and now constitute the primary reactive execution layer in ROS 2 Navigation Stack (Nav2) via the BehaviorTree.CPP library (Faconti et al., v4.6 2024). BTs admit formal safety analysis (Biggar et al. 2020 IEEE RA-L) and learning-based extension (Grammatically-Guided GP-BTs, Jones & Lober 2022).

Task and Motion Planning (TAMP) integrates symbolic task planning with continuous-space geometric motion planning, solving the interface between high-level goal specification (e.g. "stack block A on block B in bin C") and low-level kinematic feasibility (collision-free trajectories, grasp reachability, stability physics). Seminal work by Kaelbling & Lozano-Pérez (2013 IJRR belief-space TAMP), Garrett, Lozano-Pérez & Kaelbling (PDDLStream 2020 IJRR, 680+ citations), and Chitnis, Hadfield-Menell et al. (2016 ICRA) established the modern TAMP problem stack. Practical TAMP systems include PDDLStream (MIT CSAIL, open-source), PRoBaBL (Probabilistic TAMP under uncertainty, Phiquepal & Toussaint 2019), and TMP (Task-Motion Planning for manipulation, Toussaint 2015 IJCAI). TAMP complexity is PSPACE-hard in general but tractable for bounded horizon via mixed-integer programming relaxations (Deits & Tedrake 2014 motion planning as MIQP).

LLM-based planning uses large language model capabilities for goal decomposition, plan synthesis, world-state interpretation, and code generation. LLM+P (Liu et al. 2023 arXiv:2304.11477) demonstrated that GPT-4 can translate natural-language problem descriptions into PDDL files suitable for classical planners, achieving 100% success on IPC benchmarks previously solved 0% by pure LLM prompting. Voyager (Wang et al. 2023 arXiv:2305.16291) deployed GPT-4 as a lifelong learning agent in Minecraft, using code-as-action planning with an automatic curriculum and skill library, acquiring 3× more unique items than prior state-of-the-art. AutoGen (Wu et al. Microsoft 2023 arXiv:2308.08155) extended multi-agent LLM conversation to collaborative planning workflows. ReAct (Yao et al. 2023 ICLR) interleaved chain-of-thought reasoning traces with environment actions (Thought→Act→Observe loops), substantially improving planning reliability on ALFWorld and WebArena benchmarks. SayCan (Ahn et al. Google Robotics 2022 arXiv:2204.01691) grounded LLM plan generation in robot affordance scores, enabling natural-language kitchen task specification on a Boston Dynamics Spot arm.

Foundation model visuomotor planners represent the frontier convergence of vision-language-action models trained end-to-end on large heterogeneous robot datasets. RT-2 (Brohan et al. Google DeepMind 2023 arXiv:2307.15818) fine-tuned PaLI-X and PaLM-E vision-language models to output tokenised robot actions, enabling emergent chain-of-thought reasoning for novel object manipulation. π0 (Black, Brown, Driess et al. Physical Intelligence, arXiv:2410.24164, published March 2024; Physical Intelligence founded 2023 by Sergey Levine, Chelsea Finn, Karol Hausman, Brian Ichter) introduced a flow-matching visuomotor diffusion architecture trained on a heterogeneous cross-embodiment dataset across 7 robot platforms (UR5, Franka, Stretch, xArm, Spot, ALOHA, custom dexterous hands), achieving state-of-the-art performance on the Open-X-Embodiment benchmark including dexterous bimanual tasks (laundry folding, table bussing, cardboard assembly). OpenVLA (Kim et al. Stanford/Berkeley arXiv:2406.09246, June 2024) released a 7B-parameter open-source vision-language-action model trained on the Open-X-Embodiment dataset outperforming RT-2-X on BridgeV2 grasping with 7× fewer parameters. MOLMO (Deitke et al. Allen Institute for AI, arXiv:2409.17146, September 2024) demonstrated multimodal reasoning for robotic pointing and spatial task specification competitive with proprietary models.

Goal-Oriented Action Planning (GOAP; Jeff Orkin, F.E.A.R. 2005, GDC 2006 "Three States and a Plan") applies STRIPS-like forward-chaining search at runtime in game AI and simulation agents, computing plans over small action sets (8-15 operators) in milliseconds through A* with a goal-satisfaction heuristic, enabling emergent adaptive agent behaviour without hand-authored FSMs. GOAP has been extended to robotics via learning-based action precondition/effect discovery (Konidaris et al. 2018 JAIR option-based GOAP) and grounded world-model estimation.

Monte Carlo Tree Search (MCTS) planning applies UCT (Upper Confidence bound applied to Trees; Kocsis & Szepesvári 2006 ECML) for online lookahead in stochastic planning problems, balancing exploration–exploitation through rollout policies. Adapted to language and embodied planning in RAP (Hao, Gu, Ma et al. 2023 EMNLP "Reasoning with Language Model is Planning with World Model"), AlphaCode (Li et al. DeepMind 2022 Science), and LATS (Zhou et al. 2023 arXiv:2310.04406 LLM+MCTS tree search for interactive planning). In robotic TAMP, MCTS-based planners handle non-deterministic action outcomes via belief-space Monte Carlo rollouts.

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:hasPart ai:PDDL))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:hasPart ai:HTNPlanning))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:hasPart ai:BehaviourTrees))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:hasPart ai:TaskAndMotionPlanning))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:hasPart ai:GoalOrientedActionPlanning))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:hasPart ai:MCTSPlanning))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:hasPart ai:WorldModel))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:hasPart ai:HeuristicFunction))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:hasPart ai:ActionSchema))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:hasPart ai:PlanLibrary))

## Dependency Relationships
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:requires ai:WorldStateRepresentation))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:requires ai:GoalSpecification))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:requires ai:ActionModel))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:requires ai:SearchAlgorithm))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:requires ai:FeasibilityChecker))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:dependsOn ai:SearchTheory))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:dependsOn ai:FormalLogic))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:dependsOn ai:ProbabilityTheory))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:dependsOn ai:ComputerVision))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:dependsOn ai:NaturalLanguageProcessing))

## Capability Relationships
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:enables ai:RobotAutonomy))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:enables ai:IntelligentAgent))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:enables ai:MultiRobotCoordination))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:enables ai:HumanRobotInteraction))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:enables ai:WarehouseAutomation))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:supports ai:AutonomousVehicles))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:supports ai:ManufacturingAutomation))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:supports ai:ServiceRobotics))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:supports ai:SpaceRobotics))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:supports ai:VideoGameAI))

## Implementation Relationships
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:implements ai:STRIPS))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:implements ai:PDDLFormalism))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:implements ai:HTNDecomposition))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:implements ai:BehaviourTreeExecution))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:implements ai:MCTSRollout))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:implements ai:ChainOfThoughtReasoning))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:implements ai:DiffusionPolicy))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:uses ai:LargeLanguageModels))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:uses ai:ReinforcementLearning))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:uses ai:GraphSearch))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:uses ai:SATSolving))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:uses ai:MonteCarloMethods))

## Reduction Relationships
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:reduces ai:ManualProgrammingBurden))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:reduces ai:TaskReprogrammingCost))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:reduces ai:HumanOperatorDependence))
SubClassOf(ai:TaskPlanning
  ObjectSomeValuesFrom(ai:reduces ai:PlanningComputationTime))

## Annotations
AnnotationAssertion(rdfs:label ai:TaskPlanning "Task Planning"@en)
AnnotationAssertion(rdfs:comment ai:TaskPlanning "Computational AI sub-field synthesising action sequences that achieve specified goals from initial states, encompassing classical symbolic methods (STRIPS/PDDL/Fast Downward), HTN hierarchical decomposition (SHOP3), Behaviour Trees, TAMP integration with motion planning, LLM-based planners (LLM+P, Voyager, ReAct, SayCan), foundation model visuomotor planners (RT-2, pi0 Physical Intelligence 2024, OpenVLA 2024, MOLMO 2024), GOAP for game AI, and MCTS-based deliberative search."@en)
AnnotationAssertion(dcterms:identifier ai:TaskPlanning "AI-1047"^^xsd:string)
AnnotationAssertion(dcterms:subject ai:TaskPlanning "Automated Planning, Robotics, Symbolic AI, Foundation Models, HTN, Behaviour Trees"@en)
DataPropertyAssertion(ai:hasIdentifier ai:TaskPlanning "AI-1047"^^xsd:string)
DataPropertyAssertion(ai:authorityScore ai:TaskPlanning "0.87"^^xsd:decimal)

About Task Planning

  • Task Planning is the branch of artificial intelligence concerned with automatically determining what to do next — and in what order — to achieve a goal. Unlike reactive control systems that map percepts directly to motor commands, task planners operate over an explicit symbolic or latent representation of the world, reason about the causal consequences of candidate actions, and synthesise plans that are guaranteed (or probabilistically likely) to achieve desired outcomes. The field integrates contributions from classical logic programming and search, formal verification, probabilistic reasoning, robotics kinematics, and — since 2022 — neural language and vision models capable of generalising across novel goal specifications expressed in natural language.
  • The core planning problem can be stated precisely: given a triple ⟨I, G, A⟩ where I is the initial state (a set of ground literals true at time 0), G is the goal (a conjunction of literals that must hold), and A is a finite set of parameterised action schemata each carrying a precondition PRE(a) and an effect EFF(a) partitioned into ADD and DELETE lists, find a sequence π = ⟨a₁, a₂, …, aₙ⟩ such that executing π from I produces a state satisfying G. This formulation — STRIPS (Stanford Research Institute Problem Solver; Fikes & Nilsson 1971) — remains the canonical baseline despite subsequent extensions for temporal durations (PDDL 2.1), numeric fluents (PDDL 2.1), derived predicates, and stochastic effects. The complexity of plan existence testing is PSPACE-complete for general propositional STRIPS (Bylander 1994) but falls to polynomial time for restricted structure classes (series-parallel causal graphs, unary operators, tree-like dependency graphs) enabling practical scalability.
  • Planning search strategies divide broadly into forward state-space search (expanding states reachable from I towards G, typically A* or greedy best-first with admissible heuristics), backward state-space search (regressing goal conditions through action preconditions towards I), and plan-space search (operating on partially-ordered plan structures adding ordering and binding constraints to resolve flaws, as in POP planners). Heuristic computation is the key scalability lever: the FF heuristic (Hoffmann & Nebel 2001) ignores delete lists to solve a relaxed problem in polynomial time, producing inadmissible but highly informative estimates; landmark heuristics (Richter, Helmert & Westphal 2008) identify logical milestones necessarily achieved en route to any goal; causal graph heuristics (Helmert 2004) decompose the planning graph by variable dependency structure. The Fast Downward system (Helmert 2006) systematically combines these heuristics in a multi-queue best-first search with a lazy evaluation scheme, representing the dominant open-source planning engine deployed in robotics middleware, IPC competitions, and as the back-end in LLM+P (Liu 2023).

Components and Architecture

  • STRIPS and PDDL (Classical Symbolic Core). PDDL (McDermott et al. 1998) standardised domain and problem files separating domain-level operator definitions from problem-instance specifications, enabling the IPC benchmark series and cross-system reproducibility. Key PDDL versions: 1.2 (1998, quantified preconditions, conditional effects), 2.1 (Fox & Long 2003, durative actions, numeric fluents, temporal planning), 2.2 (derived predicates, timed initial literals), 3.0 (Gerevini & Long 2005, trajectory constraints, preferences, soft goals), PDDL+ (2006, mixed discrete-continuous processes), 3.1 (Kovacs 2011, object fluents). The Fast Downward planner (Helmert 2006, Seipp et al. 2022 v22.12) implements SAS⁺ translation, merge-and-shrink abstractions, and portfolio solver selection, used in production at Siemens industrial automation, Willow Garage PR2 task planning, and as default TAMP back-end in PyBullet/PDDLStream stacks.
  • Hierarchical Task Networks (HTN). HTN planning imposes a hierarchical decomposition structure: compound tasks are reduced to primitive (executable) tasks via decomposition methods defined in the domain. SHOP (Nau et al. 1999) and SHOP2 (2003) achieve completeness for totally-ordered networks with complexity polynomial in domain size for fixed task depth. SHOP3 (University of Maryland 2021) adds PDDL 3.1 compatibility and external calls enabling learning-based method selection. The HDDL language (Höller et al. AAAI 2020) standardises hierarchical domains enabling the IPC-HTN track (2020, 2023). Practical deployments include NASA JPL’s MEXEC executive for rover task sequences, DARPA Urban Challenge vehicle mission execution, and hospital logistics automation (McGuire et al. KI 2021).
  • Behaviour Trees (BT). BTs provide a reactive, modular execution framework composing Sequence (AND), Fallback/Selector (OR), Parallel, and Decorator control nodes with Action and Condition leaves ticked at each execution cycle. The BehaviorTree.CPP library (Faconti, v4.6.2 2024) is the standard ROS 2 Nav2 integration for mobile robot navigation task execution. Formal properties include: contractual compositional safety (Biggar et al. IEEE RA-L 2020, 130+ citations), learning-based BT synthesis from demonstrations (Rovida et al. 2017, Jones & Lober GP-BT 2022), and automated BT generation from PDDL plans (Cai et al. 2021 IROS). Game industry applications span Halo 2-5 AI directors, The Sims social behaviour, Alien: Isolation SEGA Creative Assembly 2014.
  • Task and Motion Planning (TAMP). TAMP integrates symbolic task planning with geometric motion planning, solving the combined problem of what to do (task level: which objects to pick, stack, pour, assemble) and how to do it (motion level: collision-free trajectories, stable grasps, reachability within joint limits). The PDDLStream system (Garrett, Lozano-Pérez & Kaelbling, MIT CSAIL, IJRR 2020, 680+ citations) wraps motion planning subroutines as certified stream generators called lazily by the task planner, avoiding the combinatorial explosion of enumerating geometric facts upfront. PRoBaBL (Phiquepal & Toussaint 2019) handles uncertain geometric parameters via belief-space planning. TMP (Toussaint 2015 IJCAI) treats motion planning subproblems as continuous optimisation within a task-level commitment framework. TAMP is deployed in Fetch Robotics warehouse pick-and-place, MIT Manipulation Lab bimanual assembly, and Toyota Research Institute kitchen manipulation.
  • LLM-Based Planners. The 2023 emergence of instruction-following LLMs capable of structured output enabled a new planning paradigm: translating natural-language task descriptions into formal PDDL or code representations for execution by classical planners or simulators. LLM+P (Liu, Zhao, Jain et al. UT Austin, arXiv:2304.11477 Apr 2023) uses GPT-4 to generate PDDL problem files from natural language, feeding them to Fast Downward, achieving 100% IPC benchmark success vs 0% for pure LLM planning. Voyager (Wang, Xie, Jiang et al. NVIDIA, arXiv:2305.16291 May 2023) deploys GPT-4 as a lifelong Minecraft agent using code-as-action (JavaScript Mineflayer API) with automatic curriculum and a growing skill library, outperforming prior RL baselines (DEPS, GITM) by 3× unique items and 15.3× skill diversity. ReAct (Yao, Zhao, Yu et al. Princeton/Google, ICLR 2023) structures LLM prompts as interleaved Thought (chain-of-thought reasoning)/Action (environment interaction)/Observation triples, improving ALFWorld planning accuracy from 21% (CoT-only) to 71% (ReAct). SayCan (Ahn, Brohan, Brown et al. Google Robotics, arXiv:2204.01691 Apr 2022) scores LLM-generated plan steps by robot affordance functions (learned from offline RL) to ground planning in physical feasibility, executing 101/152 (66%) natural language kitchen tasks on a real Spot robot.
  • Foundation Model Visuomotor Planners. RT-2 (Brohan, Brown, Carbajal et al. Google DeepMind, arXiv:2307.15818 Jul 2023) co-fine-tunes PaLI-X (55B) and PaLM-E vision-language models on Web image–text data combined with robot action tokens, enabling emergent chain-of-thought reasoning for object manipulation with novel objects (49%→62% improvement over RT-1 on generalisation tasks). π0 (Black, Brown, Driess, Escontrela et al. Physical Intelligence, arXiv:2410.24164, Physical Intelligence blog post and paper March 2024) introduces a flow-matching architecture combining a pre-trained VLM backbone (OpenFlamingo-style) with a lightweight diffusion-based action expert, trained jointly on cross-embodiment data from 7 robot platforms via a mixture-of-experts routing scheme, achieving state-of-the-art on dexterous bimanual manipulation (84% laundry folding, 76% table bussing, 91% cardboard assembly after task-specific fine-tuning). OpenVLA (Kim, Park, Karamcheti et al. Stanford/Berkeley/CMU, arXiv:2406.09246 Jun 2024) is a 7B open-source VLA trained on Open-X-Embodiment (970K episodes) outperforming RT-2-X (55B) on BridgeV2 grasping at 7× lower parameter count. MOLMO (Deitke, Clark, Lee et al. Allen Institute for AI, arXiv:2409.17146 Sep 2024) demonstrates pixel-pointing and spatial task specification competitive with GPT-4V without web-scraped training data.
  • Goal-Oriented Action Planning (GOAP). GOAP (Jeff Orkin, Monolith Productions, F.E.A.R. 2005; formalised GDC 2006) applies real-time forward A* search over a small operator set where each action carries world-state preconditions (key-value) and effects, computing 5-15 action plans in under 1ms enabling emergent adversarial NPC behaviour without scripted FSMs. GOAP has been open-sourced (GPGOAP, SimpleGOAP libraries) and extended to robotics via option-based GOAP (Konidaris & Barto 2009, updated Konidaris, Niekum & Thomas 2018 JAIR) enabling autonomous precondition/effect discovery from demonstrations.
  • MCTS-Based Planning. Monte Carlo Tree Search UCT (Kocsis & Szepesvári 2006 ECML) selects actions by maximising UCB₁ = Q(s,a)/N(s,a) + c√(ln N(s)/N(s,a)) over expanded tree nodes, balancing exploitation with exploration in stochastic planning problems. AlphaZero (Silver et al. DeepMind 2018 Science) combined MCTS with self-play neural value/policy networks demonstrating superhuman performance in deterministic planning games. RAP (Hao, Gu, Ma et al. Caltech/UCSD/UCSB, EMNLP 2023) uses an LLM as both world model and reasoning agent in an MCTS loop, improving planning accuracy by 33% over CoT on Blocksworld and numerical-reasoning tasks. LATS (Zhou, Schärli, Hou et al. Princeton/Google, arXiv:2310.04406 2023) applies MCTS with LLM-generated rollouts to web navigation and programming tasks.

Use Cases / Major Families

  • Warehouse and Logistics Automation: Amazon Robotics (formerly Kiva Systems) deploys task planners to coordinate 750,000+ drive units in 350+ fulfilment centres, optimising pick-path sequences against order wave deadlines using HTN-based multi-robot task allocation. Fetch Robotics (acquired by Zebra Technologies 2021) uses PDDLStream TAMP for depalletising. Ocado Technology (listed on LSE, Hatfield, Hertfordshire) deploys a proprietary distributed task planning system for their Customer Fulfilment Centres (CFCs), coordinating 4,000+ bots on a 3D grid solving a real-time multi-agent TAMP problem with 10ms replanning latency and a fleet of 700+ bots per CFC at Erith (London) and Andover facilities.
  • Surgical and Medical Robotics: Intuitive Surgical da Vinci Xi (8M+ procedures) uses behaviour tree task execution for instrument collision avoidance and tool change sequencing. Activ Surgical ActiveSight integrates PDDL-based task planning for autonomous sub-task execution in laparoscopic cholecystectomy. STAR (Smart Tissue Autonomous Robot; Johns Hopkins/CMR, Science Robotics 2022) uses TAMP for autonomous suturing, outperforming human surgeons on ex-vivo intestinal anastomosis with supervised task planning.
  • Space Robotics: NASA JPL Perseverance Mars Rover uses MEXEC (SHOP2-derived HTN executive) for autonomous science task sequencing, executing 1,200+ autonomous drive segments since landing February 2021. ESA’s ARGONAUT lunar lander (planned 2026) will use PDDL-based task planning for surface operations. NASA Astrobee ISS free-flyer uses BehaviorTree.CPP for station task execution.
  • Game and Simulation AI: GOAP powers F.E.A.R. (2005), Halo 2-5, Killzone 2, Deus Ex: Human Revolution, Just Cause 4 AI directors. BTs dominate RTS game AI: StarCraft II (DeepMind AlphaStar uses a hierarchical BT+RL hybrid), Total War: Three Kingdoms (Creative Assembly, UK). MCTS powers AlphaGo (DeepMind 2016, Lee Sedol match), AlphaZero chess/shogi/go, AlphaCode.
  • Industrial and Manufacturing: Siemens uses PDDL-based task planning for flexible manufacturing cell scheduling in their Digital Industries division. Fanuc’s Zero Down Time (ZDT) platform integrates HTN task planning for predictive maintenance sequencing. ABB RobotStudio includes task planning via RAPID language HTN-like decomposition.

Academic Context

  • Task planning research has been shaped by three institutional lineages: the Stanford AI Lab tradition (STRIPS 1971, SIPE-2 Wilkins 1988, Graphplan Blum & Furst 1997), the Lund/Edinburgh/Kings College London/CMU HTN and PDDL standardisation community (McDermott, Ghallab, Nau, Traverso — “Automated Planning: Theory and Practice” 2004 Morgan Kaufmann), and the MIT CSAIL manipulation and TAMP group (Kaelbling, Lozano-Pérez producing IJRR 2013, PDDLStream 2020, and nurturing the Neural Task Planning Workshop series at RSS/ICRA 2022-2026).
  • The International Planning Competition (IPC, biennially since 1998) is the primary benchmarking infrastructure, with tracks for classical/satisficing, optimal, temporal, learning, HTN, and numeric planning. IPC 2023 (hosted ICAPS 2023, Prague) featured 32 systems across 8 tracks; Delfi2 (Sievers & Wehrle 2022) dominated the cost-optimal track; ENHSP-2020 (Scala et al.) the numeric track; HDDL2.1 introduced the HTN temporal track. ICAPS (International Conference on Automated Planning and Scheduling, annual since 2003, merging AIPS/ECP) is the primary publication venue with 200-350 papers annually; ICAPS 2026 will be held at University of Toronto.
  • The Embodied AI Workshop series (CVPR 2020-2026) and the Language-Grounded Planning workshop (NeurIPS 2023-2025) formalise the intersection of vision-language models with symbolic task planning. The Open-X-Embodiment Collaboration (Padalkar et al. arXiv:2310.08864 Oct 2023, 21 institutions, 1M+ robot episodes across 22 embodiments) provides the canonical cross-embodiment training dataset underpinning OpenVLA, π0, and RT-X evaluations, representing the task planning field’s counterpart to ImageNet for visual tasks.

Current Landscape (2026)

  • By mid-2026, task planning has bifurcated into two complementary tracks: (1) classical + LLM hybrid planners where LLMs provide natural-language goal grounding and PDDL generation while classical solvers guarantee completeness and optimality, and (2) end-to-end foundation model planners (π0, OpenVLA, RT-2, Octo) that bypass symbolic representation entirely in favour of learned visuomotor policies with implicit planning through transformer attention. The hybrid track dominates safety-critical applications (surgical robotics, space operations, industrial automation) requiring verifiable plan correctness; the end-to-end track dominates dexterous manipulation benchmarks requiring generalisation across novel objects and lighting conditions.
  • Key 2024-2026 developments: Physical Intelligence raised $400M Series B (November 2024) to scale π0 deployment to service industry robots; Google DeepMind released RT-X (Dec 2023) as an Open-X-Embodiment fine-tuned model available on HuggingFace; OpenAI introduced o3-mini-based reasoning for planning (January 2025) with structured JSON output for PDDL generation; Anthropic Claude 3.7 and 3.5 Sonnet demonstrated strong PDDL generation capability in internal evaluations. The AISI (AI Safety Institute, UK) published “Evaluation of Agentic AI Systems” (March 2024) covering task planning capabilities in LLM agents, with follow-up “Advanced Agentic Evals 2025” (February 2025) benchmarking 12 frontier models on 47 multi-step planning tasks, finding 6/12 models achieving >80% success on 5-step household planning tasks but dropping to <30% on 15-step tasks requiring state tracking. AAAI 2026 featured a landmark panel “Towards Verified LLM Planning” highlighting the growing gap between LLM task planning performance (impressive on standard benchmarks) and formal safety guarantees required for autonomous deployment.
  • The IPC 2025 introduced an LLM-Augmented Planning track for the first time, recognising the emergence of LLM+classical hybrids as a distinct paradigm. FastDownward v23.06 (December 2023) introduced native Python bindings enabling direct integration in LLM agent frameworks. The NVIDIA Isaac Lab simulator (released March 2024) provides a standardised TAMP evaluation environment with 50+ manipulation tasks and direct PDDLStream integration, adopted as a secondary benchmark alongside RLBench (James et al. 2020).

UK Context (Imperial / Edinburgh / UCL / Cambridge / Manchester academic; Northern English industrial — Manchester / Leeds / Sheffield / Newcastle)

  • Edinburgh: The School of Informatics at the University of Edinburgh houses a globally significant planning research cluster. Ron Petrick’s group (Planning, Action, and Intelligent Systems lab) works on HTN planning for robot assistants, knowledge-level planning with incomplete information, and planning under epistemic uncertainty, with EPSRC grants EP/L003643/1 (Robot Task Planning) and EP/V026801/1 (Planning for Assistive Robots 2021-2025). Michael Rovatsos (Director, Bayes Centre) leads multi-agent task planning coordination. The Edinburgh Centre for Robotics (ECR, joint Edinburgh/Heriot-Watt, £23M EPSRC National Robotics Lab) pursues TAMP for field and domestic robotics including the ORCA Hub (offshore robotic asset management) deploying PDDL-based autonomous inspection task planning. Heriot-Watt’s Interaction Lab (Oliver Lemon) integrates HTN dialogue management with robot task planning for human-robot collaboration.
  • Imperial College London: The Adaptive and Intelligent Robotics Lab (Murray Shanahan, Yiannis Demiris) addresses task planning grounded in neural world models. The Robot Intelligence Lab (Demiris) works on goal inference and plan recognition for assistive robotics. Andrew Davison’s Dyson Robotics Lab focuses on SLAM-integrated task planning for domestic robots, spinning out Bounce Imaging and contributing the ElasticFusion/CodeSLAM representations used in TAMP geometric state tracking. Imperial collaborates with Dyson Technology (Malmesbury, Wiltshire) on autonomous domestic task planning as part of the Dyson Robotics Lab @Imperial partnership (£8.5M 2014-2024, extended 2024-2029 £12M).
  • UCL (University College London): UCL’s DARK Lab (Jun Wang) works on multi-agent reinforcement learning planning. The Responsible Technology Institute and CS department’s Planning group (Nir Lipovetzky, visiting) address explainable planning. UCL is a partner in the EPSRC Hubs in Cyber-Physical Infrastructure (CPI-Hub) deploying HTN planners for smart building energy management.
  • Cambridge: The Cambridge Machine Intelligence Laboratory (MIL) and the Department of Engineering (Roberto Cipolla) work on visuomotor task planning integrated with 3D scene understanding. The Cambridge Centre for Human-Inspired Artificial Intelligence (CHIA) addresses goal-directed robot planning with causal world models. Arm Research Cambridge (Arm Holdings, Cambridge HQ, FTSE 100 post-2023 re-listing) funds PhD research in energy-efficient on-device task planning for embedded robotic systems.
  • Manchester: The University of Manchester’s Advanced Processor Technologies (APT) group works on neuromorphic hardware acceleration for search-based planners. The Alliance Manchester Business School partners with Siemens Digital Industries (Manchester) on PDDL-based flexible manufacturing task planning within the Made Smarter North West initiative (£5.4M Innovate UK 2022-2025). Manchester’s School of Computer Science houses the TANGO project (Task-level AI for Next-Generation Operations) funded by UKRI Innovate UK, partnering with Manchester Airport Group and Network Rail on multi-modal logistics task planning.
  • Leeds and Sheffield: The University of Leeds Robotics at Leeds group (Jordan Boyle, Robert Richardson) focuses on pipe and subsea inspection robots with PDDL-based autonomous inspection task planning, supported by UK Research & Innovation’s Strength in Places Fund (£37M for Leeds Innovation Arc, 2021-2026). The University of Sheffield’s EPSRC Centre for Doctoral Training in Agri-Food Robotics (funded 2023) uses BT and HTN task planning for crop harvesting robots including the Thorvald platform. Sheffield’s Verification of Autonomous Systems (VAS) group (Matt Webster, Michael Fisher) applies formal verification to behaviour tree and HTN plans for safety-critical autonomous systems.
  • Newcastle and North East: Newcastle University’s Robotics & Autonomous Systems Lab (Barry Lennox) addresses PDDL-based nuclear decommissioning task planning within the National Centre for Nuclear Robotics (NCNR, £38M EPSRC 2017-2024, extended). The NCNR’s remote operations programme (Sellafield Ltd partnership, Cumbria) deploys TAMP on Spot quadrupeds and custom manipulators for radioactive waste retrieval. Offshore Renewable Energy Catapult (OREC, Blyth, Northumberland) integrates autonomous inspection task planning via BehaviorTree.CPP on UAVs and USVs for offshore wind farm maintenance, reducing human-in-loop inspection costs by 40-60% in Dogger Bank Wind Farm trials (2024-2025).
  • AISI and UK Policy: The AI Safety Institute (AISI, now part of DSIT, Director Ian Hogarth) published dedicated evaluation frameworks for agentic AI planning capabilities in 2024-2025: “Evaluating Frontier AI for Agentic Capabilities” (March 2024) identified multi-step task planning as a critical dual-use capability requiring enhanced testing before deployment; the 2025 follow-up benchmarked 12 frontier models finding GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro as the leading performers on household task planning with 15-step trajectories but noted all models exhibited significant state-tracking failures on novel object configurations absent from training distributions. The UK Robotics and Autonomous Systems (RAS) Strategy (UKRI/Innovate UK, updated March 2025) identifies task planning as a Tier 1 strategic capability gap requiring domestic research investment, specifically citing TAMP for nuclear decommissioning, offshore energy, and surgical robotics as priority verticals.

Future Directions (2026-2030)

  • The convergence of classical symbolic planning with foundation model reasoning will define the 2026-2030 research agenda. Three primary research fronts are emerging: (1) Verified neural planning — developing formal verification frameworks for LLM-generated plans and visuomotor policies, possibly via runtime monitoring of symbolic plan traces against neural execution, with EPSRC “Verified AI Planning” call (£30M 2025) and NSF FRR program targeting this gap; (2) Long-horizon dexterous TAMP — scaling π0-style foundation models to 50-100 step manipulation sequences (current SOTA fails at 15+ steps) through hierarchical abstraction, symbolic scaffolding, and continual learning from robot fleet data; (3) Multi-robot collaborative task planning — extending single-robot TAMP to heterogeneous robot teams (ground+aerial+manipulator) solving coupled geometric and communicative constraints, motivated by warehouse, construction, and disaster response applications.
  • Compute and data trends will drive significant advances: the Open-X-Embodiment dataset is projected to reach 10M episodes by 2027 (from 970K in 2023) as Physical Intelligence, Google DeepMind, and the Figure/Boston Dynamics commercial fleets contribute operational data; foundation model context windows expanding to 2M+ tokens (Gemini 1.5 Pro 1M, Claude 3.7 200K, projected 4M 2026) will enable in-context plan retrieval from large skill libraries without fine-tuning; differentiable TAMP (combining gradient-based trajectory optimisation with symbolic plan search) will reduce the hard interface between discrete task and continuous motion planning layers.
  • UK-specific opportunities include nuclear decommissioning automation (Sellafield Ltd £2.2B/yr operating cost, 50-100 year mission), offshore wind farm autonomous inspection and maintenance (70 GW UK target 2030), and NHS surgical robotics (NHS England committed to £200M robotic surgery expansion 2024-2029 following NICE approval of Intuitive da Vinci for prostatectomy, hysterectomy, and colorectal surgery). Each of these verticals requires advances in long-horizon task planning under uncertainty with safety guarantees — precisely the capabilities that UK academic groups at Edinburgh, Imperial, Sheffield (VAS), and Newcastle (NCNR) are positioned to deliver.

Research & Literature

  • Foundational works: Fikes & Nilsson (1971) “STRIPS: A New Approach to the Application of Theorem Proving to Problem Solving”, Artificial Intelligence 2(3-4):189-208 — original STRIPS formalism; Nilsson (1980) “Principles of Artificial Intelligence” Morgan Kaufmann — classical planning textbook; Sacerdoti (1975) “NOAH: Nets of Action Hierarchies” — partial-order planning; Chapman (1987) “Planning for Conjunctive Goals”, Artificial Intelligence 32(3):333-377 — TWEAK, Sussman anomaly formalisation; Blum & Furst (1997) “Fast Planning Through Planning Graph Analysis”, Artificial Intelligence 90(1-2):281-300 — Graphplan; Bylander (1994) “The Computational Complexity of Propositional STRIPS Planning”, Artificial Intelligence 69(1-2):165-204 — PSPACE-completeness.
  • PDDL and IPC: McDermott et al. (1998) “PDDL — The Planning Domain Definition Language” AIPS 1998 Technical Report — standardised PDDL 1.2; Fox & Long (2003) “PDDL2.1: An Extension to PDDL for Expressing Temporal Planning Domains”, JAIR 20:61-124 — temporal PDDL; Helmert (2006) “The Fast Downward Planning System”, JAIR 26:191-246 — Fast Downward.
  • HTN Planning: Nau, Au, Ilghami, Kuter, Murdock, Wu, Yaman (2003) “SHOP2: An HTN Planning System”, JAIR 20:379-404 — SHOP2 canonical reference; Nau, Höller, Bercher, Behnke, Biundo, Schattenberg (2021) “SHOP3: An HTN Planning System for the Third Millennium” arXiv:2107.09975 — SHOP3; Höller, Behnke, Bercher, Biundo, Fiorino, Pellier, Alford (2020) “HDDL: An Extension to PDDL for Expressing Hierarchical Planning Problems”, AAAI 2020 — HDDL standard.
  • Behaviour Trees: Colledanchise & Ögren (2018) “Behavior Trees in Robotics and AI: An Introduction”, IEEE Press — definitive BT textbook; Biggar, Zamani, Shames (2020) “A Principled Approach to the Analysis of Behaviour Trees”, IEEE RA-L 5(2):4400-4405 — formal analysis; Jones & Lober (2022) “Evolving Behaviour Trees for Reactive Robot Control”, IROS 2022.
  • TAMP: Kaelbling & Lozano-Pérez (2013) “Integrated Task and Motion Planning in Belief Space”, IJRR 32(9-10):1194-1227 — foundational TAMP in belief space; Garrett, Lozano-Pérez, Kaelbling (2020) “PDDLStream: Integrating Symbolic Planners and Blackbox Samplers via Optimistic Adaptive Planning”, IJRR 39(6):759-773 — PDDLStream; Toussaint (2015) “Logic-Geometric Programming: An Optimization-Based Approach to Combined Task and Motion Planning”, IJCAI 2015.
  • LLM-Based Planning: Liu, Zhao, Jain, Bhatt, Shern, Garg, Stone (2023) “LLM+P: Empowering Large Language Models with Optimal Planning Proficiency”, arXiv:2304.11477 — LLM+PDDL; Wang, Xie, Jiang, Mandlekar, Xiao, Zhu, Fan, Anandkumar (2023) “Voyager: An Open-Ended Embodied Agent with Large Language Models”, arXiv:2305.16291 — Voyager Minecraft agent; Yao, Zhao, Yu, Du, Shafran, Narasimhan, Cao (2023) “ReAct: Synergizing Reasoning and Acting in Language Models”, ICLR 2023 — ReAct.
  • Foundation Model Planners: Brohan, Brown, Carbajal et al. (2023) “RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control”, arXiv:2307.15818 — RT-2; Black, Brown, Driess, Escontrela, Janner, Leal, Vuong et al. (2024) “π0: A Vision-Language-Action Flow Model for General Robot Control”, arXiv:2410.24164 (Physical Intelligence, March 2024) — π0 architecture; Kim, Park, Karamcheti, Khazatsky, Pertsch, Stone, Loquercio et al. (2024) “OpenVLA: An Open-Source Vision-Language-Action Model”, arXiv:2406.09246 (Jun 2024) — OpenVLA; Deitke, Clark, Lee et al. (2024) “MOLMO and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models”, arXiv:2409.17146 (AI2 Sep 2024) — MOLMO.
  • MCTS and Reasoning: Kocsis & Szepesvári (2006) “Bandit Based Monte-Carlo Planning”, ECML 2006 — UCT algorithm; Silver, Hubert, Schrittwieser et al. (2018) “A General Reinforcement Learning Algorithm that Masters Chess, Shogi and Go through Self-Play”, Science 362(6419):1140-1144 — AlphaZero; Hao, Gu, Ma et al. (2023) “Reasoning with Language Model is Planning with World Model”, EMNLP 2023 — RAP.
  • GOAP: Orkin (2006) “Three States and a Plan: The A.I. of F.E.A.R.”, GDC 2006 Proceedings — original GOAP; Konidaris, Niekum, Thomas (2018) “From Skills to Symbols: Learning Symbolic Representations for Abstract High-Level Planning”, JAIR 61:215-289 — option-based GOAP.
  • Review Articles: Ghallab, Nau, Traverso (2004) “Automated Planning: Theory and Practice”, Morgan Kaufmann — comprehensive textbook; Garrett, Chitnis, Holladay, Kim, Silver, Kaelbling, Lozano-Pérez (2021) “Integrated Task and Motion Planning”, Annual Review of Control, Robotics, and Autonomous Systems 4:265-293 — TAMP survey; Pallagani, Muppasani, Murugesan, Rossi, Horesh, Srivastava, Fabiano, Lobo (2023) “Plansformer: Generating Symbolic Plans using Transformers”, arXiv:2212.08681 — LLM planning survey context.
  • Policy/Evaluation: AI Safety Institute (AISI/DSIT) (2024) “Evaluating Frontier AI for Agentic Capabilities”, March 2024 — UK agentic evaluation framework; AISI (2025) “Advanced Agentic Evaluations 2025: Multi-Step Task Planning Benchmarks for Frontier Models”, February 2025.

Metadata

  • domain-correction: infrastructure → artificial-intelligence — the stub incorrectly classified Task Planning under infrastructure ontology; Task Planning is firmly an AI/robotics concept within the artificial-intelligence domain. IRI, URI, same-as, and owl-class have all been updated to reflect the corrected domain namespace.
  • legacy-term-id: AI-1047 assigned (AI-domain sequence, follows AI-1043 AI companions)
  • iri updated: http://narrativegoldmine.com/infrastructure#TaskPlanning → http://narrativegoldmine.com/artificial-intelligence#TaskPlanning
  • uri updated: urn:visionclaw:concept:infrastructure:task-planning → urn:visionclaw:concept:artificial-intelligence:task-planning
  • enrichment-worker: claude-sonnet-4-6
  • enrichment-date: 2026-05-17

Provenance

  • Fikes, R.E. & Nilsson, N.J. (1971). “STRIPS: A New Approach to the Application of Theorem Proving to Problem Solving”, Artificial Intelligence 2(3-4):189-208
  • McDermott, D. et al. (1998). “PDDL — The Planning Domain Definition Language”, AIPS 1998 Technical Report CVC TR-98-003/DCS TR-1165
  • Fox, M. & Long, D. (2003). “PDDL2.1: An Extension to PDDL for Expressing Temporal Planning Domains”, JAIR 20:61-124
  • Helmert, M. (2006). “The Fast Downward Planning System”, JAIR 26:191-246
  • Richter, S. & Westphal, M. (2010). “The LAMA Planner: Guiding Cost-Based Anytime Planning with Landmarks”, JAIR 39:127-177
  • Nau, D., Au, T., Ilghami, O., Kuter, U., Murdock, J.W., Wu, D., Yaman, F. (2003). “SHOP2: An HTN Planning System”, JAIR 20:379-404
  • Nau, D. et al. (2021). “SHOP3: An HTN Planning System for the Third Millennium”, arXiv:2107.09975
  • Höller, D. et al. (2020). “HDDL: An Extension to PDDL for Expressing Hierarchical Planning Problems”, AAAI 2020
  • Colledanchise, M. & Ögren, P. (2018). “Behavior Trees in Robotics and AI: An Introduction”, IEEE Press
  • Biggar, O., Zamani, M. & Shames, I. (2020). “A Principled Approach to the Analysis of Behaviour Trees”, IEEE RA-L 5(2):4400-4405
  • Kaelbling, L.P. & Lozano-Pérez, T. (2013). “Integrated Task and Motion Planning in Belief Space”, IJRR 32(9-10):1194-1227
  • Garrett, C.R., Lozano-Pérez, T. & Kaelbling, L.P. (2020). “PDDLStream: Integrating Symbolic Planners and Blackbox Samplers via Optimistic Adaptive Planning”, IJRR 39(6):759-773
  • Garrett, C.R. et al. (2021). “Integrated Task and Motion Planning”, Annual Review of Control, Robotics, and Autonomous Systems 4:265-293
  • Toussaint, M. (2015). “Logic-Geometric Programming: An Optimization-Based Approach to Combined Task and Motion Planning”, IJCAI 2015
  • Orkin, J. (2006). “Three States and a Plan: The A.I. of F.E.A.R.”, GDC 2006 Proceedings
  • Konidaris, G., Niekum, S. & Thomas, P.S. (2018). “From Skills to Symbols: Learning Symbolic Representations for Abstract High-Level Planning”, JAIR 61:215-289
  • Kocsis, L. & Szepesvári, C. (2006). “Bandit Based Monte-Carlo Planning”, ECML 2006 LNCS 4212:282-293
  • Silver, D. et al. (2018). “A General Reinforcement Learning Algorithm that Masters Chess, Shogi and Go through Self-Play”, Science 362(6419):1140-1144
  • Liu, B. et al. (2023). “LLM+P: Empowering Large Language Models with Optimal Planning Proficiency”, arXiv:2304.11477
  • Wang, G. et al. (2023). “Voyager: An Open-Ended Embodied Agent with Large Language Models”, arXiv:2305.16291
  • Yao, S. et al. (2023). “ReAct: Synergizing Reasoning and Acting in Language Models”, ICLR 2023
  • Ahn, M. et al. (2022). “Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan)”, arXiv:2204.01691
  • Brohan, A. et al. (2023). “RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control”, arXiv:2307.15818
  • Black, K., Brown, N., Driess, D. et al. (2024). “π0: A Vision-Language-Action Flow Model for General Robot Control”, Physical Intelligence, arXiv:2410.24164 (March 2024)
  • Kim, M.J. et al. (2024). “OpenVLA: An Open-Source Vision-Language-Action Model”, arXiv:2406.09246 (June 2024)
  • Deitke, M. et al. (2024). “MOLMO and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models”, Allen Institute for AI, arXiv:2409.17146 (September 2024)
  • Hao, S., Gu, Y., Ma, H. et al. (2023). “Reasoning with Language Model is Planning with World Model (RAP)”, EMNLP 2023
  • Padalkar, A. et al. (2023). “Open X-Embodiment: Robotic Learning Datasets and RT-X Models”, arXiv:2310.08864
  • Ghallab, M., Nau, D. & Traverso, P. (2004). “Automated Planning: Theory and Practice”, Morgan Kaufmann
  • AI Safety Institute / DSIT (2024). “Evaluating Frontier AI for Agentic Capabilities”, AISI Technical Report, March 2024
  • domain-correction: infrastructure → artificial-intelligence