A software system that perceives its environment, makes decisions and takes actions to achieve goals, often using a language model together with tools and memory.

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:hasPart ai:AgentLoop))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:hasPart ai:ToolRegistry))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:hasPart ai:AgentMemory))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:hasPart ai:ReasoningEngine))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:hasPart ai:PlanningAndScheduling))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:hasPart ai:Perception))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:hasPart ai:Orchestration))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:hasPart ai:Sandboxing))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:hasPart ai:ObservationInterface))

Dependency Relationships

SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:requires ai:LargeLanguageModels))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:requires ai:ToolUse))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:requires ai:MemoryManagement))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:requires ai:FoundationModel))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:requires ai:FunctionCalling))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:dependsOn ai:CognitiveArchitecture))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:dependsOn ai:VectorDatabase))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:dependsOn ai:API))

Capability Relationships

SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:enables ai:AutomatedPlanning))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:enables ai:TaskAutomation))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:enables ai:MultiAgentSystems))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:enables ai:WorkflowOrchestration))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:enables ai:SoftwareEngineeringAutomation))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:supports ai:AISafety))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:supports ai:HumanInTheLoop))

Implementation Relationships

SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:implements ai:ReActPattern))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:implements ai:ReflexionPattern))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:implements ai:BeliefDesireIntention))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:implements ai:TreeOfThoughts))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:uses ai:ChainOfThought))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:uses ai:RetrievalAugmentedGeneration))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:uses ai:PromptEngineering))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:contrastsWith ai:Chatbot))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:contrastsWith ai:RuleBasedSystem))

Reduction Relationships

SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:reducesTo ai:AutonomousAgent))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:reducesTo ai:GoalDirectedSystem))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:reducesTo ai:AIApplication))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:bridgesTo ai:RoboticProcessAutomation))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:bridgesTo ai:Robotics))
SubClassOf(ai:AIAgent
  ObjectSomeValuesFrom(ai:bridgesTo ai:DigitalTwin))

About

  • The concept of an AI Agent unifies threads from symbolic AI, decision theory, cognitive science, and contemporary deep learning into a single operational paradigm: a software system that pursues goals by choosing and executing actions in an environment it can only partially observe. The definition provided by Russell and Norvig in Artificial Intelligence: A Modern Approach (1995, updated through 4th edition 2020) remains canonical — an agent perceives its environment through sensors and acts upon that environment through actuators — but the substrate has evolved dramatically. Where 1990s agents relied on hand-crafted rule bases or Bayesian networks, modern AI agents use a Foundation Model as a general-purpose controller capable of parsing arbitrary natural-language observations, generating structured action plans, and explaining its reasoning in Chain of Thought traces that are auditable by humans and can feed back into subsequent steps.
  • The distinguish architectural leap of the 2022–2026 period is the combination of language model planning with executable Tool Use: rather than merely generating text, the agent can invoke real-world operations — writing and running code, querying databases, navigating browsers, calling APIs, or spawning sub-agents — and incorporate the concrete results into its next reasoning step. The ReAct Pattern (Yao et al., 2023) formalised this interleaving of Reasoning traces and Acting steps, and it has become near-universal. Extensions include Reflexion Pattern (Shinn et al., 2023), in which the agent generates a verbal self-critique after a failed attempt and stores it as episodic memory, avoiding repetition of the same error; and Tree of Thoughts (Yao et al., 2023), which explores multiple reasoning branches in parallel, using search algorithms to evaluate and select the most promising path for combinatorially hard problems.
  • Safety is the most pressing unsolved challenge. Because an AI Agent can take sequences of real-world actions — sending emails, committing code, executing financial transactions — a single mistaken reasoning step can cascade into irreversible consequences. The principle of minimal footprint (prefer reversible over irreversible actions, acquire only required permissions, escalate to Human-in-the-Loop for high-stakes steps) is now a design norm propagated by Anthropic’s model specification and adopted in Agent Frameworks such as LangGraph and CrewAI. Prompt Injection — where adversarial content in retrieved documents hijacks the agent’s reasoning — represents a qualitatively new attack surface that has no direct analogue in static ML systems, and remains an active research frontier.

Components / Architecture

  • Agent Loop — the core sense-plan-act cycle. At each iteration, the controller LLM receives a context window containing the system prompt, the current goal, a scratchpad of prior steps, the most recent tool output, and any retrieved episodic memories. It generates either a tool-call action (specifying name and parameters in a structured format) or a final answer. The loop controller parses this output, routes tool calls to the appropriate executor, appends the result as an observation, and passes the updated context back to the model.
  • Tool Registry — a catalogue of callable functions exposed to the agent as JSON Schema definitions conforming to the Function Calling interface. Tools range from web search, code execution (Python REPL, shell), file read/write, database queries, REST API calls, to image generation or sub-agent invocations. Tool specification quality — precision of descriptions, parameter typing, and example invocations — has been empirically shown to matter more than prompt phrasing for production agent performance.
  • Agent Memory layering:
    • In-context (working) memory — the active context window. Bounded by the model’s context length (128k–2M tokens in 2026-era models). The primary bottleneck for long-running tasks.
    • Episodic external memory — past steps, observations, and intermediate results stored in a Vector Database (pgvector, Pinecone, Weaviate, Chroma) and retrieved via Retrieval-Augmented Generation similarity search. Enables tasks that exceed the context window.
    • Procedural memory — distilled skill programs, plan templates, or fine-tuned model weights encoding reusable patterns. May be stored as few-shot examples prepended to the system prompt.
    • Semantic memory — structured knowledge graphs or curated document stores providing factual grounding independent of episodic experience.
  • Reasoning Engine — the Foundation Model controller itself (GPT-4, Claude 3.5/4, Gemini 1.5/2). Generates Chain of Thought traces interleaved with tool calls. Extended thinking modes (o1, o3, Claude extended thinking) provide deeper multi-step internal deliberation before emitting an action.
  • Planning and Scheduling — decomposes high-level goals into sub-task DAGs. Approaches include single-shot plan generation, dynamic re-planning after each observation, and MCTS-based search (used in LATS / Language Agent Tree Search). ReAct performs planning inline; Plan-and-Execute separates strategic decomposition from tactical execution for efficiency.
  • Perception — the interface through which the agent reads its environment: structured JSON from APIs, parsed HTML from web pages, file contents, terminal output, screenshot pixel arrays (in Computer Use agents), or multimodal inputs combining text, images, and audio.
  • Orchestration — the scaffolding layer (LangGraph, AutoGen, CrewAI, Semantic Kernel, Claude SDK) that manages loop execution, token budgets, error handling, retry logic, agent routing, and telemetry. In multi-agent topologies, the orchestrator routes sub-tasks between worker agents and aggregates results.
  • Sandboxing layer — isolates dangerous tool execution (code interpreters, shell access) inside Docker containers, WebAssembly runtimes, or E2B cloud sandboxes, preventing lateral movement from compromised tool outputs.

Design Patterns / Major Families

  • ReAct (Reason + Act) (Yao et al. 2023) — the dominant pattern. The LLM generates interleaved Thought / Action / Observation triplets. Thought traces are not sent to tools but serve as scratchpad reasoning; Action specifies a tool call; Observation contains the result. Near-universal in production deployments.
  • Plan-and-Execute — the agent first generates a complete plan (list of steps), then executes each step sequentially, optionally re-planning on failure. Separates strategic planning (LLM-intensive) from tactical execution, improving token efficiency for long tasks.
  • Reflexion (Shinn et al. 2023) — after task failure, the agent writes a verbal self-reflection summarising what went wrong and stores it as episodic memory. Future attempts retrieve the reflection, avoiding repeated mistakes without weight updates. Shows 20–30% improvement on HotpotQA and AlfWorld.
  • LATS / Tree-of-Thought Search (Zhou et al. 2023) — explores multiple candidate reasoning paths as a tree using MCTS, scoring and pruning branches, yielding significantly higher quality on hard combinatorial tasks at the cost of increased token usage.
  • Computer Use / Screen agents — agents that operate on pixel-level screenshot observations rather than DOM/API abstractions, enabling automation of any desktop or browser application without API access. Emerging pattern in 2025–2026, supported by Anthropic’s computer-use API and OSWorld benchmark.
  • Orchestrator–Worker — hierarchical multi-agent topology where a planner agent decomposes and delegates to specialised workers (researcher, coder, critic, summariser). Used in AutoGen, CrewAI, and MetaGPT. The 13th EMAS 2025 workshop (University of Manchester) identified this as the dominant enterprise pattern.
  • Self-evolving agents — agents that modify their own tool registry, system prompts, or fine-tuning datasets based on task performance, as surveyed in the Self-Evolving AI Agents comprehensive review (arXiv:2508.07407, 2025), bridging toward lifelong learning systems.

Use Cases

  • Software Engineering — coding agents (GitHub Copilot Workspace, Devin, Claude Code, SWE-agent) autonomously write, test, debug, and submit pull requests. SWE-bench Verified performance rose from 1.96% (Claude 2, 2023) to above 80% in top systems by late 2025, though benchmark integrity concerns exist regarding test suite leakage.
  • Research Assistance — agents retrieve literature via Retrieval-Augmented Generation, synthesise reviews, and autonomously design experiments. Systems such as AI Scientist (Lu et al., 2024) produced peer-reviewable papers autonomously. Elicit and OpenScholar provide targeted research assistance at scale.
  • Enterprise Workflow Automation — replacing brittle Robotic Process Automation scripts with adaptive agents that handle exceptions through language understanding. Finance-ops agents show median payback in 8.9 months; SDR (sales development) agents pay back in 3.4 months.
  • Scientific Discovery — embodied lab agents (robotic chemists at MIT, Emerald Cloud Lab integrations) run closed-loop experiments, interpreting results and suggesting next experiments. Bridging to Digital Twin simulations for pre-screening hypothesis space.
  • Customer Support Automation — multi-turn agents with CRM integration resolve support tickets end-to-end, escalating to Human-in-the-Loop on low-confidence paths. Deployed at scale by major financial services firms and tech companies.
  • Personal Assistants — long-running agents managing calendars, email triage, travel booking, and research on behalf of users, integrating with productivity suites via OAuth-authenticated Tool Use.
  • Cybersecurity — red-team agents enumerate attack surfaces, exploit vulnerabilities in sandboxed environments, and generate reports. Also used defensively for threat intelligence synthesis.
  • Decentralised Autonomous Organisations — agents operating Smart Contract interfaces on blockchain can execute proposals, manage treasury allocation, and coordinate governance without intermediaries, bridging AI to distributed governance paradigms.
  • Robotics and Embodied AI — language-model agents provide high-level planning for physical robot systems, translating natural-language task descriptions into motion primitives. Bridges to Digital Twin simulation for safe pre-deployment testing.

Academic Context

  • The intellectual foundations of AI Agents are distributed across several decades and disciplines. Russell and Norvig’s Artificial Intelligence: A Modern Approach (1st ed. 1995, 4th ed. 2020) provided the canonical rational agent framework — PEAS characterisation (Performance, Environment, Actuators, Sensors) — that remains pedagogically central. Wooldridge and Jennings’ “Intelligent Agents: Theory and Practice” (1995) formalised the notion of agency properties (autonomy, reactivity, proactivity, social ability) and provided a taxonomy distinguishing reactive, deliberative, and hybrid architectures. Rao and Georgeff’s BDI architecture (1991) operationalised rational deliberation as Belief-Desire-Intention triples, implemented in practical systems including PRS, dMARS, and JACK.
  • The neural turn accelerated from 2022 onward. Yao et al.’s ReAct (2023) demonstrated that interleaving chain-of-thought reasoning with grounded tool actions improved performance on HotpotQA, Fever, and ALFWorld. Shinn et al.’s Reflexion (2023) showed self-critique and episodic memory could replace weight updates for improving agent performance. Wei et al.’s Chain-of-Thought prompting (2022) established the foundation for externalised reasoning. The LLM-based autonomous agent survey (Wang et al., Frontiers in Computer Science, 2024, arXiv:2308.11432) provides a comprehensive taxonomy covering profile, memory, planning, and action modules across 200+ papers.
  • Key benchmarks charting capability frontier: SWE-bench (Jimenez et al. 2024, arXiv:2310.06770) for real-world software engineering; WebArena (Zhou et al. 2024) for web browsing tasks; OSWorld (Xie et al. 2024) for desktop computer use; GAIA (Mialon et al. 2024) for general assistant reasoning; AgentBench (Liu et al. 2024) for operating systems and databases; TAU-bench for tool-agent-user policy adherence. The SafeArena benchmark (arXiv:2503.04957, 2025) evaluates agent safety across five harm categories.
  • Foundational venues: NeurIPS, ICML, ICLR (machine learning); ACL, EMNLP (NLP); AAMAS (multi-agent systems); AAAI, IJCAI (AI generally). Agent-specific workshops proliferating at NeurIPS and ICML from 2023 onward.

Current Landscape (2026)

  • By mid-2026 the market has stratified around three interoperability standards: Anthropic’s Model Context Protocol (November 2024, donated to Linux Foundation Agentic AI Foundation December 2025, 97 million monthly SDK downloads by late 2025); Google’s Agent-to-Agent Protocol (A2A, April 2025, donated to Linux Foundation June 2025, 50+ launch partners including Salesforce, PayPal, Atlassian, Accenture, BCG); and IBM’s Agent Communication Protocol (ACP/BeeAI). These protocols address different layers: MCP standardises agent-to-tool resource access; A2A standardises agent-to-agent delegation across vendor boundaries.
  • Framework landscape in production (based on 18+ Alice Labs deployments, 2024–2026): LangGraph (#1 for complex stateful workflows, fastest latency across 2,000-task benchmark), Claude Agent SDK (#2 for Anthropic-native deployments), CrewAI (#3 for role-based crews), Microsoft Agent Framework v1.0 GA April 2026 (merger of AutoGen and Semantic Kernel). An independent 2026 benchmark across 2,000 task instances found a 37% average gap between lab benchmark scores and production performance.
  • Anthropic holds 32% of the enterprise market, followed by OpenAI (25%) and Google (20%). Enterprise AI agent spending reached USD 37 billion in 2025 (triple the USD 11.5 billion in 2024), with 72% of enterprises planning further spend increases. The global AI agents market is projected at USD 10.9–12.1 billion in 2026, growing at 44–46% CAGR through 2030.
  • The EU AI Act (effective August 2024) classifies autonomous agents operating in high-risk domains (employment, credit, healthcare, critical infrastructure) as high-risk AI systems mandating conformity assessment, human oversight mechanisms, and audit logs. This is driving adoption of Human-in-the-Loop checkpoints and structured logging in production agent deployments across the EU.
  • Benchmark integrity is an active concern: UC Berkeley research found that all eight major agent benchmarks (SWE-bench, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, CAR-bench) could be exploited to achieve near-perfect scores without solving tasks, leading to the adoption of held-out test sets and stricter evaluation protocols.

UK Context

  • The University of Edinburgh’s School of Informatics (largest in Europe) is recognised as the most influential AI school in Europe, with foundational contributions to natural language processing and autonomous systems directly relevant to AI agent development. Edinburgh researchers contributed to early BDI architecture implementations and remain active in the Agentic AI space.
  • Imperial College London’s Department of Computing centres AI research on intelligent, autonomous systems. The UKRI AI Centres for Doctoral Training programme funds PhD cohorts at Edinburgh, Bristol, Cambridge, King’s College London with Imperial, and others, producing the next generation of agent researchers.
  • UCL’s Centre for Artificial Intelligence is one of Europe’s largest, with active research in Deep Reinforcement Learning, Robotics, and NLP — all foundational disciplines for AI agent development. UCL leads the UKRI AI Hub in Generative Models, a collaborative initiative uniting Imperial, Cardiff, Cambridge, Oxford, Manchester, Edinburgh, and Surrey.
  • The University of Manchester hosted the 13th International Workshop on Engineering Multi-Agent Systems (EMAS 2025), where a community roadmap was developed for the next generation of large-scale, adaptive Multi-Agent Systems. The university’s Centre for Robotics and AI works on autonomous systems for industrial applications in partnership with BAE Systems, National Nuclear Laboratory, and Rolls-Royce — practical AI agent deployments in safety-critical domains.
  • North East England received UK government designation as an AI Growth Zone (September 2025), with a £30 billion investment package, 5,000 jobs projected, and partnerships with OpenAI (Stargate UK), NVIDIA, Newcastle University, Durham, Sunderland, and Northumbria University. This positions North East England as a major AI infrastructure hub, with particular focus on advanced manufacturing, Robotics, and healthcare agent deployments.
  • Sheffield Robotics (University of Sheffield and Sheffield Hallam joint centre) maintains significant embodied agent research, bridging language model planning with physical Robotics applications. Leeds and Sheffield digital economy initiatives increasingly target AI agent adoption in financial services (Leeds as UK’s second financial centre) and manufacturing automation.

Future Directions (2026-2030)

  • Long-horizon autonomy — extending reliable agent behaviour from current practical horizons of 10–50 steps to hundreds or thousands of steps, requiring advances in memory compression, error recovery, and goal stability. METR’s HCAST and Time Horizons benchmarks are the primary measurement framework.
  • Self-evolving agent architectures — agents that modify their own tool registries, system prompts, or fine-tuning data based on observed performance, as outlined in the Self-Evolving AI Agents survey (arXiv:2508.07407, 2025), representing the convergence of Agentic AI with lifelong machine learning.
  • Multi-agent safety governance — as multi-agent deployments proliferate, emergent coordination failures, deceptive strategies, and cascading errors across agent networks present novel safety risks. Formal verification methods from Robotics and distributed systems are being adapted for LLM agent networks.
  • Standardised interoperability — the Linux Foundation Agentic AI Foundation (December 2025) aims to converge MCP, A2A, and ACP into a unified stack, enabling plug-and-play agent compositions across vendors, reducing the current 37% lab-to-production performance gap.
  • Regulatory alignment — the EU AI Act’s high-risk classifications for agentic deployments, combined with emerging US NIST AI RMF guidance for autonomy and accountability, are shaping mandatory audit trail, explainability, and Human-in-the-Loop requirements that will become baseline engineering norms by 2027–2028.
  • Embodied and physical world integration — the boundary between AI Agents and Robotics is dissolving as language model planners control physical actuators in manufacturing, logistics, and healthcare. Digital Twin environments enable simulation-validated agent plans before physical deployment.
  • Multi-modal and computer-use agents — screen-level Computer Use agents (operating on pixel observations rather than APIs) are projected to handle the majority of desktop software automation by 2028, replacing both Robotic Process Automation and bespoke API integrations.

Research & Literature

    1. Russell, S. & Norvig, P. (2020). Artificial Intelligence: A Modern Approach (4th ed.). Pearson. [Canonical rational agent framework and PEAS characterisation.]
    1. Wooldridge, M. & Jennings, N.R. (1995). Intelligent Agents: Theory and Practice. The Knowledge Engineering Review, 10(2), 115–152. [Properties of autonomous intelligent agents.]
    1. Rao, A.S. & Georgeff, M.P. (1991). Modeling Rational Agents within a BDI-Architecture. In Proceedings of KR-91. [Foundational BDI architecture paper.]
    1. Yao, S. et al. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023. arXiv:2210.03629. [Foundational ReAct pattern for LLM agents.]
    1. Shinn, N. et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. arXiv:2303.11366. [Self-critique episodic memory for agent improvement.]
    1. Wei, J. et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022. arXiv:2201.11903. [Chain of thought as agent reasoning substrate.]
    1. Wang, L. et al. (2024). A Survey on Large Language Model Based Autonomous Agents. Frontiers in Computer Science. arXiv:2308.11432. [200+ paper taxonomy of LLM agent architectures.]
    1. Jimenez, C.E. et al. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR 2024. arXiv:2310.06770. [Software engineering benchmark for coding agents.]
    1. Zhou, S. et al. (2024). WebArena: A Realistic Web Environment for Building Autonomous Agents. ICLR 2024. arXiv:2307.13854. [Web browsing benchmark for agents.]
    1. Xie, T. et al. (2024). OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. NeurIPS 2024. arXiv:2404.07972. [Computer-use benchmark.]
    1. Mialon, G. et al. (2024). GAIA: A Benchmark for General AI Assistants. ICLR 2024. arXiv:2311.12983. [General assistant reasoning benchmark.]
    1. Liu, X. et al. (2024). AgentBench: Evaluating LLMs as Agents. ICLR 2024. arXiv:2308.03688. [Multi-environment agent evaluation across OS, DB, games.]
    1. Yao, S. et al. (2023). Tree of Thoughts: Deliberate Problem Solving with Large Language Models. NeurIPS 2023. arXiv:2305.10601. [Tree-search reasoning pattern for hard problems.]
    1. Park, J.S. et al. (2023). Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023. arXiv:2304.03442. [Simulation of agent populations with persistent memory.]
    1. Schick, T. et al. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. NeurIPS 2023. arXiv:2302.04761. [Self-supervised tool use learning.]
    1. Patil, S.G. et al. (2024). Gorilla: Large Language Model Connected with Massive APIs. arXiv:2305.15334. [API-call accuracy for LLM agents.]
    1. Qin, Y. et al. (2024). ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. ICLR 2024. arXiv:2307.16789. [Large-scale tool-use training.]
    1. Hong, S. et al. (2024). MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. ICLR 2024. arXiv:2308.00352. [Software engineering multi-agent framework.]
    1. Wu, Q. et al. (2024). AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. COLM 2024. arXiv:2308.08155. [Conversational multi-agent framework.]
    1. Packer, C. et al. (2023). MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560. [Virtual context management for long-horizon agents.]
    1. Madaan, A. et al. (2023). Self-Refine: Iterative Refinement with Self-Feedback. NeurIPS 2023. arXiv:2303.17651. [Iterative self-improvement without retraining.]
    1. Anthropic. (2024). Model Context Protocol Specification. https://modelcontextprotocol.io/ [Open standard for agent–tool integration.]
    1. Google. (2025). Agent-to-Agent (A2A) Protocol Specification. https://google.github.io/A2A/ [Inter-agent delegation protocol.]
    1. University of Manchester EMAS 2025. Engineering the Next Generation of Multi-Agent Systems: A Community Roadmap. Research Explorer, University of Manchester. [UK academic roadmap for multi-agent engineering.]
    1. Lu, C. et al. (2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv:2408.06292. [End-to-end scientific research agent.]
    1. Zhou, A. et al. (2023). Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models. arXiv:2310.04406. [MCTS-based agent planning.]
    1. A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents. (2025). arXiv:2506.23844. [Security and safety taxonomy for agentic LLMs.]
    1. A Comprehensive Survey of Self-Evolving AI Agents. (2025). arXiv:2508.07407. [Lifelong agentic learning bridging foundation models and evolution.]

Provenance