Agent Frameworks are software libraries, runtimes, and orchestration platforms that compose large language models (LLMs) with tool-use, memory, planning, and inter-agent communication into autonomous or semi-autonomous systems capable of multi-step goal pursuit, encompassing single-agent harnesse…

Semantic Classification

Content

## Compositional Relationships (Components)
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:hasPart ai:AgentRuntime))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:hasPart ai:ToolRegistry))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:hasPart ai:MemoryStore))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:hasPart ai:PlannerModule))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:hasPart ai:ExecutorModule))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:hasPart ai:PromptTemplate))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:hasPart ai:StateGraph))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:hasPart ai:HookSystem))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:hasPart ai:TraceLogger))

## Dependency Relationships
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:requires ai:LargeLanguageModel))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:requires ai:ToolUse))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:requires ai:InferenceEndpoint))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:requires ai:PersistenceLayer))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:dependsOn ai:LLMInference))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:dependsOn ai:FunctionCalling))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:dependsOn ai:JSONSchema))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:dependsOn ai:VectorDatabase))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:dependsOn ai:EmbeddingModel))

## Capability Relationships
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:enables ai:AutonomousCoding))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:enables ai:MultiAgentCollaboration))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:enables ai:ToolAugmentedReasoning))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:enables ai:WorkflowAutomation))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:enables ai:ComputerUse))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:supports ai:CodingAssistant))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:supports ai:CustomerServiceAgent))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:supports ai:ResearchAssistant))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:supports ai:RAGPipeline))

## Implementation Relationships
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:implements ai:ReActPattern))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:implements ai:ReflexionPattern))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:implements ai:ToolformerPattern))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:implements ai:PlanAndExecutePattern))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:implements ai:HierarchicalTopology))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:implements ai:MeshTopology))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:uses ai:MCPProtocol))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:uses ai:ACPProtocol))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:uses ai:A2AProtocol))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:uses ai:FunctionCalling))

## Reduction Relationships
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:reduces ai:HumanLabour))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:reduces ai:DeveloperBoilerplate))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:reduces ai:IntegrationComplexity))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:reduces ai:ContextSwitchingOverhead))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:reduces ai:PromptEngineeringEffort))

## Standardisation Relationships
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:standardizedBy ai:ModelContextProtocol))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:standardizedBy ai:AgentCommunicationProtocol))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:standardizedBy ai:AgentToAgentProtocol))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:contrastsWith ai:ClassicalWorkflowEngine))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:contrastsWith ai:RoboticProcessAutomation))
SubClassOf(ai:AgentFrameworks
  ObjectSomeValuesFrom(ai:contrastsWith ai:BPMNOrchestration))

## Data Properties (Characteristics)
DataPropertyAssertion(ai:hasIdentifier ai:AgentFrameworks "AI-1212"^^xsd:string)
DataPropertyAssertion(ai:authorityScore ai:AgentFrameworks "0.87"^^xsd:decimal)
DataPropertyAssertion(ai:githubStarsLangChain ai:AgentFrameworks "97000"^^xsd:integer)
DataPropertyAssertion(ai:githubStarsAutoGen ai:AgentFrameworks "42000"^^xsd:integer)
DataPropertyAssertion(ai:githubStarsCrewAI ai:AgentFrameworks "45900"^^xsd:integer)
DataPropertyAssertion(ai:mcpServerCount ai:AgentFrameworks "16000"^^xsd:integer)
DataPropertyAssertion(ai:mcpMonthlyDownloads ai:AgentFrameworks "97000000"^^xsd:integer)
DataPropertyAssertion(ai:sweBenchVerifiedTopScore ai:AgentFrameworks "0.772"^^xsd:decimal)
DataPropertyAssertion(ai:fortune500CrewAIAdoption ai:AgentFrameworks "0.60"^^xsd:decimal)

## Annotations
AnnotationAssertion(rdfs:label ai:AgentFrameworks "Agent Frameworks"@en)
AnnotationAssertion(rdfs:comment ai:AgentFrameworks "Software libraries, runtimes, and orchestration platforms composing LLMs with tool-use, memory, planning, and inter-agent communication into autonomous or semi-autonomous systems, spanning LangChain/LangGraph, AutoGen, CrewAI, Letta, Pydantic-AI, OpenAI Agents SDK, claude-agent-sdk, Google ADK, DSPy, LlamaIndex, Smolagents, Mastra, BAML, Inngest, Vercel AI SDK and agentic-flow, standardised by MCP/ACP/A2A/AGNTCY protocols, implementing ReAct/Reflexion/Toolformer/planner-executor patterns across hierarchical/mesh/star/pipeline topologies, evaluated against SWE-bench/GAIA/AgentBench/ToolBench/WebArena, deployed in coding assistants (Cursor, Devin, Claude Code), customer service (Klarna 700-FTE-equivalent agent), enterprise workflow (Salesforce Agentforce, ServiceNow Now Assist, Microsoft Copilot Studio) and scientific discovery (Future House)."@en)
AnnotationAssertion(dcterms:identifier ai:AgentFrameworks "AI-1212"^^xsd:string)
AnnotationAssertion(dcterms:subject ai:AgentFrameworks "LLM Agents, Orchestration, Tool Use, Multi-Agent Systems, Agent Protocols"@en)

## Property Characteristics
AsymmetricObjectProperty(ai:requires)
AsymmetricObjectProperty(ai:enables)
AsymmetricObjectProperty(ai:implements)
AsymmetricObjectProperty(ai:standardizedBy)
TransitiveObjectProperty(ai:dependsOn)
FunctionalDataProperty(ai:authorityScore)

About Agent Frameworks

  • Agent Frameworks are the software substrate that converted the bare capability of LLM tool-use (OpenAI function-calling June 2023, Anthropic tool-use November 2023) into a deployable architectural category. Where a base LLM offers a single completion endpoint, an agent framework provides the surrounding scaffolding required to plan, act, observe, remember, and recover across many such completions stitched into a coherent goal-directed loop. The category emerged in earnest with LangChain (Harrison Chase, October 2022), exploded into public consciousness with AutoGPT (Toran Bruce Richards, March 2023, 167K GitHub stars within six months) and BabyAGI (Yohei Nakajima, April 2023), and matured through 2024-2026 into a stratified ecosystem with clear production winners, established design patterns, and emerging interoperability protocols.
  • The defining problem agent frameworks solve is the gap between a single LLM call and a system: a chat completion returns tokens, but a useful agent must invoke tools, persist intermediate state across calls, recover from tool errors and hallucinated arguments, coordinate with other agents or humans, observe the real world (via browser, filesystem, APIs), and detect when it has finished or must escalate. Each of these requires engineering primitives—retry loops, durable state, schema validation, observability, sandboxing, permission control—that fall outside the LLM provider’s API surface and inside the framework’s responsibility. The frameworks differ principally in how opinionated they are about the control flow: at one extreme, LangChain Expression Language (LCEL) and Vercel AI SDK provide composition primitives; at the other extreme, CrewAI prescribes a role-playing crew with sequential or hierarchical processes. LangGraph occupies the productive middle ground with explicit state machines that compile to deterministic graphs while permitting LLM-conditioned routing.

Five Architectural Layers

The 2024-2026 generation of frameworks coheres around a roughly five-layer mental model:

  1. Inference layer — model access, often abstracted across providers (OpenAI, Anthropic, Google, Mistral, Cohere, local via Ollama / vLLM / Llama.cpp); standardised through OpenAI-compatible Chat Completions schema or, increasingly, OpenAI Responses API and Anthropic Messages API. LiteLLM, OpenRouter, and Portkey provide cross-provider routing.
  2. Tool layer — function calling with JSON Schema arguments, retrieval, code execution sandboxes, browser automation, MCP-mediated external tools. Pydantic-AI and BAML enforce strict typing at this layer.
  3. Memory layer — short-term context window management (compaction, summarisation, attention truncation), long-term vector or key-value memory (Letta self-editing memory, Mem0, Zep, MemGPT-style core/recall memory pages), episodic state across sessions (LangGraph checkpointer, OpenAI Threads, Anthropic Sessions).
  4. Orchestration layer — the agent control loop itself (ReAct loop, plan-execute loop, state graph, multi-agent broadcast), with hooks for human-in-the-loop, approval gates, and error recovery. This is the layer that distinguishes frameworks most visibly.
  5. Observability and evaluation layer — tracing (LangSmith, Arize Phoenix, Helicone, Langfuse), evaluation harnesses (RAGAS, Inspect AI, OpenAI Evals, Anthropic’s Inspect-based evals), guardrails (NVIDIA NeMo Guardrails, Guardrails AI, LLM Guard).

Canonical Control Loops

Beneath the five-layer model sit four canonical control-flow shapes that every framework instantiates in some combination:

  • ReAct loop (Yao et al. 2023): a single agent iterates Thought → Action → Observation until a termination condition (final answer, tool budget exhausted, or explicit done). The Thought is unstructured natural-language reasoning produced by the LLM; the Action is a structured tool call validated against schema; the Observation is the tool output (or error) appended back to context. ReAct is the default loop in LangChain AgentExecutor, LangGraph create_react_agent, OpenAI Agents SDK Runner.run, Pydantic-AI Agent.run, and Anthropic claude-agent-sdk. Its strengths are simplicity, debuggability, and emergent behaviour; its weaknesses are myopia (no long-horizon planning) and context-window growth proportional to trajectory length.
  • Plan-and-Execute loop: a Planner LLM first decomposes the task into a sequence (or graph) of subtasks; a separate Executor LLM (often with tool access) executes each subtask; a Replanner adjusts the plan when execution diverges. Plan-and-Execute reduces context bloat (the executor sees only the current subtask) at the cost of brittleness when plans are wrong. LangGraph’s plan-and-execute template, BabyAGI, and CrewAI’s hierarchical Process implement variants. Plan-Reflect-Execute adds an explicit reflection step between planner and executor.
  • Multi-Agent message-passing: agents communicate via structured messages (often free-text wrapped in role metadata). AutoGen’s GroupChat cycles a Manager agent that selects the next speaker based on conversation state; CrewAI sequentially or hierarchically delegates between agents; Microsoft Magentic-One uses a LedgerOrchestrator maintaining a public ledger of facts and a private one of agent capabilities, replanned as evidence accumulates.
  • State-graph execution: the agent is a finite (or recursive) state machine with typed transitions. Nodes are pure functions of state; edges are deterministic or LLM-conditioned. LangGraph is the canonical implementation; PocketFlow and LlamaIndex Workflow are minimal variants. The state-graph approach trades expressiveness for verifiability: graphs can be statically analysed for reachable states, cycle bounds can be enforced, and checkpoints align naturally with graph nodes.

Components / Architecture

Single-Agent Frameworks

LangChain (Harrison Chase, October 2022) became the default Python and TypeScript library for LLM application development. Its strength is breadth: 700+ integrations covering every conceivable LLM provider, vector store, document loader, and tool. LangChain Expression Language (LCEL, 2023) introduced declarative pipe-composition (prompt | model | parser) with automatic streaming, batching, async, and parallelism. By 2026 the repo has 97,000+ GitHub stars and 50,000+ production apps, with the v0.3 stable release in September 2024 splitting the monorepo into langchain-core, langchain-community, and provider-specific packages to address dependency bloat. LangChain’s tradeoff is its breadth-as-weakness: critics argue too many abstractions and too thin a layer over each tool, prompting many production teams to either drop LangChain entirely (Cursor famously rewrote off it) or use only LangGraph from the LangChain stack. LlamaIndex (Jerry Liu, late 2022) started as a retrieval-augmented generation library and has expanded into a full agent framework with 38,000+ stars by 2026. Its agentic-workflows abstraction (2024) introduced event-driven workflows replacing the older query-engine model. LlamaParse, the commercial document-understanding pipeline, achieves state-of-the-art table and figure extraction from PDFs and is integrated as the default for many enterprise RAG-agent deployments. Pydantic-AI (Pydantic team, May 2024) brings Pydantic’s type-safety reputation to the agent space. Function calls are strongly typed via Pydantic models; outputs are validated against return-type schemas; result-types drive the tool-call loop. Pydantic-AI’s deliberately small API surface (Agent, RunContext, ResultData) and provider-agnostic design (OpenAI, Anthropic, Google, Groq, Mistral, Ollama, plus arbitrary OpenAI-compatible endpoints) have made it the framework of choice for Python developers prioritising correctness over ecosystem breadth. Vercel AI SDK (Vercel, July 2023+) is the dominant TypeScript framework for streaming-UI LLM applications. Its streamText, generateObject, tool primitives integrate with React via useChat/useCompletion. AI Elements (released 2025) provides shadcn-style UI components for chat, citations, and tool calls. By 2026 Vercel AI SDK is bundled into the Next.js and Vercel deployment defaults, with first-class support for Mastra workflows and Inngest durable execution on top. BAML (BoundaryML 2023) takes a radically different approach: prompts are functions in a DSL (.baml files) with strongly-typed parameters and return types, compiled to Python/TypeScript/Ruby clients. BAML’s structured-output reliability (Schema-Aligned Parsing) has earned production adoption at PineconeDB, Vendia, and several Y Combinator companies. The functions-as-prompts model contrasts sharply with the data-as-prompts approach of LangChain / LlamaIndex.

Graph-State Runtimes

LangGraph (LangChain Inc., 2024) is arguably the most important agent-runtime release of 2024-2026. It models agents as state graphs where nodes are typed Python functions (an LLM call, a tool call, a conditional router, a human-in-the-loop pause) and edges are explicit transitions, optionally LLM-conditioned. Critical features: persistent checkpoints via PostgreSQL/Redis/SQLite enabling resume-after-crash and time-travel debugging; interrupt primitives for human approval; subgraph composition; streaming token-by-token and step-by-step. By 2026 Gartner cites LangGraph in 34% of agent-framework production architecture documents at 1000+ employee organisations. Production deployments include LinkedIn (job-search and recruiting agents), Uber (developer-productivity agents), Klarna (AI assistant absorbing 700 FTE-equivalent customer-service workload), Replit (Replit Agent), and Elastic. LangGraph Studio (visual IDE, 2024) and LangGraph Platform (managed deployment, 2025) commercialise the runtime. The trade-off is verbosity: a simple agent that fits in 20 LangChain lines may expand to 60-100 LangGraph lines, but the additional structure pays dividends in production debuggability and durability.

Multi-Agent Orchestrators

AutoGen (Microsoft Research, 2023) pioneered conversational multi-agent patterns where agents interact through structured chat. The original ConversableAgent / GroupChat / UserProxyAgent abstractions enabled writing multi-agent code generation, debate, and tool-using assistants in tens of lines. The v0.4 release (January 2025) reorganised the codebase into autogen-core (event-driven message-passing core), autogen-agentchat (high-level multi-agent abstractions: AssistantAgent, RoundRobinGroupChat, SelectorGroupChat), and autogen-ext (extensions for OpenAI, Azure, Docker code-execution). Microsoft Magentic-One (November 2024), built on AutoGen v0.4, is a generalist orchestrator-led multi-agent system that achieved competitive results on GAIA, AssistantBench, and WebArena. CrewAI (João Moura, 2024) prescribes a more opinionated role-playing crew abstraction: each agent has a role (e.g., “Senior Software Engineer”), goal, and backstory; tasks are assigned to agents; the crew executes sequentially, hierarchically (with a manager agent), or in parallel. CrewAI Flows (2024) added explicit state-machine control on top. With 45,900+ stars, $18M Series A (2024), and the claim of powering agents at 60% of Fortune 500 enterprises by 2026, CrewAI has become the go-to choice for teams prioritising rapid prototyping and a low-ceremony multi-agent abstraction. Critics note that the role-play framing leaks into prompts in ways that may not always serve the underlying task. OpenAI Agents SDK (March 2025) replaced the experimental Swarm (October 2024) library with a production-ready framework built atop the OpenAI Responses API. Core primitives: Agent, Runner, handoffs between agents, guardrails for input/output safety. Built-in tracing integrates with the OpenAI dashboard. The SDK is provider-locked to OpenAI but the abstraction is general; community ports to Anthropic and Google exist. Anthropic claude-agent-sdk (September 2025, formerly Claude Code SDK) exposes the same harness that powers Claude Code. Available in TypeScript and Python, it provides: tool use (built-in or via MCP servers), file-system tools (Read, Edit, Write, Glob, Grep, Bash), hook system (PreToolUse, PostToolUse, Stop, SessionStart, etc.), subagents via the Task tool, sessions with persistent context, and granular permission modes. The SDK has become the foundation for a wave of coding-agent products and custom internal agents at Anthropic enterprise customers including SK Telecom, GSK, Bridgewater, and Pfizer. Google ADK (Agent Development Kit) (April 2025, announced at Google Cloud Next) is Google’s open-source Python framework with native Vertex AI integration. It bundles agent primitives, model abstraction (Gemini-first but provider-agnostic), tool composition, and Vertex Agent Builder integration for managed deployment. ADK includes built-in agents for Gemini multimodal (vision, audio, video) and is designed to interoperate with the Google-led A2A protocol.

Specialised Frameworks

DSPy (Stanford NLP / Omar Khattab, 2023) reframes LLM application development as programming, not prompting. Programmers write declarative signatures (question -> answer); DSPy compiles these into optimised prompts via teleprompters (BootstrapFewShot, MIPROv2, COPRO) that search for example demonstrations and instruction variants improving performance against a metric. By 2026 DSPy has 19,000+ stars and is used in production at JetBlue, Replit, Hayden AI, and Databricks. DSPy’s optimisation-based philosophy contrasts with the prompt-engineering-by-hand approach of most other frameworks. Letta (formerly MemGPT, Berkeley Sky Computing Lab 2023, renamed October 2024) is the leading stateful-agent framework. Building on the MemGPT paper (Packer et al. 2023, arXiv:2310.08560), Letta agents have a hierarchical memory architecture: core memory (always in context), recall memory (conversational history, searchable), archival memory (long-term knowledge store), and a sleep-time compute mode that consolidates memories asynchronously. With $10M seed funding (2024), Letta differentiates from the stateless majority by making the agent itself a persistent entity addressable across sessions. Smolagents (HuggingFace, January 2025) is a deliberately minimal (<1000 LOC) framework where agents act by writing and executing Python code rather than calling discrete tools. The premise: code is a richer action space than JSON-tool-call schemas and benefits from LLMs’ code-pretraining. CodeAgent and ToolCallingAgent share an interface; HuggingFace Spaces integration ships pre-built tools. Mastra (Gatsby founders, November 2024) is a TypeScript-first agent and workflow framework targeting the Next.js / Vercel ecosystem. Mastra Workflows provide durable step functions; Mastra Agents wrap Vercel AI SDK; Mastra Memory integrates with Pinecone/Postgres/Upstash. Designed for serverless deployment with Inngest or Cloudflare Workers underneath. Inngest is a durable-workflow runtime (not strictly an agent framework, but increasingly used as the substrate beneath one). step.run, step.sleep, step.waitForEvent provide deterministic replay over non-deterministic LLM calls, enabling exactly-once execution semantics across multi-hour agent runs. agentic-flow / claude-flow (ruvnet) is an orchestration toolkit and command harness focused on hierarchical-mesh swarm topologies with queen-worker patterns, neural memory (RuVector-style HNSW vector stores with AgentDB), and hook-driven coordination. It targets the developer-productivity niche where a single human directs 5-50 sub-agents over a multi-hour session.

Inter-Agent and Tool Protocols

MCP — Model Context Protocol (Anthropic, 26 November 2024) is the JSON-RPC 2.0 standard for exposing tools, resources, and prompts to LLM clients. Transport layers: stdio (local servers), Server-Sent Events (HTTP+SSE for remote servers, deprecated 2025), Streamable HTTP (new default). MCP servers expose tools/list, tools/call, resources/list, resources/read, prompts/list, prompts/get. The protocol exploded from 100 servers at launch (November 2024) to 1,000+ by February 2025 to 16,000+ by late 2025, with 97 million monthly SDK downloads reported by Anthropic in December 2025. Critically, the standard was adopted by OpenAI (March 2025), Google Cloud (Vertex AI Agent Builder, March 2025), Microsoft (Copilot Studio, May 2025), and embedded in IDEs (Cursor, Windsurf, Zed, Replit, VS Code with Copilot). Reference servers from Anthropic cover GitHub, Slack, Google Drive, Postgres, Puppeteer, Filesystem, Memory; community servers cover essentially every SaaS API. ACP — Agent Communication Protocol (IBM Research, January 2025) is a REST-native protocol for inter-agent communication, with explicit capability discovery, asynchronous task delegation, and conversation threads. ACP was donated to the Linux Foundation as part of the BeeAI project (April 2025). It complements MCP: MCP covers LLM-to-tool; ACP covers agent-to-agent. A2A — Agent-to-Agent Protocol (Google, April 2025) is Google’s competing inter-agent protocol launched with 50+ partners (Atlassian, Box, Cohere, Salesforce, SAP, ServiceNow, MongoDB, Workday, Accenture, Deloitte). Core primitives: AgentCard (capability metadata served at /.well-known/agent.json), Task (work unit with lifecycle states submitted/working/input-required/completed/canceled/failed), Artifact (typed result). Donated to the Linux Foundation in June 2025. A2A and ACP are partially overlapping; consolidation discussions continue through 2026. AGNTCY Collective (Cisco, March 2025, with LangChain, Galileo, Glean, and others) defines an “Internet of Agents” reference architecture including agent identity, discovery, secure communication, and observability layers. Less mature than MCP/A2A but pitched at the highest level of inter-organisation agent interoperability.

Hook Systems, Permissions and Sandboxing

Production-grade agent frameworks have converged on a hook / middleware system for safety, observability, and policy enforcement. The Anthropic claude-agent-sdk codifies seven hook event types: PreToolUse (block, modify, or approve tool calls before execution); PostToolUse (post-process or trace tool results); Stop (intercept agent termination); SessionStart / SessionEnd (session bookkeeping); UserPromptSubmit (validate user inputs); Notification (route user-visible messages); PreCompact (state checkpointing before context compaction); SubagentStop (intercept subagent completion). The same pattern appears in LangGraph as add_messages reducers and interrupt primitives, in OpenAI Agents SDK as Guardrails with tripwire_triggered, in CrewAI as before_kickoff / after_kickoff and step callbacks, and in AutoGen as register_hook. Permission modes range from full-auto (auto-edit, bypassPermissionsAndDirectories in claude-agent-sdk) to per-tool approval (manual, plan modes), with tool-pattern allowlists/denylists at intermediate granularity. Sandboxing: code-execution agents standardly run untrusted code in containerised environments (Docker, Firecracker, E2B, Daytona, Modal sandboxes, Riza, Code Interpreter SDK), with filesystem isolation, network egress restriction, and resource caps (CPU, memory, wall-time). Audit trails combine OpenTelemetry-compatible traces (LangSmith, Arize Phoenix, Helicone, Langfuse, Datadog LLM Observability) with structured event logs persisted to durable storage for compliance.

Memory Architectures

Memory is the single most differentiating component across frameworks. Five distinct architectures coexist:

  1. Stateless (default for OpenAI Chat Completions, Anthropic Messages API, base ReAct loops): the agent’s “memory” is the conversation history passed in each call. Limits: context window size, attention dilution, cost per token.
  2. Buffer with summarisation (LangChain ConversationSummaryMemory, LlamaIndex ChatSummaryMemoryBuffer): when history exceeds a threshold, an LLM summarises older turns. Trades fidelity for length.
  3. Vector-recall memory (Mem0, Zep, LangChain VectorStoreRetrieverMemory, LlamaIndex VectorMemory): all turns embedded and indexed; relevant chunks retrieved per turn. Risks: retrieval failures lose state silently.
  4. Hierarchical OS-style memory (MemGPT/Letta, Anthropic’s claude-agent-sdk plus AgentDB / RuVector): explicit core memory (always in context), working memory (active session), recall memory (searchable conversational history), archival memory (long-term knowledge). Tool calls page memory in and out. Sleep-time compute consolidates memories asynchronously.
  5. Graph-structured memory (Cognee, MemGPT 2.0, Anthropic Memory Tool 2025): the agent maintains a knowledge graph linking entities and events; retrieval is via graph traversal rather than (or in addition to) embedding similarity. Particularly suited to long-horizon agents with rich entity structure.

Use Cases / Major Families

Coding and Software Engineering Agents

The single largest production category. Cursor (Anysphere) AI-native IDE with proprietary agent harness exceeding 1M paid seats by 2026. Cognition Devin (March 2024) marketed as an autonomous software engineer, drawing both attention and scepticism. Anthropic Claude Code (CLI agent) and claude-agent-sdk power thousands of internal coding agents at Anthropic enterprise customers. Aider (Paul Gauthier, 2023) terminal-based pair programming. Replit Agent (LangGraph-backed) generates full-stack apps from natural language. Windsurf (Codeium, 2024). GitHub Copilot Workspace (April 2024) and Copilot Coding Agent (May 2025). Sourcegraph Cody, Tabnine, Continue.dev, Cline (formerly Claude Dev), Goose (Block, 2024 open-source agent). SWE-bench Verified scores climbed from 12% (Claude 2 / SWE-agent, late 2023) to 77.2% (Claude 4.5 Sonnet with agentic harness, 2026), a 6x improvement in 30 months.

Customer Service and Support Agents

Klarna AI Assistant (LangChain / OpenAI, launched February 2024) handled 2.3M conversations in its first month, equivalent to ~700 full-time agents, delivering an estimated 4.5B valuation), Decagon (2/conversation versus $7 for human agents. ServiceNow Now Assist (May 2024) and Zendesk AI Agents (2024-2025) close the integration loop.

Enterprise Workflow and Knowledge Agents

Glean (700M) knowledge-work agents for financial services and consulting; Harvey (3B) legal AI agents deployed at PwC, Allen & Overy, A&O Shearman; **Lexis+ AI** and **Thomson Reuters CoCounsel** in legal; **AlphaSense** (4B) financial-research agents. Microsoft Copilot Studio and Microsoft 365 Copilot agents (November 2024) bring agent creation into the Microsoft enterprise tenant.

Browser / Computer-Use Agents

Anthropic Claude Computer Use (beta, October 2024) controls a desktop via screenshots and mouse/keyboard. OpenAI Operator (January 2025) browser-only computer-use agent. Google Project Mariner (December 2024) Gemini-based browser agent. Browser-Use (Magnus M. 2024) open-source library. Anthropic Claude for Chrome (2025 beta). Perplexity Comet (July 2025) agentic browser. Common evaluation: WebArena, VisualWebArena, OSWorld. Practical reliability remains the open problem: 30-60% task-success on realistic web flows by 2026.

Scientific and Research Agents

Future House (Sam Rodriques 2023, Eric Schmidt-backed) deploys Crow / Falcon / Owl agents for biology literature review, achieving superhuman performance on PaperQA2 (2024). STORM (Stanford, 2024) generates Wikipedia-style articles. OpenAI Deep Research (February 2025, derived from o3) and Google Gemini Deep Research (December 2024) consumer research agents. Anthropic Research mode and Perplexity Pages in the same category. Coding-agent variants such as AlphaEvolve (DeepMind, May 2025 LLM-based evolutionary algorithm-discovery agent improved 75-year-old algorithms for sorting and matrix multiplication, reducing Google’s compute by 0.7%).

Data-Analysis and BI Agents

Hex Magic, Mode AI, Tableau Agents, Snowflake Cortex Agents, Databricks AI/BI Genie. PandasAI and Open Interpreter (Killian Lucas, 2023) on the open-source side. Code-execution-in-sandbox is the dominant pattern, with Python-on-Jupyter-kernel or DuckDB-in-WASM as the execution substrate.

Voice and Conversational Agents

Real-time voice agents are an emerging fifth-or-sixth major use case, distinguished by sub-500ms turn latency requirements. OpenAI Realtime API (October 2024) provides speech-to-speech with tool use; Anthropic Voice (private beta 2025); ElevenLabs Conversational AI (London 2024) powering Air India and 1000+ enterprise deployments; Deepgram Aura, Cartesia Sonic, Rime AI voice synthesis; Vapi, Retell AI, Bland AI, Sindarin orchestration. The architecture sandwiches a streaming ASR (Whisper, Deepgram Nova-3) → LLM (with tool calls) → TTS, with VAD (Silero, WebRTC VAD) for turn detection. PolyAI London (Imperial spin-out) and Hume AI target the enterprise voice-agent market.

Industrial and Embodied Agents

Agentic robotics is the long-horizon frontier. Physical Intelligence (1.5B 2024) general-purpose robot brain; 1X Technologies humanoid Neo; Figure humanoid Figure 02 with OpenAI integration. NVIDIA Isaac Groot (Sim-Real foundation models 2024); DeepMind RT-2 and AutoRT (2023-2024). Frameworks: LeRobot (HuggingFace 2024) for robotic policy training; ROS 2 integration with LLM agents at planning layer. UK presence: Wayve London (Series C $1B 2024) self-driving foundation model; Dyson Robotics Lab (Imperial College); Oxford Robotics Institute.

Cybersecurity Agents

Specialised category combining defensive (SOC automation: Dropzone AI, Crogl, Prophet Security, Salem Cyber, Andesite) and offensive (red-team: XBOW, Pentera AI, Horizon3.ai NodeZero, HackerOne Hai) agents. CISA / NCSC warnings 2025 about agent-driven autonomous attack chains. Anthropic’s claude-code-security-reviewer and OpenAI’s o3 for code-security review represent the white-box review category.

Academic Context

  • Agent research at the LLM scale crystallised around four canonical papers, all 2022-2023:
    1. ReAct (Yao et al. ICLR 2023, arXiv:2210.03629, 3,000+ citations). Synergising reasoning and acting in language models. The agent interleaves Thought / Action / Observation tokens; the LLM emits thoughts in natural language and actions as structured calls. ReAct outperformed chain-of-thought-only baselines on HotpotQA, Fever, ALFWorld, and WebShop. It remains the implicit default loop in essentially every framework.
    2. Reflexion (Shinn et al. NeurIPS 2023, arXiv:2303.11366). Verbal reinforcement learning. After each trajectory, the agent reflects in natural language on what went wrong and appends the reflection to a memory buffer used in the next trial. Improved HumanEval pass@1 from 80% (GPT-4 baseline) to 91%. The pattern is now bundled into LangGraph (self-correcting agents), CrewAI (delegation with reflection), and OpenAI Agents SDK (guardrails-with-retry).
    3. Toolformer (Schick et al. NeurIPS 2023, arXiv:2302.04761). Self-supervised insertion of API calls during LLM fine-tuning. The model learned where to invoke calculator, calendar, QA, search, and translation tools by generating, executing, and self-rating call positions. Toolformer pre-dated and informed OpenAI function-calling (June 2023) and Anthropic tool-use (November 2023).
    4. Plan-and-Solve / Plan-and-Execute (Wang et al. ACL 2023, BabyAGI, AutoGPT 2023). The planner-executor split has theoretical antecedents in classical AI planning (STRIPS, HTN) but was crystallised for LLMs in this 2023 wave.
  • Supporting works span: Voyager (Wang et al. 2023) lifelong-learning Minecraft agent with iterative skill library and self-verification; Generative Agents (Park et al. UIST 2023, Stanford-Google) Smallville simulation with 25 agents demonstrating emergent social behaviour from memory-reflection-planning; Tree of Thoughts (Yao et al. NeurIPS 2023); Self-Refine (Madaan et al. NeurIPS 2023); SWE-agent (Yang et al. NeurIPS 2024) Agent-Computer Interface for software engineering; MetaGPT (Hong et al. ICLR 2024) multi-agent software-engineering company simulation; AutoGen (Wu et al. ICLR 2024 paper backing the Microsoft framework); MemGPT (Packer et al. arXiv:2310.08560) virtual context management; Tool-LLM (Qin et al. ICLR 2024) ToolBench dataset.
  • Evaluation literature consolidated around: SWE-bench (Jimenez et al. ICLR 2024, arXiv:2310.06770) with the Verified subset (OpenAI 2024) the de facto coding-agent benchmark; GAIA (Mialon et al. ICLR 2024, arXiv:2311.12983) Meta general-assistant questions; AgentBench (Liu et al. ICLR 2024) Tsinghua, eight LLM-as-agent environments (OS, DB, Knowledge Graph, Digital Card Game, Lateral Thinking Puzzle, House-holding, Web Shopping, Web Browsing); ToolBench for tool-use breadth; WebArena (Zhou et al. ICLR 2024) and VisualWebArena (Koh et al. ACL 2024); OSWorld (Xie et al. NeurIPS 2024) realistic computer-use; AssistantBench (Yoran et al. EMNLP 2024). τ-bench (Tau-Bench, Sierra 2024) measures multi-turn tool-using customer-service agents.
  • The theoretical contributions sit on classical foundations from agent-oriented AI: Russell & Norvig AIMA canonical agent taxonomy (simple reflex, model-based, goal-based, utility-based, learning); BDI (Belief-Desire-Intention, Bratman 1987, Rao & Georgeff 1995); HTN (Hierarchical Task Network) planning; STRIPS planning operators; Subsumption Architecture (Brooks 1986); Society of Mind (Minsky 1986). LLM agents replace formal symbolic state representation with natural-language state and replace search-based planning with autoregressive prediction conditioned on system prompt and observation history—at the cost of formal guarantees and the gain of unprecedented flexibility.

Empirical Findings (2024-2026)

A series of 2024-2026 papers established empirical regularities that shape framework design:

  • Long-horizon failure is the dominant failure mode. Anthropic’s Many-shot Jailbreaking (Anil et al. 2024), Sierra’s τ-bench (Yao et al. 2024), and METR’s autonomous-task evaluations consistently show that task success degrades roughly exponentially with horizon length. An agent that succeeds 95% per step succeeds only 36% across 20 steps. Frameworks respond with explicit replanning, verifier loops, and checkpointing.
  • Tool-use accuracy plateaus around model capability. Toolformer-style fine-tuning and OpenAI’s function-calling RLHF push tool-call schema-adherence above 95% for strong models, but argument hallucination on unfamiliar tools remains 10-30%. ToolBench and BFCL (Berkeley Function-Calling Leaderboard, 2024) track this metric.
  • Prompt-injection is unsolved. Greshake et al. (2023) Not What You’ve Signed Up For and the prompt-injection literature show that any agent reading untrusted content (web pages, emails, documents) is exploitable; Simon Willison coined the term and tracks live attacks at simonwillison.net. Mitigations (spotlighting, dual-LLM, signed prompts) reduce but do not eliminate risk. Anthropic’s Computer Use system card 2024 documents explicit prompt-injection acknowledgements.
  • Multi-agent does not always help. Du et al. (2023) and Xiong et al. (2024) found multi-agent debate improves on some tasks (factuality, math) but degrades on others (creative generation), and adds 3-10x cost. Single-agent loops with strong tool use often match multi-agent collaboration at lower cost.
  • Reasoning models reshape agent design. OpenAI o1 (September 2024), o3 (December 2024), Claude 3.7 Sonnet extended thinking (February 2025), DeepSeek R1 (January 2025), Gemini 2.5 Pro thinking (March 2025) demonstrate that internal chain-of-thought from RL’d reasoning models often substitutes for external scaffolding. The 2026 question: when does a reasoning model with tool use suffice, and when does explicit multi-step orchestration add value? Empirical answer (mid-2026): reasoning models eliminate planner-executor split for short-medium horizons (≤10 steps) but agentic harnesses remain essential for tasks requiring durable state, human approval, multi-day execution, or coordination across systems.

Current Landscape (2026)

  • The 2026 agent-framework market has stratified into clear tiers with distinguishable production fitness:

Tier 1 — Production Default (Enterprise)

LangGraph is the dominant production runtime for stateful agents at scale: 34% of Gartner-tracked agent-architecture documents, deployed at LinkedIn, Uber, Klarna, Elastic, Replit, Norwegian Air, and hundreds of other enterprises. The combination of explicit graph state, durable checkpointing, time-travel debugging, and human-in-the-loop interrupts maps cleanly onto enterprise audit-and-rollback requirements. OpenAI Agents SDK and claude-agent-sdk are the provider-locked Tier 1 entries for teams committed to OpenAI or Anthropic respectively. Both gained Tier 1 status in 2025 by demonstrably matching LangGraph capabilities (handoffs, guardrails, tracing, MCP). Microsoft Semantic Kernel retains a position in .NET / enterprise Microsoft shops, with Agent Framework launched as the unified successor to Semantic Kernel + AutoGen pre-1.0 (October 2024). Released v1.0 stable mid-2025.

Tier 2 — Prototyping and Multi-Agent

CrewAI dominates the rapid-prototyping segment: 45.9K stars, 60% Fortune 500 penetration claimed, accessible role-playing crew abstraction. Teams routinely prototype on CrewAI and migrate to LangGraph for production hardening. AutoGen v0.4 (Microsoft) holds research-grade multi-agent mindshare, particularly through Magentic-One. LlamaIndex agentic-workflows for RAG-heavy use cases.

Tier 3 — Specialist

DSPy for teams adopting the optimisation-as-prompting philosophy; Letta for stateful-memory-first agents; Pydantic-AI for type-strict Python; BAML for prompt-as-typed-function; Smolagents for code-action agents; Mastra for TypeScript / Next.js; Vercel AI SDK for streaming UI; Inngest for durable workflow substrate; agentic-flow for swarm orchestration.

Tier 4 — Emerging and Niche

Google ADK (April 2025) for Vertex AI-committed shops; AWS Bedrock Agents and Strands Agents (May 2025) for AWS-committed shops; HuggingFace Smolagents for HF-ecosystem. PocketFlow, Atomic Agents, PhiData (now Agno), Swarms (Kye Gomez), OpenAgents (XLang Lab) populate the long tail.

Protocol Layer

MCP (Anthropic 26 November 2024) achieved unprecedented adoption velocity: 100 servers (Nov 2024) → 1K (Feb 2025) → 16K+ (late 2025) with 97M monthly SDK downloads (Dec 2025). OpenAI, Google, Microsoft adoption by mid-2025 cemented it as the de facto LLM-to-tool standard. The June 2025 specification revision added Streamable HTTP, OAuth-based authorisation, elicitation (server-initiated questions back to user), and sampling (server-initiated LLM calls). A2A (Google April 2025) and ACP (IBM January 2025) are the two main inter-agent protocols; consolidation discussions through 2026 may merge them under Linux Foundation governance. AGNTCY (Cisco March 2025) provides higher-level Internet-of-Agents reference architecture; GoAgent, AgentNetworkProtocol (ANP), and OAS (Open Agent Specification) populate the inter-organisation niche.

Benchmark Landscape and Leaderboard Movement (mid-2026)

SWE-bench Verified leaderboard: Claude 4.5 Sonnet (Anthropic) 77.2%, GPT-5 (OpenAI) 74.9%, Gemini 2.5 Pro (Google) 71.4%, DeepSeek V4 67.8%. Climbing 6x in 30 months. GAIA: top frontier-model+agent-harness systems exceed 60% (human baseline 92%); Magentic-One reached 38%, OpenAI Deep Research 67.4% on Level 2/3 questions. OSWorld computer-use: Anthropic Claude Sonnet 4.5 with Computer Use 34.4%, OpenAI Operator 38.1%, frontier ceiling still well below human 72%. τ-bench: customer-service agent multi-turn tool use, Sierra-internal Tau-Bench-Airline leaderboard topping 60% with frontier models.

Investment and Market Sizing

Agent-framework / agent-platform investment in 2024-2026 exceeded 4.5B), Anysphere/Cursor (4B), Adept (Amazon acqui-hire 2024), Inflection (Microsoft acqui-hire 2024). Gartner forecasts the agentic AI market reaches 200B+ by 2030, with the framework / platform layer capturing 15-25% of that envelope. McKinsey 2025 projects 4.4 trillion USD economic value created annually by generative-AI agents across knowledge work by 2030.

Contrasts: What Agent Frameworks Are Not

Agent frameworks contrast pointedly with three older categories that they partially substitute and partially complement:

  • Classical Workflow Engines (Apache Airflow, Prefect, Dagster, Temporal, Argo Workflows, AWS Step Functions, Camunda BPMN): deterministic DAG executors with strong scheduling, retry, and observability guarantees but no LLM reasoning. Agent frameworks borrow durable-execution concepts (Temporal’s deterministic replay informed Inngest, Restate, Resonate); workflow engines incorporate LLM steps as nodes. The boundary blurs: Mastra Workflows, LangGraph Platform, and Inngest are explicitly hybrid. The conceptual distinction is who decides what happens next — workflow engines: the developer at design time; agent frameworks: the LLM at run time.
  • Robotic Process Automation (UiPath, Automation Anywhere, Blue Prism, Microsoft Power Automate): screen-scraping deterministic GUI automation. Brittle to UI changes, expensive to maintain, but auditable. UiPath shipped UiPath Autopilot and agent capabilities in 2024; Microsoft Power Automate added Copilot-driven AI Flows. The market consolidation question through 2026: do RPA vendors successfully transition to agent platforms, or do agent platforms (Salesforce Agentforce, ServiceNow Now Assist) eat the RPA market?
  • Rule-Based Expert Systems / Production Rule Engines (CLIPS, Drools, OpenL Tablets, IBM ODM, Red Hat Decision Manager): forward-chaining or backward-chaining inference over hand-authored rules. Highly auditable, fast, deterministic, but expensive to author and brittle to specification gaps. Agent frameworks subsume rule authoring through natural-language constraints, at the cost of probabilistic execution. Hybrid systems combine LLM agents with declarative-rule guardrails (Guardrails AI, NVIDIA NeMo Guardrails, RailsAI) — a pattern likely to dominate regulated industries (healthcare, finance, government) through 2026-2028. The deeper architectural distinction is state representation: classical engines use formal-symbolic state (typed objects, relational tuples, BPMN data objects); agent frameworks use natural-language state (conversation history, system prompts, retrieved chunks) that the LLM interprets per turn. The trade-offs are precisely the symbolic-versus-connectionist trade-offs that have animated AI since Newell and McCarthy: agent frameworks lose formal-verification tractability and gain unprecedented flexibility.

UK Context: Imperial / Edinburgh / UCL / Cambridge / Manchester

Academic Institutions

Imperial College London (Department of Computing, I-X): I-X is Imperial’s AI/X initiative cutting across departments. Murray Shanahan (Professor of Cognitive Robotics, also DeepMind Senior Research Scientist) writes prominently on LLM cognition. Marc Deisenroth’s Statistical Machine Learning Group works on RL-agent foundations. Aldo Faisal’s Behaviour Analytics Lab applies agents to clinical decision-making. Imperial’s £10M+ UKRI grants (2023-2026) include “Trustworthy AI Agents” and integration with Imperial-X agent benchmarks. University of Cambridge (Computer Laboratory and Department of Engineering): Cambridge Machine Learning Group continues the Hinton-Ghahramani-MacKay tradition. José Miguel Hernández-Lobato works on Bayesian deep learning relevant to uncertainty-aware agents. Pieter Buteneers’s spin-out works on conversational agents. Cambridge AI Centre for Doctoral Training (CDT) and Accelerate Programme for Scientific Discovery have agentic-discovery strands. University College London (UCL DARK and UCL AI Centre): DARK (Decision, Action, and Reasoning Knowledge) lab founded by Tim Rocktäschel (now DeepMind/Cohere) and Edward Grefenstette (now Anthropic) is the UK’s deepest LLM-agent research lineage. UCL produced foundational work on language-conditioned RL agents, NetHack Learning Environment (Küttler et al. NeurIPS 2020), and reasoning-action coupling. Sebastian Riedel and Pasquale Minervini work on retrieval-augmented agents. University of Edinburgh (School of Informatics, ILCC): Mirella Lapata, Frank Keller, Adam Lopez, and Ivan Titov anchor Edinburgh’s NLP-agent research. The Edinburgh-CDT in NLP and the AI Hub generate sustained agent-research throughput. Edinburgh’s robotics agent work (Sethu Vijayakumar) integrates with industrial partners. Alan Turing Institute (London + Manchester node): The national institute for data science and AI established the Manchester Alan Turing Institute node in 2024 with an agents-and-autonomy research strand. The London headquarters at the British Library hosts the AI Standards Hub (in partnership with NPL and BSI) and the Public Policy Programme working on AI agent governance. University of Manchester (Department of Computer Science, Centre for AI Fundamentals): André Freitas’s Reasoning and Explainable AI group works on agent reasoning interpretability. Manchester’s relationship with the Alan Turing Institute Manchester node and Health Innovation Manchester create healthcare-agent translation pipelines. University of Oxford (Department of Computer Science, Future of Humanity Institute closure 2024): Oxford remains an agent-safety thought-leader through the Centre for AI Safety, Oxford Internet Institute work on AI agents and labour markets, and the Computing Lab’s verification work. Yoshua Bengio’s UK affiliations and the UK AI Safety Institute (AISI, established 2023, based London) anchor the UK in agent-evaluation governance. UCL UCREL (note: UCREL is at Lancaster University, not UCL — UCREL = University Centre for Computer Corpus Research on Language; Lancaster’s corpus and computational-linguistics work feeds into NLP agent evaluation). Lancaster’s Data Science Institute (DSI) hosts industrial agent-deployment research.

UK Industry

Anthropic London (King’s Cross office, opened 2024): a major non-US Anthropic presence, with public-policy and research staff, supporting Claude and Claude Code roadmap. Cohere London (Edward Grefenstette, ex-UCL/Meta): research lab maintaining LLMs (Command R, Command R+) and Cohere North enterprise platform with agent capabilities; £500M+ Series D 2024. Stability AI London (Emad Mostaque era ended 2024; Sean Parker leadership 2024+): primarily generative-media but expanding into agent-of-the-creative-pipeline category. Synthesia London (40M Series C 2024. Faculty AI London (consultancy): builds bespoke agent systems for NHS, MoD, and FTSE 100 customers; AI-Assistant programme for HM Government. Faculty / Quantexa / DeepMind / Wayve / Isomorphic Labs: London’s AI ecosystem provides agent talent and adjacent platforms (DeepMind’s AlphaEvolve, Isomorphic’s AlphaFold-derived drug-discovery agents). BenevolentAI (London NASDAQ:BAI): drug-discovery agents (the Baricitinib COVID-19 repurposing pre-dated the LLM-agent wave but informed it). Improbable London (founded by Herman Narula): multi-agent simulation infrastructure pivoting to AI-agent applications in 2025.

Northern English Innovation Hubs

Manchester (Health Innovation Manchester, MediaCityUK, NWC AI Strategy): Manchester is the UK’s second AI cluster, anchored by the University of Manchester, Alan Turing Institute Manchester node, Health Innovation Manchester, and BBC R&D at MediaCityUK. Coding-agent and clinical-assistant deployments at Manchester NHS trusts piloted in 2025. The Christie NHS Foundation Trust deploys radiology-prioritisation agents. Leeds (Leeds Teaching Hospitals, University of Leeds, Leeds Innovation Health Hub): NHS England’s national HQ in Leeds drives healthcare-agent adoption. University of Leeds AIRC (Artificial Intelligence Research Centre) collaborates with Leeds Teaching Hospitals on pathology and surgical-video analysis agents. Sheffield (Sheffield University, AMRC, Digital Catapult Northern): Advanced Manufacturing Research Centre (AMRC) applies industrial agents to manufacturing; University of Sheffield NLP group (Diana Maynard, Mark Stevenson) continues 25-year NLP heritage that now applies to agentic information extraction. Newcastle (Newcastle University, Open Lab, NE LEP): Open Lab applies HCI-aware agent design to public-service agents; National Innovation Centre for Data co-locates agent-data integration. Liverpool (University of Liverpool, KQ Liverpool): Liverpool’s John Lennon-era humanities-meets-tech focus has expanded into agentic-content tooling for the cultural sector.

UK Regulatory and Policy

UK AI Safety Institute (AISI, established November 2023, based London/Bletchley Park): the world’s first government AI safety body, conducting pre-deployment evaluations of frontier models including agentic capabilities. Inspect (UK AISI open-source eval framework, 2024) became the de facto agent-eval harness used internally by Anthropic, OpenAI, and DeepMind for safety testing. The May 2024 AISI Bletchley Declaration and the November 2024 Paris AI Summit follow-ups codified agent-evaluation expectations for frontier model deployment. Department for Science, Innovation and Technology (DSIT): leads the UK’s pro-innovation AI regulation approach articulated in the AI White Paper (March 2023) and AI Action Plan (January 2025), with the AI Opportunities Action Plan (Matt Clifford, January 2025) shaping agent-adoption strategy. The 50-recommendation Clifford report explicitly identified agentic AI as a UK strategic priority, with funding for sovereign compute (AIRR — AI Research Resource at Bristol Isambard-AI and Cambridge Dawn) and the AI Growth Zones programme directing agent-platform investment to Culham, Northumberland, and other regions. Cabinet Office GDS (Government Digital Service): piloted Humphrey civil-service AI agents (a suite including Consult for public-consultation analysis, Parlex for parliamentary monitoring, Redbox for ministerial submissions, and Lex for legislation analysis) and Caddy chatbot for benefits advice; the i.AI incubator builds agent applications across departments under Mike Bracken’s reform legacy. Information Commissioner’s Office (ICO) published the Generative AI Consultation Series outcomes (2024-2025) including specific guidance on agents that process personal data, drawing the boundary between data-protection-by-design requirements and emergent agent behaviour. Competition and Markets Authority (CMA) AI Foundation Models update (April 2024) and follow-on report (2025) tracked agent-platform market concentration, particularly around the foundation-model layer. FCA (Financial Conduct Authority) Digital Sandbox: agent-based financial-services innovations tested with regulatory observation; the FCA’s October 2024 AI Update set out principles for agentic systems in financial markets including identity, auditability, and consumer-duty obligations. National Cyber Security Centre (NCSC): 2024-2025 advisories on AI agent risks (prompt injection, supply-chain risk in MCP servers, agentic-attack tools).

Future Directions (2026-2030)

2026-2027: Reliability and Long-Horizon Tasks

The dominant 2026-2027 research priority is making agents reliable across long-horizon tasks. Current SWE-bench Verified at 77% leaves an asymmetric failure mode: when agents fail, they often fail catastrophically (deleting files, corrupting databases, infinite loops). Verifier-prover decomposition: separating the agent that proposes actions from a verifier that checks them, with formal-methods integration for high-stakes domains. Process supervision: rewarding intermediate steps (Lightman et al. NeurIPS 2024) rather than only outcomes. Speculative-decoding-style action draft: cheap agents propose, expensive agents verify. Anthropic’s constitutional-AI principles applied to agents: explicit constraints rather than implicit RLHF.

2026-2028: Protocol Consolidation

MCP for LLM↔tool is settled. A2A versus ACP for agent↔agent remains contested through 2026 with consolidation likely by 2027 under Linux Foundation governance. Agent identity and authorisation is the unsolved problem: OAuth flows for agents, scoped delegations, audit trails, and agent-impersonation prevention. NIST and ENISA agent-security standards work begins 2026; ISO/IEC 42001 (AI management systems) extends to agents.

2027-2028: Multi-Agent Markets and Negotiation

Inter-agent marketplaces where agents discover, negotiate with, and contract other agents. Agent payments via Anthropic Stripe MCP server (2025) and similar primitives evolve into agent-native commerce. Auction theory and mechanism design return to relevance. Cisco Outshift / AGNTCY work on agent identity and observability becomes load-bearing infrastructure. Cross-organisation agent contracts likely require new legal categories.

2027-2029: Vertical Specialisation

Generalist agent frameworks (LangGraph, OpenAI Agents SDK) cap at ~70% reliability on heterogeneous tasks; vertical-specialist frameworks emerge in healthcare (deep clinical-pathway integration), legal (jurisdiction-aware reasoning), finance (regulatory compliance baked in), and scientific discovery (laboratory automation integration). Frameworks bundle domain-specific evaluation harnesses, regulator-approved prompt libraries, and vetted tool corpora.

2028-2030: Agent-Native Operating Systems

The endpoint of the trajectory is an agent-native OS layer where agents are first-class citizens of the operating system, not applications running on it. Anthropic Claude Computer Use, OpenAI Operator, and Google Project Mariner are first-generation hints. Apple’s on-device intelligence (Apple Intelligence) and Google Pixel’s Gemini Nano agentic features blur the OS / agent boundary. The 2028-2030 platform competition is between cloud-agent ecosystems (Microsoft, Google, Anthropic, OpenAI) and device-agent ecosystems (Apple, Google, Samsung) with consumer regulatory questions around delegation, consent, and liability driving policy. Microsoft’s Recall (2024 controversy), Apple’s Apple Intelligence (June 2024 announcement, mid-2025 rollout of Siri agent capabilities), and Google’s Project Astra (May 2024 → Gemini Live, 2025) point at the consumer-OS battle. On Linux the Anthropic Computer Use Reference Implementation and Open Interpreter hint at an open-source-friendly path.

2026-2028: Agentic Code Replacing Traditional Software

Cursor, Devin, Claude Code, Replit Agent, and Copilot Coding Agent collectively process tens of millions of agentic developer turns daily by 2026. The 2026-2028 inflection is the migration of these from developer-assistance (a human-in-the-loop agent suggesting changes) to substitutive agents (an agent that completes whole tickets autonomously, with human review at PR level). Cognition’s Devin, Anthropic’s claude-code in autonomous mode, and OpenAI’s Codex (May 2025 cloud-based engineering agent re-using the Codex brand from 2021) instantiate this pattern. Anthropic’s Agentic Misalignment (June 2025) and Alignment Faking (December 2024) research, plus Apollo Research’s evaluations, signal that the safety-research community treats agent-coding substitution as the highest-stakes near-term agentic deployment.

2027-2030: Education, Training, and the Labour Market

The labour-economics question — which knowledge-work jobs do agents partially or wholly displace, on what timeline, with what wage and productivity effects — is the question that determines whether the agent-framework category remains a 4.4T McKinsey envelope. Erik Brynjolfsson’s MIT economics work, Tyler Cowen’s GOAT and follow-on work, the IMF April 2024 report (40% of global employment exposed), and the UK Department for Education’s Skills Imperative 2035 framework converge on agent-augmentation in the 2025-2027 window and partial substitution in the 2028-2030 window for routine cognitive tasks. UK industrial strategy responses include the AI Opportunities Action Plan (January 2025), the Industrial Strategy Council’s 2024 review, and the Skills England (established 2024) workforce-transition mandate.

Persistent Open Problems

  • Evaluation in the wild: benchmarks saturate; real-world deployment success metrics remain hard to standardise.
  • Cost and latency: agentic loops cost 10-100x raw model inference; durable agents amortise compute differently.
  • Failure mode characterisation: prompt injection, tool-use hallucination, runaway loops, context-window exhaustion, compounding error.
  • Energy and environmental impact: agentic compute is the fastest-growing component of LLM datacentre load.
  • Labour market displacement: 30-50% of knowledge-work tasks plausibly delegable to agents by 2030 (Brookings, McKinsey, MIT studies 2024-2026) raising profound social policy questions the UK AI Opportunities Action Plan begins to address.

Research and Literature

Foundational Agent Papers (2022-2023):

  1. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023. arXiv:2210.03629.
  2. Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. arXiv:2303.11366.
  3. Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. NeurIPS 2023. arXiv:2302.04761.
  4. Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R.K.-W., & Lim, E.-P. (2023). Plan-and-Solve Prompting. ACL 2023. arXiv:2305.04091.
  5. Park, J.S., O’Brien, J.C., Cai, C.J., Morris, M.R., Liang, P., & Bernstein, M.S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023. arXiv:2304.03442.
  6. Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291.
  7. Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Zhang, C., et al. (2024). MetaGPT: Meta Programming for Multi-Agent Collaborative Framework. ICLR 2024. arXiv:2308.00352.
  8. Wu, Q., Bansal, G., Zhang, J., Wu, Y., Zhang, S., Zhu, E., Li, B., Jiang, L., Zhang, X., & Wang, C. (2024). AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. ICLR 2024. arXiv:2308.08155.
  9. Packer, C., Wooders, S., Lin, K., Fang, V., Patil, S.G., Stoica, I., & Gonzalez, J.E. (2023). MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560.
  10. Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., et al. (2023). Self-Refine: Iterative Refinement with Self-Feedback. NeurIPS 2023. arXiv:2303.17651.

Tool-Use and Function Calling: 11. Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., et al. (2024). ToolLLM: Facilitating Large Language Models to Master 16000+ Real-World APIs. ICLR 2024. arXiv:2307.16789. 12. Patil, S.G., Zhang, T., Wang, X., & Gonzalez, J.E. (2023). Gorilla: Large Language Model Connected with Massive APIs. NeurIPS 2024. arXiv:2305.15334. 13. Yang, J., Jimenez, C.E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. NeurIPS 2024. arXiv:2405.15793.

Evaluation Benchmarks: 14. Jimenez, C.E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR 2024. arXiv:2310.06770. 15. Mialon, G., Fourrier, C., Swift, C., Wolf, T., LeCun, Y., & Scialom, T. (2024). GAIA: A Benchmark for General AI Assistants. ICLR 2024. arXiv:2311.12983. 16. Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., et al. (2024). AgentBench: Evaluating LLMs as Agents. ICLR 2024. arXiv:2308.03688. 17. Zhou, S., Xu, F.F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Bisk, Y., Fried, D., Alon, U., & Neubig, G. (2024). WebArena: A Realistic Web Environment for Building Autonomous Agents. ICLR 2024. arXiv:2307.13854. 18. Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T.J., Cheng, Z., Shin, D., Lei, F., et al. (2024). OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. NeurIPS 2024. arXiv:2404.07972. 19. Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H.P.O., et al. (2021). Evaluating Large Language Models Trained on Code (HumanEval). arXiv:2107.03374.

Protocols and Standards: 20. Anthropic (2024). Introducing the Model Context Protocol. https://www.anthropic.com/news/model-context-protocol (26 November 2024). 21. Anthropic (2025). Code execution with MCP: building more efficient AI agents. https://www.anthropic.com/engineering/code-execution-with-mcp 22. Model Context Protocol Specification. https://modelcontextprotocol.io/specification (June 2025 revision: Streamable HTTP, OAuth, elicitation, sampling). 23. Google (2025). Announcing the Agent2Agent Protocol (A2A). https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability/ (April 2025; Linux Foundation donation June 2025). 24. IBM Research (2025). Agent Communication Protocol (ACP) Specification. BeeAI Project, Linux Foundation. https://agentcommunicationprotocol.dev/ 25. AGNTCY Collective (2025). Internet of Agents Reference Architecture. https://agntcy.org/

Framework Documentation and Industry Reports: 26. LangChain documentation. https://python.langchain.com / https://docs.langgraph.dev (2026 stable release notes). 27. Microsoft AutoGen v0.4 Documentation. https://microsoft.github.io/autogen/ (January 2025 release). 28. Anthropic claude-agent-sdk Documentation. https://docs.anthropic.com/en/api/agent-sdk (September 2025).

Metadata

  • Last Updated: 2026-05-16
  • Review Status: Comprehensive editorial review for Phase 6 enrichment.
  • Verification: Framework adoption metrics validated via May 2026 WebSearch; academic papers cross-referenced against arXiv canonical IDs; protocol release dates verified against vendor announcements (Anthropic MCP 26 Nov 2024, IBM ACP Jan 2025, Google A2A April 2025).
  • Regional Context: UK academic institutions (Imperial I-X, Cambridge CL, UCL DARK, Edinburgh ILCC, Alan Turing Institute London + Manchester node, Oxford), UK industry (Anthropic London, Cohere London, ElevenLabs, Synthesia, PolyAI, Faculty AI), Northern England innovation hubs (Manchester, Leeds, Sheffield, Newcastle, Liverpool) detailed.
  • Domain Validation: Domain artificial-intelligence retained (original stub already correct); IRI rewritten from generic ontology#AgentFrameworks to canonical artificial-intelligence#AgentFrameworks matching Phase 6 namespace conventions; legacy-term-id AI-1212 assigned.
  • Production-Ready: Complete OWL formal semantics with 6 axiom families, exhaustive framework taxonomy (single-agent / graph-state / multi-agent / specialist), protocol layer (MCP/ACP/A2A/AGNTCY) coverage, benchmark landscape, UK regional context, 2026-2030 forward projection.
  • Authority Score: 0.87 (mature ontology entity backed by 700+ active frameworks, three major standardisation protocols, established academic literature, $20B+ venture funding, 50K+ production deployments).

Provenance