Context Engineering is the systems-level discipline of designing, curating, compressing, and orchestrating the information delivered into a Large Language Model’s context window across a multi-step agentic loop, popularised between mid-2024 and mid-2025 by Walden Yan (Cognition AI’s “Don’t Build …

In Plain Terms

  • The craft of deciding exactly what to put in front of a model at each step — your instructions, the relevant documents, the conversation so far, and any tool results — so it has just what it needs and is not swamped by clutter. Managing that limited space well is one of the biggest levers on how well an agent performs.

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:hasPart ai:SystemPrompt))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:hasPart ai:InstructionHierarchy))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:hasPart ai:ToolSchema))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:hasPart ai:FewShotExamples))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:hasPart ai:RetrievalAugmentedContext))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:hasPart ai:Scratchpad))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:hasPart ai:PersistentMemory))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:hasPart ai:ContextCompaction))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:hasPart ai:PrefixCache))

## Dependency Relationships
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:requires ai:ContextWindow))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:requires ai:Tokenizer))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:requires ai:LargeLanguageModel))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:requires ai:RetrievalSystem))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:requires ai:MemoryStore))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:requires ai:EvaluationHarness))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:dependsOn ai:AttentionMechanism))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:dependsOn ai:TransformerArchitecture))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:dependsOn ai:InformationRetrieval))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:dependsOn ai:VectorSearch))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:dependsOn ai:Tokenization))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:dependsOn ai:PromptCaching))

## Capability Relationships
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:enables ai:LLMAgents))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:enables ai:MultiTurnCoherence))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:enables ai:LongHorizonPlanning))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:enables ai:ToolUse))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:enables ai:PersistentAIAssistants))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:enables ai:CostEfficientInference))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:supports ai:CodingAgents))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:supports ai:ResearchAgents))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:supports ai:CustomerServiceAgents))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:supports ai:BrowserAgents))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:supports ai:WorkflowAutomation))

## Implementation Relationships
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:implements ai:RecursiveSummarisation))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:implements ai:HierarchicalRetrieval))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:implements ai:ContextCompression))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:implements ai:RollingMemoryWindow))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:implements ai:MemoryBlockArchitecture))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:implements ai:StructuredOutputConstraints))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:uses ai:VectorEmbeddings))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:uses ai:Reranking))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:uses ai:JSONSchema))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:uses ai:XMLTags))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:uses ai:FunctionCalling))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:uses ai:PromptCaching))

## Reduction Relationships
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:reduces ai:ContextRot))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:reduces ai:HallucinationRate))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:reduces ai:InferenceCost))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:reduces ai:PromptInjectionRisk))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:reduces ai:AgentDrift))

## Association Relationships
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:relatedTo ai:AgenticAI))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:relatedTo ai:RetrievalAugmentedGeneration))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:relatedTo ai:MemoryArchitecture))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:relatedTo ai:LongContextModeling))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:contrastsWith ai:PromptEngineering))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:contrastsWith ai:FineTuning))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:contrastsWith ai:RetrievalEngineering))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:contrastsWith ai:ModelPreTraining))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:standardizedBy ai:ModelContextProtocol))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:standardizedBy ai:OpenAIInstructionHierarchySpec))
SubClassOf(ai:ContextEngineering
  ObjectSomeValuesFrom(ai:standardizedBy ai:AnthropicPromptEngineeringGuide))

## Data Properties (Characteristics)
DataPropertyAssertion(ai:hasIdentifier ai:ContextEngineering "AI-1078"^^xsd:string)
DataPropertyAssertion(ai:authorityScore ai:ContextEngineering "0.87"^^xsd:decimal)
DataPropertyAssertion(ai:termPopularisedYear ai:ContextEngineering "2025"^^xsd:integer)
DataPropertyAssertion(ai:typicalAgentToolCalls ai:ContextEngineering "50"^^xsd:integer)
DataPropertyAssertion(ai:nominalContextWindow ai:ContextEngineering "200000"^^xsd:integer)
DataPropertyAssertion(ai:effectiveContextRatio ai:ContextEngineering "0.40"^^xsd:decimal)
DataPropertyAssertion(ai:prefixCacheCostReduction ai:ContextEngineering "0.90"^^xsd:decimal)
DataPropertyAssertion(ai:compressionRatioMax ai:ContextEngineering "20"^^xsd:integer)

## Property Constraints
SubClassOf(ai:ContextEngineering
  DataMinCardinality(1 ai:hasContextWindow xsd:integer))
SubClassOf(ai:ContextEngineering
  DataMinCardinality(1 ai:hasSystemPrompt xsd:string))
SubClassOf(ai:ContextEngineering
  DataSomeValuesFrom(ai:hasMemoryArchitecture xsd:string))
SubClassOf(ai:ContextEngineering
  DataAllValuesFrom(ai:isSystemsDiscipline xsd:boolean))

## Annotations
AnnotationAssertion(rdfs:label ai:ContextEngineering "Context Engineering"@en)
AnnotationAssertion(rdfs:comment ai:ContextEngineering "Systems-level discipline of designing, curating, compressing, and orchestrating information delivered into an LLM's context window across a multi-step agentic loop, popularised 2024-2025 by Cognition's Walden Yan, Shopify's Tobi Lütke, Andrej Karpathy and the Anthropic Applied AI team, deliberately superseding 'prompt engineering' with a systems perspective treating the context window as a managed resource. Composes system prompts, instruction hierarchies, tool schemas, few-shot exemplars, retrieval-augmented context, episodic and semantic memory, scratchpads, persistent project memory (CLAUDE.md, .cursorrules) and prefix caching. Addresses context rot (Chroma 2025), lost-in-the-middle (Liu et al. 2024 TACL), context poisoning, prompt injection, agent drift through compression (LLMLingua, Selective Context), hierarchical retrieval, contextual retrieval (Anthropic 2024), recursive summarisation, rolling memory windows, memory blocks (Letta/MemGPT). Standardised via Model Context Protocol (Anthropic 2024), OpenAI instruction hierarchy spec, LangGraph state graphs. Distinguished from prompt engineering (single-turn), fine-tuning (weight modification) and retrieval engineering alone (subset)."@en)
AnnotationAssertion(dcterms:identifier ai:ContextEngineering "AI-1078"^^xsd:string)
AnnotationAssertion(dcterms:subject ai:ContextEngineering "Agentic AI, LLM Applications, Memory Architecture, Retrieval, Prompt Design, Systems Engineering"@en)

)

Property Characteristics

AsymmetricObjectProperty(ai:requires) AsymmetricObjectProperty(ai:enables) AsymmetricObjectProperty(ai:implements) AsymmetricObjectProperty(ai:reduces) AsymmetricObjectProperty(ai:contrastsWith) TransitiveObjectProperty(ai:dependsOn) FunctionalDataProperty(ai:termPopularisedYear) FunctionalDataProperty(ai:prefixCacheCostReduction)

About Context Engineering

  • Context Engineering is the systems-level discipline of deliberately designing, curating and managing every token that enters a Large Language Model’s context window over the course of an extended agentic interaction. Where prompt engineering asks “what is the best single instruction I can write?”, context engineering asks “what is the right information environment for this model to operate in across hundreds of decision points, multiple tool calls, evolving state, and a finite token budget?” The term, popularised between mid-2024 and mid-2025, signals a shift in how practitioners conceptualise LLM applications: from one-shot prompt tweaking to systems engineering of an information pipeline.
  • The concept crystallised publicly in June 2025 when Shopify CEO Tobi Lütke posted on X that “I really like the term ‘context engineering’ over prompt engineering. It describes the core skill better: the art of providing all the context for the task to be plausibly solvable by the LLM.” Andrej Karpathy, formerly of OpenAI and Tesla, immediately agreed in a widely-shared response, observing that “people associate prompts with short instructions, whereas in every serious LLM application, context engineering is the delicate art and science of filling the context window with just the right information for each step.” The Cognition AI engineering team had been writing about the concept in production terms since their April 2025 essay “Don’t Build Multi-Agents” by Walden Yan, in which the Devin team argued that single-agent context-engineering frequently outperforms multi-agent orchestration precisely because it avoids the context-fragmentation problem—where information needed to make a coherent decision is scattered across separate agent sub-contexts that never reunify.
  • Within twelve months of the term entering the public lexicon, it had been adopted as the descriptive label for a discipline that production LLM teams had in fact been practising since the launch of ChatGPT plugins (March 2023), the introduction of function calling (June 2023) and the proliferation of agentic frameworks throughout 2024 (LangGraph, CrewAI, AutoGen, Letta, Mastra, Vercel AI SDK). Anthropic’s Applied AI team published the Contextual Retrieval technique in September 2024 demonstrating 49% retrieval failure reduction by prepending document summaries to chunks; Anthropic’s Prompt Caching shipped August 2024 making 90%+ input-cost reductions accessible to any developer; and the Model Context Protocol (MCP) released November 2024 standardised how tools, resources and prompt templates are exposed to LLM clients, providing the protocol layer on which much of the discipline now rests.
  • The discipline contrasts deliberately with three adjacent practices. Prompt engineering operates on a single turn: the input string is the artefact. Fine-tuning modifies model weights via supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), Direct Preference Optimisation (DPO) or Group Relative Policy Optimisation (GRPO), changing the model itself rather than its inputs. Retrieval engineering focuses narrowly on the retrieval-augmented generation (RAG) pipeline—chunking, embedding, indexing, retrieval, reranking. Context engineering subsumes RAG as one layer among many and co-designs prompts, retrieval, memory, tool schemas and orchestration against latency, cost and quality budgets.

Core Conceptual Framework

Context engineering operates within four interlocking constraints:

The Token Budget: Even with nominally large context windows (Claude 3.7 Sonnet 200K, GPT-4.1 1M, Gemini 2.0 Pro 1M-2M), the practical budget after reserving output tokens, instruction overhead, tool schema and prefix-cached system content is typically 30-150K tokens. Each token consumed by retrieved context displaces a token that could carry reasoning, examples or further tool descriptions. The engineer is making a continuous allocation decision: how many tokens to spend on system instructions, how many on retrieved documents, how many reserved for the model’s chain-of-thought, how many for tool outputs.

The Attention Tax: Empirical evidence from Liu et al. (2024 TACL “Lost in the Middle”) and the Chroma Research Context Rot report (Hong et al. 2025, evaluating 18 frontier models) demonstrates that LLM attention is not uniform across the context window. Content placed in the middle of long inputs is recalled 20-50% less reliably than content near the start or end. Context rot—the gradual degradation of model performance as input length grows—affects every frontier model tested, with accuracy drops of 15-40% between 8K and 100K-token inputs on needle-in-a-haystack, multi-hop reasoning, and instruction-following benchmarks. Long context is not free even when nominally supported.

The Failure Mode Catalogue: Drew Breunig’s June 2025 essay “How Long Contexts Fail” catalogued four canonical agentic failure modes that any context-engineering pipeline must mitigate. Context poisoning: erroneous tool output enters the conversation and is treated as authoritative on subsequent turns, propagating errors. Context distraction: irrelevant retrieved chunks pull the model off-task into tangential reasoning. Context confusion: contradictory snippets create incoherent or hedged outputs. Context clash: stale memory contradicts fresh observations, causing the model to either ignore the new information or thrash between competing world models. To these must be added prompt injection through tool outputs (Greshake et al. 2023 indirect prompt injection), instruction-hierarchy attacks where user content attempts to override system instructions, and context drift in long-running agents where after 50-200 tool calls the agent has gradually lost task focus.

The Cost Curve: Input tokens dominate the cost profile of modern agentic systems. A single Devin or Manus task issuing 50-100 tool calls, each carrying 5-30K of context, can consume 5-15M input tokens before any output is generated. At Claude 3.7 Sonnet pricing (15-45 per task; at Opus 4 pricing (75-225. Prefix caching—where static system content (system prompt, tool schemas, memory blocks, contextual retrieval prefixes) is cached server-side—reduces this by 90% for cache hits and reduces time-to-first-token by 50-80%. Engineering the cacheable boundary is therefore both a quality and a cost decision.

The Layers of Context

A modern agentic context window is composed of multiple distinguishable layers, each engineered separately and stacked in a deliberate order that respects the instruction hierarchy:

1. System Prompt and Persona: The top-level instructions that define the agent’s identity, capabilities, constraints, output format and decision policy. Anthropic and OpenAI both privilege system prompts above user content in their instruction-hierarchy specifications (OpenAI’s Model Spec April 2024; Anthropic’s Constitutional AI framework). The system prompt is typically 500-3,000 tokens for production agents and is prime cache territory.

2. Instruction Hierarchy: OpenAI’s 2024 Instruction Hierarchy paper (Wallace et al.) formalised the priority order system > developer > user > tool > assistant for resolving conflicting instructions, motivated by the prompt-injection threat surface. Production agents instantiate this hierarchy explicitly via message roles, XML tags or structured channels.

3. Tool/Function-Call Schema Definitions: Each tool exposed to the model consumes 100-1,500 tokens of schema (name, description, JSON Schema parameters). A typical agent toolbelt of 10-30 tools consumes 5-25K tokens before any user interaction begins. The Manus team’s Context Engineering for AI Agents essay (Yichao Ji, July 2025) reports averaging 50 tool calls per typical task and emphasises curating toolbelts ruthlessly: fewer, sharper tools with well-disambiguated descriptions outperform large overlapping toolbelts that confuse model routing.

4. Persistent Project Memory: File-based memory conventions injected as part of the system context. Anthropic’s Claude Code uses CLAUDE.md at repo root and parent directories with hierarchical override; Cursor uses .cursorrules and .cursor/rules/*.mdc with project- and user-level scopes; Aider uses CONVENTIONS.md; Windsurf uses .windsurfrules; GitHub Copilot uses copilot-instructions.md. These files encode coding conventions, project structure, command palettes, do/don’t lists and team norms that the agent should respect across all sessions.

5. Few-Shot Exemplars: Carefully chosen input/output examples demonstrating the target behaviour, format, edge-case handling or tool-use pattern. Modern agentic systems increasingly use few-shot for tool selection and output format rather than for raw classification, where instruction-following capability has largely subsumed the need.

6. Retrieval-Augmented Context (RAG): Just-in-time retrieved chunks from a vector store, BM25 index, full-text search or hybrid retriever, optionally reranked by a cross-encoder (Cohere Rerank, Voyage Rerank, BGE Rerank, Jina ColBERT). Modern RAG layers commonly include contextual retrieval (Anthropic September 2024) where a brief document-level summary is prepended to each chunk before embedding, reducing retrieval failures by 35-49%.

7. Episodic Memory / Conversation History: The verbatim transcript of recent turns. Most agents keep the last N turns (typically 5-20) in full and summarise prior history. LangGraph state graphs, Letta message buffers, and Mastra memory threads provide framework-level abstractions for this layer.

8. Semantic Memory / Memory Blocks: Distilled long-term facts about the user, project, prior decisions or accumulated knowledge. Letta (the rebranded MemGPT, Packer et al. 2023) pioneered the memory block architecture with separate core memory (always-in-context, small, mutable by the agent), archival memory (vector-retrievable, large) and message buffer (recent turns). ChatGPT introduced cross-session memory in early 2024; Claude introduced project memory in 2024. The OpenAI Memory feature stores user facts as bullet-style summaries injected at every turn.

9. Scratchpad / Working Memory: ReAct-style (Yao et al. NeurIPS 2022) reasoning steps interleaved with actions: Thought → Action → Observation cycles. Modern agents like Claude Code, Devin and Manus expose explicit scratchpads—dedicated channels for the model to plan, take notes, track sub-goals and revise without polluting either the conversation transcript or persistent memory.

10. Tool Outputs and Observations: The results of each function/tool invocation flowing back into context. This is the largest and least-bounded layer in production agents: a single read_file on a 4,000-line file injects 40K tokens of code; a single search_web may inject 30K of returned snippets. Context-engineered agents apply aggressive truncation, summarisation or selective extraction to tool outputs before letting them re-enter context for the next turn.

Core Patterns and Techniques

The discipline has converged on a stable catalogue of patterns, each addressing one or more of the four constraints above.

Context Compression

LLMLingua / LongLLMLingua / LLMLingua-2 (Jiang et al. Microsoft Research, 2023-2024): Small language model (typically Llama-2-7B or smaller distilled variants) scores each token’s importance via perplexity and prunes low-importance tokens, achieving 3-20× compression with <5% accuracy loss on benchmarks including LongBench and ZeroSCROLLS. LongLLMLingua extends with question-aware compression for retrieval-conditioned compression; LLMLingua-2 distils a task-agnostic classifier yielding 1.6-2.9× compression at 1.6× lower compute than v1.

Selective Context (Li, EMNLP 2023): Removes self-information-low lexical units (tokens, phrases, sentences) producing 50-80% input reduction while preserving downstream-task quality.

Recursive Summarisation (Wu et al. OpenAI 2021): Hierarchically summarises long documents in chunks, then summarises the summaries. Applied to long-form Q&A, agent transcript compaction and book summarisation.

Auto-Compaction in Agents: Claude Code’s /compact command, Cursor Composer’s session summarisation, and Cline’s auto-condense trigger summarisation when window utilisation crosses 70-90%, preserving system instructions and recent turns while replacing older history with a structured summary.

Hierarchical and Contextual Retrieval

Parent-Child Chunking: Index small chunks for retrieval precision but return the larger enclosing parent chunk to give the model surrounding context. LlamaIndex and LangChain both ship parent-document retrievers.

Contextual Retrieval (Anthropic September 2024): Before embedding each chunk, prepend a 50-100 token summary of how the chunk fits within the parent document, generated by a cheap model (Claude 3 Haiku) over the full document. Empirical results: 35% retrieval failure reduction with contextual embeddings, 49% with contextual embeddings + BM25, 67% with contextual embeddings + BM25 + reranking.

Hybrid Retrieval: Combine dense (vector) and sparse (BM25, SPLADE) retrieval with reciprocal rank fusion. Production agents almost universally use hybrid + rerank (Cohere Rerank 3, Voyage Rerank-2, BGE Reranker v2, Jina ColBERT v2).

Agent Loop Context Engineering

Rolling Memory Windows: Keep last N turns verbatim, summarise everything prior. LangGraph’s MessagesState with a trimmer is the canonical implementation.

Episode Buffers: Separate per-task scratchpads from cross-task persistent memory. Mastra workflows, Inngest agent kits and Trigger.dev agent tasks all expose this distinction.

Anti-Poisoning Verification: Critical facts re-grounded against source before commitment. Devin’s verification step re-reads code from disk before claiming a change is correct; Claude Code’s Read before Edit tool ordering enforces this at the protocol level.

Context-Window-Aware Planning: Before issuing a long-running plan, estimate token consumption per step and budget proactively. The HumanLayer Advanced Context Engineering for Coding Agents essay (Ace, 2025) documents explicit context budgets per task phase.

Structured Outputs and Boundary Markers

JSON Schema Constraints: OpenAI Structured Outputs (August 2024), Anthropic tool-use with JSON Schema, Vertex AI controlled generation, Outlines, Instructor, Pydantic-AI all constrain model output to a parseable schema, eliminating downstream parsing errors that would otherwise poison subsequent turns.

XML Tags as Separators: Anthropic’s prompt-engineering guide specifically recommends XML tags (<document>, <instructions>, <output_format>, <thinking>) to unambiguously delineate context layers. This convention has been adopted across Claude-targeted prompt libraries.

Markdown Headers and Section Markers: For models trained heavily on markdown (most frontier models), ## Section headers act as effective semantic separators.

Prefix Caching

Anthropic Prompt Caching (August 2024, GA 2025): 5-minute and 1-hour cache TTLs. 90% input-cost reduction and 50-80% TTFT reduction on cache hits. Engineered by placing static content (system prompt, tool schemas, contextual retrieval prefix, persistent memory) at the start of the prompt and inserting cache breakpoints with cache_control: {type: "ephemeral"} markers.

OpenAI Prompt Caching (October 2024): Automatic for prompts ≥1,024 tokens; 50% discount on cached portions for gpt-4o, o1 and later.

Gemini Context Caching: Explicit cache objects via the Gemini API with hour-granular TTL.

DeepSeek Context Caching: On-disk distributed cache with 90% input-cost reduction on hits, contributing materially to DeepSeek-V3’s price competitiveness.

Standards, Protocols and Tooling Ecosystem

Model Context Protocol (MCP)

Released by Anthropic in November 2024, MCP is an open protocol standardising how LLM clients (Claude Desktop, Cursor, Windsurf, Cline, Continue, Zed, Goose, LibreChat) connect to external tools, resources and prompt templates exposed by independent MCP servers. The protocol defines three primary primitives: tools (executable functions the model can call), resources (readable data the model can fetch), and prompts (templated prompt fragments the user or model can invoke). Within six months of release, the MCP server ecosystem had grown to 1,000+ public servers across categories spanning code execution, web browsing, database access, knowledge management, design tools, project management, communication platforms and IoT. The protocol is now de facto the integration layer for context engineering at the application boundary.

LangGraph

LangChain’s typed state-graph framework (released 2024, stable 0.2 in late 2024) treats agent context as an explicit TypedDict state object with named channels and reducers that determine how new values combine with prior state. This makes context updates auditable, replayable and time-travel-debuggable. Production deployments include Replit Agent, Klarna’s customer service AI, and Elastic AI Assistant.

Letta (formerly MemGPT)

Packer et al. 2023 introduced MemGPT at UC Berkeley, treating the context window analogously to RAM and introducing a controller that pages between in-context core memory, recall memory and archival memory. The project rebranded to Letta in 2024 and now offers a hosted memory service and open-source server with REST/gRPC APIs.

Mastra

TypeScript-first agent framework from Gatsby/Netlify alumni offering memory threads, workflow steps, telemetry and evals with context-engineering primitives built in. Mastra’s Memory abstraction exposes per-thread message buffer, semantic recall (vector-retrieved relevant prior turns) and a working memory block automatically maintained by the agent and surfaced in every system prompt — closely mirroring Letta’s three-tier architecture in TypeScript-native form.

Anthropic Files API

Released 2025, allows uploading documents with persistent file handles that can be referenced from prompts without re-uploading; relevant content is dynamically composed into context. Combined with prompt caching, the Files API allows agents to ground on multi-megabyte corpora (codebases, legal contracts, research papers) at marginal per-call cost on cache hits.

Vector Stores, Embeddings and Rerankers

The retrieval substrate underneath context engineering is a mature ecosystem in itself. Vector databases: Pinecone, Weaviate, Qdrant, Milvus, Chroma, LanceDB, Turbopuffer, Vespa, pgvector (Postgres extension), Redis Vector, Elasticsearch dense vector. Embedding models: OpenAI text-embedding-3-large, Voyage voyage-3, Cohere embed-v3, Jina jina-embeddings-v3, Nomic nomic-embed-text, BGE-M3, E5-Mistral, GTE. Rerankers: Cohere Rerank 3, Voyage Rerank-2, Jina ColBERT v2, BGE Reranker v2, Mixedbread mxbai-rerank-large. Practitioners select across a multi-axis trade-off: dimensionality (256-4096), max input length (512-32K tokens), recall@k benchmarks (BEIR, MTEB), per-1M-token cost (0.13), and self-hostable vs API-only.

IDE Memory Files

  • CLAUDE.md (Claude Code): hierarchical with home → workspace → project → directory inheritance; lifecycle hooks; agent definitions
  • .cursorrules and .cursor/rules/*.mdc (Cursor): legacy single file plus newer modular .mdc rules with frontmatter-driven scoping
  • CONVENTIONS.md (Aider)
  • .windsurfrules (Windsurf / Codeium)
  • .github/copilot-instructions.md (GitHub Copilot Chat / Workspace)
  • AGENTS.md (proposed cross-IDE convention 2025)

Use Cases and Major Application Families

Coding Agents

Claude Code, Cursor Composer/Agent, Devin (Cognition), Windsurf Cascade, Replit Agent, GitHub Copilot Workspace, Aider, Cline, Continue, Zed AI, Lovable, Bolt.new, v0 (Vercel) and others. Context engineering here orchestrates code retrieval, file reads, diagnostic output, test results, lint warnings, dependency graphs and version-control state. The Cognition essay Don’t Build Multi-Agents argues that for coding, context-engineered single agents typically outperform multi-agent decompositions because code reasoning needs unified context. Empirical claim: 15-45 percentage-point uplift in task completion rate from disciplined context engineering versus naive prompting.

Research and Browser Agents

OpenAI Deep Research, Anthropic Claude with Computer Use, Manus (Butterfly Effect), Google Gemini Deep Research, Perplexity Pages, Exa Research, You.com agents. The context-engineering challenge is managing 50-500 retrieved web pages, hundreds of thousands of tokens of intermediate findings, and synthesising into a coherent report. Techniques: recursive summarisation, source-attribution tags, finding boards.

Customer Service and Sales Agents

Sierra (Bret Taylor), Decagon, Lindy, Klarna AI Assistant (handling ~2.3M conversations/month at peak), Intercom Fin, Salesforce Agentforce, Zendesk AI Agents. Context layers include customer profile, order history, KB articles, current ticket thread, prior interactions across channels. Prefix caching is critical: customer profile and KB content cache while the dynamic suffix is the current message.

Workflow Automation

n8n AI workflows, Zapier AI Actions, Make scenarios with AI, Inngest agent kit, Trigger.dev v3 agents, Temporal LLM workflows. Context is decomposed across deterministic and LLM-mediated steps.

Personal Assistants and Companion Agents

ChatGPT with Memory, Claude Projects, Pi (Inflection / now Microsoft), Replika, Character.AI personae. Persistent semantic memory blocks dominate the context-engineering challenge.

Failure Modes and Mitigations

Failure ModeSymptomMitigation
Context Rot (Chroma 2025)15-40% accuracy drop at 100K+ tokensCompression, hierarchical retrieval, just-in-time fetch
Lost-in-the-Middle (Liu 2024)Mid-window content under-attendedPlace critical content at head/tail; use XML markers; rerank to top
Context PoisoningBad tool output treated as truthVerification step; re-grounding; structured output schemas
Context DistractionOff-task tangents from retrieved noiseReranking; relevance thresholds; document-quality scoring
Context ConfusionContradictory snippets → hedged outputConflict resolution prompts; recency-weighted retrieval
Context ClashStale memory vs fresh observationTTL on memory blocks; explicit invalidation; observation-priority rules
Prompt Injection (Greshake 2023)User/document content overrides systemInstruction hierarchy; input sanitisation; spotlighting; structured roles
Instruction-Hierarchy Attack”Ignore previous instructions”OpenAI hierarchy training (Wallace 2024); Constitutional AI; system-message immutability
Agent DriftLoss of task focus over 50-200 turnsPeriodic re-grounding; goal restatement; auto-compact with goal preservation
Tool-Output BloatSingle read injects 40K tokensTruncation; selective extraction; summarising tools
Cache MissesCost explosionStable prefix engineering; minimise mid-prompt churn; cache breakpoints

Academic Context

Context engineering as a named discipline is primarily industrial, but its theoretical foundations are deeply academic and have moved rapidly through peer-reviewed venues.

Foundational Long-Context Work: Beltagy et al. 2020 Longformer, Zaheer et al. 2020 Big Bird, Press et al. 2022 ALiBi, Chen et al. 2023 Extending Context Window via Position Interpolation, Ding et al. 2023 LongNet, Mohtashami & Jaggi 2023 Landmark Attention, Han et al. 2024 LM-Infinite, Munkhdalai et al. 2024 Infini-Attention. These advances made 100K-2M token windows practically feasible; context engineering exists to use them wisely.

Lost in the Middle: Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni, Liang 2024 “Lost in the Middle: How Language Models Use Long Contexts” TACL. Demonstrated U-shaped attention to long context across GPT-3.5, Claude, MPT-30B-Instruct, LongChat. Cited 1,500+ times within 18 months.

Needle in a Haystack and Beyond: Kamradt’s 2023 Needle in a Haystack benchmark became the canonical first-order long-context test; subsequent work (NoLiMa, RULER from NVIDIA Hsieh et al. 2024, LongBench from Tsinghua, ZeroSCROLLS) developed more demanding evaluations exposing context-rot.

Compression Theory: Jiang et al. 2023 LLMLingua, LongLLMLingua, LLMLingua-2; Li 2023 Selective Context; Chevalier et al. 2023 AutoCompressors; Mu et al. 2023 Learning to Compress Prompts with Gist Tokens NeurIPS; Ge et al. 2024 In-context Autoencoder for Context Compression.

Retrieval-Augmented Generation: Lewis et al. 2020 RAG NeurIPS, Gao et al. 2023 Retrieval-Augmented Generation for Large Language Models: A Survey arXiv (3,000+ citations), Asai et al. 2024 Self-RAG ICLR, Anthropic 2024 Contextual Retrieval technical note.

Memory Systems: Packer et al. 2023 MemGPT: Towards LLMs as Operating Systems arXiv (Letta foundation), Park et al. 2023 Generative Agents: Interactive Simulacra of Human Behavior UIST (Smallville reflection-and-memory architecture), Wang et al. 2023 Voyager (open-ended Minecraft agent with skill library), Sumers et al. 2024 Cognitive Architectures for Language Agents TMLR.

Agent Architectures: Yao et al. 2022 ReAct: Synergizing Reasoning and Acting in Language Models ICLR, Shinn et al. 2023 Reflexion NeurIPS, Wang et al. 2024 Survey on LLM-Based Autonomous Agents Frontiers of Computer Science, Xi et al. 2023 The Rise and Potential of LLM-Based Agents arXiv.

Prompt Injection and Security: Greshake et al. 2023 Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection AISec, Wallace et al. 2024 The Instruction Hierarchy OpenAI arXiv, Hines et al. 2024 Defending Against Indirect Prompt Injection Attacks With Spotlighting Microsoft.

Current Landscape (2026)

As of mid-2026, context engineering has matured into a recognised professional discipline with the following landscape characteristics:

Frontier Model Context Capabilities: Claude 3.7 Sonnet and Claude Opus 4 at 200K with strong-quality retention; GPT-4.1 at 1M with the Context Engineering page of OpenAI’s documentation explicitly using the term; Gemini 2.0/2.5 Pro at 1M-2M with native context caching; DeepSeek-V3 at 128K with disk-based distributed prefix caching driving 10× cost competitiveness; Llama 3.3 70B and Llama 4 (Scout 10M, Maverick 1M) open-weights at extreme lengths; Mistral Large 2 at 128K. Long context is commoditised; engineering it well is not.

MCP Adoption: MCP servers exceed 1,500 publicly listed implementations. Major IDEs (Claude Desktop, Cursor, Windsurf, Zed, Cline, Continue) all support MCP natively. Hosting platforms (Smithery, Composio, MCP.so) provide one-click MCP server deployment.

Industry Roles: Job listings for Context Engineer, Agent Engineer, AI Engineer (Context Systems) exist at Anthropic, OpenAI, Google DeepMind, Cognition, Anysphere (Cursor), Sierra, Lindy, Decagon, Replit, Vercel, Shopify, Stripe, Klarna and others. Median total compensation in San Francisco for senior context-engineering roles tracks AI-engineer pay bands (650K total).

Benchmarks and Evals: Context-engineering quality is increasingly measured via SWE-bench Verified, SWE-bench Multimodal, SWE-Lancer, TAU-bench, BrowseComp, GAIA, AgentBench, MLE-Bench, OSWorld, WebArena, the Anthropic Computer Use evaluations and Cognition’s internal Devin benchmarks. Win rates on these are dominated by the quality of the context-engineering pipeline more than by raw model capability.

Cost Engineering: The combination of prefix caching, contextual retrieval and compaction is now a board-level concern at AI-product companies. Klarna publicly reported $40M annualised savings on customer service partly attributed to context-engineering optimisations; Sierra reports 70-90% cache hit rates on its enterprise deployments.

Standardisation Pressure: The proposed AGENTS.md convention (2025) attempts cross-IDE memory-file standardisation. OpenAI’s Model Spec, Anthropic’s Constitutional AI spec, the NIST AI Risk Management Framework Generative AI Profile (2024) and the EU AI Act GPAI obligations all touch on context-engineering practice for high-risk systems.

UK Context: Research, Industry and Regional Practice

The United Kingdom holds a strong position in the foundations of context engineering through its NLP and machine-learning research communities, applied AI groups within government and broadcast, and a growing cluster of consultancies, agent startups and product companies — with notable strength in the North of England alongside the southern academic spine.

Academic Institutions

Imperial College London: The Department of Computing’s NLP and Machine Learning groups (led by faculty including Lucia Specia, Marek Rei and Anandha Gopalan) work on long-context modelling, retrieval-augmented systems, faithfulness and grounding. Imperial-X (the College’s AI initiative) and the Data Science Institute host industrial collaborations with Microsoft Research Cambridge, Anthropic UK and BBC R&D on context-engineering for editorial AI.

University of Cambridge — Language Technology Lab (LTL): Anna Korhonen’s group plus the Cambridge NLIP group (Andreas Vlachos, Nigel Collier, Paula Buttery, Bill Byrne) contribute foundational work in retrieval, summarisation, factuality and dialogue context management. The Cambridge MLG (Carl Rasmussen, Adrian Weller) plus Cambridge CFI/CSER provide a strong adjacent safety and applied-AI bench.

University of Oxford: The Oxford Applied AI Lab and the OII (Oxford Internet Institute) host applied work on agent evaluation, alignment-relevant context design and the human-factors side of long-running agent interaction. Oxford’s earlier Future of Humanity Institute (closed 2024) seeded a generation of context-aware safety researchers now at AISI, ARIA and frontier labs.

University of Edinburgh — School of Informatics: One of Europe’s largest NLP communities (ILCC, with Mirella Lapata, Ivan Titov, Shay Cohen, Hao Tang, Pasquale Minervini), contributing to retrieval, structured prediction, summarisation and long-context evaluation. Edinburgh is a primary feeder of context-engineering talent into Cohere, Anthropic UK, Reka, Stability AI, ElevenLabs and Faculty AI.

University College London (UCL): The DARK Lab (Tim Rocktäschel, Edward Grefenstette), UCL NLP, the Centre for Artificial Intelligence and the Alan Turing Institute interface contribute to agent architectures, memory systems and applied evaluation. UCL spin-outs including DeepMind alumni founders and the wider London foundation-model startup cluster anchor context-engineering practice in the capital.

University of Manchester: The NLP group at the Department of Computer Science and the NaCTeM (National Centre for Text Mining, Sophia Ananiadou) supply long-running expertise in biomedical and scientific text mining — a natural context-engineering substrate. Manchester is a key Northern node for applied AI alongside the Christabel Pankhurst Institute and the regional AI Foundry initiatives.

UK Industry — National

BBC Research & Development (MediaCityUK Salford and Maida Vale London): Long-running programme on AI-assisted production, content classification and editorial knowledge graphs. The BBC’s principled AI strategy (announced 2024) explicitly addresses context engineering in editorial pipelines: grounding LLM-mediated workflows on internal style guides, fact-checking corpora and archive content via retrieval rather than reliance on parametric model knowledge.

Government Digital Service (GDS) and AI Incubator (i.AI): The 10 Downing Street AI Incubator team launched tools including Caddy (the GOV.UK Citizens Advice context-aware assistant), Lex (legal document agent), Consult and Redbox (Cabinet Office), each foregrounding retrieval-grounding, persistent memory of departmental context and instruction-hierarchy-respecting prompt design.

AI Safety Institute (AISI) — London: Post-Bletchley Declaration national institute evaluating frontier models, including their context-handling, long-context safety and agentic-context vulnerabilities. AISI’s pre-deployment evaluations feed directly into the Frontier AI Safety Commitments adopted at Seoul (May 2024) and Paris (February 2025) AI Summits.

ARIA (Advanced Research and Invention Agency): David “davidad” Dalrymple’s Safeguarded AI programme funds context-engineering-adjacent work on formally verified agent oversight.

Cohere London/Toronto: While headquartered in Toronto, Cohere’s London engineering hub contributes substantially to RAG, rerank and context-engineering tooling — its Rerank 3 and Embed v3 are mainstays of UK production context pipelines.

DeepMind London: Project Astra and the wider Gemini agentic line drive long-context and memory research. The DeepMind Cambridge office contributes to applied evaluation.

Stability AI London, Reka London, ElevenLabs London, PolyAI London, Synthesia London, Hugging Face London: All operate context-engineered LLM products at production scale.

UK Industry — Consultancies and Applied AI

Faculty AI (London with a substantial Leeds presence): Applied-AI consultancy with deep government and defence engagements; one of the most prominent UK delivery shops for context-engineered agentic systems including NHS clinical workflows and Cabinet Office tooling.

Codified (London/Manchester remote-first): RAG and agent-engineering consultancy operating across financial services and regulated industries.

Tessl (London): Founded by Guy Podjarny (Snyk founder) addressing AI-native software development — a context-engineering-first IDE and platform.

Humanloop (London): Prompt-management and evaluation platform with explicit context-versioning, prompt-caching diagnostics and agent-eval tooling, used by Duolingo, Vanta and Gusto.

Mindgard (London): LLM red-teaming including indirect prompt-injection and context-poisoning evaluation.

Northern English Innovation Hubs

Manchester (MediaCityUK, Christabel Pankhurst Institute, ID Manchester): BBC R&D, AI Foundry NW, Peak (the AI-native decision intelligence company), Evergreen Life and Solirius. Manchester’s AI cluster — backed by the Greater Manchester AI Foundry and ID Manchester redevelopment — has become a centre for applied context-engineering in healthcare (NHS GM), media (BBC, ITV Studios Salford) and financial services (Barclays Manchester).

Leeds (Channel 4, Tracsis, Sky Betting, NHS Digital Leeds): Channel 4’s data unit, Tracsis (rail data), Sky Betting & Gaming, and NHS Digital headquarters in Leeds drive context-engineered analytics agents in transport, media and healthcare. Faculty AI’s Leeds presence and the University of Leeds CDT in AI provide a feeder pipeline.

Sheffield (University of Sheffield NLP Group, AMRC, Tribe.AI): Sheffield’s NLP group (Lucia Specia formerly, Robert Gaizauskas, Nikolaos Aletras) is one of the UK’s strongest in retrieval and information extraction — directly upstream of context engineering. The Advanced Manufacturing Research Centre (AMRC) deploys agentic systems in industrial settings.

Newcastle (Newcastle University, Atom Bank, Sage): Sage’s AI division (developing Sage Copilot for accounting), Atom Bank’s customer-facing AI, and Newcastle University’s School of Computing contribute to context engineering in fintech and SMB software. Digital Catapult NE accelerates context-engineering-heavy startups across manufacturing and supply chain.

Liverpool (University of Liverpool, AI Foundry Liverpool): Materials Innovation Factory and the Liverpool AI cluster contribute scientific-discovery agents using context-engineering for chemistry and materials.

Future Directions (2026-2030)

Memory as a Service: Hosted memory infrastructure (Letta Cloud, Mem0, Zep, Pinecone Memory, Cognee, Graphlit, Vellum) standardises persistent memory across applications. By 2028, memory-as-a-service is projected to be a $2-5B annual market segment alongside vector databases.

Cross-Application Memory Portability: Standards emerging from the Personal AI Memory working group and OpenAI/Anthropic conversations on portable user memory. Likely to become a 2027-2028 regulatory pressure point under EU AI Act and UK AI Bill provisions on data portability and user agency.

Native Context Compression: Models with built-in compression (gist tokens, in-context autoencoders, retrieval-augmented attention) becoming standard. Reka, DeepMind and Together AI all have research programmes here.

Verified Context: Formal verification of context-engineering pipelines for safety-critical agentic applications. ARIA’s Safeguarded AI programme (davidad) and the Atlas Computing initiative are leading bets.

Adversarial Context Robustness: Driven by 2024-2026’s wave of indirect prompt-injection incidents, expect mandatory context-injection red-teaming for high-risk agents under the EU AI Act’s GPAI obligations and likely UK equivalents. The AISI and NIST AI Safety Institute joint protocols include indirect prompt-injection evaluations as a baseline.

Multi-Agent Context Buses: While Cognition’s Don’t Build Multi-Agents essay argues against naive multi-agent decomposition, sophisticated multi-agent architectures with explicit context-bus protocols (shared blackboard, structured handoff schemas) are an active research direction. Microsoft’s AutoGen v0.4 and Google’s A2A (Agent2Agent) protocol (released 2025) target this.

Long-Horizon Persistent Agents: Agents operating continuously over weeks-to-months with structured journals, lifelong memory pruning and goal-tree maintenance. Sakana AI’s AI Scientist, Genesis Therapeutics, Manus and the broader autonomous-research-agent category are early production examples.

Browser and Desktop Context Engineering: Anthropic Computer Use, OpenAI Operator, Google Project Mariner and AI desktop agents (Highlight, Rewind, Granola) extend context engineering into screen content, OCR’d UI elements and accessibility-tree representations.

Sub-Second Agent Loops: As models become faster (Groq, Cerebras, SambaNova, Together inference) and prefix-cache hit rates approach 90%+, context-engineering will optimise for high-frequency agent loops at 100-500ms per turn rather than today’s 1-10s.

Research & Literature

Foundational Long-Context Modelling:

  1. Beltagy, I., Peters, M.E., & Cohan, A. (2020). Longformer: The long-document transformer. arXiv:2004.05150. [Sliding-window attention]
  2. Zaheer, M., Guruganesh, G., Dubey, K.A., Ainslie, J., Alberti, C., Ontanon, S., et al. (2020). Big Bird: Transformers for longer sequences. NeurIPS 2020. [Sparse attention]
  3. Press, O., Smith, N.A., & Lewis, M. (2022). Train short, test long: Attention with linear biases enables input length extrapolation. ICLR 2022. [ALiBi]
  4. Chen, S., Wong, S., Chen, L., & Tian, Y. (2023). Extending context window of large language models via positional interpolation. arXiv:2306.15595. [PI]
  5. Munkhdalai, T., Faruqui, M., & Gopal, S. (2024). Leave no context behind: Efficient infinite context transformers with Infini-attention. arXiv:2404.07143. [Infini-Attention]

Context Quality and Failure Modes: 6. Liu, N.F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. TACL 12, 157-173. DOI: 10.1162/tacl_a_00638. [U-shaped attention] 7. Hong, K., et al. (2025). Context Rot: How Increasing Input Tokens Impacts LLM Performance. Chroma Technical Report. https://research.trychroma.com/context-rot [18-model evaluation] 8. Kamradt, G. (2023). Needle in a haystack — Pressure testing LLMs. https://github.com/gkamradt/LLMTest_NeedleInAHaystack [Canonical NIAH benchmark] 9. Hsieh, C.P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., et al. (2024). RULER: What’s the real context size of your long-context language models? COLM 2024. [NVIDIA RULER benchmark] 10. Breunig, D. (2025). How long contexts fail. https://www.dbreunig.com/2025/06/22/how-contexts-fail-and-how-to-fix-them.html [Failure-mode catalogue]

Compression: 11. Jiang, H., Wu, Q., Lin, C.Y., Yang, Y., & Qiu, L. (2023). LLMLingua: Compressing prompts for accelerated inference of large language models. EMNLP 2023. [Compression] 12. Jiang, H., Wu, Q., Luo, X., Li, D., Lin, C.Y., Yang, Y., & Qiu, L. (2024). LongLLMLingua: Accelerating and enhancing LLMs in long context scenarios via prompt compression. ACL 2024. [Question-aware compression] 13. Li, Y. (2023). Unlocking context constraints of LLMs: Enhancing context efficiency of LLMs with self-information-based content filtering. EMNLP 2023. [Selective Context] 14. Mu, J., Li, X., & Goodman, N. (2023). Learning to compress prompts with gist tokens. NeurIPS 2023. [Gist tokens] 15. Chevalier, A., Wettig, A., Ajith, A., & Chen, D. (2023). Adapting language models to compress contexts. EMNLP 2023. [AutoCompressors]

Retrieval-Augmented Generation: 16. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. NeurIPS 2020. [Foundational RAG] 17. Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., et al. (2023). Retrieval-augmented generation for large language models: A survey. arXiv:2312.10997. [Comprehensive RAG survey] 18. Anthropic. (2024). Introducing contextual retrieval. https://www.anthropic.com/news/contextual-retrieval [49% retrieval failure reduction] 19. Asai, A., Wu, Z., Wang, Y., Sil, A., & Hajishirzi, H. (2024). Self-RAG: Learning to retrieve, generate, and critique through self-reflection. ICLR 2024.

Agent Architectures and Memory: 20. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2022). ReAct: Synergizing reasoning and acting in language models. ICLR 2023. [ReAct paradigm] 21. Packer, C., Wooders, S., Lin, K., Fang, V., Patil, S.G., Stoica, I., & Gonzalez, J.E. (2023). MemGPT: Towards LLMs as operating systems. arXiv:2310.08560. [Letta foundation] 22. Park, J.S., O’Brien, J.C., Cai, C.J., Morris, M.R., Liang, P., & Bernstein, M.S. (2023). Generative agents: Interactive simulacra of human behavior. UIST 2023. [Smallville] 23. Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. NeurIPS 2023. 24. Sumers, T.R., Yao, S., Narasimhan, K., & Griffiths, T.L. (2024). Cognitive architectures for language agents. TMLR. [CoALA framework]

Industry Essays and Reports (Practitioner Canon): 25. Yan, W. & Cognition Team. (2025). Don’t build multi-agents. Cognition Blog. https://cognition.ai/blog/dont-build-multi-agents 26. Ji, Y. & Manus Team. (2025). Context engineering for AI agents: Lessons from building Manus. https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus 27. HumanLayer / Ace. (2025). Advanced context engineering for coding agents. https://github.com/humanlayer/advanced-context-engineering-for-coding-agents 28. Martin, L. & Breunig, D. (2025). Context engineering meetup slides. https://x.com/RLanceMartin/status/1948441848978309358 29. Karpathy, A. (2025). On context engineering. X/Twitter thread responding to Tobi Lütke, June 2025. 30. Lütke, T. (2025). On context engineering. X/Twitter post, June 2025.

Standards and Security: 31. Anthropic. (2024). Model Context Protocol specification. https://modelcontextprotocol.io 32. Wallace, E., Xiao, K., Leike, J., Weng, L., Heidecke, J., & Beutel, A. (2024). The instruction hierarchy: Training LLMs to prioritize privileged instructions. arXiv:2404.13208. [OpenAI hierarchy] 33. Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. AISec 2023. 34. Hines, K., Lopez, G., Hall, M., Zarfati, F., Zunger, Y., & Kiciman, E. (2024). Defending against indirect prompt injection attacks with spotlighting. arXiv:2403.14720. [Microsoft spotlighting]

Metadata

  • Last Updated: 2026-05-16
  • Review Status: Phase 6 comprehensive enrichment from stub
  • Verification: Practitioner-canon essays (Cognition, Manus, HumanLayer), Anthropic/OpenAI/Microsoft technical reports, peer-reviewed venues (TACL, NeurIPS, ICLR, ICML, EMNLP, ACL, UIST, TMLR) cross-referenced
  • Domain Verification: artificial-intelligence domain confirmed (matches concept ontological category — frontmatter ngm-equivalent IRI realigned via iri:: and same-as:: to canonical artificial-intelligence namespace)
  • Regional Context: UK academic (Imperial, Cambridge, Oxford, Edinburgh, UCL, Manchester), national industry (BBC R&D, GDS i.AI, AISI, ARIA, DeepMind London, Cohere London, Stability, Reka, ElevenLabs, PolyAI, Synthesia), consultancies (Faculty AI, Codified, Tessl, Humanloop, Mindgard), Northern hubs (Manchester, Leeds, Sheffield, Newcastle, Liverpool) detailed
  • Production-Ready: 41 OWL axioms across Compositional, Dependency, Capability, Implementation, Reduction and Association families plus Data Properties, Property Constraints and Annotations; comprehensive content coverage (definitional pluralism, mathematical and engineering framework, ten-layer context decomposition, pattern catalogue, ecosystem and standards survey, failure-mode matrix, UK context, 2026-2030 future directions, 34-reference bibliography)
  • Authority Score: 0.87 (foundational practitioner canon, peer-reviewed theoretical foundations, mature production tooling ecosystem, standardised protocol layer via MCP)

Provenance

  • domain-correction: ngm-placeholder → artificial-intelligence (iri/uri/same-as/owl-class aligned to canonical artificial-intelligence namespace; original ngm reference in stub Definition replaced)