Documentation Generation is the automated production of human-readable technical documentation—API references, code comments, user guides, and release notes—using large language models and natural language generation pipelines. By coupling static analysis, code execution traces, and prompt engineering, these systems reduce the documentation burden on developers while improving consistency and coverage.

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:hasPart ai:CodeSummarisation))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:hasPart ai:DocstringGeneration))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:hasPart ai:APIReferenceGeneration))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:hasPart ai:AbstractSyntaxTree))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:hasPart ai:CallGraphAnalysis))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:hasPart ai:ReleaseNoteGeneration))

Dependency Relationships

SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:requires ai:LargeLanguageModels))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:requires ai:CodeCorpus))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:requires ai:PromptEngineering))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:requires ai:StaticAnalysis))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:dependsOn ai:Transformer))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:dependsOn ai:NaturalLanguageProcessing))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:dependsOn ai:Tokenization))

Capability Relationships

SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:enables ai:AIAugmentedSoftwareEngineering))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:enables ai:DeveloperProductivity))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:enables ai:SoftwareMaintenance))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:enables ai:ProgramComprehension))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:enables ai:KnowledgeGraphConstruction))

Implementation Relationships

SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:implements ai:TextGeneration))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:implements ai:RetrievalAugmentedGeneration))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:implements ai:NaturalLanguageGeneration))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:implements ai:CodeUnderstanding))

Reduction Relationships

SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:reducesTo ai:TextGeneration))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:reducesTo ai:CodeSummarisation))

Support Relationships

SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:supports ai:SoftwareEngineering))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:supports ai:OpenSourceDevelopment))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:supports ai:ContinuousIntegration))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:supports ai:APIDesign))

Usage Relationships

SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:uses ai:AbstractSyntaxTree))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:uses ai:StaticAnalysis))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:uses ai:RetrievalAugmentedGeneration))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:uses ai:PromptEngineering))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:uses ai:FineTuning))

Contrast Relationships

SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:contrastsWith ai:ManualDocumentation))
SubClassOf(ai:DocumentationGeneration
  ObjectSomeValuesFrom(ai:contrastsWith ai:LiterateProgramming))

About

Documentation Generation refers to the automated construction of technical documentation from software artefacts — source code, schemas, configurations, and runtime traces — using AI and natural language processing techniques. Historically, documentation was written entirely by hand, a labour-intensive process that consistently lagged code development: studies repeatedly confirm that a majority of open-source codebases are poorly documented or out-of-date, and that developer time spent reading and understanding undocumented code constitutes the single largest category of wasted engineering effort. A 2024 GitHub study found that developers spend approximately 58% of their time reading and understanding code rather than writing it, a proportion that AI-assisted documentation generation can directly reduce by surfacing intent and behaviour inline.

The field began its modern trajectory in the early 2010s with information-retrieval approaches that matched code tokens to similar, already-documented snippets, using lexical overlap to rank and adapt retrieved descriptions. These early systems — exemplified by Portfolio (McMillan et al., 2011) and HN (Horn et al., 2013) — demonstrated that even naive similarity-based documentation could be useful, but they were brittle against vocabulary mismatch and failed to capture semantic rather than lexical relatedness. The graduated to statistical sequence-to-sequence models that learned to summarise method bodies into short English descriptions, with Iyer et al.’s CODE-NN (2016) establishing the first attention-based neural baseline and Hu et al.’s SBT (2018) encoding Abstract Syntax Tree paths to capture structural semantics. The decisive inflection arrived with the introduction of large pre-trained code-language models — CodeBERT (2020), Codex (2021, GitHub Copilot), CodeT5 (2021), UniXcoder (2022), and StarCoder (2023) — which enabled fluent, contextually grounded docstring and prose generation at scale, with BLEU scores on CodeSearchNet doubling within two years of pre-trained model introduction.

The architecture of contemporary documentation generation systems is layered and modular. At the syntactic layer, language-specific parsers (tree-sitter for multi-language support, ANTLR-based grammars, Roslyn for C#, JavaParser for Java) build Abstract Syntax Tree representations of source files, extracting the structural skeleton of classes, methods, parameters, return types, decorators, exception specifications, and type annotations regardless of surface formatting or whitespace conventions. This AST representation is language-agnostic in form, enabling documentation pipelines to operate uniformly across Python, Java, TypeScript, JavaScript, C++, Go, Rust, and other languages without language-specific prompting logic. At the semantic layer, call-graph and control-flow analysis tools (CodeQL, pyan, Understand) build inter-procedural dependency graphs revealing how functions compose and which callers invoke a given method, providing the wider context that isolated method-level summarisation misses: a function named process_batch generates very different documentation when its callers reveal it is processing HTTP request payloads rather than database rows. At the discourse layer, Large Language Models receive explicitly constructed prompts encoding AST metadata, inlined code snippets, adjacent documentation from enclosing classes and called functions, and project-level dependency graph context, then generate descriptive prose constrained to match the signature’s type annotations, inferred postconditions, and applicable docstring style conventions (Google Python style, NumPy, reST/Sphinx, JSDoc, Rustdoc).

Retrieval-augmented variants additionally query indexed documentation corpora, issue trackers, commit histories, and Stack Overflow Q&A to ground generation in real developer intent rather than surface pattern statistics alone. These systems retrieve semantically similar documented functions using dense vector search over CodeBERT or GraphCodeBERT embeddings, prepend retrieved examples as few-shot demonstrations, and let the LLM adapt the style and structure of the retrieved documentation to the specific function under documentation. This retrieval augmentation substantially reduces hallucination of incorrect parameter semantics or impossible postconditions, because the model is anchored to documented precedents from the same codebase or similar projects.

From a software-engineering perspective, Documentation Generation occupies a critical position in the modern Continuous Integration and developer-experience stack. Platforms such as Mintlify, Swimm, and ReadTheDocs now offer AI-powered writing assistants that auto-draft documentation stubs when new public APIs are committed, flag stale documentation when referenced code changes, and expose Model Context Protocol (MCP) servers so that AI coding tools can query up-to-date documentation during a task without hallucinating API signatures. This feedback loop — where generated documentation feeds back into agentic code generation — is reshaping the relationship between documentation and software artefacts from a passive record into an active executable interface. As of 2026, Mintlify reports that AI agents account for approximately 45% of documentation site traffic, and the adoption of the llms.txt convention allows sites to expose structured documentation specifically formatted for LLM consumption, analogous to robots.txt for search engine crawlers. The OpenAPI 3.1 specification remains the dominant machine-readable API documentation standard, and toolchains like Speakeasy and Fern auto-generate both SDKs and documentation simultaneously from OpenAPI specs, collapsing the traditionally separate documentation authoring step.

The quality evaluation of generated documentation is an active research sub-problem. BLEU (Bilingual Evaluation Understudy) and ROUGE-L are the most widely used automatic metrics, measuring n-gram overlap between generated and reference documentation, but they are known to correlate imperfectly with human judgements of adequacy, correctness, and fluency. Human evaluation studies typically rate documentation along dimensions of adequacy (does it say what the function does?), coverage (does it mention all parameters?), and coherence (is it readable and self-consistent?). The DocAgent (2025) system addresses quality through rubric-guided self-evaluation loops: a chain-of-thought planning agent generates an initial draft, an evaluation agent scores it against a rubric, and a refinement agent revises; the multi-agent loop terminates when the rubric score exceeds a threshold or a fixed iteration budget is exhausted.

Components / Architecture

The architecture of a production documentation generation system comprises the following tightly coupled subsystems:

  • Abstract Syntax Tree (AST) parsers — language-specific tools (tree-sitter for multi-language support, ANTLR-based grammars, Roslyn for C# and .NET, JavaParser for Java, rustc AST for Rust) extract structural metadata: class hierarchies, method signatures, parameter names and types, return type annotations, access modifiers, thrown exceptions, and decorators or attributes. Tree-sitter is particularly important because it produces concrete syntax trees for 40+ languages with a single uniform API and error-tolerant parsing that handles partially complete or syntactically malformed files in real codebases.

  • Call-graph and dependency analysis — static analysers (Understand by SciTools, CodeQL from GitHub, pyan for Python, Roslyn Workspaces for C#) build caller-callee graphs that provide inter-procedural context. Documentation generation systems use these graphs to: (i) identify all call sites of a function and infer usage patterns; (ii) resolve which callers pass null arguments or specific constant literals; (iii) generate “see also” cross-references that link related functions semantically rather than lexically; and (iv) surface the architectural role of a module relative to the surrounding system.

  • Code embedding models — pre-trained code-language models (CodeBERT, GraphCodeBERT, UniXcoder, CodeT5+, StarEncoder) produce dense vector representations of code snippets and documentation strings, enabling semantic similarity search over the codebase’s own documentation corpus. Retrieved similar documented functions are prepended as few-shot in-context examples to steer the generator toward appropriate style, level of detail, and technical vocabulary.

  • Prompt construction pipelines — template engines (Jinja2, Handlebars, custom Python f-strings) assemble structured prompts from AST-extracted metadata, inlined code snippets, retrieved few-shot examples, and docstring skeleton fragments specifying the target style (Google Python style guide, NumPy style, reST/Sphinx format, JSDoc, Rustdoc, Doxygen). The prompt skeleton explicitly lists all parameter names and type annotations, forcing the model to cover all parameters explicitly and reducing omission errors.

  • LLM inference backends — large code models generate docstring prose, parameter descriptions, raises and returns clauses, and usage examples. API-accessed models used in 2026: GPT-4o (OpenAI), Claude 3.7 Sonnet (Anthropic), Gemini 1.5 Pro (Google DeepMind). Open-weight alternatives: DeepSeek-Coder-V2 (236B MoE), CodeLlama-70B, StarCoder2-15B, Qwen2.5-Coder-32B. Choosing between API and open-weight depends on data-sensitivity requirements (on-premise models for proprietary codebases) and cost per function at scale.

  • Post-processing and validation pipeline — generated text is post-processed through: (i) parameter coverage checker (all parameter names from the AST signature must appear in the generated docstring); (ii) style linter (enforcing indentation, section ordering, punctuation); (iii) semantic consistency checker (return type description must be consistent with the declared return type annotation); (iv) readability scorer (Flesch-Kincaid or similar); (v) quality rubric evaluator for documentation rating along adequacy, coverage, and coherence dimensions.

  • Staleness detection and CI integration — Swimm’s “code-coupled documentation” approach embeds AST path hashes (computed from function signatures and structural paths) into documentation artefacts; any code change that alters the hashed paths triggers a stale-documentation warning in CI pipelines. GitHub Actions and GitLab CI support integration, so documentation staleness is surfaced as a check failure on pull requests, preventing merges with stale documentation.

  • MCP documentation servers — Mintlify and similar platforms auto-host Model Context Protocol (MCP) server endpoints for every documentation site, enabling AI coding tools (Cursor, Claude Code, Windsurf, GitHub Copilot Extensions) to perform structured semantic queries against live, up-to-date documentation during code generation tasks. This closes the loop between documentation generation and code generation: the same documentation generated by the pipeline is immediately available as grounding context for the next round of AI-assisted coding.

  • Multi-agent orchestration — DocAgent-style (2025) systems add a planning agent that analyses the entire repository structure, prioritises documentation targets by public API surface and missing coverage, and assigns individual functions to documentation sub-agents that run in parallel, with a central evaluation agent checking quality before committing generated docs to the repository.

    Formal Treatment and Evaluation Metrics

    The documentation generation task is formally defined as a conditional text generation problem: given a code snippet C (typically a function or method body), produce a natural language description D that faithfully captures the function’s purpose, parameters, return behaviour, and side effects. Early systems measured quality with BLEU-4 (the geometric mean of 1- through 4-gram precision with a brevity penalty), borrowed from machine translation evaluation. ROUGE-L (the longest common subsequence F-measure between generated and reference documentation) provides complementary coverage measurement. METEOR and CIDEr are also used in some evaluations.

    The CodeSearchNet benchmark (Husain et al., 2019) standardised evaluation across Python, JavaScript, Ruby, Go, Java, and PHP, providing held-out test sets of function-docstring pairs mined from GitHub. On this benchmark, BLEU-4 scores progressed from approximately 14 (RNN seq2seq baseline in 2018) to approximately 20-22 with CodeBERT-based models (2020-2021) to over 25 with CodeT5 and fine-tuned GPT variants (2022-2023). However, BLEU-4 saturation has raised concerns that the benchmark may be solved at the metric level without capturing the human-perceived quality improvement in generated documentation.

    The Code2Doc dataset (2024) introduced a quality-first curated evaluation protocol: automated filters based on documentation length, uniqueness, and absence of boilerplate (e.g., machine-generated “This method does X” stubs) reduced a raw GitHub corpus to 13,358 high-quality function-documentation pairs across five languages. Models fine-tuned on this quality-filtered dataset showed 29% BLEU and 24% ROUGE-L improvement over models trained on unfiltered corpora of equivalent size, validating that data quality rather than data quantity is the primary determinant of documentation generation performance at current model scales.

    Human evaluation remains the gold standard. DocAgent (2025) uses a rubric with the following dimensions: Completeness (all parameters documented, return value described, exceptions noted), Accuracy (description matches actual code behaviour), Clarity (documentation is unambiguous and grammatically correct), Conciseness (documentation does not include irrelevant information), and Usability (a developer unfamiliar with the codebase can correctly use the function from the documentation alone). Each dimension is scored 1-5; a minimum aggregate score of 4.0 is required for committed documentation.

    Use Cases / Major Families

    Inline docstring generation

    The most prevalent and commercially deployed use case: an IDE assistant or CI bot automatically drafts or completes docstrings when a developer adds a new function or method. GitHub Copilot, Tabnine, JetBrains AI Assistant, Cursor, and Codeium all support inline documentation suggestions. The quality depends heavily on context window size and the richness of surrounding code: methods with informative naming, explicit type annotations, and typed signatures in well-structured files generate far better docstrings than short, ambiguously named helpers in loosely typed codebases. Models with larger context windows (128K+) can incorporate enclosing class documentation, interface definitions, and example usage from test files to produce substantially more accurate inline docs.

    Repository-level documentation

    Tools such as RepoAgent, RepoSummary, and DocAgent (2025) operate at repository scale — parsing the entire dependency DAG and generating a hierarchical documentation site covering modules, classes, cross-cutting architectural concerns, and inter-component relationships. RepoAgent (2024) indexes codebases via AST parsing and caller-callee analysis, feeds LLMs with structured prompts encompassing the project’s full dependency graph, and generates a structured documentation tree with module-level overviews, class inventories, and cross-reference indices. DocAgent extends this with multi-agent planning: a strategist agent surveys the repository, ranks public APIs by importance and documentation gap, and dispatches documentation agents in parallel, achieving higher throughput than sequential single-agent approaches on large monorepos.

    API reference generation

    Given an OpenAPI 3.1 or JSON Schema specification, LLM pipelines produce natural-language descriptions of REST endpoints, request/response parameters, authentication requirements, error codes, and rate limits, often augmented with generated usage examples in multiple programming languages (curl, Python requests, JavaScript fetch, Go net/http). Platforms like Fern and Speakeasy auto-publish versioned, searchable API documentation sites from machine-readable specs and simultaneously generate typed SDKs in six or more languages, collapsing the documentation and SDK authoring steps into a single automated pipeline. This is particularly valuable for API-first companies where documentation and SDK freshness directly affects developer adoption.

    Release note generation

    Commit history, pull request descriptions, Jira ticket summaries, and code diff analyses are fed to LLMs to produce human-readable release notes, categorised by feature additions, bug fixes, performance improvements, deprecations, and breaking changes. GitHub’s Copilot Workspace includes a release-note drafting capability, and tools like What’s New, Release Drafter, and proprietary pipelines at large software companies automate this for every sprint. The quality of release note generation depends critically on the discipline of commit message writing in the upstream development process: repositories using conventional commits (feat:, fix:, chore:, BREAKING CHANGE:) provide structured signals that dramatically improve categorisation accuracy.

    Architecture documentation

    System-level prose — Architecture Decision Records (ADRs), component interaction diagrams as Mermaid or PlantUML markup, runbook snippets, deployment topology descriptions — can be drafted from infrastructure-as-code (Terraform, Pulumi, CDK), Kubernetes manifests, docker-compose files, and system topology metadata by specialised LLM pipelines. Backstage (Spotify’s developer portal framework, now a CNCF project) integrates LLM-generated architecture documentation into its software catalog, providing searchable technical summaries of every registered service without manual authoring.

    Multilingual technical writing

    Enterprises with global developer bases use documentation generation pipelines to produce localised documentation simultaneously in multiple natural languages, leveraging the multilingual capabilities of models such as GPT-4o, Gemini 1.5 Ultra, and Claude 3.7 Sonnet. The English source documentation is generated first, then translated and culturally adapted for Japanese, Korean, Chinese, German, French, Spanish, and other markets. This pipeline is particularly valuable for developer tools companies targeting Asian markets, where localised documentation is a significant adoption driver.

    Test and example generation alongside documentation

    Increasingly, documentation generation pipelines simultaneously generate test cases and code examples as part of the documentation artefact, producing a richer developer resource that demonstrates correct usage, common error cases, and performance characteristics. Tools like pytest-doctests can execute the code examples embedded in generated docstrings as part of CI, providing a lightweight correctness check on the generated documentation itself.

    Knowledge base and ontology population

    At the enterprise level, documentation generation feeds structured knowledge extraction pipelines: generated documentation is parsed to extract entities (class names, method names, parameter names), relationships (calls, extends, implements, uses), and constraints (preconditions, postconditions, invariants), which populate internal knowledge graphs or ontologies. This enables semantic search over codebases, automated impact analysis of API changes, and generation of dependency-aware change summaries for security auditing.

    Academic Context

    Documentation Generation has attracted sustained attention across software engineering, NLP, and program analysis communities. Foundational work on code summarisation appeared at ICSE, ASE, and ICPC conferences from 2010 onwards, establishing the basic task definitions and early evaluation methodologies. Sridhara et al. (2010, ASE) introduced one of the first systems that automatically generated summaries for Java methods by identifying the “primary action” (the most semantically significant statement) and constructing natural language descriptions from it using a context-free grammar over code tokens. McBurney and McMillan (2014) showed that incorporating caller context — information about how and where a function is called — dramatically improves summary quality compared to analysing the function body in isolation, a finding that motivated subsequent call-graph integration in documentation systems.

    Iyer et al. (2016, ACL/EMNLP) introduced CODE-NN, the first attention-based neural model for code summarisation, trained on Stack Overflow code snippets and their accompanying accepted answers. This established the neural sequence-to-sequence paradigm for the task and set the first reproducible neural baseline. Hu et al. (2018, ICSE) proposed the structure-based traversal (SBT) approach that serialises Abstract Syntax Tree paths into sequences using a specific left-to-right DFS traversal, providing structural inductive bias to the sequence model. Wan et al. (2018, ASE) extended this with deep reinforcement learning, using BLEU as a reward signal to guide summarisation towards higher metric scores. Ahmad et al. (2020, ACL) demonstrated that vanilla Transformer architectures without AST-specific modifications substantially outperformed RNN-based approaches on the CodeSearchNet benchmark, suggesting that the pre-training objective matters more than architecture specialisation for standard code summarisation. The CodeSearchNet benchmark (Husain et al., 2019) provided a standard evaluation corpus of 2.3 million code-documentation pairs across Python, JavaScript, Ruby, Go, Java, and PHP from GitHub repositories, enabling reproducible community comparison for the first time.

    The pre-trained model era opened with CodeBERT (Feng et al., 2020, EMNLP), a bimodal model trained simultaneously on programming language and natural language using masked language modelling and replaced token detection objectives over CodeSearchNet, which substantially improved BLEU scores on code summarisation and code search tasks. GraphCodeBERT (Guo et al., 2021, ICLR) extended CodeBERT by incorporating data-flow graph edges extracted from code as additional input, capturing semantic dependencies (variable definitions and uses) beyond surface token sequences, and achieving further BLEU improvements. CodeT5 (Wang et al., 2021, EMNLP) framed both code summarisation and code generation as a unified text-to-text task using a T5-style encoder-decoder architecture, pre-trained on CodeSearchNet plus a C/C++ corpus; by conditioning generation on explicit identifier-awareness tokens, it reduced hallucination of non-existent parameter names. UniXcoder (Guo et al., 2022, ACL) introduced cross-modal contrastive learning, aligning code, AST, and comment representations in a shared embedding space and enabling retrieval-augmented documentation generation within a single architecture.

    Multi-agent documentation systems emerged prominently in 2025: DocAgent (Liang et al., 2025, arXiv:2504.08725) demonstrated that chain-of-thought planning agents using rubric-guided evaluation loops outperform single-pass generation, with a planning agent decomposing the documentation task, a writing agent producing the initial draft, an evaluation agent scoring against a multi-dimensional rubric, and a refinement agent revising iteratively; this multi-agent loop achieved human-adequacy scores above 4.2/5 on Python and TypeScript benchmarks. Code2Doc (arXiv:2512.18748, 2024) contributed a quality-first curated dataset of 13,358 high-quality function-documentation pairs across Python, Java, TypeScript, JavaScript, and C++ extracted from widely used open-source repositories, with automated filters for uniqueness, length, and absence of AI-generated boilerplate; models fine-tuned on this quality corpus showed 29% BLEU and 24% ROUGE-L gains versus models trained on equivalent-size unfiltered data, validating data quality as the primary leverage point. Repository-level summarisation approaches were benchmarked at LLM4Code@ICSE 2025 (arXiv:2604.06793), establishing hierarchical summarisation — where module summaries are generated by aggregating class summaries, which in turn aggregate method summaries — as superior to flat all-at-once summarisation for large enterprise monorepos. The study also showed that local LLMs (DeepSeek-Coder-V2 quantised to 8-bit) approached GPT-4o quality at approximately one-fifth the inference cost for documentation of well-typed, standardly structured code.

    Current Landscape (2026)

    By 2026, documentation generation has become a standard capability embedded in developer toolchains rather than a standalone research artefact. GitHub Copilot (GPT-4o backend) generates docstrings as a first-class feature of VS Code, JetBrains, and Visual Studio. Mintlify reports that AI agents account for approximately 45% of traffic on documentation sites it hosts, reflecting the decisive shift where generated docs serve machines as much as humans. Mintlify auto-hosts an MCP server for every docs site, enabling Cursor, Claude Code, and Windsurf to query current documentation during task execution without hallucinating signatures.

    The Code2Doc dataset (December 2024) established quality-first principles: models trained on filtered, high-quality pairs achieve up to 29% BLEU and 24% ROUGE-L gains over models trained on raw, noisy corpora. Multi-agent pipelines (DocAgent, RepoAgent, RepoSummary) now routinely achieve human-rated adequacy scores above 4/5 on repository-level documentation for well-typed Python and TypeScript codebases.

    The competitive landscape spans: open-weight (DeepSeek-Coder-V2, CodeLlama-70B, StarCoder2, Qwen2.5-Coder); API-accessible (GPT-4o, Claude 3.7 Sonnet, Gemini 1.5 Pro); IDE-integrated (GitHub Copilot, JetBrains AI Assistant, Cursor, Tabnine); and platform (Mintlify, Swimm, ReadTheDocs AI, Fern, Speakeasy, Document360). Swimm’s “code-coupled documentation” concept has influenced a new generation of staleness-aware documentation tools that embed structural fingerprints into doc files and trigger stale warnings in CI pipelines.

    Standards are emerging: the llms.txt convention (analogous to robots.txt) allows documentation sites to expose structured content specifically for LLM consumption, and platforms including Fern, Mintlify, and Readme.com adopted it in 2025. The OpenAPI 3.1 specification remains the primary machine-readable API documentation standard; toolchains like Speakeasy and Fern auto-generate SDKs and documentation simultaneously from OpenAPI specs.

    UK Context

    The UK software engineering and AI research communities have contributed meaningfully to documentation generation and its underpinning program analysis techniques. UCL’s Programming Principles, Logic and Verification group has produced foundational work in program verification and code comprehension that informs documentation quality standards. The University of Edinburgh’s ILCC (Institute for Language, Cognition and Computation) has active threads in code-language models and software NLP. At King’s College London, software engineering research groups have studied developer documentation practices and maintenance costs at scale.

    In the Northern English technology cluster, Sheffield’s Digital Catapult activities and Leeds’ research in human-computer interaction have examined the ergonomics of AI-assisted writing tools, relevant to documentation assistants. Manchester’s Alliance Manchester Business School has studied the productivity economics of documentation quality in enterprise software. Newcastle University’s Open Lab has examined the social and workflow dimensions of documentation as a knowledge artefact in collaborative development.

    On the industry side, UK-headquartered companies including Arm, AVEVA (Cambridge), and Sage Group deploy documentation generation tooling to maintain large C/C++, Python, and .NET codebases with lean technical writing teams. The UKRI-funded AI for Science and Technology programme has included grants examining AI-augmented research software engineering, in which documentation generation is a key deliverable. The UK’s Government Digital Service (GDS) mandates that public-facing digital services publish comprehensive API documentation, creating institutional demand for automated generation tooling in the public sector.

    Key Terminology Glossary

  • Code Summarisation — the sub-task of generating a short natural language description (typically 1-3 sentences) of a single function or method body, without necessarily covering all parameters or return semantics. Distinguished from full docstring generation by scope and format constraints.

  • Docstring — a structured string literal embedded immediately after the declaration of a function, class, or module in source code, conventionally formatted in a recognised style (Google, NumPy, reST/Sphinx, JSDoc, Rustdoc, Doxygen) and extracted by documentation toolchains to build API reference sites.

  • Abstract Syntax Tree (AST) — a tree data structure that represents the syntactic structure of source code without surface-level formatting details (whitespace, comments, parentheses positions). Each node represents a language construct (function declaration, assignment, binary expression, import statement). AST parsing is the entry point for all structural analysis in documentation generation pipelines.

  • Call graph — a directed graph in which nodes represent functions or methods and directed edges represent call relationships (A → B means A calls B). Documentation generation systems use call graphs to enrich method-level documentation with information about callers (who uses this function?) and callees (what does this function depend on?), contextualising the function within the broader system architecture.

  • CodeSearchNet — a benchmark and dataset published by GitHub (Husain et al., 2019) containing 2.3 million function-docstring pairs across Python, JavaScript, Ruby, Go, Java, and PHP, mined from public GitHub repositories. The benchmark defines a standardised evaluation protocol for code summarisation and code search, including held-out test sets stratified by language. It has been the primary benchmark for documentation generation research from 2019 to the present.

  • BLEU score — Bilingual Evaluation Understudy, originally developed for machine translation quality measurement; adapted for code documentation evaluation by measuring n-gram precision (up to 4-gram) between generated and reference documentation with a brevity penalty for short outputs. BLEU-4 is the conventional variant reported in code summarisation papers. Scores on CodeSearchNet Python typically range from 14 (early RNN models) to 25+ (strong LLM-based models), but the metric is known to correlate imperfectly with human judgements of documentation quality, particularly for creative reformulations that are semantically equivalent but lexically distant from the reference.

  • ROUGE-L — Recall-Oriented Understudy for Gisting Evaluation, longest common subsequence variant. Captures recall of reference documentation content in generated text; complementary to BLEU-4 which emphasises precision. Commonly reported alongside BLEU-4 in documentation generation evaluations, with ROUGE-L typically ranging from 30-55 on CodeSearchNet across languages and models.

  • Retrieval-Augmented Generation (RAG) — a generation paradigm in which a retrieval system first fetches relevant reference examples from an indexed corpus, and the retrieved examples are provided as in-context demonstrations to the generative model. In documentation generation, RAG retrieves semantically similar already-documented functions from the codebase or from external documentation corpora, using dense vector search over code embeddings, and prepends them as few-shot examples to the generation prompt.

  • llms.txt — a proposed web standard (analogous to robots.txt) that allows documentation sites to expose a structured, LLM-consumable version of their content at the /llms.txt endpoint. Adopted by Mintlify, Fern, and Readme.com in 2025, it provides a flat, markdown-formatted view of documentation that AI coding assistants can consume efficiently without needing to parse and render HTML documentation sites.

  • Model Context Protocol (MCP) — an open protocol (initially developed by Anthropic, adopted broadly in 2024-2025) that enables AI coding assistants to communicate with external tools and services, including documentation servers, via a standardised structured query interface. MCP documentation server endpoints allow AI agents to perform semantic searches over documentation, retrieve specific API reference entries, and query parameter documentation by name, enabling grounded code generation without hallucinating API signatures.

  • Code2Doc — a quality-first curated dataset (December 2024, arXiv:2512.18748) of 13,358 high-quality function-documentation pairs across Python, Java, TypeScript, JavaScript, and C++, assembled from widely used open-source repositories using automated quality filters that remove noisy, duplicated, or AI-generated boilerplate documentation. Models fine-tuned on Code2Doc achieve 29% BLEU and 24% ROUGE-L improvement over models trained on equivalent-size unfiltered corpora.

  • Docstring staleness — the condition where a function’s documentation describes behaviour that differs from the current code implementation due to code changes that were not reflected in the documentation. Staleness is a chronic problem in real codebases: studies show that 20-30% of docstrings in active repositories describe outdated parameter names, deprecated return values, or removed functionality. Swimm’s code-coupled documentation architecture uses AST structural fingerprints embedded in documentation files to detect staleness automatically.

    Docstring Style Standards and Format Families

    Docstring conventions vary substantially across programming language ecosystems, and documentation generation systems must be capable of targeting the correct style for each language. The major families are:

    Python has the richest ecosystem of competing but formally documented styles. The Google Python Style Guide defines sections for Args (parameter name, type, description), Returns (type, description), Raises (exception type, condition), Yields (type, description, for generators), and Example (code block). NumPy style (used by NumPy, SciPy, Matplotlib, and the scientific Python stack) uses tabular parameter lists with underlined section headers, providing more structured formatting that renders well in both raw text and generated HTML. The reST/Sphinx style (the oldest Python docstring format, used by the Python standard library and Django) uses :param name: description and :type name: type directives compatible with Sphinx’s autodoc parser. PEP 257 provides minimal conventions (module, function, class, and method docstrings) that all three styles comply with. Documentation generation tools must detect which style the existing codebase uses (via sampling and pattern matching) and generate consistent output.

    JavaScript and TypeScript use JSDoc (/** … */ block comments with @param {type} name - description, @returns {type} description, @throws {ErrorType} condition tags). JSDoc is consumed by tsdoc (for TypeScript-specific constructs), JSDoc 4.x, and documentation generation tools (TypeDoc for TypeScript, JSDoc for JavaScript). Modern TypeScript codebases increasingly rely on the type system itself (TypeScript interface declarations, zod schemas) rather than prose type annotations in JSDoc, so documentation generation must extract type information from the TypeScript compiler’s type checker rather than just parsing JSDoc tags.

    Java uses Javadoc (/** … */ with @param name description, @return description, @throws ExceptionType description, @deprecated, @since, @see cross-reference tags), consumed by the javadoc tool and IDEs. Java documentation generation has the longest history, with McBurney et al. (2014) and Moreno et al. (2013) establishing foundational approaches specifically for Java method summarisation.

    Rust uses Rustdoc (/// for item-level doc comments, //! for module-level), with Markdown content rendered by the rustdoc tool into HTML documentation. Rustdoc automatically runs code blocks in doc comments as doctests via cargo test, providing executable documentation; documentation generation systems must ensure that generated code examples are syntactically valid Rust and produce the documented output when compiled.

    C++ uses Doxygen (@brief for one-line summary, @param [in|out|in,out] name description, @return description, @throws ExceptionType condition, @note, @warning, @code/@endcode for examples), consumed by Doxygen, Breathe/Sphinx for Python-style HTML output, and Helix QAC for safety-critical applications (ISO 26262, DO-178C). C++ documentation generation is particularly challenging due to template metaprogramming, SFINAE, constexpr, and other features that are difficult to describe in natural language.

    Go uses Godoc (// comment immediately preceding a declaration, with no formal markup tags; the first sentence becomes the package/function summary). Go’s documentation philosophy emphasises brevity; documentation generation systems must match this minimalist style and avoid over-documentation of simple, self-documenting code.

    Documentation generation pipelines increasingly support multi-language repositories (polyglot monorepos) where Python, TypeScript, Go, and C++ components coexist. This requires language detection per file, style selection per language, and cross-language cross-reference generation for APIs that expose the same functionality across multiple language bindings (a REST API documented in OpenAPI, with language-specific SDKs in Python, TypeScript, and Go, each requiring appropriate docstrings).

    Ecosystem and Standards Landscape (2026)

    The documentation generation ecosystem has converged around several standards and platform categories that deserve detailed treatment.

    The OpenAPI Specification (formerly Swagger), now at version 3.1, is the dominant machine-readable format for HTTP API contracts. An OpenAPI document specifies endpoints, HTTP methods, request parameters (path, query, header, cookie, request body), response schemas, authentication schemes, and error codes in a machine-readable JSON or YAML format. Documentation generation pipelines consume OpenAPI specs to produce natural-language narrative around the structured schema information — turning "type": "string", "format": "date-time" into “ISO 8601 UTC timestamp in the format YYYY-MM-DDTHH:MM:SSZ” with appropriate contextual caveats. The planned OpenAPI 4.0 (Moonwalk) is expected to better support event-driven APIs (AsyncAPI integration) and more expressive JSON Schema constructs.

    The AsyncAPI specification fills a parallel role for event-driven systems — message brokers (Kafka, RabbitMQ, AWS SNS/SQS), WebSocket endpoints, and pub/sub systems. Documentation generation for AsyncAPI schemas produces documentation for message event types, channel names, binding parameters, and message payload schemas, with similar LLM-based approaches to those used for REST APIs.

    Docstring style standards are programming-language-specific but increasingly converging. Python has the most formal standards: Google Style (Args/Returns/Raises/Yields/Example sections), NumPy Style (tabular parameter lists with type and description columns), and reST/Sphinx (the oldest, used by CPython and the Python standard library). JavaScript and TypeScript use JSDoc (/** … / block comments with @param, @returns, @throws tags). Rust uses Rustdoc (//! for module docs, /// for item docs, with markdown content and automated doctest execution). Java uses Javadoc (@param, @return, @throws). C++ uses Doxygen (/// or /* */ with @brief, @param, @return). Documentation generation pipelines must target the appropriate style for each language and optionally enforce style consistency across a mixed-language monorepo.

    Hugging Face Code Models hub provides the central repository for open-weight code-language models used in documentation generation. As of mid-2026, the most capable open-weight models for documentation generation include DeepSeek-Coder-V2 (236B MoE, Apache 2.0 licence), StarCoder2-15B (with The Stack v2 training data), Qwen2.5-Coder-32B (strong coding performance), and CodeLlama-70B (Meta AI, with instruction-tuned variant optimised for code explanation). These models are deployable on-premise on servers with 4-8 H100 GPUs, enabling documentation generation for proprietary codebases where sending code to external API providers is prohibited by information security policy.

    CI/CD integration patterns for documentation generation have standardised around: (1) documentation stub generation as a PR check — when a function with no existing docstring is added to a PR, the CI system auto-comments a draft docstring for developer review; (2) staleness detection as a required check — PR merges that modify function signatures without updating docstrings fail the CI pipeline; (3) documentation coverage reporting — percentage of public functions with complete docstrings tracked as a project quality metric alongside test coverage; (4) documentation deployment — merged documentation changes automatically trigger documentation site rebuild and deployment via Mintlify, ReadTheDocs, or GitHub Pages.

    Benchmark Datasets for Documentation Generation Research

    The field has converged around several standard benchmark datasets, each with distinct characteristics:

  • CodeSearchNet (Husain et al., 2019) — 2.3 million code-documentation pairs from GitHub across Python (457K), JavaScript (148K), Ruby (53K), Go (347K), Java (405K), and PHP (897K). The standard evaluation split (CodeSearchNet Challenge) provides train/valid/test sets for both code summarisation (given a function, generate the docstring) and code search (given a natural language query, retrieve the most relevant function). The CodeSearchNet Python test set has 14,918 function-docstring pairs and is the most widely reported benchmark for documentation generation. BLEU-4 and ROUGE-L are the standard metrics; human evaluation studies confirm that BLEU-4 above approximately 20 on CodeSearchNet Python corresponds to documentation that developers rate as “adequate” or better on a 5-point scale.

  • Code2Doc (2024) — 13,358 high-quality function-documentation pairs across Python, Java, TypeScript, JavaScript, and C++, filtered for uniqueness, length, and quality from widely used open-source repositories. The quality-filtered nature of this dataset makes it better than CodeSearchNet for fine-tuning purposes but smaller for pre-training. Models fine-tuned on Code2Doc achieve 29% BLEU and 24% ROUGE-L improvement over models trained on equivalent-size unfiltered corpora.

  • CONCODE (Iyer et al., 2018) — 100K Java code-documentation pairs from GitHub, paired at the class environment level (documentation of a class method with the full class context provided), enabling evaluation of context-aware documentation generation. CONCODE has a natural language description → code direction (code generation) as well as code → description direction (summarisation).

  • TL-CodeSum — a dataset of 90,102 pairs of GitHub issue titles and corresponding Java method summaries; provides a different granularity of documentation (issue-level versus method-level) and tests whether models can generate issue-resolution documentation.

  • HumanEval-Docs (evaluated in multi-agent documentation studies, 2025) — derives documentation benchmarks from the HumanEval function programming benchmark; evaluates whether generated documentation is accurate enough for downstream code generation to reproduce the original function.

    Future Directions (2026-2030)

  • Executable documentation — documentation that embeds live code examples runnable in browser sandboxes (WASM-powered Python or JavaScript runtimes in the browser), auto-updated by the documentation generation pipeline on every release and guaranteed to produce the documented output when run. This closes the common documentation failure mode where example code is correct at authoring time but silently breaks as the API evolves. Starlette’s interactive docs (powered by Swagger UI) point in this direction, but full WASM-powered execution with CI-validated outputs remains a research target.

  • Multi-modal documentation — combining generated prose with auto-generated architecture diagrams (Mermaid, D2, PlantUML), sequence diagrams derived from execution traces, data flow diagrams generated from call-graph analysis, and screencast-style video walkthroughs generated by multimodal LLMs from code execution demos. Tools like Mintlify already support Mermaid diagram embedding; the research frontier is generating the diagram markup automatically from code structure analysis.

  • Causal documentation — going beyond “what this function does” to “why this design decision was made”, drawing on git history (commit messages, PR descriptions, review comments), RFC threads, ADRs, and issue tracker context to surface design rationale alongside API signatures. This form of documentation captures institutional knowledge that is currently lost when original authors leave a project.

  • Domain-adapted documentation models — fine-tuning specialised documentation LLMs per programming language ecosystem to respect ecosystem-specific conventions: Rust documentation must discuss ownership, borrowing, and lifetime semantics; Haskell documentation involves type-level programming concepts; CUDA kernel documentation requires discussion of thread hierarchy, shared memory, and synchronisation primitives; Erlang/OTP documentation must address supervision trees and message passing. General-purpose LLMs hallucinate ecosystem-specific conventions at rates that specialised models can substantially reduce.

  • Bidirectional doc-code consistency enforcement — systems that not only generate documentation from code but bidirectionally flag when code drifts from its documented contract, treating the docstring as an executable specification. Integration with property-based testing frameworks (Hypothesis, QuickCheck) could automatically validate that generated usage examples actually produce the documented outputs.

  • Standards convergence — anticipated convergence around MCP-native documentation endpoints, llms.txt, and OpenAPI 4.0 as an integrated developer-documentation ecosystem for agentic AI where documentation sites are first-class citizens of the AI coding assistant’s tool registry rather than passive web pages.

  • Low-resource programming language support — extending documentation generation to under-documented programming languages (Coq, Agda, Idris, Lean, K Framework DSLs, domain-specific languages for hardware description, query languages) using few-shot adaptation of multilingual code models. This is particularly relevant for formal verification tools where the proof assistant’s documentation is the primary interface for new users.

  • Self-updating documentation at deployment time — documentation generation systems that monitor production telemetry (API call patterns, error rates, common parameter combinations) and automatically annotate documentation with empirically observed usage patterns, correcting discrepancies between the documented “typical use case” and the actual observed usage in production.

    Quality, Evaluation, and the Hallucination Problem

    A fundamental challenge in documentation generation is hallucination: the tendency of large language models to generate plausible-sounding but factually incorrect documentation. In the code documentation context, hallucination takes specific forms: generating parameter descriptions that refer to parameters that do not exist in the function signature, claiming a function returns a value when it returns None, asserting that a function is thread-safe when it is not, or generating code examples that are syntactically invalid or produce incorrect results. Hallucination in documentation is particularly dangerous because incorrect documentation propagates errors: a developer who reads incorrect documentation writes incorrect code, and the error manifests at runtime rather than at development time.

    The primary mitigations for documentation hallucination in practice are: (i) signature grounding — constraining generation by explicitly providing all parameter names and type annotations from the AST in the prompt, with post-generation validation that all provided parameter names appear in the generated documentation; (ii) retrieval augmentation — providing similar existing, correct documentation as few-shot in-context examples that anchor the generation to realistic patterns; (iii) rubric-guided evaluation — using a separate LLM evaluator that checks the generated documentation against a factual rubric (does parameter X appear in the signature? does the return description match the declared return type?); and (iv) doctest validation — for languages that support executable documentation (Python doctests, Rust doctests), automatically compiling and running generated code examples as part of the pipeline and failing on incorrect output.

    The Code2Doc work (2024) established that data quality is the primary lever for reducing hallucination rates in fine-tuned documentation models: the 13,358 high-quality pairs in Code2Doc were selected by filtering for uniqueness (cosine similarity below 0.85 between documentation strings), length (400-1000 characters), absence of AI-generated boilerplate patterns (detected via regex heuristics for known ChatGPT-style phrases), and human-readable quality (no truncated sentences, no formatting artefacts). Models fine-tuned on this filtered dataset hallucinate at substantially lower rates than models fine-tuned on raw mined corpora, even at equivalent corpus size.

    Multi-agent evaluation pipelines (DocAgent, 2025) address hallucination through iterative refinement: an initial documentation draft is evaluated by a dedicated evaluation agent that cross-checks each factual claim against the code, and a refinement agent corrects identified errors. The multi-agent loop terminates when the evaluation agent assigns a passing score on a rubric that explicitly checks for: all parameters mentioned, return value described, no non-existent parameters referenced, code examples syntactically valid. This rubric-guided approach reduces hallucination rates by 60-70% compared to single-pass generation on the same model, at the cost of approximately 3x the inference budget.

    Integration with Agentic Software Development Workflows

    The emergence of agentic AI development tools — systems in which AI agents autonomously plan, write, test, refactor, and document code across multi-step workflows — has elevated documentation generation from a developer convenience to an architectural primitive. In agentic pipelines, documentation serves two distinct audiences simultaneously: human developers who read it to understand and extend the codebase, and AI agents that consume it as grounding context for subsequent autonomous actions.

    The GitHub Copilot Workspace (2024-2025) represents the most widely deployed agentic development environment. When Copilot Workspace receives a task (typically specified as a GitHub issue), it autonomously reads relevant files, proposes a plan including which files to create or modify, and then executes the plan. The quality of its code generation and refactoring decisions depends critically on the documentation quality of the APIs it must call: functions with accurate docstrings specifying return types, parameter constraints, and side effects are correctly used; functions with missing or inaccurate documentation are misused, producing subtle bugs. This creates a direct economic incentive for organisations using agentic tools to maintain high-quality documentation, closing the gap between documentation as a “nice to have” and documentation as a correctness requirement.

    Claude Code (Anthropic, 2025) represents a different architectural approach: a terminal-resident AI assistant that maintains a persistent understanding of a codebase across a session. When tasked with implementing a feature or debugging an issue, Claude Code queries the project’s documentation (if MCP-exposed), reads relevant source files, and generates code that is consistent with the existing API contracts and conventions. The MCP-enabled documentation query capability means that Claude Code can ask the project’s documentation server “what does function X return when called with an empty list?” and receive a structured, reliable answer rather than hallucinating from pre-training knowledge — a capability that directly depends on the documentation server having been populated by a documentation generation pipeline.

    Cursor (2024-2025), a VSCode fork with deep LLM integration, introduced the concept of “Cursor Rules” — project-level instruction files that guide the AI’s code generation behaviour, and ”@ mentions” that pull specific documentation pages into the AI’s context window during a coding task. When developers @ mention an SDK documentation page during a Cursor chat, the AI reads the documentation and generates API-compatible code. Documentation generation systems that produce clean, accurately structured documentation directly improve Cursor’s generation quality in a measurable way.

    The emerging concept of documentation-driven development (DDD) inverts the traditional code-first, documentation-second workflow: documentation is authored (or generated from specifications) before implementation, and code generation agents are then given the documentation as their implementation specification. This is analogous to test-driven development (TDD) but at the documentation level, and it leverages the fact that LLMs can generate correct code from accurate documentation more reliably than they can generate documentation from ambiguous code. Tools exploring this workflow include Devin (Cognition AI), SWE-agent, and various research systems for autonomous software engineering.

    The relationship between documentation generation and code generation also reveals an interesting bidirectionality: code generation models benefit from exposure to documentation during pre-training (learning what good documented APIs look like), and documentation generation models benefit from exposure to code (learning how to describe code in natural language). Models like CodeT5, UniXcoder, and the recent family of unified code-language models explicitly train on both directions simultaneously, with shared encoder representations that benefit both tasks.

    Research & Literature

    1. McBurney, P. W., & McMillan, C. (2014). Automatic documentation generation via source code summarization of method context. ICPC 2014.
    2. Iyer, S., Konstas, I., Cheung, A., & Zettlemoyer, L. (2016). Summarizing source code using a neural attention model. ACL 2016.
    3. Hu, X., Li, G., Xia, X., Lo, D., & Jin, Z. (2018). Deep code comment generation. ICSE 2018.
    4. Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., & Brockschmidt, M. (2019). CodeSearchNet challenge: Evaluating the state of semantic code search. arXiv:1909.09436.
    5. Feng, Z., Guo, D., Tang, D., Duan, N., Feng, X., Gong, M., et al. (2020). CodeBERT: A pre-trained model for programming and natural languages. EMNLP 2020.
    6. Guo, D., Ren, S., Lu, S., Feng, Z., Tang, D., Liu, S., et al. (2021). GraphCodeBERT: Pre-training code representations with data flow. ICLR 2021.
    7. Wang, Y., Wang, W., Joty, S., & Hoi, S. C. H. (2021). CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. EMNLP 2021.
    8. Guo, D., Lu, S., Duan, N., Wang, Y., Zhou, M., & Yin, J. (2022). UniXcoder: Unified cross-modal pre-training for code representation. ACL 2022.
    9. Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., et al. (2021). Evaluating large language models trained on code. arXiv:2107.03374 [Codex / GitHub Copilot].
    10. Li, R., Allal, L. B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., et al. (2023). StarCoder: May the source be with you! arXiv:2305.06161.
    11. Lozhkov, A., Li, R., Allal, L. B., Cassano, F., Lamy-Poirier, J., Tazi, N., et al. (2024). StarCoder 2 and The Stack v2. arXiv:2402.19173.
    12. Liang, J., et al. (2025). DocAgent: A multi-agent system for automated code documentation generation. arXiv:2504.08725.
    13. Code2Doc: A quality-first curated dataset for code documentation. (2024). arXiv:2512.18748.
    14. Gu, X., Zhang, H., & Kim, S. (2016). Deep API learning. FSE 2016.
    15. Zhang, J., Wang, X., Zhang, H., Sun, H., Wang, K., & Liu, X. (2019). A novel neural source code representation based on abstract syntax tree. ICSE 2019.
    16. Wan, Y., Zhao, Z., Yang, M., Xu, G., Ying, H., Wu, J., & Yu, P. S. (2018). Improving automatic source code summarization via deep reinforcement learning. ASE 2018.
    17. Ahmad, W. U., Chakraborty, S., Ray, B., & Chang, K.-W. (2020). A transformer-based approach for source code summarization. ACL 2020.
    18. Mintlify (2025). Documentation is your AI interface. Mintlify Blog. https://www.mintlify.com/blog/docs-as-ai-interface
    19. When LLMs meet API documentation: Can retrieval augmentation aid code generation just as it helps developers? (2025). arXiv:2503.15231.
    20. Evaluating repository-level software documentation via question answering. LLM4Code@ICSE 2025. arXiv:2604.06793.
    21. Automated and context-aware code documentation leveraging advanced LLMs. (2025). arXiv:2509.14273.
    22. McMillan, C., Grechanik, M., Poshyvanyk, D., Xie, Q., & Fu, C. (2011). Portfolio: Finding relevant functions and their usage. ICSE 2011.
    23. Moreno, L., Aponte, J., Sridhara, G., Marcus, A., Pollock, L., & Vijay-Shanker, K. (2013). Automatic generation of natural language summaries for Java classes. ICPC 2013.
    24. Sridhara, G., Hill, E., Muppaneni, D., Pollock, L., & Vijay-Shanker, K. (2010). Towards automatically generating summary comments for Java methods. ASE 2010.
    25. Documentation Matters: Human-centered AI system to assist data science code documentation. (2021). arXiv:2102.12592.
    26. Holistic Evaluation of State-of-the-Art LLMs for Code Generation. (2024). arXiv:2512.18131.
    27. Allamanis, M., Barr, E. T., Devanbu, P., & Sutton, C. (2018). A survey of machine learning for big code and naturalness. ACM Computing Surveys, 51(4).

Provenance