Document summarisation is the natural language processing task of producing a concise, faithful representation of the salient information in one or more source documents. It encompasses extractive approaches, which select and concatenate important spans, and abstractive approaches, which generate new text that paraphrases the content. Modern systems are built predominantly on transformer-based large language models and are evaluated for informativeness, coherence, and factual consistency.
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:hasPart ai:ExtractiveSummariser))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:hasPart ai:AbstractiveSummariser))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:hasPart ai:FaithfulnessEvaluator))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:hasPart ai:SalienceDetector))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:hasPart ai:ChunkingStrategy))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:hasPart ai:QueryFocusedSummariser))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:hasPart ai:MultiDocumentFusion))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:hasPart ai:RelevanceRanker))
Dependency Relationships
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:requires ai:LanguageModel))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:requires ai:Tokenisation))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:requires ai:AttentionMechanism))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:requires ai:ModelEvaluation))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:dependsOn ai:LargeLanguageModels))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:dependsOn ai:DeepLearning))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:dependsOn ai:TrainingData))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:dependsOn ai:Transformer))
Capability Relationships
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:enables ai:KnowledgeManagement))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:enables ai:InformationRetrieval))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:enables ai:QuestionAnswering))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:enables ai:RetrievalAugmentedGeneration))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:enables ai:LiteratureReview))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:enables ai:ExecutiveBriefGeneration))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:enables ai:MultiDocumentSynthesis))
Implementation Relationships
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:implements ai:Transformer))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:implements ai:SequenceToSequenceLearning))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:implements ai:ReinforcementLearningFromHumanFeedback))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:implements ai:AttentionMechanism))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:implements ai:FineTuning))
Reduction Relationships
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:reducesTo ai:TextGeneration))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:reducesTo ai:ContentCompression))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:reducesTo ai:SalienceDetection))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:reducesTo ai:RelevanceRanking))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:reducesTo ai:InformationSelection))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:reducesTo ai:NLPGenerationTask))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:reducesTo ai:TextAbstractionProblem))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:reducesTo ai:DocumentCompressionTask))
SubClassOf(ai:DocumentSummarisation
ObjectSomeValuesFrom(ai:reducesTo ai:SalienceSelectionProblem))
About
Document Summarisation has been studied as a computational task since the earliest days of natural language processing research. Luhn’s 1958 IBM Journal paper “The Automatic Creation of Literature Abstracts” introduced the foundational extractive approach: sentences are scored by the frequency and clustering of significant terms, and the highest-scoring sentences are concatenated as an abstract. This word-frequency salience heuristic — conceptually related to TF-IDF in Information Retrieval — underpins extractive summarisers to this day. The Edmundson (1969) extractive summariser added structural features (title words, cue phrases, sentence position) to frequency scoring, establishing the multi-feature extractive paradigm. For four decades, document summarisation research was dominated by extractive methods because they guaranteed that output sentences were grammatically correct, factually grounded in the source, and interpretable — properties that early natural language generation systems could not reliably provide.
The abstractive paradigm became computationally viable with the advent of neural sequence-to-sequence models. Rush, Chopra and Weston’s “A Neural Attention Model for Abstractive Sentence Summarisation” (EMNLP 2015) demonstrated that an attentional neural encoder-decoder could produce grammatical, compressed headlines from news sentences that outperformed extractive baselines on ROUGE metrics. Nallapati et al. (2016) extended this to full document summarisation using CNNs and RNNs with the CNN/DailyMail dataset — still a primary benchmark — demonstrating abstractive generation across multiple sentences. See et al.’s “Get to the Point” (ACL 2017) introduced the pointer-generator network that combined copy (extract) and generate (abstract) operations, addressing the vocabulary coverage limitation of pure generation models and substantially reducing hallucination by allowing the model to copy source phrases directly.
The transformer revolution transformed document summarisation as it did all of NLP. BERT (Devlin et al., 2019) provided rich pre-trained encoders that improved extractive summarisation through better sentence encoding. The sequence-to-sequence transformer models — BART (Lewis et al., 2020) and T5 (Raffel et al., 2020) — pre-trained with denoising objectives on large corpora and fine-tuned on summarisation benchmarks, established the new state of the art for abstractive summarisation. PEGASUS (Zhang et al., 2020), pre-trained with a gap-sentences generation (GSG) objective specifically designed to approximate the summarisation task, achieved state-of-the-art ROUGE scores on XSum and CNN/DailyMail with very little fine-tuning data, demonstrating that pre-training objective alignment to the target task is a powerful inductive bias. By 2021-2023, decoder-only Large Language Models including GPT-3, GPT-4, Claude, and Gemini demonstrated powerful zero-shot and few-shot summarisation via instruction following, matching or surpassing fine-tuned BART and PEGASUS on many benchmarks without any task-specific training — at the cost of significantly higher inference cost and greater hallucination risk for long, technical documents.
The central technical challenge of the contemporary era is faithfulness and factual consistency. Research by Maynez et al. (ACL 2020) established that approximately 25% of summaries generated by state-of-the-art abstractive models contain hallucinated content — information not present in or contradicted by the source document — even when ROUGE scores are competitive with extractive systems. ROUGE’s n-gram overlap metric, which measures lexical similarity between generated and reference summaries, is systematically blind to hallucination: a fluent, coherent summary that invents plausible but false facts can achieve high ROUGE scores. This finding triggered a research programme in faithfulness evaluation (FactCC, SummaC, FactScore, BERTScore) and faithfulness-improving training methods (faithful fine-tuning with contrastive learning, RLHF alignment, self-consistency decoding) that continues through 2026.
Components / Architecture
The modern document summarisation pipeline comprises the following architectural components:
- Document Ingestion and Chunking: Long documents exceeding the context window (typically 4K-128K tokens for contemporary LLMs) are chunked into manageable segments. Chunking strategies include fixed-size windows with overlap, recursive paragraph splitting, and semantic chunking using Embedding similarity to detect topical boundaries. The chunking strategy is a critical hyperparameter: chunks too small lose cross-sentence context; chunks too large exceed the context window or dilute salience signals.
- Extractive Component: Sentence or passage scoring using TF-IDF salience, position features (lead bias: news articles concentrate key information in early sentences), graph-based methods (TextRank treats sentences as nodes and cosine similarity as edges, applying PageRank to identify central sentences), or Embedding-based clustering (k-means over sentence embeddings to identify representative centroid sentences). Extractive summaries are factually grounded by construction but may be incoherent across selected sentences and miss synthesised insights absent from individual sentences.
- Abstractive Generation (Seq2Seq / Decoder): Encoder-decoder models (BART, T5, PEGASUS) encode the full source document and decode a summary auto-regressively via Attention Mechanism over the encoder states. Decoder-only Large Language Models (GPT-4, Claude, Gemini, Llama-3) generate summaries via prompted text completion. Key architectural variants:
- BART (Bidirectional and Auto-Regressive Transformer): Pre-trained with text infilling (masking spans with a single mask token), token deletion, sentence permutation, and document rotation noise. Fine-tuned on CNN/DailyMail achieves ROUGE-1 44.16. Best-in-class for news summarisation.
- PEGASUS: Pre-trained with Gap Sentences Generation (GSG) — whole sentences are masked and the model generates them from surrounding context, closely approximating the abstractive summarisation objective. Achieves ROUGE-1 47.21 on XSum with only 1,000 fine-tuning examples via transfer learning.
- T5: Treats all NLP tasks (including summarisation) as text-to-text transformations with task-specific prefixes (“summarize: …”). Flexible multi-task training enables strong few-shot transfer. T5-11B achieves ROUGE-1 43.52 on CNN/DailyMail.
- LED (Longformer Encoder-Decoder): Extends BART with sparse local+global attention patterns (Longformer) in the encoder, enabling direct summarisation of documents up to 16,384 tokens without chunking. Suitable for academic papers, legal documents, and medical records.
- Multi-Document Fusion: When summarising multiple documents (news cluster, literature review, multiple retrieved RAG passages), redundancy must be detected and eliminated while cross-document complementarity is captured. Approaches include: fusion-in-decoder (concatenate all documents and pass to a single decoder), iterative refinement (summarise each document then summarise summaries), and graph-based multi-document fusion (cross-document sentence similarity graph with centrality-based selection).
- Map-Reduce for Very Long Documents: When total document length exceeds model context, a map-reduce strategy applies: chunk the document into segments, summarise each chunk independently (map), then summarise the concatenated intermediate summaries (reduce). Multiple reduce iterations may be applied for very long documents. This strategy trades coherence across chunk boundaries for tractable memory usage and is widely used in production LLM summarisation pipelines.
- Faithfulness Evaluation and Post-Processing: Generated summaries are assessed for factual consistency using NLI-based models (SummaC), question-generation-and-answering frameworks (QAEval, FactCC), or LLM-as-judge prompting (asking an LLM to identify which claims in the summary are not supported by the source). Summaries failing faithfulness thresholds are regenerated with more conservative prompting (lower temperature, citation instructions) or fall back to extractive summaries.
Summarisation Paradigm Families
- Extractive Summarisation: Selects verbatim sentences or spans from the source document. Factual accuracy guaranteed by construction; can be incoherent across discontinuous sentences; misses synthesised insights. Methods: TF-IDF scoring, TextRank, BertSum (BERT-based sentence classification), SumBasic, LexRank. Best for legal and medical documents where faithfulness is paramount and verbatim phrasing carries legal significance.
- Abstractive Summarisation: Generates new text compressing and paraphrasing the source. Coherent, concise, and capable of synthesis; prone to hallucination; requires faithfulness evaluation. Methods: BART, PEGASUS, T5, LED, GPT-4, Claude. Best for news, consumer-facing summaries, and meeting notes where readability outweighs verbatim precision.
- Hybrid Summarisation: Combines extractive and abstractive components. Pointer-generator networks (copy or generate per token); extract-then-abstract (extract key sentences, then abstractively rephrase); and re-ranking of abstractive candidates by an extractive faithfulness scorer. Best balance of faithfulness and fluency for most production applications.
- Query-Focused Summarisation: Generates a summary specifically addressing a user query rather than the document’s overall content. Relevant for Retrieval-Augmented Generation, Question Answering, and aspect-based analysis. Implemented by conditioning the abstractive generator on both the document and the query via concatenation or cross-attention.
- Multi-Document Summarisation: Synthesises information from multiple source documents, managing redundancy and integrating complementary information. PRIMERA (Gu et al., 2022) uses entity-centred pre-training with multi-document inputs; PEGASUS-XL with saliency-guided scoring handles multi-document abstractive summarisation. Applications: news cluster summarisation, scientific literature review, multi-source RAG synthesis.
- Aspect-Based and Structured Summarisation: Generates summaries structured around predefined aspects or templates (PICO for clinical trials: Population, Intervention, Comparison, Outcome; background/methods/results/conclusions for scientific abstracts). Particularly valuable in regulated domains where summary structure must conform to reporting standards.
- Dialogue and Meeting Summarisation: Specialised for conversational text (meeting transcripts, customer service interactions, chat history). Unique challenges include turn-taking structure, coreference, informal language, and topic segmentation. AMI and ICSI meeting corpora; SamSum conversational summarisation dataset. Models including BART fine-tuned on SamSum achieve ROUGE-1 53.28.
Use Cases / Major Families
- Knowledge Management and Executive Briefing: Organisations process volumes of reports, research papers, regulatory filings, and news that vastly exceed human reading capacity. Document summarisation enables systematic compression of incoming information into actionable executive briefs, regulatory digests, and knowledge repository entries. Law firms use document summarisation for contract review acceleration; investment banks use it for earnings call and analyst report summarisation; pharmaceutical companies use it for clinical trial literature monitoring.
- Retrieval-Augmented Generation Context Preparation: In RAG pipelines, retrieved documents are often too long to fit within the LLM context window or contain much irrelevant content. Summarising retrieved passages before injecting them as LLM context improves answer faithfulness by reducing distraction from off-topic content and concentrating the relevant information. The 2026 production RAG setup described by FutureAGI combines a retrieval layer with a summarisation model and an evaluation harness scoring groundedness, completeness, and refusal correctness.
- Legal Document Review: Contract analysis (identifying obligation clauses, termination conditions, liability provisions), e-discovery (summarising documents flagged by keyword search for attorney review), court judgment summarisation, and regulatory compliance document review. UK legal AI companies including Luminance (London) deploy extractive and hybrid summarisation to reduce attorney review time by 50-90%. A comprehensive 2025 survey (arXiv:2501.17830) analysed the specific challenges of legal case judgment summarisation, identifying structural complexity, lengthy documents, and specialised legal terminology as primary obstacles for general-purpose LLMs.
- Clinical and Biomedical Summarisation: Patient record summarisation for handover notes, clinical trial report synthesis for systematic reviews, drug label summarisation, and radiology report summarisation. The NHS, NICE, and NHS England AI Lab are active adopters. Clinical domain summarisation requires extremely high faithfulness — medication errors or missed diagnoses from hallucinated summaries carry patient safety implications. Specialised clinical BERT models (BioMedBERT, ClinicalBERT) provide better extractive sentence scoring for medical text than general-purpose encoders.
- News and Media Monitoring: Automatic generation of news article abstracts, multi-article cluster summaries for news aggregators, and personalised topic digests. CNN/DailyMail, XSum, and Newsroom are standard benchmarks. Production systems for media monitoring (Bloomberg Terminal, Reuters Connect, Financial Times archive) use PEGASUS and BART variants fine-tuned on domain-specific news corpora.
- Meeting and Conversation Summarisation: Automatic generation of meeting minutes from transcripts (Microsoft Teams, Zoom, Google Meet AI summarisation features), customer service interaction summaries for CRM logging, and conversational AI session digests. This is one of the highest-volume commercial applications in 2025-2026, driven by the proliferation of hybrid work tools with built-in AI summarisation. Microsoft 365 Copilot, Google Workspace Duet AI, and Zoom AI Companion all provide meeting summarisation as a standard enterprise feature.
- Scientific Literature Review: Accelerating systematic literature reviews by summarising abstracts and full papers by research question, intervention, and outcome. Tools including Elicit, Consensus, and Semantic Scholar’s TLDR feature use LLM-based document summarisation to help researchers navigate the exponentially growing scientific literature. The Question Answering use case for scientific literature (e.g. “what does the literature say about the efficacy of drug X for condition Y?”) requires multi-document summarisation across hundreds of retrieved papers.
Academic Context
Document summarisation research is disseminated primarily through ACL, EMNLP, NAACL, EACL, and COLING conferences, as well as journals including Computational Linguistics, IEEE Transactions on Knowledge and Data Engineering, and ACM TOIS. The ACL Anthology (aclanthology.org) provides open access to virtually all NLP research including summarisation literature.
Foundational extractive summarisation research spans Luhn (1958) through Kupiec, Pedersen and Chen (1995) automated abstract generation to Mihalcea and Tarau’s TextRank (EMNLP 2004). The shift to neural methods began with Rush et al. (EMNLP 2015) on neural attention for headline generation. Nallapati et al. (CoNLL 2016) introduced SummaRuNNer, an extractive/abstractive RNN model with the CNN/DailyMail dataset. See et al.’s pointer-generator network (ACL 2017) remains a frequently cited baseline. The pre-trained language model era was initiated by Liu and Lapata’s BertSum (EMNLP 2019) for extractive BERT-based summarisation, followed by the BART (Lewis et al., ACL 2020), PEGASUS (Zhang et al., ICML 2020), and T5 (Raffel et al., JMLR 2020) sequence-to-sequence giants.
Faithfulness research crystallised with Maynez et al. “On Faithfulness and Factuality in Abstractive Summarisation” (ACL 2020) and Kryscinski et al.’s FactCC factual consistency evaluation (EMNLP 2020). Long-document summarisation challenges were addressed by Beltagy, Peters and Cohan’s Longformer (arXiv 2020) and its encoder-decoder variant LED. Huang et al.’s “The Factual Inconsistency Problem in Abstractive Text Summarization” and SummaC (Laban et al., TACL 2022) consolidated faithfulness evaluation methodology.
UK academic contributions include EdinburghNLP’s research on long document summarisation, supported by UKRI CDT in NLP (EP/S022481/1) and ELIAI at the University of Edinburgh. Work on “Enhancing Long Document Long Form Summarisation with Self-Planning” (arXiv:2512.17179, 2025) emerged from Edinburgh researchers. The University of Warwick and Imperial College London conduct biomedical NLP research relevant to clinical summarisation. Sheffield NLP group contributes multilingual summarisation research. King’s College London’s NLP research group investigates dialogue summarisation for clinical applications.
Current Landscape (2026)
By 2026 document summarisation has transitioned from an academic NLP task to a commodity feature embedded in productivity software, enterprise knowledge management platforms, and RAG infrastructure. Microsoft 365 Copilot, Google Workspace Duet AI, Zoom AI Companion, and Notion AI all include document summarisation as a standard feature, collectively reaching hundreds of millions of enterprise users. Production systems in 2026 use decoder-only Large Language Models (Claude Opus, GPT-4, Gemini Pro) for high-quality single-document and moderate-length multi-document summarisation via direct prompting, with map-reduce chunking strategies for documents exceeding model context windows.
The FutureAGI 2026 production guide documents the standard architecture: a retrieval layer fetches relevant document chunks; a summarisation model (gpt-4o, claude-opus, or gemini-pro in long-context mode) generates summaries; and an evaluation harness scores groundedness (faithfulness), completeness, and refusal correctness, with automatic regeneration or fallback to extractive summarisation when faithfulness scores fall below threshold. Models like claude-opus-4 and gemini handle 200K+ token contexts natively, reducing the need for chunking strategies for many single-document use cases.
Hallucination remains the primary quality challenge. Research from 2024-2026 shows that approximately 25% of state-of-the-art model summaries still contain unsupported claims when evaluated on technical documents, despite RLHF alignment. The RAGAS evaluation framework (Shahul et al., 2023) has become the production standard for RAG faithfulness evaluation, decomposing answer claims into atomic statements and checking each against retrieved context — a methodology directly applicable to summarisation faithfulness assessment. The Lynx open-source hallucination detection model outperforms RAGAS on hallucination detection for long-context cases. Production deployments increasingly use LLM-as-judge faithfulness scoring (asking a separate LLM instance to assess whether each claim in the summary is supported by the source) as an automated quality gate.
Novel architectures in 2026 include PEGASUS-XL with saliency-guided scoring for multi-document abstractive summarisation (Nature Scientific Reports, 2025), demonstrating improved factual accuracy by incorporating explicit saliency constraints into the generation objective. A Frontiers in AI 2026 paper introduced a bird-flocking-inspired LLM framework combining multiple LLM agents with a bird-flocking consensus algorithm to generate summaries with greater factual accuracy by preventing generation of unsupported information through inter-agent consistency enforcement. A comprehensive 2025 survey (arXiv:2409.02413) reviewed the state of abstractive text summarisation, cataloguing advances in faithfulness, long-document handling, multilingual summarisation, and domain-specific adaptation.
UK Context
United Kingdom institutions are active contributors to document summarisation research across both academic NLP and applied AI deployment. The University of Edinburgh’s NLP group (EdinburghNLP, edinburghnlp.inf.ed.ac.uk) — one of Europe’s largest academic NLP groups — conducts research in language generation, abstractive summarisation, and long-document processing, supported by UKRI CDT in Natural Language Processing (grant EP/S022481/1, University of Edinburgh and Heriot-Watt University consortium) and ELIAI (Edinburgh Laboratory for Integrated Artificial Intelligence, EPSRC-funded). Recent Edinburgh work on long document long-form summarisation with self-planning (arXiv:2512.17179, 2025) addresses the coherence challenges of generating extended summaries from lengthy source documents.
Imperial College London’s Department of Computing includes NLP and ML researchers working on clinical NLP applications including patient record summarisation and biomedical literature synthesis. King’s College London NLP group investigates dialogue summarisation for clinical handover applications within the NHS context. The University of Manchester’s text mining group (National Centre for Text Mining, NaCTeM) has long-running projects on biomedical text summarisation for clinical decision support and drug discovery. Sheffield NLP group contributes to multilingual summarisation research with direct relevance to UK’s diverse language communities. UCL Computer Science includes text generation and summarisation in its NLP curriculum.
In applied industry, Luminance (London-based legal AI) and Synthesia (London, AI video generation with scripted summaries) represent UK commercial applications of document summarisation technology. NHS England’s AI Lab funds clinical NLP projects including radiology report summarisation and discharge summary generation at NHS trusts across England. The NHS Clinical Decision Support programme evaluates LLM-based summarisation of NICE guideline clusters for point-of-care clinical use. Northern English academic medical centres — Manchester University NHS Foundation Trust, Leeds Teaching Hospitals NHS Trust, and Newcastle Hospitals — are participating in NHS AI summarisation pilot programmes for clinical handover and discharge note generation, with safety evaluation protocols mandating faithfulness scoring before deployment.
Future Directions (2026–2030)
Document summarisation research and deployment will advance along several converging trajectories through 2030. Factual faithfulness improvement is the primary research priority: combining retrieval augmentation (ensuring summaries are grounded in retrieved evidence rather than parametric memory), uncertainty quantification (summaries with uncertainty estimates highlighting low-confidence claims), and constitutional AI methods (rule-based constraints preventing hallucination of numeric facts, dates, and named entities) will push faithfulness rates toward 98%+ for structured document types. Structured and citation-preserving summarisation — where each summary claim is linked to the specific source sentence it derives from — will become the standard for legal, medical, and scientific applications, combining extractive traceability with abstractive fluency.
Long-context native summarisation will displace map-reduce chunking strategies as context windows expand: models with 1M+ token contexts will summarise entire books, legal case files, or clinical records in a single pass without information loss at chunk boundaries. Multi-modal document summarisation — handling PDFs with embedded charts, tables, equations, and images natively — will become standard as multimodal Large Language Models mature. Personalised and aspect-targeted summarisation, where the user specifies their role, expertise level, or specific interest dimensions and the system adapts summary depth and terminology accordingly, will enable truly user-adaptive knowledge distillation. Continuous summarisation pipelines that maintain living summaries updated incrementally as new documents arrive — rather than generating summaries on demand — will enable always-current knowledge management for fast-moving domains such as news, regulatory updates, and scientific literature. Cross-lingual document summarisation, producing English summaries of documents in Welsh, Urdu, Arabic, or any of the 100+ languages covered by BGE-M3, will extend the technology to UK’s diverse linguistic communities and global enterprise deployments without language barriers.
Formal Summarisation Algorithms
Modern document summarisation follows distinct computational procedures depending on the paradigm:
Extractive Summarisation (TextRank):
- Sentence segmentation: split source document D into sentences S = {s₁, s₂, …, sₙ}.
- Sentence embedding: encode each sentence sᵢ as a vector vᵢ using a pre-trained sentence encoder (Embedding Model) or TF-IDF weighted term vector.
- Similarity graph construction: build graph G = (V, E) where nodes V = S and edge weight w(sᵢ, sⱼ) = cosine_similarity(vᵢ, vⱼ) for all i ≠ j.
- PageRank-style centrality scoring: iteratively update scores r(sᵢ) = (1-d) + d × Σⱼ [w(sᵢ,sⱼ) / Σₖ w(sⱼ,sₖ)] × r(sⱼ) until convergence (typically 50-100 iterations, damping factor d=0.85).
- Select top-k highest-scoring sentences as the extractive summary, maintaining their original order for coherence.
- Optional: apply MMR (Maximum Marginal Relevance) to reduce redundancy among selected sentences by penalising candidates that are highly similar to already-selected sentences.
Abstractive Summarisation (BART/PEGASUS seq2seq):
- Tokenise source document D into token sequence X = (x₁, …, xₘ) using BPE or SentencePiece tokeniser. Apply truncation or chunking if m exceeds model context window (typically 1024 for BART, 512 for PEGASUS, up to 128K for long-context LLMs).
- Encoder forward pass: compute contextualised representations H = Encoder(X), where H = (h₁, …, hₘ) are bidirectional self-attention representations of all input tokens.
- Decoder auto-regressive generation: at each step t, generate token yₜ = argmax P(y | y₁,…,y_{t-1}, H) where P is computed by cross-attention over encoder states H followed by a softmax projection over vocabulary V.
- Apply beam search (beam size 4-8) or sampling (temperature, top-p nucleus sampling) to generate high-quality, diverse summaries. Beam search maximises likelihood; sampling increases diversity at the cost of some coherence.
- Post-process: apply length constraints (min/max_length), no-repeat-ngram-size to prevent repetition, and force_words constraints for domain-specific required terms.
Map-Reduce for Long Documents:
- Chunk document D into segments {C₁, …, Cₖ} of max_chunk_tokens tokens with overlap of overlap_tokens.
- Map: generate intermediate summary Sᵢ = Summarise(Cᵢ) independently for each chunk.
- Reduce: concatenate intermediate summaries S_concat = [S₁; S₂; …; Sₖ] and apply a second summarisation: S_final = Summarise(S_concat).
- If S_concat still exceeds context window, apply additional reduce iterations recursively.
- Optional refinement: apply a coherence-focused refinement pass over S_final to smooth junctures between chunk summaries.
Faithfulness Evaluation (RAGAS-style):
- Decompose generated summary Y into atomic claims C = {c₁, …, cₙ} using a claim extraction LLM prompt.
- For each claim cᵢ: check entailment against source document D using an NLI model (SummaC, TRUE) or LLM-as-judge prompt asking “Is claim cᵢ supported by the source document?“.
- Faithfulness score = |{cᵢ : supported(cᵢ, D)}| / |C|.
- Completeness score = recall of key source facts in Y, estimated by generating questions from source and checking if Y answers them (QAEval approach).
Evaluation Metrics and Benchmarks
Document summarisation evaluation has evolved from purely lexical metrics toward comprehensive faithfulness and quality assessment:
Lexical Overlap Metrics:
-
ROUGE-N (Recall-Oriented Understudy for Gisting Evaluation): Lin (2004). Measures n-gram overlap between generated and reference summaries. ROUGE-1 (unigrams), ROUGE-2 (bigrams), ROUGE-L (longest common subsequence). ROUGE-1/2/L standard reporting on CNN/DailyMail and XSum. BART achieves ROUGE-1 44.16 on CNN/DailyMail; PEGASUS achieves 47.21 on XSum. Critical limitation: cannot detect hallucination — a fluent hallucinated summary can achieve high ROUGE if its vocabulary matches the reference.
-
BLEU (Bilingual Evaluation Understudy): Precision-focused n-gram overlap. More commonly used for machine translation but reported for summarisation in multilingual settings.
Semantic Similarity Metrics:
-
BERTScore: Zhang et al. (2019). Computes precision, recall, and F1 between token-level BERT embeddings of generated and reference summaries, capturing semantic similarity beyond exact n-gram overlap. More robust than ROUGE to paraphrase and synonym-based generation. BERTScore F1 correlates better with human judgements than ROUGE on abstractive summaries.
-
MoverScore: Zhao et al. (2019). Earth mover’s distance between token embedding distributions of generated and reference summaries. Accounts for semantic similarity at token level with soft matching.
Faithfulness Evaluation:
-
FactCC: Kryscinski et al. (EMNLP 2020). NLI-based factual consistency checker trained on automatically constructed consistent/inconsistent summary pairs. Binary classification: consistent or inconsistent with source.
-
SummaC: Laban et al. (TACL 2022). Sentence-level NLI-based inconsistency detection with aggregation strategies (SummaC-ZS and SummaC-Conv). Current state-of-the-art NLI-based faithfulness evaluator.
-
QAEval: Question-generation-and-answering approach. Generate questions from the source document, answer them from both source and summary, and measure agreement. Low agreement indicates hallucination or omission.
-
FactScore: Min et al. (ACL 2023). Decomposes long-form generated text into atomic facts and checks each against retrieved Wikipedia evidence. Most fine-grained factuality evaluation available; standard for open-domain summarisation faithfulness.
-
RAGAS Groundedness: Shahul et al. (2023). In RAG context, measures the fraction of answer claims supported by retrieved context passages. Directly applicable to RAG-based document summarisation evaluation.
Benchmark Datasets:
-
CNN/DailyMail: ~300K news article-highlight pairs. Most widely used English summarisation benchmark. Highlights are multi-sentence bullet-point abstracts averaging 56 words. BART ROUGE-1 44.16, ROUGE-2 21.28, ROUGE-L 40.90.
-
XSum (Extreme Summarisation): Narayan et al. (2018). 226K BBC articles with one-sentence journalist-written summaries. Requires abstractive compression of an entire article into a single sentence — a harder abstractive task than CNN/DailyMail. PEGASUS ROUGE-1 47.21, ROUGE-2 24.56, ROUGE-L 39.25.
-
SamSum: 16K messenger-style dialogues with human-written summaries. Standard benchmark for conversation summarisation. BART fine-tuned achieves ROUGE-1 53.28.
-
ArXiv/PubMed: Cohan et al. (2018). Long scientific documents (avg. 6K tokens) with abstract-style summaries. Standard benchmark for long-document summarisation.
-
MultiNews: Fabbri et al. (2019). 56K news cluster-summary pairs for multi-document summarisation.
-
MeetingBank: Meeting transcript summarisation benchmark from city council proceedings, covering structured long-form dialogue with topical organisation.
Key Terminology Glossary
- Extractive Summarisation: A summarisation approach that selects and concatenates verbatim sentences or spans from the source document. Guarantees factual accuracy by construction; limited in expressive compression.
- Abstractive Summarisation: A summarisation approach that generates new text compressing and paraphrasing source content. Produces fluent, concise output but may hallucinate.
- Hallucination: The generation of plausible-sounding text claims not supported by or contradicted by the source document. The primary quality failure mode of abstractive neural summarisation.
- Faithfulness / Factual Consistency: The property that all claims in a generated summary are entailed by and supported by the source document. Measured by FactCC, SummaC, FactScore, RAGAS groundedness.
- Salience: The relative importance of a sentence or passage in representing the key information of a document. Extractive summarisers score salience; abstractive models implicitly learn to generate salient content.
- ROUGE (Recall-Oriented Understudy for Gisting Evaluation): The dominant lexical overlap evaluation metric for summarisation, measuring n-gram recall/precision between generated and reference summaries.
- Lead Bias: The empirical observation that news articles and many structured documents concentrate key information in the first few sentences. Extractive summarisers exploiting positional features leverage this; it also explains why extractive lead-3 baselines are competitive on CNN/DailyMail.
- Pointer-Generator Network: A hybrid seq2seq architecture (See et al., 2017) that at each generation step can either generate a vocabulary word or copy a source token via a soft copy mechanism, combining abstractive fluency with extractive faithfulness.
- Coverage Mechanism: An extension of the pointer-generator that tracks which source tokens have been attended to, penalising repeated attention to the same positions to reduce repetition in generated summaries.
- Map-Reduce Summarisation: A strategy for summarising documents exceeding model context windows by splitting into chunks (map), summarising each independently, and summarising the intermediates (reduce).
- Query-Focused Summarisation: Generating a summary tailored to a specific question or information need rather than the document’s overall content. Essential for RAG and aspect-based analysis.
- Multi-Document Summarisation: Producing a single coherent summary from multiple source documents, managing cross-document redundancy and complementarity.
Limitations and Failure Modes
Document summarisation systems exhibit systematic failure modes that production deployments must explicitly address through evaluation and mitigation strategies.
Hallucination and Factual Inconsistency: The most critical failure mode of abstractive neural summarisation. Approximately 25% of summaries from state-of-the-art models contain unsupported claims (Maynez et al., ACL 2020), including entity name errors (wrong person names, wrong numbers, wrong dates), relation errors (correct entities but wrong relationships), and entirely fabricated facts. RLHF alignment and instruction tuning reduce but do not eliminate hallucination. Production mitigation: faithfulness evaluation scoring gates (regenerate if faithfulness < 0.9), citation-preserving generation (instruct model to cite source spans), and fallback to extractive summarisation for high-stakes domains.
Context Window Limitations: Standard BART and PEGASUS models support only 1024 input tokens, forcing aggressive truncation for many real-world documents. Truncation silently drops content from the middle and end of long documents (following the lead bias heuristic of prioritising beginning-of-document tokens), potentially omitting critical information in documents where key content appears late. LED/LongT5 models extend context to 16K-32K; native long-context LLMs handle 128K-1M tokens but at higher inference cost.
ROUGE Metric Misalignment: ROUGE rewards lexical similarity to reference summaries but does not penalise hallucination, repetition, or incoherence. A model that produces fluent but factually incorrect summaries can outperform a faithful extractive system on ROUGE, misaligning optimisation targets with real quality. ROUGE is necessary but insufficient as a sole evaluation metric; production systems must complement ROUGE with faithfulness (SummaC, FactScore) and coherence (human evaluation or LLM-as-judge) metrics.
Reference Summary Quality Dependence: Evaluation accuracy depends on the quality and coverage of human reference summaries, which may themselves be incomplete, inconsistent, or reflect individual annotator preferences. CNN/DailyMail “highlights” are bullet-point lists rather than coherent summaries; XSum journalist summaries are sometimes subjective or include context not in the article. This limits the reliability of automatic evaluation on these benchmarks as proxies for real-world summarisation quality.
Domain Generalisation: Models fine-tuned on CNN/DailyMail (news) perform poorly on legal contracts, medical records, or scientific papers without domain adaptation. Legal text includes extensive cross-references, defined terms, and conditional logic that standard summarisers fail to handle; medical text requires clinical ontology awareness; scientific text involves mathematical formalism and domain-specific argument structure. Domain-specific fine-tuning on in-domain annotated data, or few-shot prompting of large general-purpose Large Language Models, is required for satisfactory performance in specialised domains.
Research & Literature
- Luhn, H. P. (1958). The Automatic Creation of Literature Abstracts. IBM Journal of Research and Development, 2(2), 159–165.
- Edmundson, H. P. (1969). New Methods in Automatic Extracting. Journal of the ACM, 16(2), 264–285.
- Mihalcea, R., & Tarau, P. (2004). TextRank: Bringing Order into Texts. EMNLP 2004.
- Rush, A. M., Chopra, S., & Weston, J. (2015). A Neural Attention Model for Abstractive Sentence Summarization. EMNLP 2015. arXiv:1509.00685.
- Nallapati, R., Zhou, B., Gulcehre, C., & Xiang, B. (2016). Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond. CoNLL 2016. arXiv:1602.06023.
- See, A., Liu, P. J., & Manning, C. D. (2017). Get to the Point: Summarization with Pointer-Generator Networks. ACL 2017. arXiv:1704.04368.
- Liu, Y., & Lapata, M. (2019). Text Summarization with Pretrained Encoders (BertSum). EMNLP 2019. arXiv:1908.08345.
- Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., … & Zettlemoyer, L. (2020). BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. ACL 2020. arXiv:1910.13461.
- Zhang, J., Zhao, Y., Saleh, M., & Liu, P. J. (2020). PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization. ICML 2020. arXiv:1912.08777.
- Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., … & Liu, P. J. (2020). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5). JMLR, 21(140), 1–67. arXiv:1910.10683.
- Maynez, J., Narayan, S., Bohnet, B., & McDonald, R. (2020). On Faithfulness and Factuality in Abstractive Summarization. ACL 2020. arXiv:2005.00661.
- Beltagy, I., Peters, M. E., & Cohan, A. (2020). Longformer: The Long-Document Transformer. arXiv:2004.05150.
- Kryscinski, W., McCann, B., Xiong, C., & Socher, R. (2020). Evaluating the Factual Consistency of Abstractive Text Summarization (FactCC). EMNLP 2020. arXiv:1910.12840.
- Laban, P., Schnabel, T., Bennett, P. N., & Hearst, M. A. (2022). SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization. TACL 2022. arXiv:2111.09525.
- Gu, Y., et al. (2022). PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document Summarization. ACL 2022. arXiv:2110.08499.
- Shahul, E., James, J., Mäder, L., Lanchier, N., Fröbel, N., & Bhatt, J. (2023). Ragas: Automated Evaluation of Retrieval Augmented Generation. arXiv:2309.15217.
- Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W., Koh, P. W., … & Hajishirzi, H. (2023). FActScoring: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. ACL 2023. arXiv:2305.14251.
- Abstractive Text Summarization: State of the Art, Challenges, and Improvements. (2024). arXiv:2409.02413.
- A Comprehensive Survey on Legal Summarization: Challenges and Future Directions. (2025). arXiv:2501.17830.
- Enhancing Long Document Long Form Summarisation with Self-Planning. (2025). arXiv:2512.17179. [Edinburgh NLP].
- PEGASUS-XL with saliency-guided scoring and long-input encoding for multi-document abstractive summarization. (2025). Scientific Reports. https://www.nature.com/articles/s41598-025-11062-2
- A bird-inspired AI framework for advanced large text summarization. (2026). Frontiers in Artificial Intelligence. https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2026.1703769/full
- Document Summarization with LLMs in 2026: A Production Guide. (2026). FutureAGI. https://futureagi.com/blog/revolutionizing-document-management-llm-2025/
- Evaluating RAG Systems in 2025: RAGAS Deep Dive, Giskard Showdown. (2025). Cohorte Engineering Blog. https://cohorte.co/blog/evaluating-rag-systems-in-2025-ragas-deep-dive-giskard-showdown-and-the-future-of-context
- Summarizing Business News: Evaluating BART, T5, and PEGASUS for Effective Summarization. (2025). IIETA. https://www.iieta.org/download/file/fid/132319
- RAG Evaluation 2026 - Faithfulness, Relevancy, and Context Metrics. (2026). Benchmarking Agents. https://benchmarkingagents.com/rag-eval/
- UKRI CDT in Natural Language Processing. University of Edinburgh Research Explorer. https://www.research.ed.ac.uk/en/organisations/ukri-cdt-in-natural-language-processing/
- EdinburghNLP Research Group. University of Edinburgh, School of Informatics. https://edinburghnlp.inf.ed.ac.uk/