A large language model (LLM) is a deep neural network — almost universally based on the Transformer Architecture — trained via self-supervised next-token prediction on web-scale corpora of text (and often code, mathematics, and structured data), resulting in a system that assigns a probability distribution over token sequences and can generate coherent, contextually appropriate continuations. Scale — both in parameter count (billions to hundreds of billions) and training tokens (trillions) — is the defining characteristic that distinguishes LLMs from earlier, smaller language models, and is the proximate cause of qualitative capability jumps such as in-context learning, instruction following, chain-of-thought reasoning, and emergent generalisation across domains. LLMs are typically released as base pretrained models that are subsequently aligned to human preferences through supervised fine-tuning and reinforcement learning from human feedback, producing the instruction-following assistants widely deployed in consumer and enterprise applications. The paradigm has become the de facto foundation for natural language processing, code synthesis, autonomous agent planning, and multimodal AI systems.
Overview
- LLMs represent the convergence of three trends that crossed critical thresholds simultaneously: the Transformer Architecture providing a parallelisable, scalable compute primitive; massive increases in available training data from the web; and GPU/TPU hardware enabling previously infeasible training runs.
- The term “large” is relative and has shifted over time — early GPT (2018, 117M parameters) was large by the standards of its era; frontier models as of 2025 operate at the scale of hundreds of billions of parameters and trillions of training tokens.
- Scaling laws (empirically established by Kaplan et al. 2020 and refined by Hoffmann et al. 2022 — the Chinchilla result) show that model capability scales predictably as a power law with compute, parameter count, and data volume, enabling principled allocation of training budgets.
- Emergent abilities — capabilities that appear abruptly as scale increases and are absent at smaller scale — have been observed for arithmetic, multi-step reasoning, and code generation, though the mechanism remains debated.
- LLMs are the canonical instance of Foundation Model: a single large pretrained model that serves as the starting point for many downstream tasks via Fine Tuning or prompting.
Key Components
- Transformer Architecture — The dominant architectural substrate: stacked blocks of multi-head self-attention and feed-forward sub-layers, enabling parallel computation over full sequences and long-range dependency capture.
- Attention Mechanism — Scaled dot-product attention computes, for each token, a weighted sum over all other tokens in the Context Window, with weights derived from learned query-key interactions across multiple parallel heads.
- Embeddings — Continuous dense vector representations mapping discrete token IDs into a high-dimensional latent space where geometric relationships encode semantic and syntactic structure.
- Tokenisation — Byte-pair encoding (BPE) or SentencePiece algorithms decompose raw Unicode text into a fixed vocabulary of subword units, balancing coverage with sequence length; vocabulary sizes typically range from 32,000 to 256,000 tokens.
- Context Window — The maximum number of tokens the model can attend over in a single forward pass. Has grown from 2,048 (GPT-2) to 128,000+ (GPT-4o, Gemini 1.5) to 1M+ tokens in recent systems via techniques such as RoPE, ALiBi, and sliding-window attention.
- Positional encoding — Since self-attention is permutation-invariant, token order is injected explicitly via sinusoidal embeddings or learnable positional encodings; rotary positional embeddings (RoPE) have become dominant for long-context models.
- Vocabulary projection head — A linear layer projecting the final hidden state of each token onto the vocabulary dimension, followed by softmax normalisation to produce the next-token probability distribution minimised against the training cross-entropy.
Training Pipeline
Pretraining
- Pretraining on a large, diverse Training Data corpus — typically a blend of web crawl data (Common Crawl), books, code repositories, scientific literature, and curated high-quality sources — using next-token prediction as the self-supervised objective.
- Decoder-only autoregressive models (GPT family, Llama, Mistral, Qwen) predict each token given all preceding tokens; encoder-only masked models (BERT family) predict masked tokens given bidirectional context.
- Self-Supervised Learning removes the need for human annotation at pretraining scale; the model acquires world knowledge, linguistic structure, and reasoning patterns from distributional statistics alone.
- Compute Infrastructure demands are extreme: frontier pretraining runs consume tens of thousands of GPU Computing accelerators over months, with energy consumption measured in gigawatt-hours.
- Scaling Laws provide empirical guidance: compute-optimal runs allocate roughly equal scaling to model size and training tokens (Chinchilla: ~20 training tokens per parameter for a compute-optimal base model).
Alignment and Fine-Tuning
- Supervised Fine-Tuning (SFT) — The pretrained base model is fine-tuned on curated instruction-response demonstrations produced by human annotators or distilled from stronger models, teaching the model to follow instructions and maintain a helpful conversational register.
- Reinforcement Learning from Human Feedback (RLHF) — A reward model trained on human preference comparisons between model outputs is used to score generations; proximal policy optimisation (PPO) or direct preference optimisation (DPO) updates the language model policy to maximise expected reward, aligning outputs with human values.
- Constitutional AI — An Anthropic technique that encodes a set of principles into the training loop, allowing the model to critique and revise its own outputs without relying solely on human annotators, improving scalability of alignment.
- Parameter-efficient fine-tuning — LoRA (Low-Rank Adaptation), QLoRA, prefix tuning, and adapter layers reduce fine-tuning cost by updating a small fraction of parameters, enabling adaptation on consumer hardware.
Inference
- Tokens are generated autoregressively: the model produces a probability distribution over the vocabulary at each step; a token is sampled via temperature scaling, top-p (nucleus) sampling, or beam search; the sampled token is appended to the context and the process repeats.
- KV-cache stores previously computed key-value pairs across attention heads, amortising quadratic attention cost over the decoding sequence.
- Quantisation (INT8, INT4, GPTQ, AWQ, GGUF) compresses weight precision to reduce memory footprint and increase throughput, enabling deployment on smaller accelerators or consumer GPUs.
Architectural Variants
- Decoder-only autoregressive — The dominant family for open-ended generation: GPT-4, Claude, Llama 3, Mistral, Gemini, Qwen, Falcon, Command R. Attends only to past tokens via a causal mask.
- Encoder-only — BERT, RoBERTa, DeBERTa; bidirectional context enables strong classification and embedding tasks but not open-ended generation. Still widely used for semantic search and Retrieval-Augmented Generation retrieval encoders.
- Encoder-decoder (seq2seq) — T5, BART, mT5; suitable for translation, summarisation, and structured prediction; the encoder reads the full input bidirectionally and the decoder generates outputs autoregressively.
- Mixture of Experts — Mixtral, DeepSeek-MoE, Grok: a router selects a sparse subset of expert sub-networks per token, scaling total parameter count while keeping active parameters (and compute) fixed per token. Enables trillion-parameter models without proportional inference cost.
- Multimodal LLMs — GPT-4o, Gemini, Claude, LLaVA, Qwen-VL extend the text backbone with vision encoders (CLIP, SigLIP) and audio/video tokenisers, enabling image, audio, and video understanding alongside text generation.
Applications and Use Cases
- Conversational AI — Chat assistants (ChatGPT, Claude, Gemini, Copilot) provide interactive question-answering, drafting, summarisation, and tutoring across consumer and enterprise contexts.
- Code Generation — GitHub Copilot, Cursor, Codestral, StarCoder, and DeepSeek-Coder generate, complete, explain, and refactor code in dozens of programming languages, substantially accelerating software development workflows.
- Retrieval-Augmented Generation — LLMs grounded with a retrieval system over a Knowledge Graph or document corpus reduce hallucination and extend effective knowledge to post-training facts, enabling accurate knowledge-intensive QA.
- AI Agent orchestration — LLMs serve as the reasoning and planning backbone for autonomous agents that call tools, execute code, browse the web, and coordinate with other agents in Multi-Agent System architectures.
- Natural Language Processing — Machine translation, named entity recognition, sentiment analysis, document classification, and information extraction pipelines increasingly rely on LLM feature extraction or fine-tuned heads.
- Scientific discovery — Protein sequence modelling (ESM), drug discovery, literature synthesis, hypothesis generation, and experimental design assistance in biology, chemistry, and materials science.
- Legal and medical applications — Contract review, clinical note summarisation, medical coding, regulatory document analysis; specialist fine-tuned variants (Med-PaLM, BioGPT) optimise for domain accuracy.
- Education and tutoring — Personalised explanations, worked examples, Socratic dialogue, and automated feedback on student writing at scale.
Risks and Limitations
- Hallucination — LLMs generate plausible-sounding but factually incorrect content because they optimise for coherent token sequences rather than grounded truth; mitigation strategies include Retrieval-Augmented Generation, grounding constraints, and uncertainty calibration.
- Bias and toxicity — Training corpora reflect societal biases; without careful Instruction Tuning and safety filtering, models may produce harmful, discriminatory, or offensive outputs.
- AI Safety concerns — Misalignment between model objectives and human values, susceptibility to adversarial prompting (jailbreaking), and potential for misuse in disinformation, cyberattacks, or automated harmful content generation.
- Environmental cost — Frontier pretraining and large-scale inference impose substantial energy and water footprints; the field is investigating efficiency improvements via distillation, Quantisation, and sparse models.
- Opacity and interpretability — The internal representations and reasoning processes of LLMs remain poorly understood, complicating debugging, auditing, and regulatory compliance.
- Context window limits — Despite rapid improvement, very long documents and multi-session memory remain challenges; retrieval and summarisation hierarchies partially compensate.
Standards & Context
- The EU AI Act Regulatory Instrument (2024) classifies general-purpose AI models (GPAIs) above a training compute threshold of 10²⁵ FLOPs as “systemic-risk” systems subject to mandatory transparency, incident reporting, and adversarial testing obligations; frontier LLMs fall squarely within scope.
- NIST AI Risk Management Framework (AI RMF 1.0, 2023) provides voluntary governance guidance applicable to LLM deployment, covering identification, measurement, and mitigation of AI-related risks.
- Model cards (Mitchell et al.) and datasheets for datasets (Gebru et al.) are community norms for documenting intended use, limitations, evaluation results, and training data provenance, adopted by major LLM providers.
- The Chinchilla scaling laws (Hoffmann et al., 2022) and Kaplan et al. (2020) power-law scaling results are the empirical foundation for compute-budget allocation decisions in frontier training runs.
- Responsible Scaling Policies (Anthropic ASL, OpenAI Preparedness Framework) establish voluntary capability thresholds that trigger additional safety evaluations before continued scaling.
Current Landscape (2026)
- Reasoning models became a distinct category through 2025 — OpenAI’s o1/o3, Anthropic’s extended-thinking Claude, and DeepSeek-R1 (which claimed frontier reasoning at roughly $6M training cost) — normalising inference-time “thinking” budgets rather than single-pass generation.
- The frontier flagship race intensified across late 2025 into 2026: Google’s Gemini 3 Pro (18 November 2025), OpenAI’s GPT-5.2 (11 December 2025) and Anthropic’s Claude Opus 4.5 (October 2025), followed by Opus 4.6 with adaptive thinking on 5 February 2026, with 1M-token context windows moving from premium to default.
- Coding and agentic benchmarks converged near saturation, with SWE-bench Verified scores clustering around 76–81% (Claude Opus 4.5 leading at ~80.9%) and AIME 2025 maths hitting 100% for several models, shifting evaluation toward harder frontier suites like ARC-AGI-2, Humanity’s Last Exam and Terminal-Bench.
- Open-weight models closed much of the gap: DeepSeek scaled its Mixture-of-Experts line from V3 (December 2024) through V3.2 with Sparse Attention to a V4 Preview (April 2026), while Alibaba’s Qwen3/Qwen3.5 (Apache 2.0), Meta’s Llama 4 Scout (10M-token context), GLM-5, Kimi K2 and Mistral Large 3 gave regulated enterprises viable self-hosted options — parity on coding, but a persistent ~5–18 point closed-model lead on general knowledge and long-context retrieval.
- The Model Context Protocol (MCP) emerged as the de facto standard for agent tool connectivity — a provider-agnostic JSON-RPC interface adopted by AWS, Azure, GCP, VS Code and JetBrains through 2025 — alongside the Agent-to-Agent (A2A) protocol for multi-agent coordination, accelerating enterprise agentic deployments.
- Regulation tightened materially: the EU AI Act’s obligations for general-purpose AI (GPAI) model providers entered into application on 2 August 2025 (transparency, technical documentation, copyright policy, training-data summaries), with models above 10^25 FLOP presumed to carry “systemic risk”; the Commission’s enforcement powers, including fines up to €35M or 7% of global turnover, took effect on 2 August 2026.
- Open challenges as of 2026 include residual hallucination (best models around 6% on targeted evals), inference cost and energy, model provenance and data-sovereignty scrutiny (especially for Chinese-origin open weights), and the shift from single-model selection to per-workload routing across mixed open/closed fleets.
References
-
- European Commission (2025). Guidelines for providers of general-purpose AI models under the AI Act. https://digital-strategy.ec.europa.eu/en/policies/guidelines-gpai-providers
-
- European Commission (2026). AI Act — Regulatory framework for AI (Regulation (EU) 2024/1689). https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai
-
- Introl (2026). GPT-5.2 vs Gemini 3 Pro: 2026 Benchmark Comparison. https://introl.com/blog/gpt-5-2-vs-gemini-3-benchmark-comparison-2026
-
- Hidekazu Konishi (2026). Open-Weights LLM Release History and Timeline — Llama, Mistral, Qwen, DeepSeek. https://hidekazu-konishi.com/entry/open_weights_llm_release_history_and_timeline.html
-
- SimuPro (2026). LLM & AI Developments 2025–2026. https://simupro.nl/guides/llm-ai-developments-2025-2026/