Fine-tuning is the post-pre-training adaptation of a large neural network — typically a transformer-based Foundation Models or Large Language Models — to narrower task distributions or behavioural objectives by continuing gradient-based optimisation on comparatively small, curated dataset…
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:hasPart ai:SupervisedFineTuning))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:hasPart ai:ReinforcementLearningFromHumanFeedback))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:hasPart ai:DirectPreferenceOptimisation))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:hasPart ai:RewardModel))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:hasPart ai:ParameterEfficientFineTuning))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:hasPart ai:InstructionTuning))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:hasPart ai:ConstitutionalAI))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:hasPart ai:SafetyFineTuning))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:hasPart ai:KnowledgeDistillation))
## Dependency Relationships
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:requires ai:PreTrainedModel))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:requires ai:TrainingData))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:requires ai:LossFunction))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:requires ai:GradientDescent))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:requires ai:ComputeInfrastructure))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:dependsOn ai:TransformerArchitecture))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:dependsOn ai:ScalingLaws))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:dependsOn ai:ReinforcementLearning))
## Capability Relationships
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:enables ai:InstructionFollowing))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:enables ai:AlignmentWithHumanValues))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:enables ai:DomainSpecialisedLLM))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:enables ai:SafetyAlignment))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:enables ai:CodeGeneration))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:supports ai:MedicalAI))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:supports ai:LegalAI))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:supports ai:MultiTaskLearning))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:supports ai:AgentCapability))
## Implementation Relationships
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:implements ai:CausalLanguageModelling))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:implements ai:ProximalPolicyOptimisation))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:implements ai:DirectPreferenceOptimisation))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:implements ai:LowRankAdaptation))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:implements ai:Backpropagation))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:implements ai:ElasticWeightConsolidation))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:implements ai:PrefixTuning))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:implements ai:PromptTuning))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:implements ai:Kahneman-TverskyOptimisation))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:implements ai:OddsRatioPreferenceOptimisation))
## Reduction Relationships
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:reduces ai:TrainableParameterCount))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:reduces ai:GPUMemoryRequirement))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:reduces ai:AlignmentTaxOnCapability))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:reduces ai:CatastrophicForgetting))
SubClassOf(ai:TrainingAndFineTuning
ObjectSomeValuesFrom(ai:reduces ai:AnnotationCost))
About Fine-Tuning
Fine-tuning is the second major phase of the modern LLM development pipeline, picking up where Model Training (self-supervised pre-training) leaves off. Whereas pre-training endows a model with broad world knowledge through exposure to hundreds of billions or trillions of tokens, fine-tuning redirects that latent knowledge toward specific behavioural objectives — following instructions, preferring safer outputs, mastering a narrow domain — using orders-of-magnitude fewer labelled examples and compute cycles.
The distinction from pre-training is both practical and conceptual. Pre-training learns what the world looks like via next-token prediction; fine-tuning learns how the model should behave, typically via supervised examples of target behaviour or feedback signals expressing human preferences. This division of labour reflects an empirical finding: a sufficiently large pre-trained model already encodes the knowledge required for most downstream tasks; fine-tuning serves primarily to unlock, direct, and align that knowledge rather than to instil new facts.
Fine-tuning’s economic importance is substantial. For an organisation deploying a 70B-parameter LLaMA or Mistral AI Open-Weight Model Family model, a full fine-tuning run touching all weights requires ~5.6 TB of GPU memory (bf16 parameters + gradients + Adam states) and thousands of GPU-hours. LoRA DoRA etc (LoRA, r=16) reduces this to ~140 GB — achievable on four A100-40GB GPUs — while recovering 95–98% of full fine-tuning quality on most benchmarks. QLoRA further compresses the frozen base to 4-bit NF4, enabling a 65B-parameter fine-tune on a single 48 GB GPU.
The four fine-tuning tracks — SFT, RLHF/DPO/KTO/ORPO alignment, PEFT, and domain/continual adaptation — are typically composed in sequence: a raw base model undergoes SFT to become instruction-following, then one or more alignment rounds to become preference-optimised, with PEFT applied throughout for compute efficiency, and domain adaptation layered on top for specialist deployments.
Components and Architecture
Supervised Fine-Tuning (Instruction Tuning)
SFT is the foundational step converting a raw language model pre-trained on web text into a model that follows natural-language instructions. The training signal is the causal language modelling loss applied exclusively over response tokens (not instruction tokens), encouraging the model to generate correct, helpful completions given a structured prompt:
L_SFT(θ) = −(1/|y|) Σᵢ log P_θ(yᵢ | x, y₁,…,yᵢ₋₁)
where x is the instruction/context and y = (y₁,…,y|y|) is the desired response. Masking loss over instruction tokens prevents the model from memorising the prompt phrasing.
Dataset construction is the dominant quality bottleneck. Early public datasets (Alpaca, 52K GPT-3.5-generated instruction-response pairs, Stanford 2023) were inexpensive but noisy. Higher-quality approaches curate from human-written data (OpenAssistant, 161K human turns), model distillation from stronger teachers (WizardLM, Orca, using GPT-4 to generate complex reasoning traces), or hybrid pipelines combining human seed tasks with model-augmented responses (FLAN collection, Wei et al. 2022, 1,836 tasks across 62 NLP benchmarks). By 2025, leading SFT datasets include Open AI’s internal instruction-response curation (scale undisclosed), Meta’s Llama-3.1-Instruct training blend, and community datasets on Hugging Face such as OpenHermes-2.5 (1M examples), Tulu-3 (RLHF-augmented), and Magpie (synthetic self-distillation at 1M+ scale).
FLAN-T5 (Chung et al. 2022) demonstrated that multi-task instruction tuning at scale generalises strongly to unseen tasks (zero-shot MMLU 55.1% vs 44.5% without FLAN). InstructGPT (Ouyang et al. 2022) showed that SFT followed by RLHF produced outputs preferred by human annotators over 175B base GPT-3 by 85% margin despite using only 1.3B fine-tuned parameters — a seminal demonstration that alignment, not scale, drives helpfulness.
RLHF: PPO Pipeline
The canonical three-stage RLHF pipeline (Christiano et al. 2017, refined in InstructGPT):
- Supervised Fine-Tuning (SFT): As above, producing π_SFT.
- Reward Model (RM) Training: Human annotators rank model outputs for the same prompt (chosen ≻ rejected). A separate transformer encodes (prompt, response) and outputs a scalar reward R_φ(x, y). Trained with Bradley-Terry ranking loss: L_RM = −E[log σ(R_φ(x, y_c) − R_φ(x, y_r))] over (chosen, rejected) pairs.
- RL Fine-Tuning via PPO: The SFT model is further fine-tuned using PPO to maximise the reward model signal while staying close to π_SFT via a KL-penalty term:
J(θ) = E_{xD, yπ_θ}[R_φ(x,y)] − β · KL(π_θ(y|x) ‖ π_SFT(y|x))
The KL-divergence coefficient β (typically 0.01–0.1) prevents reward hacking — exploitation of weaknesses in the reward model without improving true quality. PPO is an on-policy algorithm requiring rollout generation at each update step, making it computationally expensive (3–5× the cost of SFT for equivalent token throughput). OpenAI’s PPO infrastructure for InstructGPT used 32 A100 GPUs per rollout worker.
Reward hacking is a pervasive failure mode: the policy learns to produce outputs that score highly on the reward model without genuinely improving quality, often via verbosity, sycophantic hedging, or distribution shift away from the RM’s training domain. Mitigation strategies include reward model ensembles, online reward model updates (RLHF online), and constitutional constraints.
DPO, KTO, and ORPO: Reward-Free Preference Optimisation
Direct Preference Optimisation (DPO, Rafailov et al. 2023) observed that the optimal policy under the RLHF objective has a closed-form solution relating the reward to the log-ratio of the optimal and reference policies:
r*(x, y) = β · log [π*(y|x) / π_ref(y|x)] + β · log Z(x)
Substituting into the Bradley-Terry loss eliminates the need for an explicit reward model, yielding a single supervised loss over (prompt, chosen, rejected) triplets:
L_DPO(θ) = −E[log σ(β · log(π_θ(y_c|x)/π_ref(y_c|x)) − β · log(π_θ(y_r|x)/π_ref(y_r|x)))]
DPO is 2–3× cheaper than PPO (no rollout generation, no reward model inference) and numerically more stable. It became the dominant alignment algorithm for open-source models through 2023–2024: Llama-2-Chat, Zephyr-β, Mistral-7B-Instruct-v0.2, and Qwen-2-Instruct all use DPO or closely related variants.
KTO (Kahneman-Tversky Optimisation, Ethayarajh et al. 2023) extends DPO to unpaired feedback — single (prompt, response, binary good/bad signal) examples without explicit chosen-rejected pairs — using a loss inspired by prospect theory’s asymmetric treatment of losses and gains. This reduces annotation cost by ~50% since annotators mark individual outputs rather than ranking pairs.
ORPO (Odds Ratio Preference Optimisation, Hong et al. 2024) further merges SFT and preference alignment into a single training pass by augmenting the SFT loss with an odds-ratio penalty discouraging rejected responses:
L_ORPO = L_SFT − λ · E[log σ(log(odds_θ(y_c|x) / odds_θ(y_r|x)))]
where odds_θ(y|x) = P_θ(y|x) / (1 − P_θ(y|x)). ORPO eliminates the reference model requirement, reducing memory by ~50% and training time by ~30% versus DPO while matching DPO quality on AlpacaEval 2.0.
IPO (Identity Preference Optimisation, Azar et al. 2024) addresses DPO’s tendency to overfit by regularising via a squared residual loss rather than a log-sigmoid margin.
Constitutional AI
Constitutional AI Training Methodology (CAI, Bai et al. 2022, Anthropic) replaces human preference annotation for harmlessness training with a set of explicit ethical principles (“the constitution”) from which AI-generated preference data is synthesised:
- Supervised Constitutional Revision (SCC): Sample red-team prompts, generate an initial response, then prompt the model to critique and revise the response against constitution principles. Use revised responses as SFT examples.
- RLAIF: Train a preference model on (original, revised) response pairs using the model’s own constitutional scoring rather than human preference judgements. Apply PPO or DPO using this AI-generated reward signal.
CAI scales harmlessness training without proportionally scaling human annotation costs and is the stated basis for Anthropic’s Claude model alignment stack. By 2025, RLAIF (Reinforcement Learning from AI Feedback) using synthetic preference generation is widely adopted: Gemini’s alignment stack (Google DeepMind), Llama-3’s self-generated critique pipeline, and Qwen-2.5’s DPO data synthesis all employ variants.
Knowledge Distillation Fine-Tuning
Knowledge Distillation applied in the fine-tuning context produces smaller student models that match the quality of larger teacher models on specific tasks. Two main paradigms:
Black-box distillation (self-instruct / model distillation): Use outputs of a strong teacher (GPT-4, Claude-3.5-Sonnet) as SFT targets for a smaller student. This is the basis of Alpaca (GPT-3.5→LLaMA-7B), WizardLM (GPT-4→LLaMA-13B), and Orca (GPT-4 with explanation traces→LLaMA-13B). Quality gains of 1–3 MMLU points per doubling of distillation data budget are typical.
White-box / logit distillation: Minimise KL divergence between teacher and student output distributions token-by-token, preserving richer calibration information than hard labels. Used in MiniLLM (Gu et al. 2023), DistiLLM, and Hugging Face’s distilgpt2/distilbert pipelines. Requires teacher logits (not just final outputs), so applicable primarily within an organisation’s own model family.
Distillation-based fine-tuning is central to edge and mobile deployment where a 7B or 3B student must approximate a 70B teacher for domain-specific tasks at a fraction of the inference cost.
Use Cases / Major Families
Instruction Tuning at Scale
The FLAN (Fine-tuned Language Net, Wei et al. 2022) and T0 (Sanh et al. 2021) projects demonstrated that instruction-tuning across hundreds of task templates dramatically improves zero-shot generalisation. FLAN-PaLM (540B) achieved state-of-the-art on a majority of 62 benchmarks. Subsequent work scaled instruction diversity: Super-NaturalInstructions (Wang et al. 2022, 1,616 tasks), Tulu (Ivison et al. 2023), and Magpie (Xu et al. 2024, 1M LLaMA-3-generated self-instruct examples).
OpenAI’s GPT-4o fine-tuning API (launched July 2024) allows organisations to fine-tune on up to 10 million tokens per training job (up from 1 million for GPT-3.5), with automatic hyperparameter selection and integrated evaluation. Pricing as of 2025: 3.75/1M input tokens during inference. This democratises instruction-tuned commercial models for domain-specific deployments without public data disclosure.
PEFT for Resource-Constrained Environments
LoRA DoRA etc covers LoRA, QLoRA, and DoRA in depth. From a fine-tuning workflow perspective, the dominant PEFT recipe by 2025 is: 4-bit QLoRA base quantisation + LoRA adapters targeting all attention projection matrices (q_proj, k_proj, v_proj, o_proj) and often MLP layers (gate_proj, up_proj, down_proj), rank r = 16–64, lora_alpha = 2r, lora_dropout = 0.05. The Hugging Face PEFT + Accelerate + TRL (Transformer Reinforcement Learning) stack is the de facto open-source toolchain, with trl.SFTTrainer and trl.DPOTrainer abstracting multi-GPU orchestration. Parameter-Efficient Fine-Tuning is a closely related ontology node covering the broader PEFT taxonomy.
Domain-Specific Fine-Tuning
Medical: Medical AI deployments fine-tune on clinical notes (MIMIC-IV, PubMed, clinical trials). Med-PaLM 2 (Google, 2023) achieved 86.5% on US Medical Licensing Exam (USMLE) — above passing threshold — via SFT + ensemble refinement. MedLLaMA, BioMedLM (Stanford), Clinical-Camel (UAB), and Meditron (EPFL/Yale, 2023, LLaMA-2 fine-tuned on 48.1B medical tokens) demonstrate the medical vertical. UK NHS AI Lab evaluations (2024) benchmark domain-fine-tuned models on UK-specific clinical coding (ICD-10, SNOMED CT) and triage tasks; Medical AI and Medical Imaging AI pages contain deployment specifics.
Legal: Legal-Bert, LegalBench (Harvard Law / Stanford, 2023, 162 legal tasks), Lawformer (Chinese civil law), and Lexis+ AI (LexisNexis, GPT-4 fine-tuned on case law) address contract analysis, precedent retrieval, and statutory interpretation. UK-specific: Legal Framework considerations around fine-tuned models used in legal research (SRA Ethics guidance 2024).
Code: Code Generation is among the strongest fine-tuning applications. StarCoder 2 (BigCode/Hugging Face, 2024, 15B parameters, 4T+ tokens code pre-training + instruction tuning), CodeLlama (Meta), and DeepSeek-Coder V2 (236B MoE, strong fine-tuning) dominate. GitHub Copilot uses a proprietary fine-tuned Codex descendant. OpenAI’s o3-mini demonstrates that reasoning-focused RLHF (process reward models scoring intermediate reasoning steps) can dramatically improve code generation — HumanEval+ 92.7%.
Science / Multimodal: Galactica (Taylor et al. 2022, Meta, 120B science fine-tune), SciGLM, and ChemLLM address scientific reasoning. Multimodal fine-tuning via LLaVA (Liu et al. 2023), Idefics 2, and InternVL extends instruction tuning to vision-language tasks.
Continual Pre-Training and Domain-Adaptive Pre-Training (DAPT)
Distinct from pure fine-tuning, Continual Pre-Training (CPT) continues the self-supervised language modelling objective on a domain-specific corpus before SFT, allowing the model to absorb new facts rather than merely learn to follow instructions within a domain. Gururangan et al. (2020, ACL) showed 3–7 MMLU-equivalent point gains from DAPT on biomedical, CS, news, and reviews corpora before downstream fine-tuning. Meditron used CPT on 48.1B medical tokens before instruction tuning; Code Llama used CPT on 500B code tokens as the foundation for instruction tuning.
Key challenge: catastrophic forgetting — CPT on narrow domain data degrades general capabilities. Mitigations: mixing in a replay fraction (5–15%) of general-domain tokens (Scialom et al. 2022), progressive layer unfreezing (fine-tune outer layers first, gradually unfreeze inner layers), and EWC++ (Elastic Weight Consolidation with Fisher information regularisation penalising weight drift proportional to parameter importance).
Full Fine-Tuning vs PEFT Trade-offs
Full fine-tuning (FFT) updates all N model parameters, allowing maximal task adaptation but at steep compute and memory cost. PEFT methods trade parameter expressivity for efficiency. The empirically observed trade-off landscape (Hu et al. 2021, Zhang et al. 2023) as of 2025:
| Method | Trainable params | GPU memory (70B model) | Quality gap vs FFT |
|---|---|---|---|
| Full FT | 70B | ~5.6 TB (bf16) | Baseline |
| LoRA r=64 | ~500M (0.7%) | ~320 GB | 1–2% MMLU |
| QLoRA r=16 | ~125M (0.2%) | ~48 GB (4-bit base) | 2–4% MMLU |
| Prefix Tuning | ~10M (0.01%) | ~280 GB | 5–10% MMLU |
| Prompt Tuning | ~1M (0.001%) | ~280 GB | 10–20% MMLU |
FFT is preferred when: (a) compute budget is unconstrained, (b) the target task is substantially out-of-distribution from the pre-training corpus, (c) the model must maintain calibrated uncertainty on tail distributions. PEFT is preferred for: (a) multiple adapters sharing a base model (hub-and-spoke deployments), (b) on-premise hardware constraints, (c) rapid iteration across task variants, (d) model serving with hot-swappable adapters (merged at inference or served via LoRA DoRA etc adapter routing).
Academic Context
The fine-tuning literature crystallised around three landmark papers: Radford et al. (2018) GPT introducing unidirectional LM fine-tuning for classification; Devlin et al. (2018) BERT establishing bidirectional masked LM pre-training + downstream SFT as the dominant NLP paradigm; and Howard & Ruder (2018) ULMFiT formalising discriminative fine-tuning (different learning rates per layer) and gradual unfreezing as engineering principles.
Preference alignment emerged as a distinct research thread via Christiano et al. (2017) RLHF for Atari and summarisation, amplified by Ziegler et al. (2019) applying RLHF to GPT-2 summarisation, and canonised by Ouyang et al. (2022) InstructGPT scaling RLHF to 175B GPT-3. Rafailov et al. (2023) DPO then simplified the pipeline, triggering a wave of alignment algorithms (KTO, IPO, ORPO, SimPO, TDPO, REBEL, RLOO).
The PEFT thread was formalised by Houlsby et al. (2019) Adapter Tuning, Li & Liang (2021) Prefix Tuning, Lester et al. (2021) Prompt Tuning, and Hu et al. (2021) LoRA — with LoRA’s zero-inference-latency property making it dominant. Dettmers et al. (2023) QLoRA democratised fine-tuning by enabling 65B-parameter runs on consumer hardware. Continual learning connections are drawn through Kirkpatrick et al. (2017) EWC and Rolnick et al. (2019) experience replay applied to LLM settings.
The intrinsic dimensionality hypothesis (Aghajanyan et al. 2020) provides theoretical grounding: fine-tuning objectives lie on low-dimensional manifolds (d* ≈ 200 for BERT, d* ≈ 1,000 for GPT-2) within the full parameter space, explaining why PEFT methods with much smaller effective dimensionality can approach full fine-tuning performance.
Current Landscape (2026)
Commercial APIs: OpenAI fine-tuning API supports GPT-4o and GPT-4o-mini (July 2024 launch), gpt-3.5-turbo-0125, with reinforcement fine-tuning (RFT) preview announced January 2025 for reasoning-task specialisation using process reward models. Anthropic offers fine-tuning of Claude models through a partner programme (as of 2025), with public fine-tuning API expected later in 2026 based on roadmap signals. Google’s Vertex AI Gemini tuning API supports supervised tuning and RLHF for Gemini 1.5 Flash/Pro. AWS Bedrock offers fine-tuning for Titan, Llama, Mistral, and Cohere models via SageMaker.
Open-source ecosystem: Axolotl (multi-backend SFT/DPO wrapper, 8K+ GitHub stars), LLaMA-Factory (unified fine-tuning toolkit, 40K+ stars), Unsloth (2–5× faster QLoRA via hand-written Triton kernels, 20K+ stars), and TRL (Hugging Face, 10K+ stars) form the dominant community toolchain. GGUF + llama.cpp enables fine-tuning inference on consumer hardware.
Alignment research 2024–2026: Self-play fine-tuning (SPIN, Chen et al. 2024) bootstraps preference data from the model’s own outputs without human annotation. RLVR (Reinforcement Learning with Verifiable Rewards, applied in DeepSeek-R1, Kimi k1.5, QwQ) uses process reward models (PRMs) scoring intermediate reasoning steps, dramatically improving chain-of-thought quality on MATH (90%+ pass@1 on MATH-500 for reasoning-optimised models). Constitutional AI 2.0 (Anthropic 2024) scales RLAIF with more nuanced multi-party constitutional principles. Safety and alignment covers the broader alignment landscape.
PEFT evolution 2024–2026: GaLore (Zhao et al. 2024) projects gradients to low-rank subspace during pre-training (not just fine-tuning), enabling full-parameter pre-training at LoRA memory cost. Flora (Hao et al. 2024) uses random projection for memory-efficient full fine-tuning. Mixture-of-LoRA (MoLoRA) routes input tokens to multiple LoRA adapters, approaching MoE expressivity at LoRA cost. LoRA DoRA etc contains detailed coverage of the expanding PEFT taxonomy.
Evaluation: Benchmark saturation (MMLU, HellaSwag) has driven adoption of more discriminating suites: MMLU-Pro, GPQA Diamond, LiveBench, BIG-Bench Hard, MT-Bench, AlpacaEval 2.0, and Arena-Hard. Evaluation benchmarks and leaderboards covers the benchmark landscape. AISI (UK AI Safety Institute) evaluation frameworks for fine-tuned model capability uplift and alignment verification inform UK government procurement of fine-tuned models from 2025.
UK Context
Edinburgh: The University of Edinburgh’s ILCC (Institute for Language, Cognition and Computation) produces leading PEFT and instruction-tuning research. Professor Ivan Titov’s group works on parameter-efficient adaptation for low-resource NLP. The Llama-2 Edinburgh fine-tune (2023) for Scottish Gaelic demonstrated cultural domain adaptation. The Alan Turing Institute Edinburgh node collaborates on alignment evaluation for UK public-sector LLM deployments.
Imperial College London: The Data Science Institute and Computing Department contribute to RLHF theory (Professor Marc Deisenroth’s group on model-based RL applied to preference learning) and PEFT for scientific computing. Imperial’s collaboration with Wellcome Sanger Institute on genomic LLM fine-tuning (Nucleotide Transformer fine-tuning for UK Biobank variant interpretation) exemplifies the medical fine-tuning vertical.
Cambridge: The Computational and Biological Learning Lab (CBL) under Professor Zoubin Ghahramani (now VP at Google DeepMind) contributed foundational Bayesian approaches to model adaptation. Cambridge’s AutoML group (Professor Neil Lawrence, now at Amazon) investigates meta-learning for efficient fine-tuning. The Cambridge Centre for AI in Medicine fine-tunes models on NHS Cambridge University Hospitals data under NHS Data Access Agreement protocols.
Oxford: The Oxford Future of Humanity Institute (successor body post-2024 restructuring) and the Oxford Internet Institute contribute to Constitutional AI and AI safety research. The Oxford Robotics Institute fine-tunes vision-language models for embodied agent control.
Manchester and Northern England: The Alan Turing Institute’s Manchester node, in partnership with Manchester University Hospitals NHS Foundation Trust, pilots fine-tuned models for radiology reporting under NHSX guidelines. The University of Sheffield NLP group (Professor Nikos Aletras) contributes to instruction-tuning evaluation methodology. Newcastle University’s Digital Institute investigates continual fine-tuning for industrial IoT sensor data interpretation under the AISI Northern England AI programme. Leeds Teaching Hospitals NHS Trust participates in federated fine-tuning pilots for clinical NLP under NHS DigiTrials, avoiding central data aggregation.
AISI (UK AI Safety Institute): AISI’s evaluation team (Frontier AI Taskforce successor) performs bespoke fine-tuning evaluations assessing capability uplift from fine-tuning on dangerous-knowledge datasets — a key concern as fine-tuning APIs lower barriers to CBRN-relevant capability acquisition. The AISI Model Evaluation and Threat Research (METR) collaboration produces public fine-tuning safety benchmarks consumed by UK government procurement frameworks. UK DSIT’s AI Opportunities Action Plan (January 2025) explicitly funds NHS and public-sector fine-tuning of foundation models on sovereign UK data.
Future Directions (2026–2030)
Process reward models and RLVR: The 2025 breakthrough of DeepSeek-R1 and Kimi k1.5 demonstrated that RLVR — fine-tuning with verifiable rewards (unit tests for code, symbolic verifiers for maths) — produces reasoning models substantially outperforming SFT+DPO baselines on formal problem domains. Expect process reward model training to become standard in the 2026–2028 alignment stack, extending to scientific research assistance, legal reasoning, and medical diagnosis.
Federated and privacy-preserving fine-tuning: Differential privacy SGD (DP-SGD) + LoRA enables fine-tuning on sensitive organisational data with provable (ε, δ)-differential privacy guarantees. NHS, financial services, and legal verticals will adopt federated LoRA fine-tuning (where adapters are trained locally and aggregated centrally without sharing raw data) as the governance-compliant path to domain adaptation. Google’s FedAvg + LoRA and Flower framework are leading implementations.
Multimodal fine-tuning: Instruction tuning for vision-language models (LLaVA-NeXT, InternVL 2.5, Qwen2-VL) is maturing. Audio-language fine-tuning (Qwen-Audio, Gemini audio) and video-language fine-tuning (VideoLLaMA 2, InternVideo2) are expanding the fine-tuning paradigm beyond text. Unified multimodal fine-tuning recipes — single LoRA adapter targeting both vision encoder projections and LLM attention layers — will become standard by 2027.
Test-time adaptation and online fine-tuning: Fine-tuning at inference time from a few in-context examples (meta-learning / gradient-based in-context learning, Akyürek et al. 2022) blurs the boundary between fine-tuning and prompting. LISA (Layerwise Importance Sampled AdamW, Pan et al. 2024) performs efficient online fine-tuning during deployment, adapting continuously to user distributions without retraining cycles.
Alignment stability and catastrophic forgetting: Long-term deployment exposes fine-tuned models to distributional drift. Replay-based continual learning, elastic fine-tuning, and periodic re-alignment against updated preference datasets will become operational requirements for production LLM services by 2027. AISI and NIST AI RMF will likely mandate re-evaluation after fine-tuning modifications to safety-critical deployments.
Research and Literature
Foundational Works:
- Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving Language Understanding by Generative Pre-Training. OpenAI Blog. [GPT fine-tuning baseline]
- Devlin, J., Chang, M.W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL 2019. arXiv:1810.04805. [BERT SFT paradigm]
- Howard, J., & Ruder, S. (2018). Universal Language Model Fine-Tuning for Text Classification. ACL 2018. arXiv:1801.06146. [ULMFiT discriminative fine-tuning]
- Christiano, P., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep Reinforcement Learning from Human Preferences. NeurIPS 2017. arXiv:1706.03741. [RLHF origin]
- Ziegler, D.M., Stiennon, N., Wu, J., Brown, T.B., Radford, A., Amodei, D., … & Irving, G. (2019). Fine-Tuning Language Models from Human Preferences. arXiv:1909.08593. [RLHF text summarisation]
Alignment Core: 6. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., … & Lowe, R. (2022). Training Language Models to Follow Instructions with Human Feedback. NeurIPS 2022. arXiv:2203.02155. [InstructGPT, RLHF at scale] 7. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C.D., & Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. arXiv:2305.18290. [DPO] 8. Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., … & Kaplan, J. (2022). Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv:2204.05862. [Anthropic RLHF] 9. Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., … & Clark, J. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073. [CAI / RLAIF] 10. Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., & Kiela, D. (2024). KTO: Model Alignment as Prospect Theoretic Optimization. ICML 2024. arXiv:2402.01306. [KTO] 11. Hong, J., Lee, N., & Thorne, J. (2024). ORPO: Monolithic Preference Optimization without Reference Model. arXiv:2403.07691. [ORPO] 12. Azar, M.G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., & Munos, R. (2024). A General Theoretical Paradigm to Understand Learning from Human Feedback. AISTATS 2024. arXiv:2310.12036. [IPO]
PEFT: 13. Hu, E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., … & Chen, W. (2021). LoRA: Low-Rank Adaptation of Large Language Models. ICLR 2022. arXiv:2106.09685. [LoRA] 14. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. NeurIPS 2023. arXiv:2305.14314. [QLoRA] 15. Liu, S.H., Zeng, Z., Liu, W., Chen, B., Shi, Z., Zhang, Y., … & Cheng, Y. (2024). DoRA: Weight-Decomposed Low-Rank Adaptation. ICML 2024. arXiv:2402.09353. [DoRA] 16. Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., … & Gelly, S. (2019). Parameter-Efficient Transfer Learning for NLP. ICML 2019. arXiv:1902.00751. [Adapter Tuning] 17. Li, X., & Liang, P. (2021). Prefix-Tuning: Optimizing Continuous Prompts for Generation. ACL 2021. arXiv:2101.00190. [Prefix Tuning] 18. Aghajanyan, A., Zettlemoyer, L., & Gupta, S. (2020). Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning. arXiv:2012.13255. [Theoretical grounding for PEFT] 19. Zhao, J., Zhang, Z., Chen, B., Wang, Z., Anandkumar, A., & Tian, Y. (2024). GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection. ICML 2024. arXiv:2403.03507. [GaLore]
Instruction Tuning & Distillation: 20. Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A.W., Lester, B., … & Le, Q. (2022). Finetuned Language Models Are Zero-Shot Learners. ICLR 2022. arXiv:2109.01652. [FLAN] 21. Chung, H.W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., … & Wei, J. (2022). Scaling Instruction-Finetuned Language Models. JMLR 2024. arXiv:2210.11416. [FLAN-T5/PaLM] 22. Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., … & Hashimoto, T.B. (2023). Stanford Alpaca: An Instruction-Following LLaMA Model. Stanford CRFM Blog. [Alpaca self-instruct] 23. Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., & Awadallah, A. (2023). Orca: Progressive Learning from Complex Explanation Traces of GPT-4. arXiv:2306.02707. [Orca distillation]
Domain Fine-Tuning: 24. Singhal, K., Azizi, S., Tu, T., Mahdavi, S.S., Wei, J., Chung, H.W., … & Natarajan, V. (2023). Large Language Models Encode Clinical Knowledge. Nature, 620, 172–180. arXiv:2212.13138. [Med-PaLM] 25. Chen, Z., Cano, A.H., Romanou, A., Bonnet, A., Matoba, K., Salvi, F., … & Jaggi, M. (2023). MEDITRON-70B: Scaling Medical Pretraining for Large Language Models. arXiv:2311.16079. [Meditron] 26. Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., & Smith, N.A. (2020). Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks. ACL 2020. arXiv:2004.10964. [DAPT]
Continual Learning: 27. Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A.A., … & Hadsell, R. (2017). Overcoming Catastrophic Forgetting in Neural Networks. PNAS, 114(13), 3521–3526. arXiv:1612.00796. [EWC] 28. Chen, Z., & Liu, B. (2018). Lifelong Machine Learning (2nd ed.). Morgan & Claypool. ISBN 978-1-68173-434-3. [Comprehensive continual learning reference]
Provenance
- domain-correction: infrastructure → artificial-intelligence (concept is an AI/ML process, not infrastructure; IRI, URI, owl-class, same-as corrected accordingly)