Checkpoints are serialised snapshots of a machine-learning model’s complete trainable state — weight tensors, optimiser moment accumulators, learning-rate schedules, gradient scalers, random-number-generator seeds, and epoch/step counters — persisted to durable storage at regular intervals during…

In Plain Terms

  • Saved snapshots of a model taken partway through training, like hitting ‘save’ in a long game. If something crashes you can pick up from the last snapshot instead of starting over, and you can compare or reuse earlier versions.

Semantic Classification

  • domain-correction: infrastructure → artificial-intelligence (stub had wrong domain; concept is a core ML/AI training artifact with secondary blockchain bridge; corrected 2026-05-17)

Content

Compositional Relationships (Components)

SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:hasPart ai:WeightTensor))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:hasPart ai:OptimiserState))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:hasPart ai:LearningRateSchedule))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:hasPart ai:RNGState))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:hasPart ai:StepCounter))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:hasPart ai:EMAWeights))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:hasPart ai:ShardMetadata))

## Dependency Relationships
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:requires ai:PersistentStorage))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:requires ai:SerialisationFormat))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:requires ai:TrainingLoop))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:requires ai:CheckpointFrequencyPolicy))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:requires ai:ModelArchitecture))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:dependsOn ai:DistributedTraining))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:dependsOn ai:FileSystem))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:dependsOn ai:Serialisation))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:dependsOn ai:FaultTolerance))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:dependsOn ai:ModelRegistry))

## Capability Relationships
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:enables ai:TrainingResumption))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:enables ai:ModelVersioning))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:enables ai:TransferLearning))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:enables ai:FineTuning))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:enables ai:ModelServing))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:enables ai:CheckpointAveraging))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:enables ai:DistributedTrainingRecovery))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:enables ai:ExperimentReproducibility))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:supports ai:MLOps))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:supports ai:DiffusionModels))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:supports ai:LargeLanguageModels))

## Implementation Relationships
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:implements ai:Safetensors))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:implements ai:PyTorchDCP))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:implements ai:TensorFlowSavedModel))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:implements ai:DeepSpeedUCP))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:implements ai:FSDP))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:implements ai:HuggingFaceHub))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:implements ai:MLflow))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:uses ai:PickleSerialisation))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:uses ai:ONNX))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:uses ai:GGUF))

## Reduction Relationships
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:reducesRisk ai:TrainingDataLoss))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:reducesRisk ai:ModelDrift))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:reducesRisk ai:HardwareFailureLoss))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:reducesRisk ai:ArbitraryCodeExecution))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:reducesRisk ai:IrreproducibleExperiment))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:contrasts-with ai:StreamingInference))
SubClassOf(ai:Checkpoints
  ObjectSomeValuesFrom(ai:contrasts-with ai:StatelessExecution))

About Checkpoints

  • Checkpoints are the primary mechanism by which machine-learning practitioners preserve and exchange trained model state. The word “checkpoint” appears across both the ML and blockchain ecosystems with related but distinct semantics: in ML it is a serialised parameter snapshot; in blockchain it is a trusted anchor block hash or epoch-boundary vote commitment. This page focuses on the ML meaning, with blockchain usage addressed as a secondary bridge domain.
  • The ML checkpoint lifecycle spans three phases: (1) write — the training loop periodically calls a serialisation routine to flush the current state dict to storage; (2) store — the file is retained in a model registry, object store, or shared filesystem with versioning metadata; (3) load — at resume time or serving time the weights are deserialised back into a model instance and optionally post-processed (averaged, quantised, compiled, merged).
  • The frequency of writing and the number of retained checkpoints are configurable; common policies include: every-N-steps retention (write every 500-5000 steps depending on training speed and storage budget); keep-last-K deletion (maintain only the most recent K checkpoints to bound storage consumption); and best-metric retention (keep only the checkpoint achieving the lowest validation loss or highest evaluation accuracy regardless of recency).
  • Modern MLOps platforms treat checkpoints as first-class versioned artifacts tracked alongside code, data, and hyperparameters via tools such as Weights and Biases, MLflow, and DVC, ensuring every production model can be traced back to the exact training run, data snapshot, and code commit that produced it.

Mathematical Framework for Checkpoint-Based Techniques

EMA Weight Averaging

  • The exponential moving average (EMA) of model parameters is defined by the recurrence θ_ema(t) ← β · θ_ema(t-1) + (1-β) · θ(t) where θ(t) is the raw parameter vector at training step t and β ∈ (0,1) is the EMA decay coefficient. As t → ∞ this converges to the time-weighted average θ_ema = (1-β) · Σ_{k=0}^{∞} β^k · θ(t-k), with effective window length 1/(1-β): β=0.999 produces an effective window of 1000 steps.
  • Bias correction: At the beginning of training the EMA is initialised at zero or at the initial weights, producing a biased estimate. The bias-corrected estimate is θ̂_ema(t) = θ_ema(t) / (1 - β^t), the same correction used in the Adam optimiser. The timm ModelEmaV3 implementation applies bias correction during the warm-up phase (first 1/(1-β) steps) and then transitions to the standard EMA update.
  • Optimal decay selection: For diffusion model training, β values of 0.9999 are common, corresponding to an effective window of 10,000 steps. For Fine-tuning runs spanning only 1,000-5,000 steps, β values of 0.99 (window of 100 steps) are more appropriate because the fine-tuned weights diverge rapidly from their initial values. Post-hoc EMA (Karras et al. 2024) eliminates the need to pre-commit to β by storing snapshot EMA weights at multiple time points during training.
  • EMA and the bias-variance tradeoff: Higher β produces lower variance (smoother trajectory) but higher bias (slower adaptation to new training signal); lower β produces lower bias but higher variance (noisy tracking). The optimal β minimises a weighted combination of bias and variance and can be derived analytically as a function of the loss landscape’s local curvature, though this is rarely computed in practice — empirical tuning over β ∈ {0.99, 0.999, 0.9999} is the standard approach.

Checkpoint Averaging as Polyak-Ruppert Estimation

  • Checkpoint averaging computes θ_avg = (1/K) · Σ_{k=t-K+1}^{t} θ(k) where K is the number of checkpoints included in the average and θ(k) is the parameter vector at checkpoint step k. This is equivalent to uniform-weight Polyak-Ruppert averaging restricted to the final K iterates of training.
  • Theoretical guarantees: Under standard stochastic convex optimisation assumptions, Polyak-Ruppert averaging yields an iterate with expected excess risk O(1/√(KT)) vs O(1/√T) for the final iterate alone, representing an improvement of √K in statistical efficiency. For non-convex neural network training the guarantees are weaker but empirical improvements of 0.3-1.5 BLEU / 0.1-0.5 accuracy are consistently observed.
  • Stochastic Weight Averaging (SWA): Extends checkpoint averaging by using a cyclical learning rate schedule to ensure that the averaged checkpoints span diverse regions of the loss landscape rather than all clustering near the same local minimum. SWA applies the cosine annealing schedule η(t) = η_min + 0.5·(η_max - η_min)·(1 + cos(π·t_mod/T_cycle)) with cycle length T_cycle typically set to 1-5 epochs, collecting one checkpoint per cycle at the minimum learning rate point.
  • Batch normalisation update: Models averaged via SWA require a forward pass through the training data to update BatchNorm running statistics, because the averaged weights may be far from any single training trajectory and the running statistics accumulated during training may be mismatched to the averaged parameter values. This adds one epoch of forward computation (no gradients) to the SWA pipeline.
  • Practical SWA / checkpoint averaging recipes:
    • Neural machine translation (Fairseq): python scripts/average_checkpoints.py --inputs checkpoint_last5.pt checkpoint_last4.pt ... --output averaged.pt; average over the final 5-10 checkpoints of training; typically improves BLEU by 0.3-1.2 without any retraining
    • Image classification (timm): python train.py --swa --swa-lr 0.005 --swa-epoch-start 75 --swa-anneal-epochs 20; SWA applied in the final 20 epochs with constant low LR, then swa_utils.update_bn(loader, model) to fix BatchNorm
    • LLM fine-tuning (custom): average last 3-8 adapter checkpoint files before merging into the base model; reduces fine-tuning variance particularly for small datasets (<5K examples)
    • Diffusion training: maintain 3-5 short-EMA checkpoints (β=0.95, 0.99, 0.999) in parallel throughout training; reconstruct target EMA post-training via Karras et al. linear mixture formula

Checkpoint Frequency Policies

  • Checkpoint frequency policy selection involves a fundamental tradeoff between recovery cost (proportional to steps since last checkpoint × step compute time) and storage overhead (proportional to checkpoint size × frequency × retention count):
    • Fixed interval: Write every N steps, typically N = 200-2000 for medium-scale training (8-64 GPUs) and N = 50-500 for large-scale preemptible cloud training; the optimal N minimises the expected total wasted compute given failure rate λ (failures/step) — the expected waste is λ·N·T_step/2 where T_step is seconds per step, balanced against checkpoint overhead I/C where I is checkpoint write time and C is cycle length
    • Adaptive interval: Dynamically adjust N based on observed hardware failure rate, shortening the checkpoint interval on clusters with elevated GPU fault rates and lengthening it on stable hardware; implemented in Azure ML, AWS SageMaker HyperPod, and Google Cloud TPU training pipelines
    • Keep-last-K: Maintain only the most recent K checkpoints on disk (typically K = 3-5), deleting the oldest when a new one is written; bounds storage at K × checkpoint_size and is the default in HuggingFace Trainer (save_total_limit=5)
    • Best-and-last: Keep the single best validation metric checkpoint plus the most recent N checkpoints; ensures rollback capability while retaining the optimal model for deployment; standard in PyTorch Lightning ModelCheckpoint(save_last=True, save_top_k=3)
    • Milestone checkpoints: Permanently retain checkpoints at human-designated milestones (10%, 25%, 50%, 75%, 100% of training budget) regardless of the keep-last-K policy; enables retrospective analysis of training dynamics and intermediate model release for community evaluation

Safetensors Format Internals

  • The safetensors binary format consists of a fixed 8-byte little-endian uint64 header length field, followed by the JSON header of exactly that byte length, followed by the tensor data bytes in the declared order. The JSON header is a dictionary mapping tensor name strings to metadata objects of the form {"dtype": "F32", "shape": [768, 768], "data_offsets": [0, 2359296]} where data_offsets gives the [start, end) byte range within the tensor data section (relative to the end of the JSON header).
  • Security properties: Because the format is entirely declarative (the JSON header fully determines what bytes go where with no code execution), a conforming parser cannot be tricked into running arbitrary Python code. The Rust reference implementation performs bounds-checking on every data_offsets range to prevent integer overflow attacks; the Python bindings (safetensors pip package) delegate all parsing to the Rust core via PyO3.
  • Memory mapping: The safetensors.load_file(path, device="cpu") API uses mmap system call to map the file into the process address space, then creates tensor views into the mapped memory without copying bytes. For a 7B-parameter BF16 model (~14 GB), this reduces cold-load time from ~40 seconds (pickle, requiring full deserialisation into Python objects) to ~4 seconds (safetensors, just establishing the mmap and constructing view tensors).
  • Multi-file sharding protocol: When a model is split into multiple .safetensors shard files, the model.safetensors.index.json file provides the routing table: a top-level JSON object with "metadata" (total file size) and "weight_map" (dictionary mapping parameter name → shard filename). The transformers load_sharded_checkpoint utility reads the index, identifies which shards contain the requested parameters, and loads only the necessary shards — enabling partial model loading for layer-wise inference or selective fine-tuning.

Distributed Checkpoint Internals (PyTorch DCP)

  • PyTorch Distributed Checkpoint (DCP) organises saved checkpoints as a directory containing per-rank shard files and a metadata.json global index:
    • Each rank writes __0_0.distcp (rank 0 shard file), __1_0.distcp (rank 1 shard file), etc., where filenames encode __<rank>_<file_index>.distcp
    • The global metadata.json contains a state_dict_metadata section describing, for each tensor key, its global shape, dtype, and a list of ChunkStorageMetadata records identifying which rank file contains which byte slice of the global tensor
    • At load time, DCP reads metadata.json and distributes load assignments across ranks such that each rank reads the byte slices it needs for its local DTensor allocation, with no rank needing to read the full checkpoint file
    • Resharding at load time reconstructs the mapping from global-tensor-slice to local-rank-allocation for a target topology (e.g. 8 ranks → 32 ranks) by solving a bin-packing assignment of global slices to local ranks, then issuing peer-to-peer torch.distributed.recv / send operations to distribute bytes that landed on incorrect ranks after the initial parallel reads
  • The dcp.save(state_dict, checkpoint_id=path) API accepts any Stateful object (implementing state_dict() / load_state_dict()) and automatically handles sharding; models wrapped with FSDP2 expose their local DTensor shards through the standard state_dict() interface, making DCP compatible with FSDP2 without any special-casing.
  • Two-level checkpointing for I/O bottleneck relief:
    • Level 1 (fast): write checkpoint to node-local NVMe SSD at training speed; each node writes only its own shards to /nvme/checkpoint_tmp/, completing in seconds
    • Level 2 (durable): asynchronously copy Level 1 checkpoint to shared network storage (Lustre, GPFS, S3) using rsync or aws s3 cp in a background subprocess; takes 10-120 seconds depending on network bandwidth
    • The Level 1 checkpoint is the live recovery target (fast resume from local disk on same node); Level 2 is the durable recovery target (resume on different hardware after node failure)
    • NVIDIA’s TorchTitan and IBM’s FMS both implement two-level checkpointing; ByteRobust extends this to in-memory checkpointing (Level 0 in GPU pinned memory) for sub-second checkpoint overhead
  • Checkpoint integrity verification:
    • SHA-256 hashing of checkpoint files before and after transfer to detect corruption during network transfer or storage write errors
    • Hugging Face Hub automatically computes and stores SHA-256 hashes for every uploaded file, verified at download time by the Hub client
    • For safetensors files, the JSON header’s data_offsets ranges are validated against the actual file size by the Rust parser, detecting truncation or corruption that would not be caught by simple file-size checks

Components / Architecture

State Dict Components

  • A complete PyTorch checkpoint dictionary typically contains the following components:
    • model_state_dict — parameter and buffer tensors indexed by dotted module name (e.g. encoder.layer.0.attention.self.query.weight)
    • optimizer_state_dict — momentum vectors, Adam second-moment estimates (exp_avg, exp_avg_sq), and step counts per parameter group
    • lr_scheduler_state_dict — last epoch, base learning rates, warmup state, cycle parameters
    • scaler_state_dict — gradient scaler for mixed-precision training with torch.cuda.amp.GradScaler
    • rng_states — CPU and CUDA RNG seeds enabling exact bit-for-bit training resumption
    • epoch / global_step counters for resume position
    • best_metric — the validation metric value at which this checkpoint was saved, used by best-model retention logic
    • config — model architecture hyperparameters (optional but strongly recommended for reproducibility)
  • The minimum viable checkpoint for fault-tolerant resumption is the model and optimiser state dicts; all other components improve reproducibility but are not required for correctness.

Serialisation Formats

  • Four formats are in active production use as of 2025-2026:
    • .pt / .pth (PyTorch pickle) — fastest to write and read within a trusted environment, single-file, unsafe for shared distribution due to arbitrary code-execution risk via pickle deserialisation; legacy format, now deprecated for public sharing
    • .safetensors — flat binary with JSON header, mmap-safe, zero code-execution risk, default on Hugging Face Hub for all major models, 5-10× faster load than pickle for large models, now governed by the PyTorch Foundation (donated April 2026)
    • SavedModel — TensorFlow directory bundle containing the computation graph as a saved_model.pb protobuf plus variable shards; cross-language safe and serves directly via TensorFlow Serving; ecosystem-locked to TF/Keras
    • GGUF (GPT-Generated Unified Format) — community-developed format used by llama.cpp for quantised LLM weights, embedding quantisation metadata (bits per weight, block size, quant type) in a self-describing header; primary format for consumer-hardware LLM inference
  • ONNX is an exchange format rather than a training checkpoint format, primarily used for cross-framework inference deployment after training is complete.

Sharding for Large Models

  • Models exceeding a single GPU’s memory require sharded serialisation:
    • Hugging Face Hub splits models into shards of configurable size (default 5 GB per shard) and provides a model.safetensors.index.json manifest mapping parameter names to shard filenames; the transformers library’s save_pretrained and from_pretrained methods handle shard assembly transparently
    • PyTorch DCP natively saves per-rank shards under a directory with a metadata.json durable index; shards can be loaded back into a different number of ranks via the resharding path
    • DeepSpeed ZeRO-3 saves each rank’s parameter slice independently plus a single optimizer-state directory; the zero_to_fp32.py utility reconstructs full FP32 parameters from ZeRO-3 shards for evaluation or fine-tuning
    • NVIDIA NeMo uses the NVIDIA Distributed Checkpoint (NDC) format layered on DCP for both FP8 and BF16 checkpoints in large-scale pre-training (e.g. Megatron-Core runs)

EMA State

  • Diffusion Models universally maintain an EMA copy alongside training weights:
    • The EMA state dict contains a shadow copy of every parameter tensor, updated each training step without gradient accumulation using the exponential decay rule θ_ema ← α · θ_ema + (1-α) · θ
    • At inference time only the EMA weights are loaded; the raw training weights are used only for computing gradients during fine-tuning or continued training
    • The timm library’s ModelEmaV3 class and the HuggingFace diffusers EMAModel class provide standard implementations with configurable decay schedules and support for EMA warm-up over the first N steps
    • Post-hoc EMA (Karras et al. 2024) stores periodic snapshots of short-EMA weight states during training; after training completes, any target EMA profile (any decay α) can be reconstructed as a linear mixture of the stored snapshots, enabling EMA hyperparameter search to be decoupled from the training compute budget

Distributed Checkpoint Coordination

  • Under FSDP each rank holds a shard of each parameter tensor as a DTensor (distributed tensor):
    • torch.distributed.checkpoint’s dcp.save / dcp.load API serialises each rank’s local DTensor slices in parallel to a shared directory, completing in O(max_rank_write_time) rather than O(sum_rank_write_time)
    • A global metadata.json file records the global tensor shape, dtype, and per-rank byte offset for each parameter, enabling the load path to reconstruct any target sharding configuration
    • FSDP2 (torch.distributed.fsdp.fully_shard, PyTorch 2.4+) integrates more cleanly with DCP and removes the constraint that full state dicts must be materialised on rank-0 before saving, which was a memory bottleneck in FSDP1
    • DeepSpeed Universal Checkpointing (UCP) extends cross-parallelism support: a ZeRO-3 checkpoint can be loaded into a tensor-parallel + pipeline-parallel configuration without manual resharding scripts, enabling elastic resource management (scaling training from 512 to 2048 GPUs mid-job)

Use Cases / Major Families

Training Resumption after Failure

  • Failure modes in large-scale ML training:
    • GPU memory errors (ECC uncorrectable errors, CUDA illegal memory access) — rate ~0.1-0.5% per GPU per day at scale, translating to 5-25 failures/day in a 5,000-GPU job
    • NVLink / NVSwitch faults — fabric-level failures causing entire NVLink islands (8-32 GPUs) to drop; recovery requires node replacement rather than process restart
    • InfiniBand transients — temporary fabric congestion causing NCCL collective communication timeouts; often recoverable without node replacement but requires process restart from last checkpoint
    • SLURM preemption — scheduled job preemption for higher-priority workloads on shared academic HPC clusters; typically gives 30-120 seconds’ warning, sufficient to write a checkpoint if the save routine is asynchronous
    • Power outage / data-centre cooling failure — rare but catastrophic without off-node checkpoint persistence; requires checkpoints on shared network storage rather than node-local NVMe
  • Clusters running multi-week LLM pre-training (GPT-4 scale, 1024+ GPUs) experience multiple hardware failures per day due to GPU memory errors, NVLink faults, InfiniBand transients, and SLURM preemption. Without checkpointing every failed run would restart from epoch 0, wasting potentially weeks of compute worth millions of dollars.
  • Checkpoint frequency trades storage I/O overhead against recovery cost: writing every 100-500 steps at sub-minute granularity costs 0.5-2% training throughput but reduces replay on failure to minutes rather than weeks. For a 70B-parameter model in BF16, each checkpoint is ~140 GB; at NVMe SSD bandwidth of 5-10 GB/s this takes 14-28 seconds per rank, representing a significant I/O overhead when writing from 8,000 ranks simultaneously.
  • ByteDance’s ByteRobust system (2025) achieves every-step checkpointing with under 0.9% overhead by storing snapshots in fast GPU-attached NVMe memory before asynchronously flushing to persistent storage, achieving 97% effective training time ratio across a 9,600-GPU, three-month training job.
  • FlashRecovery (2025) reduces failure recovery to under 150 seconds through active failure detection and partial re-computation strategies, outperforming checkpoint-based recovery by 4-10× on detection-to-resume latency.

Best-Model Selection

  • Validation loss oscillates during training; the final step checkpoint is rarely the best generalising model due to late-stage overfitting on the training data distribution. Keeping the checkpoint with the lowest validation perplexity or highest evaluation accuracy is standard practice.
  • Tools: Weights and Biases ModelCheckpoint callback, HuggingFace Trainer’s load_best_model_at_end=True, PyTorch Lightning’s ModelCheckpoint(monitor="val/loss", mode="min", save_top_k=3).
  • For LLM pre-training where validation loss decreases monotonically, best-model selection is replaced by best-step selection based on downstream benchmark performance (MMLU, HumanEval, HellaSwag) evaluated at periodic checkpoints.
  • Checkpoint selection strategies by training regime:
    • Supervised classification fine-tuning: monitor val_accuracy or val_f1; typical improvement 0.5-2% over final step
    • Language model fine-tuning: monitor val_loss (perplexity); final checkpoint frequently overfits if fine-tuning data is small (<10K examples), with best checkpoint 20-50% into training
    • Reinforcement Learning from Human Feedback (RLHF): monitor win-rate against a reference policy on a held-out preference test set; checkpoint selection more complex due to reward hacking dynamics
    • LLM pre-training: evaluate checkpoints at 10%, 25%, 50%, 75%, 100% of compute budget on benchmark suites; final pre-training checkpoint is almost always best for pre-trained capabilities, with fine-tuning applied afterwards
    • Diffusion model training: monitor FID (Fréchet Inception Distance) on a validation image set; EMA weights are always used for FID evaluation; raw training weights are never evaluated directly in diffusion pipelines

Checkpoint Averaging

  • Averaging the weight vectors of the last K checkpoints produces a parameter set with lower variance than any individual checkpoint, consistently improving evaluation metrics without any additional training compute:
    • Neural machine translation: Fairseq’s --average-last-n-checkpoint flag averages the final 5-10 checkpoints, improving BLEU by 0.3-1.5 points on WMT benchmarks
    • LLM fine-tuning: averaging the last 3-8 fine-tuning checkpoints reduces variance from stochastic gradient noise, particularly beneficial for small fine-tuning datasets (<10K examples)
    • Stochastic Weight Averaging (SWA) formalises this as a cyclical learning rate (cosine with restarts) that forces the model to traverse a wide, flat region of the loss landscape, then averages the checkpoints encountered at each cycle minimum
    • Model Soups (Wortsman et al. 2022) demonstrated that averaging checkpoints from independently fine-tuned models with different hyperparameter settings produces merged models outperforming any individual member on ImageNet and CLIP evaluation suites
  • A 2025 analysis found that EMA with decay 0.2 applied over the last 6 checkpoints effectively restores curriculum learning benefits lost during long training runs, often surpassing traditional warmup-stable-decay methods.

Transfer Learning and Fine-tuning

  • Pre-trained checkpoints from Large-Scale Pretrained Foundation Model (BERT, RoBERTa, LLaMA, Mistral, GPT-2, Gemma, Qwen) published on Hugging Face Hub enable practitioners to initialise domain-specific models with zero or minimal training data:
    • AutoModel.from_pretrained("bert-base-uncased") downloads the .safetensors checkpoint, verifies the JSON header, deserialises weights via mmap, and loads them into the appropriate architecture class
    • The Hub’s model.safetensors.index.json enables lazy loading of only the required parameter shards for architectures where only specific layers are fine-tuned
  • Fine-tuning techniques including LoRA, QLoRA, and prefix-tuning checkpoint only the delta weights (adapter tensors) rather than the full model, reducing checkpoint size from gigabytes (full model) to megabytes (adapter only). PEFT library checkpoints save adapter_config.json + adapter_model.safetensors, typically 10-500 MB for a 7-70B base model.
  • Checkpoint anatomy for parameter-efficient fine-tuning:
    • LoRA checkpoints contain low-rank adapter matrices A (d × r) and B (r × k) for each adapted linear layer, where r is the rank (typically 4-64); the merged weight is W + α/r · B·A where α is the LoRA scaling factor
    • QLoRA checkpoints additionally store the quantised base model weights (4-bit NF4 format via bitsandbytes) alongside the full-precision LoRA adapters; the base model is frozen during training and the adapter alone is saved in the checkpoint
    • Prompt-tuning checkpoints save only the continuous prompt embeddings (typically 8-128 virtual tokens × embedding dimension), typically 100KB-5MB per task
    • IA³ checkpoints save per-layer scaling vectors rather than full rank decompositions, achieving 10-100× smaller checkpoints than LoRA for comparable fine-tuning performance on certain tasks
  • Cross-organisation checkpoint sharing protocols: Hugging Face Hub provides SHA-256 hash verification for every file, Git-LFS storage for large binaries, access-controlled repositories (gated models requiring accepted usage terms, e.g. LLaMA-4, Gemma 3), and model cards (README.md + model_card_data.yaml) recording training data, base model, evaluation results, and intended use — creating a standardised checkpoint exchange protocol used by 500,000+ model repositories as of 2026.

Diffusion Model Inference

  • Diffusion Models ship two checkpoint artefacts: the EMA weights used for inference and the training weights used for continued fine-tuning:
    • The Stable Diffusion ecosystem historically used .ckpt (PyTorch pickle) files distributed via CivitAI; security researchers demonstrated arbitrary code execution via malicious .ckpt files in 2023, triggering a community migration to .safetensors
    • FLUX.1, SDXL-Turbo, Stable Diffusion 3.5, and all major Stability AI releases now ship exclusively in .safetensors format; CivitAI enforces .safetensors for new model uploads in its verification programme
    • The diffusers library’s StableDiffusionPipeline.from_pretrained() loads EMA weights automatically; the --resume_from_checkpoint flag in training scripts loads training weights with optimiser state for continued training

Blockchain Checkpoints

  • Bitcoin’s legacy checkpoint mechanism compiled hardcoded block hashes into chainparams.cpp to prevent long-range reorganisation attacks before headers-first sync was mature; the last hardcoded checkpoint (block 295,000) was added in 2014, and the practice is now considered obsolete.
  • Ethereum’s beacon chain formalises checkpoints as (epoch, block_root) pairs:
    • Casper FFG justifies a checkpoint when ≥ 2/3 of staked ETH attests to it in a vote
    • A checkpoint is finalised when a subsequent checkpoint is also justified, providing cryptoeconomic finality: reversing a finalised checkpoint would require burning ≥ 1/3 of all staked ETH (currently ~$30-50B), making reversion economically prohibitive
    • Ethereum’s weak-subjectivity checkpoint sync allows new nodes to obtain a recent finalised (epoch, block_root) from trusted community endpoints (maintained at eth-clients.github.io/checkpoint-sync-endpoints) and begin syncing from that anchor rather than from genesis block, reducing sync time from days to hours while remaining immune to long-range attacks

Checkpoint-Based Model Analysis and Interpretability

  • Periodic checkpoint snapshots enable post-hoc analysis of training dynamics without re-running training:
    • Loss trajectory analysis: Plotting training/validation loss across checkpoint steps reveals phenomena such as loss spikes (transient divergence at large learning rate), double descent (loss initially increases before decreasing as model capacity grows), and grokking (sudden delayed generalisation)
    • Representation similarity analysis: Computing CKA (Centred Kernel Alignment) between layer activations at checkpoint steps k and k+N quantifies how rapidly each layer’s representations are changing; earlier layers typically stabilise first, while later task-specific layers continue evolving throughout training
    • Gradient norm tracking: Monitoring the L2 norm of the gradient for each parameter tensor at checkpoint steps identifies which modules are learning most actively at each training phase; attention heads often show high gradient norms early in fine-tuning before stabilising
    • Weight norm trajectory: Tracking ||W||_F (Frobenius norm) for each layer’s weight matrix over checkpoint steps reveals implicit regularisation dynamics — weight norms typically grow early in training and plateau as gradient descent approaches a local minimum
    • Probing classifiers at intermediate checkpoints: Training lightweight linear classifiers on intermediate layer activations at each checkpoint step reveals when specific capabilities (e.g. syntactic structure, factual knowledge) emerge in the model’s representations; used extensively in BERTology and mechanistic interpretability research

Model Serving and Versioning

  • Model Registry platforms promote checkpoints through lifecycle stages — Development → Staging → Production → Archived — with approval gates requiring human sign-off:
    • MLflow Model Registry stores checkpoint artifacts with linked training run metadata (metrics, params, code commit, data version)
    • Weights and Biases Artifact Registry provides lineage graphs linking model checkpoints to upstream dataset artifacts and downstream evaluation runs
    • SageMaker Model Registry and Vertex AI Model Registry integrate with cloud serving infrastructure for automated deployment of approved checkpoints
    • MLflow 3.0 (2025) extended registry semantics to generative AI agents, tracking not only weight artifacts but prompt templates, fine-tuned LoRA adapters, retrieval configurations, and evaluation run metadata linked by git commit hash

Security Landscape for Checkpoint Distribution

  • The transition from pickle-based to safetensors-based checkpoint distribution was catalysed by a series of security incidents demonstrating practical exploitation of pickle deserialisation:
    • 2022-2023: Multiple malicious .ckpt files were discovered on CivitAI and Hugging Face Hub containing code that spawned reverse shells or exfiltrated API tokens when loaded with torch.load
    • The fickling tool (released by Trail of Bits, 2022) automated auditing of pickle files for malicious bytecode, finding that approximately 0.3-0.5% of sampled public checkpoint files contained suspicious but non-malicious pickle opcodes (e.g. REDUCE operations for benign utility classes)
    • Hugging Face responded by deploying pickle scanning infrastructure on Hub uploads and by setting safetensors as the default upload format for new models in 2023
    • PyTorch added weights_only=True as the default in torch.load (PyTorch 2.6+), which restricts deserialisation to safe tensor types and raises an error on any pickle opcode that could execute arbitrary code
  • The model merging security surface introduces new risks: mergekit and similar tools operate on .safetensors shards and are therefore immune to pickle attacks, but the provenance of source checkpoints being merged must still be validated — a merged model inherits any backdoors or poisoned weights present in its source checkpoints if those were inserted at the weight level (trigger-pattern backdoors that are not detectable from the weight statistics alone).

Academic Context

  • The theoretical basis for checkpoint-based fault tolerance in distributed systems derives from Chandy-Lamport distributed snapshots (1985), which established that a consistent global snapshot can be taken from a distributed computation without stopping all processes simultaneously — a property that holds for iterative optimisation where the state is a parameter vector and the process is stochastic gradient descent.
  • Polyak-Ruppert averaging (1992) proved that the time-averaged iterate of SGD converges at the optimal rate; EMA is a computationally efficient approximation that forgoes storage of the full trajectory while retaining most of the statistical benefit of full averaging. The formal equivalence between Polyak-Ruppert averaging and EMA in the large-decay-constant limit was established by subsequent work in stochastic approximation theory.
  • The flat minima hypothesis (Hochreiter and Schmidhuber 1997) provides the geometric intuition for why checkpoint averaging improves generalisation: checkpoints along a training trajectory that collectively span a broad, flat basin of the loss landscape produce an average that sits at the centre of the basin, which is more robust to perturbation than any individual sharp minimum encountered during training.
  • The Universal Checkpointing paper (Lian et al. 2024, arXiv:2406.18820) provides formal treatment of checkpoint transformation under arbitrary parallelism strategy changes, proving the transformation is lossless for any model where parameters are partitioned without overlap across ranks — a property satisfied by ZeRO, FSDP, and pipeline parallelism.
  • The safetensors security model was independently audited by Trail of Bits (2023), finding no vulnerabilities in the file format itself while noting that consumers must validate the JSON header to prevent integer overflow in byte-offset arithmetic — a defence implemented in the reference implementation’s Rust parser.
  • Model merging theory (Wortsman et al. 2022, Yadav et al. 2023 TIES-merging, Yu et al. 2023 DARE) demonstrates that linear combinations of fine-tuned checkpoint parameters often lie in the same loss basin as their sources, enabling ensemble-like multi-task performance from a single merged parameter set without running multiple models at inference time.
  • Linear mode connectivity (Frankle et al. 2020, Entezari et al. 2022) is the theoretical underpinning of checkpoint merging: two models fine-tuned from the same pre-trained checkpoint are often linearly connected in parameter space — meaning the linear interpolation path between their weight vectors contains no high-loss barrier — a property absent for randomly initialised models and what makes checkpoint averaging effective for models sharing a common pre-training origin.
  • The lottery ticket hypothesis (Frankle and Carlin 2018) relates to checkpoints through “winning tickets” — sparse subnetworks identified at early training checkpoints that, when re-initialised to their initial values and retrained in isolation, match the full network’s performance. Checkpoint-based iterative magnitude pruning (IMP) is the standard protocol: train to convergence, save checkpoint, prune smallest-magnitude weights globally, reload initial checkpoint with pruned weights zeroed, repeat until target sparsity is reached.
  • Training dynamics analysis via checkpoints: A growing research programme uses periodic checkpoint snapshots to study how neural network representations evolve during training — tracking singular value spectra of weight matrices, representational similarity across layers (CKA, SVCCA metrics), and neuron specialisation via probing classifiers. This “checkpoint archaeology” approach provides mechanistic insights into phenomena including grokking (sudden delayed generalisation 100-1000 steps after apparent convergence), loss spikes (transient divergence followed by recovery without intervention), and phase transitions in capability emergence at specific training compute thresholds.

Current Landscape (2026)

  • As of May 2026, the safetensors format dominates model distribution on Hugging Face Hub: LLaMA-4, Qwen-3, Deepseek-R1, Gemma 3, Mistral Large, and Claude 3.5 (Haiku weights, when released) all ship exclusively in .safetensors shards. Pickle-based .bin files remain available for legacy compatibility but are no longer the default download artifact.
  • The PyTorch Foundation governance transfer of safetensors (April 2026) signals movement toward an open, cross-framework checkpoint standard under the Linux Foundation. PyTorch 2.6+ integrates native safetensors support into torch.serialization, replacing the insecure torch.load(weights_only=False) pattern.
  • PyTorch DCP has stabilised in PyTorch 2.3-2.5 and is the default checkpointing mechanism in TorchTitan (Meta’s production LLM pre-training framework), NVIDIA NeMo 2.0, IBM FMS, and Lightning Fabric. FSDP2 (torch.distributed.fsdp.fully_shard) replaces the legacy FSDP1 API with a cleaner DCP integration.
  • HuggingFace Accelerate 1.x (2024-2025) unified checkpoint saving across FSDP, DeepSpeed, and single-GPU setups through a common accelerator.save_state / accelerator.load_state API, abstracting the backend-specific sharding logic.
  • At the largest training scales (10,000+ GPUs for frontier LLMs), checkpoint I/O is a major throughput bottleneck: writing a 70B-parameter model in BF16 produces ~140 GB of weight data per checkpoint; production systems mitigate this by writing to node-local NVMe and asynchronously flushing to shared storage (NVIDIA’s “two-level checkpointing”), using all-gather communication to consolidate full-precision weights on designated checkpoint ranks.
  • Model merging from checkpoints has matured from research curiosity into production technique: SLERP, TIES-merging, DARE, and linear mode connectivity methods are implemented in mergekit (2024-2025, 15,000+ GitHub stars), which operates directly on .safetensors shards without loading full models into GPU memory.
  • MLOps tooling in 2026 universally supports checkpoint lineage tracking: DVC, MLflow, Weights and Biases, and Comet all link checkpoint artifacts to the exact git commit, dataset version, hyperparameter set, and evaluation metrics that produced them, enabling automated rollback when a new production checkpoint degrades downstream benchmarks.
  • The Hugging Face Hub model-cards standard now mandates checkpoint provenance fields — base model, training data, training framework version, and evaluation results — for any model requesting the “Hub Verified” badge, enforcing documentation discipline across the open-source model ecosystem.
  • Checkpoint storage economics at scale (2025-2026):
    • A frontier LLM checkpoint (70B params, BF16): ~140 GB per checkpoint; at 28/month per training run
    • Checkpointing every 500 steps for a 500,000-step training run generates 1,000 checkpoints; with keep-last-5 policy the live storage footprint stays at ~700 GB (0.01/GB transfer: $1,400 egress cost for distributed reads)
    • Safetensors’ mmap enables “streaming” checkpoint reads from object storage (S3, GCS, Azure Blob) via fsspec without downloading the full file; this enables model evaluation directly against cloud-stored checkpoints during hyperparameter sweeps
    • Quantised checkpoints (GGUF 4-bit, 70B model) reduce size to ~40 GB, reducing storage cost by 3-4× at the cost of slight quality degradation relative to full-precision BF16 weights
  • Multi-framework interoperability in 2026: The KerasHub (Keras 3.x) from_preset() API loads safetensors weights from Hugging Face Hub into Keras models regardless of the source framework (PyTorch vs JAX vs TensorFlow), enabling cross-framework weight reuse without manual conversion scripts. NVIDIA’s NeMo 2.0 provides nemo_to_hf and hf_to_nemo conversion utilities for bidirectional checkpoint exchange between the NeMo Megatron format and Hugging Face safetensors layout.

UK Context (Imperial / Edinburgh / UCL / Cambridge / Manchester academic; Northern English industrial)

Academic Leadership

  • The University of Edinburgh’s School of Informatics (ILCC — Institute for Language, Cognition and Computation) has been a leading contributor to checkpoint-efficient Neural Machine Translation training, with researchers developing the checkpoint-averaging conventions adopted by Fairseq and OpenNMT; Edinburgh’s work on curriculum learning and checkpoint-based learning rate restarts contributed to the cosine-with-warm-restarts schedule universally used in LLM pre-training.
  • Imperial College London’s Department of Computing hosts the Data Science Institute, where checkpoint-based fault-tolerance research for distributed genomics ML pipelines has been conducted in collaboration with the NHS Genomics England programme, ensuring training runs on sensitive clinical data can resume without reprocessing restricted datasets.
  • UCL’s Centre for Artificial Intelligence (Prof. David Silver’s group) uses checkpoint archiving extensively in reinforcement learning to save policy snapshots for replay buffer management, behaviour cloning baselines, and evaluation at intermediate training steps — critical for the multi-week training runs required for game-playing agents.
  • The University of Cambridge Computer Laboratory’s Systems Research Group has published on checkpoint-aware storage system design, including work on checkpoint co-placement in distributed file systems to minimise I/O interference between simultaneous checkpoint writes from different training jobs.
  • The University of Manchester’s Advanced Research Computing cluster (CSF4, 60,000+ CPU cores, GPU partition) supports training runs using SLURM preemption, requiring all compute jobs to checkpoint every 2-4 hours for fair-share scheduling; Manchester’s machine translation and computational biology groups have made checkpoint management tooling contributions to the PyTorch Lightning community.
  • Alan Turing Institute (ATI) checkpoint governance recommendations (published 2024): permanently archive the final production checkpoint, the best-validation-metric checkpoint, and milestone checkpoints at 25%/50%/75% of training budget for any research result intended for publication, ensuring peer reviewers and replication studies can access the same model states used in the original work.
  • UK Research and Innovation (UKRI) data management requirements: UKRI grants now require that model checkpoints for published findings be deposited in a persistent repository (Zenodo, institutional repository, or domain-specific archive) with a DOI and minimum 10-year retention, treating checkpoints as research data subject to FAIR (Findable, Accessible, Interoperable, Reusable) data principles — driving standardised format adoption (safetensors, GGUF) in UK academic ML groups.
  • NHS AI Lab model assurance framework (2024): AI models deployed in clinical settings must maintain checkpoint lineage records linking the deployed model version to training data version with patient cohort description, model architecture specification, training hyperparameters, evaluation results on external validation set, and equality impact assessment — with records retained for the lifetime of clinical deployment plus 10 years.

Northern English Industrial Applications

  • The Hartree Centre (Science and Technology Facilities Council, Daresbury, Cheshire) operates the SCARF HPC cluster and has published checkpoint optimisation guidelines for UK research HPC users; Hartree provides training and consultancy on distributed checkpointing strategies to UK pharmaceutical and manufacturing SMEs adopting ML. Hartree’s PEARL (Platform for Extreme-scale Analytics and Research on Lustre) storage system supports parallel checkpoint writes at >10 GB/s aggregate bandwidth, supporting multi-node training jobs from UK industrial partners including AstraZeneca, GSK, and Rolls-Royce.
  • The N8 Research Partnership (eight Northern English universities: Newcastle, Leeds, Sheffield, Manchester, Durham, York, Lancaster, Liverpool) shares checkpoint management best practices across the N8 CIR Bede GPU cluster (32 IBM Power9 nodes, 4 NVIDIA V100 GPUs each), hosting collaborative deep-learning projects in climate modelling, materials science, and medical imaging requiring robust fault-tolerant training. The N8 has published a “Checkpoint Best Practices for HPC” guide tailored to SLURM preemptible environments widely used across Northern English academic computing.
  • Sheffield’s INSIGNEO Institute for in silico Medicine uses checkpoint-based training for personalised cardiac simulation models in collaboration with Sheffield Teaching Hospitals NHS Trust, where training runs span multiple weeks on patient-specific finite-element mesh data. Checkpoint versioning enables auditable model updates as new patient cohorts are incorporated into training sets, satisfying UK MHRA AIaMD documentation requirements.
  • Leeds Institute for Data Analytics (LIDA) applies checkpoint versioning to multi-site clinical trial ML pipelines involving NHS trusts across Yorkshire and Humber, ensuring regulatory auditability under UK MHRA guidance on AI as a Medical Device (AIaMD) — every production inference model must be traceable to an approved checkpoint with documented training provenance, evaluation metrics, and dataset version.
  • The Connected Places Catapult (headquartered in London with Northern England programmes in Newcastle and Leeds) applies checkpoint management to urban mobility prediction models trained on continuous smart city sensor streams, where preemptible cloud compute requires robust resume capability. Newcastle’s urban data observatory provides continuous sensor feeds (pedestrian counters, traffic cameras, air quality sensors) requiring online learning pipelines with hourly checkpoint cycles.
  • Industrial applications in Northern England:
    • AstraZeneca (Macclesfield, Cheshire): uses checkpoint-based molecular property prediction models for ADME (absorption, distribution, metabolism, excretion) property screening; checkpoint versioning required for regulatory submissions under ICH M7 mutagenicity assessment guidelines
    • BAE Systems (Warton, Lancashire; Barrow-in-Furness): applies checkpoint-based computer vision for automated defect detection in aircraft composite manufacturing; checkpoint provenance required for AS9100 aerospace quality management certification
    • Siemens Digital Industries (Lincoln): uses ML model checkpoints for predictive maintenance of industrial turbines; checkpoint versioning enables model updates when new failure modes are observed without losing accumulated training on existing failure patterns
    • National Grid (Warwick): applies checkpoint-based energy demand forecasting models updated nightly with rolling training windows; checkpoint retention policies balance storage cost against the need for retrospective investigation when forecast errors occur

Future Directions (2026-2030)

Checkpoint-Free Recovery

  • Research active in 2025-2026 proposes log-based recovery that reconstructs model state from the sequence of gradient updates without materialising periodic weight snapshots, analogous to write-ahead logging in databases. “All is Not Lost: LLM Recovery without Checkpoints” (arXiv:2506.15461, 2025) demonstrates recovery from complete node failure without stored checkpoints by replaying the gradient update sequence from persistent gradient logs.
  • FlashRecovery (2025) demonstrates near-real-time recovery via active failure detection and partial re-computation, reducing recovery overhead to under 150 seconds — approaching the latency of in-memory checkpoint restoration without the storage overhead.
  • If validated at frontier scale (100,000+ GPUs), checkpoint-free approaches could eliminate the checkpoint I/O bottleneck entirely, recovering the 1-2% throughput lost to periodic writes.

Post-hoc EMA Tuning

  • Karras et al.’s post-hoc EMA framework (2024) is expected to become standard practice for Diffusion Models training: rather than committing to a single EMA decay at training time, practitioners will store short-EMA snapshots and reconstruct any desired profile post-training. Tools implementing this are expected in diffusers and timm by 2027.

Universal Checkpoint Exchange Standard

  • The donation of safetensors to the PyTorch Foundation signals movement toward an open cross-framework checkpoint standard. Expected near-term developments:
    • Native safetensors support in TensorFlow/Keras (partially delivered via KerasHub 2024 weight import)
    • An ONNX-to-safetensors round-trip specification enabling lossless exchange between training and inference representations
    • Standardised quantisation metadata headers (bits per weight, block size, quant scheme) in safetensors, absorbing GGUF’s use case into the main format

Model Merging as First-Class Operation

  • Checkpoint merging is maturing from a post-hoc research technique into a first-class MLOps operation. By 2028, model registries are expected to natively support merge provenance tracking — linking a merged model checkpoint to its constituent source checkpoints with documented merge algorithm, weights, and evaluation results. This enables regulatory auditability for merged models deployed in high-risk applications.

Regulatory Auditability

  • UK MHRA, EU AI Act, and US FDA guidance on AI/ML-based Software as a Medical Device increasingly requires complete checkpoint provenance — linking every production inference model to training data version, code commit, evaluation results, and approval workflow — creating demand for:
    • Checkpoint management platforms with immutable audit logs (append-only object store buckets with cryptographic hash verification)
    • Cryptographic signing of checkpoint artifacts (GPG or Sigstore signatures on .safetensors files, verifiable against a public key registry)
    • Third-party attestation services certifying that a deployed checkpoint matches the approved training lineage
  • The EU AI Act (effective August 2024, enforcement from August 2026 for high-risk systems) mandates that high-risk AI systems maintain complete technical documentation including “the specifications, the versions and modifications of the system that have been used for its development” — interpreted by legal experts to require checkpoint versioning with signed provenance for any AI system deployed in high-risk domains (medical devices, critical infrastructure, law enforcement, education).
  • The UK government’s AI Safety Institute (AISI) published “Frontier AI Safety Commitments” in 2023 requiring major AI developers to provide safety evaluations at key training checkpoints before deployment; this institutionalised checkpoint-based safety evaluation as a regulatory requirement for frontier AI models in the UK.

Emerging Checkpoint Formats and Protocols

  • GGUF v3 extensions (2025): The GGUF format, maintained by the llama.cpp community, extended its metadata header in 2025 to support: multimodal model architecture descriptions (vision encoders alongside language decoders), Mamba/state-space model architectural metadata, grouped-query attention configuration, and RoPE (Rotary Position Embedding) configuration — making GGUF increasingly self-describing and reducing the need for separate architecture configuration files alongside checkpoint weights.
  • MLflow AI Gateway checkpoint proxy (2025): MLflow 3.0 introduced an AI Gateway that proxies checkpoint loading requests through a policy engine, enforcing access controls (restricting which users can load which checkpoint versions), rate-limiting checkpoint download bandwidth, and logging all checkpoint access events to an immutable audit trail — addressing enterprise governance requirements for checkpoint access control in regulated industries.
  • Hugging Face Hub webhooks for checkpoint events (2024): The Hub’s webhook API now emits events on checkpoint upload, update, and deletion, enabling downstream CI/CD systems (GitHub Actions, GitLab CI, Jenkins) to trigger automated evaluation pipelines whenever a new checkpoint version is published, closing the loop between training and evaluation in automated MLOps workflows.
  • Checkpoint compression: Research in 2024-2025 explored applying lossless compression (Zstandard, LZ4) and lossy compression (quantisation, pruning) to checkpoint files to reduce storage and transfer overhead. Lossless compression of BF16 weight tensors typically achieves 15-30% size reduction; lossy 8-bit quantisation of the checkpoint (quantising for storage but dequantising for continued training) achieves 50% reduction with <0.1% performance degradation on most benchmarks.

Quantum and Neuromorphic Contexts

  • As quantum-classical hybrid training matures, the checkpoint semantics will need to extend to quantum circuit parameter state (angle registers for variational quantum algorithms) alongside classical neural network weights. Early frameworks including PennyLane and Qiskit Machine Learning are beginning to add checkpoint utilities for variational quantum eigensolver (VQE) and QAOA parameter snapshots, using the same serialisation principles as classical ML checkpoints but adapted to the complex-valued parameter space of quantum circuits.
  • Key differences from classical ML checkpoints in the quantum context:
    • Parameter space is complex-valued (angles in [0, 2π] or arbitrary real numbers for variational gate parameters), requiring appropriate serialisation of complex64/complex128 tensors
    • Circuit topology (gate connectivity, qubit layout) must be co-serialised with parameters, as the same angle vector applied to a different circuit topology produces a different quantum state
    • Noise model provenance becomes critical: a VQE checkpoint trained under one hardware noise calibration may perform poorly when loaded under a different noise environment, requiring noise model snapshots alongside parameter snapshots
    • Quantum circuit checkpoints are typically orders of magnitude smaller than classical ML checkpoints (kilobytes for 10-100 qubit circuits vs gigabytes for LLMs) but the checkpoint frequency requirements are similar (every N training steps of the classical optimiser controlling the variational parameters)

Federated Learning Checkpoints

  • Federated learning introduces additional checkpoint coordination complexity beyond standard distributed training:
    • Federated global model checkpoints: The aggregated global model (produced by FedAvg, FedProx, or similar aggregation rules) is checkpointed on the central server after each communication round; this checkpoint represents the current consensus model
    • Local client checkpoints: Each client device maintains a local checkpoint of its personalised model delta (the difference between the global model and its locally fine-tuned version); these are never shared with the server but must be persisted to enable resumption after client disconnection
    • Differential privacy accounting: Under DP-SGD (differentially private stochastic gradient descent) training, the checkpoint must record the accumulated privacy budget expenditure (epsilon, delta parameters) alongside the weights — a new requirement with no parallel in centralised training
    • Cross-silo vs cross-device federated checkpointing: In cross-silo federated learning (10-100 trusted institution nodes), standard DCP-style parallel checkpoint writes are feasible; in cross-device federated learning (millions of unreliable mobile devices), only the global server checkpoint is durable — client checkpoints are best-effort and may be lost without warning
    • UK implementation: The NHS Federated Analytics Programme (coordinated by NHSX/NHS England) uses federated model checkpoints across NHS trusts for privacy-preserving clinical prediction models, with checkpoint governance managed via the NHS Federated Data Platform (FDP)

Research & Literature

Core References

  • Chandy, K.M. and Lamport, L. (1985). “Distributed Snapshots: Determining Global States of Distributed Systems.” ACM Transactions on Computer Systems 3(1):63-75.
    • Foundational distributed fault-tolerance theory; establishes that consistent global snapshots are achievable without global pause, underpinning all distributed ML checkpoint designs.
  • Polyak, B.T. and Ruppert, D. (1992). “Acceleration of stochastic approximation by averaging.” SIAM Journal on Control and Optimization 30(4):838-855.
    • Theoretical basis for checkpoint/weight averaging; proves time-averaged SGD iterates achieve optimal O(1/√T) convergence rate.
  • Hochreiter, S. and Schmidhuber, J. (1997). “Flat Minima.” Neural Computation 9(1):1-42.
    • Generalisation theory supporting checkpoint averaging; flat minima hypothesis explains why checkpoint averages generalise better than individual training iterates.
  • Frankle, J. and Carlin, M. (2018). “The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.” ICLR 2019.
    • Introduces checkpoint-based iterative magnitude pruning (IMP); establishes winning ticket re-initialisation from early training checkpoints.
  • Izmailov, P. et al. (2018). “Averaging Weights Leads to Wider Optima and Better Generalization.” UAI 2018.
    • Original Stochastic Weight Averaging (SWA) paper; formalises cyclical-LR checkpoint collection and averaging for reaching wider, flatter minima.
  • Wortsman, M. et al. (2022). “Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.” ICML 2022.
    • Demonstrates checkpoint averaging across independently fine-tuned checkpoints; greedy soup achieves ImageNet 90.94% top-1, outperforming individual members.
  • Frankle, J. et al. (2020). “Linear Mode Connectivity and the Lottery Ticket Hypothesis.” ICML 2020.
    • Demonstrates linear mode connectivity between checkpoints from same pre-training origin; theoretical foundation for checkpoint merging effectiveness.
  • Entezari, R. et al. (2022). “The Role of Permutation Invariance in Linear Mode Connectivity of Neural Networks.” ICLR 2022.
    • Shows permutation-aligned models always exhibit linear mode connectivity; provides algorithm for aligning checkpoint parameters before merging to eliminate loss barriers.

Distributed Training and Fault Tolerance

  • Zhao, Y. et al. (2023). “PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.” arXiv:2304.11277.
    • Production FSDP distributed checkpointing with DCP; documents checkpoint I/O performance at 1024+ GPU scale.
  • Lian, X. et al. (2024). “Universal Checkpointing: Efficient and Flexible Checkpointing for Large Scale Distributed Training.” arXiv:2406.18820v2.
    • DeepSpeed UCP: lossless checkpoint transformation across arbitrary parallelism strategy combinations (ZeRO + TP + PP + SP); zero additional save overhead.
  • Zhao, S. et al. (2025). “FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs.” arXiv:2509.03047.
    • Sub-150-second failure recovery via active failure detection and partial re-computation; 4-10× faster than checkpoint-based recovery.
  • Zhang, Y. et al. (2025). “Robust LLM Training Infrastructure at ByteDance.” arXiv:2509.16293.
    • ByteRobust every-step checkpointing with <0.9% overhead; 97% effective training time ratio on 9,600-GPU job.
  • “All is Not Lost: LLM Recovery without Checkpoints.” arXiv:2506.15461. (2025).
    • Checkpoint-free gradient-log recovery; reconstructs model state from gradient update sequence without periodic weight snapshots.

EMA and Weight Averaging

  • Karras, T. et al. (2024). “Analyzing and Improving the Training Dynamics of Diffusion Models.” arXiv:2312.02696.
    • Post-hoc EMA: reconstructs arbitrary EMA profiles post-training from stored short-EMA snapshots via linear mixture; decouples EMA tuning from training compute budget.
  • Huang, J. et al. (2024). “Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits.” arXiv:2411.18704.
    • Systematic study of EMA decay hyperparameter sensitivity; EMA improves label noise robustness, calibration, and transfer learning beyond raw generalisation.

Model Merging

  • Yadav, P. et al. (2023). “TIES-Merging: Resolving Interference When Merging Models.” NeurIPS 2023.
    • TIES algorithm: trim (prune low-magnitude delta weights), elect sign (resolve parameter-wise sign conflicts), disjoint merge (average only non-conflicting parameters); outperforms naive averaging on multi-task benchmarks.
  • Yu, L. et al. (2023). “DARE: Language Model Weights Can Be Merged by Large-Scale Sparsification.” arXiv:2311.03099.
    • Sparsity-based checkpoint merging; randomly zeros 90-99% of delta weights before merging, reducing interference while preserving task-specific capability.
  • Entezari, R. et al. (2022). “The Role of Permutation Invariance in Linear Mode Connectivity of Neural Networks.” ICLR 2022.
    • Linear mode connectivity for checkpoint merging; models with permuted neurons can be re-aligned to enable loss-barrier-free interpolation.

Serialisation Formats and Security

  • Hugging Face. (2022). “Safetensors: A Simple Format for Storing Tensors Safely.” huggingface.co/docs/safetensors.
    • Security-first checkpoint serialisation specification; eliminates pickle arbitrary code execution via declarative JSON-header + flat-binary layout.
  • Trail of Bits. (2023). “Security Audit of Safetensors.” Independent security audit. trailofbits.com.
    • Confirms no format-level vulnerabilities; identifies integer overflow defence requirement in byte-offset bounds-checking, implemented in Rust reference parser.
  • PyTorch Foundation. (2026). “Hugging Face Contributes Safetensors to PyTorch Foundation.” pytorch.org.
    • April 2026 governance transfer; trademark and repo now held by Linux Foundation, format APIs and Hub compatibility unchanged.
  • PyTorch. (2024). “HuggingFace Safetensors Support in PyTorch Distributed Checkpointing.” pytorch.org/blog.
    • Native safetensors integration into PyTorch DCP; enables secure parallel distributed checkpoint writes in safetensors format.
  • PyTorch. (2024). “Performant Distributed checkpointing in Production with IBM.” pytorch.org/blog.
    • Documents IBM’s FSDP-scale checkpoint performance improvements; order-of-magnitude speedup via sharded checkpointing vs rank-0-consolidation.
  • PyTorch. (2025). “Getting Started with Distributed Checkpoint (DCP).” docs.pytorch.org/tutorials.
    • Official DCP API documentation including resharding examples and FSDP2 integration patterns.
  • KerasHub / Google Developers. (2024). “Load Model Weights from Safetensors into KerasHub.” developers.googleblog.com.
    • Cross-framework safetensors loading into JAX/TensorFlow/PyTorch via KerasHub from_preset(); demonstrates format’s multi-framework reach.

Blockchain Checkpoints

  • Bitcoin Wiki. “Checkpoint Lockin.” en.bitcoin.it/wiki/Checkpoint_Lockin.
    • Documents Bitcoin’s historical hardcoded checkpoint practice (discontinued 2014); explains SPV security model and why checkpoints are now considered unnecessary with headers-first sync.
  • Ethereum Foundation. (2022). “Weak Subjectivity.” ethereum.org/developers/docs/consensus-mechanisms/pos/weak-subjectivity.
    • Ethereum beacon chain checkpoint semantics; Casper FFG justification and finalisation via epoch-boundary validator votes.
  • Ethereum Consensus Specs. “Phase0 Weak Subjectivity.” github.com/ethereum/consensus-specs.
    • Formal SSZ specification of the checkpoint type and the weak subjectivity period calculation.

MLOps and Versioning

  • Dubey, A. et al. (2024). “The Llama 3 Herd of Models.” arXiv:2407.21783.
    • Meta’s checkpoint management and model versioning at frontier scale; documents checkpoint frequency policy for Llama 3 pre-training (70B and 405B parameter models).
  • Abdin, M. et al. (2024). “Phi-3 Technical Report.” Microsoft Research. arXiv:2404.14219.
    • Checkpoint management for efficient LLM training at smaller scale; documents how best-checkpoint selection was used to identify optimal stopping point in Phi-3-mini training.
  • MLflow. (2025). “MLflow 3.0: Model Registry for Generative AI.” mlflow.org.
    • Extends checkpoint lineage tracking to generative AI: prompt templates, LoRA adapters, retrieval configs, and evaluation metadata linked to checkpoint artifacts.
  • Atlan. (2025). “AI Model Versioning Best Practices: MLOps Guide for Enterprises.” atlan.com.
    • Enterprise guide to checkpoint versioning, registry promotion policies, and regulatory compliance for production ML pipelines.
  • NVIDIA. (2024). “TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.” arXiv:2410.06511.
    • Documents TorchTitan’s DCP-based checkpoint integration; two-level checkpoint strategy and async save for frontier LLM pre-training.
  • Shoeybi, M. et al. (2019). “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.” arXiv:1909.08053.
    • Original Megatron-LM checkpoint format for tensor-parallel training; establishes multi-shard checkpoint conventions subsequently extended by NeMo and DeepSpeed UCP.
  • Rajbhandari, S. et al. (2020). “ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.” SC20.
    • Introduces ZeRO-1/2/3 parameter partitioning; establishes the distributed checkpoint problem that DCP and UCP subsequently solve.
  • Nakamoto, S. (2008). “Bitcoin: A Peer-to-Peer Electronic Cash System.” bitcoin.org.
    • Original Bitcoin whitepaper; SPV (Simplified Payment Verification) section describes the precursor to blockchain checkpointing; hardcoded checkpoints added subsequently in Bitcoin Core implementation.

Software Ecosystem

PyTorch Checkpoint APIs

  • torch.save / torch.load: The fundamental single-file checkpoint API in PyTorch:
    • torch.save(state_dict, path) — serialises using pickle protocol 4 by default; torch.save(state_dict, path, _use_new_zipfile_serialization=True) enables zip-based format with per-tensor metadata
    • torch.load(path, map_location=device) — deserialises; weights_only=True parameter (default in PyTorch 2.6+) restricts to safe tensor-only deserialisation without arbitrary pickle opcode execution
    • torch.load(path, mmap=True) — enables memory-mapped loading for .pt files where supported, partially recovering safetensors-level load performance for trusted files
  • torch.distributed.checkpoint (DCP): The distributed-first checkpoint API:
    • dcp.save(state_dict, checkpoint_id=checkpoint_dir) — saves sharded state dict to a directory; each rank writes its local shards in parallel
    • dcp.load(state_dict, checkpoint_id=checkpoint_dir) — loads into the current model/optimiser state dict, handling resharding automatically via the global metadata index
    • dcp.async_save(state_dict, checkpoint_id=checkpoint_dir) — non-blocking save that returns a Future; training can continue while the checkpoint is serialised asynchronously in a background thread pool
    • StorageWriter / StorageReader — pluggable storage backends; built-in implementations for local filesystem and fsspec (enabling S3, GCS, HDFS); custom backends can be implemented for specialised storage systems

HuggingFace transformers Checkpoint APIs

  • PreTrainedModel.save_pretrained(save_dir): Saves model weights plus config.json:
    • Automatically shards into files ≤5 GB (configurable via max_shard_size)
    • Writes model.safetensors for single-file models or model-00001-of-00003.safetensors + model.safetensors.index.json for sharded models
    • Falls back to pytorch_model.bin if safetensors is not available (deprecated path)
  • AutoModel.from_pretrained(model_id): Loads from Hub or local path:
    • Resolves model_id to Hub URL, checks for model.safetensors.index.json, downloads required shards
    • Supports revision= parameter for loading specific git commits / tagged checkpoint versions
    • torch_dtype=torch.bfloat16 parameter downcasts weights at load time, halving memory consumption for 32-bit checkpoints
    • device_map="auto" enables device-map-based model parallelism for models too large for a single GPU
  • Trainer.train(resume_from_checkpoint=path): Training resume API:
    • Restores model weights, optimiser state, LR scheduler, and training step counter from a checkpoint directory
    • Skips already-processed batches from the resumed epoch using the DataLoader’s skip_batches utility
    • Requires that the checkpoint directory contain both pytorch_model.bin (or .safetensors) and optimizer.pt / scheduler.pt for full resume

MLflow Model Registry Integration

  • Artifact logging: mlflow.pytorch.log_model(model, "model", registered_model_name="Checkpoints-Demo") saves the PyTorch checkpoint as an MLflow artifact with automatic schema detection
  • Version promotion: client.transition_model_version_stage(name, version, stage) promotes checkpoints through Development → Staging → Production lifecycle stages
  • Lineage tracking: Every logged checkpoint is linked to its parent run (mlflow.start_run() context), which records parameters, metrics, tags, and the git commit hash via mlflow.set_tag("mlflow.source.git.commit", commit_hash)
  • Model cards: MLflow 3.0 supports model_config YAML attached to checkpoint artifacts, encoding intended use, bias evaluation results, and training data metadata in a machine-readable format aligned with Hugging Face Hub model card conventions

Weights and Biases Checkpoint Integration

  • Artifact API: wandb.Artifact("checkpoint", type="model") creates a versioned artifact; artifact.add_file(checkpoint_path) attaches the checkpoint file; run.log_artifact(artifact) registers it with lineage to the current run
  • Checkpoint callbacks: WandbModelCheckpoint(filepath, monitor="val_loss", save_best_only=True) automatically saves Keras/TF checkpoints to W&B artifacts on metric improvement
  • Model registry: W&B’s model registry provides a central catalogue of promoted checkpoints with linked evaluation runs, dataset versions, and deployment records; teams can configure automated evaluation pipelines triggered on artifact version creation via W&B Launch

Metadata

  • domain-correction: infrastructure → artificial-intelligence
  • note: Stub contained Stable Diffusion model links (CivitAI, Reddit, HuggingFace model pages) from a Logseq bookmark dump — these were contaminated content, not ontology assertions. Original stub domain was wrongly infrastructure; corrected to artificial-intelligence. The blockchain checkpoint meaning is addressed under Use Cases as a secondary domain bridge.

Glossary of Key Terms

  • state dict: A Python dictionary mapping parameter/buffer names to tensor values, returned by model.state_dict(); the fundamental unit of a checkpoint in PyTorch
  • .pt / .pth: PyTorch checkpoint file extensions using Python pickle serialisation; .pt typically denotes a full checkpoint dict; .pth often denotes just model weights
  • .safetensors: The Safetensors format file extension; a flat binary with JSON header, no code execution risk, mmap-capable; default format on Hugging Face Hub
  • .ckpt: Ambiguous extension used by both TensorFlow (legacy checkpoint format) and PyTorch pickle files; context-dependent; avoided in new code in favour of .safetensors
  • SavedModel: TensorFlow directory-based checkpoint format containing a serialised computation graph (saved_model.pb) and variable shard files; cross-language safe
  • GGUF: Community format for quantised LLM weights used by llama.cpp; self-describing header including quantisation scheme metadata; primary format for consumer-hardware inference
  • EMA (Exponential Moving Average): Shadow parameter copy updated as θ_ema ← α·θ_ema + (1-α)·θ; inference model in diffusion pipelines; provides smoother generalisation than raw training weights
  • DCP (Distributed Checkpoint): PyTorch’s torch.distributed.checkpoint API for parallel per-rank shard writes; enables topology-change-robust checkpoint loading
  • UCP (Universal Checkpointing): DeepSpeed extension of DCP enabling checkpoint transformation across arbitrary parallelism strategy combinations (ZeRO + TP + PP + SP)
  • SWA (Stochastic Weight Averaging): Training technique using cyclical learning rates to collect checkpoints spanning a flat loss basin, then averaging them for improved generalisation
  • TIES-merging: Checkpoint merging algorithm that resolves sign conflicts between merged parameters by voting before averaging, reducing parameter interference
  • DARE: Checkpoint merging via random sparsification of delta weights (zeroing 90-99%) before merging, reducing cross-checkpoint interference
  • Post-hoc EMA: Karras et al. 2024 technique for reconstructing any EMA profile post-training via linear mixture of stored short-EMA snapshots
  • Weak subjectivity checkpoint: Ethereum Smart Contract Platform beacon chain mechanism providing syncing nodes with a recent finalised (epoch, block_root) anchor for fast-sync without full chain replay
  • Casper FFG: Ethereum Smart Contract Platform’s Friendly Finality Gadget overlay protocol that justifies and finalises checkpoint pairs (epoch, block_root) via stake-weighted validator attestations

Provenance

  • PyTorch Distributed Checkpoint Documentation (docs.pytorch.org, 2024-2025)
  • HuggingFace Safetensors Documentation (huggingface.co/docs/safetensors, 2022-2026)
  • PyTorch Blog: Safetensors + PyTorch Foundation Donation (pytorch.org/blog, April 2026)
  • PyTorch Blog: Performant Distributed Checkpointing with IBM (pytorch.org/blog, 2024)
  • Universal Checkpointing paper (arXiv:2406.18820, Lian et al. 2024)
  • Post-hoc EMA paper (arXiv:2312.02696, Karras et al. 2024)
  • EMA dynamics paper (arXiv:2411.18704, Huang et al. 2024)
  • ByteRobust LLM training infrastructure (arXiv:2509.16293, Zhang et al. 2025)
  • FlashRecovery paper (arXiv:2509.03047, Zhao et al. 2025)
  • Checkpoint-free recovery (arXiv:2506.15461, 2025)
  • Model Soups / checkpoint averaging (Wortsman et al. ICML 2022)
  • TIES-Merging (Yadav et al. NeurIPS 2023)
  • Ethereum Weak Subjectivity documentation (ethereum.org)
  • Bitcoin Wiki Checkpoint Lockin (en.bitcoin.it)
  • MLflow 3.0 announcement (mlflow.org, 2025)
  • Atlan AI Model Versioning Best Practices (atlan.com, 2025)
  • KerasHub safetensors blog post (developers.googleblog.com, 2024)
  • WebSearch research conducted May 2026
  • enrichment-notes: Domain corrected from infrastructure to artificial-intelligence. Blockchain checkpoint meaning retained as secondary bridge. Stub body (CivitAI/Reddit links) was a Logseq bookmarks dump; entirely replaced with ontology-conformant content. Research conducted via WebSearch May 2026.
  • related-uk-infrastructure: Hartree Centre (Daresbury), N8 CIR Bede GPU cluster, JASMIN (STFC), Edinburgh ARCHER2 HPE Cray EX
  • related-uk-policy: NHS AI Lab Model Assurance Framework (2024), UKRI Research Data Management Policy, UK MHRA AIaMD Guidance, UK AI Safety Institute Frontier AI Safety Commitments (2023)