The set of techniques and engineering practices that reduce the energy consumed by training and running artificial intelligence systems while preserving acceptable task performance, spanning model compression, efficient hardware selection, workload scheduling, and data-centre operation.
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:ModelQuantisation))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:ModelPruning))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:KnowledgeDistillation))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:NeuralArchitectureSearch))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:MixedPrecisionTraining))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:GradientCheckpointing))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:FlashAttention))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:WorkloadScheduling))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:SparseMoERouting))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:SpeculativeDecoding))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:SparseAttention))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:hasPart ai:ParameterEfficientFineTuning))
Dependency Relationships
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:requires ai:HardwareAcceleration))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:requires ai:GPUComputing))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:dependsOn ai:DeepLearning))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:dependsOn ai:CarbonAwareComputing))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:dependsOn ai:PowerUsageEffectiveness))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:dependsOn ai:NeuromorphicComputing))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:requires ai:EnergyMeasurementInfrastructure))
Capability Relationships
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:enables ai:EdgeAI))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:enables ai:OnDeviceInference))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:enables ai:EnvironmentalSustainability))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:supports ai:AIGovernance))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:supports ai:RegulatoryCompliance))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:enables ai:EdgeComputing))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:supports ai:AIEthicsChecklist))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:enables ai:CarbonAwareComputing))
Implementation Relationships
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:implements ai:GreenAI))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:implements ai:EfficientTransformers))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:uses ai:FederatedLearning))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:uses ai:TransferLearning))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:implements ai:ModelCompression))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:uses ai:RenewableEnergy))
Reduction Relationships
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:reducesTo ai:ModelCompression))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:reducesTo ai:ParameterEfficientFineTuning))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:reducesTo ai:KnowledgeDistillation))
SubClassOf(ai:AIEnergyOptimisation
ObjectSomeValuesFrom(ai:reducesTo ai:ModelQuantisation))
About
AI Energy Optimisation — referred to interchangeably as Green AI or Efficient AI in the research literature — addresses one of the most pressing systemic challenges in contemporary computing: the rapidly escalating electricity demand associated with training, fine-tuning, and deploying Artificial Intelligence systems at scale. According to the International Energy Agency’s April 2025 “Energy and AI” World Energy Outlook Special Report, AI-focused Data Centre capacity more than tripled in the eighteen months to early 2026, and the largest technology company capital expenditure exceeded USD 400 billion in 2025 alone, expected to jump a further 75% in 2026. The electricity consumed by data centres surged 50% in 2025, making AI infrastructure one of the fastest-growing contributors to global electricity demand. Without structured optimisation at every layer of the AI stack, the aggregate carbon footprint of AI could reach 32.6–79.7 million tonnes of CO₂ annually by 2030 — comparable in scale to the aviation industry’s total annual emissions.
The urgency of AI energy optimisation stems from a structural asymmetry in the AI development landscape: model capability scales roughly with the cube root of compute (Kaplan et al., 2020; Hoffmann et al., 2022), but the energy and carbon consequences of unconstrained scaling are linear or super-linear with parameter count and training token budget. This means that the most direct path to improved AI capabilities — simply training larger models on more data using more powerful compute — is also the most ecologically expensive. The field of AI energy optimisation seeks to decouple these two trajectories, identifying techniques that deliver capability improvements without proportional energy cost increases. In practical terms, the goal is to shift the performance-per-watt frontier: to obtain the same or better task performance at lower absolute energy expenditure, whether during the training phase that produces the model weights, the fine-tuning phase that adapts them to specific tasks, or the inference phase that executes the model millions or billions of times in production.
The practical methods constituting AI Energy Optimisation operate across two broad axes: model-side efficiency and infrastructure-side efficiency. Model-side techniques — including Model Quantisation, Model Pruning, Knowledge Distillation, and Neural Architecture Search — reduce the computational work required per forward pass, directly lowering energy per token generated or per classification made. At the most fundamental level, this involves reducing the bit-width of numerical representations used for weights and activations, since lower-precision arithmetic requires less silicon area and consumes less dynamic power per operation. FP32 (32-bit floating point) remains the standard for training precision but is increasingly being replaced by BF16 and FP16 for forward passes, with INT8 quantisation becoming standard for deployment and INT4 now viable for many tasks through post-training quantisation techniques like GPTQ and AWQ. Infrastructure-side techniques — including Carbon-Aware Computing workload scheduling, geographic placement near low-carbon grid regions, cooling optimisation to improve Power Usage Effectiveness, and co-design with dedicated Hardware Acceleration chips — reduce the energy overhead of running any given compute operation. The two axes compound: a model compressed to half its parameter count running on hardware with half the PUE and powered by 100% renewable energy delivers approximately four times the environmental benefit of addressing either axis alone. In practice, Google’s Ironwood TPU represents this hardware efficiency trajectory — at 30x the energy efficiency per FLOP of their first public TPU — while their fleet-wide Carbon-Aware Computing scheduling achieved a 44x reduction in carbon per prompt over 2024–2025.
The mathematical framework underlying many Model Quantisation approaches exploits the observation that neural network weights, once trained, exhibit heavy-tailed distributions with many near-zero values. Quantisation maps the continuous weight distribution to a discrete grid of representable values, minimising the loss in representational fidelity (measured by perplexity increase for language models or accuracy drop for classifiers) for a given reduction in bits per weight. The key challenge is that standard rounding-based quantisation introduces errors that compound across layers; state-of-the-art methods (GPTQ, AWQ, SmoothQuant) use Hessian-based compensation, activation-aware scaling, or second-order optimisation to correct for this compounding. Research on StarCoder2 published in 2024 (arXiv:2411.12758) demonstrates that INT8 quantisation achieves median power reduction of 39% versus FP16 with negligible accuracy loss, while the GGUF Q4_K_M format — used by default in Ollama and llama.cpp — reduces a 14.39 GB FP16 model to under 1.5 GB whilst achieving 47.9 tokens/second on CPU inference. The “era of 1-bit LLMs” (Ma et al., 2024 — BitNet b1.58) pushes this further still, demonstrating that ternary weight representations ({-1, 0, +1}) can match standard precision models on language modelling benchmarks when trained from scratch with appropriate scaling, with potentially transformative implications for inference energy requirements.
Model Pruning applies a complementary approach: rather than reducing the precision of each weight, it removes weights entirely. Unstructured pruning identifies and zeros individual weights based on magnitude or gradient-based saliency criteria, achieving high theoretical compression ratios but requiring sparse computation support in hardware to realise energy savings. Structured pruning — removing entire attention heads, convolutional filters, or Transformer layers — produces dense models that are immediately compatible with standard GPU Computing hardware without requiring sparse matrix support. The Lottery Ticket Hypothesis (Frankle & Carlin, 2019) demonstrated that dense networks contain sparse sub-networks (“winning tickets”) that can be trained in isolation to full accuracy, providing theoretical grounding for aggressive pruning. In practice, a three-stage pipeline combining magnitude pruning, INT8 quantisation, and Huffman encoding achieves 35x–49x storage reduction without measurable accuracy degradation, as demonstrated in robotics deployments achieving 75% model size reduction and 50% power reduction while maintaining 97% task accuracy.
Knowledge Distillation offers a third compression pathway that is conceptually distinct from both quantisation and pruning: rather than modifying the representation or topology of a trained model, it trains an entirely new, smaller model (the “student”) to mimic the output distribution of a larger model (the “teacher”). The key insight, established by Hinton et al. (2015), is that the teacher’s soft probability distribution over output classes contains vastly more information than the hard one-hot ground-truth label: near-zero probabilities reflect the model’s implicit representations of class similarity, enabling the student to learn semantic structure from the teacher’s “dark knowledge” rather than from raw labels alone. DistilBERT (Sanh et al., 2019) demonstrated this at scale, achieving 97% of BERT’s GLUE benchmark performance with 40% fewer parameters and 60% faster inference — a result that has made distillation the dominant pathway for deploying language model capabilities on resource-constrained hardware.
The discipline intersects with AI Governance and Regulatory Compliance as emerging regulations increasingly mandate energy reporting as a component of responsible AI deployment. The EU AI Act’s Article 51, applicable from August 2025 for frontier models designated as having systemic risk, requires providers to measure and report their model’s energy consumption during training and deployment, maintain an energy efficiency policy, and implement measures to reduce energy footprint. The UK government’s AI Safety Institute and DSIT AI strategy both reference sustainability as a cross-cutting governance dimension, and the ISO/IEC JTC 1/SC 42 working group on AI sustainability is developing standardised measurement methodologies that will form part of the next revision of IEC 42001. The UK’s AI Energy Disclosure guidance, published by DSIT in early 2026, specifies reporting formats for organisations deploying AI in public services. Measuring and improving energy efficiency is therefore no longer purely a cost-optimisation concern but a legal and reputational imperative for organisations deploying Large Language Model systems at scale — and an integral module in any rigorous AI Ethics Checklist for high-impact AI systems.
Components / Architecture
Model Compression Methods:
-
Model Quantisation — reducing numerical precision of weights and activations from FP32 or FP16 to INT8, INT4, or lower. The three principal approaches are: (a) post-training quantisation (PTQ), applied after training without requiring modified training procedures, with methods like GPTQ (Frantar et al., 2022) using second-order weight correction to compensate for quantisation error; (b) quantisation-aware training (QAT), which simulates quantisation effects during training by inserting “fake quantisation” nodes into the computational graph, achieving higher accuracy than PTQ at the cost of additional training compute; and (c) dynamic quantisation, which quantises weights statically but activations dynamically at inference time, requiring no calibration data. Research on StarCoder2 (arXiv:2411.12758) demonstrates that INT8 quantisation reduces median power by 39% versus FP16 baseline. The GGUF format with Q4_K_M quantisation reduces a 14.39 GB FP16 model to under 1.5 GB with performance competitive with INT8 quality, achieving 47.9 tokens/second on CPU. A three-stage pipeline combining magnitude pruning, INT8 quantisation, and Huffman encoding achieves 35x–49x storage reduction without measurable accuracy degradation. The AWQ (Activation-aware Weight Quantisation) method (Lin et al., 2023) achieves near-INT4 quality by identifying and protecting the small fraction (1%) of salient weights that account for disproportionate output variance.
-
Model Pruning — removing weights, neurons, heads, or entire layers based on saliency criteria. Unstructured (element-wise) pruning achieves the highest compression ratios but requires sparse matrix computation support for hardware-level energy savings; structured pruning (removing entire convolutional filters, attention heads, or MLP blocks) produces dense, immediately deployable compressed models. The Lottery Ticket Hypothesis (Frankle & Carlin, 2019) demonstrated that dense networks contain sparse sub-networks with identical initialisation that can be trained to full accuracy in isolation, providing theoretical grounding for post-hoc pruning approaches. Movement pruning (Sanh et al., 2020), which prunes based on weight movement magnitude during fine-tuning, achieves better compression-accuracy trade-offs than magnitude pruning for fine-tuned language models. Practical deployments combining structured pruning (removing 20–40% of attention heads and 30–50% of MLP neurons) with INT4 quantisation have achieved 75% model size reduction and 50% power reduction while maintaining 97% task accuracy in robotics and automation applications.
-
Knowledge Distillation — training a compact student model on the soft-target probability distribution (logit outputs) of a larger teacher model, transferring the teacher’s learned representations without requiring access to the teacher’s internal activations. The student loss is a linear combination of cross-entropy against hard ground-truth labels and KL divergence against teacher soft probabilities, with a temperature hyperparameter T controlling the softness of the teacher distribution (higher T spreads probability mass across more classes, revealing more of the teacher’s uncertainty structure). DistilBERT (Sanh et al., 2019) achieves 97% of BERT’s GLUE benchmark performance with 40% fewer parameters and 60% faster inference — the canonical demonstration of distillation efficiency. Task-specific distillation typically outperforms general distillation, and layer-to-layer distillation (aligning intermediate representations in addition to final outputs) further improves student quality. Transfer Learning and distillation are complementary: a student initialised from a relevant pre-trained checkpoint trains faster and achieves higher distillation fidelity than a randomly initialised student.
-
Neural Architecture Search — automated search over architecture design spaces to discover computationally frugal topologies. Differentiable NAS (DARTS) allows gradient-based search over continuous relaxations of discrete architecture choices, making the search itself more efficient. Hardware-aware NAS (ProxylessNAS, Once-for-All) directly incorporates hardware efficiency metrics (latency, energy, memory footprint) into the search objective alongside accuracy, enabling architectures that are simultaneously accurate and efficient on specified target hardware platforms. EfficientNet-B0 (Tan & Le, 2019, found by NAS with accuracy/efficiency compound objective) achieves 77.1% ImageNet top-1 at 5.3M parameters, versus ResNet-50’s 76.0% at 25M parameters — demonstrating NAS’s ability to identify Pareto-optimal architectures. The CE-NAS framework (NeurIPS 2024) applies reinforcement learning to adjust GPU resource allocation during the search process based on real-time grid carbon intensity, representing the first integration of operational carbon data into the NAS search procedure itself.
Training-Time Optimisations:
-
Mixed Precision Training — using FP16 or BF16 for activations and gradients in forward and backward passes, with FP32 master weights maintained for gradient accumulation and parameter updates. The BF16 format (Brain Float 16, developed by Google) is preferred over FP16 for training on Ampere and later NVIDIA GPUs and all TPU generations, because BF16’s wider dynamic range (8 exponent bits versus 5 for FP16) prevents gradient underflow without requiring loss scaling. BF16 training delivers 2–3x throughput improvement on modern GPUs via Tensor Core utilisation with negligible accuracy impact versus FP32. NVIDIA’s Transformer Engine (H100 GPUs) extends this to FP8 training (8-bit floating point with 5-bit mantissa), demonstrated to achieve FP16-equivalent training quality at 2x further speedup, representing the current frontier in mixed-precision training efficiency.
-
Gradient Checkpointing — trading additional compute for memory reduction by recomputing selected intermediate activations during the backward pass rather than storing them from the forward pass. Standard backpropagation stores all activations produced in the forward pass to compute gradients efficiently, resulting in memory usage that scales with model depth and sequence length. Gradient checkpointing selects a subset of “checkpoint” activations to store and recomputes the remainder from the nearest checkpoint during backward pass, reducing peak activation memory by up to 10x at the cost of approximately 33% more total backward compute. This is particularly valuable for very long sequence lengths (document processing, protein structure prediction) where activation memory would otherwise exceed available GPU VRAM, enabling larger effective batch sizes and better GPU utilisation on a given hardware budget.
-
Parameter-Efficient Fine-Tuning (PEFT) — a family of methods including LoRA (Low-Rank Adaptation, Hu et al., 2021), QLoRA (quantised LoRA, Dettmers et al., 2023), prefix tuning, prompt tuning, and adapter layers, all of which update fewer than 1% of model parameters during task-specific fine-tuning. LoRA represents weight update matrices as low-rank decompositions (W + ΔW = W + AB where A ∈ ℝ^{d×r} and B ∈ ℝ^{r×k} with r << min(d,k)), reducing the number of trainable parameters by factors of 1,000–10,000 for typical language model fine-tuning. QLoRA combines 4-bit quantisation of the base model with BF16 LoRA adapters, enabling fine-tuning of 65B models on a single 48GB A100 GPU — an accessibility improvement that has made fine-tuning practically available to researchers and organisations without large-scale GPU clusters, while also reducing fine-tuning energy costs by 50–90% versus full fine-tuning.
Inference-Time Optimisations:
-
Flash Attention — IO-aware attention computation using a tiling strategy that keeps the computation within GPU on-chip SRAM (shared memory), avoiding repeated read-write cycles to HBM (High Bandwidth Memory). Standard attention materialises the full n×n attention matrix in HBM; Flash Attention computes attention in tiles that fit in SRAM, reading and writing each position in Q, K, V once rather than O(n) times for the softmax normalisation. This reduces HBM memory bandwidth consumption by 2–4x, enabling 2–4x wall-clock speedup and enabling context lengths up to 256K tokens without linear memory cost growth. Flash Attention 2 (Dao, 2023) improved GPU occupancy through better parallelisation of the sequence dimension; Flash Attention 3 (Shah et al., 2024) further exploits H100 warp specialisation and overlapping of compute and data movement.
-
Speculative Decoding — a throughput technique that uses a small, fast “draft” model to generate a sequence of candidate tokens that the large “target” model then verifies in a single forward pass (accepting some prefix and rejecting the remainder). The target model’s parallel verification of k candidate tokens replaces k sequential autoregressive steps, achieving 2–3x decoding throughput while guaranteeing the same output distribution as standard sampling from the target model alone. The draft model need be only 5–10x smaller than the target to achieve significant throughput gains; effective draft models can be purpose-trained or derived from the target by early-exit or layer-skipping.
-
Sparse Attention — approaches that restrict the full O(n²) attention computation to a subset of token pairs. Locality-based approaches (sliding window, dilated attention) attend to local neighbourhoods; content-based approaches (BigBird, Longformer, Reformer) add structured global tokens or use LSH-based approximate nearest-neighbour attention; hybrid approaches combine dense local and sparse global attention. Sparse attention reduces FLOPs for long sequences from O(n²) to O(n log n) or O(n), enabling practical processing of document-length inputs on fixed hardware budgets.
Infrastructure and Systems:
-
Carbon-Aware Computing — workload scheduling approaches that temporally and geographically optimise job placement to minimise grid carbon intensity at execution time. Real-time carbon intensity data sources include the National Grid ESO Carbon Intensity API (GB grid, 5-minute resolution, 48-hour forecasts), Electricity Maps (global, 30-minute resolution), and WattTime (US grid regions). Google’s fleet-wide Carbon-Aware Computing scheduling achieves a 44x reduction in total carbon footprint per prompt via a combination of renewable energy procurement, hardware efficiency improvements, and temporal scheduling. Microsoft Azure’s carbon-aware scheduling feature routes batch training jobs to regions with lowest marginal carbon intensity. The Greenhouse Gas Protocol’s Scope 2 and Scope 3 accounting standards apply to AI workload carbon, distinguishing market-based (using PPAs and RECs) from location-based (using grid average carbon intensity) carbon accounting.
-
Power Usage Effectiveness (PUE) — the standard metric for data centre energy overhead, defined as total facility energy divided by IT equipment energy. PUE of 1.0 represents perfect efficiency (100% of facility energy used directly by IT equipment); typical legacy data centres operate at PUE 1.5–2.0; hyperscale cloud providers achieve 1.1–1.2 through advanced cooling (cold aisle/hot aisle containment, liquid cooling, outside air economisation) and power delivery optimisation. Google’s fleet-wide average PUE of 1.09 represents near-theoretical efficiency; Microsoft’s new data centres in the Nordic region use outdoor air cooling to achieve PUE below 1.1. Water Usage Effectiveness (WUE) is an emerging complementary metric measuring water consumed for cooling per unit of IT energy, relevant to data centres in water-stressed regions.
-
Workload Scheduling — intelligent batching, request routing, and model tiering strategies that maximise GPU utilisation (reducing idle power waste) and match query complexity to model capability. Continuous batching (PagedAttention, implemented in vLLM) dramatically improves GPU utilisation for LLM serving by dynamically adding new requests to in-flight batches, increasing throughput by 2–4x versus static batching. Model tiering routes simple queries to smaller, faster models and complex queries to larger models, achieving average inference energy savings proportional to the fraction of queries handleable by the smaller tier.
-
Federated Learning — distributing training across edge devices or institutional data holders to avoid centralised data centre aggregation of sensitive data. While Federated Learning removes the need for data transfer to a central server (reducing network energy and enabling privacy-preserving training), it introduces communication overhead (frequent gradient or model parameter exchange between edge nodes and the federated server) and typically requires more training rounds to converge than centralised training due to statistical heterogeneity (non-IID data distribution across participating nodes). Compression of federated gradients (quantisation, sparsification, error feedback) is essential to make Federated Learning energy-competitive with centralised training for large models.
Use Cases / Major Families
Cloud-Scale LLM Serving: The most economically consequential application domain for AI energy optimisation is large-scale inference serving of frontier Large Language Model systems. At billion-query-per-day scale, the per-token energy cost multiplies into meaningful absolute figures: benchmarking by arXiv:2505.09598 (“How Hungry is AI?”) quantified per-query energy consumption across major model families and query types. Choosing between model variants — e.g., selecting a 7B distilled model over a 70B foundation model for appropriate tasks — can reduce energy per query by up to 70x; choosing optimal geographic hosting based on grid carbon intensity can reduce carbon footprint by 50x; implementing optimisation techniques (quantisation, batching, caching) can cut operational costs by 60–80%. The practical implication for organisations deploying AI through API services is that the carbon footprint of their AI usage is substantially determined by vendor efficiency choices that are not typically surfaced in standard cost-per-token pricing metrics. This creates an emerging market for carbon-aware LLM API routing, where downstream applications route queries to the most efficient provider for a given carbon intensity target.
Edge Deployment and On-Device Inference: On-Device Inference on mobile phones, IoT sensors, wearable devices, and automotive controllers requires aggressive model compression to operate within device power envelopes — typically 1–10 watts for smartphones, 100–500 milliwatts for IoT nodes, and 5–30 watts for automotive ADAS processors. The MobileNet family (Howard et al., 2017; Sandler et al., 2018) pioneered depth-wise separable convolutions to reduce FLOPs by approximately 8–9x versus standard convolutions; EfficientNet-B0 (Tan & Le, 2019) applied compound scaling to achieve 77.1% ImageNet top-1 accuracy (comparable to ResNet-50 at 76.0%) with only 5.3M parameters versus 25M. GGUF-formatted LLMs running via llama.cpp on consumer CPU hardware achieve 47.9 tokens/second for Q4_K_M 7B models, making locally-hosted assistant AI practical on a laptop with no GPU and a sub-30W power draw — a dramatic departure from the cloud-GPU paradigm and a significant advance for privacy-preserving and offline-capable AI deployment.
Training Run Optimisation: Research laboratories and frontier AI developers applying compute-optimal scaling laws (Hoffmann et al., 2022 — Chinchilla) have demonstrated that the industry was substantially over-computing (training very large models for too few tokens) in the era of GPT-3-scale work. Chinchilla established that for a compute budget C, the optimal parameter count N and training token count D satisfy N ∝ D ∝ C^0.5, meaning that a 70B model trained on 1.4T tokens (as with LLaMA-2) is compute-optimal in a way that a 175B model trained on 300B tokens is not. Applying these scaling laws directly translates to energy efficiency: given a fixed energy budget, compute-optimal allocation produces better-performing models than naive parameter scaling. Mixed Precision Training (FP16/BF16 for activations, FP32 master weights) typically delivers 2–3x throughput improvement with no accuracy impact. Gradient Checkpointing reduces activation memory by up to 10x at the cost of approximately 33% more backward-pass compute — a frequently optimal trade-off when GPU memory is the binding constraint. Parameter-Efficient Fine-Tuning via LoRA reduces fine-tuning energy by 50–90% versus full fine-tuning while reaching 90–97% of full fine-tuning task performance on standard benchmarks.
Healthcare and Industrial IoT: Remote diagnostic devices — including wearable ECG monitors, point-of-care diagnostic tablets, and surgical robotics systems — deploy compressed AI models under strict power budgets dictated by battery life requirements and medical device certification constraints. The UK’s NHS AI Lab has funded edge AI research for rural primary care settings where connectivity is intermittent; compressed diagnostic models operating entirely on-device preserve patient privacy and function without network access. Smart factory sensors and predictive maintenance systems in manufacturing facilities operated by BAE Systems, Rolls-Royce, and similar Northern England industrial employers deploy quantised anomaly detection models on microcontrollers with milliwatt budgets. Knowledge Distillation from cloud-hosted domain specialist models to on-device student models is the dominant deployment pathway: a cloud model trained on comprehensive clinical or manufacturing data teaches a student that can run on an ARM Cortex-M4 microcontroller.
Scientific Computing and National Facilities: Particle physics experiments at CERN’s Large Hadron Collider, and astrophysics data pipelines at the Square Kilometre Array (SKA, with UK construction involvement), deploy quantised AI models for real-time event classification at data rates measured in petabytes per second. At these data rates, inference must be performed on FPGAs or ASICs with power envelopes in the tens of watts; post-training quantisation to INT4 or binary representations is often required to meet latency and power specifications. The UK’s STFC (Science and Technology Facilities Council) operates particle physics computing facilities (GridPP) where energy efficiency in ML inference is a first-class operational constraint; the Hartree Centre provides AI energy optimisation consultancy to UK research facilities.
Financial Services and Regulatory Compliance: UK financial institutions subject to the FCA’s AI regulatory guidelines are increasingly required to document the energy and carbon footprint of algorithmic trading and credit scoring systems as part of ESG reporting obligations. Goldman Sachs, JP Morgan, and Barclays have published AI energy metrics in their sustainability reports since 2024. The combination of inference volume and regulatory scrutiny makes financial AI one of the highest-priority domains for systematic energy optimisation via quantisation, model tiering (routing low-complexity queries to smaller models), and Carbon-Aware Computing batch scheduling during overnight settlement runs.
Academic Context
The formal academic study of AI’s energy impact has a surprisingly brief history relative to the age of Deep Learning research. For most of the deep learning era (2012–2019), computational efficiency was considered a secondary concern; researchers optimised for accuracy, and hardware improvements were assumed to absorb any efficiency shortfall via Moore’s Law. This assumption was challenged by Strubell et al. (2019) “Energy and Policy Considerations for Deep Learning in NLP” (ACL 2019), which first quantified the CO₂ emissions of training NLP models at scale. Their most striking finding was that training a single Transformer with full Neural Architecture Search hyperparameter optimisation cost approximately 626,155 lbs of CO₂ equivalent — roughly five times the lifetime carbon cost of an average American car including manufacture. This result was widely reproduced and prompted immediate discussion in both research and policy communities about whether the prevailing compute-maximising culture of AI research was environmentally sustainable.
Schwartz et al. (2020), writing in Communications of the ACM, codified the response by coining the term “Green AI” and proposing efficiency (measured as accuracy per unit compute) as a first-class research criterion alongside raw accuracy. Their “Red AI” / “Green AI” distinction — between research that prioritises performance at any compute cost and research that prioritises performance per unit resource — has since become a standard framing in the field. The practical implication is that benchmark leaderboards should report not only accuracy but also training compute cost, inference FLOPs, and energy consumption, enabling researchers and practitioners to make informed accuracy-efficiency trade-offs.
The energy-efficiency implications of scaling laws were sharpened by Hoffmann et al. (2022), the Chinchilla paper from DeepMind, which demonstrated that the largest models of the time (GPT-3, Gopher, Megatron) were substantially undertrained relative to their parameter counts. The compute-optimal scaling law N ∝ D ∝ C^0.5 implies that for a given compute budget, the optimal strategy is to train a smaller model on more data — a result with direct energy efficiency implications: Chinchilla (70B parameters, 1.4T tokens) outperforms Gopher (280B parameters, 300B tokens) at 4x lower parameter count, meaning that careful scaling law application delivers better capability per training FLOP. This insight catalysed a wave of smaller-but-better models (LLaMA, Mistral, Phi) that dominate practical deployments in 2025–2026.
Operator-level efficiency research has produced some of the most impactful practical contributions. Dao et al. (2022) introduced Flash Attention, reformulating self-attention computation to exploit GPU memory hierarchy properties by tiling the computation to fit in on-chip SRAM rather than materialising the full O(n²) attention matrix in HBM (high-bandwidth memory). This IO-aware reformulation reduces memory bandwidth consumption by 2–4x and achieves equivalent computation with no approximation — a rare result of hardware-aware algorithmic design delivering exact speedup. Flash Attention 2 (Dao, 2023) and Flash Attention 3 (Shah et al., 2024) extended the approach to multi-head attention with improved GPU utilisation, collectively making long-context inference viable at reasonable energy cost. Frantar et al. (2022) published GPTQ, demonstrating that post-training quantisation of 175B models to INT4 is achievable with perplexity increases of less than 0.5 perplexity points versus the FP16 baseline through second-order weight correction, enabling practical deployment of frontier models on consumer-grade hardware for the first time.
The measurement infrastructure for AI energy research has matured substantially. The Green Algorithms project (Lannelongue, Grealey & Inouye, 2021, University of Cambridge MRC Biostatistics Unit) provides a web-based tool and Python library for estimating the carbon footprint of computational workloads based on hardware type, runtime, location, and grid carbon intensity, and has been adopted as a reporting requirement by Nature journals for computationally intensive submissions. The CodeCarbon Python package (Courty et al., 2023) enables inline energy measurement during training and inference loops. MLCommons’ MLPerf Inference benchmark tracks accuracy-per-watt across hardware platforms, establishing a standardised basis for comparing the energy efficiency of different hardware and software stack combinations.
Key research groups currently active in AI energy optimisation include: Stanford HAI (Human-Centred AI) energy benchmarking and policy studies; MIT CSAIL efficient inference and hardware co-design; Google Brain’s efficient neural network team (producers of EfficientNet, MobileNet, and PaLM efficiency analyses); Microsoft Research’s DeepSpeed team (ZeRO optimiser, 1-bit Adam); Hugging Face’s model efficiency team (PEFT, AutoGPTQ, optimum library); the University of Edinburgh’s School of Informatics (neuromorphic computing and quantum materials for efficient computing); the University of Cambridge’s MRC Biostatistics Unit (Green Algorithms); Imperial College London’s AI group (attention mechanism optimisation); and the Alan Turing Institute’s AI for Science programme (compute-efficient scientific ML).
Current Landscape (2026)
As of mid-2026, AI energy optimisation has transitioned from a niche research concern to a mainstream operational and regulatory priority. The scale of the challenge is illustrated by the IEA’s 2025 central scenario, which projects data-centre electricity consumption reaching 945 TWh by 2030, representing approximately 3% of global electricity demand — comparable to the total electricity consumption of France. AI’s share of data-centre power is estimated at 5–15% currently, potentially reaching 35–50% by 2030 as foundation model deployment continues to scale. The capital expenditure trajectory is similarly striking: the largest technology companies collectively spent more than USD 400 billion on AI infrastructure in 2025 and are projected to spend approximately USD 700 billion in 2026, with a substantial fraction of this directed at AI-capable Data Centre construction and Hardware Acceleration procurement.
In response to these scaling pressures, the leading cloud providers have made substantial efficiency commitments that are now partially measurable against published metrics. Google reports that its Ironwood TPU generation is 30x more energy-efficient per FLOP than the first publicly available TPU (TPU v1, 2016), and that its fleet-wide median energy per prompt fell by 33x over 2024–2025 — a figure attributed to improved model efficiency (Gemini Nano and distilled variants replacing direct use of larger models), hardware improvements (TPU v5 and Ironwood deployment), and fleet-wide Carbon-Aware Computing scheduling (PUE 1.09 fleet average, versus the industry average of approximately 1.5). Google’s energy per query per median prompt fell by 33x and total carbon footprint fell by 44x over the same period, as renewable energy procurement and efficiency improvements compounded. Microsoft’s commitments include adding 10.5 GW of new renewable energy to the grid (enough to power approximately 8 million homes), alongside AI-assisted materials research to accelerate carbon-free energy technologies. Meta has published inference energy metrics for its Llama model series, enabling community-driven optimisation research.
The open-source quantisation ecosystem has matured rapidly in 2025–2026. The GGUF format and Q4_K_M quantisation scheme have become a practical standard for portable model deployment, with Ollama defaulting to Q4_K_M for local model serving. The EfficientLLM survey (arXiv:2505.13840, 2025) systematically benchmarks efficiency techniques across 40+ methods, providing the most comprehensive overview of the current state of the art. Hybrid pipelines combining structured Model Pruning with INT4 Model Quantisation achieve 75–93.75% model size reduction across edge robotics, automotive, and mobile applications, with demonstrated deployments maintaining 97%+ task accuracy. The CE-NAS framework (NeurIPS 2024 workshop) demonstrates automated carbon-aware Neural Architecture Search using reinforcement learning to adjust GPU resource allocation against live grid carbon intensity signals, representing the first system to close the loop between model design and real-time grid carbon data.
Regulatory momentum is accelerating on multiple fronts. The EU AI Act’s Article 51 systemic-risk provisions, applied from August 2025, require providers of general-purpose AI models that are designated as having systemic risk (a designation based on training compute exceeding 10^25 FLOPs) to measure and report their model’s training and inference energy footprint, implement an energy efficiency policy, and make this information available to downstream providers and the EU AI Office. The UK’s DSIT published AI energy disclosure guidance in early 2026, specifying reporting formats for organisations deploying AI in public services under the UK government’s AI Procurement Framework. The ISO/IEC JTC 1/SC 42 committee is developing a supplementary standard to IEC 42001 specifically addressing AI sustainability metrics, including standardised energy measurement methodologies for training and inference, scope 2 and scope 3 carbon attribution, and water usage effectiveness (WUE) for water-cooled facilities. The US NIST AI RMF GenAI Profile (March 2024) includes energy and environmental impact as a risk category, and the EU’s science advisory body (SAPEA) published a report in 2026 calling for mandatory AI energy labelling analogous to EU appliance energy efficiency classes. Together, these regulatory developments are transforming energy efficiency from an internal cost optimisation into a reportable metric, public disclosure obligation, and potential basis for regulatory intervention.
UK Context
The United Kingdom occupies a distinctive position in the global AI energy optimisation landscape: it hosts world-class academic research in both Deep Learning efficiency and energy systems, operates significant national computing infrastructure with known environmental footprints, and faces specific industrial deployment contexts — particularly in manufacturing, healthcare, and energy — where edge AI efficiency is operationally critical.
At the academic level, the University of Edinburgh’s School of Informatics runs a dedicated PhD programme on AI-Driven Discovery of Quantum Materials for Neuromorphic and Energy-Efficient Computing, funded through EPSRC and the Royal Academy of Engineering. This programme, hosted in the Institute for Condensed Matter and Complex Systems, investigates materials-level approaches to energy-efficient computation — exploring properties of novel oxide compounds and 2D materials (graphene derivatives, MoS₂) with neuromorphic potential, using AI to accelerate materials discovery and characterisation. The programme accesses ARCHER2 (the UK’s national Tier-1 supercomputing service at Edinburgh, operated by EPCC) and CIRRUS (Tier-2 HPC at Edinburgh). ARCHER2 operates at approximately 52 MW peak power consumption and has implemented carbon intensity-aware scheduling in partnership with National Grid ESO to shift discretionary workloads to low-carbon periods, achieving a 15–20% reduction in scope 2 emissions relative to flat scheduling.
The University of Cambridge’s MRC Biostatistics Unit produced the Green Algorithms project (Lannelongue, Grealey & Inouye, 2021), which provides the standard carbon-footprint estimation tool for computational science workloads. Nature family journals adopted Green Algorithms reporting as a submission requirement for computationally intensive manuscripts in 2023; IEEE Transactions on Neural Networks and Learning Systems and ICML 2024 followed with similar requirements. Cambridge’s Department of Engineering also contributes to Efficient Transformers research, with publications on linearised attention and sparse attention mechanisms with direct energy efficiency implications.
Imperial College London’s AI research groups, including the Intelligent Systems and Networks group and the Data Science Institute, have contributed to Parameter-Efficient Fine-Tuning methods and attention mechanism optimisation. UCL’s AI Centre has published on energy-efficient generative model architectures and has been active in developing UK-context ESG reporting frameworks for AI systems. The Alan Turing Institute (ATI), operating from the British Library in London with affiliated institutes at universities across the UK, runs an AI Energy and Sustainability working group coordinating research across member institutions. The ATI’s Data-Centric Engineering programme, in partnership with Lloyd’s Register Foundation, addresses efficient AI for structural health monitoring in infrastructure — a context where sensor node power constraints directly dictate compression requirements.
Northern England represents both a significant research base and a set of distinctive industrial deployment contexts. The Hartree Centre (Daresbury, Cheshire), operated by STFC and embedded within the UK’s Sci-Tech Daresbury campus, is the UK’s national centre for AI and high-performance computing applied to industrial challenges. The Hartree Centre’s Baskerville supercomputer (University of Birmingham, 2021) and Lynx GPU cluster provide GPU resource specifically for UK industry AI R&D; the Centre provides AI energy optimisation consultancy to UK SMEs and large enterprises in chemicals, pharmaceuticals, aerospace, and energy. AstraZeneca’s UK computational chemistry operations (Macclesfield) and BAE Systems’ AI research programmes (Warton, Lancashire; Rochester, Kent) are among the industrial users engaging with efficient AI deployment for high-stakes applications.
Greater Manchester’s advanced manufacturing ecosystem — including the National Graphene Institute (NGI) at University of Manchester and the Henry Royce Institute for advanced materials (pan-UK, headquartered at Manchester) — connects directly to AI energy optimisation through materials science applications: graphene-based memristors and 2D material neuromorphic devices under investigation at NGI are candidate substrates for ultra-low-power AI inference. The Manchester-based National Robotics Innovation Centre and Robotics and AI Hub (RAIN) deploy compressed AI models on autonomous systems under power-constrained operational environments.
Leeds hosts the STFC-funded ATLAS computing grid node at the University of Leeds, which participates in particle physics data processing alongside CERN; energy-efficient AI inference for LHC data is an active optimisation challenge here. Sheffield’s Advanced Manufacturing Research Centre (AMRC) at Catcliffe deploys On-Device Inference for defect detection and predictive maintenance in aerospace manufacturing contexts where line-side hardware runs on shared factory power circuits with strict maximum draw requirements. Newcastle’s NICD (National Innovation Centre for Data) works with Northern SMEs on practical efficient ML deployment, frequently encountering edge hardware with sub-10W thermal envelopes and ARM Cortex-M class processors as the target deployment platform.
The UK’s National Grid ESO publishes a real-time and forecast Carbon Intensity API (carbonintensity.org.uk), which has become an infrastructure component for Carbon-Aware Computing implementations within the UK. Several UK cloud-hosted AI services (including NHS AI Lab workloads and Alan Turing Institute research compute) now use this API to implement time-shifting of training workloads to low-carbon periods — typically between 2am–6am when wind generation is high and demand is low across the GB grid.
Future Directions (2026–2030)
Neuromorphic Hardware at Commercial Scale: Intel’s Loihi 2 neuromorphic processor and Heidelberg University’s BrainScaleS-2 system demonstrate energy improvements of 1,000–10,000x per inference versus CMOS GPU architectures for spike-coded workloads, exploiting event-driven computation that consumes power only when neurons fire. IBM’s NorthPole chip (2023) demonstrated 25x better energy efficiency than contemporary GPUs on ResNet-50 inference by eliminating off-chip memory access through a fully on-chip memory architecture. Commercial neuromorphic deployment currently requires substantial software ecosystem development; the University of Edinburgh and Manchester contribute to the EU Human Brain Project’s neuromorphic software frameworks (PyNN, SpiNNaker 2 SDK). The trajectory suggests that neuromorphic accelerators may become competitive for specific inference workloads (speech recognition, sensor fusion, anomaly detection) within the 2026–2028 timeframe, particularly as SpiNNaker 2 (led by TU Dresden with Manchester involvement) targets commercial deployment.
Mixture-of-Experts (MoE) at Scale and Routing Efficiency: Sparse MoE architectures — in which a gating network routes each input token to a subset of “expert” sub-networks (typically 2 out of 8, or 2 out of 64) — decouple model capacity from inference cost. Mistral AI’s Mixtral 8x7B model activates only 13B parameters per forward pass despite having 47B total parameters, achieving performance comparable to 70B dense models at 70B inference cost. Google’s Switch Transformer demonstrated linear scaling of model capacity with constant FLOPs per token. Active research frontiers include improving routing efficiency (minimising expert load imbalance that wastes allocated compute), extending MoE to Federated Learning settings, and enabling dynamic expert sets that grow as domain knowledge is added without retraining the routing network. The energy implications are significant: routing 95% of tokens to 2 of 64 experts means that 94% of parameters are never accessed for any given forward pass, enabling very large model capacity with inference energy comparable to a much smaller dense model.
State-Space Models and Architectural Alternatives to Transformers: Mamba (Gu & Dao, 2023) and related structured state-space models (S4, H3) offer O(n) sequence processing versus Transformer’s O(n²) attention, which has direct energy implications for long-context tasks: at sequence length 8,192 tokens, a Mamba-class model processes the sequence in approximately 1/64th the FLOPs of a standard attention Transformer. Hybrid architectures interleaving SSM and attention layers (Jamba, Zamba) are being explored to capture the complementary strengths of both: attention’s associative retrieval capability and SSMs’ efficient long-range state tracking. These architectural innovations, if they mature to match Transformer performance on all benchmark dimensions, could reduce long-context inference energy by 1–2 orders of magnitude, with particular relevance to document processing, code analysis, and genomics applications.
Carbon-Aware Training Scheduling and Climate-Aligned AI Infrastructure: Real-time carbon intensity APIs — including Electricity Maps (global coverage), National Grid ESO Carbon Intensity API (GB grid), and ENTSO-E Transparency Platform (EU) — are being integrated directly into distributed training orchestration frameworks. Hugging Face Accelerate (version 0.24+) includes a carbon intensity callback; PyTorch Lightning has introduced a CarbonCallback plugin; Microsoft’s Azure training service supports carbon-aware region selection. The theoretical maximum benefit of temporal workload shifting (scheduling all flexible training compute to occur only during periods of 100% renewable generation) is estimated at 50–80% carbon reduction on the GB grid without hardware changes. Geographic shifting (routing to the lowest-carbon available data centre region) provides a further 2–3x benefit where multi-region infrastructure is available.
Standardised Energy Benchmarking and Carbon Accounting: MLCommons MLPerf Inference 2026–2027 cycle plans to extend accuracy-per-watt metrics to include scope 2 and scope 3 carbon emissions, covering supply chain embodied carbon in hardware as well as operational electricity carbon. ISO/IEC JTC 1/SC 42 WG3 is developing AI sustainability metrics for inclusion in the ISO/IEC 42001 revision expected 2026–2027, specifying standardised energy measurement methodologies for training (total training energy, training energy intensity per parameter, training energy per benchmark point), inference (energy per 1,000 queries at specified batch size and hardware type), and model lifecycle (embodied carbon amortisation, end-of-life hardware disposal). The EU AI Act’s technical standards mandate (CEN-CENELEC mandate M/589) includes energy measurement standards development.
Post-Training Quantisation Beyond INT4: Research into 2-bit and 1.58-bit quantisation (BitNet b1.58, Ma et al., 2024) suggests the theoretical possibility of near-binary weight models maintaining competitive language modelling performance when trained from scratch at appropriate scale. The energy implications are substantial: binary weight inference requires only bit-wise operations (XNOR for multiply, popcount for accumulate), reducing energy per multiply-accumulate by approximately 2 orders of magnitude versus FP16. While the accuracy-efficiency frontier at sub-4-bit is not yet competitive with INT4 for arbitrary task-specific fine-tuning, BitNet architectures demonstrate the existence of capable binary weight models, and the trajectory of post-training quantisation research (from FP32→INT8 in 2020, to INT4 in 2022, to INT2 in 2024) suggests that practical sub-4-bit deployment may be achievable within the 2026–2028 period. This would transform the energy economics of inference for smaller organisations, enabling capable AI on CPU-only hardware with power draws under 5 watts.
Research & Literature
-
Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. ACL 2019. https://arxiv.org/abs/1906.02629
-
Schwartz, R., Dodge, J., Smith, N. A., & Etzioni, O. (2020). Green AI. Communications of the ACM, 63(12), 54–63. https://doi.org/10.1145/3381831
-
Hoffmann, J., Borgeaud, S., Mensch, A., et al. (2022). Training compute-optimal large language models (Chinchilla). arXiv:2203.15556. https://arxiv.org/abs/2203.15556
-
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Ré, C. (2022). FlashAttention: Fast and memory-efficient exact attention with IO-awareness. NeurIPS 2022. https://arxiv.org/abs/2205.14135
-
Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2022). GPTQ: Accurate post-training quantization for generative pre-trained transformers. arXiv:2210.17323. https://arxiv.org/abs/2210.17323
-
Hinton, G., Vanhoucke, V., & Dean, J. (2015). Distilling the knowledge in a neural network. arXiv:1503.02531. https://arxiv.org/abs/1503.02531
-
Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. ICML 2019. https://arxiv.org/abs/1905.11946
-
Howard, A., et al. (2019). Searching for MobileNetV3. ICCV 2019. https://arxiv.org/abs/1905.02244
-
Lannelongue, L., Grealey, J., & Inouye, M. (2021). Green algorithms: Quantifying the carbon footprint of computation. Advanced Science, 8(12), 2100707. https://doi.org/10.1002/advs.202100707
-
International Energy Agency. (2025). Energy and AI: World Energy Outlook Special Report. IEA, Paris. https://www.iea.org/reports/key-questions-on-energy-and-ai
-
Patterson, D., et al. (2021). Carbon considerations for large model development. arXiv:2104.10350. https://arxiv.org/abs/2104.10350
-
Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. NeurIPS 2023. https://arxiv.org/abs/2305.14314
-
Ma, S., et al. (2024). The era of 1-bit LLMs: All large language models are in 1.58 bits. arXiv:2402.17764. https://arxiv.org/abs/2402.17764
-
Kim, S., et al. (2023). SqueezeLLM: Dense-and-sparse quantization. ICML 2023. https://arxiv.org/abs/2306.07629
-
Chen, T., Xu, B., Zhang, C., & Guestrin, C. (2016). Training deep nets with sublinear memory cost (gradient checkpointing). arXiv:1604.06174. https://arxiv.org/abs/1604.06174
-
Hu, E., et al. (2021). LoRA: Low-rank adaptation of large language models. ICLR 2022. https://arxiv.org/abs/2106.09685
-
CE-NAS: Carbon-Efficient Neural Architecture Search. OpenReview, NeurIPS 2024 Workshop. https://openreview.net/forum?id=v6W55lCkhN
-
An exploration of the effect of quantisation on energy consumption and inference time of StarCoder2. (2024). arXiv:2411.12758. https://arxiv.org/abs/2411.12758
-
How hungry is AI? Benchmarking energy, water, and carbon footprint of LLM inference. (2025). arXiv:2505.09598. https://arxiv.org/abs/2505.09598
-
Leviathan, Y., Kalman, M., & Matias, Y. (2023). Fast inference from transformers via speculative decoding. ICML 2023. https://arxiv.org/abs/2211.17192
-
Child, R., et al. (2019). Generating long sequences with sparse transformers. arXiv:1904.10509. https://arxiv.org/abs/1904.10509
-
Promwad. (2025). AI model compression: Pruning and quantization strategies for real-time devices. https://promwad.com/news/ai-model-compression-real-time-devices-2025
-
Green AI techniques for reducing energy consumption in AI systems. (2025). ScienceDirect / Energy and AI. https://www.sciencedirect.com/science/article/pii/S2590005625002796
-
Green AI: A systematic review and meta-analysis of its definitions, lifecycle models, hardware and measurement attempts. (2025). arXiv:2511.07090. https://arxiv.org/abs/2511.07090
-
MIT Sloan Management Review. (2025). AI has high data center energy costs — but there are solutions. https://mitsloan.mit.edu/ideas-made-to-matter/ai-has-high-data-center-energy-costs-there-are-solutions
-
Presenc AI. (2026). AI Data Center Energy Consumption Statistics 2026. https://presenc.ai/research/ai-data-center-energy-consumption-2026
-
TTMS. (2026). AI data centers energy consumption 2024–2026: Trends, projections, environmental impact. https://ttms.com/growing-energy-demand-of-ai-data-centers-2024-2026/
-
University of Edinburgh. (2025). AI-driven discovery of quantum materials for neuromorphic and energy-efficient computing. https://www.ph.ed.ac.uk/phd-projects/ai-driven-discovery-of-quantum-materials-for-neuromorphic-and-energy-efficient-computing
-