Time orecasting (TSF) is the domain of statistical and machine learning mods that predict future values of ordered sequential observations indexed by time, encompassing classical statistical approaches (ARIMA — AutoRegressive Integrated Moving Average — modelling the s a linear combination of its…

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:hasPart ai:PointForecast))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:hasPart ai:IntervalForecast))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:hasPart ai:ProbabilisticForecast))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:hasPart ai:ForecastHorizon))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:hasPart ai:LookbackWindow))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:hasPart ai:Covariates))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:hasPart ai:ReconciliationStep))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:hasPart ai:EvaluationMetric))

## Dependency Relationships
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:requires ai:HistoricalTimeSeriesData))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:requires ai:StationarityAnalysis))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:requires ai:CrossValidationStrategy))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:requires ai:ForecastEvaluationProtocol))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:requires ai:ComputeInfrastructure))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:dependsOn ai:ProbabilityTheory))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:dependsOn ai:StatisticalInference))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:dependsOn ai:OptimisationTheory))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:dependsOn ai:DeepLearning))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:dependsOn ai:SignalProcessing))

## Capability Relationships
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:enables ai:DemandPlanning))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:enables ai:RiskQuantification))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:enables ai:AnomalyDetection))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:enables ai:CapacityPlanning))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:enables ai:DecisionSupport))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:enables ai:ScenarioAnalysis))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:supports ai:EnergyForecasting))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:supports ai:SupplyChainManagement))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:supports ai:FinancialRiskManagement))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:supports ai:ClinicalDecisionSupport))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:supports ai:SmartGridOperations))

## Implementation Relationships
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:implements ai:ARIMA))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:implements ai:ExponentialSmoothing))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:implements ai:LSTMRecurrence))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:implements ai:TransformerAttention))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:implements ai:StateSpaceModels))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:implements ai:QuantileRegression))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:implements ai:HierarchicalReconciliation))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:uses ai:StationarityTesting))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:uses ai:AutocorrelationAnalysis))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:uses ai:FourierTransform))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:uses ai:GradientDescent))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:uses ai:BayesianInference))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:uses ai:EnsembleMethods))

## Reduction Relationships
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:reduces ai:ForecastingError))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:reduces ai:InventoryHoldingCost))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:reduces ai:OperationalUncertainty))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:reduces ai:ManualForecastingLabour))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:reduces ai:DowntimeRisk))

## Association Relationships
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:relatedTo ai:AnomalyDetection))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:relatedTo ai:Econometrics))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:relatedTo ai:SignalProcessing))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:relatedTo ai:OperationsResearch))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:contrastsWith ai:RegressionAnalysis))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:contrastsWith ai:Classification))
SubClassOf(ai:TimeSeriesForecasting
  ObjectSomeValuesFrom(ai:contrastsWith ai:CausalInference))

## Data Properties (Characteristics)
DataPropertyAssertion(ai:hasIdentifier ai:TimeSeriesForecasting "AI-1087"^^xsd:string)
DataPropertyAssertion(ai:authorityScore ai:TimeSeriesForecasting "0.87"^^xsd:decimal)
DataPropertyAssertion(ai:mCompetitionSeries ai:TimeSeriesForecasting "6"^^xsd:integer)
DataPropertyAssertion(ai:m5SeriesCount ai:TimeSeriesForecasting "42840"^^xsd:integer)
DataPropertyAssertion(ai:chronosTrainingPoints ai:TimeSeriesForecasting "100000000000"^^xsd:integer)
DataPropertyAssertion(ai:moiraiLOTSAObservations ai:TimeSeriesForecasting "27000000000"^^xsd:integer)
DataPropertyAssertion(ai:timesFMParameters ai:TimeSeriesForecasting "200000000"^^xsd:integer)

## Property Constraints
SubClassOf(ai:TimeSeriesForecasting
  DataMinCardinality(1 ai:hasForecastHorizon xsd:integer))
SubClassOf(ai:TimeSeriesForecasting
  DataMinCardinality(1 ai:hasLookbackWindow xsd:integer))
SubClassOf(ai:TimeSeriesForecasting
  DataAllValuesFrom(ai:isProbabilistic xsd:boolean))
SubClassOf(ai:TimeSeriesForecasting
  DataSomeValuesFrom(ai:hasEvaluationMetric xsd:string))

## Annotations
AnnotationAssertion(rdfs:label ai:TimeSeriesForecasting "Time Series Forecasting"@en)
AnnotationAssertion(rdfs:comment ai:TimeSeriesForecasting "Domain of statistical and machine learning methods predicting future values of temporally ordered sequential observations, spanning classical ARIMA/ETS/Prophet, deep learning DeepAR/N-BEATS/N-HiTS/TFT, foundation models Chronos/TimesFM/TimeGPT/Moirai/Lag-Llama, Mamba/state-space architectures, hierarchical MinT reconciliation, neural ODEs, anomaly detection, and M-competition benchmarking (M1-M6), with production deployment in energy, retail, finance, clinical and smart-city domains."@en)
AnnotationAssertion(dcterms:identifier ai:TimeSeriesForecasting "AI-1087"^^xsd:string)
AnnotationAssertion(dcterms:subject ai:TimeSeriesForecasting "Forecasting, Time Series, ARIMA, Deep Learning, Foundation Models, Probabilistic Forecasting"@en)

)

Property Characteristics

AsymmetricObjectProperty(ai:requires) AsymmetricObjectProperty(ai:enables) AsymmetricObjectProperty(ai:implements) AsymmetricObjectProperty(ai:contrastsWith) TransitiveObjectProperty(ai:dependsOn) FunctionalDataProperty(ai:hasForecastHorizon) FunctionalDataProperty(ai:authorityScore)

About Time Series Forecasting

  • Time Series Forecasting (TSF) is among the oldest and most practically consequential subfields of statistics and machine learning, concerned with predicting future values of observations recorded sequentially over time.
  • A time series is a sequence {y_1, y_2, …, y_T} indexed by discrete time steps t = 1, …, T (or by continuous time in irregular-sampling settings), where temporal ordering is fundamental — shuffling the observations destroys the signal.
  • This distinguishes TSF from cross-sectional prediction: the past values of the target, its autocorrelation structure, seasonality, trend, and covariate history all constitute predictive information not available in i.i.d. data.
  • The goal is to produce forecasts ŷ_{T+h} for horizons h = 1, …, H, either as:
    • Point estimates: conditional mean or median of the predictive distribution
    • Quantile forecasts: estimates at specified coverage levels (10th, 50th, 90th percentile) for asymmetric loss settings such as inventory optimisation
    • Full distributional forecasts: parametric likelihoods (Gaussian, Student-t, Negative Binomial) or non-parametric sample paths for risk quantification
    • Probability-of-exceedance outputs: binary threshold-crossing probabilities for alarm systems
  • The choice of forecast type is dictated by the downstream decision: inventory optimisation requires quantile forecasts to balance stockout against holding costs; grid balancing requires probabilistic distributions over wind/solar power; credit risk monitoring needs conditional default probabilities given macroeconomic scenarios.
  • The field spans a century of methodological development — from the Yule (1927) autoregressive model and Slutsky (1937) moving-average model, through the Box-Jenkins ARIMA framework (1970), through the ETS state space unification (Hyndman et al. 2008), to the deep learning revolution (LSTNet 2018, DeepAR 2020) and the foundation model era (2024–2026).
  • Empirical evidence from the M-Competition series has repeatedly shown that ensembles of simple methods outperform complex single models — a finding that shaped both research priorities and practitioner intuition, even as modern deep learning substantially shifted the state-of-the-art when sufficient cross-series data is available.

Core Mathematical Framework

  • A univariate time series model posits a data-generating process y_t = f(y_{t-1}, y_{t-2}, …, ε_t) where ε_t ~ i.i.d.(0, σ²) is white noise.
  • ARIMA(p,d,q) specifies: (1 - φ_1 B - … - φ_p B^p)(1-B)^d y_t = (1 + θ_1 B + … + θ_q B^q) ε_t
    • B is the backshift operator: B^k y_t = y_{t-k}
    • d differencing operations achieve stationarity (Augmented Dickey-Fuller or KPSS test guides selection)
    • p, q selected by minimising AIC or BIC over a grid search (Hyndman-Khandakar algorithm automates this in R forecast)
    • SARIMA(p,d,q)(P,D,Q)_s adds multiplicative seasonal AR/MA and differencing at lag s
    • Model diagnostics: Ljung-Box Q test for autocorrelation of residuals; Jarque-Bera normality test; AIC = 2k - 2 ln(L) penalising complexity; BIC = k ln(n) - 2 ln(L) penalising more severely at large n
  • State Space / ETS: Reformulating ARIMA and ETS as latent-state linear systems:
    • Observation: y_t = Z_t’ α_t + ε_t
    • State transition: α_{t+1} = T_t α_t + R_t η_t
    • Kalman filter recursion yields optimal linear prediction and maximum likelihood parameter estimation
    • Harvey (1989), de Jong (1991), Durbin and Koopman (2001) unified ETS into a principled probabilistic framework
    • ETS(A,A,A) — additive errors, additive trend, additive seasonality — equivalent to classical Holt-Winters additive method
    • ETS(M,Ad,M) — multiplicative errors, additive damped trend, multiplicative seasonality — Hyndman et al. (2002) “damped” variant consistently top-performing on M3 benchmarks
  • Stationarity and Pre-processing:
    • Augmented Dickey-Fuller (ADF) test: H₀: unit root present (non-stationary); reject at p < 0.05 to confirm stationarity
    • KPSS test: H₀: stationary; complementary to ADF; recommended to run both to classify series as I(0), I(1), or ambiguous
    • Box-Cox transformation λ: y_t^(λ) = (y_t^λ - 1)/λ; automatic λ selection via profile likelihood; λ=0 gives log transform; stabilises variance
    • Seasonal adjustment: X-13ARIMA-SEATS (US Census Bureau) and TRAMO-SEATS (Bank of Spain) for official statistics; classical STL decomposition (Cleveland et al. 1990) for fast robust seasonality extraction
  • Probabilistic Forecast Evaluation:
    • CRPS (Continuous Ranked Probability Score) = E[|Y - y|] - (1/2) E[|Y - Y’|] where Y, Y’ are independent draws from the predictive distribution; strictly proper scoring rule rewarding both accuracy and calibration; lower is better
    • For parametric forecasts: CRPS can be computed analytically for Gaussian (mean μ, std σ): CRPS = σ{(y-μ)/σ [2Φ((y-μ)/σ) - 1] + 2φ((y-μ)/σ) - 1/√π}
    • Pinball / Quantile Loss at level τ: L_τ(y, ŷ) = τ(y - ŷ) if y ≥ ŷ, else (1-τ)(ŷ - y); used in M5 Winkler score evaluation
    • Winkler Interval Score for (1-α)% prediction interval [l, u]: W = (u - l) + (2/α)(l - y) 1[y < l] + (2/α)(y - u) 1[y > u]; rewards narrow calibrated intervals
    • MASE (Mean Absolute Scaled Error): MAE of forecast / MAE of seasonal naïve baseline; scale-independent, enabling cross-series comparison; MASE < 1 beats the seasonal naïve baseline
    • SMAPE (Symmetric Mean Absolute Percentage Error): 200 × |y - ŷ| / (|y| + |ŷ|); bounded in [0%, 200%]; used in M3 and M4 competitions; criticised for undefined behaviour near zero
    • WRMSSE (Weighted Root Mean Squared Scaled Error): M5 competition metric weighting series by their sales revenue contribution; amplifies errors on high-value items
  • Cross-Validation for Time Series:
    • Standard k-fold CV violates temporal order (future data leaking into training); must use temporal ordering
    • Walk-forward validation (rolling origin evaluation): train on first k observations, test on k+1 through k+h, advance origin by 1, repeat; computationally intensive but unbiased
    • Expanding window: training window grows monotonically; appropriate when more data always helps
    • Fixed rolling window: training window slides with fixed length; appropriate when recent data more relevant (concept drift)
    • Block cross-validation: maintain blocks of consecutive observations; reduces but does not eliminate leakage for correlated series

Components and Architecture

Classical Statistical Family

  • ARIMA / SARIMA: Box-Jenkins framework; automated via auto.arima() (R forecast) and statsmodels.tsa.statespace.SARIMAX (Python).
    • Strengths: interpretable, fast, well-understood uncertainty intervals, no training data requirement beyond the univariate series
    • Weaknesses: linearity assumption, limited covariate integration, poor scalability to millions of series, assumes Gaussian errors
  • ETS (Error-Trend-Seasonality): 30 possible (E,T,S) model combinations; AIC-optimal selection.
    • State space exponential smoothing with smoothing parameters (α, β, γ) controlling adaptation rates
    • Dominant for monthly business series; won many M3 and M4 series categories
    • Implemented in R ets(), Python statsforecast ETS()
  • Prophet (Taylor and Letham 2018, Meta): Analyst-friendly additive decomposition model.
    • Trend: piecewise linear or logistic growth with automatic changepoint detection via Laplace prior on changepoint magnitudes
    • Seasonality: Fourier series at daily/weekly/yearly frequencies; truncated at K=10 terms by default
    • Holidays: sparse additive indicator regressors with optional window effects
    • Estimated via Stan’s L-BFGS optimiser; designed for business analysts without forecasting expertise
    • Widely deployed in business dashboards (Meta, LinkedIn, Uber) despite mediocre M-competition rank on generic statistical series
  • Theta Method (Assimakopoulos and Nikolopoulos 2000): Won the M3 competition.
    • Decomposes series into theta lines (θ=0: regression line, θ=2: original series amplifying local trends)
    • Combines forecasts from two theta lines; shown to be equivalent to SES with drift
    • Remains a strong baseline in M4 and beyond due to simplicity vs accuracy balance
  • TBATS / BATS: Exponential smoothing with Box-Cox transformation, ARMA errors, trend and seasonal components.
    • Handles multiple non-integer seasonalities (e.g. daily/weekly overlap) common in electricity demand
    • Computationally expensive; state space with Kalman filter estimation
  • VAR (Vector Autoregression): Multivariate extension of AR to multiple related series.
    • y_t = A_1 y_{t-1} + … + A_p y_{t-p} + ε_t where y_t is a k-dimensional vector
    • Used in macro-economic forecasting and energy market analysis; Granger causality testing embedded within VAR framework

Deep Learning Architectures

  • DeepAR (Salinas et al. 2020, Amazon)
    • Architecture: multi-layer LSTM autoregressive decoder; at each step t, hidden state h_t = LSTM(z_{t-1}, h_{t-1}, x_t) where z_{t-1} is the previous target value and x_t is the covariate vector
    • Output head: Gaussian likelihood (μ_t, σ_t) = Linear(h_t) for real-valued series; Negative Binomial (μ_t, α_t) for count data
    • Global training: all series from the dataset share encoder weights, enabling cross-series knowledge transfer critical for cold-start items with no history
    • Deployed at Amazon scale: product demand forecasting across 400M+ active ASINs; inventory optimisation at AWS Forecast service
    • Key innovation: probabilistic output head provides calibrated quantile forecasts without requiring post-hoc quantile regression
  • N-BEATS (Oreshkin et al. 2020, Element AI)
    • Architecture: stack of doubly-residual blocks, each producing backcast (residual subtracted from input) and forecast (contribution to output)
    • Each block: FC layers → basis expansion θ_f, θ_b → forecast g_f(θ_f), backcast g_b(θ_b) via basis functions (polynomial: {1, t, t², …} or Fourier: {cos(2πt/P), sin(2πt/P), …})
    • “Generic” variant: arbitrary learned basis; “Interpretable” variant: polynomial trend + Fourier seasonality basis, providing decomposition interpretable by analysts
    • Won M4 competition when trained only on M4 data; no recurrence, no convolution, no attention
    • Training on public datasets (M4, M3, Tourism) alone yields SOTA; further boosted by dataset-specific fine-tuning
  • N-HiTS (Challu et al. 2023)
    • Extension of N-BEATS with hierarchical interpolation and multi-rate input sampling
    • Coarse block: MaxPooled(8) input captures trend; Medium block: MaxPooled(4) captures medium seasonality; Fine block: raw input captures noise
    • Interpolation layer upsamples block outputs to forecast horizon, enabling each block to specialise on its frequency band
    • Achieves best-in-class on long-horizon benchmarks (ETTh1 96-step, ETTm1, Exchange Rate, ILI)
    • 50% fewer FLOPs vs N-BEATS at equivalent accuracy; preferred for resource-constrained deployment
  • Temporal Fusion Transformer (Lim et al. 2021, Google)
    • Gated Residual Networks (GRN): gated skip connections controlling information flow throughout architecture
    • Variable Selection Networks: learned importance weights for each input variable (static covariates, observed history, known future inputs)
    • LSTM encoder/decoder: processes local temporal patterns whilst preserving positional context
    • Multi-head self-attention: learns long-range temporal dependencies across all time steps in the input window
    • Quantile output: simultaneously predicts 10th, 50th, 90th percentiles via shared penultimate representation
    • Interpretable attention maps identify which past time steps drive each forecast horizon — valuable for audit and debugging in regulated industries
    • Deployed at Google, Ocado, NHS England analytics platforms
  • PatchTST (Nie et al. 2023)
    • Applies Vision Transformer (ViT) architecture to time series: divides series into non-overlapping patches of length P (default 16)
    • Each patch treated as a token; positional encoding applied; Transformer encoder over patch sequence
    • Channel independence: each variate (channel) processed independently — prevents spurious cross-variate leakage on noisy datasets
    • Reduces effective sequence length from T to T/P: quadratic Transformer complexity becomes O((T/P)²), enabling longer lookback windows (336, 720 time steps) at manageable cost
    • Outperforms Autoformer, FEDformer, Informer on ETT, Weather, ECL, Traffic long-horizon benchmarks
  • iTransformer (Liu et al. 2024, Tsinghua)
    • Inverts the standard Transformer’s attentional axes: applies self-attention across variates (channels) at the same time step, rather than across time steps
    • Each variate’s entire time series embedded as a single token; FFN processes temporal dependencies within each variate’s token
    • Captures inter-variate correlations more explicitly than temporal attention which mixes variate and temporal information
    • Achieves SOTA on multivariate long-horizon benchmarks (Traffic, Weather, ECL, ETT in 2024 evaluation)
    • Particularly effective for multivariate datasets with strong cross-variate dependencies (e.g. sensor arrays, electricity consumption per household)
  • DLinear (Zeng et al. 2023): The “Are Transformers Effective?” benchmark paper showed a simple linear layer — decompose trend and residual, apply separate linear projections — matched or outperformed Autoformer/FEDformer/Informer on multiple benchmarks.
    • Published the sobering finding that architectural complexity does not guarantee improvement over linear baselines on standard datasets
    • Prompted re-evaluation of benchmark validity, dataset diversity, and proper ablation practices in the TSF community
    • iTransformer and PatchTST subsequently restored transformer credibility on broader, more diverse evaluations

Foundation / Zero-Shot Models (2024-2026)

  • Chronos (Ansari et al. 2024, Amazon)
    • Tokenisation: real-valued series normalised by mean scaling (μ = mean(|y|)) then quantised into B=4096 discrete bins via linear binning
    • Architecture: T5 encoder-decoder (Raffel et al. 2020) treating forecasting as next-token prediction over the quantised vocabulary
    • Pre-training corpus: ~800K real series from Monash, M5, FRED, GluonTS repositories; augmented with synthetic series sampled from Gaussian processes with varied kernels (RBF, Matérn, Periodic, linear)
    • Model sizes: Chronos-Mini (21M), Chronos-Small (46M), Chronos-Base (200M), Chronos-Large (710M) — all trained identically, larger models consistently outperform smaller
    • Zero-shot evaluation on 42 held-out datasets: median CRPS competitive with or better than supervised per-dataset models on 21 of 42 datasets, including weather, energy, finance, and retail series
    • The synthetic data augmentation proved critical: models trained without it showed 15-25% higher CRPS on out-of-domain series
  • TimesFM (Das et al. 2024, Google DeepMind)
    • Architecture: decoder-only Transformer (200M parameters) inspired by PaLM/GPT structure, adapted for time series via input/output patching
    • Input: non-overlapping patches of length 32; each patch embedded via linear projection plus positional encoding
    • Output: causal generation of forecast patches; multiple forecast patches generated sequentially for longer horizons
    • Pre-training corpus: 100 billion time-points from Google Trends (2006-2023), Wikipedia pageviews, first-party Google data, and public sources
    • Zero-shot MASE on 37 diverse benchmarks: outperforms Prophet, ARIMA, and classical ETS on 29/37 series; near-supervised performance on many categories without any fine-tuning
    • Fine-tuning available: 100-1000 domain observations sufficient to match or exceed supervised domain-specific baselines
    • Deployed internally at Google for product demand and ad campaign performance forecasting; TimesFM 2.0 (2026) extends corpus to 1 trillion data-points and adds multivariate support
  • TimeGPT-1 (Garza and Mergenthaler-Canseco 2024, Nixtla)
    • Transformer trained on the world’s largest disclosed proprietary time series corpus: 100B+ data-points spanning 50 different domains
    • Commercial API (api.nixtla.io): REST and Python SDK nixtlats supporting zero-shot inference, fine-tuning via parameter-efficient LoRA adaptation, probabilistic interval generation, anomaly detection, counterfactual scenario modelling
    • Integrated into Azure Marketplace as “Azure AI Forecasting” service (2025); free developer tier (50K API calls/month) and enterprise tiers
    • Fine-tuning latency: 2-5 minutes on 1K domain observations; significant improvement over zero-shot on domain-specific series
    • Anomaly detection: residuals between forecast and observed; automatic threshold calibration via conformal prediction coverage guarantees
    • TimeGPT-2 (Q2 2026 preview): 8192-step context window vs 1024 in TimeGPT-1; improved multivariate cross-variate attention
  • Moirai (Woo et al. 2024, Salesforce)
    • Architecture: Masked Encoder Transformer (BERT-style bidirectional attention) applied to patched time series tokens
    • LOTSA (Large Open Time Series Archive): 27B observations across 9 domains — energy (electricity, gas), finance (stocks, FX), transport (traffic), web (Wikipedia pageviews), climate (weather stations), healthcare (medical monitors), air quality (PM2.5, ozone), economics (macro indicators), nature (ecological sensor arrays)
    • Any-variate attention: variable-length multivariate input handled by per-variate patch tokens plus a learnt variate-identity embedding — supports univariate through high-dimensional multivariate at inference time without re-training
    • Frequency-specific output: separate linear heads for each frequency family (hourly, daily, weekly, monthly, yearly) with appropriate seasonality priors
    • Output distribution: Student-t / Negative Binomial / Gaussian mixture selected per series based on likelihood maximisation
    • Zero-shot CRPS: Moirai-Large (311M) competitive with supervised baselines on 29 benchmarks; better than Chronos-Large on weekly and monthly frequencies
    • Moirai-MoE (2025): Mixture-of-Experts routing at FFN level, 16 experts per layer, top-2 routing — reduces effective FLOPs 40% at same parameter budget
  • Lag-Llama (Rasul et al. 2024, ServiceNow/Mila/McGill)
    • Open-source foundation model adapting LLaMA-7B architecture to univariate probabilistic forecasting
    • Input: lagged values at multiple fixed lags (1, 2, 3, 4, 5, 6, 7, 11, 12, 23, 24, 47, 48, 95, 96, 143, 144, 167, 168, 335, 336, 719, 720) capturing short, medium and long-range dependencies across hourly/daily/weekly patterns
    • Time covariates: day-of-week, day-of-month, week-of-year, month-of-year, hour-of-day as learnt embeddings
    • Output head: Student-t distribution (μ, σ, ν) parameters via linear projection from final transformer hidden state
    • Pre-training: 352 datasets from Monash Time Series Archive; 1.3B parameters trained on a single node with 8xA100 GPUs for 4 days
    • Zero-shot CRPS: SOTA on several hourly and daily Monash benchmarks; enables few-shot fine-tuning with 5-10 labelled observations per series

State Space and Mamba Family

  • S4 (Gu, Goel, Re 2022, Stanford)
    • Structured State Space Sequence model based on continuous-time ODE: ẋ(t) = Ax(t) + Bu(t), y(t) = Cx(t) + Du(t)
    • HiPPO matrix initialisation: A initialised as “High-order Polynomial Projection Operator” ensuring history projected onto Legendre polynomial basis, enabling faithful long-range memory
    • Diagonal-Plus-Low-Rank (DPLR) parameterisation enables O(N log N) FFT-based training instead of O(N²) quadratic state expansion
    • Achieves near-perfect 96.6% accuracy on Path-X 16K sequence task (memorising path of 16,000 steps) where LSTMs and Transformers fail outright
    • Competitive on Time Series Classification (UEA/UCR benchmarks), sequential CIFAR-10, speech modelling
  • Mamba (Gu and Dao 2023)
    • Selective SSM: replaces S4’s fixed (A, B, C) matrices with input-dependent (Δ_t, B_t, C_t) = f(x_t) functions — the “selection mechanism”
    • The selection allows the model to contextually decide which historical information to retain vs compress; resolves the fixed-parameter bottleneck limiting S4’s adaptivity
    • Hardware-aware parallel scan: custom CUDA kernel fusing the recurrent scan into a single GPU memory pass, achieving linear-time training and inference with constant memory in sequence length
    • Outperforms Transformers of equivalent size on language modelling (The Pile benchmark); linear scaling to 1M+ sequence length
    • TimeMamba, S-Mamba, MambaFormer (2024) adapt selective scanning to multivariate time series: channel-mixing variants outperform iTransformer on several Traffic and Weather benchmarks whilst using 50% less GPU memory

Hierarchical Forecasting and Reconciliation

  • Problem: Forecasts at different aggregation levels (store/region/national, SKU/category/total) are typically generated independently and may be incoherent — store-level sum may not equal the national forecast.
  • Bottom-Up Aggregation: Forecast at finest granularity; aggregate upward by summation.
    • Preserves item-level signal; loses top-level smoothing
    • Preferred when bottom-level data is rich and top-level aggregates add little information
  • Top-Down Disaggregation: Forecast at aggregate level; allocate proportionally using historical shares.
    • Smooths out item-level noise; propagates aggregate bias to disaggregate levels
    • Preferred when aggregate-level series are longer and more regular
  • Middle-Out: Forecast at an intermediate level; aggregate up and disaggregate down simultaneously.
  • MinT — Minimum Trace Reconciliation (Wickramasuriya et al. 2019):
    • Formulate hierarchical structure as constraint: S β̂ = y̅ where S is the summing matrix, β̂ are base forecasts, y̅ are reconciled forecasts
    • Optimal reconciliation matrix P minimises trace of covariance of reconciled forecast errors: P* = (S’ Ŵ^{-1} S)^{-1} S’ Ŵ^{-1}
    • Ŵ estimated as: OLS (identity covariance), WLS (variance-scaled), MinT-Shrink (shrinkage estimator for large hierarchies), MinT-Sample (full sample covariance on training windows)
    • Implemented in R packages hts, fabletools; Python statsforecast (>1.5.0)
    • Consistently outperforms bottom-up and top-down across M5 hierarchy levels; improvement most pronounced at intermediate aggregation levels
  • ML Reconciliation (2022-2024): End-to-end neural approaches learn reconciliation jointly with base forecasting model.
    • ERM-Reconciliation: Empirical Risk Minimisation with coherence-promoting regularisation term
    • DeepHierarchicalForecasting: separate LSTM per hierarchy level with cross-level attention
    • Improvement over MinT on M5 Walmart hierarchy: 5-12% further CRPS reduction at store/item level

Anomaly Detection in Time Series

  • Statistical Methods:
    • CUSUM (Page 1954): cumulative sum control chart detecting sustained shifts in mean; C_t = max(0, C_{t-1} + (x_t - μ_0 - k)) where k is allowance parameter
    • EWMA: exponentially weighted moving average; control chart with adaptive variance estimate
    • Seasonal-Hybrid ESD (Twitter AnomalyDetection, 2015): decompose series, apply Extreme Studentised Deviate test on residuals; handles multiple seasonalities
    • BOCPD — Bayesian Online Changepoint Detection (Adams and MacKay 2007): maintains posterior over changepoint location via sequential Bayes; computationally O(T²) but can be approximated
  • Deep Learning Methods:
    • LSTM-AE / LSTM-VAE: train autoencoder on normal data; at inference time, high reconstruction MSE signals anomaly
    • OmniAnomaly (Su et al. 2019): robust VRNN + normalising flow handling stochastic temporal dependencies; evaluated on SMAP/MSL/SMD spacecraft telemetry
    • Anomaly Transformer (Zhou et al. 2022): replace reconstruction loss with association discrepancy between local adjacent attention and global attention patterns; anomalies exhibit distinctive association discrepancy patterns
  • Foundation Model Anomaly Detection:
    • TimeGPT-1: forecast future timesteps; compare prediction intervals against observed; statistical test on residuals with conformal prediction coverage guarantees
    • Chronos: generate M sample paths; compute empirical prediction interval; flag observations outside 99th percentile
    • No labelled anomaly data required — zero-shot anomaly detection enables deployment without curating ground-truth anomaly labels

Use Cases / Major Families

  • Energy and Utilities: Electricity load forecasting 1h-7 days ahead for grid balancing.
    • National Grid ESO (UK) dispatches generation based on 30-minute load forecasts; TFT-based probabilistic load model deployed 2025 replacing linear regression, achieving 15% CRPS improvement
    • RTE France, EirGrid Ireland, ENTSO-E European grid: SARIMA and GBM ensembles at operational centres
    • Renewable generation forecasting: wind power quantile forecasts via Gaussian process models and numerical weather prediction (NWP) integration; solar irradiance ML downscaling
    • Day-ahead price forecasting for electricity markets: deep learning price spike prediction critical for trading desks and renewable energy hedging
  • Retail and Supply Chain: Demand forecasting across millions of SKUs.
    • Walmart M5 benchmark: 42,840 daily series across stores, departments, items; LightGBM with lag features won most M5 Accuracy track submissions; deep learning N-HiTS competitive at weekly/monthly aggregation
    • Amazon demand forecasting: DeepAR deployed across 400M+ ASINs; probabilistic forecasts feed inventory replenishment and warehouse automation systems
    • Tesco (UK): seasonal demand spikes, promotions uplift modelling, waste reduction in fresh produce via 14-day probabilistic forecasts
    • Out-of-stock prediction: anomaly detection on sales velocity velocity to flag stockouts before they propagate to customer-facing shortages
  • Finance and Economics: Volatility and risk forecasting.
    • GARCH (Bollerslev 1986) and GJR-GARCH remain industry standards for daily equity and FX volatility; ML hybrids (GARCH-LSTM, HAR-LSTM) extend to intraday volatility
    • M6 Financial Forecasting Competition (2024): monthly return point forecasts plus portfolio weight allocation for 50 US equity series; novel evaluation combining MASE accuracy with portfolio Sharpe ratio
    • Credit risk early warning: delinquency rate prediction over 12-month horizons using macro-economic leading indicators as external regressors in TFT
    • Nowcasting: high-frequency indicators (Google Trends, satellite imagery, payment transaction flows) used to estimate GDP, inflation, unemployment in near-real-time ahead of official statistics releases
  • Healthcare and Epidemiology: NHS and public health applications.
    • NHS bed occupancy and emergency department attendances: NHS England SitRep short-term forecasts (1-7 day) using SARIMA and ensemble ML; NHSX Analytics Platform exploring TFT for longer-horizon capacity planning
    • COVID-19 epidemic curve projection: SAGE SPI-M-O ensemble of SIR compartmental, renewal equation, and ML models; ensemble median used for policy; quantile forecasts for uncertainty communication
    • Sepsis early warning: vital sign time-series LSTM classifiers detecting early physiological deterioration 6-12h ahead of clinical diagnosis; deployed in Leeds Teaching Hospitals NHS Trust (Aire sepsis model)
    • WHO Global Influenza Surveillance: seasonal flu incidence forecasting via ensemble of SARIMA, compartmental and ML models across 50 countries
  • Predictive Maintenance and IoT: Industrial sensor time series anomaly detection and degradation forecasting.
    • Turbine blade fatigue monitoring: RNN anomaly detection on vibration spectra (GE Digital Predix, Siemens MindSphere) enabling maintenance scheduling before unplanned outage
    • Aircraft engine degradation (NASA C-MAPSS benchmark): remaining useful life (RUL) prediction from multivariate sensor streams; LSTM and Transformer models achieving RMSE 12-15 cycles vs 40+ for classical methods
    • Rolls-Royce TotalCare engine health monitoring: anomaly detection on fuel flow, EGT (exhaust gas temperature), vibration series predicting shop visits; AOG (aircraft on ground) event reduction enabling £50M+ annual savings
    • Manufacturing process monitoring: multivariate LSTM on CNC machine sensor data detecting tool wear patterns (SINUMERIK Analytics, Siemens)
  • Traffic and Urban Mobility: Spatio-temporal time series forecasting.
    • Road link speed forecasting: Graph Neural Network extensions (ST-GCN, DCRNN) modelling spatial road network topology alongside temporal dependencies; Uber Movement uses historical speed predictions for routing ETA
    • Public transport demand: Transport for London ridership (Oyster card tap-in data) for rolling stock planning; Network Rail passenger volumes for capacity management
    • METR-LA and PEMS-BAY benchmarks: standard multivariate graph-based time series benchmarks measuring MAE of 15/30/60-minute-ahead traffic speed forecasts
    • Last-mile logistics: DHL, DPD delivery demand density forecasting per postcode sector for driver allocation optimisation
  • Climate Science: Long-horizon probabilistic forecasting.
    • Temperature anomaly projection: ML post-processing (quantile mapping, distributional regression) of CMIP6 GCM output for regional climate projections
    • Extreme precipitation: deep learning downscaling of NWP coarse grid to station-level; Met Office/ECMWF-DeepMind collaboration (GraphCast, NeuralGCM approaches)
    • Seasonal-to-decadal forecasting: ensemble mean bias correction; S2S (Subseasonal-to-Seasonal Prediction) competitions benchmarking ML against physics-based NWP at 2-6 week horizon

Academic Context

  • M-Competition Legacy: The empirical forecasting competition series initiated by Spyros Makridakis in 1982 has shaped the field more than any theoretical development.
    • M1 (1001 series, 1982): simple exponential smoothing competitive with Parzen spectral methods; first empirical evidence against the “complex is better” assumption
    • M3 (3003 series, 2000): simple methods (Theta, Robust-Trend, Damped ETS) outperform complex models; spawned the “simple methods hypothesis”
    • M4 (100,000 series, 2018): ES-RNN hybrid (Smyl, Uber) won — a combination of Holt-Winters with a neural correction network — by approximately 3% SMAPE over the best statistical entry; pure ML did not win outright
    • M5 (Makridakis et al. 2022, Walmart retail): 42,840 daily item-store sales series; LightGBM gradient boosting dominated the leaderboard alongside LSTM hybrids; M5 Uncertainty Track evaluated CRPS across 9 quantiles; N-BEATS/N-HiTS competitive at weekly aggregation
    • M6 (2024): financial forecasting under realistic portfolio-construction constraints (monthly returns, portfolio weights summing to 1); novel combined accuracy-portfolio score rewarding both MASE accuracy and mean portfolio Sharpe ratio
  • Key Theoretical Contributions:
    • Yule (1927) — autoregressive model for sunspot number prediction
    • Slutsky (1937) — moving-average processes from random shocks
    • Muth (1960) — optimal forecasting with exponential smoothing
    • Box and Jenkins (1970) — ARIMA unified framework; cornerstone of industrial forecasting for 30 years
    • Holt (1957) and Winters (1960) — double/triple exponential smoothing for trend and seasonality
    • Assimakopoulos and Nikolopoulos (2000) — Theta method; M3 competition winner
    • Hyndman et al. (2008) — ETS state space unification; R forecast package foundation
    • Harvey (1989) — structural time series via Kalman filter
    • Hamilton (1994) — “Time Series Analysis”; graduate econometrics standard
    • Makridakis, Wheelwright, Hyndman (1998) — “Forecasting: Methods and Applications” 3rd ed.; practitioner bible
    • Hyndman and Athanasopoulos (2021) — “Forecasting: Principles and Practice” 3rd ed.; open-access standard reference (fpp3)
  • Deep Learning Turning Points:
    • LSTNet (Lai et al. 2018): first strong LSTM/CNN combination outperforming ARIMA on large traffic/electricity datasets; established deep learning as viable for multivariate TSF
    • DeepAR (Salinas et al. 2020): global cross-series training with probabilistic output; industry deployment benchmark
    • Informer (Zhou et al. 2021): ProbSparse attention for long-sequence Transformers; AAAI Best Paper — later found to be partially replicated by simpler baselines (DLinear)
    • N-BEATS (Oreshkin et al. 2020): pure MLP stack achieving SOTA without recurrence; resurrected MLP research credibility
    • DLinear (Zeng et al. 2023): provoked the “transformer effectiveness” debate; catalysed more rigorous benchmark practices
    • Chronos and TimesFM (both 2024): established foundation model zero-shot forecasting as competitive with supervised methods

Current Landscape (2026)

  • The TSF landscape in 2026 is defined by the maturation of foundation models, consolidation of probabilistic evaluation standards, and rising practical adoption of zero-shot deployment patterns.
  • Foundation Model Production Deployment: TimeGPT-1 (Nixtla/Azure), Chronos (Amazon Forecast), TimesFM (Google Cloud) provide commercial APIs enabling enterprises to deploy zero-shot forecasting without training data collection or model maintenance.
  • Transformer vs Linear Debate Resolved: The community reached consensus (2024-2025) that architecture complexity matters only when sufficient cross-series diversity and scale are available. Foundation models (vast pre-training) and properly designed transformers (PatchTST, iTransformer) outperform linear baselines at scale; per-series fitting still benefits from simpler methods.
  • Mamba Emergence: TimeMamba and S-Mamba demonstrate Mamba competitive with or better than iTransformer/PatchTST on several benchmarks whilst using 40-50% less GPU memory, positioning selective SSMs as the likely next dominant architecture for IoT/edge deployment.
  • Probabilistic Standard: CRPS has replaced MAPE as the community-standard evaluation metric following the M5 Uncertainty Track; most new model papers (2023+) report CRPS alongside point metrics.
  • Key 2025-2026 Industry Deployments:
    • NHS England Federated Analytics Platform: Chronos zero-shot across 100+ Trusts for system-wide demand planning without centralising patient data
    • National Grid ESO: TFT probabilistic load forecasting replacing linear regression baseline; 15% CRPS improvement in day-ahead grid balancing
    • Salesforce Einstein Analytics: Moirai-MoE for customer revenue forecasting at enterprise scale
    • Google Ads: TimesFM 2.0 for campaign performance prediction at 1-trillion-data-point scale
    • JP Morgan Chase: proprietary TSF ensemble for credit card spend forecasting and fraud velocity anomaly detection at transaction level
    • Airbus Supply Chain: N-HiTS-based 90-day demand forecasting for 80,000+ aircraft spare parts, replacing monthly ARIMA batch runs with daily automated reforecasting
    • Met Office: NeuralGCM integration with ML post-processing (quantile regression forests) for 10-day precipitation ensemble calibration; operational deployment alongside ECMWF ensemble
  • Open-Source Momentum: The NumFOCUS-sponsored sktime project passed 7,000 GitHub stars in 2025; Nixtla StatsForecast surpassed 4,000 stars; Darts crossed 8,000 stars. The combined open-source TSF ecosystem now covers 200+ models accessible through unified sklearn-compatible APIs, dramatically lowering the barrier to production forecasting.
  • Probabilistic Forecast Adoption: CRPS-based evaluation has expanded from research benchmarks into enterprise procurement; major forecasting vendors (SAP Integrated Business Planning, Oracle Demand Management Cloud, Blue Yonder) added probabilistic/quantile output capabilities in 2024-2025 releases, driven by demand from retailers and energy companies requiring calibrated inventory safety stock calculations.

UK Context

  • Imperial College London: Professor Niall Adams (Statistics) — Time Series Research Group; expertise in streaming data anomaly detection, sequential changepoint methods applied to cybersecurity and financial surveillance; Data Science Institute hosts industry-academic partnerships in financial time series.
  • University of Oxford Statistics Department: Professor Yee Whye Teh — Bayesian nonparametric methods including Gaussian process regression for irregularly sampled series; Professor Chris Holmes — probabilistic forecasting applied to clinical trial interim analysis and NHS demand modelling; Neural CDE work (Kidger et al. 2020) emerging from Oxford Numerical Methods group.
  • Lancaster University STOR-i CDT: Centre for Doctoral Training in Statistics and Operational Research with industry sponsorship — explicitly focuses on time series forecasting in partnership with JBA Risk Management (flood probability), Procter & Gamble (retail demand), Shell (energy demand), GCHQ (signals intelligence). Professor Idris Eckley — leader in changepoint detection methodology (PELT, CROPS algorithms in R changepoint package).
  • University of Manchester: Professor David Lowe’s group — ML forecasting for energy systems; Manchester Institute for Data Science and Artificial Intelligence (IDSAI) hosting projects in traffic flow and public health forecasting; sktime library development team (Tony Bagnall, Matthew Middlehurst).
  • University of Edinburgh: Professor Chris Williams (informatics) — Gaussian processes for time series; Centre for Statistics delivering cross-Scotland time series teaching via SMSTC; EPSRC-funded projects on probabilistic forecasting for renewable energy grids.
  • UCL: Professor Ricardo Silva — probabilistic time series and causal discovery; UCL AI Centre hosting NHS-partnered demand forecasting projects.
  • Alan Turing Institute: National centre hosting the sktime open-source library development team (Franz Kiraly, Tony Bagnall, Lovkesh Agar); flagship project on probabilistic forecasting for National Grid ESO; Data Study Groups engaging with NHS England, Network Rail, and National Grid time series challenges annually.
  • Northern England Industrial Deployment:
    • National Grid ESO Hinckley Point Operations Centre: TFT-based probabilistic load forecasting replacing linear regression, 15% CRPS improvement; ancillary services trading using probabilistic wind/solar forecasts
    • EDF Energy (West Burton, Heysham nuclear stations): ML-assisted generation scheduling via demand-generation balance forecasting
    • Yorkshire Water and United Utilities: LSTM-based water demand forecasting (drought planning, treatment works capacity) and sewer overflow prediction via sensor anomaly detection
    • Northern Rail: LSTM models for passenger demand forecasting on Liverpool-Manchester-Leeds corridor; rolling stock allocation optimisation
    • Manufacturing Technology Centre (MTC, Coventry): time series anomaly detection on factory floor CNC sensor data for predictive maintenance; Aerospace Technology Institute funded project with Rolls-Royce
    • JBA Risk Management (Skipton, North Yorkshire): probabilistic flood risk forecasting for the insurance industry using ensemble hydrological model time series output; Monte Carlo simulation over climate uncertainty combined with river flow LSTM surrogates to produce 1-in-200-year flood loss distributions
    • Electricity System Operator (NESO, from 2024): short-term demand and generation forecasting underpinning every balancing mechanism trade; probabilistic half-hourly load forecasts (Bayesian structural time series + gradient boosting hybrids) drive £2-5B annual balancing market; wind/solar aggregated generation forecasts for constraint management across the GB transmission network
    • Arup (Engineering Consultancy, UK-wide): LSTM-based water demand forecasting for smart meter analytics across UK water companies; ML anomaly detection on pipe pressure sensor networks for leakage detection; Manchester-based data science team delivering real-time forecasting infrastructure for utilities and urban transport authorities

Future Directions (2026-2030)

  • Multivariate Foundation Models: Current foundation models primarily handle univariate series; learning cross-variate dependencies at pre-training scale is the next frontier.
    • Moirai’s any-variate attention and TimesFM 2.0 multivariate capabilities represent early steps
    • 2027-2028 models expected to treat entire correlated datasets (e.g. all smart meters in a city) as a single multivariate context, enabling cross-variate dependency transfer zero-shot
  • Causal Time Series Forecasting: Integrating causal inference (do-calculus, instrumental variables, difference-in-differences) with deep forecasting for intervention-robust predictions.
    • Critical for policy evaluation: what would hospital admissions be if vaccination coverage increased by 10%?
    • DeepIV (Hartford et al. 2017), CausalFoundation (2025) represent early efforts; active Lancaster/Edinburgh/UCL research agenda
  • Federated Forecasting: Training foundation models across distributed data silos without centralising sensitive data.
    • Essential for NHS multi-Trust deployment, cross-bank financial stability monitoring, multi-retailer supply chain
    • Alan Turing Institute Federated Analytics Programme targeting NHS deployment 2026-2028
  • Continuous-Time Foundation Models: Neural CDEs and Neural ODE-based foundation models handling irregular sampling natively.
    • Critical for healthcare wearable data (Apple Watch continuous monitoring), financial tick data (microsecond resolution), IoT sensor streams with dropout
    • Kidger’s Neural CDE framework (Oxford 2020) expected to underpin 2026-2027 clinical time series foundation models
  • Uncertainty-Aware Hierarchical Reconciliation: Reconciling full distributional forecasts (not just point forecasts) across hierarchy levels.
    • Ensuring coherence of prediction intervals at all aggregation levels whilst accounting for cross-hierarchy correlation
    • Active research at Monash University (Panagiotelis, Hyndman group) and Lancaster STOR-i
  • Mamba for Very Long Contexts: Linear-time complexity positions Mamba as architecture of choice for extremely long context windows (10,000-100,000 steps).
    • Relevant for minute-resolution electricity grid data (525,600 points/year), tick-level financial data, continuous wearable sensor streams
    • Quadratic Transformer attention is computationally infeasible at these scales; Mamba enables 100× longer lookback
  • Physics-Informed Hybrid Forecasting: Embedding domain ODE/PDE structure as inductive biases in neural forecasters.
    • Weather prediction (ECMWF NeuralGCM, DeepMind GraphCast), epidemiological SIR compartments, power system AC optimal flow physics
    • Improves extrapolation beyond training distribution; crucial for climate scenario analysis and pandemic response
    • Met Office / ECMWF-DeepMind collaboration active 2025-2028; UK Research and Innovation (UKRI) AI for Weather Programme
  • Explainable Forecasting: As foundation-model forecasts enter regulated domains (credit scoring, clinical decision support, grid dispatch), regulatory pressure (EU AI Act Articles 13-14) mandates human-interpretable explanations.
    • TFT attention maps provide built-in variable importance; post-hoc SHAP (Shapley Additive exPlanations) applied to model outputs; counterfactual explanations (“what input change would have produced a different forecast?“)
    • Research into forecast confidence attribution — identifying which historical observations most contribute to each forecast step
    • UK FCA (Financial Conduct Authority) guidance on AI model explainability in financial forecasting requiring documented methodology and audit trails
  • Real-Time Streaming Forecasting: Production systems increasingly require sub-second reforecasting as new observations arrive (e.g. electricity grid imbalance, fraud detection, trading systems).
    • Online learning variants of ARIMA (recursive least squares updates, forgetting factors) enable O(1) per-step update
    • Warm-starting deep learning models: fine-tune only final layers on recent data; Moirai and Chronos provide contextual window inference without retraining
    • Apache Kafka + Flink pipeline integration; AWS Kinesis Data Analytics for streaming time series at 100K+ events per second

Research and Literature

  • Foundational Books
    • Box, G.E.P., Jenkins, G.M., Reinsel, G.C., Ljung, G.M. (2015). “Time Series Analysis: Forecasting and Control.” 5th ed. Wiley. Definitive ARIMA reference.
    • Hyndman, R.J. and Athanasopoulos, G. (2021). “Forecasting: Principles and Practice.” 3rd ed. OTexts (open-access). Primary reference for ETS, hierarchical forecasting, fabletools; used at Lancaster STOR-i, Imperial, Manchester.
    • Brockwell, P.J. and Davis, R.A. (2002). “Introduction to Time Series and Forecasting.” 2nd ed. Springer. Graduate-level statistical theory.
    • Hamilton, J.D. (1994). “Time Series Analysis.” Princeton University Press. Econometrics perspective: ARIMA, cointegration, GARCH, VAR.
    • Durbin, J. and Koopman, S.J. (2012). “Time Series Analysis by State Space Methods.” 2nd ed. Oxford. Kalman filter; structural time series.
    • Shumway, R.H. and Stoffer, D.S. (2017). “Time Series Analysis and Its Applications.” 4th ed. Springer (open-access).
  • Key Academic Papers
    • Salinas, D. et al. (2020). “DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks.” International Journal of Forecasting 36(3):1181-1191. Amazon Research.
    • Oreshkin, B.N. et al. (2020). “N-BEATS: Neural basis expansion analysis for interpretable time series forecasting.” ICLR 2020.
    • Ansari, A.F. et al. (2024). “Chronos: Learning the Language of Time Series.” TMLR 2024. Amazon.
    • Das, A. et al. (2024). “A decoder-only foundation model for time-series forecasting.” ICML 2024. Google DeepMind.
    • Garza, A. and Mergenthaler-Canseco, M. (2023/2024). “TimeGPT-1.” arXiv:2310.03589. Nixtla.
    • Woo, G. et al. (2024). “Moirai: Unified Training of Universal Time Series Forecasting Transformers.” ICML 2024. Salesforce.
    • Rasul, K. et al. (2024). “Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting.” arXiv:2310.08278.
    • Lim, B. et al. (2021). “Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting.” International Journal of Forecasting 37(4):1748-1764.
    • Gu, A. et al. (2022). “Efficiently Modeling Long Sequences with Structured State Spaces (S4).” ICLR 2022.
    • Gu, A. and Dao, T. (2023). “Mamba: Linear-Time Sequence Modeling with Selective State Spaces.” arXiv:2312.00752.
    • Challu, C. et al. (2023). “N-HiTS: Neural Hierarchical Interpolation for Time Series Forecasting.” AAAI 2023.
    • Nie, Y. et al. (2023). “A Time Series is Worth 64 Words: Long-term Forecasting with Transformers.” ICLR 2023. (PatchTST)
    • Liu, Y. et al. (2024). “iTransformer: Inverted Transformers Are Effective for Time Series Forecasting.” ICLR 2024.
    • Zeng, A. et al. (2023). “Are Transformers Effective for Time Series Forecasting?” AAAI 2023. (DLinear)
    • Wickramasuriya, S. et al. (2019). “Optimal Forecast Reconciliation Using a Unifying Framework.” JBES 37(2):225-246. (MinT)
    • Makridakis, S. et al. (2020). “The M4 Competition: 100,000 time series and 61 forecasting methods.” IJF 36(1):54-74.
    • Makridakis, S. et al. (2022). “M5 accuracy competition: Results, findings, and conclusions.” IJF 38(4):1346-1364.
    • Smyl, S. (2020). “A hybrid method of exponential smoothing and recurrent neural networks for time series forecasting.” IJF 36(1):75-85. (M4 Winner)
    • Gneiting, T. and Raftery, A.E. (2007). “Strictly Proper Scoring Rules, Prediction, and Estimation.” JASA 102(477):359-378. (CRPS theory)
    • Taylor, S.J. and Letham, B. (2018). “Forecasting at Scale.” The American Statistician 72(1):37-45. (Prophet)
    • Chen, R.T.Q. et al. (2018). “Neural Ordinary Differential Equations.” NeurIPS 2018.
    • Kidger, P. et al. (2020). “Neural Controlled Differential Equations for Irregular Time Series.” NeurIPS 2020. Oxford.
    • Zhou, H. et al. (2021). “Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting.” AAAI 2021 Best Paper.
    • Zhou, J. et al. (2022). “Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy.” ICLR 2022.
    • Lai, G. et al. (2018). “Modeling Long- and Short-Term Temporal Patterns with Deep Neural Networks (LSTNet).” SIGIR 2018.
    • Assimakopoulos, V. and Nikolopoulos, K. (2000). “The theta model: a decomposition approach to forecasting.” IJF 16(4):521-530.

Production Software Ecosystem

  • Nixtla StatsForecast: CPU-optimised classical models; benchmarked at 1M series per second on a single machine; includes AutoARIMA, AutoETS, AutoCES (Complex Exponential Smoothing), MSTL (Multiple Seasonal-Trend decomposition using LOESS), Theta, TBATS; unified pandas-compatible API; from statsforecast.models import AutoARIMA, AutoETS.
  • Nixtla NeuralForecast: GPU-accelerated deep learning; includes N-BEATS, N-HiTS, TFT, PatchTST, iTransformer, NHITS, TSMixer; same API as StatsForecast; distributed training via Ray; conformal prediction intervals via standard conformal_interval parameter.
  • Amazon GluonTS: Research standard for probabilistic forecasting; DeepAR reference implementation; GaussianProcess, N-BEATS, Transformer; dataset loaders for M3/M4/M5/electricity/traffic benchmarks; MXNet and PyTorch backends.
  • PyTorch Forecasting: Production-grade PyTorch Lightning wrapper; TFT, N-BEATS, DeepAR, NHiTS, RecurrentNetwork; integrated hyperparameter tuning via Optuna; TemporalFusionTransformer.from_dataset(dataset, learning_rate=0.03).
  • Darts (Unit8): Unified deterministic and probabilistic forecasting; 30+ models including ARIMA, Prophet, LSTM, TFT, Chronos, Lag-Llama; sklearn-compatible .fit() / .predict() API; backtest() utility for walk-forward evaluation.
  • sktime (Alan Turing Institute): Widest sklearn-compatible interface for 200+ forecasters; integrates statsmodels, pmdarima, prophet, deep learning backends; pipeline composition: ForecastingPipeline([('imputer', Imputer()), ('forecaster', AutoARIMA())]).
  • R forecast / fable / fabletools: Rob Hyndman’s R packages implementing ARIMA, ETS, hierarchical reconciliation; fabletools provides tidy (tidyverts) interface; model(ARIMA(value), ETS(value)) |> forecast(h=12).
  • R hts / thief: Hierarchical time series reconciliation; MinT, top-down, bottom-up, middle-out; temporal hierarchical forecasting (THIEF) reconciles across aggregation levels simultaneously.
  • TSlib (Tsinghua / Wu et al. 2023): The de-facto deep learning benchmark repository; contains reference implementations of PatchTST, iTransformer, TimesNet, Autoformer, FEDformer, Informer, DLinear, Crossformer; standardised datasets (ETT-h1/h2/m1/m2, Weather, Electricity, Traffic, Exchange); evaluation scripts producing directly comparable RMSE/MAE numbers.
  • Monash Time Series Forecasting Archive (Godahewa et al. 2021): 30+ diverse datasets (M3/M4, NN5, Tourism, Electricity, Solar, Rideshare, Hospital, Fred-MD, Dominick, Favorita) with standardised splits; enables fair cross-benchmark evaluation of new algorithms.

Evaluation Benchmarks and Competitions

  • M-Competition Series (Makridakis et al.):
    • M1 (1982): 1001 series across annual/quarterly/monthly/other frequencies; Box-Jenkins complex methods not superior to simple alternatives
    • M2 (1993): 29 companies providing real business forecasting problems with practitioner involvement; Theta and regression methods performed well
    • M3 (2000): 3003 series; Theta won; triggered research into combination forecasting and ensemble methods; M3 results still used as baseline comparison in 2026
    • M4 (2018): 100,000 series; 61 competing methods; Smyl’s ES-RNN (Uber) won overall SMAPE; hybrid statistical+ML approaches dominated top-10; pure deep learning not in top-3
    • M5 Accuracy (2020/2022): 42,840 Walmart item-store daily series; 909 competing teams; LightGBM with feature engineering won both M5 Accuracy and M5 Uncertainty; N-BEATS at weekly aggregation competitive; top ensemble combined LightGBM + LSTM + TFT
    • M6 Financial (2024): 50 monthly return series; novel joint metric: 50% accuracy (MASE) + 50% portfolio performance (Sharpe ratio); winning teams combined statistical forecasting with portfolio optimisation (Black-Litterman, CVaR minimisation)
  • Long-Horizon Benchmarks (ETT, Weather, ECL, Traffic):
    • ETT-h1/h2 (Electricity Transformer Temperature hourly): 17,420 training observations; predict 96/192/336/720 steps ahead; most widely used deep learning TSF benchmark since Informer (2021)
    • Weather (Max Planck Institute meteorological): 52,696 observations of 21 weather indicators; 96-720 step prediction
    • ECL (Electricity Consuming Load): 26,304 hourly observations from 321 Portuguese clients; 96-720 step prediction
    • METR-LA / PEMS-BAY: Road sensor traffic speed; 207/325 sensors; spatial graph topology available; benchmarks for graph-based spatio-temporal methods
  • Monash Time Series Forecasting Archive: 30+ standardised datasets across diverse domains; used for zero-shot evaluation of foundation models (Chronos, TimesFM, Moirai); enables standardised CRPS comparison across methods

Metadata

  • Enrichment date: 2026-05-17T10:00:00Z
  • Worker model: claude-sonnet-4-6
  • Domain corrected: No domain change required (was correctly artificial-intelligence)
  • Legacy term ID: AI-1087 (assigned during enrichment)
  • Source stub lines: 43
  • Research notes cached: _enrich/research-cache/Time Series Forecasting.json

Provenance

  • Salinas, D., Flunkert, V., Gasthaus, J., and Januschowski, T. (2020). “DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks.” International Journal of Forecasting 36(3):1181-1191.
  • Oreshkin, B.N., Carpov, D., Chapados, N., and Bengio, Y. (2020). “N-BEATS: Neural basis expansion analysis for interpretable time series forecasting.” ICLR 2020. Element AI / Mila.
  • Ansari, A.F. et al. (2024). “Chronos: Learning the Language of Time Series.” TMLR 2024. Amazon.
  • Das, A., Kong, W., Sen, R., and Zhou, Y. (2024). “A decoder-only foundation model for time-series forecasting.” ICML 2024. Google DeepMind.
  • Garza, A. and Mergenthaler-Canseco, M. (2023). “TimeGPT-1.” arXiv:2310.03589. Nixtla.
  • Woo, G. et al. (2024). “Moirai: Unified Training of Universal Time Series Forecasting Transformers.” ICML 2024. Salesforce.
  • Rasul, K. et al. (2024). “Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting.” arXiv:2310.08278. ServiceNow / Mila.
  • Lim, B., Arik, S.O., Loeff, N., and Pfister, T. (2021). “Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting.” International Journal of Forecasting 37(4):1748-1764.
  • Gu, A., Goel, K., and Re, C. (2022). “Efficiently Modeling Long Sequences with Structured State Spaces.” ICLR 2022.
  • Gu, A. and Dao, T. (2023). “Mamba: Linear-Time Sequence Modeling with Selective State Spaces.” arXiv:2312.00752.
  • Challu, C. et al. (2023). “N-HiTS: Neural Hierarchical Interpolation for Time Series Forecasting.” AAAI 2023.
  • Nie, Y. et al. (2023). “A Time Series is Worth 64 Words.” ICLR 2023. (PatchTST)
  • Liu, Y. et al. (2024). “iTransformer: Inverted Transformers Are Effective for Time Series Forecasting.” ICLR 2024.
  • Zeng, A. et al. (2023). “Are Transformers Effective for Time Series Forecasting?” AAAI 2023. (DLinear)
  • Wickramasuriya, S., Athanasopoulos, G., and Hyndman, R.J. (2019). “Optimal Forecast Reconciliation.” Journal of Business and Economic Statistics 37(2):225-246.
  • Makridakis, S. et al. (2020). “The M4 Competition.” International Journal of Forecasting 36(1):54-74.
  • Makridakis, S. et al. (2022). “M5 accuracy competition.” International Journal of Forecasting 38(4):1346-1364.
  • Smyl, S. (2020). “A hybrid method of exponential smoothing and recurrent neural networks.” IJF 36(1):75-85.
  • Hyndman, R.J. and Athanasopoulos, G. (2021). “Forecasting: Principles and Practice.” 3rd ed. OTexts.
  • Gneiting, T. and Raftery, A.E. (2007). “Strictly Proper Scoring Rules.” JASA 102(477):359-378.
  • Taylor, S.J. and Letham, B. (2018). “Forecasting at Scale.” The American Statistician 72(1):37-45.
  • Chen, R.T.Q. et al. (2018). “Neural Ordinary Differential Equations.” NeurIPS 2018.
  • Kidger, P. et al. (2020). “Neural Controlled Differential Equations for Irregular Time Series.” NeurIPS 2020. Oxford.
  • Zhou, H. et al. (2021). “Informer.” AAAI 2021.
  • Zhou, J. et al. (2022). “Anomaly Transformer.” ICLR 2022.
  • Lai, G. et al. (2018). “LSTNet.” SIGIR 2018.
  • Assimakopoulos, V. and Nikolopoulos, K. (2000). “The theta model.” IJF 16(4):521-530.