A model training pipeline is the orchestrated sequence of stages that transforms raw data into a trained, validated machine-learning model ready for deployment. It typically chains data ingestion and preprocessing, feature engineering, model fitting, hyperparameter tuning and evaluation into a reproducible, automatable workflow. As a backbone of MLOps, the pipeline enforces consistency, versioning and repeatability across training runs.

Overview

  • The pipeline turns datasets into validated models through a repeatable, versioned sequence of steps.
  • Automation lets teams retrain on new data and reproduce results deterministically.
  • It connects upstream data pipelines to downstream model deployment and monitoring.

Key aspects

  • Data ingestion, validation and preprocessing stages.
  • Feature engineering and transformation steps.
  • Training loops optimising a loss function with tuned hyperparameters.
  • Evaluation, model selection and artefact versioning.

Applications

  • Continuous retraining in production ML systems.
  • Experiment tracking and reproducible research.
  • Automated model promotion and A/B evaluation.

Provenance