A model training pipeline is the orchestrated sequence of stages that transforms raw data into a trained, validated machine-learning model ready for deployment. It typically chains data ingestion and preprocessing, feature engineering, model fitting, hyperparameter tuning and evaluation into a reproducible, automatable workflow. As a backbone of MLOps, the pipeline enforces consistency, versioning and repeatability across training runs.
Overview
- The pipeline turns datasets into validated models through a repeatable, versioned sequence of steps.
- Automation lets teams retrain on new data and reproduce results deterministically.
- It connects upstream data pipelines to downstream model deployment and monitoring.
Key aspects
- Data ingestion, validation and preprocessing stages.
- Feature engineering and transformation steps.
- Training loops optimising a loss function with tuned hyperparameters.
- Evaluation, model selection and artefact versioning.
Applications
- Continuous retraining in production ML systems.
- Experiment tracking and reproducible research.
- Automated model promotion and A/B evaluation.