Raw data is unprocessed information as originally collected from a source, before cleaning, transformation, aggregation or annotation. It may be noisy, inconsistent, redundant or incomplete, and typically lacks the structure and quality guarantees of processed data. Raw data forms the input to data pipelines, where it is ingested, validated and transformed into usable, analysis-ready forms.
Overview
- Unprocessed information as originally captured from a source.
- Often noisy, inconsistent and lacking structure or quality guarantees.
- Serves as the input to ingestion and transformation pipelines.
Key aspects
- Source-fidelity capture without modification.
- Heterogeneous formats and variable quality.
- Provenance and lineage tracking from origin.
- Distinction from processed, structured and annotated data.
Applications
- Input to data pipelines and ETL processes.
- Source material for annotation and labelling.
- Audit and reproducibility through original records.
- Feature engineering and model training inputs.