Big data denotes datasets whose volume, velocity and variety exceed the capacity of conventional single-machine tools, demanding distributed storage and parallel computation. It is characterised by horizontally scalable architectures, schema-flexible stores, and batch or streaming processing frameworks that move computation to where data resides. The term also names the discipline of extracting value from such datasets through analytics, mining and machine learning at scale.

Overview

  • Big data emerged as digitisation produced data faster than storage and compute could keep pace using vertical scaling alone. The response was to partition data across clusters and bring computation to the data.
  • The defining “three Vs” — volume, velocity, variety — are often extended with veracity and value, capturing the uncertainty of raw inputs and the need to justify analytical effort.
  • Architecturally, big data systems favour shared-nothing clusters, replicated distributed storage, and processing engines that tolerate node failure transparently.
  • The field reshaped data engineering practice: pipelines became distributed, storage moved towards lakes and lakehouses, and elastic cloud resources replaced fixed on-premise capacity for many workloads.

Key aspects

  • Volume: petabyte-scale collections that exceed the addressable storage of any single server.
  • Velocity: high-throughput, low-latency streams requiring continuous ingestion and processing.
  • Variety: structured, semi-structured and unstructured sources combined within one system.
  • Distribution: data and computation partitioned across many nodes for parallelism and fault tolerance.
  • Storage layers: distributed file systems, object stores, NoSQL databases, data lakes and warehouses.
  • Processing models: batch frameworks for throughput and stream frameworks for timeliness, often unified.

Applications

  • Powering recommendation, search and personalisation across web-scale platforms.
  • Training and serving machine-learning models on very large corpora.
  • Real-time monitoring of telemetry, logs and IoT sensor streams.
  • Fraud detection, risk modelling and predictive analytics in finance.
  • Genomics, climate modelling and other compute-intensive scientific analyses.
  • Consolidating enterprise data into lakes and warehouses for organisation-wide analytics.

Provenance