Apache Spark is an open-source unified analytics engine for large-scale data processing across clusters of machines. It exposes high-level APIs for batch processing, structured queries, stream processing and machine learning, and accelerates workloads by keeping intermediate data in memory between operations. Spark abstracts distributed datasets as fault-tolerant collections and schedules computations as directed acyclic graphs of stages, making it a foundational tool for big-data engineering and analytics.

Overview

  • Spark provides a single engine and a coherent set of APIs for the major shapes of data work: batch transformations, SQL-style structured queries, near-real-time streaming and distributed machine learning.
  • Its performance advantage comes largely from keeping intermediate results in memory across operations rather than writing them to disk between every step, which suits iterative algorithms and interactive analytics.
  • Computations are represented as directed acyclic graphs of transformations on resilient distributed datasets, allowing the scheduler to optimise execution and recover lost partitions by recomputation rather than replication.

Key aspects

  • Resilient distributed datasets and DataFrames model partitioned, fault-tolerant collections spread across a cluster.
  • A lazy execution model builds a logical plan that is optimised before any computation runs.
  • In-memory caching accelerates iterative and interactive workloads.
  • Built-in libraries cover SQL, structured streaming, graph processing and machine learning.

Applications

  • Large-scale ETL and data-pipeline construction over data lakes and warehouses.
  • Interactive analytics and ad-hoc querying across very large datasets.
  • Near-real-time stream processing for event and log analytics.
  • Distributed feature engineering and model training for machine learning.

Provenance