Apache Spark is an open-source unified analytics engine for large-scale data processing across clusters of machines. It exposes high-level APIs for batch processing, structured queries, stream processing and machine learning, and accelerates workloads by keeping intermediate data in memory between operations. Spark abstracts distributed datasets as fault-tolerant collections and schedules computations as directed acyclic graphs of stages, making it a foundational tool for big-data engineering and analytics.
Overview
- Spark provides a single engine and a coherent set of APIs for the major shapes of data work: batch transformations, SQL-style structured queries, near-real-time streaming and distributed machine learning.
- Its performance advantage comes largely from keeping intermediate results in memory across operations rather than writing them to disk between every step, which suits iterative algorithms and interactive analytics.
- Computations are represented as directed acyclic graphs of transformations on resilient distributed datasets, allowing the scheduler to optimise execution and recover lost partitions by recomputation rather than replication.
Key aspects
- Resilient distributed datasets and DataFrames model partitioned, fault-tolerant collections spread across a cluster.
- A lazy execution model builds a logical plan that is optimised before any computation runs.
- In-memory caching accelerates iterative and interactive workloads.
- Built-in libraries cover SQL, structured streaming, graph processing and machine learning.
Applications
- Large-scale ETL and data-pipeline construction over data lakes and warehouses.
- Interactive analytics and ad-hoc querying across very large datasets.
- Near-real-time stream processing for event and log analytics.
- Distributed feature engineering and model training for machine learning.