A training paradigm in which a single differentiable model is optimised to map raw inputs directly to final task outputs — pixels to steering angles, waveforms to transcripts, text to text — with all intermediate representations learned jointly by gradient descent rather than specified as hand-engineered features or separately built pipeline stages; it trades the modularity, interpretability, and testability of engineered pipelines for the ability to discover representations that human designers would not, and dominates wherever data and compute are abundant.

Semantic Classification

Content

Definition

End-to-end learning trains one differentiable system to perform an entire task, from raw sensory input to final output, optimising every internal representation against the task loss. The alternative it displaced is the staged pipeline: hand-crafted Feature Engineering followed by a shallow classifier, or, in robotics, a chain of separately engineered modules — perception, localisation, planning, control — each with its own specification, interface, and test suite.

The paradigm’s case is empirical: given enough data, jointly learned intermediate representations consistently beat hand-designed ones, because the optimiser co-adapts every stage to the final objective and exploits regularities designers never articulated. Speech recognition abandoned phoneme pipelines for sequence-to-sequence models; machine translation abandoned alignment-and-phrase-table systems; computer vision abandoned SIFT and HOG for learned convolutional features. NVIDIA’s PilotNet (2016) made the paradigm vivid in driving — a network mapping camera pixels directly to steering commands — and modern large language models are the paradigm at its extreme, a single network subsuming what were once dozens of NLP components.

The costs are equally real, and they explain why safety-critical autonomy still often prefers a modular Perception System feeding explicit planning. End-to-end systems are hard to verify compositionally: there is no intermediate interface at which to write a specification, failures are diagnosed by data archaeology rather than unit tests, credit assignment for errors is opaque, and the system may exploit spurious shortcuts in the training distribution. The current frontier blends the two views — differentiable modular architectures, auxiliary intermediate losses, and interpretable bottlenecks — seeking end-to-end optimisation without monolithic opacity.

Current Landscape

  • Autonomous driving: the field’s live controversy, and the paradigm has hardened into product. Tesla’s FSD transitioned to a true single-network end-to-end stack at v12 — replacing “thousands of lines of rule-based C++” with one neural net — and extended it in v13 (higher-resolution video, ~36 Hz temporal sampling) and v14 (multi-second temporal reasoning, audio awareness, a possible mixture-of-models), the model used in the 2025 Austin robotaxi trials.

  • Convergence toward hybrids: a 2026 survey (“The Era of End-to-End Autonomy”) frames “supervised E2E driving” (FSD Supervised / L2++) as the emerging category several manufacturers plan to ship from 2026. Waymo’s December 2025 Waymo Foundation Model is explicitly neither pure end-to-end nor modular: it supports full end-to-end backpropagation while materialising structured representations (objects, roadgraph) for inference-time safety validation, exactly the “interpretable bottleneck” middle ground.

  • Foundation models: large language and multimodal models are end-to-end learning at maximal scale, with Representation Learning emerging implicitly from a single objective; instruction tuning and RLHF extend the differentiable path through behavioural alignment.

  • Robotics: end-to-end visuomotor policies (RT-2, diffusion policies, ALOHA-style imitation) map camera input to actuation, challenging the classical sense–plan–act decomposition catalogued in robotics core concepts.

  • Engineering practice: data curation has replaced feature design as the main human lever — the effort formerly spent on features now goes into datasets, augmentation, and evaluation harnesses, alongside auxiliary losses and probing to recover some of the lost inspectability.

    Sources:

  • https://arxiv.org/html/2603.16050

  • https://waymo.com/blog/2025/12/demonstrably-safe-ai-for-autonomous-driving/