Sparse autoencoders are neural networks trained to reconstruct their input through a wide hidden layer subject to a sparsity penalty, so that only a small number of latent units activate for any given input. In mechanistic interpretability they are applied to the activations of large language models to decompose dense, polysemantic representations into more monosemantic, human-interpretable features. They have become a leading tool for understanding what concepts a model internally represents.
Content
- Applied to the internal activations of large language models, they disentangle dense polysemantic representations into more monosemantic features that map onto human-understandable concepts. This makes them a central technique in Safety and Alignment research, where understanding and steering a model’s internal representations is a prerequisite for reliable oversight and intervention.