Mechanistic interpretability is a research area that seeks to reverse-engineer the internal computations of neural networks into human-understandable algorithms. It studies features, circuits, and representations within model weights and activations to explain how specific behaviours arise. The field aims to make models transparent enough to predict, audit, and align, supporting AI safety.
Content
- Techniques include activation patching, sparse autoencoders for feature disentanglement, and circuit analysis that traces how attention heads and MLP layers compose to implement a task. The goal is faithful, causal explanations rather than post-hoc rationalisations, enabling auditing of deception, capability, and failure modes in frontier models.