An informal, evolving glossary of mechanistic interpretability terms, methods, and metrics.
Why each attention head and MLP neuron writes its own additive term into the transformer residual stream, following Anthropic’s Mathematical Framework for Transformer Circuits.
What I mean by hard things, what I don’t, why the neuroscience behind willpower makes a compelling case for doing them anyway, and how that played out over 100 days.
A GPT-2–style walkthrough of causal self-attention—from a small worked example through single-head and multi-head formulations and their computational graphs.
An exploration of Sparse Autoencoders as a tool for decomposing polysemantic neural network representations into interpretable, monosemantic features.