Reading Group
Introduction
This is a small reading group working through foundational and recent papers in mechanistic interpretability; the effort to reverse-engineer the internal computations of neural networks into human-understandable algorithms. Each session centers on one paper, presented in depth, with an emphasis on building intuition for the methods (activation patching, circuit analysis, causal tracing) as well as the specific findings.
Sessions are sequenced deliberately: earlier papers introduce the tools and vocabulary (attention head taxonomy, patching, attribution) that later papers depend on, building toward more complex circuit-discovery techniques and current work on superposition and sparse autoencoders.
We expect to continue beyond the papers currently listed below – this is simply what’s been planned so far.
Useful Links
- Reading Group Calendar
- Presenter Form (fill this out if you’re interested in presenting)
- Paper Suggestion Form
Schedule
Sessions are currently held every Wednesday from 7:30 to 8:30 p.m. Pacific Time. To register, please visit the Reading Group Calendar.
This schedule is tentative and subject to change.