Reading Group

A paper-by-paper exploration of mechanistic interpretability, from foundational methods to modern circuit discovery.

Introduction

This is a small reading group working through foundational and recent papers in mechanistic interpretability; the effort to reverse-engineer the internal computations of neural networks into human-understandable algorithms. Each session centers on one paper, presented in depth, with an emphasis on building intuition for the methods (activation patching, circuit analysis, causal tracing) as well as the specific findings.

Sessions are sequenced deliberately: earlier papers introduce the tools and vocabulary (attention head taxonomy, patching, attribution) that later papers depend on, building toward more complex circuit-discovery techniques and current work on superposition and sparse autoencoders.

We expect to continue beyond the papers currently listed below – this is simply what’s been planned so far.

Schedule

Sessions are currently held every Wednesday from 7:30 to 8:30 p.m. Pacific Time. To register, please visit the Reading Group Calendar.

This schedule is tentative and subject to change.

Session Number Date Paper Slides Interactive Visualizations
1 Jun 10, 2026 How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model Coming Soon Coming Soon
2 Jun 17, 2026 In-context Learning and Induction Heads Coming Soon Coming Soon
3 Jun 24, 2026 Axiomatic Attribution for Deep Networks Coming Soon Coming Soon
4 Jul 1, 2026 Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 small Coming Soon Coming Soon
5 Jul 8, 2026 A Circuit for Python Docstrings in a 4-Layer Attention-Only Transformer Coming Soon Coming Soon
Jul 15, 2026 Cancelled — no session
6 Jul 22, 2026 A Mathematical Framework for Transformer Circuits Coming Soon Coming Soon
7 Jul 29, 2026 Towards Automated Circuit Discovery for Mechanistic Interpretability Coming Soon Coming Soon
8 Aug 5, 2026 Attribution Patching: Activation Patching At Industrial Scale Coming Soon Coming Soon
9 Aug 12, 2026 Attribution Patching Outperforms Automated Circuit Discovery Coming Soon Coming Soon
10 Aug 19, 2026 Toward Monosemanticity: Decomposing Language Models With Dictionary Learning Coming Soon Coming Soon