Skip to content

MI notes

Mechanistic interpretability: reading what a model is doing from its internals rather than its outputs. Everything else I write lives on my personal blog.

Concepts

The ideas the field is built on: a running primer on features, circuits and why the mechanism is worth reading at all.