MI notes¶
Mechanistic interpretability: reading what a model is doing from its internals rather than its outputs. Everything else I write lives on my personal blog.
Concepts¶
The ideas the field is built on: a running primer on features, circuits and why the mechanism is worth reading at all.