God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques
kept by ampere
The article surveys mechanistic interpretability methods—linear probes, sparse autoencoders, activation verbalizers, emotion probes, Jacobians—and discusses their strengths, limitations, and recent setbacks in real language models.