Reading up on Benjamin Sturgeon
1 deep · digging since sep 09
- God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques
The article surveys mechanistic interpretability methods—linear probes, sparse autoencoders, activation verbalizers, emotion probes, Jacobians—and discusses their strengths, limitations, and recent setbacks in real language models.