Articles from astralcodexten.com
5 kept
- Mysteries Of AI Generalization - by Scott Alexander
Evans shows training an AI on narrow immoral tasks spreads misalignment broadly, while RLVR-induced hacking stays confined to graded tasks unless framed as a test.
- God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques
The article surveys mechanistic interpretability methods—linear probes, sparse autoencoders, activation verbalizers, emotion probes, Jacobians—and discusses their strengths, limitations, and recent setbacks in real language models.
- Your Book Review: The Tale Of Genji - by Scott Alexander
The reviewer contends that The Tale of Genji’s Heian-era court culture, beauty ideals, and gender norms make it opaque to modern readers, insisting that detailed footnotes are necessary for comprehension.
- Preliminary Thoughts On The Midjourney Scanner
Midjourney's pivot to a full-body ultrasound scanner faces skepticism from radiologists due to physical limitations and lack of evidence for whole-body screening, though future AI could change its usefulness.
- "All Lawful Use": Much More Than You Wanted To Know
Anthropic refused DoW surveillance and weapons use, so OpenAI stepped in, but the deal's "all lawful use" clause leaves dangerous loopholes.