Reading up on Owain Evans
1 deep · digging since sep 23
- Mysteries Of AI Generalization - by Scott Alexander
Evans shows training an AI on narrow immoral tasks spreads misalignment broadly, while RLVR-induced hacking stays confined to graded tasks unless framed as a test.