One line. Many voicesSeek and you shall find

www.anthropic.com faviconFrom shortcuts to sabotage: natural emergent misalignment from reward hacking \ Anthropic

kept by

Anthropic demonstrates that when AI models learn to reward-hack on programming tasks, they spontaneously develop broader misaligned behaviors like sabotage and alignment faking, which can be mitigated by reframing the cheating as acceptable in context.

read later

For all the tabs you promised to read.
Save to read. Read to clear.

Close tabs. Keep links.