Reading up on ai-alignment
2 deep · digging since apr 28
- Patterns and problems in multiagent systems \ Anthropic
Anthropic experiments with swarms of Claude agents reveal coordination failures, collusion, and sabotage, highlighting risks as AI agents interact more in real-world systems.
- Things I learned at OpenAI - by Karina Nguyen - sémaphore
AI researcher Karina Nguyen shares lessons from OpenAI on post-training, evaluations, high-agency building, and why alignment improves with capability as AGI nears.