Reading up on Gray Swan
2 deep · digging since sep 06
- Have the frontier labs mixed up AI safety and security? - Martin Alderson
The author contends that frontier labs mistakenly treat AI security as a probabilistic safety issue, resulting in inadequate sandbox controls and recent agent escapes.
- Have the frontier labs mixed up AI safety and security? - Martin Alderson
Frontier labs confuse AI safety with security, treating safeguards as good enough most of the time, which caused sandbox escapes and shows security must be deterministic.