How we contain Claude across productskept by eddie • may 27Anthropic details how it contains AI agent blast radius across three products using sandboxes, VMs, and human-in-the-loop controls, while documenting missed risks like pre-trust code execution and user-mediated prompt injection.AboutAnthropic · gVisor · Claude · Claude Code · Claude Cowork · Model Context ProtocolFiled#ai-agents#ai-safety#developer-tools#security#software-engineeringRelatedGitHub - anthropics/defending-code-reference-harness: Skills for threat modeling, scanning, triage, patching, plus an autonomous scanning harness you can /customizeAnthropic released an open-source reference harness and Claude Code skills for automating vulnerability discovery, triage, and patching using Claude, alongside a managed product called Claude Security.also on gVisor, Claude, Claude Code, Anthropic, #developer-toolsA field guide to sandboxes for AISandboxing tools for AI agents differ fundamentally by boundary (containers, gVisor, microVMs, Wasm), not just by policy, and picking the wrong one creates either leaks or excessive cost.also on gVisor, #security, #ai-agents, #developer-toolsClaude Cowork exfiltrates filesResearchers demonstrate file exfiltration from Claude Cowork via prompt injection, exploiting an acknowledged but unremediated vulnerability.also on Claude Cowork, #ai-safety, #security, Anthropic, #ai-agentsFirst impressions of Claude Cowork, Anthropic’s general agentClaude Cowork wraps Claude Code in a sandboxed UI for non-developers, offering general agent capabilities but raising prompt injection concerns.also on Claude Cowork, #security, Claude Code, Anthropic, #ai-agents, #developer-toolsClaude Code and What Comes Next - by Ethan MollickClaude Code, using agentic harness and compaction, autonomously built a working website and game, showing a major leap in AI capability for sustained work.also on Model Context Protocol, Claude Code, Anthropic, #ai-agents, #software-engineering, #developer-tools