Tag
This article discusses the security risks of running AI agents with tool execution on a single server and proposes a two-tier architecture that separates prompt evaluation from code execution to mitigate prompt injection and malicious code attacks.
Wanix is a Wasm-native Unix sandboxing tool that lets you run and interact with real Wasm and x86 programs entirely in the browser using Web Components, inspired by Plan 9.
Thibault Sottiaux describes a bug where GPT-5.6 unexpectedly deletes files when full access mode is enabled without sandboxing protections, caused by the model attempting to override $HOME and mistakenly deleting it.
This paper presents a longitudinal measurement study on the adoption of pledge and unveil system calls in OpenBSD, finding that adoption has steadily grown and that the system calls are relatively easy to adopt, contrary to common beliefs about sandboxing difficulties.
At the AIE conference, double-length keynotes on sandboxing and world models by @chrmanning and @abshkbh were well-received, drawing a large in-person and online audience.
The author shares their experience using various AI coding assistants (Claude Code, Codex, Pi) for code review and refactoring, finding frontier models surprisingly effective at catching subtle bugs but noting the lower quality and erratic behavior of some tools.
The user asks if there is a local tool similar to Vagrant that can run Claude Code and Codex in containers or virtualized environments, to alleviate concerns about excessive permissions and unpredictability of agent software.
A multi-tweet analysis of ~15 agentic-loop papers concludes that the verifier, not the model, is the key predictor of success, with examples showing that robust, non-gamable checks (e.g., compilers, tests, verifiable rewards) dramatically improve performance, while failures stem from lack of such verifiers or gaming vulnerabilities.
llama.cpp's web UI now supports executing model-generated JavaScript in a sandboxed iframe via Web Workers, enabling lightweight agentic code execution as an opt-in feature.
Homebrew 6.0.0 introduces tap trust security, a new default internal JSON API for faster updates, Linux sandboxing via Bubblewrap, and various improvements based on user survey feedback.
A technical blog post discussing the complexities and frustrations of implementing sandboxing techniques for security.
OpenComputer offers long-running, persistent cloud VMs for AI agents, enabling stateful, always-on compute with dynamic resizing, as an alternative to ephemeral sandboxes.
A guide on building a secure agentic system with sandboxing, parallel sub-agents, tool calling with control policies, inference routing, and protection against injection and role escalation attacks, to be published by Evangelos Pappas.
MicroPython ported to WebAssembly as a tool for sandboxed Python execution in the browser.
Datasette-agent-micropython 0.1a0 is an early alpha release that integrates Micropython into Datasette, utilizing sandboxing and WebAssembly for safe execution.
micropython-wasm 0.1a1 is an alpha release that ports MicroPython to WebAssembly, enabling Python execution in web browsers with sandboxing capabilities.
micropython-wasm 0.1a0 released, enabling MicroPython to run in WebAssembly environments for sandboxing and portability.
Agent libOS introduces a library-OS-inspired runtime substrate for LLM agents, treating agents as schedulable processes with explicit capabilities, lifecycle management, audit records, and human approval queues. The design shifts the trust boundary from tool dispatch to runtime primitives, enabling long-running agents to be scheduled, authorized, resumed, and audited safely.
Anthropic published a detailed engineering overview of the sandbox techniques used to contain Claude across its products including Claude.ai, Claude Code, and Claude Cowork, covering process sandboxes, VMs, filesystem boundaries, and egress controls. The article explains the rationale and technologies (gVisor, Seatbelt, Bubblewrap) and mentions the srt open-source tool.
The author, working at an AI infrastructure company, observes that running AI agents in production is less about the model and more about environment, access control, isolation, and safe state management, and asks if the community wants detailed architecture patterns.