BraveGuard: From Open-World Threats to Safer Computer-Use Agents
Summary
BraveGuard is a self-evolving defense framework that trains guard models using open-world threat signals and realistic agent trajectories to improve safety detection in computer-use agents, achieving significant accuracy gains on the AgentHazard benchmark.
View Cached Full Text
Cached at: 06/04/26, 03:41 AM
Paper page - BraveGuard: From Open-World Threats to Safer Computer-Use Agents
Source: https://huggingface.co/papers/2606.01166 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
BraveGuard is a self-evolving defense framework that trains guard models using open-world threat signals and realistic agent trajectories to improve safety detection in computer-use agents.
Computer-use agentsextend language models from text generation to sustained interaction with files, terminals, browsers, and external tools. This shift createssafety risksthat are difficult to detect from isolated prompts or final responses, because harm often emerges only through multi-step execution traces whose individual actions appear locally benign. We introduce BraveGuard, a self-evolving defense framework for trainingguard modelsfromopen-world threat signalsand realisticagent trajectories. BraveGuard mines recent research sources to identify emerging risks and attack patterns, instantiates them asexecutable computer-use tasks, collects agent rollouts, and derivestrajectory-level supervisionfor guard model training. As new threats and validation failures appear, the pipeline can be repeated, yielding anadaptive defense looprather than a static, benchmark-driven training process. We instantiate BraveGuard by training multipleguard backbones, includingQwen3-GuardandLlama-Guardvariants, and evaluate the resulting guards on trajectory-level agent-safety benchmarks. BraveGuard consistently improvessafety detectionacross computer-use trajectories. OnAgentHazard, it substantially improves detection accuracy over off-the-shelfguard models, with accuracy increasing from 38.79% to 82.38% under the averaged guard-model setting. These results show that guard supervision grounded in open-world threat discovery and realistic agent execution can improve safety monitoring beyond fixed taxonomies and synthetic prompt-level data. BraveGuard offers a scalable path toward adaptive defenses forcomputer-use agentsfacing evolving real-world risks.
View arXiv pageView PDFGitHub27Add to collection
Get this paper in your agent:
hf papers read 2606\.01166
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
#### Yunhao-Feng/BraveGuard Text Generation• Updatedabout 21 hours ago • 5
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.01166 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.01166 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
HazardAuditor: From Executable Threats to Safer Computer-Use Agents
HazardAuditor introduces an execution-grounded framework for supervising safety in computer-use agents, with Guard Policy Optimization improving safety outcomes by up to 16.5% accuracy over prior methods.
OSGuard: A Benchmark for Safety in Computer-Use Agents
OSGuard is a dual-granularity benchmark for evaluating safety in computer-use agents under benign user instructions, featuring action-level judgments and risk-augmented execution suites to detect unsafe shortcuts.
AdaGuard: An Adaptive Guard Model with User-defined Policies
AdaGuard introduces a family of guard models (0.6B, 4B, 8B) that assess language model agent trajectories under user-defined policies, using a new dataset and reinforcement learning algorithm for adaptive safety.
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
Introduces SeerGuard, a consequence-aware safety framework for mobile GUI agents that uses a safety-augmented world model to assess risks before execution, improving safety-utility scores.
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
DreamGuard is a proactive runtime guardrail for LLM agents that uses a risk-aware world model to track latent state and predict future risks, enabling interventions before unsafe actions execute. It outperforms baselines on benchmarks and online evaluation with 25ms latency.