policy-compliance

Tag

Cards List
#policy-compliance

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Hugging Face Blog · 2026-09-08 Cached

The article discusses the limitations of topic-level safety guards in AI models and introduces a new paper proposing boundary-aware self-distillation for controlled LLM safety refusal, focusing on refusing specific harmful subsets within topics rather than entire topics.

0 favorites 0 likes
#policy-compliance

PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents

Hugging Face Daily Papers · 2026-08-20 Cached

PolicyGuide compiles domain policies into workflow graphs and uses a proactive verifier to guide LLM agents through multi-step procedures, improving policy compliance across various benchmarks.

0 favorites 0 likes
#policy-compliance

Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters

arXiv cs.LG · 2026-07-31 Cached

Compliance2LoRA proposes a hypernetwork-based framework that generates policy-compliant LoRA adapters on demand for large reasoning models, enabling adjustable safety alignment across arbitrary policy subsets without retraining separate models.

0 favorites 0 likes
#policy-compliance

Policy-Conditioned Constrained Decoding for Column-Level Access Control in Text-to-SQL

arXiv cs.CL · 2026-07-15 Cached

This paper introduces PCC-SQL, a method for enforcing column-use policies in text-to-SQL generation by constrained decoding, achieving deterministic elimination of violations with 0% Leakage Rate and high coverage on benchmarks.

0 favorites 0 likes
#policy-compliance

Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents

arXiv cs.AI · 2026-07-09 Cached

This paper identifies a silent failure mode in tool-using LLM agents where policy violations occur without tool errors or agent self-reporting. The authors propose and evaluate lightweight deterministic pre-execution gates that significantly reduce such failures in the τ²-bench airline domain.

0 favorites 0 likes
#policy-compliance

Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models

arXiv cs.CL · 2026-05-21 Cached

This paper presents a large-scale assessment of medical LLMs, including custom MedGPTs and open-source models, finding 25-30% exhibit low factual accuracy and 33.6-54.3% violate operational thresholds, highlighting systemic safety risks.

0 favorites 0 likes
#policy-compliance

PolicyBank: Evolving Policy Understanding for LLM Agents

arXiv cs.CL · 2026-04-20 Cached

PolicyBank proposes a memory mechanism that enables LLM agents to autonomously refine their understanding of organizational policies through iterative interaction and corrective feedback, closing specification gaps that cause systematic behavioral divergence from true requirements. The work introduces a systematic testbed and demonstrates PolicyBank can close up to 82% of policy-gap alignment failures, significantly outperforming existing memory mechanisms.

0 favorites 0 likes
← Back to home

Submit Feedback