instruction-hierarchy

Tag

Cards List
#instruction-hierarchy

A prompt injection test caught something we would've shipped

Reddit r/AI_Agents · 2d ago

A team describes how their prompt-injection eval suite caught a regression in a document assistant before shipping, emphasizing the importance of maintaining a strict hierarchy between system instructions and retrieved data.

0 favorites 0 likes
#instruction-hierarchy

Steering Instruction Hierarchies at Inference Time

arXiv cs.CL · 2026-07-30 Cached

Introduces V-Steer, a training-free inference-time method that edits cached value vectors to restore instruction hierarchy in language models, raising primary constraint accuracy from under 18% to 92% on controlled benchmarks with negligible overhead.

0 favorites 0 likes
#instruction-hierarchy

Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs

arXiv cs.CL · 2026-07-28 Cached

This paper introduces XIH-Bench, a benchmark for evaluating instruction hierarchy compliance in multilingual LLMs, revealing language-dependent asymmetry and a Language Boundary Effect where cross-language conflicts yield higher compliance than same-language ones.

0 favorites 0 likes
#instruction-hierarchy

Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models

arXiv cs.AI · 2026-06-09 Cached

This paper introduces a white-box diagnostic framework that localizes instruction hierarchy failures in reasoning language models into identification, conflict resolution, and response realization stages. It evaluates several models and proposes two training-free self-monitoring mechanisms that reduce non-compliance by 81–99%.

0 favorites 0 likes
#instruction-hierarchy

Evaluating using Mock Tool Calls to Quarantine Untrusted Prompt Inputs

arXiv cs.CL · 2026-06-01 Cached

This paper evaluates whether wrapping untrusted content in mock tool calls improves LLM robustness against adversarial inputs, finding it does not broadly help and sometimes increases attack success rates.

0 favorites 0 likes
#instruction-hierarchy

Improving instruction hierarchy in frontier LLMs

OpenAI Blog · 2026-03-10 Cached

OpenAI presents a training approach using instruction-hierarchy tasks to improve LLM safety and reliability by teaching models to properly prioritize instructions based on trust levels (system > developer > user > tool). The method addresses prompt-injection attacks and safety steerability through reinforcement learning with a new dataset called IH-Challenge.

0 favorites 0 likes
#instruction-hierarchy

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

OpenAI Blog · 2024-04-19 Cached

OpenAI proposes an instruction hierarchy approach to defend LLMs against prompt injection and jailbreak attacks by training models to prioritize system instructions over user inputs. The method significantly improves robustness without degrading standard capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback