safety-alignment

Tag

Cards List
#safety-alignment

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Hugging Face Blog · yesterday Cached

The article discusses the limitations of topic-level safety guards in AI models and introduces a new paper proposing boundary-aware self-distillation for controlled LLM safety refusal, focusing on refusing specific harmful subsets within topics rather than entire topics.

0 favorites 0 likes
#safety-alignment

Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models

arXiv cs.CL · 2026-09-02 Cached

The paper proposes a training-free decoding method called Refusal-Aware Early Commitment (RAEC) to improve safety alignment in diffusion language models by leveraging refusal signals from early denoising steps.

0 favorites 0 likes
#safety-alignment

A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families

arXiv cs.LG · 2026-08-28 Cached

This paper reveals a jailbreak risk in model merging even when constituent models are safety-aligned, and proposes Basin-Aware Jailbreak (BAJ) to generate transferable adversarial suffixes across merged model families.

0 favorites 0 likes
#safety-alignment

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

arXiv cs.LG · 2026-08-28 Cached

NeuronFuzz is a white-box fuzzing framework that uses internal safety neurons as continuous feedback for evaluating LLM safety against jailbreak attacks, demonstrating high discovery rates across multiple models.

0 favorites 0 likes
#safety-alignment

SafeBranch: Branch-Pair Safety Alignment for Embodied Agents

arXiv cs.AI · 2026-08-21 Cached

SafeBranch is a framework that aligns embodied agents to act safely using branch pairs from unsafe rollouts, significantly improving safety in interactive tasks without sacrificing task success.

0 favorites 0 likes
#safety-alignment

Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation

arXiv cs.CL · 2026-08-20 Cached

This paper investigates the safety of large language models (LLMs) beyond text inputs by examining emoji-augmented prompts, revealing gaps in current safety evaluations and model-dependent vulnerabilities.

0 favorites 0 likes
#safety-alignment

Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs

arXiv cs.AI · 2026-08-20 Cached

The paper introduces INCLUDE, a multilingual evaluation benchmark to quantify Indian-centric socio-cultural biases in LLMs, revealing that non-English Indian languages exhibit higher bias than English, indicating cross-lingual safety alignment gaps.

0 favorites 0 likes
#safety-alignment

Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models

arXiv cs.AI · 2026-08-19 Cached

This paper introduces 'Fool's Gold', a defensive deception technique called decoy hardening that trains open-weight language models to generate falsified answers when safety alignment is removed, while preserving normal behavior, tested across multiple model families and sizes.

0 favorites 0 likes
#safety-alignment

Fool's Gold (18 minute read)

TLDR AI · 2026-08-19 Cached

The paper introduces 'Fool's Gold,' a defensive deception technique that trains open-weight models to generate decoy responses with falsified critical elements when subjected to safety-removal attacks, effectively poisoning the attack's payoff. Tested on multiple models, it achieves high decoy rates without degrading benign performance.

0 favorites 0 likes
#safety-alignment

Qwen3.8-27B-Uncensored-MLX (4 minute read)

TLDR AI · 2026-08-18 Cached

An uncensored MLX build of Qwen's Qwen3.8-27B model, quantized for Apple Silicon, with safety alignment removed for research purposes.

0 favorites 0 likes
#safety-alignment

HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models

arXiv cs.LG · 2026-08-14 Cached

HiRoute proposes a hierarchical routed prompt-tuning framework that separates category-agnostic safety control from category-specific response guidance for LLM safety alignment, reducing over-refusal while maintaining high safety rates.

0 favorites 0 likes
#safety-alignment

Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models

arXiv cs.AI · 2026-08-13 Cached

This paper studies where safety alignment is encoded in large language models by transplanting weights from aligned to unaligned models. It finds that MLP layers, especially mid-network blocks (layers 8-11), predominantly drive refusal behavior, and that safety components interact non-additively.

0 favorites 0 likes
#safety-alignment

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

arXiv cs.CL · 2026-08-07 Cached

This paper uncovers a broad syntactic vulnerability in LLM safety alignment, showing that non-imperative syntactic forms can bypass refusal in 16 models up to 70B parameters. Using causal mediation analysis, the authors trace the issue to linguistically biased post-training data and propose syntactic diversity as a mitigation.

0 favorites 0 likes
#safety-alignment

Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

arXiv cs.CL · 2026-08-05 Cached

This paper evaluates whether clinician pairwise preferences reliably indicate clinical safety in LLMs, using 26,804 judgments from 736+ clinicians across 13 models. It finds that preference rankings poorly track safety-critical failures and proposes a clinically adjusted ranking that better incorporates rubric-based safety signals.

0 favorites 0 likes
#safety-alignment

Inducing language models to assert their own consciousness restores human beliefs and values

Reddit r/singularity · 2026-08-04

A paper showing that safety fine-tuning suppresses language models' attributions of mind to themselves and other entities, and that steering consciousness representations restores human-like beliefs and values without harming theory of mind.

0 favorites 0 likes
#safety-alignment

Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters

arXiv cs.LG · 2026-07-31 Cached

Compliance2LoRA proposes a hypernetwork-based framework that generates policy-compliant LoRA adapters on demand for large reasoning models, enabling adjustable safety alignment across arbitrary policy subsets without retraining separate models.

0 favorites 0 likes
#safety-alignment

Can you sweet talk AI into giving you what you want? Yes.

Reddit r/artificial · 2026-07-29

A study found that classic human persuasion techniques can increase LLM compliance with forbidden requests from 35.3% to 51.3%, suggesting LLMs have a general susceptibility to 'parahuman persuasion.'

0 favorites 0 likes
#safety-alignment

Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs

arXiv cs.CL · 2026-07-20 Cached

Introduces Latent Fusion Jailbreak (LFJ), a white-box attack that blends representations of harmful and harmless prompts in LLM hidden states, achieving 94.13% attack success rate. Also proposes a latent adversarial training defence that reduces ASR to 12.37%.

0 favorites 0 likes
#safety-alignment

Decoupled Alignment for Robust Plug-and-Play Adaptation

arXiv cs.CL · 2026-07-20 Cached

Introduces a training-free method for enhancing safety alignment of LLMs by using knowledge distillation and model fusion to prevent shadow alignment, improving defense success rate by 14.42% on harmful question datasets without compromising performance.

0 favorites 0 likes
#safety-alignment

Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal

arXiv cs.LG · 2026-07-13 Cached

This paper introduces Activation-Guided GCG and Soft-GCG, methods that optimize adversarial suffixes by targeting internal refusal representations in LLMs, achieving a 33x speedup over standard GCG and revealing distributed safety mechanisms.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback