boundary-aware-training

Tag

Cards List
#boundary-aware-training

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Hugging Face Blog · yesterday Cached

The article discusses the limitations of topic-level safety guards in AI models and introduces a new paper proposing boundary-aware self-distillation for controlled LLM safety refusal, focusing on refusing specific harmful subsets within topics rather than entire topics.

0 favorites 0 likes
← Back to home

Submit Feedback