adversarial-defense

Tag

Cards List
#adversarial-defense

Hybrid Adversarial Defence for Natural Language Understanding Tasks

arXiv cs.CL · 2026-06-04 Cached

Researchers from Southampton and Manchester propose a hybrid adversarial defence framework for LLMs that combines entropy-based, uncertainty-based, and geometric-based models to simultaneously address hallucination and adversarial vulnerability in NLU tasks, achieving up to 64.92% improvement in adversarial robustness and 62.27% reduction in attack success rate.

0 favorites 0 likes
#adversarial-defense

Protecting Language Models Against Unauthorized Distillation through Trace Rewriting

arXiv cs.CL · 2026-04-20 Cached

This paper proposes methods for protecting large language models against unauthorized knowledge distillation by rewriting reasoning traces to degrade training usefulness while preserving correctness, and embedding verifiable watermarks in distilled student models. The approach uses instruction-based and gradient-based rewriting techniques to achieve anti-distillation effects without compromising teacher model performance.

0 favorites 0 likes
← Back to home

Submit Feedback