Tag
This paper introduces a reference-based method to detect whether an LLM was distilled from a specific teacher model, using membership inference. The approach achieves near-perfect accuracy in controlled settings and provides new evidence about potential distillation relationships involving QwQ, DeepSeek-R1, and GPT-OSS.
MIT researchers developed a technique to audit AI models for their capability to generate child sexual abuse material without producing illegal outputs, achieving 100% accuracy in tests. This method examines hidden model representations to infer whether a model has been fine-tuned for harmful content, providing a scalable way for platforms and law enforcement to detect unsafe models.
This paper introduces SAR, a lightweight LoRA adapter that enables fine-tuned language models to self-report hidden behaviors like backdoors, outperforming existing methods by detecting all implanted behaviors and halving hallucination rates.
Cyclic denoising is introduced as a novel extraction attack that reveals ultrastable memorized training images in diffusion models by repeatedly noising and denoising samples. The technique requires no gradients or weight inspection and has implications for privacy auditing.
This paper introduces idSCD, a white-box method that uses semantic correlation descriptors to identify whether a dataset was used in training a model, outperforming existing baselines across multiple settings.
This paper introduces I-SAFE, a post-hoc distributional auditing framework for scientific AI models using Wasserstein Coherence Metrics, which reveals structural differences in model outputs that accuracy-based evaluation fails to capture. Demonstrated on drug-target interaction prediction, the framework is model-agnostic and applicable to any domain with structured inputs and external priors.