model-auditing

Tag

Cards List
#model-auditing

Reference-Based Distillation Detection in LLMs

arXiv cs.LG · 2026-07-14 Cached

This paper introduces a reference-based method to detect whether an LLM was distilled from a specific teacher model, using membership inference. The approach achieves near-perfect accuracy in controlled settings and provides new evidence about potential distillation relationships involving QwQ, DeepSeek-R1, and GPT-OSS.

0 favorites 0 likes
#model-auditing

New method aims to keep kids safe from illegal AI-generated content

MIT News — Artificial Intelligence · 2026-07-13 Cached

MIT researchers developed a technique to audit AI models for their capability to generate child sexual abuse material without producing illegal outputs, achieving 100% accuracy in tests. This method examines hidden model representations to infer whether a model has been fine-tuned for harmful content, providing a scalable way for platforms and law enforcement to detect unsafe models.

0 favorites 0 likes
#model-auditing

Revealing Hidden Model Behaviors with Task-Specific Self-Reports

arXiv cs.CL · 2026-07-07 Cached

This paper introduces SAR, a lightweight LoRA adapter that enables fine-tuned language models to self-report hidden behaviors like backdoors, outperforming existing methods by detecting all implanted behaviors and halving hallucination rates.

0 favorites 0 likes
#model-auditing

Cyclic Denoising Reveals Ultrastable Memories in Diffusion Models

arXiv cs.LG · 2026-06-24 Cached

Cyclic denoising is introduced as a novel extraction attack that reveals ultrastable memorized training images in diffusion models by repeatedly noising and denoising samples. The technique requires no gradients or weight inspection and has implications for privacy auditing.

0 favorites 0 likes
#model-auditing

idSCD: Identifying Training Datasets through Semantic Correlation Descriptors

arXiv cs.LG · 2026-06-01 Cached

This paper introduces idSCD, a white-box method that uses semantic correlation descriptors to identify whether a dataset was used in training a model, outperforming existing baselines across multiple settings.

0 favorites 0 likes
#model-auditing

I-SAFE: Wasserstein Coherence Metrics for Structural Auditing of Scientific AI Models

arXiv cs.LG · 2026-05-22 Cached

This paper introduces I-SAFE, a post-hoc distributional auditing framework for scientific AI models using Wasserstein Coherence Metrics, which reveals structural differences in model outputs that accuracy-based evaluation fails to capture. Demonstrated on drug-target interaction prediction, the framework is model-agnostic and applicable to any domain with structured inputs and external priors.

0 favorites 0 likes
← Back to home

Submit Feedback