Tag
The paper proposes AdvRole, an adversarial closed-loop curriculum framework for training role-playing agents in large language models, which evolves scenario pools to improve performance on benchmarks.
DART proposes a distributional adversarial training framework for recurrent reasoning models, improving their robustness and performance on structured problems by using a local target distribution instead of single-point supervision.
RL-ADA is a co-evolutionary framework that replaces human labels with world feedback for training adversarially robust enterprise dialogue agents, using reinforcement learning to improve performance without annotation data.
RL-FAT is a reinforcement learning framework for fair adversarial training that improves robustness while reducing class-wise disparities. Experiments demonstrate competitive accuracy and better fairness compared to standard methods.
This paper proposes a novel adversarial training method that eliminates input gradients by using low-rank Householder expansions, reducing computational cost while achieving robustness comparable to standard techniques for small perturbations.
This paper introduces AnTrap, a benchmark for evaluating the robustness of Android GUI agents against runtime anomalies, revealing universal vulnerabilities and differentiating between learnable traps and intrinsic reasoning limitations.
This paper proposes MeanFlow-Transfer (MF-T) and Continuous Adversarial MeanFlow (CAMF) to unify the adaptation and acceleration of pretrained diffusion and flow models, enabling high-quality few-step generation on new domains with limited data.
The paper introduces a failure-aware adversarial retrieval-augmented framework using contextual bandits to improve robustness in natural language understanding, with significant improvements on benchmarks like SNLI, ANLI, and MultiNLI.
PixRestore is a VAE-free pixel-space diffusion transformer for unified image restoration, achieving high fidelity and efficiency via flow matching and adversarial fine-tuning to a one-step generator.
Introduces Adversarial Causal Intervention Falsification (ACIF), a sequential game where a structural causal generator proposes observational and interventional distributions while an adversarial experimentalist selects interventions to falsify it. The paper provides theoretical guarantees, including finite-sample convergence and model-selection, bridging causal generative modeling, active discovery, and experimental design.
This paper proposes computation-efficient strategies for latent adversarial training (LAT) of LLMs, using low-rank representation fine-tuning and circuit-guided surrogate models to reduce per-step FLOPs by 48.1% while requiring only 0.0118% trainable parameters.
RAGuard is a layered defense framework for Retrieval-Augmented Generation (RAG) systems that uses adversarial fine-tuning of the retriever and a label-free filter (ZKIP) to achieve zero attack success against corpus poisoning, maintaining high retrieval accuracy.
This paper introduces GPT-Red, an automated red-teaming agent trained via self-play at scale to discover novel prompt injection attacks against frontier LLMs, and uses it to adversarially train GPT-5.6, achieving the largest documented LLM safety training run.
This paper proposes Merge-Adversarial Training to make text watermarks in open-source LLMs survive model merging, outperforming baselines while preserving downstream capabilities.
Fence proposes using Small Language Models trained on high-quality synthetic data as specialized guardrails for LLM applications, demonstrating performance gains over prompt-based LLM guardrails.
OpenAI introduces GPT-Red, an automated red-teaming model that finds prompt injection vulnerabilities at scale and is used to adversarially train models like GPT-5.6 Sol, achieving 6x fewer failures on hard prompt injection benchmarks.
Introduces Latent Personality Alignment (LPA), a lightweight adversarial training method that uses 66 psychometric personality statements to achieve near-zero attack success rates on jailbreak attacks without degrading utility, requiring only minutes on a single GPU.
This paper introduces Evidential Adversarial Training (EV-AT), a method that improves the robustness-uncertainty trade-off in classifiers by combining an evidence-based loss with robust evidence alignment, achieving state-of-the-art results on selective classification benchmarks.
This paper from MIT proposes an adversarial generator-discriminator framework that combines verifiable rewards with a learned signal from human demonstrations to address issues like diversity collapse, unnatural responses, and reward hacking in RLVR training of language models.
GRAPE is a training framework that progressively exposes parameter space during adversarial training, achieving higher robust accuracy with fewer parameters compared to fixed-structure methods on CIFAR-10.