Tag
AdversaBench introduces an automated LLM red-teaming pipeline that uses five mutation operators and a three-judge panel with a meta-judge tiebreaker to confirm failures, revealing that attack difficulty varies by category and that adversarial prompts transfer from smaller to larger models.
The author introduces 'Bracket', an open-source tool that automates hyperparameter search for diffusion model fine-tuning using parallel training trials and VLM-based scoring to objectively determine the best configuration.