Tag
SAGE is a framework that mitigates exploration and compounding biases in long-horizon reasoning for large language models through algebraic sparsification and hyperbolic structural guidance, achieving significant improvements across benchmarks.
The study systematically assesses algorithmic fairness in machine learning models for predicting treatment retention in medication for opioid use disorder, finding performance gaps across patient subgroups and evaluating bias mitigation techniques with trade-offs.
FairLMs is a Python library designed to streamline fairness research in language models by providing metrics, mitigation methods, and diagnostics with explicit capability declarations.
MUtE introduces a novel framework for optimal concept erasure and counterfactual interventions in text representations, enhancing algorithmic fairness and enabling counterfactual text generation in NLP models.
This paper identifies popularity bias in large language models, where they systematically favor popular but incorrect options in multiple-choice questions, and proposes a mitigation technique called PopDebias.
This paper proposes Subgraph Filtering for Fair Graph Neural Networks (SF-GNN), a framework that mitigates structural bias in GNNs by filtering edges to improve fairness-accuracy trade-offs.
This paper introduces Think-Probe-Respond, a method to improve large language models' judgment of research idea novelty by probing latent judgments and reducing bias towards medium novelty ratings, achieving a 22.30% performance improvement.
This paper introduces an adaptive triggering framework for bias correction in LLM chain-of-thought reasoning, using online change-point detection to optimize intervention timing with white-box and black-box signals, improving accuracy and reducing interventions.
GGSS reduces demographic bias in generative vision-language models by steering visual tokens along geodesic arcs with an adaptive gate during inference, preserving visual-language accuracy.
This paper proposes Counterfactual Ensemble Decoding (CED) to mitigate social biases in large vision-language models by constructing multi-group counterfactual perspectives and integrating them during decoding, achieving substantial bias reduction while preserving model capabilities.
This paper investigates how large language models implicitly personalize outputs based on demographic cues, locating an internal activation signal that tracks these shifts and showing that removing this signal can suppress the behavior.
This paper proposes HEIMAT, a heuristic-style automatic debiasing framework for language models that uses heuristic prompts to reveal biases and fine-tunes the model to reduce bias while preserving NLU performance.
This paper presents Fairness Pruning, a structural intervention method that locates demographic bias in GLU-MLP layers of large language models by identifying differentially activated neurons. Zeroing a very small number of neurons disrupts bias processing while preserving reasoning and general knowledge, suggesting bias and capabilities rely on dissociable circuits.
SkillSight is a training-free retrieval framework that calibrates shared background in skill descriptions to improve skill retrieval accuracy for LLM agents, achieving up to 20.21 percentage point improvement in Recall@10 over dense retrievers.
FairSelect is a toolkit for systematically evaluating fairness mitigation strategies across preprocessing, inprocessing, and postprocessing stages, supporting intersectional subgroup evaluation and comparison of fairness-utility tradeoffs.
This paper audits and mitigates dialect bias in large language models, showing they systematically prefer Standard American English over African American English. The authors introduce activation steering, a training-free method that reduces bias significantly while preserving fluency, and release the largest real-AAE parallel corpus to date.
The paper introduces CO-ALIGN, a bias mitigation method for text-to-image diffusion models that aligns concept graphs in the text encoder and denoiser, achieving 30% fairness improvement and 11.4 FID gain while reducing incoherent outputs by 88%.
Introduces Reward-Gated Test-Time Adaptation (RG-TTA), a reinforcement learning framework that selectively applies debiasing to CLIP models based on input bias sensitivity, resolving the fairness-utility trade-off.
Proposes a multimodal framework for fair Mild Cognitive Impairment detection from speech, using unlearning via gradient reversal to reduce demographic bias and improve performance across subgroups.
Introduces Face-Fairness (FF), a plug-and-play framework for bias mitigation in deepfake detection, featuring Face-Feature Tuning (FFT) as the first demographic label-free fairness method that improves group accuracy and reduces performance gaps across demographics.