bias-mitigation

Tag

Cards List
#bias-mitigation

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

Hugging Face Daily Papers ↗ · 2026-09-24 Cached

SAGE is a framework that mitigates exploration and compounding biases in long-horizon reasoning for large language models through algebraic sparsification and hyperbolic structural guidance, achieving significant improvements across benchmarks.

0 favorites 0 likes
#bias-mitigation

Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder

arXiv cs.LG ↗ · 2026-09-22 Cached

The study systematically assesses algorithmic fairness in machine learning models for predicting treatment retention in medication for opioid use disorder, finding performance gaps across patient subgroups and evaluating bias mitigation techniques with trade-offs.

0 favorites 0 likes
#bias-mitigation

FairLMs: A Turnkey Library for Fairness in Language Models

arXiv cs.LG ↗ · 2026-09-21 Cached

FairLMs is a Python library designed to streamline fairness research in language models by providing metrics, mitigation methods, and diagnostics with explicit capability declarations.

0 favorites 0 likes
#bias-mitigation

MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions

arXiv cs.LG ↗ · 2026-09-11 Cached

MUtE introduces a novel framework for optimal concept erasure and counterfactual interventions in text representations, enhancing algorithmic fairness and enabling counterfactual text generation in NLP models.

0 favorites 0 likes
#bias-mitigation

Large Language Models Systematically Favor Popular Options: Evidence and Mitigation Across MCQs

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper identifies popularity bias in large language models, where they systematically favor popular but incorrect options in multiple-choice questions, and proposes a mitigation technique called PopDebias.

0 favorites 0 likes
#bias-mitigation

Subgraph Filtering for Fair Graph Neural Networks

arXiv cs.LG ↗ · 2026-08-28 Cached

This paper proposes Subgraph Filtering for Fair Graph Neural Networks (SF-GNN), a framework that mitigates structural bias in GNNs by filtering edges to improve fairness-accuracy trade-offs.

0 favorites 0 likes
#bias-mitigation

Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

arXiv cs.CL ↗ · 2026-08-27 Cached

This paper introduces Think-Probe-Respond, a method to improve large language models' judgment of research idea novelty by probing latent judgments and reducing bias towards medium novelty ratings, achieving a 22.30% performance improvement.

0 favorites 0 likes
#bias-mitigation

Adaptive Triggering for Bias Correction in LLM Reasoning

arXiv cs.CL ↗ · 2026-08-27 Cached

This paper introduces an adaptive triggering framework for bias correction in LLM chain-of-thought reasoning, using online change-point detection to optimize intervention timing with white-box and black-box signals, improving accuracy and reducing interventions.

0 favorites 0 likes
#bias-mitigation

GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models

Hugging Face Daily Papers ↗ · 2026-08-26 Cached

GGSS reduces demographic bias in generative vision-language models by steering visual tokens along geodesic arcs with an adaptive gate during inference, preserving visual-language accuracy.

0 favorites 0 likes
#bias-mitigation

Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding

arXiv cs.CL ↗ · 2026-08-25 Cached

This paper proposes Counterfactual Ensemble Decoding (CED) to mitigate social biases in large vision-language models by constructing multi-group counterfactual perspectives and integrating them during decoding, achieving substantial bias reduction while preserving model capabilities.

0 favorites 0 likes
#bias-mitigation

Locating and Controlling Implicit Personalization in Large Language Models

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper investigates how large language models implicitly personalize outputs based on demographic cues, locating an internal activation signal that tracks these shifts and showing that removing this signal can suppress the behavior.

0 favorites 0 likes
#bias-mitigation

A Heuristic Perspective on Debiasing Language Models

arXiv cs.CL ↗ · 2026-08-04 Cached

This paper proposes HEIMAT, a heuristic-style automatic debiasing framework for language models that uses heuristic prompts to reveal biases and fine-tunes the model to reduce bias while preserving NLU performance.

0 favorites 0 likes
#bias-mitigation

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

arXiv cs.CL ↗ · 2026-07-31 Cached

This paper presents Fairness Pruning, a structural intervention method that locates demographic bias in GLU-MLP layers of large language models by identifying differentially activated neurons. Zeroing a very small number of neurons disrupts bias processing while preserving reasoning and general knowledge, suggesting bias and capabilities rely on dissociable circuits.

0 favorites 0 likes
#bias-mitigation

SkillSight: Seeing Through Shared Descriptions for Accurate Skill Retrieval

arXiv cs.AI ↗ · 2026-07-22 Cached

SkillSight is a training-free retrieval framework that calibrates shared background in skill descriptions to improve skill retrieval accuracy for LLM agents, achieving up to 20.21 percentage point improvement in Recall@10 over dense retrievers.

0 favorites 0 likes
#bias-mitigation

FairSelect: A Systematic Evaluation of Multi-Level and Intersectional Algorithmic Fairness

arXiv cs.LG ↗ · 2026-07-13 Cached

FairSelect is a toolkit for systematically evaluating fairness mitigation strategies across preprocessing, inprocessing, and postprocessing stages, supporting intersectional subgroup evaluation and comparison of fairness-utility tradeoffs.

0 favorites 0 likes
#bias-mitigation

LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering

arXiv cs.CL ↗ · 2026-07-09 Cached

This paper audits and mitigates dialect bias in large language models, showing they systematically prefer Standard American English over African American English. The authors introduce activation steering, a training-free method that reduces bias significantly while preserving fluency, and release the largest real-AAE parallel corpus to date.

0 favorites 0 likes
#bias-mitigation

Efficient bias mitigation in T2I diffusion models using Concept Graphs

arXiv cs.AI ↗ · 2026-07-07 Cached

The paper introduces CO-ALIGN, a bias mitigation method for text-to-image diffusion models that aligns concept graphs in the text encoder and denoiser, achieving 30% fairness improvement and 11.4 FID gain while reducing incoherent outputs by 88%.

0 favorites 0 likes
#bias-mitigation

Selective Test-Time Debiasing for CLIP via Reward Gating

arXiv cs.CL ↗ · 2026-07-02 Cached

Introduces Reward-Gated Test-Time Adaptation (RG-TTA), a reinforcement learning framework that selectively applies debiasing to CLIP models based on input bias sensitivity, resolving the fairness-utility trade-off.

0 favorites 0 likes
#bias-mitigation

Fair Cognitive Impairment Detection Through Unlearning

arXiv cs.LG ↗ · 2026-06-18 Cached

Proposes a multimodal framework for fair Mild Cognitive Impairment detection from speech, using unlearning via gradient reversal to reduce demographic bias and improve performance across subgroups.

0 favorites 0 likes
#bias-mitigation

Toward Calibrated, Fair, and accurate Deepfake Detection

arXiv cs.LG ↗ · 2026-06-10 Cached

Introduces Face-Fairness (FF), a plug-and-play framework for bias mitigation in deepfake detection, featuring Face-Feature Tuning (FFT) as the first demographic label-free fairness method that improves group accuracy and reduces performance gaps across demographics.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback