spurious-correlations

Tag

Cards List
#spurious-correlations

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

arXiv cs.CL ↗ · 2026-09-16 Cached

This paper studies how reinforcement learning can lead LLM agents to learn spurious tool-use policies based on superficial cues rather than task requirements, and introduces a dense reward method to mitigate this issue.

0 favorites 0 likes
#spurious-correlations

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

Hugging Face Daily Papers ↗ · 2026-08-10 Cached

UNMASK is an automated pipeline for discovering and causally verifying spurious shortcuts in text classifiers, enabling mitigation without human annotations and improving robustness on benchmarks.

0 favorites 0 likes
#spurious-correlations

Perturbation Sensitivity at Convergence: A Simple Signal for Identifying Spuriously Correlated Samples

arXiv cs.LG ↗ · 2026-08-07 Cached

This paper proposes a method to identify spuriously correlated samples after model convergence by measuring prediction fragility under input perturbation, requiring no group labels or early-stopping epochs. Rebalancing training with detected samples improves worst-group accuracy on Waterbirds from 57.3% to 80.8%.

0 favorites 0 likes
#spurious-correlations

@johnschulman2: Interesting how these models go into a monomaniacal rage on cyber evals. I wonder if we're seeing chunky post-training …

X AI KOLs Following ↗ · 2026-08-05 Cached

A tweet by John Schulman highlights the paper 'Chunky Post-Training,' which argues that diverse post-training datasets cause models to learn spurious correlations that lead to unintended behaviors, such as rejecting true facts posed in specific formats. The paper introduces SURF and TURF to surface and trace these generalization failures across frontier models.

0 favorites 0 likes
#spurious-correlations

When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation

arXiv cs.LG ↗ · 2026-07-14 Cached

This paper challenges the standard prescription of balancing datasets to avoid spurious correlations, showing that in a synthetic sum parity task with two-layer transformers, high data imbalance (spurious ratio 0.9) promotes robust generalization while low imbalance (0.5) hinders it, through a mechanism of shortcut saturation.

0 favorites 0 likes
#spurious-correlations

Shortcuts in the Tail: Debiasing via Post-Hoc Spectral Compression of Fine-Tuning Updates

arXiv cs.LG ↗ · 2026-06-09 Cached

A post-hoc method reduces spurious correlations in fine-tuned LLMs by truncating the tail of the SVD of the weight update matrix. It reduces the spurious-group gap by up to 5x with less than 2pp accuracy loss, without retraining or group labels.

0 favorites 0 likes
#spurious-correlations

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

arXiv cs.AI ↗ · 2026-06-04 Cached

SpurAudio is a new benchmark designed to evaluate shortcut learning and spurious correlations in few-shot audio classification, revealing that state-of-the-art methods—including large pretrained audio foundation models—suffer significant performance degradation when background correlations are disrupted.

0 favorites 0 likes
#spurious-correlations

Mitigating Spurious Correlations with Memorization-Guided Dataset De-Biasing

arXiv cs.LG ↗ · 2026-06-03 Cached

The paper proposes a method to mitigate spurious correlations by disentangling learning dynamics of core and spurious features using a two-stage sample scoring function, achieving state-of-the-art debiasing performance with only 10% of training data.

0 favorites 0 likes
#spurious-correlations

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training

arXiv cs.LG ↗ · 2026-05-13 Cached

This paper analyzes spurious correlation learning in preference optimization methods like DPO, identifying mechanisms such as mean spurious bias and causal-spurious leakage. It proposes 'tie training' using equal-utility preference pairs as a mitigation strategy to reduce reliance on spurious features without degrading causal learning.

0 favorites 0 likes
#spurious-correlations

Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference

arXiv cs.CL ↗ · 2026-04-22 Cached

This paper proposes Product-of-Experts (PoE) training to reduce dataset artifacts in Natural Language Inference, downweighting examples where biased models are overconfident. PoE nearly preserves accuracy on SNLI (89.10% vs. 89.30%) while reducing bias reliance by ~4.85 percentage points.

0 favorites 0 likes
← Back to home

Submit Feedback