counterfactual

Tag

Cards List
#counterfactual

C$^{3}$T: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees

arXiv cs.CL · 2026-09-03 Cached

The paper presents C3T, a thread-structured temporal model for predicting sentiment shifts in social media conversation trees using counterfactual causal reasoning, and introduces the CaSiRe dataset for causal sentiment reasoning.

0 favorites 0 likes
#counterfactual

Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses

arXiv cs.AI · 2026-08-14 Cached

Introduces DECAF, a method that decomposes perturbation responses into evidence, contradiction, and fragility components, improving interpretability over raw response magnitude and achieving strong results across vision benchmarks.

0 favorites 0 likes
#counterfactual

Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge

arXiv cs.AI · 2026-08-11 Cached

This paper introduces TKFQA, a counterfactual benchmark of 10,130 QA pairs over tables, texts, and knowledge graphs for evaluating LLM factuality consistency and order-robust reasoning, and proposes ORLF, a training framework that improves reasoning-chain accuracy and reduces input-order sensitivity.

0 favorites 0 likes
#counterfactual

Evidence-RL: Towards Evidence-intensive Visual Reasoning

Hugging Face Daily Papers · 2026-08-08 Cached

This paper introduces Counterfactual Evidence Disentanglement (CED), a training-time method that makes vision-language models rely on concrete image evidence rather than language priors or shortcuts, improving visual reasoning grounding across benchmarks.

0 favorites 0 likes
#counterfactual

C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

arXiv cs.AI · 2026-08-07 Cached

Introduces C3PO, a benchmark of 3,404 samples for evaluating cross-modal composition and counterfactual reasoning in multimodal LLMs. It finds modality dominance causes most failures, with even the best model (Gemini-3.1-Pro) far below human accuracy.

0 favorites 0 likes
#counterfactual

CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models

arXiv cs.AI · 2026-08-06 Cached

Introduces CARGO-VL, a group-relative optimization framework for vision-language models that improves handling of conflicting image-text evidence and unsupported-answer avoidance via counterfactual consistency and risk-constrained control, along with the XMC conflict training resource.

0 favorites 0 likes
#counterfactual

DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series

arXiv cs.LG · 2026-07-31 Cached

DoTime is a synthetic benchmark generator for interventional and counterfactual time series, providing scalable TSCM-based data generation with exact ground truth, released as a PyPI package with evaluation suites. It enables training and benchmarking causal foundation models on non-observational time series data.

0 favorites 0 likes
#counterfactual

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

Hugging Face Daily Papers · 2026-07-30 Cached

Introduces Visual Attribution Distillation (VAD), a counterfactual target-reconstruction method for multimodal on-policy distillation that estimates the visually-attributable part of teacher corrections. Outperforms existing approaches across fine-grained visual benchmarks at 4B and 9B scales.

0 favorites 0 likes
#counterfactual

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning

Hugging Face Daily Papers · 2026-07-30 Cached

This paper proposes Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for highly sensitive tokens and renormalizes token-level advantages for long-CoT mathematical reasoning. It consistently outperforms GRPO baselines, while also revealing that privileged token-shift directions are unreliable and mostly reflect counterfactual sensitivity rather than learning value.

0 favorites 0 likes
#counterfactual

TokenMem: Faithful Knowledge Injection for Frozen LLMs

arXiv cs.AI · 2026-07-28 Cached

TokenMem injects knowledge into frozen LLMs via a dedicated cross-attention channel, training a thin gating adapter through two-phase curriculum to improve knowledge compliance under counterfactual knowledge, achieving 69-70% KC compared to 20-52% for vanilla RAG.

0 favorites 0 likes
#counterfactual

Concept-based Visual Counterfactual Explanations with Diffusion Models

arXiv cs.AI · 2026-07-28 Cached

Introduces C-VCE, a diffusion framework that builds an interpretable concept bottleneck layer into the generative model, enabling human-guided visual counterfactual explanations without relying on external noise-robust classifiers.

0 favorites 0 likes
#counterfactual

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs

arXiv cs.AI · 2026-07-15 Cached

This paper introduces a method for ensuring LLMs report their true beliefs by using counterfactual report coordinates that resist pressure but remain responsive to genuine evidence. The approach achieves high performance on a benchmark, demonstrating a causal certificate for internal incentive compatibility.

0 favorites 0 likes
#counterfactual

Counterfactual Residual Data Augmentation for Regression

arXiv cs.LG · 2026-06-30 Cached

Proposes Counterfactual Residual Data Augmentation (CRDA) for tabular regression, leveraging residual invariance under feature perturbations to generate realistic training samples, achieving significant MSE reduction on benchmarks.

0 favorites 0 likes
#counterfactual

Counterfactual Optimization of Baseball Pitch Sequences and Estimation of Its Impact on Season-Level Statistics

arXiv cs.LG · 2026-06-17 Cached

This paper uses a Transformer-based model on MLB Statcast data to counterfactually optimize baseball pitch sequences, finding that optimizing both final and setup pitches can improve season-level statistics like K/9 by over 1.0.

0 favorites 0 likes
#counterfactual

A Definition of Good Explanations and the Challenges Explaining LLM Outputs

arXiv cs.AI · 2026-06-16 Cached

This paper proposes a definition of good explanations based on counterfactuals and prior beliefs, and discusses the inherent difficulties in explaining LLM outputs under this definition.

0 favorites 0 likes
#counterfactual

WorldKernel: A World Model is the Coupling Kernel of Admissible Possible Worlds

arXiv cs.AI · 2026-06-10 Cached

The paper identifies a failure mode where predictors collapse to a point on unidentified counterfactual couplings and proposes a framework using a positive semidefinite coupling kernel to bound counterfactuals, showing that prediction cannot represent uncertainty over cross-world couplings and that enforcing kernel constraints yields tractable bounds.

0 favorites 0 likes
#counterfactual

Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents

arXiv cs.AI · 2026-06-09 Cached

Introduces CICL, a decision-aware context layer that selects and compresses evidence for tool-using LLM agents by treating context as a decision-time intervention, using counterfactual-inspired scoring and typed memory cards under a token budget. Experiments on SWE-bench and RepoBench show concrete gains in retrieval accuracy and action criticality.

0 favorites 0 likes
#counterfactual

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

arXiv cs.LG · 2026-06-01 Cached

This paper introduces the Causal Sensitivity Score (CSS), an interventional metric that evaluates whether clinical LLMs and agents appropriately update their recommendations when patient inputs change along clinically meaningful dimensions. It reveals hidden capability profiles not captured by standard coverage-based metrics, exposing safety blind spots and structural responsiveness deficits.

0 favorites 0 likes
#counterfactual

COFT: Counterfactual-Conformal Decoding for Fair Chain-of-Thought Reasoning in Large Language Models

arXiv cs.CL · 2026-06-01 Cached

COFT is a training-free decoding method that applies token-level fairness control and conformal calibration to reduce bias in chain-of-thought reasoning of large language models, achieving 30-55% bias reduction with minimal computational overhead.

0 favorites 0 likes
#counterfactual

From Pixels to Concepts: Do Segmentation Models Understand What They Segment?

Hugging Face Daily Papers · 2026-05-10 Cached

Introduces CAFE, a benchmark for evaluating whether promptable segmentation models truly understand concepts by using counterfactual attribute manipulation, revealing that accurate mask prediction does not guarantee faithful semantic grounding.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback