Tag
This paper introduces CLAM, a method for estimating localized causal effects from coarse-resolution data by jointly learning causal mechanisms and a disaggregation mapping, with applications in public health and environmental policy.
Introduces MedPIC-Bench, a benchmark with counterfactual questions to evaluate whether LLMs correctly apply medication-safety rules when patient-specific conditions change; across 28 LLMs, accuracy drops significantly on counterfactual questions, revealing a common failure to revise judgments.
StyleForge introduces a scene-level structured selection framework for fixed-layout indoor furniture styling, using a dynamic hypergraph style field and counterfactual style preference learning to improve furniture retrieval and style coherence on 3D-FRONT.
CG-World is a large-scale world-state dataset and protocol derived from industrial computer graphics pipelines, explicitly recording multimodal world states, interventions, and counterfactual branches to support world model research. It demonstrates improvements in geometry-conditioned video generation, action prediction, and closed-loop transfer of vision-language-action policies.
Elias Bareinboim announces his group's multiple ICML 2025 papers on causal AI, covering relational world models, robust offline RL, counterfactual identification, and causal game theory, highlighting the need for causal knowledge in AI reasoning.
NormWorlds-CF is a solver-verified benchmark for counterfactual normative reasoning. The paper proposes MR-GRPO, a reward mechanism that improves structured reasoning beyond final answers, showing that answer-only accuracy can be misleading in normative tasks.
PragReST is a self-supervised framework that improves LLM pragmatic reasoning by generating counterfactual reasoning traces and training models via supervised fine-tuning and reinforcement learning, achieving significant gains on pragmatic benchmarks without human-labeled data.
This paper surveys evaluation methods for world models and argues for a decision-making-centric framework that prioritizes counterfactual reasoning, planning, and policy optimization over visual quality. It introduces an L0–L7 evaluation ladder and a benchmark protocol to align evaluation with claimed utility.
CRAFT is a unified counterfactual reasoning framework that improves tabular question answering and fact verification by constructing both original and counterfactual statements, extracting evidence from bidirectional reasoning paths, and integrating them via a weighted mechanism. Experiments show consistent improvements over baselines on WikiTQ and TabFact datasets.
Introduces Discrete-WAM, a unified discrete latent vision-action world policy that enables compositional causal reasoning and counterfactual reasoning in autonomous driving through aligned discrete tokens and a shared discrete diffusion framework.
CRONOS is a benchmark that evaluates counterfactual physical consistency in video prediction models by intervening on viewpoint, scene, object category, and appearance while keeping physical event types fixed. It reveals substantial failures in current video generators.
Introduces Implicit Behavior Policy Optimization (IBPO), a counterfactual comparison-based credit assignment framework that improves training stability and performance in multi-step reasoning tasks for large language models by converting sparse terminal rewards into step-sensitive learning signals.
该论文提出并评估了一类称为事件图基质的因果推理世界模型,通过确定性重放在类型化RDF事件日志上进行反事实查询,在多个基准上优于基线模型,同时保证了可检查性和可重放一致性。
This paper introduces NoisyCoconut, an inference-time method that improves LLM reliability by injecting noise into latent trajectories to generate diverse reasoning paths. The approach enables models to abstain when uncertain, significantly reducing error rates in mathematical reasoning tasks without requiring retraining.
This paper proposes a Three-in-One world model using Deep Boltzmann Machines for marketing intervention, combining energy-based consistency, outcome prediction, and counterfactual inference.