Tag
LatentMAS is a new multi-agent collaboration method where agents directly transfer reasoning states in latent space without text encoding/decoding, achieving a 13.3% accuracy improvement, 4.3x speed, and 83.7% reduction in token usage. It requires no extra training and can be plugged into existing LLMs. It has been accepted as an ICML 2026 Spotlight.
This paper proposes the Time-Reparameterized Cumulative Intensity Extrapolation (TR-CIE) sampler for discrete flow matching, which improves sampling quality under limited function evaluations by rescaling the time grid and reusing cached model outputs, with theoretical analysis and experiments on text and image generation.
A paper on few queries has been accepted at the ICML DL4C workshop.
This paper proposes a Machine-Learned Comorbidity Index (MLCI) that uses diagnosis codes and nonlinear learning to improve risk adjustment across multiple clinical outcomes, outperforming traditional mortality-centric indices.
This ICML 2026 spotlight position paper identifies a failure mode in image-generation alignment where aesthetic preference optimization overrides explicit user intent, terming it 'reversed alignment' and testing on anti-aesthetic prompts.
The paper reveals that latent reasoning in transformer-based reasoning models (TRMs) functions as a policy improvement operator, and proposes an algorithm that enhances learning and inference efficiency by up to 18x.
This paper introduces relational structural causal models, extending structural causal models to settings with varying objects and relations. It provides theoretical results for identification and proposes relational neural causal models that outperform non-relational baselines on simulated traffic scenes.
This paper proposes a definition of good explanations based on counterfactuals and prior beliefs, and discusses the inherent difficulties in explaining LLM outputs under this definition.
Introduces TruDi, a method that enables training diffusion policies in massively parallel on-policy reinforcement learning by using a trust-region optimization rule to enforce KL constraints, achieving strong performance across 73 tasks.
Forgis Labs presents a family of foundation models for time series sensor data in industrial settings, with five papers accepted to ICML 2026 workshops, enabling event prediction and natural language explanation from raw sensor streams.
This paper introduces DRIVE, a unified Transformer-based framework for offline auto-bidding that decouples candidate action generation from decision making, combining distributional action modeling, retrieval-augmented candidate generation, and value-based evaluation to improve bidding performance under budget and cost constraints.
This paper proposes Constraint-Sensitive Policy Optimization (CSPO), a first-order primal-dual method for safe reinforcement learning that incorporates local constraint sensitivity to improve safety recovery and reduce oscillations near safety boundaries, achieving higher constrained returns on navigation and locomotion benchmarks.
This paper demonstrates that data selection in low-resource verification regimes, where verifiers only have access to fragmented and biased slices of the target distribution, can paradoxically accelerate model collapse by pruning globally relevant tail modes. The authors provide theoretical proof and propose a collaborative proxy reference mechanism as a mitigation strategy.
This paper introduces DyCon, a training-free framework that uses step-level embeddings to model evolving task difficulty and dynamically control reasoning depth in Large Reasoning Models, effectively reducing overthinking and improving efficiency without sacrificing accuracy.
This paper introduces OCLGen, a compute-efficient test-time search algorithm that integrates generative planning models with a classical Open-Closed List framework, improving solution quality across combinatorial planning domains.
This paper introduces Parameterized Diffusion Policy (PDP), a framework that makes diffusion policies controllable by conditioning on low-dimensional latent parameters, enabling smooth behavior interpolation and adaptation without retraining. It demonstrates improved performance on complex multimodal robot tasks in simulation and real-world experiments.
Proposes a node-level spectral energy formulation for detecting camouflaged anomalies in graphs, extending to spatio-temporal settings with energy-driven message passing. Demonstrates effectiveness on large-scale benchmarks.
LithoGRPO introduces a novel framework that combines flow matching with GRPO-based reinforcement learning for fast and high-quality inverse lithography mask optimization, achieving state-of-the-art performance while maintaining efficient generation.
This paper investigates memorization in diffusion models and finds that they preferentially memorize prototypical examples with common substrings, even after deduplication, and that early stopping leads to an overproduction of common motifs, dubbed 'slop'.
This paper introduces a formal definition of causal pathways for rare events and discusses testable implications, bridging simple verbal explanations with detailed causal models.