Tag
The paper introduces CaLR, a framework that reformulates reasoning as constrained latent optimization using causal topology to enhance diffusion language models, achieving state-of-the-art performance on complex benchmarks.
Hugging Face shares findings from a community hackathon where 1,200 participants used coding agents to reproduce 2,226 ICML papers, highlighting issues in review rigor and the potential for AI agents to scale verification.
A blog post arguing that LLMs could potentially achieve scientific breakthroughs like General Relativity through deductive, non-abductive routes, countering Tom Zahavy's 'LLMs can't jump' position paper.
Hugging Face hosted a live broadcast discussing how AI agents reproduced ICML 2026 papers.
This paper investigates subliminal learning in language models, showing that biases can transfer from teacher to student via seemingly random synthetic data. The authors find that adding Gaussian noise to weights increases transfer, and that students inherit not just the semantic bias but also the type of intervention used, with implications for training safety and data auditing.
This paper proposes a dataset-centric meta-evaluation framework that audits LLM benchmarks at the sample level across five latent dimensions, exposing internal heterogeneity and enabling criterion-driven composition of benchmark subsets for targeted model evaluation.
Researchers present a paper at ICML arguing that a fundamental flaw in how LLMs identify instructions makes them impossible to fully secure against attacks, demonstrating successful exploits against models from OpenAI, Anthropic, Alibaba, and DeepSeek.
VeriSimpl introduces a solver-LLM framework that uses simplification-based verification to ensure correct translation of natural language optimization problems into solver formulations, achieving improved accuracy over existing methods.
DecodeShare proposes a method to identify a low-dimensional subspace consistently shared across tasks in LLM decode-time hidden states and shows that disturbing this subspace degrades decision performance more than random or prefill-derived subspaces, with implications for activation steering.
This paper introduces C-MTP, a direct supervision method for training continuous chain-of-thought models that compresses reasoning traces into latent representations. The method performs competitively on simple tasks but reveals that both direct and indirect supervision methods struggle with complex long reasoning traces, showing about 65% performance drop.
Introduces RELIC, a framework for learning interpretable and composable skills in multi-agent planning via revealed principles, enabling privacy-preserving coordination and cross-agent skill transfer without sharing code.
New research shows that LLMs can develop their own biases from experience and stereotype job applicants more than humans, raising concerns about AI in hiring.
Introduces M+Adam, a novel optimizer combining additive and multiplicative updates to enable effective low-precision training of LLMs using BF16, FP8, or FP4 master weights, avoiding failure modes of standard optimizers.
Arvind Narayanan's ICML 2026 keynote argues that the 'AI as Normal Technology' framework is useful for understanding AI's impacts, rejects the notion of sudden job loss from AI, and envisions a future of human-AI co-superintelligence requiring significant adaptation.
A paper titled 'Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity' has been accepted to ICML. It proposes a simple prompt-engineering trick for more diverse sampling, sparking debate over whether such work belongs at a top-tier ML conference.
Pengrui Han's paper received the Best Paper Award at the ICML Combining Theory and Benchmarks Workshop, with congratulations from Anima Anandkumar.
ConOrd proposes a contrastive learning framework for ordinal regression that integrates contrastive learning and order learning, achieving state-of-the-art performance on facial age estimation, image quality assessment, and video quality assessment.
This paper reformulates rank estimation with noisy ordinal labels as a stochastic ordering problem and proposes a learning framework (SOL) that captures ordinal label uncertainty through discriminative and stochastic order losses, achieving reliable rank estimation under various noise types.
This paper introduces MemExplainer, a method to explain predictions of Temporal Graph Networks (TGNs) by attributing contributions through topology attribution trees and memory backtracking trees, using Layer-wise Relevance Propagation (LRP) for faithful explanations.
Flex-Forcing introduces a unified framework for video diffusion that supports both autoregressive and bidirectional generation modes, offering flexible control for video generation tasks.