icml

Tag

Cards List
#icml

@icmlconf: Announcing the #ICML2026 Awards! Including Outstanding Papers (research paper & position paper, winner & honorable ment…

X AI KOLs Timeline ↗ · 2026-07-05 Cached

ICML 2026 announces its award winners, including Outstanding Paper awards (research and position papers) and the Test of Time Award, with details on the selection process.

0 favorites 0 likes
#icml

@niclane7: Just in time for ICML week, we are sharing our take on a key question for recursive self-improving AI. How can AI keep …

X AI KOLs Timeline ↗ · 2026-07-05 Cached

The Red Queen Gödel Machine enables recursive self-improvement in AI by co-evolving the agent and evaluator, achieving better coding performance with fewer tokens.

0 favorites 0 likes
#icml

@jchudnov: Pass@k and self-consistency work great for math and code; sample more and verify. So we asked: can the same trick scale…

X AI KOLs Following ↗ · 2026-07-05 Cached

A new paper shows that scaling inference compute via methods like self-consistency improves LLM accuracy in math and code but fails to improve truthfulness in domains without external verifiers, as model errors are too correlated.

0 favorites 0 likes
#icml

@adiba_ejaz: How can causal (and statistical) models generalize to novel combinations of interacting objects? Our work w/ @eliasbare…

X AI KOLs Timeline ↗ · 2026-07-04 Cached

This paper presented at ICML explores how causal and statistical models can generalize to novel combinations of interacting objects, with a poster session scheduled at the conference.

0 favorites 0 likes
#icml

Dispersion loss counteracts embedding condensation in small language models

Hacker News Top ↗ · 2026-07-03 Cached

This paper observes that token embeddings in small language models condense into a narrow cone-like subspace, a phenomenon termed embedding condensation, and proposes a dispersion loss to counteract it, improving generalization.

0 favorites 0 likes
#icml

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

arXiv cs.CL ↗ · 2026-07-03 Cached

This paper proposes SpeechCombine, an instruction-following speech language model trained without any instruction tuning, using only speech pre-training and a weight combination strategy that transfers text LLM capabilities to the speech domain.

0 favorites 0 likes
#icml

How Should Transformers Encode Numeric Values in Electronic Health Records?

arXiv cs.LG ↗ · 2026-07-03 Cached

This paper systematically compares discrete, continuous, and hybrid value encoding strategies for transformers in electronic health record data, finding that hybrid token-based approaches with binning provide robust performance and are recommended as a practical default.

0 favorites 0 likes
#icml

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

arXiv cs.AI ↗ · 2026-07-02 Cached

This paper introduces OpenAgent, a problem setting for tool-use agents in open-world scenarios with distributional shifts, and proposes Perturbation-Augmented Fine-Tuning to improve robustness. Experiments reveal that both SFT and RL agents degrade under environmental shifts.

0 favorites 0 likes
#icml

@LiorOnAI: Most world models predict what happens next. Sora predicts pixels, JEPA compresses observations. NEO tries to figure ou…

X AI KOLs Following ↗ · 2026-07-01 Cached

NEO is a new type of world model that learns to discover reusable building blocks of explanation from raw observations without supervision or language, selected as an ICML 2026 oral presentation.

0 favorites 0 likes
#icml

Revisiting the Volume Hypothesis

arXiv cs.LG ↗ · 2026-07-01 Cached

This paper revisits the volume hypothesis, which posits that generalization in over-parameterized networks is mainly due to the larger volume of good-generalizing regions in weight space rather than SGD's implicit bias. Through experiments with binary networks, the authors show that the generalization advantage of gradient learning over random sampling diminishes as training data size grows, potentially resolving contradictory prior findings.

0 favorites 0 likes
#icml

@t0m1ab: Heading to ICML 2026 in Seoul next week with @romfbr31 to present Hibiki-Zero[https://kyutai.org/blog/2026-02-12-hibiki…

X AI KOLs Following ↗ · 2026-06-30 Cached

Kyutai presents Hibiki-Zero, a real-time speech-to-speech translation model, at ICML 2026 in Seoul, with an oral presentation scheduled for July 8.

0 favorites 0 likes
#icml

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

arXiv cs.CL ↗ · 2026-06-30 Cached

This paper presents the first comprehensive empirical study of safety impacts of benign multilingual fine-tuning on LLMs, showing that safety outcomes vary drastically by language and that assessing only English is insufficient.

0 favorites 0 likes
#icml

@MSFTResearch: AI agents can't remember past conversations. They must constantly reload or retrieve context, which grows less efficien…

X AI KOLs Following ↗ · 2026-06-29 Cached

Memora is a scalable memory system for AI agents that decouples storage from retrieval, enabling long-horizon tasks with up to 98% fewer context tokens while setting new state-of-the-art on benchmarks. The paper is published at ICML 2026.

0 favorites 0 likes
#icml

@HazanPrinceton: Just in time for our tutorial at ICML next week, Annie posted an update to our universal sequence preconditioning paper…

X AI KOLs Timeline ↗ · 2026-06-29 Cached

This paper update presents a universal sequence preconditioning method achieving dimension-free regret bounds for marginally stable linear dynamical systems, using second-order VAW algorithm and Faber polynomials.

0 favorites 0 likes
#icml

@MihaelaVDS: Can LLMs keep learning new skills without updating their weights? Modern LLMs can already master & combine many skills.…

X AI KOLs Timeline ↗ · 2026-06-29 Cached

Introduces 'skill neologisms', a method for enabling LLMs to learn new skills without weight updates, addressing catastrophic forgetting. Presented at ICML.

0 favorites 0 likes
#icml

Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge

arXiv cs.AI ↗ · 2026-06-29 Cached

This paper introduces LaViD, a framework that transfers semantic knowledge from a language-only LLM to a vision student model by generating multiple-choice questions as conceptual signatures, achieving superior fine-grained classification performance and robustness.

0 favorites 0 likes
#icml

@askalphaxiv: Looking for a fun weekend read? Introducing the Illustrated ICML We indexed all 6000+ ICML papers and built a visual wa…

X AI KOLs Timeline ↗ · 2026-06-26 Cached

A visual web app that indexes over 6000 ICML papers, allowing users to explore the paper landscape by topic.

0 favorites 0 likes
#icml

CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs

arXiv cs.CL ↗ · 2026-06-26 Cached

CAT-Q introduces a post-training ternary quantization method for LLMs that uses learnable modulation and softened ternarization, achieving superior performance over BitNet 1.58-bit while using only 512 calibration samples and scaling to 235B parameters.

0 favorites 0 likes
#icml

LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography

arXiv cs.AI ↗ · 2026-06-26 Cached

LithoDreamer is the first physics-informed World Model framework for computational lithography, modeling the multi-stage lithography process as a decision-driven system. It achieves state-of-the-art performance in forward evolution and inverse planning for semiconductor manufacturing.

0 favorites 0 likes
#icml

@prakashkagitha: As #ICML2026 is excitingly near, here are some of the most-cited papers and most-starred git repos: Code and workflows …

X AI KOLs Timeline ↗ · 2026-06-25 Cached

A tweet highlighting the most-cited papers and most-starred GitHub repos related to ICML 2026, with publicly available code and workflows to map papers to Semantic Scholar citation data and GitHub repos.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback