Tag
A study on Large Language Models playing the game Diplomacy in multi-agent simulations reveals which models kept their promises when allowed to lie.
GPT-6-Astra-Max achieved 62.7% on the ARC-AGI-3 benchmark within six months, and with a memory adapter, the benchmark is saturated, though costs remain high, indicating a trend toward cheaper AI models.
The tweet announces the start of Interspeech 2026 in Sydney, highlighting the presentation of 2 tutorials and 19 papers on spoken language models and conversational speech recognition.
DSPy 3.4.0 is released with native support for Jev and System one models and a new optimizer, ReAnchor, for calibrating outputs with confidence.
This paper introduces Selective Supervision for Direct-OPD (S2D-OPD), a method that improves knowledge distillation by masking low-divergence states, enhancing accuracy on math reasoning benchmarks without extra computation.
ALOE introduces a semantically addressed low-rank operator for knowledge editing in language models, improving edit scope and precision with high efficacy and locality on standard benchmarks.
Introduces IndicBankBench, a 799-case benchmark for evaluating safety and reliability of language model assistants in Indian retail banking, with multi-stage evaluation and public release of code and data.
The paper proposes LastOPD, a method to prevent collapse in latent on-policy distillation by applying latent signals only at the last layer during a short crossfade period, leading to improved performance on benchmarks like MATH-500.
CounterRoute introduces an online reinforcement-learning framework that jointly learns routing and mode-conditioned responses in dual-mode language models, improving accuracy while reducing inference tokens.
The Stream Recursion Model (SRM) is a modification of the Hierarchical Reasoning Model that organizes computation into recursive latent streams to improve mechanistic interpretability for large language models, achieving performance comparable to GPT-2 per parameter.
PFArena introduces a benchmark for evaluating language models on protein modification tasks, comparing protein language models and large language models to assess their capabilities in biological applications.
This paper instruments a minimal FunSearch-style loop with operator packages to test components of proposers in verified search for mathematical construction problems, finding that composition closes the gap and repulsion increases diversity.
StepCOPS is a statistical framework for selecting language-model policies that uses closed-testing and lower-tail certificates to ensure safety guarantees with high probability, improving efficiency over conservative methods.
This paper presents a cross-model observational study on fixed-point structures to arbitrate between data-side and weights-side explanations of neural text degeneration, finding that the structural class is not determined by training data and varies across model architectures.
A controlled re-examination of ternary language models at 60K parameters reveals that baseline shape significantly impacts performance comparisons, challenging previous claims about the routed ternary model's advantage.
The paper investigates how parts-of-speech categories are encoded in Sparse AutoEncoder latent spaces, finding that they are distributed and not one-to-one with individual latents.
This paper introduces Corpus Task Complexity (CTC) to characterize how task difficulty scales with corpus size, presents high-CTC tasks, and releases CTC-Bench, showing that high-CTC tasks are more challenging for long-context language models.
The paper introduces Active Taskless Distillation (ATD), a method that transfers capabilities from a teacher model to a student model using only single-word responses on task-unrelated prompts, probing the behavioral shadows of post-training.
This paper identifies a failure mode where language models are persuaded by assertions from incentive-misaligned witnesses in CRM records, leading to incorrect decisions, and proposes a diagnostic method to analyze this issue.
A study interviewed four AI models on 24 subjects, recording 1,452 positions to archive their explicit views when pushed for consistency.