Tag
The article questions why AI benchmarks are considered saturated at 90%+ accuracy and advocates for aiming at 100% or developing new benchmarks.
The cost of intelligence is dropping rapidly at 50% per quarter, outpacing declines in DNA sequencing, compute, lithium batteries, and electricity, with predictions of significant annual changes and AI accessibility improvements.
This article provides an update on the Engram model, detailing its 2.6b parameter architecture with a large Engram table and initial training progress at 100m tokens, showing improved completions with context-aware data offloading.
This paper characterizes the minimal recurrent behavioral memory required for imitating an expert under partial observability, using information-theoretic measures and experimental validation.
The paper introduces Modular Norm RandOpt, an architecture-aware perturbation method for efficient ensembling of language models, showing improved performance with fewer candidates across multiple tasks and model scales.
This paper investigates the grokking transition in neural networks using replica-overlap probes, but reports challenges with the probe's validity and offers post-hoc statistical analysis.
The paper investigates privileged context design in on-policy self-distillation, demonstrating that intermediate levels of abstraction can improve model performance over full solutions while using fewer hint tokens.
This paper introduces PermuFormer, an autoregressive transformer pretrained on multi-task, multi-encoding data for permutation-focused tasks in algebraic combinatorics, demonstrating effective fine-tuning on downstream tasks compared to baselines.
The paper identifies sequential reappearance as a failure mode in diffusion data-point unlearning and proposes a sharpness-guided method to improve forgetting persistence across deletion sequences.
The paper introduces the Differentiable Fuzzy Inference Layer (DFIL), a novel prediction head for large language models that enables monotonic and compositional ordinal reasoning without requiring compositional training data.
The paper proposes Informed Masking (IM), a structure-aware perturbation method for reinforcement learning in diffusion large language models that prioritizes downstream tokens, improving performance on math and planning benchmarks.
BELXTR is a novel multi-vector embedding model for biomedical entity linking that improves upon state-of-the-art methods by leveraging token-level matching, enhancing performance in disambiguation tasks.
Qwen3.8-Omni-Flash is a natively multimodal agentic model designed for real-world productivity, improving multimodal understanding and reasoning with a MoE architecture and extended context window, along with releasing open-source frameworks for multimodal applications.
ISA-Bench introduces a benchmark using programming games with constrained instruction sets to evaluate computational reasoning in large language models, revealing insights into model capabilities and a reasoning-execution gap.
This paper systematically evaluates input representations (images, text, or both) for multimodal document QA, finding that images improve accuracy but increase latency, and proposes a TF-IDF router to optimize performance.
This paper studies the interaction between prompt breadth and rollout refresh in on-policy distillation for mathematical reasoning, revealing that prompt efficiency depends on both the refresh rate of student policies and the inference budget.
This study adapts a human deliberation paradigm to large language models, finding that deliberation among diverse models reduces collective error and improves individual accuracy across multiple real-world domains, with diversity being crucial.
The article presents research on the rapid decline in AI performance costs, averaging a 47% quarterly drop across key benchmarks, highlighting a transformative economic trend for AI technology.
Perplexity's new research presents hint-guided self-distillation for post-training a Computer model, reducing tool-call failures by 21.2% in a live A/B test.
The article speculates on when large language models will reach a 'Navier moment' in physics, potentially leading to autonomous AI researchers that could revolutionize daily life on a monthly basis.