Tag
The paper introduces a teacher-guided curriculum learning approach for Reinforcement Learning with Verifiable Rewards (RLVR) to efficiently train language models on initially unsolvable mathematical problems, achieving significant data efficiency and expanding reasoning boundaries.
The paper establishes a theoretical framework showing that data and memory resources in autoregressive prediction are governed by a single predictive-energy spectrum, providing a joint scaling law.
This paper investigates structural transfer from non-language data to improve data efficiency in language modeling, finding that while it reduces next-token prediction loss, it does not reliably enhance downstream linguistic performance.
DE-Venus is a unified framework for data-efficient Reinforcement Learning with Verifiable Rewards in large language models, reducing annotation and training costs while maintaining performance.
The paper shows that multilayer perceptrons naturally develop monosemantic specialized neurons that improve data efficiency by learning local low-dimensional representations instead of a global one in regression problems with clustered data.
The article highlights the data efficiency gap where children outperform AI models in learning, explores research to close this gap, and covers emerging careers in space travel alongside other tech developments like AI data center backlash and humanoid robot records.
Children learn language with far less data than AI models like LLMs, and understanding this data efficiency gap could lead to more efficient AI systems and insights into human cognition.
This paper presents C-Guard, a constitution-grid instrument for data-efficient RL alignment of safety guards, using C-LIM learnability scores to prune, densify, amend, or expand training data. It reduces over-refusal from 22.4% to 12.8% and improves data efficiency.
Introduces ReliableTableQA, a framework for training LLMs to annotate statistical reliability of tabular QA results, showing that a small SFT set is sufficient and GRPO only helps when SFT is under-trained.
ILLUME-X is a unified multimodal model for free-form interleaved text-image generation, featuring improved data efficiency, stable training, and a comprehensive evaluation metric called ILScore. It outperforms previous models on tasks like style transfer, image decomposition, and storytelling.
CLI-Universe is a synthesis engine that generates verifiable terminal-agent tasks via multi-dimensional capability taxonomy and evidence-guided research, producing a distilled dataset of 6,000 trajectories. Fine-tuning Qwen3-32B on this dataset achieves 33.4% on Terminal-Bench 2.0, setting a new state-of-the-art for open-source models at or below 32B parameters.
This paper proposes offline preference-based trajectory evaluation for agentic systems, which compares trajectories via temporal preferences rather than binary success metrics. It shows that this approach reduces ties from roughly 75% to 35%, improving discriminative power and data efficiency across diverse benchmarks.
APEX introduces a dynamic data selection strategy for automatic prompt optimization, stratifying datasets into easy, hard, and mixed tiers to improve data efficiency, achieving significant performance gains over initial prompts on multiple benchmarks.
This paper proposes a novel active learning framework that leverages foundation model priors to jointly address class imbalance and label noise, achieving over 50% annotation savings compared to baselines across image and text domains.
This paper proposes a framework for applying tabular foundation models to industrial time series for prognostics and health management, demonstrating strong performance and data efficiency across multiple PHM tasks.
This paper proposes a domain-aware coreset construction pipeline that enables a tabular foundation model to predict flood depth with only 0.7% of the training data, achieving 98.5% of the supervised reference accuracy and allowing transfer across watersheds without retraining.
This paper empirically measures the symmetry–data exchange rate predicted by equivariance theory, finding that wrong-group symmetry constraints are actively harmful, augmentation with test-time orbit averaging matches equivariant architectures, and the theoretical |G|-fold sample complexity reduction is only weakly confirmed with wide confidence intervals. The study is explicitly exploratory and not pre-registered.
This paper proposes Decoupled Residual Denoising Diffusion Models (DRDD) for unified and data-efficient image-to-image translation, decoupling noise diffusion for domain harmonization from residual diffusion for semantic mapping.
This paper investigates the 'small-vs-large gap', where training on fewer samples with more repetitions can lead to faster learning and compute savings compared to using larger datasets, attributing the speedup to layer-wise growth enabled by sampling biases. The findings suggest that smaller datasets with repetition can be proactively leveraged as favorable inductive biases, particularly in reasoning tasks.
This paper introduces a context optimization method that uses active information seeking via Wikipedia search and browser tools, combined with a search-based training procedure, to achieve robust performance improvements across diverse domains without updating model weights.