Tag
Memorizon decouples long-span supervision from attention in streaming world models by having each scored chunk retrieve a shared bounded top-K latent bank via camera co-visibility, enabling consistent revisits beyond the context window at minimal extra cost. Retrieval improves revisit consistency over a sliding-window baseline by up to 30%, and results show the model genuinely uses the retrieved latents.
The paper introduces CARE, a condition-aware representation regularization framework for diffusion models that improves sample quality and training efficiency by dynamically modulating feature distributions based on condition similarity. Empirically, it achieves significant reductions in FID and faster convergence for both class-to-image and text-to-image tasks.
Linum AI introduces JiT-DDT, a novel encoder-decoder architecture that trains text-to-image models 3.6× faster than previous methods while generating images with higher resolution.
Xing4.0-29B-A4B is an open-source 29B-parameter large language model optimized for agent tasks and Ascend NPU, featuring a MoE architecture and achieving high training efficiency with competitive benchmark results.
FlashREINFORCE introduces a critic-free, single-rollout asynchronous reinforcement learning framework for agentic language models, improving stability and efficiency by learning from each trajectory immediately.
LeanGRPO eliminates redundant recomputation in diffusion RL methods by introducing recompute-free training schedules, achieving up to 1.83x speedup while preserving optimization objectives.
This paper proposes a method to close the MLP reachability gap in Low-Rank Clone distillation by training the full deployed matrix, resulting in significant improvements in token efficiency and model performance at no additional inference cost.
The paper investigates on-policy distillation of large language models, demonstrating that a single training query can achieve substantial state coverage and alignment, suggesting the method is algorithm-starved rather than data-starved.
This paper examines whether larger batch sizes can reduce wall-clock time-to-target in reinforcement learning for large language models by separating algorithmic and systems-level effects, providing a decision rule for optimization.
This paper scales the Muon optimizer for Diffusion Transformers from 1.3B to 15B parameters, introducing Periodic Row-wise Muon to reduce computational overhead while preserving generative quality improvements over AdamW.
The paper proposes DeltaMomentum, a key-value based momentum update rule that adapts forgetting rates based on input frequency, demonstrating faster convergence in neural network training across various scales.
Data-DPO is a target model-oriented data selection method for LLM supervised fine-tuning that learns data preferences through one-step probing and combines them with quality scores and diversity, outperforming baselines on Vision-Flan and LLaVA-CoT datasets.
This paper develops a statistical framework for optimal noise-level allocation in diffusion model training, showing that the optimal schedule is atomic in the coupled regime and follows a square-root entropy proxy in the independent-learner regime, with experiments confirming these predictions.
LongStraw is an architecture-aware execution stack for million-token RL post-training under a fixed GPU budget, demonstrating execution capacity up to 2.1M positions on Qwen and GLM models.
Introduces K-ABENA, a selective gradient computation framework that uses compensated loss-based sample exclusion with unbiased gradient estimation, proving convergence guarantees and showing compute savings of 28-54% without performance degradation across various datasets.
LongCat-2.0 is a 1.6 trillion parameter mixture-of-experts model trained entirely on custom AI ASICs.
The paper introduces the context-ready transformer, a recurrent architecture that pre-contextualizes tokens before the transformer block, achieving significant inference speedups (e.g., 1.7x on A100) while matching or exceeding standard transformer performance with fewer layers.
This paper proposes Nemotron-Labs-Diffusion-Image, a masked discrete diffusion model for high-resolution text-to-image synthesis, introducing a token-editing mechanism and grouped cross-entropy objective to improve token refinement and training efficiency.
This paper introduces LISA, a regularization method that aligns the intermediate features of a side network with an approximated likelihood score to improve training efficiency and the quality of visual-condition controllable generation in score-based generative models.
Introduces Holistic Data Scheduler (HDS), a reinforcement learning-based framework that dynamically adjusts data mixtures during LLM pre-training using a multi-objective reward function, achieving 44% fewer iterations to reach target perplexity and a 7.2% improvement on MMLU.