training-efficiency

Tag

Cards List
#training-efficiency

Memorizon: Training World Models Beyond Their Context Window

Hugging Face Daily Papers ↗ · 5d ago Cached

Memorizon decouples long-span supervision from attention in streaming world models by having each scored chunk retrieve a shared bounded top-K latent bank via camera co-visibility, enabling consistent revisits beyond the context window at minimal extra cost. Retrieval improves revisit consistency over a sliding-window baseline by up to 30%, and results show the model genuinely uses the retrieved latents.

0 favorites 0 likes
#training-efficiency

CARE: Condition-Aware Representation Regularization for Diffusion Models

arXiv cs.LG ↗ · 2026-09-25 Cached

The paper introduces CARE, a condition-aware representation regularization framework for diffusion models that improves sample quality and training efficiency by dynamically modulating feature distributions based on condition similarity. Empirically, it achieves significant reductions in FID and faster convergence for both class-to-image and text-to-image tasks.

0 favorites 0 likes
#training-efficiency

Training Text-to-Image Models 3.6× Faster

Hacker News Top ↗ · 2026-09-16 Cached

Linum AI introduces JiT-DDT, a novel encoder-decoder architecture that trains text-to-image models 3.6× faster than previous methods while generating images with higher resolution.

0 favorites 0 likes
#training-efficiency

XingChen-AGI/Xing4.0-29B-A4B

Hugging Face Models Trending ↗ · 2026-09-16 Cached

Xing4.0-29B-A4B is an open-source 29B-parameter large language model optimized for agent tasks and Ascend NPU, featuring a MoE architecture and achieving high training efficiency with competitive benchmark results.

0 favorites 0 likes
#training-efficiency

@askalphaxiv: “FlashREINFORCE: Critic-Free Single-Rollout Asynchronous RL for Agentic LMs” This paper argues that scalable agent RL d…

X AI KOLs Timeline ↗ · 2026-09-15 Cached

FlashREINFORCE introduces a critic-free, single-rollout asynchronous reinforcement learning framework for agentic language models, improving stability and efficiency by learning from each trajectory immediately.

0 favorites 0 likes
#training-efficiency

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

arXiv cs.LG ↗ · 2026-09-04 Cached

LeanGRPO eliminates redundant recomputation in diffusion RL methods by introducing recompute-free training schedules, achieving up to 1.83x speedup while preserving optimization objectives.

0 favorites 0 likes
#training-efficiency

Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation

arXiv cs.LG ↗ · 2026-09-03 Cached

This paper proposes a method to close the MLP reachability gap in Low-Rank Clone distillation by training the full deployed matrix, resulting in significant improvements in token efficiency and model performance at no additional inference cost.

0 favorites 0 likes
#training-efficiency

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

The paper investigates on-policy distillation of large language models, demonstrating that a single training query can achieve substantial state coverage and alignment, suggesting the method is algorithm-starved rather than data-starved.

0 favorites 0 likes
#training-efficiency

When Do Larger Batches Help Scale LLM Reinforcement Learning?

arXiv cs.LG ↗ · 2026-09-01 Cached

This paper examines whether larger batch sizes can reduce wall-clock time-to-target in reinforcement learning for large language models by separating algorithmic and systems-level effects, providing a decision rule for optimization.

0 favorites 0 likes
#training-efficiency

Scaling Muon for Diffusion Transformers

arXiv cs.LG ↗ · 2026-08-24 Cached

This paper scales the Muon optimizer for Diffusion Transformers from 1.3B to 15B parameters, introducing Periodic Row-wise Muon to reduce computational overhead while preserving generative quality improvements over AdamW.

0 favorites 0 likes
#training-efficiency

DeltaMomentum: A Key-Value based Anisotropic Momentum Update via Delta Rule

arXiv cs.LG ↗ · 2026-08-21 Cached

The paper proposes DeltaMomentum, a key-value based momentum update rule that adapts forgetting rates based on input frequency, demonstrating faster convergence in neural network training across various scales.

0 favorites 0 likes
#training-efficiency

Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training

arXiv cs.LG ↗ · 2026-08-19 Cached

Data-DPO is a target model-oriented data selection method for LLM supervised fine-tuning that learns data preferences through one-step probing and combines them with quality scores and diversity, outperforming baselines on Vision-Flan and LLaVA-CoT datasets.

0 favorites 0 likes
#training-efficiency

From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime

arXiv cs.LG ↗ · 2026-07-24 Cached

This paper develops a statistical framework for optimal noise-level allocation in diffusion model training, showing that the optimal schedule is atomic in the coupled regime and follows a square-root entropy proxy in the independent-learner regime, with experiments confirming these predictions.

0 favorites 0 likes
#training-efficiency

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

Hugging Face Daily Papers ↗ · 2026-07-16 Cached

LongStraw is an architecture-aware execution stack for million-token RL post-training under a fixed GPU budget, demonstrating execution capacity up to 2.1M positions on Qwen and GLM models.

0 favorites 0 likes
#training-efficiency

K-ABENA: K-Adaptive Backpropagation with Error-based N-exclusion Algorithm : (Compensated Loss-Based Sample Exclusion with Unbiased Gradient Estimation)

arXiv cs.LG ↗ · 2026-07-08 Cached

Introduces K-ABENA, a selective gradient computation framework that uses compensated loss-based sample exclusion with unbiased gradient estimation, proving convergence guarantees and showing compute savings of 28-54% without performance degradation across various datasets.

0 favorites 0 likes
#training-efficiency

LongCat-2.0

Product Hunt ↗ · 2026-07-07

LongCat-2.0 is a 1.6 trillion parameter mixture-of-experts model trained entirely on custom AI ASICs.

0 favorites 0 likes
#training-efficiency

The Context-Ready Transformer

arXiv cs.CL ↗ · 2026-06-29 Cached

The paper introduces the context-ready transformer, a recurrent architecture that pre-contextualizes tokens before the transformer block, achieving significant inference speedups (e.g., 1.7x on A100) while matching or exceeding standard transformer performance with fewer layers.

0 favorites 0 likes
#training-efficiency

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

Hugging Face Daily Papers ↗ · 2026-06-29 Cached

This paper proposes Nemotron-Labs-Diffusion-Image, a masked discrete diffusion model for high-resolution text-to-image synthesis, introducing a token-editing mechanism and grouped cross-entropy objective to improve token refinement and training efficiency.

0 favorites 0 likes
#training-efficiency

LISA: Likelihood Score Alignment for Visual-condition Controllable Generation

Hugging Face Daily Papers ↗ · 2026-06-25 Cached

This paper introduces LISA, a regularization method that aligns the intermediate features of a side network with an approximated likelihood score to improve training efficiency and the quality of visual-condition controllable generation in score-based generative models.

0 favorites 0 likes
#training-efficiency

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning

Hugging Face Daily Papers ↗ · 2026-06-23 Cached

Introduces Holistic Data Scheduler (HDS), a reinforcement learning-based framework that dynamically adjusts data mixtures during LLM pre-training using a multi-objective reward function, achieving 44% fewer iterations to reach target perplexity and a 7.2% improvement on MMLU.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback