training-speedup

Tag

Cards List
#training-speedup

Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics

Hugging Face Daily Papers ↗ · 5d ago Cached

This paper introduces ActFirst-OPD, a framework that accelerates on-policy distillation for multi-turn language agents by decoupling action execution from full reasoning, achieving significant training speedups while maintaining performance across benchmarks.

0 favorites 0 likes
#training-speedup

A Computational Comparison of Fourier Spectral Differentiation and Spatial Automatic Differentiation in Periodic Physics-Informed Neural Networks

arXiv cs.LG ↗ · 2026-09-03 Cached

This paper compares Fourier spectral differentiation and spatial automatic differentiation in periodic physics-informed neural networks, finding that Fourier methods achieve significant training speedups and memory reductions without compromising accuracy.

0 favorites 0 likes
#training-speedup

Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases

arXiv cs.LG ↗ · 2026-05-21 Cached

This paper investigates the 'small-vs-large gap', where training on fewer samples with more repetitions can lead to faster learning and compute savings compared to using larger datasets, attributing the speedup to layer-wise growth enabled by sampling biases. The findings suggest that smaller datasets with repetition can be proactively leveraged as favorable inductive biases, particularly in reasoning tasks.

0 favorites 0 likes
#training-speedup

@AnjneyMidha: very cool a 2-3x speed up in training by essentially letting the model learn more flexibly in its early stages than rig…

X AI KOLs Following ↗ · 2026-05-14 Cached

A new training method achieves 2-3x speedup by allowing models to learn more flexibly in early stages, akin to homeschooling vs. factory education.

0 favorites 0 likes
#training-speedup

@hardmaru: The human brain is incredibly efficient because it only activates the specific neurons needed for a thought. Modern LLM…

X AI KOLs Timeline ↗ · 2026-05-08 Cached

This paper introduces TwELL and Hybrid sparse formats with custom CUDA kernels to efficiently leverage unstructured sparsity in LLMs, achieving over 20% faster training and inference on H100 GPUs while reducing energy and memory usage.

0 favorites 0 likes
← Back to home

Submit Feedback