Tag
This paper introduces ActFirst-OPD, a framework that accelerates on-policy distillation for multi-turn language agents by decoupling action execution from full reasoning, achieving significant training speedups while maintaining performance across benchmarks.
This paper compares Fourier spectral differentiation and spatial automatic differentiation in periodic physics-informed neural networks, finding that Fourier methods achieve significant training speedups and memory reductions without compromising accuracy.
This paper investigates the 'small-vs-large gap', where training on fewer samples with more repetitions can lead to faster learning and compute savings compared to using larger datasets, attributing the speedup to layer-wise growth enabled by sampling biases. The findings suggest that smaller datasets with repetition can be proactively leveraged as favorable inductive biases, particularly in reasoning tasks.
A new training method achieves 2-3x speedup by allowing models to learn more flexibly in early stages, akin to homeschooling vs. factory education.
This paper introduces TwELL and Hybrid sparse formats with custom CUDA kernels to efficiently leverage unstructured sparsity in LLMs, achieving over 20% faster training and inference on H100 GPUs while reducing energy and memory usage.