fsdp2

Tag

Cards List
#fsdp2

KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models

arXiv cs.CL · 5d ago Cached

KDFlow is a novel knowledge distillation framework for large language models that uses a decoupled architecture with SGLang for teacher inference and FSDP2 for student training, achieving 1.44x to 6.36x speedup over existing frameworks.

0 favorites 0 likes
#fsdp2

HRM-Text: Trained on only 1k$ and 40b tokens with brain inspired hierarchical latent architecture

Reddit r/singularity · 2026-05-19 Cached

HRM-Text is a 1B parameter text generation model that uses a brain-inspired hierarchical recurrent architecture to achieve efficient pretraining with only 40B tokens and ~$1000, enabling accessible foundation model training with dramatically reduced compute and data requirements.

0 favorites 0 likes
← Back to home

Submit Feedback