Tag
DynaResize is a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training stages during RL-based LLM post-training, improving throughput by 66.5% and reducing execution time by 33% over static configurations.
DeepSeek released DSpark, a speculative decoding method that boosts throughput by 51% to 400% for V4 Flash & Pro, along with the open-source DeepSpec codebase for training and evaluating draft models.
Ray Serve LLM achieves 4.4x and 24.8x throughput improvements on prefill- and decode-heavy workloads via direct streaming, a new vLLM V2 executor backend, and HAProxy ingress, now available in Ray 2.56 in partnership with Google Cloud and vLLM.
ServiceNow releases SuperApriel-15B-Instruct, a single 15B checkpoint offering 8 mixer presets that trade between 1× and 10.7× decode throughput while maintaining up to 96% quality on 32K contexts.