throughput-optimization

Tag

Cards List
#throughput-optimization

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training

arXiv cs.AI · 2026-07-28 Cached

DynaResize is a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training stages during RL-based LLM post-training, improving throughput by 66.5% and reducing execution time by 33% over static configurations.

0 favorites 0 likes
#throughput-optimization

@danielhanchen: DeepSeek just released DSpark for V4 Flash & Pro, a new speculative decoding method boosting throughput by 51% to 400%!…

X AI KOLs Timeline · 2026-06-27 Cached

DeepSeek released DSpark, a speculative decoding method that boosts throughput by 51% to 400% for V4 Flash & Pro, along with the open-source DeepSpec codebase for training and evaluating draft models.

0 favorites 0 likes
#throughput-optimization

@raydistributed: Ray Serve LLM now offers 4.4x higher request throughput on prefill-heavy workloads, and 24.8x higher request throughput…

X AI KOLs Following · 2026-06-18 Cached

Ray Serve LLM achieves 4.4x and 24.8x throughput improvements on prefill- and decode-heavy workloads via direct streaming, a new vLLM V2 executor backend, and HAProxy ingress, now available in Ray 2.56 in partnership with Google Cloud and vLLM.

0 favorites 0 likes
#throughput-optimization

ServiceNow-AI/SuperApriel-15B-Instruct · Hugging Face

Reddit r/LocalLLaMA · 2026-04-22 Cached

ServiceNow releases SuperApriel-15B-Instruct, a single 15B checkpoint offering 8 mixer presets that trade between 1× and 10.7× decode throughput while maintaining up to 96% quality on 32K contexts.

0 favorites 0 likes
← Back to home

Submit Feedback