Tag
A new GPU-native parallel optimizer, ChiSao, for multimodal black-box functions that uses convergence-anticonvergence oscillation to find all modes. It achieves 100% mode recovery and up to 34x speedup over baselines on benchmark functions.
Achieved 400 tokens per second with DeepSeek-V4-Flash-FP8 using 8 parallel aggregates on local hardware, marking a significant milestone for local LLM inference.