Tag
Describes using ncnn's Vulkan backend for vendor-agnostic ML inference on production edge devices, achieving 10x speedup over CPU ONNX for face detection and embedding models.
Onepot AI can synthesize and deliver custom molecules in just 5 days by combining robotic synthesis with large-scale ML inference on Anyscale, dramatically accelerating drug discovery.
This benchmark compares Gemma 4's Multi-Token Prediction (MTP) and z-lab's DFlash speculative decoding methods on a single H100 GPU, showing MTP faster for dense models and DFlash faster for MoE models.