model-compilation

Tag

Cards List
#model-compilation

@vikhyatk: Got sick of hand-tuning GPU kernels, so we built a compiler. Photon 2.0 compiles Moondream, Qwen 3.5, and Gemma 4 into …

X AI KOLs Timeline ↗ · 2026-08-03 Cached

Photon 2.0 is a new inference engine and compiler that compiles models like Moondream, Qwen 3.5, and Gemma 4 into megakernels, claiming up to 2.3x throughput over vLLM and SGLang for physical AI workloads.

0 favorites 0 likes
#model-compilation

@dair_ai: NEW paper worth reading. A full agentic workflow can be distilled into model weights and run at roughly 100x lower infe…

X AI KOLs Following ↗ · 2026-05-22 Cached

This paper demonstrates that agentic workflows can be distilled into small fine-tuned models, achieving near-frontier quality while reducing inference cost by two orders of magnitude compared to orchestration approaches.

0 favorites 0 likes
← Back to home

Submit Feedback