Tag
AsmEvo is an agentic assembly-level optimizer for AMD GPU kernels that improves performance by proposing low-level edits and verifying functional equivalence against original binaries, achieving speedups up to 3.88x on MI308X GPUs.
The author shares their work over 6-8 months in ML systems and AI infrastructure, including a lightweight Python LLM inference engine (tachyon) that achieves 600+ tokens/s on consumer hardware with continuous batching and prefix caching, alongside blog posts on CUDA/CUTE DSL and collective communication, and contributions to SGLang and vLLM.
A new book from CMU's Machine Learning Systems course teaches modern GPU programming for ML systems, covering Blackwell architecture, GEMM, and FlashAttention using the TIRx Python DSL.
A curated video-guided curriculum and comprehensive list of resources for learning ML systems and LLM infrastructure, including papers, courses, and tutorials.
Kyle Kingsbury discusses emerging roles of human accountability in ML systems, including content moderators, legal representatives, and compliance officers who may bear responsibility for AI system failures.