metal-runtime

Tag

Cards List
#metal-runtime

BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal

arXiv cs.CL · 2026-07-02 Cached

BaseRT is a native Metal inference runtime for LLMs on Apple Silicon, achieving up to 1.56x higher decode throughput than llama.cpp and 1.35x higher than MLX across tested models.

0 favorites 0 likes
← Back to home

Submit Feedback