Tag
A performance throttle in the Apple M3 Neural Engine was discovered when weight sizes are multiples of 1 MiB, causing throughput drops. By avoiding the problematic DMA path, token throughput for models like Llama 3.2 and Qwen3-8B was significantly improved.