Tag
The author implemented a vectorized FMA for Rust's fearless_simd library, leading to performance improvements and the discovery of bugs in Rust and musl libc standard libraries.
A systematic exploration of FP32 matrix multiplication optimization on AMD Zen 3, achieving 85.30 GFLOPS (63.5% of theoretical peak) using AVX2/FMA intrinsics, surpassing naive implementation by 56.5x and matching optimized libraries.