Cached at:
09/20/26, 03:30 AM
# The last mile of a long road: faster NumPy in the browser
Source: [https://notebook.link/blog/the-last-mile-faster-numpy/](https://notebook.link/blog/the-last-mile-faster-numpy/)
For a long time, running NumPy in the browser meant running it without an accelerated BLAS\. Matrix multiplications fell back to plain loops \(portable, but blind to cache and SIMD\)\.
That just changed\. The[Emscripten\-forge](https://emscripten-forge.org/)NumPy package now links[OpenBLAS](https://www.openblas.net/)in WebAssembly, and at`n = 1024`square`np\.matmul`jumps to about**30\.92×**faster \(`float32`\) and**14\.90×**faster \(`float64`\)\. The next OpenBLAS release, already available as an experimental package on Emscripten\-forge, with kernels contributed by QuantStack, pushes it further, and an optional[Relaxed SIMD](https://github.com/WebAssembly/relaxed-simd)build adds another step on engines that support it\.
## What is Emscripten\-forge?[](https://notebook.link/blog/the-last-mile-faster-numpy/#what-is-emscripten-forge)
[Emscripten\-forge](https://emscripten-forge.org/)is a software distribution for WebAssembly\. Together with[conda\-forge](https://conda-forge.org/), it rebuilds the conda ecosystem for WebAssembly, so compilers, runtimes, and shared libraries ship as redistributable**conda packages**with a coherent ABI \(not Python wheels, not R packages, but native libraries that any language ecosystem can link\)\.
Bringing a scientific language to the browser is hard enough that, so far, it has mostly happened one language at a time\.[Pyodide](https://pyodide.org/)pioneered it for Python\. Later on,[WebR](https://docs.r-wasm.org/webr/latest/)did it for R, on the same premise\. Each is a self\-contained, language\-specific distribution\. Emscripten\-forge takes a different shape: a language\-agnostic distribution\. It builds on the foundational work of the Pyodide and WebR projects, and goes further, covering not only Python and R, but also C\+\+, OCaml, Lua, compiler toolchains, and many more tools \(and shipping with a package manager\)\.
## The missing compiler: Fortran on wasm32[](https://notebook.link/blog/the-last-mile-faster-numpy/#the-missing-compiler-fortran-on-wasm32)
Fortran sits at the core of foundational scientific packages:[LAPACK](https://www.netlib.org/lapack/), historically much of[SciPy](https://scipy.org/)\(although this is changing as SciPy has made strides to become Fortran\-free\), and R's compiled statistical routines are all Fortran underneath\. Bringing any of them to WebAssembly therefore means bringing a Fortran compiler that can target`wasm32`\.
For a while, Pyodide worked around the absence of such a compiler: its SciPy build was produced with[f2c](https://www.netlib.org/f2c/), a Fortran\-to\-C translator, so the Fortran source was transpiled to C and compiled with the existing Emscripten toolchain\. That workaround enabled SciPy to run in the browser, but it was not enough for R, whose distribution could not sidestep a real Fortran compiler\. That constraint is what started the work on[Flang](https://flang.llvm.org/)\(LLVM's Fortran frontend\) for WebAssembly, in the context of bringing[R to the browser on Emscripten\-forge](https://medium.com/jupyter-blog/r-in-the-browser-announcing-our-webassembly-distribution-9450e9539ed5)\.
## OpenBLAS in WebAssembly[](https://notebook.link/blog/the-last-mile-faster-numpy/#openblas-in-webassembly)
[OpenBLAS](https://www.openblas.net/)is a widely used, optimized implementation of[BLAS](https://www.netlib.org/blas/)\(Basic Linear Algebra Subprograms\)\. BLAS is organized in three levels:**Level 1**covers vector–vector operations \(for example AXPY and dot products\);**Level 2**covers matrix–vector operations \(for example GEMV, as in`A @ x`\);**Level 3**covers matrix–matrix operations \(for example GEMM, as in`np\.matmul`\), which dominate large dense linear algebra\. OpenBLAS also ships[LAPACK](https://www.netlib.org/lapack/)\(Linear Algebra PACKage\): higher\-level routines for solving linear systems, factorizations, and eigenproblems, and most`np\.linalg`calls ultimately delegate to those LAPACK entry points\.
Building a Fortran toolchain for WebAssembly was itself a multi\-person effort\.[Isabel Paredes](https://github.com/IsabelParedes)and[Serge Guelton](https://github.com/serge-sans-paille)wrote the initial Flang patches for WebAssembly, inspired by[George Stagg's work on Fortran in WebAssembly](https://gws.phd/posts/fortran_wasm/);[Serge Guelton](https://github.com/serge-sans-paille)also upstreamed parts of that work into LLVM\.[Isabel Paredes](https://github.com/IsabelParedes)now maintains the up\-to\-date recipe on Emscripten\-forge,[`flang\_emscripten\-wasm32`](https://github.com/emscripten-forge/recipes/tree/main/recipes/recipes/flang_emscripten-wasm32)\.
But even with a working Flang, the toolchain alone was not enough to bring OpenBLAS to the browser\. The first OpenBLAS and LAPACK builds could not be used immediately: WebAssembly enforces stricter calling conventions than native targets, and additional patches were needed to make the library compile, archive, and link under Emscripten\.[Ian Thomas](https://github.com/ianthomas23)adapted OpenBLAS for Flang and Emscripten; those patches live in the[Emscripten\-forge OpenBLAS recipe](https://github.com/emscripten-forge/recipes/tree/main/recipes/recipes_emscripten/openblas)\(the bridge between upstream OpenBLAS and the NumPy builds measured in this post\)\.
The rest of this post walks the last mile: what changed when NumPy could finally call BLAS, and what OpenBLAS 0\.3\.34 and the upcoming 0\.3\.35 deliver\.
## The last mile: NumPy links OpenBLAS[](https://notebook.link/blog/the-last-mile-faster-numpy/#the-last-mile-numpy-links-openblas)
The final step was linking NumPy to an accelerated BLAS\. The previous[Emscripten\-forge](https://emscripten-forge.org/)NumPy package \(`2\.5\.2`on[`emscripten\-forge\-4x`](https://prefix.dev/channels/emscripten-forge-4x/packages/numpy)\) shipped without OpenBLAS, so`np\.matmul`,`@`, and`np\.linalg`fell back to portable C loops: naive implementations blind to cache hierarchy and SIMD\. Matrix products spent most of their time on memory traffic rather than arithmetic\. This is our**no\-BLAS**baseline\.
As of[`emscripten\-forge/recipes\#6310`](https://github.com/emscripten-forge/recipes/pull/6310), the default NumPy on[`emscripten\-forge\-4x`](https://prefix.dev/channels/emscripten-forge-4x/packages/numpy)\(`2\.5\.3`, build 3\) links OpenBLAS**0\.3\.34**with WASM SIMD \(`TARGET=WASM128\_GENERIC`\)\. This is now a stable, main\-channel release\. Calls to`np\.matmul`,`@`, and`np\.linalg\.solve`now dispatch to the accelerated BLAS in WebAssembly\.
The packaging model is what makes this powerful\. Emscripten\-forge is, in effect, a**conda\-forge for WebAssembly**: packages are built from recipes, published on conda channels, and installed with a package manager under a shared ABI: the same workflow as on`linux\-64`, but targeting`emscripten\-wasm32`\. On that model, OpenBLAS is a separate conda package\. NumPy dynamically links`libopenblas`at runtime, so you can upgrade OpenBLAS without rebuilding NumPy\. This benefits the entire ecosystem: Scientific Python projects like SciPy and scikit\-learn, as well as non\-Python stacks such as[xtensor\-blas](https://github.com/xtensor-stack/xtensor-blas), all share the same library\. PyPI NumPy wheels, by contrast, vendor a BLAS snapshot at build time\. The same NumPy 2\.5\.3 package will thus automatically use OpenBLAS**0\.3\.35**once it's published\.
## OpenBLAS 0\.3\.34 versus no BLAS[](https://notebook.link/blog/the-last-mile-faster-numpy/#openblas-0334-versus-no-blas)
Relative to the same Emscripten\-forge stack without a BLAS implementation, square`np\.matmul`at`n = 1024`reaches about**30\.92×**\(`float32`\) and**14\.90×**\(`float64`\) \(28\.0 / 14\.6 GFLOPS versus 0\.90 / 0\.98 GFLOPS\)\. Full size grids and a single\-thread linux\-64 reference are in[Appendix F](https://notebook.link/blog/the-last-mile-faster-numpy/#Appendix-F-Full-benchmark-results-including-linux-64); focused graphs and tables for this comparison are in[Appendix B](https://notebook.link/blog/the-last-mile-faster-numpy/#Appendix-B-OpenBLAS-0334-versus-no-BLAS)\.

Geometric mean over both dtypes and the size grid \(OpenBLAS 0\.3\.34 time in the denominator\)\. Matrix products benefit most\. Some`np\.linalg\.\*`APIs see smaller gains \(1\.08–1\.65× versus no BLAS\) because they delegate to LAPACK, whose implementation in OpenBLAS is not yet optimized or specialized for WebAssembly\.`A @ x`and vector`np\.dot`remain near 1×: OpenBLAS 0\.3\.34 does not provide a fast column\-major GEMV for that layout\. OpenBLAS 0\.3\.35 adds that kernel\.
Without BLAS,`np\.matmul`throughput decreases with`n`\(naive GEMM, cache misses\)\. With OpenBLAS 0\.3\.34 it increases with`n`\. Operations implemented on top of GEMM \(`@`, two\-dimensional`np\.dot`,`np\.tensordot`,`np\.linalg\.multi\_dot`\) follow the same trend\. Speedup versus`n`and the associated numbers are in[Appendix B](https://notebook.link/blog/the-last-mile-faster-numpy/#Appendix-B-OpenBLAS-0334-versus-no-BLAS)\.
On`np\.linalg`at`n = 1024`\(`float32`\), median times drop by about**1\.2–2\.1×**versus no BLAS \(`solve`2\.11×,`cholesky`1\.88×,`qr`2\.03×,`inv`1\.66×,`eigh`1\.60×,`svd`1\.21×;`float64`agrees to within a few percent\)\. For matrix–vector products,`x @ A`already uses a Level\-2 kernel in 0\.3\.34 \(**14\.86×**at`float32`,`n = 1024`\), while`A @ x`stays near the no\-BLAS ~4 GFLOPS \(**1\.02×**\)\.
## OpenBLAS 0\.3\.35 will be even faster[](https://notebook.link/blog/the-last-mile-faster-numpy/#openblas-0335-will-be-even-faster)
The next OpenBLAS release includes WASM SIMD work already on`develop`\. The numbers below use the`openblas\-experimental`packages published by[`emscripten\-forge/recipes\#6761`](https://github.com/emscripten-forge/recipes/pull/6761)\(`0\.3\.35\.dev0`,`develop`@[`539bb47`](https://github.com/OpenMathLib/OpenBLAS/commit/539bb47f020d3278a05805f86c1ac63326b8a3d1), still`TARGET=WASM128\_GENERIC`\): portable**`simd128\_hbf81ecf\_3`**\(default\) and opt\-in**`relaxed\_simd\_h13a9a80\_3`**\.[WebAssembly Relaxed SIMD](https://github.com/WebAssembly/relaxed-simd)is an engine extension that allows slightly looser floating\-point semantics \(notably fused multiply\-add\) in exchange for faster vector math; OpenBLAS can use those opcodes when the build enables them\. Unless noted,**0\.3\.35**in this post means the portable SIMD128 build\.
QuantStack contributed the WASM SIMD kernels upstream to OpenBLAS \([Julien Jerphanion](https://github.com/jjerphan),[Matthias Meschede](https://github.com/MMesch)\), with review and integration by[Martin Kroeker](https://github.com/martin-frbg)\. That work covers Level\-3 GEMM/TRMM, Level\-1/2 AXPY and GEMV, numerical CBLAS tests for Node/Emscripten, and the optional Relaxed SIMD path\. Portable SIMD128 remains the default: Relaxed SIMD FMA stays off unless you opt into the`relaxed\_simd`build\. The headline`float32`GEMM step from 0\.3\.34 is an 8×4 microkernel; the optional Relaxed SIMD variant is covered in the next subsection\.
Versus WASM OpenBLAS 0\.3\.34 at`n = 1024`,`np\.matmul`improves by**1\.79×**\(`float32`, 28\.0 → 50\.0 GFLOPS\) and**1\.41×**\(`float64`, 14\.6 → 20\.5 GFLOPS\)\. Graphs and numbers for this comparison are in[Appendix C](https://notebook.link/blog/the-last-mile-faster-numpy/#Appendix-C-OpenBLAS-0335-SIMD128-versus-OpenBLAS-0334); the complete grid is in[Appendix F](https://notebook.link/blog/the-last-mile-faster-numpy/#Appendix-F-Full-benchmark-results-including-linux-64)\.

Geometric mean over both dtypes and the size grid \(OpenBLAS 0\.3\.35 time in the denominator\)\. The largest relative gains versus WASM 0\.3\.34 are not only in GEMM:**`A @ x`**jumps from the no\-BLAS ~4 GFLOPS to**23\.1 GFLOPS**\(`float32`,`n = 1024`\) via the GEMV kernels \(**5\.52×**versus 0\.3\.34;`x @ A`is**1\.73×**\)\. Vector`np\.dot`improves similarly\. Square`np\.matmul`is**1\.58×**in geometric mean \(and**1\.79×**at`n = 1024``float32`from the 8×4 SGEMM\)\. On`np\.linalg`at`n = 1024``float32`, median times improve by about**1\.3–1\.9×**versus 0\.3\.34 \(`solve`1\.38×,`cholesky`1\.32×,`qr`1\.49×,`inv`1\.33×,`eigh`1\.64×,`svd`1\.90×\)\.
### Optional: Relaxed SIMD on top of 0\.3\.35[](https://notebook.link/blog/the-last-mile-faster-numpy/#optional-relaxed-simd-on-top-of-0335)
Engines that implement[WebAssembly Relaxed SIMD](https://github.com/WebAssembly/relaxed-simd)can load the`relaxed\_simd`OpenBLAS build \(`WASM\_RELAXED\_SIMD=1`\), which uses FMA\-style`v\_muladd`from[\#6001](https://github.com/OpenMathLib/OpenBLAS/pull/6001)/[\#6020](https://github.com/OpenMathLib/OpenBLAS/pull/6020)\. On Chromium 153 in this bench, that is another**1\.17×**\(`float32`\) and**1\.43×**\(`float64`\) on`np\.matmul`at`n = 1024`versus portable 0\.3\.35 \(50\.0 → 58\.7 GFLOPS and 20\.5 → 29\.4 GFLOPS\)\. Geometric\-mean gains versus SIMD128 are largest on matrix products \(~1\.3×\); Level\-2 and most of`np\.linalg`move by about**1\.0–1\.2×**\. Graphs and numbers for this comparison are in[Appendix D](https://notebook.link/blog/the-last-mile-faster-numpy/#Appendix-D-OpenBLAS-0335-Relaxed-SIMD-versus-OpenBLAS-0335-SIMD128)\.

Chrome ≥114 and Firefox ≥146 expose the feature; Safari needs the JavaScriptCore flag`useWebAssemblyRelaxedSIMD`or the module fails to instantiate\. The portable`simd128`package remains the default for that reason \([`emscripten\-forge/recipes\#6761`](https://github.com/emscripten-forge/recipes/pull/6761)down\-prioritizes the Relaxed SIMD variant\)\. Current engine support is tracked at[webassembly\.org/features](https://webassembly.org/features/)\.
### From no BLAS to OpenBLAS 0\.3\.35[](https://notebook.link/blog/the-last-mile-faster-numpy/#from-no-blas-to-openblas-0335)
Putting the steps together \(linking OpenBLAS 0\.3\.34, then picking up the 0\.3\.35 SIMD128 kernels, then optionally Relaxed SIMD\) is the speedup path from the previous no\-BLAS Emscripten\-forge NumPy\.

Geometric mean over both dtypes and the size grid \(no\-BLAS time in the numerator\)\. Each API shows three bars:**OpenBLAS 0\.3\.34**,**OpenBLAS 0\.3\.35 \(SIMD128\)**, and**OpenBLAS 0\.3\.35 \(Relaxed SIMD\)**\. Matrix products gain the most; Level\-2`A @ x`and vector`np\.dot`finally move once the 0\.3\.35 GEMV kernels land;`np\.linalg\.\*`sits in a smaller but still clear band above 1×, limited by LAPACK routines that are not yet WASM\-specialized in OpenBLAS\. NumPy as published on Emscripten\-forge already delivers the OpenBLAS 0\.3\.34 half of that path; portable 0\.3\.35 \(`simd128\_hbf81ecf\_3`from[`emscripten\-forge/recipes\#6761`](https://github.com/emscripten-forge/recipes/pull/6761)\) is on the experimental channel today, with the optional Relaxed SIMD build \(`relaxed\_simd\_h13a9a80\_3`\) for supporting engines\. End\-to\-end`np\.matmul`at`n = 1024`reaches about**64\.89×**\(`float32`\) and**30\.07×**\(`float64`\) versus no BLAS with Relaxed SIMD\. Speedup versus`n`for this end\-to\-end comparison is in[Appendix E](https://notebook.link/blog/the-last-mile-faster-numpy/#Appendix-E-All-changes-OpenBLAS-0335-Relaxed-SIMD-versus-no-BLAS); the complete size grid, including linux\-64, is in[Appendix F](https://notebook.link/blog/the-last-mile-faster-numpy/#Appendix-F-Full-benchmark-results-including-linux-64)\.
## Using the packages[](https://notebook.link/blog/the-last-mile-faster-numpy/#using-the-packages)
NumPy linked against OpenBLAS is on the main[`emscripten\-forge\-4x`](https://prefix.dev/channels/emscripten-forge-4x)channel \(as of[`emscripten\-forge/recipes\#6310`](https://github.com/emscripten-forge/recipes/pull/6310)\)\. No experimental channel is required for that combination: install`numpy`and it pulls in OpenBLAS**0\.3\.34**\.
As an example, an environment for[JupyterLite](https://jupyterlite.readthedocs.io/)or[notebook\.link](https://notebook.link/):
```
name: numpy-openblaschannels: - https://prefix.dev/emscripten-forge-4x - https://prefix.dev/conda-forgedependencies: - xeus-python - numpy
```
OpenBLAS**0\.3\.35**\(`0\.3\.35\.dev0`\) is on[`emscripten\-forge\-4x\-experimental`](https://prefix.dev/channels/emscripten-forge-4x-experimental)as the dual\-variant packages from[`emscripten\-forge/recipes\#6761`](https://github.com/emscripten-forge/recipes/pull/6761)\(`openblas\-0\.3\.35\.dev0\-simd128\_hbf81ecf\_3`and`openblas\-0\.3\.35\.dev0\-relaxed\_simd\_h13a9a80\_3`\)\. Add the experimental channel when you want the newer OpenBLAS without rebuilding NumPy, and pin the build string explicitly if you need Relaxed SIMD:
```
name: numpy-openblas-035channels: - https://prefix.dev/emscripten-forge-4x - https://prefix.dev/emscripten-forge-4x-experimental - https://prefix.dev/conda-forgedependencies: - xeus-python - numpy - openblas * *simd128* # SIMD128 portable build (default) # - openblas * *relaxed_simd* # Relaxed SIMD build (not as portable, and deprioritized)
```
The portable`simd128`default does**not**require[WebAssembly Relaxed SIMD](https://github.com/WebAssembly/relaxed-simd)\. The`relaxed\_simd`package does: Chrome ≥114 and Firefox ≥146 work; Safari needs the JavaScriptCore flag`useWebAssemblyRelaxedSIMD`or the environment fails to load\.
## Conclusion[](https://notebook.link/blog/the-last-mile-faster-numpy/#conclusion)
NumPy on Emscripten\-forge can now call an accelerated BLAS in WebAssembly: linking OpenBLAS 0\.3\.34 already delivers large gains on matrix products and`np\.linalg`, and the upcoming 0\.3\.35 kernels \(with an optional Relaxed SIMD build\) push further\. Because that work landed upstream in OpenBLAS and as redistributable conda packages, the same improvements can benefit other WebAssembly Python stacks, including[Pyodide](https://pyodide.org/), when they adopt an OpenBLAS build and the Fortran toolchain pieces from Emscripten\-forge\.
## About the authors[](https://notebook.link/blog/the-last-mile-faster-numpy/#about-the-authors)
- [Julien Jerphanion](https://github.com/jjerphan)\(QuantStack\) contributed the WebAssembly kernels upstream to OpenBLAS, designed and ran the benchmarks in this post, added numerical tests upstream, and packaged and tested NumPy and SciPy in the[Emscripten\-forge OpenBLAS recipe](https://github.com/emscripten-forge/recipes/tree/main/recipes/recipes_emscripten/openblas)\.
- [Matthias Meschede](https://github.com/MMesch)\(QuantStack\) contributed to OpenBLAS to ensure that the WebAssembly kernels were indeed used, and provided help with package management\.
- [Ian Thomas](https://github.com/ianthomas23)\(QuantStack\) authored the initial patches to OpenBLAS that enabled its use by SciPy and other Emscripten\-forge packages, some of which may be upstreamed into the OpenBLAS project\.
## Acknowledgements[](https://notebook.link/blog/the-last-mile-faster-numpy/#acknowledgements)
Beyond the authors of this post, thanks to the people and communities who built the road this post only measures at the end:
- [Isabel Paredes](https://github.com/IsabelParedes)and[Serge Guelton](https://github.com/serge-sans-paille), for the Flang patches that made Fortran\-on\-`wasm32`practical \(building on[George Stagg's work](https://gws.phd/posts/fortran_wasm/)\);[Serge Guelton](https://github.com/serge-sans-paille)for upstreaming parts of that work into LLVM; and[Isabel Paredes](https://github.com/IsabelParedes)for maintaining[`flang\_emscripten\-wasm32`](https://github.com/emscripten-forge/recipes/tree/main/recipes/recipes/flang_emscripten-wasm32)on Emscripten\-forge\.
- [Martin Kroeker](https://github.com/martin-frbg), core[OpenMathLib / OpenBLAS](https://github.com/OpenMathLib/OpenBLAS)maintainer, for reviewing and integrating our contributions\.
- [Sylvain Corlay](https://github.com/sylvaincorlay),[Isabel Paredes](https://github.com/IsabelParedes), and[Antoine Pitrou](https://github.com/pitrou), for comments that improved this post\.
---
## Appendix A: Benchmark settings[](https://notebook.link/blog/the-last-mile-faster-numpy/#appendix-a-benchmark-settings)
WebAssembly numbers are median\-time speedups\. Tables also give median time and GFLOPS\. The**linux\-64**column is the same API grid on this machine\.
SettingValueHostDell XPS 15 9530, 13th Gen Intel Core i7\-13700H \(6P\+8E, 20 threads, AVX2, up to 5\.0 GHz\), 64 GB RAM, Fedora Linux 44WebAssemblyChromium 153; OpenBLAS`USE\_THREAD=0 TARGET=WASM128\_GENERIC`; packages from[`emscripten\-forge/recipes\#6761`](https://github.com/emscripten-forge/recipes/pull/6761):`simd128\_hbf81ecf\_3`\(default\) and`relaxed\_simd\_h13a9a80\_3`\(`WASM\_RELAXED\_SIMD=1`\)linux\-64[conda\-forge](https://conda-forge.org/)`linux\-64`NumPy 2\.5\.2 and OpenBLAS 0\.3\.34 \(`DYNAMIC\_ARCH`, Haswell kernel\), one thread \(`OPENBLAS\_NUM\_THREADS=1`\), pinned to a P\-coreMatrix sizes64, 128, 256, 512, 1024Vector sizes10⁴, 10⁵, 10⁶dtypes`float32`,`float64`Statisticmedian of 10 samples after 5 warmupsScripts[github\.com/jjerphan/numpy\-wasm\-openblas\-bench](https://github.com/jjerphan/numpy-wasm-openblas-bench)## Appendix B: OpenBLAS 0\.3\.34 versus no BLAS[](https://notebook.link/blog/the-last-mile-faster-numpy/#appendix-b-openblas-0334-versus-no-blas)
OpenBLAS**0\.3\.34**relative to the no\-BLAS Emscripten\-forge NumPy baseline\.
### Speedup versus input size[](https://notebook.link/blog/the-last-mile-faster-numpy/#speedup-versus-input-size)

APIndtype0\.3\.34 vs no BLAS`np\.matmul`64`float32`4\.54×`np\.matmul`64`float64`2\.56×`np\.matmul`128`float32`6\.88×`np\.matmul`128`float64`3\.56×`np\.matmul`256`float32`7\.67×`np\.matmul`256`float64`4\.17×`np\.matmul`512`float32`12\.04×`np\.matmul`512`float64`14\.16×`np\.matmul`1024`float32`30\.92×`np\.matmul`1024`float64`14\.90×`np\.linalg\.solve`64`float32`1\.03×`np\.linalg\.solve`64`float64`1\.06×`np\.linalg\.solve`128`float32`1\.45×`np\.linalg\.solve`128`float64`1\.48×`np\.linalg\.solve`256`float32`1\.95×`np\.linalg\.solve`256`float64`1\.91×`np\.linalg\.solve`512`float32`1\.98×`np\.linalg\.solve`512`float64`1\.95×`np\.linalg\.solve`1024`float32`2\.11×`np\.linalg\.solve`1024`float64`2\.10×`np\.linalg\.cholesky`64`float32`0\.96×`np\.linalg\.cholesky`64`float64`0\.99×`np\.linalg\.cholesky`128`float32`1\.22×`np\.linalg\.cholesky`128`float64`1\.20×`np\.linalg\.cholesky`256`float32`1\.51×`np\.linalg\.cholesky`256`float64`1\.51×`np\.linalg\.cholesky`512`float32`1\.59×`np\.linalg\.cholesky`512`float64`1\.59×`np\.linalg\.cholesky`1024`float32`1\.88×`np\.linalg\.cholesky`1024`float64`1\.85×`np\.linalg\.inv`64`float32`1\.10×`np\.linalg\.inv`64`float64`1\.11×`np\.linalg\.inv`128`float32`1\.29×`np\.linalg\.inv`128`float64`1\.26×`np\.linalg\.inv`256`float32`1\.63×`np\.linalg\.inv`256`float64`1\.62×`np\.linalg\.inv`512`float32`1\.56×`np\.linalg\.inv`512`float64`1\.64×`np\.linalg\.inv`1024`float32`1\.66×`np\.linalg\.inv`1024`float64`1\.65×`np\.linalg\.eigh`64`float32`1\.17×`np\.linalg\.eigh`64`float64`1\.16×`np\.linalg\.eigh`128`float32`1\.25×`np\.linalg\.eigh`128`float64`1\.23×`np\.linalg\.eigh`256`float32`1\.46×`np\.linalg\.eigh`256`float64`1\.42×`np\.linalg\.eigh`512`float32`1\.48×`np\.linalg\.eigh`512`float64`1\.50×`np\.linalg\.eigh`1024`float32`1\.60×`np\.linalg\.eigh`1024`float64`1\.58×`x @ A`64`float32`1\.68×`x @ A`64`float64`1\.26×`x @ A`128`float32`3\.15×`x @ A`128`float64`1\.93×`x @ A`256`float32`4\.04×`x @ A`256`float64`2\.21×`x @ A`512`float32`5\.12×`x @ A`512`float64`7\.23×`x @ A`1024`float32`14\.86×`x @ A`1024`float64`7\.53×`A @ x`64`float32`0\.88×`A @ x`64`float64`0\.94×`A @ x`128`float32`0\.97×`A @ x`128`float64`0\.96×`A @ x`256`float32`0\.99×`A @ x`256`float64`1\.00×`A @ x`512`float32`1\.03×`A @ x`512`float64`1\.05×`A @ x`1024`float32`1\.02×`A @ x`1024`float64`1\.06×## Appendix C: OpenBLAS 0\.3\.35 \(SIMD128\) versus OpenBLAS 0\.3\.34[](https://notebook.link/blog/the-last-mile-faster-numpy/#appendix-c-openblas-0335-simd128-versus-openblas-0334)
Portable OpenBLAS**0\.3\.35**\(SIMD128\) relative to WASM**0\.3\.34**\.
### Speedup versus input size[](https://notebook.link/blog/the-last-mile-faster-numpy/#speedup-versus-input-size-1)

APIndtype0\.3\.35 vs 0\.3\.34`np\.matmul`64`float32`1\.55×`np\.matmul`64`float64`1\.37×`np\.matmul`128`float32`1\.86×`np\.matmul`128`float64`1\.47×`np\.matmul`256`float32`1\.86×`np\.matmul`256`float64`1\.41×`np\.matmul`512`float32`1\.79×`np\.matmul`512`float64`1\.41×`np\.matmul`1024`float32`1\.79×`np\.matmul`1024`float64`1\.41×`np\.linalg\.solve`64`float32`1\.24×`np\.linalg\.solve`64`float64`1\.27×`np\.linalg\.solve`128`float32`1\.35×`np\.linalg\.solve`128`float64`1\.29×`np\.linalg\.solve`256`float32`1\.34×`np\.linalg\.solve`256`float64`1\.37×`np\.linalg\.solve`512`float32`1\.36×`np\.linalg\.solve`512`float64`1\.38×`np\.linalg\.solve`1024`float32`1\.38×`np\.linalg\.solve`1024`float64`1\.38×`np\.linalg\.cholesky`64`float32`1\.33×`np\.linalg\.cholesky`64`float64`1\.37×`np\.linalg\.cholesky`128`float32`1\.22×`np\.linalg\.cholesky`128`float64`1\.20×`np\.linalg\.cholesky`256`float32`1\.25×`np\.linalg\.cholesky`256`float64`1\.24×`np\.linalg\.cholesky`512`float32`1\.25×`np\.linalg\.cholesky`512`float64`1\.28×`np\.linalg\.cholesky`1024`float32`1\.32×`np\.linalg\.cholesky`1024`float64`1\.37×`np\.linalg\.inv`64`float32`1\.22×`np\.linalg\.inv`64`float64`1\.23×`np\.linalg\.inv`128`float32`1\.27×`np\.linalg\.inv`128`float64`1\.28×`np\.linalg\.inv`256`float32`1\.28×`np\.linalg\.inv`256`float64`1\.29×`np\.linalg\.inv`512`float32`1\.37×`np\.linalg\.inv`512`float64`1\.37×`np\.linalg\.inv`1024`float32`1\.33×`np\.linalg\.inv`1024`float64`1\.41×`np\.linalg\.eigh`64`float32`1\.36×`np\.linalg\.eigh`64`float64`1\.33×`np\.linalg\.eigh`128`float32`1\.48×`np\.linalg\.eigh`128`float64`1\.46×`np\.linalg\.eigh`256`float32`1\.55×`np\.linalg\.eigh`256`float64`1\.55×`np\.linalg\.eigh`512`float32`1\.62×`np\.linalg\.eigh`512`float64`1\.64×`np\.linalg\.eigh`1024`float32`1\.64×`np\.linalg\.eigh`1024`float64`1\.66×`x @ A`64`float32`1\.18×`x @ A`64`float64`1\.38×`x @ A`128`float32`1\.47×`x @ A`128`float64`1\.80×`x @ A`256`float32`1\.93×`x @ A`256`float64`2\.14×`x @ A`512`float32`1\.95×`x @ A`512`float64`1\.82×`x @ A`1024`float32`1\.73×`x @ A`1024`float64`1\.98×`A @ x`64`float32`1\.58×`A @ x`64`float64`1\.67×`A @ x`128`float32`3\.78×`A @ x`128`float64`2\.81×`A @ x`256`float32`6\.83×`A @ x`256`float64`3\.74×`A @ x`512`float32`6\.80×`A @ x`512`float64`3\.37×`A @ x`1024`float32`5\.52×`A @ x`1024`float64`3\.60×## Appendix D: OpenBLAS 0\.3\.35 \(Relaxed SIMD\) versus OpenBLAS 0\.3\.35 \(SIMD128\)[](https://notebook.link/blog/the-last-mile-faster-numpy/#appendix-d-openblas-0335-relaxed-simd-versus-openblas-0335-simd128)
Opt\-in`relaxed\_simd`OpenBLAS \(`WASM\_RELAXED\_SIMD=1`, Relaxed SIMD\) versus the portable`simd128`0\.3\.35 default on Chromium 153\.
### Speedup versus input size[](https://notebook.link/blog/the-last-mile-faster-numpy/#speedup-versus-input-size-2)

APIndtypeRelaxed SIMD vs SIMD128`np\.matmul`64`float32`1\.33×`np\.matmul`64`float64`1\.41×`np\.matmul`128`float32`1\.18×`np\.matmul`128`float64`1\.43×`np\.matmul`256`float32`1\.18×`np\.matmul`256`float64`1\.43×`np\.matmul`512`float32`1\.18×`np\.matmul`512`float64`1\.43×`np\.matmul`1024`float32`1\.17×`np\.matmul`1024`float64`1\.43×`np\.linalg\.solve`64`float32`1\.06×`np\.linalg\.solve`64`float64`1\.07×`np\.linalg\.solve`128`float32`1\.16×`np\.linalg\.solve`128`float64`1\.19×`np\.linalg\.solve`256`float32`1\.16×`np\.linalg\.solve`256`float64`1\.26×`np\.linalg\.solve`512`float32`1\.15×`np\.linalg\.solve`512`float64`1\.25×`np\.linalg\.solve`1024`float32`1\.25×`np\.linalg\.solve`1024`float64`1\.31×`np\.linalg\.cholesky`64`float32`0\.98×`np\.linalg\.cholesky`64`float64`0\.98×`np\.linalg\.cholesky`128`float32`1\.09×`np\.linalg\.cholesky`128`float64`1\.14×`np\.linalg\.cholesky`256`float32`1\.05×`np\.linalg\.cholesky`256`float64`1\.18×`np\.linalg\.cholesky`512`float32`1\.18×`np\.linalg\.cholesky`512`float64`1\.16×`np\.linalg\.cholesky`1024`float32`1\.03×`np\.linalg\.cholesky`1024`float64`1\.22×`np\.linalg\.inv`64`float32`1\.10×`np\.linalg\.inv`64`float64`1\.15×`np\.linalg\.inv`128`float32`1\.21×`np\.linalg\.inv`128`float64`1\.24×`np\.linalg\.inv`256`float32`1\.29×`np\.linalg\.inv`256`float64`1\.28×`np\.linalg\.inv`512`float32`1\.14×`np\.linalg\.inv`512`float64`1\.28×`np\.linalg\.inv`1024`float32`1\.18×`np\.linalg\.inv`1024`float64`1\.32×`np\.linalg\.eigh`64`float32`1\.02×`np\.linalg\.eigh`64`float64`1\.06×`np\.linalg\.eigh`128`float32`1\.10×`np\.linalg\.eigh`128`float64`1\.13×`np\.linalg\.eigh`256`float32`1\.05×`np\.linalg\.eigh`256`float64`1\.17×`np\.linalg\.eigh`512`float32`1\.21×`np\.linalg\.eigh`512`float64`1\.20×`np\.linalg\.eigh`1024`float32`1\.14×`np\.linalg\.eigh`1024`float64`1\.22×`x @ A`64`float32`1\.05×`x @ A`64`float64`1\.00×`x @ A`128`float32`1\.04×`x @ A`128`float64`1\.03×`x @ A`256`float32`1\.07×`x @ A`256`float64`1\.07×`x @ A`512`float32`1\.13×`x @ A`512`float64`1\.12×`x @ A`1024`float32`1\.18×`x @ A`1024`float64`1\.03×`A @ x`64`float32`1\.26×`A @ x`64`float64`1\.13×`A @ x`128`float32`1\.10×`A @ x`128`float64`1\.04×`A @ x`256`float32`1\.23×`A @ x`256`float64`1\.06×`A @ x`512`float32`1\.11×`A @ x`512`float64`0\.99×`A @ x`1024`float32`1\.29×`A @ x`1024`float64`1\.10×## Appendix E: All changes: OpenBLAS 0\.3\.35 \(Relaxed SIMD\) versus no BLAS[](https://notebook.link/blog/the-last-mile-faster-numpy/#appendix-e-all-changes-openblas-0335-relaxed-simd-versus-no-blas)
End\-to\-end speedup from the no\-BLAS Emscripten\-forge NumPy baseline to OpenBLAS**0\.3\.35**with Relaxed SIMD \(the full path: linking OpenBLAS, portable 0\.3\.35 kernels, then Relaxed SIMD\)\.
### Speedup versus input size[](https://notebook.link/blog/the-last-mile-faster-numpy/#speedup-versus-input-size-3)

APIndtypeRelaxed SIMD vs no BLAS`np\.matmul`64`float32`9\.35×`np\.matmul`64`float64`4\.92×`np\.matmul`128`float32`15\.10×`np\.matmul`128`float64`7\.46×`np\.matmul`256`float32`16\.85×`np\.matmul`256`float64`8\.44×`np\.matmul`512`float32`25\.49×`np\.matmul`512`float64`28\.50×`np\.matmul`1024`float32`64\.89×`np\.matmul`1024`float64`30\.07×`np\.linalg\.solve`64`float32`1\.34×`np\.linalg\.solve`64`float64`1\.43×`np\.linalg\.solve`128`float32`2\.26×`np\.linalg\.solve`128`float64`2\.27×`np\.linalg\.solve`256`float32`3\.04×`np\.linalg\.solve`256`float64`3\.29×`np\.linalg\.solve`512`float32`3\.10×`np\.linalg\.solve`512`float64`3\.36×`np\.linalg\.solve`1024`float32`3\.64×`np\.linalg\.solve`1024`float64`3\.78×`np\.linalg\.cholesky`64`float32`1\.25×`np\.linalg\.cholesky`64`float64`1\.33×`np\.linalg\.cholesky`128`float32`1\.61×`np\.linalg\.cholesky`128`float64`1\.65×`np\.linalg\.cholesky`256`float32`1\.98×`np\.linalg\.cholesky`256`float64`2\.22×`np\.linalg\.cholesky`512`float32`2\.35×`np\.linalg\.cholesky`512`float64`2\.36×`np\.linalg\.cholesky`1024`float32`2\.54×`np\.linalg\.cholesky`1024`float64`3\.08×`np\.linalg\.inv`64`float32`1\.47×`np\.linalg\.inv`64`float64`1\.56×`np\.linalg\.inv`128`float32`1\.99×`np\.linalg\.inv`128`float64`1\.99×`np\.linalg\.inv`256`float32`2\.69×`np\.linalg\.inv`256`float64`2\.67×`np\.linalg\.inv`512`float32`2\.44×`np\.linalg\.inv`512`float64`2\.89×`np\.linalg\.inv`1024`float32`2\.63×`np\.linalg\.inv`1024`float64`3\.08×`np\.linalg\.eigh`64`float32`1\.63×`np\.linalg\.eigh`64`float64`1\.63×`np\.linalg\.eigh`128`float32`2\.04×`np\.linalg\.eigh`128`float64`2\.03×`np\.linalg\.eigh`256`float32`2\.37×`np\.linalg\.eigh`256`float64`2\.59×`np\.linalg\.eigh`512`float32`2\.90×`np\.linalg\.eigh`512`float64`2\.95×`np\.linalg\.eigh`1024`float32`2\.99×`np\.linalg\.eigh`1024`float64`3\.20×`x @ A`64`float32`2\.09×`x @ A`64`float64`1\.74×`x @ A`128`float32`4\.80×`x @ A`128`float64`3\.61×`x @ A`256`float32`8\.33×`x @ A`256`float64`5\.06×`x @ A`512`float32`11\.30×`x @ A`512`float64`14\.78×`x @ A`1024`float32`30\.22×`x @ A`1024`float64`15\.35×`A @ x`64`float32`1\.76×`A @ x`64`float64`1\.76×`A @ x`128`float32`4\.05×`A @ x`128`float64`2\.82×`A @ x`256`float32`8\.31×`A @ x`256`float64`3\.95×`A @ x`512`float32`7\.79×`A @ x`512`float64`3\.50×`A @ x`1024`float32`7\.31×`A @ x`1024`float64`4\.17×## Appendix F: Full benchmark results \(including linux\-64\)[](https://notebook.link/blog/the-last-mile-faster-numpy/#appendix-f-full-benchmark-results-including-linux-64)
Absolute throughput and speedups for OpenBLAS**0\.3\.34**, OpenBLAS**0\.3\.35**\(SIMD128\), the optional**0\.3\.35 Relaxed SIMD**build, and a single\-thread conda\-forge**linux\-64**OpenBLAS 0\.3\.34 reference on the same machine\. Speedup is baseline time divided by featured time \(`\>1`means the featured stack is faster\)\.
### Geometric\-mean speedups by API[](https://notebook.link/blog/the-last-mile-faster-numpy/#geometric-mean-speedups-by-api)
API0\.3\.34 vs no BLAS0\.3\.34 vs linux\-640\.3\.35 vs 0\.3\.340\.3\.35 vs no BLAS0\.3\.35 vs linux\-64Relaxed SIMD vs 0\.3\.35Relaxed SIMD vs no BLASRelaxed SIMD vs linux\-64`np\.matmul`7\.68×0\.25×1\.58×12\.13×0\.39×1\.31×15\.92×0\.51×`np\.tensordot`7\.36×0\.26×1\.55×11\.45×0\.40×1\.28×14\.64×0\.51×`np\.linalg\.multi\_dot`7\.48×0\.25×1\.59×11\.87×0\.40×1\.30×15\.41×0\.51×`np\.linalg\.matrix\_power`7\.42×0\.25×1\.59×11\.79×0\.39×1\.31×15\.40×0\.51×`x @ A`3\.70×0\.38×1\.71×6\.33×0\.65×1\.07×6\.78×0\.70×`np\.linalg\.solve`1\.65×0\.45×1\.33×2\.20×0\.61×1\.18×2\.60×0\.72×`np\.linalg\.inv`1\.43×0\.42×1\.30×1\.87×0\.55×1\.22×2\.28×0\.66×`np\.linalg\.cholesky`1\.40×0\.46×1\.28×1\.79×0\.59×1\.10×1\.96×0\.65×`np\.linalg\.qr`1\.33×0\.39×1\.60×2\.12×0\.63×1\.13×2\.40×0\.71×`np\.linalg\.eigh`1\.38×0\.43×1\.52×2\.10×0\.65×1\.13×2\.37×0\.73×`np\.linalg\.lstsq`1\.24×0\.47×1\.66×2\.06×0\.77×1\.06×2\.19×0\.82×`np\.linalg\.svd`1\.08×0\.46×1\.66×1\.80×0\.77×1\.03×1\.86×0\.80×`A @ x`0\.99×0\.18×3\.56×3\.52×0\.65×1\.13×3\.97×0\.73×`np\.dot`0\.97×0\.26×3\.21×3\.11×0\.83×1\.01×3\.15×0\.84×### `np\.matmul`across sizes \(GFLOPS\)[](https://notebook.link/blog/the-last-mile-faster-numpy/#npmatmul-across-sizes-gflops)
ndtypeno BLASOpenBLAS 0\.3\.34OpenBLAS 0\.3\.350\.3\.35 Relaxed SIMDlinux\-640\.3\.34 vs no BLAS0\.3\.35 vs 0\.3\.340\.3\.35 vs no BLASRelaxed SIMD vs 0\.3\.35Relaxed SIMD vs no BLAS0\.3\.34 vs linux\-640\.3\.35 vs linux\-64Relaxed SIMD vs linux\-6464`float32`4\.74**21\.5****33\.3****44\.3**84\.74\.54×1\.55×7\.03×1\.33×9\.35×0\.25×0\.39×0\.52×64`float64`5\.01**12\.8****17\.5****24\.7**44\.12\.56×1\.37×3\.50×1\.41×4\.92×0\.29×0\.40×0\.56×128`float32`3\.57**24\.5****45\.6****53\.9**110\.46\.88×1\.86×12\.78×1\.18×15\.10×0\.22×0\.41×0\.49×128`float64`3\.76**13\.4****19\.7****28\.0**52\.93\.56×1\.47×5\.24×1\.43×7\.46×0\.25×0\.37×0\.53×256`float32`3\.38**25\.9****48\.0****56\.9**120\.57\.67×1\.86×14\.23×1\.18×16\.85×0\.21×0\.40×0\.47×256`float64`3\.42**14\.3****20\.2****28\.9**55\.54\.17×1\.41×5\.90×1\.43×8\.44×0\.26×0\.36×0\.52×512`float32`2\.29**27\.5****49\.4****58\.3**118\.612\.04×1\.79×21\.61×1\.18×25\.49×0\.23×0\.42×0\.49×512`float64`1\.02**14\.5****20\.4****29\.2**56\.814\.16×1\.41×19\.95×1\.43×28\.50×0\.26×0\.36×0\.51×1024`float32`0\.90**28\.0****50\.0****58\.7**120\.430\.92×1\.79×55\.30×1\.17×64\.89×0\.23×0\.42×0\.49×1024`float64`0\.98**14\.6****20\.5****29\.4**57\.214\.90×1\.41×20\.99×1\.43×30\.07×0\.25×0\.36×0\.51×### `np\.linalg`at`n = 1024`\(median time, ms\)[](https://notebook.link/blog/the-last-mile-faster-numpy/#nplinalg-at-n--1024-median-time-ms)
calldtypeno BLASOpenBLAS 0\.3\.34OpenBLAS 0\.3\.350\.3\.35 Relaxed SIMDlinux\-640\.3\.34 vs no BLAS0\.3\.35 vs 0\.3\.340\.3\.35 vs no BLASRelaxed SIMD vs 0\.3\.35Relaxed SIMD vs no BLAS0\.3\.34 vs linux\-640\.3\.35 vs linux\-64Relaxed SIMD vs linux\-64`np\.linalg\.solve``float32`122**58****42****34**212\.11×1\.38×2\.91×1\.25×3\.64×0\.36×0\.49×0\.61×`np\.linalg\.solve``float64`118**56****41****31**212\.10×1\.38×2\.90×1\.31×3\.78×0\.37×0\.51×0\.67×`np\.linalg\.cholesky``float32`68**36****27****27**161\.88×1\.32×2\.48×1\.03×2\.54×0\.46×0\.60×0\.62×`np\.linalg\.cholesky``float64`66**36****26****22**171\.85×1\.37×2\.52×1\.22×3\.08×0\.47×0\.64×0\.78×`np\.linalg\.qr``float32`552**272****183****142**962\.03×1\.49×3\.01×1\.29×3\.89×0\.35×0\.52×0\.68×`np\.linalg\.qr``float64`549**259****184****141**962\.12×1\.41×2\.98×1\.31×3\.89×0\.37×0\.52×0\.68×`np\.linalg\.inv``float32`357**215****161****136**701\.66×1\.33×2\.22×1\.18×2\.63×0\.33×0\.44×0\.52×`np\.linalg\.inv``float64`364**221****157****118**721\.65×1\.41×2\.33×1\.32×3\.08×0\.33×0\.46×0\.61×`np\.linalg\.eigh``float32`793**497****303****265**1711\.60×1\.64×2\.62×1\.14×2\.99×0\.34×0\.57×0\.65×`np\.linalg\.eigh``float64`778**491****297****243**1651\.58×1\.66×2\.62×1\.22×3\.20×0\.34×0\.56×0\.68×`np\.linalg\.svd``float32`569**471****248****253**2051\.21×1\.90×2\.29×0\.98×2\.25×0\.44×0\.83×0\.81×`np\.linalg\.svd``float64`571**487****258****235**2151\.17×1\.89×2\.21×1\.10×2\.43×0\.44×0\.83×0\.92×### Matrix–vector`@`across sizes \(GFLOPS\)[](https://notebook.link/blog/the-last-mile-faster-numpy/#matrixvector--across-sizes-gflops)
callndtypeno BLASOpenBLAS 0\.3\.34OpenBLAS 0\.3\.350\.3\.35 Relaxed SIMDlinux\-640\.3\.34 vs no BLAS0\.3\.35 vs 0\.3\.340\.3\.35 vs no BLASRelaxed SIMD vs 0\.3\.35Relaxed SIMD vs no BLAS0\.3\.34 vs linux\-640\.3\.35 vs linux\-64Relaxed SIMD vs linux\-64`x @ A`64`float32`2\.944\.94**5\.84****6\.14**12\.01\.68×1\.18×1\.99×1\.05×2\.09×0\.41×0\.49×0\.51×`x @ A`64`float64`2\.953\.72**5\.12****5\.12**9\.571\.26×1\.38×1\.74×1\.00×1\.74×0\.39×0\.54×0\.54×`x @ A`128`float32`3\.129\.81**14\.4****15\.0**27\.53\.15×1\.47×4\.62×1\.04×4\.80×0\.36×0\.52×0\.54×`x @ A`128`float64`3\.035\.85**10\.6****10\.9**18\.61\.93×1\.80×3\.49×1\.03×3\.61×0\.31×0\.57×0\.59×`x @ A`256`float32`3\.2713\.2**25\.6****27\.2**42\.04\.04×1\.93×7\.81×1\.07×8\.33×0\.31×0\.61×0\.65×`x @ A`256`float64`3\.237\.13**15\.3****16\.3**25\.12\.21×2\.14×4\.73×1\.07×5\.06×0\.28×0\.61×0\.65×`x @ A`512`float32`2\.6913\.8**26\.9****30\.4**43\.25\.12×1\.95×10\.00×1\.13×11\.30×0\.32×0\.62×0\.70×`x @ A`512`float64`1\.027\.35**13\.4****15\.0**14\.17\.23×1\.82×13\.17×1\.12×14\.78×0\.52×0\.95×1\.06×`x @ A`1024`float32`0\.9514\.2**24\.4****28\.8**30\.214\.86×1\.73×25\.68×1\.18×30\.22×0\.47×0\.81×0\.95×`x @ A`1024`float64`1\.007\.50**14\.9****15\.3**14\.37\.53×1\.98×14\.94×1\.03×15\.35×0\.53×1\.04×1\.07×`A @ x`64`float32`2\.912\.56**4\.06****5\.12**10\.70\.88×1\.58×1\.39×1\.26×1\.76×0\.24×0\.38×0\.48×`A @ x`64`float64`2\.892\.72**4\.53****5\.10**9\.790\.94×1\.67×1\.57×1\.13×1\.76×0\.28×0\.46×0\.52×`A @ x`128`float32`4\.053\.94**14\.9****16\.4**24\.20\.97×3\.78×3\.67×1\.10×4\.05×0\.16×0\.61×0\.68×`A @ x`128`float64`4\.013\.85**10\.8****11\.3**19\.90\.96×2\.81×2\.70×1\.04×2\.82×0\.19×0\.55×0\.57×`A @ x`256`float32`4\.264\.22**28\.8****35\.4**38\.50\.99×6\.83×6\.77×1\.23×8\.31×0\.11×0\.75×0\.92×`A @ x`256`float64`4\.194\.20**15\.7****16\.6**28\.61\.00×3\.74×3\.75×1\.06×3\.95×0\.15×0\.55×0\.58×`A @ x`512`float32`4\.124\.26**28\.9****32\.1**42\.61\.03×6\.80×7\.01×1\.11×7\.79×0\.10×0\.68×0\.75×`A @ x`512`float64`4\.074\.28**14\.4****14\.2**15\.21\.05×3\.37×3\.54×0\.99×3\.50×0\.28×0\.95×0\.94×`A @ x`1024`float32`4\.084\.18**23\.1****29\.8**29\.51\.02×5\.52×5\.66×1\.29×7\.31×0\.14×0\.78×1\.01×`A @ x`1024`float64`4\.024\.25**15\.3****16\.7**14\.91\.06×3\.60×3\.80×1\.10×4\.17×0\.28×1\.02×1\.12×