@Modular: .@dMatrix_AI had matrix multiplication running in Mojo on their Corsair accelerator 6 days after getting access to the …
Summary
At ModCon 2026, Modular showcased how its Mojo/Max technology stack can deploy the Gemma 4 31B model to Amazon Trainium, Google TPU, and other accelerators with virtually no changes, using small C interfaces and operator extensions — with dMatrix getting matrix multiplication running on its Corsair accelerator just 6 days after receiving access to the Mojo compiler.
View Cached Full Text
Cached at: 10/02/26, 08:47 PM
.@dMatrix_AI had matrix multiplication running in Mojo on their Corsair accelerator 6 days after getting access to the Mojo compiler.
At ModCon 2026, our Chief Scientist Abdul Dakkak walked through why this was possible with the Modular stack: https://t.co/GQd3sPwPIA
TL;DR:Modular’s Chief Scientist Abdul used cases like Trainium and TPUs to show that with just a small C interface and a few operators, the Mojo/Max stack can be extended to brand-new hardware in a matter of days — with virtually no changes to the model or the serving layer.
How Modular Extended Max Beyond GPUs
A talk by Modular’s Chief Scientist, Abdul, centered on a single question: as AI accelerators grow increasingly fragmented, how can the same software, the same model code, run reliably on more and more hardware? His answer was Modular’s modular tech stack, along with a methodology for bringing up new hardware.
Background: The Fragmentation of AI Software
Today’s AI software is deeply fragmented. Every accelerator comes with its own compiler, operator libraries, runtime, and its own separate development path. Hardware vendors integrate with one framework after another, and their contributions often devolve into a zoo of forks.
Ultimately, users bear the brunt of this mess — stitching together disparate components, taping things together just to deploy their own models. It’s hardly a pleasant experience.
Modular’s mission is to provide a unified compute layer — stable enough that users can deploy consistently, yet flexible enough to absorb new models, new modalities, and new hardware.
The Three-Layer Stack: Mojo, Max, Endpoints
The Modular platform is built from modular components, and the entire stack distills into three layers:
- Mojo — the portable programming language. Every Max operator that runs today is written in Mojo.
- Max — the model serving and framework layer: how you write models and serve them.
- Endpoints — exposing everything below as API endpoints, deployable either on Modular cloud or in the user’s own environment.
A key separation principle: bringing in new hardware should mainly affect the bottom of the stack — the rest of the stack stays untouched. This means the model and serving layers never need to be rebuilt for each individual accelerator.
The Methodology for Bringing Up New Hardware
1. Max Driver: Making Max Aware of the Hardware
To use any device, Max needs to know how to communicate with it — and that communication happens through the Max Driver. It provides the connection layer, handling things like memory transfer: how data is copied from the CPU to the accelerator, back from the accelerator to the CPU, function launches, synchronization operations, and so on.
All a team introducing new hardware needs to do is implement that small C interface (exposed via a C API). Once implemented, Max can load it in as a plugin.
2. Optional Device Driver Module Abstractions
At this point Max can see the hardware and communicate with it — but there’s no programming model and no performance yet. Sitting on top of the driver, the device driver module provides shared utilities and library abstractions that work across hardware.
These abstractions are optional — Modular wants you to use them, but doesn’t force it. There’s a real need for an escape hatch: when you want to fully exploit what the hardware does best, you need to expose target-hardware-specific capabilities. While this cuts against the abstraction philosophy and is rare, when you hit it you still want to stay in the same language rather than switching to another one.
Furthermore, all of these abstractions live in libraries, not hard-coded into the compiler. This lets Modular evolve quickly while still writing competitive operators that stand up to vendor implementations.
3. Validation: Three Cases of Escalating Difficulty
Modular validates this methodology in three ways:
- Their own team on Trainium;
- An external team on TPU;
- A chip vendor integrating its own stack.
In each case, the questions asked are: what changed? What was shared? How fast was the target reached? Those lessons feed back into the stack to make it more general.
The validation goal was running the Gemma 4 31B dense model (developed by Google; the visual encoder and speculative decoder were omitted from the demo) end to end. The key point: the model was written once, deployed to every target in the demo, and also runs on the six additional hardware platforms mentioned in the keynote. The model uses a modern architecture with sliding window attention, fully connected blocks, sampling operations, and more.
Case One: Trainium
Amazon Trainium is a matrix multiplication engine delivering up to 630 TFLOPs per chip (BF16), with roughly 96 GB of high-speed HBM memory at about 3 TB/s bandwidth, and 16 chips per node.
Trainium exposed three aspects that the stack needed to generalize:
- Static shapes: Max supports both dynamic and static shapes (batch size and sequence length are dynamic in LLMs), but static information wasn’t previously preserved throughout the operator implementation. The fix was to ensure static shape info is carried all the way through.
- Tiles: Trainium operates in tiles, while the prior abstractions were based on vectors (SIMD vectors). The stack was extended to support both 1-D SIMD vectors and 2-D tiled tensors.
- Key operator optimizations: Once the model ran, what remained was precise tuning of the two performance-critical operators: attention and matrix multiplication.
The model serving and model definition were entirely shared — it had already been validated on the B200, so porting to Trainium could be confirmed as a faithful implementation of the model. The MatMul code in the demo carries the same abstractions as on CPU and GPU, but specialized for Trainium: shared structure, without hiding hardware details.
Porting Timeline
- May: Two engineers took on the foundational module bring-up; vector addition ran in about a week.
- Team then scaled up to build the required Max operations, and by late July Gemma was functionally running, including collectives — since the model spans multiple chips and actually used tensor parallelism with a degree of 4.
- Max Serve required almost no changes, taking very little time.
- The final phase was performance optimization.
Performance Results
Matrix multiplication: Comparing the slide-sized Mojo MatMul code with existing best public implementations on Gemma 4’s shapes, it was consistently faster — at parity at worst, and in some cases up to ~1.6× faster than state of the art. Beating vendor implementations, usually written in assembly, is very hard.
Optimization velocity: With performance at the moment the model became functional (Day 0) as the 1× baseline, by Day 5:
- Time to first token improved ~50×
- Latency per output token improved ~7×
The point isn’t the numbers themselves, but the velocity: once the model runs, the shared infrastructure — tooling, profiling capabilities, related skills — lets teams optimize fast, identify bottlenecks, and iterate.
What Changed, What Didn’t
Changed: New communication capabilities in the device driver for Trainium, plus key operator implementations for matrix multiplication and attention.
Unchanged: the model definition, the Max Serve layer, everything from chunk prefill to speculative decoders — all exactly the same. Nothing unrelated to operators had to change. Modular says it will continue working with AWS to bring Trainium support into Max.
Case Two: An External Team Onboarding TPU (Htech)
Trainium showed the platform can scale to fundamentally different architectures — but that was built by Modular itself. The tougher test: could an external team do the same?
Modular partnered with Htech to verify exactly this. Htech compiler engineer Mihilo described the experience from an external partner’s perspective — especially with only limited support available from Modular’s engineering team. Htech is an engineering company with over 250 employees and a proven track record in onboarding new hardware platforms; it partnered with Modular a few months ago with the goal of bringing TPU support to Mojo Max.
TPU Architecture Characteristics
The TPU is a domain-specific accelerator built around large matrix multiplication units — an in-order execution machine, similar to a VLIW CPU, but with very wide registers. The v6e TPU used here includes:
- A single tensor core: two 256×256 matrix multiply systolic arrays (MXUs)
- One 8×128 vector unit (VPU)
- A single scalar core responsible for orchestration (DMA scheduling, address computation)
Compute units can only access data in scratchpads — VM (Vector Memory) and SPM (Scalar Memory), fully software-managed. Data must be explicitly DMA’d between scratchpads and HBM, where the “user” in this context effectively means the compiler.
Compared to GPUs: a GPU is a massively parallel machine with hundreds of SMs and thousands of threads, hiding memory latency dynamically in hardware. The TPU is an in-order machine with a small number of compute units, excellent at matrix math, but with a different memory hierarchy.
The Core Challenge: Software Stack Differences
GPUs have a dedicated LLVM backend; TPUs don’t. They use a domain-specific compiler — XLA — and a runtime called PGRT, which for example forbids pointer arithmetic outright and restricts the use of conventional control-flow constructs. Meanwhile, Mojo currently only lowers to LLVM IR — a major problem when the target hardware has no LLVM backend.
As an alternative to LLVM IR, the XLA compiler’s entry point is Mosaic — a collection of ML dialects, including upstream dialects like arith, func, and mem, plus a TPU-specific dialect for expressing TPU hardware capabilities. In the MatMul operator implemented by Mosaic, parameters are represented as shaped memrefs, loaded via vector dialect load ops, with the computation itself expressed through TPU dialect operations.
A notable trend: in domain-specific accelerators, ML is becoming the unifying layer, replacing LLVM IR. The Mojo developers were aware of this and provided a path for integration by exposing ML interop tooling.
Solution: The Mosaic Lowering Pipeline
Mojo offers ML interop through ML extensions, allowing TPU operations to be used directly inside operators. But such operators can’t be lowered to LLVM IR, so the Mojo compiler implementation had to be modified — deviating from the default lowering path to create a new Mosaic lowering pipeline, thereby compiling operators into Mosaic.
Mihilo also pointed out an unresolved issue: the contract between the host-side and device-side tensor implementations in Mojo is just a plain pointer with no shape information, which sits uneasily with the shaped memrefs that Mosaic requires. The original talk cut off at this point without a final conclusion.
Key Takeaways
Modular’s methodology can be distilled into a well-bounded path: implement a small C interface so Max can see the hardware → cover cross-hardware capabilities with optional library abstractions → apply hardware specialization only to critical operators like attention and matrix multiplication. The model definition, the serving layer, and the rest of the infrastructure stay unchanged, which is why the same Mojo/Max code can span NVIDIA, AMD, Qualcomm, Trainium, and more, running with the same commands. The Trainium case shows that once the model is functionally running, the shared tooling and profiling capabilities let teams drive large performance improvements within days.
Source: https://www.youtube.com/watch?v=hYjKbTAGOo0
Similar Articles
@dMatrix_AI: #ModCon2026 highlighted the momentum behind heterogeneous computing and open AI infrastructure. With @Qualcomm, d-Matri…
At ModCon 2026, Modular announced that Mojo 1.0 is fully open source under Apache 2.0, Modular Cloud is publicly available, and the platform now supports various AI accelerators including AWS Trainium, Google TPUs, and Qualcomm hardware.
@Modular: Today's AI software is fragmented. Every accelerator has its own compiler, kernel libraries, runtime, and development p…
Modular presents a unified compute layer to address AI software fragmentation, demonstrating faster hardware enablement through collaborations with HTEC and d-Matrix to extend MAX and Mojo to new accelerators like Google TPU.
@Modular: Mojo has minimal boilerplate, a strict type system, and compile-time validation of code, all things that make it well-s…
Modular releases open-source Mojo agent skills to help AI coding agents produce correct, idiomatic Mojo code, including a demo translating CUDA kernel code to Mojo. Mojo's minimal boilerplate, strict type system, and compile-time validation make it suitable for agentic workflows.
@Modular: Day two of AI Engineer World's Fair @aiDotEngineer! Meet the team at booth U-G28 to talk all things Mojo, MAX, and Modu…
Modular is at the AI Engineer World's Fair demonstrating Mojo, MAX, and Modular Cloud with fast image generation using FLUX.2 on DGX Spark, and real-time video generation.
@Modular: We're the #1 trending repo on @github today. The Mojo compiler went open source at ModCon this week, and developers not…
Modular announces that the Mojo compiler has gone open source, the platform now supports six hardware architectures including NVIDIA, AMD, and new ASICs, and Modular Cloud is officially launched.