@Modular: Today's AI software is fragmented. Every accelerator has its own compiler, kernel libraries, runtime, and development p…
Summary
Modular presents a unified compute layer to address AI software fragmentation, demonstrating faster hardware enablement through collaborations with HTEC and d-Matrix to extend MAX and Mojo to new accelerators like Google TPU.
View Cached Full Text
Cached at: 09/08/26, 11:39 PM
Today’s AI software is fragmented. Every accelerator has its own compiler, kernel libraries, runtime, and development path. Developers absorb the cost of this fragmentation.
In his ModCon 2026 tech talk, Abdul Dakkak, Chief Scientist at Modular, presents our alternative: a unified compute layer, flexible enough to extend to new models, modalities, and hardware.
Abdul walks through our recipe for bringing up new hardware, demonstrated with three examples:
-
AWS Trainium, via our own team. Gemma 4 31B end to end.
-
Google TPU v6e, through our partner @HTECgroup. Their team brought it up with no LLVM backend to lower to, no prior knowledge of Mojo or MAX internals, and minimal support from us.
-
d-Matrix Corsair, implemented by @dMatrix_AI’s own team on their own stack. Repo access Tuesday, working matmul by the following Monday.
Thanks to HTEC and d-Matrix for their collaboration, and to Mihailo for presenting HTEC’s learnings.
Read about the HTEC collaboration: https://htec.com/insights/media-coverage/htec-and-modular-extend-max-and-mojo-to-google-tpu-cutting-ai-hardware-enablement-from-years-to-months/…
Watch the full talk: https://youtube.com/watch?v=hYjKbTAGOo0…
HTEC and Modular Extend MAX and Mojo to Google TPU, Cutting AI Hardware Enablement from Years to Months | HTEC
Source: https://htec.com/insights/media-coverage/htec-and-modular-extend-max-and-mojo-to-google-tpu-cutting-ai-hardware-enablement-from-years-to-months/ Advanced compiler and runtime engineering validates a faster, reusable path for bringing AI workloads to new silicon without rebuilding the software stack
PALO ALTO, CA,September 8 2026– HTEC, a global AI-first engineering and technology services company, and Modular have demonstrated Modular’s AI software stack running on Google’s TPU architecture. Presented at ModCon 2026, the work shows that support for a new silicon target can be delivered by extending an existing software foundation rather than rebuilding it for each accelerator.
The project brought Modular’s MAX and Mojo ecosystem to Google TPU in months, a class of hardware enablement effort that has traditionally taken years. It provides practical validation of an alternative model for AI infrastructure: a portable software layer, combined with specialized compiler and runtime engineering, that can make new compute architectures usable faster.
As demand for AI compute grows, organizations are looking beyond a single hardware ecosystem to improve performance and scalability, while reducing exposure to cost, capacity and supply constraints. The barrier is often software. Each new accelerator can require extensive work across compilers, runtimes and inference systems before AI workloads can be used effectively.
Modular is addressing that barrier with technology designed to make AI workloads portable across silicon. HTEC applied its specialized compiler engineering and runtime systems expertise to extend that foundation to Google TPU, working from ambitious technical requirements and with significant engineering autonomy.
Darko Todorovic, CTO at HTECcommented: “We proved that Modular’s MAX and Mojo ecosystem could reach a new silicon target in months rather than years. That validates both Modular’s architecture and a practical path for organizations to adopt new AI hardware faster.”
Chris Lattner, CEO, Modular, added: “AI cannot afford a full software rebuild for every new accelerator. HTEC’s work on Google TPU shows how a common software foundation can bring new silicon into the AI ecosystem faster and at scale. That fundamentally changes the equation for silicon companies: instead of rebuilding an entire software stack, they can focus their engineering on what makes their hardware unique.”
Beyond the Google TPU milestone, the project validates HTEC’s long-term investment in AI hardware enablement and Modular’s approach to hardware-agnostic AI software. The experience further strengthens HTEC’s ability to support future semiconductor and AI inference initiatives across the industry.
About HTEC
HTEC Group Inc.is a global technology and engineering firm headquartered in Silicon Valley, providing AI-first, full-stack solutions that move organizations from strategy to production in a repeatable way. Combining premium engineering talent with digital strategy, design, and venture-building capabilities, HTEC supports the entire product lifecycle—from ideation to deployment.
With more than 15 years of experience applying AI and machine learning in enterprise environments, HTEC brings deep expertise across the technology stack—from silicon to application, including data engineering, model optimization, and AI application development. This ensures that emerging technologies translate into measurable ROI, scalable growth, and enduring business impact. By bridging the gap between vision and execution, HTEC helps the world’s leading companies define what’s possible—and scale it successfully.
Media contact: [email protected]
Similar Articles
@dMatrix_AI: #ModCon2026 highlighted the momentum behind heterogeneous computing and open AI infrastructure. With @Qualcomm, d-Matri…
At ModCon 2026, Modular announced that Mojo 1.0 is fully open source under Apache 2.0, Modular Cloud is publicly available, and the platform now supports various AI accelerators including AWS Trainium, Google TPUs, and Qualcomm hardware.
@Modular: Portability across CPUs, GPUs, and accelerators sounds simple until you've tried to build it. At #ModCon2026, @clattner…
Modular CEO Chris Lattner and Qualcomm CEO Cristiano Amon discussed portability across CPUs, GPUs, and accelerators at ModCon2026, emphasizing their collaboration to improve software for heterogeneous compute following Qualcomm's acquisition of Modular.
@PyTorch: New chips are shipping faster than ever before, but the ecosystem is being held back by having to rewrite and then re-d…
A keynote at PyTorch Conference North America will showcase an open software platform for heterogeneous compute powered by Mojo and MAX, addressing ecosystem challenges in deploying AI models across diverse hardware.
@Modular: Fragmentation creates a tax across the entire AI ecosystem. At @PyTorch Conference this year, @clattner_llvm presents a…
The PyTorch Conference North America in 2026 will feature a talk by Chris Lattner on a unified software stack to address fragmentation in the AI ecosystem, with registration details provided.
@clattner_llvm: I'm excited to share that Qualcomm is acquiring Modular: this will accelerate our path to unifying accelerated compute …
Qualcomm has agreed to acquire Modular, an AI infrastructure company focused on unifying accelerated compute with an open platform. This acquisition aims to advance Qualcomm's developer-first AI strategy and expand its software capabilities for AI from edge to cloud.