Tag
Memory now accounts for 63% of AI accelerator costs, up from 52% in early 2024, shifting industry priorities toward data movement optimization over process shrinks.
Samsung unveiled zHBM and zNAND-O memory technologies at the FMS 2026 conference, designed to stack memory directly on AI accelerators for up to 8x higher performance and improved power efficiency in AI systems.
Modular CEO Chris Lattner and Qualcomm CEO Cristiano Amon discussed portability across CPUs, GPUs, and accelerators at ModCon2026, emphasizing their collaboration to improve software for heterogeneous compute following Qualcomm's acquisition of Modular.
Mojo, a programming language designed for AI accelerators and GPUs, is now fully open source under the Apache 2.0 license, with its compiler and toolchain available on GitHub for development and contribution.
This repository presents a proof of concept for a hardware-native optical timing-frozen control plane engine that uses fluid dynamics modeling in GPU registers to reduce jitter in optical data centers for distributed AI architectures.
Chinese companies are increasingly allocating AI accelerator budgets to domestic suppliers like Huawei and Hygon, reducing reliance on Nvidia amid US-China tensions. A Bloomberg survey shows 46% of spending will go to domestic products in the next year, up from 30%.
John Carmack comments on memory cost and capacity issues for AI accelerators, noting that model inference can have deterministic memory access patterns, contrasting with game rendering.
QuixiAI released QuixiCore, a family of native high-performance AI kernel libraries for modern accelerators, with standalone implementations for CUDA, Metal, ROCm, XPU, and Gaudi backends, all sharing a common contract but no shared code.
At least seven Chinese companies are shipping H100/H200-class AI accelerators, most having recently IPO'd, with several founded by former NVIDIA/AMD architects. Huawei's Ascend 950 targets H200-class performance, and China's domestic market share is rising as NVIDIA's declines.
A user asks about buying Chinese AI accelerators/GPUs for inference, specifically looking for Huawei alternatives to Nvidia, with support for vLLM or Llama.cpp.
KForge is a cross-platform framework that uses two collaborating LLM-based agents to automatically generate and optimize high-performance compute kernels for diverse AI accelerators, achieving significant speedups on NVIDIA B200 and Intel Arc B580 hardware.
This paper introduces TRAM, a method that jointly optimizes approximate multiplier structures and AI model parameters to reduce power consumption in AI accelerators while maintaining accuracy.
AccelOpt is a self-improving LLM agentic system that autonomously optimizes AI accelerator kernels through iterative generation and optimization memory, achieving 49-61% peak throughput improvements on AWS Trainium while being 26x cheaper than Claude Sonnet 4.