Tag
Z.ai completed a 1-gigawatt data center powered entirely by Chinese-made chips, expanding computing infrastructure for training its advanced GLM models.
Chinese factories are repurposing consumer RTX 5090 GPUs into 128GB server cards as a workaround to export controls on dedicated AI chips, creating a grey-market competitor to official enterprise offerings.
An alpha version of a Real Steel-like robot has been unveiled, reminiscent of the movie's boxing robots.
Apoorv Shankar raised $5.5M to build an AI hardware interface that understands user intent, aiming to replace touchscreens and keyboards.
This article explains how systolic arrays handle over 95% of AI chip compute, detailing their design, modes of operation, and why they are efficient for matrix multiplication.
Booster Robotics announces the Booster T2, a bipedal humanoid robot with 2070 TFLOPS of on-device computing power, enabling edge autonomy for perception, decision, and motion.
The article deeply deconstructs how CXL technology breaks the AI memory wall, analyzes memory bottlenecks from training to inference stages, and the application prospects of CXL in data centers.
SK Hynix, a major supplier of memory chips for Nvidia's AI accelerators, had a successful US IPO, opening at $170 per share and reaching a $1 trillion valuation amid surging demand for DRAM and HBM driven by AI data center buildout.
Nvidia's stock has fallen 15% from its peak as GPU shortage eases, while memory demand surges, making companies like Micron the new AI bottleneck.
The article argues that monopolies in the semiconductor supply chain (ASML, TSMC, etc.) create a bottleneck for AI development and hardware affordability, and suggests that breaking these monopolies could democratize computing and accelerate technological singularity.
Facing US export controls, Chinese AI startup DeepSeek is developing its own inference chips to reduce reliance on both Nvidia and Huawei, joining a trend of AI companies designing custom silicon.
A small team spent 9 months designing and assembling Prototype Version 0 of AI, software, and hardware in Paris at UMA_Robots.
A prediction that high-end consumer hardware may achieve Mythos-class AI capability within roughly two years, based on current trends.
Groq founder Jonathan Ross discusses how West Coast VCs missed investing in Groq due to herd mentality, contrasting with East Coast VCs who do independent analysis. The conversation also covers Groq's $20 billion partnership with NVIDIA.
Jetha Chan disassembled the Attention kernel for the SM100 data center GPU, identified and fixed inefficiencies, resulting in significant performance improvements.
Ahmad Osman predicts that within 18 months, a GPU like the RTX 5090 will be able to host intelligence equivalent to GLM 5.2.
Recent advancements in AI hardware, including custom chips from OpenAI, Etched, Amazon, and SambaNova, mark a significant shift towards specialized ASICs for AI workloads, promising major efficiency gains and challenging Nvidia's dominance.
Anthropic is developing its own AI inference chip and is in early talks with Samsung for its 2nm process, a move towards vertical integration in the AI race.
A user considers buying four Ascend GX10s to run GLM5.2, citing performance numbers like 400-500 tok/s prompt processing and ~15 tok/s output at 128k context, and plans for future open-source models.
Etched emerges from stealth with AI inference hardware, claiming state-of-the-art throughput and efficiency, posing serious competition to Nvidia.