Tag
Nvidia's Jetson GPUs will be deployed on a Lunar Outpost rover destined for the moon, likely becoming the first GPU on the lunar surface to enable advanced autonomous navigation.
AMD will invest up to $5 billion in Anthropic as part of a partnership where Anthropic deploys AMD's Instinct MI450 GPUs for AI infrastructure. The companies also plan a multi-year engineering collaboration using Anthropic's Claude model.
Discusses a Latent Space podcast episode where Anjney Midha explains why AI labs with unlimited GPUs still fail, drawing on his experience at amppublic and a16z.
A benchmark comparison of 15 older GPUs considered e-waste, testing their performance on modern workloads.
NVIDIA's full-stack inference software, codesigned with hardware, has reduced token costs by up to 5x on the Blackwell platform in just one month, enabling lower cost per token for AI factories. Companies like Baseten, Cognition, Deep Infra, and Together AI are using the stack to optimize inference performance.
Valve is working with Intel and Nvidia to expand SteamOS support to more GPUs and handhelds, with initial firmware for Intel handhelds and ongoing driver work for Nvidia.
A researcher suggests it's time to buy more GPUs and build a local AI stack, referencing Qwen 3.5 27B and GLM 5.2 as models that cancel the threat of a permanent underclass.
A discussion on whether foundational AI research can be done without access to high-performance computing, given that early work like 'Attention is all you need' used consumer GPUs.
In a tweet, Sarah Hooker argues that GPUs are ill-suited for the long-tail distribution of real-world data, suggesting a need for alternative AI hardware.
The author argues that running local LLMs has become inaccessible due to high hardware costs, contrasting with earlier days when consumer GPUs sufficed, and expresses frustration with the perceived lack of democratic access.
CMU Software Engineering Institute publishes an overview of ML training infrastructure, covering hardware considerations like GPU vs CPU and memory requirements.
The article discusses how agentic AI may shift the computing focus back to CPUs from GPUs, citing OpenAI's CFO and Ark Invest's CEO. It argues that inference for agents involves orchestration and general-purpose tasks that CPUs handle better.
Hugging Face shares community hardware statistics showing the distribution of GPUs, CPUs, and Apple Silicon among its users.
Modal explains the four key ingredients they developed to spin up serverless GPU inference replicas in seconds instead of minutes, enabling efficient GPU allocation for variable AI workloads.
The article breaks down memory bandwidth as the critical metric for local AI hardware performance, comparing current GPUs and unified memory systems from NVIDIA, Apple, AMD, Intel, and others across different performance tiers.