Tag
Discusses methods to run CUDA software on non-Nvidia GPUs, enabling use of Nvidia's AI ecosystem on other hardware.
The article analyzes the new ACE specification from the x86 Ecosystem Advisory Group, which extends Intel's AMX for AI matrix multiplication with fixed tile sizes and outer product instructions, comparing it to Arm's SME.
Chinese researchers have developed an all-optical interconnect system that links standard electronic chips, boosting AI distributed inference speeds by over 100 times while using just one-ninth of typical computational resources. The breakthrough, published in National Science Review, uses silicon photonic transceiver chips and FPGAs to achieve dramatic efficiency gains.
A curated list of resources for mastering GPU engineering for AI systems, covering CUDA, ROCm, optimization tools, multi-GPU orchestration, and distributed training.
A curated GitHub list of resources for learning GPU engineering, covering architecture, kernel programming, optimization, distributed systems, and AI acceleration with books, frameworks, profilers, and interview prep.
OpenAI and Broadcom unveiled Jalapeño, a custom LLM-optimized inference chip that promises substantially better performance per watt than current state-of-the-art, designed from the ground up for current and future AI models.
Tensordyne introduces Napier, an inference system using logarithmic math on silicon, claiming massive efficiency gains for MoE and reasoning models, with air-cooled racks.
Anthropic internal data shows Claude is accelerating AI development, potentially leading to a path of recursive self-improvement. The process of AI autonomously building a more powerful successor is happening faster than anticipated.
Anthropic's Institute publishes analysis on progress toward recursive self-improvement, showing AI is already accelerating AI development—engineers ship 8x more code per quarter—and projecting that AI systems capable of fully autonomous self-improvement could arrive sooner than most institutions are prepared for.
Terence Tao discusses how AI is dramatically accelerating math research, reducing the years of education needed to contribute to the frontier.
OWC announces Stack AI, a Thunderbolt 5 device that combines AI acceleration with storage expansion to enable running larger local AI models on Windows and Linux, reducing cloud dependency and improving privacy.
Chinese developer @yadong_xie notes that within a single year AI has shifted from merely helping write code to fully taking over rendering, calling global tech progress a 100× speed-up.
OpenAI releases a research paper documenting early experiments where GPT-5 accelerates scientific discovery across mathematics, physics, biology, and other fields, including cases where the model helped identify disease mechanisms, solve open problems, and improve optimization algorithms in collaboration with leading universities and national laboratories.
OpenAI has raised $6.6B in new funding at a $157B post-money valuation to accelerate frontier AI research, increase compute capacity, and expand access to advanced AI tools. The funding aims to scale the benefits of AI globally, with over 250 million weekly ChatGPT users already leveraging the platform.