Tag
This article explains how CUDA Kernels work on GPUs, covering the basics of CUDA, GPU architecture, threads, blocks, and their role in AI and machine learning.
This article explains how GPUs organize computational work into kernels, grids, blocks, and warps to achieve massive parallelism and hide memory latency through streaming multiprocessors.
The article argues that AI's true power lies in mass collaboration through multi-agent systems, exemplified by OpenAI's swarm architecture solving the Navier-Stokes problem rapidly.
This article explains how to use NVIDIA Warp and MjWarp to scale robotics simulations on GPUs, enabling parallel environments for accelerated learning workflows.
This article describes the automated optimization of the MBX molecular simulation program using compiler vectorization and parallel computation to enhance performance through SIMD instructions.
Diego Almeida, fondateur de Typesafe AI, présente JEV, un nouveau modèle de base optimisé pour l'automatisation en temps réel. L'article explore son architecture de calcul parallèle et ses capacités dans des cas pratiques comme les simulations de jeux et la navigation par drone.
Bend is a fast, parallel programming language that uses proofs to block AI mistakes, designed for AI-assisted development with C-like speed and CUDA parallelism.
Victor Taelin humorously comments on the job titles associated with users of the Bend programming language, such as metal benders and vulkan benders.
OpenAI used around 10,000 AI agents in parallel to tackle the Navier-Stokes equations, suggesting that scientific progress could scale significantly with compute power.
The article investigates improvements to Cargo's scheduler by analyzing its behavior through benchmarks of Rust build tasks and exploring alternative scheduling algorithms.
The article details a validated hardware configuration using 16 RTX 5060 Ti GPUs with PLX switches to run the Deepseek V4 Flash model, achieving specific performance metrics for context handling and throughput.
KernelArc is a multi-agent framework that uses strategy-specialized agents to autonomously optimize GPU kernels across heterogeneous workloads, achieving top rankings on NVIDIA GPU benchmarks.
Rex is a statically typed, pure functional workflow language designed for scientific computing and data processing, featuring parallel execution, content-addressable storage, and Docker isolation.
The article discusses the end of free performance gains from CPU speed increases due to physical limits, and emphasizes the fundamental shift in software development towards concurrency with the rise of multicore processors.
The article explains CUDA shared memory swizzling techniques to optimize GPU memory access patterns, with code examples demonstrating performance improvements.
The paper proposes a parallel architecture that assembles static gradient methods to achieve adaptivity in stochastic gradient descent, simplifying convergence analysis while retaining parameter adaptivity.
Y Combinator hosted a Paper Club where researchers presented innovations in multi-GPU kernel optimization, including ParallelKittens, a CUDA framework that simplifies development of overlapped multi-GPU kernels and achieves significant speedups across workloads.
An explainer comparing five AI hardware architectures (CPU, GPU, TPU, NPU, LPU) with visual diagrams, covering their tradeoffs in flexibility, parallelism, and memory access for AI workloads.
An updated beginner-friendly tutorial on CUDA programming, covering how to write a simple kernel to add arrays on the GPU.
This article provides a detailed walkthrough of what happens when a CUDA kernel is compiled and executed on an NVIDIA GPU, covering compilation to PTX and SASS, and the underlying hardware interaction.