Tag
Researchers from Nanyang Technological University, Cornell, and Bristol introduce PowerZooJax, an open-source JAX-based benchmark suite for reinforcement learning in power system operation, providing five constrained MDP tasks (generation, transmission, distribution, DERs, data center microgrids) that run fully on GPU with substantial speedups over CPU-based simulators.
The article explains subtle differences in struct layouts between Rust and WebGPU, and introduces a unit testing approach using naga to verify memory offsets for safer GPU compute shader configuration.
The article discusses using the Qwen 27B model on a 4090 GPU to create motion graphics, highlighting local AI capabilities inspired by Opus 5.5.
An agentic CUDA kernel optimizer that automates GPU implementation generation through iterative code generation, correctness checks, benchmarking, and refinement, powered by LangGraph and OpenAI models.
vgpu is a TypeScript library for WebGPU, featuring typed shader imports, a tiny GPU-first API, and compatibility across browser, headless Node, and test environments.
A promotional post for TensorTonic, an online platform for learning machine learning through coding practice, starting with an explanation of the RBF kernel.
A user successfully runs the Qwen 3.8 27B AI model on a mixed setup of RTX 3060 and 5060 Ti GPUs using tensor parallelism with exllamav3, achieving around 50 tokens per second with MTP enabled.
This article explains in detail MoE (Mixture of Experts) inference engineering, corrects misconceptions about activated parameters and deployment costs, and delves into technical details such as router selection, runtime grouping, GPU execution, memory management, and expert parallelism.
The author demonstrates forcing AMD MI50 GPUs to work with ROCm 6.4.3 and building a Vulkan-based neural network training stack, challenging the consensus on legacy hardware support.
Periodic Labs trained an open-weight AI model using 1,300 NVIDIA H200 GPUs and internal lab data, surpassing GPT-6 Astra on their internal benchmarks.
The author is building a dual GPU system for running large language models, evaluating AMD Radeon AI Pro R9700 versus NVIDIA RTX 3090 options while aiming to reduce AI subscription costs.
Kanverse GPU Borrow is a product that enables devices to temporarily borrow GPU compute from other machines with explicit user authorization, using Blink Bridge and orchestrated by GPT-6 Astra.
Researchers from Cognition used the AI agent Devin to factor RSA-260, setting a new record in the RSA Factoring Challenge and demonstrating the potential of autonomous software engineering in cryptographic research.
Modal is a serverless platform that operates on top of the GPU layer, simplifying GPU access for developers.
A promotional post for Modal, an AI infrastructure platform featuring SDK, autoscaling, and production-ready tools, along with an announcement for the Runtime conference for AI engineers.
A user announces they have solved all attention-related problems on LeetGPU and compiled them into a resource for easy access.
The article describes experiments on split-K matrix multiplications, revealing that answer variability depends on block layouts and split counts, with tests on a B200 GPU showing output changes under different conditions.
The article describes a story contest inviting people to imagine a future where GPUs are abundant, making AI accessible to everyone by 2040, and explores potential societal impacts.
This article provides a practical guide on running 8x NVIDIA RTX PRO 6000 Blackwell GPUs for AI workloads, emphasizing high-concurrency inference, model fleets, and 70B fine-tuning, while comparing performance to more expensive data center solutions.
SlopTV is an AI livestream that uses Minimax H3 on dual 5090 GPUs to generate infinite content from YouTube chat comments.