parallel-computing

Tag

Cards List
#parallel-computing

@amitiitbhu: Just published: How do CUDA Kernels work? Read here:

X AI KOLs Timeline ↗ · 5d ago Cached

This article explains how CUDA Kernels work on GPUs, covering the basics of CUDA, GPU architecture, threads, blocks, and their role in AI and machine learning.

0 favorites 0 likes
#parallel-computing

@akshay_pachaar: How work is organized inside a GPU. A GPU does not treat a large computation as one job. It keeps dividing that job int…

X AI KOLs Timeline ↗ · 6d ago Cached

This article explains how GPUs organize computational work into kernels, grids, blocks, and warps to achieve massive parallelism and hide memory latency through streaming multiprocessors.

0 favorites 0 likes
#parallel-computing

Why the Real Power of AI Isn't Better Thinking—It's Mass Collaboration

Reddit r/artificial ↗ · 2026-09-24 Cached

The article argues that AI's true power lies in mass collaboration through multi-agent systems, exemplified by OpenAI's swarm architecture solving the Navier-Stokes problem rapidly.

0 favorites 0 likes
#parallel-computing

How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows

Hugging Face Blog ↗ · 2026-09-23 Cached

This article explains how to use NVIDIA Warp and MjWarp to scale robotics simulations on GPUs, enabling parallel environments for accelerated learning workflows.

0 favorites 0 likes
#parallel-computing

Automated optimization of a molecular simulation program

Hacker News Top ↗ · 2026-09-22 Cached

This article describes the automated optimization of the MBX molecular simulation program using compiler vectorization and parallel computation to enhance performance through SIMD instructions.

0 favorites 0 likes
#parallel-computing

Diego Almeida, fondateur de Typesafe AI, présente JEV,

Reddit r/artificial ↗ · 2026-09-18

Diego Almeida, fondateur de Typesafe AI, présente JEV, un nouveau modèle de base optimisé pour l'automatisation en temps réel. L'article explore son architecture de calcul parallèle et ses capacités dans des cas pratiques comme les simulations de jeux et la navigation par drone.

0 favorites 0 likes
#parallel-computing

Bend

Hacker News Top ↗ · 2026-09-17 Cached

Bend is a fast, parallel programming language that uses proofs to block AI mistakes, designed for AI-assisted development with C-like speed and CUDA parallelism.

0 favorites 0 likes
#parallel-computing

@VictorTaelin: you may not like Bend but you can't deny its users get the best job titles. metal benders, vulkan benders, proof bender…

X AI KOLs Following ↗ · 2026-09-15

Victor Taelin humorously comments on the job titles associated with users of the Bend programming language, such as metal benders and vulkan benders.

0 favorites 0 likes
#parallel-computing

@VraserX: The wild part about OpenAI’s Navier–Stokes result isn’t that one AI had a brilliant idea. They threw around 10,000 agen…

X AI KOLs Timeline ↗ · 2026-09-10 Cached

OpenAI used around 10,000 AI agents in parallel to tackle the Navier-Stokes equations, suggesting that scientific progress could scale significantly with compute power.

0 favorites 0 likes
#parallel-computing

Could Cargo's scheduler be better?

Lobsters Hottest ↗ · 2026-08-31 Cached

The article investigates improvements to Cargo's scheduler by analyzing its behavior through benchmarks of Rust build tasks and exploring alternative scheduling algorithms.

0 favorites 0 likes
#parallel-computing

The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches

Reddit r/LocalLLaMA ↗ · 2026-08-20

The article details a validated hardware configuration using 16 RTX 5060 Ti GPUs with PLX switches to run the Deepseek V4 Flash model, achieving specific performance metrics for context handling and throughput.

0 favorites 0 likes
#parallel-computing

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

arXiv cs.AI ↗ · 2026-08-19 Cached

KernelArc is a multi-agent framework that uses strategy-specialized agents to autonomously optimize GPU kernels across heterogeneous workloads, achieving top rankings on NVIDIA GPU benchmarks.

0 favorites 0 likes
#parallel-computing

Show HN: Rex, a parallel functional language for scientific workflows

Hacker News Top ↗ · 2026-08-17 Cached

Rex is a statically typed, pure functional workflow language designed for scientific computing and data processing, featuring parallel execution, content-addressable storage, and Docker isolation.

0 favorites 0 likes
#parallel-computing

The Free Lunch Is Over: A Fundamental Turn Toward Concurrency in Software (2005)

Lobsters Hottest ↗ · 2026-08-15 Cached

The article discusses the end of free performance gains from CPU speed increases due to physical limits, and emphasizes the fundamental shift in software development towards concurrency with the rise of multicore processors.

0 favorites 0 likes
#parallel-computing

CUDA Shared Memory Swizzling

Hacker News Top ↗ · 2026-08-13 Cached

The article explains CUDA shared memory swizzling techniques to optimize GPU memory access patterns, with code examples demonstrating performance improvements.

0 favorites 0 likes
#parallel-computing

Adaptivity via a Parallel Architecture for Stochastic Gradient Methods Adaptivity via a Parallel Architecture for Stochastic Gradient Methods Adaptivity via a Parallel Architecture for Stochastic Gradient Methods

arXiv cs.LG ↗ · 2026-08-03 Cached

The paper proposes a parallel architecture that assembles static gradient methods to achieve adaptivity in stochastic gradient descent, simplifying convergence analysis while retaining parameter adaptivity.

0 favorites 0 likes
#parallel-computing

@ycombinator: At our latest YC Paper Club, researchers and builders presented on multi-GPU kernels, intelligence per watt, heterogene…

X AI KOLs Timeline ↗ · 2026-07-29 Cached

Y Combinator hosted a Paper Club where researchers presented innovations in multi-GPU kernel optimization, including ParallelKittens, a CUDA framework that simplifies development of overlapped multi-GPU kernels and achieves significant speedups across workloads.

0 favorites 0 likes
#parallel-computing

@akshay_pachaar: CPU vs GPU vs TPU vs NPU vs LPU, explained visually: 5 hardware architectures power AI today. Each one makes a fundamen…

X AI KOLs Following ↗ · 2026-07-23 Cached

An explainer comparing five AI hardware architectures (CPU, GPU, TPU, NPU, LPU) with visual diagrams, covering their tradeoffs in flexibility, parallelism, and memory access for AI workloads.

0 favorites 0 likes
#parallel-computing

@v0xium: If you are looking for an article going in details regarding basics of CUDA, please spend an hour reading this. Link to…

X AI KOLs Timeline ↗ · 2026-07-20 Cached

An updated beginner-friendly tutorial on CUDA programming, covering how to write a simple kernel to add arrays on the GPU.

0 favorites 0 likes
#parallel-computing

@kalyan_kpl: What happens when you run a CUDA Kernel A CUDA kernel is a specialized code written to execute parallel computations on…

X AI KOLs Timeline ↗ · 2026-07-11 Cached

This article provides a detailed walkthrough of what happens when a CUDA kernel is compiled and executed on an NVIDIA GPU, covering compilation to PTX and SASS, and the underlying hardware interaction.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback