Tag
PyTorch blog publishes a Ray-focused guide to PyTorch Conference North America 2026 in San Jose, highlighting keynotes and sessions on co-evolving Ray and Kubernetes, elastic training stacks, and scalable RL from speakers at Anyscale, Google, LinkedIn, Pinterest, and Uber.
This article demonstrates how to pool RAM from old devices to run a larger local AI model and build a private chatbot using custom knowledge from PDFs, ensuring all data remains local.
The article explores the idea of building a community-owned AI to compete with corporate frontiers, referencing open-source software models and suggesting technologies like Petals for distributed GPU computation, while highlighting organizational over technical hurdles.
DeepSeek Elastic Compute (DSec) is a sandbox infrastructure for effective agentic training of large language models at scale, featuring elastic execution with unified SDK and integration with reinforcement learning frameworks.
The article explores the concept of decentralized AI as an alternative to centralized systems, reviewing research on federated learning and distributed computing that could enable collective ownership of AI.
PyTorch 2.14 release introduces fault tolerance as a first-class c10d concept with in-place process-group reconfiguration, alongside features like NVGEMM kernels, a new nccl2 backend, and native linear algebra for Apple Silicon.
NVIDIA announced PAIR at IFA, a tool that automatically links local network systems to distribute inference requests for efficient AI agent operation.
Nvidia has launched PAIR, an open-source tool that links idle home computers to create a personal AI data center for distributed local AI inference tasks, supporting agentic workflows.
This paper proposes a manifold-aware encoding strategy for general coded computing that exploits the intrinsic low-dimensional geometry of high-dimensional data to improve straggler resilience in distributed systems. Experiments show reduced mean squared recovery error in neural network inference and polynomial evaluation tasks.
PyTorch 2.14 introduces updates to Inductor, PyTorch Distributed, dynamic shapes, Apple Silicon support, and more, with a live Q&A and PyTorchCon announcement.
This paper introduces General Coded Computing (GCC), a learning-theoretic framework for mitigating stragglers in distributed computing systems, providing theoretical performance guarantees and experimental validation on deep neural networks.
An interactive, explorable explanation of various parallelization schemes for training transformers, adapted from academic content on scaling models.
This article explores the limitations of Pubsub systems, focusing on challenges in distributed computing environments.
apra-fleet is a tool that lets you run a fleet of AI agents across your machines.
Running coding agents across multiple machines reveals that managing the control plane is more challenging than the agents themselves, offering insights into distributed agent orchestration.
A technical note on performing RL post-training across 14 Macs distributed in 4 countries, highlighting distributed compute for reinforcement learning.
Prime Intellect engineers demonstrated a method to train reasoning models in 30 minutes using distributed RL over the open internet, utilizing Prime-RL, LLM judges, and multi-cloud GPUs, enabling open models to compete with closed labs without owning data centers.
Mesh LLM is a distributed AI computing platform that pools idle GPUs across multiple machines to run large language models, exposing a single OpenAI-compatible API. It leverages iroh's peer-to-peer networking to enable private, decentralized inference without a central server.
The article revisits the SETI@home distributed computing model as a potential solution to alleviate the demand for new database centers by utilizing idle office and home PCs for processing tasks.
This paper investigates distributed sketching for OLS regression, where sketches are built from partitioned subsets rather than the whole dataset, reducing computational cost. The authors characterize the exact excess loss of the averaged estimator and show it matches that of whole-data sketching when subset covariance divergence is small.