distributed-inference

Tag

Cards List
#distributed-inference

ESP32S3 cluster running 1.58-bit (BitNet) Language model

Hacker News Top ↗ · 2d ago Cached

This project implements a distributed pipeline inference engine that runs a 0.5B BitNet language model on a cluster of seven ESP32S3 microcontrollers using 1.58-bit quantization and SPI daisy-chain communication.

0 favorites 0 likes
#distributed-inference

Resource-Efficient Distributed Recursive Gaussian Processes

arXiv cs.LG ↗ · 2026-09-24 Cached

This paper develops two resource-efficient distributed recursive Gaussian process algorithms for multi-output regression in multi-agent systems, reducing communication overhead while maintaining estimation accuracy.

0 favorites 0 likes
#distributed-inference

@ai_xiaomu: What makes Mac Studio M5 impressive isn't the 512G config, which M3 already has. The differences are: 1.2TB/s memory bandwidth, 36-core CPU + 80-core GPU with independent NPU, up to 512GB unified memory, Thunderbolt 5 + RDMA, supporting distributed AI inference across multiple Macs…

X AI KOLs Following ↗ · 2026-08-26 Cached

The highlights of Mac Studio M5 include its memory bandwidth, CPU/GPU core count, independent NPU, and support for distributed AI inference across multiple Macs. AI computing power is 4.3 times that of M3 Ultra, making it suitable for local model workstations.

0 favorites 0 likes
#distributed-inference

Inside Our Distributed LLM Inference Research for Intel PCs

Reddit r/artificial ↗ · 2026-08-20 Cached

The article presents research on distributed LLM inference for Intel PC fleets, focusing on pipeline-parallel sharded inference using OpenVINO with performance optimizations for heterogeneous hardware.

0 favorites 0 likes
#distributed-inference

Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri

Hacker News Top ↗ · 2026-08-14 Cached

Lumabri lets users run huge mixture-of-experts models on a P2P swarm using the Colibri engine, allowing any machine to join and chat without downloading the full model up front. It is pure C, dependency-free, and works on CPU first with optional GPU acceleration.

0 favorites 0 likes
#distributed-inference

Cascadia Launches Distributed AI Inference for Intel Hardware

Reddit r/artificial ↗ · 2026-08-13

Cascadia has launched a distributed AI inference system designed for Intel hardware, enabling scalable and efficient inference workloads.

0 favorites 0 likes
#distributed-inference

EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference

arXiv cs.LG ↗ · 2026-08-11 Cached

This paper presents EasyBalance, a cross-layer load balancing strategy for distributed Mixture-of-Experts (MoE) inference that schedules and jointly executes workloads from different layers to mitigate GPU idling without modifying expert-device mappings, reducing idle time by over 40% in experiments.

0 favorites 0 likes
#distributed-inference

@TheTechDiggest: [OpenSource - Distributed AI & Mesh LLM Inference] Buying an expensive enterprise GPU isn't the only way to run massive…

X AI KOLs Timeline ↗ · 2026-08-08 Cached

A tweet thread introduces mesh-llm, an open-source tool that pools local network devices into a unified, OpenAI-compatible API for running large LLMs without expensive enterprise GPUs.

0 favorites 0 likes
#distributed-inference

China’s optical chip breakthrough boosts AI speed 100-fold using fraction of compute power

Reddit r/ArtificialInteligence ↗ · 2026-07-13 Cached

Chinese researchers have developed an all-optical interconnect system that links standard electronic chips, boosting AI distributed inference speeds by over 100 times while using just one-ninth of typical computational resources. The breakthrough, published in National Science Review, uses silicon photonic transceiver chips and FPGAs to achieve dramatic efficiency gains.

0 favorites 0 likes
#distributed-inference

@antirez: Based on what I'm saying with GLM 5.2 implementation inside DwarfStar, there is 90% of probability I'll merge the branc…

X AI KOLs Following ↗ · 2026-06-24

Antirez announces high probability of merging a branch implementing GLM 5.2 in DwarfStar, which could become the best model for 512GB Mac Studio and potentially run on distributed 128GB MacBooks with 2-bit quantization.

0 favorites 0 likes
#distributed-inference

Someone just ran a 744B parameter model at 30 tok/s across 6 consumer GPUs in 6 different US states over the open internet

Reddit r/ArtificialInteligence ↗ · 2026-06-20

A researcher debuted Shard, achieving 30 tok/s inference on a 744B parameter model distributed across 6 consumer GPUs over the open internet, a 15-20x improvement over previous methods.

0 favorites 0 likes
#distributed-inference

@MiaAI_lab: A PR to vLLM to allow TP=3 for MiniMax M3 His NVFP4 quant is 260GB - lukealonso/MiniMax-M3-NVFP4 Hopefully this will wo…

X AI KOLs Timeline ↗ · 2026-06-14 Cached

A pull request to vLLM adds support for tensor parallelism degree 3 for MiniMax M3 with its NVFP4 quantization, enabling the model to run on 3x DGX Sparks with 87GB memory each.

0 favorites 0 likes
#distributed-inference

@m_sirovatka: KV Cache re-use is the most important thing for agentic rollouts. We've integrated Mooncake Store into prime-rl with vL…

X AI KOLs Following ↗ · 2026-06-02 Cached

vLLM integrates Mooncake Store for distributed KV cache reuse, enabling cross-node prefix caching to efficiently serve agentic workloads with high token reuse.

0 favorites 0 likes
#distributed-inference

@dorsa_rohani: This paper might be the bible of distributed inference atp

X AI KOLs Timeline ↗ · 2026-05-23 Cached

A tweet recommending a paper that is described as the bible of distributed inference.

0 favorites 0 likes
#distributed-inference

Clustering Raspberry Pis together to learn distributed training/inference

Reddit r/LocalLLaMA ↗ · 2026-05-14

A blog post guides readers through setting up a Raspberry Pi cluster for distributed training and inference, part of a series aimed at making distributed AI accessible using affordable hardware.

0 favorites 0 likes
#distributed-inference

@antirez: Announcing with gratitude that @audreyt just gifted me an M5 Max 128GB MacBook Pro! It will let me develop DwarfStar4 (…

X AI KOLs Timeline ↗ · 2026-05-12

antirez announces receiving an M5 Max 128GB MacBook Pro from audreyt to develop DwarfStar4 and experiment with distributed inference across M3 Max and M5 Max hardware.

0 favorites 0 likes
#distributed-inference

Federation of Experts: Communication Efficient Distributed Inference for Large Language Models

Hugging Face Daily Papers ↗ · 2026-05-07 Cached

Federation of Experts (FoE) restructures mixture-of-experts blocks into clusters that process KV heads independently, eliminating inter-node communication bottlenecks and improving inference throughput and latency by up to 5.2x while maintaining generation quality.

0 favorites 0 likes
#distributed-inference

2x 512gb ram M3 Ultra mac studios

Reddit r/LocalLLaMA ↗ · 2026-04-21

A user shares their $25k hardware setup of two 512GB RAM M3 Ultra Mac Studios for running large language models locally, having tested DeepSeek V3 Q8 and GLM 5.1 Q4 via the exo distributed inference backend, while awaiting Kimi 2.6 MLX optimization.

0 favorites 0 likes
← Back to home

Submit Feedback