gpu-computing

Tag

Cards List
#gpu-computing

@pochenai: Inspired by @percyliang's CS336, I built a static performance model for LLM inference — no dynamic batching/chunking, j…

X AI KOLs Following ↗ · 2026-08-28 Cached

Inspired by CS336, a static performance model for LLM inference provides analytical bounds for VRAM, time-to-first-token, and throughput, covering various configurations and calibrated against public benchmarks.

0 favorites 0 likes
#gpu-computing

@0x0SojalSec: Your phone got power with superintelligent, locally Qwen 3.8 27B on a 3090, over Tailscale. agent designing a camera ap…

X AI KOLs Timeline ↗ · 2026-08-27 Cached

A user demonstrates running the Qwen 3.8 27B model locally on a 3090 GPU via Tailscale, with an AI agent designing a camera app featuring live preview, object detection, and high token generation speeds.

0 favorites 0 likes
#gpu-computing

@TheAhmadOsman: GLM 5.3 Flash and Qwen 3.8 Flash Next are great examples of Local AI progression This is the good timeline

X AI KOLs Timeline ↗ · 2026-08-26 Cached

This tweet discusses the advancement in local AI, highlighting how models like GLM 5.3 Flash and Qwen 3.8 Flash Next can now be run on single GPUs, improving on previous hardware requirements.

0 favorites 0 likes
#gpu-computing

Open Source Kernel in Qwen3.6-35B-A3B for AMD MI350X: 78,498 output tok/s on 8 GPUs

Reddit r/LocalLLaMA ↗ · 2026-08-26

The article describes the optimization of AMD MI350X GPUs for running the Qwen3.6-35B-A3B LLM, achieving high output token throughput and open-sourcing the kernel to improve performance.

0 favorites 0 likes
#gpu-computing

A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU

arXiv cs.LG ↗ · 2026-08-26 Cached

This paper presents a feature-major codebook layout for memory-efficient sparse-binary self-organizing maps, enabling the training of large-scale maps up to 1.05 million neurons on a single GPU with significant speed improvements for MEDLINE data.

0 favorites 0 likes
#gpu-computing

Today I merged the first feature branch written entirely by my 4060Ti 16GB!

Reddit r/LocalLLaMA ↗ · 2026-08-25

A developer shares their experience using the quantized Qwen 3.8 27B model on a 4060Ti GPU to autonomously write and merge a feature branch, demonstrating significant progress in local AI for coding tasks.

0 favorites 0 likes
#gpu-computing

TorchMorph: CUDA-accelerated Morphological Transforms

Hugging Face Daily Papers ↗ · 2026-08-25 Cached

TorchMorph is a CUDA-accelerated PyTorch extension that provides GPU-optimized morphological and distance transform operators, achieving significant speed-ups over CPU implementations like SciPy with a compatible API.

0 favorites 0 likes
#gpu-computing

Hot Chips 2026: CUDA Targets RISC-V – By Chester Lam

Hacker News Top ↗ · 2026-08-24 Cached

At Hot Chips 2026, Nvidia announced plans to extend CUDA support to RISC-V CPUs, outlining specific hardware requirements like server-grade CPUs, ACPI, and PCIe coherency to enable efficient GPU compute.

0 favorites 0 likes
#gpu-computing

Ling Tiny, King of Speed

Reddit r/LocalLLaMA ↗ · 2026-08-23

Ling Tiny has replaced Gemma4-12B as an auxiliary model on a 4060Ti GPU, delivering phenomenal speed; users should avoid MTP and use the vLLM fork for BailingMoE3.

0 favorites 0 likes
#gpu-computing

@cyrilXBT: UNREAL. Qwen 3.8 27B is now happily running on an RTX 4060 locally with only 8GB of VRAM. You’re looking at a 64k conte…

X AI KOLs Timeline ↗ · 2026-08-23 Cached

A user shares that the Qwen 3.8 27B model runs locally on an RTX 4060 with only 8GB VRAM using Unsloth's IQ4_XS quantization, achieving a 64k context window and impressive performance metrics.

0 favorites 0 likes
#gpu-computing

I did it! I'm free! It's been 7 hours since I used claudecode

Reddit r/LocalLLaMA ↗ · 2026-08-21

After their Claude Code subscription expired, the user switched to using local models like Qwen3.8-27b and Pi, comparing their performance with Claude Sonnet 5 in coding tasks and app development.

0 favorites 0 likes
#gpu-computing

NVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.

Reddit r/LocalLLaMA ↗ · 2026-08-19

The author achieved performance parity between four 2017 Tesla V100 GPUs and a modern RTX 5090 when running the Qwen 3.8 model with NVFP4 precision, using custom software optimizations like the QPN kernel for efficient inference.

0 favorites 0 likes
#gpu-computing

@no_stp_on_snek: Check out Buun's work, he cookin.

X AI KOLs Following ↗ · 2026-08-19 Cached

A user highlights Buun's work on optimizing AI models, achieving high-speed inference of Qwen 3.6 on a single 3090 GPU and developing DFlash2 for Qwen 3.8.

0 favorites 0 likes
#gpu-computing

@fedebruzzone7: From a category theory perspective: Categorical Foundations for CuTe Layouts https://arxiv.org/abs/2601.05972

X AI KOLs Timeline ↗ · 2026-08-18 Cached

This paper introduces a categorical framework to formalize the layout algebra in NVIDIA's CUTLASS library, defining categories and morphisms to characterize tensor layouts, and provides a Python implementation with proofs of compatibility.

0 favorites 0 likes
#gpu-computing

5090: Windows or Linux for Qwen3.8.27b

Reddit r/LocalLLaMA ↗ · 2026-08-16

User seeks advice on the best operating system (Windows or Linux) and inference server to run the Qwen3.8.27b model on a dedicated AI rig with RTX 5090 and 96GB RAM for optimal performance.

0 favorites 0 likes
#gpu-computing

@TheAhmadOsman: Prediction We’re gonna get Kimi K3 equivalent intelligence running on a single RTX PRO 6000 in less than 18 months How?…

X AI KOLs Following ↗ · 2026-08-15 Cached

A Twitter user predicts that AI intelligence comparable to Kimi K3 will run on a single RTX PRO 6000 GPU within 18 months, later noting that Opus 4.6 Max quality already fits on a single RTX 5090.

0 favorites 0 likes
#gpu-computing

SpaceX officially closes its Cursor acquisition

TechCrunch AI ↗ · 2026-08-15 Cached

SpaceX has officially closed its $60 billion acquisition of AI coding startup Cursor, combining Cursor's AI capabilities with SpaceX's vast GPU computing infrastructure to advance AI development.

0 favorites 0 likes
#gpu-computing

@matthewjgunton: There are 3 basic levels in the NVIDIA software stack CUDA C++, PTX, and SASS Understanding all 3 helps you know why CU…

X AI KOLs Timeline ↗ · 2026-08-07 Cached

An educational tweet explaining the three levels of NVIDIA's software stack (CUDA C++, PTX, SASS) and how CUDA's abstraction creates a moat, while mentioning Luminal's automatic compiler search.

0 favorites 0 likes
#gpu-computing

@omarsar0: TokTier makes tokenization stateful for agentic serving. Across 153,951 real agent calls with a 94.1% prompt-cache hit …

X AI KOLs Following ↗ · 2026-08-03 Cached

TokTier is a stateful tokenization service that cuts tokenization overhead in agentic LLM serving by reusing stable token boundaries and running exact CPU/GPU tokenization, reducing time-to-first-token by 16–34% under vLLM.

0 favorites 0 likes
#gpu-computing

Operator-Aware Mixed-Precision Tolerance Calibration for Tensor Kernels

arXiv cs.LG ↗ · 2026-07-21 Cached

This paper proposes an operator-aware calibration method for absolute tolerances in tensor kernel correctness tests, using error distribution data to set tighter thresholds. The method achieves a 9.3% absolute recall improvement in bug detection with minimal false positives, demonstrated on the gpuemu corpus.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback