gpu

Tag

Cards List
#gpu

16GB (and in many cases 12GB) is the max vram most people will ever reasonably have

Reddit r/LocalLLaMA · 3d ago

The article discusses how 16GB of VRAM is the realistic high-end limit for most users due to financial constraints, but recent AI model improvements like Qwen 27B quants are enabling more capabilities on such hardware, with hopes for future architectural innovations.

0 favorites 0 likes
#gpu

@shao__meng: https://x.com/shao__meng/status/2101835798316495007

X AI KOLs Timeline · 3d ago Cached

Baseten's 'Inference Engineering' is a systematic book that explains AI inference optimization techniques from CUDA to production deployment, helping engineers efficiently run open-source models in production environments.

0 favorites 0 likes
#gpu

@TheAhmadOsman: The future of inference isn’t necessarily in any of the current hardware providers btw No disrespect to the incumbents,…

X AI KOLs Timeline · 3d ago Cached

A tweet suggests that future AI inference hardware may not come from current providers like NVIDIA, highlighting acquisitions of startups such as Groq because GPUs are not optimally designed for inference.

0 favorites 0 likes
#gpu

@TheAhmadOsman: Any potential hire will sign on the spot if you offer them a DGX Station / GB300 as a sign-on bonus btw

X AI KOLs Timeline · 4d ago

A tweet suggests that offering NVIDIA DGX Station or GB300 as a sign-on bonus would instantly attract hires in the tech industry.

0 favorites 0 likes
#gpu

@TheAhmadOsman: Inference Engineering simply is - Encodings that can be loaded - Operations on said encodings implemented by backends -…

X AI KOLs Timeline · 5d ago Cached

A tweet explains the core components of inference engineering: loadable encodings, backend operations, and native hardware arithmetic, with a discussion on kernels and GPU efficiency.

0 favorites 0 likes
#gpu

AMD Plans 10% Price Hike Across GPUs, Chipsets, and Possibly CPUs

Reddit r/LocalLLaMA · 6d ago Cached

AMD is planning a 10% price increase across its GPU and chipset products, with potential hikes for CPUs as well.

0 favorites 0 likes
#gpu

Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions

arXiv cs.AI · 2026-09-16 Cached

This paper investigates placement policies for KV cache across GPU, CPU, and SSD tiers to optimize LLM serving for long-lived sessions, evaluating policies like recency and reuse-frequency and finding workload-specific recommendations for migration and prefetch.

0 favorites 0 likes
#gpu

@seclink: A bit interesting, learn a bit...

X AI KOLs Following · 2026-09-16 Cached

Version 1.0 of auto-gpu-kernel has been released, a meta-harness tool that autonomously generates high-performance GPU kernels.

0 favorites 0 likes
#gpu

Don’t buy a $9K RTX 5090.... instead.

Reddit r/LocalLLaMA · 2026-09-15

The article advises against buying an RTX 5090 for $9K by suggesting to purchase it cheaper in Taiwan for about $4K, including travel expenses and a vacation.

0 favorites 0 likes
#gpu

Radeon AI PRO R9700 - Which one to choose?

Reddit r/LocalLLaMA · 2026-09-15

A user seeks advice on choosing between different vendor versions of the Radeon AI PRO R9700 GPU, considering future expansion and cooling options.

0 favorites 0 likes
#gpu

Nvidia's RTX 5090 vanishes from online retail in the US — third-party sellers now demand as much as $9,500 for Nvidia's fastest GPU

Reddit r/LocalLLaMA · 2026-09-14 Cached

The Nvidia RTX 5090 GPU has vanished from US online retail, with third-party sellers demanding up to $9,500, driven by AI demand and raising scam risks.

0 favorites 0 likes
#gpu

For the GPU poor. K2 Horizon 7B ranks between qwen 3.6 27B and qwen 3.6 35BA3b on the Artificial Analysis Intelligence Index.

Reddit r/LocalLLaMA · 2026-09-14

K2 Horizon 7B, a compact AI model, achieves impressive benchmark scores rivaling larger models and shows solid performance in initial testing for tasks like compiling llama.cpp for CUDA.

0 favorites 0 likes
#gpu

@reprompting: reading about tritonBLAS today https://arxiv.org/pdf/2512.04226

X AI KOLs Timeline · 2026-09-14 Cached

The paper presents tritonBLAS, an analytical model for optimizing GPU GEMM kernel parameters without runtime autotuning, achieving near-optimal performance and significantly reducing compilation time.

0 favorites 0 likes
#gpu

RTX PRO 5500 Blackwell (84GB) released

Reddit r/LocalLLaMA · 2026-09-14 Cached

NVIDIA announces the RTX PRO 5500 Blackwell Workstation Edition, a professional GPU with 84GB GDDR7 memory built for enterprise AI and graphics workloads.

0 favorites 0 likes
#gpu

Should I sell my RTX 5090 for a Mac Studio M5 Ultra 96GB?

Reddit r/LocalLLaMA · 2026-09-13

A user asks whether trading an RTX 5090 for a Mac Studio M5 Ultra is a sensible upgrade for coding, based on memory bandwidth and cost differences.

1 favorites 1 likes
#gpu

nvidia rtx 5090 with 96gb of vram.

Reddit r/LocalLLaMA · 2026-09-11

A China-modified Nvidia RTX 5090 with 96GB of VRAM is available on Alibaba for under $4,000, offering three times more memory at 65% of the original cost.

0 favorites 0 likes
#gpu

What GPUs will give me GOOD speeds and on DSV4 Flash and similar models, and not have to run a mega quantized version? Budget around $15k-ish.

Reddit r/LocalLLaMA · 2026-09-11

A user in a tech forum is seeking advice on GPUs for running AI models like DSV4 Flash locally with good performance, comparing AMD and NVIDIA options within a $15k budget.

0 favorites 0 likes
#gpu

CUDA/HIP: Flash Attention tuning (gfx1201) by pwilkin · Pull Request #28102 · ggml-org/llama.cpp

Reddit r/LocalLLaMA · 2026-09-11 Cached

This pull request introduces Flash Attention tuning optimizations for CUDA/HIP in the llama.cpp project, targeting gfx1201 hardware to enhance inference performance.

0 favorites 0 likes
#gpu

Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)

Reddit r/LocalLLaMA · 2026-09-09 Cached

Apple has announced the A20 Pro SoC for the iPhone 18 Pro, featuring a 7-core GPU with 40% faster graphics performance, a 32-core Neural Engine, and 50% more memory bandwidth, all manufactured on a 2 nm process.

0 favorites 0 likes
#gpu

Experimenting with an adaptive memory governor for PyTorch on an 8GB GPU — would love some feedback

Reddit r/LocalLLaMA · 2026-09-09

The author is experimenting with an adaptive memory governor for PyTorch to prevent CUDA OOM errors on 8GB GPUs, sharing code and seeking community feedback.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback