Tag
A user details their first attempt at tuning the Qwen3.8-27B model with Q4_K_M quantization on an RTX 5080 16GB, achieving about 13.2 tokens per second at 50-61K context by selectively offloading FFN tensors to CPU to improve performance.
Best Buy has raised the price of the Asus ROG Astral RTX 5080 OC to $2,099, exceeding the RTX 5090's MSRP, reflecting ongoing component shortages and GPU price hikes.
NVIDIA is expanding its GeForce NOW cloud gaming service with a new GeForce RTX 5080-powered server in Toronto, bringing improved performance and lower latency to members in Canada. The update also includes new games and native touch controls for NTE: Neverness to Everness.
A setup using RTX 5080 and RTX 3090 GPUs achieves 80 tokens per second on the Qwen 3.6 27B Q8 model.
Detailed benchmarks of Qwen3.6 35B MoE on RTX 5080 16GB show that MTP (Multi-Token Prediction) does not improve inference speed at 128k context due to VRAM constraints; the best configuration is Q4_K_XL without MTP, achieving ~56 tok/s generation at 128k context.
NVIDIA announces 16 games joining GeForce NOW cloud streaming in May, including new AAA titles like Forza Horizon 6 and 007 First Light, and expands RTX 5080-class performance across the library for Ultimate members.