RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
Summary
A setup using RTX 5080 and RTX 3090 GPUs achieves 80 tokens per second on the Qwen 3.6 27B Q8 model.
Similar Articles
[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw
Released GSQ-RCO quantized and expert-pruned versions of Qwen3.8-Flash-Next, achieving BF16-level performance at reduced bit-widths and enabling deployment on smaller hardware.
@elonmusk: Not bad
Grok 4.7 xHigh, an AI model from X, ranks first in the Artificial Analysis Cyber Index, outperforming other leading models like GPT-6 in enterprise cyber defense.
Searching for 3.8 35B: Qwen3.6-35B-A3B (Testing 5 Finetunes vs. Base)
The article benchmarks the Qwen3.6-35B-A3B base model and five finetunes, finding that only Occamy-1.0 is competitive with the base model, while others like Tiel underperform significantly.
Source: Inference provider Modal Labs closing in on $750M round at $15.75B valuation
Modal Labs, an AI inference infrastructure provider, is closing in on a $750 million funding round at a $15.75 billion valuation, more than tripling its valuation from four months ago amid soaring demand for AI inference services.
Qwen next 3.8 and 3.8 27b Vs Sonnet 5.5 low and Sonnet 5.5 medium.
The article discusses how local AI models like Qwen-Next 3.8 are now competitive with larger models such as Sonnet 5.5, based on personal experience and comparison.