Tag
Tweet highlights running the Qwen 3.8 27B model locally on an RTX 5090 system with 32GB VRAM, achieving 115 tokens/sec, and notes the official BF16 checkpoint is 55.6GB.
The article details an experiment achieving 50 tokens per second inference with Qwen3.8-27B at 256K context on a 24GB GPU using Multi-Token Prediction and custom optimizations.
The author experimented with the 27B model on VLLM and created a 'high' reasoning mode by blending prompts from low and xhigh modes, resulting in more efficient and enjoyable reasoning output.
User reports successfully running a 27B parameter model quantized to 1-bit on a Jetson Orin NX 16GB edge device, expressing amazement at the feasibility.
Ternary Bonsai 27B, a large language model, is demonstrated running locally on an NVIDIA RTX 5090 GPU, requiring under 6GB of memory and enabling end-to-end agentic workflows on consumer hardware.
PrismML announces Bonsai 27B, a multimodal model based on Qwen3.6 27B that can run on a phone, enabling local multi-step reasoning, tool use, and long-context workflows.
The author built an autonomous development pipeline and benchmarked it by running the same project using a local 27B model on a modified RTX 4090 versus cheap cloud LLM APIs.
Introduces a new 27B post-trained model that distills positives from Fable and Kimi 2.7 Coder, with links to download.