@TheAhmadOsman: Been playing with @PrismML's new model that turned Qwen 3.5 27B into a sub-4GB and sub-6GB weights and I am impressed C…
Summary
PrismML released a compressed version of Qwen 3.5 27B that fits in sub-4GB and sub-6GB memory, enabling impressive local AI performance.
View Cached Full Text
Cached at: 07/16/26, 04:03 AM
Been playing with @PrismML’s new model that turned Qwen 3.5 27B into a sub-4GB and sub-6GB weights and I am impressed
Cannot believe how far Opensource and Local AI have come since Christmas (~8 months ago)
Similar Articles
@sudoingX: this lab took qwen 3.6 27b, the model i've been calling king of the 24gb tier all month, and crushed it down to 3.9gb. …
PrismML announces Bonsai 27B, a binary-quantized version of Qwen3.6 27B that runs on a phone using only 1.125 bits per weight, claiming 89.5% intelligence retention. The model is being independently tested by @sudoingX to verify performance.
Compressed Version of Qwen-3.6-27B coming from PrismML - Khosla-Backed Startup Claims Breakthrough With Largest-Ever AI Model on an iPhone
PrismML, a Khosla-backed startup, releases a compressed version of Qwen-3.6-27B, claiming it's the largest AI model ever to run on an iPhone.
@UnslothAI: Qwen3.8-27B is coming! Will run locally on 17GB RAM/VRAM setups.
Alibaba announces Qwen3.8-27B open-weights release, capable of running locally on 17GB RAM/VRAM, alongside the larger Qwen3.8-Max.
Prism-ML Bonsai Qwen 3.6 27B
Prism ML released Ternary-Bonsai-27B, a ternary-quantized version of Qwen3.6-27B that retains 95% of FP16 intelligence at a ~7.2 GB footprint, enabling full 27B-class reasoning on laptops and single GPUs with speeds up to 26 tok/s on Apple M5 Pro.
Qwen 3.8 27b is out. Big news for local AI
Qwen 3.8 27b, a sub-30 billion parameter AI model, has been released and is suitable for local inference on consumer hardware like RTX 3090 or M4 Pro, potentially replacing cloud-based AI subscriptions and shifting workflows locally.