Qwen3.8-27B took a serious hit to *knowledge* vs 3.6
Summary
The author finds that Qwen3.8-27B has weaker knowledge recall compared to Qwen3.6 based on personal benchmarks and offline tests, suggesting it may not be suitable for airgapped knowledge retrieval.
Similar Articles
A hunch: Qwen3.8-27B's general knowledge got pruned (good, if true)
The author shares a hunch that Qwen3.8-27B has pruned general knowledge to improve coding and agentic skills, based on reduced knowledge of a specific German town compared to earlier Qwen models.
Qwen 3.8 27b is strong even at Q3_xxs
The user finds Qwen 3.8 27b in Q3 quantization highly effective for coding tasks with fast inference speeds, outperforming previous models, despite minor issues in general conversations.
Quantization hurts knowledge nonlinearly - Qwen3.6 27B case study
A case study on how quantization affects factual knowledge in Qwen3.6 27B, showing that knowledge loss scales nonlinearly and obscure facts degrade most at low bit widths, unlike benchmark scores.
Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores
This article reports benchmark results showing that quantization has little effect on knowledge benchmarks (GPQA) but significantly degrades agentic performance (Terminal-Bench 2) for Qwen 3.6 models, with further observations on timeout settings and run variability.
Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
Qwen 3.8 27B is a powerful open-source 27B parameter vision-capable LLM from Alibaba's Qwen research lab, praised for its benchmarks but criticized for defaulting to excessive reasoning effort, which slows down performance on consumer hardware.