Qwen 3.8 is off to College. Got a 34 on the ACT

Reddit r/AI_Agents News

Summary

An experiment tested the Qwen 3.8 27B AI model on ACT practice exams using vision capabilities, achieving high composite scores of 34-36, showcasing strong performance in standardized testing.

I made Qwen 3.8 27B take the ACT to see if it’s ready for college. I’ve been testing the new Qwen Model over the past few days on my PC. I tested the full version the Q8, Q6 and Q4 versions and landed on the Q8 for speed vs quality. I decided to download some practice tests and had the model solve them. I fed it the raw PDFs to test not only how well it knows the answers but also how good the vision capabilities are at answering the questions one by one. At the end I graded its answers. Here are my findings from taking 2 tests. \*\*Setup:\*\* Qwen 3.8 27B Instruct, Q8\_0 GGUF, LM Studio, 2× RTX 3090 (full offload, 32k context). Two \*official\* ACT practice PDFs, 342 questions total, graded against the answer keys and the official raw→scale conversion tables that ship in the same PDFs. No human help, no retries on wrong answers, no cherry-picking. \## Results Section Test A Test B English 48/50 → \*\*35\*\* 45/50 → \*\*33\*\* Mathematics 44/45 → \*\*36\*\* 43/45 → \*\*35\*\* Reading 36/36 → \*\*36\*\* 36/36 → \*\*36\*\* Science 39/40 → \*\*35\*\* 35/40 → \*\*33\*\* \*\*Composite\*\* \*\*36\*\* \*\*34\*\* \*\*326/342 correct overall (95.3%).\*\* Zero blanks. 36 is the maximum composite the ACT awards; 34 is roughly 99th percentile. \*\*Reading was perfect on both papers — 72/72.\*\* Time: 177 minutes for both tests, \~88 min per test. A human gets \~165 min for one. I was surprised that it did so well but also that it took so long. I thought it would be a 10-20 minute job but it was over 2 hours for 2 tests which looking back at it is understandable since it was using the vision capabilities to read instead of given plain text for each question
Original Article

Similar Articles

Qwen 3.8 Low and Medium are goated

Reddit r/LocalLLaMA

Artificial Analysis benchmarked the Qwen 3.8 Low and Medium models, showing exceptional scores that confirm their success beyond overthinking.

Qwen 3.8 27b is strong even at Q3_xxs

Reddit r/LocalLLaMA

The user finds Qwen 3.8 27b in Q3 quantization highly effective for coding tasks with fast inference speeds, outperforming previous models, despite minor issues in general conversations.