Simple Bench - QWEN 3.8 27b has a common sense almost like GPT 5.0 Pro??

Reddit r/singularity News

Summary

A benchmark called Simple Bench shows that the QWEN 3.8 27b model has common sense capabilities almost comparable to GPT 5.0 Pro.

WTF They really cooked. https://simple-bench.com/
Original Article

Similar Articles

Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores

Reddit r/LocalLLaMA

This article reports benchmark results showing that quantization has little effect on knowledge benchmarks (GPQA) but significantly degrades agentic performance (Terminal-Bench 2) for Qwen 3.6 models, with further observations on timeout settings and run variability.

Qwen 3.8 27B SlopCodeBench results

Reddit r/LocalLLaMA

The article presents benchmark results for the Qwen 3.8 27B model on SlopCodeBench, showing poor performance on strict checkpoints but fair results on core ones, indicating it may not be suitable for autonomous code management without direction.