@analogalok: Operating System powered by Qwen 3.8 27B at 1950 tokens/sec! here is what 1,950 tokens/second @Alibaba_Qwen's 3.8 27b a…
Summary
A demonstration of an operating system powered by Qwen 3.8 27B running at 1950 tokens per second on Cerebras hardware, where the model weights act as the runtime for real-time software generation.
View Cached Full Text
Cached at: 09/15/26, 05:37 AM
Operating System powered by Qwen 3.8 27B at 1950 tokens/sec!
here is what 1,950 tokens/second @Alibaba_Qwen’s 3.8 27b actually looks like on @cerebras:
i wrote a minimal python web server that turns cerebras inference into a live operating system.
zero apps on disk.
when you double click an icon:
- python proxies a raw SSE stream from qwen 27b at 1,950 tok/s
- calculator compiles & mounts in 11s
- full canvas paint studio with brush engine compiles in 10s.
at 2,000 tokens/second, software is just an on demand hallucination that runs instantly.
the model weights ARE the operating system runtime.
what else would you build at 1,950 tokens/second?
Alok (@analogalok): qwen 3.8 27b cerebras inference math:
speed: 1,850 tok/s cost: $0.99/M in | $1.49/M out
intelligence drops like a database query.
standard 60Hz screens refresh every 16.6ms.
- local 3090/4090 (4-bit): 75 tok/s (1.2 tokens/frame)
- cerebras qwen 3.8 27b: 1,850 tok/s (31
Similar Articles
@RoundtableSpace: Qwen 3.8 27B running locally on an RTX 5090 beats Opus 4.8 in personal benchmarks at up to 200 tokens per second with n…
Qwen 3.8 27B, a 27-billion parameter AI model, outperforms Opus 4.8 in personal benchmarks when running locally on an RTX 5090 GPU, achieving up to 200 tokens per second without internet or API access.
Qwen 3.8 27B available on Cerebras at 1500 tokens/s
Cerebras has announced the availability of the Qwen 3.8 27B model on its inference platform with a speed of 1500 tokens per second, detailing its model compression techniques such as quantization and pruning.
Qwen 3.8 27B is faster than expected
A user reports that the Qwen 3.8 27B model achieves 50-60 tokens per second on dual 5060 TI cards, showing unexpected speed improvements over previous versions like Qwen 3.6.
@rohanpaul_ai: Qwen 3.6 27B on a MacBook Pro M5 Max 64GB hitting 34tokens per sec, locally with atomic[.]chat 90% acceptance rate, i.e…
Qwen 3.6 27B achieves 34 tokens/sec on a MacBook Pro M5 Max 64GB locally with 90% draft acceptance, enabled by TurboQuant, GGUF, and llama.cpp, showcasing a major advancement in laptop-based AI inference.
Alibaba's RISC-V CPU, XuanTie C950, Runs Qwen-3.8 27B at 30 tps
Alibaba's XuanTie C950 RISC-V chip, fabricated by TSMC, now runs the Qwen-3.8 27B AI model natively at 30 tps, demonstrating vertical integration and enabling efficient edge AI deployment.