qwen3

Tag

Cards List
#qwen3

Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment

arXiv cs.AI · 3d ago Cached

This research report evaluates post-training ternarization of the Qwen3-4B model, achieving a 1.641-bit effective weight representation with substantial storage compression, while noting a performance trade-off and unresolved deployment acceleration issues.

0 favorites 0 likes
#qwen3

Qwen3.8 Flash AP Quants

Reddit r/LocalLLaMA · 3d ago

The post announces new quantized versions of the Qwen3.8 Flash model that outperform other high-quality quants through modified KLD measurement and optimized prefill performance.

0 favorites 0 likes
#qwen3

ContextPilot-14B (Hugging Face Repository)

TLDR AI · 6d ago Cached

ContextPilot-14B is a Qwen3-14B checkpoint that teaches language-model agents proactive context management via fine-grained reinforcement learning for long-horizon tasks.

0 favorites 0 likes
#qwen3

Qwen3.8-Flash-Next NVFP4 Day-3 support for 4xV100

Reddit r/LocalLLaMA · 2026-08-30

RadixArk/Qwen3.8-Flash-Next-NVFP4 is now supported in SGLang-V100, enabling full context operation on 4xV100 GPUs with performance metrics showing high throughput and context handling up to 256k tokens.

0 favorites 0 likes
#qwen3

ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

Hugging Face Models Trending · 2026-08-28 Cached

This repository provides GGUF quantizations of the Qwen3.8-27B AI model using GSQ and RCO methods for efficient deployment in standard tools.

0 favorites 0 likes
#qwen3

Local agentic coding Benchmark : Qwen3.8-Flash-Next NVFP4 vs 27B (and the others...)

Reddit r/LocalLLaMA · 2026-08-28

The benchmark reveals that Qwen3.8-Flash-Next-NVFP4 achieves nearly the highest scores among local models, with superior efficiency in request handling and token generation, and notable speed despite being undertrained.

0 favorites 0 likes
#qwen3

It's unbelievable! I used the mmap function in llama.cpp to fit Qwen3.8-Flash-Next IQ3_XSS into 16G+64G RAM, and the speed still reached 26t/s.

Reddit r/LocalLLaMA · 2026-08-28

A user successfully used the mmap function in llama.cpp to fit the Qwen3.8-Flash-Next IQ3_XSS model into 16GB+64GB RAM, achieving a speed of 26 tokens per second, which outperforms a larger non-MOE 30B model.

0 favorites 0 likes
#qwen3

Qwen3.8 27B C8 at 972 TG / 5,680 PP on 4x MI100 rig ($6.5k) using my new INT8 vLLM fork

Reddit r/LocalLLaMA · 2026-08-26

A new vLLM fork introduces comprehensive INT8 optimization for Qwen3.8 27B on older AMD MI100 GPUs, achieving up to 972 tokens per second throughput with rigorous accuracy validation.

0 favorites 0 likes
#qwen3

Fully quantized NVFP4 Qwen3.8-27B with QUASAR QAD

Reddit r/LocalLLaMA · 2026-08-26 Cached

A fully quantized 4-bit NVFP4 version of the Qwen3.8-27B AI model, trained with the QUASAR method to maintain high quality while reducing model size.

0 favorites 0 likes
#qwen3

How we built a SOTA search engine using PostgreSQL, pgvector, and Qwen3 embeddings [P]

Reddit r/MachineLearning · 2026-08-25

A technical breakdown of how Papers with Code built a state-of-the-art hybrid search engine combining keyword and semantic search using PostgreSQL, pgvector, and Qwen3 embeddings, powered by Hugging Face infrastructure.

0 favorites 0 likes
#qwen3

@servasyy_ai: https://x.com/servasyy_ai/status/2091416214283379123

X AI KOLs Timeline · 2026-08-23 Cached

This article details the local deployment guide for the Qwen3.8 27B model, covering two routes for Mac and Nvidia graphics cards, and provides real-world performance data to help users run this model on consumer-grade hardware.

0 favorites 0 likes
#qwen3

[MASSIVE TINY RELEASE] - Supra2-Medium-Base - a tiny 25M parameters model competing heavily with our previous 50M model!

Reddit r/LocalLLaMA · 2026-08-20

Supra2-Medium-Base is a new 25M parameter AI model based on qwen3 architecture, trained from scratch, that benchmarks competitively against a larger 50M parameter model.

0 favorites 0 likes
#qwen3

Qwen3.8 27b just exceeded my expectations on svg generation :D

Reddit r/LocalLLaMA · 2026-08-20

A user tested the Qwen3.8 27B model on SVG generation with a complex prompt and found the results exceeded expectations, highlighting the model's capability in creative tasks.

0 favorites 0 likes
#qwen3

EschaLabs/Qwen3.8-27B-Escha-W2

Hugging Face Models Trending · 2026-08-20 Cached

Escha Labs has released a 2-bit quantized version of the Qwen3.8-27B AI model, enabling it to run on a single 24GB consumer GPU with up to 64k context while maintaining performance comparable to FP8 references.

0 favorites 0 likes
#qwen3

OrcaRouter's uncensored Qwen3.8-27B still caveats 27–56% of harmful answers

Reddit r/ArtificialInteligence · 2026-08-20

OrcaRouter's uncensored Qwen3.8-27B derivative reduces harmful-prompt refusal to 0-6% but still caveats 27-56% of answers, with mixed performance metrics compared to the base model.

0 favorites 0 likes
#qwen3

Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation

arXiv cs.LG · 2026-08-20 Cached

This paper introduces a candidate-free control for single-pass test-time aggregation to evaluate the value of candidate context in AI reasoning. It finds that candidate conditioning improves accuracy when multiple candidates are correct but lowers it when all candidates are wrong.

0 favorites 0 likes
#qwen3

Rethinking Privileged Information in On-Policy Self-Distillation

arXiv cs.LG · 2026-08-20 Cached

The paper investigates whether performance gains in on-policy self-distillation come from learning privileged reference information or recovering existing reasoning behavior, finding that the correct reference does not consistently benefit performance across various conditions.

0 favorites 0 likes
#qwen3

@10xmylife: In the overall performance of Agents, Qwen3.8-27B ranks 7th. As more powerful models become open-sourced, when open-source models have sufficient capability and local deployment costs are acceptable, what impact will this have on OpenAI and Anthropic?

X AI KOLs Following · 2026-08-19 Cached

Qwen3.8-27B ranks seventh in overall Agent performance, and the article explores the potential impact on OpenAI and Anthropic as open-source models advance in capability.

0 favorites 0 likes
#qwen3

@inco_ai: Hope you enjoy our first release! (and more to come)

X AI KOLs Timeline · 2026-08-18 Cached

Inco AI announces DFlash 2, a next-generation decoding technique that achieves up to 4.6× speed increase for Qwen3.8-27B on M5 Max MacBook Pro, improving efficiency without altering output.

0 favorites 0 likes
#qwen3

@Lonely__MH: Unleashed! The uncensored version of Qwen3.8-27B with safety restrictions removed is here! Kudos to the community for the speed! Deeply optimized for Mac M chips! I see everyone discussing the DGX Spark deployment for ling-3.0-flash, and many people's first reaction is that the compute power is too expensive to buy. Since cloud costs are high...

X AI KOLs Timeline · 2026-08-18 Cached

Qwen3.8-27B uncensored version released, optimized for Mac M chips, supports local deployment, retains multimodal capabilities and safety research features, with simplified installation steps.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback