qwen3

Tag

Cards List
#qwen3

Two 96 GB Ascend cards crun Qwen3.8-flash-next hardware notes, vLLM work, benchmarks, and what is next

Reddit r/LocalLLaMA ↗ · yesterday

A deep-dive guide on running Qwen3.8 Flash-Next on two 96 GB Huawei Atlas 300I Duo cards via vLLM Ascend, detailing the two-chip-per-card layout, passive cooling, and software optimizations that brought the machine from incoherent ~1 tok/s output to coherent ~30 tok/s single-request and ~61 tok/s aggregate decode, including a full GPQA Diamond evaluation.

0 favorites 0 likes
#qwen3

Clef: Open Weights decision model by Cloudflare

Reddit r/LocalLLaMA ↗ · 2d ago Cached

Cloudflare released Clef, an open-weights 27B multimodal decision model that takes a state and a schema of typed questions as input and returns probabilities for each option in a single forward pass, with no free-form generation or output parsing. It is post-trained from Qwen3.8-27B, ships on Hugging Face with a smaller Clef-Flash variant, and is compatible with the Jev/SystemOne API.

0 favorites 0 likes
#qwen3

BAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B

Reddit r/LocalLLaMA ↗ · 3d ago Cached

BAAI releases AREX-2, a 27B long-horizon self-improving agent model based on Qwen3.8 that learns to propose, measure, reflect, and revise solutions over multiple test-time rounds, achieving strong results on coding, ML engineering, and deep research benchmarks with only 27B parameters.

0 favorites 0 likes
#qwen3

Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine

Reddit r/LocalLLaMA ↗ · 3d ago

The article highlights the Strata inference engine, which significantly outperforms llama.cpp for running Qwen3.8 models on a laptop with 12GB VRAM and 64GB RAM, achieving up to 50 tokens per second for text generation and 1500 tokens per second for prompt processing.

0 favorites 0 likes
#qwen3

Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper introduces ActFirst-OPD, a framework that accelerates on-policy distillation for multi-turn language agents by decoupling action execution from full reasoning, achieving significant training speedups while maintaining performance across benchmarks.

0 favorites 0 likes
#qwen3

I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s

Reddit r/LocalLLaMA ↗ · 5d ago

The user built a low-cost setup using five ex-mining BC-250 boards to run the Qwen3-Coder-Next AI model, achieving around 40 tokens per second at 30k context with plans to expand.

0 favorites 0 likes
#qwen3

LLMs are General Asynchronous Agents

Hugging Face Daily Papers ↗ · 5d ago Cached

The paper proposes a general asynchronous LLM framework that lets users define inference coroutines with overlapping memory states, demonstrating that Qwen3.x models can handle streaming video, video games, and monitoring tasks asynchronously without task-specific training.

0 favorites 0 likes
#qwen3

FreedomIntelligence/HuatuoGPT-3-27B · Hugging Face

Reddit r/LocalLLaMA ↗ · 2026-09-24 Cached

HuatuoGPT-3-27B is a medical language model built on Qwen3.8-27B using One-stage Policy Optimization (OnePO), a reinforcement learning method for domain adaptation without supervised fine-tuning.

0 favorites 0 likes
#qwen3

is this good? 262k Qwen3.8:27B-Q4_K_M

Reddit r/LocalLLaMA ↗ · 2026-09-19

The user has implemented nvfp4 KV cache support for the Qwen3.8 model on a heterogeneous GPU setup using custom CUDA kernels and quantization to optimize performance.

0 favorites 0 likes
#qwen3

153 tok/s on 1x AMD Radeon R9700 running Qwen3.8 27b NVFP4, 470 tok/s @ 8 conc requests, Prefill @ 3,619 tok/s

Reddit r/LocalLLaMA ↗ · 2026-09-17

Optimized performance for running the Qwen3.8 27b NVFP4 model on a single AMD Radeon R9700 GPU, achieving up to 153 tokens per second in decode and 3,619 tokens per second in prefill with improved concurrency.

0 favorites 0 likes
#qwen3

SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale

arXiv cs.CL ↗ · 2026-09-17 Cached

This controlled study compares supervised fine-tuning and reinforcement learning methods for training tool-calling agents across different datasets and model scales, finding that SFT with LoRA is strongest in-distribution while RL shows slight advantages in cross-dataset transfer.

0 favorites 0 likes
#qwen3

Decided to build a game, and test the ceiling of Qwen3.8 27b

Reddit r/LocalLLaMA ↗ · 2026-09-14

A user built a game to test the performance limits of the Qwen3.8 27b AI model using an overclocked RTX 3090, and plans to release it as open-source on GitHub for community contribution.

0 favorites 0 likes
#qwen3

What pi.dev plugin do you suggest for context, compaction and memory management of local models?

Reddit r/LocalLLaMA ↗ · 2026-09-12

The user describes struggles with context, compaction, and memory management for local AI models using pi.dev plugins and seeks suggestions for solutions that handle varying model context windows and VRAM limitations.

0 favorites 0 likes
#qwen3

2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s

Reddit r/LocalLLaMA ↗ · 2026-09-12

The user describes a hardware configuration with dual RTX 3090 GPUs and an AMD EPYC CPU for running the Qwen3-Flash-Next model using llama.cpp, achieving 38 tokens per second in single-stream inference, and seeks advice on whether to add a third GPU or upgrade the CPU to improve performance, especially for running multiple parallel agents.

0 favorites 0 likes
#qwen3

Adding emotion control tags to Qwen3-TTS

Reddit r/LocalLLaMA ↗ · 2026-09-09

The article describes fine-tuning Qwen3-TTS to add emotion control tags, overcoming training challenges like codec prefix inconsistencies and generation concurrency issues, and discovering that emotion can be manipulated via affine transformations in speaker embeddings.

0 favorites 0 likes
#qwen3

@ashxhart: Been giving my M3 Ultra Studio some love and tinkering with MLX. Started the afternoon at 42 tok/s. Now: 73 tok/s, and …

X AI KOLs Timeline ↗ · 2026-09-08 Cached

An author improved MLX performance on an M3 Ultra Studio to achieve 73 tok/s for a 4-bit Qwen3.8-Flash-Next model, which shows intelligence scores matching GPT-5.6 Sol, highlighting the growing potential of local AI.

0 favorites 0 likes
#qwen3

ukisai/Swift-Qwen3.8-27b

Hugging Face Models Trending ↗ · 2026-09-08 Cached

Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B, reducing thinking tokens by 58.3% while maintaining near-identical performance.

0 favorites 0 likes
#qwen3

@HuggingPapers: FlowBalance: verifier-grounded self-improvement for reasoning models Improves math reasoning by +2.12 avg over GRPO on …

X AI KOLs Following ↗ · 2026-09-08 Cached

FlowBalance introduces a verifier-grounded self-improvement technique that improves math reasoning performance by an average of 2.12 over GRPO on the Qwen3-8B model, offering faster training, enhanced stability, and greater solution diversity.

0 favorites 0 likes
#qwen3

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

This paper proposes Feedback-Enriched Environments (FEEs) to bootstrap self-evolving agents in long-horizon tasks by enriching feedback for reinforcement learning, showing improved stability and performance on benchmarks using Qwen3 models and RL algorithms.

0 favorites 0 likes
#qwen3

Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment

arXiv cs.AI ↗ · 2026-09-03 Cached

This research report evaluates post-training ternarization of the Qwen3-4B model, achieving a 1.641-bit effective weight representation with substantial storage compression, while noting a performance trade-off and unresolved deployment acceleration issues.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback