Tag
A deep-dive guide on running Qwen3.8 Flash-Next on two 96 GB Huawei Atlas 300I Duo cards via vLLM Ascend, detailing the two-chip-per-card layout, passive cooling, and software optimizations that brought the machine from incoherent ~1 tok/s output to coherent ~30 tok/s single-request and ~61 tok/s aggregate decode, including a full GPQA Diamond evaluation.
Cloudflare released Clef, an open-weights 27B multimodal decision model that takes a state and a schema of typed questions as input and returns probabilities for each option in a single forward pass, with no free-form generation or output parsing. It is post-trained from Qwen3.8-27B, ships on Hugging Face with a smaller Clef-Flash variant, and is compatible with the Jev/SystemOne API.
BAAI releases AREX-2, a 27B long-horizon self-improving agent model based on Qwen3.8 that learns to propose, measure, reflect, and revise solutions over multiple test-time rounds, achieving strong results on coding, ML engineering, and deep research benchmarks with only 27B parameters.
The article highlights the Strata inference engine, which significantly outperforms llama.cpp for running Qwen3.8 models on a laptop with 12GB VRAM and 64GB RAM, achieving up to 50 tokens per second for text generation and 1500 tokens per second for prompt processing.
This paper introduces ActFirst-OPD, a framework that accelerates on-policy distillation for multi-turn language agents by decoupling action execution from full reasoning, achieving significant training speedups while maintaining performance across benchmarks.
The user built a low-cost setup using five ex-mining BC-250 boards to run the Qwen3-Coder-Next AI model, achieving around 40 tokens per second at 30k context with plans to expand.
The paper proposes a general asynchronous LLM framework that lets users define inference coroutines with overlapping memory states, demonstrating that Qwen3.x models can handle streaming video, video games, and monitoring tasks asynchronously without task-specific training.
HuatuoGPT-3-27B is a medical language model built on Qwen3.8-27B using One-stage Policy Optimization (OnePO), a reinforcement learning method for domain adaptation without supervised fine-tuning.
The user has implemented nvfp4 KV cache support for the Qwen3.8 model on a heterogeneous GPU setup using custom CUDA kernels and quantization to optimize performance.
Optimized performance for running the Qwen3.8 27b NVFP4 model on a single AMD Radeon R9700 GPU, achieving up to 153 tokens per second in decode and 3,619 tokens per second in prefill with improved concurrency.
This controlled study compares supervised fine-tuning and reinforcement learning methods for training tool-calling agents across different datasets and model scales, finding that SFT with LoRA is strongest in-distribution while RL shows slight advantages in cross-dataset transfer.
A user built a game to test the performance limits of the Qwen3.8 27b AI model using an overclocked RTX 3090, and plans to release it as open-source on GitHub for community contribution.
The user describes struggles with context, compaction, and memory management for local AI models using pi.dev plugins and seeks suggestions for solutions that handle varying model context windows and VRAM limitations.
The user describes a hardware configuration with dual RTX 3090 GPUs and an AMD EPYC CPU for running the Qwen3-Flash-Next model using llama.cpp, achieving 38 tokens per second in single-stream inference, and seeks advice on whether to add a third GPU or upgrade the CPU to improve performance, especially for running multiple parallel agents.
The article describes fine-tuning Qwen3-TTS to add emotion control tags, overcoming training challenges like codec prefix inconsistencies and generation concurrency issues, and discovering that emotion can be manipulated via affine transformations in speaker embeddings.
An author improved MLX performance on an M3 Ultra Studio to achieve 73 tok/s for a 4-bit Qwen3.8-Flash-Next model, which shows intelligence scores matching GPT-5.6 Sol, highlighting the growing potential of local AI.
Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B, reducing thinking tokens by 58.3% while maintaining near-identical performance.
FlowBalance introduces a verifier-grounded self-improvement technique that improves math reasoning performance by an average of 2.12 over GRPO on the Qwen3-8B model, offering faster training, enhanced stability, and greater solution diversity.
This paper proposes Feedback-Enriched Environments (FEEs) to bootstrap self-evolving agents in long-horizon tasks by enriching feedback for reinforcement learning, showing improved stability and performance on benchmarks using Qwen3 models and RL algorithms.
This research report evaluates post-training ternarization of the Qwen3-4B model, achieving a 1.641-bit effective weight representation with substantial storage compression, while noting a performance trade-off and unresolved deployment acceleration issues.