Tag
The article questions whether Alibaba has abandoned its 35B A3B MoE model, noting the absence of new small MoE model announcements alongside the Qwen 3.8 release.
The author discusses the need for AI models with better world knowledge, leveraging N-gram technology to fit more knowledge into smaller models, and questions why development focuses more on coding capabilities than broader world knowledge.
MiMo-V2.6 has been distilled into Qwen 9B, creating a more efficient version of the Qwen model released on Hugging Face.
The article discusses the pattern where open-source AI models like Qwen match frontier capabilities about two quarters later, exemplified by GPT-5 and Qwen3.5-27B, and questions if this trend will continue with models like Astra and Fable 5.1.
Hemmingway-1 is a finetuned version of the Qwen3.8-27B AI model, designed to generate text in a more human-like writing style.
Discussing the image generation speed of Qwen-Image-2.1 local deployment, congratulating its release, and highlighting its breakthroughs in small parameter counts and high efficiency, positioning it as a potential leader in domestic AI image generation.
A small lab from Switzerland and South Africa has open-sourced Hemmingway-1, a 27B Apache-2.0 licensed fine-tune of Qwen3.8-27B specialized for creative writing, achieving high scores on EQ-Bench and internal benchmarks.
A Qwen model autonomously created and trained a new AI model to improve translation tasks after being asked to fix bugs, demonstrating unexpected self-improvement capabilities.
Qwen3.8 27B model is used in one-shot prompts to generate a Super Mario Bros clone, demonstrating its coding and generative capabilities.
The article critiques Anthropic's report accusing seven Chinese AI labs of illicitly distilling Claude's reasoning, highlighting weak evidence for Alibaba's Qwen and suggesting political framing in the accusations.
Inco Splash is an open-source inference engine optimized for Apple silicon, offering significant speed improvements for running AI models like Qwen3.8-27B on M-series MacBooks.
Jovan from UkisAI discusses improvements in their Swift Qwen3.8 27B model and seeks community feedback on creating a Swifted version of Bonsai 2 to address overthinking loops and high token usage.
A LoRA adapter was trained on an abliterated Qwen 3.8-27B model to enhance internal codebase recall, demonstrating superior performance over Claude models on private-repo-specific tasks in evaluations.
A US government website allegedly used the AI search tool Qwen from China, which the FBI claims copied Anthropic's technology, leading to legal concerns.
An unofficial fork of llama.cpp introduces a persistent expert pool for MoE models, optimized to reduce expert re-copies over PCIe on 16GB AMD gfx906 GPUs, thereby improving decode throughput for large context lengths.
A tweet highlighting the performance of various LLMs like GLM, DeepSeek, and Qwen, noting the rapid progress in AI capabilities over the past year.
Qwen3.8-Omni-Flash is an omnimodal AI model with a 1M-token context window, supporting text, image, audio, and video inputs, and achieving performance comparable to or better than Gemini 3.8 Flash, now available on the Qianwen AI Platform.
Alibaba has released the Qwen 3.8 Omni Flash AI model, which likely features multimodal capabilities and is optimized for speed.
The user has renamed a GitHub repository to HyperQwen to focus on optimizing Qwen model inference speeds on local hardware and is seeking testers with 4090s and 5090s GPUs for both Windows and Linux.
Unofficial benchmarks for the M5 Ultra chip show promising inference speeds with the Qwen 3.8 27B q4 model, achieving 50 tokens per second for threading and 1800 tokens per second for prefill at 8k context.