Jovan from UkisAI thanks the community for the success of Swift Qwen 3.8 27B, an open-source model that reduces token usage by 58.3% and increases speed by 1.95x through efficient thinking patterns. Future improved checkpoints and models are planned.
Hey everyone, Jovan from UkisAI here, a small lab building the tech to make tiny frontier LLMs possible (and doing it open-source!) The purpose of this post is simply to thank the community for all the amazing finetunes, quantizations and overall improvements over our original release which made our model get attention and the support for us to continue building in this direction! If it weren't for you guys going out of the way to contribute we wouldn't have half the results of this. For context: Swift Qwen 3.8 27B is our first open-source model release. It is proof of how penalizing pathological overthinking patterns inside of small LLMs can bring their token usage down -58.3% and speed x1.95 without losing accuracy by not training them to think shorter directly but rather to think more efficiently. We are continuing to build and are about to drop: - Swift1.5 Qwen3.8 27B (an improved checkpoint of the model with some training bugs fixed and more RL) - Swift Qwen3.8 Flash Next in the upcoming week week, we are now running the benchmark suite to not give out premature or incomplete results. This time we ran even more benchmarks as you guys suggested, including more coding and long horizon! It would be amazing if those of you who tried Swift would let us know what quants, features, changes you want to see in our upcoming model releases so we can do it better this time as we didn't even think about half of the stuff you guys were requesting last time :) Let the era of non-slop finetunes begin! EDIT: Links - https://huggingface.co/ukisai/Swift-Qwen3.8-27b https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF https://huggingface.co/bartowski/ukisai_Swift-Qwen3.8-27b-GGUF
UkisAI has post-trained the Qwen 3.8 27B model to reduce unnecessary thinking tokens by 58% and achieve 1.95x speed up with less than 1% accuracy loss, open-sourcing the model and offering a free research API.
Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B, reducing thinking tokens by 58.3% while maintaining near-identical performance.
A user thanks the community for assistance in setting up the Qwen 3.8 27B AI model, which now works with HomeAssistant and includes vision capabilities, enabling local AI use.
Qwen releases Qwen3.8-27B, an open-weights 27B dense vision-language model with major gains in coding, professional work, and long-horizon agentic tasks, available in FP8 with flexible thinking control.