A developer shares their experience using the quantized Qwen 3.8 27B model on a 4060Ti GPU to autonomously write and merge a feature branch, demonstrating significant progress in local AI for coding tasks.
Howdy folks, You might (or likely not) know me around here with shilling of Pi harness, and the use of Pi as productivity assistant and KB manager. Lately, I have also been telling anyone who listens to try Qwen 3.8 27B IQ3_K_XXS by Unsloth, which fits fully on 4060Ti 16GB with nearly 100k context at Q8, if you drop both mmproj and MTP. This setup nets you average 800tk/s prefill and 20tk/s decode (drop to 17tk/s average as the model passes 64k context depth.). Despite using this setup extensively for personal assistance, I actually didn't intend to use it for coding at all, because I don't think it can handle actual codebase with proper architectural and engineering conventions that it must follow. However, today, I decided "what the heck", and let it work in the same way I let Minimax M3 works on my codebase. The workflow in my codebase is like this: I provide the feature description and a hint about the general area where the relevant code could be The agent must investigate the code base and existing architectural decision record (ADR), and return to me with a clear implementation plan with justification. Interactive refinement of the plan. After receiving approval, the agent would write down ADR and plan for me to do final review. After the final approval, the agent can continue coding until it implements the code and passes all the QA gates. I do final code review. If I'm happy, I'll allow the agent to commit. For context, the project I'm working on is in the attached image. It's an extensible framework that I can drop on all of my machines, and then control remotely from VPN so that I can download models, change llama-swap configuration, deploy comfyui, deploy aitoolkit, and other goodies that I usually use. The goal is to be able to run headless (to save VRAM), without having to SSH into the server and muck around. Anyhow, that's not the point. The task for the agent to implement is like this: My framework ships with source code a set of recipes to configure various model families for 16GB VRAM. It supports no other way for end-user customisation. It makes life pretty hard for me to tinker without pushing update to git. So, I want the ability to load extra config from the correct XDG path that the llama-swap module already use. And this extra config could override the built-in defaults. I want to expose this reload ability to UI. The agent must NOT break the existing plug-in architecture with clear separation of concern that I enforce. The task started at around 1200. The agent was given full autonomy to code at around 1220. It finishes around 1300. The available context detected by pi was 98k. There were 3 compaction from the beginning to the end. The agent only failed text edit a few times, whilst consistently pushing parallel tool calls per turn. By 1310, I finished manual code review and approved the merge. The larger context and why I shared this: I'm utterly shocked by the progress in both model training and quantization! I have built a few harnesses in the past, and I have been pushing for people around me to "go local". But, personally, I knew that "go local" was only usable for generic chatting, not agent with tool calls, unless you have huge stack of GPUs. This conclusion was reached when I tried and pretty much failed to get GPT-OSS 20B to run a multi-agent workflow, where each subagent edits a shared WIP file to put its own contribution in. The edit keeps failing, or one agent would decide to use bash to overwrite the entire file with its own content. The Qwen 30B-A3B Q4_K_XL was also helpless when running in Aider in the sense that it just cannot edit code file correctly. Enter this 27B IQ3_K_XXS. It plans. It calls tool in parallel. It edits multiple spots in the same file at once, successfully. All while being brain damaged by IQ3XXS quant. It feels utmost unreal when I approve this merge. Years after buying this 4060Ti, I finally managed to get the LLM to do things that I hope it would do. And now, I want more. R9700, here I come. Random P/Ss: Most funny thing in the session: the 27B thinks "damn it! I send the two edits to the wrong file". The strangest thing I ask 27B to do: make me a menu and groceries list for the week. It takes my journal, planned time table, review my recipes in the KB, and reason it way toward a full 7 days menu based on my preference and available recipes, and then synthesis a groceries list from those, adjusted to the right portion count. I'm grabbing Gemma 31B Q3XXS and Muse Glimmer Q3XXS now. Let's see if they are any good.
A user is testing the newly announced Ternary Bonsai 2 27B AI model, a smaller and quantized version of Qwen3.8 27B, on an NVIDIA 5070 Ti GPU and finds its performance impressive.
Developer achieves productive local agentic coding with Qwen3.6-35B 4-bit MLX and pi.dev tool, completing real tickets efficiently on current hardware.
A user demonstrates Qwen 3.6 running autonomously on an AMD 7900 XTX GPU, locally creating an Android app — described as a sci-fi reality achieved today.
The author shares their experience running Qwen3.6 35B-A3B locally on an ASUS Zenbook Pro 14, achieving 27 TPS at 32k context, marking a personal milestone towards fully local AI for privacy.
Qwen 3.6 27B runs fast on 16 GB VRAM thanks to 'Pure Quant' technology, achieving 40 tokens/s with MTP and supporting 64k contexts, enabling local AI on consumer GPUs like RTX 4060 Ti.