GPT Live clone on an RTX 3060

Reddit r/LocalLLaMA Tools

Summary

A developer tested their local voice assistant project, Fulloch, which uses Qwen models on an RTX 3060 to replicate GPT Live functionality, showing impressive performance with open-source tools.

I wanted to see how my fully local home voice assistant compared to the latest GPT Live, so I tested it using the same conversation used in their "Improved Intelligence" demo. In this video they ask the AI to see if a flight route is feasible and while it is figuring that out they continue to ask it questions about what they can eat at each destination. The models I ran are (all squeezed into 12 GB VRAM): Speech recognition: Qwen3 1.7B ASR PyTorch LLM: Qwen3.5-9B-UD-Q4_K_XL GGUF with 12K context Voice: Pocket TTS PyTorch So I copied the exact query and threw it at my Fulloch project. This blog post has the video of the interaction and breaks down how it did. The final report and searches it did are also linked in that blog post. The video has sped up two sections where I had to wait for the 9B model to finish thinking through the task, but it did the whole thing in under six and a half minutes. In the end it couldn't find a suitable flight route but it gave good food and restaurant recommendations and did it all pretty quickly. I am still impressed with how well the Qwen3.5 9B model does with these sorts of tasks with such a small footprint. If you want to try it out yourself the source code and pre-compiled docker images can be found at https://github.com/liampetti/fulloch.
Original Article

Similar Articles

GPT‑Live

Hacker News Top

OpenAI announces GPT-Live, a new full-duplex voice model that enables more natural, real-time conversations by allowing simultaneous listening and speaking, with GPT-5.5 as the backend model.

Testing Qwen 3.8 27B running locally on a single 5090

Reddit r/ArtificialInteligence

The article demonstrates the capabilities of running the Qwen 3.8 27B AI model locally on a single 5090 GPU, using Row-Bot to generate a rich animation showcasing tasks from language synthesis to physics simulation.