Ling Tiny 3.0 is a glimpse of the future

Reddit r/LocalLLaMA Models

Summary

The author runs the Ling Tiny 3.0 AI model on a 2017 laptop without GPU, achieving 10 tokens per second and completing tasks like code generation, showcasing the potential for edge intelligence on existing hardware.

I've been playing around with Ling 3.0 Tiny, which is an 8 billion parameter model (MoE, 1B active). And I've had a lot of poignant thoughts as a result. Just for fun, I got it running with llama.cpp on an old laptop. This is a laptop from 2017 with a 7th gen i5 and 8 gigs of RAM, like barely even usable for modern tasks. No VRAM, no GPU. Well, I got Pi running on it and asked it to make a script to scan the local network for all available models on llama.cpp servers. It started chugging along at about 10 tokens per second. And 20 minutes later, it was done. It had several back and forth turns with writing code, running it, getting feedback and iterating. It's a simple task, yes. But it's a task that would have taken me an hour or two to do in 2020. It's just incredible that such a potato hardware is actually accomplishing something useful on a reasonable timeline. One billion active parameters is so small that an old CPU can run at 10 tokens a second with basically no optimization effort. I.e. I just built llama.cpp and ran the first Q6 quant I found. I guess my point is, do you remember that feeling a year or two ago when you looked down at your expensive GPU rig and thought, wow, the computer writes the code itself now? It actually feels like something, like it's intelligent somehow. Well, now that's starting to happen for every potato casual computing device that's been made since 2015. Obviously more expensive rigs will be always be much more power efficient and cost efficient and fast at producing tokens. So it may never be practical to actually use old potato hardware. But maybe it will make sense. There are all kinds of things that a mildly intelligent computer could do in the background. So it may be a new beginning for edge intelligence. No new compute required. Just everything that already exists can suddenly start doing intelligent tasks. Not sure. I guess I'm just saying I'm amazed I've had that weird sensation when looking at my GPUs and feeling like there's something more than just bits in there, but for an old CPU. (no I don't think it's conscious. Not talking about that.)
Original Article

Similar Articles

ling 3.0 flash/tiny base models

Reddit r/LocalLLaMA

InclusionAI has open-sourced the Ling-3.0 series, featuring highly efficient language models with sparse MoE architecture and hybrid linear attention, providing checkpoints at various training stages to support research and innovation.

Is Ling 3 tiny underrated for its size?

Reddit r/LocalLLaMA

A user discusses the Ling 3 tiny model's benchmarks, comparing it to Qwen3.5 9b and questioning if other open-source models are being overlooked.