Base-10's Charlie O'Neill on why Kimi and GLM are "almost objectively" better than Opus 5

Reddit r/LocalLLaMA News

Summary

The article recommends a YouTube episode featuring Charlie O'Neill from Baseten, discussing how Kimi and GLM models are superior to Opus 5 and addressing skepticism about AI progress and AGI.

Edit: Spelled Baseten not Base-10 Full episode of this available at https://www.youtube.com/watch?v=PrSf7IOYu-I It's interesting to see how Dwarkesh has had to come around to the evidence that we are well on our way to creating AGI and even RSI in the last few months, despite historically being very skeptical. I highly recommend people interested in large language models check out this particular episode, because it dispels a lot of mythology about stuff like plateaus from lack of data etc. For those who thought we were hitting a wall a year ago, it turns out there was a ton of low hanging fruit and the researchers in this episode discuss what that fruit was. They also extrapolate these trends into the future. It's funny this subreddit is becoming rather skeptical of AI progress, which to put diplomatically, I think is based on a lack of information and too much time on Reddit.
Original Article

Similar Articles

IS GLM 5.2, Kimi 2.7 still worth it?

Reddit r/LocalLLaMA

A discussion questioning whether older AI models like GLM 5.2 and Kimi 2.7 remain relevant for coding now that newer models such as Kimi K3, Qwen 3.8 Max, and DeepSeek V4 Pro are arriving.

Can Kimi K3 solve the same problems that Claude Fable can?

Reddit r/LocalLLaMA

A discussion questioning whether open-source models like Kimi K3 or GLM can replicate the mathematical and cybersecurity problem-solving achievements recently demonstrated by closed-source models from OpenAI and Anthropic.

UPDATE: "Gentle Coding" is mathematically proven. 1,500+ test runs show major gain for Kimi K2.6 and even more for GLM-5.1! GPT 5.4/5.5 and Claude Sonnet 3.5/Opus 4.6 also better, with ZERO REGRESSION ACROSS THE BOARD.

Reddit r/LocalLLaMA

The 'Gentle Coding' technique is empirically validated across 1,500+ tests, showing significant improvements (zero regression) for multiple models including Kimi K2.6, GLM-5.1, GPT 5.4/5.5, and Claude Sonnet 3.5/Opus 4.6 by reducing looping and hallucinations.