All oneshots from Kimi-K3, looks better than opus4.8.
Summary
A user ran Kimi K3 through 34 oneshot prompts and found it outperformed Opus 4.8 in HTML/screenshot/gif generation evals while being far cheaper.
Similar Articles
Kimi K2.6 is a legit Opus 4.7 replacement
A user reports that Kimi K2.6 is a strong alternative to Claude Opus 4.7, capable of handling ~85% of tasks at comparable quality while offering vision and browser-use capabilities, suggesting frontier models may not always offer unique advantages.
@eliebakouch: kimi K2.6 vs K2.5, mythos, opus 4.7, and cursor composer 2 (based on K2.5) on every benchmark i could find tl;dr: it's …
Kimi K2.6 shows strong performance gains over K2.5 and rivals like Mythos and Opus 4.7 across multiple benchmarks.
We Tested DeepSeek V4 Pro and Flash Against Claude Opus 4.7 and Kimi K2.6 (11 minute read)
DeepSeek released V4 Pro and V4 Flash under MIT license on April 24, 2026. In benchmarks against Claude Opus 4.7 and Kimi K2.6, V4 Pro scored 77/100 at $2.25, placing between Opus 4.7 (91) and Kimi K2.6 (68), while V4 Flash scored 60/100 at $0.02, the cheapest in the comparison, with a 75% discount on V4 Pro through May 31.
@nutlope: https://x.com/nutlope/status/2067281915887943890
A comparison experiment shows that Kimi K2.7 Code generates landing pages at about 94% lower cost than Claude Fable 5 with similar quality, especially when given design context via an MCP server.
I got Kimi-k3 running.....
User successfully runs the Kimi-k3 model using llama.cpp on high-end hardware, achieving low tokens per second (0.41 prompt eval, 0.23 generation).