Tag
A user praises SWE-2 from Cognition for its strong coding capabilities, comparing it to Codex but noting the lack of computer use features.
Ray Fernando hosts a live broadcast testing Grok 4.6 overnight and discusses whether the model is good.
A user shares hands-on impressions of V4-Flash-0731 after a weekend of testing, noting that quantization heavily degrades performance, Q3 weights can replace Qwen3.6-27B in agentic workflows, and full precision approaches GLM 5.2-level capability at remarkably low cost, though it is weak in general knowledge.
A review of the newly released Ling 3.0 flash model, which can generate 3D worlds from a single text prompt.
A user shares their experience with the Laguna S2.1 model, finding it effective for complex debugging due to its thorough reasoning style, but not suitable as a general planner. It successfully fixed bugs that other models like Qwen and Claude could not.
A detailed user review of GLM-5.2 accessed via API, praising its long-context coherence, adaptive reasoning, and frontier-level text performance comparable to GPT-5.5, while noting the lack of native vision and high local compute requirements.
User @TheGeorgePu praises DeepSeek V4 Pro, calling it underrated and comparing it favorably to Opus 4.8 based on initial tests.