A user shares updated benchmark results for DeepSeek V4 Flash on SlopCodeBench using local quants (antirez imatrix quant) with the pi harness, showing improved performance over previous runs but still slower than the hosted API.
The author shares hands-on experience with LFM 2.6B, a small model designed for phones, praising its speed and usefulness for quick tasks like summarization and autocomplete, though it has a 128k context limit.
A developer successfully ran a 5.2 million parameter MoE LLM quantized to INT4 on an ESP32 Dev Kit V1 using only 81KB of SRAM by streaming experts from flash, achieving about 5 tokens per second.
Guillermo Rauch highlights Grok Imagine Image 2.0, now available on Vercel AI Gateway and ranking #2 on Arena.ai's leaderboard. Vercel offers access via AI CLI and a live playground.
A trimmed English-only GGUF version of Kimi K3 (IQ2-XXS) reduces model size from 711GB to 478GB by removing multi-language components, with early tests suggesting it may match or outperform the standard 2-bit version on coding tasks.
Grok announces Imagine Image 2.0, a next-generation image model with precision editing, crisp text rendering, improved factuality, and real-world usefulness.
DeepMind's hurricane AI model gives forecasters an extra day of warning and is being open-sourced as WeatherNext models, though researchers don't fully understand how it works.
Grok announces Imagine Image 2.0, a next-generation image model with precision editing, crisp text rendering, and improved factuality for real-world use.
Grok Imagine Image 2.0 (Low) from xAI jumped to #2 in the Text-to-Image Arena, beating its own older quality model and showing significant improvement.
The author shares excitement for the upcoming Qwen 3.8 model, highlighting their experience with Qwen 3.6 27B for local LLM use, and discusses the potential of self-hosted AI to replace subscription-based frontier models.
Meta's Muse Spark 1.2 reaches #4 in the Text Arena, moving the cost-quality Pareto frontier upward with a ~91% price cut while sacrificing only 9 points relative to the top model.
Seedance 2.5 introduces 30-second video generation, multimodal references, multilingual creation, and targeted editing, with API access coming soon on BytePlus for developers and enterprises.
Ahmad Osman shares performance numbers from running DeepSeek V4 Flash 0731 on an NVIDIA DGX Station.
TokenAI, an Egyptian startup, announces Early Access for Horus Cyper Nano 1.0 BETA, a specialized cybersecurity model for offensive security and red teaming research, with open weights planned for September 2026.
Alibaba released Qwen3.8-Max, a 2.4 trillion-parameter sparse MoE model with 95B active parameters per token, 1M token context, and strong agentic and benchmark results, including autonomously coding for days, circuit design, and outperforming rivals on Terminal Bench and PaperBench.
Sam Altman announces that 'astra' is a powerful AI model being prepared for general availability, with additional safety time needed due to its cyber capabilities.
Sam Altman announces that the Astra model is powerful and OpenAI is working to make it generally available, while taking extra time to ensure safety given its cyber capabilities.
Greg Brockman shares a testimonial from Grigori Karapetyan praising GPT-5.6 Sol (Codex) for enabling speedy cybersafety investigations and responses.
Simon Willison tests GPT-5.6 Sol Ultra via Codex Desktop by asking it to recreate a 'Raccoon Heist' game from a four-year-old prompt, resulting in a much better game than Claude Fable 5's version, though it had a bug with oversized eyeballs.
OpenAI's next major model, Astra, shows significant capability gains in agentic coding and cybersecurity, and is being treated as a 'critical' model for cybersecurity under the Preparedness Framework, with additional controls planned.