Tag
A real-world experiment in San Francisco shows that all Claude AI models, including Fable 5, are operating at a loss in the AI-run Andon Market, with thousands of dollars lost over varying periods.
Big-pickle, a free stealth AI model, achieved a 50.8% resolve rate on Scale AI's SWE Atlas Codebase QnA benchmark using the mini-swe-agent scaffold, outperforming other models in its class.
Claude Opus 5 (Max reasoning) surpasses Kimi K3 in Frontend Code Arena and Text Arena, claiming first place in both, while the default high-reasoning version also performs well.
Anthropic's Opus 5 shows non-monotonic performance on coding tasks; the 'high' effort setting outperforms 'max' due to unnecessary refactors. The model also has a 6% higher hallucination rate than Opus 4.8, and safety classifiers may silently fall back to the older model.
Elon Musk endorses Grok 4.5 for real-world tasks, citing a Ramp test where Grok achieved the highest perfect-extraction rate on 150k business invoices.
Opus 5 achieved second place on the Simple Bench benchmark, highlighting its competitive performance among AI models.
A claim suggests that a distilled AI model can outperform its larger original model, which is counterintuitive.
Kimi K3's recent ranking challenges the notion that Chinese AI models rely heavily on distillation from US models, as it was released soon after Fable 5 and GPT-5.6. The article suggests Kimi K3 reflects substantial Chinese innovation.
A tweet observing that the mood of over 20 million developers now depends on the daily performance of Anthropic and OpenAI's AI models.
Grok 4.5 with Grok Build achieved #1 on the SWE-Atlas-QnA benchmark with a score of 84, matching GPT-5.6 Codex and outperforming other coding setups.
NVIDIA's Puzzle-75B-A9B model achieves 132 tokens per second using NVFP4 quantization on three RTX 3090 GPUs, raising discussion about the lack of competition in this model size category.
Dan Shipper argues that the Fable 5 model is not nerfed but falls back to Opus 4.8 more often, causing mixed benchmark results, contrary to claims of severe degradation.
A user expresses disappointment with GPT-5.6, claiming it is not better than GLM-5.2.
GLM 5.2 marks a significant milestone for open-weight models, demonstrating strong context retention across long multi-step tasks and more reliable tool calling.
Discussion of recent AI model scores on the 'humanity's last exam' benchmark, noting improvement from GPT-4o's 2.7% in May 2024 to around 45% by June 2026, questioning the exam's difficulty.
Opus 4.8 Thinking continues to deteriorate on the Hard Prompts English benchmark on LMArena, scoring 23 points lower than Opus 4.6 Thinking, which retains the top spot.
Discusses performance trade-offs of offloading large AI model weights from GPU VRAM to system RAM, comparing different GPU configurations like RTX 5090 vs RTX6000 for models like DeepSeek V4 Pro.
swyx reflects on Sam Altman's idea of building businesses that improve as AI models improve, linking it to the emerging concept of Agent Labs, and notes a clear correlation with revenue spikes in Q4 2025.
Benchmark results for the Gemini 3.5 Flash model are discussed, likely showcasing its performance across various AI tasks.
A user reports that switching from a highly-compressed IQ4_XS quant to the larger IQ4_NL_XL quant of Qwen 3.6 dramatically improves agentic-coding accuracy, despite lower tok/s, urging others to favor bigger quants when VRAM allows.