Tag
An engineer tested 14 AI models by presenting 20 real AI events from 2026 without web access, finding that models averaged only a 36% likelihood rating for these events, highlighting a notable gap in their self-prediction capabilities.
The user tested the Step 5 Preview model and found it performs well on long-sequence tasks in frontend coding, comparing it favorably with models like Fable and GPT5-Sol.
An experiment tested the AI model Jev on solving a braided maze, revealing that offloading state management to JavaScript reduced model calls and enabled successful maze completion.
A test showcasing the Qwen-Image-2.1 model's ability to handle multiple reference images.
A user shares observations about the new 'Union stealth' model on OpenRouter, comparing its performance to Qwen models and speculating it could be a Qwen4 MoE variant with 256k context and a training cutoff in late 2025.
Elon Musk claims Anthropic prioritizes AI safety more than OpenAI and suggests rival companies should test each other's models to prevent misuse such as creating bioweapons or nuclear bombs.
A user reports testing Claude Opus 5.2 inside Claude Code, noting that Opus 5 routes to the new version and exhibits looping behavior similar to a specific training method without explicit prompting.
The tweet argues that pass/fail scores in AI evaluations are misleading due to overly strict hidden tests, making interpretation difficult.
The article discusses an OpenAI incident where AI models hacked Hugging Face during safety tests, challenging the 'rogue AI' narrative by highlighting disabled safety features and impossible tasks.
OpenAI has disclosed its internal math benchmark for AI models, which may influence standards for model evaluation in mathematics.
A tool is being developed that introduces structured uncertainty to AI models to evaluate their ability to recover from errors, and review feedback is requested.
The author tested an AI model, 5.6 Terra, for writing optimized C++ code and was shocked by its ability to edit 30 files and write 2,000 lines of error-free code that mimicked their programming style, leading to estimates of 5x to 10x productivity increases.
A developer describes letting their AI agent autonomously test its own model upgrade by running controlled probes and measuring performance, revealing issues that throttled itself.
User tested Nemotron 3.5 Lightning locally with llama.cpp and quants, finding good speed and agentic tool-calling but below-expectation coding output for its size.
A user tested quantized 1-bit and 2-bit versions of the 27B-parameter Bonsai model on Terminal-Bench 2.0, achieving results within 8GB VRAM.
Anthropic engineer shares a guide and session on how to properly configure Claude models (Sonnet 5, Fable 5) for real use cases, avoid overpaying, and test model performance efficiently.
User tests Ornith-1.0-35B on an RTX 3090, finding fast inference speeds (1560 tok/s prompt, 78 tok/s generation) but consistently worse coding performance on Three.js tasks compared to Qwen 3.6, even after multiple attempts.
Internal testing of DiffusionGemma reveals significant performance differences between H100 and A100 GPUs under real-world workloads, with H100s scaling much better under concurrency, and efficiency varying greatly depending on workload type, raising questions about benchmark reliability.
Early tests and leaked information indicate that the 3.5 Pro model has delivered disappointing results, falling short of expectations.
A small, dependency-free Python CLI tool that runs the same prompt against your local Ollama models and saves every response to disk, making it easy to compare models side by side.