Tag
Sam Altman expresses surprise that OpenAI's models have become proficient at design tasks.
Leading theoretical physicist Yuji Tachikawa reports that Claude Fable solved a research problem that had stumped him and his collaborators for six months.
GPT-5.6 significantly outperforms published state-of-the-art on a fundamental mathematical problem about gradient flow length, achieving exponential improvements. This marks a major advance in AI's ability to reason about complex mathematical questions.
The author describes being deeply impressed and unsettled by GPT 5.6 and Codex, highlighting the model's ability to decompose tasks, recall errors, and propose an optimized process with specific efficiency gains.
Peter Gostev compares Fable and GPT-5.6-Sol, describing Fable as more intelligent and eloquent but less reliable, while GPT-5.6-Sol is a diligent workhorse that excels at executing tasks and maintaining code patterns. He provides detailed observations on UI, writing, robustness, and other features.
Ahmad Osman predicts that within 18 months, a GPU like the RTX 5090 will be able to host intelligence equivalent to GLM 5.2.
Frontier models have become so proficient at generating web apps that the author now benchmarks them on building a WebGPU water renderer from scratch, comparing Opus-4.8 and Fable-5 orchestrator with GPT-5.5 implementer.
Corey Quinn apologizes for underestimating OpenAI's Codex, praising its significant improvement and capabilities.
The author questions whether the reported capabilities of the Fable 5 AI model are genuine or part of a psychological operation, citing lack of evidence and suspicious timing from AWS and NSA claims.
Anthropic demonstrates that AI systems can now perform world-modeling, as evidenced by the Fable standoff experiment.
This article discusses whether it is realistically possible to achieve AI capabilities comparable to Claude or Codex using locally-run models, exploring the current state of open-source alternatives and their limitations.
A Reddit user questions why some people dismiss AI capabilities despite their own positive experiences with AI solving complex problems, suggesting a disconnect between public perception and actual AI performance.
A speculative question about whether a super intelligent AI could learn to modify human biology.
The author reflects on how Claude has surpassed them in creative guidance and reasoning, catching them out with better judgment and understanding.
The tweet comments on the concept of jagged intelligence, noting that code is an exception to the pattern where fable is not significantly better.
Claude Fable 5, a new Mythos-class AI model from Anthropic, is claimed to excel at coding but still lacks design capabilities.
OpenAI reports early signs of recursive self-improvement in current AI systems, a potentially significant development in AI capabilities.
A user reports that Claude is excellent at generating optimized travel routes on Google Maps, personalizing directions for walking, driving, or taxi, and found it perfect for planning a trip to Tokyo.
A detailed critique of the METR AI time horizons graph reveals numerous severe methodological errors, including biased human baselines, unmeasured data, and test-training contamination, undermining its conclusions about AI capabilities.
The author observes that AI has increasingly made coding a solved problem.