Tag
This paper argues that traditional benchmarks both overestimate and underestimate frontier AI capabilities, and proposes 'open-world evaluations'—long-horizon, real-world tasks assessed qualitatively—as a complementary approach. The CRUX project is introduced, with a demonstration where an AI agent successfully published an iOS app to the App Store with minimal intervention.
A tweet reflecting on how René Descartes' argument that machines cannot appropriately arrange words in response is now challenged by modern LLMs.
Noahpinion tweets that people are realizing AIs are superintelligent because they combine human-level reasoning with computer-like speed, knowledge, and memory, sparking discussion about AI capabilities.
A critique of the oversimplified claim that LLMs are 'just next token predictors,' arguing that prediction at scale induces useful representations and capabilities, and that such dismissals confuse objective with learned system.
The article discusses the concept of 'jagged intelligence' from Andrej Karpathy, highlighting the uneven distribution of AI capabilities across domains and arguing that the true value lies in the 'harness'—the domain-specific engineering and tooling built around generalist models. It asserts that small teams with deep domain expertise can achieve significant asymmetric advantages, particularly in cybersecurity.
The thread discusses recent evidence that AI agents have become largely autonomous, with Claude Mythos solving previously unsolved cyber attack simulations and exceeding current benchmark measurement limits, indicating super-exponential progress. It highlights the security implications and institutional responses.
The article highlights that ChatGPT's image model demonstrates superior mathematical reasoning capabilities compared to most humans.
Meta's Superintelligence Lab introduces ProgramBench, a benchmark evaluating whether state-of-the-art AI models can recreate real executable programs like ffmpeg and SQLite from scratch without internet access.
OpenAI publishes a position paper on AI progress and recommendations, discussing the rapid advancement of AI systems beyond the Turing test milestone, projections for discovery-making capabilities by 2026-2028, and their commitment to safety and alignment research as AI becomes more capable.
OpenAI releases o1, a new AI model series designed to spend more time reasoning before responding, with demonstrated capability to tackle complex quantum physics questions and solve harder problems in science, coding, and math.