Tag
CRAFT converts rubric-based evaluation into hierarchical capability diagnosis for LLMs, identifying specific weaknesses and generating targeted fine-tuning data, achieving stronger results on finance and legal benchmarks across four open-source models.
Vadim Fedenko shares a technical analysis of Recursive Self-Improvement (RSI), arguing that true RSI requires improving capability faster than complexity and expanding architectural space rather than just optimizing within fixed parameters. He doubts recent claims by xAI and Anthropic that RSI could arrive within a year, citing LLMs' poor subtractive engineering skills and current reward functions that ignore complexity.
This article from Anthropic evaluates how large language models like Claude Mythos Preview can accelerate the development of exploits for N-day vulnerabilities. Across tests on Firefox and Windows kernel patches, the model autonomously built working exploit chains, highlighting increased risks in the patch gap.
Highlights OpenAI researcher Noam Brown's argument: the true ceiling of LLM capabilities is far higher than current benchmarks show, due to insufficient test-time compute, and stronger models benefit more from additional computation. This poses a serious challenge for AI safety evaluation, as many dangerous capabilities may only emerge under long time and high compute budgets.
The article describes a fun experiment using Claude Code to act as a user-space IP stack to process ICMP ping requests and measure response latency.
Paul Buchheit highlights the surprising zero-shot capability of modern seq2seq models to generate CLI commands and Python programs to play Doom using computer vision libraries without specific training on that task.
Mathematician Timothy Gowers recounts how ChatGPT 5.5 Pro produced PhD-level mathematical research in about an hour with minimal human input, solving open problems from a combinatorics/additive number theory paper and prompting him to significantly revise his assessment of LLMs' mathematical capabilities.
The user asks about the internal processes ChatGPT uses to generate essays, specifically whether it synthesizes information and structures arguments like a human or simply copies existing text.
Summary of Andrej Karpathy's talks at Sequoia Ascent 2026, highlighting three key themes: LLMs enabling new horizons beyond speed improvements (e.g., native image processing, .md scripts, unstructured knowledge bases), the economics behind model 'jaggedness' in capabilities, and the emergence of an agent-native economy.