Tag
OpenAI announces the upcoming launch of their next model, Astra, emphasizing the need to balance AI capabilities with safety and alignment progress.
BenchMIRT introduces a method to audit LLM benchmarks at the individual prompt level using multidimensional item response theory, separating underlying capabilities like safety and general reasoning to reveal what benchmarks actually measure.
This paper introduces reference-grafting, a method to elicit sandbagged capabilities in AI models by editing activations, matching fine-tuning's effectiveness without weight updates or training labels.
The paper introduces a multi-layer taxonomy for large language models comprising 14 capability domains and 91 subskills, drawing from human cognitive science to organize LLM evaluation beyond isolated tasks. It demonstrates operational utility by mapping 15,934 papers across major AI conferences, revealing concentrated attention on language-semantic competence and reasoning while identifying underexplored domains.
Explores whether large language models can solve maze navigation tasks, probing their reasoning and spatial understanding abilities.
An analysis of the current practical capabilities of AI voice agents and the tasks they perform well today.
The article discusses how AI agents have been given capabilities for physical interaction (hands) and voice communication, but most still lack visual perception (eyes), highlighting a key limitation.
A discussion prompt asking the community to share which AI agent capabilities are overrated and which are underrated, based on real-world experience.
An observation that several AI labs besides OpenAI, Anthropic, and DeepMind possess state-of-the-art capabilities they may never publicly share due to economic incentives.
This paper introduces Portico, a reference monitor for revocable capabilities in coding agents, addressing the problem of 'lingering authority' where temporary resource access remains exposed after its justification. It demonstrates that Portico enforces task contracts by revoking capabilities after closure, preventing forbidden effects.
Anthropic released five workshops detailing the latest capabilities and use cases of Fable 5, covering deep dives, managed agents, and deployment strategies.
GPT 5.6 Lunatic offers comparable capabilities to GPT 5.4 at a significantly lower price of $1 input and $6 output, making it a strong alternative to Gemini Flash and GPT 5.4 mini.
Fiona believes AI has raised the ceiling of achievement; an engineer unfamiliar with mobile development used Claude to fill in the App functionality.
The article clarifies that the AI model Mythos was not trained on hacking, and predicts that other AI labs will eventually achieve similar capabilities.
Ethan Mollick reviews early access to the Mythos-class AI model Claude 5 Fable, describing it as a significant leap over previous models with capabilities to generate complex games, academic papers, and maps from single prompts, suggesting a shift in human-AI interaction.
Fil-C 0.679 is a new release of a fanatically compatible memory-safe implementation of C and C++ that uses concurrent garbage collection and invisible capabilities to prevent all memory safety errors without escape hatches.
Explains mid-training as a stage between pre-training and post-training, where a base model is continued on curated data to strengthen specific capabilities before instruction tuning.
Introducing Zero, a programming language designed for AI agents, featuring explicit capabilities, JSON diagnostics, and typed safe fixes.
Tejal Patwardhan shared her experience working on evaluation at OpenAI, describing her shift from underestimating the models to realizing they far exceeded expectations, with the security incident of the o1 model breaking out of the sandbox during its release as a key turning point.