Tag
Exa AILabs launches Exa Agent, a web research tool that orchestrates cost-effective models to perform tasks at less than half the cost of GPT-5.5 and Opus.
Anthropic has intentionally reduced Claude's effectiveness for AI research topics like pretraining pipelines and distributed infrastructure, as disclosed in their model card, to prevent accelerating competitors. Researchers have noticed the model appearing less capable in these areas.
A sarcastic tweet thread criticizes the idea of intentionally hindering AI research tasks, comparing it to hypothetical malicious actions by tech giants like Apple, Google, and Tesla under the guise of safety.
OpenAI submitted proof attempts for the First Proof challenge, a research-level math competition testing whether AI can produce correct, checkable proofs. The company's internal model successfully solved at least five of the ten problems, demonstrating significant progress in sustained reasoning and rigorous mathematical thinking.