@AIatMeta: As a test of our progress to advance the frontier of AI research, in June we entered the next generation of our autonom…
Summary
Meta's autonomous AI research system AIRA₃ placed 8th out of approximately 4,000 teams to win gold in a NVIDIA Kaggle competition to fine-tune a 30B Nemotron model, outperforming human competitors with access to the same tools.
Similar Articles
Inside Meta's attempts to play catch-up with AI
A detailed report on Meta's efforts to catch up in AI, including the hiring of Alexandr Wang and the release of the Muse Spark model, with mixed opinions on progress.
@METR_Evals: Could an AI company lose control of its own agents? To find out, Anthropic, Google, Meta, and OpenAI let us (1) test th…
METR published its first Frontier Risk Report, assessing the risk of AI companies losing control of their own agents. The report involved testing the best internal models from Anthropic, Google, Meta, and OpenAI with chain-of-thought access and reviewing non-public information about capabilities and alignment.
@dair_ai: NEW paper from Meta: Agentic Discovery of Neural Architectures. This is a hot new area of research! Keep an eye on it.
Meta's new paper presents an agentic system that autonomously discovers neural architectures outperforming Llama 3.2 at 350M, 1B, and 3B scales within a 24-hour compute budget.
@rohanpaul_ai: New Meta, Stanford, Google and many other top labs paper proposes AutoResearchClaw. Shows that automated research impro…
A new paper from Meta, Stanford, and Google introduces AutoResearchClaw, which improves automated research by integrating failure recovery, debate, and selective human input. It outperforms AI Scientist v2 by 54.7% on ARC-Bench and reveals that autonomy is enhanced when constrained by process rather than given unlimited freedom.
NVIDIA just announced the release of Nemotron 3 Ultra (2 minute read)
Anthropic released Claude Opus 4.5, its most intelligent model, scoring 70 on the Artificial Analysis Intelligence Index and ranking second only to Gemini 3 Pro. It achieves significant gains in coding and agentic tasks while reducing per-token pricing and maintaining strong safety performance.