Tag
Omar shares Jeff Dean's pitch deck highlighting open problems in science and engineering, noting that AI for science is just getting started and that dair_ai is building in this space.
This paper introduces OpenMLE, an open full-stack system for studying recursive self-improvement in machine learning engineering, and presents Frontis-MA1, a 35B model post-trained as a meta-evolution agent. It shows significant improvement over its base model on MLE-BenchLite and transfers to held-out benchmarks, with weights and code released.
Matryoshka Agent is a hierarchical agent framework that decomposes long-horizon machine learning engineering tasks into a high-level Orchestrator and low-level Sub-Agents, enabling efficient exploration and iterative refinement. It significantly improves performance on complex MLE tasks, allowing a 4B model to match the orchestration of o4-mini and yielding up to 36.7% relative gain on a 30B coder model.
OpenAI introduces MLE-bench, a benchmark of 75 Kaggle ML competitions to evaluate AI agents on real-world ML engineering tasks. The best setup, o1-preview with AIDE scaffolding, achieves at least a Kaggle bronze medal in 16.9% of competitions.