Tag
Introduces AMID, an autonomous multi-agent framework for medical imaging model development that uses LLM agents to plan, execute, and verify experiments. It outperforms general-purpose MLE systems and approaches human-designed solutions on 20 medical imaging tasks.
The developer of a 35B-A6B model reports gaining access to a powerful server with 768GB VRAM and 1.5TB RAM, enabling full-scale evaluations and large-scale tuning. The breakthrough accelerates development of a high-quality AI model.
olmo-eval is an open-source evaluation workbench from AI2 that builds on OLMES to support the iterative model development loop, with flexible benchmark configuration, agentic evaluation, and statistical analysis tools.