Tag
This paper introduces BatteryLake, a governed data lakehouse that uses LLM agents for evidence-grounded metadata extraction and schema mapping, with human-in-the-loop verification, to curate heterogeneous battery aging datasets and release an open benchmark.
Prime Intellect released verifiers v1 and prime-rl 0.7.0, an RL training tool with full support for verifiers, multiple algorithms like GRPO and OPD, and performance improvements.
OpenProver is an open-source system for LLM-driven automated theorem proving using Lean 4, featuring a Planner-Worker-Verifier architecture and both autonomous and interactive modes. It enables reproducible evaluation and human-AI synergy in mathematical proof search.
A 35B-parameter MoE model with only 3B active parameters matches or surpasses 100B-class models using post-training RL, achieving significant efficiency gains.
MAD-OPD utilizes a multi-teacher debate mechanism to break through the single-teacher distillation ceiling, enabling small models to surpass large teacher models in tool invocation and code generation tasks.
Alexandr Wang stated that Muse Spark 1.1 is an industry-competitive agentic and coding model, capable of competing with GPT-5.5 and Opus 4.8 in multiple benchmarks. It is now available via the Meta Model API and Meta AI.
Amazon has announced an agentic version of Alexa with long-term memory and integration with over 1,000 apps, enhancing its capabilities as a personal AI assistant.
Muse Spark 1.1, a new agentic and coding model from Meta, achieves +5% improvement on HealthBench-Pro, outperforming all competitors except Fable and Mythos.
Meta AI released Muse Spark 1.1, a multimodal reasoning model designed for agentic tasks.
Tess-4-27B is a 27B reasoning model built on Qwen3.6-27B, post-trained on 64K-token long-context agentic traces with weight-scaled reasoning. It is designed for efficient, honest, and agentic task execution, available in open-source formats.
Meta released Muse Image, an agentic image generation model that plans, searches the web, writes code, and edits before rendering.
Meta introduced Muse Image and previewed Muse Video, an agentic image and video generation system that enables precise edits, multiple references, and integration with Instagram context, turning media generation into a full creative operating system.
Tencent releases Hy3, a 295B MoE AI model that rivals trillion-scale flagships, open-sourced under Apache 2.0 with a free API trial.
A benchmark comparing 8 local models on a classic medieval European fantasy role-playing and agentic task found that Qwen3.6-27B performed better than its size would suggest.
Anthropic released Claude Sonnet 5, its most capable Sonnet model yet, featuring improved reasoning, coding, and tool use, with reduced hallucination. Priced at $2/$10 per million tokens through August.
Anthropic released Claude Sonnet 5 with improved reasoning, tool use, and coding, but its updated tokenizer maps text to more tokens (up to 1.35×), increasing effective cost per task despite the same listed price; introductory pricing applies until August 31, 2026.
NVIDIA proposes HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution, achieving 100% benchmark completion across several hardware design suites.
Jon Udell argues for reframing 'human in the loop' as 'human agent in the loop,' where humans invite AI agents into collaborative processes rather than being subordinated to machine-driven loops.
Introduces a paper accepted at PACT 2025, proposing the ComPilot framework, which uses off-the-shelf LLMs as optimization agents to automatically optimize complex loop nests without fine-tuning, achieving a geometric mean speedup of 3.54x, surpassing the SOTA Pluto.
Meet Gemma 4 12B Agentic Fable5, a locally-run GGUF model designed for coding, terminal tasks, and agentic workflows, with 206k downloads.