Tag
This paper identifies a scalability bottleneck in RL-trained automatic research agents—environment execution dominates training cost—and proposes World Model RL (WMRL) with online debiasing and inverse-variance denoising to replace real execution, achieving 3–4x training speedups and better generalization.
Google is reportedly testing a dedicated Agent management UI in AI Studio, allowing users to browse, create, and configure managed agents per Google Cloud project. The unreleased interface includes a workbench tab, artifact management, and an editor, bringing graphical management to agent definitions.
Google releases Gemini 3.7 Flash, a new AI model focused on coding and agentic workflows, with a 50% introductory price cut on API tokens until end of 2026. The model improves planning, instruction following, and error recovery to reduce retries and manual oversight.
Google introduces Gemini 3.7 Flash, its most intelligent workhorse model for coding and agents, with significant improvements in software engineering, web development, and knowledge work, offered at half the introductory price of Gemini 3.6 Flash.
Google introduces Gemini 3.7 Flash, an update to its workhorse model offering significant gains in coding, knowledge work, and web development at half the original cost of 3.6 Flash.
DeepSeek launches V4-Pro and V4-Flash with flexible reasoning effort, native OpenAI Responses API support, and optimized agent workflows for Codex, available via API and app/web.
A tweet lists 15 AI engineering projects to build for the 2026 hiring season, from terminal coding agents and MCP servers to multimodal document agents and agent-to-agent commerce, each demonstrating key skills for hiring managers.
This paper introduces SHAPER, a self-evolving framework for embodied agents that keeps model parameters frozen and improves performance by evolving reusable skills and context-code harnesses through target-environment rollouts. Evaluated on VLABench and ESI-Bench, it proposes a practical alternative to fine-tuning when training is expensive or unavailable.
This paper introduces CompInt, an evaluation suite for measuring how well context compaction preserves user-issued session constraints in LLM systems. It finds current compactors retain only 17% of constraints on average and proposes an SC-aware extractor that achieves over 90% retention without modifying the compactor or LLM.
PlayWorld is a benchmark for evaluating interactive video world models using multi-modal agent players pursuing long-horizon objectives. It assesses geometry consistency, interaction fidelity, and state evolution, revealing that current models struggle with spatial consistency and persistent state evolution.
Kent C. Dodds shares a Kody community package that provides an SDK for the Devin v3 API, enabling automated management of Devin sessions, knowledge, playbooks, and schedules.
A developer shares an approach where teams use Claude Code for writing code and Codex for verification, focusing on detailed specs and overnight AI agent runs.
Introduces CliniCARE-Bench, a deployment-oriented benchmark for evaluating AI agents on clinical audit tasks over longitudinal EHR data, with 25 clinician-validated scenarios and 750 patient cases. It assesses verdict accuracy, evidence grounding, policy adherence, and calibrated abstention, finding that raw accuracy overstates investigation quality.
This paper introduces DSAgentBench, the first benchmark for evaluating autonomous agents on complete, multi-tool data-science workflows in real computer environments. Results show that even the strongest agent (Claude-4.6-Sonnet) achieves only 56.70% task success, while open-source agents remain below 1%, revealing a substantial capability gap.
SkillZip is a new method for compressing the accumulated skills of self-evolving agents without evaluation rollouts, by finding a minimal faithful structural explanation that reuses repeated rules and procedures while preserving rare exceptions.
The author critiques legacy workflow tools like Zapier and n8n as unsuitable for the AI era, and pitches a new AI-native automation platform with features like talk-to-build canvas, auto-healing pipelines, dynamic agent swarms, WhatsApp human-in-the-loop, and white-label publishing.
A small startup shares its production multi-agent architecture, where an orchestrator routes tasks to specialized worker agents that monitor ads, product reviews, churn, and SEO, all coordinated via Slack channels.
An opinion piece arguing that humanising LLM outputs via prompt instructions is the wrong abstraction—agents should exchange high-fidelity data and only compress into human-friendly prose at the final boundary.
The author, after trying Qwen's latest app, points out that for Agents to reach mainstream users, three bottlenecks need to be resolved: cost, response speed, and the way humans interact with AI.
LangChain is promoting Interrupt, its agent-focused conference with events in NYC and London in 2026, where builders and engineers can connect.