agents

Tag

Cards List
#agents

Scaling Automatic Research Agents via World Models

arXiv cs.LG · 15h ago Cached

This paper identifies a scalability bottleneck in RL-trained automatic research agents—environment execution dominates training cost—and proposes World Model RL (WMRL) with online debiasing and inverse-variance denoising to replace real execution, achieving 3–4x training speedups and better generalization.

0 favorites 0 likes
#agents

Google tests Agent management UI on AI Studio (2 minute read)

TLDR AI · 19h ago Cached

Google is reportedly testing a dedicated Agent management UI in AI Studio, allowing users to browse, create, and configure managed agents per Google Cloud project. The unreleased interface includes a workbench tab, artifact management, and an editor, bringing graphical management to agent definitions.

0 favorites 0 likes
#agents

Google's Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut (13 minute read)

TLDR AI · 19h ago Cached

Google releases Gemini 3.7 Flash, a new AI model focused on coding and agentic workflows, with a 50% introductory price cut on API tokens until end of 2026. The model improves planning, instruction following, and error recovery to reduce retries and manual oversight.

0 favorites 0 likes
#agents

Gemini 3.7 Flash

Hacker News Top · yesterday Cached

Google introduces Gemini 3.7 Flash, its most intelligent workhorse model for coding and agents, with significant improvements in software engineering, web development, and knowledge work, offered at half the introductory price of Gemini 3.6 Flash.

0 favorites 0 likes
#agents

@sundarpichai: Our Flash models are workhorses that offer performance at a great price. So, we're shipping updates fast to get them in…

X AI KOLs Timeline · yesterday Cached

Google introduces Gemini 3.7 Flash, an update to its workhorse model offering significant gains in coding, knowledge work, and web development at half the original cost of 3.6 Flash.

0 favorites 0 likes
#agents

DeepSeek API Pricing Update

Hacker News Top · yesterday Cached

DeepSeek launches V4-Pro and V4-Flash with flexible reasoning effort, native OpenAI Responses API support, and optimized agent workflows for Codex, available via API and app/web.

0 favorites 0 likes
#agents

@suraj_sharma14: Top 15 AI Engineer projects for the 2026 hiring season. If you can build these. You're hired. Project 1: Terminal Codin…

X AI KOLs Timeline · yesterday Cached

A tweet lists 15 AI engineering projects to build for the 2026 hiring season, from terminal coding agents and MCP servers to multimodal document agents and agent-to-agent commerce, each demonstrating key skills for hiring managers.

0 favorites 0 likes
#agents

Self-Evolving Embodied Agents via Skill-Harness Evolution

arXiv cs.CL · yesterday Cached

This paper introduces SHAPER, a self-evolving framework for embodied agents that keeps model parameters frozen and improves performance by evolving reusable skills and context-code harnesses through target-environment rollouts. Evaluated on VLABench and ESI-Bench, it proposes a practical alternative to fine-tuning when training is expensive or unavailable.

0 favorites 0 likes
#agents

Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction

arXiv cs.CL · yesterday Cached

This paper introduces CompInt, an evaluation suite for measuring how well context compaction preserves user-issued session constraints in LLM systems. It finds current compactors retain only 17% of constraints on average and proposes an SC-aware extractor that achieves over 90% retention without modifying the compactor or LLM.

0 favorites 0 likes
#agents

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

Hugging Face Daily Papers · yesterday Cached

PlayWorld is a benchmark for evaluating interactive video world models using multi-modal agent players pursuing long-horizon objectives. It assesses geometry consistency, interaction fidelity, and state evolution, revealing that current models struggle with spatial consistency and persistent state evolution.

0 favorites 0 likes
#agents

@kentcdodds: Playing with @DevinAI so naturally I had it create a Devin package on @kodykoala https://heykody.app/@kentcdodds/devin…

X AI KOLs Following · 3d ago Cached

Kent C. Dodds shares a Kody community package that provides an SDK for the Devin v3 API, enabling automated management of Devin sessions, knowledge, playbooks, and schedules.

0 favorites 0 likes
#agents

@svpino: Claude Code to write your code and Codex to verify it. I met with a team that's been doing this for a few weeks now. I …

X AI KOLs Following · 3d ago Cached

A developer shares an approach where teams use Claude Code for writing code and Codex for verification, focusing on detailed specs and overnight AI agent runs.

0 favorites 0 likes
#agents

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

arXiv cs.AI · 3d ago Cached

Introduces CliniCARE-Bench, a deployment-oriented benchmark for evaluating AI agents on clinical audit tasks over longitudinal EHR data, with 25 clinician-validated scenarios and 750 patient cases. It assesses verdict accuracy, evidence grounding, policy adherence, and calibrated abstention, finding that raw accuracy overstates investigation quality.

0 favorites 0 likes
#agents

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

Hugging Face Daily Papers · 3d ago Cached

This paper introduces DSAgentBench, the first benchmark for evaluating autonomous agents on complete, multi-tool data-science workflows in real computer environments. Results show that even the strongest agent (Claude-4.6-Sonnet) achieves only 56.70% task success, while open-source agents remain below 1%, revealing a substantial capability gap.

0 favorites 0 likes
#agents

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

Hugging Face Daily Papers · 3d ago Cached

SkillZip is a new method for compressing the accumulated skills of self-evolving agents without evaluation rollouts, by finding a minimal faithful structural explanation that reuses repeated rules and procedures while preserving rare exceptions.

0 favorites 0 likes
#agents

Zapier and n8n are fundamentally broken for the AI Era. So, I’m building an "AI-Native" alternative. Need your brutal feedback! 🚀

Reddit r/AI_Agents · 3d ago

The author critiques legacy workflow tools like Zapier and n8n as unsuitable for the AI era, and pitches a new AI-native automation platform with features like talk-to-build canvas, auto-healing pipelines, dynamic agent swarms, WhatsApp human-in-the-loop, and white-label publishing.

0 favorites 0 likes
#agents

How our small startup runs a multi-agent setup in production (orchestrator + workers, all piloted from slack)

Reddit r/AI_Agents · 3d ago

A small startup shares its production multi-agent architecture, where an orchestrator routes tasks to specialized worker agents that monitor ads, product reviews, churn, and SEO, all coordinated via Slack channels.

0 favorites 0 likes
#agents

Humanising LLM Outputs Is Dumb

Hacker News Top · 4d ago Cached

An opinion piece arguing that humanising LLM outputs via prompt instructions is the wrong abstraction—agents should exchange high-fidelity data and only compress into human-friendly prose at the final boundary.

0 favorites 0 likes
#agents

@changgaowei: After trying out the latest version of the Qwen app, which can connect to various Agents through Qwen, I found that three bottlenecks still need to be solved before Agents can reach the general public. 1. Cost — it must be cheap enough. 2. Response speed — the experience must be smooth. Don't make people wait. 3. Human-AI interaction — the native...

X AI KOLs Following · 4d ago

The author, after trying Qwen's latest app, points out that for Agents to reach mainstream users, three bottlenecks need to be resolved: cost, response speed, and the way humans interact with AI.

0 favorites 0 likes
#agents

@LangChain: Interrupt is the place to connect with agent builders. Whether-in between sessions, at our workshops, or during our rec…

X AI KOLs Timeline · 4d ago Cached

LangChain is promoting Interrupt, its agent-focused conference with events in NYC and London in 2026, where builders and engineers can connect.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback