Tag
GDPevo is a benchmark for evaluating agent self-evolution on real business tasks, covering CRM, ERP, finance, healthcare, legal, and data-centric workflows. The authors release an automated pipeline and find that self-evolution improves held-out accuracy by up to 16.44 percentage points, though agents remain well below an oracle ceiling.
Alibaba Cloud integrates TinyFish's Mako execution engine into its desktop agent to deliver a web-native, workflow-trained intelligence layer for handling complex multi-step tasks with lower latency and fewer stalls.
OpenAI's GPT-5.6 update improves artifact quality across presentations, documents, and spreadsheets, enhancing template compatibility and enterprise workflow integration.
A tutorial from Google on building long-running AI agents that can pause for days, survive restarts, and resume without losing context using the Agent Development Kit (ADK), with code and step-by-step guidance for enterprise workflows like new hire onboarding.
Anchor is a task-generation pipeline that addresses artifact drift in AI agent benchmarks by jointly producing instructions, environments, solutions, and verifiers from a single constraint optimization specification, yielding consistent and auditable evaluation tasks for enterprise workflows. The paper introduces ERP-Bench, a benchmark of 300 long-horizon tasks in a production ERP system, showing that frontier models satisfy explicit constraints in 26.1% of trials but reach optimal solutions in only 17.4%.