enterprise-workflows

Tag

Cards List
#enterprise-workflows

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

Hugging Face Daily Papers · 2026-08-04 Cached

GDPevo is a benchmark for evaluating agent self-evolution on real business tasks, covering CRM, ERP, finance, healthcare, legal, and data-centric workflows. The authors release an automated pipeline and find that self-evolution improves held-out accuracy by up to 16.44 percentage points, though agents remain well below an oracle ceiling.

0 favorites 0 likes
#enterprise-workflows

@alibaba_cloud: We’re excited to integrate TinyFish’s Mako execution engine into Alibaba Cloud’s desktop agent—bringing a web-native, w…

X AI KOLs Timeline · 2026-07-24 Cached

Alibaba Cloud integrates TinyFish's Mako execution engine into its desktop agent to deliver a web-native, workflow-trained intelligence layer for handling complex multi-step tasks with lower latency and fewer stalls.

0 favorites 0 likes
#enterprise-workflows

@OpenAI: GPT‑5.6 improves artifact quality across presentations, documents, and spreadsheets, and works better with your templat…

X AI KOLs · 2026-07-09 Cached

OpenAI's GPT-5.6 update improves artifact quality across presentations, documents, and spreadsheets, enhancing template compatibility and enterprise workflow integration.

0 favorites 0 likes
#enterprise-workflows

@googledevs: Most agent tutorials stop at stateless agents. Real workflows run for weeks. Build long-running AI agents that pause fo…

X AI KOLs Following · 2026-05-29 Cached

A tutorial from Google on building long-running AI agents that can pause for days, survive restarts, and resume without losing context using the Agent Development Kit (ADK), with code and step-by-step guidance for enterprise workflows like new hire onboarding.

0 favorites 0 likes
#enterprise-workflows

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation

arXiv cs.AI · 2026-05-27 Cached

Anchor is a task-generation pipeline that addresses artifact drift in AI agent benchmarks by jointly producing instructions, environments, solutions, and verifiers from a single constraint optimization specification, yielding consistent and auditable evaluation tasks for enterprise workflows. The paper introduces ERP-Bench, a benchmark of 300 long-horizon tasks in a production ERP system, showing that frontier models satisfy explicit constraints in 26.1% of trials but reach optimal solutions in only 17.4%.

0 favorites 0 likes
← Back to home

Submit Feedback