ibm-research

Tag

Cards List
#ibm-research

Model Routing Is Simple. Until It Isn’t.

Hugging Face Blog · 2026-07-15 Cached

IBM Research explains why model routing in agentic systems is more complex than a simple classification problem, highlighting how caching and hidden factors like actual workload cost and task difficulty estimation make routing a systems optimization challenge.

0 favorites 0 likes
#ibm-research

ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

Hugging Face Blog · 2026-06-30 Cached

IBM Research introduces ScarfBench, an open benchmark for evaluating AI agents on cross-framework Java migration tasks, focusing on Spring, Jakarta EE, and Quarkus. The benchmark assesses whether migrated applications build, deploy, and preserve behavior, unlike traditional code generation benchmarks.

0 favorites 0 likes
#ibm-research

Beyond LLMs: Why Scalable Enterprise AI Adoption Depends on Agent Logic

Hugging Face Blog · 2026-06-01 Cached

IBM Research explores how agent logic—software primitives like knowledge graphs and program analysis—can guide LLM-based agents to efficiently handle complex enterprise workflows, reducing hallucinations and costs while improving outcomes.

0 favorites 0 likes
#ibm-research

The Open Agent Leaderboard

Hugging Face Blog · 2026-05-18 Cached

IBM Research launches the Open Agent Leaderboard, an open benchmark and evaluation framework for comparing full AI agent systems based on quality and cost, aiming to measure generality across diverse tasks.

1 favorites 1 likes
← Back to home

Submit Feedback