Tag
IBM Research explains why model routing in agentic systems is more complex than a simple classification problem, highlighting how caching and hidden factors like actual workload cost and task difficulty estimation make routing a systems optimization challenge.
IBM Research introduces ScarfBench, an open benchmark for evaluating AI agents on cross-framework Java migration tasks, focusing on Spring, Jakarta EE, and Quarkus. The benchmark assesses whether migrated applications build, deploy, and preserve behavior, unlike traditional code generation benchmarks.
IBM Research explores how agent logic—software primitives like knowledge graphs and program analysis—can guide LLM-based agents to efficiently handle complex enterprise workflows, reducing hallucinations and costs while improving outcomes.
IBM Research launches the Open Agent Leaderboard, an open benchmark and evaluation framework for comparing full AI agent systems based on quality and cost, aiming to measure generality across diverse tasks.