Tag
A paper describes using JEV as a gatekeeper for reasoning tasks, acting as a cost-effective first-pass judge that escalates to LLMs or humans when confidence is low.
The article discusses methods to evaluate the anonymous AI model Space Bunny for agent workflows, focusing on stateful testing, error recovery, and consistency to ensure reliability.
Introduces IndicBankBench, a 799-case benchmark for evaluating safety and reliability of language model assistants in Indian retail banking, with multi-stage evaluation and public release of code and data.
This study evaluates the cross-region and cross-event generalization of the Wavelet Diffusion Model for precipitation downscaling using U.S. regional data, showing that models trained on limited regions can perform competitively in unseen areas.
PFArena introduces a benchmark for evaluating language models on protein modification tasks, comparing protein language models and large language models to assess their capabilities in biological applications.
RECLAIM is a benchmark for AI agents to reproduce claims from machine learning papers, with difficulty tiers based on available code, data, and weights. It shows that agents struggle to reproduce results, with best success rates of 41% in the easiest tier.
StepCOPS is a statistical framework for selecting language-model policies that uses closed-testing and lower-tail certificates to ensure safety guarantees with high probability, improving efficiency over conservative methods.
This paper evaluates the causal link between explanations and model predictions in vision-language reasoning through generation order interventions, finding that larger models are required for rationale-first reasoning and that answer-first generation reduces format-related errors.
This paper identifies a failure mode where language models are persuaded by assertions from incentive-misaligned witnesses in CRM records, leading to incorrect decisions, and proposes a diagnostic method to analyze this issue.
A finance benchmark named DAYJOB by Surge AI evaluates AI agents on completing financial forecasting tasks, with detailed criteria for pass/fail responses focusing on errors in net sales calculations and revenue growth projections.
The article explains how to integrate the Jev AI model into an agent orchestration framework for intelligent model routing, risk interception in auto mode, and automated evaluation, referencing a LangChain video for further details.
A fireside chat at Interrupt NYC featuring the CEO of RogoAI discussing the evaluation of agentic AI at scale with LinkedIn.
LangSmith has launched Custom Apps, now generally available, allowing users to build and publish custom interfaces for their agent data within the platform to enhance workflows like annotation and experiment comparison.
Anthropic reported that using about 950 Claude agents over 21 hours identified a potential biological lead in enzyme research, but the result is preliminary and raises questions about validating AI agent workflows in science.
The author shares lessons learned from a long-running Django benchmark, highlighting fixes in the evaluation workflow stability and updated results showing Flash Next as the top performer with reasoning effort levels now properly evaluated.
FWBench introduces a benchmark for evaluating how language models select and use time-series forecasts to make cost-constrained decisions, comparing hosted and local configurations on electricity and cycle-hire datasets with efficient budget usage by GPT-6 Astra.
This paper defines a hierarchy of faithfulness criteria for knowledge base completion models and evaluates current embedding models, finding they are not logically faithful.
This paper introduces NAF-Bench to study how large language models adhere to specified negation semantics, finding that frontier models like o4-mini perform well while open-source models lag, and suggesting improvements via solver delegation or fine-tuning.
This paper evaluates small language models beyond answer accuracy in knowledge graph question answering by isolating graph navigation capabilities, showing significant differences in path fidelity and the need for broader evaluation metrics.
MolDesignBench is a new benchmark for evaluating LLM-based agents in scenario-grounded molecular design, comprising 2K instances with implicit and explicit constraints, revealing that current frontier LLMs achieve low success rates, especially in reasoning about implicit constraints and infeasibility detection.