Tag
LLM-as-a-Verifier introduces a probabilistic verification framework that computes continuous scores from LLM logits, scaling across granularity, repeated evaluation, and criteria decomposition. It achieves state-of-the-art results on multiple agentic benchmarks and provides dense feedback for RL.
Proposes a new RANSAC scoring function that marginalizes the inlier scale analytically, removing the need for user-supplied parameters. The method achieves state-of-the-art accuracy on a benchmark of nearly 70,000 image pairs.
This article from Chrome Developers details the agentic browsing category in Lighthouse, which evaluates how ready websites are to interact with AI agents. The category focuses on data collection and providing actionable signals rather than a traditional numerical score, and includes audits for WebMCP, accessibility, and content stability.
This paper introduces BioConCal, a supervised scorer that uses inference-time panel and candidate features to rank biomedical entity candidates surfaced by LLM panels, significantly improving over raw agreement for curator triage.
A developer shares practical lessons from building an AI lead qualification agent, highlighting that the hardest issues were not AI-related but involved vague answers, routing logic, Slack noise, CRM structure, and handling low-fit leads.