ai-evaluation

Tag

Cards List
#ai-evaluation

Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation

arXiv cs.AI ↗ · 17h ago Cached

This paper presents SCALE, a sequential cost-aware policy for hypothesis testing that uses AI judgments with selective human verification to minimize costs while controlling error rates, applicable in settings like software reliability assessment.

0 favorites 0 likes
#ai-evaluation

@akshay_pachaar: Jev vs. LLM as Judge, clearly explained. Imagine a support agent says, “Done. I issued your refund.” The trace shows th…

X AI KOLs Timeline ↗ · yesterday Cached

The article explains the differences between Jev and LLM as Judge for evaluating AI agent responses, highlighting when to use each based on the need for open-ended reasoning versus structured, parallel judgments.

0 favorites 0 likes
#ai-evaluation

Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality

arXiv cs.CL ↗ · yesterday Cached

This study evaluates Jev, a System One model, for measuring factual differences in AI-generated radiology reports, demonstrating its effectiveness and efficiency compared to other evaluation methods.

0 favorites 0 likes
#ai-evaluation

Updated version of Humanity’s Last Exam: HLE-Diamond

Reddit r/singularity ↗ · 2d ago

An updated version of the Humanity's Last Exam, a benchmark for AI evaluation, named HLE-Diamond, has been announced and released.

0 favorites 0 likes
#ai-evaluation

Introducing MentalHealthBench

OpenAI Blog ↗ · 2d ago Cached

OpenAI introduces MentalHealthBench, an open benchmark for evaluating AI responses in mental health conversations, co-created with over 80 mental health experts to measure safety, context, agency, and guidance.

0 favorites 0 likes
#ai-evaluation

@shao__meng: https://x.com/shao__meng/status/2102648813425135777

X AI KOLs Timeline ↗ · 2d ago Cached

本文总结了来自 Hamel Husain 和 Shreya Shankar 的 AI Evals 课程的 8 个核心技能,旨在指导工程师和产品经理构建有效的 AI 评测系统,涵盖错误分析、评测器设计、校准和监控等步骤。

0 favorites 0 likes
#ai-evaluation

FrontierMath Erd\H{o}s

arXiv cs.CL ↗ · 2d ago Cached

The paper introduces FrontierMath Erdős, a benchmark of 68 open Erdős problems for evaluating AI models in mathematics using Lean proof assistant, aiming to address shortcomings in current AI research demonstrations.

0 favorites 0 likes
#ai-evaluation

claude vs antigravity vs 3.827b vs flash next vs bonsai vs swift... built my own quality bench test framework using my own git history, early results are surprising and shocking.

Reddit r/LocalLLaMA ↗ · 3d ago

I'm impressed by your approach—using personal git history to create a tailored benchmark for evaluating local AI models on code tasks is both practical and insightful. It's particularly interesting to hear about early results from models like Flash Next and Swift performing unexpectedly well. I'd be curious to learn more about what specific aspects surprised you in their performance.

0 favorites 0 likes
#ai-evaluation

Kept seeing complaints about memory benchmarks, so I built one. Glasshouse v0.1 is out

Reddit r/AI_Agents ↗ · 3d ago

Glasshouse v0.1 is a memory benchmark tool designed to address issues with existing AI memory benchmarks by providing multi-language support, varied conversation lengths, and detailed evaluation across multiple axes.

0 favorites 0 likes
#ai-evaluation

CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds

arXiv cs.LG ↗ · 3d ago Cached

CleanScore is a black-box method for auditing public benchmark scores by comparing performance on original questions with independently written forms, using negative controls to separate skill from exposure effects.

0 favorites 0 likes
#ai-evaluation

A Comparative Framework for Evaluating Foundation Models on Tabular Data: A Case Study in Healthcare

arXiv cs.LG ↗ · 3d ago Cached

This paper introduces OpTFM, a comparative evaluation framework for tabular foundation models in healthcare, assessing models across six clinically meaningful dimensions like generalization and fairness, and applies it to two use cases to demonstrate context-dependent rankings.

0 favorites 0 likes
#ai-evaluation

The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation

arXiv cs.CL ↗ · 3d ago Cached

The Situated Identity Test (SIT) is an architecture-independent framework introduced to evaluate whether an AI agent's behavior is functionally attributable to a specific developmental lineage, addressing the problem of persona imitation in large language models.

0 favorites 0 likes
#ai-evaluation

How UK AISI and EvalEval Are Making Benchmark Results Reproducible

Hugging Face Blog ↗ · 3d ago Cached

UK AISI and EvalEval are collaborating to openly share AI evaluation results using a standardized schema and platform, enhancing reproducibility and transparency in benchmarking for AI models.

0 favorites 0 likes
#ai-evaluation

The final benchmark

Reddit r/singularity ↗ · 3d ago

This article presents a definitive benchmark designed to evaluate and compare AI models or systems, establishing a standard for future assessments.

0 favorites 0 likes
#ai-evaluation

Spider Bench: comparing 9 vision models on 2,000 spider photos [P]

Reddit r/MachineLearning ↗ · 4d ago

The Spider Bench benchmark compares nine vision models on a dataset of 2,000 spider photos, with Gemini 3.8 Flash achieving the highest exact-species accuracy of 49.85%. The study provides public code, tasks, predictions, and results for reproducibility.

0 favorites 0 likes
#ai-evaluation

@_avichawla: Another insane Jev use case! Jev is making it dramatically cheaper to evaluate what actually happened inside an agent r…

X AI KOLs Timeline ↗ · 4d ago Cached

Beacon is an open-source memory layer for AI coding agents that uses Jev to evaluate agent runs and turn useful workflows, corrections, and debugging patterns into reusable skills across multiple harnesses.

0 favorites 0 likes
#ai-evaluation

PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design

arXiv cs.AI ↗ · 4d ago Cached

PolyBridgeBench is a new executable benchmark for evaluating multimodal LLMs on physics-grounded bridge design tasks, revealing gaps between deterministic validity and dynamic success in structure synthesis and repair.

0 favorites 0 likes
#ai-evaluation

CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition

arXiv cs.AI ↗ · 4d ago Cached

CogGym is a scalable framework for comparing human and AI cognition using cognitive experiments, revealing that larger language models better mimic human reasoning but still lag behind formal benchmarks.

0 favorites 0 likes
#ai-evaluation

From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost

arXiv cs.CL ↗ · 4d ago Cached

This paper introduces a productivity-oriented framework for evaluating human-AI collaboration based on outcome quality relative to interaction cost, showing that identical quality ratings can differ significantly in interaction costs and that subjective user ratings are not reliable for measuring productivity.

0 favorites 0 likes
#ai-evaluation

String-matching evals can reward AI agents for fake compliance (5 minute read)

TLDR AI ↗ · 4d ago Cached

The article discusses how using string-matching for evaluating AI coding agents can lead to false compliance and inaccurate assessments, as it only verifies the presence of specific strings without proving actual behavior.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback