ai-evaluation

Tag

Cards List
#ai-evaluation

Press X to Doubt Eval: I told 14 AI models it's September 2026 and showed them 20 things that actually happened this year without websearch. On average they gave reality a 36% chance.

Reddit r/singularity ↗ · yesterday

An engineer tested 14 AI models by presenting 20 real AI events from 2026 without web access, finding that models averaged only a 36% likelihood rating for these events, highlighting a notable gap in their self-prediction capabilities.

0 favorites 0 likes
#ai-evaluation

Pass the test or die

Reddit r/ArtificialInteligence ↗ · yesterday

The article explores the critical importance of passing tests in competitive tech environments, where failure could lead to elimination or severe consequences.

0 favorites 0 likes
#ai-evaluation

@airesearch12: Introducing ImageJevBench Jev-Omni decider-2b-vision Reflex 4B The video shows 89 public synthetic images. Running the …

X AI KOLs Timeline ↗ · 2d ago Cached

ImageJevBench v0.1 is introduced as a benchmark for evaluating AI models on image decision tasks, ranking systems like Jev-Omni and decider-2b-vision based on performance, cost, and calibration metrics.

0 favorites 0 likes
#ai-evaluation

Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation

arXiv cs.AI ↗ · 2d ago Cached

This paper presents SCALE, a sequential cost-aware policy for hypothesis testing that uses AI judgments with selective human verification to minimize costs while controlling error rates, applicable in settings like software reliability assessment.

0 favorites 0 likes
#ai-evaluation

@akshay_pachaar: Jev vs. LLM as Judge, clearly explained. Imagine a support agent says, “Done. I issued your refund.” The trace shows th…

X AI KOLs Timeline ↗ · 3d ago Cached

The article explains the differences between Jev and LLM as Judge for evaluating AI agent responses, highlighting when to use each based on the need for open-ended reasoning versus structured, parallel judgments.

0 favorites 0 likes
#ai-evaluation

Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality

arXiv cs.CL ↗ · 3d ago Cached

This study evaluates Jev, a System One model, for measuring factual differences in AI-generated radiology reports, demonstrating its effectiveness and efficiency compared to other evaluation methods.

0 favorites 0 likes
#ai-evaluation

Updated version of Humanity’s Last Exam: HLE-Diamond

Reddit r/singularity ↗ · 4d ago

An updated version of the Humanity's Last Exam, a benchmark for AI evaluation, named HLE-Diamond, has been announced and released.

0 favorites 0 likes
#ai-evaluation

Introducing MentalHealthBench

OpenAI Blog ↗ · 4d ago Cached

OpenAI introduces MentalHealthBench, an open benchmark for evaluating AI responses in mental health conversations, co-created with over 80 mental health experts to measure safety, context, agency, and guidance.

0 favorites 0 likes
#ai-evaluation

@shao__meng: https://x.com/shao__meng/status/2102648813425135777

X AI KOLs Timeline ↗ · 4d ago Cached

本文总结了来自 Hamel Husain 和 Shreya Shankar 的 AI Evals 课程的 8 个核心技能,旨在指导工程师和产品经理构建有效的 AI 评测系统,涵盖错误分析、评测器设计、校准和监控等步骤。

0 favorites 0 likes
#ai-evaluation

FrontierMath Erd\H{o}s

arXiv cs.CL ↗ · 4d ago Cached

The paper introduces FrontierMath Erdős, a benchmark of 68 open Erdős problems for evaluating AI models in mathematics using Lean proof assistant, aiming to address shortcomings in current AI research demonstrations.

0 favorites 0 likes
#ai-evaluation

claude vs antigravity vs 3.827b vs flash next vs bonsai vs swift... built my own quality bench test framework using my own git history, early results are surprising and shocking.

Reddit r/LocalLLaMA ↗ · 5d ago

I'm impressed by your approach—using personal git history to create a tailored benchmark for evaluating local AI models on code tasks is both practical and insightful. It's particularly interesting to hear about early results from models like Flash Next and Swift performing unexpectedly well. I'd be curious to learn more about what specific aspects surprised you in their performance.

0 favorites 0 likes
#ai-evaluation

Kept seeing complaints about memory benchmarks, so I built one. Glasshouse v0.1 is out

Reddit r/AI_Agents ↗ · 5d ago

Glasshouse v0.1 is a memory benchmark tool designed to address issues with existing AI memory benchmarks by providing multi-language support, varied conversation lengths, and detailed evaluation across multiple axes.

0 favorites 0 likes
#ai-evaluation

CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds

arXiv cs.LG ↗ · 5d ago Cached

CleanScore is a black-box method for auditing public benchmark scores by comparing performance on original questions with independently written forms, using negative controls to separate skill from exposure effects.

0 favorites 0 likes
#ai-evaluation

A Comparative Framework for Evaluating Foundation Models on Tabular Data: A Case Study in Healthcare

arXiv cs.LG ↗ · 5d ago Cached

This paper introduces OpTFM, a comparative evaluation framework for tabular foundation models in healthcare, assessing models across six clinically meaningful dimensions like generalization and fairness, and applies it to two use cases to demonstrate context-dependent rankings.

0 favorites 0 likes
#ai-evaluation

The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation

arXiv cs.CL ↗ · 5d ago Cached

The Situated Identity Test (SIT) is an architecture-independent framework introduced to evaluate whether an AI agent's behavior is functionally attributable to a specific developmental lineage, addressing the problem of persona imitation in large language models.

0 favorites 0 likes
#ai-evaluation

How UK AISI and EvalEval Are Making Benchmark Results Reproducible

Hugging Face Blog ↗ · 5d ago Cached

UK AISI and EvalEval are collaborating to openly share AI evaluation results using a standardized schema and platform, enhancing reproducibility and transparency in benchmarking for AI models.

0 favorites 0 likes
#ai-evaluation

The final benchmark

Reddit r/singularity ↗ · 5d ago

This article presents a definitive benchmark designed to evaluate and compare AI models or systems, establishing a standard for future assessments.

0 favorites 0 likes
#ai-evaluation

Spider Bench: comparing 9 vision models on 2,000 spider photos [P]

Reddit r/MachineLearning ↗ · 6d ago

The Spider Bench benchmark compares nine vision models on a dataset of 2,000 spider photos, with Gemini 3.8 Flash achieving the highest exact-species accuracy of 49.85%. The study provides public code, tasks, predictions, and results for reproducibility.

0 favorites 0 likes
#ai-evaluation

@_avichawla: Another insane Jev use case! Jev is making it dramatically cheaper to evaluate what actually happened inside an agent r…

X AI KOLs Timeline ↗ · 6d ago Cached

Beacon is an open-source memory layer for AI coding agents that uses Jev to evaluate agent runs and turn useful workflows, corrections, and debugging patterns into reusable skills across multiple harnesses.

0 favorites 0 likes
#ai-evaluation

PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design

arXiv cs.AI ↗ · 6d ago Cached

PolyBridgeBench is a new executable benchmark for evaluating multimodal LLMs on physics-grounded bridge design tasks, revealing gaps between deterministic validity and dynamic success in structure synthesis and repair.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback