evaluation

Tag

Cards List
#evaluation

A hierarchy of faithfulness criteria for knowledge base completion

arXiv cs.AI ↗ · 3d ago Cached

This paper defines a hierarchy of faithfulness criteria for knowledge base completion models and evaluates current embedding models, finding they are not logically faithful.

0 favorites 0 likes
#evaluation

Not What You Meant: Can LLMs Follow a Specified Negation Semantics?

arXiv cs.AI ↗ · 3d ago Cached

This paper introduces NAF-Bench to study how large language models adhere to specified negation semantics, finding that frontier models like o4-mini perform well while open-source models lag, and suggesting improvements via solver delegation or fine-tuning.

0 favorites 0 likes
#evaluation

The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA

arXiv cs.CL ↗ · 3d ago Cached

This paper evaluates small language models beyond answer accuracy in knowledge graph question answering by isolating graph navigation capabilities, showing significant differences in path fidelity and the need for broader evaluation metrics.

0 favorites 0 likes
#evaluation

MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

arXiv cs.AI ↗ · 3d ago Cached

MolDesignBench is a new benchmark for evaluating LLM-based agents in scenario-grounded molecular design, comprising 2K instances with implicit and explicit constraints, revealing that current frontier LLMs achieve low success rates, especially in reasoning about implicit constraints and infeasibility detection.

0 favorites 0 likes
#evaluation

PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models

arXiv cs.CL ↗ · 3d ago Cached

PRISM-VLM is a multi-axis discriminative benchmark that evaluates compact vision-language models along seven axes, including task quality and behavioral robustness, to provide more reliable separation and insights compared to single-axis benchmarks. It aims to release an open pipeline for the community to improve AI evaluation methods.

0 favorites 0 likes
#evaluation

PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks

arXiv cs.AI ↗ · 3d ago Cached

PotARCin is a multi-dimensional benchmark that extends ARC to evaluate abstract reasoning skills across five dimensions, showing that standard evaluations may not fully capture AI models' true capabilities.

0 favorites 0 likes
#evaluation

Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models

arXiv cs.CL ↗ · 3d ago Cached

The paper introduces TACT, a benchmark for evaluating turn-taking in full-duplex spoken dialogue models that uses intent-conditioned continuous scoring to replace binary rules, demonstrating better alignment with human judgments.

0 favorites 0 likes
#evaluation

Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure

arXiv cs.CL ↗ · 3d ago Cached

This paper proposes a method to estimate the causal effect of benchmark exposure on AI model performance, moving beyond traditional overlap techniques for more robust evaluation.

0 favorites 0 likes
#evaluation

When Learned Context Planning Fails to Beat Strong Retrieval: A Controlled Study of Planning, Routing, and Reranking for Long-Context QA

arXiv cs.CL ↗ · 3d ago Cached

This paper presents a controlled study on learned context planning for long-context multiple-choice QA, demonstrating that it does not robustly outperform strong retrieval methods like BM25 and hybrid retrieval under various experimental settings.

0 favorites 0 likes
#evaluation

Do y'all remember the snake model evaluation test?

Reddit r/LocalLLaMA ↗ · 3d ago

The article reflects on the rapid progress in AI models from three years ago, comparing early models like Bard that struggled with coding tests to current models like Qwen 27B and Opus 5.5, and speculates on future advancements.

0 favorites 0 likes
#evaluation

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Hugging Face Daily Papers ↗ · 3d ago Cached

The paper introduces ExplorationBench, a benchmark for evaluating AI systems' exploration abilities in verifiable alien worlds, addressing challenges in scientific discovery by providing executable rules and preventing recall from pre-training data.

0 favorites 0 likes
#evaluation

The AI you test in the afternoon may not be the AI you test at night, even with the same name

Reddit r/artificial ↗ · 4d ago

Testing Claude Opus 5 twice in one day revealed significant response differences, likely due to background settings rather than model changes, highlighting the need to document all settings when comparing AI tools.

0 favorites 0 likes
#evaluation

@OpenAI: Most mental health benchmarks focus on emergency situations. MentalHealthBench is designed to cover the full spectrum o…

X AI KOLs ↗ · 4d ago Cached

MentalHealthBench is designed to cover the full spectrum of mental health conversations for AI, from everyday support to acute crisis scenarios, as announced by OpenAI.

0 favorites 0 likes
#evaluation

GPT-6 Astra has gained the ability to drive a car

Hacker News Top ↗ · 4d ago Cached

DrivingBench evaluates whether frontier AI models like GPT-6 Astra can drive a real car by controlling a Toyota Corolla on a predefined course, tracking metrics such as progress, distance, and token cost.

0 favorites 0 likes
#evaluation

How do you decide which AI agents are worth keeping in production?

Reddit r/AI_Agents ↗ · 4d ago

The article explores methods for evaluating AI agents in production to decide whether to retain, improve, or shut them down, citing research on metrics like cost, reliability, human effort, and business outcomes.

0 favorites 0 likes
#evaluation

Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation

arXiv cs.AI ↗ · 4d ago Cached

This paper examines the limitations of general LLM rankings, including benchmark saturation and commercial incentives, and advocates for task-specific evaluation methods to improve reliability.

0 favorites 0 likes
#evaluation

CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine

arXiv cs.AI ↗ · 4d ago Cached

CraftBench-UE presents a deterministic evaluation harness for coding agents in Unreal Engine, enabling assessment of gameplay features through build, asset, and runtime checks without relying on LLM judges.

0 favorites 0 likes
#evaluation

ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures

arXiv cs.AI ↗ · 4d ago Cached

ISA-Bench introduces a benchmark using programming games with constrained instruction sets to evaluate computational reasoning in large language models, revealing insights into model capabilities and a reasoning-execution gap.

0 favorites 0 likes
#evaluation

ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch

arXiv cs.CL ↗ · 4d ago Cached

The paper details ufakzeka-1, a 151M-parameter Turkish language model built from scratch with a total cost of about $286, describing the training pipeline, evaluation methods, and key findings on small-model training limitations.

0 favorites 0 likes
#evaluation

Understanding Reliability in LLM-based Human Behavior Simulation

arXiv cs.CL ↗ · 4d ago Cached

This paper proposes ReliMap, a framework for evaluating reliability in LLM-based human behavior simulations across individual and population levels, emphasizing coordinated improvements in model capacity, profile completeness, and coverage.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback