evaluation

Tag

Cards List
#evaluation

MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

arXiv cs.AI ↗ · yesterday Cached

MolDesignBench is a new benchmark for evaluating LLM-based agents in scenario-grounded molecular design, comprising 2K instances with implicit and explicit constraints, revealing that current frontier LLMs achieve low success rates, especially in reasoning about implicit constraints and infeasibility detection.

0 favorites 0 likes
#evaluation

PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models

arXiv cs.CL ↗ · yesterday Cached

PRISM-VLM is a multi-axis discriminative benchmark that evaluates compact vision-language models along seven axes, including task quality and behavioral robustness, to provide more reliable separation and insights compared to single-axis benchmarks. It aims to release an open pipeline for the community to improve AI evaluation methods.

0 favorites 0 likes
#evaluation

PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks

arXiv cs.AI ↗ · yesterday Cached

PotARCin is a multi-dimensional benchmark that extends ARC to evaluate abstract reasoning skills across five dimensions, showing that standard evaluations may not fully capture AI models' true capabilities.

0 favorites 0 likes
#evaluation

Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models

arXiv cs.CL ↗ · yesterday Cached

The paper introduces TACT, a benchmark for evaluating turn-taking in full-duplex spoken dialogue models that uses intent-conditioned continuous scoring to replace binary rules, demonstrating better alignment with human judgments.

0 favorites 0 likes
#evaluation

Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure

arXiv cs.CL ↗ · yesterday Cached

This paper proposes a method to estimate the causal effect of benchmark exposure on AI model performance, moving beyond traditional overlap techniques for more robust evaluation.

0 favorites 0 likes
#evaluation

When Learned Context Planning Fails to Beat Strong Retrieval: A Controlled Study of Planning, Routing, and Reranking for Long-Context QA

arXiv cs.CL ↗ · yesterday Cached

This paper presents a controlled study on learned context planning for long-context multiple-choice QA, demonstrating that it does not robustly outperform strong retrieval methods like BM25 and hybrid retrieval under various experimental settings.

0 favorites 0 likes
#evaluation

Do y'all remember the snake model evaluation test?

Reddit r/LocalLLaMA ↗ · 2d ago

The article reflects on the rapid progress in AI models from three years ago, comparing early models like Bard that struggled with coding tests to current models like Qwen 27B and Opus 5.5, and speculates on future advancements.

0 favorites 0 likes
#evaluation

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Hugging Face Daily Papers ↗ · 2d ago Cached

The paper introduces ExplorationBench, a benchmark for evaluating AI systems' exploration abilities in verifiable alien worlds, addressing challenges in scientific discovery by providing executable rules and preventing recall from pre-training data.

0 favorites 0 likes
#evaluation

The AI you test in the afternoon may not be the AI you test at night, even with the same name

Reddit r/artificial ↗ · 2d ago

Testing Claude Opus 5 twice in one day revealed significant response differences, likely due to background settings rather than model changes, highlighting the need to document all settings when comparing AI tools.

0 favorites 0 likes
#evaluation

@OpenAI: Most mental health benchmarks focus on emergency situations. MentalHealthBench is designed to cover the full spectrum o…

X AI KOLs ↗ · 2d ago Cached

MentalHealthBench is designed to cover the full spectrum of mental health conversations for AI, from everyday support to acute crisis scenarios, as announced by OpenAI.

0 favorites 0 likes
#evaluation

GPT-6 Astra has gained the ability to drive a car

Hacker News Top ↗ · 2d ago Cached

DrivingBench evaluates whether frontier AI models like GPT-6 Astra can drive a real car by controlling a Toyota Corolla on a predefined course, tracking metrics such as progress, distance, and token cost.

0 favorites 0 likes
#evaluation

How do you decide which AI agents are worth keeping in production?

Reddit r/AI_Agents ↗ · 2d ago

The article explores methods for evaluating AI agents in production to decide whether to retain, improve, or shut them down, citing research on metrics like cost, reliability, human effort, and business outcomes.

0 favorites 0 likes
#evaluation

Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation

arXiv cs.AI ↗ · 2d ago Cached

This paper examines the limitations of general LLM rankings, including benchmark saturation and commercial incentives, and advocates for task-specific evaluation methods to improve reliability.

0 favorites 0 likes
#evaluation

CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine

arXiv cs.AI ↗ · 2d ago Cached

CraftBench-UE presents a deterministic evaluation harness for coding agents in Unreal Engine, enabling assessment of gameplay features through build, asset, and runtime checks without relying on LLM judges.

0 favorites 0 likes
#evaluation

ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures

arXiv cs.AI ↗ · 2d ago Cached

ISA-Bench introduces a benchmark using programming games with constrained instruction sets to evaluate computational reasoning in large language models, revealing insights into model capabilities and a reasoning-execution gap.

0 favorites 0 likes
#evaluation

ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch

arXiv cs.CL ↗ · 2d ago Cached

The paper details ufakzeka-1, a 151M-parameter Turkish language model built from scratch with a total cost of about $286, describing the training pipeline, evaluation methods, and key findings on small-model training limitations.

0 favorites 0 likes
#evaluation

Understanding Reliability in LLM-based Human Behavior Simulation

arXiv cs.CL ↗ · 2d ago Cached

This paper proposes ReliMap, a framework for evaluating reliability in LLM-based human behavior simulations across individual and population levels, emphasizing coordinated improvements in model capacity, profile completeness, and coverage.

0 favorites 0 likes
#evaluation

IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law

arXiv cs.AI ↗ · 2d ago Cached

IntLawNER is a new named entity recognition dataset and benchmark for international law, covering gold-annotated sentences from legal texts and evaluating model performance with few-shot improvements.

0 favorites 0 likes
#evaluation

Same Quantity, Different Answer: Numerical Representation Invariance in Language Models

arXiv cs.CL ↗ · 2d ago Cached

This paper evaluates numerical representation invariance in language models, finding that evaluator interface issues can mimic reasoning failures and identifying model-specific errors like unit conversion problems in Mistral Small 4.

0 favorites 0 likes
#evaluation

Training Object Permanence in World Models

Hugging Face Daily Papers ↗ · 3d ago Cached

This paper introduces WROP, a dataset for training object permanence in world models, and evaluates 14 video models, releasing PWM-WROP, a 16B world model that ranks among top continuation models.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback