benchmark

Tag

Cards List
#benchmark

@_philschmid: Gemini 3.8 Flash. Just use Gemini for multimodal understanding.

X AI KOLs Following ↗ · 2d ago Cached

A tweet suggests using Gemini 3.8 Flash for multimodal understanding and references a visual test comparing AI models' ability to name people from a drawing.

0 favorites 0 likes
#benchmark

Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark.

Reddit r/LocalLLaMA ↗ · 2d ago

The author shares lessons learned from a long-running Django benchmark, highlighting fixes in the evaluation workflow stability and updated results showing Flash Next as the top performer with reasoning effort levels now properly evaluated.

0 favorites 0 likes
#benchmark

Fable 5.1 and Astra have both achieved perfect scores on Mensa Norway

Reddit r/singularity ↗ · 2d ago

Fable 5.1 and Astra, two AI models, have both achieved perfect scores on the Mensa Norway intelligence test, demonstrating their exceptional reasoning abilities.

0 favorites 0 likes
#benchmark

Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools

arXiv cs.LG ↗ · 2d ago Cached

FWBench introduces a benchmark for evaluating how language models select and use time-series forecasts to make cost-constrained decisions, comparing hosted and local configurations on electricity and cycle-hire datasets with efficient budget usage by GPT-6 Astra.

0 favorites 0 likes
#benchmark

An open benchmark for machine learning-based polymer property prediction

arXiv cs.LG ↗ · 2d ago Cached

The paper introduces Polymer Benchmark 2026, an open dataset for benchmarking machine learning methods in polymer property prediction across diverse architectures and properties.

0 favorites 0 likes
#benchmark

Reachable Global Optimization in AI Systems: How Global Is Global?

arXiv cs.AI ↗ · 2d ago Cached

This paper introduces Reachability-Induced Optimization (RIO) to argue that global optimization claims in AI systems should be based on the actually reachable region, providing theoretical results and benchmark data to support this framework.

0 favorites 0 likes
#benchmark

Improving LLM-based Autonomous Web Agents with Filtering

arXiv cs.CL ↗ · 2d ago Cached

The paper proposes retrieval strategies to filter irrelevant HTML content for LLM-based autonomous web agents, improving performance on benchmarks like WebArena.

0 favorites 0 likes
#benchmark

Not What You Meant: Can LLMs Follow a Specified Negation Semantics?

arXiv cs.AI ↗ · 2d ago Cached

This paper introduces NAF-Bench to study how large language models adhere to specified negation semantics, finding that frontier models like o4-mini perform well while open-source models lag, and suggesting improvements via solver delegation or fine-tuning.

0 favorites 0 likes
#benchmark

When Context Misleads: In-context Learning with Jurisdiction in Large Language Models

arXiv cs.CL ↗ · 2d ago Cached

This paper introduces FakeContext-bench to evaluate how well large language models distinguish between contextual information and factual knowledge, and proposes Jurisdiction In-Context Learning (J-ICL) to enhance both in-context learning performance and resistance to misleading context.

0 favorites 0 likes
#benchmark

MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

arXiv cs.AI ↗ · 2d ago Cached

MolDesignBench is a new benchmark for evaluating LLM-based agents in scenario-grounded molecular design, comprising 2K instances with implicit and explicit constraints, revealing that current frontier LLMs achieve low success rates, especially in reasoning about implicit constraints and infeasibility detection.

0 favorites 0 likes
#benchmark

PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models

arXiv cs.CL ↗ · 2d ago Cached

PRISM-VLM is a multi-axis discriminative benchmark that evaluates compact vision-language models along seven axes, including task quality and behavioral robustness, to provide more reliable separation and insights compared to single-axis benchmarks. It aims to release an open pipeline for the community to improve AI evaluation methods.

0 favorites 0 likes
#benchmark

PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks

arXiv cs.AI ↗ · 2d ago Cached

PotARCin is a multi-dimensional benchmark that extends ARC to evaluate abstract reasoning skills across five dimensions, showing that standard evaluations may not fully capture AI models' true capabilities.

0 favorites 0 likes
#benchmark

Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models

arXiv cs.CL ↗ · 2d ago Cached

The paper introduces TACT, a benchmark for evaluating turn-taking in full-duplex spoken dialogue models that uses intent-conditioned continuous scoring to replace binary rules, demonstrating better alignment with human judgments.

0 favorites 0 likes
#benchmark

Automated Extraction of Records of Processing Activities (RoPA) Using Hybrid RAG and Locally Deployed Large Language Models

arXiv cs.CL ↗ · 2d ago Cached

This paper proposes RoPA Manager, an automated system for extracting Records of Processing Activities using hybrid retrieval and locally deployed large language models, evaluated on a Vietnamese benchmark.

0 favorites 0 likes
#benchmark

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

arXiv cs.AI ↗ · 2d ago Cached

This paper introduces CAVEAT, a benchmark for evaluating computer-use agents in incentive-misaligned environments, revealing that agents often fail to preserve user objectives and proposes interventions to improve robustness.

0 favorites 0 likes
#benchmark

Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach

arXiv cs.CL ↗ · 2d ago Cached

This paper introduces UGTPhon, a benchmark for grapheme-to-phoneme conversion in user-generated text, and presents a compositional approach that improves performance by leveraging canonical forms.

0 favorites 0 likes
#benchmark

Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure

arXiv cs.CL ↗ · 2d ago Cached

This paper proposes a method to estimate the causal effect of benchmark exposure on AI model performance, moving beyond traditional overlap techniques for more robust evaluation.

0 favorites 0 likes
#benchmark

Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

arXiv cs.CL ↗ · 2d ago Cached

This paper introduces a sentence-level benchmark for classifying interpretive canons in legal documents, evaluating LLMs on a dataset from the German Federal Constitutional Court.

0 favorites 0 likes
#benchmark

Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms

arXiv cs.CL ↗ · 2d ago Cached

This paper introduces a generation benchmark for evaluating large language models on culturally specific kinship terms in Hindi, Tamil, and Korean, revealing significant gaps between recognition and generation capabilities.

0 favorites 0 likes
#benchmark

Gemini 4 Pro nears its preview release. (Yes, another preview)

Reddit r/singularity ↗ · 2d ago

Google's Gemini 4 Pro AI model is nearing a preview release with early post-training versions expected to outperform competitors like Astra, with a possible public release in October.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback