Tag
A tweet suggests using Gemini 3.8 Flash for multimodal understanding and references a visual test comparing AI models' ability to name people from a drawing.
The author shares lessons learned from a long-running Django benchmark, highlighting fixes in the evaluation workflow stability and updated results showing Flash Next as the top performer with reasoning effort levels now properly evaluated.
Fable 5.1 and Astra, two AI models, have both achieved perfect scores on the Mensa Norway intelligence test, demonstrating their exceptional reasoning abilities.
FWBench introduces a benchmark for evaluating how language models select and use time-series forecasts to make cost-constrained decisions, comparing hosted and local configurations on electricity and cycle-hire datasets with efficient budget usage by GPT-6 Astra.
The paper introduces Polymer Benchmark 2026, an open dataset for benchmarking machine learning methods in polymer property prediction across diverse architectures and properties.
This paper introduces Reachability-Induced Optimization (RIO) to argue that global optimization claims in AI systems should be based on the actually reachable region, providing theoretical results and benchmark data to support this framework.
The paper proposes retrieval strategies to filter irrelevant HTML content for LLM-based autonomous web agents, improving performance on benchmarks like WebArena.
This paper introduces NAF-Bench to study how large language models adhere to specified negation semantics, finding that frontier models like o4-mini perform well while open-source models lag, and suggesting improvements via solver delegation or fine-tuning.
This paper introduces FakeContext-bench to evaluate how well large language models distinguish between contextual information and factual knowledge, and proposes Jurisdiction In-Context Learning (J-ICL) to enhance both in-context learning performance and resistance to misleading context.
MolDesignBench is a new benchmark for evaluating LLM-based agents in scenario-grounded molecular design, comprising 2K instances with implicit and explicit constraints, revealing that current frontier LLMs achieve low success rates, especially in reasoning about implicit constraints and infeasibility detection.
PRISM-VLM is a multi-axis discriminative benchmark that evaluates compact vision-language models along seven axes, including task quality and behavioral robustness, to provide more reliable separation and insights compared to single-axis benchmarks. It aims to release an open pipeline for the community to improve AI evaluation methods.
PotARCin is a multi-dimensional benchmark that extends ARC to evaluate abstract reasoning skills across five dimensions, showing that standard evaluations may not fully capture AI models' true capabilities.
The paper introduces TACT, a benchmark for evaluating turn-taking in full-duplex spoken dialogue models that uses intent-conditioned continuous scoring to replace binary rules, demonstrating better alignment with human judgments.
This paper proposes RoPA Manager, an automated system for extracting Records of Processing Activities using hybrid retrieval and locally deployed large language models, evaluated on a Vietnamese benchmark.
This paper introduces CAVEAT, a benchmark for evaluating computer-use agents in incentive-misaligned environments, revealing that agents often fail to preserve user objectives and proposes interventions to improve robustness.
This paper introduces UGTPhon, a benchmark for grapheme-to-phoneme conversion in user-generated text, and presents a compositional approach that improves performance by leveraging canonical forms.
This paper proposes a method to estimate the causal effect of benchmark exposure on AI model performance, moving beyond traditional overlap techniques for more robust evaluation.
This paper introduces a sentence-level benchmark for classifying interpretive canons in legal documents, evaluating LLMs on a dataset from the German Federal Constitutional Court.
This paper introduces a generation benchmark for evaluating large language models on culturally specific kinship terms in Hindi, Tamil, and Korean, revealing significant gaps between recognition and generation capabilities.
Google's Gemini 4 Pro AI model is nearing a preview release with early post-training versions expected to outperform competitors like Astra, with a possible public release in October.