Tag
This paper introduces CAVEAT, a benchmark for evaluating computer-use agents in incentive-misaligned environments, revealing that agents often fail to preserve user objectives and proposes interventions to improve robustness.
This paper introduces UGTPhon, a benchmark for grapheme-to-phoneme conversion in user-generated text, and presents a compositional approach that improves performance by leveraging canonical forms.
This paper proposes a method to estimate the causal effect of benchmark exposure on AI model performance, moving beyond traditional overlap techniques for more robust evaluation.
This paper introduces a sentence-level benchmark for classifying interpretive canons in legal documents, evaluating LLMs on a dataset from the German Federal Constitutional Court.
This paper introduces a generation benchmark for evaluating large language models on culturally specific kinship terms in Hindi, Tamil, and Korean, revealing significant gaps between recognition and generation capabilities.
Google's Gemini 4 Pro AI model is nearing a preview release with early post-training versions expected to outperform competitors like Astra, with a possible public release in October.
The tweet discusses using AI model Opus 5.5 to discover optimizations for shell startup time via the ShellPerfBench tool, recommending users enhance their .zshrc files.
This paper introduces RGBD20K, a large-scale benchmark dataset for RGB-D semantic segmentation with 160 fine-grained categories and 20,000 image pairs, featuring high-quality annotations and a novel score-purified fusion method.
This paper presents RLCDAlignBench, a benchmark for evaluating Jev, a model trained with reinforcement learning for calibrated decisions, as a zero-shot detector for ten AI alignment failures with high performance and cost efficiency.
The paper introduces ExplorationBench, a benchmark for evaluating AI systems' exploration abilities in verifiable alien worlds, addressing challenges in scientific discovery by providing executable rules and preventing recall from pre-training data.
WanPE is a 397B-parameter prompt enhancement model that improves cinematic planning in text-to-video generation, showing significant human preference boosts over raw prompts and competitive performance with commercial offerings.
FLUX 3 Action is an open-weight 7B AI model that sets a new standard on the RoboLab benchmark, outperforming previous open models while being faster and more efficient, with applications in robotics and other environments requiring visual action planning.
MentalHealthBench is designed to cover the full spectrum of mental health conversations for AI, from everyday support to acute crisis scenarios, as announced by OpenAI.
VoiceArena launched Diarization Bench, a benchmark comparing 12 diarization systems on real-world audio, highlighting how accuracy drops with overlap and the impact of collar effect on error rates.
DrivingBench evaluates whether frontier AI models like GPT-6 Astra can drive a real car by controlling a Toyota Corolla on a predefined course, tracking metrics such as progress, distance, and token cost.
OpenAI introduces MentalHealthBench, an open benchmark for evaluating AI responses in mental health conversations, co-created with over 80 mental health experts to measure safety, context, agency, and guidance.
Open-Jev-27B-v1.1 is an open-source AI model with a LoRA adapter, achieving 85.28% accuracy on JevBench and featuring interactive demos for various tasks.
SambaGraph introduces a spatio-temporal graph dataset and benchmark for modeling soccer tactical responses, curated from 2022 FIFA World Cup data to classify actions and retrieve defensive examples.
TicTacBench is a new benchmark for evaluating coding agents' timing closure capabilities in RTL design, revealing that current agents have significant room for improvement and proposing TicTacSkill to enhance performance.
CraftBench-UE presents a deterministic evaluation harness for coding agents in Unreal Engine, enabling assessment of gameplay features through build, asset, and runtime checks without relying on LLM judges.