benchmark

Tag

Cards List
#benchmark

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

arXiv cs.AI ↗ · 4d ago Cached

This paper introduces CAVEAT, a benchmark for evaluating computer-use agents in incentive-misaligned environments, revealing that agents often fail to preserve user objectives and proposes interventions to improve robustness.

0 favorites 0 likes
#benchmark

Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach

arXiv cs.CL ↗ · 4d ago Cached

This paper introduces UGTPhon, a benchmark for grapheme-to-phoneme conversion in user-generated text, and presents a compositional approach that improves performance by leveraging canonical forms.

0 favorites 0 likes
#benchmark

Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure

arXiv cs.CL ↗ · 4d ago Cached

This paper proposes a method to estimate the causal effect of benchmark exposure on AI model performance, moving beyond traditional overlap techniques for more robust evaluation.

0 favorites 0 likes
#benchmark

Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

arXiv cs.CL ↗ · 4d ago Cached

This paper introduces a sentence-level benchmark for classifying interpretive canons in legal documents, evaluating LLMs on a dataset from the German Federal Constitutional Court.

0 favorites 0 likes
#benchmark

Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms

arXiv cs.CL ↗ · 4d ago Cached

This paper introduces a generation benchmark for evaluating large language models on culturally specific kinship terms in Hindi, Tamil, and Korean, revealing significant gaps between recognition and generation capabilities.

0 favorites 0 likes
#benchmark

Gemini 4 Pro nears its preview release. (Yes, another preview)

Reddit r/singularity ↗ · 4d ago

Google's Gemini 4 Pro AI model is nearing a preview release with early post-training versions expected to outperform competitors like Astra, with a possible public release in October.

0 favorites 0 likes
#benchmark

@rauchg: ShellPerfBench. I'm really neurotic about the startup time of a new shell session. Opus 5.5 found a lot of really great…

X AI KOLs Timeline ↗ · 4d ago

The tweet discusses using AI model Opus 5.5 to discover optimizations for shell startup time via the ShellPerfBench tool, recommending users enhance their .zshrc files.

0 favorites 0 likes
#benchmark

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper introduces RGBD20K, a large-scale benchmark dataset for RGB-D semantic segmentation with 160 fine-grained categories and 20,000 image pairs, featuring high-quality annotations and a novel score-purified fusion method.

0 favorites 0 likes
#benchmark

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper presents RLCDAlignBench, a benchmark for evaluating Jev, a model trained with reinforcement learning for calibrated decisions, as a zero-shot detector for ten AI alignment failures with high performance and cost efficiency.

0 favorites 0 likes
#benchmark

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Hugging Face Daily Papers ↗ · 4d ago Cached

The paper introduces ExplorationBench, a benchmark for evaluating AI systems' exploration abilities in verifiable alien worlds, addressing challenges in scientific discovery by providing executable rules and preventing recall from pre-training data.

0 favorites 0 likes
#benchmark

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Hugging Face Daily Papers ↗ · 4d ago Cached

WanPE is a 397B-parameter prompt enhancement model that improves cinematic planning in text-to-video generation, showing significant human preference boosts over raw prompts and competitive performance with commercial offerings.

0 favorites 0 likes
#benchmark

Flux 3 Action: open weights 7B World Action Model

Reddit r/singularity ↗ · 5d ago

FLUX 3 Action is an open-weight 7B AI model that sets a new standard on the RoboLab benchmark, outperforming previous open models while being faster and more efficient, with applications in robotics and other environments requiring visual action planning.

0 favorites 0 likes
#benchmark

@OpenAI: Most mental health benchmarks focus on emergency situations. MentalHealthBench is designed to cover the full spectrum o…

X AI KOLs ↗ · 5d ago Cached

MentalHealthBench is designed to cover the full spectrum of mental health conversations for AI, from everyday support to acute crisis scenarios, as announced by OpenAI.

0 favorites 0 likes
#benchmark

@rohanpaul_ai: VoiceArena launched Diarization Bench to compare 12 diarization systems on the same real-world audio and scoring rules.…

X AI KOLs Following ↗ · 5d ago Cached

VoiceArena launched Diarization Bench, a benchmark comparing 12 diarization systems on real-world audio, highlighting how accuracy drops with overlap and the impact of collar effect on error rates.

0 favorites 0 likes
#benchmark

GPT-6 Astra has gained the ability to drive a car

Hacker News Top ↗ · 5d ago Cached

DrivingBench evaluates whether frontier AI models like GPT-6 Astra can drive a real car by controlling a Toyota Corolla on a predefined course, tracking metrics such as progress, distance, and token cost.

0 favorites 0 likes
#benchmark

Introducing MentalHealthBench

OpenAI Blog ↗ · 5d ago Cached

OpenAI introduces MentalHealthBench, an open benchmark for evaluating AI responses in mental health conversations, co-created with over 80 mental health experts to measure safety, context, agency, and guidance.

0 favorites 0 likes
#benchmark

@Zefan_Cai: Spot the wrong target before you hit Confirm. Try the Open-Jev-27B-v1.1 demo: click between two examples and see how th…

X AI KOLs Following ↗ · 5d ago Cached

Open-Jev-27B-v1.1 is an open-source AI model with a LoRA adapter, achieving 85.28% accuracy on JevBench and featuring interactive demos for various tasks.

0 favorites 0 likes
#benchmark

SambaGraph: Action-Reaction Spatio-Temporal Graphs for Soccer Tactical Response Modeling

arXiv cs.LG ↗ · 5d ago Cached

SambaGraph introduces a spatio-temporal graph dataset and benchmark for modeling soccer tactical responses, curated from 2022 FIFA World Cup data to classify actions and retrieve defensive examples.

0 favorites 0 likes
#benchmark

TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents

arXiv cs.AI ↗ · 5d ago Cached

TicTacBench is a new benchmark for evaluating coding agents' timing closure capabilities in RTL design, revealing that current agents have significant room for improvement and proposing TicTacSkill to enhance performance.

0 favorites 0 likes
#benchmark

CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine

arXiv cs.AI ↗ · 5d ago Cached

CraftBench-UE presents a deterministic evaluation harness for coding agents in Unreal Engine, enabling assessment of gameplay features through build, asset, and runtime checks without relying on LLM judges.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback