benchmark

Tag

Cards List
#benchmark

Gemini 4 Pro nears its preview release. (Yes, another preview)

Reddit r/singularity ↗ · 4d ago

Google's Gemini 4 Pro AI model is nearing a preview release with early post-training versions expected to outperform competitors like Astra, with a possible public release in October.

0 favorites 0 likes
#benchmark

@rauchg: ShellPerfBench. I'm really neurotic about the startup time of a new shell session. Opus 5.5 found a lot of really great…

X AI KOLs Timeline ↗ · 4d ago

The tweet discusses using AI model Opus 5.5 to discover optimizations for shell startup time via the ShellPerfBench tool, recommending users enhance their .zshrc files.

0 favorites 0 likes
#benchmark

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper introduces RGBD20K, a large-scale benchmark dataset for RGB-D semantic segmentation with 160 fine-grained categories and 20,000 image pairs, featuring high-quality annotations and a novel score-purified fusion method.

0 favorites 0 likes
#benchmark

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper presents RLCDAlignBench, a benchmark for evaluating Jev, a model trained with reinforcement learning for calibrated decisions, as a zero-shot detector for ten AI alignment failures with high performance and cost efficiency.

0 favorites 0 likes
#benchmark

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Hugging Face Daily Papers ↗ · 4d ago Cached

The paper introduces ExplorationBench, a benchmark for evaluating AI systems' exploration abilities in verifiable alien worlds, addressing challenges in scientific discovery by providing executable rules and preventing recall from pre-training data.

0 favorites 0 likes
#benchmark

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Hugging Face Daily Papers ↗ · 4d ago Cached

WanPE is a 397B-parameter prompt enhancement model that improves cinematic planning in text-to-video generation, showing significant human preference boosts over raw prompts and competitive performance with commercial offerings.

0 favorites 0 likes
#benchmark

Flux 3 Action: open weights 7B World Action Model

Reddit r/singularity ↗ · 4d ago

FLUX 3 Action is an open-weight 7B AI model that sets a new standard on the RoboLab benchmark, outperforming previous open models while being faster and more efficient, with applications in robotics and other environments requiring visual action planning.

0 favorites 0 likes
#benchmark

@OpenAI: Most mental health benchmarks focus on emergency situations. MentalHealthBench is designed to cover the full spectrum o…

X AI KOLs ↗ · 4d ago Cached

MentalHealthBench is designed to cover the full spectrum of mental health conversations for AI, from everyday support to acute crisis scenarios, as announced by OpenAI.

0 favorites 0 likes
#benchmark

@rohanpaul_ai: VoiceArena launched Diarization Bench to compare 12 diarization systems on the same real-world audio and scoring rules.…

X AI KOLs Following ↗ · 4d ago Cached

VoiceArena launched Diarization Bench, a benchmark comparing 12 diarization systems on real-world audio, highlighting how accuracy drops with overlap and the impact of collar effect on error rates.

0 favorites 0 likes
#benchmark

GPT-6 Astra has gained the ability to drive a car

Hacker News Top ↗ · 4d ago Cached

DrivingBench evaluates whether frontier AI models like GPT-6 Astra can drive a real car by controlling a Toyota Corolla on a predefined course, tracking metrics such as progress, distance, and token cost.

0 favorites 0 likes
#benchmark

Introducing MentalHealthBench

OpenAI Blog ↗ · 5d ago Cached

OpenAI introduces MentalHealthBench, an open benchmark for evaluating AI responses in mental health conversations, co-created with over 80 mental health experts to measure safety, context, agency, and guidance.

0 favorites 0 likes
#benchmark

@Zefan_Cai: Spot the wrong target before you hit Confirm. Try the Open-Jev-27B-v1.1 demo: click between two examples and see how th…

X AI KOLs Following ↗ · 5d ago Cached

Open-Jev-27B-v1.1 is an open-source AI model with a LoRA adapter, achieving 85.28% accuracy on JevBench and featuring interactive demos for various tasks.

0 favorites 0 likes
#benchmark

SambaGraph: Action-Reaction Spatio-Temporal Graphs for Soccer Tactical Response Modeling

arXiv cs.LG ↗ · 5d ago Cached

SambaGraph introduces a spatio-temporal graph dataset and benchmark for modeling soccer tactical responses, curated from 2022 FIFA World Cup data to classify actions and retrieve defensive examples.

0 favorites 0 likes
#benchmark

TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents

arXiv cs.AI ↗ · 5d ago Cached

TicTacBench is a new benchmark for evaluating coding agents' timing closure capabilities in RTL design, revealing that current agents have significant room for improvement and proposing TicTacSkill to enhance performance.

0 favorites 0 likes
#benchmark

CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine

arXiv cs.AI ↗ · 5d ago Cached

CraftBench-UE presents a deterministic evaluation harness for coding agents in Unreal Engine, enabling assessment of gameplay features through build, asset, and runtime checks without relying on LLM judges.

0 favorites 0 likes
#benchmark

ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures

arXiv cs.AI ↗ · 5d ago Cached

ISA-Bench introduces a benchmark using programming games with constrained instruction sets to evaluate computational reasoning in large language models, revealing insights into model capabilities and a reasoning-execution gap.

0 favorites 0 likes
#benchmark

IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law

arXiv cs.AI ↗ · 5d ago Cached

IntLawNER is a new named entity recognition dataset and benchmark for international law, covering gold-annotated sentences from legal texts and evaluating model performance with few-shot improvements.

0 favorites 0 likes
#benchmark

Peerify: Benchmarking Peer-Review Claim Verification

arXiv cs.CL ↗ · 5d ago Cached

Peerify is a pipeline for automatically verifying peer-review claims against manuscript evidence, using a benchmark of 800 claims from NeurIPS 2024 and ICLR 2024, demonstrating that retrieval-centered verification outperforms entailment baselines.

0 favorites 0 likes
#benchmark

Same Quantity, Different Answer: Numerical Representation Invariance in Language Models

arXiv cs.CL ↗ · 5d ago Cached

This paper evaluates numerical representation invariance in language models, finding that evaluator interface issues can mimic reasoning failures and identifying model-specific errors like unit conversion problems in Mistral Small 4.

0 favorites 0 likes
#benchmark

Training Object Permanence in World Models

Hugging Face Daily Papers ↗ · 5d ago Cached

This paper introduces WROP, a dataset for training object permanence in world models, and evaluates 14 video models, releasing PWM-WROP, a 16B world model that ranks among top continuation models.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback