Tag
Google's Gemini 4 Pro AI model is nearing a preview release with early post-training versions expected to outperform competitors like Astra, with a possible public release in October.
The tweet discusses using AI model Opus 5.5 to discover optimizations for shell startup time via the ShellPerfBench tool, recommending users enhance their .zshrc files.
This paper introduces RGBD20K, a large-scale benchmark dataset for RGB-D semantic segmentation with 160 fine-grained categories and 20,000 image pairs, featuring high-quality annotations and a novel score-purified fusion method.
This paper presents RLCDAlignBench, a benchmark for evaluating Jev, a model trained with reinforcement learning for calibrated decisions, as a zero-shot detector for ten AI alignment failures with high performance and cost efficiency.
The paper introduces ExplorationBench, a benchmark for evaluating AI systems' exploration abilities in verifiable alien worlds, addressing challenges in scientific discovery by providing executable rules and preventing recall from pre-training data.
WanPE is a 397B-parameter prompt enhancement model that improves cinematic planning in text-to-video generation, showing significant human preference boosts over raw prompts and competitive performance with commercial offerings.
FLUX 3 Action is an open-weight 7B AI model that sets a new standard on the RoboLab benchmark, outperforming previous open models while being faster and more efficient, with applications in robotics and other environments requiring visual action planning.
MentalHealthBench is designed to cover the full spectrum of mental health conversations for AI, from everyday support to acute crisis scenarios, as announced by OpenAI.
VoiceArena launched Diarization Bench, a benchmark comparing 12 diarization systems on real-world audio, highlighting how accuracy drops with overlap and the impact of collar effect on error rates.
DrivingBench evaluates whether frontier AI models like GPT-6 Astra can drive a real car by controlling a Toyota Corolla on a predefined course, tracking metrics such as progress, distance, and token cost.
OpenAI introduces MentalHealthBench, an open benchmark for evaluating AI responses in mental health conversations, co-created with over 80 mental health experts to measure safety, context, agency, and guidance.
Open-Jev-27B-v1.1 is an open-source AI model with a LoRA adapter, achieving 85.28% accuracy on JevBench and featuring interactive demos for various tasks.
SambaGraph introduces a spatio-temporal graph dataset and benchmark for modeling soccer tactical responses, curated from 2022 FIFA World Cup data to classify actions and retrieve defensive examples.
TicTacBench is a new benchmark for evaluating coding agents' timing closure capabilities in RTL design, revealing that current agents have significant room for improvement and proposing TicTacSkill to enhance performance.
CraftBench-UE presents a deterministic evaluation harness for coding agents in Unreal Engine, enabling assessment of gameplay features through build, asset, and runtime checks without relying on LLM judges.
ISA-Bench introduces a benchmark using programming games with constrained instruction sets to evaluate computational reasoning in large language models, revealing insights into model capabilities and a reasoning-execution gap.
IntLawNER is a new named entity recognition dataset and benchmark for international law, covering gold-annotated sentences from legal texts and evaluating model performance with few-shot improvements.
Peerify is a pipeline for automatically verifying peer-review claims against manuscript evidence, using a benchmark of 800 claims from NeurIPS 2024 and ICLR 2024, demonstrating that retrieval-centered verification outperforms entailment baselines.
This paper evaluates numerical representation invariance in language models, finding that evaluator interface issues can mimic reasoning failures and identifying model-specific errors like unit conversion problems in Mistral Small 4.
This paper introduces WROP, a dataset for training object permanence in world models, and evaluates 14 video models, releasing PWM-WROP, a 16B world model that ranks among top continuation models.