benchmark

Tag

Cards List
#benchmark

MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs

arXiv cs.CL ↗ · 2026-09-21 Cached

MME-Safety is a rigorously verified benchmark for evaluating the safety of Multimodal Large Language Models, featuring a four-dimensional annotation schema and a hierarchical framework to assess risk scenarios, harm severity, and modality-specific stealth levels.

0 favorites 0 likes
#benchmark

VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering

arXiv cs.CL ↗ · 2026-09-21 Cached

VisPath introduces a visual-intent-guided path reasoning framework for multimodal knowledge graph question answering, achieving significant improvements over baselines on a new benchmark, VisPath-Bench, and existing datasets.

0 favorites 0 likes
#benchmark

PhysioBench: A Unified Benchmark for Physiological Signal Question Answering

arXiv cs.CL ↗ · 2026-09-21 Cached

PhysioBench introduces a unified benchmark for physiological signal question answering, harmonizing 22 datasets into 61.4 million questions across 30 tasks to evaluate the performance of various AI models.

0 favorites 0 likes
#benchmark

HappyWorld-Bench

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

HappyWorld-Bench is a comprehensive benchmark that evaluates the reliability of world models under interaction and modification across video, spatial, and embodied tracks.

0 favorites 0 likes
#benchmark

VideoGen-Agent: Reinforcing Video Generation Agents

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

The paper presents VideoGen-Agent, a reinforcement learning-based multimodal agent that coordinates tools for video generation, significantly improving performance on the new VABench benchmark.

0 favorites 0 likes
#benchmark

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

The paper introduces GameHorizon Suite, a unified data and evaluation framework for assessing AI models' capabilities in gameplay across multiple temporal horizons, featuring an annotation pipeline, large-scale dataset, and reproducible benchmark.

0 favorites 0 likes
#benchmark

Hemmingway-1, an Apache-2.0 27B creative-writing fine-tune (Qwen3.8-27B base, EQ-Bench 4 1330)[R]

Reddit r/MachineLearning ↗ · 2026-09-20

A small lab from Switzerland and South Africa has open-sourced Hemmingway-1, a 27B Apache-2.0 licensed fine-tune of Qwen3.8-27B specialized for creative writing, achieving high scores on EQ-Bench and internal benchmarks.

0 favorites 0 likes
#benchmark

@QuixiAI: Nice! Can't wait to see the community LoRAs

X AI KOLs Timeline ↗ · 2026-09-20 Cached

The tweet expresses anticipation for community-created LoRAs following the announcement of an open-weight 7B model that reportedly surpasses Nano Banana 2.0 in performance.

0 favorites 0 likes
#benchmark

@FinanceYF5: 1/ Jev + Mercury 2.5 nearly "punched through" the WebMCP benchmark: All 49/49 tasks completed. Model cost is approximat…

X AI KOLs Following ↗ · 2026-09-20 Cached

Jev + Mercury 2.5 AI model completed all tasks on the WebMCP benchmark with significantly lower cost compared to GPT-6 Astra variants.

0 favorites 0 likes
#benchmark

@_avichawla: https://x.com/_avichawla/status/2101563610644496464

X AI KOLs Timeline ↗ · 2026-09-20 Cached

The article explains how to build a local decision engine using open-source LLMs and SGLang, enabling efficient scoring and probability distributions for fixed choices without full text generation, compared to systems like Jev.

0 favorites 0 likes
#benchmark

I gave Jev, Laya, finetuned ModernCE and Qwen3.5 the controls to Doom

Reddit r/LocalLLaMA ↗ · 2026-09-20

An experiment where AI models including Jev, Laya, finetuned ModernCE-base-nli, and Qwen3.5-4B are given control in the game Doom, with performance metrics measured for kill counts and survival time in different scenarios.

0 favorites 0 likes
#benchmark

OmniEcho: Spatial Audio Understanding for Embodied Agents

Hugging Face Daily Papers ↗ · 2026-09-20 Cached

OmniEcho introduces a spatially aware omni-modal model and a new benchmark for spatial audio-visual perception and navigation in embodied agents, achieving state-of-the-art performance.

0 favorites 0 likes
#benchmark

PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing

Hugging Face Daily Papers ↗ · 2026-09-20 Cached

PackLab introduces a comprehensive framework for robotic bin packing using multimodal large language models, including a simulation platform, a specialized model, and a benchmark that outperforms traditional methods.

0 favorites 0 likes
#benchmark

@rohanpaul_ai: A new benchmark called JevBench just dropped. for models whose output is a bounded software decision rather than open-e…

X AI KOLs Timeline ↗ · 2026-09-19 Cached

JevBench is a new benchmark that evaluates AI models on bounded software decisions by combining intelligence, calibration, speed, and cost, as announced by @rohanpaul_ai.

0 favorites 0 likes
#benchmark

@dair_ai: Banger paper from MIT and Sakana AI. They show that self-improving coding agents work. The best part is that their appr…

X AI KOLs Timeline ↗ · 2026-09-19 Cached

The paper introduces Self-Improvement via Fast Tree-search (SIFT), a framework that uses an LLM-as-a-judge to efficiently evaluate self-modifications in coding agents, achieving better benchmark performance with significantly reduced CPU hours and API costs.

0 favorites 0 likes
#benchmark

@lateinteraction: incidentally and on a more serious note, @dianetc_ and i have wondered for some time if RL for reasoning followed by a …

X AI KOLs Following ↗ · 2026-09-19 Cached

The article discusses a paper titled 'Reasoning-Intensive Regression' that proposes MENTAT, a lightweight method combining batch-reflective prompt optimization with neural ensemble learning to improve numerical score prediction from text in AI tasks, showing up to 65% improvement over baselines.

0 favorites 0 likes
#benchmark

@paul_cal: Did a sceptical deep dive on some of the claimed issues w Humanity's Last Exam and... yep, all q's I looked at are defi…

X AI KOLs Timeline ↗ · 2026-09-19 Cached

A skeptical deep dive finds numerous errors in the Humanity's Last Exam benchmark, with the official o3-mini grader incorrectly marking correct answers as wrong.

0 favorites 0 likes
#benchmark

Apple M6 Pro Achieves the Highest Single-Core CPU Score in Geekbench 7

Hacker News Top ↗ · 2026-09-19

Apple's M6 Pro chip has achieved the highest single-core CPU score in the Geekbench 7 benchmark, demonstrating its leading performance capabilities.

0 favorites 0 likes
#benchmark

@BohuTANG: I'm doing something similar too. Jev is still too slow in this kind of scenario.

X AI KOLs Timeline ↗ · 2026-09-19 Cached

A discussion about using Jev for query optimization, with one user sharing a 12% speed improvement in Postgres queries and another commenting on Jev's performance issues.

0 favorites 0 likes
#benchmark

@googledevs: See how AI models can help you with multi-day engineering workflows with Android Bench 2.0. The updated benchmark evalu…

X AI KOLs Following ↗ · 2026-09-18 Cached

Android Bench 2.0 is an updated AI evaluation framework that assesses models on long-horizon engineering tasks for Android, such as building apps from scratch and migrating codebases, with continuous scoring.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback