benchmark

Tag

Cards List
#benchmark

MoCA: Implicit Social Context Analysis

arXiv cs.CL · 2d ago Cached

Introduces MoCA (Implicit Social Context Analysis), a new task and benchmark for modeling implicit social scenarios across affection, intent, and stance, along with a Conflict-Driven Abductive Reasoning (CoDAR) framework. Experiments show state-of-the-art multimodal LLMs struggle on this task, while CoDAR improves performance but still lags behind human reasoning.

0 favorites 0 likes
#benchmark

M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

arXiv cs.CL · 2d ago Cached

This paper introduces M3R-Bench, a unified evidence-grounded benchmark for multimodal metaphor understanding with 1,000 image-text instances, and proposes M3R-Reasoner, an 8B-parameter model combining curriculum-based reasoning supervision and reinforcement learning that outperforms larger proprietary models.

0 favorites 0 likes
#benchmark

DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory Graph

arXiv cs.CL · 2d ago Cached

This paper introduces DREAM, a structured memory framework for LLM-based role-playing agents that uses an Event-aware Memory Graph to maintain temporal and causal coherence, and proposes the TCM benchmark for evaluation.

0 favorites 0 likes
#benchmark

Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning

arXiv cs.CL · 2d ago Cached

This paper introduces a three-condition experimental framework and a benchmark of 24,300 prompts to study how biased user turns modulate cognitive bias expression in frontier LLMs under multi-turn interactions. It finds that biased conversational context amplifies bias in most models, while explicit bias cues can trigger alignment-related suppression.

0 favorites 0 likes
#benchmark

PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

arXiv cs.CL · 2d ago Cached

Introduces PoolBench, a benchmark that isolates pooling strategies in decoder-only LLM concept representation, evaluating 19 strategies across 17 concepts and 3 models to provide a standardized protocol for comparing pooling choices.

0 favorites 0 likes
#benchmark

MS-MLB: An Open Machine Learning Benchmark for Blood-Based MS Classification

arXiv cs.LG · 2d ago Cached

This paper introduces MS-MLB, an open machine learning benchmark for classifying multiple sclerosis from whole blood RNA expression data using the public GSE17048 cohort. It provides a reproducible, leakage-controlled evaluation pipeline and reports Gradient Boosting as the top performer.

0 favorites 0 likes
#benchmark

Unified Agent: Managing Interactions across Devices

arXiv cs.AI · 2d ago Cached

This paper introduces Unified Agent, a stateful AI agent that maintains compact interaction state across devices and time to handle cross-device, cross-time user requests, outperforming existing agent designs in a new benchmark.

0 favorites 0 likes
#benchmark

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

arXiv cs.AI · 2d ago Cached

Introduces SkillTV-Bench, a 681-case benchmark for evaluating skill-aware trajectory verification in LLM agents, along with SkillTV-Evolve, a method that externalizes verification knowledge as a reusable JudgeSkill, improving judge accuracy by 14.8 percentage points.

0 favorites 0 likes
#benchmark

EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents

arXiv cs.AI · 2d ago Cached

This paper introduces EcoAgent-Bench, a 304-task benchmark for evaluating LLM agents' economic decision-making under explicit budgets and priced actions, testing four cost-related decisions across five task families.

0 favorites 0 likes
#benchmark

C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

arXiv cs.AI · 2d ago Cached

Introduces C3PO, a benchmark of 3,404 samples for evaluating cross-modal composition and counterfactual reasoning in multimodal LLMs. It finds modality dominance causes most failures, with even the best model (Gemini-3.1-Pro) far below human accuracy.

0 favorites 0 likes
#benchmark

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

arXiv cs.AI · 2d ago Cached

This paper introduces LUNAR, a benchmark for evaluating how large language models personalize responses from longitudinal app interaction histories across daily-life domains such as clothing, food, housing, and mobility. Experiments on 19 mainstream LLMs reveal that effective personalization depends on evidence selection and cross-domain integration, and that stronger personalization can come at the cost of privacy protection.

0 favorites 0 likes
#benchmark

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

arXiv cs.AI · 2d ago Cached

This paper introduces SearchAuditBench, a benchmark of 1,243 failed long-horizon search-agent trajectories with expert annotations, and SearchAuditor, a multi-perspective auditing framework that localizes, attributes, and repairs agent failures. Experiments show SearchAuditor outperforms baselines, achieving a 32.3% end-to-end pass rate with frontier models like GPT-5.5.

0 favorites 0 likes
#benchmark

@tobi: Let's all agree that this is the correct and final eval for agi https://jackhopkins.github.io/factorio-learning-environ…

X AI KOLs Timeline · 2d ago Cached

Factorio Learning Environment v0.3.0 is an open-source platform for evaluating AI agents in Factorio, adding headless scaling, OpenAI Gym compatibility, and Claude Code integration for live demonstrations.

0 favorites 0 likes
#benchmark

Qwen3.8 Max now ranked as the best overall model by agentic index

Hacker News Top · 2d ago Cached

Qwen3.8 Max is now ranked as the best overall model on Artificial Analysis's agentic index, surpassing other leading AI models in independent evaluations.

0 favorites 0 likes
#benchmark

I measured what 13 search APIs actually cost to run inside an agent. The pricing page is the smaller half of the bill

Reddit r/AI_Agents · 2d ago

A developer benchmarks 13 search API configurations inside an AI agent, revealing that hidden token costs from reading payloads can dominate the total bill and vary by up to 67x across providers.

0 favorites 0 likes
#benchmark

2 x 5070ti Qwen 27B full config / stats

Reddit r/LocalLLaMA · 2d ago

Technical post sharing performance stats for running Qwen 27B on 2x RTX 5070 Ti GPUs with vLLM cu129-nightly, achieving up to 94-87 tps decode and 170k GPU KV cache.

0 favorites 0 likes
#benchmark

I compared even more parsers on 14 PDF-parsing capabilities using different types

Reddit r/LocalLLaMA · 2d ago

A benchmark comparing 8 PDF parsers across 14 capabilities, finding Chandra the most accurate while noting trade-offs like speed; LightOnOCR-1B impresses for its size but hallucinates on illegible text.

0 favorites 0 likes
#benchmark

32 total local models tested head to head

Reddit r/LocalLLaMA · 2d ago

A head-to-head evaluation of 32 local language models on a fact-extraction corpus finds that most models are statistically indistinguishable, with LFM2.5 models performing significantly worse despite larger sizes.

0 favorites 0 likes
#benchmark

How come artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCode

Reddit r/LocalLLaMA · 2d ago

A discussion about how Artificial Analysis ranks Gemma 4 above Qwen3.6 27b on the SciCode benchmark, questioning whether the ranking reflects real-world coding ability or reveals a benchmarking issue.

0 favorites 0 likes
#benchmark

Auto-fit vs tuned MoE offload: 564 → 1330 pp tok/s, unchanged decode (Qwen3.6-35B-A3B Q6 / RTX 3090)

Reddit r/LocalLLaMA · 2d ago

A developer benchmarks Qwen3.6-35B-A3B Q6 on an RTX 3090, showing that offloading eight MoE expert layers to CPU and increasing batch sizes improves prompt processing by 2.36× (564→1330 tok/s) with no decode speed regression, using evolutionary search to find the tuning config.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback