benchmark

Tag

Cards List
#benchmark

enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think

Reddit r/LocalLLaMA · 8h ago

Enabling PCI-E peer-to-peer (P2P) for consumer Nvidia GPUs with patched drivers and vLLM environment variables yields roughly 25% prefill throughput improvement for free, as demonstrated by benchmarks.

0 favorites 0 likes
#benchmark

@jakevin7: https://x.com/jakevin7/status/2086031167040426488

X AI KOLs Timeline · 20h ago Cached

Maka is an open-source Agent Harness. Through mechanisms such as log-as-runtime, context pruning, and thinking feedback, it cuts the cost of the same DeepSeek task to 1/8 of OpenCode, while achieving a higher pass rate on Terminal-Bench at lower cost.

0 favorites 0 likes
#benchmark

@rohanpaul_ai: Grok Imagine Image 2.0 (Low) from @SpaceXAI jumped 12 places to #2, beating its own older quality model. The previous i…

X AI KOLs Following · 23h ago Cached

Grok Imagine Image 2.0 (Low) from xAI jumped to #2 in the Text-to-Image Arena, beating its own older quality model and showing significant improvement.

0 favorites 0 likes
#benchmark

Qwen 35B-A3B MoE vs 27B dense in local coding tests: ~4× faster, much smaller quality gap than I expected

Reddit r/LocalLLaMA · yesterday

A local experiment comparing Qwen 35B-A3B MoE and Qwen 27B dense on coding-maintenance tasks, finding the MoE model ~3.9× faster with a smaller quality gap than expected.

0 favorites 0 likes
#benchmark

@rohanpaul_ai: Meta's Muse Spark 1.2 reaches #4 while moving Text Arena's cost-quality Pareto frontier upward. you surrender 9 points …

X AI KOLs Following · yesterday Cached

Meta's Muse Spark 1.2 reaches #4 in the Text Arena, moving the cost-quality Pareto frontier upward with a ~91% price cut while sacrificing only 9 points relative to the top model.

0 favorites 0 likes
#benchmark

@TheAhmadOsman: Some numbers from running DeepSeek V4 Flash 0731 on a DGX Station

X AI KOLs Timeline · yesterday Cached

Ahmad Osman shares performance numbers from running DeepSeek V4 Flash 0731 on an NVIDIA DGX Station.

0 favorites 0 likes
#benchmark

@DeFiMinty: Can AI systems keep generating harder training tasks for AI agents? New research from Tencent suggests they can. Recurs…

X AI KOLs Following · yesterday Cached

New research from Tencent introduces Recursive Synthetic Terminal Tasks (RST), a method that progressively generates harder, verifiable training tasks for AI agents. Starting from 639 tasks, it produced 37,484 verified tasks across 15 rounds, and reinforcement learning with these tasks improved Qwen3.5-27B from 22.7% to 32.0% on Terminal-Bench Hard.

0 favorites 0 likes
#benchmark

@samsja19: great benchmark

X AI KOLs Following · yesterday Cached

Introducing VGI-Bench, a multimodal benchmark that probes 12 distinct visual and audio-visual skills with 550 human-curated questions to expose failures in today's video benchmarks.

0 favorites 0 likes
#benchmark

DeepSeek V4 Flash 0731

Hacker News Top · yesterday Cached

DeepSeek V4 Flash 0731 presents its results on the ARC-AGI benchmark, highlighting progress in abstract reasoning for AI models.

0 favorites 0 likes
#benchmark

TutorMoments: Do AI tutors know when to help and when to hold back?

Hugging Face Blog · yesterday Cached

Ai2 introduces TutorMoments, a replay-based evaluation framework and dataset for measuring whether LLMs can balance when to help and when to hold back in one-on-one math tutoring. Preliminary results show models tend to over-help, and prompt engineering only partially closes the gap to human tutors.

0 favorites 0 likes
#benchmark

I'm excited for Intel after testing the XPS 13

Jeff Geerling · yesterday Cached

Jeff Geerling reviews the Dell XPS 13 with Intel Core 5 320, highlighting its impressive power efficiency and Linux compatibility, and expresses optimism about Intel's low-end processor direction.

0 favorites 0 likes
#benchmark

@joelniklaus: Codex is overoptimised for large models: it ranks 2nd out of 10 for GLM 5.2 but drops to 9th place for Gemma-4! Almost …

X AI KOLs Timeline · yesterday Cached

A study comparing 10 coding agent harnesses on SWE-bench Pro shows harness choice dramatically affects pass@1 and cost, revealing vendor harnesses are overfitted to large models and do not transfer to smaller ones.

0 favorites 0 likes
#benchmark

@VraserX: Seedance 2.5 is finally here. I put it head-to-head against Seedance 2.0 using the exact same prompts across four compl…

X AI KOLs Following · yesterday Cached

A user shares a head-to-head comparison of Seedance 2.5 versus Seedance 2.0 across four cinematic genres, noting significant differences in output quality.

0 favorites 0 likes
#benchmark

LabyrinthBench: a local-focused, judge-free LLM benchmark that measures context recall under interference for multi-step agentic tasks.

Reddit r/LocalLLaMA · yesterday

LabyrinthBench is a local, judge-free LLM benchmark that deterministically measures context recall under interference for multi-step agentic tasks. Initial experiments show that wiping context and re-injecting gate answers helped 7 of 9 models but backfired on 2, highlighting the complexity of context management strategies.

0 favorites 0 likes
#benchmark

@reach_vb: luna-maxxing @ 80% lower costs!

X AI KOLs Timeline · yesterday Cached

ARC Prize re-tested OpenAI's GPT-5.6 Luna on ARC-AGI after an 80% price cut, confirming similar performance at a much lower cost per task.

0 favorites 0 likes
#benchmark

A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies

arXiv cs.LG · 2d ago Cached

This paper proposes a unified definition of uncertainty as pointwise posterior risk and introduces a theory-backed benchmark using semi-synthetic datasets to directly compute oracle epistemic and aleatoric uncertainty, enabling fine-grained evaluation beyond proxy tasks.

0 favorites 0 likes
#benchmark

FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India

arXiv cs.CL · 2d ago Cached

Presents FormBharo, a voice agent that uses LLMs with rule-based controls to fill structured forms over phone calls for low-literacy Hindi-speaking users in India, piloted with ARMMAN. The paper also introduces FormVoiceAgentBench, a benchmark of 3,760 multi-turn conversation tests, and shows that end-to-end evaluation is necessary since component-level performance does not predict full form completion.

0 favorites 0 likes
#benchmark

EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?

arXiv cs.CL · 2d ago Cached

EpiBench is a new closed-book, sequence-based benchmark for evaluating how well LLMs understand epitopes across five antibody-drug-discovery tasks, finding that current models capture partial signals but struggle with antibody-specific reasoning.

0 favorites 0 likes
#benchmark

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

arXiv cs.CL · 2d ago Cached

This paper introduces MameLoshnLM, the first open-source 8B-parameter Yiddish language model, along with the Oytser pretraining corpus and Kashes evaluation benchmark. It demonstrates that continued pretraining on high-quality Yiddish data outperforms general multilingual models, highlighting the value of dedicated low-resource language modeling.

0 favorites 0 likes
#benchmark

MoCA: Implicit Social Context Analysis

arXiv cs.CL · 2d ago Cached

Introduces MoCA (Implicit Social Context Analysis), a new task and benchmark for modeling implicit social scenarios across affection, intent, and stance, along with a Conflict-Driven Abductive Reasoning (CoDAR) framework. Experiments show state-of-the-art multimodal LLMs struggle on this task, while CoDAR improves performance but still lags behind human reasoning.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback