benchmark

Tag

Cards List
#benchmark

We got 100% on ARC-AGI-3 ft09 with zero model calls. The failures are more interesting.

Reddit r/artificial · 1h ago

An experimental reasoning system at Orivael scored 100% on ARC-AGI-3 ft09 with zero model calls, revealing that its failures stem from incorrect environment representations rather than planning errors.

0 favorites 0 likes
#benchmark

Updated benchmark: Deepseek V4 Flash on SlopCodeBench (local)

Reddit r/LocalLLaMA · 14h ago

A user shares updated benchmark results for DeepSeek V4 Flash on SlopCodeBench using local quants (antirez imatrix quant) with the pi harness, showing improved performance over previous runs but still slower than the hosted API.

0 favorites 0 likes
#benchmark

enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think

Reddit r/LocalLLaMA · 23h ago

Enabling PCI-E peer-to-peer (P2P) for consumer Nvidia GPUs with patched drivers and vLLM environment variables yields roughly 25% prefill throughput improvement for free, as demonstrated by benchmarks.

0 favorites 0 likes
#benchmark

@jakevin7: https://x.com/jakevin7/status/2086031167040426488

X AI KOLs Timeline · yesterday Cached

Maka is an open-source Agent Harness. Through mechanisms such as log-as-runtime, context pruning, and thinking feedback, it cuts the cost of the same DeepSeek task to 1/8 of OpenCode, while achieving a higher pass rate on Terminal-Bench at lower cost.

0 favorites 0 likes
#benchmark

@rohanpaul_ai: Grok Imagine Image 2.0 (Low) from @SpaceXAI jumped 12 places to #2, beating its own older quality model. The previous i…

X AI KOLs Following · yesterday Cached

Grok Imagine Image 2.0 (Low) from xAI jumped to #2 in the Text-to-Image Arena, beating its own older quality model and showing significant improvement.

0 favorites 0 likes
#benchmark

Qwen 35B-A3B MoE vs 27B dense in local coding tests: ~4× faster, much smaller quality gap than I expected

Reddit r/LocalLLaMA · yesterday

A local experiment comparing Qwen 35B-A3B MoE and Qwen 27B dense on coding-maintenance tasks, finding the MoE model ~3.9× faster with a smaller quality gap than expected.

0 favorites 0 likes
#benchmark

@rohanpaul_ai: Meta's Muse Spark 1.2 reaches #4 while moving Text Arena's cost-quality Pareto frontier upward. you surrender 9 points …

X AI KOLs Following · yesterday Cached

Meta's Muse Spark 1.2 reaches #4 in the Text Arena, moving the cost-quality Pareto frontier upward with a ~91% price cut while sacrificing only 9 points relative to the top model.

0 favorites 0 likes
#benchmark

@TheAhmadOsman: Some numbers from running DeepSeek V4 Flash 0731 on a DGX Station

X AI KOLs Timeline · yesterday Cached

Ahmad Osman shares performance numbers from running DeepSeek V4 Flash 0731 on an NVIDIA DGX Station.

0 favorites 0 likes
#benchmark

@DeFiMinty: Can AI systems keep generating harder training tasks for AI agents? New research from Tencent suggests they can. Recurs…

X AI KOLs Following · yesterday Cached

New research from Tencent introduces Recursive Synthetic Terminal Tasks (RST), a method that progressively generates harder, verifiable training tasks for AI agents. Starting from 639 tasks, it produced 37,484 verified tasks across 15 rounds, and reinforcement learning with these tasks improved Qwen3.5-27B from 22.7% to 32.0% on Terminal-Bench Hard.

0 favorites 0 likes
#benchmark

@samsja19: great benchmark

X AI KOLs Following · yesterday Cached

Introducing VGI-Bench, a multimodal benchmark that probes 12 distinct visual and audio-visual skills with 550 human-curated questions to expose failures in today's video benchmarks.

0 favorites 0 likes
#benchmark

DeepSeek V4 Flash 0731

Hacker News Top · 2d ago Cached

DeepSeek V4 Flash 0731 presents its results on the ARC-AGI benchmark, highlighting progress in abstract reasoning for AI models.

0 favorites 0 likes
#benchmark

TutorMoments: Do AI tutors know when to help and when to hold back?

Hugging Face Blog · 2d ago Cached

Ai2 introduces TutorMoments, a replay-based evaluation framework and dataset for measuring whether LLMs can balance when to help and when to hold back in one-on-one math tutoring. Preliminary results show models tend to over-help, and prompt engineering only partially closes the gap to human tutors.

0 favorites 0 likes
#benchmark

I'm excited for Intel after testing the XPS 13

Jeff Geerling · 2d ago Cached

Jeff Geerling reviews the Dell XPS 13 with Intel Core 5 320, highlighting its impressive power efficiency and Linux compatibility, and expresses optimism about Intel's low-end processor direction.

0 favorites 0 likes
#benchmark

@joelniklaus: Codex is overoptimised for large models: it ranks 2nd out of 10 for GLM 5.2 but drops to 9th place for Gemma-4! Almost …

X AI KOLs Timeline · 2d ago Cached

A study comparing 10 coding agent harnesses on SWE-bench Pro shows harness choice dramatically affects pass@1 and cost, revealing vendor harnesses are overfitted to large models and do not transfer to smaller ones.

0 favorites 0 likes
#benchmark

@VraserX: Seedance 2.5 is finally here. I put it head-to-head against Seedance 2.0 using the exact same prompts across four compl…

X AI KOLs Following · 2d ago Cached

A user shares a head-to-head comparison of Seedance 2.5 versus Seedance 2.0 across four cinematic genres, noting significant differences in output quality.

0 favorites 0 likes
#benchmark

LabyrinthBench: a local-focused, judge-free LLM benchmark that measures context recall under interference for multi-step agentic tasks.

Reddit r/LocalLLaMA · 2d ago

LabyrinthBench is a local, judge-free LLM benchmark that deterministically measures context recall under interference for multi-step agentic tasks. Initial experiments show that wiping context and re-injecting gate answers helped 7 of 9 models but backfired on 2, highlighting the complexity of context management strategies.

0 favorites 0 likes
#benchmark

@reach_vb: luna-maxxing @ 80% lower costs!

X AI KOLs Timeline · 2d ago Cached

ARC Prize re-tested OpenAI's GPT-5.6 Luna on ARC-AGI after an 80% price cut, confirming similar performance at a much lower cost per task.

0 favorites 0 likes
#benchmark

A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies

arXiv cs.LG · 2d ago Cached

This paper proposes a unified definition of uncertainty as pointwise posterior risk and introduces a theory-backed benchmark using semi-synthetic datasets to directly compute oracle epistemic and aleatoric uncertainty, enabling fine-grained evaluation beyond proxy tasks.

0 favorites 0 likes
#benchmark

FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India

arXiv cs.CL · 2d ago Cached

Presents FormBharo, a voice agent that uses LLMs with rule-based controls to fill structured forms over phone calls for low-literacy Hindi-speaking users in India, piloted with ARMMAN. The paper also introduces FormVoiceAgentBench, a benchmark of 3,760 multi-turn conversation tests, and shows that end-to-end evaluation is necessary since component-level performance does not predict full form completion.

0 favorites 0 likes
#benchmark

EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?

arXiv cs.CL · 2d ago Cached

EpiBench is a new closed-book, sequence-based benchmark for evaluating how well LLMs understand epitopes across five antibody-drug-discovery tasks, finding that current models capture partial signals but struggle with antibody-specific reasoning.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback