benchmark

Tag

Cards List
#benchmark

M5 ultra AI test results

Reddit r/LocalLLaMA · 5h ago

M5 ultra AI test results show prompt processing speeds up to 4-4.5 times faster and token generation 1.5x faster than M3 ultra, but with doubled power consumption, increased fan noise, and higher temperatures.

0 favorites 0 likes
#benchmark

LLM Ass Bench

Hacker News Top · 6h ago Cached

LLM Ass Bench is a benchmark or tool for evaluating Large Language Models, with a focus on prompts.

0 favorites 0 likes
#benchmark

Every Benchmark Comparison for Opus 5.5 and GPT-6 Astra/Sol

Reddit r/singularity · 7h ago

Opus 5.5 outperforms GPT-6 Astra and Sol in benchmark comparisons, but at a higher cost, as shown in the provided image.

0 favorites 0 likes
#benchmark

Opus 5.5 Cost vs Performance on Terminal-Bench 4.0

Reddit r/singularity · 10h ago

This article evaluates the cost-effectiveness and performance of the Opus 5.5 AI model on the Terminal-Bench 4.0 benchmark.

0 favorites 0 likes
#benchmark

Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index, along with a 20% price cut and larger cache hit discount

Reddit r/ArtificialInteligence · 10h ago

Claude Opus 5.5 has achieved the top position on the Artificial Analysis Intelligence Index, along with a 20% price reduction and enhanced cache hit discounts.

0 favorites 0 likes
#benchmark

LinearSolveBench: new benchmark for linear solvers [P]

Reddit r/MachineLearning · 11h ago

LinearSolveBench is a new benchmark for evaluating linear solvers in C, focusing on speed, accuracy, and generality for large sparse systems to promote algorithmic advances.

0 favorites 0 likes
#benchmark

Show HN: JevBench, a reproducible benchmark for typed decision models

Hacker News Top · 13h ago Cached

JevBench v1.3.0 is a reproducible benchmark for Jev-class decision models, evaluating and ranking 52 systems based on intelligence, calibration, speed, and cost.

0 favorites 0 likes
#benchmark

GPT-6 Astra vs. Anthropic’s new Mythos do the Pelican SVG test

Reddit r/singularity · 18h ago

A tweet compares GPT-6 Astra and Anthropic’s new Mythos model on the Pelican SVG test, with a link to an external source.

0 favorites 0 likes
#benchmark

MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis

Hacker News Top · 22h ago Cached

An analysis of the MiMo-v2.6-Pro AI model's intelligence, performance, and price using Artificial Analysis's benchmarks and indexes.

0 favorites 0 likes
#benchmark

Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark

arXiv cs.LG · 22h ago Cached

This study benchmarks coding agents on reproducing Eurostat statistics, finding that semantic validation and a retry budget are crucial for reliability, not just execution diagnostics.

0 favorites 0 likes
#benchmark

Dissecting Training-Free Uncertainty Estimation in Multimodal Large Language Models

arXiv cs.CL · 22h ago Cached

A systematic study benchmarking training-free uncertainty quantification strategies for multimodal Large Language Models, categorizing methods into token-level, verbalized, and semantic approaches and finding optimal strategies depend on response length.

0 favorites 0 likes
#benchmark

PII-TRACE: A Benchmark for Context-Aware PII Detection in Multi-Turn LLM Conversations

arXiv cs.CL · 22h ago Cached

Introduces PII-TRACE, the first benchmark for context-aware PII detection in multi-turn LLM conversations, and PII-Tracer, a compact detector that achieves high entity-level coverage.

0 favorites 0 likes
#benchmark

SCoR: A Hierarchical Framework for Forecasting Relations Between Scientific Concepts

arXiv cs.CL · 22h ago Cached

This paper introduces SCoR, a hierarchical framework for forecasting relations between scientific concepts, with a benchmark and model that improve research-direction discovery by predicting typed relations.

0 favorites 0 likes
#benchmark

Is Imagination Derived from Hallucination? A Cross-Taxonomy Evaluation of Imagination and Hallucination in Large Language Models

arXiv cs.CL · 22h ago Cached

The paper introduces Whiteboard, the first benchmark for evaluating imagination in large language models by cross-referencing it with hallucination, and reveals a counterintuitive negative correlation between the two across 79 state-of-the-art LLMs.

0 favorites 0 likes
#benchmark

The final benchmark

Reddit r/singularity · yesterday

This article presents a definitive benchmark designed to evaluate and compare AI models or systems, establishing a standard for future assessments.

0 favorites 0 likes
#benchmark

@devindesktop: Grok 4.7 is now available in Devin Desktop and Devin CLI!

X AI KOLs Following · yesterday Cached

Grok 4.7 AI model is now available in Devin Desktop and CLI, with evaluation results showing strong performance on hard backend engineering tasks.

0 favorites 0 likes
#benchmark

@browser_use: Grok 4.7 just dropped. Still chasing DeepSeek Long Horizon Browser Use Benchmark v2 > GPT-6 Astra: 80.6 > DeepSeek V4.1…

X AI KOLs Timeline · yesterday Cached

Grok 4.7 has been released and shows improved performance over Grok 4.6 on the Long Horizon Browser Use Benchmark v2, but still lags significantly behind DeepSeek V4.1 Flash and GPT-6 Astra.

0 favorites 0 likes
#benchmark

@rohanpaul_ai: Forward Deployed Engineers have a compounding problem: they spend months learning a company’s systems, and hidden depen…

X AI KOLs Timeline · yesterday Cached

Codos launches a virtual Chief AI Officer that addresses context loss in Forward Deployed Engineers by deploying company-wide memory and automation agents, with preliminary benchmark scores showing high performance on EnterpriseRAG-Bench.

0 favorites 0 likes
#benchmark

@omarsar0: StepFun’s new Step 5 Preview model is impressive! Had a chance to test it early. I've been testing it as a coding agent…

X AI KOLs Following · yesterday Cached

StepFun's new Step 5 Preview model is tested as a coding agent, demonstrating competitive performance with models like GLM 5.3 and excelling in long-horizon tasks due to its effective stopping behavior.

0 favorites 0 likes
#benchmark

Putting the question before the context took my local Qwen from 89% to 100% on a decision benchmark, and from ~400 ms to ~80 ms

Reddit r/LocalLLaMA · yesterday

Switching the order of question and context in prompts for local Qwen models improved accuracy from 89% to 100% and reduced latency from ~400 ms to ~80 ms on a decision benchmark.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback