benchmark

Tag

Cards List
#benchmark

Splash fork optimised for M5 Max: ~1.5× faster (1.25× single request)

Reddit r/LocalLLaMA ↗ · 7h ago Cached

Splish is an unofficial fork of Splash that optimizes Metal kernels for Apple M5 Max chips, delivering up to 1.5× faster AI inference speeds for models like Qwen3.8-27B without compromising quality.

0 favorites 0 likes
#benchmark

@yibie: https://x.com/yibie/status/2104076142500069608

X AI KOLs Timeline ↗ · 8h ago Cached

This article demonstrates how to convert GLM-5.3-Flash into a Jev-style System 1 decision model, achieving typed decisions through a single forward pass. Benchmark tests show that it performs comparably to specialized models in terms of accuracy and speed.

0 favorites 0 likes
#benchmark

Tauon: A new optimizer outperforming Muon on GPT-Mini (lower loss, ~8.5% faster step time) [P]

Reddit r/MachineLearning ↗ · 10h ago

Tauon is a new optimizer using polynomial and orthogonalization techniques that outperforms Muon and AdamW in initial benchmarks on a small GPT-Mini model, showing lower loss and faster step time.

0 favorites 0 likes
#benchmark

@BenjaminDEKR: BenchBench: multimodal LLMs compete to assemble Ikea furniture

X AI KOLs Following ↗ · 13h ago

BenchBench is a benchmark that challenges multimodal LLMs to compete in assembling Ikea furniture, testing their practical AI capabilities.

0 favorites 0 likes
#benchmark

Sonnet 5.5, Which Already Supposedly Beats GPT-6 Sol, Has Had a Last-Minute Upgrade With Release Expected Monday

Reddit r/singularity ↗ · 16h ago

Sonnet 5.5, which is claimed to outperform GPT-6 Sol, has received a last-minute upgrade with a release expected on Monday.

0 favorites 0 likes
#benchmark

VSArena — an open benchmark for AI agents in 3D embodied environments

Reddit r/ArtificialInteligence ↗ · 21h ago

VSArena is an open benchmark for evaluating AI agents in interactive 3D environments, offering remote execution and a public leaderboard to assess perception, reasoning, and action without physical robots.

0 favorites 0 likes
#benchmark

Turning GLM-5.3-Flash into a Jev-like decision model

Hacker News Top ↗ · 22h ago Cached

This article demonstrates how to adapt the GLM-5.3-Flash LLM to function as a typed decision model similar to Jev, achieving comparable accuracy and speed while enabling decisions on images in a single forward pass.

0 favorites 0 likes
#benchmark

An open-source alternative to Jev

Reddit r/artificial ↗ · 23h ago

jevos is an open-source tool that processes text with yes/no questions in a single forward pass on a laptop CPU, with performance benchmarks compared to Jev and Lay.

0 favorites 0 likes
#benchmark

@cline: Pixel Canary (stealth model) has just released and is free in Cline. It's tied with GPT-6 Astra and beats Kimi K3 on Ne…

X AI KOLs Timeline ↗ · yesterday Cached

Pixel Canary, a stealth AI model, has been released for free in Cline. It matches the performance of GPT-6 Astra and outperforms Kimi K3 on the Next.js Agent Evals benchmark for web and mobile development tasks.

0 favorites 0 likes
#benchmark

@BenKoska: We're #1 on NVIDIA's SOL-ExecBench across all four tracks. Across 235 real production kernels run on B200s we beat out …

X AI KOLs Timeline ↗ · yesterday Cached

A team announced they achieved first place on NVIDIA's SOL-ExecBench benchmark across all tracks, beating 583 teams with real production kernels on B200 GPUs.

0 favorites 0 likes
#benchmark

Kev 4B topped out in every Tetris game I ran. Mica v0.1 4B cleared about 4x more lines and survived two of them to the end

Reddit r/LocalLLaMA ↗ · yesterday

Mica v0.1 4B outperforms Kev 4B in Tetris games by clearing about 4x more lines and surviving to the end in some cases, using only board descriptions without search or lookahead.

0 favorites 0 likes
#benchmark

@stanfordnlp: For people who came of age in the 1980s playing Hack (and Rogue before it), before the age of modern real-time strategy…

X AI KOLs Timeline ↗ · yesterday Cached

An LLM agent named GPT 6 Astra has achieved ascension in the game NetHack, marking a first recorded instance of an LLM beating the complex game through autonomous tool-building.

0 favorites 0 likes
#benchmark

@VraserX: Xiaomi's MiMo-V2.6-Pro has MIT-licensed weights and is listed on OpenRouter at $0.87 per million output tokens. It does…

X AI KOLs Timeline ↗ · yesterday Cached

Xiaomi's MiMo-V2.6-Pro, an AI model with MIT-licensed weights, is listed on OpenRouter at $0.87 per million output tokens, highlighting its cost-effectiveness compared to expensive frontier models.

0 favorites 0 likes
#benchmark

Jev vs. Kev: open-source Jev alternative tested side by side

Reddit r/LocalLLaMA ↗ · yesterday Cached

The article compares TypeSafe's Jev model with the open-source Kev alternative, testing their accuracy, token usage, and speed on a new dataset to avoid data leakage, finding similar performance but differences in implementation.

0 favorites 0 likes
#benchmark

@omarsar0: Recommended benchmark. I expect voice to become one of the main ways people interact with robots and physical AI. That …

X AI KOLs Timeline ↗ · yesterday Cached

BRIDGE ASR 2.0 is a benchmark that evaluates speech recognition models on real two-person conversations in 21 languages, focusing on code-switching and multilingual performance to enhance voice interaction in AI and robotics.

0 favorites 0 likes
#benchmark

AGI (/j) Gemini 4 pro's NEW checkpoint (reupload cause I messed up image last time)

Reddit r/singularity ↗ · yesterday

A new checkpoint for Gemini 4 pro is announced, showing detailed outputs in 6 minutes and outperforming Gemini 3.8 flash on ArenaAI.

0 favorites 0 likes
#benchmark

The Last Human Gate: Forward Deployed Engineering for Governance Automation

arXiv cs.AI ↗ · 2d ago Cached

This paper presents a task-substitution framework for automating enterprise governance reviews using AI agents and software, with benchmarks like DGF-Bench showing that models such as Gemini 3.8 Flash can achieve high success rates in replacing human execution for specified review tasks.

0 favorites 0 likes
#benchmark

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

arXiv cs.AI ↗ · 2d ago Cached

Introduces IndicBankBench, a 799-case benchmark for evaluating safety and reliability of language model assistants in Indian retail banking, with multi-stage evaluation and public release of code and data.

0 favorites 0 likes
#benchmark

RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

arXiv cs.AI ↗ · 2d ago Cached

RECLAIM is a benchmark for AI agents to reproduce claims from machine learning papers, with difficulty tiers based on available code, data, and weights. It shows that agents struggle to reproduce results, with best success rates of 41% in the easiest tier.

0 favorites 0 likes
#benchmark

TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment

arXiv cs.AI ↗ · 2d ago Cached

TWIST is a proposed benchmark suite for evaluating intervention quality in conversational memory systems, featuring human-validated tracks to detect failures like unresolved tensions and stale facts, revealing trade-offs between recall and specificity that traditional metrics miss.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback