benchmark

Tag

Cards List
#benchmark

An open-source alternative to Jev

Reddit r/artificial ↗ · 1h ago

jevos is an open-source tool that processes text with yes/no questions in a single forward pass on a laptop CPU, with performance benchmarks compared to Jev and Lay.

0 favorites 0 likes
#benchmark

@cline: Pixel Canary (stealth model) has just released and is free in Cline. It's tied with GPT-6 Astra and beats Kimi K3 on Ne…

X AI KOLs Timeline ↗ · 15h ago Cached

Pixel Canary, a stealth AI model, has been released for free in Cline. It matches the performance of GPT-6 Astra and outperforms Kimi K3 on the Next.js Agent Evals benchmark for web and mobile development tasks.

0 favorites 0 likes
#benchmark

@BenKoska: We're #1 on NVIDIA's SOL-ExecBench across all four tracks. Across 235 real production kernels run on B200s we beat out …

X AI KOLs Timeline ↗ · 16h ago Cached

A team announced they achieved first place on NVIDIA's SOL-ExecBench benchmark across all tracks, beating 583 teams with real production kernels on B200 GPUs.

0 favorites 0 likes
#benchmark

Kev 4B topped out in every Tetris game I ran. Mica v0.1 4B cleared about 4x more lines and survived two of them to the end

Reddit r/LocalLLaMA ↗ · 17h ago

Mica v0.1 4B outperforms Kev 4B in Tetris games by clearing about 4x more lines and surviving to the end in some cases, using only board descriptions without search or lookahead.

0 favorites 0 likes
#benchmark

@stanfordnlp: For people who came of age in the 1980s playing Hack (and Rogue before it), before the age of modern real-time strategy…

X AI KOLs Timeline ↗ · 21h ago Cached

An LLM agent named GPT 6 Astra has achieved ascension in the game NetHack, marking a first recorded instance of an LLM beating the complex game through autonomous tool-building.

0 favorites 0 likes
#benchmark

@VraserX: Xiaomi's MiMo-V2.6-Pro has MIT-licensed weights and is listed on OpenRouter at $0.87 per million output tokens. It does…

X AI KOLs Timeline ↗ · 21h ago Cached

Xiaomi's MiMo-V2.6-Pro, an AI model with MIT-licensed weights, is listed on OpenRouter at $0.87 per million output tokens, highlighting its cost-effectiveness compared to expensive frontier models.

0 favorites 0 likes
#benchmark

Jev vs. Kev: open-source Jev alternative tested side by side

Reddit r/LocalLLaMA ↗ · 22h ago Cached

The article compares TypeSafe's Jev model with the open-source Kev alternative, testing their accuracy, token usage, and speed on a new dataset to avoid data leakage, finding similar performance but differences in implementation.

0 favorites 0 likes
#benchmark

@omarsar0: Recommended benchmark. I expect voice to become one of the main ways people interact with robots and physical AI. That …

X AI KOLs Timeline ↗ · yesterday Cached

BRIDGE ASR 2.0 is a benchmark that evaluates speech recognition models on real two-person conversations in 21 languages, focusing on code-switching and multilingual performance to enhance voice interaction in AI and robotics.

0 favorites 0 likes
#benchmark

AGI (/j) Gemini 4 pro's NEW checkpoint (reupload cause I messed up image last time)

Reddit r/singularity ↗ · yesterday

A new checkpoint for Gemini 4 pro is announced, showing detailed outputs in 6 minutes and outperforming Gemini 3.8 flash on ArenaAI.

0 favorites 0 likes
#benchmark

The Last Human Gate: Forward Deployed Engineering for Governance Automation

arXiv cs.AI ↗ · yesterday Cached

This paper presents a task-substitution framework for automating enterprise governance reviews using AI agents and software, with benchmarks like DGF-Bench showing that models such as Gemini 3.8 Flash can achieve high success rates in replacing human execution for specified review tasks.

0 favorites 0 likes
#benchmark

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

arXiv cs.AI ↗ · yesterday Cached

Introduces IndicBankBench, a 799-case benchmark for evaluating safety and reliability of language model assistants in Indian retail banking, with multi-stage evaluation and public release of code and data.

0 favorites 0 likes
#benchmark

RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

arXiv cs.AI ↗ · yesterday Cached

RECLAIM is a benchmark for AI agents to reproduce claims from machine learning papers, with difficulty tiers based on available code, data, and weights. It shows that agents struggle to reproduce results, with best success rates of 41% in the easiest tier.

0 favorites 0 likes
#benchmark

TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment

arXiv cs.AI ↗ · yesterday Cached

TWIST is a proposed benchmark suite for evaluating intervention quality in conversational memory systems, featuring human-validated tracks to detect failures like unresolved tensions and stale facts, revealing trade-offs between recall and specificity that traditional metrics miss.

0 favorites 0 likes
#benchmark

EnSiTa - A Trilingual Multi-Domain Parallel Dataset and Benchmark for Domain-Specific Machine Translation

arXiv cs.CL ↗ · yesterday Cached

EnSiTa is a trilingual multi-domain parallel dataset and benchmark for English, Sinhala, and Tamil, featuring human post-edited training data and extensive experiments on domain-specific machine translation to address low-resource language challenges.

0 favorites 0 likes
#benchmark

BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech

arXiv cs.CL ↗ · yesterday Cached

The paper introduces BanglaTurn, a benchmark corpus and Whisper-based model for end-of-turn detection in Bangla speech, achieving 84.33% accuracy compared to a 69.28% baseline.

0 favorites 0 likes
#benchmark

COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages

arXiv cs.CL ↗ · yesterday Cached

This paper presents COILD, an Indic-centric parallel corpus with over 1.16 million sentence pairs for 20 Indian language pairs and a domain-centric benchmark, demonstrating improvements in machine translation when fine-tuning multilingual models.

0 favorites 0 likes
#benchmark

A finance benchmark asks agents to finish the whole assignment (18 minute read)

TLDR AI ↗ · yesterday Cached

A finance benchmark named DAYJOB by Surge AI evaluates AI agents on completing financial forecasting tasks, with detailed criteria for pass/fail responses focusing on errors in net sales calculations and revenue growth projections.

0 favorites 0 likes
#benchmark

Dailychained PLX 88096 switches, Quad RTX 5070 Ti + Quad RTX 5060 Ti

Reddit r/LocalLLaMA ↗ · yesterday

A user shares their experience building a multi-GPU system with RTX 5070 Ti and 5060 Ti for running AI models like Qwen3.8-27B-FP8 using vLLM on Linux, detailing hardware setup, benchmarks, and challenges.

0 favorites 0 likes
#benchmark

@svpino: Cutting down noise before sending the audio to a speech-to-text model makes a huge improvement. Voice isolation is the …

X AI KOLs Timeline ↗ · yesterday Cached

Krisp released an open benchmark and dataset showing that voice isolation reduces word error rates in speech-to-text models by 73%, with significant improvements across workplace and call-center recordings.

0 favorites 0 likes
#benchmark

Claude Opus 5.5 tops SimpleBench with its 88.4% score.

Reddit r/singularity ↗ · yesterday

Claude Opus 5.5 achieves a top score of 88.4% on the SimpleBench benchmark, indicating significant performance in AI evaluation.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback