Tag
jevos is an open-source tool that processes text with yes/no questions in a single forward pass on a laptop CPU, with performance benchmarks compared to Jev and Lay.
Pixel Canary, a stealth AI model, has been released for free in Cline. It matches the performance of GPT-6 Astra and outperforms Kimi K3 on the Next.js Agent Evals benchmark for web and mobile development tasks.
A team announced they achieved first place on NVIDIA's SOL-ExecBench benchmark across all tracks, beating 583 teams with real production kernels on B200 GPUs.
Mica v0.1 4B outperforms Kev 4B in Tetris games by clearing about 4x more lines and surviving to the end in some cases, using only board descriptions without search or lookahead.
An LLM agent named GPT 6 Astra has achieved ascension in the game NetHack, marking a first recorded instance of an LLM beating the complex game through autonomous tool-building.
Xiaomi's MiMo-V2.6-Pro, an AI model with MIT-licensed weights, is listed on OpenRouter at $0.87 per million output tokens, highlighting its cost-effectiveness compared to expensive frontier models.
The article compares TypeSafe's Jev model with the open-source Kev alternative, testing their accuracy, token usage, and speed on a new dataset to avoid data leakage, finding similar performance but differences in implementation.
BRIDGE ASR 2.0 is a benchmark that evaluates speech recognition models on real two-person conversations in 21 languages, focusing on code-switching and multilingual performance to enhance voice interaction in AI and robotics.
A new checkpoint for Gemini 4 pro is announced, showing detailed outputs in 6 minutes and outperforming Gemini 3.8 flash on ArenaAI.
This paper presents a task-substitution framework for automating enterprise governance reviews using AI agents and software, with benchmarks like DGF-Bench showing that models such as Gemini 3.8 Flash can achieve high success rates in replacing human execution for specified review tasks.
Introduces IndicBankBench, a 799-case benchmark for evaluating safety and reliability of language model assistants in Indian retail banking, with multi-stage evaluation and public release of code and data.
RECLAIM is a benchmark for AI agents to reproduce claims from machine learning papers, with difficulty tiers based on available code, data, and weights. It shows that agents struggle to reproduce results, with best success rates of 41% in the easiest tier.
TWIST is a proposed benchmark suite for evaluating intervention quality in conversational memory systems, featuring human-validated tracks to detect failures like unresolved tensions and stale facts, revealing trade-offs between recall and specificity that traditional metrics miss.
EnSiTa is a trilingual multi-domain parallel dataset and benchmark for English, Sinhala, and Tamil, featuring human post-edited training data and extensive experiments on domain-specific machine translation to address low-resource language challenges.
The paper introduces BanglaTurn, a benchmark corpus and Whisper-based model for end-of-turn detection in Bangla speech, achieving 84.33% accuracy compared to a 69.28% baseline.
This paper presents COILD, an Indic-centric parallel corpus with over 1.16 million sentence pairs for 20 Indian language pairs and a domain-centric benchmark, demonstrating improvements in machine translation when fine-tuning multilingual models.
A finance benchmark named DAYJOB by Surge AI evaluates AI agents on completing financial forecasting tasks, with detailed criteria for pass/fail responses focusing on errors in net sales calculations and revenue growth projections.
A user shares their experience building a multi-GPU system with RTX 5070 Ti and 5060 Ti for running AI models like Qwen3.8-27B-FP8 using vLLM on Linux, detailing hardware setup, benchmarks, and challenges.
Krisp released an open benchmark and dataset showing that voice isolation reduces word error rates in speech-to-text models by 73%, with significant improvements across workplace and call-center recordings.
Claude Opus 5.5 achieves a top score of 88.4% on the SimpleBench benchmark, indicating significant performance in AI evaluation.