here it is: Benchmark-Yourself app - compete against open source LLMs and get your score - 5 benchmarks available - Add your results to your CV or linkedIn (if you dare)... or just paste them below for community shaming.
Summary
A web app that allows users to benchmark their own performance against open source LLMs on five benchmarks, with the option to add results to a CV or LinkedIn.
Similar Articles
Homebench – Benchmark local LLMs for speed, memory, and quality
Homebench is a zero-config terminal tool that benchmarks locally-run LLMs for speed, memory, and quality, presenting a live leaderboard. It supports Ollama, LM Studio, llama.cpp, vLLM, and OpenAI-compatible servers.
Open-source LLM benchmark runs 147 coding tasks every 4 hours, 5-trial median with 95% CI, and uses CUSUM for change-point detection. Curious what people think of the methodology
An open-source LLM benchmark with 147 coding tasks runs every 4 hours, using 5-trial median with 95% confidence intervals and CUSUM for change-point detection, sparking discussion on its methodology.
@ms_aifrontiers: Most LLM benchmark scores are predictable before you ever run them. New from the MS AI Frontiers team: BenchPress. The …
The MS AI Frontiers team introduces BenchPress, a method that uses matrix completion to predict LLM benchmark scores from just five probes, showing the score matrix is effectively rank-2.
Meta-Benchmarks for Financial-Services LLM Evaluation
This paper presents a meta-benchmarking framework that aggregates 452 existing public benchmarks into 41 work activities and 38 banking business domains, enabling more precise LLM evaluation and governance for financial services institutions.
Benchmarking LLM Judges for Mobile Agent Evaluation
This paper introduces MobileJudgeBench, a benchmark with 931 human-annotated trajectories for systematically evaluating LLM-based judges on mobile agent tasks. It finds that simple baseline judges with sampled screenshots rival purpose-built methods, with the LLM backbone being the primary driver of quality.