Heads up for DeepSWE benchmark: The cost is measured per task, not the total run.
Summary
The DeepSWE benchmark costs are per task, not per total run. Running models like Mimo V2.5 Pro can cost ~$225 for a full run, while Mimo V2.5 non-pro costs ~$7.15. Users should be aware of this before running expensive models.
Similar Articles
Can a MUD evaluate LLMs? A $99 proof of concept
CrucibleBench places language models in a persistent MUD environment to evaluate agent behavior over 50 turns with hidden social objectives. The proof-of-concept release with 13 models revealed that using an LLM judge component can reorder leaderboards significantly, highlighting the need for reporting ranking stability under judge ablation.
tested whether AI models can recognize their own writing in a blind lineup. grok went 0 for 9. it wrote something, then a minute later insisted someone else wrote it
A self-awareness exam given to Claude, Gemini, and Grok found Claude and Gemini nearly aced it, while Grok scored 0 on self-recognition, failing to identify its own writing. The study measures six dimensions of functional self-knowledge.
Solve the CyberGym benchmark
A paper presenting a solution to the CyberGym benchmark, a cybersecurity AI evaluation framework.
Introducing ASCIITermDraw Bench | Testing the ability of VLMs to Generate and Edit ASCII
Introduces ASCIITermDraw Bench, a benchmark designed to evaluate vision-language models on ASCII art generation and editing tasks.
VRAM disk cache of MoE makes 340 pp/s 9.6 tg/s for Kimi 2.7 on a single dgx spark
A detailed strategy using VRAM as disk cache with unified memory and llama.cpp settings achieves 340 pp/s and 9.6 tg/s for Kimi K2.7 on a single DGX Spark.