model-performance

Tag

Cards List
#model-performance

Terminal Bench v4 scores

Reddit r/LocalLLaMA · 10h ago

Terminal Bench v4 scores show GLM-5.3 leading among open models, with GLM-5.3-Flash topping flash models, and Kimi-K3 underperforming, as the benchmark is considered by some to better reflect model intelligence.

0 favorites 0 likes
#model-performance

GPT-6 Astra 2,340 Elo on ChessBench - Ranked #11

Reddit r/singularity · 13h ago Cached

The ChessBench AI chess leaderboard has been updated with new model performances, including GPT-6 Astra achieving 2,340 Elo and ranking #11, along with other models like Claude Fable 5.1 and Gemini 3.8 Flash.

0 favorites 0 likes
#model-performance

GPT 6 Astra beat Fallout 2 in 22 hours with a Vision-only harness

Reddit r/singularity · 2026-09-04

GPT-6 Astra, using a vision-only harness, completed Fallout 2 in 22 hours. Clad3815, known for streaming Pokémon games with AI, had early access to the model.

0 favorites 0 likes
#model-performance

@JinjingLiang: Codex / Claude / Grok all down. Never imagined `agy` to be so load-bearing...

X AI KOLs Timeline · 2026-09-03 Cached

The post highlights Gemini 3.8 Flash's superior performance over Opus 5 on DeepSWE-bench, emphasizing its capabilities in coding and agentic tasks, and notes its accessibility through the `agy` harness in Orca.

0 favorites 0 likes
#model-performance

Simple Bench - QWEN 3.8 27b has a common sense almost like GPT 5.0 Pro??

Reddit r/singularity · 2026-09-03

A benchmark called Simple Bench shows that the QWEN 3.8 27b model has common sense capabilities almost comparable to GPT 5.0 Pro.

0 favorites 0 likes
#model-performance

@ArizePhoenix: it now leads CyberGym at 84.5%, and https://Z.ai is running a public disclosure ledger for the 2,400+ real-world vulner…

X AI KOLs Following · 2026-09-03 Cached

An AI model has achieved 84.5% on CyberGym, leading the benchmark, and Z.ai is running a public disclosure ledger for over 2,400 real-world vulnerabilities surfaced by the model.

0 favorites 0 likes
#model-performance

All currently popular local models in one table + Opus 4.8 results

Reddit r/LocalLLaMA · 2026-09-01

This article provides a comparative table of popular local AI models across various benchmarks, including agentic, coding, general, and multimodal tasks, to help users choose models based on hardware specs and use cases.

0 favorites 0 likes
#model-performance

@rohanpaul_ai: Another example that the harness, more than the model itself, decides how far intelligence actually gets. With the API …

X AI KOLs Timeline · 2026-08-28 Cached

The tweet highlights how Atomic Agent, a model-agnostic agent layer, improves the performance of GLM 5.3 by executing model actions and preserving state, nearly doubling token usage for only a 77-cent cost increase.

0 favorites 0 likes
#model-performance

@rohanpaul_ai: This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed fee…

X AI KOLs Following · 2026-08-28 Cached

A paper evaluates eight leading AI models on long-horizon tasks, finding that even the best-performing model achieves only 27.3% of human performance, highlighting significant limitations for dependable long-horizon AI execution.

0 favorites 0 likes
#model-performance

Qwen3.8-27b q8 KV cache does seem to actually hurt model performance

Reddit r/LocalLLaMA · 2026-08-28

The article discusses how on-the-fly KV cache quantization can reduce long-context model performance due to compounding errors, based on experiments with Qwen3.8-27B.

0 favorites 0 likes
#model-performance

A 27b model beating latest frontier models was not on my 2026 bingo card

Reddit r/LocalLLaMA · 2026-08-26

A user shares that a 27b AI model is beating latest frontier models and provides personal experience with Qwen 3.8 for agentic tasks and Qwen 3.7 flash for overall tasks.

0 favorites 0 likes
#model-performance

Qwen 3.8 27B Aider score

Reddit r/LocalLLaMA · 2026-08-24

A user benchmarks Qwen 3.8 27B using Aider and finds it scores 72.9, matching Gemini 2.5 Pro and outperforming other state-of-the-art models on a MacBook with local inference via vLLM.

0 favorites 0 likes
#model-performance

@narens: Benchmaxxed

X AI KOLs Following · 2026-08-23 Cached

Gemini 3.7 flash outperforms Fable 5, Opus 5, and GPT-5.6 on the Analyst Agent benchmark by Artificial Analysis.

0 favorites 0 likes
#model-performance

@karanC_12: Ox Alpha just finished the full DeepSWE run. Final score: ~63% For reference: • DeepSeek V4 Pro → 63% • Grok 4.6 → 65% …

X AI KOLs Timeline · 2026-08-23 Cached

Ox Alpha, a free model with 1M context, achieved a ~63% score on the DeepSWE benchmark, matching frontier mid-tier models and suggesting potential for local AI agent development.

0 favorites 0 likes
#model-performance

I benchmarked Ox Alpha on SWE-bench Verified Mini (50 tasks): 96% resolved. Now I’m skeptical of myself.

Reddit r/LocalLLaMA · 2026-08-21

The author benchmarked Ox Alpha on SWE-bench Verified Mini, achieving a 96% resolution rate, but raises concerns about the score's validity due to data contamination, small sample size, and non-comparable baselines.

0 favorites 0 likes
#model-performance

Training AI on AI slop

Reddit r/ArtificialInteligence · 2026-08-21

The article discusses the problem of training AI models on low-quality data generated by other AI systems, known as 'slop,' and its implications for model reliability and performance.

0 favorites 0 likes
#model-performance

@gfodor: Based on my experience, it doesn’t seem like there has been a single day where Anthropic actually has the best coding m…

X AI KOLs Following · 2026-08-20

The author shares their experience that Anthropic's models are not the best for coding, yet Claude Code is widely used, questioning why this is the case.

0 favorites 0 likes
#model-performance

Qwen 3.8 27B SlopCodeBench results

Reddit r/LocalLLaMA · 2026-08-19

The article presents benchmark results for the Qwen 3.8 27B model on SlopCodeBench, showing poor performance on strict checkpoints but fair results on core ones, indicating it may not be suitable for autonomous code management without direction.

0 favorites 0 likes
#model-performance

@Argona0x: whoever leaked this has bigger balls than sense someone at Anthropic hired 80 AI helpers onto one project, gave them tw…

X AI KOLs Timeline · 2026-08-19 Cached

A leaked experiment at Anthropic revealed that using 80 AI helpers on a project led to poor results unless they worked independently, with newer models performing better due to isolation. The author recommends assigning each AI helper a single, owned job to improve productivity.

0 favorites 0 likes
#model-performance

The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning

arXiv cs.CL · 2026-08-19 Cached

The IOL-AI Challenge is an open-science competition using unseen problems from the International Linguistics Olympiad 2026 to evaluate AI models on linguistic reasoning, showing that performance depends more on decoding and output handling than model scale.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback