benchmark-results

Tag

Cards List
#benchmark-results

@QuixiAI: ChatGPT failed Claude failed Deepseek failed Qwen3.8 passed GLM 5.3 passed Grok passed

X AI KOLs Timeline · 2026-08-27 Cached

The tweet reports that ChatGPT, Claude, and Deepseek failed a test or benchmark, while Qwen3.8, GLM 5.3, and Grok passed, based on a linked source.

0 favorites 0 likes
#benchmark-results

LOL, my timeline was flooded with this guy's shock. Someone just bypassed the safety guardrails of Qwen3.8-27B. OrcaRouter had previously released a Qwen3.8-27B-Uncensored-FP8 on Hugging Face. They used Abliteration…

X AI KOLs Timeline · 2026-08-21 Cached

By using Abliteration to modify the weights, someone published an uncensored version of Qwen3.8-27B, drastically lowering the refusal rate and preserving model performance, which has sparked debates on AI safety.

0 favorites 0 likes
#benchmark-results

I benchmarked Ox Alpha on SWE-bench Verified Mini (50 tasks): 96% resolved. Now I’m skeptical of myself.

Reddit r/LocalLLaMA · 2026-08-21

The author benchmarked Ox Alpha on SWE-bench Verified Mini, achieving a 96% resolution rate, but raises concerns about the score's validity due to data contamination, small sample size, and non-comparable baselines.

0 favorites 0 likes
#benchmark-results

FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills

Hugging Face Daily Papers · 2026-08-20 Cached

FlowEvo is a training-free framework that enables large language model agents to co-evolve reusable skills and workflows at inference time, achieving state-of-the-art accuracy and efficiency across benchmarks like ALFWorld, HumanEval, and GSM8K.

0 favorites 0 likes
#benchmark-results

DeepSeek v4 Flash has a nice bump in Capability

Reddit r/LocalLLaMA · 2026-07-31

DeepSeek V4 Flash shows significant benchmark gains in preview updates, trading blows with GPT-5.6 Terra on agentic coding tasks.

0 favorites 0 likes
#benchmark-results

Artificial Analysis: Muse Spark 1.1 Results

Reddit r/singularity · 2026-07-10

Artificial Analysis reports benchmark results for the Muse Spark 1.1 AI model, providing performance metrics.

0 favorites 0 likes
#benchmark-results

ZAYA1-8B Technical Report

arXiv cs.AI · 2026-05-08 Cached

This technical report introduces ZAYA1-8B, a mixture-of-experts reasoning model trained on AMD hardware that achieves competitive performance on math and coding benchmarks using under 1B active parameters. It also details Markovian RSA, a novel test-time compute method for aggregating parallel reasoning traces.

0 favorites 1 likes
#benchmark-results

Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet

Anthropic Engineering · 2026-05-08 Cached

Anthropic's updated Claude 3.5 Sonnet achieves a new state-of-the-art 49% on the SWE-bench Verified benchmark, demonstrating significant capabilities in autonomous software engineering tasks.

0 favorites 0 likes
← Back to home

Submit Feedback