benchmark-results

Tag

Cards List
#benchmark-results

Artificial Analysis: Muse Spark 1.1 Results

Reddit r/singularity · 2026-07-10

Artificial Analysis reports benchmark results for the Muse Spark 1.1 AI model, providing performance metrics.

0 favorites 0 likes
#benchmark-results

ZAYA1-8B Technical Report

arXiv cs.AI · 2026-05-08 Cached

This technical report introduces ZAYA1-8B, a mixture-of-experts reasoning model trained on AMD hardware that achieves competitive performance on math and coding benchmarks using under 1B active parameters. It also details Markovian RSA, a novel test-time compute method for aggregating parallel reasoning traces.

0 favorites 1 likes
#benchmark-results

Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet

Anthropic Engineering · 2026-05-08 Cached

Anthropic's updated Claude 3.5 Sonnet achieves a new state-of-the-art 49% on the SWE-bench Verified benchmark, demonstrating significant capabilities in autonomous software engineering tasks.

0 favorites 0 likes
← Back to home

Submit Feedback