performance-benchmarks

Tag

Cards List
#performance-benchmarks

Ornith-1.0-35B GGUF update: native MTP speculative-decode graft + full serving/TTFT/long-context numbers (llama.cpp, tp=1)

Reddit r/LocalLLaMA · 2026-06-28

An update on the Ornith-1.0-35B GGUF model introduces a native MTP speculative-decode graft for faster inference on a single GPU, achieving ~1.3-1.35x decode speedup while maintaining near-identical token distribution. Benchmark numbers for throughput, TTFT, and long-context performance across multiple quants are provided.

0 favorites 0 likes
#performance-benchmarks

Why Chinese AI Models Are Reshaping the Economics of AI

Reddit r/AI_Agents · 2026-06-03

Chinese AI models like DeepSeek and Qwen deliver competitive performance at 5x–20x lower cost than Western counterparts, reshaping the economics of AI and driving multi-model deployment strategies.

0 favorites 0 likes
← Back to home

Submit Feedback