Ornith-1.5-35B-A3B Q4 running 60tk/s on 4070Ti.
Summary
A user reports achieving 60 tokens per second with the Ornith-1.5-35B-A3B MOE model on an NVIDIA RTX 4070 Ti, demonstrating efficient local inference optimization.
Similar Articles
@DogukanUrker: Ornith-1.5-9B at Q5 on a single RTX 3060: 200k context at ~52 tok/s, ~1700 tok/s prefill. 11.8 of the 12GB, zero cpu of…
The post details running the Ornith-1.5-9B AI model on an RTX 3060 with 200k context, achieving high inference speeds using advanced quantization techniques.
Ornith-397B running at Q4 on a single RTX PRO 6000 Blackwell 96GB - 2,354 tok/s prefill, ~20–24 tok/s decode
Krasis, a MoE-focused runtime, enables running the 397B-parameter Ornith model on a single RTX PRO 6000 Blackwell 96GB GPU with ~20-24 tok/s decode by dynamically managing expert residency in VRAM.
Ornith-1.5-35B-A3B-NInfer - 250 tok/s, 5-8k prefill, 5090
Ornith-1.5-35B-A3B-NInfer is a derived, quantized inference artifact for the Qwen 3.5 MoE multimodal model, optimized for high-speed serving via the NInfer engine.
@ItsmeAjayKV: Quick update: I tried Ornith-1.0-35B-Q5_K_M on my 3090, and i have mixed feelings. The good: it's really fast. I measur…
User tests Ornith-1.0-35B on an RTX 3090, finding fast inference speeds (1560 tok/s prompt, 78 tok/s generation) but consistently worse coding performance on Three.js tasks compared to Qwen 3.6, even after multiple attempts.
@sudoingX: i was running Ornith new 35b moe on llama.cpp with a Q4 quant, 4 bit, small, fast. it hit ~78 tok/s. then i swapped eng…
A 35B MoE agentic coding model called Ornith runs near lossless at FP8 on a single DGX Spark, achieving 3M token context and ~36 tok/s, with speculative decoding expected to boost speed further.