Fixed the MTP head on Ornith1.5 35B A3B. +3% TPS -33% wall clock
Summary
A user fixed the MTP head on the Ornith1.5 35B A3B model, achieving a 33% reduction in wall clock time for tasks, making it significantly faster for local HAM radio applications.
Similar Articles
If you are wondering why Ornith 1.5 35B A3B with MTP is so slow, this is why
The Ornith 1.5 35B A3B model's MTP tensors appear to be uninitialized, causing poor speculative decoding performance, and grafting the trained head from Qwen3.6-35B-A3B improves speed by 29%.
I added MTP to local SoTA Agentic Coding Model Ornith 35B FP8 E4M3
Introduced MTP speculative decoding to the Ornith 35B coding model in FP8 precision, achieving approximately 18% faster inference with minimal extra VRAM.
Ornith-1.0-35B GGUF update: native MTP speculative-decode graft + full serving/TTFT/long-context numbers (llama.cpp, tp=1)
An update on the Ornith-1.0-35B GGUF model introduces a native MTP speculative-decode graft for faster inference on a single GPU, achieving ~1.3-1.35x decode speedup while maintaining near-identical token distribution. Benchmark numbers for throughput, TTFT, and long-context performance across multiple quants are provided.
@Italianclownz: Converted Qwen 3.6 35b a3b to ROCmfp4 and this is flying. Used the mtp version bc this ROCmfp4 can also incorporate the…
Converted the Qwen 3.6 35b a3b model to ROCmfp4 format, leveraging MTP benefits for improved performance on AMD hardware.
@malikwas1f: Ornith-1.0-35B: a Qwen3.6-35B-A3B coding fine-tune that edges the base on real coding (aider 15/30 vs 13) — full 262K a…
Announces Ornith-1.0-35B, a coding fine-tune of Qwen3.6-35B-A3B that slightly outperforms the base model on aider benchmarks. Also promotes the club-3090 repository for running LLMs on RTX 3090s.