tensor-level-allocation

Tag

Cards List
#tensor-level-allocation

Qwen 3.5 4B IQ2_XS: +16.67% Reasoning Performance From Tensor-Level Allocation

Reddit r/LocalLLaMA · 2026-08-22

ByteOtter replicates tensor-level allocation on Qwen 3.5 4B, achieving a 16.67% relative improvement in reasoning performance with only a 0.412% increase in model size, marking the first cross-family application outside Gemma.

0 favorites 0 likes
#tensor-level-allocation

Gemma 4 12B Q3: +8.55% Coding Performance From Tensor-Level Quantization Allocation

Reddit r/LocalLLaMA · 2026-08-13

A developer created a task-aware GGUF quantization pipeline that uses tensor-level bit allocation to improve Gemma 4 12B Q3 coding performance by 8.55% over a hand-tuned imatrix while increasing model size by only 0.119%.

0 favorites 0 likes
← Back to home

Submit Feedback