efficient-deployment

Tag

Cards List
#efficient-deployment

ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level

arXiv cs.LG · 2026-07-16 Cached

ExTernD introduces an expanded-rank ternary decomposition for post-training LLM quantization, enabling accuracy approaching bf16 by using a factored representation with free inner rank. It matches Q4_K accuracy at 5.2-5.5 effective bits per weight on models like Gemma-4 and Qwen3.5.

0 favorites 0 likes
#efficient-deployment

Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?

arXiv cs.AI · 2026-07-15 Cached

The paper proposes Light-MER, a lightweight multimodal emotion recognition framework that uses knowledge distillation from an 8B teacher model to a sub-1B student, achieving state-of-the-art performance with significantly higher inference efficiency, challenging the necessity of models larger than 1B parameters.

0 favorites 0 likes
#efficient-deployment

It Takes a MAESTRO To Prune Bad Experts

arXiv cs.CL · 2026-07-10 Cached

This paper introduces Maestro, a structured pruning framework for Mixture-of-Experts language models that uses Markov chains to model expert activation trajectories, achieving globally aware pruning and outperforming baselines by up to 10.61% under 50% compression.

0 favorites 0 likes
#efficient-deployment

FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models

arXiv cs.LG · 2026-06-29 Cached

FlexMoE proposes a one-for-all nested intra-expert pruning method for MoE language models, enabling multiple deployable subnetworks from a single training run with minimal performance loss.

0 favorites 0 likes
#efficient-deployment

CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs

arXiv cs.CL · 2026-06-26 Cached

CAT-Q introduces a post-training ternary quantization method for LLMs that uses learnable modulation and softened ternarization, achieving superior performance over BitNet 1.58-bit while using only 512 calibration samples and scaling to 235B parameters.

0 favorites 0 likes
#efficient-deployment

TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts

arXiv cs.LG · 2026-06-10 Cached

TENP proposes a structured pruning framework for Mixture-of-Experts LLMs that retains important experts and applies neuron pruning to less important ones, achieving high sparsity with minimal accuracy loss on Qwen and DeepSeek models.

0 favorites 0 likes
#efficient-deployment

Joint Structural Pruning and Mixed-Precision Quantization for LLM Compression

arXiv cs.AI · 2026-06-09 Cached

A novel end-to-end framework for LLM compression that jointly optimizes structural pruning and mixed-precision quantization, achieving significant perplexity reductions and speedups over state-of-the-art methods, especially at ultra-low bit precisions.

0 favorites 0 likes
#efficient-deployment

BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization

arXiv cs.LG · 2026-06-02 Cached

BitsMoE introduces a spectral-energy-guided bit allocation framework for quantizing Mixture-of-Experts LLMs, achieving substantial accuracy improvements and speedups under ultra-low-bit quantization.

0 favorites 0 likes
#efficient-deployment

@berryxia: Apple has been betting on on-device models all along! Unified architecture memory is the natural habitat for on-device models! Unified memory means memory is VRAM. We are seeing more and more excellent on-device models emerge. OpenBMB released MiniCPM-V 4.6, a 1.3B multimodal model. After reading it…

X AI KOLs Timeline · 2026-05-12

OpenBMB released MiniCPM-V 4.6, a 1.3B parameter multimodal model. Using high-resolution visual processing and efficient compression, it achieves fast inference on consumer hardware and mobile phones, outperforming larger models. It is fully open-source and supports multiple inference and quantization frameworks.

0 favorites 0 likes
← Back to home

Submit Feedback