efficient-deployment

Tag

Cards List
#efficient-deployment

Pro-Router: Token-Aware Progressive Model Routing with Adaptive Edge-Cloud Collaboration for Efficient Multimodal LLM Inference

arXiv cs.AI · 2026-09-01 Cached

Pro-Router introduces a token-aware progressive model routing method for efficient multimodal LLM inference, leveraging adaptive edge-cloud collaboration to improve throughput and reduce costs.

0 favorites 0 likes
#efficient-deployment

Memory Efficient Tabular Foundation Models

arXiv cs.LG · 2026-07-31 Cached

This paper investigates memory requirements for tabular foundation models like TabPFN and shows that model compression (e.g., INT4 quantization) can reduce memory footprint up to 7.6x with minimal accuracy loss, improving practical deployment efficiency.

0 favorites 0 likes
#efficient-deployment

ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level

arXiv cs.LG · 2026-07-16 Cached

ExTernD introduces an expanded-rank ternary decomposition for post-training LLM quantization, enabling accuracy approaching bf16 by using a factored representation with free inner rank. It matches Q4_K accuracy at 5.2-5.5 effective bits per weight on models like Gemma-4 and Qwen3.5.

0 favorites 0 likes
#efficient-deployment

Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?

arXiv cs.AI · 2026-07-15 Cached

The paper proposes Light-MER, a lightweight multimodal emotion recognition framework that uses knowledge distillation from an 8B teacher model to a sub-1B student, achieving state-of-the-art performance with significantly higher inference efficiency, challenging the necessity of models larger than 1B parameters.

0 favorites 0 likes
#efficient-deployment

It Takes a MAESTRO To Prune Bad Experts

arXiv cs.CL · 2026-07-10 Cached

This paper introduces Maestro, a structured pruning framework for Mixture-of-Experts language models that uses Markov chains to model expert activation trajectories, achieving globally aware pruning and outperforming baselines by up to 10.61% under 50% compression.

0 favorites 0 likes
#efficient-deployment

FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models

arXiv cs.LG · 2026-06-29 Cached

FlexMoE proposes a one-for-all nested intra-expert pruning method for MoE language models, enabling multiple deployable subnetworks from a single training run with minimal performance loss.

0 favorites 0 likes
#efficient-deployment

CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs

arXiv cs.CL · 2026-06-26 Cached

CAT-Q introduces a post-training ternary quantization method for LLMs that uses learnable modulation and softened ternarization, achieving superior performance over BitNet 1.58-bit while using only 512 calibration samples and scaling to 235B parameters.

0 favorites 0 likes
#efficient-deployment

TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts

arXiv cs.LG · 2026-06-10 Cached

TENP proposes a structured pruning framework for Mixture-of-Experts LLMs that retains important experts and applies neuron pruning to less important ones, achieving high sparsity with minimal accuracy loss on Qwen and DeepSeek models.

0 favorites 0 likes
#efficient-deployment

Joint Structural Pruning and Mixed-Precision Quantization for LLM Compression

arXiv cs.AI · 2026-06-09 Cached

A novel end-to-end framework for LLM compression that jointly optimizes structural pruning and mixed-precision quantization, achieving significant perplexity reductions and speedups over state-of-the-art methods, especially at ultra-low bit precisions.

0 favorites 0 likes
#efficient-deployment

BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization

arXiv cs.LG · 2026-06-02 Cached

BitsMoE introduces a spectral-energy-guided bit allocation framework for quantizing Mixture-of-Experts LLMs, achieving substantial accuracy improvements and speedups under ultra-low-bit quantization.

0 favorites 0 likes
#efficient-deployment

@berryxia: Apple has been betting on on-device models all along! Unified architecture memory is the natural habitat for on-device models! Unified memory means memory is VRAM. We are seeing more and more excellent on-device models emerge. OpenBMB released MiniCPM-V 4.6, a 1.3B multimodal model. After reading it…

X AI KOLs Timeline · 2026-05-12

OpenBMB released MiniCPM-V 4.6, a 1.3B parameter multimodal model. Using high-resolution visual processing and efficient compression, it achieves fast inference on consumer hardware and mobile phones, outperforming larger models. It is fully open-source and supports multiple inference and quantization frameworks.

0 favorites 0 likes
← Back to home

Submit Feedback