Tag
Released GSQ-RCO quantized and expert-pruned versions of Qwen3.8-Flash-Next, achieving BF16-level performance at reduced bit-widths and enabling deployment on smaller hardware.
ToMoE is a method that converts dense large language models into mixture-of-experts models through dynamic structural pruning, achieving strong performance without weight updates and outperforming state-of-the-art pruning and MoE techniques.
This paper analyzes how prompt template selection during knowledge distillation affects the safety alignment of student large language models, finding that chat templates lead to greater degradation compared to non-chat templates across multiple models and benchmarks.
The paper proposes LT-OPD, an on-policy self-distillation framework for extreme visual token reduction in multimodal large language models, significantly improving performance under low token budgets while reducing KV-cache usage and FLOPs.
This paper investigates how post-training weight compression of Whisper ASR models widens demographic disparities in transcription errors, leading to increased correction time for marginalized speakers.
This paper proposes a quantization-robust unlearning framework for large language models, using loss landscape analysis to ensure effective forgetting while maintaining model utility after compression.
The paper introduces Magnitude Profile (MP) scoring, a calibration-free method for pruning attention heads in transformers, achieving better perplexity on models like OPT-6.7B and RoBERTa-large compared to existing methods, with zero forward passes or calibration data.
The paper introduces a method to prune encoder layers in the Whisper ASR model, reducing encoder size by 18.5% and recovering performance through unlabeled data distillation, with code and a pre-trained model released for adoption.
DeltaTensors compresses fine-tuned model checkpoints by storing weight differences from the base model, reducing storage requirements significantly while allowing reconstruction with minimal quality loss.
The paper proposes a correlation-aware structured pruning method for large language models, explicitly modeling cross-unit dependencies to improve pruning decisions and achieve competitive accuracy-efficiency trade-offs.
The paper introduces GeoPair, a training-free framework for transformer compression that optimizes cross-layer factorizations while preserving activation geometries, achieving state-of-the-art results across diverse architectures.
This tweet recommends a Stanford course on efficient generative language models, covering techniques from pre-training to inference to balance performance and cost with limited compute.
The paper presents a physics-inspired approach to pruning LLM blocks by modeling block removal as a constrained binary optimization problem mapped to an Ising glass, achieving significant compression gains without benchmarking each configuration.
Multiverse Computing released Quasar 1.1 438B, an AI coding model built using quantum-generated data, with improved efficiency and reduced political censorship.
Ternary-Bonsai-2-27B is a 27B parameter AI model compressed into 2-bit ternary format, running on llama.cpp with CUDA and Metal support, facilitating lightweight on-device AI deployment.
A tweet posing a question about who is distilling the Astra model for robotics, indicating community interest in AI model optimization for deployment in robotic systems.
The paper introduces a layer-wise curriculum learning method for efficient LLM compression, achieving state-of-the-art performance with significant reductions in GPU memory usage and training time.
A user is testing the newly announced Ternary Bonsai 2 27B AI model, a smaller and quantized version of Qwen3.8 27B, on an NVIDIA 5070 Ti GPU and finds its performance impressive.
Ternary Bonsai 2 is a 27B parameter model derived from Qwen3.8-27B that uses ternary weights to achieve a size under 6GB while retaining 98.2% of its intelligence, enabling it to run in-browser on WebGPU.
Edge0 is a streaming MoE inference engine that uses trained routing prediction to serve 35B Mixture-of-Experts models from SSD, achieving near-fp16 performance on consumer hardware with low memory usage.