Tag
This review paper surveys the application of large models in battery prognostics and health management, addressing long-standing challenges and proposing a roadmap for future research in this domain.
This article explains embedding models and their role in LLM inference, covering architectures, traffic profiles, and optimization techniques like quantization.
This paper presents a cascaded multi-granularity pruning framework for deploying LLMs on Industrial IoT edge devices, achieving up to 13.8x compression with minimal accuracy loss on MHA+GELU architectures while exposing a collapse on GQA+SwiGLU designs.
This paper analyzes residual scaling in looped (weight-tied) transformers, showing that weight sharing requires stronger scaling (1/N) than standard residual networks, and derives a factored parameterization that enables hyperparameter transfer across loop counts without retuning.