Good Summarization SLMs for < 2000 tokens
Summary
A novice asks for recommendations on small language models and prompting strategies to build an employee note summarization engine under 2000 tokens, after experiencing hallucinations with Qwen2.5-7B-Instruct.
Similar Articles
Large Language Models for Token-Efficient and Semantic-Preserving Opinion Summarization
This paper presents a framework for opinion summarization using LLMs that combines multidimensional classification and stratified sampling to reduce token usage while preserving semantic diversity and balance across viewpoints.
The death of SLMs?
The author reflects on whether small language models under 27B are being overshadowed by larger models like Qwen 3.5 and Gemma 4, and asks the community for capable SLMs for agentic coding tasks.
Trained my first small language model
The author trained a small language model to replace Gemini Flash for a summarization task, achieving 97% accuracy with 0.06s latency, suitable for deployment in an internal app.
Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization
This paper proposes a parameter-efficient vocabulary adaptation method for LLM-based text summarization in specialized domains, augmenting pretrained tokenizers with domain-specific tokens and selectively replacing under-trained ones to reduce training time by 35-55% and parameter counts by up to 37%.
LexLattice: Multilingual Extractive Summarization via Neural Cellular Automata on Document Hierarchies
LexLattice introduces a multilingual extractive summarization method using neural cellular automata on document hierarchies, achieving state-of-the-art ROUGE scores on the EUR-Lex-Sum dataset with a compact model surpassing large instruction-tuned baselines.