Tag
This paper presents SimpleOPD, a method for on-policy distillation from long-context reasoning teachers to short-context students, overcoming tokenizer mismatch and training instability to enhance mathematical reasoning capabilities.
This paper proposes a tokenizer-agnostic modification to DeepSeek's Engram conditional memory module, replacing XOR-based n-gram hashing with polynomial hashing to enable compatibility across different tokenizers while maintaining comparable performance.
This paper proposes Byte-Level Distillation (BLD), a simple method for cross-tokenizer knowledge transfer in language models by operating at a shared byte-level interface, achieving competitive or superior performance compared to more complex existing approaches across 1B-8B parameter models.