Tag
Shibai-700M-Base is a 700M parameter LLaMA-based model pre-trained on 18B tokens of English, math, and Python code, achieving coherent generation despite being under-trained compared to larger models, and notable for its efficient training on a single RTX 5070 Ti GPU.
CHIAR-Former uses spectral entropy-based routing to dynamically select between DCT, RBF, and self-attention operators, achieving improved efficiency on large text datasets while maintaining performance through hybrid attention mechanisms.