wikitext

Tag

Cards List
#wikitext

I pre-trained a 700m on 18B tokens optimized for Python and Wikitext | TheOneWhoWill/Shibai-700M-Base · Hugging Face

Reddit r/LocalLLaMA · 3d ago Cached

Shibai-700M-Base is a 700M parameter LLaMA-based model pre-trained on 18B tokens of English, math, and Python code, achieving coherent generation despite being under-trained compared to larger models, and notable for its efficient training on a single RTX 5070 Ti GPU.

0 favorites 0 likes
#wikitext

Chiaroscuro Attention: Spending Compute in the Dark

Hugging Face Daily Papers · 2026-06-06 Cached

CHIAR-Former uses spectral entropy-based routing to dynamically select between DCT, RBF, and self-attention operators, achieving improved efficiency on large text datasets while maintaining performance through hybrid attention mechanisms.

0 favorites 0 likes
← Back to home

Submit Feedback