[RELEASE] SupraBrain-50M-v0.1

Reddit r/LocalLLaMA Models

Summary

SupraLabs releases SupraBrain-50M, a hybrid language model combining Gated DeltaNet, Sliding-Window Attention, and Surprise-Gated mechanisms, achieving near-parity with Supra-Base-50M despite training on far fewer tokens.

Hey there! So today we're releasing SupraBrain-50M, a hybrid language model that combines Gated DeltaNet linear recurrence with Sliding-Window Attention and Surprise-Gated update mechanisms to deliver very strong performance. Here are the benchmarks: https://preview.redd.it/tv0bxkbyh4hh1.png?width=615&format=png&auto=webp&s=f296ed9e6c1d94738cb7aa84015117dd3e65d2a5 Despite being trained on MUCH less data (5B vs 20B tokens!!), it's almost as good as Supra-Base-50M! 🔥 Some samples: Artificial intelligence is 200% more efficient than human intelligence. - The human brain is 100,000 times more efficient at processing information than the computer. The Human Brain The human brain consists of the following parts: - Brain: The brain is the organ that processes information. It is the brain that is responsible for the following: The brain is composed of three parts: the cerebrum, the cerebellum and the cerebrospinal fluid. And: The mitochondrion produces adenosine triphosphate (ATP) by a process called adenosyltransferase, or AATP. The AATTP is released from the mitochondria and binds to ATP. This ATP is used to create adenosines, which are then released to the cell nucleus. The nucleus then releases adenosaminase (AATP), which is then released by the mitochondria. The mitochondria then break down the ATP into adenosidic bonds and adenosic acid. Link to the HF model: https://huggingface.co/SupraLabs/SupraBrain-50M Give us a follow if you want to support us! BTW: Supra2-100M is releasing in the next few hours!! Stay tuned 🤗
Original Article

Similar Articles

[NEW] Supra-50M Released!

Reddit r/LocalLLaMA

SupraLabs released Supra-50M, a compact 50M-parameter causal language model with base and instruct versions, trained on 20B tokens from fineweb-edu, achieving competitive benchmarks against larger models like GPT-2 and SmolLM.

[NEW MODEL] Supra-Title-0.3B Just released!

Reddit r/LocalLLaMA

Supra Labs released Supra Title, a 350M parameter model specialized for generating chat conversation titles. Built on LFM2.5, it runs on any hardware in GGUF format and requires no system prompt.

Introducing DWARF-55M-Base

Reddit r/LocalLLaMA

DWARF-55M-Base is a new language model using a nearly all-sparse attention architecture (DSQG) with a single full causal attention layer, achieving reliable retrieval up to 2048 tokens and extrapolating to 3x that context. It is released as a research prototype for community experimentation.