[BIG DATASET RELEASE] - SupraLabs/reasoning-corpus-4K-5M-v1 - Train your tiny SLMs to think!
Summary
SupraLabs releases reasoning-corpus-4K-5M-v1, a 5M-sample reasoning dataset for training small language models (SLMs), featuring chain-of-thought traces and ChatML format, hosted on Hugging Face.
Similar Articles
@AdinaYakup: OpenBMB just released an impressive SFT dataset UltraData-SFT-2605 15M+ high quality samples Deep Thinking + Non-thinki…
OpenBMB releases UltraData-SFT-2605, a large-scale dataset with over 15 million high-quality samples for supervised fine-tuning (SFT) of reasoning LLMs, covering deep thinking, non-thinking, math, code, knowledge, instruction following, and multilingual data.
Multilingual Tiny (3.7B) Reasoning MoE pretrained from scratch on a consumer-grade GPU
A multilingual 3.7B parameter reasoning MoE model has been pretrained from scratch on a consumer-grade GPU over several months, with support for 13 languages and available on Hugging Face.
A Very Big Video Reasoning Suite
This paper introduces the Very Big Video Reasoning (VBVR) dataset and benchmark, a large-scale resource with over one million video clips across 200 reasoning tasks, enabling systematic study of spatiotemporal reasoning and showing early signs of emergent generalization.
GigaChat-3.5-Reasoning
This article presents a collection on Hugging Face featuring GigaChat 3.5 Reasoning models, which are large AI models optimized for reasoning tasks in various formats.
Domyn-Small: A European 10B Reasoning Language Model
Domyn-Small is a 10-billion-parameter open-weight reasoning language model released under MIT license, trained on 9 trillion tokens and optimized for reasoning, instruction following, and tool use. It achieves strong accuracy-efficiency balance against peers like Qwen3.5-9B and OLMo-3-7B-Think.