@AdinaYakup: OpenBMB just released an impressive SFT dataset UltraData-SFT-2605 15M+ high quality samples Deep Thinking + Non-thinki…
Summary
OpenBMB releases UltraData-SFT-2605, a large-scale dataset with over 15 million high-quality samples for supervised fine-tuning (SFT) of reasoning LLMs, covering deep thinking, non-thinking, math, code, knowledge, instruction following, and multilingual data.
View Cached Full Text
Cached at: 05/30/26, 04:05 AM
OpenBMB just released an impressive SFT dataset
UltraData-SFT-2605 📊
✨ 15M+ high quality samples ✨ Deep Thinking + Non-thinking data ✨ Math/ Code/ Knowledge/ IF/ Multilingual coverage ✨ Built for reasoning LLM post-training ✨ Full data pipeline: filtering/ https://t.co/RUYIwqTGiQ
Similar Articles
@AtriaASI: We’re grateful to OpenBMB for supporting Atria Dawn Preview with UltraData-SFT-Agent-2609. The dataset’s high-quality a…
OpenBMB releases UltraData-SFT-Agent-2609, a dataset of 500,000 samples for agent instruction-tuning, used in the post-training of the MiniCPM5-2B model.
[BIG DATASET RELEASE] - SupraLabs/reasoning-corpus-4K-5M-v1 - Train your tiny SLMs to think!
SupraLabs releases reasoning-corpus-4K-5M-v1, a 5M-sample reasoning dataset for training small language models (SLMs), featuring chain-of-thought traces and ChatML format, hosted on Hugging Face.
@cjzafir: Fine-tune your first AI model today. Run GPT4o level model and run on your phone or laptop. @OpenBMB released 15M sampl…
OpenBMB released UltraData-SFT-2605, a 15M-sample high-quality SFT dataset for fine-tuning AI models like MiniCPM5-1B to run on phones or laptops.
@AdinaYakup: Ultra-FineWeb-L1 Open English web corpus for LLM pre-training from @OpenBMB https://huggingface.co/datasets/openbmb/Ult…
OpenBMB releases Ultra-FineWeb-L1, an open English web corpus with over 1 trillion tokens for LLM pre-training, derived from Common Crawl and featuring advanced cleaning with Trafilatura 2.0.
Learning to Reason with Insight for Informal Theorem Proving
This paper proposes DeepInsightTheorem, a hierarchical dataset and Progressive Multi-Stage SFT training strategy to improve LLMs' informal theorem proving by teaching them to identify and apply core techniques through insight-aware reasoning.