@AdinaYakup: OpenBMB just released an impressive SFT dataset UltraData-SFT-2605 15M+ high quality samples Deep Thinking + Non-thinki…

X AI KOLs Following Tools

Summary

OpenBMB releases UltraData-SFT-2605, a large-scale dataset with over 15 million high-quality samples for supervised fine-tuning (SFT) of reasoning LLMs, covering deep thinking, non-thinking, math, code, knowledge, instruction following, and multilingual data.

OpenBMB just released an impressive SFT dataset UltraData-SFT-2605 📊 ✨ 15M+ high quality samples ✨ Deep Thinking + Non-thinking data ✨ Math/ Code/ Knowledge/ IF/ Multilingual coverage ✨ Built for reasoning LLM post-training ✨ Full data pipeline: filtering/ https://t.co/RUYIwqTGiQ
Original Article
View Cached Full Text

Cached at: 05/30/26, 04:05 AM

OpenBMB just released an impressive SFT dataset

UltraData-SFT-2605 📊

✨ 15M+ high quality samples ✨ Deep Thinking + Non-thinking data ✨ Math/ Code/ Knowledge/ IF/ Multilingual coverage ✨ Built for reasoning LLM post-training ✨ Full data pipeline: filtering/ https://t.co/RUYIwqTGiQ

Similar Articles

Learning to Reason with Insight for Informal Theorem Proving

arXiv cs.CL

This paper proposes DeepInsightTheorem, a hierarchical dataset and Progressive Multi-Stage SFT training strategy to improve LLMs' informal theorem proving by teaching them to identify and apply core techniques through insight-aware reasoning.