Tag
The article argues that separating data and AI strategies creates technical debt and proposes a unified approach with pillars and a roadmap to integrate data infrastructure with AI capabilities for production scalability.
This paper proposes a difficulty-aware SFT-then-RL framework for training small language models (≤3B parameters) on reasoning tasks, arguing that data difficulty should be strategically aligned with the distinct roles of SFT (learning new skills) and RL (consolidating partial skills). The authors introduce a Bridge mechanism for hard SFT samples and Critique Fine-Tuning for RL failures, showing consistent improvements across five reasoning benchmarks.