Tag
This paper presents Search-on-Graph-R1 (SoG-R1), which trains an 8B LLM to navigate knowledge graphs by first scaffolding a frontier teacher with gold SPARQL queries to produce grounded trajectories, then applying supervised fine-tuning and reinforcement learning. The compact model surpasses frozen frontier systems on WebQSP, CWQ, and GrailQA, notably achieving the best results on CWQ among compared methods.