GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation
Summary
GraphGen is a knowledge-graph-guided framework for generating synthetic QA data to improve supervised fine-tuning of LLMs, targeting knowledge gaps with multi-hop sampling and style-controlled generation. Experiments show it outperforms conventional synthetic data methods.
View Cached Full Text
Cached at: 08/03/26, 01:29 PM
Paper page - GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation
Source: https://huggingface.co/papers/2505.20416 Published on May 26, 2025
Abstract
GraphGen is a knowledge graph-guided framework that addresses challenges in synthetic data generation for LLMs by constructing fine-grained knowledge graphs, targeting high-value knowledge gaps, and employing multi-hop sampling and style-controlled generation.
Fine-tuningforlarge language models(LLMs) typically requires substantial amounts of high-quality supervised data, which is both costly and labor-intensive to acquire. Whilesynthetic data generationhas emerged as a promising solution, existing approaches frequently suffer from factual inaccuracies, insufficient long-tail coverage, simplistic knowledge structures, and homogenized outputs. To address these challenges, we introduce GraphGen, aknowledge graph-guided framework designed for three key question-answering (QA) scenarios: atomic QA, aggregated QA, andmulti-hop QA. It begins by constructing a fine-grainedknowledge graphfrom the source text. It then identifiesknowledge gapsin LLMs using theexpected calibration errormetric, prioritizing the generation of QA pairs that target high-value, long-tail knowledge. Furthermore, GraphGen incorporatesmulti-hop neighborhood samplingto capture complex relational information and employs style-controlled generation to diversify the resulting QA data. Experimental results on knowledge-intensive tasks underclosed-book settingsdemonstrate that GraphGen outperforms conventional synthetic data methods, offering a more reliable and comprehensive solution to the data scarcity challenge in supervisedfine-tuning. The code and data are publicly available at https://github.com/open-sciencelab/GraphGen.
View arXiv pageView PDFProject pageGitHub1.17kAdd to collection
Get this paper in your agent:
hf papers read 2505\.20416
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2505.20416 in a model README.md to link it from this page.
Datasets citing this paper1
#### chenzihong/GraphGen-Data Preview• UpdatedJun 16, 2025 • 27 • 2
Spaces citing this paper1
Collections including this paper1
Similar Articles
Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation
This paper presents a synthetic data generation method for fine-tuning small LLMs to convert natural language to Cypher queries for property graphs, achieving competitive performance with large proprietary models while enabling local deployment and data sovereignty.
Enhancing Metacognitive AI: Knowledge-Graph Population with Graph-Theoretic LLM Enrichment
MetaKGEnrich is a fully automated pipeline that uses graph metrics to detect knowledge gaps in LLM applications, retrieves web evidence, and improves answer quality by 80-87% across three benchmark datasets.
SelfGraphRAG: Bridging the Supervision Gap in Graph-Based RAG with Synthetic QA Generation
SelfGraphRAG introduces a framework that generates synthetic question-answer pairs from knowledge graphs to address the supervision gap in graph-based retrieval-augmented generation, enhancing retrieval precision and reasoning performance.
Goal-Conditioned Supervised Learning for LLM Fine-Tuning
This paper proposes goal-conditioned supervised learning (GCSL) as an offline fine-tuning framework for LLMs, which treats feedback as an explicit goal and trains models via supervised learning with a novel goal formulation and natural-language goal representations. Evaluated on non-toxic generation, code generation, and recommendation, it outperforms standard offline baselines.
RSF-GLLM: Bridging the Semantic Gap in Multi-Hop Knowledge Graph QA via Recurrent Soft-Flow and Decoupled LLM Generation
This paper introduces RSF-GLLM, a framework that decouples differentiable graph reasoning from LLM generation to address the semantic gap in multi-hop knowledge graph question answering, achieving competitive performance with superior inference efficiency.