Semi-Supervised Text-Attributed Graph Distillation
Summary
Proposes STAD, a semi-supervised framework for distilling text-attributed graphs using Wasserstein distance, dual-pathway encoders, and LLM-based text synthesis to achieve a state-of-the-art performance-compression trade-off.
View Cached Full Text
Cached at: 07/24/26, 05:01 AM
# Semi-Supervised Text-Attributed Graph Distillation Source: [https://arxiv.org/html/2607.20477](https://arxiv.org/html/2607.20477) ,Samir MoustafaCeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences Universität ViennaViennaAustria[samir\.moustafa@univie\.ac\.at](https://arxiv.org/html/2607.20477v1/mailto:[email protected]),Renchi Yang†Hong Kong Baptist UniversityHong KongChina[renchi@hkbu\.edu\.hk](https://arxiv.org/html/2607.20477v1/mailto:[email protected])andTsz Nam ChanShenzhen UniversityShenzhenChina[edisonchan@szu\.edu\.cn](https://arxiv.org/html/2607.20477v1/mailto:[email protected]) ###### Abstract\. Text\-Attributed Graphs\(TAGs\) have emerged as an expressive data model for integrating graph topology with rich textual semantics\. Existing representation learning methods over TAGs suffer from severe scalability bottlenecks, particularly together withLarge Language Models\(LLMs\)\. While data distillation offers a promising data\-centric solution, existing methods fail to capture the complex interplay between graph and text modalities, struggle with the label scarcity inherent in semi\-supervised settings, and lack the ability to produce the human\-readable textual attributes required for downstream LLM\-based tasks\. To address these challenges, we proposeSTAD, a unified semi\-supervised framework guided by theWasserstein Distance\(WSD\)\. Grounded in our empirical findings on real TAGs,STADintroduces a graph\-text collaborative encoding module that utilizes dual\-pathway encoders \(graph\-aware and \-free\) within a collaborative self\-training scheme to harvest reliable pseudo\-labels and fuse complementary graph\-text features\. Furthermore, we develop a theoretically grounded WSD\-based graph sketching algorithm and a cost\-effective LLM text synthesis module, which leverages cluster\-based keyword extraction to generate coherent, human\-readable summaries for condensed nodes\. Extensive experiments on benchmark datasets demonstrate thatSTADachieves a state\-of\-the\-art performance\-compression trade\-off in terms of both GNN\- and LLM\-based downstream tasks, enabling effective and efficient TAG learning or analytics\. Data Distillation; Text\-Attributed Graphs; Semi\-Supervised Classification; Text Synthesis ††submissionid:41300footnotetext:†\\daggerCorresponding author## 1\.Introduction Text\-Attributed Graphs\(TAGs\) have emerged as a popular graph data paradigm, integrating rich semantic text with complex relations among textual entities\(Yanet al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib13); Fenget al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib20); Chenet al\.,[2024b](https://arxiv.org/html/2607.20477#bib.bib6); Zhanget al\.,[2024b](https://arxiv.org/html/2607.20477#bib.bib25)\)\. By bridging the gap between graph data and natural language, TAGs unlock the potential ofLarge Language Models\(LLMs\) for diverse cross\-domain applications\(Liet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib22); Liuet al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib26); Jinet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib23); Wanget al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib21)\), from drug discovery\(Liet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib66)\)to fraud detection\(Yanget al\.,[2025b](https://arxiv.org/html/2607.20477#bib.bib67)\), etc\. Through exploiting the synergy between two modalities in TAGs, recent hybrid frameworks combiningGraph Neural Networks\(GNNs\)\(Kipf and Welling,[2017](https://arxiv.org/html/2607.20477#bib.bib15); Veličkovićet al\.,[2018](https://arxiv.org/html/2607.20477#bib.bib16); Hamiltonet al\.,[2017](https://arxiv.org/html/2607.20477#bib.bib17)\)with LLMs have achieved remarkable results for learning tasks over TAGs\(Zhaoet al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib32); Heet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib28); Zhanget al\.,[2024a](https://arxiv.org/html/2607.20477#bib.bib31); Chenet al\.,[2024b](https://arxiv.org/html/2607.20477#bib.bib6); Zhuet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib27); Tanget al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib29); Chenet al\.,[2024a](https://arxiv.org/html/2607.20477#bib.bib30); Sunet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib33); Wuet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib24)\)\. Despite the advancements made, these approaches still suffer from a critical scalability bottleneck in practical deployment, as real\-world TAGs often scale to hundreds of thousands of nodes and edges\(Chenet al\.,[2024b](https://arxiv.org/html/2607.20477#bib.bib6)\)\. The integration with LLMs further significantly increases the computational and financial overhead demanded by these methods, which, in turn, severely limits their applicability in large\-scale industrial and scientific scenarios\. Data distillationoffers a promising data\-centric solution to this scalability challenge in large\-scale learning tasks\(Sachdeva and McAuley,[2023](https://arxiv.org/html/2607.20477#bib.bib34); Lei and Tao,[2023](https://arxiv.org/html/2607.20477#bib.bib35)\)\. Its goal is to synthesize a condensed, small\-scale dataset that preserves the training dynamics and information of the original large\-scale data\. In recent years, substantial efforts\(Jinet al\.,[2022](https://arxiv.org/html/2607.20477#bib.bib38); Yanget al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib39); Liuet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib41); Laiet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib40)\)have been made towardsgraph distillation\(a\.k\.a\. graph condensation\) andtext distillation\(Li and Li,[2021](https://arxiv.org/html/2607.20477#bib.bib56); Sucholutsky and Schonlau,[2021](https://arxiv.org/html/2607.20477#bib.bib42); Maekawaet al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib48); Taoet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib59)\), where the former seeks to condense the original graphs with nodal attributes into small ones, while the latter focuses on compressing text into concise semantic representations\. However, extending these successes to TAGs still presents unique, non\-trivial challenges\. Firstly, the naive combination of graph and text distillation fails to capture the complex interplay between graph and textual modalities\(Jinet al\.,[2022](https://arxiv.org/html/2607.20477#bib.bib38); Zhouet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib65)\)\. Graph distillation methods typically treat text as static numerical features, failing to maintain semantic consistency, while text distillation overlooks structural features\. Second, existing graph condensation methods simply construct continuous features that lack interpretability as node attributes\. But for TAGs, it is necessary to generate human\-readable textual attributes such that the distilled TAGs remain compatible with downstream LLM\-based tasks, e\.g\., prompt\-based learning or explanation generation\(Heet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib28)\)\. On top of that, real\-world TAGs often present severe label scarcity, where standard distillation techniques that often rely on abundant supervision to align gradients or distributions, falter or even fail\. To overcome the aforementioned challenges, this paper proposesSTAD\(Semi\-supervisedText\-Attributed GraphDistillation\), a unified framework designed to distill TAGs into compact, human\-readable graphs under semi\-supervised settings\. Our approach is grounded in our quantitative empirical studies, where we identify that \(i\) graph\-aware models \(GNNs\) and graph\-free models \(MLPs\) capture complementary information, and simply optimizing one modality is insufficient for effective condensation of TAGs, and \(ii\) theWasserstein Distance\(WSD\)\(Solomonet al\.,[2015a](https://arxiv.org/html/2607.20477#bib.bib49); Yanget al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib39)\)between the distributions of original and condensed graphs is strongly correlated with downstream performance\. Inspired by these observations, we first propose agraph\-text collaborative encodingmodule, which employsdual\-pathwayencoders \(graph\-aware and \-free pathways\) within a collaborative self\-training scheme as complementary learners, where pseudo\-supervision is iteratively generated for better feature fusion, thereby mitigating label scarcity and cross\-modality misalignment\. Subsequently, building on these fused features, a theoretically\-grounded graph sketching aiming at minimizing the WSD between distribution shifts between original and condensed TAGs is leveraged for condensing both graph topology, attributes, and labels\. Furthermore, we construct textual attributes for distilled graphs through ourKeywords\-based LLM text synthesis\. Instead of optimizing abstract embeddings or directly harnessing LLMs for text summarization, which is powerful but costly\(Zhanget al\.,[2025b](https://arxiv.org/html/2607.20477#bib.bib47),[a](https://arxiv.org/html/2607.20477#bib.bib44)\),STADresorts to a three\-stage pipeline, where keywords are first extracted from condensed clusters, and LLMs are then cost\-effectively harnessed to generate coherent candidate text from these keywords\. Lastly, a WSD\-validation\-guided selection strategy is adopted to identify the synthetic texts that best preserve the distributional characteristics of the original data\. In sum, our major contributions in this paper are as follows: - •We pioneer the systematic exploration of TAG distillation\. We propose the use of WSD as a reliable metric to quantify distribution shifts in TAGs, offering a theoretical basis for assessing distillation quality in graph\-text contexts\. - •We developSTADframework, which integrates graph\-text collaborative encoding to resolve label scarcity and modality misalignment and modality fusion, and WSD\-guided graph sketching and keywords\-based LLM text synthesis to create high\-quality condensed graphs and textual attributes\. - •Our extensive experiments on multiple benchmark TAG datasets demonstrate thatSTADconsistently achieves a state\-of\-the\-art performance\-compression trade\-off in both GNN\- and LLM\-based node classification tasks, compared to existing graph or text distillation solutions\. ## 2\.Related Work Graph Distillation\.Graph Distillation \(or Graph Condensation\) is a data\-centric solution to reduce the computational and storage costs of GNNs on large\-scale graphs, whose goal is to synthesize compact graphs, maintaining the key information of original graphs for competitive model performance\.GCond\(Jinet al\.,[2022](https://arxiv.org/html/2607.20477#bib.bib38)\)is a simple gradient\-matching based method which synthesizes graphs by matching the GNN’s gradients on the synthesis and original graphs\.SGDD\(Yanget al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib39)\)proposes a structure\-broadcasting scheme to preserve the original structural information in the gradient matching process\.GCDM\(Liuet al\.,[2022](https://arxiv.org/html/2607.20477#bib.bib50)\)andGDEM\(Liuet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib41)\)are distribution matching\-based methods, withGCDMaligning the receptive field distributions of synthetic and original graphs via MMD, and GDEM focusing on matching the eigenbasis and spectral distributions of the two graphs\.SFGC\(Zhenget al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib51)\)andGCSR\(Liuet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib41)\)are trajectory matching\-based distillation methods, which align the long\-term learning dynamics of GNNs on synthetic and original graphs\.CGC\(Gaoet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib55)\)andClustGDD\(Laiet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib40)\)are clustering\-based methods, withCGCconducting class partition to derive condensed node features and topology in closed form, andClustGDDoptimizes FID between synthetic and original graphs via clustering\. Text Distillation\.Text distillation is the process of synthesizing compact text datasets for NLP tasks\.TDD\(Sucholutsky and Schonlau,[2021](https://arxiv.org/html/2607.20477#bib.bib42)\)updates network parameters via gradient descent with distilled samples and soft labels, computes loss on real text, and optimizes the distilled component by loss gradients to capture text semantics, then decodes distilled text using the original word embedding dictionary\.DD4TC\(Li and Li,[2021](https://arxiv.org/html/2607.20477#bib.bib56)\)uses the gradients from training on the original texts to be back\-propagated to the distillation texts\.ASD\(Maekawaet al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib48)\)simultaneously optimizes distilled input embeddings, soft labels, and attention labels using gradient descent, and guides model parameter updates by combining task loss and transformer attention loss\.DiLM\(Maekawaet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib60)\)uses a GPT\-2 model to generate candidate distilled texts, and learn a generation probability via matching the learner’s gradients on both original and distilled texts\.DaLLME\(Taoet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib59)\)conducts k\-centroid on text embeddings and translates the clustered embedding center back to readable text via a T5\(Raffelet al\.,[2020](https://arxiv.org/html/2607.20477#bib.bib63)\)model finetuned on embedding\-text pairs constructed from the original dataset\. Although many distillation methods exist for graph and text, no dedicated scheme has been designed for TAGs\. Single\-modal distillation cannot fully exploit the other modality, motivating us to propose a collaborative graph\-text distillation scheme for TAGs\. TAG Representation Learning\.LMs \(including LLMs\), combined with traditional GNNs, have leveraged their abilities on TAG representation learning based on rich semantic information of raw texts in TAGs\. LMs can serve as TAG encoders by integrating graph topology and texts into the embedding space\(Yanet al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib13); Chenet al\.,[2024b](https://arxiv.org/html/2607.20477#bib.bib6); Zhuet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib27)\)\.ENGINE\(Zhuet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib27)\)combines the inter\-layer embeddings of LLMs for text encoding\. LMs can serve as TAG augmenters, e\.g\.,TAPE\(Heet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib28)\)makes LLMs to directly generate predictions and explanations to enrich the text of nodes\.CTGL\(Zhouet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib65)\)quantitatively compares the complementarity between graph structure and textual attributes, serially augments and encodes TAG using GNNs and LMs, and fuses graph structure and text in a contrastive way\. LMs can be the predictors to give final predictions of TAG learning tasks\(Tanget al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib29); Chenet al\.,[2024a](https://arxiv.org/html/2607.20477#bib.bib30)\)\.LLaGA\(Chenet al\.,[2024a](https://arxiv.org/html/2607.20477#bib.bib30)\)puts the serialized graph structure and node texts into an LLM to generate predictions and natural language descriptions\.\(Wuet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib24)\)gives an extensive survey of different types of LLMs’ performance on TAG node classification\. However, these methods will suffer from training efficiency problems when the TAGs scale up, which motivates the exploration of TAG distillation\. We In this work, we define LLM4TAG as large language models tailored for TAG learning tasks\(Suet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib9)\), and use this standardized terminology consistently in the experimental section\. ## 3\.Preliminaries ### 3\.1\.Notations and Problem Formulation Notations\.Let𝒢=\(𝒱,ℰ,𝒮\)\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\},\\mathcal\{S\}\)be aText\-Attributed Graph\(TAG\), where𝒱\\mathcal\{V\}denotes the set ofNNnodes \(N=\|𝒱\|N=\|\\mathcal\{V\}\|\),ℰ\\mathcal\{E\}denotes the set ofMMedges \(M=\|ℰ\|M=\|\\mathcal\{E\}\|\),𝒮=\{si\}vi∈𝒱\\mathcal\{S\}=\\\{s\_\{i\}\\\}\_\{v\_\{i\}\\in\\mathcal\{V\}\}is the collection of text sequences for nodes\. We denote the neighbor set ofviv\_\{i\}as𝒩\(vi\)\\mathcal\{N\}\(v\_\{i\}\)with degreed\(vi\)=\|𝒩\(vi\)\|d\(v\_\{i\}\)=\|\\mathcal\{N\}\(v\_\{i\}\)\|\.𝐀∈\{0,1\}N×N\\mathbf\{A\}\\in\\\{0,1\\\}^\{N\\times N\}is used to represent the adjacency matrix of𝒢\\mathcal\{G\}\(where𝐀i,j=1\\mathbf\{A\}\_\{i,j\}=1if\(vi,vj\)∈ℰ\(v\_\{i\},v\_\{j\}\)\\in\\mathcal\{E\}, otherwise𝐀i,j=0\\mathbf\{A\}\_\{i,j\}=0\) and𝐃∈ℝN×N\\mathbf\{D\}\\in\\mathbb\{R\}^\{N\\times N\}symbolizes the diagonal degree matrix\. The normalized adjacency matrix is𝐀~=𝐃−12𝐀𝐃−12\\tilde\{\\mathbf\{A\}\}=\\mathbf\{D\}^\{\-\\frac\{1\}\{2\}\}\\mathbf\{A\}\\mathbf\{D\}^\{\-\\frac\{1\}\{2\}\}\. Let𝐗∈ℝN×D\\mathbf\{X\}\\in\\mathbb\{R\}^\{N\\times D\}be the text embeddings of nodes, wherein𝐗i\\mathbf\{X\}\_\{i\}\(theii\-th row\) denotes the vectorized encoding ofsis\_\{i\}for nodevi∈𝒱v\_\{i\}\\in\\mathcal\{V\}via a text encoderg:𝒮→𝐗g:\\mathcal\{S\}\\to\\mathbf\{X\}\. By default, we adopt a pretrained Sentence\-BERT \(SBERT\)\(Reimers and Gurevych,[2019](https://arxiv.org/html/2607.20477#bib.bib37)\)asggthroughout this paper\. Let𝒴\\mathcal\{Y\}be the task\-specific class set, andK=\|𝒴\|K=\|\\mathcal\{Y\}\|is the number of node classes\. We denote by𝐘∈\{0,1\}N×K\\mathbf\{Y\}\\in\\\{0,1\\\}^\{N\\times K\}the ground\-truth label matrix of𝒢\\mathcal\{G\}where𝐘i,k=1\\mathbf\{Y\}\_\{i,k\}=1ifyk∈𝒴y\_\{k\}\\in\\mathcal\{Y\}\(1≤k≤K1\\leq k\\leq K\) isviv\_\{i\}’s ground\-truth class label and0otherwise\. Under semi\-supervised settings, only a small set of nodes𝒱tr⊆𝒱\\mathcal\{V\}\_\{\\text\{tr\}\}\\subseteq\\mathcal\{V\}\(\|𝒱tr\|≪N\|\\mathcal\{V\}\_\{\\text\{tr\}\}\|\\ll N\) are labeled, referred to as training samples\. Problem Formulation\.The overarching goal ofText\-Attributed Graph Distillation\(TAGD\) is to distill a compact, high\-fidelity compressed TAG𝒢′=\(𝒱′,ℰ′,𝒮′\)\\mathcal\{G\}^\{\\prime\}=\(\\mathcal\{V\}^\{\\prime\},\\mathcal\{E\}^\{\\prime\},\\mathcal\{S\}^\{\\prime\}\)from the original larger𝒢\\mathcal\{G\}, where𝒱′\\mathcal\{V\}^\{\\prime\}\(N′=\|𝒱′\|≪NN^\{\\prime\}=\|\\mathcal\{V\}^\{\\prime\}\|\\ll N\) andℰ′\\mathcal\{E\}^\{\\prime\}\(M′=\|ℰ′\|≪MM^\{\\prime\}=\|\\mathcal\{E\}^\{\\prime\}\|\\ll M\) stand for the node and edge sets of𝒢′\\mathcal\{G\}^\{\\prime\}, respectively, and𝒮′=\{si′\}vi′∈𝒱′\\mathcal\{S\}^\{\\prime\}=\\\{s^\{\\prime\}\_\{i\}\\\}\_\{v^\{\\prime\}\_\{i\}\\in\\mathcal\{V\}^\{\\prime\}\}represents human\-readable text sequences for condensed nodes therein\. Let𝐀′\\mathbf\{A\}^\{\\prime\}and𝐘′\\mathbf\{Y\}^\{\\prime\}be the adjacency matrix and node labels of𝒢′\\mathcal\{G\}^\{\\prime\}, respectively\. Formally, the objective of the TAGD problem can be formulated as: \(1\)min𝒢′f\(ℳϕ𝒢′\(𝐀,𝒮\),𝐘\)s\.t\.ϕ𝒢′=argminϕf\(ℳϕ\(𝐀′,𝒮′\),𝐘′\),\\begin\{gathered\}\\min\_\{\\mathcal\{G\}^\{\\prime\}\}\\,\{f\}\\left\(\\mathcal\{M\}\_\{\\boldsymbol\{\\phi\}\_\{\\mathcal\{G\}^\{\\prime\}\}\}\(\\mathbf\{A\},\\mathcal\{S\}\),\\mathbf\{Y\}\\right\)\\\\ \\text\{s\.t\.\}\\,\\boldsymbol\{\\phi\}\_\{\\mathcal\{G\}^\{\\prime\}\}=\\arg\\min\_\{\\boldsymbol\{\\phi\}\}f\\left\(\\mathcal\{M\}\_\{\\boldsymbol\{\\phi\}\}\(\\mathbf\{A\}^\{\\prime\},\\mathcal\{S\}^\{\\prime\}\),\\mathbf\{Y\}^\{\\prime\}\\right\),\\\\ \\end\{gathered\}whereℳϕ\\mathcal\{M\}\_\{\\boldsymbol\{\\phi\}\}denotes a general model parameterized byϕ\\boldsymbol\{\\phi\}, andf\(⋅\)f\(\\cdot\)is the task\-specific criteria\. In this paper, we consider the semi\-supervised node classification setting and thedistillation ratiois defined asr=N′\|𝒱tr\|≤1r=\\frac\{N^\{\\prime\}\}\{\|\\mathcal\{V\}\_\{\\text\{tr\}\}\|\}\\leq 1\. ### 3\.2\.2\-Wasserstein Distance \(WSD\) The2\-Wasserstein Distance\(WSD\) inoptimal transport\(Solomonet al\.,[2015b](https://arxiv.org/html/2607.20477#bib.bib52)\)is a classic distance function used to measure shift between two probability distributions𝒫\\mathcal\{P\}and𝒬\\mathcal\{Q\}: WSD\(𝒫,𝒬\)=infγ∈Γ\(𝒫,𝒬\)∫𝒳×𝒳‖𝐩−𝐪‖dγ\(𝐩,𝐪\)\\textsf\{WSD\}\(\\mathcal\{P\},\\mathcal\{Q\}\)=\\inf\_\{\\gamma\\in\\Gamma\(\\mathcal\{P\},\\mathcal\{Q\}\)\}\\int\_\{\\mathcal\{X\}\\times\\mathcal\{X\}\}\\\|\\mathbf\{p\}\-\\mathbf\{q\}\\\|\\,\\text\{d\}\\gamma\(\\mathbf\{p\},\\mathbf\{q\}\)whereΓ\(𝒫,𝒬\)\\Gamma\(\\mathcal\{P\},\\mathcal\{Q\}\)is the set of joint distributions with marginals𝒫\\mathcal\{P\}and𝒬\\mathcal\{Q\},𝒳\\mathcal\{X\}is the feature space, and‖𝐩−𝐪‖\\\|\\mathbf\{p\}\-\\mathbf\{q\}\\\|is the Euclidean distance\.SGDD\(Yanget al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib39)\)uses WSD as an alternative to the Laplacian Energy Distribution shift coefficient for graph structure optimization\.ClustGDD\(Laiet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib40)\)finds a correlation betweenFréchet Inception Distance\(FID\) and accuracy of GNNs on original and synthetic graphs, where FID between two multivariate Gaussian distributionsN\(μ𝒫,Σ𝒫\)N\(\\mu\_\{\\mathcal\{P\}\},\\Sigma\_\{\\mathcal\{P\}\}\)andN\(μ𝒬,Σ𝒬\)N\(\\mu\_\{\\mathcal\{Q\}\},\\Sigma\_\{\\mathcal\{Q\}\}\), given by FID\(𝒫,𝒬\)=‖μ𝒫−μ𝒬‖22\+Tr\(Σ𝒫\+Σ𝒬−2\(Σ𝒫Σ𝒬\)1/2\),\\text\{FID\}\(\\mathcal\{P\},\\mathcal\{Q\}\)=\\\|\\mu\_\{\\mathcal\{P\}\}\-\\mu\_\{\\mathcal\{Q\}\}\\\|\_\{2\}^\{2\}\+\\operatorname\{Tr\}\\big\(\\Sigma\_\{\\mathcal\{P\}\}\+\\Sigma\_\{\\mathcal\{Q\}\}\-2\(\\Sigma\_\{\\mathcal\{P\}\}\\Sigma\_\{\\mathcal\{Q\}\}\)^\{1/2\}\\big\),corresponds to a special case of WSD when the underlying distributions are multivariate Gaussians\. Next, we will compute WSD between original and synthetic textual attributes, graph structures under various distillation ratios, and comprehensively investigate the correlations between WSD and the model performance\. Figure 1\.Complementary coverage of GA vs\. GF on TAGs\.Figure 2\.WSD and node classification accuracy on condensed TAGs vs\. distillation ratio\. ### 3\.3\.Preliminary Studies We empirically investigate the TAGD task over four real datasets \(Cora,Photo,History, andWikiCS\) pertaining to two aspects before entering into the algorithmic design of our proposedSTAD\. Graph\-Aware vs\. Graph\-Free\.Since both graph structures and textual attributes encode rich semantics in TAGs, a natural question is their roles in the downstream task, i\.e\., semi\-supervised node classification\. To this end, we compare thegraph\-aware\(GA\) model, i\.e\., GNNs using𝐗\\mathbf\{X\}and𝐀\\mathbf\{A\}for node classification, against thegraph\-free\(GF\) model, i\.e\., MLPs using pure text attribute of𝒮\\mathcal\{S\}\. Let𝒞GA\\mathcal\{C\}\_\{\\text\{GA\}\}and𝒞GF\\mathcal\{C\}\_\{\\text\{GF\}\}be the sets of node samples whose labels are correctly predicted by the GA and GF models, respectively\. Accordingly,𝒞GA∪𝒞GF¯\\overline\{\\mathcal\{C\}\_\{\\text\{GA\}\}\\cup\\mathcal\{C\}\_\{\\text\{GF\}\}\}stands for the complementary set containing node samples whose labels cannot be correctly predicted by both models\. Fig\.[1](https://arxiv.org/html/2607.20477#S3.F1)reports the proportions of different sets on the four tested datasets\. It can be observed that although𝒞GA\\mathcal\{C\}\_\{\\text\{GA\}\}and𝒞GF\\mathcal\{C\}\_\{\\text\{GF\}\}have a big overlap \(43\.9%43\.9\\%\-64\.4%64\.4\\%\), their differences are still considerable\. Firstly, this observation underscores the indispensability of both structural and textual information and their complementary strengths in TAGs\. Additionally, the𝒞GF∖𝒞GA\\mathcal\{C\}\_\{\\text\{GF\}\}\\setminus\\mathcal\{C\}\_\{\\text\{GA\}\}part implies that using GNNs with𝐗\\mathbf\{X\}in the GA model fails to fully capture the textual semantics in𝒮\\mathcal\{S\}, necessitating new feature encoders\. Relation between the WSD and Accuracy\.Next, we empirically analyze the distributional shifts between attribute matrices of the original TAG𝒢\\mathcal\{G\}and the condensed TAG𝒢′\\mathcal\{G\}^\{\\prime\}, their topological structures \(i\.e\., adjacency matrices\), their fused representations via GNNs, as well as their relations to node classification performance\. Fig\.[2](https://arxiv.org/html/2607.20477#S3.F2)displays the WSD values between𝐗\\mathbf\{X\}and𝐗′\\mathbf\{X\}^\{\\prime\},𝐀\\mathbf\{A\}and𝐀′\{\\mathbf\{A\}\}^\{\\prime\},GNN\(𝐀,𝐗\)\\textsf\{GNN\}\(\\mathbf\{A\},\\mathbf\{X\}\)andGNN\(𝐀′,𝐗′\)\\textsf\{GNN\}\(\{\\mathbf\{A\}\}^\{\\prime\},\\mathbf\{X\}^\{\\prime\}\), and node classification accuracies \(Acc\) when varying the distillation ratios, i\.e\.,r∈\[0\.05,0\.25,0\.5,1\.0\]r\\in\[0\.05,0\.25,0\.5,1\.0\]on the four real datasets \(Cora,Photo,History, andWikiCS\)\. The empirical results reveal a key trend\.WSD\(𝐗,𝐗′\)\\textsf\{WSD\}\(\{\\mathbf\{X\}\},\{\\mathbf\{X\}\}^\{\\prime\}\),WSD\(𝐀,𝐀′\)\\textsf\{WSD\}\(\{\\mathbf\{A\}\},\{\\mathbf\{A\}\}^\{\\prime\}\), andWSD\(GNN\(𝐀,𝐗\),GNN\(𝐀′,𝐗′\)\)\\textsf\{WSD\}\(\\textsf\{GNN\}\(\\mathbf\{A\},\\mathbf\{X\}\),\\textsf\{GNN\}\(\{\\mathbf\{A\}\}^\{\\prime\},\\mathbf\{X\}^\{\\prime\}\)\), increase constantly asrris decreased \(i\.e\., their distributional shifts enlarge\), while the classification performance keeps going down\. This observation suggests that maximizing the node classification performance over the TAG𝒢\\mathcal\{G\}in Eq\. \([1](https://arxiv.org/html/2607.20477#S3.E1)\) can be equivalently transformed into the problem of reducing the discrepancy between𝒢\\mathcal\{G\}and𝒢′\\mathcal\{G\}^\{\\prime\}in terms of both graph structures and attributes, e\.g\., minimizingWSD\(𝐗,𝐗′\)\\textsf\{WSD\}\(\\mathbf\{X\},\\mathbf\{X\}^\{\\prime\}\)andWSD\(𝐀,𝐀′\)\\textsf\{WSD\}\(\\mathbf\{A\},\\mathbf\{A\}^\{\\prime\}\)jointly\. To reduce bias, we compute the mean and std of the results from different baseline distillation methods\. ## 4\.Methodology This section presents ourSTADmethod for TAGD under semi\-supervised settings\. As illustrated in Fig\.[3](https://arxiv.org/html/2607.20477#S4.F3),STADincludes three key modules for feature encoding in TAGs, graph sketching, and synthesis of textual attributes\. We begin with elucidating ourgraph\-text collaborative encodingmodel for generating node feature vectors𝐇\\mathbf\{H\}in §[4\.1](https://arxiv.org/html/2607.20477#S4.SS1), followed by constructing the sketching matrix𝐒\\mathbf\{S\}for graph condensation in §[4\.2](https://arxiv.org/html/2607.20477#S4.SS2)\. §[4\.3](https://arxiv.org/html/2607.20477#S4.SS3)delineates our cost\-effective approach for synthesizing textual attributes of nodes in𝒢′\\mathcal\{G\}^\{\\prime\}with LLMs\. For the interest of space, we defer the complete algorithm and provide the cost analysis to Appendix[A](https://arxiv.org/html/2607.20477#A1)\. ### 4\.1\.Graph\-Text Collaborative Encoding As per our analysis in §[3\.3](https://arxiv.org/html/2607.20477#S3.SS3), we propose to adoptdual\-pathway encodersand leveragecollaborative self\-trainingto unearth the complementary information underlying the graph and text features for enhanced feature encoding\. Dual\-Pathway Encoders\.Since the GNN encoder will lose distinct features underlying𝒮\\mathcal\{S\}, as revealed in §[3\.3](https://arxiv.org/html/2607.20477#S3.SS3),STADincludes two paths of encoders: a GA \(or graph\) encoder via GCNs, and a GF \(or text\) encoder via an MLP layer\. Specifically, \(2\)𝐇GA=GCN\(𝐀,𝐗\),𝐇GF=MLP\(𝐗\),\\mathbf\{H\}^\{\\text\{\\tiny GA\}\}=\\textsf\{GCN\}\(\\mathbf\{A\},\\mathbf\{X\}\),\\ \\mathbf\{H\}^\{\\text\{\\tiny GF\}\}=\\textsf\{MLP\}\(\\mathbf\{X\}\),whoseii\-th rows𝐇iGA∈ℝh\\mathbf\{H\}\_\{i\}^\{\\text\{\\tiny GA\}\}\\in\\mathbb\{R\}^\{h\}and𝐇iGF∈ℝh\\mathbf\{H\}\_\{i\}^\{\\text\{\\tiny GF\}\}\\in\\mathbb\{R\}^\{h\}signify the feature vectors of nodeviv\_\{i\}output by the dual\-pathway encoders\. Subsequently,STADcalculates attention weights by \(3\)αiGA=σ\(𝐖att⋅𝐇iGA\+𝐛att\),αiGF=1−αiGA,\\alpha^\{\\text\{\\tiny GA\}\}\_\{i\}=\\sigma\\left\(\\mathbf\{W\}\_\{\\text\{att\}\}\\cdot\\mathbf\{H\}\_\{i\}^\{\\text\{\\tiny GA\}\}\+\\mathbf\{b\}\_\{\\text\{att\}\}\\right\),\\ \\alpha^\{\\text\{\\tiny GF\}\}\_\{i\}=1\-\\alpha^\{\\text\{\\tiny GA\}\}\_\{i\},where𝐖att∈ℝh\\mathbf\{W\}\_\{\\text\{att\}\}\\in\\mathbb\{R\}^\{h\}and𝐛att∈ℝ\\mathbf\{b\}\_\{\\text\{att\}\}\\in\\mathbb\{R\}are learnable parameters, andσ\(⋅\)\\sigma\(\\cdot\)denotes the sigmoid function\. Accordingly, the fused predictedlabel distributions\(or soft labels\) of each nodeviv\_\{i\}can be derived via a weighted element\-wise summation \(⊙\\odotis the element\-wise multiplication\) of the feature vectors from both encoders: \(4\)𝐏i=softmax\(αiGA⊙𝐇iGA\+αiGF⊙𝐇iGF\)\.\\mathbf\{P\}\_\{i\}=\\textsf\{softmax\}\(\\alpha^\{\\text\{\\tiny GA\}\}\_\{i\}\\odot\\mathbf\{H\}\_\{i\}^\{\\text\{\\tiny GA\}\}\+\\alpha^\{\\text\{\\tiny GF\}\}\_\{i\}\\odot\\mathbf\{H\}\_\{i\}^\{\\text\{\\tiny GF\}\}\)\.Inspired by the adaptive mechanism in\(Boet al\.,[2021](https://arxiv.org/html/2607.20477#bib.bib11); Luanet al\.,[2022](https://arxiv.org/html/2607.20477#bib.bib10)\),STADlearns attention weights to automatically fuse graph and text features\. In particular, the soft labels derived from the GA and GF encoders can be represented as follows: 𝐏iGA=softmax\(𝐇iGA\),𝐏iGF=softmax\(𝐇iGF\)\.\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}\_\{i\}=\\textsf\{softmax\}\(\\mathbf\{H\}\_\{i\}^\{\\text\{\\tiny GA\}\}\),\\ \\mathbf\{P\}^\{\\text\{\\tiny GF\}\}\_\{i\}=\\textsf\{softmax\}\(\\mathbf\{H\}\_\{i\}^\{\\text\{\\tiny GF\}\}\)\. Figure 3\.The Overview ofSTADCollaborative Self\-Training\.To fully leverage scarce labeled data, we introducecollaborative self\-training\(CoST\) that treats GA and GF encoders as complementary learners, as they capture distinct features predictive of different labels\. Let𝐘^\\hat\{\\mathbf\{Y\}\}be the label matrix consisting of both true labels and hard pseudo\-labels obtained based on label distribution𝐏\\mathbf\{P\}, and𝐘^GA\\hat\{\\mathbf\{Y\}\}^\{\\text\{\\tiny GA\}\}and𝐘^GF\\hat\{\\mathbf\{Y\}\}^\{\\text\{\\tiny GF\}\}are counterparts from𝐏GA\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}and𝐏GF\\mathbf\{P\}^\{\\text\{\\tiny GF\}\}, respectively\. CoST trains the model with such predictions and pseudo\-labels by jointly optimizing the following composite loss: \(5\)ℒ\\displaystyle\\mathcal\{L\}=CE\(𝐏,𝐘^\)\+CE\(𝐏GF,𝐘^GF\)\+CE\(𝐏GA,𝐘^GA\),\\displaystyle=\\textsf\{CE\}\(\\mathbf\{P\},\\hat\{\\mathbf\{Y\}\}\)\+\\textsf\{CE\}\(\\mathbf\{P\}^\{\\text\{\\tiny GF\}\},\\hat\{\\mathbf\{Y\}\}^\{\\text\{\\tiny GF\}\}\)\+\\textsf\{CE\}\(\\mathbf\{P\}^\{\\text\{\\tiny GA\}\},\\hat\{\\mathbf\{Y\}\}^\{\\text\{\\tiny GA\}\}\),whereCE\(⋅,⋅\)\\textsf\{CE\}\(\\cdot,\\cdot\)denotes the cross\-entropy loss\. Next, we elaborate on the constructions of𝐘^\\hat\{\\mathbf\{Y\}\},𝐘^GA\\hat\{\\mathbf\{Y\}\}^\{\\text\{\\tiny GA\}\}, and𝐘^GF\\hat\{\\mathbf\{Y\}\}^\{\\text\{\\tiny GF\}\}\. Unlike conventional self\-training\(Lee and others,[2013](https://arxiv.org/html/2607.20477#bib.bib69); Liet al\.,[2018](https://arxiv.org/html/2607.20477#bib.bib70)\), where the pseudo\-labels from a single model are prone to error accumulation, CoST employs aconsensus verificationbetween GA and GF modules to filter uncertain pseudo\-labels\. For ease of exposition, we represent by𝒰\\mathcal\{U\},𝒰GA\\mathcal\{U\}^\{\\text\{\\tiny GA\}\}, and𝒰GF\\mathcal\{U\}^\{\\text\{\\tiny GF\}\}the labeled node sets of𝐘^\\hat\{\\mathbf\{Y\}\},𝐘^GA\\hat\{\\mathbf\{Y\}\}^\{\\text\{\\tiny GA\}\}, and𝐘^GF\\hat\{\\mathbf\{Y\}\}^\{\\text\{\\tiny GF\}\}, respectively\. Initially,𝒰=𝒰GA=𝒰GF=𝒱tr\\mathcal\{U\}=\\mathcal\{U\}^\{\\text\{\\tiny GA\}\}=\\mathcal\{U\}^\{\\text\{\\tiny GF\}\}=\\mathcal\{V\}\_\{\\text\{tr\}\}\. We iteratively expand these labeled node sets by identifying pseudo\-labeled data from𝒱∖𝒰\\mathcal\{V\}\\setminus\\mathcal\{U\}\. Specifically, in each epoch, we obtain a setℛ\\mathcal\{R\}of candidate nodes for pseudo\-labeling based on the GA pathway𝐏GA\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}by selecting theρ\\rhoproportion of unlabeled nodes with the highestconfidence scoreswhere the confidence score ofviv\_\{i\}is denoted bymax1≤k≤K𝐏i,kGA\\max\_\{1\\leq k\\leq K\}\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}\_\{i,k\}\. Notice that the value ofρ\\rhois incremented by 0\.1 per epoch until reaching a maximum of 1\.0 as in\(Karisani,[2023](https://arxiv.org/html/2607.20477#bib.bib46)\)\. Subsequently, the candidate setℛ\\mathcal\{R\}is partitioned into aconsensussetℛCS\\mathcal\{R\}^\{\\text\{\\tiny CS\}\}and adisagreementsetℛDS\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}by the predictions in the GF pathway𝐏GF\\mathbf\{P\}^\{\\text\{\\tiny GF\}\}: ℛCS=\{vi∈ℛ\|yiGA=yiGF\},ℛDS=\{vi∈ℛ\|yiGA≠yiGF\}=ℛ∖ℛCS\\mathcal\{R\}^\{\\text\{\\tiny CS\}\}=\\\{v\_\{i\}\\in\\mathcal\{R\}\|y^\{\\text\{\\tiny GA\}\}\_\{i\}=y^\{\\text\{\\tiny GF\}\}\_\{i\}\\\},\\ \\mathcal\{R\}^\{\\text\{\\tiny DS\}\}=\\\{v\_\{i\}\\in\\mathcal\{R\}\|y^\{\\text\{\\tiny GA\}\}\_\{i\}\\neq y^\{\\text\{\\tiny GF\}\}\_\{i\}\\\}=\\mathcal\{R\}\\setminus\\mathcal\{R\}^\{\\text\{\\tiny CS\}\}whereyiGA=argmax1≤k≤K𝐏i,kGAy^\{\\text\{\\tiny GA\}\}\_\{i\}=\\underset\{1\\leq k\\leq K\}\{\\operatorname\{arg\}\\,\\operatorname\{max\}\}\\;\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}\_\{i,k\}andyiGF=argmax1≤k≤K𝐏i,kGFy^\{\\text\{\\tiny GF\}\}\_\{i\}=\\underset\{1\\leq k\\leq K\}\{\\operatorname\{arg\}\\,\\operatorname\{max\}\}\\;\\mathbf\{P\}^\{\\text\{\\tiny GF\}\}\_\{i,k\}represent the predicted labels forviv\_\{i\}by both pathways\. Intuitively, in the consensus setℛCS\\mathcal\{R\}^\{\\text\{\\tiny CS\}\}, both pathways produce the same label predictions, and thus, they can be directly used as pseudo\-labels, i\.e\., 𝒰=𝒰∪ℛCS,𝒰GA=𝒰GA∪ℛCS,and𝒰GF=𝒰GF∪ℛCS\.\\mathcal\{U\}=\\mathcal\{U\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny CS\}\},\\mathcal\{U\}^\{\\text\{\\tiny GA\}\}=\\mathcal\{U\}^\{\\text\{\\tiny GA\}\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny CS\}\},\\text\{and\}\\ \\mathcal\{U\}^\{\\text\{\\tiny GF\}\}=\\mathcal\{U\}^\{\\text\{\\tiny GF\}\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny CS\}\}\.As for the disagreement setℛDS\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}, we further divide it into two subsetsℛGADS\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GA\}\}andℛGFDS\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GF\}\}, where each nodeviv\_\{i\}in the former set has a higher confidence scoreωGA=max1≤k≤K𝐏i,kGA\\omega^\{\\text\{\\tiny GA\}\}=\\max\_\{1\\leq k\\leq K\}\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}\_\{i,k\}from the GA pathway compared to the GF one \(ωGA=max1≤k≤K𝐏i,kGA\\omega^\{\\text\{\\tiny GA\}\}=\\max\_\{1\\leq k\\leq K\}\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}\_\{i,k\}\), while the opposite is true for the latter\. The pseudo\-labels of the nodes inℛGADS\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GA\}\}andℛGFDS\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GF\}\}are thus determined by𝐏GA\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}and𝐏GF\\mathbf\{P\}^\{\\text\{\\tiny GF\}\}, respectively\. Accordingly, the labeled node sets are expanded as follows: 𝒰=𝒰∪ℛGADS∪ℛGFDS,𝒰GA=𝒰GA∪ℛGADS,and𝒰GF=𝒰GF∪ℛGFDS\.\\mathcal\{U\}=\\mathcal\{U\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GA\}\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GF\}\},\\mathcal\{U\}^\{\\text\{\\tiny GA\}\}=\\mathcal\{U\}^\{\\text\{\\tiny GA\}\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GA\}\},\\ \\text\{and\}\\ \\mathcal\{U\}^\{\\text\{\\tiny GF\}\}=\\mathcal\{U\}^\{\\text\{\\tiny GF\}\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GF\}\}\. ### 4\.2\.WSD\-based Graph Sketching After the obtainment of soft labels𝐏\\mathbf\{P\}, the next step inSTADis to construct asparsesketching matrix𝐒∈ℝN×N′\\mathbf\{S\}\\in\\mathbb\{R\}^\{N\\times N^\{\\prime\}\}such that the sketches of𝐀\\mathbf\{A\},𝐗\\mathbf\{X\}, and label matrix of𝒢\\mathcal\{G\}can be used as the adjacency matrix𝐀′∈ℝN′×N′\\mathbf\{A\}^\{\\prime\}\\in\\mathbb\{R\}^\{N^\{\\prime\}\\times N^\{\\prime\}\}, attribute matrix𝐗′∈ℝN′×D\\mathbf\{X\}^\{\\prime\}\\in\\mathbb\{R\}^\{N^\{\\prime\}\\times D\}, and label matrix𝐘′∈ℝN′\\mathbf\{Y\}^\{\\prime\}\\in\\mathbb\{R\}^\{N^\{\\prime\}\}of the condensed TAG𝒢′\\mathcal\{G\}^\{\\prime\}, respectively\. More precisely, \(6\)𝐀′=𝐒⊤𝐀𝐒,𝐗′=𝐒⊤𝐗\.\\mathbf\{A\}^\{\\prime\}=\\mathbf\{S\}^\{\\top\}\\mathbf\{A\}\\mathbf\{S\},\\ \\mathbf\{X\}^\{\\prime\}=\\mathbf\{S\}^\{\\top\}\\mathbf\{X\}\.and for each nodevi∈𝒢′v\_\{i\}\\in\\mathcal\{G\}^\{\\prime\}, its synthetic label is constructed as follows: \(7\)𝐘i,k′=\{1ifk=argmax1≤k≤K\(𝐒⊤𝐏\)i,:0otherwise\.\\mathbf\{Y\}^\{\\prime\}\_\{i,k\}=\\begin\{cases\}1&\\text\{if \}k=\\underset\{1\\leq k\\leq K\}\{\\operatorname\{arg\}\\,\\operatorname\{max\}\}\\;\{\(\\mathbf\{S\}^\{\\top\}\\mathbf\{P\}\)\}\_\{i,:\}\\\\ 0&\\text\{otherwise\}\\end\{cases\}\. Akin to the count\-sketch\(Clarkson and Woodruff,[2017](https://arxiv.org/html/2607.20477#bib.bib45)\), we impose aone\-hot constraintover𝐒\\mathbf\{S\}for sparsity, wherein each row in𝐒\\mathbf\{S\}contains a single non\-zero entry indicating the hashing or sampling of the node into a bucket\. Sketching Objective\.Inspired by our empirical observations in §[3\.3](https://arxiv.org/html/2607.20477#S3.SS3), a desired𝒢′\\mathcal\{G\}^\{\\prime\}obtained by Eq\. \([6](https://arxiv.org/html/2607.20477#S4.E6)\) should have small WSD distances from𝒢\\mathcal\{G\}in terms of textual attributes and graph topology, leading to the following objectives: \(8\)min𝐒WSD\(f\(𝐀\),f\(𝐀′\)\),andmin𝐒WSD\(𝐗,𝐗′\),\\min\_\{\\mathbf\{S\}\}\{\\textsf\{WSD\}\(f\(\\mathbf\{A\}\),f\(\\mathbf\{A\}^\{\\prime\}\)\)\},\\ \\text\{and\}\\ \\min\_\{\\mathbf\{S\}\}\{\\textsf\{WSD\}\(\\mathbf\{X\},\\mathbf\{X\}^\{\\prime\}\)\},wheref\(⋅\)f\(\\cdot\)stands for an embedding function that maps𝐀i∈ℝN\\mathbf\{A\}\_\{i\}\\in\\mathbb\{R\}^\{N\}and𝐀i′∈ℝN′\\mathbf\{A\}^\{\\prime\}\_\{i\}\\in\\mathbb\{R\}^\{N^\{\\prime\}\}to vectors of the same dimensions\. Since Fig\.[2](https://arxiv.org/html/2607.20477#S3.F2)shows thatWSD\(GNN\(𝐀,𝐗\),GNN\(𝐀′,𝐗′\)\)\\textsf\{WSD\}\(\\textsf\{GNN\}\(\\mathbf\{A\},\\mathbf\{X\}\),\\textsf\{GNN\}\(\\mathbf\{A\}^\{\\prime\},\\mathbf\{X\}^\{\\prime\}\)\)exhibits the same trends as that of the single modality, and Fig\.[1](https://arxiv.org/html/2607.20477#S3.F1)uncovers thatGNN\(𝐀,𝐗\)\\textsf\{GNN\}\(\\mathbf\{A\},\\mathbf\{X\}\)fails to fully capture the features inMLP\(𝐗\)\\textsf\{MLP\}\(\\mathbf\{X\}\), we thus transform Eq\. \([8](https://arxiv.org/html/2607.20477#S4.E8)\) into the joint minimization ofWSD\(GNN\(𝐀,𝐗\),GNN\(𝐀′,𝐗′\)\)\\textsf\{WSD\}\(\\textsf\{GNN\}\(\\mathbf\{A\},\\mathbf\{X\}\),\\textsf\{GNN\}\(\\mathbf\{A\}^\{\\prime\},\\mathbf\{X\}^\{\\prime\}\)\)andWSD\(MLP\(𝐗\),MLP\(𝐗′\)\)\\textsf\{WSD\}\(\\textsf\{MLP\}\(\\mathbf\{X\}\),\\textsf\{MLP\}\(\\mathbf\{X\}^\{\\prime\}\)\)\. ###### Lemma 4\.1\. For any matrix𝐌\\mathbf\{M\}, the minimization ofWSD\(𝐌,𝐒⊤𝐌\)\\textsf\{\{WSD\}\}\(\\mathbf\{M\},\\mathbf\{S\}^\{\\top\}\\mathbf\{M\}\)is equivalent tomin𝐒1N∑k=1N′∑i=1N𝟙𝐒k,i≠0⋅‖𝐌i−∑j=1N𝐒k,j𝐌j‖22\\min\_\{\\mathbf\{S\}\}\{\\frac\{1\}\{N\}\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\sum\_\{i=1\}^\{N\}\\mathbb\{1\}\_\{\\mathbf\{S\}\_\{k,i\}\\neq 0\}\\cdot\\left\\\|\\mathbf\{M\}\_\{i\}\-\\sum\_\{j=1\}^\{N\}\\mathbf\{S\}\_\{k,j\}\\mathbf\{M\}\_\{j\}\\right\\\|\_\{2\}^\{2\}\}\. Since𝐒\\mathbf\{S\}has a single non\-zero entry per row, the sketching matrix𝐒\\mathbf\{S\}can be represented as a set of disjoint clusters\{𝒞1,𝒞2,…,𝒞N′\}\\\{\\mathcal\{C\}\_\{1\},\\mathcal\{C\}\_\{2\},\\ldots,\\mathcal\{C\}\_\{N^\{\\prime\}\}\\\}where𝒞k=\{vi∈𝒱\|𝐒k,i\>0\}\\mathcal\{C\}\_\{k\}=\\\{v\_\{i\}\\in\\mathcal\{V\}\|\\ \\mathbf\{S\}\_\{k,i\}\>0\\\}\. Accordingly,𝟙𝐒k,i≠0\\mathbb\{1\}\_\{\\mathbf\{S\}\_\{k,i\}\\neq 0\}can be seen as a cluster indicator and∑j=1N𝐒k,j𝐌j=∑vj∈𝒞k𝐒k,j𝐌j\\sum\_\{j=1\}^\{N\}\\mathbf\{S\}\_\{k,j\}\\mathbf\{M\}\_\{j\}=\\sum\_\{v\_\{j\}\\in\\mathcal\{C\}\_\{k\}\}\{\\mathbf\{S\}\_\{k,j\}\\mathbf\{M\}\_\{j\}\}can be perceived as the weighted centroid of the cluster𝒞k\\mathcal\{C\}\_\{k\}\. The objective in Lemma[4\.1](https://arxiv.org/html/2607.20477#S4.Thmtheorem1)can be rewritten as min\{𝒞1,𝒞2,…,𝒞N′\}1N∑k=1N′∑vi∈𝒞k‖𝐌i−∑vj∈𝒞k𝐒k,j𝐌j‖22\.\\min\_\{\\\{\\mathcal\{C\}\_\{1\},\\mathcal\{C\}\_\{2\},\\ldots,\\mathcal\{C\}\_\{N^\{\\prime\}\}\\\}\}\{\\frac\{1\}\{N\}\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\sum\_\{v\_\{i\}\\in\\mathcal\{C\}\_\{k\}\}\\left\\\|\\mathbf\{M\}\_\{i\}\-\\sum\_\{v\_\{j\}\\in\\mathcal\{C\}\_\{k\}\}\\mathbf\{S\}\_\{k,j\}\\mathbf\{M\}\_\{j\}\\right\\\|\_\{2\}^\{2\}\}\.In other words, Lemma[4\.1](https://arxiv.org/html/2607.20477#S4.Thmtheorem1)111The proof can be found in Appendix[B](https://arxiv.org/html/2607.20477#A2)implies that the construction of sketching matrix𝐒\\mathbf\{S\}towards optimizing our above WSD\-based objective can be equivalently perceived as the clustering over feature vectorsGCN\(𝐀,𝐗\)=𝐇GA\\textsf\{GCN\}\(\\mathbf\{A\},\\mathbf\{X\}\)=\\mathbf\{H\}^\{\\text\{\\tiny GA\}\}andMLP\(𝐗\)=𝐇GF\\textsf\{MLP\}\(\\mathbf\{X\}\)=\\mathbf\{H\}^\{\\text\{\\tiny GF\}\}\. According to Eq\. \([4](https://arxiv.org/html/2607.20477#S4.E4)\),𝐏\\mathbf\{P\}adaptively combines𝐇GA\\mathbf\{H\}^\{\\text\{\\tiny GA\}\}and𝐇GF\\mathbf\{H\}^\{\\text\{\\tiny GF\}\}via attention weights\. As such, instead of joint clustering over both features, we simply cluster vectors in𝐏\\mathbf\{P\}, leading to the following optimization objective for deriving the sketching matrix𝐒\\mathbf\{S\}: \(9\)min\{𝒞1,𝒞2,…,𝒞N′\}1N∑k=1N′∑vi∈𝒞k‖𝐏i−∑vj∈𝒞k𝐒k,j𝐏j‖22\.\\min\_\{\\\{\\mathcal\{C\}\_\{1\},\\mathcal\{C\}\_\{2\},\\ldots,\\mathcal\{C\}\_\{N^\{\\prime\}\}\\\}\}\{\\frac\{1\}\{N\}\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\sum\_\{v\_\{i\}\\in\\mathcal\{C\}\_\{k\}\}\\left\\\|\\mathbf\{P\}\_\{i\}\-\\sum\_\{v\_\{j\}\\in\\mathcal\{C\}\_\{k\}\}\\mathbf\{S\}\_\{k,j\}\\mathbf\{P\}\_\{j\}\\right\\\|\_\{2\}^\{2\}\}\. Construction of𝐒\\mathbf\{S\}\.The objective function in Eq\. \([9](https://arxiv.org/html/2607.20477#S4.E9)\) is awithin\-cluster sum of squares, if we set the weight𝐒k,j\\mathbf\{S\}\_\{k,j\}of each nodevjv\_\{j\}to1\|𝒞k\|\\frac\{1\}\{\|\\mathcal\{C\}\_\{k\}\|\}, which can be solved by the well\-known K\-Means algorithm\(Lloyd,[1982](https://arxiv.org/html/2607.20477#bib.bib43)\)\. However, the class labels of nodes are obtained byhardmax discretizationof softmax distributions𝐏\\mathbf\{P\}or𝐒⊤𝐏\\mathbf\{S\}^\{\\top\}\\mathbf\{P\}\(see Eq\. \([7](https://arxiv.org/html/2607.20477#S4.E7)\)\)\. The clustering of softmax distributions is likely to group nodes with distinct labels into the same cluster\. To alleviate this issue, we additionally employ agreedy reassignmentscheme as a post\-refinement of𝐒\\mathbf\{S\}\. More concretely, we first evaluate the affinity between each class labelyk∈𝒴y\_\{k\}\\in\\mathcal\{Y\}\(1≤k≤K1\\leq k\\leq K\) and each cluster𝒞j∈\{𝒞1,𝒞2,…,𝒞N′\}\\mathcal\{C\}\_\{j\}\\in\\\{\\mathcal\{C\}\_\{1\},\\mathcal\{C\}\_\{2\},\\ldots,\\mathcal\{C\}\_\{N^\{\\prime\}\}\\\}by computing the fraction of the total population of labelyky\_\{k\}that falls into𝒞j\\mathcal\{C\}\_\{j\}: \|\{vi∈𝒞j\|𝐘i,k=1\}\|\|\{vi∈𝒱\|𝐘i,k=1\}\|\.\\frac\{\|\\\{v\_\{i\}\\in\\mathcal\{C\}\_\{j\}\|\\ \\mathbf\{Y\}\_\{i,k\}=1\\\}\|\}\{\|\\\{v\_\{i\}\\in\\mathcal\{V\}\|\\ \\mathbf\{Y\}\_\{i,k\}=1\\\}\|\}\.We then pick, for each class, the class\-cluster pair with the highest affinity value, and then select the remaining class\-cluster pairs with the highest affinity values among all theK×N′K\\times N^\{\\prime\}pairs until obtaining top\-N′N^\{\\prime\}pairs in total\. The nodes in each of these classes will be finally assigned to the corresponding cluster\. For each nodeviv\_\{i\}without a final assignment in the above step, its cluster is updated to the one whose centroid has the minimum distance toviv\_\{i\}’s label distribution among the selected clusters of the same predicted class\. ### 4\.3\.Keywords\-based Text Synthesis with LLMs It is necessary to generate corresponding readable texts of each node in𝒢′\\mathcal\{G\}^\{\\prime\}, which can be utilized in downstream LLM\-based tasks and make it easier for human understanding\. Notice that the sparse sketching matrix𝐒\\mathbf\{S\}obtained in the preceding section essentially partitions the nodes in𝒢\\mathcal\{G\}into disjoint clusters\{𝒞1,𝒞2,…,𝒞N′\}\\\{\\mathcal\{C\}\_\{1\},\\mathcal\{C\}\_\{2\},\\ldots,\\mathcal\{C\}\_\{N^\{\\prime\}\}\\\}, each of which corresponds to a node in𝒢′\\mathcal\{G\}^\{\\prime\}\. Intuitively, the textual attributes of each nodevi′∈𝒢′v^\{\\prime\}\_\{i\}\\in\\mathcal\{G\}^\{\\prime\}should be a summarization of the text sequences of nodes in its corresponding cluster𝒞i\\mathcal\{C\}\_\{i\}\. Traditional text distillation maps the synthetic embeddings to discrete synthetic texts\(Morriset al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib64)\)viavector to text\(V2T\), which may involve poor cross\-architecture generalization, high tuning overhead, and low interpretability\. To address these limitations, we directly invoke an LLM to summarize the original texts corresponding to nodes within each cluster, obtaining the synthetic texts without intermediate embedding\-to\-text mapping\. However, directly processing clustered texts via LLMs is both highly time\-consuming and financially costly\. To cost\-effectively leverage the massive knowledge and remarkable generative abilities of LLMs for the summarization fulfilling the above goals,STADresorts to a three\-phase approach\. Cluster\-based Keyword Extraction\.Instead of simply assembling all the text sequences in each cluster𝒞i\\mathcal\{C\}\_\{i\}as the input to LLMs, which not only comprises substantial noisy and redundant information but also entails a high query cost, our solution is to first sift out the keywords from the text sequences𝒮i=\{s1,s2,…,s\|𝒞i\|\}\\mathcal\{S\}\_\{i\}=\\\{s\_\{1\},s\_\{2\},\\dots,s\_\{\|\\mathcal\{C\}\_\{i\}\|\}\\\}of each cluster𝒞i\\mathcal\{C\}\_\{i\}\. Towards this end, we define theintra\-clusterTF\-IDF for each wordwwas follows: \(10\)TF\-IDF\(w,𝒮i\)=TF\(w,𝒮i\)×IDF\(w,𝒮i\),\\text\{TF\-IDF\}\(w,\\mathcal\{S\}\_\{i\}\)=\\text\{TF\}\(w,\\mathcal\{S\}\_\{i\}\)\\times\\text\{IDF\}\(w,\\mathcal\{S\}\_\{i\}\),whereTF\(w,𝒮i\)\\text\{TF\}\(w,\\mathcal\{S\}\_\{i\}\)andIDF\(w,𝒮i\)\\text\{IDF\}\(w,\\mathcal\{S\}\_\{i\}\)are formulated by TF\(w,𝒮i\)=1\|𝒞i\|∑k=1\|𝒞i\|\#\(w,sk\)maxw′∈sk\#\(w′,sk\),IDF\(w,𝒮i\)=log\(\|𝒞i\|1\+df\(w,𝒮i\)\)\.\\begin\{gathered\}\\text\{TF\}\(w,\\mathcal\{S\}\_\{i\}\)=\\frac\{1\}\{\|\\mathcal\{C\}\_\{i\}\|\}\\sum\_\{k=1\}^\{\|\\mathcal\{C\}\_\{i\}\|\}\\frac\{\\\#\(w,s\_\{k\}\)\}\{\\underset\{w^\{\\prime\}\\in s\_\{k\}\}\{\\max\}\\\#\(w^\{\\prime\},s\_\{k\}\)\},\\ \\text\{IDF\}\(w,\\mathcal\{S\}\_\{i\}\)=\\log\\left\(\\frac\{\|\\mathcal\{C\}\_\{i\}\|\}\{1\+\\text\{df\}\(w,\\mathcal\{S\}\_\{i\}\)\}\\right\)\.\\end\{gathered\}Particularly,\#\(w,sj\)\\\#\(w,s\_\{j\}\)is the frequency of wordwwin textsjs\_\{j\}, anddf\(w,𝒮i\)\\text\{df\}\(w,\\mathcal\{S\}\_\{i\}\)denotes the number of texts in𝒮i\\mathcal\{S\}\_\{i\}containing wordww\. Accordingly, we next pick the top\-2ξ2\\xiwords with the highest intra\-cluster TF\-IDF values in𝒮i\\mathcal\{S\}\_\{i\}, followed by an additional filtering step to prune theξ\\xiwords with the smallest Euclidean distances \(SBERT embedding vectors\) to the corresponding attribute vector𝐗i′\\mathbf\{X\}^\{\\prime\}\_\{i\}of𝒞i\\mathcal\{C\}\_\{i\}\. Finally, we obtain a list ofξ\\xidistinct words𝒲i=\{wi,1,wi,2,…,wi,ξ\}\\mathcal\{W\}\_\{i\}=\\\{w\_\{i,1\},w\_\{i,2\},\\dots,w\_\{i,\\xi\}\\\}as the keywords for cluster𝒞i\\mathcal\{C\}\_\{i\}or condensed nodevi′v^\{\\prime\}\_\{i\}in𝒢′\\mathcal\{G\}^\{\\prime\}\. Notably, our keyword extraction jointly considers words’ statistical importance \(quantified by intra\-cluster TF\-IDF\) and semantic relevance \(retained via embedding distance\)\. Using compact keywords instead of full sentences allows LLMs to capture more entities within the same input length\. Keywords as input can help with interpretable text synthesis and enable flexible output by adjusting keyword order or composition\. Candidate Text Generation with LLMs\.Based on keywords𝒲i\\mathcal\{W\}\_\{i\}, we prompt anLLMto generate human\-readable summary textsi,j′s^\{\\prime\}\_\{i,j\}as a candidate text forvi′∈𝒢′v^\{\\prime\}\_\{i\}\\in\\mathcal\{G\}^\{\\prime\}: \(11\)si,j′=LLM\(𝒫,𝒲i,yi′\),s^\{\\prime\}\_\{i,j\}=\\textsf\{LLM\}\(\\mathcal\{P\},\\mathcal\{W\}\_\{i\},y^\{\\prime\}\_\{i\}\),where𝒫\\mathcal\{P\}signifies the prompt specifying the task instruction and a formal description of𝒢\\mathcal\{G\}’s domain, task, and attribute characteristics, andyi′y^\{\\prime\}\_\{i\}is the semantic label ofvi′v^\{\\prime\}\_\{i\}derived in Eq\. \([6](https://arxiv.org/html/2607.20477#S4.E6)\)\. In particular,STADprompts the LLMs to generate a set ofqq\(typicallyq=3q=3\) candidate text\{si,1′,si,2′,…,si,q′\}\\\{s^\{\\prime\}\_\{i,1\},s^\{\\prime\}\_\{i,2\},\\dots,s^\{\\prime\}\_\{i,q\}\\\}forvi′v^\{\\prime\}\_\{i\}\. WSD\-Validation\-guided Approximate Selection\.Since each node in𝒢′\\mathcal\{G\}^\{\\prime\}hasqqcandidate text sequences, it is still prohibitively expensive to find the best synthetic textual attributes due to theO\(qN′\)O\(q^\{N^\{\\prime\}\}\)possible combinations\. As a workaround,STADsamplesnsn\_\{s\}combinations of the candidate text for nodes in𝒢′\\mathcal\{G\}^\{\\prime\}uniformly at random, yieldingnsn\_\{s\}candidate graphs with distinct text sequences\{s1\(j\)′,s2\(j\)′,…,sN′\(j\)′\}j=1ns\\\{s^\{\(j\)\\prime\}\_\{1\},s^\{\(j\)\\prime\}\_\{2\},\\dots,s^\{\(j\)\\prime\}\_\{N^\{\\prime\}\}\\\}\_\{j=1\}^\{n\_\{s\}\}\. Based thereon, each set of text sequences is converted into an attribute matrix𝐗\(j\)′\\mathbf\{X\}^\{\(j\)\\prime\}via a pre\-trained text encoder \(i\.e\., SBERT\)\. We select the top\-kk\(oftenk=5k=5\) candidates with the smallest WSD to𝐗\\mathbf\{X\}, evaluate their downstream task performance on the original dataset, and choose the one achieving the highest validation accuracy as the final textual attributes for𝒢′\\mathcal\{G\}^\{\\prime\}\. Table 1\.Statistics of the seven TAG datasets\.Name\#Nodes\#EdgesDomain\#ClassesCora270810556CS Citation7CiteSeer31868450CS Citation6DBLP14376431326CS Citation4Computers872291256548E\-commerce10Photo48362873782E\-commerce12History41551503180E\-commerce12WikiCS11701431726Knowledge10Table 2\.Node classification accuracy over distillation ratios\. Best and runner\-up per cell, in green and light green, respectively\. Using Qwen3\-max for text synthesis\.DatasetrrDistillation MethodsFull MethodsGCondGCDMGDEMClustGDDCGCTCondTCDMASDClustTDDDaLLMESTAD\(𝐗′\\mathbf\{X\}^\{\\prime\}\)STAD\(𝒮′\\mathcal\{S\}^\{\\prime\}\)GCNMLPCora1\.072\.8±\\pm2\.076\.2±\\pm1\.678\.7±\\pm0\.877\.9±\\pm1\.474\.9±\\pm0\.566\.1±\\pm2\.672\.9±\\pm2\.173\.8±\\pm1\.575\.1±\\pm1\.466\.8±\\pm2\.079\.5±\\pm1\.876\.9±\\pm0\.577\.7±\\pm2\.072\.8±\\pm1\.00\.571\.5±\\pm2\.373\.7±\\pm2\.072\.9±\\pm2\.077\.7±\\pm1\.774\.1±\\pm0\.863\.1±\\pm2\.761\.3±\\pm1\.871\.7±\\pm1\.973\.9±\\pm1\.464\.8±\\pm2\.379\.1±\\pm1\.176\.1±\\pm0\.80\.2563\.5±\\pm5\.668\.9±\\pm2\.876\.3±\\pm1\.176\.6±\\pm1\.372\.6±\\pm1\.163\.1±\\pm1\.866\.0±\\pm1\.566\.9±\\pm1\.267\.8±\\pm3\.560\.0±\\pm1\.179\.1±\\pm0\.976\.5±\\pm0\.50\.0545\.8±\\pm5\.247\.7±\\pm6\.473\.1±\\pm1\.674\.2±\\pm2\.474\.9±\\pm0\.955\.1±\\pm2\.348\.0±\\pm5\.346\.9±\\pm7\.069\.2±\\pm4\.452\.1±\\pm1\.677\.5±\\pm0\.673\.9±\\pm0\.7Citeseer1\.075\.1±\\pm0\.675\.4±\\pm1\.375\.7±\\pm0\.876\.1±\\pm0\.673\.7±\\pm0\.966\.0±\\pm3\.372\.5±\\pm0\.972\.6±\\pm1\.072\.4±\\pm0\.870\.3±\\pm3\.476\.3±\\pm0\.476\.3±\\pm0\.674\.0±\\pm2\.171\.9±\\pm1\.50\.569\.2±\\pm2\.274\.8±\\pm0\.973\.5±\\pm1\.773\.9±\\pm1\.573\.5±\\pm1\.564\.3±\\pm0\.570\.4±\\pm1\.570\.6±\\pm1\.372\.7±\\pm0\.867\.5±\\pm2\.376\.3±\\pm1\.175\.4±\\pm0\.60\.2566\.6±\\pm9\.070\.2±\\pm1\.175\.6±\\pm0\.672\.9±\\pm2\.173\.1±\\pm0\.958\.4±\\pm1\.364\.7±\\pm3\.067\.5±\\pm3\.768\.2±\\pm0\.763\.1±\\pm3\.576\.0±\\pm1\.675\.3±\\pm1\.50\.0548\.3±\\pm7\.651\.5±\\pm9\.076\.8±\\pm0\.571\.7±\\pm3\.075\.1±\\pm0\.558\.3±\\pm2\.755\.7±\\pm6\.250\.6±\\pm3\.266\.9±\\pm0\.563\.0±\\pm3\.575\.9±\\pm0\.573\.2±\\pm0\.9DBLP1\.063\.0±\\pm5\.877\.4±\\pm0\.676\.7±\\pm0\.378\.4±\\pm0\.476\.2±\\pm0\.861\.9±\\pm1\.268\.8±\\pm1\.168\.7±\\pm0\.569\.4±\\pm1\.167\.6±\\pm7\.378\.3±\\pm0\.378\.5±\\pm0\.477\.9±\\pm0\.568\.0±\\pm0\.80\.567\.5±\\pm4\.275\.3±\\pm1\.074\.3±\\pm2\.077\.8±\\pm0\.463\.1±\\pm3\.658\.1±\\pm2\.265\.7±\\pm1\.464\.9±\\pm1\.567\.9±\\pm2\.167\.1±\\pm2\.278\.4±\\pm0\.378\.1±\\pm0\.30\.2561\.9±\\pm12\.171\.6±\\pm3\.776\.3±\\pm0\.575\.3±\\pm1\.552\.5±\\pm10\.156\.8±\\pm2\.760\.5±\\pm2\.261\.2±\\pm2\.367\.0±\\pm2\.265\.3±\\pm1\.378\.3±\\pm0\.477\.8±\\pm0\.50\.0570\.3±\\pm4\.055\.7±\\pm5\.775\.9±\\pm0\.774\.6±\\pm2\.874\.4±\\pm4\.152\.1±\\pm2\.846\.0±\\pm4\.743\.0±\\pm3\.667\.0±\\pm2\.454\.4±\\pm0\.677\.2±\\pm0\.276\.4±\\pm0\.8Computers1\.055\.8±\\pm10\.966\.7±\\pm0\.761\.7±\\pm1\.568\.2±\\pm3\.163\.0±\\pm1\.038\.6±\\pm1\.549\.8±\\pm2\.049\.4±\\pm2\.252\.0±\\pm1\.748\.3±\\pm1\.573\.5±\\pm1\.570\.2±\\pm1\.372\.2±\\pm1\.948\.9±\\pm1\.30\.557\.4±\\pm3\.762\.6±\\pm3\.262\.7±\\pm0\.968\.1±\\pm2\.855\.1±\\pm1\.437\.1±\\pm1\.443\.4±\\pm2\.844\.2±\\pm3\.050\.5±\\pm3\.345\.5±\\pm2\.774\.2±\\pm1\.070\.3±\\pm1\.10\.2548\.0±\\pm6\.353\.9±\\pm5\.154\.2±\\pm4\.161\.9±\\pm3\.550\.5±\\pm3\.537\.9±\\pm1\.739\.2±\\pm2\.139\.8±\\pm1\.042\.9±\\pm3\.041\.4±\\pm3\.271\.5±\\pm2\.868\.2±\\pm0\.80\.0539\.9±\\pm7\.735\.8±\\pm3\.960\.7±\\pm4\.161\.3±\\pm5\.857\.9±\\pm0\.736\.4±\\pm1\.830\.0±\\pm2\.530\.0±\\pm1\.843\.3±\\pm6\.027\.7±\\pm2\.265\.4±\\pm3\.762\.6±\\pm2\.5Photo1\.061\.7±\\pm2\.062\.2±\\pm2\.661\.4±\\pm1\.669\.5±\\pm1\.865\.7±\\pm1\.244\.3±\\pm2\.252\.5±\\pm2\.653\.5±\\pm2\.452\.0±\\pm5\.750\.5±\\pm3\.471\.2±\\pm0\.968\.7±\\pm1\.270\.0±\\pm0\.951\.3±\\pm2\.50\.554\.8±\\pm7\.459\.5±\\pm1\.461\.5±\\pm2\.670\.1±\\pm3\.050\.2±\\pm5\.144\.5±\\pm2\.549\.8±\\pm3\.751\.1±\\pm1\.449\.7±\\pm8\.346\.5±\\pm3\.170\.3±\\pm1\.168\.3±\\pm0\.80\.2549\.3±\\pm5\.850\.0±\\pm3\.154\.6±\\pm1\.068\.8±\\pm2\.246\.4±\\pm2\.745\.6±\\pm2\.746\.2±\\pm5\.148\.3±\\pm3\.442\.0±\\pm7\.847\.0±\\pm1\.769\.9±\\pm1\.868\.9±\\pm1\.30\.0544\.0±\\pm4\.440\.6±\\pm2\.359\.3±\\pm3\.059\.1±\\pm3\.765\.7±\\pm0\.645\.3±\\pm1\.741\.7±\\pm0\.441\.4±\\pm0\.946\.8±\\pm5\.640\.5±\\pm3\.463\.4±\\pm1\.764\.6±\\pm1\.4History1\.076\.5±\\pm1\.074\.2±\\pm1\.770\.9±\\pm2\.174\.2±\\pm3\.074\.5±\\pm0\.972\.1±\\pm1\.069\.4±\\pm2\.172\.9±\\pm0\.774\.6±\\pm1\.374\.2±\\pm2\.276\.8±\\pm1\.678\.1±\\pm1\.170\.6±\\pm3\.662\.8±\\pm3\.70\.574\.6±\\pm2\.871\.4±\\pm1\.873\.5±\\pm1\.274\.2±\\pm3\.171\.9±\\pm1\.669\.4±\\pm2\.865\.6±\\pm2\.868\.2±\\pm3\.772\.1±\\pm1\.575\.4±\\pm2\.778\.6±\\pm0\.877\.8±\\pm0\.80\.2572\.2±\\pm1\.468\.2±\\pm4\.866\.6±\\pm2\.572\.6±\\pm1\.870\.4±\\pm1\.667\.3±\\pm1\.062\.0±\\pm5\.063\.4±\\pm1\.870\.4±\\pm2\.674\.2±\\pm2\.377\.4±\\pm2\.878\.3±\\pm1\.30\.0565\.9±\\pm6\.058\.9±\\pm3\.571\.9±\\pm1\.474\.8±\\pm2\.854\.6±\\pm3\.552\.0±\\pm4\.558\.0±\\pm3\.158\.0±\\pm2\.370\.2±\\pm3\.864\.8±\\pm5\.875\.9±\\pm1\.773\.4±\\pm2\.8WikiCS1\.050\.7±\\pm16\.275\.4±\\pm1\.671\.7±\\pm0\.670\.8±\\pm2\.169\.7±\\pm1\.555\.0±\\pm1\.366\.8±\\pm1\.166\.2±\\pm1\.366\.3±\\pm1\.662\.5±\\pm2\.777\.6±\\pm0\.474\.9±\\pm0\.975\.9±\\pm1\.266\.3±\\pm1\.50\.558\.4±\\pm4\.571\.1±\\pm2\.571\.4±\\pm1\.166\.6±\\pm3\.057\.1±\\pm1\.652\.6±\\pm2\.761\.3±\\pm2\.861\.6±\\pm1\.363\.5±\\pm1\.659\.8±\\pm2\.776\.8±\\pm0\.773\.7±\\pm0\.90\.2543\.9±\\pm4\.668\.6±\\pm3\.368\.8±\\pm1\.067\.2±\\pm1\.547\.3±\\pm3\.350\.6±\\pm1\.755\.6±\\pm2\.353\.9±\\pm1\.357\.9±\\pm1\.854\.0±\\pm3\.376\.5±\\pm0\.974\.0±\\pm0\.80\.0549\.3±\\pm3\.140\.2±\\pm3\.452\.0±\\pm3\.463\.7±\\pm3\.971\.4±\\pm2\.147\.0±\\pm1\.239\.0±\\pm4\.138\.0±\\pm4\.460\.2±\\pm2\.248\.5±\\pm3\.476\.6±\\pm0\.371\.9±\\pm2\.5 ## 5\.Experiments We conduct experiments to address the following questions\. First, how does theSTADperform on semi\-supervised TAGD tasks? Second, how can𝒢′\\mathcal\{G\}^\{\\prime\}generated bySTADhelp LL4TAG methods? Third, how do the modules inSTADcontribute to the overall performance? ### 5\.1\.Experimental Settings Dataset\.Table[1](https://arxiv.org/html/2607.20477#S4.T1)summarizes TAG datasets: citation networksCora,Citeseer\(Yanget al\.,[2016](https://arxiv.org/html/2607.20477#bib.bib8)\),DBLP\(Jiet al\.,[2010](https://arxiv.org/html/2607.20477#bib.bib12)\); E\-commerce co\-view networksComputers,Photo,History\(Yanet al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib13)\); andWikiCS\(Yanet al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib13)\)\(Wikipedia CS pages\)\. Baselines\.We compare with the following recent baselines: - •Graph distillation:GCond\(Jinet al\.,[2022](https://arxiv.org/html/2607.20477#bib.bib38)\),GCDM\(Liuet al\.,[2022](https://arxiv.org/html/2607.20477#bib.bib50)\),GDEM\(Liuet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib41)\),ClustGDD\(Laiet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib40)\),CGC\(Gaoet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib55)\); - •Text distillation:ASD\(Maekawaet al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib48)\),DaLLME\(Taoet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib59)\)\. We also include topology\-removed variantsTCond,TCDM,ClustTDD\(replacing𝐀,𝐀′\\mathbf\{A\},\\mathbf\{A\}^\{\\prime\}with identity matrices𝐈,𝐈′\\mathbf\{I\},\\mathbf\{I\}^\{\\prime\}, respectively\)\. STADhas two variants:STAD\(𝐗′\\mathbf\{X\}^\{\\prime\}\) andSTAD\(𝒮′\\mathcal\{S\}^\{\\prime\}\)\.STAD\(𝐗′\\mathbf\{X\}^\{\\prime\}\) trains a GCN on𝐗′\\mathbf\{X\}^\{\\prime\},𝐀′\\mathbf\{A\}^\{\\prime\},𝐘′\\mathbf\{Y\}^\{\\prime\}and evaluates it on𝐗\\mathbf\{X\},𝐀\\mathbf\{A\},𝐘\\mathbf\{Y\}\. ForSTAD\(𝒮′\\mathcal\{S\}^\{\\prime\}\),𝒮′\\mathcal\{S\}^\{\\prime\}is remapped to𝐗′′∈ℝN′×D\\mathbf\{X\}^\{\\prime\\prime\}\\in\\mathbb\{R\}^\{N^\{\\prime\}\\times D\}via SBERT, then a GCN is trained on𝐗′′\\mathbf\{X\}^\{\\prime\\prime\},𝐀′\\mathbf\{A\}^\{\\prime\},𝐘′\\mathbf\{Y\}^\{\\prime\}and evaluated on𝐗\\mathbf\{X\},𝐀\\mathbf\{A\},𝐘\\mathbf\{Y\}\. Implementation Details\.We use five random splits with 20 nodes per class \(or all if fewer exist\)\. GCN and MLP are two\-layer \(hidden 64, lr 0\.01, weight decay0\.00050\.0005\)\. Text synthesis uses Qwen3\-max \(API\); LLM4TAG training/inference uses Qwen3\-1\.7B\(Yanget al\.,[2025a](https://arxiv.org/html/2607.20477#bib.bib68)\)\. ### 5\.2\.Node Classification Performance Evaluation Table[2](https://arxiv.org/html/2607.20477#S4.T2)validates our design: WSD\-guided dual\-path encoding and keyword\-based text synthesis yield strong, stable compression\. Graph distillation generally outperforms text\-only distillation, consistent with full models \(GCN vs\. MLP\)\. Withr=0\.05r=0\.05,GCondandGCDMfall belowTCondandTCDM, likely due to over\-smoothing on low\-quality synthetic topology\. Gradient\-based methods \(GCond,TCond\) are unstable asrrdecreases, e\.g\.,GCondonCoradrops from 72\.8%±\\pm2\.0 to 45\.8%±\\pm5\.2\. Clustering\-based methods \(ClustGDD,CGC,ClustTDD,DaLLME\) are more competitive and stable, andGDEMachieves good performance via spectral alignment\.STAD\(𝐗′\\mathbf\{X\}^\{\\prime\}\) andSTAD\(𝒮′\\mathcal\{S\}^\{\\prime\}\) achieve overall best accuracy\. OnWikiCSatr=0\.5r\{=\}0\.5,STAD\(𝐗′\\mathbf\{X\}^\{\\prime\}\) reaches 76\.8%, outperforming full GCN \(75\.9%\) and the next\-best distillation \(71\.4%\)\.STADon distilled data sometimes exceeds full\-data training, which we attribute to effective extraction of discriminative information and noise reduction\.STAD\(𝐗′\\mathbf\{X\}^\{\\prime\}\) usually beatsSTAD\(𝒮′\\mathcal\{S\}^\{\\prime\}\) because𝐗′\\mathbf\{X\}^\{\\prime\}preserves semantics via direct aggregation;𝐗′′\\mathbf\{X\}^\{\\prime\\prime\}from𝒮′\\mathcal\{S\}^\{\\prime\}incurs information loss from SBERT’s 256\-token limit\.STAD\(𝒮′\\mathcal\{S\}^\{\\prime\}\) achieves relative better performance on TAGs with large scale\. This observation arises because clusters with richer text provide more context and statistical evidence, yielding more representative, informative keywords\. Table 3\.Node classification accuracy \(mean %±\\pmstd\) with semi\-supervised LLM4TAG methods withr=0\.5r=0\.5\.DatasetDistillationMethodBaselinesGNN\-LLMENGINETAPECoraFull Methods77\.8±\\pm1\.678\.8±\\pm1\.177\.5±\\pm1\.0GCond\+V2T29\.4±\\pm2\.435\.0±\\pm7\.033\.5±\\pm6\.0GCDM\+V2T45\.1±\\pm7\.947\.1±\\pm6\.358\.7±\\pm6\.8GDEM\+V2T27\.4±\\pm5\.925\.9±\\pm6\.010\.4±\\pm3\.3ClustGDD\+V2T55\.9±\\pm5\.641\.8±\\pm6\.365\.0±\\pm2\.0DaLLME53\.0±\\pm5\.157\.9±\\pm5\.061\.5±\\pm2\.1STAD70\.9±\\pm4\.273\.9±\\pm1\.670\.4±\\pm2\.6HistoryFull Methods65\.4±\\pm4\.869\.1±\\pm2\.469\.0±\\pm4\.9GCond\+V2T40\.6±\\pm21\.448\.0±\\pm17\.922\.1±\\pm3\.7GCDM\+V2T48\.4±\\pm15\.946\.4±\\pm19\.744\.6±\\pm15\.3GDEM\+V2T46\.1±\\pm15\.640\.6±\\pm19\.514\.6±\\pm12\.1ClustGDD\+V2T70\.9±\\pm3\.163\.0±\\pm3\.768\.3±\\pm4\.1DaLLME64\.6±\\pm6\.273\.8±\\pm2\.172\.7±\\pm1\.1STAD76\.6±\\pm1\.874\.6±\\pm4\.374\.3±\\pm2\.2Table 4\.Node classification accuracy \(mean %±\\pmstd\.\) with training\-free LLM4TAG methods \(ICL\)\.DatasetMethodsDirectSTAD\+DirectNS\.STAD\+NS\.Cora59\.9±\\pm0\.262\.3±\\pm0\.660\.5±\\pm0\.764\.1±\\pm0\.7History13\.6±\\pm0\.023\.5±\\pm2\.314\.8±\\pm0\.032\.5±\\pm4\.6 ### 5\.3\.LLM4TAG Performance Evaluation We train three LLM4TAG methods,GNN\-LLM\(Wuet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib24)\),ENGINE\(Zhuet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib27)\), andTAPE\(Heet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib28)\), withr=0\.5r=0\.5\. To ensure fair comparison, we use V2T\(Morriset al\.,[2023](https://arxiv.org/html/2607.20477#bib.bib64)\)to map synthetic attribute matrices to text\. Table[3](https://arxiv.org/html/2607.20477#S5.T3)shows thatSTADachieves the best transfer performance to these frameworks\. OnCora,STADreaches 70\.9%, 73\.9%, and 70\.4% on the three methods, respectively, outperforming all baselines and narrowing the gap to the full\-data baseline\. OnHistory,STADattains 76\.6%, 74\.6%, and 74\.3%, surpassing the full\-data baseline forENGINEandTAPE\. Human\-readable synthetic texts preserve the semantics required by LLM4TAG models; in contrast,GCondandGDEMwith V2T perform poorly, likely due to embedding misalignment caused by V2T\. We also useSTAD\(𝒮′\\mathcal\{S\}^\{\\prime\}\) for training\-free LLM4TAG, which includes two modes: Direct \(i\.e\., LLMs directly give predictions based on node texts\) and Neighbor Summarization \(NS, i\.e\., LLMs first summarize the text of a node’s neighbors, then make predictions based on the node’s own text and the neighbor summary\)\. We applySTAD\(𝒮′\\mathcal\{S\}^\{\\prime\}\) viain\-context learning\(ICL\): prompts are constructed based on the distances between input embeddings and synthetic node attributes\. Table[4](https://arxiv.org/html/2607.20477#S5.T4)shows consistent performance gains: onCora, we observe \+2\.4% \(Direct\) and \+3\.6% \(Nbr\.Sum\.\); onHistory, the gains are \+9\.9% and \+17\.7%\. These results show that synthetic texts are effective ICL exemplars\. ### 5\.4\.Ablation Study We conduct ablation studies withr=0\.05r=0\.05to validate the effectiveness of modules in our framework in Table[5](https://arxiv.org/html/2607.20477#S5.T5)\. \(w/o: module removing, w: another module as alternative\.\) Removing the GA pathway leads to significant performance degradation, particularly onPhoto, where accuracy drops by 11\.5%, indicating the critical role of graph structure\. Disabling CoST consistently reduces performance across all datasets, confirming that it effectively enhances the dual\-pathway architecture\. Removing our reassignment after K\-Means hurts results, validating the necessity of reassignment for more pure clusters\. We omit graph\-text collaborative encoding and employ raw attributes from SBERT for sketching\. The performance decline verifies the necessity of the encoding module\. Furthermore, applying vanilla self\-training to dual\-path encoders and baseline methods \(e\.g\., GCDM\+ST\) yields inferior performance compared toSTAD\(𝐗′\\mathbf\{X\}^\{\\prime\}\), demonstrating the superiority of Graph\-Text Collaborative Encoding in semi\-supervised distillation\. We evaluate alternative text generation strategies\. Variants without LLM summarization, with random words, or with key sentences all underperform our keywords\-based approach, with random words causing severe degradation up to 15\.5% on Photo\. We build summarization prompts without synthetic label injection, causing performance drops and instability across most datasets\. This proves that synthetic label constraints enable LLMs to produce more category\-aware summaries\. Direct generation without candidate set construction also reduces accuracy, validating the importance of candidate selection\. Finally, we select candidates only via WSD\. Results show both WSD scores and validation accuracy are essential for optimal candidate selection\. Table 5\.Ablation study: accuracy \(mean %±\\pmstd\)\.ModulesCoraPhotoHistoryWikiCSw/o GA75\.3±\\pm1\.052\.0±\\pm5\.670\.8±\\pm3\.965\.9±\\pm1\.5w/o GF76\.6±\\pm1\.661\.4±\\pm1\.871\.8±\\pm2\.675\.8±\\pm0\.9w/o CoST75\.7±\\pm0\.960\.6±\\pm0\.664\.8±\\pm4\.176\.4±\\pm0\.7w/o Reassignment75\.2±\\pm2\.261\.8±\\pm1\.475\.4±\\pm4\.068\.1±\\pm3\.3w SBERT Clustering73\.8±\\pm0\.859\.6±\\pm1\.868\.9±\\pm5\.771\.8±\\pm0\.7w vanilla ST77\.4±\\pm1\.063\.4±\\pm0\.773\.2±\\pm2\.076\.2±\\pm0\.5GCDM\+ST77\.3±\\pm2\.270\.8±\\pm1\.970\.1±\\pm3\.274\.9±\\pm3\.2GDEM\+ST73\.1±\\pm2\.254\.7±\\pm1\.872\.0±\\pm4\.250\.8±\\pm1\.2ClustGDD\+ST74\.3±\\pm1\.262\.8±\\pm4\.574\.5±\\pm3\.162\.8±\\pm4\.9STAD\(𝐗′\\mathbf\{X\}^\{\\prime\}\)77\.5±\\pm1\.363\.5±\\pm1\.775\.9±\\pm1\.776\.6±\\pm0\.3w/o LLM Summary71\.8±\\pm1\.362\.0±\\pm1\.073\.3±\\pm2\.070\.4±\\pm1\.5w/o Syn Label73\.5±\\pm1\.059\.9±\\pm1\.967\.2±\\pm6\.572\.7±\\pm1\.6w Random Words69\.7±\\pm2\.149\.1±\\pm3\.363\.0±\\pm7\.767\.3±\\pm1\.9w Key Sentence72\.6±\\pm1\.058\.9±\\pm2\.470\.6±\\pm2\.968\.2±\\pm3\.7w/o Candidates set72\.5±\\pm1\.562\.1±\\pm5\.571\.6±\\pm1\.071\.5±\\pm0\.9w WSD\-only73\.2±\\pm1\.263\.3±\\pm1\.272\.2±\\pm1\.770\.7±\\pm2\.6STAD\(𝒮′\\mathcal\{S\}^\{\\prime\}\)73\.9±\\pm0\.764\.6±\\pm1\.473\.4±\\pm2\.871\.9±\\pm2\.5 ### 5\.5\.Evaluation of LLMs for Text Synthesis Table[6](https://arxiv.org/html/2607.20477#S5.T6)compares LLM performance and cost for text synthesis222https://www\.aliyun\.com/, https://www\.deepseek\.com/, https://openai\.comwithr=0\.05r=0\.05\. Qwen3\-max achieves the best results on both datasets \(73\.9% on Cora, 73\.4% on History\) at moderate cost, while smaller Qwen3 variants \(1\.7B–8B\) offer competitive accuracy with significantly lower expense\. Scaling model size within the Qwen3 series does not consistently improve performance, with 8B outperforming 32B on Cora\. GPT\-5 achieve comparable accuracy to top performers but at 5–7×\\timeshigher price\. All methods yield similar low token usage, compared to the full graphs333Token number estimated by https://pypi\.org/project/tiktoken/\. Overall, cost\-efficient models match or exceed expensive alternatives, suggesting model selection matters more than scale for this task\. Table 6\.Accuracy and I/O price of various LLMs\.ModelI/OCoraHistory\($/1M\)Acc\#Tokens \(M\)Acc\#Tokens \(M\)Qwen3\-1\.7B0\.04/0\.1766\.2±\\pm2\.60\.022672\.6±\\pm1\.70\.0379Qwen3\-8B0\.07/0\.2871\.0±\\pm0\.70\.022872\.4±\\pm1\.90\.0377Qwen3\-32B0\.28/1\.1270\.5±\\pm0\.80\.023171\.8±\\pm2\.00\.0366Qwen3\-max0\.35/1\.4073\.9±\\pm0\.70\.023773\.4±\\pm2\.80\.0383DeepSeek\-V3\.20\.28/0\.4271\.3±\\pm1\.30\.022373\.1±\\pm2\.80\.0372GPT\-4\.1\-mini0\.40/1\.6072\.2±\\pm1\.10\.022869\.1±\\pm1\.70\.0371GPT\-4\.12\.00/8\.0072\.0±\\pm1\.20\.023072\.8±\\pm2\.90\.0369GPT\-5\-mini0\.25/2\.070\.0±\\pm1\.50\.022172\.2±\\pm3\.90\.0351GPT\-51\.25/10\.072\.2±\\pm1\.70\.023273\.9±\\pm1\.70\.0376\#Tokens for entire TAGs \(M\)0\.44612\.5 ### 5\.6\.In\-depth Analysis of CoST and Sketching Figure[4](https://arxiv.org/html/2607.20477#S5.F4)tracks accuracy across CoST iterations, i\.e\., self\-training steps with progressively increasing pseudo\-label ratio, for the GA and GF pathways, and the fused final prediction\. GA consistently outperforms GF while final prediction exceeds both, and performance gaps narrow over iterations as joint optimization on pseudo\-labels enables bidirectional knowledge transfer\. All curves rise rapidly, then plateau or decline, with optimal iterations varying across datasets\. We thus iterate until unlabeled nodes are exhausted and select the checkpoint with the highest validation performance\. Here we define the overall purity of clusters in sketching as∑j=1N′maxy∈𝒴\|\{vi∈𝒞j:yi=y\}\|N\\frac\{\\sum\_\{j=1\}^\{N^\{\\prime\}\}\\max\_\{y\\in\\mathcal\{Y\}\}\|\\\{v\_\{i\}\\in\\mathcal\{C\}\_\{j\}:y\_\{i\}=y\\\}\|\}\{N\}, which measures the extent to which clusters contain nodes of a single class\. As shown in Figure[5](https://arxiv.org/html/2607.20477#S5.F5), reassignment consistently improves purity,e\.g\.,WikiCSexhibits the largest relative improvement \(from 0\.64 to 0\.79\), indicating that reassignment is effective for datasets with initially noisy cluster structures\. These results validate that our reassignment after clustering successfully refines cluster quality, improving sketching\. 𝐏GA\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}𝐏GF\\mathbf\{P\}^\{\\text\{\\tiny GF\}\}𝐏\\mathbf\{P\} 0112233445566778899707075758080Acc\(%\)\(\(a\)\)01122334455667788996060656570707575Acc\(%\)\(\(b\)\) Figure 4\.Accuracy by𝐏\\mathbf\{P\},𝐏GA\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}, and𝐏GF\\mathbf\{P\}^\{\\text\{\\tiny GF\}\}when varying \#iterations\.CoraHistoryPhotoWikiCS0\.60\.60\.650\.650\.70\.70\.750\.750\.80\.80\.850\.85Purity ValueBeforeAfterFigure 5\.Cluster purity before/after greedy reassignment\. ### 5\.7\.Hyperparameter Analysis Figure[6](https://arxiv.org/html/2607.20477#S5.F6)examines four key hyperparameters\. Increasing input keywordsξ\\xifrom 64 to 512 improvesCoragradually to 74\.5%, whileHistoryremains stable around 73\.3% and then drops to 70\.6% at 512, suggesting excessive keywords may introduce noise\. Output tokens limitsκ\\kappashow similar divergence:Coraimproves to 73\.9% atκ=256\\kappa=256whileHistoryat 73\.4% withκ=12\\kappa=12\. For the number of candidate summaries per clusterqq\. Our method is robust onCorawith 74\.0±\\pm0\.2% acrossq∈\{2,3,4,5\}q\\in\\\{2,3,4,5\\\}\.Historydrops to 70\.4% atq=4q=4but recovers atq=5q=5; For number of candidate synthetic graphsnsn\_\{s\},Coramaintains 72\.8%–74\.1% across\{10,50,100,200\}\\\{10,50,100,200\\\}\.Historyimproves from 72\.9% to 75\.5%, demonstrating larger pools benefit complex textual datasets\. CoraHistory 6412825651268687070727274747676Acc68687070727274747676Acc\(\(a\)\)6412825651268687070727274747676Acc68687070727274747676Acc\(\(b\)\)234568687070727274747676Acc68687070727274747676Acc\(\(c\)\)105010020068687070727274747676Acc68687070727274747676Acc\(\(d\)\) Figure 6\.Hyperparameter analysis: effect ofξ\\xi,κ\\kappa,qq, andnsn\_\{s\}on accuracy \(mean %\)±\\pmstd\.\)\. ### 5\.8\.Visualization As shown in Figure[7](https://arxiv.org/html/2607.20477#S5.F7), t\-SNE visualizations of node attributes reveal that original data features exhibit entangled class boundaries, while synthetic features generated by our distillation method form clearly separated clusters\. This result confirms that our distillation strategy not only preserves but also sharpens class\-discriminative structural information, even with substantial compression ratios\. \(a\)Cora\(𝐗\\mathbf\{X\}\)\(b\)Cora\(𝐗′\\mathbf\{X\}^\{\\prime\}\)\(c\)History\(𝐗\\mathbf\{X\}\)\(d\)History\(𝐗′\\mathbf\{X\}^\{\\prime\}\)abcdFigure 7\.t\-SNE of node attributes: original \(𝐗\\mathbf\{X\}\) vs\. synthetic \(𝐗′\\mathbf\{X\}^\{\\prime\}\) onCoraandHistory\. ### 5\.9\.Conclusion This paper proposesSTADfor semi\-supervised TAG distillation\. Guided by WSD,STADintroduces three modules: graph\-text collaborative encoding, WSD\-based graph sketching, and keywords\-based LLM text synthesis\.STADachieves the state\-of\-the\-art performance compared to current graph and text distillation methods on various real\-world datasets\. The synthetic TAGs obtained fromSTADenable downstream LLM4Graph methods to be efficient and effective\. A potential direction for future work is to explore multi\-modal graph distillation for continual learning\. ## 6\.Acknowledgments Renchi Yang is supported by the Guangdong and Hong Kong Universities ”1\+1\+1” Joint Research Collaboration Scheme, project No\.: 2025A0505000002, the NSFC \(No\. 62302414\), the Hong Kong RGC ECS grant \(No\. 22202623\), and YCRG \(No\. C2003\-23Y\)\. Tsz Nam Chan is supported by the Natural Science Foundation of China under grants 23IAA00610 and 62572326\. ## References - D\. Bo, X\. Wang, C\. Shi, and H\. Shen \(2021\)Beyond low\-frequency information in graph convolutional networks\.Proceedings of the AAAI Conference on Artificial Intelligence35\(5\),pp\. 3950–3957\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/16514),[Document](https://dx.doi.org/10.1609/aaai.v35i5.16514)Cited by:[§4\.1](https://arxiv.org/html/2607.20477#S4.SS1.p2.11)\. - R\. Chen, T\. Zhao, A\. K\. Jaiswal, N\. Shah, and Z\. Wang \(2024a\)LLaGA: large language and graph assistant\.InInternational Conference on Machine Learning,pp\. 7809–7823\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1),[§2](https://arxiv.org/html/2607.20477#S2.p4.1)\. - Z\. Chen, H\. Mao, J\. Liu, Y\. Song, B\. Li, W\. Jin, B\. Fatemi, A\. Tsitsulin, B\. Perozzi, H\. Liu,et al\.\(2024b\)Text\-space graph foundation models: comprehensive benchmarks and new insights\.Advances in Neural Information Processing Systems37,pp\. 7464–7492\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1),[§2](https://arxiv.org/html/2607.20477#S2.p4.1)\. - K\. L\. Clarkson and D\. P\. Woodruff \(2017\)Low\-rank approximation and regression in input sparsity time\.Journal of the ACM \(JACM\)63\(6\),pp\. 1–45\.Cited by:[§4\.2](https://arxiv.org/html/2607.20477#S4.SS2.p2.2)\. - J\. Feng, H\. Liu, L\. Kong, M\. Zhu, Y\. Chen, and M\. Zhang \(2024\)TAGLAS: an atlas of text\-attributed graph datasets in the era of large graph and language models\.arXiv preprint arXiv:2406\.14683\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - X\. Gao, G\. Ye, T\. Chen, W\. Zhang, J\. Yu, and H\. Yin \(2025\)Rethinking and accelerating graph condensation: a training\-free approach with class partition\.InProceedings of the ACM on Web Conference 2025,pp\. 4359–4373\.Cited by:[§2](https://arxiv.org/html/2607.20477#S2.p1.1),[1st item](https://arxiv.org/html/2607.20477#S5.I1.i1.p1.1)\. - W\. Hamilton, Z\. Ying, and J\. Leskovec \(2017\)Inductive representation learning on large graphs\.Advances in neural information processing systems30\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - X\. He, X\. Bresson, T\. Laurent, A\. Perold, Y\. LeCun, and B\. Hooi \(2024\)Harnessing explanations: llm\-to\-lm interpreter for enhanced text\-attributed graph representation learning\.InICLR,Cited by:[§C\.2](https://arxiv.org/html/2607.20477#A3.SS2.p1.3),[§1](https://arxiv.org/html/2607.20477#S1.p1.1),[§1](https://arxiv.org/html/2607.20477#S1.p2.1),[§2](https://arxiv.org/html/2607.20477#S2.p4.1),[§5\.3](https://arxiv.org/html/2607.20477#S5.SS3.p1.1)\. - M\. Ji, Y\. Sun, M\. Danilevsky, J\. Han, and J\. Gao \(2010\)Graph regularized transductive classification on heterogeneous information networks\.InJoint European Conference on Machine Learning and Knowledge Discovery in Databases,pp\. 570–586\.Cited by:[§5\.1](https://arxiv.org/html/2607.20477#S5.SS1.p1.1)\. - B\. Jin, G\. Liu, C\. Han, M\. Jiang, H\. Ji, and J\. Han \(2024\)Large language models on graphs: a comprehensive survey\.IEEE Transactions on Knowledge and Data Engineering\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - W\. Jin, L\. Zhao, S\. Zhang, Y\. Liu, J\. Tang, and N\. Shah \(2022\)Graph condensation for graph neural networks\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p2.1),[§2](https://arxiv.org/html/2607.20477#S2.p1.1),[1st item](https://arxiv.org/html/2607.20477#S5.I1.i1.p1.1)\. - P\. Karisani \(2023\)Neural networks against \(and for\) self\-training: classification with small labeled and large unlabeled sets\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 12148–12162\.Cited by:[§4\.1](https://arxiv.org/html/2607.20477#S4.SS1.p5.10)\. - T\. N\. Kipf and M\. Welling \(2017\)Semi\-supervised classification with graph convolutional networks\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=SJU4ayYgl)Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - Y\. Lai, T\. Zhang, and R\. Yang \(2025\)Simple yet effective graph distillation via clustering\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\. 2,pp\. 1229–1240\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p2.1),[§2](https://arxiv.org/html/2607.20477#S2.p1.1),[§3\.2](https://arxiv.org/html/2607.20477#S3.SS2.p1.9),[1st item](https://arxiv.org/html/2607.20477#S5.I1.i1.p1.1)\. - D\. Leeet al\.\(2013\)Pseudo\-label: the simple and efficient semi\-supervised learning method for deep neural networks\.InWorkshop on challenges in representation learning, ICML,Vol\.3,pp\. 896\.Cited by:[§4\.1](https://arxiv.org/html/2607.20477#S4.SS1.p4.11)\. - S\. Lei and D\. Tao \(2023\)A comprehensive survey of dataset distillation\.IEEE Transactions on Pattern Analysis and Machine Intelligence46\(1\),pp\. 17–32\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p2.1)\. - Q\. Li, Z\. Han, and X\. Wu \(2018\)Deeper insights into graph convolutional networks for semi\-supervised learning\.InProceedings of the AAAI conference on artificial intelligence,Vol\.32\.Cited by:[§4\.1](https://arxiv.org/html/2607.20477#S4.SS1.p4.11)\. - T\. Li, Z\. Fang, X\. Zhang, K\. Tang, H\. Chen, Z\. Jiang, T\. Zhao, R\. Xu, F\. Cheng, X\. Li,et al\.\(2025\)DrugLM: a unified framework to enhance drug\-target interaction predictions by incorporating textual embeddings via language models\.bioRxiv,pp\. 2025–07\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - Y\. Li and W\. Li \(2021\)Data distillation for text classification\.arXiv preprint arXiv:2104\.08448\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p2.1),[§2](https://arxiv.org/html/2607.20477#S2.p2.1)\. - Y\. Li, Z\. Li, P\. Wang, J\. Li, X\. Sun, H\. Cheng, and J\. X\. Yu \(2024\)A survey of graph meets large language model: progress and future directions\.InProceedings of the Thirty\-Third International Joint Conference on Artificial Intelligence,pp\. 8123–8131\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - J\. Liu, C\. Yang, Z\. Lu, J\. Chen, Y\. Li, M\. Zhang, T\. Bai, Y\. Fang, L\. Sun, P\. S\. Yu,et al\.\(2023\)Towards graph foundation models: a survey and beyond\.arXiv preprint arXiv:2310\.11829\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - M\. Liu, S\. Li, X\. Chen, and L\. Song \(2022\)Graph condensation via receptive field distribution matching\.arXiv preprint arXiv:2206\.13697\.Cited by:[§2](https://arxiv.org/html/2607.20477#S2.p1.1),[1st item](https://arxiv.org/html/2607.20477#S5.I1.i1.p1.1)\. - Y\. Liu, D\. Bo, and C\. Shi \(2024\)Graph distillation with eigenbasis matching\.InInternational Conference on Machine Learning,pp\. 30702–30717\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p2.1),[§2](https://arxiv.org/html/2607.20477#S2.p1.1),[1st item](https://arxiv.org/html/2607.20477#S5.I1.i1.p1.1)\. - S\. Lloyd \(1982\)Least squares quantization in pcm\.IEEE transactions on information theory28\(2\),pp\. 129–137\.Cited by:[§4\.2](https://arxiv.org/html/2607.20477#S4.SS2.p6.12)\. - S\. Luan, C\. Hua, Q\. Lu, J\. Zhu, M\. Zhao, S\. Zhang, X\. Chang, and D\. Precup \(2022\)Revisiting heterophily for graph neural networks\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 1362–1375\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/092359ce5cf60a80e882378944bf1be4-Paper-Conference.pdf)Cited by:[§4\.1](https://arxiv.org/html/2607.20477#S4.SS1.p2.11)\. - A\. Maekawa, N\. Kobayashi, K\. Funakoshi, and M\. Okumura \(2023\)Dataset distillation with attention labels for fine\-tuning BERT\.InACL Short Papers,A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 119–127\.External Links:[Link](https://aclanthology.org/2023.acl-short.12/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-short.12)Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p2.1),[§2](https://arxiv.org/html/2607.20477#S2.p2.1),[2nd item](https://arxiv.org/html/2607.20477#S5.I1.i2.p1.2)\. - A\. Maekawa, S\. Kosugi, K\. Funakoshi, and M\. Okumura \(2025\)Dilm: distilling dataset into language model for text\-level dataset distillation\.Journal of Natural Language Processing32\(1\),pp\. 252–282\.Cited by:[§2](https://arxiv.org/html/2607.20477#S2.p2.1)\. - J\. Morris, V\. Kuleshov, V\. Shmatikov, and A\. M\. Rush \(2023\)Text embeddings reveal \(almost\) as much as text\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 12448–12460\.Cited by:[§4\.3](https://arxiv.org/html/2607.20477#S4.SS3.p1.7),[§5\.3](https://arxiv.org/html/2607.20477#S5.SS3.p1.1)\. - C\. Raffel, N\. Shazeer, A\. Roberts, K\. Lee, S\. Narang, M\. Matena, Y\. Zhou, W\. Li, and P\. J\. Liu \(2020\)Exploring the limits of transfer learning with a unified text\-to\-text transformer\.Journal of machine learning research21\(140\)\.Cited by:[§2](https://arxiv.org/html/2607.20477#S2.p2.1)\. - N\. Reimers and I\. Gurevych \(2019\)Sentence\-bert: sentence embeddings using siamese bert\-networks\.InEMNLP\-IJCNLP,pp\. 3982–3992\.Cited by:[§3\.1](https://arxiv.org/html/2607.20477#S3.SS1.p1.36)\. - N\. Sachdeva and J\. McAuley \(2023\)Data distillation: a survey\.Transactions on Machine Learning Research\.Note:Survey CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=lmXMXP74TO)Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p2.1)\. - J\. Solomon, F\. de Goes, G\. Peyré, M\. Cuturi, A\. Butscher, A\. Nguyen, T\. Du, and L\. Guibas \(2015a\)Convolutional wasserstein distances: efficient optimal transportation on geometric domains\.ACM Trans\. Graph\.34\(4\)\.External Links:ISSN 0730\-0301,[Link](https://doi.org/10.1145/2766963),[Document](https://dx.doi.org/10.1145/2766963)Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p3.1)\. - J\. Solomon, F\. De Goes, G\. Peyré, M\. Cuturi, A\. Butscher, A\. Nguyen, T\. Du, and L\. Guibas \(2015b\)Convolutional wasserstein distances: efficient optimal transportation on geometric domains\.ACM Transactions on Graphics \(ToG\)34\(4\),pp\. 1–11\.Cited by:[§3\.2](https://arxiv.org/html/2607.20477#S3.SS2.p1.2)\. - G\. Su, H\. Wang, J\. Wang, W\. Zhang, Y\. Zhang, and J\. Pei \(2025\)Large language models meet text\-attributed graphs: a survey of integration frameworks and applications\.arXiv preprint arXiv:2510\.21131\.Cited by:[§2](https://arxiv.org/html/2607.20477#S2.p4.1)\. - I\. Sucholutsky and M\. Schonlau \(2021\)Soft\-label dataset distillation and text dataset distillation\.InIJCNN,pp\. 1–8\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p2.1),[§2](https://arxiv.org/html/2607.20477#S2.p2.1)\. - S\. Sun, Y\. Ren, J\. Chen, and C\. Ma \(2025\)Large language models as topological structure enhancers for text\-attributed graphs\.InInternational Conference on Database Systems for Advanced Applications,pp\. 106–122\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - J\. Tang, Y\. Yang, W\. Wei, L\. Shi, L\. Su, S\. Cheng, D\. Yin, and C\. Huang \(2024\)Graphgpt: graph instruction tuning for large language models\.InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 491–500\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1),[§2](https://arxiv.org/html/2607.20477#S2.p4.1)\. - Y\. Tao, L\. Kong, A\. Kan, and L\. Callot \(2024\)Textual dataset distillation via language model embedding\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 12557–12569\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p2.1),[§2](https://arxiv.org/html/2607.20477#S2.p2.1),[2nd item](https://arxiv.org/html/2607.20477#S5.I1.i2.p1.2)\. - P\. Veličković, G\. Cucurull, A\. Casanova, A\. Romero, P\. Liò, and Y\. Bengio \(2018\)Graph attention networks\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=rJXMpikCZ)Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - Z\. Wang, S\. Liu, Z\. Zhang, T\. Ma, C\. Zhang, and Y\. Ye \(2025\)Can llms convert graphs to text\-attributed graphs?\.InNAACL,pp\. 1412–1432\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - X\. Wu, Y\. Shen, F\. Ge, C\. Shan, Y\. Jiao, X\. Sun, and H\. Cheng \(2025\)When do llms help with node classification? a comprehensive analysis\.InForty\-second International Conference on Machine Learning,Cited by:[§C\.2](https://arxiv.org/html/2607.20477#A3.SS2.p1.3),[§1](https://arxiv.org/html/2607.20477#S1.p1.1),[§2](https://arxiv.org/html/2607.20477#S2.p4.1),[§5\.3](https://arxiv.org/html/2607.20477#S5.SS3.p1.1)\. - H\. Yan, C\. Li, R\. Long, C\. Yan, J\. Zhao, W\. Zhuang, J\. Yin, P\. Zhang, W\. Han, H\. Sun,et al\.\(2023\)A comprehensive study on text\-attributed graphs: benchmarking and rethinking\.Advances in Neural Information Processing Systems36\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1),[§2](https://arxiv.org/html/2607.20477#S2.p4.1),[§5\.1](https://arxiv.org/html/2607.20477#S5.SS1.p1.1)\. - A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025a\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§5\.1](https://arxiv.org/html/2607.20477#S5.SS1.p4.1)\. - B\. Yang, K\. Wang, Q\. Sun, C\. Ji, X\. Fu, H\. Tang, Y\. You, and J\. Li \(2023\)Does graph distillation see like vision dataset counterpart?\.Advances in Neural Information Processing Systems36,pp\. 53201–53226\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p2.1),[§1](https://arxiv.org/html/2607.20477#S1.p3.1),[§2](https://arxiv.org/html/2607.20477#S2.p1.1),[§3\.2](https://arxiv.org/html/2607.20477#S3.SS2.p1.9)\. - C\. Yang, H\. Liu, D\. Wang, Z\. Zhang, C\. Yang, and C\. Shi \(2025b\)Flag: fraud detection with llm\-enhanced graph neural network\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\. 2,Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - Z\. Yang, W\. Cohen, and R\. Salakhudinov \(2016\)Revisiting semi\-supervised learning with graph embeddings\.InInternational conference on machine learning,pp\. 40–48\.Cited by:[§5\.1](https://arxiv.org/html/2607.20477#S5.SS1.p1.1)\. - D\. C\. Zhang, M\. Yang, R\. Ying, and H\. W\. Lauw \(2024a\)Text\-attributed graph representation learning: methods, applications, and challenges\.InCompanion Proceedings of the ACM Web Conference 2024,pp\. 1298–1301\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - H\. Zhang, P\. S\. Yu, and J\. Zhang \(2025a\)A systematic survey of text summarization: from statistical methods to large language models\.ACM Computing Surveys57\(11\),pp\. 1–41\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p3.1)\. - J\. Zhang, J\. Chen, M\. Yang, A\. Feng, S\. Liang, J\. Shao, and R\. Ying \(2024b\)DTGB: a comprehensive benchmark for dynamic text\-attributed graphs\.Advances in Neural Information Processing Systems37,pp\. 91405–91429\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - Y\. Zhang, H\. Jin, D\. Meng, J\. Wang, and J\. Tan \(2025b\)A comprehensive survey on automatic text summarization with exploration of llm\-based methods\.Neurocomputing,pp\. 131928\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p3.1)\. - J\. Zhao, M\. Qu, C\. Li, H\. Yan, Q\. Liu, R\. Li, X\. Xie, and J\. Tang \(2023\)Learning on large\-scale text\-attributed graphs via variational inference\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=q0nmYciuuZN)Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p1.1)\. - X\. Zheng, M\. Zhang, C\. Chen, Q\. V\. H\. Nguyen, X\. Zhu, and S\. Pan \(2023\)Structure\-free graph condensation: from large\-scale graphs to condensed graph\-free data\.Advances in Neural Information Processing Systems36,pp\. 6026–6047\.Cited by:[§2](https://arxiv.org/html/2607.20477#S2.p1.1)\. - C\. Zhou, J\. Du, H\. Zhou, H\. Chen, F\. Huang, and X\. Huang \(2025\)Text\-attributed graph learning with coupled augmentations\.InProceedings of the 31st International Conference on Computational Linguistics,pp\. 10865–10876\.Cited by:[§1](https://arxiv.org/html/2607.20477#S1.p2.1),[§2](https://arxiv.org/html/2607.20477#S2.p4.1)\. - Y\. Zhu, Y\. Wang, H\. Shi, and S\. Tang \(2024\)Efficient tuning and inference for large language models on textual graphs\.InProceedings of the Thirty\-Third International Joint Conference on Artificial Intelligence,IJCAI ’24\.External Links:ISBN 978\-1\-956792\-04\-1,[Link](https://doi.org/10.24963/ijcai.2024/634),[Document](https://dx.doi.org/10.24963/ijcai.2024/634)Cited by:[§C\.2](https://arxiv.org/html/2607.20477#A3.SS2.p1.3),[§1](https://arxiv.org/html/2607.20477#S1.p1.1),[§2](https://arxiv.org/html/2607.20477#S2.p4.1),[§5\.3](https://arxiv.org/html/2607.20477#S5.SS3.p1.1)\. ## Appendix AComplete Algorithm and Analysis This section provides a formal algorithmic description of theSTADframework along with a detailed complexity analysis of its three core modules\. Algorithm Overview\.Algorithms[1](https://arxiv.org/html/2607.20477#algorithm1),[2](https://arxiv.org/html/2607.20477#algorithm2), and[3](https://arxiv.org/html/2607.20477#algorithm3)present the completeSTADpipeline, which takes a large\-scale TAG𝒢\\mathcal\{G\}as input and produces a compact synthetic TAG𝒢′\\mathcal\{G\}^\{\\prime\}with human\-readable text attributes\. Notation\.We follow Section[3\.1](https://arxiv.org/html/2607.20477#S3.SS1)for basic TAG notations:N=\|𝒱\|N=\|\\mathcal\{V\}\|,M=\|ℰ\|M=\|\\mathcal\{E\}\|,DDthe feature dimension,KKthe number of classes\. For Module 1 \(§[4\.1](https://arxiv.org/html/2607.20477#S4.SS1)\): dual\-pathway encodersGCNandMLP, hidden dimensionhh, attention parameters𝐖att∈ℝh\\mathbf\{W\}\_\{\\text\{att\}\}\\in\\mathbb\{R\}^\{h\}and𝐛att∈ℝ\\mathbf\{b\}\_\{\\text\{att\}\}\\in\\mathbb\{R\}, soft labels𝐏GA\\mathbf\{P\}^\{\\text\{\\tiny GA\}\},𝐏GF\\mathbf\{P\}^\{\\text\{\\tiny GF\}\},𝐏\\mathbf\{P\}, confidence scoresmax1≤k≤K𝐏i,kGA\\max\_\{1\\leq k\\leq K\}\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}\_\{i,k\}, pseudo\-label ratioρ\\rhowith increment step0\.10\.1, consensus setℛCS\\mathcal\{R\}^\{\\text\{\\tiny CS\}\}and disagreement subsetsℛGADS\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GA\}\},ℛGFDS\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GF\}\}, labeled sets𝒰\\mathcal\{U\},𝒰GA\\mathcal\{U\}^\{\\text\{\\tiny GA\}\},𝒰GF\\mathcal\{U\}^\{\\text\{\\tiny GF\}\}\. For Module 2 \(§[4\.2](https://arxiv.org/html/2607.20477#S4.SS2)\): target condensed sizeN′N^\{\\prime\}, sketching matrix𝐒∈ℝN×N′\\mathbf\{S\}\\in\\mathbb\{R\}^\{N\\times N^\{\\prime\}\}, clusters\{𝒞1,…,𝒞N′\}\\\{\\mathcal\{C\}\_\{1\},\\ldots,\\mathcal\{C\}\_\{N^\{\\prime\}\}\\\}, K\-Means iterations\. For Module 3 \(§[4\.3](https://arxiv.org/html/2607.20477#S4.SS3)\): keyword countξ\\xi, intra\-cluster TF\-IDF, keywords𝒲i\\mathcal\{W\}\_\{i\}, candidate countqq\(typicallyq=3q=3\), prompt𝒫\\mathcal\{P\}, sample countnsn\_\{s\}\(typicallyns=100n\_\{s\}=100\), top\-kkvalidation \(oftenk=5k=5\), text encoder \(SBERT\)\. Input:TAG 𝒢\\mathcal\{G\}with 𝐀\\mathbf\{A\}, 𝐗\\mathbf\{X\}, 𝐘\\mathbf\{Y\}, and initial labeled set 𝒱tr\\mathcal\{V\}\_\{\\text\{tr\}\}\. Output:Soft labels 𝐏\\mathbf\{P\}and labeled sets 𝒰\\mathcal\{U\}, 𝒰GA\\mathcal\{U\}^\{\\text\{\\tiny GA\}\}, 𝒰GF\\mathcal\{U\}^\{\\text\{\\tiny GF\}\}\. 0\.5exStep 1 — Dual\-pathway encoding\.; 𝐇GA←GCN\(𝐀,𝐗\)\\mathbf\{H\}^\{\\text\{\\tiny GA\}\}\\leftarrow\\textsf\{GCN\}\(\\mathbf\{A\},\\mathbf\{X\}\); 𝐇GF←MLP\(𝐗\)\\mathbf\{H\}^\{\\text\{\\tiny GF\}\}\\leftarrow\\textsf\{MLP\}\(\\mathbf\{X\}\); 𝐏GA←softmax\(𝐇GA\)\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}\\leftarrow\\textsf\{softmax\}\(\\mathbf\{H\}^\{\\text\{\\tiny GA\}\}\); 𝐏GF←softmax\(𝐇GF\)\\mathbf\{P\}^\{\\text\{\\tiny GF\}\}\\leftarrow\\textsf\{softmax\}\(\\mathbf\{H\}^\{\\text\{\\tiny GF\}\}\); for*i=1i=1toNN*do αiGA←σ\(𝐖att⋅𝐇iGA\+𝐛att\)\\alpha^\{\\text\{\\tiny GA\}\}\_\{i\}\\leftarrow\\sigma\(\\mathbf\{W\}\_\{\\text\{att\}\}\\cdot\\mathbf\{H\}\_\{i\}^\{\\text\{\\tiny GA\}\}\+\\mathbf\{b\}\_\{\\text\{att\}\}\); αiGF←1−αiGA\\alpha^\{\\text\{\\tiny GF\}\}\_\{i\}\\leftarrow 1\-\\alpha^\{\\text\{\\tiny GA\}\}\_\{i\}; 𝐏i←softmax\(αiGA⊙𝐇iGA\+αiGF⊙𝐇iGF\)\\mathbf\{P\}\_\{i\}\\leftarrow\\textsf\{softmax\}\(\\alpha^\{\\text\{\\tiny GA\}\}\_\{i\}\\odot\\mathbf\{H\}\_\{i\}^\{\\text\{\\tiny GA\}\}\+\\alpha^\{\\text\{\\tiny GF\}\}\_\{i\}\\odot\\mathbf\{H\}\_\{i\}^\{\\text\{\\tiny GF\}\}\); end for 0\.5exStep 2 — Collaborative self\-training\.; Initialize 𝒰←𝒱tr\\mathcal\{U\}\\leftarrow\\mathcal\{V\}\_\{\\text\{tr\}\}, 𝒰GA←𝒱tr\\mathcal\{U\}^\{\\text\{\\tiny GA\}\}\\leftarrow\\mathcal\{V\}\_\{\\text\{tr\}\}, 𝒰GF←𝒱tr\\mathcal\{U\}^\{\\text\{\\tiny GF\}\}\\leftarrow\\mathcal\{V\}\_\{\\text\{tr\}\}; ρ←0\.1\\rho\\leftarrow 0\.1; while*ρ≤1\.0\\rho\\leq 1\.0*do Recompute 𝐇GA\\mathbf\{H\}^\{\\text\{\\tiny GA\}\}, 𝐇GF\\mathbf\{H\}^\{\\text\{\\tiny GF\}\}, 𝐏GA\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}, 𝐏GF\\mathbf\{P\}^\{\\text\{\\tiny GF\}\}, 𝐏\\mathbf\{P\}; ℛ←\\mathcal\{R\}\\leftarrowtop\- ρ\\rhoproportion of unlabeled nodes 𝒱∖𝒰\\mathcal\{V\}\\setminus\\mathcal\{U\}by confidence maxk𝐏i,kGA\\max\_\{k\}\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}\_\{i,k\}; ℛCS←\{vi∈ℛ∣argmax𝑘𝐏i,kGA=argmax𝑘𝐏i,kGF\}\\mathcal\{R\}^\{\\text\{\\tiny CS\}\}\\leftarrow\\\{v\_\{i\}\\in\\mathcal\{R\}\\mid\\underset\{k\}\{\\operatorname\{arg\}\\,\\operatorname\{max\}\}\\;\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}\_\{i,k\}=\\underset\{k\}\{\\operatorname\{arg\}\\,\\operatorname\{max\}\}\\;\\mathbf\{P\}^\{\\text\{\\tiny GF\}\}\_\{i,k\}\\\}; ℛDS←ℛ∖ℛCS\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\\leftarrow\\mathcal\{R\}\\setminus\\mathcal\{R\}^\{\\text\{\\tiny CS\}\}; ℛGADS←\{vi∈ℛDS∣maxk𝐏i,kGA\>maxk𝐏i,kGF\}\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GA\}\}\\leftarrow\\\{v\_\{i\}\\in\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\\mid\\max\_\{k\}\\mathbf\{P\}^\{\\text\{\\tiny GA\}\}\_\{i,k\}\>\\max\_\{k\}\\mathbf\{P\}^\{\\text\{\\tiny GF\}\}\_\{i,k\}\\\}; ℛGFDS←ℛDS∖ℛGADS\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GF\}\}\\leftarrow\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\\setminus\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GA\}\}; 𝒰←𝒰∪ℛCS∪ℛGADS∪ℛGFDS\\mathcal\{U\}\\leftarrow\\mathcal\{U\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny CS\}\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GA\}\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GF\}\}; 𝒰GA←𝒰GA∪ℛCS∪ℛGADS\\mathcal\{U\}^\{\\text\{\\tiny GA\}\}\\leftarrow\\mathcal\{U\}^\{\\text\{\\tiny GA\}\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny CS\}\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GA\}\}; 𝒰GF←𝒰GF∪ℛCS∪ℛGFDS\\mathcal\{U\}^\{\\text\{\\tiny GF\}\}\\leftarrow\\mathcal\{U\}^\{\\text\{\\tiny GF\}\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny CS\}\}\\cup\\mathcal\{R\}^\{\\text\{\\tiny DS\}\}\_\{\\text\{\\tiny GF\}\}; Update 𝐖att\\mathbf\{W\}\_\{\\text\{att\}\},GCN,MLPby minimizing ℒ=CE\(𝐏,𝐘^\)\+CE\(𝐏GF,𝐘^GF\)\+CE\(𝐏GA,𝐘^GA\)\\mathcal\{L\}=\\textsf\{CE\}\(\\mathbf\{P\},\\hat\{\\mathbf\{Y\}\}\)\+\\textsf\{CE\}\(\\mathbf\{P\}^\{\\text\{\\tiny GF\}\},\\hat\{\\mathbf\{Y\}\}^\{\\text\{\\tiny GF\}\}\)\+\\textsf\{CE\}\(\\mathbf\{P\}^\{\\text\{\\tiny GA\}\},\\hat\{\\mathbf\{Y\}\}^\{\\text\{\\tiny GA\}\}\); ρ←ρ\+0\.1\\rho\\leftarrow\\rho\+0\.1; end while return 𝐏\\mathbf\{P\}, 𝒰\\mathcal\{U\}, 𝒰GA\\mathcal\{U\}^\{\\text\{\\tiny GA\}\}, 𝒰GF\\mathcal\{U\}^\{\\text\{\\tiny GF\}\}; Algorithm 1Graph\-Text Collaborative EncodingComplexity Analysis\.We analyze the time complexity of each module\. LetNNbe the number of nodes,MMthe number of edges,DDthe feature dimension of𝐗\\mathbf\{X\},LLthe number of GCN layers,hhthe hidden dimension,KKthe number of classes,T=10T=10the number of self\-training epochs \(ρ\\rhofrom0\.10\.1to1\.01\.0with step0\.10\.1\),IkmI\_\{\\text\{km\}\}the number of K\-Means iterations,WWthe average text length,VVthe total vocabulary size \(TF\-IDF\),𝒯LLM\\mathcal\{T\}\_\{\\text\{LLM\}\}the cost of one LLM call,IsinkI\_\{\\text\{sink\}\}the number of Sinkhorn iterations for WSD computation, andCCthe size of the sampled node subset used in WSD \(N′<C<NN^\{\\prime\}<C<N\)\. Module 1: Graph\-Text Collaborative Encoding\.The GA pathway \(GCN withLLlayers\) costs𝒪\(L\(MD\+Nh2\)\)\\mathcal\{O\}\(L\(MD\+Nh^\{2\}\)\)for sparse matrix multiplication and transformations\. The GF pathway \(MLP\) costs𝒪\(NDh\)\\mathcal\{O\}\(NDh\)\. Attention fusion requires𝒪\(Nh\)\\mathcal\{O\}\(Nh\)\. Collaborative self\-training repeats forTTiterations, each involving confidence selection𝒪\(NlogN\)\\mathcal\{O\}\(N\\log N\), consensus verification𝒪\(N\)\\mathcal\{O\}\(N\), and backpropagation𝒪\(L\(MD\+Nh2\)\)\\mathcal\{O\}\(L\(MD\+Nh^\{2\}\)\)\. Total for Module 1: \(12\)𝒯M1=𝒪\(T⋅L\(MD\+Nh2\)\+TNlogN\)\\mathcal\{T\}\_\{\\text\{M1\}\}=\\mathcal\{O\}\(T\\cdot L\(MD\+Nh^\{2\}\)\+TN\\log N\) Module 2: WSD\-based Graph Sketching\.K\-Means clustering onNNsoft label vectors of dimensionKKintoN′N^\{\\prime\}clusters costs𝒪\(N⋅N′⋅K⋅Ikm\)\\mathcal\{O\}\(N\\cdot N^\{\\prime\}\\cdot K\\cdot I\_\{\\text\{km\}\}\)\. Greedy reassignment involves computing class\-cluster affinities𝒪\(NK\)\\mathcal\{O\}\(NK\)and distance computations𝒪\(NKN′\)\\mathcal\{O\}\(NKN^\{\\prime\}\)\. Graph compression via sparse matrix operations costs𝒪\(ND\+M\+N′2\)\\mathcal\{O\}\(ND\+M\+N^\{\\prime 2\}\)\. Total for Module 2: \(13\)𝒯M2=𝒪\(NN′K⋅Ikm\+ND\+M\+N′2\)\\mathcal\{T\}\_\{\\text\{M2\}\}=\\mathcal\{O\}\(NN^\{\\prime\}K\\cdot I\_\{\\text\{km\}\}\+ND\+M\+N^\{\\prime 2\}\) Input:Soft labels 𝐏\\mathbf\{P\}, 𝐀\\mathbf\{A\}, 𝐗\\mathbf\{X\}, 𝐘\\mathbf\{Y\}, and target size N′N^\{\\prime\}\. Output:Sketching matrix 𝐒\\mathbf\{S\}and condensed 𝐀′\\mathbf\{A\}^\{\\prime\}, 𝐗′\\mathbf\{X\}^\{\\prime\}, 𝐘′\\mathbf\{Y\}^\{\\prime\}\. 0\.5exStep 1 —KK\-Means clustering on soft labels\.; \{𝒞1,𝒞2,…,𝒞N′\}←K\-Means\(𝐏,N′\)\\\{\\mathcal\{C\}\_\{1\},\\mathcal\{C\}\_\{2\},\\ldots,\\mathcal\{C\}\_\{N^\{\\prime\}\}\\\}\\leftarrow\\textsf\{K\-Means\}\(\\mathbf\{P\},N^\{\\prime\}\); Construct 𝐒\\mathbf\{S\}s\.t\. 𝐒k,i=1\|𝒞k\|\\mathbf\{S\}\_\{k,i\}=\\frac\{1\}\{\|\\mathcal\{C\}\_\{k\}\|\}if vi∈𝒞kv\_\{i\}\\in\\mathcal\{C\}\_\{k\}, else 0; 0\.5exStep 2 — Greedy reassignment post\-processing\.; for*k=1k=1toKK*do for*j=1j=1toN′N^\{\\prime\}*do Compute affinity \|\{vi∈𝒞j\|𝐘i,k=1\}\|\|\{vi∈𝒱\|𝐘i,k=1\}\|\\frac\{\|\\\{v\_\{i\}\\in\\mathcal\{C\}\_\{j\}\|\\ \\mathbf\{Y\}\_\{i,k\}=1\\\}\|\}\{\|\\\{v\_\{i\}\\in\\mathcal\{V\}\|\\ \\mathbf\{Y\}\_\{i,k\}=1\\\}\|\}; end for end for \{yki,𝒞ki\}i=1N′←\\\{y\_\{k\_\{i\}\},\\mathcal\{C\}\_\{k\_\{i\}\}\\\}\_\{i=1\}^\{N^\{\\prime\}\}\\leftarrowthe pairs with the highest affinity values; Assign nodes with label ykiy\_\{k\_\{i\}\}to 𝒞ki∀1≤i≤N′\\mathcal\{C\}\_\{k\_\{i\}\}\\ \\forall\{1\\leq i\\leq N^\{\\prime\}\}; for*each nodeviv\_\{i\}w/o labels\{yki\}i=1N′\\\{y\_\{k\_\{i\}\}\\\}\_\{i=1\}^\{N^\{\\prime\}\}*do Add viv\_\{i\}to 𝒞k\\mathcal\{C\}\_\{k\}s\.t\. min1≤k≤N′‖𝐏i−Mean\(𝒞k\)‖2\\min\_\{1\\leq k\\leq N^\{\\prime\}\}\\\|\\mathbf\{P\}\_\{i\}\-\\textsf\{Mean\}\(\\mathcal\{C\}\_\{k\}\)\\\|\_\{2\}; end for Update 𝐒\\mathbf\{S\}accordingly; 0\.5exStep 3 — Graph compression\.; 𝐀′←𝐒⊤𝐀𝐒\\mathbf\{A\}^\{\\prime\}\\leftarrow\\mathbf\{S\}^\{\\top\}\\mathbf\{A\}\\mathbf\{S\}; 𝐗′←𝐒⊤𝐗\\mathbf\{X\}^\{\\prime\}\\leftarrow\\mathbf\{S\}^\{\\top\}\\mathbf\{X\}; 𝐘i,k′←\{1ifk=argmax1≤k′≤K\(𝐒⊤𝐏\)i,k′0otherwise\\mathbf\{Y\}^\{\\prime\}\_\{i,k\}\\leftarrow\\begin\{cases\}1&\\text\{if \}k=\\underset\{1\\leq k^\{\\prime\}\\leq K\}\{\\operatorname\{arg\}\\,\\operatorname\{max\}\}\\;\{\(\\mathbf\{S\}^\{\\top\}\\mathbf\{P\}\)\}\_\{i,k^\{\\prime\}\}\\\\ 0&\\text\{otherwise\}\\end\{cases\}; return 𝐒\\mathbf\{S\}, 𝐀′\\mathbf\{A\}^\{\\prime\}, 𝐗′\\mathbf\{X\}^\{\\prime\}, 𝐘′\\mathbf\{Y\}^\{\\prime\}; Algorithm 2WSD\-based Graph SketchingModule 3: Keywords\-based Text Synthesis\.TF\-IDF computation acrossN′N^\{\\prime\}clusters with total vocabulary sizeVVcosts𝒪\(VW⋅N′\+N⋅W\)\\mathcal\{O\}\(VW\\cdot N^\{\\prime\}\+N\\cdot W\)\. SBERT\-based keyword filtering costs𝒪\(N′ξD\)\\mathcal\{O\}\(N^\{\\prime\}\\xi D\)\. LLM generation forN′N^\{\\prime\}nodes withqqcandidates each costs𝒪\(N′q⋅𝒯LLM\)\\mathcal\{O\}\(N^\{\\prime\}q\\cdot\\mathcal\{T\}\_\{\\text\{LLM\}\}\)\. WSD computation fornsn\_\{s\}candidate graphs using Sinkhorn algorithm costs𝒪\(ns\(N\+N′\)2Isink\)\\mathcal\{O\}\(n\_\{s\}\(N\+N^\{\\prime\}\)^\{2\}I\_\{\\text\{sink\}\}\)\. In practice, sampling a subset of original nodes with sizeCC\(N′<C<NN^\{\\prime\}<C<N\) is enough, so that the Sinkhorn algorithm can be𝒪\(ns\(C\+N′\)2Isink\)\\mathcal\{O\}\(n\_\{s\}\(C\+N^\{\\prime\}\)^\{2\}I\_\{\\text\{sink\}\}\)\. Total for Module 3: \(14\)𝒯M3=𝒪\(VWN′\+N′ξD\+N′q⋅𝒯LLM\+ns\(C\+N′\)2Isink\)\\mathcal\{T\}\_\{\\text\{M3\}\}=\\mathcal\{O\}\(VWN^\{\\prime\}\+N^\{\\prime\}\\xi D\+N^\{\\prime\}q\\cdot\\mathcal\{T\}\_\{\\text\{LLM\}\}\+n\_\{s\}\(C\+N^\{\\prime\}\)^\{2\}I\_\{\\text\{sink\}\}\) Overall Complexity\.The total time complexity is𝒯STAD=𝒯M1\+𝒯M2\+𝒯M3\\mathcal\{T\}\_\{\\texttt\{STAD\}\{\}\}=\\mathcal\{T\}\_\{\\text\{M1\}\}\+\\mathcal\{T\}\_\{\\text\{M2\}\}\+\\mathcal\{T\}\_\{\\text\{M3\}\}\. The space complexity is dominated by storing the adjacency matrix \(𝒪\(M\)\\mathcal\{O\}\(M\)\), feature matrices \(𝒪\(ND\+N′h\)\\mathcal\{O\}\(ND\+N^\{\\prime\}h\)\), soft labels \(𝒪\(NK\)\\mathcal\{O\}\(NK\)\), and candidate texts \(𝒪\(N′qW\)\\mathcal\{O\}\(N^\{\\prime\}qW\)\), yielding𝒮STAD=𝒪\(M\+ND\+N′h\+NK\+N′qW\)\\mathcal\{S\}\_\{\\texttt\{STAD\}\{\}\}=\\mathcal\{O\}\(M\+ND\+N^\{\\prime\}h\+NK\+N^\{\\prime\}qW\)\. In practice, sinceN′≪NN^\{\\prime\}\\ll N, Module 3 operates efficiently despite LLM calls\. The dominant cost in Module 1 is the GCN operations scaling linearly with edgesMM\. Module 2’s clustering is efficient whenN′N^\{\\prime\}is small\. The WSD validation in Module 3 uses a smallnsn\_\{s\}\(e\.g\.,100100\) to balance quality and cost\. Input:Sketching matrix 𝐒\\mathbf\{S\}, clusters \{𝒞1,…,𝒞N′\}\\\{\\mathcal\{C\}\_\{1\},\\ldots,\\mathcal\{C\}\_\{N^\{\\prime\}\}\\\}, text 𝒮\\mathcal\{S\}, labels 𝐘′\\mathbf\{Y\}^\{\\prime\}, integers ξ\\xi, qq, nsn\_\{s\}, kk, and prompt 𝒫\\mathcal\{P\}\. Output:Synthetic texts 𝒮′=\{s1′,…,sN′′\}\\mathcal\{S\}^\{\\prime\}=\\\{s^\{\\prime\}\_\{1\},\\ldots,s^\{\\prime\}\_\{N^\{\\prime\}\}\\\}for 𝒢′\\mathcal\{G\}^\{\\prime\}\. 0\.5exStep 1 — Cluster\-based keyword extraction\.; for*i=1i=1toN′N^\{\\prime\}*do 𝒮i←\{sj∣vj∈𝒞i\}\\mathcal\{S\}\_\{i\}\\leftarrow\\\{s\_\{j\}\\mid v\_\{j\}\\in\\mathcal\{C\}\_\{i\}\\\}; TF\-IDF\(w,𝒮i\)=TF\(w,𝒮i\)×IDF\(w,𝒮i\)\\text\{TF\-IDF\}\(w,\\mathcal\{S\}\_\{i\}\)=\\text\{TF\}\(w,\\mathcal\{S\}\_\{i\}\)\\times\\text\{IDF\}\(w,\\mathcal\{S\}\_\{i\}\)∀w∈𝒮i\\forall\{w\\in\\mathcal\{S\}\_\{i\}\}; Select top\- 2ξ2\\xiwords by intra\-cluster TF\-IDF; Prune ξ\\xiwords with lowest embedding distance to 𝐗i′\\mathbf\{X\}^\{\\prime\}\_\{i\}; Obtain keywords 𝒲i=\{wi,1,…,wi,ξ\}\\mathcal\{W\}\_\{i\}=\\\{w\_\{i,1\},\\ldots,w\_\{i,\\xi\}\\\}; end for 0\.5exStep 2 — Candidate text generation with LLMs\.; for*i=1i=1toN′N^\{\\prime\}*do yi′←argmax𝑘𝐘i,k′y^\{\\prime\}\_\{i\}\\leftarrow\\underset\{k\}\{\\operatorname\{arg\}\\,\\operatorname\{max\}\}\\;\\mathbf\{Y\}^\{\\prime\}\_\{i,k\}; \{si,1′,…,si,q′\}←LLM\(𝒫,𝒲i,yi′\)\\\{s^\{\\prime\}\_\{i,1\},\\ldots,s^\{\\prime\}\_\{i,q\}\\\}\\leftarrow\\textsf\{LLM\}\(\\mathcal\{P\},\\mathcal\{W\}\_\{i\},y^\{\\prime\}\_\{i\}\); end for 0\.5exStep 3 — WSD\-validation\-guided selection\.; for*l=1l=1tonsn\_\{s\}*do 𝒮l′←\{si,ji′∣i∈\[N′\],ji∼Uniform\(\[q\]\)\}\\mathcal\{S\}^\{\\prime\}\_\{l\}\\leftarrow\\\{s^\{\\prime\}\_\{i,j\_\{i\}\}\\mid i\\in\[N^\{\\prime\}\],j\_\{i\}\\sim\\text\{Uniform\}\(\[q\]\)\\\}; 𝐗l′←SBERT\(𝒮l′\)\\mathbf\{X\}^\{\\prime\}\_\{l\}\\leftarrow\\text\{SBERT\}\(\\mathcal\{S\}^\{\\prime\}\_\{l\}\); Compute WSDl←WSD\(𝐗,𝐗l′\)\\text\{WSD\}\_\{l\}\\leftarrow\\textsf\{WSD\}\(\\mathbf\{X\},\\mathbf\{X\}^\{\\prime\}\_\{l\}\); end for Select top\- kkcandidates with the lowest WSD; Choose 𝒮′\\mathcal\{S\}^\{\\prime\}with the best downstream task performance; return 𝒮′\\mathcal\{S\}^\{\\prime\}; Algorithm 3Keywords\-based Text Synthesis with LLMs ## Appendix BTheoretical Proof ###### Proof of Lemma[4\.1](https://arxiv.org/html/2607.20477#S4.Thmtheorem1)\. We establish this equivalence through the geometric form of the 2\-Wasserstein distance in the context of graph sketching\. Step 1: Distribution construction\.Let𝒫=1N∑i=1Nδ𝐌i\\mathcal\{P\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\delta\_\{\\mathbf\{M\}\_\{i\}\}denote the empirical distribution of the original features, whereδ𝐌i\\delta\_\{\\mathbf\{M\}\_\{i\}\}is the Dirac delta distribution centered at𝐌i\\mathbf\{M\}\_\{i\}\. The sketching matrix𝐒∈ℝN×N′\\mathbf\{S\}\\in\\mathbb\{R\}^\{N\\times N^\{\\prime\}\}induces a hard clustering partition\{𝒞k\}k=1N′\\\{\\mathcal\{C\}\_\{k\}\\\}\_\{k=1\}^\{N^\{\\prime\}\}of the node set\[N\]\[N\], where𝒞k=\{i∈\[N\]:𝐒k,i≠0\}\\mathcal\{C\}\_\{k\}=\\\{i\\in\[N\]:\\mathbf\{S\}\_\{k,i\}\\neq 0\\\}represents the set of nodes assigned to thekk\-th condensed node\. By construction of𝐒\\mathbf\{S\}in Algorithm[2](https://arxiv.org/html/2607.20477#algorithm2), each rowkksatisfies𝐒k,i=1\|𝒞k\|\\mathbf\{S\}\_\{k,i\}=\\frac\{1\}\{\|\\mathcal\{C\}\_\{k\}\|\}ifi∈𝒞ki\\in\\mathcal\{C\}\_\{k\}, and𝐒k,i=0\\mathbf\{S\}\_\{k,i\}=0otherwise\. Consequently, the compressed empirical distribution is𝒬=∑k=1N′\|𝒞k\|Nδ𝝁k\\mathcal\{Q\}=\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\frac\{\|\\mathcal\{C\}\_\{k\}\|\}\{N\}\\delta\_\{\\boldsymbol\{\\mu\}\_\{k\}\}, where the centroid𝝁k=∑j=1N𝐒k,j𝐌j=1\|𝒞k\|∑j∈𝒞k𝐌j\\boldsymbol\{\\mu\}\_\{k\}=\\sum\_\{j=1\}^\{N\}\\mathbf\{S\}\_\{k,j\}\\mathbf\{M\}\_\{j\}=\\frac\{1\}\{\|\\mathcal\{C\}\_\{k\}\|\}\\sum\_\{j\\in\\mathcal\{C\}\_\{k\}\}\\mathbf\{M\}\_\{j\}\. Step 2: Voronoi property verification\.We verify that the partition\{𝒞k\}k=1N′\\\{\\mathcal\{C\}\_\{k\}\\\}\_\{k=1\}^\{N^\{\\prime\}\}satisfies the Voronoi property under the squared Euclidean metric\. For any sample𝐌i∈𝒞k\\mathbf\{M\}\_\{i\}\\in\\mathcal\{C\}\_\{k\}and any other clusterk′≠kk^\{\\prime\}\\neq k, the K\-Means clustering in Algorithm[2](https://arxiv.org/html/2607.20477#algorithm2)\(Step 1\) assigns each node to the nearest centroid, ensuring: ‖𝐌i−𝝁k‖22≤‖𝐌i−𝝁k′‖22\.\\\|\\mathbf\{M\}\_\{i\}\-\\boldsymbol\{\\mu\}\_\{k\}\\\|\_\{2\}^\{2\}\\leq\\\|\\mathbf\{M\}\_\{i\}\-\\boldsymbol\{\\mu\}\_\{k^\{\\prime\}\}\\\|\_\{2\}^\{2\}\. Step 3: Optimal transport plan\.Consider the transport planγ∗∈ℝN×N′\\gamma^\{\*\}\\in\\mathbb\{R\}^\{N\\times N^\{\\prime\}\}induced by the clustering: γik∗=\{1N,if𝐌i∈𝒞k,0,otherwise\.\\gamma^\{\*\}\_\{ik\}=\\begin\{cases\}\\frac\{1\}\{N\},&\\text\{if \}\\mathbf\{M\}\_\{i\}\\in\\mathcal\{C\}\_\{k\},\\\\ 0,&\\text\{otherwise\}\.\\end\{cases\} This plan is feasible: \(i\)∑k=1N′γik∗=1N\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\gamma^\{\*\}\_\{ik\}=\\frac\{1\}\{N\}for allii\(marginal of𝒫\\mathcal\{P\}\), and \(ii\)∑i=1Nγik∗=\|𝒞k\|N\\sum\_\{i=1\}^\{N\}\\gamma^\{\*\}\_\{ik\}=\\frac\{\|\\mathcal\{C\}\_\{k\}\|\}\{N\}for allkk\(marginal of𝒬\\mathcal\{Q\}\)\. Step 4: Optimality via Voronoi property\.For any feasible transport planγ\\gamma, the Voronoi property yields the pointwise lower bound for eachii: ∑k=1N′γik‖𝐌i−𝝁k‖22≥∑k=1N′γik‖𝐌i−𝝁k\(i\)‖22=1N‖𝐌i−𝝁k\(i\)‖22,\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\gamma\_\{ik\}\\\|\\mathbf\{M\}\_\{i\}\-\\boldsymbol\{\\mu\}\_\{k\}\\\|\_\{2\}^\{2\}\\geq\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\gamma\_\{ik\}\\\|\\mathbf\{M\}\_\{i\}\-\\boldsymbol\{\\mu\}\_\{k\(i\)\}\\\|\_\{2\}^\{2\}=\\frac\{1\}\{N\}\\\|\\mathbf\{M\}\_\{i\}\-\\boldsymbol\{\\mu\}\_\{k\(i\)\}\\\|\_\{2\}^\{2\},wherek\(i\)k\(i\)denotes the cluster index of𝐌i\\mathbf\{M\}\_\{i\}\. Summing over alliigives: ∑i=1N∑k=1N′γik‖𝐌i−𝝁k‖22≥1N∑i=1N‖𝐌i−𝝁k\(i\)‖22=∑i=1N∑k=1N′γik∗‖𝐌i−𝝁k‖22\.\\sum\_\{i=1\}^\{N\}\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\gamma\_\{ik\}\\\|\\mathbf\{M\}\_\{i\}\-\\boldsymbol\{\\mu\}\_\{k\}\\\|\_\{2\}^\{2\}\\geq\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\\|\\mathbf\{M\}\_\{i\}\-\\boldsymbol\{\\mu\}\_\{k\(i\)\}\\\|\_\{2\}^\{2\}=\\sum\_\{i=1\}^\{N\}\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\gamma^\{\*\}\_\{ik\}\\\|\\mathbf\{M\}\_\{i\}\-\\boldsymbol\{\\mu\}\_\{k\}\\\|\_\{2\}^\{2\}\. Thusγ∗\\gamma^\{\*\}is optimal, and the squared WSD equals: WSD2\(𝒫,𝒬\)=1N∑k=1N′∑𝐌i∈𝒞k‖𝐌i−𝝁k‖22\.\\textsf\{WSD\}^\{2\}\(\\mathcal\{P\},\\mathcal\{Q\}\)=\\frac\{1\}\{N\}\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\sum\_\{\\mathbf\{M\}\_\{i\}\\in\\mathcal\{C\}\_\{k\}\}\\\|\\mathbf\{M\}\_\{i\}\-\\boldsymbol\{\\mu\}\_\{k\}\\\|\_\{2\}^\{2\}\. Step 5: Equivalence to sketching objective\.Rewriting using𝟙\[𝐒k,i≠0\]=𝟙\[i∈𝒞k\]\\mathbb\{1\}\_\{\[\\mathbf\{S\}\_\{k,i\}\\neq 0\]\}=\\mathbb\{1\}\_\{\[i\\in\\mathcal\{C\}\_\{k\}\]\}and substituting𝝁k=∑j=1N𝐒k,j𝐌j\\boldsymbol\{\\mu\}\_\{k\}=\\sum\_\{j=1\}^\{N\}\\mathbf\{S\}\_\{k,j\}\\mathbf\{M\}\_\{j\}: WSD2\(𝐌,𝐒⊤𝐌\)\\displaystyle\\textsf\{WSD\}^\{2\}\(\\mathbf\{M\},\\mathbf\{S\}^\{\\top\}\\mathbf\{M\}\)=1N∑k=1N′∑i=1N𝟙\[𝐒k,i≠0\]⋅‖𝐌i−𝝁k‖22\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\sum\_\{i=1\}^\{N\}\\mathbb\{1\}\_\{\[\\mathbf\{S\}\_\{k,i\}\\neq 0\]\}\\cdot\\left\\\|\\mathbf\{M\}\_\{i\}\-\\boldsymbol\{\\mu\}\_\{k\}\\right\\\|\_\{2\}^\{2\}=1N∑k=1N′∑i=1N𝟙\[𝐒k,i≠0\]⋅‖𝐌i−∑j=1N𝐒k,j𝐌j‖22\.\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{k=1\}^\{N^\{\\prime\}\}\\sum\_\{i=1\}^\{N\}\\mathbb\{1\}\_\{\[\\mathbf\{S\}\_\{k,i\}\\neq 0\]\}\\cdot\\left\\\|\\mathbf\{M\}\_\{i\}\-\\sum\_\{j=1\}^\{N\}\\mathbf\{S\}\_\{k,j\}\\mathbf\{M\}\_\{j\}\\right\\\|\_\{2\}^\{2\}\. Therefore, minimizingWSD\(𝐌,𝐒⊤𝐌\)\\textsf\{WSD\}\(\\mathbf\{M\},\\mathbf\{S\}^\{\\top\}\\mathbf\{M\}\)is equivalent to the stated minimization problem, which corresponds precisely to the K\-Means clustering objective in the feature space\. ∎ ## Appendix CTAG Distillation Time and LLM4TAG training Time ### C\.1\.TAG Distillation Time Comparison Here we give a simple comparison betweenSTADandClustGDDin terms of distillation time, another efficient graph distillation method, atr=0\.05r=0\.05onCoraandHistory\. We selectClustGDDas the baseline because it achieves both high efficiency and accuracy, being orders of magnitude faster than other gradient\-based or distribution\-matching distillation methods while maintaining competitive performance\. Experimental results demonstrate the efficiency trade\-offs of theSTADmethod: when utilizing an LLM for text synthesis inSTAD\(𝒮′\)\\texttt\{STAD\}\{\}\(\\mathcal\{S\}^\{\\prime\}\), it consumes191\.67191\.67seconds and291\.24291\.24seconds onCoraandHistorydatasets respectively, which is slower thanSTAD\(𝐗′\)\\texttt\{STAD\}\{\}\(\\mathbf\{X\}^\{\\prime\}\)\. This overhead primarily stems from calling the LLM to generate𝒮′\\mathcal\{S\}^\{\\prime\}, which incurs high computational complexity\. Meanwhile, both variants lag far behindClustGDD\(6\.796\.79seconds and10\.3110\.31seconds\)\. However,STADachieves higher accuracy and better interpretability thanClustGDD, as demonstrated in previous experiments\. Therefore, a time gap of the same order of magnitude or even one order of magnitude larger is acceptable\. Table 7\.Distillation running time comparison \(seconds\)DatasetSTAD\(𝐗′\)\\texttt\{STAD\}\{\}\(\\mathbf\{X\}^\{\\prime\}\)STAD\(𝒮′\)\\texttt\{STAD\}\{\}\(\\mathcal\{S\}^\{\\prime\}\)ClustGDDCora56\.0191\.76\.79History83\.1291\.210\.31 ### C\.2\.Time Cost of LLM4TAG Methods on𝒢\\mathcal\{G\}and𝒢′\\mathcal\{G\}^\{\\prime\} We compare the preprocessing and training times \(in seconds\) of three LLM4TAG methods,GNN\-LLM\(Wuet al\.,[2025](https://arxiv.org/html/2607.20477#bib.bib24)\),ENGINE\(Zhuet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib27)\), andTAPE\(Heet al\.,[2024](https://arxiv.org/html/2607.20477#bib.bib28)\)on the original graph𝒢\\mathcal\{G\}and the distilled graph𝒢′\\mathcal\{G\}^\{\\prime\}at distillation ratior=0\.5r=0\.5\.GNN\-LLMextracts LLM last\-layer representation to train GNNs;ENGINEfuses intermediate\-layer LLM representations for GNN training;TAPEprompts LLMs to generate predictions and natural language explanations for textual graph data, incurring substantially higher computational costs due to autoregressive generation\. Table 8\.Preprocessing time \(seconds\) of semi\-supervised LLM4TAG methods atr=0\.5r=0\.5on𝒢\\mathcal\{G\}\(Full Methods\) and𝒢′\\mathcal\{G\}^\{\\prime\}\(STAD\)\.DatasetDistillationMethodBaselinesGNN\-LLMENGINETAPECoraFull Methods22\.5111\.02784\.7STAD3\.96\.850\.0HistoryFull Methods385\.21743\.132756\.0STAD5\.68\.575\.0Preprocessing Time Analysis\.The preprocessing time is substantially reduced on the distilled graph𝒢′\\mathcal\{G\}^\{\\prime\}compared to the original graph𝒢\\mathcal\{G\}for all three methods\. The reduction is most pronounced forTAPE, which requires invoking LLMs to generate predictions and explanations for each node\. OnCora,TAPE’s preprocessing time drops from 2,784\.7s to 50\.0s \(55\.7×\\timesspeedup\); on the largerHistorydataset, it decreases from 32,756\.0s to 75\.0s \(436\.7×\\timesspeedup\)\. This dramatic improvement stems fromSTADreducing the graph size, thereby minimizing the number of expensive LLM API calls\.GNN\-LLMandENGINEalso benefit from distillation, though with more modest reductions, as they only require forward passes through the LLM to obtain embeddings or intermediate representations, avoiding the costly autoregressive generation process\. Due to the early stopping,GNN\-LLMspends less time on the larger datasetHistory\. Table 9\.Training time \(seconds\) of semi\-supervised LLM4TAG methods atr=0\.5r=0\.5on𝒢\\mathcal\{G\}\(Full Methods\) and𝒢′\\mathcal\{G\}^\{\\prime\}\(STAD\)\.DatasetDistillationMethodBaselinesGNN\-LLMENGINETAPECoraFull Methods4\.210\.17\.1STAD3\.310\.04\.3HistoryFull Methods3\.5121\.336\.4STAD3\.286\.435\.0Training Time Analysis\.The training time reductions on𝒢′\\mathcal\{G\}^\{\\prime\}are comparatively modest across all methods\. ForGNN\-LLMandTAPE, training times remain similar between𝒢\\mathcal\{G\}and𝒢′\\mathcal\{G\}^\{\\prime\}because the downstream GNN training is already efficient even on the full graph\.ENGINEshows a more noticeable reduction onHistory\(from 121\.3s to 86\.4s\), likely due to its architecture\-specific fusion mechanism benefiting from the smaller distilled graph\. Table 10\.Dataset statistics before/after distillation atr=0\.05r=0\.05: nodes, edges, storage \(MB\), compression ratio, and accuracy\.DatasetStageNodesEdges𝐗\\mathbf\{X\}\(MB\)𝐀\\mathbf\{A\}\(MB\)𝐘\\mathbf\{Y\}\(MB\)𝒮\\mathcal\{S\}\(MB\)Total \(MB\)Compression RatioAcc \(%\)CoraOriginal270852783\.96680\.10070\.02072\.30366\.3918–77\.7±\\pm2\.0Synthetic7470\.01030\.00090\.00010\.00450\.0158403\.39×\\times77\.5±\\pm0\.6CiteseerOriginal318684504\.66700\.16120\.02436\.218111\.0706–74\.0±\\pm2\.1Synthetic6360\.00880\.00070\.00000\.00690\.0164673\.20×\\times75\.9±\\pm0\.5DBLPOriginal1437643132621\.05868\.22690\.10972\.433931\.8291–77\.9±\\pm0\.5Synthetic4160\.00590\.00030\.00000\.00230\.00853739\.87×\\times77\.2±\\pm0\.2ComputersOriginal87229721081127\.776913\.75350\.665542\.7394184\.9353–72\.2±\\pm1\.9Synthetic10990\.01460\.00190\.00010\.01230\.02896397\.35×\\times65\.4±\\pm3\.7PhotoOriginal4836250093970\.84289\.55470\.369037\.8577118\.6242–70\.0±\\pm0\.9Synthetic121390\.01760\.00270\.00010\.01400\.03443451\.25×\\times64\.6±\\pm1\.4HistoryOriginal4155135857460\.86576\.83930\.317057\.3151125\.3371–70\.6±\\pm3\.6Synthetic111060\.01610\.00200\.00010\.01100\.02924291\.29×\\times75\.9±\\pm1\.7WikiCSOriginal1170143172617\.14018\.23450\.0893145\.8948171\.3587–75\.9±\\pm1\.2Synthetic101000\.01460\.00190\.00010\.00620\.02287505\.11×\\times76\.6±\\pm0\.3 ## Appendix DStatistics of Original and Synthetic TAGs Table[10](https://arxiv.org/html/2607.20477#A3.T10)summarizes the dataset statistics and performance before and after distillation at the reduction ratior=0\.05r=0\.05\. We report the number of nodes and edges, storage usage \(in MB\) for attributes, adjacency matrices, labels, and text, as well as total storage consumption, compression ratio, and classification accuracy \(Acc, %\)\. It can be observed that our distillation method achieves extremely high compression ratios across all seven datasets\. The compression ratios range from 403\.39×\\timesfor the Cora dataset to 7505\.11×\\timesfor the WikiCS dataset, indicating a significant reduction in data volume\. Specifically, the total storage of the original datasets, which ranges from several MB to nearly 200 MB, is compressed to only about 0\.01–0\.04 MB for the synthetic datasets\. Meanwhile, the number of nodes is drastically reduced from thousands to tens \(4–12 nodes for synthetic datasets\), and the number of edges is reduced from thousands to hundreds \(16–139 edges for synthetic datasets\), realizing efficient compression of graph structure\. Despite the extreme compression of graph structure and data volume, the synthetic datasets largely preserve the predictive performance of the original graphs\. For most datasets, the classification accuracy of the synthetic datasets is close to that of the original datasets with only minor fluctuations\. Notably, the synthetic datasets of Citeseer, History, and WikiCS achieve comparable or slightly higher accuracy than their original counterparts, which demonstrates that our distillation method can effectively retain the critical structural, attribute, and semantic information required for downstream classification tasks\. For the Computers and Photo datasets, there is a slight drop in accuracy of the synthetic datasets compared to the original ones, but the accuracy remains at a reasonable level, which is acceptable given the extremely high compression ratio\. In summary, the experimental results show that our distillation method can achieve efficient compression of graph datasets while maintaining good performance, verifying the effectiveness and practicality of the method\. ## Appendix EPrompts and Results ### E\.1\.Prompts for text synthesis inSTAD First, we present the complete prompt templates used in the keywords\-based LLM text synthesis inSTAD, including the system initialization template, user task instruction template, and corresponding dataset background descriptions\. System Prompt Template\.This template serves as the initial system instruction to define the role and global context for the researcher\. It provides the basic background of the TAG dataset and lays the foundation for subsequent document summarization tasks\. Dataset Background Descriptions\.To complement the placeholders’dataset\_name’and’sys\_dataset’in the system prompt template, detailed background descriptions of typical datasets used in this task are provided below\. TakingSYS\_CORAandSYS\_HISTORYas examples, we clearly present the structure and category information of each dataset\. User Prompt Template\.Building upon the system prompt, this user prompt template further provides specific task instructions, strict mandatory requirements, and standardized output formats for clustered document summarization\. It clarifies the key constraints and operational details to ensure the validity of the generated outputs\. SystemPromptTemplate: """ WehaveperformedclusteringontheText\-AttributedGraphdatasetandextractedkeywordsforthetextsineachcluster\. Youareaprofessionalresearchertaskedwithsummarizingclustereddocumentsbasedonkeywords\. FollowALLinstructionsstrictly\. \#Text\-AttributedGraphdatasetBackground\(GlobalContext\) Thistaskprocessesdocumentsfromthe"\{dataset\_name\}"dataset: \{sys\_dataset\} """ SYS\_CORA: """ Itisanacademiccitationnetworkinthefieldofmachinelearning,consistingof2708nodes\(eachnoderepresentsamachinelearning\-relatedpaper\),10556edges\(eachedgerepresentsthecitationrelationshipbetweenpapers\),andallnodesaredividedinto7categorylabels,namelyRule\_Learning,Neural\_Networks,Case\_Based,Genetic\_Algorithms,Theory,Reinforcement\_Learning,Probabilistic\_Methods\. """ SYS\_HISTORY: """ ItisabookrecommendationnetworkextractedfromtheAmazon\-Booksdataset,consistingof41551nodes\(eachnoderepresentsabookwiththesecond\-levellabelHistory\),503180edges\(eachedgerepresentstwobooksarefrequentlyco\-purchasedorco\-viewed\),andallnodesaredividedinto12categorylabels\(thethree\-levellabelsofthebooks\):World,Americas,Asia,Military,Europe,Russia,Africa,AncientCivilizations,MiddleEast,HistoricalStudy&EducationalResources,Australia&Oceania,Arctic&Antarctica\. """ UserPromptTemplate: """ TaskInstructions: Clustereddocuments'keywords\(prioritizefirst3\-5ascore\):\{center\_key\} Possibledocuments'category:\{center\_label1\} StrictMandatoryRequirements\(VIOLATION=INVALIDOUTPUT\): 1\.TokenLimit:Totaloutput\(including\[\[LABEL\]\],\[\[SUMMARY\]\],spaces,punctuation\)mustNOTexceed\{tok\_lim\}tokens\. 2\.OutputFormat\(EXACTLYasshown,NOextralines/words/symbols/linebreaks/markdown\): \[\[LABEL\]\]<selectedcategoryname\> \[\[SUMMARY\]\]<singleshortparagraphsummary\> 3\.\[\[LABEL\]\]Rule: \-SelectONLYONEcategoryfromthegivencategory\(\{center\_label1\}\)\. \-OutputONLYthecategoryname\(noadditionaltext,punctuation,orexplanations\)\. 4\.\[\[SUMMARY\]\]Rules\(ALLmustbesatisfied\): \-Length:Max\{tok\_lim\}tokens\(after\[\[SUMMARY\]\]:\)\. \-Content: \-Integratesthefirst3\-5corekeywordsnaturallyintothesummary\(avoidkeywordstacking\);incorporateotherkeywordsonlyiftheyfitthecontextwithoutexceedingtokenlimits; \-Clearlyandconciselysummarizesthecoretheme/keyfindings/primaryfocusoftheclustereddocuments\(1\-2corepointsonly,notrivialdetails\); \-Format:Singleshortparagraph\(nobulletpoints,nolinebreaks,nolists,nomarkdown\)\. OutputNOW\(strictlyfollowallrulesabove\): """ ### E\.2\.Keywords and Synthesis inSTAD To elaborate on the specific results of keyword extraction and cluster summarization, we first analyze theCoraandHistoryunder the condition ofr=0\.05r=0\.05\. We here present the keywords and corresponding cluster summarizations\. Focusing on the classNeural Networkin theCora, we extracted 256 keywords from its corresponding cluster, which are listed as follows\. Keywords: """ classifiers,neural,perceptrons,classifier,backpropagation,supervised,classification,features,recognition,feature,networks,neuron,learns,learning,neurons,regularization,algorithms,cortex,perceptron,adaptive,nodes,models,ensemble,clustering,prediction,generalization,handwritten,machines,computational,segmentation,detection,wavelet,overfitting,connectionist,inputs,computing,feedforward,multilayer,representations,net,predict,trained,sparse,algorithm,synaptic,modeling,optimization,network,brain,patterns,mlp,signals,learned,knowledge,layers,generative,model,receptive,gradient,processing,kernel,stochastic,implementation,\.\.\. """ We can see that the keywords are with high related to the class, shown the effectiveness of our keywords extraction\. Based on the extracted keywords of theNeural Networkcluster, we generated three summarizations as synthetic text to capture the core themes of the cluster, as shown below\. Candidate1: """ \[\[LABEL\]\]Neural\_Networks \[\[SUMMARY\]\]Theclusterfocusesonneuralnetworks,particularlyclassifiersandperceptrons,emphasizingsupervisedlearningandbackpropagationfortaskslikeclassificationandpatternrecognition\.Corethemesincludenetworkarchitecturessuchasmultilayerperceptrons\(MLPs\),regularizationtechniquestocombatoverfitting,andfeaturerepresentationlearning\.Additionaltopicsspangradient\-basedoptimization,activationfunctions,generalization,andapplicationsinvision,signalprocessing,andcomputationalmodelinginspiredbybiologicalneuronsandcorticalstructures\. """ Candidate2: """ \[\[LABEL\]\]Neural\_Networks \[\[SUMMARY\]\]Theclustereddocumentsfocusonneuralnetworks,particularlyexploringclassifiers,neuralarchitectures,andperceptronsinsupervisedlearningcontexts\.Corethemesincludebackpropagationfortrainingmultilayernetworks,featureextractionforclassificationtasks,andstrategiestoimprovegeneralizationwhilemitigatingoverfitting\.Additionaltopicsspanregularization,adaptivelearning,andapplicationsinpatternrecognition,suchashandwrittendigitdetectionandsignalprocessing,oftenleveragingfeedforwardandrecurrentnetworkmodels\. """ Candidate3: """ \[\[LABEL\]\]Neural\_Networks \[\[SUMMARY\]\]Thisclusterfocusesonneuralnetworks,emphasizingclassifiers,neuralarchitectures,andperceptrons,withsignificantattentiontobackpropagationandsupervisedlearningforclassificationtasks\.Corethemesincludenetworkmodelslikemultilayerperceptrons\(MLPs\),featurerepresentation,generalization,andtechniquestomitigateoverfittingthroughregularization\.Additionaltopicsspanlearningalgorithms,adaptivesystems,andapplicationsinpatternrecognition,imageprocessing,andcomputationalneuroscience,oftenleveragingfeedforwardandrecurrentstructures,activationfunctions,andgradient\-basedoptimizationtoimprovepredictionaccuracyandmodelrobustness\. """ Similarly, forAncient CivilizationsinHistory, the keywords and corresponding summarizations are Keywords: """ antiquity,archaeologists,antiquities,archeology,egyptology,archaeology,archaeological,ancient,archaeologist,egyptologist,egyptians,mesopotamia,tombs,historia,histories,historians,archaelogy,monuments,archeologist,tomb,inscriptions,byzantium,babylon,egyptian,romans,byzantine,greeks,anthropology,excavated,antique,coptic,egypt,egypts,grecoroman,history,civilization,bookdescription,hammurabi,museum,paganism,rome,bce,books,chronicle,greek,johannes,depictions,temple,roman,book,emperors,tutankhamun,paintings,cairo,mysteries,turks,statuette,deities,myths,oetempledestructiona,artifact,goddesses,literary,empire,autobiography,\.\.\. """ Candidate1: """ \[\[LABEL\]\]AncientCivilizations \[\[SUMMARY\]\]Thisclustercentersonantiquity,archaeologists,andantiquities,withastrongfocusonarchaeologyandegyptology\.ThedocumentsexploreancientcivilizationssuchasEgypt,Mesopotamia,andtheGreco\-Romanworld,examiningarchaeologicalfindings,tombs,inscriptions,andartifacts\.Keythemesincludethestudyofancientculturesthroughexcavatedmaterials,historicalinterpretationsbyrenownedscholars,andinsightsintoreligiouspractices,monuments,anddailylifeinantiquity\. """ Candidate2: """ \[\[LABEL\]\]AncientCivilizations \[\[SUMMARY\]\]Theclustercentersonantiquity,archaeologists,andantiquities,withastrongfocusonarchaeologyandegyptology\.DocumentsexploreancientcivilizationslikeEgypt,Mesopotamia,andRomethrougharchaeologicalfindings,tombs,inscriptions,andartifacts,highlightingscholarlyresearchondeities,emperors,andhistoricalnarratives\. """ Candidate3: """ \[\[LABEL\]\]AncientCivilizations \[\[SUMMARY\]\]Theclustercentersonantiquity,archaeologists,andantiquities,withastrongfocusonarchaeologyandegyptology\.DocumentsexploreancientcivilizationslikeEgypt,Mesopotamia,andRomethrougharchaeologicalfindings,tombs,inscriptions,andartifacts,highlightingscholarlyresearchondeities,emperors,andhistoricalnarratives\. """ The reason why the summaries forCoraandHistorydatasets differ in length is that the optimal output token limit varies for different downstream classification tasks\. ### E\.3\.Synthetic Texts fromSTAD,DaLLMEComparison We here give a comparison between texts synthesized bySTADandDaLLME\. A qualitative analysis reveals the superiority ofSTAD, which adopts a keyword\-guided LLM framework for cluster summarization, overDaLLME\.DaLLMEdirectly maps cluster representations to natural language without explicit semantic guidance\. As demonstrated in theCoraandHistoryexamples,STADgenerates domain\-consistent summaries with strong topic alignment and fidelity to domain knowledge by grounding generation in cluster keywords, effectively preserving fine\-grained terminology and capturing the full semantic scope of each document cluster\.DaLLME, even with text\-pair fine\-tuning, is constrained by the inherent information bottlenecks of its V2T paradigm, leading to outputs that are relatively less comprehensive\. Here are samples ofCora: STAD: """ \[\[LABEL\]\]Neural\_Networks \[\[SUMMARY\]\]Theclusterfocusesonneuralnetworks,particularlyclassifiersandperceptrons,emphasizingsupervisedlearningandbackpropagationfortaskslikeclassificationandpatternrecognition\.Corethemesincludenetworkarchitecturessuchasmultilayerperceptrons\(MLPs\),regularizationtechniquestocombatoverfitting,andfeaturerepresentationlearning\.Additionaltopicsspangradient\-basedoptimization,activationfunctions,generalization,andapplicationsinvision,signalprocessing,andcomputationalmodelinginspiredbybiologicalneuronsandcorticalstructures\. """ DaLLME: """ "AnalysisofNeuralNetworks,":Inthispaperwepresentamethodforevaluatingtheperformanceofneuralnetworksinavarietyofcontexts\.Themethodisbasedontheideathatneuralnetworkscanbetrainedtoperformagiventaskinagiventimeframe """ Here are samples ofHistory STAD: \[\[LABEL\]\]AncientCivilizations \[\[SUMMARY\]\]Thisclustercentersonantiquity,archaeologists,andantiquities,withastrongfocusonarchaeologyandegyptology\.ThedocumentsexploreancientcivilizationssuchasEgypt,Mesopotamia,andtheGreco\-Romanworld,examiningarchaeologicalfindings,tombs,inscriptions,andartifacts\.Keythemesincludethestudyofancientculturesthroughexcavatedmaterials,historicalinterpretationsbyrenownedscholars,andinsightsintoreligiouspractices,monuments,anddailylifeinantiquity\. DaLLME: featurenode\.BookDescription:“Thisbookisamust\-readforanyoneinterestedinthehistoryoftheancientworld\.”—JournalofAmericanHistory“—Journ ### E\.4\.Details of In\-Context Learning based onSTAD TheSTADbased ICL prompt generation takes the following steps\. First, it loads test node attributes, synthetic attribute and labels fromSTAD\. Unique categories of the current dataset are extracted from the loaded labels, forming the actual category list to fill the placeholder in the predefined templates\. Then we compute Euclidean distances between each test node’s attribute vector and the synthetic attribute vectors\. For each category, the algorithm further conducts statistical analysis on key distance metrics, including the average distance between the test node and all synthetic nodes in the category, the closest distance to any synthetic node in the category, and the index of the closest synthetic node\.STADbased ICL is effective because it naturally fuses rich semantic priors with structural and feature similarity from the distilled dataset, which preserves the core information of the original graph\. By incorporating category\-aware distance guidance and semantic context into prompts, ICL enables the model to directly reason over compact yet representative knowledge without fine\-tuning, leading to strong generalization and accurate node classification\. We here give the prompts of two training\-free LLMs node classification Direct andneighbor summary\(NS\) , with and without the enhance ofin\-context learning\(ICL\) based on ourSTAD\. We takeCoraas an samples\. System Prompt Template\. 'Youareanaccuratetextgraphnodeclassifier\.Yourtaskistocategorizetext\-describednodesintospecifiedcategories\.Hereisthescenario:\\n' Usr Prompt Template\.For Direct and Direct\+STAD, the user prompts are as follow, Direct: """ \{inputtext\} Question:Whichofthefollowingsub\-categoriesofAIdoesthispaperbelongto?Herearethe7categories:Rule\_Learning,Neural\_Networks,Case\_Based,Genetic\_Algorithms,Theory,Reinforcement\_Learning,Probabilistic\_Methods\.Replyonlyonecategorythatyouthinkthispapermightbelongto\.Onlyreplythecategoryphrasewithoutanyotherexplanationwords\. Answer: """ Diectly\+STAD: """ \{inputtext\} Question:Whichofthefollowingsub\-categoriesofAIdoesthispaperbelongto?Herearethe7categories:Rule\_Learning,Neural\_Networks,Case\_Based,Genetic\_Algorithms,Theory,Reinforcement\_Learning,Probabilistic\_Methods\.Replyonlyonecategorythatyouthinkthispapermightbelongto\.Onlyreplythecategoryphrasewithoutanyotherexplanationwords\. Answer: Note:Thecloserthedistance,thehighertheprobabilityofbelongingtothisclass\. Contextinformationbycategory: \-Rule\_Learning:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Neural\_Networks:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Case\_Based:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Genetic\_Algorithms:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Theory:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Reinforcement\_Learning:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Probabilistic\_Methods:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} Answer: """ We show the pipeline of NS\. Firstly, we give the prompt to for neighbor context summarization, NeighborContextSummarization: """ Thefollowinglistrecordssomepaperrelatedtothecurrentone,withrelationshipbeingcitation\. Neighbor'stextinformation:\{neighbor\_1\_text\} Neighbor'stextinformation:\{neighbor\_2\_text\} Neighbor'stextinformation:\{neighbor\_3\_text\} \.\.\. Pleasesummarizetheinformationabovewithashortparagraph,findsomecommonpointswhichcanreflectthecategoryofthispaper\. """ Then the query for node classification based on the neighbor context is NS: """ \{inputtext\} \{neighborcontextsummary\} HereIgiveyouthecontentofthenodeitselfandthesummaryinformationofits1st\-orderneighbors\. Therelationbetweenthenodeanditsneighborsis'citation'\.Question:Basedontheseinforamtion, Whichofthefollowingsub\-categoriesofAIdoesthispaper\(thisnode\)belongto?Herearethe7categories:Rule\_Learning,Neural\_Networks,Case\_Based,Genetic\_Algorithms,Theory,Reinforcement\_Learning,Probabilistic\_Methods\. Replyonlyonecategorythatyouthinkthispapermightbelongto\.Onlyreplythecategorynamewithoutanyotherwords\.\\n\\nAnswer: """ Similar to Direct \+STAD, the query for NS \+STADis NS\+STAD: """ \{inputtext\} \{neighborcontextsummary\} HereIgiveyouthecontentofthenodeitselfandthesummaryinformationofits1st\-orderneighbors\. Therelationbetweenthenodeanditsneighborsis'citation'\.Question:Basedontheseinforamtion, Whichofthefollowingsub\-categoriesofAIdoesthispaper\(thisnode\)belongto?Herearethe7categories:Rule\_Learning,Neural\_Networks,Case\_Based,Genetic\_Algorithms,Theory,Reinforcement\_Learning,Probabilistic\_Methods\. Replyonlyonecategorythatyouthinkthispapermightbelongto\.Onlyreplythecategorynamewithoutanyotherwords\. Answer: Note:Thecloserthedistance,thehighertheprobabilityofbelongingtothisclass\. Contextinformationbycategory: \-Rule\_Learning:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Neural\_Networks:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Case\_Based:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Genetic\_Algorithms:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Theory:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Reinforcement\_Learning:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} \-Probabilistic\_Methods:Averagedistance=\{\},Closestdistance=\{\},Closestsampleinfo:\{\} Answer: """
Similar Articles
TAG-DLM: Diffusion Language Models for Text-Attributed Graph Learning
TAG-DLM unifies textual reasoning and graph message passing within a masked diffusion language model, enabling joint reasoning over text and graph topology for node classification and link prediction tasks.
TAKE: Trajectory-Aware Knowledge Estimation for Text Dataset Distillation
This paper introduces TAKE (Trajectory-Aware Knowledge Estimation), a text dataset distillation framework that uses influence functions and optimal transport to reduce datasets to as little as 0.1% of their original size while preserving downstream task fidelity.
TERGAD: Structure-Aware Text-Enhanced Representations for Graph Anomaly Detection
TERGAD is a novel data augmentation framework that uses large language models to translate node-level topological properties into semantic narratives, then fuses these with original node attributes via a gated dual-branch autoencoder for graph anomaly detection, achieving state-of-the-art results on six datasets.
Beyond the Golden Teacher: Enhancing Graph Learning through LLM-GNN Co-teaching
This paper proposes LLM-GNN Co-Teaching, a bidirectional framework for few-shot graph learning on text-attributed graphs. The LLM and GNN exchange confident pseudo-labels and use round-based preference optimization (RPL-PO) to mutually improve, outperforming prior methods on benchmarks.
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
Researchers extend MeanFlow one-step image generation from class labels to flexible text inputs by integrating highly-discriminative LLM-based text encoders, enabling efficient text-conditioned synthesis with improved performance.