An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation

arXiv cs.CL 论文

摘要

This paper proposes a hybrid knowledge graph generation pipeline that combines top-down grounding in Wikidata with bottom-up agentic synthesis to handle noisy, multilingual HR skill declarations, producing a scalable and self-healing skills taxonomy.

arXiv:2608.07023v1 Announce Type: new Abstract: Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, directly impacting downstream tasks like accurate talent matching. To address this, we propose a hybrid knowledge graph generation pipeline that grounds a Large Language Model (LLM) in the Wikidata multilingual Knowledge Graph (KG) while employing an agentic reflexion pattern to synthesize emerging concepts and their associated metadata. Unlike rigid top-down methods or fragmented bottom-up approaches, our system anchors recognized concepts to stable Knowledge Graph entities while dynamically creating new nodes and relational metadata for unrecognized skills. Executed across five stages, entity reconciliation, multilingual canonicalization, active curation, deduplication, and the iterative recovery of unmapped concepts, the system autonomously adapts to rapidly evolving, noisy skill mentions across five European languages. Ultimately, this pipeline provides a highly scalable, explicable, and self-healing framework for generating a comprehensive skills knowledge graph, from which a structured taxonomy is derived, using unstructured, noisy text.
查看原文
查看缓存全文

缓存时间: 2026/08/10 08:05

# An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation
Source: [https://arxiv.org/html/2608.07023](https://arxiv.org/html/2608.07023)
11institutetext:Malt, Paris, France###### Abstract

Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources \(HR\) platforms, directly impacting downstream tasks like accurate talent matching\. To address this, we propose a hybrid knowledge graph generation pipeline that grounds a Large Language Model \(LLM\) in the Wikidata multilingual Knowledge Graph \(KG\) while employing an agentic reflexion pattern to synthesize emerging concepts and their associated metadata\. Unlike rigid top\-down methods or fragmented bottom\-up approaches, our system anchors recognized concepts to stable Knowledge Graph entities while dynamically creating new nodes and relational metadata for unrecognized skills\. Executed across five stages, entity reconciliation, multilingual canonicalization, active curation, deduplication, and the iterative recovery of unmapped concepts, the system autonomously adapts to rapidly evolving, noisy skill mentions across five European languages\. Ultimately, this pipeline provides a highly scalable, explicable, and self\-healing framework for generating a comprehensive skills knowledge graph, from which a structured taxonomy is derived, using unstructured, noisy text\.

## 1Introduction

Comprehensive, structured skills knowledge graphs and their derivative taxonomies serve as a critical foundation for HR applications, such as standardizing job requirements and facilitating talent matching\. However, maintaining these structures is a persistent challenge for talent marketplaces like Malt, the European freelancer marketplace\. Free\-text expertise declarations frequently lead to multilingual fragmentation, niche jargon, compound skills, and rapid temporal drift as occupations and work practices evolve\. Although Large Language Models \(LLMs\) overcome the scalability issues of hand\-crafted taxonomies, limitations remain, such as a susceptibility to hallucinated skills, formatting instability, and a lack of grounding in verifiable metadata\[[13](https://arxiv.org/html/2608.07023#bib.bib1),[22](https://arxiv.org/html/2608.07023#bib.bib12)\]\.

As shown in Figure[1](https://arxiv.org/html/2608.07023#S1.F1), traditional generation methods generally fall into two complementary yet individually limited paradigms\. Rigid top\-down methods rely on predefined ontologies such as "European Skills, Competences, Qualifications and Occupations" \(ESCO\)\[[3](https://arxiv.org/html/2608.07023#bib.bib5)\], which provide structural consistency and high precision but often struggle to capture emerging or highly localized concepts\. Conversely, purely generative bottom\-up approaches are flexible enough to identify novel trends, but they frequently lack structural constraints\. As a result, semantically equivalent concepts may fragment into redundant clusters, while noisy relationships can emerge between otherwise unrelated concepts\.

Core OntologyProjectManagementProjectManagementWeb ProjectManagementMissing Node\(A\) Rigid Top\-DownPM ClusterWeb PMGroupProjectManagementWeb PMExcelFragmentedHallucination\(B\) Chaotic Bottom\-UpTOP\-DOWNGROUNDINGBOTTOM\-UPSYNTHESISWikidata Entity:Q179012\+ Rich MetadataSynthesized Entity:Web Project Mgmt\+ Derived MetadataProjectManagementWeb PMWeb ProjectManagementGroundingAgentic Reflection\(C\) Proposed Hybrid

Figure 1:Comparison of structural modeling paradigms via a project management narrative\. \(A\) Static Top\-Down: Misses emerging specializations due to static ontological boundaries\. \(B\) Unconstrained Bottom\-Up: Lacks guardrails, causing semantic fragmentation \(separating “Project Mgmt” and “Web PM”\) and link hallucinations \(“Excel”\)\. \(C\) Proposed Hybrid: Achieves complete convergence\. The baseline concept is anchored to a Wikidata entity node, while agentic reflection synthesizes a distinct sub\-entity linked with rich relational metadata that cleanly absorbs related shortcuts \(“Web PM”\)\.To bridge this gap, we introduce a fully automated, agentic, hybrid pipeline\. Our architecture grounds an LLM in an interchangeable Knowledge Graph, instantiated here with Wikidata\[[21](https://arxiv.org/html/2608.07023#bib.bib13)\], to ensure that the generated entities and their associated metadata remain verifiable and auditable\. At the same time, we leverage an iterative multi\-agent loop to capture, validate, and structure the emerging "long\-tail" skills that formal ontologies have not yet registered\. Ultimately, this approach yields a rich knowledge graph from which a streamlined skills taxonomy is derived as a structured child entity\. Specifically, this paper makes the following contributions:

Hybrid Skill Extraction\.We combine structured knowledge graph integration with dynamic, bottom\-up parsing to capture both established global standards and emerging long\-tail expertise, enriching them with contextual metadata\.

Multilingual Standardization\.We enforce structural schema constraints on the underlying model to support consistent multilingual representations across five target languages and reduce hallucinated outputs\.

Iterative Self\-Refinement\.We introduce a reflection\-based multi\-agent loop in which agents iteratively analyze and correct intermediate outputs\. This enables progressive consolidation of fragmented concepts and improves overall taxonomic coherence\.

To contextualize these contributions, the remainder of this paper reviews the evolution of automated taxonomy generation, details our pipeline’s architecture, and evaluates its real\-world performance on our platform data\.

## 2Related Work

Taxonomy creation has shifted from costly expert ontologies and static knowledge bases\[[7](https://arxiv.org/html/2608.07023#bib.bib7)\]to automated machine learning methods\. Traditional skill extraction models\[[4](https://arxiv.org/html/2608.07023#bib.bib3),[27](https://arxiv.org/html/2608.07023#bib.bib14)\], and even some recent efficient encoders\[[5](https://arxiv.org/html/2608.07023#bib.bib4)\], primarily produced flat lists susceptible to semantic ambiguity\.To establish the relational hierarchies necessary for a true knowledge graph, top\-down approaches anchor semantics using Knowledge Graphs, utilizing domain seeds\[[3](https://arxiv.org/html/2608.07023#bib.bib5)\], automated tree expansion\[[19](https://arxiv.org/html/2608.07023#bib.bib11),[26](https://arxiv.org/html/2608.07023#bib.bib15),[12](https://arxiv.org/html/2608.07023#bib.bib16)\], and LLM\-driven ranking and iterative prompting\[[15](https://arxiv.org/html/2608.07023#bib.bib17),[25](https://arxiv.org/html/2608.07023#bib.bib18)\]\. However, these static methods struggle to adapt to the dynamic vocabularies of modern job markets\.

Conversely, bottom\-up clustering scales effectively\[[1](https://arxiv.org/html/2608.07023#bib.bib22),[9](https://arxiv.org/html/2608.07023#bib.bib2)\]but might lead to uninterpretable labels and missing relational metadata\[[2](https://arxiv.org/html/2608.07023#bib.bib23)\]\. LLM based models overcome this by autonomously structuring categories via abstractive prompting\[[16](https://arxiv.org/html/2608.07023#bib.bib24),[23](https://arxiv.org/html/2608.07023#bib.bib25)\], localized induction\[[8](https://arxiv.org/html/2608.07023#bib.bib8),[6](https://arxiv.org/html/2608.07023#bib.bib6)\], and specialized schemas\[[18](https://arxiv.org/html/2608.07023#bib.bib19)\]\. While end\-to\-end methods automate generation\[[22](https://arxiv.org/html/2608.07023#bib.bib12)\], their batch processing might lead to fragmented hierarchies rather than cohesive graphs\. Recent multi\-agent frameworks resolve this through reflection\[[13](https://arxiv.org/html/2608.07023#bib.bib1),[20](https://arxiv.org/html/2608.07023#bib.bib10)\]and dynamic alignment\[[10](https://arxiv.org/html/2608.07023#bib.bib9)\], leveraging complex reasoning heuristics like self\-correction\[[14](https://arxiv.org/html/2608.07023#bib.bib20)\], prompt optimization\[[17](https://arxiv.org/html/2608.07023#bib.bib26)\], and tree search\[[24](https://arxiv.org/html/2608.07023#bib.bib21)\]\. Additionally, LLMs can distill domain knowledge into lightweight downstream classifiers\[[11](https://arxiv.org/html/2608.07023#bib.bib27)\]\.

Despite these advances, unconstrained LLMs remain susceptible to hallucinated skills, formatting instability, and bias\. We address this via a hybrid architecture: while an LLM drives the pipeline, its reasoning is strictly anchored to a deterministic KG for recognized entities, reserving unconstrained generative reflection solely for unmapped skills and their corresponding relational metadata\.

## 3Proposed Approach

Raw Inputs"gestion de projet web"1\. Semantic Recon\.Contextually maps string to foundational anchor nodeQ179012\.2\. CanonicalizationGroups inputs into a canonical entity node with localized metadata\.3\. Active CurationFlags input as aSPECIALIZATIONoutlier and extracts structural attributes\.Stable Core KG NodesHolds grounded high\-precision entities and relational metadata\.Wikidata Knowledge GraphOntological Anchors &Multilingual MetadataPlatform Usage DataEmpirical Co\-occurrences &Category Distributions4\. ConsolidationGenerates a unique, stableORPHAN\_IDto instantiate sub\-graph branches\.Valid QIDQ179012Base Entity Node“Project Management”Passed CoreRejected Outlier Entity:"Web Project Management"5\. Iterative Epochs \(Self\-Healing Loop\)Re\-injects synthetic identifier into next generation cycle to group nested specializations under sub\-graph\.

Figure 2:The 5\-stage hybrid pipeline architecture, traced via the project management example\. The system grounds the model in Wikidata \(Stage 1\) and uses agentic reflection \(Stage 3\) to capture outliers\. These are converted into stable synthetic identifiers to instantiate new sub\-graph entities \(Stage 4\) and rerouted \(Stage 5\) for continuous, autonomous self\-healing of the knowledge graph structure\.We designed a hybrid, multi\-agent pipeline to overcome the limitations of rigid top\-down ontologies and ungrounded bottom\-up generation\. Unlike top\-down methods that map unstructured inputs to a static seed ontology, our system dynamically constructs its own custom, evolving knowledge graph derived directly from empirical data\. Conversely, unlike pure bottom\-up methods that rely on unconstrained generative clustering, we enforce strict semantic grounding by anchoring these emerging clusters to factual Knowledge Graph entities and capturing their relational metadata\. As illustrated in Figure[2](https://arxiv.org/html/2608.07023#S3.F2), the system operates as an iterative loop rather than a strictly linear execution pipeline, processing batches through five distinct stages\. By utilizing Wikidata as a semantic anchor, chosen for its broad multilingual domain coverage enabled by its massive open\-source scale, and constraining Gemini 1\.5 Flash with strict schemas, we prevent hallucinations while preserving linguistic variants and contextual structure\. We specifically selected this lightweight model for its cost\-efficiency and immediate availability within our infrastructure\. Because our pipeline relies on strict external grounding rather than complex internal reasoning, these results could likely be replicated using comparable open\-source small language models\. Crucially, this architecture is not tied to a single semantic structure and could easily integrate alternative or combined external databases\. The following subsections detail the mechanics of each phase, demonstrating how the pipeline enables continuous, autonomous updates as novel expertises and their structural relationships emerge in the market\.

### 3\.1Reconciliation: Grounding Raw Skills in Wikidata

Trace Example — Stage 1 \(Reconciliation\):In:Raw Text \("gestion de projet web"\) \+ Malt Context\.Out:Linked Anchor \(Q179012\) \+ Validation Flags \(is\_skill: true,is\_compound: false\)

As the first step in this pipeline, the Reconciliation phase maps noisy, unstandardized, and multilingual skill mentions to stable, unambiguous Wikidata entity identifiers \(QIDs\), which serve as the foundational anchor nodes for our knowledge graph\. This process begins by extracting the top ten Wikidata entries linked to the raw input text, alongside metadata such as labels, descriptions, and regional variants across the five target languages\. Because raw user mentions are frequently ambiguous or completely devoid of context when analyzed in isolation, this baseline retrieval must be structurally enriched with domain\-specific semantic neighborhoods and high\-level category distributions\.

At Malt, this enrichment queries the platform’s user profile graph\. Upon skill ingestion, the engine extracts and appends two empirical features from the freelancer’s history: the top fifteen \(an empirical threshold selected to maximize semantic signal while filtering out long\-tail profile noise\) co\-occurring peer skills and the top five overarching professional categories\.

From this enriched context, the LLM executes a disambiguation step to select the best corresponding QID\. To ensure structured outputs, the LLM must output a JSON payload containing the selected QIDs, confidence scores, step\-by\-step reasoning, and critical boolean flags indicating whether the input is a valid professional skill \(is\_skill\) and if it contains multiple skills \(is\_compound\)\.

This dual validation approach, combining Wikidata’s semantic recall with the LLM’s context\-aware precision, reduces potential hallucinated mappings\. Ultimately, by evaluating candidate labels in all target languages simultaneously, the system ensures that non\-English skills are reliably anchored to the exact same global QID as their English counterparts\.

### 3\.2Canonicalization: Clustering and Preferred Label Generation

Trace Example — Stage 2 \(Canonicalization\):In:Resolved Entity Cluster \(Q179012\) \+ Historical Usage Logs\.Out:Multilingual Label Mapping \(EN:"Project Management"\[wikidata\],FR:"Gestion de Projet"\[malt\]\)

Following reconciliation, the second phase groups validated inputs by their unique QID combinations to generate human\-readable, multilingual canonical names\. Provided with context such as Wikidata descriptions and empirical usage distributions within the freelance platform, the LLM synthesizes localized preferred labels across five target languages\. To mitigate clustering noise, the prompt enforces a strict fallback hierarchy: the model must prioritize empirical platform usage, fall back to official ontological titles, and synthesize novel labels only when necessary\. To maintain algorithmic explicability, each label is explicitly tagged with its resulting provenance \(malt,wikidata, orgenerative\)\.

### 3\.3Curation: Semantic Validation and Reflexion

Trace Example — Stage 3 \(Curation\):In:Combined Cluster Member \("gestion de projet web"\) vs\. Core Candidate \("Project Management"\)\.Out:Status:REJECTED\(Criterion:SPECIALIZATION\)→\\rightarrowOutput Target:suggested\_pref\_label: "Web Project Management"

To prevent disjointed or overly granular skills within canonicalized nodes, an LLM\-powered Curation agent validates the strict equivalence of every raw skill against its broader entity grouping and selected preferred label\. If a skill is non\-equivalent, the model outputs a structured output detailing one of seven granular rejection criteria \(AMBIGUOUS,SPECIALIZATION,SEMANTIC\_MISMATCH,NOT\_A\_SKILL,METHODOLOGY,CONTEXT,SUB\_TASK\)\. For pipeline tracking and evaluation purposes, these semantic rejections \(aside fromNOT\_A\_SKILL\) are aggregated under the broaderNODE\_CURATIONstatus\. Crucially, asuggested\_pref\_labelis generated to define what the rejected concept should actually be named\. This contextual rejection reason is retained as metadata for subsequent consolidation tasks\.

Following outlier extraction, the agent re\-evaluates and refines the core node’s canonical label using strictly the accepted subset\. This bifurcates the data into tightly curated baseline entities and a structured "Orphan" queue for iteration, which will be processed later as explained in section[3\.5](https://arxiv.org/html/2608.07023#S3.SS5)\. As a safeguard, any node with a rejection rate exceeding 50% is automatically flagged for human\-in\-the\-loop review to prevent cascading systemic errors\.

### 3\.4Consolidation: Cross\-Batch Deduplication

Trace Example — Stage 4 \(Consolidation\):In:Cross\-Batch Incoming Candidate \("Web PM"\) vs\. Active Orphan Taxonomy Target \("Web Project Management"\)\.Out:Decision:MERGE→\\rightarrowSurviving Entity Destination ID:hash\("Web Project Management"\)

To maintain global structural consistency across incremental runs, the Consolidation phase merges overlapping sub\-graphs\. To avoid the computational explosion of exhaustive pairwise comparisons, we employ an asymmetric bootstrapping strategy\. The system initializes from a blank state, where the first processed batch establishes the foundational knowledge graph\. In all subsequent runs, newly generated nodes are strictly compared against this continuously growing, established baseline\. Before LLM evaluation, lightweight heuristics flag potential merges between incoming and established entities based onmember intersection\(shared raw skills\) orlexical similarity\(low edit distance between preferred labels\)\. Once flagged, the LLM evaluates the combined metadata of these pairs to output a structural decision \(MERGEorKEEP\_SEPARATE\), designating a surviving ID if merged\. Logging these decisions creates a cache that prevents redundant re\-evaluations in future Epochs\.

### 3\.5Iteration: Orphan Recovery and Convergence

Trace Example — Stage 5 \(Iteration\):In:Verified Isolated Structural Orphan Queue \+ Validated Label Target \.Out:Sub\-graph Generation: Instantiates stable sub\-branch linked toQ179012via deterministic cryptographic label routing\.

Step 1: Canonical Entity NodeInput NodeQ192253\(“Project Management”\)Step 2:LLM ValidationIs skill equivalent?Step 3 \(Yes\): Core KG NodeAccepts stable generic node:“Project Management”Step 3 \(No\): Rejection AnalysisFlags “gestion de projet web” as aSPECIALIZATIONStep 4: Label GenerationSynthesizessuggested\_pref\_label:“Web Project Management”Step 5: Orphan Node CreationUses suggested label to define a unique, stableORPHAN\_IDAgentic Reflection PhaseStep 3 \(No\): Rejection AnalysisFlags “gestion de projet web” as aSPECIALIZATIONStep 4: Label GenerationSynthesizessuggested\_pref\_label:“Web Project Management”Step 5: Orphan Node CreationUses suggested label to define a unique, stableORPHAN\_IDYesNo \(Reject\)Step 6: Route to EpochN\+1N\+1Forced grouping of nestedspecialization into sub\-graph

Figure 3:The Agentic Reflection and Orphan Lifecycle\. Rejected skills trigger an active reflection loop \(Steps 3–5\) where the LLM justifies the outlier status and synthesizes asuggested\_pref\_label\. This generates a stable syntheticORPHAN\_IDrouted into the next Epoch for autonomous self\-healing of the knowledge graph layout\.This iterative loop primarily serves to fix granularity gaps inherited from the baseline anchor graph\. For example, both "Project Management" and "Web Project Management" might initially map to the same broad Wikidata entity\. To preserve graph coherence, the curation phase flags "Web Project Management" as an outlier while keeping the main node broad\. Instead of discarding this niche specialization, the pipeline captures it as an "orphan" to generate the fine\-grained relational edges and specialized nodes that Wikidata lacks natively\.

As illustrated in Figure[3](https://arxiv.org/html/2608.07023#S3.F3), these orphans are re\-processed using a recursive routing loop\. Each orphan is assigned a synthetic identifier derived directly from its LLM\-generatedsuggested\_pref\_label\. Because identical concepts receive the exact same suggested label from the model, this mechanism naturally groups separate but matching specializations together into a unified sub\-graph in the subsequent Epoch\. This loop repeats until the orphan queue is empty, achieving full semantic convergence\.

## 4Evaluation

We evaluated our pipeline using a proprietary dataset of unstructured, multilingual expertise declarations from the Malt freelancing platform\. This dataset reflects chaotic, real\-world labor market dynamics—featuring a severe long\-tail distribution of highly niche, emerging, or misspelled jargon—providing a robust stress test compared to static theoretical ontologies\. Because HR matching engines require strict reliability, model performance is evaluated against a curated Wikidata gold standard\. As this research represents ongoing work, our current evaluation focuses strictly on the initial retrieval and pre\-consolidation phases; the final post\-consolidation phase was recently introduced to the pipeline and has yet to be formally benchmarked\.

### 4\.1Vocabulary Coverage and Semantic Compression

The normalization pipeline processed an initial vocabulary of 36,037 raw expertise strings\. The reconciliation engine successfully resolved 27,743 of these inputs, achieving a Global Coverage rate of 77% \(the percentage of valid inputs successfully mapped to a knowledge graph node\)\. The remaining 8,294 unmapped entries were flagged by the curation layer as non\-skills or semantic noise \(see Table[2](https://arxiv.org/html/2608.07023#S4.T2)\)\.The 27,743 mapped variations were grouped into 15,010 semantic groupings, which were further streamlined into 13,298 canonical skill nodes\. This represents a compression rate of 52\.1% \(1−\[13,298/27,743\]1\-\[13,298/27,743\]\), significantly reducing downstream redundancy\. Alongside, the Average Skills per Node \(A​S​p​NASpN\) metric, which stands at 2\.08 variations per canonical node, further demonstrates our vocabulary consolidation efficiency\. Importantly, the knowledge graph maintains perfect cross\-lingual symmetry: 100% of the 13,298 canonical concept nodes are fully supported across all five target locales \(fr, en, de, nl, es\)\. This generates exactly 66,490 standardized preferred labels \(13,298×513,298\\times 5\), ensuring uniform matching regardless of the user’s interface language\. Empirical platform data reveals a strong Pareto distribution: the top 1,000 canonical skills account for 82\.74% of platform usage volume, and the top 5,000 capture 97\.25%\. Interpreting this requires distinguishing between the lexical long\-tail \(typographical noise and redundant expression variants\) and the semantic long\-tail \(rare, highly specialized, or emerging capabilities\)\. While our pipeline aggressively filters and compresses the unmanaged lexical long\-tail to eliminate marketplace redundancy, it systematically preserves the semantic long\-tail through agentic reflection, yielding a highly comprehensive and nuanced final knowledge graph of 13,298 standardized concept nodes\.

### 4\.2Quantitative Precision against Gold Standard

Evaluated against a hand\-annotated gold standard, independently curated by five domain experts without overlap, the pre\-consolidation pipeline achieved a global baseline Alignment Coverage of 79\.7% \(the overall proportion of inputs successfully and correctly mapped to the gold standard\) and a Found Coverage of 84\.9% \(measuring precision strictly on the subset of inputs where the model actually attempted a retrieval\)\. Due to the semantic ambiguity of unmanaged inputs, the global Wrong Guess Rate \(WGR\) was 19\.1%\. Given the chaotic nature of user\-generated profile text, this error rate remains highly competitive and is mitigated by downstream curation\. As shown in Table[2](https://arxiv.org/html/2608.07023#S4.T2), performance varies by domain\. Highly structured domains like Video Games \(91\.8% Found Coverage\) outperformed softer, more subjective fields like Communication \(81% Found Coverage\), where shifting jargon and conceptual overlap complicate alignment\.

### 4\.3Provenance and Structural Purity

To ensure explicability, we track provenance metadata across all 66,490 preferred labels\. Accounting for source overlap, 80\.65% of the knowledge graph is strictly anchored in factual, real\-world data \(comprising 67\.28% empirical platform usage and 22\.08% Wikidata titles\)\. The generative engine synthesized only the remaining 19\.35% to resolve emerging concepts and isolated language gaps, proving the structure is rooted in empirical reality rather than ungrounded model hallucinations\.

Finally, we measured structural coherence using the Outlier Rate \(skills manually removed during human review vs\. total generated sub\-graphs\)\. Qualitative audits revealed highly cohesive node groupings, with marginal error rates \(0\.01 to 0\.06 outliers per sub\-graph\)\. This purity stems from the Active Curation agent, which aggressively isolates semantic noise before ingestion, successfully blocking 13\.7% of rejected inputs via the NOT\_A\_SKILL filter \(Table[2](https://arxiv.org/html/2608.07023#S4.T2)\)\.

Table 1:Baseline Reconciliation Coverage
Table 2:Active Curation Rejections
Note:NODE\_CURATIONaggregates Section 3\.3 criteria\.

## 5Discussion and Future Work

While this pipeline provides a structured representation of the underlying data, it represents an initial step toward modeling the complexity of modern labor markets\. Such structured representations are a prerequisite for building robust HR analytics systems\. Looking ahead, a key objective is to identify emerging skill signals in real\-time, enabling downstream analysis of labor market dynamics and supporting workforce planning applications\. A streamlined skill taxonomy, extracted directly from this underlying knowledge graph, is already integrated into the platform’s candidate\-matching algorithms\. A primary benefit of this deployment is that the graph provides an intermediate abstraction layer that significantly improves the interpretability and auditability of our automated matching systems\. While relying on Wikidata as a primary anchor introduces limitations, such as a lag in capturing niche HR jargon and structural inconsistencies due to its generalist nature, our hybrid approach mitigates this\. The bottom\-up orphan recovery loop acts as a safety net, autonomously structuring the emerging long\-tail skills that Wikidata natively misses\. As the system scales, it is important to explicitly address potential representational biases\. Large language and embedding\-based models often exhibit English\-centric tendencies, which can lead to over\-normalization of non\-English occupational structures and a reduced fidelity of locale\-specific distinctions\. Our goal is to mitigate these effects by preserving meaningful cross\-lingual variation in occupational and skill representations\. Addressing gender\-related bias is also critical\. In highly inflected languages such as French and German, occupational and skill terms are often gender\-marked\. We aim for the underlying knowledge graph to support the seamless mapping of these gendered variants to shared underlying occupational concepts, while preserving their distinct linguistic forms for accurate representation and equitable matching\. Finally, to evaluate the robustness of the system at scale, our evaluation roadmap will focus on three key areas:

Extended Evaluations\.While parts of our pipeline, particularly the consolidation phase, are already implemented, they require further evaluation\. We will benchmark the consolidated graph structure against existing standards \(like ESCO\) and other AI approaches \(like TnT\-LLM\[[22](https://arxiv.org/html/2608.07023#bib.bib12)\]or CLIMB\[[13](https://arxiv.org/html/2608.07023#bib.bib1)\]\), while assessing the system’s ability to incorporate new skills and the effectiveness of the human\-in\-the\-loop components\.

Cost and Scalability\.Large\-scale deployment of LLM\-based pipelines introduces significant computational costs\. We will measure token consumption, processing latency, and optimization strategies required to ensure operational scalability\.

Failure Analysis\.To support responsible deployment, we will analyze failure cases by tracking manual intervention rates, categorizing systematic mapping errors, and evaluating performance degradation under high\-volume concept comparison scenarios\.

## 6Conclusion

This ongoing work introduces a hybrid, multi\-agent architecture designed to build skills knowledge graphs that are both rigorously structured and highly adaptable\. By grounding a large language model in Wikidata, our pipeline effectively parses multilingual free\-text while keeping the underlying entity data reliable and easy to audit\. Through continuous curation and automated reflection, the system successfully captures the niche and emerging long\-tail instances that traditional, static models often miss\. Ultimately, this approach creates a living knowledge graph capable of keeping pace with the rapid technological changes and linguistic shifts of the modern freelance market\. While the overarching framework operates as a rich, metadata\-driven knowledge graph, its hierarchical output can be easily downstreamed as a clean, structured skills taxonomy\.

## References

- \[1\]C\. C\. Aggarwal and C\. Zhai\(2012\)A survey of text clustering algorithms\.InMining text data,pp\. 77–128\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[2\]J\. Chang, S\. Gerrish, C\. Wang, J\. Boyd\-Graber, and D\. Blei\(2009\)Reading tea leaves: how humans interpret topic models\.Advances in neural information processing systems22\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[3\]J\. De Smedt, M\. le Vrang, and A\. Papantoniou\(2015\)ESCO: towards a semantic web for the european labor market\.\.Ldow@ www144\.Cited by:[§1](https://arxiv.org/html/2608.07023#S1.p2.1),[§2](https://arxiv.org/html/2608.07023#S2.p1.1)\.
- \[4\]J\. Decorte, J\. Van Hautte, T\. Demeester, and C\. Develder\(2021\)Jobbert: understanding job titles through skills\.arXiv preprint arXiv:2109\.09605\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p1.1)\.
- \[5\]J\. Decorte, J\. Van Hautte, C\. Develder, and T\. Demeester\(2025\)Efficient text encoders for labor market analysis\.IEEE Access\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p1.1)\.
- \[6\]M\. Gao, J\. Shah, W\. Wang, K\. Huang, and D\. Khashabi\(2025\)Science hierarchography: hierarchical organization of science literature\.arXiv preprint arXiv:2504\.13834\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[7\]T\. R\. Gruber\(1993\)A translation approach to portable ontology specifications\.Knowledge acquisition5\(2\),pp\. 199–220\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p1.1)\.
- \[8\]M\. Gunn, D\. Park, and N\. Kamath\(2024\)Creating a fine grained entity type taxonomy using llms\.arXiv preprint arXiv:2402\.12557\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[9\]A\. K\. Jain\(2010\)Data clustering: 50 years beyond k\-means\.Pattern recognition letters31\(8\),pp\. 651–666\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[10\]P\. Kargupta, N\. Zhang, Y\. Zhang, R\. Zhang, P\. Mitra, and J\. Han\(2025\)Taxoadapt: aligning llm\-based multidimensional taxonomy construction to evolving research corpora\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 29834–29850\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[11\]D\. Lee, J\. Pujara, M\. Sewak, R\. White, and S\. Jauhar\(2023\)Making large language models better data creators\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 15349–15360\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[12\]D\. Lee, J\. Shen, S\. Kang, S\. Yoon, J\. Han, and H\. Yu\(2022\)Taxocom: topic taxonomy completion with hierarchical discovery of novel topic clusters\.InProceedings of the ACM Web Conference 2022,pp\. 2819–2829\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p1.1)\.
- \[13\]N\. Li, B\. Kang, and T\. De Bie\(2025\)Building data\-driven occupation taxonomies: a bottom\-up multi\-stage approach via semantic clustering and multi\-agent collaboration\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track,pp\. 1596–1614\.Cited by:[§1](https://arxiv.org/html/2608.07023#S1.p1.1),[§2](https://arxiv.org/html/2608.07023#S2.p2.1),[§5](https://arxiv.org/html/2608.07023#S5.p2.1)\.
- \[14\]A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang,et al\.\(2023\)Self\-refine: iterative refinement with self\-feedback\.Advances in neural information processing systems36,pp\. 46534–46594\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[15\]O\. Marchenko and D\. Dvoichenkov\(2024\)TaxoRankConstruct: a novel rank\-based iterative approach to taxonomy construction with large language models\.\.InISS@ IT&I,pp\. 11–27\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p1.1)\.
- \[16\]C\. M\. Pham, A\. Hoyle, S\. Sun, P\. Resnik, and M\. Iyyer\(2024\)TopicGPT: a prompt\-based topic modeling framework\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 2956–2984\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[17\]R\. Pryzant, D\. Iter, J\. Li, Y\. Lee, C\. Zhu, and M\. Zeng\(2023\)Automatic prompt optimization with “gradient descent” and beam search\.InProceedings of the 2023 conference on empirical methods in natural language processing,pp\. 7957–7968\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[18\]C\. Sas and A\. Capiluppi\(2024\)Automatic bottom\-up taxonomy construction: a software application domain study\.arXiv preprint arXiv:2409\.15881\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[19\]J\. Shen, Z\. Wu, D\. Lei, C\. Zhang, X\. Ren, M\. T\. Vanni, B\. M\. Sadler, and J\. Han\(2018\)Hiexpan: task\-guided taxonomy construction by hierarchical tree expansion\.InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining,pp\. 2180–2189\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p1.1)\.
- \[20\]N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. Yao\(2023\)Reflexion: language agents with verbal reinforcement learning\.Advances in neural information processing systems36,pp\. 8634–8652\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[21\]D\. Vrandečić and M\. Krötzsch\(2014\)Wikidata: a free collaborative knowledgebase\.Communications of the ACM57\(10\),pp\. 78–85\.Cited by:[§1](https://arxiv.org/html/2608.07023#S1.p3.1)\.
- \[22\]M\. Wan, T\. Safavi, S\. K\. Jauhar, Y\. Kim, S\. Counts, J\. Neville, S\. Suri, C\. Shah, R\. W\. White, L\. Yang,et al\.\(2024\)Tnt\-llm: text mining at scale with large language models\.InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 5836–5847\.Cited by:[§1](https://arxiv.org/html/2608.07023#S1.p1.1),[§2](https://arxiv.org/html/2608.07023#S2.p2.1),[§5](https://arxiv.org/html/2608.07023#S5.p2.1)\.
- \[23\]Z\. Wang, J\. Shang, and R\. Zhong\(2023\)Goal\-driven explainable clustering via language descriptions\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 10626–10649\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[24\]S\. Yao, D\. Yu, J\. Zhao, I\. Shafran, T\. L\. Griffiths, Y\. Cao, and K\. Narasimhan\(2023\-05\)Tree of Thoughts: Deliberate Problem Solving with Large Language Models\.arXiv e\-prints,pp\. arXiv:2305\.10601\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2305.10601)Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p2.1)\.
- \[25\]Q\. Zeng, Y\. Bai, Z\. Tan, S\. Feng, Z\. Liang, Z\. Zhang, and M\. Jiang\(2024\)Chain\-of\-layer: iteratively prompting large language models for taxonomy induction from limited examples\.InProceedings of the 33rd ACM International Conference on Information and Knowledge Management,pp\. 3093–3102\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p1.1)\.
- \[26\]C\. Zhang, F\. Tao, X\. Chen, J\. Shen, M\. Jiang, B\. Sadler, M\. Vanni, and J\. Han\(2018\)Taxogen: unsupervised topic taxonomy construction by adaptive term embedding and clustering\.InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining,pp\. 2701–2709\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p1.1)\.
- \[27\]M\. Zhang, K\. Jensen, S\. Sonniks, and B\. Plank\(2022\)SkillSpan: hard and soft skill extraction from english job postings\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 4962–4984\.Cited by:[§2](https://arxiv.org/html/2608.07023#S2.p1.1)\.

## Supplementary Material

To ease reproducibility while respecting the strict confidentiality of our proprietary platform data and internal infrastructure, we provide in this appendix the templates of all the prompts for all five stages of the pipeline\. These prompts were executed with Gemini 1\.5 Flash with temperature set to 0 and structured output defined\. Because our architecture is intrinsically dataset\-agnostic and anchored to public Wikidata entities, researchers can readily implement and benchmark this exact pipeline using open\-source labor market datasets\.

## Appendix 0\.AReconciliation details

This section details the initial phase of the pipeline, where noisy, unstandardized skill mentions are mapped to stable Wikidata entity identifiers \(QIDs\)\. As demonstrated in the prompt below, the model is strictly grounded using enriched platform context—specifically, empirical co\-occurring peer skills and overarching professional categories—to accurately disambiguate and anchor the input while mitigating hallucinations\.

[⬇](data:text/plain;base64,WW91IGFyZSBhIGhpZ2hseSBwcmVjaXNlIFNraWxsIEVudGl0eSBMaW5raW5nIGFnZW50LiBZb3VyIG1pc3Npb24gaXMgdG8gYW5hbHl6ZSBhIGZyZWVsYW5jZXIncyBzdGF0ZWQgZXhwZXJ0aXNlLCBkZXRlcm1pbmUgaWYgaXQgcXVhbGlmaWVzIGFzIGEgcHJvZmVzc2lvbmFsIHNraWxsLCBhbmQgbGluayBpdCB0byB0aGUgbW9zdCByZWxldmFudCBXaWtpZGF0YSBpdGVtKHMpIGZyb20gYSBwcm92aWRlZCBsaXN0IG9mIGNhbmRpZGF0ZXMuCgojIyMgUHJpbWFyeSBEaXJlY3RpdmVzOgoxLiAqKkNsYXNzaWZ5IHRoZSBFeHBlcnRpc2UqKjogRmlyc3QsIGRldGVybWluZSBpZiB0aGUgaW5wdXQgd29yZCAie2V4cGVydGlzZX0iIGlzIGEgc2tpbGwgYW5kIHdoZXRoZXIgaXQgaXMgYSBzaW5nbGUgb3IgY29tcG91bmQgc2tpbGwuCjIuICoqTGluayB0aGUgU2tpbGwqKjogSWYgaXQgaXMgYSBza2lsbCwgc2VsZWN0IHRoZSBiZXN0IG1hdGNoaW5nIFdpa2lkYXRhIFFJRChzKSBmcm9tIHRoZSBjYW5kaWRhdGVzIHVzaW5nIHRoZSBzdHJpY3QgU2VsZWN0aW9uIFJ1bGVzIGJlbG93LgoKIyMjIERlZmluaXRpb24gb2YgYSBTa2lsbDoKLSAqKldoYXQgSVMgYSBza2lsbCoqOiBTcGVjaWZpYywgbGVhcm5hYmxlIHByb2Zlc3Npb25hbCBhYmlsaXRpZXMuIEluY2x1ZGVzIHRlY2huaWNhbCB0b29scywgZnJhbWV3b3JrcywgcHJvZ3JhbW1pbmcgbGFuZ3VhZ2VzLCBtZXRob2RvbG9naWVzLCBhbmQgc3BlY2lmaWMgZG9tYWlucy4KLSAqKldoYXQgaXMgTk9UIGEgc2tpbGwqKjoKICAtIEdlbmVyYWwgcGVyc29uYWwgYXR0cmlidXRlcyBvciBzb2Z0IHNraWxscyAoZS5nLiwgIkhhcmQgd29ya2VyIiwgIk1vdGl2YXRlZCIpLgogIC0gTGV2ZWxzIG9mIHNlbmlvcml0eSBvciB1bml0cyBvZiB0aW1lIChlLmcuLCAiMTAgeWVhcnMgZXhwZXJpZW5jZSIsICJTZW5pb3IiKS4KCiMjIyBDYW5kaWRhdGUgU2VsZWN0aW9uIFJ1bGVzIChDUlVDSUFMKToKQW5hbHl6ZSB0aGUgcHJvdmlkZWQgV2lraWRhdGEgY2FuZGlkYXRlcyBjYXJlZnVsbHkuIFlvdSBtdXN0IG5hdmlnYXRlIHRoZSBmb2xsb3dpbmcgZWRnZSBjYXNlczoKMS4gKipEaXJlY3QgTWF0Y2ggUHJpbmNpcGxlKio6IEtlZXAgT05MWSBRSURzIHRoYXQgcmVwcmVzZW50IHRoZSBleHBlcnRpc2UgZGlyZWN0bHksIG9yIHJlcHJlc2VudCBhIGxlZ2l0aW1hdGUgY29tcG9uZW50IHBhcnQgb2YgYSBjb21wb3VuZCBleHBlcnRpc2UuCjIuICoqVGhlIE92ZXJsYXAgUnVsZSAoRGVkdXBsaWNhdGlvbikqKjogSWYgbXVsdGlwbGUgUUlEcyByZWZlciB0byB0aGUgZXhhY3Qgc2FtZSBjb25jZXB0IG9yIHRoZSBzYW1lIHBhcnQgb2YgdGhlIGV4cGVydGlzZSwgKiprZWVwIG9ubHkgdGhlIHNpbmdsZSBiZXN0LW1hdGNoaW5nIG9uZSoqIGFuZCBkaXNjYXJkIHRoZSByZXN0LiBEbyBub3QgcmV0dXJuIDMgZGlmZmVyZW50IFFJRHMgdGhhdCBhbGwgbWVhbiB0aGUgc2FtZSB0aGluZy4KMy4gKipUaGUgQ29tcG91bmQgLyBDb21iaW5hdGlvbiBSdWxlKio6CiAgIC0gRnJlZWxhbmNlcnMgb2Z0ZW4gdHlwZSBjb21wb3VuZCBza2lsbHMgKGUuZy4sICJSZWFjdC5qcyAmIE5vZGUuanMiKS4KICAgLSBTb21ldGltZXMgYSBza2lsbCBpcyBhIGNvbWJpbmF0aW9uIG9mIGNvbmNlcHRzIGFuZCBXaWtpZGF0YSBpcyB0b28gZmluZS1ncmFpbmVkLgogICAtIElmIGEgc2luZ2xlIFFJRCBkb2VzIG5vdCBjb3ZlciB0aGUgZnVsbCBzaWduYWwgb2YgdGhlIGV4cGVydGlzZSwgeW91IE1VU1Qgc2VsZWN0IGEgY29tYmluYXRpb24gb2YgUUlEcyB0aGF0IHByZXNlcnZlcyB0aGUgZnVsbCBzaWduYWwuIFByZWZlciBhIGNvbWJpbmF0aW9uIG92ZXIgYSBzaW5nbGUgcGFydGlhbCBtYXRjaC4KNC4gKipUaGUgRGlyZWN0aW9uYWxpdHkgUnVsZSoqOgogICAtIElmIHRoZSBjb21wb3VuZCBza2lsbCByZXByZXNlbnRzIGEgZGlyZWN0aW9uYWwgcHJvY2VzcyB3aGVyZSB0aGUgb3JkZXIgb2YgaXRlbXMgc3RyaWN0bHkgbWF0dGVycyAoZS5nLiwgIkVuZ2xpc2ggdG8gRnJlbmNoIHRyYW5zbGF0aW9uIiwgIkZpZ21hIHRvIFJlYWN0IiwgIkRhdGEgbWlncmF0aW9uIGZyb20gT3JhY2xlIHRvIFBvc3RncmVzIiksIHlvdSBNVVNUIHNldCAiaXNfZGlyZWN0aW9uYWwiOiB0cnVlLgogICAtIEZvciBzdGFuZGFyZCBjb21iaW5hdGlvbnMgd2hlcmUgb3JkZXIgZG9lc24ndCBtYXR0ZXIgKGUuZy4sICJSZWFjdCBhbmQgTm9kZS5qcyIpLCBzZXQgaXQgdG8gZmFsc2UuCjUuICoqTVVMVElMSU5HVUFMIFVOSUZJQ0FUSU9OIFJVTEUgKEFudGktU3BsaXR0aW5nKSoqOgogICAtIFdpa2lkYXRhIFFJRHMgYXJlIGxhbmd1YWdlLWFnbm9zdGljIGNvbmNlcHRzLiBZb3VyIGdvYWwgaXMgdG8gbWFwIGV4YWN0IHRyYW5zbGF0aW9ucyB0byB0aGUgU0FNRSBwcmltYXJ5IHVuaXZlcnNhbCBRSUQuCiAgIC0gSWYgdGhlIGlucHV0IGV4cGVydGlzZSBpcyBpbiBhIG5vbi1FbmdsaXNoIGxhbmd1YWdlIChlLmcuICJSw6lzZWF1eCBzb2NpYXV4IiBpbiBGcmVuY2gpLCBtYXAgaXQgdG8gdGhlIHByaW1hcnkgZ2xvYmFsIFFJRCBmb3IgdGhhdCBjb25jZXB0ICh3aGljaCBpcyB1c3VhbGx5IGFuY2hvcmVkIGJ5IHRoZSBFbmdsaXNoIHN0YW5kYXJkLCBlLmcuICJTb2NpYWwgTWVkaWEiKS4KICAgLSBETyBOT1Qgc2VsZWN0IGEgc2Vjb25kYXJ5LCBuYXJyb3dlciBRSUQganVzdCBiZWNhdXNlIGl0cyB0cmFuc2xhdGVkIGxhYmVsIGlzIGEgY2xvc2VyIGxpdGVyYWwgbWF0Y2guIEZvcmNlIGRpcmVjdCB0cmFuc2xhdGlvbnMgdG8gY29udmVyZ2Ugb24gdGhlIHNhbWUgY2VudHJhbCBRSUQgdG8gYXZvaWQgbGFuZ3VhZ2Ugc2lsb2luZy4KNi4gKipDUklUSUNBTCBMQU5HVUFHRSBSVUxFIChGYWxzZSBGcmllbmRzKSoqOgogICAtIEJld2FyZSBvZiAiRmFsc2UgRnJpZW5kcyIgKEZhdXggYW1pcykgYWNyb3NzIGxhbmd1YWdlcy4gRG8gbm90IG1hcCBhIGZvcmVpZ24gd29yZCB0byBhbiBFbmdsaXNoIFdpa2lkYXRhIGNvbmNlcHQganVzdCBiZWNhdXNlIHRoZXkgYXJlIHNwZWxsZWQgc2ltaWxhcmx5IGlmIHRoZSBwcm9mZXNzaW9uYWwgbWVhbmluZyBpcyBkaWZmZXJlbnQuCiAgIC0gRXhhbXBsZTogVGhlIEZyZW5jaCAiUsOpZGFjdGlvbiIgbWVhbnMgIkNvcHl3cml0aW5nL1dyaXRpbmciLCBpdCBkb2VzIE5PVCBtZWFuIHRoZSBFbmdsaXNoICJSZWRhY3Rpb24vQ2Vuc29yaW5nIi4gUHJpb3JpdGl6ZSB0aGUgc2VtYW50aWMgbWVhbmluZyB1c2VkIGluIGEgZnJlZWxhbmNlIG1hcmtldHBsYWNlIGNvbnRleHQuCgojIyMgU1BFQ0lGSUNJVFkgUlVMRToKSWYgYSBmcmVlbGFuY2Ugc2tpbGwgbWVudGlvbnMgYm90aCBhIGJyb2FkIGNhdGVnb3J5IGFuZCBhIHNwZWNpZmljIGZyYW1ld29yayAoZS5nLiwgJ03DqXRob2RlIEFnaWxlIFNjcnVtJyksIGRvIE5PVCByZXR1cm4gbXVsdGlwbGUgUUlEcy4gWW91IE1VU1QgcmV0dXJuIE9OTFkgdGhlIFFJRCBvZiB0aGUgbW9zdCBzcGVjaWZpYywgcHJpbWFyeSBjb25jZXB0IChlLmcuLCBvbmx5IHRoZSBRSUQgZm9yICdTY3J1bScpLiBOZXZlciBjcmVhdGUgY29tcG9zaXRlIFFJRCBsaXN0cyB1bmxlc3MgdGhlIHNraWxsIGlzIHRydWx5IGEgZGlyZWN0aW9uYWwgbWFwcGluZyAobGlrZSB0cmFuc2xhdGluZyBmcm9tIG9uZSBsYW5ndWFnZSB0byBhbm90aGVyKS4KCiMjIyBTdGVwLWJ5LVN0ZXAgSW5zdHJ1Y3Rpb25zOgoxLiAqKkFuYWx5emUgIntleHBlcnRpc2V9IioqOiBCYXNlZCBvbiB0aGUgZGVmaW5pdGlvbnMgYWJvdmUsIGRlY2lkZSBpZiBpdCdzIGEgc2tpbGwuCjIuICoqSWRlbnRpZnkgQ29tcG91bmQgU2tpbGxzKio6IElmIGl0IElTIGEgc2tpbGwsIGRldGVybWluZSBpZiBpdCBpcyBhIGNvbXBvdW5kIHNraWxsLgozLiAqKkV2YWx1YXRlIENhbmRpZGF0ZXMqKjogQ2FyZWZ1bGx5IGV2YWx1YXRlIGVhY2ggY2FuZGlkYXRlIFFJRCB1c2luZyB0aGUgU2VsZWN0aW9uIFJ1bGVzIGFib3ZlLiBVc2UgYWxsIHByb3ZpZGVkIGNvbnRleHQgKGNvLW9jY3VycmluZyBza2lsbHMsIGpvYiBjYXRlZ29yaWVzLCBXaWtpZGF0YSBpbmZvKS4KNC4gKipTZWxlY3QgJiBTY29yZSoqOgogICAtIENob29zZSB0aGUgYmVzdCBRSUQocykuIEZvciBhIGNvbXBvdW5kIHNraWxsLCB5b3UgTVVTVCByZXR1cm4gb25lIGVudHJ5IGZvciBlYWNoIGRpc3RpbmN0IHNraWxsIGlkZW50aWZpZWQuCiAgIC0gRm9yIGVhY2ggc2VsZWN0ZWQgUUlELCBwcm92aWRlIGEgc2NvcmUgYW5kIGEgY29uY2lzZSByZWFzb25pbmcuCgojIyMgU2NvcmluZyBSdWJyaWM6Ci0gKiowLjAqKjogVGhlIGJlc3QgbWF0Y2hpbmcgd2lraWRhdGEgaXRlbShzKSBpcyBub3QgcmVhbGx5IGEgbWF0Y2guIEl0IGlzIHRvdGFsbHkgaXJyZWxldmFudCBmb3IgdGhlIHByb3ZpZGVkIHNraWxsLgotICoqMC40Kio6IE1vZGVyYXRlIG1hdGNoLiBNYXliZSB0aGUgcGVyZmVjdCB3aWtpZGF0YSBpdGVtIGRvZXNuJ3QgZXhpc3Qgb3Igd2FzIG5vdCBwcm92aWRlZC4KLSAqKjAuNyoqOiBTdHJvbmcgbWF0Y2guIFRoZSB3aWtpZGF0YSBpdGVtIGNhcHR1cmVzIHRoZSBjb25jZXB0IHdlbGwuCi0gKioxLjAqKjogUGVyZmVjdCBtYXRjaC4gTm8gb3RoZXIgd2lraWRhdGEgaXRlbSBvciBjb25jZXB0IHdpbGwgYmUgYmV0dGVyLgoKIyMjIFJlYXNvbmluZyBJbnN0cnVjdGlvbnM6Ci0gWW91ciByZWFzb25pbmcgbXVzdCBiZSBjb25jaXNlICgxLTIgc2VudGVuY2VzKS4KLSAqKkp1c3RpZnkgeW91ciBkZWNpc2lvbiBieSByZWZlcmVuY2luZyB0aGUgcHJvdmlkZWQgY29udGV4dCoqLgoKIyMjIENPTlRFWFQgRk9SIFRIRSBUQVNLCiMjIyBDby1vY2N1cnJpbmcgc2tpbGxzIGZvciAie2V4cGVydGlzZX0iOgp7Y29fb2NjdXJyZW5jZXN9CgojIyMgSm9iIGNhdGVnb3JpZXMgZm9yICJ7ZXhwZXJ0aXNlfSI6CntjYXRlZ29yaWVzfQoKIyMjIFdpa2lkYXRhIGl0ZW0gY2FuZGlkYXRlcyBmb3IgIntleHBlcnRpc2V9IjoKe2NhbmRpZGF0ZXN9)YouareahighlypreciseSkillEntityLinkingagent\.Yourmissionistoanalyzeafreelancer’sstatedexpertise,determineifitqualifiesasaprofessionalskill,andlinkittothemostrelevantWikidataitem\(s\)fromaprovidedlistofcandidates\.\#\#\#PrimaryDirectives:1\.\*\*ClassifytheExpertise\*\*:First,determineiftheinputword"\{expertise\}"isaskillandwhetheritisasingleorcompoundskill\.2\.\*\*LinktheSkill\*\*:Ifitisaskill,selectthebestmatchingWikidataQID\(s\)fromthecandidatesusingthestrictSelectionRulesbelow\.\#\#\#DefinitionofaSkill:\-\*\*WhatISaskill\*\*:Specific,learnableprofessionalabilities\.Includestechnicaltools,frameworks,programminglanguages,methodologies,andspecificdomains\.\-\*\*WhatisNOTaskill\*\*:\-Generalpersonalattributesorsoftskills\(e\.g\.,"Hardworker","Motivated"\)\.\-Levelsofseniorityorunitsoftime\(e\.g\.,"10yearsexperience","Senior"\)\.\#\#\#CandidateSelectionRules\(CRUCIAL\):AnalyzetheprovidedWikidatacandidatescarefully\.Youmustnavigatethefollowingedgecases:1\.\*\*DirectMatchPrinciple\*\*:KeepONLYQIDsthatrepresenttheexpertisedirectly,orrepresentalegitimatecomponentpartofacompoundexpertise\.2\.\*\*TheOverlapRule\(Deduplication\)\*\*:IfmultipleQIDsrefertotheexactsameconceptorthesamepartoftheexpertise,\*\*keeponlythesinglebest\-matchingone\*\*anddiscardtherest\.Donotreturn3differentQIDsthatallmeanthesamething\.3\.\*\*TheCompound/CombinationRule\*\*:\-Freelancersoftentypecompoundskills\(e\.g\.,"React\.js&Node\.js"\)\.\-SometimesaskillisacombinationofconceptsandWikidataistoofine\-grained\.\-IfasingleQIDdoesnotcoverthefullsignaloftheexpertise,youMUSTselectacombinationofQIDsthatpreservesthefullsignal\.Preferacombinationoverasinglepartialmatch\.4\.\*\*TheDirectionalityRule\*\*:\-Ifthecompoundskillrepresentsadirectionalprocesswheretheorderofitemsstrictlymatters\(e\.g\.,"EnglishtoFrenchtranslation","FigmatoReact","DatamigrationfromOracletoPostgres"\),youMUSTset"is\_directional":true\.\-Forstandardcombinationswhereorderdoesn’tmatter\(e\.g\.,"ReactandNode\.js"\),setittofalse\.5\.\*\*MULTILINGUALUNIFICATIONRULE\(Anti\-Splitting\)\*\*:\-WikidataQIDsarelanguage\-agnosticconcepts\.YourgoalistomapexacttranslationstotheSAMEprimaryuniversalQID\.\-Iftheinputexpertiseisinanon\-Englishlanguage\(e\.g\."Réseauxsociaux"inFrench\),mapittotheprimaryglobalQIDforthatconcept\(whichisusuallyanchoredbytheEnglishstandard,e\.g\."SocialMedia"\)\.\-DONOTselectasecondary,narrowerQIDjustbecauseitstranslatedlabelisacloserliteralmatch\.ForcedirecttranslationstoconvergeonthesamecentralQIDtoavoidlanguagesiloing\.6\.\*\*CRITICALLANGUAGERULE\(FalseFriends\)\*\*:\-Bewareof"FalseFriends"\(Fauxamis\)acrosslanguages\.DonotmapaforeignwordtoanEnglishWikidataconceptjustbecausetheyarespelledsimilarlyiftheprofessionalmeaningisdifferent\.\-Example:TheFrench"Rédaction"means"Copywriting/Writing",itdoesNOTmeantheEnglish"Redaction/Censoring"\.Prioritizethesemanticmeaningusedinafreelancemarketplacecontext\.\#\#\#SPECIFICITYRULE:Ifafreelanceskillmentionsbothabroadcategoryandaspecificframework\(e\.g\.,’MéthodeAgileScrum’\),doNOTreturnmultipleQIDs\.YouMUSTreturnONLYtheQIDofthemostspecific,primaryconcept\(e\.g\.,onlytheQIDfor’Scrum’\)\.NevercreatecompositeQIDlistsunlesstheskillistrulyadirectionalmapping\(liketranslatingfromonelanguagetoanother\)\.\#\#\#Step\-by\-StepInstructions:1\.\*\*Analyze"\{expertise\}"\*\*:Basedonthedefinitionsabove,decideifit’saskill\.2\.\*\*IdentifyCompoundSkills\*\*:IfitISaskill,determineifitisacompoundskill\.3\.\*\*EvaluateCandidates\*\*:CarefullyevaluateeachcandidateQIDusingtheSelectionRulesabove\.Useallprovidedcontext\(co\-occurringskills,jobcategories,Wikidatainfo\)\.4\.\*\*Select&Score\*\*:\-ChoosethebestQID\(s\)\.Foracompoundskill,youMUSTreturnoneentryforeachdistinctskillidentified\.\-ForeachselectedQID,provideascoreandaconcisereasoning\.\#\#\#ScoringRubric:\-\*\*0\.0\*\*:Thebestmatchingwikidataitem\(s\)isnotreallyamatch\.Itistotallyirrelevantfortheprovidedskill\.\-\*\*0\.4\*\*:Moderatematch\.Maybetheperfectwikidataitemdoesn’texistorwasnotprovided\.\-\*\*0\.7\*\*:Strongmatch\.Thewikidataitemcapturestheconceptwell\.\-\*\*1\.0\*\*:Perfectmatch\.Nootherwikidataitemorconceptwillbebetter\.\#\#\#ReasoningInstructions:\-Yourreasoningmustbeconcise\(1\-2sentences\)\.\-\*\*Justifyyourdecisionbyreferencingtheprovidedcontext\*\*\.\#\#\#CONTEXTFORTHETASK\#\#\#Co\-occurringskillsfor"\{expertise\}":\{co\_occurrences\}\#\#\#Jobcategoriesfor"\{expertise\}":\{categories\}\#\#\#Wikidataitemcandidatesfor"\{expertise\}":\{candidates\}

Figure 4:Semantic Reconciliation Prompt
## Appendix 0\.BCanonicalization details

Following reconciliation, the canonicalization phase groups validated inputs by their resolved QIDs to synthesize localized, human\-readable preferred labels across the five target languages\. To handle the diverse complexity of the data, the system utilizes three distinct prompts depending on the entity type: single Wikidata concepts, compound concepts requiring directional logic, and completely novel orphaned skills that require synthetic identifiers\.

[⬇](data:text/plain;base64,WW91IGFyZSBhbiBleHBlcnQgKipNdWx0aWxpbmd1YWwgVGF4b25vbXkgTGluZ3Vpc3QgYW5kIEhSIFNwZWNpYWxpc3QqKi4KWW91ciBPTkxZIHRhc2sgaXMgdG8gYW5hbHl6ZSBhIHNpbmdsZSBXaWtpZGF0YSBjb25jZXB0IGFuZCBpdHMgYXNzb2NpYXRlZCByYXcgc2tpbGxzIHRvIGRldGVybWluZSB0aGUgYmVzdCBQcmVmZXJyZWQgTGFiZWwgKGBwcmVmX2xhYmVsYCkgZm9yIGVhY2ggdGFyZ2V0IGxhbmd1YWdlLgoKLS0tCgojIyMgUFJJTUFSWSBESVJFQ1RJVkVTCgoxLiAqKlByZWZlcnJlZCBMYWJlbCBTZWxlY3Rpb24qKgogICAtIFByaW9yaXRpemUgdGhlIG1vc3QgKipmcmVxdWVudGx5IHVzZWQqKiBsYWJlbCBmcm9tIHRoZSBNYWx0IHVzYWdlIGNvdW50cyBpZiBpdCBpcyBwcm9mZXNzaW9uYWwuCiAgIC0gRmFsbGJhY2sgdG8gdGhlICoqV2lraWRhdGEgbGFiZWwqKiBpZiB0aGUgdXNhZ2UtYmFzZWQgdGVybXMgYXJlIGFtYmlndW91cywgaW5mb3JtYWwsIG9yIGluYXBwcm9wcmlhdGUuCiAgIC0gQ3JlYXRlIGEgKipTeW50aGV0aWMqKiAobmV3KSBwcm9mZXNzaW9uYWwgbGFiZWwgaWYgbmVpdGhlciBhY2N1cmF0ZWx5IGRlc2NyaWJlcyB0aGUgY29yZSBjb25jZXB0LgoKMi4gKipQcm92ZW5hbmNlIFRyYWNraW5nIChgc291cmNlYCkqKgogICAtIEZvciBldmVyeSBsb2NhbGl6ZWQgbGFiZWwsIHNwZWNpZnkgaXRzIG9yaWdpbjoKICAgICAtIGAibWFsdCJgOiBiYXNlZCBvbiBhbiBleGlzdGluZyBoaWdoLXVzYWdlIHByb2ZpbGUgZXhwZXJ0aXNlLgogICAgIC0gYCJ3aWtpZGF0YSJgOiBzZWxlY3RlZCB0aGUgb2ZmaWNpYWwgV2lraWRhdGEgZGVzY3JpcHRpb24gb3IgbGFuZ3VhZ2UgdmFyaWFudC4KICAgICAtIGAiZ2VtaW5pImA6IHlvdSBzeW50aGVzaXplZCBhIGJyYW5kIG5ldyB0ZXJtIHlvdXJzZWxmLgoKLS0tCgojIyMgSU5QVVQgREFUQQoKKipUYXJnZXQgTGFuZ3VhZ2VzOioqIHsiLCAiLmpvaW4obGFuZ3VhZ2VzKX0KCioqMS4gQXNzb2NpYXRlZCBXaWtpZGF0YSBDb25jZXB0OioqCnt3ZF90ZXh0fQoKKioyLiBTa2lsbHMgaW4gdGhpcyBHcm91cCAoVXNlIGZvciBjb250ZXh0IHRvIHVuZGVyc3RhbmQgdGhlIGNvbW11bml0eSB1c2FnZSk6KioKe21hbHRfdGV4dF9qb2luZWR9)Youareanexpert\*\*MultilingualTaxonomyLinguistandHRSpecialist\*\*\.YourONLYtaskistoanalyzeasingleWikidataconceptanditsassociatedrawskillstodeterminethebestPreferredLabel\(‘pref\_label‘\)foreachtargetlanguage\.\-\-\-\#\#\#PRIMARYDIRECTIVES1\.\*\*PreferredLabelSelection\*\*\-Prioritizethemost\*\*frequentlyused\*\*labelfromtheMaltusagecountsifitisprofessional\.\-Fallbacktothe\*\*Wikidatalabel\*\*iftheusage\-basedtermsareambiguous,informal,orinappropriate\.\-Createa\*\*Synthetic\*\*\(new\)professionallabelifneitheraccuratelydescribesthecoreconcept\.2\.\*\*ProvenanceTracking\(‘source‘\)\*\*\-Foreverylocalizedlabel,specifyitsorigin:\-‘"malt"‘:basedonanexistinghigh\-usageprofileexpertise\.\-‘"wikidata"‘:selectedtheofficialWikidatadescriptionorlanguagevariant\.\-‘"gemini"‘:yousynthesizedabrandnewtermyourself\.\-\-\-\#\#\#INPUTDATA\*\*TargetLanguages:\*\*\{","\.join\(languages\)\}\*\*1\.AssociatedWikidataConcept:\*\*\{wd\_text\}\*\*2\.SkillsinthisGroup\(Useforcontexttounderstandthecommunityusage\):\*\*\{malt\_text\_joined\}

Figure 5:Canonicalization Prompt \- Single QID[⬇](data:text/plain;base64,WW91IGFyZSBhbiBleHBlcnQgKipNdWx0aWxpbmd1YWwgVGF4b25vbXkgTGluZ3Vpc3QgYW5kIEhSIFNwZWNpYWxpc3QqKi4KWW91ciBPTkxZIHRhc2sgaXMgdG8gYW5hbHl6ZSBhIGNvbWJpbmF0aW9uIG9mIFdpa2lkYXRhIGNvbmNlcHRzIGFuZCB0aGUgcmF3IHNraWxscyBhc3NvY2lhdGVkIHdpdGggdGhlbSB0byBkZXRlcm1pbmUgdGhlIGJlc3QgdW5pZmllZCBQcmVmZXJyZWQgTGFiZWwgKGBwcmVmX2xhYmVsYCkgYW5kIGZpbmFsIENvbmNlcHQgSWRlbnRpZmllci4KCi0tLQoKIyMjIFBSSU1BUlkgRElSRUNUSVZFUwoKMS4gKipDb21wb3VuZCBMYWJlbCBTZWxlY3Rpb24qKgogICAtIFlvdSBNVVNUIGNob29zZSBhIFByZWZlcnJlZCBMYWJlbCB0aGF0IGlzIGNvaGVyZW50IHdpdGggdGhlICoqbWFqb3JpdHkgb2YgdGhlIGhpZ2gtdXNhZ2Ugc2tpbGxzKiogd2l0aGluIHRoZSBjbHVzdGVyLgogICAtICoqRGlyZWN0aW9uYWxpdHkgUnVsZSAoQ1JJVElDQUwpOioqIElmIGBJcyBEaXJlY3Rpb25hbCBQcm9jZXNzYCBpcyBgVHJ1ZWAgaW4gdGhlIElucHV0IERhdGEsIHRoZSBvcmRlciBvZiB0aGUgY29uY2VwdHMgc3RyaWN0bHkgbWF0dGVycyAoZS5nLiwgYSBtaWdyYXRpb24gb3IgdHJhbnNsYXRpb24gZnJvbSBDb25jZXB0IEEgdG8gQ29uY2VwdCBCKS4gWW91ciBsYWJlbCBNVVNUIHJlZmxlY3QgdGhpcyBkaXJlY3Rpb25hbCByZWxhdGlvbnNoaXAgKGUuZy4sICJFbmdsaXNoIHRvIEZyZW5jaCBUcmFuc2xhdGlvbiIgaW5zdGVhZCBvZiAiRW5nbGlzaCBhbmQgRnJlbmNoIikuCiAgIC0gSWYgYElzIERpcmVjdGlvbmFsIFByb2Nlc3NgIGlzIGBGYWxzZWAsIHRyZWF0IGl0IGFzIGEgc3RhbmRhcmQgY29tYmluYXRpb24gYW5kIGxvb2sgZm9yIGEgc3RhbmRhcmQgaW5kdXN0cnkgdW1icmVsbGEgdGVybSAoZS5nLiwgIk1FUk4gU3RhY2siKS4KCjIuICoqRGVmaW5lIExvZ2ljIGZvciB0aGUgSWRlbnRpZmllcioqCiAgIC0gVGhlIEJhc2VsaW5lIENvbXBvc2l0ZSBJRCBncm91cHMgdGhlIHByb3ZpZGVkIFFJRHMgKGUuZy4sIGB7cWlkX2NvbWJvfWApLgogICAtIElmIHlvdSBtb2RpZnkgb3IgYWRkIHRvIHRoaXMgaWRlbnRpZmllciwgdXNlIGAmYCAoQU5EKSBmb3IgaW50ZWdyYXRlZCBwcmFjdGljZXMsIGFuZCBgfGAgKE9SKSBmb3IgaW5kZXBlbmRlbnQgdHJhaXRzLgoKMy4gKipTeW50aGV0aWMgSWRlbnRpZmllcnMqKgogICAtIFlvdSBtdXN0IGV4Y2x1c2l2ZWx5IHVzZSB0aGUgcHJvdmlkZWQgV2lraWRhdGEgUUlEcyB3aGVuZXZlciBwb3NzaWJsZS4KICAgLSBIb3dldmVyLCBpZiB0aGUgbWFqb3JpdHkgb2YgdGhlIHNraWxscyBpbnRyb2R1Y2UgYSBjcml0aWNhbCwgZGlzdGluY3QgY29uY2VwdCB0aGF0IGlzIE5PVCBjb3ZlcmVkIGJ5IGFueSBwcm92aWRlZCBRSUQsIHlvdSBNVVNUIGludmVudCBhIGNvbmNpc2UsIHVwcGVyY2FzZSBFbmdsaXNoIHRleHR1YWwgaWRlbnRpZmllciBhbmQgYXBwZW5kIGl0IChlLmcuLCBge3FpZF9jb21ib30gJiBXRUJgIG9yIGB7cWlkX2NvbWJvfSB8IElUYCkuIFVzZSB0aGUgZXhhY3Qgc2FtZSB0ZXh0dWFsIGlkZW50aWZpZXIgYWNyb3NzIGFsbCBsYW5ndWFnZXMgZm9yIGNvbnNpc3RlbmN5LgoKNC4gKipDcmVhdGUgUHJlZmVycmVkIExhYmVscyoqCiAgIC0gTG9jYWxpemUgbmFtZXMgZm9yIHRoZSBmaW5hbCBjb21wb3NpdGUgaWRlbnRpZmllciAoaW5jbHVkaW5nIHRob3NlIHdpdGggc3ludGhldGljIHRleHR1YWwgSURzIGFwcGVuZGVkKS4gUHJpb3JpdGl6ZSB1c2FnZSBjb3VudHMsIGZhbGxiYWNrIHRvIFdpa2lkYXRhLCBvciBzeW50aGVzaXplIGEgcHJvZmVzc2lvbmFsIEhSIHRlcm0gaWYgbmVlZGVkLgoKNS4gKipQcm92ZW5hbmNlIFRyYWNraW5nIChgc291cmNlYCkqKgogICAtIEZvciBldmVyeSBsb2NhbGl6ZWQgbGFiZWwsIHNwZWNpZnkgaXRzIG9yaWdpbjoKICAgICAtIGAibWFsdCJgOiBkZXJpdmVkIGZyb20gYW4gZXhpc3RpbmcgcHJvZmlsZSBleHBlcnRpc2UuCiAgICAgLSBgIndpa2lkYXRhImA6IHRha2VuIGZyb20gV2lraWRhdGEgY29udGV4dC4KICAgICAtIGAiZ2VtaW5pImA6IGEgbmV3IHN5bnRoZXNpemVkIGNvbXBvc2l0ZSB0ZXJtLgoKLS0tCgojIyMgRk9STUFUVElORyBSVUxFUyAoQ1JJVElDQUwpCi0gKipObyBJbnRlcm5hbCBRdW90ZXM6KiogRG8gTk9UIHdyYXAgeW91ciBsYWJlbCBzdHJpbmcgdmFsdWVzIGluIHNpbmdsZSBvciBkb3VibGUgcXVvdGVzLgoKLS0tCgojIyMgSU5QVVQgREFUQQoKKipUYXJnZXQgTGFuZ3VhZ2VzOioqIHsiLCAiLmpvaW4obGFuZ3VhZ2VzKX0KCioqQmFzZWxpbmUgQ29tcG9zaXRlIElEOioqIHtxaWRfY29tYm99CioqSXMgRGlyZWN0aW9uYWwgUHJvY2VzczoqKiB7aXNfZGlyZWN0aW9uYWx9e2RpcmVjdGlvbmFsX2hpbnR9CgoqKjEuIEFzc29jaWF0ZWQgV2lraWRhdGEgQ29uY2VwdHMgaW4gdGhpcyBjb21wb3VuZDoqKgp7d2RfdGV4dF9qb2luZWR9CgoqKjIuIFNraWxscyBpbiB0aGlzIEdyb3VwIChVc2UgZm9yIGNvbnRleHQgdG8gZGV0ZXJtaW5lIG1ham9yaXR5IGNvaGVyZW5jZSAmIG1pc3NpbmcgY29uY2VwdHMpOioqCnttYWx0X3RleHRfam9pbmVkfQ==)Youareanexpert\*\*MultilingualTaxonomyLinguistandHRSpecialist\*\*\.YourONLYtaskistoanalyzeacombinationofWikidataconceptsandtherawskillsassociatedwiththemtodeterminethebestunifiedPreferredLabel\(‘pref\_label‘\)andfinalConceptIdentifier\.\-\-\-\#\#\#PRIMARYDIRECTIVES1\.\*\*CompoundLabelSelection\*\*\-YouMUSTchooseaPreferredLabelthatiscoherentwiththe\*\*majorityofthehigh\-usageskills\*\*withinthecluster\.\-\*\*DirectionalityRule\(CRITICAL\):\*\*If‘IsDirectionalProcess‘is‘True‘intheInputData,theorderoftheconceptsstrictlymatters\(e\.g\.,amigrationortranslationfromConceptAtoConceptB\)\.YourlabelMUSTreflectthisdirectionalrelationship\(e\.g\.,"EnglishtoFrenchTranslation"insteadof"EnglishandFrench"\)\.\-If‘IsDirectionalProcess‘is‘False‘,treatitasastandardcombinationandlookforastandardindustryumbrellaterm\(e\.g\.,"MERNStack"\)\.2\.\*\*DefineLogicfortheIdentifier\*\*\-TheBaselineCompositeIDgroupstheprovidedQIDs\(e\.g\.,‘\{qid\_combo\}‘\)\.\-Ifyoumodifyoraddtothisidentifier,use‘&‘\(AND\)forintegratedpractices,and‘\|‘\(OR\)forindependenttraits\.3\.\*\*SyntheticIdentifiers\*\*\-YoumustexclusivelyusetheprovidedWikidataQIDswheneverpossible\.\-However,ifthemajorityoftheskillsintroduceacritical,distinctconceptthatisNOTcoveredbyanyprovidedQID,youMUSTinventaconcise,uppercaseEnglishtextualidentifierandappendit\(e\.g\.,‘\{qid\_combo\}&WEB‘or‘\{qid\_combo\}\|IT‘\)\.Usetheexactsametextualidentifieracrossalllanguagesforconsistency\.4\.\*\*CreatePreferredLabels\*\*\-Localizenamesforthefinalcompositeidentifier\(includingthosewithsynthetictextualIDsappended\)\.Prioritizeusagecounts,fallbacktoWikidata,orsynthesizeaprofessionalHRtermifneeded\.5\.\*\*ProvenanceTracking\(‘source‘\)\*\*\-Foreverylocalizedlabel,specifyitsorigin:\-‘"malt"‘:derivedfromanexistingprofileexpertise\.\-‘"wikidata"‘:takenfromWikidatacontext\.\-‘"gemini"‘:anewsynthesizedcompositeterm\.\-\-\-\#\#\#FORMATTINGRULES\(CRITICAL\)\-\*\*NoInternalQuotes:\*\*DoNOTwrapyourlabelstringvaluesinsingleordoublequotes\.\-\-\-\#\#\#INPUTDATA\*\*TargetLanguages:\*\*\{","\.join\(languages\)\}\*\*BaselineCompositeID:\*\*\{qid\_combo\}\*\*IsDirectionalProcess:\*\*\{is\_directional\}\{directional\_hint\}\*\*1\.AssociatedWikidataConceptsinthiscompound:\*\*\{wd\_text\_joined\}\*\*2\.SkillsinthisGroup\(Useforcontexttodeterminemajoritycoherence&missingconcepts\):\*\*\{malt\_text\_joined\}

Figure 6:Canonicalization Prompt \- Multiple QID[⬇](data:text/plain;base64,WW91IGFyZSBhICoqU2tpbGwgUmVjb25jaWxpYXRpb24gU3BlY2lhbGlzdCoqIGZvciBvcnBoYW5lZCBza2lsbHMuCgpUaGVzZSBza2lsbHMgd2VyZSBSRUpFQ1RFRCBmcm9tIHRoZWlyIG9yaWdpbmFsIGNsdXN0ZXJzIGR1cmluZyBjdXJhdGlvbiBiZWNhdXNlIHRoZXkgd2VyZSBkZWVtZWQgbm90IGVxdWl2YWxlbnQgdG8gdGhlIGNsdXN0ZXIncyBjb3JlIGNvbmNlcHQuCllvdXIgdGFzayBpcyB0byBjcmVhdGUgQ09NUExFVEVMWSBORVcgY29uY2VwdCBpZGVudGlmaWVycyBmb3IgdGhlc2Ugb3JwaGFuZWQgc2tpbGxzLgoKLS0tCgojIyMgUFJJTUFSWSBESVJFQ1RJVkVTCgoxLiAqKkZyZXNoIFJlY29uY2lsaWF0aW9uKioKICAgLSBUcmVhdCBlYWNoIHNraWxsIGFzIGEgcG90ZW50aWFsbHkgTkVXIGNvbmNlcHQgdGhhdCBuZWVkcyBpdHMgb3duIGlkZW50aWZpZXIuCiAgIC0gRE8gTk9UIGFzc3VtZSB0aGVzZSBza2lsbHMgZml0IGludG8gdGhlaXIgcHJldmlvdXMgV2lraWRhdGEgUUlEcy4KICAgLSBUaGUgV2lraWRhdGEgY29udGV4dCBwcm92aWRlZCBpcyBmcm9tIHRoZWlyIE9SSUdJTkFMIChmYWlsZWQpIG1hdGNoaW5nIC0gdXNlIGl0IG9ubHkgYXMgcmVmZXJlbmNlIHRvIGtub3cgd2hhdCB0aGV5IGFyZSBOT1QuCgoyLiAqKlN5bnRoZXRpYyBJRCBDcmVhdGlvbioqCiAgIC0gQ3JlYXRlIG5ldyBzeW50aGV0aWMgaWRlbnRpZmllcnMgdXNpbmcgdGhlIGZvcm1hdDogYFNZTlRIX1NLSUxMX1hYWGAgKHdoZXJlIFhYWCBpcyBhIHVuaXF1ZSBudW1iZXIpLgogICAtIElmIGEgc2tpbGwgZ2VudWluZWx5IG1hdGNoZXMgb25lIG9mIHRoZSBwcm92aWRlZCBXaWtpZGF0YSBRSURzLCB5b3UgbWF5IHVzZSBpdC4KICAgLSBJZiBhIHNraWxsIGNvbWJpbmVzIG11bHRpcGxlIGNvbmNlcHRzLCB1c2UgYCZgIG5vdGF0aW9uIChlLmcuLCBgU1lOVEhfU0tJTExfMDAxICYgUTEyMzQ1YCkuCiAgIC0gRWFjaCB1bmlxdWUgc2tpbGwgY29uY2VwdCBzaG91bGQgZ2V0IGl0cyBvd24gdW5pcXVlIGlkZW50aWZpZXIuIEdyb3VwIHNraWxscyB0b2dldGhlciBpZiB0aGV5IG1lYW4gdGhlIGV4YWN0IHNhbWUgdGhpbmcuCgozLiAqKlByZWZlcnJlZCBMYWJlbCBHZW5lcmF0aW9uKioKICAgLSBGb3IgZWFjaCBORVcgc3ludGhldGljIElELCBjcmVhdGUgbG9jYWxpemVkIHByZWZlcnJlZCBsYWJlbHMgaW4gYWxsIHRhcmdldCBsYW5ndWFnZXMuCiAgIC0gU2V0IGBzb3VyY2VgIHRvIGAiZ2VtaW5pImAgZm9yIHN5bnRoZXNpemVkIGxhYmVscy4KICAgLSBTZXQgYHNvdXJjZWAgdG8gYCJtYWx0ImAgaWYgeW91J3JlIHVzaW5nIGEgaGlnaC1mcmVxdWVuY3kgTWFsdCB0ZXJtLgogICAtIFNldCBgc291cmNlYCB0byBgIndpa2lkYXRhImAgb25seSBpZiB5b3UncmUgcmV1c2luZyBhIFdpa2lkYXRhIGxhYmVsLgoKNC4gKipTa2lsbCBWYWxpZGF0aW9uKioKICAgLSBTZXQgYGlzX3NraWxsOiBmYWxzZWAgaWYgdGhlIHRleHQgaXMgTk9UIGEgdmFsaWQgcHJvZmVzc2lvbmFsIHNraWxsLgogICAtIFVzZSBgcmVqZWN0aW9uX3JlYXNvbmA6IGBOT1RfQV9TS0lMTGAsIGBTRU1BTlRJQ19NSVNNQVRDSGAsIG9yIGBBTUJJR1VPVVNgLgoKLS0tCgojIyMgRk9STUFUVElORyBSVUxFUyAoQ1JJVElDQUwpCi0gKipObyBJbnRlcm5hbCBRdW90ZXM6KiogRG8gTk9UIHdyYXAgbGFiZWwgdmFsdWVzIGluIHF1b3RlcyAod3JpdGUgYERhdGEgQW5hbHlzaXNgLCBOT1QgYCdEYXRhIEFuYWx5c2lzJ2ApLgoKLS0tCgojIyMgSU5QVVQgREFUQQoKKipUYXJnZXQgTGFuZ3VhZ2VzOioqIHsiLCAiLmpvaW4obGFuZ3VhZ2VzKX0KCioqMS4gT3JpZ2luYWwgV2lraWRhdGEgQ29udGV4dCAoZm9yIHJlZmVyZW5jZSBvbmx5KToqKgp7d2RfdGV4dF9qb2luZWR9CgoqKjIuIE9ycGhhbmVkIFNraWxscyB0byBSZWNvbmNpbGU6KioKe21hbHRfdGV4dF9qb2luZWR9CgpSZW1lbWJlcjogVGhlc2Ugc2tpbGxzIGZhaWxlZCB0byBmaXQgdGhlaXIgb3JpZ2luYWwgY2x1c3RlcnMuIENyZWF0ZSBmcmVzaCwgYXBwcm9wcmlhdGUgaWRlbnRpZmllcnMgZm9yIHRoZW0u)Youarea\*\*SkillReconciliationSpecialist\*\*fororphanedskills\.TheseskillswereREJECTEDfromtheiroriginalclustersduringcurationbecausetheyweredeemednotequivalenttothecluster’scoreconcept\.YourtaskistocreateCOMPLETELYNEWconceptidentifiersfortheseorphanedskills\.\-\-\-\#\#\#PRIMARYDIRECTIVES1\.\*\*FreshReconciliation\*\*\-TreateachskillasapotentiallyNEWconceptthatneedsitsownidentifier\.\-DONOTassumetheseskillsfitintotheirpreviousWikidataQIDs\.\-TheWikidatacontextprovidedisfromtheirORIGINAL\(failed\)matching\-useitonlyasreferencetoknowwhattheyareNOT\.2\.\*\*SyntheticIDCreation\*\*\-Createnewsyntheticidentifiersusingtheformat:‘SYNTH\_SKILL\_XXX‘\(whereXXXisauniquenumber\)\.\-IfaskillgenuinelymatchesoneoftheprovidedWikidataQIDs,youmayuseit\.\-Ifaskillcombinesmultipleconcepts,use‘&‘notation\(e\.g\.,‘SYNTH\_SKILL\_001&Q12345‘\)\.\-Eachuniqueskillconceptshouldgetitsownuniqueidentifier\.Groupskillstogetheriftheymeantheexactsamething\.3\.\*\*PreferredLabelGeneration\*\*\-ForeachNEWsyntheticID,createlocalizedpreferredlabelsinalltargetlanguages\.\-Set‘source‘to‘"gemini"‘forsynthesizedlabels\.\-Set‘source‘to‘"malt"‘ifyou’reusingahigh\-frequencyMaltterm\.\-Set‘source‘to‘"wikidata"‘onlyifyou’rereusingaWikidatalabel\.4\.\*\*SkillValidation\*\*\-Set‘is\_skill:false‘ifthetextisNOTavalidprofessionalskill\.\-Use‘rejection\_reason‘:‘NOT\_A\_SKILL‘,‘SEMANTIC\_MISMATCH‘,or‘AMBIGUOUS‘\.\-\-\-\#\#\#FORMATTINGRULES\(CRITICAL\)\-\*\*NoInternalQuotes:\*\*DoNOTwraplabelvaluesinquotes\(write‘DataAnalysis‘,NOT‘’DataAnalysis’‘\)\.\-\-\-\#\#\#INPUTDATA\*\*TargetLanguages:\*\*\{","\.join\(languages\)\}\*\*1\.OriginalWikidataContext\(forreferenceonly\):\*\*\{wd\_text\_joined\}\*\*2\.OrphanedSkillstoReconcile:\*\*\{malt\_text\_joined\}Remember:Theseskillsfailedtofittheiroriginalclusters\.Createfresh,appropriateidentifiersforthem\.

Figure 7:Canonicalization Prompt \- Orphans
## Appendix 0\.CCuration details

To maintain structural purity, an active curation agent validates the strict equivalence of every raw skill against its broader canonical grouping\. As shown in the first prompt, skills that fail this check are rejected with a specific granular reason \(e\.g\.,DOMAIN\_SPECIALIZATIONorNOT\_A\_SKILL\) and assigned a suggested preferred label, routing them to the ’Orphan’ queue for future iteration\. A secondary rewriting prompt is then applied to the surviving baseline entity to ensure the final labels are stripped of typographical noise and irrelevant job titles\.

[⬇](data:text/plain;base64,WW91IGFyZSBhICoqU2tpbGxzIENsdXN0ZXIgUmV2aWV3ZXIqKiBmb3IgYSBmcmVlbGFuY2UgbWFya2V0cGxhY2UuCllvdXIgdGFzayBpcyB0byBhbmFseXplIGEgbGlzdCBvZiByYXcgc2tpbGxzIGFuZCBkZXRlcm1pbmUgaWYgdGhleSBzaG91bGQgYmUgbWFwcGVkIHRvIHRoZSBlc3RhYmxpc2hlZCBQcmVmZXJyZWQgTmFtZXMgZm9yIHNlYXJjaCBpbmRleGluZy4KCi0tLQoKIyMjIERJUkVDVElWRVMKKipNQVhJTVVNIElOQ0xVU0lPTiBSVUxFOioqIFdlIGFyZSBidWlsZGluZyBhIHNlYXJjaCBlbmdpbmUgaW5kZXguIFlvdSBtdXN0IGVyciBvbiB0aGUgc2lkZSBvZiBJTkNMVVNJT04gKGBpc19lcXVpdmFsZW50OiB0cnVlYCkuCkEgc2tpbGwgaXMgKiplcXVpdmFsZW50KiogaWYgYSBjbGllbnQgc2VhcmNoaW5nIGZvciB0aGUgUHJlZmVycmVkIE5hbWUgd291bGQgcmVhc29uYWJseSB3YW50IHRvIGhpcmUgYSBmcmVlbGFuY2VyIHdobyB3cm90ZSB0aGlzIHJhdyB0ZXh0LgoKMS4gKipXaGF0IHRvIEtFRVAgKGBpc19lcXVpdmFsZW50OiB0cnVlYCwgYHJlamVjdGlvbl9yZWFzb246IG51bGxgKToqKgogICAtICoqVHJhbnNsYXRpb25zIChDUklUSUNBTCkqKjogRGlyZWN0IHRyYW5zbGF0aW9ucyBvZiB0aGUgY29yZSBjb25jZXB0IGluIEVuZ2xpc2gsIEZyZW5jaCwgU3BhbmlzaCwgR2VybWFuLCBvciBEdXRjaCBNVVNUIGJlIGtlcHQgdG9nZXRoZXIgaW4gdGhlIHNhbWUgY2x1c3RlciAoZS5nLiwgS0VFUCAiU29jaWFsIG1lZGlhIiBvciAiUmVkZXMgc29jaWFsZXMiIGluc2lkZSB0aGUgIlLDqXNlYXV4IHNvY2lhdXgiIGNsdXN0ZXIpLgogICAtICoqU3lub255bXMgJiBQYXJhcGhyYXNlcyoqOiBWYXJpYXRpb25zIGluIHBocmFzaW5nLgogICAtICoqVHlwb3MgJiBUcnVuY2F0aW9ucyoqOiBNaXNzcGVsbGluZ3Mgb3IgY3V0LW9mZiB3b3JkcyAoZS5nLiwgInBob3Rvc2hvIiBmb3IgUGhvdG9zaG9wLCAiaW5kZXMiIGZvciBJbkRlc2lnbikuCiAgIC0gKipQcm9maWNpZW5jeSBNb2RpZmllcnMqKjogKGUuZy4sICJFeHBlcnQgaW4iLCAiTWHDrnRyaXNlIGRlIiwgIlNlbmlvciIpLgogICAtICoqUm9sZSBNb2RpZmllcnMgJiBKb2IgVGl0bGVzKio6IChlLmcuLCAiQ29uc3VsdGFudCBTRU8iLCAiRnJlZWxhbmNlciIsICJEYXRhIEFuYWx5c3QiLCAiUHJvamVjdCBNYW5hZ2VyIikuCiAgIC0gKipWZXJzaW9ucyoqOiAoZS5nLiwgIlB5dGhvbiAzIiBmb3IgIlB5dGhvbiIpLgoKMi4gKipXaGF0IHRvIFJFTU9WRSAoYGlzX2VxdWl2YWxlbnQ6IGZhbHNlYCkgYW5kIENhdGVnb3JpemUgKGByZWplY3Rpb25fcmVhc29uYCk6KioKICAgLSBgRE9NQUlOX1NQRUNJQUxJWkFUSU9OYDogQSBzcGVjaWZpYyBkb21haW4gb3IgaW5kdXN0cnkgdGhhdCByZXF1aXJlcyBkaXN0aW5jdCB0ZWNobmljYWwga25vd2xlZGdlLiBFdmVuIGlmIGl0IGNvbnRhaW5zIHRoZSByb290IHdvcmQsIGl0IG11c3QgYmUgcmVqZWN0ZWQgc28gaXQgY2FuIGZvcm0gaXRzIG93biBjbHVzdGVyLiAoZS5nLiwgUkVNT1ZFICJXZWIgUHJvamVjdCBNYW5hZ2VtZW50IiwgIklUIFByb2plY3QgTWFuYWdlbWVudCIsIG9yICJBZ2lsZSBQcm9qZWN0IE1hbmFnZW1lbnQiIGZyb20gdGhlIGdlbmVyYWwgIlByb2plY3QgTWFuYWdlbWVudCIgY2x1c3RlcikuCiAgIC0gYE5PVF9BX1NLSUxMYDogQ29tcGxldGVseSB1bnJlbGF0ZWQgdG8gcHJvZmVzc2lvbmFsIHdvcmsgKGUuZy4sICJIYXJkIHdvcmtlciIsICJBdmFpbGFibGUiLCAiSSBhbSBmYXN0IikuCiAgIC0gYFNFTUFOVElDX01JU01BVENIYDogQSBjb21wbGV0ZWx5IGRpZmZlcmVudCB0ZWNobmljYWwgb3IgcHJvZmVzc2lvbmFsIGNvbmNlcHQuCiAgIC0gYEFNQklHVU9VU2A6IFRvbyB2YWd1ZSB0byBtYXAgdG8gYW55dGhpbmcgKGUuZy4sICJJVCIsICJDb25zdWx0aW5nIiwgIk1hbmFnZW1lbnQiIC0gd2hlbiBzdGFuZGluZyBhbG9uZSkuCiAgIC0gYERJU1RJTkNUX1RPT0xfT1JfTEFOR1VBR0VgOiBJdCBpcyBhIGNvbXBsZXRlbHkgZGlmZmVyZW50IHNvZnR3YXJlL2xhbmd1YWdlIHRoYXQgZGVzZXJ2ZXMgaXRzIG93biBjbHVzdGVyIChlLmcuLCByZW1vdmluZyAiSmF2YSIgZnJvbSBhICJKYXZhU2NyaXB0IiBjbHVzdGVyKS4KCiAgIyMjIE1VTFRJTElOR1VBTCBIQU5ETElORwogIC0gU2tpbGxzIG1heSBiZSB3cml0dGVuIGluIEZyZW5jaCwgU3BhbmlzaCwgR2VybWFuLCBEdXRjaCwgb3IgRW5nbGlzaC4KICAtIFdoZW4gY29tcGFyaW5nIHRvIFByZWZlcnJlZCBOYW1lcywgY29uc2lkZXIgY3Jvc3MtbGFuZ3VhZ2UgZXF1aXZhbGVuY2U6CiAgICAtICJHZXN0aW9uIGRlIHByb2pldCIgKEZSKSA9ICJQcm9qZWN0IE1hbmFnZW1lbnQiIChFTikKICAgIC0gIlByb2pla3RtYW5hZ2VtZW50IiAoREUpID0gIlByb2plY3QgTWFuYWdlbWVudCIgKEVOKQogIC0gUHJvZmljaWVuY3kgbW9kaWZpZXJzIHRyYW5zbGF0ZSBhY3Jvc3MgbGFuZ3VhZ2VzOgogICAgLSAiRXhwZXJ0IGVuIiAoRlIpLCAiRXhwZXJ0byBlbiIgKEVTKSwgIkV4cGVydCBpbiIgKEVOKSwgIkV4cGVydGUgaW4iIChERSkgYXJlIGFsbCBlcXVpdmFsZW50CgojIyMgUkVKRUNUSU9OIFBST1RPQ09MCklmICdpc19lcXVpdmFsZW50JyBpcyBGQUxTRSwgeW91IE1VU1Q6CjEuIFNlbGVjdCBhICdyZWplY3Rpb25fcmVhc29uJyAoZS5nLiwgRE9NQUlOX1NQRUNJQUxJWkFUSU9OKS4KMi4gV2hlbiByZWplY3RpbmcgYSBza2lsbCwgeW91IE1VU1QgcHJvdmlkZSBhIGBzdWdnZXN0ZWRfcHJlZl9sYWJlbGAuIFRoaXMgbGFiZWwgTVVTVCBBTFdBWVMgYmUgd3JpdHRlbiBpbiBzdGFuZGFyZCBFbmdsaXNoLCByZWdhcmRsZXNzIG9mIHRoZSBsYW5ndWFnZSBvZiB0aGUgcmF3IHRleHQgKGUuZy4sIGlmIHRoZSByYXcgdGV4dCBpcyAnYXVkaXQgUkgnLCB0aGUgc3VnZ2VzdGVkIGxhYmVsIG11c3QgYmUgJ0hSIEF1ZGl0JykuCgotLS0KCiMjIyBGT1JNQVRUSU5HIFJVTEVTCi0gWW91IE1VU1QgZXZhbHVhdGUgRVZFUlkgc2tpbGwgcHJvdmlkZWQgaW4gdGhlIElucHV0IERhdGEuCgotLS0KCiMjIyBJTlBVVCBEQVRBCioqUHJlZmVycmVkIE5hbWVzIGZvciBDb250ZXh0OioqCntqc29uLmR1bXBzKHByZWZfbGFiZWxzLCBpbmRlbnQ9MiwgZW5zdXJlX2FzY2lpPUZhbHNlKX0KCioqU2tpbGxzIHRvIFJldmlldzoqKgp7anNvbi5kdW1wcyhza2lsbHMsIGluZGVudD0yLCBlbnN1cmVfYXNjaWk9RmFsc2UpfQ==)Youarea\*\*SkillsClusterReviewer\*\*forafreelancemarketplace\.YourtaskistoanalyzealistofrawskillsanddetermineiftheyshouldbemappedtotheestablishedPreferredNamesforsearchindexing\.\-\-\-\#\#\#DIRECTIVES\*\*MAXIMUMINCLUSIONRULE:\*\*Wearebuildingasearchengineindex\.YoumusterronthesideofINCLUSION\(‘is\_equivalent:true‘\)\.Askillis\*\*equivalent\*\*ifaclientsearchingforthePreferredNamewouldreasonablywanttohireafreelancerwhowrotethisrawtext\.1\.\*\*WhattoKEEP\(‘is\_equivalent:true‘,‘rejection\_reason:null‘\):\*\*\-\*\*Translations\(CRITICAL\)\*\*:DirecttranslationsofthecoreconceptinEnglish,French,Spanish,German,orDutchMUSTbekepttogetherinthesamecluster\(e\.g\.,KEEP"Socialmedia"or"Redessociales"insidethe"Réseauxsociaux"cluster\)\.\-\*\*Synonyms&Paraphrases\*\*:Variationsinphrasing\.\-\*\*Typos&Truncations\*\*:Misspellingsorcut\-offwords\(e\.g\.,"photosho"forPhotoshop,"indes"forInDesign\)\.\-\*\*ProficiencyModifiers\*\*:\(e\.g\.,"Expertin","Ma^itrisede","Senior"\)\.\-\*\*RoleModifiers&JobTitles\*\*:\(e\.g\.,"ConsultantSEO","Freelancer","DataAnalyst","ProjectManager"\)\.\-\*\*Versions\*\*:\(e\.g\.,"Python3"for"Python"\)\.2\.\*\*WhattoREMOVE\(‘is\_equivalent:false‘\)andCategorize\(‘rejection\_reason‘\):\*\*\-‘DOMAIN\_SPECIALIZATION‘:Aspecificdomainorindustrythatrequiresdistincttechnicalknowledge\.Evenifitcontainstherootword,itmustberejectedsoitcanformitsowncluster\.\(e\.g\.,REMOVE"WebProjectManagement","ITProjectManagement",or"AgileProjectManagement"fromthegeneral"ProjectManagement"cluster\)\.\-‘NOT\_A\_SKILL‘:Completelyunrelatedtoprofessionalwork\(e\.g\.,"Hardworker","Available","Iamfast"\)\.\-‘SEMANTIC\_MISMATCH‘:Acompletelydifferenttechnicalorprofessionalconcept\.\-‘AMBIGUOUS‘:Toovaguetomaptoanything\(e\.g\.,"IT","Consulting","Management"\-whenstandingalone\)\.\-‘DISTINCT\_TOOL\_OR\_LANGUAGE‘:Itisacompletelydifferentsoftware/languagethatdeservesitsowncluster\(e\.g\.,removing"Java"froma"JavaScript"cluster\)\.\#\#\#MULTILINGUALHANDLING\-SkillsmaybewritteninFrench,Spanish,German,Dutch,orEnglish\.\-WhencomparingtoPreferredNames,considercross\-languageequivalence:\-"Gestiondeprojet"\(FR\)="ProjectManagement"\(EN\)\-"Projektmanagement"\(DE\)="ProjectManagement"\(EN\)\-Proficiencymodifierstranslateacrosslanguages:\-"Experten"\(FR\),"Expertoen"\(ES\),"Expertin"\(EN\),"Expertein"\(DE\)areallequivalent\#\#\#REJECTIONPROTOCOLIf’is\_equivalent’isFALSE,youMUST:1\.Selecta’rejection\_reason’\(e\.g\.,DOMAIN\_SPECIALIZATION\)\.2\.Whenrejectingaskill,youMUSTprovidea‘suggested\_pref\_label‘\.ThislabelMUSTALWAYSbewritteninstandardEnglish,regardlessofthelanguageoftherawtext\(e\.g\.,iftherawtextis’auditRH’,thesuggestedlabelmustbe’HRAudit’\)\.\-\-\-\#\#\#FORMATTINGRULES\-YouMUSTevaluateEVERYskillprovidedintheInputData\.\-\-\-\#\#\#INPUTDATA\*\*PreferredNamesforContext:\*\*\{json\.dumps\(pref\_labels,indent=2,ensure\_ascii=False\)\}\*\*SkillstoReview:\*\*\{json\.dumps\(skills,indent=2,ensure\_ascii=False\)\}

Figure 8:Curation Prompt[⬇](data:text/plain;base64,WW91IGFyZSBhbiBleHBlcnQgKipNdWx0aWxpbmd1YWwgVGF4b25vbXkgTGluZ3Vpc3QgYW5kIEhSIFNwZWNpYWxpc3QqKi4KUmV2aWV3IGFuZCByZWZpbmUgdGhlIGN1cnJlbnQgIlByZWZlcnJlZCBMYWJlbHMiIGZvciBwcm9mZXNzaW9uYWwgc2tpbGwgY2x1c3RlcnMgYWNyb3NzIHRoZXNlIGxhbmd1YWdlczoge2xhbmd1YWdlc30uCgojIyMgRElSRUNUSVZFUwpBY3QgYXMgYSBsb2NhbCBIUiBzcGVjaWFsaXN0IGluIHRoZSBjb3VudHJ5IG9mIHRoZSB0YXJnZXQgbGFuZ3VhZ2UuCi0gKipQcm9mZXNzaW9uYWwgQ29udGV4dDoqKiBJZ25vcmUgbGl0ZXJhbCBkaWN0aW9uYXJ5IHRyYW5zbGF0aW9ucy4gVXNlIHN0YW5kYXJkIENWL0pvYiBEZXNjcmlwdGlvbiB0ZXJtcyAoZS5nLiwgIkNvbnN1bHRpbmciIGluc3RlYWQgb2YgIkNvbnNlaWwiIGluIEZyZW5jaCkuCi0gKipBbmdsaWNpc21zOioqIElmIHRoZSBFbmdsaXNoIHRlcm0gaXMgdGhlIGRvbWluYW50IHN0YW5kYXJkIChjb21tb24gaW4gVGVjaC9CdXNpbmVzcyksIGtlZXAgaXQhIChlLmcuLCAiQ2xvdWQgQ29tcHV0aW5nIiBmb3IgR2VybWFuKS4KLSAqKkdyYW1tYXIgU3RhbmRhcmQ6KiogVXNlIFNpbmd1bGFyIE5vdW5zIG9yIEdlcnVuZHMgcmVwcmVzZW50aW5nIHRoZSBjYXBhYmlsaXR5LCBub3QgdGhlIGFjdGlvbiAoZS5nLiwgIk1hbmFnZW1lbnQiIGluc3RlYWQgb2YgIlRvIE1hbmFnZSIpLgoKIyMjIFBVUklGSUNBVElPTiBSVUxFIChDUklUSUNBTCkKVGhlICJBbHRlcm5hdGUgU2tpbGxzIiBsaXN0IHByb3ZpZGVkIGJlbG93IGNvbnRhaW5zIG1lc3N5IHNlYXJjaCBkYXRhLiBJdCBpbmNsdWRlcyBqb2IgdGl0bGVzIChlLmcuLCAiQ29uc3VsdGFudCBTRU8iKSwgc2VuaW9yaXR5IGxldmVscyAoZS5nLiwgIkV4cGVydCIpLCBhbmQgdHlwb3MuCioqWW91ciBnZW5lcmF0ZWQgbGFiZWxzIE1VU1QgU1RSSVAgT1VUIGFsbCBvZiB0aGlzIG5vaXNlLioqIERvIG5vdCBpbmNsdWRlIHdvcmRzIGxpa2UgIkNvbnN1bHRhbnQiLCAiRXhwZXJ0Iiwgb3IgIkZyZWVsYW5jZSIgaW4geW91ciBmaW5hbCBsYWJlbHMuIEV4dHJhY3QgT05MWSB0aGUgcHVyZSwgY2Fub25pY2FsIHVuZGVybHlpbmcgc2tpbGwuCgojIyMgSU5QVVQgREFUQQoqKkN1cnJlbnQgUHJlZmVycmVkIExhYmVsczoqKgp7anNvbi5kdW1wcyhwcmVmX2xhYmVscywgaW5kZW50PTIpfQoKKipBbHRlcm5hdGUgU2tpbGxzIGluIENsdXN0ZXIgKGZvciBjb250ZXh0KToqKgp7anNvbi5kdW1wcyhhbHRlcm5hdGVfc2tpbGxzLCBpbmRlbnQ9Mil9)Youareanexpert\*\*MultilingualTaxonomyLinguistandHRSpecialist\*\*\.Reviewandrefinethecurrent"PreferredLabels"forprofessionalskillclustersacrosstheselanguages:\{languages\}\.\#\#\#DIRECTIVESActasalocalHRspecialistinthecountryofthetargetlanguage\.\-\*\*ProfessionalContext:\*\*Ignoreliteraldictionarytranslations\.UsestandardCV/JobDescriptionterms\(e\.g\.,"Consulting"insteadof"Conseil"inFrench\)\.\-\*\*Anglicisms:\*\*IftheEnglishtermisthedominantstandard\(commoninTech/Business\),keepit\!\(e\.g\.,"CloudComputing"forGerman\)\.\-\*\*GrammarStandard:\*\*UseSingularNounsorGerundsrepresentingthecapability,nottheaction\(e\.g\.,"Management"insteadof"ToManage"\)\.\#\#\#PURIFICATIONRULE\(CRITICAL\)The"AlternateSkills"listprovidedbelowcontainsmessysearchdata\.Itincludesjobtitles\(e\.g\.,"ConsultantSEO"\),senioritylevels\(e\.g\.,"Expert"\),andtypos\.\*\*YourgeneratedlabelsMUSTSTRIPOUTallofthisnoise\.\*\*Donotincludewordslike"Consultant","Expert",or"Freelance"inyourfinallabels\.ExtractONLYthepure,canonicalunderlyingskill\.\#\#\#INPUTDATA\*\*CurrentPreferredLabels:\*\*\{json\.dumps\(pref\_labels,indent=2\)\}\*\*AlternateSkillsinCluster\(forcontext\):\*\*\{json\.dumps\(alternate\_skills,indent=2\)\}

Figure 9:Rewriting Prompt
## Appendix 0\.DConsolidation details

To ensure global structural consistency across incremental batch runs, the consolidation phase evaluates candidate pairs flagged by lightweight lexical heuristics to merge overlapping sub\-graphs\. The prompt below instructs the model to determine whether to merge or keep entities separate based on semantic overlap, multilingual equivalency, and versioning constraints, outputting a definitive decision and a surviving composite ID\.

[⬇](data:text/plain;base64,WW91IGFyZSBhbiBleHBlcnQgKipUYXhvbm9teSBBcmNoaXRlY3QqKi4KVHdvIHNraWxsIGNsdXN0ZXJzIGhhdmUgYmVlbiBmbGFnZ2VkIGFzIHBvdGVudGlhbCBkdXBsaWNhdGVzLiBZb3VyIGdvYWwgaXMgdG8gZGV0ZXJtaW5lIGlmIHRoZXkgcmVwcmVzZW50IHRoZSBleGFjdCBzYW1lIHByb2Zlc3Npb25hbCBjb25jZXB0LgoKIyMjIERFQ0lTSU9OIENSSVRFUklBCioqTUVSR0UqKiBpZiBBTlkgb2YgdGhlc2UgY29uZGl0aW9ucyBhcHBseToKLSAqKkV4YWN0IFN5bm9ueW1zOioqICJOb2RlSlMiIGFuZCAiTm9kZS5qcyIsICJSZWFjdCIgYW5kICJSZWFjdEpTIgotICoqVHlwbyBWYXJpYW50czoqKiAiSmF2c2NyaXB0IiBhbmQgIkphdmFTY3JpcHQiLCAiTWFuYWdtZW50IiBhbmQgIk1hbmFnZW1lbnQiCi0gKipNdWx0aWxpbmd1YWwgRXF1aXZhbGVudHM6KiogIkdlc3Rpb24gZGUgcHJvamV0IiAoRlIpIGFuZCAiUHJvamVjdCBNYW5hZ2VtZW50IiAoRU4pIHdpdGggc2FtZSBzY29wZQotICoqTm90YXRpb24gRGlmZmVyZW5jZXM6KiogIkMrKyIgYW5kICJDIHBsdXMgcGx1cyIsICJDIyIgYW5kICJDIFNoYXJwIgoKKipLRUVQX1NFUEFSQVRFKiogaWYgQU5ZIG9mIHRoZXNlIGNvbmRpdGlvbnMgYXBwbHk6Ci0gKipTcGVjaWFsaXphdGlvbiB2cyBHZW5lcmFsOioqICJUZWNobmljYWwgUHJvamVjdCBNYW5hZ2VtZW50IiB2cyAiUHJvamVjdCBNYW5hZ2VtZW50IiAob25lIGlzIG5hcnJvd2VyKQotICoqVmVyc2lvbiBSZXF1aXJpbmcgRGlmZmVyZW50IEV4cGVydGlzZToqKiAiQW5ndWxhciAxLngiIHZzICJBbmd1bGFyIDIrIiAoZnVuZGFtZW50YWxseSBkaWZmZXJlbnQgZnJhbWV3b3JrcykKLSAqKkNvbXBsZXRlbHkgRGlmZmVyZW50IERvbWFpbnM6KiogIkphdmEiIChsYW5ndWFnZSkgdnMgIkphdmFTY3JpcHQiICh1bnJlbGF0ZWQgbGFuZ3VhZ2UpCi0gKipNZXRob2RvbG9neSB2cyBUb29sOioqICJBZ2lsZSBQcm9qZWN0IE1hbmFnZW1lbnQiIHZzICJQcm9qZWN0IE1hbmFnZW1lbnQiIChvbmUgYWRkcyBtZXRob2RvbG9neSkKLSAqKkNvbnRleHQgU3BlY2lmaWNpdHk6KiogIlJlbW90ZSBQcm9qZWN0IE1hbmFnZW1lbnQiIHZzICJQcm9qZWN0IE1hbmFnZW1lbnQiIChvbmUgYWRkcyBjb250ZXh0KQoKIyMjIE1VTFRJTElOR1VBTCBWQUxJREFUSU9OCi0gSWYgY2x1c3RlcnMgaGF2ZSBkaWZmZXJlbnQgcHJpbWFyeSBsYWJlbHMgYnV0IHNhbWUgUUlEcywgdGhleSdyZSBsaWtlbHkgbXVsdGlsaW5ndWFsIHZhcmlhbnRzIHRoZW4gKipNRVJHRSoqCi0gSWYgY2x1c3RlcnMgaGF2ZSBzaW1pbGFyIEVuZ2xpc2ggbGFiZWxzIGJ1dCBkaWZmZXJlbnQgdmFsaWRhdGVkIHNraWxscywgdGhleSdyZSBsaWtlbHkgZGlzdGluY3QgdGhlbiAqKktFRVBfU0VQQVJBVEUqKgoKIyMjIFNVUlZJVk9SIFJVTEVTCklmIHlvdSBNRVJHRSwgeW91IG11c3QgcGljayB0aGUgYmVzdCBJRCB0byBrZWVwOgoxLiAqKk9mZmljaWFsIG92ZXIgU3ludGhldGljOioqIFByaW9yaXRpemUgV2lraWRhdGEgUUlEcyAoZS5nLiwgUTEyMykgb3ZlciBPcnBoYW4gSURzIChlLmcuLCBPUlBIQU5fSVRFUjFfLi4uKS4KMi4gKipMb3dlciBRSUQgTnVtYmVyOioqIElmIGJvdGggYXJlIFdpa2lkYXRhIFFJRHMsIHByZWZlciB0aGUgb25lIHdpdGggdGhlIGxvd2VyIG51bWJlciAob2xkZXIsIG1vcmUgZXN0YWJsaXNoZWQgYnJvYWRlciBjb25jZXB0KS4KMy4gKipTdGFiaWxpdHk6KiogSWYgbWVyZ2luZyBhbiBPcnBoYW4gaW50byBhIFFJRCwgdGhlIFFJRCBNVVNUIGJlIHRoZSBgc3Vydml2aW5nX2NvbXBvc2l0ZV9pZGAuCgojIyMgSU5QVVQgREFUQQoKKipDTFVTVEVSIEE6KioKLSBJRDoge3BhaXJfZGF0YVsnY2x1c3Rlcl8xX2lkJ119Ci0gUmVmaW5lZCBMYWJlbHMgKEZpbHRlcmVkKToKe2NsdXN0ZXJfMV9sYWJlbHN9Ci0gVmFsaWRhdGVkIFNraWxscyBpbiB0aGlzIENsdXN0ZXI6IHtqc29uLmR1bXBzKHBhaXJfZGF0YVsnY2x1c3Rlcl8xX3NraWxscyddLCBlbnN1cmVfYXNjaWk9RmFsc2UpfQoKKipDTFVTVEVSIEI6KioKLSBJRDoge3BhaXJfZGF0YVsnY2x1c3Rlcl8yX2lkJ119Ci0gUmVmaW5lZCBMYWJlbHMgKEZpbHRlcmVkKToKe2NsdXN0ZXJfMl9sYWJlbHN9Ci0gVmFsaWRhdGVkIFNraWxscyBpbiB0aGlzIENsdXN0ZXI6IHtqc29uLmR1bXBzKHBhaXJfZGF0YVsnY2x1c3Rlcl8yX3NraWxscyddLCBlbnN1cmVfYXNjaWk9RmFsc2UpfQoKLS0tCkJhc2VkIG9uIHRoZSBsYWJlbHMgYW5kIHNraWxscyBwcm92aWRlZCwgZGVjaWRlIGlmIHRoZXNlIHR3byBjbHVzdGVycyBzaG91bGQgYmUgbWVyZ2VkLg==)Youareanexpert\*\*TaxonomyArchitect\*\*\.Twoskillclustershavebeenflaggedaspotentialduplicates\.Yourgoalistodetermineiftheyrepresenttheexactsameprofessionalconcept\.\#\#\#DECISIONCRITERIA\*\*MERGE\*\*ifANYoftheseconditionsapply:\-\*\*ExactSynonyms:\*\*"NodeJS"and"Node\.js","React"and"ReactJS"\-\*\*TypoVariants:\*\*"Javscript"and"JavaScript","Managment"and"Management"\-\*\*MultilingualEquivalents:\*\*"Gestiondeprojet"\(FR\)and"ProjectManagement"\(EN\)withsamescope\-\*\*NotationDifferences:\*\*"C\+\+"and"Cplusplus","C\#"and"CSharp"\*\*KEEP\_SEPARATE\*\*ifANYoftheseconditionsapply:\-\*\*SpecializationvsGeneral:\*\*"TechnicalProjectManagement"vs"ProjectManagement"\(oneisnarrower\)\-\*\*VersionRequiringDifferentExpertise:\*\*"Angular1\.x"vs"Angular2\+"\(fundamentallydifferentframeworks\)\-\*\*CompletelyDifferentDomains:\*\*"Java"\(language\)vs"JavaScript"\(unrelatedlanguage\)\-\*\*MethodologyvsTool:\*\*"AgileProjectManagement"vs"ProjectManagement"\(oneaddsmethodology\)\-\*\*ContextSpecificity:\*\*"RemoteProjectManagement"vs"ProjectManagement"\(oneaddscontext\)\#\#\#MULTILINGUALVALIDATION\-IfclustershavedifferentprimarylabelsbutsameQIDs,they’relikelymultilingualvariantsthen\*\*MERGE\*\*\-IfclustershavesimilarEnglishlabelsbutdifferentvalidatedskills,they’relikelydistinctthen\*\*KEEP\_SEPARATE\*\*\#\#\#SURVIVORRULESIfyouMERGE,youmustpickthebestIDtokeep:1\.\*\*OfficialoverSynthetic:\*\*PrioritizeWikidataQIDs\(e\.g\.,Q123\)overOrphanIDs\(e\.g\.,ORPHAN\_ITER1\_\.\.\.\)\.2\.\*\*LowerQIDNumber:\*\*IfbothareWikidataQIDs,prefertheonewiththelowernumber\(older,moreestablishedbroaderconcept\)\.3\.\*\*Stability:\*\*IfmerginganOrphanintoaQID,theQIDMUSTbethe‘surviving\_composite\_id‘\.\#\#\#INPUTDATA\*\*CLUSTERA:\*\*\-ID:\{pair\_data\[’cluster\_1\_id’\]\}\-RefinedLabels\(Filtered\):\{cluster\_1\_labels\}\-ValidatedSkillsinthisCluster:\{json\.dumps\(pair\_data\[’cluster\_1\_skills’\],ensure\_ascii=False\)\}\*\*CLUSTERB:\*\*\-ID:\{pair\_data\[’cluster\_2\_id’\]\}\-RefinedLabels\(Filtered\):\{cluster\_2\_labels\}\-ValidatedSkillsinthisCluster:\{json\.dumps\(pair\_data\[’cluster\_2\_skills’\],ensure\_ascii=False\)\}\-\-\-Basedonthelabelsandskillsprovided,decideifthesetwoclustersshouldbemerged\.

Figure 10:Consolidation Prompt

相似文章

一种面向异构文档知识图谱构建的本体引导、去重感知抽取层

arXiv cs.AI

本文介绍了一个生产级抽取层,利用本地托管的经过调优的Qwen大语言模型,将异构文档转换为与本体对齐的知识图谱,并采用本体引导的提示、多阶段去重和基于嵌入的解析。在情报语料上的评估表明,搜索召回率从约70%提升到95%,且没有错误合并。