GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus
Summary
This paper releases GLAN-QnA-KR, a 303,581-row Korean instruction-QA corpus generated via a seedless taxonomy-driven pipeline using Phi-3.5-MoE-instruct, with near-zero duplication and low benchmark contamination.
View Cached Full Text
Cached at: 07/24/26, 05:16 AM
# GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus
Source: [https://arxiv.org/html/2607.20443](https://arxiv.org/html/2607.20443)
###### Abstract
We releaseGLAN\-QnA\-KR, a 303,581\-row openly redistributable Korean instruction\-QA corpus produced via the seedless taxonomy\-driven GLAN synthesis pipeline\[[7](https://arxiv.org/html/2607.20443#bib.bib8)\]with Microsoft’s Phi\-3\.5\-MoE\-instruct as the producer model \(generation: 2024\-12; release: 2024\-12; licence: OpenRAIL\)\. The corpus spans a flat taxonomy of 1,084 English\-labelled disciplines paired with Korean question/answer text, a 100–900 difficulty scale, and a median of 313 question characters and 1,098 answer characters per record\. Two properties are atypical for synthetic instruction data at this scale: \(i\) exact duplicate questions number only 1 in 303,581 rows and character\-trigram near\-duplicate clusters at Jaccard≥0\.9\\geq 0\.9number zero in a 5,000\-sample probe, and \(ii\) a two\-layer contamination audit against KMMLU, KoBEST \(five sub\-tasks\), and HAE\-RAE\-Bench shows a maximum test\-vs\-corpus question\-level character\-trigram Jaccard of0\.163with zero test items at Jaccard≥0\.7\\geq 0\.7, and a maximum multilingual\-E5 cosine of0\.901with a single test item at cosine≥0\.90\\geq 0\.90and zero at≥0\.95\\geq 0\.95, across 20,000 sampled GLAN questions and seven evaluation sets\. At the time of release, this is, to our knowledge, the largest single\-pipeline synthetic Korean instruction corpus verifiable on the Hugging Face Hub and the only Korean≥\\geq100k\-row corpus built under a seedless taxonomy\-driven protocol\. This note documents the generation protocol, corpus statistics, the contamination audit, and the licensing boundary in a form suitable for downstream citation\.
## 1Introduction
Open Korean instruction\-tuning data remains dominated by two patterns: \(i\)translationof English SFT corpora \(KoAlpaca, KULLM\-v2, KOR\-OpenOrca\-Platypus, KOpen\-platypus\), which inherits upstream style and benchmark contamination risk; and \(ii\)aggregationof heterogeneous native sources into a single corpus \(KoCommercial\), which is Korean\-native but not a single synthesis pipeline\. A third pattern—*seedless taxonomy\-driven synthesis*in the style of Li et al\.\[[7](https://arxiv.org/html/2607.20443#bib.bib8)\]—has not been publicly instantiated at scale for Korean\. We releaseGLAN\-QnA\-KR, a 303,581\-row Korean realisation of that third pattern, together with a reproducible statistics manifest, a two\-layer contamination audit against public Korean benchmarks \(max character\-trigram Jaccard=0\.163=0\.163with zero test items at≥0\.7\\geq 0\.7; max multilingual\-E5 cosine=0\.901=0\.901with a single item at≥0\.90\\geq 0\.90and zero at≥0\.95\\geq 0\.95\), and an explicit licensing boundary\.
The dataset was generated in December 2024 using Microsoft’s Phi\-3\.5\-MoE\-instruct\[[9](https://arxiv.org/html/2607.20443#bib.bib10)\]as the producer model over a taxonomy of 1,084 English\-labelled disciplines inherited from the GLAN protocol, producing Korean question/answer pairs along a 100–900 difficulty scale\. The pipeline is conceptually compatible with the open reference implementation atAzure/synthetic\-qa\-generation/glan\-instruct\[[8](https://arxiv.org/html/2607.20443#bib.bib2)\], which defaults to\-\-language Korean; however, that repository is sample code and*not*the exact generator used for this corpus, and the two should not be read as identical\. The Azure reference defaults target GPT\-4o; the present corpus uses Phi\-3\.5\-MoE\-instruct, and the prompt templates, taxonomy expansion order, and post\-filtering differ\.
### What this note is and is not\.
We do*not*claim a new synthesis method—GLAN is upstream\[[7](https://arxiv.org/html/2607.20443#bib.bib8)\]—and we do*not*claim state\-of\-the\-art Korean instruction quality, since no standardised Korean SFT quality benchmark exists at the time of writing\. We do document, with measured numbers and a reproducible audit: \(i\) corpus schema and 1,084\-way taxonomy distribution, \(ii\) length and difficulty statistics, \(iii\) dedup cleanliness, \(iv\) train/eval contamination against KMMLU\[[11](https://arxiv.org/html/2607.20443#bib.bib12)\], KoBEST\[[6](https://arxiv.org/html/2607.20443#bib.bib7)\], and HAE\-RAE\-Bench\[[12](https://arxiv.org/html/2607.20443#bib.bib11)\], and \(v\) the licence and intended\-use boundary\. The goal is to provide a citable reference for downstream Korean SFT work that uses this corpus, and to publish a contamination\-audit protocol that future open Korean synthetic corpora can reuse\.
### Contributions\.
\(1\) A documented data statement forGLAN\-QnA\-KR, including generation protocol, producer model, taxonomy scope, and licence\. \(2\) Corpus statistics \(length, difficulty, 1,084\-way discipline distribution, duplicate audit\)\. \(3\) A two\-layer contamination audit against the three standard public Korean benchmarks \(KMMLU, KoBEST’s five sub\-tasks, HAE\-RAE\-Bench\)—character\-trigram Jaccard and multilingual\-E5 cosine similarity\. The maximum observed per\-test\-question character\-trigram Jaccard is0\.1630\.163with zero test items at≥0\.8\\geq 0\.8; the maximum observed multilingual\-E5 cosine is0\.9010\.901with a single test item at≥0\.90\\geq 0\.90and zero at≥0\.95\\geq 0\.95\. \(4\) A positioning table against every open Korean instruction corpus≥\\geq10k rows \(Table[1](https://arxiv.org/html/2607.20443#S2.T1)\)\. \(5\) An explicit licensing and intended\-use discussion that separates the OpenRAIL corpus licence from the producer\-model terms\.
## 2Related Resources
### GLAN and the synthetic\-instruction genealogy\.
Synthetic instruction data spans a genealogy from Self\-Instruct\[[14](https://arxiv.org/html/2607.20443#bib.bib13)\]\(seed\-expansion from 175 hand\-written instructions\) and WizardLM / Evol\-Instruct\[[16](https://arxiv.org/html/2607.20443#bib.bib16)\]\(seed\-rewriting along deepen / concretise / complicate axes\) to document\-grounded variants \(OSS\-Instruct\[[15](https://arxiv.org/html/2607.20443#bib.bib15)\], Orca\-2\[[10](https://arxiv.org/html/2607.20443#bib.bib9)\]\) and, separately, to taxonomy\-driven seedless pipelines such as GLAN\[[7](https://arxiv.org/html/2607.20443#bib.bib8)\]and the Phi\-series textbook\-synthesis reports\[[4](https://arxiv.org/html/2607.20443#bib.bib6),[1](https://arxiv.org/html/2607.20443#bib.bib1)\]\. The distinguishing axis for our work is*seedless*: GLAN samples discipline→\\tosyllabus→\\toclass session→\\tohomework Q/A from a curated taxonomy with no input instruction corpus, so the diversity and contamination profile of the output are governed by the taxonomy itself and the producer model, not by an upstream instruction pool\.
### Open Korean instruction corpora\.
Table[1](https://arxiv.org/html/2607.20443#S2.T1)compares the main open Korean instruction corpora≥\\geq10k rows\. We group them along three axes:scale,origin\(Korean\-native synthesis vs\. translation of an English corpus vs\. aggregation\), andlicence tier\(permissive vs\. NC vs\. unstated\)\. Only three Korean instruction datasets exceed 100k rows: KULLM\-v2 \(152,630; translated\), KoCommercial \(175,454; aggregated native\), andGLAN\-QnA\-KR\(303,581; seedless synthetic\)\. Among these,GLAN\-QnA\-KRis the only single\-pipeline synthetic corpus, the only one using Phi\-3\.5\-MoE as producer, and the only one released under OpenRAIL without an NC rider\.
Table 1:Open Korean instruction\-tuning corpora≥\\geq10k rows on the Hugging Face Hub\. “Origin” distinguishes Korean\-native synthesis from translation of an English corpus and from aggregation of heterogeneous sources\. Sizes and licences verified against each dataset card on 2026\-05\-12\. KoVicuna and KoAlpaca\-v1\.0 are omitted because their cards were not publicly retrievable at audit time\.
### Korean benchmarks \(for contamination audit, not training\)\.
KMMLU\[[11](https://arxiv.org/html/2607.20443#bib.bib12)\]\(35,030 Korean exam questions across 45 subjects, CC\-BY\-ND\-4\.0\), KoBEST\[[6](https://arxiv.org/html/2607.20443#bib.bib7)\]\(five\-task Korean NLU benchmark\), and HAE\-RAE\-Bench\[[12](https://arxiv.org/html/2607.20443#bib.bib11)\]are the three standard open Korean evaluation sets that downstream SFT consumers ofGLAN\-QnA\-KRare likely to evaluate on\. No dedicated Korean\-native Math\-KR synthetic reasoning corpus at≥\\geq10k scale is publicly documented as of May 2026\.
### Positioning\.
GLAN\-QnA\-KRsits in the niche of*large\-scale, openly redistributable, seedless\-synthetic*Korean instruction corpora\. It is not a replacement for translation\-based corpora when English\-English pair recovery is desired; it is a reference open Korean synthetic instruction set for downstream SFT work that wants a single\-pipeline, contamination\-audited, permissively licensed training signal\.
## 3Generation Protocol
### Producer model\.
microsoft/Phi\-3\.5\-MoE\-instruct\[[9](https://arxiv.org/html/2607.20443#bib.bib10)\], accessed via Azure OpenAI\-compatible inference\. No fine\-tuning of the producer was performed; sampling uses the producer’s default instruction\-following posture\.
### Taxonomy\.
The taxonomy follows the GLAN protocol\[[7](https://arxiv.org/html/2607.20443#bib.bib8)\]: a flat list of 1,084 English\-labelled disciplines that span STEM \(e\.g\.,Machine Learning,Discrete Mathematics,Cell Biology\), applied health \(e\.g\.,Pharmacology,Health Informatics,Biostatistics\), business and management \(Project Management,Customer Relationship Management,Event Management\), and education and cognition \(Cognitive Development,Assessment and Evaluation,Educational Technology\)\. No single discipline dominates the corpus: the top\-weighted subject \(Grant Writing and Fundraising\) has 826 rows \(0\.27% of the corpus\), and the top 20 subjects together cover only 3\.6% of the corpus\.
### Generation\.
For each discipline, the producer samples a class session and then generates a \(question, answer\) pair conditioned on the session topic and a difficulty level drawn from\{100,200,…,900\}\\\{100,200,\\ldots,900\\\}\. Crucially, this is a*seedless*pipeline: no input instruction corpus is used, so the corpus’s topical coverage and style are determined by the taxonomy and by the producer’s learned Korean distribution, not by an upstream seed pool\. The prompt templates elicit Korean\-language question and answer text while retaining English\-languagesubjectandsubtopicstags, producing a cross\-lingual label / body structure\.
### Post\-filtering\.
The released corpus contains only records with non\-emptyquestionandanswerfields; no further content\-level filtering \(length caps, perplexity filtering, or moderation filtering\) was applied\. Duplicate removal was implicit in the sampling procedure rather than applied as a separate pass, which is visible in the extreme dedup cleanliness reported in §[4](https://arxiv.org/html/2607.20443#S4)\.
### Reference implementation\.
An open reference implementation of GLAN for Korean is available atAzure/synthetic\-qa\-generation/glan\-instruct\[[8](https://arxiv.org/html/2607.20443#bib.bib2)\]\. That repository defaults to\-\-language Koreanandgpt\-4oand provides a compatible starting point; it is*not*the exact generator used for this corpus\. The key protocol differences from the Azure reference are: \(i\)GLAN\-QnA\-KRuses Phi\-3\.5\-MoE\-instruct as the producer, and \(ii\) prompt templates and taxonomy traversal order differ\. We release the released corpus without also releasing the proprietary prompt templates\.
## 4Corpus Statistics
Table 2:Length statistics per text field, in characters\.question,subtopics, andanswerstatistics are computed on the first 100k rows;subjectuses the full 303,581 rows\. Distributions are long\-tailed in bothquestionandanswer\.### Schema and split\.
A singletrainsplit of 303,581 records with five columns:question\(Korean string\),subject\(English discipline label, 1,084 distinct values\),level\(integer in\{100,200,…,900\}\\\{100,200,\\ldots,900\\\}\),subtopics\(English comma\-separated sub\-tags\), andanswer\(Korean string\)\. The mismatched language between the English taxonomy labels \(subject,subtopics\) and the Korean content \(question,answer\) is a deliberate cross\-lingual design choice inherited from the GLAN taxonomy, not a cleanup bug; downstream trainers must decide whether to translate or mask the English tags during SFT\.
### Length and compression\.
Table[2](https://arxiv.org/html/2607.20443#S4.T2)reports per\-field character length statistics\. Questions have a median of 313 characters and answers a median of 1,098 characters—a median per\-record answer\-to\-question length ratio of roughly3\.5×3\.5\\times\. Both distributions are long\-tailed, with 99\.3% of questions fitting in the first histogram bucket \(31–1,206 characters\) and a hard tail reaching 11,787 characters; answers reach 177,014 characters at the extreme\. A few essay\-length answers will dominate any naive length\-packed SFT sampler, and we recommend downstream users clip or bucket theanswerfield explicitly\.
### Difficulty distribution\.
Thelevelfield is bell\-shaped and mid\-biased: levels 300–500 together cover 46% of the corpus \(47,890 / 50,642 / 42,232\), while level 100 is smallest at 19,408 rows\. Usinglevelfor difficulty\-aware training requires no additional re\-weighting for mid\-band coverage; extreme\-level SFT ablations should either stratify or up\-sample the 100 / 900 tails\.
### Discipline distribution\.
The 1,084\-waysubjectdistribution is*flat*: the top\-weighted subject \(Grant Writing and Fundraising\) contributes only 0\.27% of records, and the top 20 subjects together contribute 3\.6%\. No macro\-domain column is provided in the raw schema; grouping by STEM / applied health / business / humanities requires an external discipline\-to\-macro\-domain map, which we recommend downstream users construct once and reuse\.
### Dedup cleanliness\.
Two findings are atypical for synthetic instruction data at 303k scale\. First, exact duplicatequestiontext numbers only1 in 303,581 rows\(1 duplicate group, 1 redundant row\)\. Second, in a random 5,000\-question probe, character\-trigram Jaccard near\-duplicate clusters at threshold 0\.9 number0\. The taxonomy\-driven seedless generation procedure therefore produces essentially non\-overlapping prompts at this scale, without an explicit deduplication pass\.
### Language composition\.
In a random 200\-question probe: 159 questions \(79\.5%\) are Korean\-heavy \(≥\\geq80% Hangul characters\), 41 \(20\.5%\) are mixed Korean and Latin \(typically Korean body plus English technical acronyms; e\.g\., “IWRM” inside a Korean paragraph\), and 0 are English\-dominant\.
## 5Contamination Audit
Synthetic instruction corpora are a recurring source of train/eval contamination: if a taxonomy\-driven producer happens to reproduce eval\-set questions at generation time, downstream SFT on the corpus can silently inflate eval scores\[[2](https://arxiv.org/html/2607.20443#bib.bib3),[3](https://arxiv.org/html/2607.20443#bib.bib4)\]\. We release a audit against the three standard public Korean benchmarks that downstream SFT consumers ofGLAN\-QnA\-KRare likely to evaluate on\.
### Protocol\.
For each benchmark test split we \(i\) extract the question text, \(ii\) compute the character\-trigram Jaccard similarity between each test question and each of 20,000 GLAN questions sampled uniformly from the 303k corpus, and \(iii\) report the median and mean of the*per\-test\-question maximum*Jaccard, together with the counts and rates of test questions with maximum Jaccard≥0\.7\\geq 0\.7and≥0\.8\\geq 0\.8\. Character\-trigram Jaccard is language\-agnostic and robust to whitespace and punctuation normalisation; thresholds are chosen to match community practice for Korean\-specific near\-duplicate auditing given the absence of a canonical KR protocol\[[3](https://arxiv.org/html/2607.20443#bib.bib4)\]\.
### Results\.
Table[3](https://arxiv.org/html/2607.20443#S5.T3)reports the audit\. Across all seven evaluation sets, the per\-test\-question maximum character\-trigram Jaccard against a 20,000\-question GLAN sample is small: median Jaccard ranges from0\.0000\.000\(KoBEST/wic\) to0\.0530\.053\(HAE\-RAE\-Bench\), and the single largest observation in the entire audit is0\.1630\.163\(one KMMLU question against its nearest GLAN neighbour\)\. Zero test items on any benchmark reach the operational near\-duplicate threshold of Jaccard≥0\.8\\geq 0\.8\.
Table 3:Contamination audit: per\-test\-question maximum character\-trigram Jaccard similarity against a 20,000\-question sample ofGLAN\-QnA\-KR\(3\-grams over whitespace\-normalised Korean text\)\. Lower is better\. J≥\\geq0\.8 is the operational near\-duplicate threshold; zero test items across all seven benchmarks reach it, and the single largest observation anywhere is 0\.163\. KMMLU is sampled across eight representative subjects \(Math, Computer\-Science, Biology, Chemistry, Economics, Korean\-History, Law, Psychology\); KoBEST audits the natural prompt field per task; HAE\-RAE\-Bench aggregates seven subsets \(general knowledge, history, loan words, standard nomenclature, rare words, date understanding, reading comprehension\)\.
### Interpretation\.
GLAN\-QnA\-KRis substantially contamination\-free against these seven standard Korean evaluation sets at the0\.80\.8near\-duplicate threshold, and downstream Korean SFT scores on KMMLU, KoBEST, and HAE\-RAE\-Bench can be read without a contamination\-adjustment for training onGLAN\-QnA\-KR\. We attribute the cleanness to the seedless taxonomy\-driven generation protocol: because no input instructions are used as seeds, the corpus’s textual forms are driven by Phi\-3\.5\-MoE’s own Korean generation distribution conditioned on English\-language discipline labels, which has minimal lexical overlap with the Korean human\-written prompts surfaced by these benchmarks\.
### Semantic paraphrase audit\.
Character\-trigram Jaccard detects lexical near\-duplicates but not semantic paraphrase\. We complement the lexical audit with a multilingual\-embedding pass usingintfloat/multilingual\-e5\-base\[[13](https://arxiv.org/html/2607.20443#bib.bib14)\]: we embed both sides \(each test question and each of the 20,000 GLAN questions\) with L2\-normalised E5 embeddings using the requiredquery/passageprefixes, and for each test question we report the maximum cosine against the 20,000\-question bank \(Table[4](https://arxiv.org/html/2607.20443#S5.T4)\)\. The maximum cosine observed anywhere in the seven eval sets is0\.901\(a single KoBEST/boolq paragraph\); exactly one test item reaches cosine≥0\.90\\geq 0\.90, and zero reach≥0\.95\\geq 0\.95\. The cos≥0\.85\\geq 0\.85rate ranges from 2\.0% \(KoBEST/copa, KoBEST/wic\) to 38\.0% \(HAE\-RAE\-Bench\), but this should be read against the baseline similarity of multilingual\-e5\-base across unrelated Korean text: the*median*max\-cosine is already in the 0\.81–0\.84 band on every benchmark, meaning cos≈0\.85\\approx 0\.85is roughly one noise\-band above baseline and is not by itself evidence of contamination; the≥0\.90\\geq 0\.90threshold, where only one item sits across all seven audits, is the operationally meaningful line\.
Table 4:Semantic paraphrase audit: per\-test\-question maximum cosine similarity ofintfloat/multilingual\-e5\-baseembeddings against the same 20,000\-question GLAN sample used in Table[3](https://arxiv.org/html/2607.20443#S5.T3)\. The median max\-cosine is already 0\.81–0\.84 across all benchmarks, reflecting the baseline similarity level of multilingual\-e5 on unrelated Korean prose; cos≥0\.95\\geq 0\.95is the operationally meaningful near\-duplicate threshold \(zero items across all seven audits\), with a single boundary hit at cos=0\.901=0\.901on KoBEST/boolq\.A producer model that generatessemanticallyequivalent test questions in different Korean wording would raise the cos≥0\.90\\geq 0\.90or≥0\.95\\geq 0\.95rate; neither is raised here, which, together with the lexical audit, supports readingGLAN\-QnA\-KRas substantially contamination\-free against the seven audited Korean evaluation splits\. The one cos=0\.901=0\.901KoBEST/boolq item is a known boundary case that downstream consumers can exclude from any contamination\-sensitive eval\.
## 6License and Intended Use
### Corpus artifact\.
The packaging \(schema, the released 303,581 rows, the statistics manifest produced by this work\) is released underOpenRAILas declared on the Hugging Face dataset card\[[5](https://arxiv.org/html/2607.20443#bib.bib5)\]\. OpenRAIL permits research and commercial use subject to its downstream\-use restrictions \(harm\-oriented use cases are prohibited\)\.
### Producer\-model terms\.
Thequestionandanswerfields are outputs of Microsoft’s Phi\-3\.5\-MoE\-instruct\[[9](https://arxiv.org/html/2607.20443#bib.bib10)\]and are subject to that model’s terms of use at generation time\. Downstream consumers should verify the current Phi\-3\.5 licence and acceptable\-use policy when training on these outputs, particularly for commercial deployment\.
### Recommended disclosure for downstream work\.
Papers that fine\-tune onGLAN\-QnA\-KRshould \(i\) cite this note and the upstream GLAN paper\[[7](https://arxiv.org/html/2607.20443#bib.bib8)\], \(ii\) cite Phi\-3\.5\-MoE\[[9](https://arxiv.org/html/2607.20443#bib.bib10)\]as the producer model, and \(iii\) disclose that the training signal is synthetic and produced by a proprietary instruction\-tuned model, rather than by human annotators\.
## 7Reproducibility
The dataset is hosted at
[https://huggingface\.co/datasets/daekeun\-ml/GLAN\-qna\-kr\-300k](https://huggingface.co/datasets/daekeun-ml/GLAN-qna-kr-300k)\[[5](https://arxiv.org/html/2607.20443#bib.bib5)\]
The corpus statistics reported in §[4](https://arxiv.org/html/2607.20443#S4)and the two\-layer contamination audit in §[5](https://arxiv.org/html/2607.20443#S5)are reproducible from the released dataset alone: the lexical audit uses character\-trigram Jaccard over whitespace\-normalised questions against a 20,000\-question random sample ofGLAN\-QnA\-KR, and the semantic audit uses L2\-normalisedintfloat/multilingual\-e5\-baseembeddings with the requiredquery/passageprefixes at the same sample size\. We do not release the author’s internal audit scripts as a separate artifact; the procedures above are described in enough detail to be re\-implemented\. A compatible reference generator for GLAN\-for\-Korean, not used for this specific corpus, is available atAzure/synthetic\-qa\-generation\[[8](https://arxiv.org/html/2607.20443#bib.bib2)\]\.
## 8Limitations
1. 1\.Synthetic\-only signal\.Everyquestion/answerpair is produced by a single instruction\-tuned model\. The corpus inherits that model’s Korean style, factual biases, and failure modes\. Downstream SFT on this corpus exclusively will reproduce those biases; mixing with human\-written Korean instruction data is recommended\.
2. 2\.Single\-producer dependence\.Phi\-3\.5\-MoE\-instruct is the sole producer\. A checkpoint\-level change in Microsoft’s released Phi\-3\.5 would make exact re\-generation of this corpus impossible\.
3. 3\.No SFT\-gain evaluation\.This note does*not*report a controlled downstream SFT evaluation \(e\.g\., fine\-tuning a small Korean LLM onGLAN\-QnA\-KRand reporting KMMLU / KoBEST / HAE\-RAE delta\)\. Anecdotal downstream use by the author on small Korean LLMs suggests the corpus is*usable*as a mid\-scale Korean SFT signal; we do not quantify this here\.
4. 4\.Cross\-lingual label/body split\.Englishsubjectandsubtopicstags with Koreanquestion/answerbodies complicates certain prompt formats and pretrained\-tokenizer comparisons\.
5. 5\.Contamination audit encoder dependence\.Both the lexical \(character\-trigram Jaccard\) and semantic \(multilingual\-e5\-basecosine\) audits detect surface\-form and embedding\-level matches, respectively\. A producer model that paraphrases test questions into a form that our encoder embeds far from the original, or that reproduces only rare factual surface forms \(named entities, numeric answers\) not captured by question\-level similarity, would still evade detection\. A stronger guarantee would require answer\-level semantic audit and an encoder\-ensemble check; we leave that as future work\.
6. 6\.Taxonomy inheritance\.The 1,084\-discipline taxonomy is inherited from Microsoft’s GLAN\[[7](https://arxiv.org/html/2607.20443#bib.bib8)\]and is English\-centric; it over\-represents Western academic disciplines \(e\.g\.,Grant Writing and Fundraising,Leadership in Hospitality\) and under\-represents Korean\-specific domains \(Korean history, Korean law, Korean literature\)\.
7. 7\.Answer\-length tail\.The 177,014\-character answer tail will dominate naive length\-packed SFT samplers; downstream users must clip or bucket theanswerfield\.
8. 8\.Licence downstream dependence\.The corpus is OpenRAIL, but the producer\-model terms govern commercial deployment of SFT checkpoints trained on the corpus\. Downstream users must verify the current Phi\-3\.5 terms of use\.
## Acknowledgments
We thank the authors of GLAN\[[7](https://arxiv.org/html/2607.20443#bib.bib8)\]for the upstream synthesis protocol, the maintainers of the open Korean benchmarks \(KMMLU, KoBEST, HAE\-RAE\-Bench\) whose existence makes the contamination audit possible, and the maintainers of the Hugging Face Hub for continuing to host Korean open\-source NLP artifacts\.
## References
- \[1\]M\. Abdinet al\.\(2024\)Phi\-3 technical report: a highly capable language model locally on your phone\.arXiv preprint arXiv:2404\.14219\.Cited by:[§2](https://arxiv.org/html/2607.20443#S2.SS0.SSS0.Px1.p1.3)\.
- \[2\]T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.\(2020\)Language models are few\-shot learners\.NeurIPS\.Cited by:[§5](https://arxiv.org/html/2607.20443#S5.p1.1)\.
- \[3\]Y\. Elazar, A\. Bhagia, I\. Magnusson, A\. Ravichander, D\. Schwenk, A\. Suhr, P\. Walsh, D\. Groeneveld, L\. Soldaini, S\. Singh, H\. Hajishirzi, N\. A\. Smith, and J\. Dodge\(2023\)What’s in my big data?\.arXiv preprint arXiv:2310\.20707\.Cited by:[§5](https://arxiv.org/html/2607.20443#S5.SS0.SSS0.Px1.p1.2),[§5](https://arxiv.org/html/2607.20443#S5.p1.1)\.
- \[4\]S\. Gunasekar, Y\. Zhang, J\. Aneja, C\. C\. T\. Mendes, A\. Del Giorno, S\. Gopi, M\. Javaheripi, P\. Kauffmann, G\. de Rosa, O\. Saarikivi,et al\.\(2023\)Textbooks are all you need\.arXiv preprint arXiv:2306\.11644\.Cited by:[§2](https://arxiv.org/html/2607.20443#S2.SS0.SSS0.Px1.p1.3)\.
- \[5\]D\. Kim\(2024\)GLAN\-qna\-kr\-300k dataset card\.Note:[https://huggingface\.co/datasets/daekeun\-ml/GLAN\-qna\-kr\-300k](https://huggingface.co/datasets/daekeun-ml/GLAN-qna-kr-300k)Cited by:[§6](https://arxiv.org/html/2607.20443#S6.SS0.SSS0.Px1.p1.1),[§7](https://arxiv.org/html/2607.20443#S7.p1.2)\.
- \[6\]D\. Kim, M\. Jang, D\. S\. Kwon, and E\. Davis\(2022\)KoBEST: Korean balanced evaluation of significant tasks\.InCOLING,Cited by:[§1](https://arxiv.org/html/2607.20443#S1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2607.20443#S2.SS0.SSS0.Px3.p1.1)\.
- \[7\]H\. Li, Q\. Dong, Z\. Chen, H\. Wang, H\. Meng, J\. Liu, J\. Lv, Y\. Liu, H\. Xu, Z\. Sun,et al\.\(2024\)Synthetic data \(almost\) from scratch: generalized instruction tuning for language models\.arXiv preprint arXiv:2402\.13064\.Cited by:[§1](https://arxiv.org/html/2607.20443#S1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2607.20443#S1.p1.5),[§2](https://arxiv.org/html/2607.20443#S2.SS0.SSS0.Px1.p1.3),[§3](https://arxiv.org/html/2607.20443#S3.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2607.20443#S6.SS0.SSS0.Px3.p1.1),[item 6](https://arxiv.org/html/2607.20443#S8.I1.i6.p1.1),[Acknowledgments](https://arxiv.org/html/2607.20443#Sx1.p1.1)\.
- \[8\]Microsoft Azure\(2024\)Azure/synthetic\-qa\-generation: generate synthetic qnas from real\-world data\.Note:[https://github\.com/Azure/synthetic\-qa\-generation](https://github.com/Azure/synthetic-qa-generation)Cited by:[§1](https://arxiv.org/html/2607.20443#S1.p2.1),[§3](https://arxiv.org/html/2607.20443#S3.SS0.SSS0.Px5.p1.1),[§7](https://arxiv.org/html/2607.20443#S7.p1.3)\.
- \[9\]Microsoft\(2024\)Phi\-3\.5\-MoE\-instruct model card\.Note:[https://huggingface\.co/microsoft/Phi\-3\.5\-MoE\-instruct](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct)Cited by:[§1](https://arxiv.org/html/2607.20443#S1.p2.1),[§3](https://arxiv.org/html/2607.20443#S3.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2607.20443#S6.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2607.20443#S6.SS0.SSS0.Px3.p1.1)\.
- \[10\]A\. Mitra, L\. Del Corro, S\. Mahajan, A\. Codas, C\. Simoes, S\. Agarwal, X\. Chen, A\. Razdaibiedina, E\. Jones, K\. Aggarwal,et al\.\(2023\)Orca 2: teaching small language models how to reason\.arXiv preprint arXiv:2311\.11045\.Cited by:[§2](https://arxiv.org/html/2607.20443#S2.SS0.SSS0.Px1.p1.3)\.
- \[11\]G\. Son, H\. Lee, S\. Kim, S\. Kim, N\. Muennighoff, T\. Choi, C\. Park, K\. M\. Yoo, and S\. Biderman\(2024\)KMMLU: measuring massive multitask language understanding in Korean\.arXiv preprint arXiv:2402\.11548\.Cited by:[§1](https://arxiv.org/html/2607.20443#S1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2607.20443#S2.SS0.SSS0.Px3.p1.1)\.
- \[12\]G\. Son, H\. Lee, S\. Kim, J\. Kim, N\. Muennighoff, T\. Choi, C\. Park, K\. M\. Yoo, and S\. Biderman\(2023\)HAE\-RAE bench: evaluation of Korean knowledge in language models\.arXiv preprint arXiv:2309\.02706\.Cited by:[§1](https://arxiv.org/html/2607.20443#S1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2607.20443#S2.SS0.SSS0.Px3.p1.1)\.
- \[13\]L\. Wang, N\. Yang, X\. Huang, L\. Yang, R\. Majumder, and F\. Wei\(2024\)Multilingual E5 text embeddings: a technical report\.arXiv preprint arXiv:2402\.05672\.Cited by:[§5](https://arxiv.org/html/2607.20443#S5.SS0.SSS0.Px4.p1.5)\.
- \[14\]Y\. Wang, Y\. Kordi, S\. Mishra, A\. Liu, N\. A\. Smith, D\. Khashabi, and H\. Hajishirzi\(2022\)Self\-instruct: aligning language models with self\-generated instructions\.arXiv preprint arXiv:2212\.10560\.Cited by:[§2](https://arxiv.org/html/2607.20443#S2.SS0.SSS0.Px1.p1.3)\.
- \[15\]Y\. Wei, Z\. Wang, J\. Liu, Y\. Ding, and L\. Zhang\(2023\)Magicoder: source code is all you need\.arXiv preprint arXiv:2312\.02120\.Cited by:[§2](https://arxiv.org/html/2607.20443#S2.SS0.SSS0.Px1.p1.3)\.
- \[16\]C\. Xu, Q\. Sun, K\. Zheng, X\. Geng, P\. Zhao, J\. Feng, C\. Tao, and D\. Jiang\(2023\)WizardLM: empowering large language models to follow complex instructions\.arXiv preprint arXiv:2304\.12244\.Cited by:[§2](https://arxiv.org/html/2607.20443#S2.SS0.SSS0.Px1.p1.3)\.Similar Articles
Annotating Korean adnominal ending constructions in corpus data: Beyond relative-clause identification
This paper proposes a corpus-based typology for Korean adnominal ending constructions and implements a construction-sensitive annotation layer for the KLUE dependency treebank, finding that relative-clause-like uses account for only 39.4% of analyzed instances.
Generating training datasets for legal chatbots in Korean
This paper presents a method for generating large-scale, labeled training datasets for legal chatbots in Korean using Local Grammar Graphs, achieving 91% F1-score with a DIET classifier.
KG2Cypher: Data-Centric Pipeline for Building Enterprise Text-to-Cypher Systems
KG2Cypher presents a data-centric pipeline for building enterprise text-to-Cypher systems from existing knowledge graphs. It uses LLMs to generate natural language question-Cypher pairs, validated by an LLM judge and human review, and achieves significant performance improvements on Korean enterprise datasets with LoRA-based fine-tuning.
KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms
The paper introduces KoNeoBench, a curated benchmark dataset for evaluating large language models' understanding of Korean neologisms, based on 1,785 entries from online news since 2020. It reveals limitations in current LLMs in handling recent lexical changes and specific Korean linguistic properties.
AI_LectureNote: A Retrospective Pilot Study of a Post-ASR Workflow for English-Script Rendering and Semantic Drift in Korean-English Medical Lectures
A retrospective pilot study evaluating AI_LectureNote, a post-ASR workflow for Korean-English medical lectures, showing that while it improves English-script rendering, it introduces semantic drift and polarity failures in the transcripts.