Generating training datasets for legal chatbots in Korean
Summary
This paper presents a method for generating large-scale, labeled training datasets for legal chatbots in Korean using Local Grammar Graphs, achieving 91% F1-score with a DIET classifier.
View Cached Full Text
Cached at: 05/11/26, 07:03 AM
# Generating training datasets for legal chatbots in Korean Source: [https://arxiv.org/abs/2605.07432](https://arxiv.org/abs/2605.07432) [View PDF](https://arxiv.org/pdf/2605.07432) > Abstract:Chatbots are robots that can communicate with humans using text or voice signals\. Legal chatbots improve access to justice, since legal representation and legal advice by lawyers come with a high cost that excludes disadvantaged and vulnerable people\. However, capturing the diversity of actual user input in datasets for deep\-learning dialog systems \(chatbots\) is a technical challenge\. Diversity requires large volumes of data, which must also be labelled in order to classify the user's intent, while the cost of labelling datasets increases with volume\. Instead of labelling large volumes of authentic data from users, our approach consists in jointly generating large volumes of utterances and high\-quality labels\. The generator of labelled datasets is based on language resources that take the form of local grammar graphs \(LGG\), which capture and generalize the vocabulary and local syntax observed by linguists in text\. The LGGs associate labels to the utterances according to a domain\-specific classification system\. We tested this approach by implementing LIGA, a legal chatbot in Korean\. The chatbot answers users' conversational queries on legal situations by providing information on similar legal cases, made publicly available by the Korean government\. We generated labelled utterances from the LGGs with the aid of the open\-source Unitex platform\. This process produced 700 million utterances\. We trained a DIET classifier on a dataset made of these utterances, and the trained model reached 91% f1\-score performance\. We implemented a chatbot called LIGA, which uses the results of the model to select a link to a web page that documents similar legal cases\. ## Submission history From: Eric Laporte \[[view email](https://arxiv.org/show-email/9dabeb0b/2605.07432)\] **\[v1\]**Fri, 8 May 2026 08:32:56 UTC \(825 KB\)
Similar Articles
Korean Culture into LLM Alignment: Toward Cultural Coherence
This paper proposes a dataset generation pipeline to align large language models with Korean cultural norms using DPO fine-tuning, improving cultural safety without degrading general performance.
KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms
The paper introduces KoNeoBench, a curated benchmark dataset for evaluating large language models' understanding of Korean neologisms, based on 1,785 entries from online news since 2020. It reveals limitations in current LLMs in handling recent lexical changes and specific Korean linguistic properties.
Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services
This paper introduces DasanCallDial, a large-scale Korean benchmark dataset for dialogue-level ASR error correction, and proposes the DCSC framework, achieving state-of-the-art performance in text-only post-editing.
GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus
This paper releases GLAN-QnA-KR, a 303,581-row Korean instruction-QA corpus generated via a seedless taxonomy-driven pipeline using Phi-3.5-MoE-instruct, with near-zero duplication and low benchmark contamination.
KG2Cypher: Data-Centric Pipeline for Building Enterprise Text-to-Cypher Systems
KG2Cypher presents a data-centric pipeline for building enterprise text-to-Cypher systems from existing knowledge graphs. It uses LLMs to generate natural language question-Cypher pairs, validated by an LLM judge and human review, and achieves significant performance improvements on Korean enterprise datasets with LoRA-based fine-tuning.