HKJudge: A Legal Discourse-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule
Summary
HKJudge is the first sentence-level expert-annotated legal discourse corpus for Hong Kong criminal judgments, featuring a two-tier discourse schema and benchmark evaluations of BERT-based and LLM models.
View Cached Full Text
Cached at: 06/08/26, 09:20 AM
# A Legal Discourse-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule
Source: [https://arxiv.org/html/2606.06679](https://arxiv.org/html/2606.06679)
Xi Xuan1, Wenxin Zhang2, Yufei Zhou1, King\-kui Sin1, Chunyu Kit1 1City University of Hong Kong2University of Chinese Academy of Sciences \{xixuan3, ctckit\}@cityu\.edu\.hk
###### Abstract
Court judgments are central to legal practice and jurisprudence, yet discourse analysis of Hong Kong judgments has received limited attention, owing largely to the absence of expert\-annotated corpora\. We introduce the Hong Kong Judgment Discourse Dataset \(HKJudge\), the first sentence\-level expert\-annotated legal discourse corpus\. HKJudge includes criminal judgments across all five levels of HK’s court hierarchy, comprising∼\\sim290k sentences and∼\\sim6\.5 million tokens, fully annotated by legal linguistics experts\. We design a two\-tier discourse schema that captures what facts a court finds, how it reasons, and what it rules\. At the sentence level, each sentence is assigned one of 26 rhetorical roles\. At the span level, sentences are further annotated with three sentencing elements \(charge, imprisonment term, fine\)\. Ten legal linguistics annotators produced the annotations with an inter\-annotator agreement ofκ=0\.8\\kappa=0\.8\. We formulate two tasks on HKJudge, termed rhetorical role classification and legal element extraction, and provide the first benchmark evaluation of four BERT\-based models, two open\-source LLMs under zero\-shot and fine\-tuning settings, and four commercial LLMs on both tasks\. Our work demonstrates the value of sentence\-level discourse annotation for modeling the structure of HK judgments and provides a rich data foundation for future work on legal judgment prediction\. The HKJudge dataset and code are available at111https://github\.com/xuanxixi/HKJudge\.
HKJudge: A Legal Discourse\-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule
Xi Xuan1, Wenxin Zhang2, Yufei Zhou1, King\-kui Sin1, Chunyu Kit11City University of Hong Kong2University of Chinese Academy of Sciences\{xixuan3, ctckit\}@cityu\.edu\.hk
## 1Introduction
Court judgments are among the most important legal genres for the legal profession, in both legal practice and jurisprudence\(Chenget al\.,[2008a](https://arxiv.org/html/2606.06679#bib.bib65)\)\. They are performative speech acts whose fundamental function is to adjudicate, and they simultaneously serve declaratory, justificatory, and legitimating purposes within a single document\(Maley,[2014](https://arxiv.org/html/2606.06679#bib.bib66)\)\. Making judgments tractable for downstream NLP tasks, including legal search\(Werner,[1981](https://arxiv.org/html/2606.06679#bib.bib70); Moet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib73)\), case analysis\(Liet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib4)\), and legal judgment prediction \(LJP\)\(Gillman,[2001](https://arxiv.org/html/2606.06679#bib.bib74); Heet al\.,[2024](https://arxiv.org/html/2606.06679#bib.bib72); Dancy and Zalnieriute,[2026](https://arxiv.org/html/2606.06679#bib.bib75)\), requires modeling knowledge of the generic structures of legal documents\. Such structural modeling reduces search space, facilitates the identification of rhetorical segments, thereby facilitating the working efficiency of court judgments\(Saravanan,[2010](https://arxiv.org/html/2606.06679#bib.bib15); Hanet al\.,[2018](https://arxiv.org/html/2606.06679#bib.bib12); Kalamkaret al\.,[2022](https://arxiv.org/html/2606.06679#bib.bib41)\)\.
While substantial progress has been made in modeling court judgment structure for Indian\(Ghosh,[2019](https://arxiv.org/html/2606.06679#bib.bib37); Kalamkaret al\.,[2022](https://arxiv.org/html/2606.06679#bib.bib41); Nigamet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib40)\), European\(Rosas,[2007](https://arxiv.org/html/2606.06679#bib.bib109); Held and Habernal,[2026](https://arxiv.org/html/2606.06679#bib.bib110)\), United States\(Robinson,[2013](https://arxiv.org/html/2606.06679#bib.bib111); Williams,[2022](https://arxiv.org/html/2606.06679#bib.bib112); Shuet al\.,[2024](https://arxiv.org/html/2606.06679#bib.bib113)\)and Chinese mainland\(Xiaoet al\.,[2018](https://arxiv.org/html/2606.06679#bib.bib39); Liebmanet al\.,[2020](https://arxiv.org/html/2606.06679#bib.bib13); Feiet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib14)\)jurisdictions, comparable resources for Hong Kong \(HK\) case law remain scarce\. HK court judgments, produced within a bilingual common\-law jurisdiction with its own appellate hierarchy\(Chenget al\.,[2008a](https://arxiv.org/html/2606.06679#bib.bib65); Yu,[2023](https://arxiv.org/html/2606.06679#bib.bib23); Xuan and others,[2024b](https://arxiv.org/html/2606.06679#bib.bib104)\), follow drafting conventions and discourse structures that differ from those of the corpora discussed above, particularly in citation practice\(Cheng,[2015](https://arxiv.org/html/2606.06679#bib.bib31)\), sentencing discourse\(Yu,[2025](https://arxiv.org/html/2606.06679#bib.bib30); Xuanet al\.,[2026c](https://arxiv.org/html/2606.06679#bib.bib108)\), and bilingual reasoning\(Cheng and He,[2016](https://arxiv.org/html/2606.06679#bib.bib29); Xuan and others,[2025](https://arxiv.org/html/2606.06679#bib.bib105)\)\. Direct transfer of models trained on other jurisdictions is therefore unreliable\.
Figure 1:Overview of the HKJudge annotation process\. Stage 1 \(left\): an anonymized Hong Kong District Court criminal judgment\. Stage 2 \(center\): each sentence is labeled with one of 26 rhetorical roles \(see Appendix[C](https://arxiv.org/html/2606.06679#A3)for full definitions\), grouped into four categories: Fact \(F\), Inference \(I\), Result \(R\), and Others \(O\); twelve representative labels are shown\. Stage 3 \(right\): summary of the HKJudge dataset \(4,000 documents,∼\\sim290k sentences, 26 labels\) and three span\-level element types: Charge, Term, and Fine, extracted from R\-tagged sentences\.Previous research in this domain has highlighted the importance of annotated datasets for training effective models\. However, currently there is no publicly available dataset for the HK JLP task that isfully annotated by legal linguistics experts\. Many existing studies have relied on relatively small annotated datasets with only coarse\-grained, three\-level labels offacts,reasoning, andruling, so that charges and prior records share the same label, and case\-law reasoning is not distinguished from that citing an Ordinance, limiting their effectiveness for LJP systems in real\-world scenarios\. The few existing HK legal dataset resources each have important limitations\. For instance, the Legal\-NLP Dataset of\(Sen,[2023](https://arxiv.org/html/2606.06679#bib.bib10)\)was constructed from HKLII judgments using regular expressions and semantic parsing, without expert annotation of rhetorical structure\. The HKCFA Judgment 97\-22 dataset\(Xuan and Kit,[2026](https://arxiv.org/html/2606.06679#bib.bib11)\)targets legal translation rather than discourse analysis, and covers only Court of Final Appeal judgments\. The LegalHK dataset\(Shiet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib86)\), from the LegalReasoner framework, relies on GPT\-4 to extract structured information from judicial documents withonly partial manual review by judicial experts, excludes appellate cases from the court of appeal and the court of final appeal, and captures only three coarse functional blocks without sentence\-level rhetorical roles\.
We see the need to introduce a unified mode of study that can quickly incorporate new areas and applications of law\. In this work, wedevelop a uniform discourse schema for characterising a HK court judgment\. Discourse analysis, the study of how texts are organized into functional segments above the sentence level\(Gill,[2000](https://arxiv.org/html/2606.06679#bib.bib48); Jotyet al\.,[2019](https://arxiv.org/html/2606.06679#bib.bib47); Gee,[2025](https://arxiv.org/html/2606.06679#bib.bib49)\), has been successfully applied to areas like news events\(Nakshatriet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib46)\), dialogue understanding\(Koet al\.,[2023](https://arxiv.org/html/2606.06679#bib.bib44); Xuanet al\.,[2026a](https://arxiv.org/html/2606.06679#bib.bib107)\), web documents\(Liuet al\.,[2023a](https://arxiv.org/html/2606.06679#bib.bib43)\), legal documents\(Sovranoet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib42)\), and synthetic\-content characterization\(Xuan and others,[2024a](https://arxiv.org/html/2606.06679#bib.bib103); Xuanet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib100); Linet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib99); Liet al\.,[2026](https://arxiv.org/html/2606.06679#bib.bib101); Xuanet al\.,[2026b](https://arxiv.org/html/2606.06679#bib.bib102); Zhang and others,[2026](https://arxiv.org/html/2606.06679#bib.bib106)\)\. In legal domain,\(Sovranoet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib42)\)effectively use discourse analysis for legal question answering, improving state\-of\-the\-art without fine\-tuning or re\-training the language models on the regulations at hand\.
In this work, we develop a legal discourse schema to address this need\. At its core, our schema seeks to answer three questions about each judgment: \(1\)whatfacts the court finds, \(2\)howit reasons, and \(3\)whatit rules\. We show that both pretrained encoders and LLMs struggle to model this schema, whereas legal linguistics experts label it with high inter\-annotator agreement\. In sum, this paper makes three key contributions:
1. 1\.Introducing, Annotating and Modeling a Legal Discourse Schema:We develop a HK legal discourse schema, consisting of 3 span\-level and 26 sentence\-level rhetorical role classes, some of which are shown in Figure[1](https://arxiv.org/html/2606.06679#S1.F1)\. We construct court judgments dataset annotated by legal linguistics experts, with 148,600 spans and 292,240 rhetorical role annotations\. We show that our schema can be labeled with high inter\-annotator agreement\. Additionally, we show LLM models \(zero\-shot and fine\-tuned\) outperform BERT\-based models\.
2. 2\.Web Scraping Public HK Case Law:We build web scrapers and collect over 4,000 judgments spanning 1968–2024 from five Hong Kong courts, namely the Court of Final Appeal, the Court of Appeal, the Court of First Instance, the District Court, and the Magistrates’ Courts\. Hong Kong judgments are subject to HKSAR Government copyright but publicly available for private and academic use\.222[https://www\.judiciary\.hk/en/other\_information/disclaimer\.html](https://www.judiciary.hk/en/other_information/disclaimer.html)\.Our scrapers comply with the access policies of the Judiciary’s Legal Reference System333[https://legalref\.judiciary\.hk/](https://legalref.judiciary.hk/)and use rate\-limited requests\.
3. 3\.Benchmarking BERT\-based and LLM\-based Methods on Court Judgments Annotation:We evaluate four BERT\-based methods, open\-source LLMs, and commercial LLMs \(including GPT\-4, Claude, and Gemini\) under zero\-shot and fine\-tuned settings\. Although fine\-tuning yields substantial gains, all LLMs still fall noticeably short of human expert annotators and commercial LLMs, highlighting the value of expert annotation and pointing to open challenges in legal LLM reasoning\.
## 2A Legal Discourse Schema
Court judgments serve multiple functions, including adjudication, declaration, justification, and legitimation\(Maley,[2014](https://arxiv.org/html/2606.06679#bib.bib66)\)\. Modeling their structure at the discourse level therefore provides an effective entry point into legal reasoning\(Carlsonet al\.,[2003](https://arxiv.org/html/2606.06679#bib.bib50); Prasadet al\.,[2017](https://arxiv.org/html/2606.06679#bib.bib81)\), and supports downstream tasks including legal search\(Moet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib73)\), case summarization\(Liet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib4)\), and legal judgment prediction\(Aletraset al\.,[2016](https://arxiv.org/html/2606.06679#bib.bib88); Maliket al\.,[2021](https://arxiv.org/html/2606.06679#bib.bib82); Shiet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib86); Dancy and Zalnieriute,[2026](https://arxiv.org/html/2606.06679#bib.bib75)\)\. Hong Kong court judgments in particular can be segmented along the rhetorical roles of heading, introduction, facts, analysis, and conclusion\(Chenget al\.,[2008b](https://arxiv.org/html/2606.06679#bib.bib84)\)\.
As shown in Figure[1](https://arxiv.org/html/2606.06679#S1.F1), modeling the different rhetorical roles of a legal doctrine asdiscourse unitsand how they interact can be an effective way of discerning meaning\(Carlsonet al\.,[2003](https://arxiv.org/html/2606.06679#bib.bib50); Prasadet al\.,[2017](https://arxiv.org/html/2606.06679#bib.bib81)\)\. Identifying these parts poses a basic test of a model’s legal reasoning and unlocks practical applications, asHendryckset al\.\([2021](https://arxiv.org/html/2606.06679#bib.bib83)\)demonstrated in contract review\. We accordingly introduce a schema that captures this distinction at two levels, starting with sentence\-level annotations and extending to span\-level extraction of three result elements, termedcharge, the offence of conviction;imprisonment term, the length of custodial sentence; andfine, the monetary penalty\.
### 2\.1Discourse\-level Schema
The 4 discourse functions we identify in our legal discourse schema areFact,Inference,Result, andOther\. The first three functions correspond to the three layers of information carried by every judgment, whereas Other is a residual class covering procedural or formulaic sentences that fall outside the preceding three\. The first three functions, Fact, Inference, and Result, capture how a judgment proceeds from established facts, through the court’s reasoning, to the ruling, inspired by the contrastive analysis of HK court judgment structure done byChenget al\.\([2008b](https://arxiv.org/html/2606.06679#bib.bib84)\)\. We describe each category in turn\.
- •AFact \(F\)sentence typically reports information presented to the court without expressing the court’s own evaluation\. We distinguish 15 sub\-tags by procedural origin and evidentiary status: F0\-charge, F1\-issue, F2\-event, F3\-supplement, F4\-previous\_info, F5\-previous\_record, F6\-argument, F7\-jury, F8\-other, F9\-admission, F10\-assertion, F11\-question, F12\-answer, F13\-objection, and F14\-instructions2jury\. F11–F14 generally originate inside the courtroom, while F2–F5 generally originate outside it\. \(e\.g\.*“The defendant is 24 and has 2 conviction records, which include 2 ‘Theft’ offences, 3 ‘Robbery’ offences and 1 ‘Attempted Robbery’ offence\.”*is tagged F3\-supplement\.\)
- •AnInference \(I\)sentence is one in which the court itself reasons toward its decision\. We distinguish 8 sub\-tags by the authority appealed to: I1\-case\_law, I2\-ordinance, I3\-legislation, I4\-conventional\_practice, I5\-jury, I6\-assertion, I7\-other, and I8\-question\. The boundary with Fact tends to rest on whose voice is speaking, since a party’s argument and a judge’s reasoning can be lexically similar\. \(e\.g\.*“It does not identify a purpose which it thinks would be beneficial and then construe the statute to fit it\.”*is tagged I1\-case\_law\.\)
- •AResult \(R\)sentence states the disposition of the case, under two sub\-tags: R, the final determination, and R\-other, supplementary content attached to it \(clarifications, calculations, statements of consequence\)\. The boundary can be subtle, since explanatory material may intervene between successive operative rulings within a single paragraph\. \(e\.g\.*“For that reason, and in light of the concession made by the appellant in relation to the third respondent, the appeal must be dismissed against all of the respondents\.”*is tagged R\.\)
We give full definitions of the rhetorical role sub\-tags in Appendix[C](https://arxiv.org/html/2606.06679#A3)\. Judgments rendered by the court of appeal and the court of final appeal embed the lower court’s Facts, Inferences, and Results into their own text, tagged F4\-previous\_info\. Some of these sentences retain a secondary discourse function such as I8\-question, and we allow both tags in such cases\.
### 2\.2Span\-Level Schema
We define 3 span\-level element types during our annotation process, applied to sentences tagged Result\. These cover what the court decides, termedchargedenoting the offence of conviction;imprisonment termdenoting the length of custodial sentence; andfinedenoting the monetary penalty\. The three types are mutually exclusive at the span level, with each span identifying one sentencing decision\. Charges and their penalties often appear in the same R sentence, producing spans of more than one type\. We do not annotate spans for other sentencing options such as suspended sentences, community service orders, or disqualification orders, since these surface infrequently in our data; the containing sentence is retained under R at the sentence level\.
## 3Dataset Construction
In this section, we describe how we operationalized the schema discussed in Section[2](https://arxiv.org/html/2606.06679#S2)\. We scrape a dataset of Hong Kong criminal case laws from 1968 to 2024 across five court levels, which we discuss in Section[3\.1](https://arxiv.org/html/2606.06679#S3.SS1)\. We build an annotation framework, described in Section[3\.2](https://arxiv.org/html/2606.06679#S3.SS2), and enlist ten annotators, who collectively annotate∼\\sim290k law sentences\.
### 3\.1Dataset Source and Web Scraping
The corpus contains over 4,000 Hong Kong criminal judgments from 1968 to 2024, spanning all five court levels with criminal jurisdiction\. We scrape these judgments from the Hong Kong Legal Information Institute \(HKLII\),444[https://www\.hklii\.hk/databases](https://www.hklii.hk/databases)a public\-access platform that aggregates court judgments released by the Hong Kong Judiciary\. Raw judgments data has a mix of court of final appeal judgments \(20%\), court of appeal judgments \(20%\), court of first instance judgments \(20%\), district court judgments \(20%\) and magistrates’ courts judgments \(20%\)\. We audit the HKLII output against the Judiciary’s Legal Reference System \(LRS\),555[https://legalref\.judiciary\.hk/](https://legalref.judiciary.hk/)and find judgments that HKLII does not cover, has not updated, or renders as image\-based PDFs with high OCR error rates;666For example, older judgments fromHKMagCon HKLII are stored as scanned PDFs; we extract text directly from the LRS PDF instead\.for these we fall back to LRS PDFs read throughpdfplumber\.777[https://github\.com/jsvine/pdfplumber](https://github.com/jsvine/pdfplumber)
Hong Kong judgments are subject to HKSAR Government copyright but are publicly available for private and academic use\.888[https://www\.judiciary\.hk/en/other\_information/disclaimer\.html](https://www.judiciary.hk/en/other_information/disclaimer.html)Although HKLII and LRS websites are publicly accessible, they employ a range of mechanisms \(e\.g\. timeouts, dynamically generated URLs, cookie\-based access\) that make them difficult to scrape\. To circumvent these, our scrapers are robust and mimic human web\-browsing behavior\. We develop a generalized scraper for Hong Kong court judgment public\-access websites using scrapy999[https://scrapy\.org/](https://scrapy.org/)and selenium\-webdriver\.101010[https://www\.selenium\.dev/](https://www.selenium.dev/)In order to scrape HKLII, we launch three Google Compute Engine \(GCE\) instances for a total of 40 compute hours\.111111We will release our code for scraping with Docker images created to perform these scrapes\. Given the difficulty in creating this dataset, we believe these routines constitute a considerable resource for academic inquiries into Hong Kong case law\.
### 3\.2Annotation Process
We recruited 10 annotators from a pool of 30 annotator candidates, all graduate students majoring in legal linguistics, selected on the basis of their strong academic backgrounds and familiarity with legal processes\. We then trained all of the annotators for multiple rounds, until they were achieving above an 80% accuracy in both discourse and span identification tasks, based on a gold\-label set that we constructed\. After reaching this agreement level, we began accepting completed tasks from annotators\. We had multiple rounds of conferencing throughout the period of annotation where we discussed edge\-cases, and maintained a WeChat channel throughout the annotation process that was continually monitored\. The annotation process spanned from October 2025 to May 2026\. Together, the annotators annotated∼\\sim290k sentences, with a 10% overlap, from which we calculated aκ=\.8\\kappa=\.8\.
We found that our annotators could learn to identify different discourse and span levels in most contexts quite easily\. Appendix[D](https://arxiv.org/html/2606.06679#A4)presents an example of legal linguistics expert annotation on a Court of Final Appeal judgment \(FACC 22/2018\)\. However, most of the error and ambiguity of the annotation process derived from when to distinguish F9\-admission from F10\-assertion \(e\.g\. “Mr\. WU accepted that the defendant had caused PW1 to lose his properties, but the defendant did not retain any” can be read either as a counsel’s neutral assertion or as an admission on behalf of the defendant\)\. The decision usually depends on many factors, e\.g\. whether the speaker has authority to bind the defendant\. Despite many rounds of training, annotators still sometimes struggled with borderline cases; in such circumstances, they consulted senior legal linguistics experts for adjudication\.
Figure 2:Distribution of Rhetorical Roles within the HKJudge Dataset\.
### 3\.3Dataset Description and Statistics
As shown in Table[3\.3](https://arxiv.org/html/2606.06679#S3.SS3), the court judgments we annotate average1,631\.91\{,\}631\.9tokens and73\.173\.1sentences per document\. The judgments we focused on are criminal cases; see Appendix[B](https://arxiv.org/html/2606.06679#A2)for an HKCFI Judgment Example \(Case No\. HCCC 12/2021\)\. The HKJudge corpus contains4,0004\{,\}000documents,292,240292\{,\}240annotated sentences \(of which285,847285\{,\}847are unique\), and6,527,6006\{,\}527\{,\}600tokens\. Multi\-tagged sentences account for 1\.97% of the corpus\. As shown in Figures[2](https://arxiv.org/html/2606.06679#S3.F2),[3](https://arxiv.org/html/2606.06679#S4.F3), and[4](https://arxiv.org/html/2606.06679#S4.F4), sentence\-level annotations are distributed across different rhetorical roles, termedfactaccounting for 59\.99% of all annotations,inferencefor 27\.96%,othersfor 8\.40%, andresultfor 3\.64%\. The HKJudge dataset is available at121212https://github\.com/xuanxixi/HKJudge\.
HKJudge Dataset Overall StatisticsDocuments4,000Unique sentences285,847Sentence–tag pairs292,240Total tokens6,527,600Avg\. sentences per document73\.1Avg\. tokens per document1,631\.9Avg\. tokens per sentence22\.3Multi\-labeled sentences5,757 \(1\.97%\)
Distribution Across Hong Kong CourtsCourt\# Docs\# Sents\# TokensCourt of Final Appeal80056,2431,256,187Court of Appeal80059,4181,325,962Court of First Instance80058,7911,312,854District Court80057,8731,293,716Magistrates’ Courts80059,9151,338,881Discourse\-level DistributionCategory\# SentsPct\. \(%\)\# TokensF \(Fact\)175,31759\.993,848,629I \(Inference\)81,73127\.962,168,542R \(Result\)10,6433\.64212,758O \(Others\)24,5498\.40297,671
Table 1:Dataset statistics for the HKJudge dataset\.## 4Legal Discourse and Entity Modeling
We frame two tasks using the data we collect:Rhetorical Role ClassificationandElement Extraction\. Each sentence in a judgment document is labeled with one of four top\-level categories: Fact \(F\), Inference \(I\), Result \(R\), or Other \(O\)\. T1 is a sentence classification task that assigns each F, I, or O sentence to one or more sub\-categories from our annotation scheme \(§[C](https://arxiv.org/html/2606.06679#A3)\)\. T2 is a generative extraction task that prompts large language models to identify charges, imprisonment terms, and fines from R sentences\. We will first describe these tasks, then discuss methods, with a particular focus on how we use this setup to interrogate the reasoning capabilities of large language models\.
### 4\.1Task Formulation
F and I sentences describewhathappened andhowthe court reasoned\. R sentences statewhatthe court ruled, including charges, imprisonment terms, and fines\. O sentences are residual\. We model F, I, and O as classification, and R as element extraction\.
#### Task 1: Rhetorical Role Classification\.
The goal of this task is to develop models capable of performing semantic segmentation on court judgments by identifying and classifying rhetorical roles \(RR\)\. LetC=\{c1,c2,…,cn\}C=\\\{c\_\{1\},c\_\{2\},\\ldots,c\_\{n\}\\\}represent a collection of court judgments, whereci∈Cc\_\{i\}\\in Cconsists of a sequence of sentencesSi=\{si1,si2,…,sim\}S\_\{i\}=\\\{s\_\{i1\},s\_\{i2\},\\ldots,s\_\{im\}\\\}, withmmrepresenting the number of sentences in court judgmentcic\_\{i\}\. The task is to assign a rhetorical role labelyij∈Yy\_\{ij\}\\in Yto each sentencesijs\_\{ij\}, whereYYis the predefined set of 26 rhetorical role labels defined in Appendix[C](https://arxiv.org/html/2606.06679#A3), organized into four top\-level categories: Fact \(F\), Inference \(I\), Result \(R\), and Other \(O\)\. Formally, the task can be described as:
f:Si→Yf:S\_\{i\}\\rightarrow Y\(1\)whereYYis defined as:
Y=YF∪YI∪YR∪YOY=Y\_\{F\}\\cup Y\_\{I\}\\cup Y\_\{R\}\\cup Y\_\{O\}\(2\)whereffis a function that maps each sentencesijs\_\{ij\}in a judgmentcic\_\{i\}to its corresponding rhetorical role labelyijy\_\{ij\}\. Thus, the goal is to find:
f\(sij\)=yij,∀sij∈Si,yij∈Yf\(s\_\{ij\}\)=y\_\{ij\},\\quad\\forall s\_\{ij\}\\in S\_\{i\},\\quad y\_\{ij\}\\in Y\(3\)The input to the system is a court judgmentcic\_\{i\}, and the output is rhetorical role labels corresponding to each sentence in the court judgment:
f\(Si\)=\{yi1,yi2,…,yim\},yij∈Yf\(S\_\{i\}\)=\\\{y\_\{i1\},y\_\{i2\},\\ldots,y\_\{im\}\\\},\\quad y\_\{ij\}\\in Y\(4\)We benchmark various large language models using accuracy and macro\-F1 scores\.
#### Task 2: Legal Element Extraction\.
Given a judgment sentencesis\_\{i\}labeled as R, we extract three element types defined by the Hong Kong sentencing framework\(Young,[2016](https://arxiv.org/html/2606.06679#bib.bib92); Xueet al\.,[2024](https://arxiv.org/html/2606.06679#bib.bib51)\): charge \(𝖢𝗁𝖺𝗋𝗀𝖾\\mathsf\{Charge\}\), imprisonment term \(𝖳𝖾𝗋𝗆\\mathsf\{Term\}\), and fine \(𝖥𝗂𝗇𝖾\\mathsf\{Fine\}\)\. The task output is formalized as:
Ei=f\(si\)=\{\(tj,vj\)\}j=1ki,E\_\{i\}=f\(s\_\{i\}\)=\\\{\(t\_\{j\},v\_\{j\}\)\\\}\_\{j=1\}^\{k\_\{i\}\},\(5\)wheref\(⋅\)f\(\\cdot\)denotes the extraction function implemented by a large language model,tj∈𝒯=\{𝖢𝗁𝖺𝗋𝗀𝖾,𝖳𝖾𝗋𝗆,𝖥𝗂𝗇𝖾\}t\_\{j\}\\in\\mathcal\{T\}=\\\{\\mathsf\{Charge\},\\mathsf\{Term\},\\mathsf\{Fine\}\\\}indicates the legal element type, andvjv\_\{j\}is the textual span extracted fromsis\_\{i\}\. The cardinalitykik\_\{i\}varies across sentences, as a single sentence may convey multiple charges or penalties, or contain no extractable element \(Ei=∅E\_\{i\}=\\emptyset\)\. We also benchmark large language models on element extraction using precision and macro\-F1 scores\. We count a prediction\(tj,vj\)\(t\_\{j\},v\_\{j\}\)as correct if its type matches the gold type and its span shares at least 80% of tokens with the gold span \(after removing stop words and punctuation\) with length no more than twice the gold span\.
Figure 3:Discourse function distribution across five court levels\. Higher courts \(HKCFA, HKCA\) allocate a larger share toInference, matching their emphasis on legal reasoning and precedent\.Figure 4:Distribution of annotated sentence lengths in HKJudge Dataset\. Mean \(purple dashed\) and median \(pink dashed\) are indicated\.ModelRhetorical Role ClassificationLegal Element ExtractionAccuracyAUCPrecisionMacro\-F1AccuracyAUCPrecisionMacro\-F1BERT\-based methodsLegalBERT64\.235±1\.79371\.926±1\.58161\.547±1\.86461\.924±1\.83258\.314±1\.98766\.082±1\.84355\.471±2\.05256\.243±2\.011NeuralJudge63\.521±1\.84271\.043±1\.62560\.874±1\.91361\.273±1\.87657\.482±2\.04165\.317±1\.89254\.628±2\.10355\.391±2\.054ML\-LJP64\.923±1\.76272\.583±1\.54861\.832±1\.82162\.213±1\.78958\.832±1\.95466\.728±1\.81255\.741±2\.01856\.521±1\.976JurBERT63\.913±1\.81571\.624±1\.60261\.218±1\.88761\.634±1\.85457\.962±2\.01465\.731±1\.87155\.142±2\.07855\.924±2\.035Commercial LLMsGPT\-473\.532±1\.48778\.214±1\.34870\.583±1\.52470\.921±1\.49868\.421±1\.68773\.582±1\.61266\.518±1\.71266\.832±1\.684Claude\-3\.5\-Sonnet73\.804±1\.47578\.421±1\.34170\.612±1\.51870\.931±1\.49368\.742±1\.67273\.821±1\.59867\.521±1\.70267\.831±1\.674Claude\-Opus\-477\.152±1\.32881\.532±1\.21471\.832±1\.41272\.134±1\.38572\.031±1\.52476\.842±1\.45268\.072±1\.56468\.354±1\.538Gemini\-2\.5\-Pro76\.842±1\.34281\.218±1\.22872\.143±1\.39872\.421±1\.37271\.823±1\.53876\.524±1\.46868\.342±1\.55268\.621±1\.524
Table 2:Performance of BERT\-based methods and commercial LLMs on rhetorical role classification and legal element extraction\.Boldnumbers indicate the best score andunderlinednumbers represent the second best within each category\. Red and blue rows highlight the best and second\-best models in each group, respectively\.
### 4\.2Baselines
We conduct experiments on both BERT\-based and LLM\-based methods\. For BERT\-based methods:
- •LegalBERT\(Chalkidiset al\.,[2020](https://arxiv.org/html/2606.06679#bib.bib63)\)pre\-trains BERT on legal documents from scratch\.
- •NeuralJudge\(Yueet al\.,[2021](https://arxiv.org/html/2606.06679#bib.bib64)\)enhances pre\-trained BERT with LJP\-specific fine\-tuning\.
- •ML\-LJP\(Liuet al\.,[2023b](https://arxiv.org/html/2606.06679#bib.bib61)\)integrates contrastive learning and Graph Attention Networks to model law article interactions\.
- •JurBERT\(Masalaet al\.,[2024](https://arxiv.org/html/2606.06679#bib.bib62)\)extends LegalBERT with a Sliding Encoder for improved long\-context understanding\.
For LLM\-based methods, we compare both open\-source and commercial LLMs:
- •LLaMA 3\.1\(Grattafioriet al\.,[2024](https://arxiv.org/html/2606.06679#bib.bib60)\)andQwen 2\.5\(Huiet al\.,[2024](https://arxiv.org/html/2606.06679#bib.bib59)\)represent state\-of\-the\-art open\-source language models\.
- •GPT 4\(Sanderson,[2023](https://arxiv.org/html/2606.06679#bib.bib57)\),Claude 3\.5 Sonnet\(Benzon,[2025](https://arxiv.org/html/2606.06679#bib.bib55)\),Claude Opus 4\(Joshi,[2026](https://arxiv.org/html/2606.06679#bib.bib56)\), andGemini 2\.5 Pro\(Comaniciet al\.,[2025](https://arxiv.org/html/2606.06679#bib.bib54)\)demonstrate strong performance as proprietary models\.
ModelRhetorical Role ClassificationLegal Element ExtractionAccuracyAUCPrecisionMacro\-F1AccuracyAUCPrecisionMacro\-F1LLaMA\-3\.1\-8B61\.594±2\.13767\.281±1\.94258\.328±2\.11258\.621±2\.08456\.218±2\.28362\.142±2\.11853\.421±2\.29353\.742±2\.261\+ Fine\-tuning64\.473±1\.87671\.421±1\.72861\.231±1\.88461\.521±1\.85760\.274±2\.06567\.418±1\.88757\.312±2\.04857\.642±2\.018LLaMA\-3\.1\-70B70\.231±1\.62474\.582±1\.48367\.428±1\.68767\.752±1\.65865\.962±1\.81270\.231±1\.73863\.142±1\.85263\.421±1\.825\+ Fine\-tuning72\.371±1\.54776\.318±1\.41268\.423±1\.61268\.741±1\.58368\.213±1\.72872\.487±1\.65864\.218±1\.78164\.542±1\.753Qwen\-2\.5\-7B66\.014±1\.82370\.341±1\.71462\.918±1\.88763\.291±1\.85461\.812±1\.94866\.218±1\.84258\.871±2\.01459\.218±1\.983\+ Fine\-tuning71\.423±1\.61273\.821±1\.52764\.412±1\.65864\.702±1\.62867\.321±1\.75270\.142±1\.68760\.218±1\.81460\.541±1\.785Qwen\-2\.5\-14B69\.842±1\.68773\.421±1\.56366\.987±1\.72467\.302±1\.69565\.421±1\.84269\.318±1\.76862\.918±1\.87663\.241±1\.848\+ Fine\-tuning71\.742±1\.60575\.143±1\.48766\.873±1\.64267\.213±1\.61468\.582±1\.71871\.842±1\.65262\.821±1\.78163\.142±1\.753Qwen\-2\.5\-72B70\.752±1\.61275\.218±1\.47869\.218±1\.62869\.547±1\.59866\.521±1\.79870\.842±1\.72465\.083±1\.81265\.421±1\.785\+ Fine\-tuning72\.987±1\.52476\.831±1\.40269\.348±1\.58769\.682±1\.56268\.823±1\.68772\.918±1\.61465\.214±1\.75265\.531±1\.728
Table 3:Performance of open\-source LLMs \(LLaMA\-3\.1 and Qwen\-2\.5 series\) under both zero\-shot and fine\-tuning training\.Boldnumbers indicate the best score andunderlinednumbers represent the second best across all models\. Red and blue rows highlight the best and second\-best results, respectively\.
### 4\.3Evaluation Metrics
To evaluate the performance of models, we adopt a set of standard metrics including Accuracy, AUC, Precision, and Macro\-F1, which are commonly used in classification and element extraction tasks\. For each sentence in the dataset, the predicted label \(26 rhetorical roles for classification and 3 span\-level element types for extraction\) is considered correct if it matches the label assigned by the human expert annotator\. To ensure statistical reliability, we report two times the standard deviation for all metrics using 1,000 bootstrap runsEfron and others \([1986](https://arxiv.org/html/2606.06679#bib.bib52)\)on the test dataset\.
### 4\.4Implementation Details
For all BERT\-based baselines and open\-source LLMs, we fine\-tune them on the HKJudge dataset, using the pre\-segmented sentences of original court judgments as input and their rhetorical roles as labels\. Open\-source and commercial LLMs are additionally evaluated in the zero\-shot setting without task\-specific training\. We fine\-tune LLaMA\-3\.1 and Qwen\-2\.5 using the LLaMA\-Factory framework\(Zhenget al\.,[2024](https://arxiv.org/html/2606.06679#bib.bib53)\), with the AdamW optimizer, learning rate1×10−51\\times 10^\{\-5\}, weight decay 0\.01, and cosine learning rate schedule\. The number of training epochs is set to 3\.
## 5Results and Analysis
In this section, we present the results of our experiments on rhetorical role classification and legal element extraction, and analyze the performance of the different models\. Tables[2](https://arxiv.org/html/2606.06679#S4.T2)and[3](https://arxiv.org/html/2606.06679#S4.T3)summarize the evaluation metrics for BERT\-based methods, commercial LLMs, and open\-source LLMs under zero\-shot and fine\-tuning settings\.
#### Rhetorical Role Classification
Among the evaluated BERT\-based methods, ML\-LJP attains the highest overall performance on rhetorical role classification, and the remaining encoder baselines, including LegalBERT, NeuralJudge, and JurBERT, follow closely with only marginal differences\. The ability of ML\-LJP to capture relationships between law articles through its multi\-law\-aware contrastive representation contributes to its performance, since rhetorical roles in legal documents are not assigned in isolation but follow conventional patterns of citation and reasoning that depend on the surrounding statutory context\.
In contrast, the open\-source LLMs benefit substantially from scale\. Fine\-tuning LLMs further improves performance across all open\-source variants, although the marginal gain from fine\-tuning decreases as the base model grows larger\. The Qwen\-2\.5\-72B model with fine\-tuning attains the strongest open\-source result on rhetorical role classification, exceeding the ML\-LJP by a clear margin, which highlights the advantage of large instruction\-tuned decoders over encoder\-only architectures for discourse\-level legal tasks\.
Among the commercial LLMs, Claude\-Opus\-4 leads on accuracy and AUC, while Gemini\-2\.5\-Pro leads on precision and macro\-F1\. The gap between the strongest open\-source model and the commercial systems is smaller than the gap between the BERT\-family baselines and that strongest open\-source model, which suggests that the principal bottleneck on rhetorical role classification is the reasoning supporting the tag decision rather than the decision itself, and that this reasoning capability scales with model capacity and instruction tuning\.
#### Legal Element Extraction
Performance decreases across most of the evaluated model families when moving from rhetorical role classification to legal element extraction, indicating that span\-level extraction of charges, imprisonment terms, and fines is the more challenging task on HKJudge\. The BERT\-family ranking on extraction is largely preserved from the classification setting, with ML\-LJP remaining the strongest in this group, although the absolute scores degrade more substantially than they do on classification\.
For the open\-source LLMs, fine\-tuning produces larger gains on extraction than on classification, with the improvements observed for LLaMA\-3\.1\-8B and Qwen\-2\.5\-72B among the largest single\-task gains across our experiments\. This is consistent with our hypothesis that extraction depends more heavily on task\-specific supervision than classification does, since the surface conventions of HK sentencing spans, particularly the ordinance citation format and the phrasing of suspended and concurrent terms, are unlikely to be adequately represented in general instruction tuning\. The commercial LLMs retain the lead on extraction, although their advantage over the strongest fine\-tuned open\-source model is reduced relative to the classification setting, which nonetheless demonstrates the potential of commercial LLMs for span\-level legal reasoning tasks\.
## 6Conclusion and Future Work
We presentedHKJudge, the first sentence\-level discourse corpus of Hong Kong court judgments fully annotated by legal linguistics experts\. We benchmarked four BERT\-based encoders, two open\-source LLMs under zero\-shot and fine\-tuning settings, and four commercial LLMs on rhetorical role classification and legal element extraction\. Across both tasks, performance increases monotonically from BERT\-based encoders to fine\-tuned open\-source LLMs to commercial LLMs\. ML\-LJP achieves the highest scores among encoders, fine\-tuned Qwen\-2\.5\-72B leads the open\-source LLMs, and Claude\-Opus\-4 and Gemini\-2\.5\-Pro lead the commercial LLMs\.
We highlight three findings from these results\. First, the performance gap between the strongest open\-source and commercial models is smaller than the gap between BERT\-based baselines and the strongest open\-source LLM, indicating that the principal bottleneck on rhetorical role classification is the legal reasoning that supports each tag assignment rather than the assignment itself, and that this reasoning capability scales with model size and instruction tuning\. Second, fine\-tuning yields larger gains on legal element extraction than on rhetorical role classification, consistent with the surface conventions of Hong Kong sentencing spans, including ordinance citation formats and the phrasing of suspended and concurrent terms, being underrepresented in general\-purpose instruction tuning\. Third, all evaluated models fall noticeably short of expert annotators, indicating open challenges for legal LLM reasoning over legal discourse\.
For future work, we will use the HKJudge dataset proposed in this paper to explore and address legal judgment prediction in Hong Kong, a task that supports legal professionals \(practitioners, law firms, judicial bodies, policymakers, and government departments\), improving judicial efficiency and justice, and enabling citizens to anticipate case outcomes without costly legal consultation\. Together with the dataset and benchmark released in this work, we hope HKJudge will serve as a step toward legal discourse modeling for LegalAI\.
## Limitations
Our research focuses on Hong Kong criminal case law, which is governed by the common\-law tradition and exhibits a bilingual drafting practice with highly standardized rhetorical conventions\. Consequently, the discourse schema and trained models developed in this work may not be directly applicable to judgments from civil\-law jurisdictions, monolingual common\-law systems, or non\-criminal legal areas such as civil and family proceedings\. The results of our study, therefore, may not cover all countries or types of legal documents\.
In addition, the boundary between certain Fact sub\-categories \(notably F9\-admission vs\. F10\-assertion\) remains subject to interpretive judgment by the annotator\. Our span\-level schema is also restricted to three sentencing elements \(charge, imprisonment term, and fine\), leaving alternative outcomes such as suspended sentences, community service orders, and disqualification orders for future extension\.
## Ethics Statement
Our dataset and evaluation benchmark contain no personal, sensitive, or private information; they consist solely of publicly available data\.
## Acknowledgments
The work described in this paper was fully supported by a grant from the Research Grants Council of HKSAR, China \(Project No\. CityU 11602524\)\. We also thank the expert annotators in legal linguistics for their valuable contributions\.
## References
- Predicting judicial decisions of the european court of human rights: a natural language processing perspective\.PeerJ computer science2,pp\. e93\.Cited by:[§2](https://arxiv.org/html/2606.06679#S2.p1.1)\.
- W\. L\. Benzon \(2025\)What miriam yevick saw: the nature of intelligence and the prospects for ai, a dialog with claude 3\.5 sonnet\.A Dialog with Claude3\.Cited by:[2nd item](https://arxiv.org/html/2606.06679#S4.I2.i2.p1.1)\.
- L\. Carlson, D\. Marcu, and M\. E\. Okurowski \(2003\)Building a discourse\-tagged corpus in the framework of rhetorical structure theory\.InCurrent and New Directions in Discourse and Dialogue,pp\. 85–112\.Cited by:[§2](https://arxiv.org/html/2606.06679#S2.p1.1),[§2](https://arxiv.org/html/2606.06679#S2.p2.1)\.
- I\. Chalkidis, M\. Fergadiotis, P\. Malakasiotis, N\. Aletras, and I\. Androutsopoulos \(2020\)LEGAL\-bert: the muppets straight out of law school\.InFindings of the association for computational linguistics: EMNLP 2020,pp\. 2898–2904\.Cited by:[1st item](https://arxiv.org/html/2606.06679#S4.I1.i1.p1.1)\.
- K\. K\. Cheng \(2015\)Moral discourse in hong kong’s chinese criminal proceedings\.The Chinese Journal of Comparative Law3\(2\),pp\. 375–389\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- L\. Cheng and L\. He \(2016\)Revisiting judgment translation in hong kong\.Semiotica2016\(209\),pp\. 59–75\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- L\. Cheng, K\. K\. Sin,et al\.\(2008a\)A discursive approach to legal texts: court judgments as an example\.The Asian ESP Journal4\(1\),pp\. 14–28\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p1.1),[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- L\. Cheng, K\. Sin, and Y\. Zheng \(2008b\)Contrastive analysis of chinese and american court judgments\.US\-China Law Review5,pp\. 56\.Cited by:[§2\.1](https://arxiv.org/html/2606.06679#S2.SS1.p1.1),[§2](https://arxiv.org/html/2606.06679#S2.p1.1)\.
- G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein, O\. Ram, D\. Zhang, E\. Rosen,et al\.\(2025\)Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.arXiv preprint arXiv:2507\.06261\.Cited by:[2nd item](https://arxiv.org/html/2606.06679#S4.I2.i2.p1.1)\.
- T\. Dancy and M\. Zalnieriute \(2026\)AI and transparency in judicial decision making\.Oxford journal of legal studies46\(1\),pp\. 1–34\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p1.1),[§2](https://arxiv.org/html/2606.06679#S2.p1.1)\.
- B\. Efronet al\.\(1986\)Bootstrap methods for standard errors, confidence intervals, and other measures of statistical accuracy\.Statistical science,pp\. 54–75\.Cited by:[§4\.3](https://arxiv.org/html/2606.06679#S4.SS3.p1.1)\.
- Z\. Fei, S\. Zhang, X\. Shen, D\. Zhu, X\. Wang, J\. Ge, and V\. Ng \(2025\)Internlm\-law: an open\-sourced chinese legal large language model\.InProceedings of the 31st International Conference on Computational Linguistics,pp\. 9376–9392\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- J\. P\. Gee \(2025\)An introduction to discourse analysis: theory and method\.routledge\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- S\. Ghosh \(2019\)Identification of rhetorical roles of sentences in indian legal judgments\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- R\. Gill \(2000\)Discourse analysis\.Qualitative researching with text, image and sound1,pp\. 172–190\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- H\. Gillman \(2001\)What’s law got to do with it? judicial behavioralists test the “legal model” of judicial decision making\.Law & social inquiry26\(2\),pp\. 465–504\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.\(2024\)The llama 3 herd of models\.InNeural Information Processing Systems,Cited by:[1st item](https://arxiv.org/html/2606.06679#S4.I2.i1.p1.1)\.
- Z\. Han, V\. K\. Bhatia, and Y\. Ge \(2018\)The structural format and rhetorical variation of writing chinese judicial opinions: a genre analytical approach\.Pragmatics28\(4\),pp\. 463–488\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p1.1)\.
- Z\. He, P\. Cao, C\. Wang, Z\. Jin, Y\. Chen, J\. Xu, H\. Li, K\. Liu, and J\. Zhao \(2024\)Agentscourt: building judicial decision\-making agents with court debate simulation and legal knowledge augmentation\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 9399–9416\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p1.1)\.
- L\. Held and I\. Habernal \(2026\)LaCour\!: enabling research on argumentation in hearings of the european court of human rights: l\. held and i\. habernal\.34\(2\),pp\. 311–334\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- D\. Hendrycks, C\. Burns, A\. Chen, and S\. Ball \(2021\)CUAD: an expert\-annotated nlp dataset for legal contract review\.InThirty\-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track \(Round 1\),Cited by:[§2](https://arxiv.org/html/2606.06679#S2.p2.1)\.
- B\. Hui, J\. Yang, Z\. Cui, J\. Yang, D\. Liu, L\. Zhang, T\. Liu, J\. Zhang, B\. Yu, K\. Lu,et al\.\(2024\)Qwen2\. 5\-coder technical report\.arXiv preprint arXiv:2409\.12186\.Cited by:[1st item](https://arxiv.org/html/2606.06679#S4.I2.i1.p1.1)\.
- S\. Joshi \(2026\)Architectural advances and performance benchmarks of large language models in light of anthropic’s claude opus 4\.6\.Cited by:[2nd item](https://arxiv.org/html/2606.06679#S4.I2.i2.p1.1)\.
- S\. Joty, G\. Carenini, R\. Ng, and G\. Murray \(2019\)Discourse analysis and its applications\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts,pp\. 12–17\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- P\. Kalamkar, A\. Tiwari, A\. Agarwal, S\. Karn, S\. Gupta, V\. Raghavan, and A\. Modi \(2022\)Corpus for automatic structuring of legal documents\.InProceedings of the Thirteenth Language Resources and Evaluation Conference,pp\. 4420–4429\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p1.1),[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- W\. Ko, Y\. Wu, C\. Dalton, D\. Srinivas, G\. Durrett, and J\. J\. Li \(2023\)Discourse analysis via questions and answers: parsing dependency structures of questions under discussion\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 11181–11195\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- H\. Li, J\. Chen, J\. Yang, Q\. Ai, W\. Jia, Y\. Liu, K\. Lin, Y\. Wu, G\. Yuan, Y\. Hu,et al\.\(2025\)Legalagentbench: evaluating llm agents in legal domain\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 2322–2344\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p1.1),[§2](https://arxiv.org/html/2606.06679#S2.p1.1)\.
- Z\. Li, X\. Xuan, S\. Song, and B\. Jin \(2026\)FASTQR: fast, accurate and stable quantile regression for time\-series analysis via adaptive huber smoothing\.InICASSP 2026 \- 2026 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 4261–4265\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP55912.2026.11462804)Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- B\. L\. Liebman, M\. E\. Roberts, R\. E\. Stern, and A\. Z\. Wang \(2020\)Mass digitization of chinese court decisions: how to use text as data in the field of chinese law\.Journal of Law and Courts8\(2\),pp\. 177–201\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- Z\. Lin, J\. Wang, R\. Li, F\. Shen, and X\. Xuan \(2025\)PrimeK\-net: multi\-scale spectral learning via group prime\-kernel convolutional neural networks for single channel speech enhancement\.InICASSP 2025 \- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 1–5\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP49660.2025.10890034)Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- P\. Liu, H\. Lin, M\. Liao, H\. Xiang, X\. Han, and L\. Sun \(2023a\)WebDP: understanding discourse structures in semi\-structured web documents\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 10235–10258\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- Y\. Liu, Y\. Wu, Y\. Zhang, C\. Sun, W\. Lu, F\. Wu, and K\. Kuang \(2023b\)Ml\-ljp: multi\-law aware legal judgment prediction\.InProceedings of the 46th international ACM SIGIR conference on research and development in information retrieval,pp\. 1023–1034\.Cited by:[3rd item](https://arxiv.org/html/2606.06679#S4.I1.i3.p1.1)\.
- Y\. Maley \(2014\)The language of the law\.InLanguage and the Law,pp\. 11–50\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p1.1),[§2](https://arxiv.org/html/2606.06679#S2.p1.1)\.
- V\. Malik, R\. Sanjay, S\. K\. Nigam, K\. Ghosh, S\. K\. Guha, A\. Bhattacharya, and A\. Modi \(2021\)ILDC for cjpe: indian legal documents corpus for court judgment prediction and explanation\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 4046–4062\.Cited by:[§2](https://arxiv.org/html/2606.06679#S2.p1.1)\.
- M\. Masala, T\. Rebedea, and H\. Velicu \(2024\)Improving legal judgement prediction in romanian with long text encoders\.InProceedings of the 3rd Annual Meeting of the Special Interest Group on Under\-resourced Languages@ LREC\-COLING 2024,pp\. 126–132\.Cited by:[4th item](https://arxiv.org/html/2606.06679#S4.I1.i4.p1.1)\.
- F\. Mo, K\. Mao, Z\. Zhao, H\. Qian, H\. Chen, Y\. Cheng, X\. Li, Y\. Zhu, Z\. Dou, and J\. Nie \(2025\)A survey of conversational search\.ACM Transactions on Information Systems43\(6\),pp\. 1–50\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p1.1),[§2](https://arxiv.org/html/2606.06679#S2.p1.1)\.
- N\. S\. Nakshatri, N\. Mehta, S\. Liu, S\. Chen, D\. Hopkins, D\. Roth, and D\. Goldwasser \(2025\)Talking point based ideological discourse analysis in news events\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 575–594\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- S\. K\. Nigam, T\. Dubey, G\. Sharma, N\. Shallum, K\. Ghosh, and A\. Bhattacharya \(2025\)Legalseg: unlocking the structure of indian legal judgments through rhetorical role classification\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 1129–1144\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- R\. Prasad, B\. Webber, and A\. Joshi \(2017\)The penn discourse treebank: an annotated corpus of discourse relations\.InHandbook of linguistic annotation,pp\. 1197–1217\.Cited by:[§2](https://arxiv.org/html/2606.06679#S2.p1.1),[§2](https://arxiv.org/html/2606.06679#S2.p2.1)\.
- N\. Robinson \(2013\)Structure matters: the impact of court structure on the indian and us supreme courts\.61\(1\),pp\. 173–208\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- A\. Rosas \(2007\)The european court of justice in context: forms and patterns of judicial dialogue\.1,pp\. 121\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- K\. Sanderson \(2023\)GPT\-4 is here: what scientists think\.Nature615\(7954\),pp\. 773\.Cited by:[2nd item](https://arxiv.org/html/2606.06679#S4.I2.i2.p1.1)\.
- M\. Saravanan \(2010\)Identification of rhetorical roles for segmentation and summarization of a legal judgment\.Artificial Intelligence and Law18\(1\),pp\. 45–76\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p1.1)\.
- S\. Sen \(2023\)Analyzing hong kong’s legal judgments from a computational linguistics point\-of\-view\.arXiv preprint arXiv:2305\.02558\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p3.1)\.
- W\. Shi, H\. Zhu, J\. Ji, M\. Li, J\. Zhang, R\. Zhang, J\. Zhu, J\. Xu, S\. Han, and Y\. Guo \(2025\)Legalreasoner: step\-wised verification\-correction for legal judgment reasoning\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7297–7313\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p3.1),[§2](https://arxiv.org/html/2606.06679#S2.p1.1)\.
- D\. Shu, H\. Zhao, X\. Liu, D\. Demeter, M\. Du, and Y\. Zhang \(2024\)Lawllm: law large language model for the us legal system\.InProceedings of the 33rd ACM International Conference on information and knowledge management,pp\. 4882–4889\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- F\. Sovrano, M\. Palmirani, S\. Sapienza, and V\. Pistone \(2025\)DiscoLQA: zero\-shot discourse\-based legal question answering on european legislation\.Artificial Intelligence and Law33\(2\),pp\. 323–359\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- W\. Werner \(1981\)Corporation law in search of its future\.Columbia Law Review81\(8\),pp\. 1611–1666\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p1.1)\.
- R\. C\. Williams \(2022\)Jurisdiction as power\.89\(7\),pp\. 1719–1792\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- C\. Xiao, H\. Zhong, Z\. Guo, C\. Tu, Z\. Liu, M\. Sun, Y\. Feng, X\. Han, Z\. Hu, H\. Wang,et al\.\(2018\)Cail2018: a large\-scale legal dataset for judgment prediction\.arXiv preprint arXiv:1807\.02478\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- X\. Xuan, D\. Carbone, W\. Zhang, R\. Pandey, and T\. H\. Kinnunen \(2026a\)WST\-x series: wavelet scattering transform for interpretable speech deepfake detection\.External Links:2602\.02980,[Link](https://arxiv.org/abs/2602.02980)Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- X\. Xuan and C\. Kit \(2026\)TransLaw: a large\-scale dataset and multi\-agent benchmark simulating professional translation of hong kong case law\.External Links:2507\.00875,[Link](https://arxiv.org/abs/2507.00875)Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p3.1)\.
- X\. Xuan, X\. Liu, W\. Zhang, Y\. Lin, X\. Lin, and T\. Kinnunen \(2026b\)WaveSP\-net: learnable wavelet\-domain sparse prompt tuning for speech deepfake detection\.InICASSP 2026 \- 2026 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 18047–18051\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP55912.2026.11461768)Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- X\. Xuanet al\.\(2024a\)Conformer\-based speaker recognition model for real\-time multi\-scenarios\.Computer Engineering and Applications60\(7\),pp\. 147–156\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- X\. Xuanet al\.\(2024b\)Efficient real\-time multi\-scenario speaker recognition with mel\-spectrogram\-based hybrid tdnn for edge system\.InINTERSPEECH 2024\-Young Female\* Researchers in Speech Workshop \(YFRSW 2024\),Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- X\. Xuanet al\.\(2025\)Multilingual Source Tracing of Speech Deepfakes: A First Benchmark\.In5th Symposium on Security and Privacy in Speech Communication,pp\. 27–34\.External Links:[Document](https://dx.doi.org/10.21437/SPSC.2025-5)Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- X\. Xuan, W\. Zhang, Z\. Li, J\. Williams, V\. Hautamäki, and T\. H\. Kinnunen \(2026c\)Disentangling Speaker Traits for Deepfake Source Verification via Chebyshev Polynomial and Riemannian Metric Learning\.InInterspeech 2026,Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- X\. Xuan, Z\. Zhu, W\. Zhang, Y\. Lin, and T\. Kinnunen \(2025\)Fake\-mamba: real\-time speech deepfake detection using bidirectional mamba as self\-attention’s alternative\.In2025 IEEE Automatic Speech Recognition and Understanding Workshop \(ASRU\),Vol\.,pp\. 1–8\.External Links:[Document](https://dx.doi.org/10.1109/ASRU65441.2025.11434679)Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- Z\. Xue, H\. Liu, Y\. Hu, Y\. Qian, Y\. Wang, K\. Kong, C\. Wang, Y\. Liu, and W\. Shen \(2024\)LEEC for judicial fairness: a legal element extraction dataset with extensive extra\-legal labels\.\.InIJCAI,pp\. 7527–7535\.Cited by:[§4\.1](https://arxiv.org/html/2606.06679#S4.SS1.SSS0.Px2.p1.4)\.
- S\. N\. Young \(2016\)Sentencing\.InUnderstanding criminal justice in Hong Kong,pp\. 286–306\.Cited by:[§4\.1](https://arxiv.org/html/2606.06679#S4.SS1.SSS0.Px2.p1.4)\.
- W\. Yu \(2023\)Negotiation of justice: the discursive construction of attitudinal positioning in bilingual legal judgments of hksar v kwan wan ki\.International Journal of Legal Discourse8\(2\),pp\. 299–333\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- W\. Yu \(2025\)Linguistic tension in the postcolonial judicial landscape: a case study of legal bilingualism in hong kong sar\.International Journal for the Semiotics of Law\-Revue internationale de Sémiotique juridique,pp\. 1–20\.Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p2.1)\.
- L\. Yue, Q\. Liu, B\. Jin, H\. Wu, K\. Zhang, Y\. An, M\. Cheng, B\. Yin, and D\. Wu \(2021\)Neurjudge: a circumstance\-aware neural framework for legal judgment prediction\.InProceedings of the 44th international ACM SIGIR conference on research and development in information retrieval,pp\. 973–982\.Cited by:[2nd item](https://arxiv.org/html/2606.06679#S4.I1.i2.p1.1)\.
- W\. Zhanget al\.\(2026\)Robust rumor detection against noise\.NeurocomputingEur\. J\. Legal Stud\.Artificial Intelligence and LawThe American Journal of Comparative LawThe University of Chicago Law Review,pp\. 132741\.External Links:ISSN 0925\-2312,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.neucom.2026.132741)Cited by:[§1](https://arxiv.org/html/2606.06679#S1.p4.1)\.
- Y\. Zheng, R\. Zhang, J\. Zhang, Y\. Ye, and Z\. Luo \(2024\)Llamafactory: unified efficient fine\-tuning of 100\+ language models\.InProceedings of the 62nd annual meeting of the association for computational linguistics \(volume 3: system demonstrations\),pp\. 400–410\.Cited by:[§4\.4](https://arxiv.org/html/2606.06679#S4.SS4.p1.1)\.
## Appendix AUse of AI Assistants
We used Claude Opus 4\.7 and Sonnet 4\.6 for coding, shortening texts and editing LaTeX more efficiently\.
## Appendix BHKCFI Judgment Example \(Case No\. HCCC 12/2021\)

Figure 5:Example of the first page of a Hong Kong Court of First Instance judgment \(\[2021\] HKCFI 1919, Case No\. HCCC 12/2021\)\.
## Appendix CThe HKJudge Legal Discourse Annotation Scheme
CategoryRhetorical Role TagDescriptionF \(Fact\)—What kinds of information are presented in a court hearing and/or recorded in a judgment?F0F0\-chargeCharge\(s\) or offence\(s\) in a criminal case; a special case ofF1\-issue\.F1F1\-issueMarking the key issue\(s\) that a judgment is intended to judge\.F2F2\-eventDescriptions of time, location, individuals involved, causes, processes, and outcomes of an event; constitutes the narrative of the core incident under consideration\.F3F3\-supplementSupplementary information regarding the core incident, such as background details of relevant individuals \(typically in mitigation arguments\) and events, attributes of objects, etc\.F4F4\-previous\_infoApplicable*exclusively to appeal cases*; citations or paraphrasing of prior court rulings or reasoning of the*same case*\.F5F5\-previous\_recordHistorical records of verdict, convictions and/or sentence from previous*unrelated*trial\(s\)\.F6F6\-argumentNeutral reporting of disputed legal issues or contentions advanced by the appellant or defendant \(complaint, ground of appeal, application, purpose\)\.F7F7\-juryStatements of fact or view provided by the jury\.F8F8\-otherMiscellaneous factual content not covered by other F sub\-categories\.F9F9\-admissionDefendant’s admission, confession, or guilty plea, in either direct or indirect quote\.F10F10\-assertionClaims or statements by defendant\(s\), their counsels, or both sides, which*may or may not be factual*\.F11F11\-questionQuestions asked or challenges raised during hearing or other court events\.F12F12\-answerAnswers during hearing or other court events\.F13F13\-objectionObjection of either side \(usually to a question raised to the defendant\) and relevant info \(reason, sustained/overruled, outcome\)\.F14F14\-instructions2juryJudge’s instructions given to the jury, in either direct or indirect quote\.I \(Inference\)—How does the court reason towards its decision?I1I1\-case\_lawReferences to prior judicial precedents \(common law\) employed during reasoning\. Includes content from previous judgments, in direct or indirect quote\.I2I2\-ordinanceCitations of statutory laws, regulations, or ordinances used in the reasoning process\.I3I3\-legislationCitations of legislative documents, processes, organisations, or relevant info thatI2\-ordinancedoes not cover\.I4I4\-conventional\_practiceEstablished customary practices \(non\-statutory\) referenced during reasoning, such as reductions in sentencing \(e\.g\., one\-third reduction\)\.I5I5\-juryContent related to the jury within the reasoning process\.I6I6\-assertionJudge’s conclusive statement about the current case, such as assertion, concluding evaluation, result, etc\., during inference\.¶I7I7\-otherMiscellaneous inferential content not covered by other I sub\-categories\.I8I8\-questionQuestion raised by the judge as part of reasoning \(vs\.F8\-questionfor questioning a party\)\.R \(Result\)—What does the court decide?–RFinal judgment determinations for a case\.–R\-otherSupplementary info adhered to a court determination \(explanation, interpretation, clarification, calculation, consequence, effects\)\.O \(Others\)–OSentences that do not fit any of the above categories \(e\.g\., “Court adjourns”, appearance records\)\.Table 4:TheHKJudgesentence\-level rhetorical role annotation scheme, comprising 26 tags grouped into four legal discourse functions: Fact \(F\), Inference \(I\), Result \(R\), and Other \(O\)\. The scheme was designed by legal linguistics experts and used as the reference during annotation\.## Appendix DExample of Legal Linguistics Expert Annotation on Court of Final Appeal Judgment \(FACC 22/2018\)
Figure 6:Example of sentence\-level rhetorical role annotation inHKJudge, illustrated on a Court of Final Appeal judgment \(FACC 22/2018\)\. Each line is prefixed with its assigned rhetorical role tag\.Similar Articles
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
RankJudge is a benchmark generator that creates paired multi-turn conversations with injected flaws to evaluate LLM judges on their ability to correctly identify better and worse responses in complex dialogues.
SenseJudge: Human-Centric Preference-Driven Judgment Framework
SenseJudge is a human-centric framework for customizable LLM judging that adapts to diverse user preferences, outperforming existing methods. It also introduces SenseBench, a benchmark derived from real-world multi-turn interactions.
JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation
JudgeArena is an open-source framework that unifies major LLM-judge benchmarks under a single interface, enabling systematic study of judge choices and offering open-model judges that match or outperform closed models, with the ability to simulate LMArena Elo scores.
LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification
Researchers release LegalBench-BR, the first public benchmark for evaluating LLMs on Brazilian legal text classification, showing LoRA-fine-tuned BERTimbau dramatically outperforms GPT-4o mini and Claude 3.5 Haiku.
Judge Circuits
This paper investigates the internal mechanisms of LLM-as-a-judge, finding a shared Latent Evaluator sub-graph in mid-to-late MLPs across models that handles abstract judging, while format-specific terminal branches map the judgment to output tokens, revealing the cause of format-induced inconsistency.