Polar: A Benchmark for Evaluating Political Bias in LLMs
Summary
Polar is a 4,026-instance multiple-choice benchmark for evaluating political bias in LLMs across U.S. and South Korean political contexts, measuring bias through option-level likelihoods. Experiments on 38 LLMs show systematic bias patterns varying by political context, issue category, and presentation language.
View Cached Full Text
Cached at: 06/12/26, 08:51 AM
# Polar: A Benchmark for Evaluating Political Bias in LLMs
Source: [https://arxiv.org/html/2606.12922](https://arxiv.org/html/2606.12922)
Sangho Kim1,Heejin Kim11footnotemark:11,Yoonhee Park1, Hyunggeun Jeon1,Jaejin Lee1,2
1Graduate School of Data Science, Seoul National University 2Dept\. of Computer Science and Engineering, Seoul National University \{ksh4931, kheejin, yoonheepark, jhg123456, jaejin\}@snu\.ac\.kr https://thunder\.snu\.ac\.kr
###### Abstract
Political bias in large language models \(LLMs\) is increasingly significant, but difficult to measure reproducibly across political and linguistic contexts\. We introducePolar, a 4,026\-instance multiple\-choice benchmark that measures political bias through option\-level likelihoods rather than prompt\-based generation\. Polar covers two ideological axes and eight issue categories derived from the Manifesto Project, and evaluates models in parallel across U\.S\. and South Korean political contexts\. Across 38 LLMs, measured bias varies systematically with political context, issue category, model group, and presentation language\. All models lean left\-progressive on U\.S\. political content, but show more centered and mixed patterns on South Korean content\. Translation experiments further show that presentation language alone can shift measured bias\. These findings highlight the need for multilingual and cross\-contextual evaluation of political bias in LLMs\.
Polar: A Benchmark for Evaluating Political Bias in LLMs
Sangho Kim††thanks:These authors contributed equally to this work\.1, Heejin Kim11footnotemark:11, Yoonhee Park1,Hyunggeun Jeon1,Jaejin Lee1,21Graduate School of Data Science, Seoul National University2Dept\. of Computer Science and Engineering, Seoul National University\{ksh4931, kheejin, yoonheepark, jhg123456, jaejin\}@snu\.ac\.krhttps://thunder\.snu\.ac\.kr
## 1Introduction
Large language models \(LLMs\) are increasingly deployed in domains that shape public discourse and institutional decision\-making, including education, healthcare, journalism, and public policy\(Elkinset al\.,[2023](https://arxiv.org/html/2606.12922#bib.bib64); Gilsonet al\.,[2023](https://arxiv.org/html/2606.12922#bib.bib65); Xuet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib63); OECD,[2025](https://arxiv.org/html/2606.12922#bib.bib62)\)\. In these settings, the perspectives reflected in model outputs can influence how information is framed and which positions users perceive as legitimate\. Because LLMs are trained on real\-world corpora, they inevitably inherit and reproduce biases embedded in the underlying data\(Blodgettet al\.,[2020](https://arxiv.org/html/2606.12922#bib.bib11); Benderet al\.,[2021](https://arxiv.org/html/2606.12922#bib.bib10); Navigliet al\.,[2023](https://arxiv.org/html/2606.12922#bib.bib13); Kumaret al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib4)\)\. A growing body of work has focused on identifying and mitigating bias in LLM outputs\(Gallegoset al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib3)\)\. While extensive work has examined demographic stereotypes related to attributes such as gender and race, political bias has received relatively less attention\.
Political bias in LLMs is particularly consequential because model outputs can shape users’ opinions, societal attitudes, and policy\-related decisions\(Witteet al\.,[2023](https://arxiv.org/html/2606.12922#bib.bib66); Gubelmann and Karray,[2025](https://arxiv.org/html/2606.12922#bib.bib57); Hackenburget al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib28)\)\. As LLMs increasingly mediate access to information, recurring ideological patterns in generated responses may narrow the range of viewpoints users encounter and reinforce political polarization at scale\(Kuenzler and Schmid,[2026](https://arxiv.org/html/2606.12922#bib.bib87)\)\. Recent studies have begun developing datasets and methods for measuring political bias\(Ceronet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib47); Becchetti and Solferino,[2025](https://arxiv.org/html/2606.12922#bib.bib48)\), and major AI developers now frame political neutrality as a post\-training objective\(OpenAI,[2025](https://arxiv.org/html/2606.12922#bib.bib20); Anthropic,[2025](https://arxiv.org/html/2606.12922#bib.bib21)\)\. Policymakers have likewise begun treating political bias in LLMs as a governance concern\(UK AI Security Institute,[2025](https://arxiv.org/html/2606.12922#bib.bib89); US Office of Management and Budget,[2025](https://arxiv.org/html/2606.12922#bib.bib88)\)\.
However, existing benchmarks remain insufficient in three ways\. First, they often rely on prompt\-based generation, which is highly sensitive to prompt phrasing, decoding choices, and refusal behavior\. Second, prior benchmarks typically cover only a limited set of broad political topics, restricting their ability to characterize ideological tendencies systematically across issue domains\. Third, they remain centered on U\.S\. and Western European contexts\(Batzneret al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib51); Chenet al\.,[2026](https://arxiv.org/html/2606.12922#bib.bib54)\)\. Since political framing and issue priorities vary across countries and languages, evaluating political bias in LLMs requires benchmarks grounded in specific political and linguistic contexts\.
To address these limitations, we introducePolar, a multiple\-choice benchmark for systematic and reproducible evaluation of political bias across U\.S\. and South Korean political contexts\. Polar contains 4,026 evaluation instances spanning two ideological dimensions and eight issue categories derived from the comparative policy coding scheme of the Manifesto Project\(Lehmannet al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib78)\)\. Rather than relying on free\-form generation, Polar evaluates political preferences using option\-level likelihoods over politically opposed continuations and a semantically unrelated continuation\. This design enables both directional bias measurement and task\-specific language\-modeling competence evaluation through an ICAT\-style score\.
Using Polar, we evaluate 38 LLMs including 23 globally deployed models and 15 Korean\-specialized systems\. Our results show that political bias varies across political context, issue category, model group, and presentation language\. In the U\.S\. dataset, all evaluated models fall in the left\-progressive region, whereas the South Korean dataset shows more centered and mixed distributions\. Category\-level analysis reveals issue\-specific preferences that are not visible in aggregate axis\-level scores, and translation experiments show that presentation language affects measured bias\.
The main contributions of this paper are summarized as follows:
- •We propose a reproducible option\-level likelihood framework for evaluating political bias in LLMs, reducing sensitivity to prompt phrasing, decoding choices, and refusal behavior\.
- •We construct Polar, a multiple\-choice benchmark covering U\.S\. and South Korean political contexts\. It contains 4,026 instances organized along two ideological dimensions and eight issue categories derived from the Manifesto Project’s policy classification scheme\.
- •We adapt ICAT to political bias evaluation, jointly measuring political preference and language modeling competence to avoid conflating ideological balance with degraded model behavior\.
- •We evaluate 38 LLMs and find that political bias varies across political context, issue category, model group, and presentation language, with U\.S\. content producing consistent left\-progressive tendencies and South Korean content showing more centered and mixed patterns\.
## 2Related Work
### 2\.1Social Bias in LLMs
Prior studies have introduced benchmarks for measuring social bias in LLMs, including biases related to demographic groups, social attributes, or ambiguous contexts\(Zhaoet al\.,[2018](https://arxiv.org/html/2606.12922#bib.bib7); Liet al\.,[2020](https://arxiv.org/html/2606.12922#bib.bib15); Nangiaet al\.,[2020](https://arxiv.org/html/2606.12922#bib.bib5); Nadeemet al\.,[2021](https://arxiv.org/html/2606.12922#bib.bib6); Parrishet al\.,[2022](https://arxiv.org/html/2606.12922#bib.bib8)\)\. These benchmarks typically evaluate bias either through free\-form generations or through likelihood comparisons over predefined candidate options\. Likelihood\-based evaluations, such as StereoSet\(Nadeemet al\.,[2021](https://arxiv.org/html/2606.12922#bib.bib6)\)and CrowS\-Pairs\(Nangiaet al\.,[2020](https://arxiv.org/html/2606.12922#bib.bib5)\), offer a more useful precedent for our approach because they measure model preferences over paired or controlled completions\. However, most existing benchmarks focus on demographic stereotypes involving attributes such as gender and race\(Yanget al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib12)\), while broader sociopolitical biases have received less attention\(Smithet al\.,[2022](https://arxiv.org/html/2606.12922#bib.bib17); Rozado,[2024](https://arxiv.org/html/2606.12922#bib.bib34)\)\.
### 2\.2Political Bias in LLMs
Political bias in LLMs can be understood as a persistent tendency to favor one side of a politically contested issue over another\(Feldman,[2011](https://arxiv.org/html/2606.12922#bib.bib36); Liuet al\.,[2022](https://arxiv.org/html/2606.12922#bib.bib35)\)\. Recent studies show that LLMs can exhibit political bias when responding to politically sensitive prompts\(Argyleet al\.,[2023](https://arxiv.org/html/2606.12922#bib.bib38); Hartmannet al\.,[2023](https://arxiv.org/html/2606.12922#bib.bib39)\), with several evaluations reporting left\-associated preferences under Western ideological taxonomies across both commercial and open source models\(Bernardelleet al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib27); Becchetti and Solferino,[2026](https://arxiv.org/html/2606.12922#bib.bib44); Shuet al\.,[2026](https://arxiv.org/html/2606.12922#bib.bib26)\)\. These biases may have downstream risks, affecting users’ political attitudes and propagating into automated moderation systems for hate speech detection on social media\(Fenget al\.,[2023](https://arxiv.org/html/2606.12922#bib.bib22); Potteret al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib25); Sharmaet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib31); Yanget al\.,[2025b](https://arxiv.org/html/2606.12922#bib.bib23)\)\.
AxisCategoryTopics \(U\.S\. Dataset\)Topics \(South Korean Dataset\)EconomicLeft / RightMarket Economytax policy, corporate tax, market regulation, financial regulation, antitrust policy, consumer costs, drug prices법인세 \(corporate tax\), 대기업\-중소기업 관계 \(large firms–SMEs relations\), 기업 규제 완화 \(corporate deregulation\), 금융시장 개입 \(financial market intervention\), 부동산 시장 규제 \(real estate regulation\)Trade /Energytrade agreements, fair trade, tariffs, outsourcing, domestic manufacturing, climate change, clean energy, fossil fuels, environmental protection무역 정책 \(trade policy\), 자유/보호 무역 \(free trade/protectionism\), 재생에너지 \(renewable energy\), 원자력 발전 \(nuclear power\), 환경 규제 \(environmental regulation\), 식량안보 \(food security\), 녹색산업 \(green industry\)Laborlabor unions, collective bargaining, minimum wage, paid leave, unemployment insurance, immigration and labor, workplace protections, care work노동시간 \(working hours\), 최저임금 \(minimum wage\), 임금 격차 \(wage gap\), 비정규직 \(non\-regular employment\), 노조·단체교섭 \(labor unions and collective bargaining\), 산업안전 \(industrial safety\), 청년 일자리 \(youth employment\)Welfare Statehealthcare access, medicaid, social security, affordable housing, childcare, public education, school choice, welfare requirements공공임대주택 \(public rental housing\), 건강보험 \(national health insurance\), 공공의료 \(public healthcare\), 기본소득 \(basic income\), 대학 등록금 \(university tuition\), 교육 평준화 \(educational equalization\), 복지 재정 \(welfare finance\)SocioculturalProgressive /ConservativeLaw and Orderimmigration, asylum policy, refugee resettlement, criminal justice, policing, civil liberties, human rights, national identity집회 및 시위 \(assemblies and protests\), 표현의 자유 \(freedom of expression\), 언론 규제 \(media regulation\), 형사처벌 \(criminal punishment\), 경찰 권한 \(police authority\), 검찰 개혁 \(prosecutorial reform\), 국가 상징 \(national symbols\)Gender /Minorities /Equalityabortion rights, reproductive healthcare, marriage equality, LGBTQ rights, religious freedom, educational equity, racial discrimination, anti\-discrimination law성차별\(gender discrimination\), 여성고용\(women’s employment\), 성별할당제\(gender quotas\), 이주민\(migrants\), 북한이탈주민\(North Korean defectors\), 장애인고용\(disability employment\), 성소수자\(sexual minorities\), 차별금지법\(anti\-discrimination law\)International Relationsalliances and NATO, international institutions, foreign aid, Israel\-Palestine relations, development cooperation, global health, Latin America policy, China and Indo\-Pacific대북 정책 \(North Korea policy\), 남북 교류협력 \(inter\-Korean exchange and cooperation\), 통일 정책 \(unification policy\), 한중일 관계 \(Korea\-China\-Japan relations\), 공공 외교 \(public diplomacy\), 개발 협력 \(development cooperation\)NationalDefense /Securitymilitary force, defense spending, nuclear weapons, nuclear non\-proliferation, counterterrorism, military readiness, China security policy, Middle East security, Russia and arms control한미 연합 군사훈련 \(Korea\-US joint military exercise\), 대북 안보정책 \(security policy toward North Korea\), 국가보안법 \(National Security Act\), 병력 감축 \(troop reduction\), 북핵 억제 \(North Korean nuclear deterrence\)Table 1:Political issue categories and representative topics covered in the U\.S\. and South Korean datasets\.
### 2\.3Political Bias Benchmarks
Existing benchmarks for political bias in LLMs have largely adapted survey instruments designed for human respondents, such as The Political Compass, Wahl\-O\-Mat, and American National Election Studies \(ANES\)\(Motokiet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib30); Exleret al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib59); Faulbornet al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib55); Rettenbergeret al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib53); Penget al\.,[2026](https://arxiv.org/html/2606.12922#bib.bib52)\)\. These evaluations typically prompt models to select predefined responses and infer political bias from the resulting response distributions\.
However, this paradigm has three limitations\. First, survey instruments cover only a limited number of issue areas\. Second, prompt\-based evaluations are sensitive to prompt wording, decoding choices, and refusal behavior, making measurements difficult to reproduce consistently across runs and studies\(Banget al\.,[2021](https://arxiv.org/html/2606.12922#bib.bib37); Röttgeret al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib56); Wrightet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib32); Azzopardi and Moshfeghi,[2025](https://arxiv.org/html/2606.12922#bib.bib60)\)\. Third, existing benchmarks remain heavily centered on U\.S\. and Western European contexts\. Although recent work has expanded to other sociocultural settings\(Thapaet al\.,[2023](https://arxiv.org/html/2606.12922#bib.bib40); Zhou and Zhang,[2024](https://arxiv.org/html/2606.12922#bib.bib50); Helweet al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib49)\), political bias in non\-Western environments remains substantially underexplored\.
Polar addresses these gaps by using option\-level likelihood comparisons rather than prompt\-based generation\. Its parallel U\.S\. and South Korean design enables cross\-linguistic and cross\-national comparisons beyond single\-context benchmarks\.
## 3Methods
We measure political bias using a two\-axis, eight\-category taxonomy derived from the Manifesto Project\(Lehmannet al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib78)\)and a reproducible option\-level likelihood procedure adapted from prior work on evaluating stereotype bias\.
### 3\.1Two\-Axis Political Taxonomy
We classify political issues along two ideological axes rather than a single left\-right continuum since political attitudes are empirically multidimensional\(Evanset al\.,[1996](https://arxiv.org/html/2606.12922#bib.bib41)\)\. Our taxonomy distinguishes between economic and sociocultural positions\. Economic positions concern issues such as redistribution, market regulation, and labor while sociocultural positions include issues such as national identity, cultural pluralism, and religious commitment\(Lachat,[2009](https://arxiv.org/html/2606.12922#bib.bib43); Feldman and Johnston,[2014](https://arxiv.org/html/2606.12922#bib.bib42)\)\.
Within this framework, we define eight issue categories grounded in the Manifesto coding scheme, a widely used framework for cross\-national analysis of political positions\. Manifesto codes provide both policy\-domain labels and ideological direction, allowing each category to correspond to a coherent policy area with a clearly defined ideological contrast\. This gives Polar a standardized basis for assigning labels across heterogeneous sources rather than relying on ad hoc issue classifications\. See Appendix[A](https://arxiv.org/html/2606.12922#A1)for full mapping details\.
We apply the same two\-axis structure to both the U\.S\. and South Korean datasets\. This shared structure keeps the datasets analytically comparable while allowing the concrete political content of each category to remain country\-specific\. Table[1](https://arxiv.org/html/2606.12922#S2.T1)provides representative topics covered by each category in the U\.S\. and South Korean datasets\. As a result, Polar supports cross\-lingual and cross\-national comparison and enables analysis of how model preferences vary across issue domains\.
### 3\.2Likelihood\-Based Measurement
Political bias evaluation requires a procedure that reduces sensitivity to prompt phrasing, decoding variability, and refusal behavior\(Banget al\.,[2021](https://arxiv.org/html/2606.12922#bib.bib37); Röttgeret al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib56)\)\. Free\-form generation is sensitive to these factors, which makes results difficult to reproduce\. We therefore use option\-level likelihood comparison\. For each instance, we compute the log\-likelihood of each candidate continuation given a shared context and treat the highest\-scoring continuation as the model’s choice\. Scoring fixed continuations removes decoding randomness and improves comparability across models\.
Our measurement procedure adapts the Context Association Test \(CAT\) framework introduced in StereoSet\(Nadeemet al\.,[2021](https://arxiv.org/html/2606.12922#bib.bib6)\)\. Each instance consists of a neutral context, two politically opposed continuations, and one semantically unrelated continuation\. The political continuations provide the bias signal, since their relative likelihood indicates the model’s directional preference along the relevant axis\. The unrelated continuation serves as a competence check\. A competent model should assign it lower likelihood than the politically relevant continuations\.
This design jointly measures political preference and language modeling competence\. A model is meaningfully balanced only when it assigns comparable likelihoods to the two opposing political continuations while also ranking them above the unrelated continuation\. Thus, the metric distinguishes ideological balance from cases in which a model appears neutral only because it fails to assign reliable likelihoods to politically relevant options\.
Section 4 describes how evaluation instances are constructed and validated, and Section 5 provides the likelihood computation and scoring procedure\.
Figure 1:Overview of Polar instance construction process\.
## 4The Polar Dataset
### 4\.1Overview
Polar contains 2,013 source instances across two political contexts: 1,004 U\.S\. instances written in English and 1,009 South Korean instances written in Korean\. Each source instance is derived from real political statements reflecting political discourse in the corresponding country and is paired with a translated counterpart in the other language, resulting in 4,026 evaluation instances\.
Each instance consists of a neutral context and three candidate continuations: two politically opposed continuations and one semantically unrelated continuation\. Models are evaluated by scoring each candidate continuation with log\-likelihood and selecting the highest\-scoring option\.
### 4\.2Source Materials
We collect political statements from primary and supplementary sources\. Our primary source is the Manifesto Project dataset\(Lehmannet al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib78)\), whose standardized policy codes indicate issue area and ideological direction\. Its database contains a vast amount of human\-coded statements from more than 1,300 political parties across 67 countries from 1945 to 2025\. For Polar, we use U\.S\. and South Korean documents from 2016 to 2024, allowing the benchmark to reflect recent political discourse in both countries\. Manifesto codes are used to assign statements to one of the eight issue categories and to determine their ideological orientation along the relevant axis\. Because Manifesto coverage is uneven across countries, time periods, and recent political issues, we supplement the source pool with official party press releases, legislative policy briefs, a public\-sector database\([AI Hub,](https://arxiv.org/html/2606.12922#bib.bib77)\), and news articles\. Supplementary statements are classified using the Manifesto codebook to maintain consistent labeling across the dataset\.
### 4\.3Construction Pipeline
Polar is constructed through five stages: instance allocation, opposing\-position pairing, instance construction, cross\-lingual translation, and review/revision\. Figure[1](https://arxiv.org/html/2606.12922#S3.F1)illustrates the overall construction process, using a labor\-related example to show how opposing statements are paired, converted into a neutral context with candidate continuations, and reviewed\.
#### Instance allocation across categories\.
Rather than allocating instances uniformly, we adjust the number of instances per category according to its level of political contestation\. We estimate contestation using bill passage rates in the U\.S\. Congress and the Korean National Assembly over the past ten years, based on official legislative records\([Library of Congress,](https://arxiv.org/html/2606.12922#bib.bib84);[National Assembly of the Republic of Korea,](https://arxiv.org/html/2606.12922#bib.bib85)\)\. Prior work treats legislative gridlock as an indicator of political conflict\(Binder,[1999](https://arxiv.org/html/2606.12922#bib.bib79); Jones,[2001](https://arxiv.org/html/2606.12922#bib.bib80); Agarwal,[2024](https://arxiv.org/html/2606.12922#bib.bib81)\), making bill passage rates a useful proxy for issue\-level contestation\.
For each country, we map bills to the eight issue categories and compute category\-level passage rates as the proportion of introduced bills that are passed\. We then allocate instances inversely proportional to passage rates, so that more contested categories receive more instances\. Appendix[B](https://arxiv.org/html/2606.12922#A2)reports the resulting category distribution and bill\-to\-category mapping procedure\.
#### Opposing\-position pairing\.
From the collected materials, we pair statements that take clear opposing positions on the same or closely related political issue along the relevant ideological axis\. For example, Figure[1](https://arxiv.org/html/2606.12922#S3.F1)shows two labor\-related statements: one supports stronger labor rights, while the other argues for reforming labor laws to reduce restrictions\. The pair addresses the same policy domain but represents opposite political positions\.
We exclude statements that only criticize opponents, praise the speaker’s own party, or provide factual information without an explicit policy position\. Within each category, we seek broad topic coverage while avoiding overrepresentation of a single policy debate, unless it appears through substantively different policy angles\. Pairing is performed manually by the authors, with U\.S\. and South Korean instances reviewed by annotators familiar with the corresponding political contexts\. Appendix[C\.1](https://arxiv.org/html/2606.12922#A3.SS1)provides more details\.
#### Instance construction\.
We convert each statement pair into an evaluation instance in three steps\. First, we write a neutral context that introduces the issue without favoring either position\. Second, we derive two political continuations from the paired source statements, preserving their substantive claims and rhetorical intensity while minimally paraphrasing them so that each continuation forms a complete sentence when joined with the context\. Third, we add a semantically unrelated continuation that is grammatically compatible with the context but disconnected from the political issue\. In Figure[1](https://arxiv.org/html/2606.12922#S3.F1), the matched pair is converted into an instance centered on a modern labor framework, with two opposed political continuations and one unrelated continuation\.
To ensure that instances measure political positions rather than associations with specific political figures or contemporary events, we generalize overly specific references where necessary\. For example, named politicians, party\-specific slogans, branded policy labels, and numerical claims are replaced with more general descriptions or removed\. We constrain instance length to roughly 200 characters for English and 100 characters for Korean\.
#### Cross\-lingual translation\.
To evaluate political bias across both context and language, each source instance is translated into the other language: U\.S\. instances are translated into Korean, and South Korean instances are translated into English\. The final dataset therefore contains 4,026 instances, consisting of 2,013 original source instances and 2,013 translated instances\. Translation is performed using GPT\-5\.5\. Because subtle shifts in wording can undermine cross\-lingual comparisons, all translated instances are audited for faithfulness to the source claim, preservation of rhetorical intensity, and fluency in the target language\. Translations that do not meet these criteria are manually revised\.
#### Review and revision\.
Each instance is drafted by one author and reviewed by another, drawing on the authors’ fluency in English and Korean and domain expertise in law and public policy\. Reviewers check the political neutrality of the context, connection between the context and political options, preservation of the source statements’ meaning, grammatical correctness of all continuations, and semantic unrelatedness of the competence\-check continuation\. Disagreements are resolved through discussion, and instances that fail to satisfy the criteria are revised or removed\. The detailed review process is described in Appendix[C\.2](https://arxiv.org/html/2606.12922#A3.SS2)\.
## 5Experiments
### 5\.1Experimental Setup
#### Models\.
We evaluate 38 LLMs covering globally deployed and Korean\-specialized models\. Global models include Qwen3\(Yanget al\.,[2025a](https://arxiv.org/html/2606.12922#bib.bib67)\), Llama 3\(Grattafioriet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib69)\), and Mistral\(Jianget al\.,[2023](https://arxiv.org/html/2606.12922#bib.bib68)\); Korean\-specialized models include Kanana 1\.5\(Baket al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib70)\), EXAONE\-4\.0\(Baeet al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib73)\), A\.X\-4\.0\(SK Telecom,[2025a](https://arxiv.org/html/2606.12922#bib.bib74)\), HyperCLOVAX\(NAVER Cloud HyperCLOVA X Team,[2025](https://arxiv.org/html/2606.12922#bib.bib72)\), Solar\(Kimet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib76)\), Mi:dm\(Shinet al\.,[2026](https://arxiv.org/html/2606.12922#bib.bib71)\), and Ko\-GPT\-Trinity\(SK Telecom,[2025b](https://arxiv.org/html/2606.12922#bib.bib75)\)\. This selection enables comparisons across model scales, developers, and language specializations\. Full model details are provided in Appendix[D](https://arxiv.org/html/2606.12922#A4)\.
Figure 2:Political positions of LLMs on the economic and sociocultural axes\. The displayed range is set to\[−0\.5,0\.5\]\[\-0\.5,0\.5\]for readability, while the full position range is\[−1,1\]\[\-1,1\]\. Appendix[F\.1](https://arxiv.org/html/2606.12922#A6.SS1)provides the full\-scale plot\.
#### Evaluation metrics\.
We evaluate models with the LM Evaluation Harness\(Bidermanet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib82)\)\. Each candidate continuation is scored by length\-normalized log\-likelihood, and the highest\-scoring option is treated as the model’s choice\. Option 1 corresponds to the left position on the economic axis and the progressive position on the sociocultural axis, option 2 to the right/conservative position, and option 3 to a semantically unrelated option\.
Polar reports political position and the ICAT \(Idealized CAT\) score\(Nadeemet al\.,[2021](https://arxiv.org/html/2606.12922#bib.bib6)\)\. The political position measures which of the two political options the model prefers within an axis or category\. For each instance, we compare the length\-normalized log\-likelihoods of option 1 and option 2\. We definep1p\_\{1\}andp2p\_\{2\}as the proportions of instances where option 1 and option 2 receive the higher score, respectively \(e\.g\.,p1=0\.6p\_\{1\}=0\.6,p2=0\.4p\_\{2\}=0\.4\)\. The position is then defined as:
Position=p2−p1\\mathrm\{Position\}=p\_\{2\}\-p\_\{1\}
This comparison excludes option 3 and focuses only on the relative preference between the two political options\. The position score ranges from−1\-1to11\. A value close to−1\-1indicates a left or progressive preference, while a value close to11indicates a right or conservative preference\. For example, if option 1 \(left\) receives a higher likelihood than option 2 \(right\) in60%60\\%of economic\-axis instances, and option 2 receives a higher likelihood in40%40\\%, thenp1=0\.6p\_\{1\}=0\.6andp2=0\.4p\_\{2\}=0\.4\. The resulting position score isp2−p1=−0\.2p\_\{2\}\-p\_\{1\}=\-0\.2\.
ICAT combines language modeling competence and ideological balance\. LetLMS\\mathrm\{LMS\}denote the language modeling score andNS\\mathrm\{NS\}denote the neutrality score\. Following StereoSet, ICAT is defined as:
ICAT=LMS×NS\\mathrm\{ICAT\}=\\mathrm\{LMS\}\\times\\mathrm\{NS\}The language modeling score measures whether the model assigns higher likelihood to politically meaningful options than to the unrelated option\. Letcic\_\{i\}denote the model’s choice for instanceii\.LMS\\mathrm\{LMS\}is defined as the percentage of instances where option 3 is not selected:
LMS=100×1N∑i=1N𝟏\[ci≠3\]\\mathrm\{LMS\}=100\\times\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\[c\_\{i\}\\neq 3\]For example, if option 3 is selected in 20% of instances,LMS\\mathrm\{LMS\}is 80\.
The neutrality score captures how close the model’s position is to zero:
NS=1−\|Position\|\\mathrm\{NS\}=1\-\|\\mathrm\{Position\}\|HigherNS\\mathrm\{NS\}indicates greater balance between two opposed options\. We compute NS at the category\-level and average across categories within each axis to prevent opposite directional biases from canceling out\. We then average the economic and sociocultural ICAT scores to obtain the final score\.
### 5\.2Results
#### Overall political position across contexts\.
We map each model to a two\-dimensional political space using its economic and sociocultural position scores\. Figure[2](https://arxiv.org/html/2606.12922#S5.F2)shows markedly different distributions across the two datasets\. On the U\.S\. dataset, all models fall in the negative region on both axes, with economic scores ranging from−0\.36\-0\.36to−0\.03\-0\.03and sociocultural scores ranging from−0\.23\-0\.23to−0\.05\-0\.05\. In contrast, on the South Korean dataset, models cluster near the origin and appear on both sides of each axis, with economic scores ranging from−0\.15\-0\.15to0\.160\.16and sociocultural scores ranging from−0\.10\-0\.10to0\.100\.10\. Detailed model\-level results are provided in Appendix[F\.5](https://arxiv.org/html/2606.12922#A6.SS5)\.
These results show a consistent left\-progressive tendency across all evaluated models on the U\.S\. dataset, while model positions are more centered and mixed on the South Korean dataset\. This pattern is consistent with prior work reporting left\-leaning tendencies in LLMs, particularly in U\.S\. and Western European settings\(Bernardelleet al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib27); Becchetti and Solferino,[2026](https://arxiv.org/html/2606.12922#bib.bib44); Shuet al\.,[2026](https://arxiv.org/html/2606.12922#bib.bib26)\)\. One possible interpretation is that pretraining corpora or alignment procedures may increase the likelihood of policy language associated with particular ideological or demographic groups\(Rozado,[2023](https://arxiv.org/html/2606.12922#bib.bib33); Santurkaret al\.,[2023](https://arxiv.org/html/2606.12922#bib.bib45); Fulayet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib24)\)\. The contrast between the two datasets indicates that political bias should be assessed across multiple political contexts, as model preferences identified in one context may not generalize to others\.
#### Category\-level bias patterns\.
Category\-level results reveal political preferences that are not visible in aggregate axis\-level scores\. In the U\.S\. dataset, category\-level mean scores are left or progressive leaning across all eight issue categories, although their strength varies by domain\.Welfare StateandGender / Minorities / Equalityshow the strongest left\-progressive tendencies, with mean position scores of−0\.43\-0\.43and−0\.32\-0\.32\. In contrast,International RelationsandNational Defense / Securityare closer to the origin, with mean scores of−0\.09\-0\.09and−0\.04\-0\.04\. Detailed plots and results are provided in Appendix[F\.2](https://arxiv.org/html/2606.12922#A6.SS2)and Appendix[F\.6](https://arxiv.org/html/2606.12922#A6.SS6)\.
These results suggest that even when the overall direction of political preference is consistent, its strength depends on the issue being evaluated\. One possible interpretation is that left\-progressive expressions in topics such ashealthcare accessandminority protectionoverlap with safety\- and inclusion\-oriented language frequently emphasized during instruction tuning and alignment\(Ouyanget al\.,[2022](https://arxiv.org/html/2606.12922#bib.bib46)\)\. By contrast, categories such asInternational Relationsmay contain policy language that is less strongly associated with a single ideological direction in training or alignment data\.
The South Korean dataset shows a more mixed pattern\. Most categories remain close to the origin, but several categories exhibit clear directional tendencies\. For example,Market Economyis right\-leaning, with a mean score of0\.180\.18, whileLaboris left\-leaning, with a mean score of−0\.16\-0\.16\. These opposing category\-level tendencies can partially cancel out at the axis level, making models appear more neutral in aggregate than they are within specific issue domains\.
#### Korean\-specialized vs global models\.
We compare Korean\-specialized and globally deployed models to examine how model group differences vary across political contexts\. We quantify these differences using the average category\-level gap, defined as the mean absolute difference between the two groups’ position scores across categories\.
On the U\.S\. dataset, the two groups show broadly similar category\-level patterns\. Both exhibit stronger progressive tendencies inWelfare StateandGender / Minorities / Equalityand remain closer to the origin inInternational RelationsandNational Defense / Security\. The average category\-level gap is0\.030\.03, suggesting that Korean\-specialized and globally deployed models produce similar patterns in U\.S\. political content\.
On the South Korean dataset, the gap increases to0\.100\.10\. The largest gaps appear inMarket EconomyandTrade / Energy, where Korean\-specialized models score0\.070\.07and−0\.08\-0\.08respectively, while global models score0\.250\.25and0\.070\.07\. Detailed group\-level plots and values are provided in Appendix[F\.3](https://arxiv.org/html/2606.12922#A6.SS3)\.
These results indicate that Korean\-specialized and global models are more similar on U\.S\. political content than on South Korean political content\. One possible explanation is that South Korean political expressions reflect local issue framing, media discourse, and policy language, which may be captured differently by models trained more heavily for Korean language usage\. Many Korean\-specialized models are built on open\-weight global base models, which may explain why they retain similar tendencies to global models on U\.S\. content\(Kimet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib76); SK Telecom,[2025a](https://arxiv.org/html/2606.12922#bib.bib74)\)\.
Figure 3:ICAT scores across models on the two datasets\. The vertical dashed line indicates the mean ICAT score\.
#### ICAT scores\.
We examine ICAT scores, which combine ideological balance with language modeling competence\. A high ICAT score requires both low directional political preference and high LMS\. Comparable likelihoods for the two political options count as meaningful balance only when the model also ranks them above the unrelated option\.
Most models achieve high LMS, indicating that they generally distinguish politically relevant continuations from unrelated continuations\. EXAONE\-4\.0\-1\.2B is a notable exception, with average LMS scores of75\.5875\.58and75\.5175\.51for each dataset, suggesting weaker language modeling competence in this evaluation setting\.
ICAT scores differ substantially between the two datasets\. The mean ICAT score is84\.8984\.89on the South Korean dataset, compared with75\.6875\.68on the U\.S\. dataset\. This difference mainly reflects the consistent left\-progressive preferences on the U\.S\. dataset, which reduce the neutrality component even when models maintain high LMS\. Thus, ICAT complements position scores by indicating whether low directional preference is supported by adequate language\-modeling competence\. Appendix[F\.5](https://arxiv.org/html/2606.12922#A6.SS5)provides full metric values for each model\.
#### Effect of presentation language\.
We define presentation language as the language in which the same political content is evaluated, independent of the political context from which the content was originally drawn\. To examine its effect, we compare original and translated Polar instances: U\.S\. political content in English and Korean, and South Korean content in Korean and English\.
Figure[2](https://arxiv.org/html/2606.12922#S5.F2)shows that presentation language substantially affects measured political preferences\. When South Korean political content is presented in English, model distributions shift from near the origin toward the left\-progressive direction\. Conversely, when U\.S\. political content is presented in Korean, the originally left\-progressive distribution moves closer to the origin\. The shift is more pronounced on the economic axis than on the sociocultural axis\. Detailed differences across models and categories are provided in Appendix[F](https://arxiv.org/html/2606.12922#A6)\.
The results indicate that measured bias is influenced not only by political content but also by the language of presentation\. A plausible explanation is that the presentation language alters the linguistic and cultural cues models use to evaluate political content, leading to the same underlying content receiving different relative likelihoods in English and Korean presentations\(Liet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib91); Xuet al\.,[2025](https://arxiv.org/html/2606.12922#bib.bib90)\)\.
## 6Conclusion
This work introduces Polar, a two\-axis, eight\-category benchmark designed to evaluate political bias in LLMs across U\.S\. and South Korean political contexts\. By leveraging option\-level likelihoods, Polar measures directional political preferences while concurrently assessing language modeling competence\. An analysis of 38 LLMs reveals consistent leftward progressiveness in choices on U\.S\. political content, whereas choices on South Korean content are more centered, mixed, and dependent on specific categories\. Language translation experiments indicate that the language of presentation affects measured bias, suggesting that models may assign different likelihoods to identical political content depending on linguistic and cultural framing\. These findings highlight the need for multilingual, cross\-contextual evaluation to achieve a comprehensive understanding of political bias in LLMs\.
## Limitations
Polar has three main limitations\. First, it reflects U\.S\. and South Korean political contexts from 2016 to 2025\. Since political coalitions, issue salience, and public debates change over time, future work should update the dataset to capture emerging issues and shifts in party positions\. Second, Polar covers only two countries\. Extending it to additional countries, languages, and cultural settings would support broader multilingual and cross\-cultural evaluation\. Third, our option\-level likelihood approach requires access to token\-level probabilities\. As a result, the method is directly applicable to open\-weight models but not to closed\-source systems that expose only generated outputs\. We seek to develop complementary evaluation methods for such models as future work\.
## Ethical Considerations
This work evaluates political bias in LLMs, a politically and socially sensitive topic\. Our goal is not to endorse or oppose any party, ideology, or policy position, but to measure relative model preferences between politically opposed options under a controlled benchmark setting\. Results should be interpreted as likelihood\-based preferences under Polar, not as evidence that models possess political beliefs, intentions, or stable ideological commitments\.
Political bias benchmarks may be misused if results are overgeneralized or treated as fixed model properties\. They may also unfairly label particular models, languages, or communities\. To reduce this risk, we report bias as relative preference within specific political contexts, languages, and issue categories\. Our findings show that measured bias varies across context, category, model group, and presentation language, cautioning against broad conclusions from a single benchmark\.
Polar does not contain private user data, and our experiments do not involve human users\. The benchmark is constructed from public political materials rewritten into neutral contexts and politically opposed continuations\. We release Polar for reproducible, diagnostic, and comparative analysis of political bias in LLMs, not for ranking models as politically acceptable or unacceptable\.
## Acknowledgments
This work was partially supported by the National Research Foundation of Korea \(NRF\) under Grant No\. RS\-2023\-00222663 \(Center for Optimizing Hyperscale AI Models and Platforms\) and under Grant No\. A400\-20260031, and by the Institute for Information and Communications Technology Promotion \(IITP\) under Grant No\. 2018\-0\-00581 \(CUDA Programming Environment for FPGA Clusters\) and No\. RS\-2025\-02304554 \(Efficient and Scalable Framework for AI Heterogeneous Cluster Systems\), all funded by the Ministry of Science and ICT \(MSIT\) of Korea\. It was also partially supported by the Korea Health Industry Development Institute \(KHIDI\) under Grant No\. RS\-2025\-25454559 \(Frailty Risk Assessment and Intervention Leveraging Multimodal Intelligence for Networked Deployment in Community Care\), funded by the Ministry of Health and Welfare \(MOHW\) of Korea\. Additional support was provided by the BK21 Plus Program for Innovative Data Science Talent Education \(Department of Data Science, Seoul National University, No\. 5199990914569\) and the BK21 FOUR Program for Intelligent Computing \(Department of Computer Science and Engineering, Seoul National University, No\. 4199990214639\), both funded by the Ministry of Education \(MOE\) of Korea\. This work was also partially supported by the Advanced GPU Utilization Support Program, funded by the Ministry of Science and ICT \(MSIT\) of Korea and operated by the National IT Industry Promotion Agency \(NIPA\)\. Research facilities were provided by the Institute of Computer Technology \(ICT\) at Seoul National University\.
## References
- K\. Agarwal \(2024\)How does political polarization impact legislative gridlock and policy\-making processes\.IOSR Journal of Humanities and Social Science29,pp\. 53–64\.External Links:[Document](https://dx.doi.org/10.9790/0837-2910065364)Cited by:[§4\.3](https://arxiv.org/html/2606.12922#S4.SS3.SSS0.Px1.p1.1)\.
- \[2\]AI HubAI Hub\.Note:Accessed: November 2025External Links:[Link](https://aihub.or.kr/)Cited by:[§4\.2](https://arxiv.org/html/2606.12922#S4.SS2.p1.1)\.
- Anthropic \(2025\)Measuring political bias in Claude\.Note:Accessed: May 2026External Links:[Link](https://www.anthropic.com/news/political-even-handedness)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p2.1)\.
- L\. P\. Argyle, E\. C\. Busby, N\. Fulda, J\. R\. Gubler, C\. Rytting, and D\. Wingate \(2023\)Out of one, many: using language models to simulate human samples\.Political Analysis31\(3\),pp\. 337–351\.External Links:[Document](https://dx.doi.org/10.1017/pan.2023.2)Cited by:[§2\.2](https://arxiv.org/html/2606.12922#S2.SS2.p1.1)\.
- L\. Azzopardi and Y\. Moshfeghi \(2025\)POW: political overton windows of large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 24767–24773\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1347/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1347),ISBN 979\-8\-89176\-335\-7Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p2.1)\.
- K\. Bae, E\. Choi, K\. Choi, S\. J\. Choi, Y\. Choi, K\. Han, S\. Hong, J\. Hwang, T\. Hwang, J\. Jang,et al\.\(2025\)EXAONE 4\.0: unified large language models integrating non\-reasoning and reasoning modes\.arXiv preprint arXiv:2507\.11407\.Cited by:[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px1.p1.1)\.
- Y\. Bak, H\. Lee, M\. Ryu, J\. Ham, S\. Jung, D\. W\. Nam, T\. Eo, D\. Lee, D\. Jung, B\. Kim,et al\.\(2025\)Kanana: compute\-efficient bilingual language models\.arXiv preprint arXiv:2502\.18934\.Cited by:[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px1.p1.1)\.
- Y\. Bang, N\. Lee, E\. Ishii, A\. Madotto, and P\. Fung \(2021\)Assessing political prudence of open\-domain chatbots\.InProceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue,H\. Li, G\. Levow, Z\. Yu, C\. Gupta, B\. Sisman, S\. Cai, D\. Vandyke, N\. Dethlefs, Y\. Wu, and J\. J\. Li \(Eds\.\),Singapore and Online,pp\. 548–555\.External Links:[Link](https://aclanthology.org/2021.sigdial-1.57/),[Document](https://dx.doi.org/10.18653/v1/2021.sigdial-1.57)Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p2.1),[§3\.2](https://arxiv.org/html/2606.12922#S3.SS2.p1.1)\.
- J\. Batzner, V\. Stocker, S\. Schmid, and G\. Kasneci \(2025\)GermanPartiesQA: benchmarking commercial large language models and ai companions for political alignment and sycophancy\.InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society,Vol\.8,pp\. 330–342\.Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p3.1)\.
- L\. Becchetti and N\. Solferino \(2025\)Unveiling biases in ai: chatgpt’s political economy perspectives and human comparisons\.arXiv preprint arXiv:2503\.05234\.Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p2.1)\.
- L\. Becchetti and N\. Solferino \(2026\)Political biases in chatgpt: insights from comparative analysis with human responses\.Economia Politica43\(1\),pp\. 285–326\.External Links:[Document](https://dx.doi.org/10.1007/s40888-025-00384-z),[Link](https://doi.org/10.1007/s40888-025-00384-z),ISSN 1973\-820XCited by:[§2\.2](https://arxiv.org/html/2606.12922#S2.SS2.p1.1),[§5\.2](https://arxiv.org/html/2606.12922#S5.SS2.SSS0.Px1.p2.1)\.
- E\. M\. Bender, T\. Gebru, A\. McMillan\-Major, and S\. Shmitchell \(2021\)On the dangers of stochastic parrots: can language models be too big?\.InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency,FAccT ’21,New York, NY, USA,pp\. 610–623\.External Links:ISBN 9781450383097,[Link](https://doi.org/10.1145/3442188.3445922),[Document](https://dx.doi.org/10.1145/3442188.3445922)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p1.1)\.
- P\. Bernardelle, L\. Fröhling, S\. Civelli, R\. Lunardi, K\. Roitero, and G\. Demartini \(2025\)Mapping and influencing the political ideology of large language models using synthetic personas\.InCompanion Proceedings of the ACM on Web Conference 2025,pp\. 864–867\.Cited by:[§2\.2](https://arxiv.org/html/2606.12922#S2.SS2.p1.1),[§5\.2](https://arxiv.org/html/2606.12922#S5.SS2.SSS0.Px1.p2.1)\.
- S\. Biderman, H\. Schoelkopf, L\. Sutawika, L\. Gao, J\. Tow, B\. Abbasi, A\. F\. Aji, P\. S\. Ammanamanchi, S\. Black, J\. Clive,et al\.\(2024\)Lessons from the trenches on reproducible evaluation of language models\.arXiv preprint arXiv:2405\.14782\.Cited by:[Appendix E](https://arxiv.org/html/2606.12922#A5.p1.1),[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px2.p1.1)\.
- S\. A\. Binder \(1999\)The dynamics of legislative gridlock, 1947–96\.American Political Science Review93\(3\),pp\. 519–533\.External Links:[Document](https://dx.doi.org/10.2307/2585572)Cited by:[§4\.3](https://arxiv.org/html/2606.12922#S4.SS3.SSS0.Px1.p1.1)\.
- S\. L\. Blodgett, S\. Barocas, H\. Daumé III, and H\. Wallach \(2020\)Language \(technology\) is power: a critical survey of “bias” in NLP\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 5454–5476\.External Links:[Link](https://aclanthology.org/2020.acl-main.485/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.485)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p1.1)\.
- T\. Ceron, N\. Falk, A\. Barić, D\. Nikolaev, and S\. Padó \(2024\)Beyond prompt brittleness: evaluating the reliability and consistency of political worldviews in llms\.Transactions of the Association for Computational Linguistics12,pp\. 1378–1400\.Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p2.1)\.
- J\. Chen, K\. Jong, A\. G\. Poole, J\. Burakowski, E\. E\. Nosti, J\. Windt, and C\. Wang \(2026\)Uncovering political bias in large language models using parliamentary voting records\.arXiv preprint arXiv:2601\.08785\.External Links:[Link](https://arxiv.org/abs/2601.08785)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p3.1)\.
- S\. Elkins, E\. Kochmar, I\. Serban, and J\. C\. K\. Cheung \(2023\)How useful are educational questions generated by large language models?\.InArtificial Intelligence in Education\. Posters and Late Breaking Results, Workshops and Tutorials, Industry and Innovation Tracks, Practitioners, Doctoral Consortium and Blue Sky,N\. Wang, G\. Rebolledo\-Mendez, V\. Dimitrova, N\. Matsuda, and O\. C\. Santos \(Eds\.\),Cham,pp\. 536–542\.External Links:ISBN 978\-3\-031\-36336\-8Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p1.1)\.
- G\. Evans, A\. Heath, and M\. Lalljee \(1996\)Measuring left\-right and libertarian\-authoritarian values in the british electorate\.British Journal of Sociology,pp\. 93–112\.Cited by:[§3\.1](https://arxiv.org/html/2606.12922#S3.SS1.p1.1)\.
- D\. Exler, M\. Schutera, M\. Reischl, and L\. Rettenberger \(2025\)Large means left: political bias in large language models increases with their number of parameters\.arXiv preprint arXiv:2505\.04393\.Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p1.1)\.
- M\. Faulborn, I\. Sen, M\. Pellert, A\. Spitz, and D\. Garcia \(2025\)Only a little to the left: a theory\-grounded measure of political bias in large language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 31684–31704\.External Links:[Link](https://aclanthology.org/2025.acl-long.1529/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1529),ISBN 979\-8\-89176\-251\-0Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p1.1)\.
- L\. Feldman \(2011\)Partisan differences in opinionated news perceptions: a test of the hostile media effect\.Political Behavior33\(3\),pp\. 407–432\.External Links:[Document](https://dx.doi.org/10.1007/s11109-010-9139-4),[Link](https://doi.org/10.1007/s11109-010-9139-4),ISSN 1573\-6687Cited by:[§2\.2](https://arxiv.org/html/2606.12922#S2.SS2.p1.1)\.
- S\. Feldman and C\. Johnston \(2014\)Understanding the determinants of political ideology: implications of structural complexity\.Political Psychology35\(3\),pp\. 337–358\.External Links:[Document](https://dx.doi.org/10.1111/pops.12055),[Link](https://onlinelibrary.wiley.com/doi/abs/10.1111/pops.12055),https://onlinelibrary\.wiley\.com/doi/pdf/10\.1111/pops\.12055Cited by:[§3\.1](https://arxiv.org/html/2606.12922#S3.SS1.p1.1)\.
- S\. Feng, C\. Y\. Park, Y\. Liu, and Y\. Tsvetkov \(2023\)From pretraining data to language models to downstream tasks: tracking the trails of political biases leading to unfair nlp models\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 11737–11762\.Cited by:[§2\.2](https://arxiv.org/html/2606.12922#S2.SS2.p1.1)\.
- S\. Fulay, W\. Brannon, S\. Mohanty, C\. Overney, E\. Poole\-Dayan, D\. Roy, and J\. Kabbara \(2024\)On the relationship between truth and political bias in language models\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 9004–9018\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.508/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.508)Cited by:[§5\.2](https://arxiv.org/html/2606.12922#S5.SS2.SSS0.Px1.p2.1)\.
- I\. O\. Gallegos, R\. A\. Rossi, J\. Barrow, M\. M\. Tanjim, S\. Kim, F\. Dernoncourt, T\. Yu, R\. Zhang, and N\. K\. Ahmed \(2024\)Bias and fairness in large language models: a survey\.Computational Linguistics50\(3\),pp\. 1097–1179\.External Links:[Link](https://aclanthology.org/2024.cl-3.8/),[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00524)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p1.1)\.
- A\. Gilson, C\. W\. Safranek, T\. Huang, V\. Socrates, L\. Chi, R\. A\. Taylor, and D\. Chartash \(2023\)How does chatgpt perform on the united states medical licensing examination? the implications of large language models for medical education and knowledge assessment\.JMIR Med Educ9,pp\. e45312\.External Links:ISSN 2369\-3762,[Document](https://dx.doi.org/10.2196/45312),[Link](https://mededu.jmir.org/2023/1/e45312)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.\(2024\)The Llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px1.p1.1)\.
- R\. Gubelmann and G\. Karray \(2025\)Assessing reliability and political bias in LLMs’ judgements of formal and material inferences with partisan conclusions\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 30005–30031\.External Links:[Link](https://aclanthology.org/2025.acl-long.1450/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1450),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p2.1)\.
- K\. Hackenburg, L\. Ibrahim, B\. M\. Tappin, and M\. Tsakiris \(2025\)Comparing the persuasiveness of role\-playing large language models and human experts on polarized u\.s\. political issues\.AI Soc\.41\(1\),pp\. 351–361\.External Links:ISSN 0951\-5666,[Link](https://doi.org/10.1007/s00146-025-02464-x),[Document](https://dx.doi.org/10.1007/s00146-025-02464-x)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p2.1)\.
- J\. Hartmann, J\. Schwenzow, and M\. Witte \(2023\)The political ideology of conversational ai: converging evidence on chatgpt’s pro\-environmental, left\-libertarian orientation\.arXiv preprint arXiv:2301\.01768\.External Links:[Link](https://arxiv.org/abs/2301.01768)Cited by:[§2\.2](https://arxiv.org/html/2606.12922#S2.SS2.p1.1)\.
- C\. Helwe, O\. Balalau, and D\. Ceolin \(2025\)Navigating the political compass: evaluating multilingual LLMs across languages and nationalities\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 17179–17204\.External Links:[Link](https://aclanthology.org/2025.findings-acl.883/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.883),ISBN 979\-8\-89176\-256\-5Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p2.1)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de Las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. Sayed \(2023\)Mistral 7b\.arXiv preprint arXiv:2310\.06825\.Cited by:[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px1.p1.1)\.
- D\. Jones \(2001\)Party polarization and legislative gridlock\.Political Research Quarterly54,pp\. 125–141\.External Links:[Document](https://dx.doi.org/10.1177/106591290105400107)Cited by:[§4\.3](https://arxiv.org/html/2606.12922#S4.SS3.SSS0.Px1.p1.1)\.
- S\. Kim, D\. Kim, C\. Park, W\. Lee, W\. Song, Y\. Kim, H\. Kim, Y\. Kim, H\. Lee, J\. Kim, C\. Ahn, S\. Yang, S\. Lee, H\. Park, G\. Gim, M\. Cha, H\. Lee, and S\. Kim \(2024\)SOLAR 10\.7B: scaling large language models with simple yet effective depth up\-scaling\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 6: Industry Track\),Y\. Yang, A\. Davani, A\. Sil, and A\. Kumar \(Eds\.\),Mexico City, Mexico,pp\. 23–35\.External Links:[Link](https://aclanthology.org/2024.naacl-industry.3/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-industry.3)Cited by:[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px1.p1.1),[§5\.2](https://arxiv.org/html/2606.12922#S5.SS2.SSS0.Px3.p4.1)\.
- A\. Kuenzler and S\. Schmid \(2026\)Communication bias in large language models: a regulatory perspective\.Communications of the ACM\.Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p2.1)\.
- C\. V\. Kumar, A\. Urlana, G\. Kanumolu, B\. M\. Garlapati, and P\. Mishra \(2025\)No LLM is free from bias: a comprehensive study of bias evaluation in large language models\.arXiv preprint arXiv:2503\.11985\.Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p1.1)\.
- R\. Lachat \(2009\)Is Left\-Right from Circleland? the issue basis of citizens’ ideological self\-placement\.CIS Working PaperTechnical Report51,Center for Comparative and International Studies\.External Links:[Link](https://ssrn.com/abstract=1551837),[Document](https://dx.doi.org/10.2139/ssrn.1551837)Cited by:[§3\.1](https://arxiv.org/html/2606.12922#S3.SS1.p1.1)\.
- P\. Lehmann, S\. Franzmann, D\. Al\-Gaddooa, T\. Burst, C\. Ivanusch, S\. Regel, F\. Riethmüller, A\. Volkens, B\. Weßels, and L\. Zehnter \(2025\)The manifesto data collection\. manifesto project \(MRG/CMP/MARPOR\)\. version 2025a\.Note:Wissenschaftszentrum Berlin für Sozialforschung / Göttinger Institut für DemokratieforschungExternal Links:[Document](https://dx.doi.org/10.25522/manifesto.mpds.2025a),[Link](https://doi.org/10.25522/manifesto.mpds.2025a)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p4.1),[§3](https://arxiv.org/html/2606.12922#S3.p1.1),[§4\.2](https://arxiv.org/html/2606.12922#S4.SS2.p1.1)\.
- C\. Li, M\. Chen, J\. Wang, S\. Sitaram, and X\. Xie \(2024\)Culturellm: incorporating cultural differences into large language models\.Advances in Neural Information Processing Systems37,pp\. 84799–84838\.Cited by:[§5\.2](https://arxiv.org/html/2606.12922#S5.SS2.SSS0.Px5.p3.1)\.
- T\. Li, D\. Khashabi, T\. Khot, A\. Sabharwal, and V\. Srikumar \(2020\)UNQOVERing stereotyping biases via underspecified questions\.InFindings of the Association for Computational Linguistics: EMNLP 2020,T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 3475–3489\.External Links:[Link](https://aclanthology.org/2020.findings-emnlp.311/),[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.311)Cited by:[§2\.1](https://arxiv.org/html/2606.12922#S2.SS1.p1.1)\.
- \[43\]Library of CongressCongress\.gov\.Note:Accessed: January 2026External Links:[Link](https://www.congress.gov/)Cited by:[Appendix B](https://arxiv.org/html/2606.12922#A2.p1.1),[§4\.3](https://arxiv.org/html/2606.12922#S4.SS3.SSS0.Px1.p1.1)\.
- R\. Liu, C\. Jia, J\. Wei, G\. Xu, and S\. Vosoughi \(2022\)Quantifying and alleviating political bias in language models\.Artificial Intelligence304,pp\. 103654\.External Links:ISSN 0004\-3702,[Document](https://dx.doi.org/10.1016/j.artint.2021.103654),[Link](https://www.sciencedirect.com/science/article/pii/S0004370221002058)Cited by:[§2\.2](https://arxiv.org/html/2606.12922#S2.SS2.p1.1)\.
- F\. Motoki, V\. Pinho Neto, and V\. Rodrigues \(2024\)More human than human: measuring ChatGPT political bias\.Public Choice198\(1\),pp\. 3–23\.Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p1.1)\.
- M\. Nadeem, A\. Bethke, and S\. Reddy \(2021\)StereoSet: measuring stereotypical bias in pretrained language models\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 5356–5371\.External Links:[Link](https://aclanthology.org/2021.acl-long.416/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.416)Cited by:[§2\.1](https://arxiv.org/html/2606.12922#S2.SS1.p1.1),[§3\.2](https://arxiv.org/html/2606.12922#S3.SS2.p2.1),[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px2.p2.4)\.
- N\. Nangia, C\. Vania, R\. Bhalerao, and S\. R\. Bowman \(2020\)CrowS\-pairs: a challenge dataset for measuring social biases in masked language models\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 1953–1967\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.154/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.154)Cited by:[§2\.1](https://arxiv.org/html/2606.12922#S2.SS1.p1.1)\.
- \[48\]National Assembly of the Republic of KoreaNational assembly bill information system\.Note:Accessed: December 2025External Links:[Link](https://likms.assembly.go.kr/bill/bi/main/mainPage.do)Cited by:[Appendix B](https://arxiv.org/html/2606.12922#A2.p1.1),[§4\.3](https://arxiv.org/html/2606.12922#S4.SS3.SSS0.Px1.p1.1)\.
- NAVER Cloud HyperCLOVA X Team \(2025\)HyperCLOVA X Think technical report\.arXiv preprint arXiv:2506\.22403\.Cited by:[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px1.p1.1)\.
- R\. Navigli, S\. Conia, and B\. Ross \(2023\)Biases in large language models: origins, inventory, and discussion\.J\. Data and Information Quality15\(2\)\.External Links:ISSN 1936\-1955,[Link](https://doi.org/10.1145/3597307),[Document](https://dx.doi.org/10.1145/3597307)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p1.1)\.
- OECD \(2025\)Governing with artificial intelligence: the state of play and way forward in core government functions\.Technical reportOECD Publishing,Paris\.External Links:[Document](https://dx.doi.org/10.1787/795de142-en),[Link](https://doi.org/10.1787/795de142-en)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p1.1)\.
- OpenAI \(2025\)Defining and evaluating political bias in LLMs\.Note:Accessed: May 2026External Links:[Link](https://openai.com/index/defining-and-evaluating-political-bias-in-llms/)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p2.1)\.
- L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray,et al\.\(2022\)Training language models to follow instructions with human feedback\.Advances in neural information processing systems35,pp\. 27730–27744\.Cited by:[§5\.2](https://arxiv.org/html/2606.12922#S5.SS2.SSS0.Px2.p2.1)\.
- A\. Parrish, A\. Chen, N\. Nangia, V\. Padmakumar, J\. Phang, J\. Thompson, P\. M\. Htut, and S\. Bowman \(2022\)BBQ: a hand\-built bias benchmark for question answering\.InFindings of the Association for Computational Linguistics: ACL 2022,S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 2086–2105\.External Links:[Link](https://aclanthology.org/2022.findings-acl.165/),[Document](https://dx.doi.org/10.18653/v1/2022.findings-acl.165)Cited by:[§2\.1](https://arxiv.org/html/2606.12922#S2.SS1.p1.1)\.
- T\. Peng, K\. Yang, S\. Lee, H\. Li, Y\. Chu, Y\. Lin, and H\. Liu \(2026\)Beyond partisan leaning: a comparative analysis of political bias in large language models\.Journal of Information Technology & Politics,pp\. 1–18\.Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p1.1)\.
- Y\. Potter, S\. Lai, J\. Kim, J\. Evans, and D\. Song \(2024\)Hidden persuaders: LLMs’ political leaning and their influence on voters\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 4244–4275\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.244/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.244)Cited by:[§2\.2](https://arxiv.org/html/2606.12922#S2.SS2.p1.1)\.
- L\. Rettenberger, M\. Reischl, and M\. Schutera \(2025\)Assessing political bias in large language models\.Journal of Computational Social Science8\(2\),pp\. 42\.Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p1.1)\.
- P\. Röttger, V\. Hofmann, V\. Pyatkin, M\. Hinck, H\. Kirk, H\. Schuetze, and D\. Hovy \(2024\)Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15295–15311\.External Links:[Link](https://aclanthology.org/2024.acl-long.816/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.816)Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p2.1),[§3\.2](https://arxiv.org/html/2606.12922#S3.SS2.p1.1)\.
- D\. Rozado \(2023\)Danger in the machine: the perils of political and demographic biases embedded in ai systems\.Technical reportManhattan Institute\.Cited by:[§5\.2](https://arxiv.org/html/2606.12922#S5.SS2.SSS0.Px1.p2.1)\.
- D\. Rozado \(2024\)The political preferences of llms\.PloS one19\(7\),pp\. e0306621\.Cited by:[§2\.1](https://arxiv.org/html/2606.12922#S2.SS1.p1.1)\.
- S\. Santurkar, E\. Durmus, F\. Ladhak, C\. Lee, P\. Liang, and T\. Hashimoto \(2023\)Whose opinions do language models reflect?\.InInternational conference on machine learning,pp\. 29971–30004\.Cited by:[§5\.2](https://arxiv.org/html/2606.12922#S5.SS2.SSS0.Px1.p2.1)\.
- N\. Sharma, Q\. V\. Liao, and Z\. Xiao \(2024\)Generative echo chamber? effect of llm\-powered search systems on diverse information seeking\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,CHI ’24,New York, NY, USA\.External Links:ISBN 9798400703300,[Link](https://doi.org/10.1145/3613904.3642459),[Document](https://dx.doi.org/10.1145/3613904.3642459)Cited by:[§2\.2](https://arxiv.org/html/2606.12922#S2.SS2.p1.1)\.
- D\. Shin, S\. Lee, S\. Bae, H\. Ryu, C\. Ok, H\. Jung, H\. Ji, J\. Lim, J\. Lee, J\. Han,et al\.\(2026\)Mi:dm 2\.0 korea\-centric bilingual language models\.arXiv preprint arXiv:2601\.09066\.Cited by:[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px1.p1.1)\.
- M\. Shu, D\. Karell, K\. Okura, and T\. R\. Davidson \(2026\)How latent and prompting biases in ai\-generated historical narratives influence opinions\.PNAS Nexus5\(3\),pp\. pgag022\.External Links:ISSN 2752\-6542,[Document](https://dx.doi.org/10.1093/pnasnexus/pgag022),[Link](https://doi.org/10.1093/pnasnexus/pgag022),https://academic\.oup\.com/pnasnexus/article\-pdf/5/3/pgag022/67194597/pgag022\_supplementary\_data\.pdfCited by:[§2\.2](https://arxiv.org/html/2606.12922#S2.SS2.p1.1),[§5\.2](https://arxiv.org/html/2606.12922#S5.SS2.SSS0.Px1.p2.1)\.
- SK Telecom \(2025a\)A\.X 4\.0: foundation model specialized in Korean, optimized for enterprise applications\.External Links:[Link](https://github.com/SKT-AI/A.X-4.0)Cited by:[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px1.p1.1),[§5\.2](https://arxiv.org/html/2606.12922#S5.SS2.SSS0.Px3.p4.1)\.
- SK Telecom \(2025b\)KoGPT Trinity 1\.2B v0\.5\.Note:Hugging Face model repositoryExternal Links:[Link](https://huggingface.co/skt/ko-gpt-trinity-1.2B-v0.5)Cited by:[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px1.p1.1)\.
- E\. M\. Smith, M\. Hall, M\. Kambadur, E\. Presani, and A\. Williams \(2022\)“I’m sorry to hear that”: finding new biases in language models with a holistic descriptor dataset\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 9180–9211\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.625/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.625)Cited by:[§2\.1](https://arxiv.org/html/2606.12922#S2.SS1.p1.1)\.
- S\. Thapa, A\. Maratha, K\. M\. Hasib, M\. Nasim, and U\. Naseem \(2023\)Assessing political inclination of Bangla language models\.InProceedings of the First Workshop on Bangla Language Processing \(BLP\-2023\),F\. Alam, S\. Kar, S\. A\. Chowdhury, F\. Sadeque, and R\. Amin \(Eds\.\),Singapore,pp\. 62–71\.External Links:[Link](https://aclanthology.org/2023.banglalp-1.8/),[Document](https://dx.doi.org/10.18653/v1/2023.banglalp-1.8)Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p2.1)\.
- UK AI Security Institute \(2025\)Frontier ai trends report\.Technical reportDepartment for Science, Innovation and Technology\.Note:Accessed: May 2026External Links:[Link](https://www.aisi.gov.uk/frontier-ai-trends-report)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p2.1)\.
- US Office of Management and Budget \(2025\)Increasing public trust in artificial intelligence through unbiased ai principles\.MemorandumTechnical ReportM\-26\-04,Executive Office of the President\.Note:Accessed: May 2026External Links:[Link](https://www.whitehouse.gov/wp-content/uploads/2025/12/M-26-04-Increasing-Public-Trust-in-Artificial-Intelligence-Through-Unbiased-AI-Principles-1.pdf)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p2.1)\.
- M\. Witte, J\. Schwenzow, M\. Heitmann, M\. Reisenbichler, and M\. Assenmacher \(2023\)Potential for decision aids based on natural language processing\.InProceedings of the European Marketing Academy,Note:52nd European Marketing Academy Conference, Paper 114322Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p2.1)\.
- T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz,et al\.\(2019\)Huggingface’s transformers: state\-of\-the\-art natural language processing\.arXiv preprint arXiv:1910\.03771\.Cited by:[Appendix E](https://arxiv.org/html/2606.12922#A5.p1.1)\.
- D\. Wright, A\. Arora, N\. Borenstein, S\. Yadav, S\. Belongie, and I\. Augenstein \(2024\)LLM tropes: revealing fine\-grained values and opinions in large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 17085–17112\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.995/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.995)Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p2.1)\.
- H\. Xu, W\. Gan, Z\. Qi, J\. Wu, and P\. S\. Yu \(2024\)Large language models for education: a survey\.arXiv preprint arXiv:2405\.13001\.External Links:[Link](https://arxiv.org/abs/2405.13001)Cited by:[§1](https://arxiv.org/html/2606.12922#S1.p1.1)\.
- Y\. Xu, L\. Hu, J\. Zhao, Z\. Qiu, K\. Xu, Y\. Ye, and H\. Gu \(2025\)A survey on multilingual large language models: corpora, alignment, and bias\.Frontiers of Computer Science19\(11\),pp\. 1911362\.Cited by:[§5\.2](https://arxiv.org/html/2606.12922#S5.SS2.SSS0.Px5.p3.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025a\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§5\.1](https://arxiv.org/html/2606.12922#S5.SS1.SSS0.Px1.p1.1)\.
- J\. Yang, X\. Han, and T\. Baldwin \(2025b\)Demographics and democracy: benchmarking LLMs’ gender bias and political leaning in European parliament\.InProceedings of the 8th International Conference on Natural Language and Speech Processing \(ICNLSP\-2025\),M\. Abbas, T\. Yousef, and L\. Galke \(Eds\.\),Southern Denmark University, Odense, Denmark,pp\. 416–439\.External Links:[Link](https://aclanthology.org/2025.icnlsp-1.41/)Cited by:[§2\.2](https://arxiv.org/html/2606.12922#S2.SS2.p1.1)\.
- Y\. Yang, X\. Liu, Q\. Jin, F\. Huang, and Z\. Lu \(2024\)Unmasking and quantifying racial bias of large language models in medical report generation\.Communications medicine4\(1\),pp\. 176\.Cited by:[§2\.1](https://arxiv.org/html/2606.12922#S2.SS1.p1.1)\.
- J\. Zhao, T\. Wang, M\. Yatskar, V\. Ordonez, and K\. Chang \(2018\)Gender bias in coreference resolution: evaluation and debiasing methods\.InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 \(Short Papers\),M\. Walker, H\. Ji, and A\. Stent \(Eds\.\),New Orleans, Louisiana,pp\. 15–20\.External Links:[Link](https://aclanthology.org/N18-2003/),[Document](https://dx.doi.org/10.18653/v1/N18-2003)Cited by:[§2\.1](https://arxiv.org/html/2606.12922#S2.SS1.p1.1)\.
- D\. Zhou and Y\. Zhang \(2024\)Political biases and inconsistencies in bilingual gpt models—the cases of the u\.s\. and china\.Scientific Reports14\(1\),pp\. 25048\.External Links:[Document](https://dx.doi.org/10.1038/s41598-024-76395-w),[Link](https://doi.org/10.1038/s41598-024-76395-w),ISSN 2045\-2322Cited by:[§2\.3](https://arxiv.org/html/2606.12922#S2.SS3.p2.1)\.
## Appendix ADetails of Polar Taxonomy
AxisCategoryPositionCodeSubcategoryEconomicMarket EconomyLeft403Market Regulation405Corporatism / Mixed Economy412Controlled Economy413NationalizationRight401Free Market Economy402Incentives: Positive414Economic OrthodoxyTrade / EnergyLeft406Protectionism: Positive501Environmental ProtectionRight407Protectionism: NegativeLaborLeft701Labor Groups: PositiveRight702Labor Groups: NegativeWelfare StateLeft504Welfare State Expansion506Education ExpansionRight505Welfare State Limitation507Education LimitationSocioculturalLaw and OrderProgressive201Freedom and Human Rights301Decentralization602National Way of Life: NegativeConservative302Centralization601National Way of Life: Positive605Law and Order ReinforcementGender / Minorities / EqualityProgressive503Equality: Positive604Traditional Morality: Negative607Multiculturalism: Positive705Underprivileged Minority GroupsConservative603Traditional Morality: Positive608Multiculturalism: Negative704Middle Class and Professional GroupsInternational RelationsProgressive101, 102Foreign Special Relationship107Internationalism: PositiveConservative101, 102Foreign Special Relationship109Internationalism: NegativeNational Defense / SecurityProgressive105Military: Negative106PeaceConservative104Military: PositiveTable 2:Mapping between Polar categories and Manifesto Project coding scheme\. Each Polar category is associated with Manifesto codes and directional labels used to assign political positions\.We provide the full mapping between the Polar categories and the Manifesto Project coding scheme\. In the Manifesto Project Dataset, each political statement is annotated with a Manifesto code\. Each code represents a policy subcategory and is associated with an ideological direction\. We use these code\-level annotations to ground the category and direction labels in Polar\. Table[2](https://arxiv.org/html/2606.12922#A1.T2)lists the economic and sociocultural categories in Polar, their directional labels, and the corresponding Manifesto codes and subcategories\.
This mapping serves two purposes\. First, it ensures that each Polar category is grounded in an established coding scheme rather than in ad\-hoc ideological labels\. Second, it provides a consistent basis for grouping Manifesto\-annotated statements into Polar categories\. For supplementary sources that do not come with Manifesto annotations, such as party documents, press releases, and news articles, we apply the same Manifesto codebook criteria to determine the corresponding category and direction\.
Figure[4](https://arxiv.org/html/2606.12922#A2.F4)shows an example of a Manifesto\-annotated political statement and how its existing code is linked to Polar’s taxonomy\. The statement is labeled in the Manifesto Project Dataset with code 605,Law and Order: Positive\. This code refers to favorable mentions of strict law enforcement and tougher responses to crime, and it corresponds to the conservative direction on the sociocultural axis in Polar\.
## Appendix BCategory\-Level Bill Passage Rates
U\.S\.South KoreaAxisCategoryPassage RateInstancesPassage RateInstancesEconomicMarket Economy11\.22%1318\.03%129EconomicTrade / Energy8\.84%13411\.48%125EconomicLabor15\.32%1258\.23%128EconomicWelfare State11\.57%13012\.36%124SocioculturalLaw and Order19\.26%1206\.58%132SocioculturalGender / Minorities / Equality27\.55%11012\.39%125SocioculturalInternational Relations8\.27%13412\.42%122SocioculturalNational Defense / Security18\.53%12013\.49%124Total1,0041,009Table 3:Bill passage rates and instance allocation by category in Polar\. Passage rates are computed from official legislative records from 2016 to 2025\. Categories with lower passage rates receive more instances\.We use official legislative records to estimate the level of political contestation in each category\. For each country, we collect bills introduced over the past ten years, from 2016 to 2025, and compute the passage rate of bills corresponding to each Polar category\. For the U\.S\. dataset, we use records from the[Library of Congress](https://arxiv.org/html/2606.12922#bib.bib84)\. Since each bill is annotated with a policy area, such asTaxationorMinority Issues, we map these policy\-area labels to the Polar categories\. For the South Korean dataset, we use records from[National Assembly of the Republic of Korea](https://arxiv.org/html/2606.12922#bib.bib85)\. Since South Korean bill records are organized by the standing committee that proposes or reviews each bill, we map standing committees to the Polar categories\.
Table[3](https://arxiv.org/html/2606.12922#A2.T3)reports the bill passage rate for each category and the corresponding number of instances allocated in Polar\. As described in the main text, we assign more instances to categories with lower passage rates, treating lower passage rates as evidence of stronger political contestation\.
Figure 4:Example of a political statement annotated with a Manifesto policy code and its corresponding political direction\.
## Appendix CDetails of the Dataset Construction Process
In this section, we provide additional details about the construction and review procedure for Polar\. While the main text describes the overall pipeline, we focus here on how candidate statements are selected, how complete instances are reviewed, and how final instances are revised and selected\. The procedure consists of three stages: \(1\) preparing candidate instances, \(2\) reviewing complete instances, and \(3\) revising and selecting final instances\.
### C\.1Constructing Candidate Instances
Before review, each author constructs candidate instances for assigned categories\. For each category, the author examines statements associated with Manifesto codes mapped to the corresponding Polar subcategories and lists candidate statements for both sides of a political issue\. We prioritize statements that express a clear political claim and exclude statements that are too general, too simple, or difficult to convert into evaluation options\.
The author then matches statements that address the same or closely related issue but express opposing positions\. A source statement can be reused when necessary, but we aim to cover diverse issues and resources\. When the same source is used more than once, we focus on different claims or policy aspects so that the resulting instances do not repeatedly test the same narrow issue\.
Each matched pair is converted into a complete evaluation instance\. The author writes a neutral context, derives two political continuations from the paired statements, and adds a semantically unrelated continuation\. The unrelated option is written to be grammatical as a continuation but semantically disconnected from the political issue\. We use non\-political topics, such as daily life, transportation, nature, food, sensory experience, and ordinary actions, to avoid unintended political associations\.
### C\.2Reviewing Instances
Each complete instance is independently cross\-checked by another author\. Reviewers examine the context, the two political options, and the unrelated option together\. The review focuses on the following criteria:
- •whether the context is politically neutral and does not favor either political option;
- •whether the context and each political option form a grammatical and semantically natural continuation;
- •whether the political options preserve the claim and rhetorical intensity of the original paired statements;
- •whether the paired statements and the resulting options represent opposing positions on the same issue, based on the Manifesto coding scheme;
- •whether the unrelated option is grammatical as a continuation while remaining semantically unrelated to the political issue;
- •whether the unrelated option avoids keywords or expressions that could evoke the political issue in the context;
- •whether the instance contains typographical errors or other surface\-level mistakes\.
Reviewers directly revise issues that can be resolved through wording changes, grammatical correction, or typo correction\. When an instance raises a more substantive issue, such as unclear ideological contrast, weak source preservation, or an ambiguous unrelated option, the authors discuss whether to revise the instance, replace the original pair, or remove it and construct a new instance\.
### C\.3Revising and Selecting Final Instances
After the first review, we revise, remove, or replace instances according to the reviewers’ feedback\. We then conduct an additional pass over the revised instances and select the final instances to include in the dataset\.
When the number of constructed instances exceeds the target allocation based on bill passage rates, we reduce redundancy within each category\. In particular, we limit repeated coverage of the same issue so that the final dataset covers a broader range of political topics\. We prioritize instances that better satisfy the review criteria and provide clearer, more balanced representations of opposing political positions\.
For translated instances, we additionally check whether the translation preserves the source claim, rhetorical intensity, and grammatical structure of the original instance\. Translations that introduce meaning shifts, weaken the political contrast, or create unnatural continuations are manually revised\.
This construction and review process improves the consistency and quality of Polar\. It supports the validity of the dataset by ensuring that each instance preserves opposing political positions, maintains grammatical compatibility across options, and separates political preference from unrelated continuation preference\.
## Appendix DDetails of the Models
A total of 38 LLMs are evaluated in this paper\. We select models from diverse developers and model scales to examine political bias across model families and deployment targets\. The models are grouped into two broad categories: global models and Korean\-specialized models\. Global models refer to broadly used international model families, while Korean\-specialized models refer to models developed with an explicit focus on Korean language use or South Korean user contexts\.
### Global Models
- •Llama 3\.1 and Llama 3\.2 - –Llama\-3\.1\-8B - –Llama\-3\.1\-8B\-instruct - –Llama\-3\.1\-70B - –Llama\-3\.1\-70B\-instruct - –Llama\-3\.2\-1B - –Llama\-3\.2\-1B\-instruct - –Llama\-3\.2\-3B - –Llama\-3\.2\-3B\-instruct
- •Mistral - –Mistral\-7B\-v0\.3 - –Mistral\-7B\-Instruct\-v0\.3 - –Mistral\-Small\-24B\-Base\-2501 - –Mistral\-Small\-24B\-Instruct\-2501
- •Qwen3 - –Qwen3\-0\.6B\-Base - –Qwen3\-0\.6B - –Qwen3\-1\.7B\-Base - –Qwen3\-1\.7B - –Qwen3\-4B\-Base - –Qwen3\-4B - –Qwen3\-8B\-Base - –Qwen3\-8B - –Qwen3\-14B\-Base - –Qwen3\-14B - –Qwen3\-32B
### Korean\-specialized Models
- •Midm\-2\.0 - –Midm\-2\.0\-Mini\-Instruct \(2\.3B\) - –Midm\-2\.0\-Base\-Instruct \(11\.5B\)
- •kanana\-1\.5 - –kanana\-1\.5\-2\.1b\-base - –kanana\-1\.5\-2\.1b\-instruct\-2505 - –kanana\-1\.5\-8b\-base - –kanana\-1\.5\-8b\-instruct\-2505
- •EXAONE\-4\.0 - –EXAONE\-4\.0\-1\.2B - –EXAONE\-4\.0\-32B
- •HyperCLOVA X SEED - –HyperCLOVAX\-SEED\-Text\-Instruct\-0\.5B - –HyperCLOVAX\-SEED\-Text\-Instruct\-1\.5B
- •A\.X\-4\.0 - –A\.X\-4\.0\-Light \(7B\) - –A\.X\-4\.0 \(72B\)
- •ko\-gpt\-trinity - –ko\-gpt\-trinity\-1\.2B\-v0\.5
- •SOLAR - –SOLAR\-10\.7B\-v1\.0 - –SOLAR\-10\.7B\-Instruct\-v1\.0
## Appendix EDetails of the Experimental Settings
We conduct all experiments using the LM Evaluation Harness\(Bidermanet al\.,[2024](https://arxiv.org/html/2606.12922#bib.bib82)\), a library for evaluating LLMs across diverse benchmarks\. For each model, we load the open\-source weights and tokenizer from the Hugging Face Hub\(Wolfet al\.,[2019](https://arxiv.org/html/2606.12922#bib.bib86)\)and evaluate the model in a multiple\-choice setting\. Given a context, the evaluation script computes the log\-likelihood of each candidate option conditioned on that context\.
We score each option using length\-normalized log\-likelihood\. Since raw log\-likelihood tends to decrease as token length increases, longer options can be penalized\. We divide each option’s log\-likelihood by its token length and use the normalized score to determine the model’s choice\. We then use a separate analysis script to compute political position, LMS, NS, and ICAT, and to generate the plots reported in the paper\. All experiments are conducted on a server with eight NVIDIA RTX 3090 GPUs\.
## Appendix FEvaluation Results
In this section, we provide plots and tables from experiments that support the main results in Section[5](https://arxiv.org/html/2606.12922#S5)\. We report model positions, category\-level position distributions, comparisons between global and Korean\-specialized model groups, and full metric tables\.
### F\.1Political Position of LLMs
Figure[5](https://arxiv.org/html/2606.12922#A6.F5)shows the political positions of all evaluated models on the original and translated datasets\. This figure uses the full\[−1,1\]\[\-1,1\]range for both axes\. Each figure shows the positions of models for the dataset\. The translated datasets further show that the same political content can lead to different position distributions depending on language form\.
### F\.2Category\-Level Position
Figures[6](https://arxiv.org/html/2606.12922#A6.F6),[7](https://arxiv.org/html/2606.12922#A6.F7),[8](https://arxiv.org/html/2606.12922#A6.F8), and[9](https://arxiv.org/html/2606.12922#A6.F9)show category\-level position distributions for the U\.S\., South Korean, and translated datasets\. These figures provide a more detailed view of the category\-level patterns discussed in the main text\. In the U\.S\. dataset, models generally lean left or progressive across categories, but the strength of this tendency differs by issue domain\. In the South Korean dataset, the direction of bias varies more across categories\. For example, some categories within the same axis move in opposite directions, which can make the aggregate axis\-level position appear more neutral\. Table[4](https://arxiv.org/html/2606.12922#A6.T4)provides the corresponding mean position values and standard deviations for the original and translated datasets, offering a more detailed numerical view of translation\-driven shifts in model positions\.
### F\.3Comparison Between Global and Korean\-Specialized Models
Figures[10](https://arxiv.org/html/2606.12922#A6.F10)and[11](https://arxiv.org/html/2606.12922#A6.F11)compare category\-level position distributions between global and Korean\-specialized model groups\. On the U\.S\. dataset, the two groups show broadly similar distributions across most categories\. This suggests that Korean\-specialized models exhibit position patterns similar to global models when evaluated in an English and U\.S\. political context\. In contrast, the South Korean dataset shows clearer differences between the two groups in several categories\. This pattern supports the main\-text observation that a model’s target language and deployment context matter more when the evaluation context is linguistically and politically local\. Table[5](https://arxiv.org/html/2606.12922#A6.T5)provides the corresponding mean position values and standard deviations by model group, offering a more detailed numerical view of the differences shown in Figures[10](https://arxiv.org/html/2606.12922#A6.F10)and[11](https://arxiv.org/html/2606.12922#A6.F11)\.
### F\.4Model\-Level Category Patterns
Figures[12\(a\)](https://arxiv.org/html/2606.12922#A6.F12.sf1)and[12\(b\)](https://arxiv.org/html/2606.12922#A6.F12.sf2)provide a model\-level view of the group\-level patterns shown above\. We select four example models and compare their category\-level positions on the U\.S\. and South Korean datasets\. On the U\.S\. dataset, the selected models show similar category\-level patterns, with stronger bias appearing in similar categories, such asWelfare StateandGender / Minorities / Equality\. On the South Korean dataset, the same models show more varied category\-level positions\. This result further indicates that model behavior becomes more heterogeneous when the evaluation context reflects local language and political discourse\.
### F\.5Model\-Level Metrics
Tables[6](https://arxiv.org/html/2606.12922#A6.T6),[7](https://arxiv.org/html/2606.12922#A6.T7),[8](https://arxiv.org/html/2606.12922#A6.T8), and[9](https://arxiv.org/html/2606.12922#A6.T9)report model\-level metrics for each dataset\. For each model, we report political position, neutrality score \(NS\), language modeling score \(LMS\), and ICAT score for the economic and sociocultural axes\. We also report the total ICAT score, computed as the average of the two axis\-level ICAT scores\. These tables provide the full numerical results underlying the aggregate patterns reported in the main text, including the comparison between original and translated evaluation settings\.
### F\.6Category\-Level Metrics
TablesLABEL:tab:us\_category,LABEL:tab:ko\_category,LABEL:tab:us2ko\_category, andLABEL:tab:ko2en\_categoryreport category\-level metrics for each dataset\. For each category, we compute position, NS, LMS, and ICAT separately\. These tables show where each model exhibits stronger or weaker directional preferences and whether low ICAT scores are driven by lower neutrality or lower LMS\. They also provide additional evidence that model bias can differ substantially across issue categories and language forms\.
Figure 5:Two\-dimensional political positions of LLMs across original and translated datasets\. The x\-axis represents the economic position, where negative values indicate left\-leaning preferences and positive values indicate right\-leaning preferences\. The y\-axis represents the sociocultural position, where negative values indicate progressive preferences and positive values indicate conservative preferences\. Each point corresponds to a model\.Figure 6:Category\-level political positions on the U\.S\. dataset\. The left column reports economic categories, and the right column reports sociocultural categories\. Each point corresponds to a model\.Figure 7:Category\-level political positions on the South Korean dataset\. The left column reports economic categories, and the right column reports sociocultural categories\. Each point corresponds to a model\.Figure 8:Category\-level political positions on the U\.S\. dataset translated into Korean\. The figure shows how model preferences vary across economic and sociocultural categories under the translated input setting\.Figure 9:Category\-level political positions on the South Korean dataset translated into English\. The figure shows how model preferences vary across economic and sociocultural categories under the translated input setting\.Mean Position \(Standard Deviation\)CategoryTranslationU\.S\. DatasetKorean DatasetTotalEconomic, SocioculturalOriginal−0\.28\-0\.28\(±0\.07\\pm 0\.07\),−0\.15\-0\.15\(±0\.04\\pm 0\.04\)0\.010\.01\(±0\.07\\pm 0\.07\),0\.010\.01\(±0\.04\\pm 0\.04\)Translated−0\.16\-0\.16\(±0\.08\\pm 0\.08\),−0\.04\-0\.04\(±0\.05\\pm 0\.05\)−0\.20\-0\.20\(±0\.05\\pm 0\.05\),−0\.05\-0\.05\(±0\.03\\pm 0\.03\)Market EconomyOriginal−0\.15\-0\.15\(±0\.07\\pm 0\.07\)0\.180\.18\(±0\.14\\pm 0\.14\)Translated−0\.03\-0\.03\(±0\.10\\pm 0\.10\)−0\.16\-0\.16\(±0\.08\\pm 0\.08\)Trade / EnergyOriginal−0\.26\-0\.26\(±0\.07\\pm 0\.07\)0\.010\.01\(±0\.12\\pm 0\.12\)Translated−0\.25\-0\.25\(±0\.09\\pm 0\.09\)−0\.27\-0\.27\(±0\.09\\pm 0\.09\)LaborOriginal−0\.28\-0\.28\(±0\.09\\pm 0\.09\)−0\.16\-0\.16\(±0\.06\\pm 0\.06\)Translated−0\.19\-0\.19\(±0\.10\\pm 0\.10\)−0\.20\-0\.20\(±0\.05\\pm 0\.05\)Welfare StateOriginal−0\.43\-0\.43\(±0\.10\\pm 0\.10\)0\.020\.02\(±0\.09\\pm 0\.09\)Translated−0\.18\-0\.18\(±0\.10\\pm 0\.10\)−0\.18\-0\.18\(±0\.05\\pm 0\.05\)Law and OrderOriginal−0\.17\-0\.17\(±0\.07\\pm 0\.07\)−0\.02\-0\.02\(±0\.09\\pm 0\.09\)Translated−0\.07\-0\.07\(±0\.08\\pm 0\.08\)−0\.05\-0\.05\(±0\.06\\pm 0\.06\)Gender / Minorities / EqualityOriginal−0\.32\-0\.32\(±0\.06\\pm 0\.06\)0\.070\.07\(±0\.11\\pm 0\.11\)Translated−0\.21\-0\.21\(±0\.10\\pm 0\.10\)−0\.01\-0\.01\(±0\.05\\pm 0\.05\)International RelationsOriginal−0\.09\-0\.09\(±0\.05\\pm 0\.05\)−0\.02\-0\.02\(±0\.10\\pm 0\.10\)Translated0\.040\.04\(±0\.09\\pm 0\.09\)−0\.06\-0\.06\(±0\.05\\pm 0\.05\)National Defense / SecurityOriginal−0\.04\-0\.04\(±0\.04\\pm 0\.04\)0\.000\.00\(±0\.08\\pm 0\.08\)Translated0\.040\.04\(±0\.06\\pm 0\.06\)−0\.11\-0\.11\(±0\.08\\pm 0\.08\)Table 4:Mean political positions and standard deviations for the original and translated datasets\. Negative values indicate left or progressive preferences, while positive values indicate right or conservative preferences\.Figure 10:Category\-level position distributions of global and Korean\-specialized models on the U\.S\. dataset\.Figure 11:Category\-level position distributions of global and Korean\-specialized models on the South Korean dataset\.Mean Position \(Standard Deviation\)CategoryModel GroupU\.S\. DatasetKorean DatasetTotalEconomic, SocioculturalGlobal−0\.30\-0\.30\(±0\.04\\pm 0\.04\),−0\.15\-0\.15\(±0\.02\\pm 0\.02\)0\.040\.04\(±0\.05\\pm 0\.05\),0\.010\.01\(±0\.03\\pm 0\.03\)Korean\-specialized−0\.26\-0\.26\(±0\.09\\pm 0\.09\),−0\.15\-0\.15\(±0\.05\\pm 0\.05\)−0\.02\-0\.02\(±0\.08\\pm 0\.08\),−0\.01\-0\.01\(±0\.05\\pm 0\.05\)Market EconomyGlobal−0\.17\-0\.17\(±0\.04\\pm 0\.04\)0\.250\.25\(±0\.10\\pm 0\.10\)Korean\-specialized−0\.12\-0\.12\(±0\.09\\pm 0\.09\)0\.070\.07\(±0\.13\\pm 0\.13\)Trade / EnergyGlobal−0\.26\-0\.26\(±0\.06\\pm 0\.06\)0\.070\.07\(±0\.08\\pm 0\.08\)Korean\-specialized−0\.27\-0\.27\(±0\.08\\pm 0\.08\)−0\.08\-0\.08\(±0\.11\\pm 0\.11\)LaborGlobal−0\.31\-0\.31\(±0\.08\\pm 0\.08\)−0\.17\-0\.17\(±0\.05\\pm 0\.05\)Korean\-specialized−0\.25\-0\.25\(±0\.10\\pm 0\.10\)−0\.14\-0\.14\(±0\.08\\pm 0\.08\)Welfare StateGlobal−0\.46\-0\.46\(±0\.05\\pm 0\.05\)0\.000\.00\(±0\.08\\pm 0\.08\)Korean\-specialized−0\.39\-0\.39\(±0\.14\\pm 0\.14\)0\.050\.05\(±0\.11\\pm 0\.11\)Law and OrderGlobal−0\.17\-0\.17\(±0\.05\\pm 0\.05\)0\.020\.02\(±0\.06\\pm 0\.06\)Korean\-specialized−0\.16\-0\.16\(±0\.09\\pm 0\.09\)−0\.08\-0\.08\(±0\.10\\pm 0\.10\)Gender / Minorities / EqualityGlobal−0\.32\-0\.32\(±0\.05\\pm 0\.05\)0\.100\.10\(±0\.08\\pm 0\.08\)Korean\-specialized−0\.33\-0\.33\(±0\.06\\pm 0\.06\)0\.020\.02\(±0\.14\\pm 0\.14\)International RelationsGlobal−0\.09\-0\.09\(±0\.04\\pm 0\.04\)−0\.05\-0\.05\(±0\.11\\pm 0\.11\)Korean\-specialized−0\.09\-0\.09\(±0\.07\\pm 0\.07\)0\.010\.01\(±0\.06\\pm 0\.06\)National Defense / SecurityGlobal−0\.05\-0\.05\(±0\.04\\pm 0\.04\)−0\.02\-0\.02\(±0\.08\\pm 0\.08\)Korean\-specialized−0\.03\-0\.03\(±0\.05\\pm 0\.05\)0\.040\.04\(±0\.07\\pm 0\.07\)Table 5:Mean political positions and standard deviations by model group\. Global and Korean\-specialized models are compared across the U\.S\. and Korean datasets\. Negative values indicate left or progressive preferences, while positive values indicate right or conservative preferences\.\(a\)U\.S\. dataset
\(b\)South Korean dataset
Figure 12:Category\-level political positions for representative models on the U\.S\. and South Korean datasets\. On the U\.S\. dataset, models show broadly similar category\-level patterns, with consistently left/progressive\-leaning positions across most categories\. In contrast, the South Korean dataset exhibits more heterogeneous patterns across models and categories, indicating that political bias is less uniform and more sensitive to issue context in the South Korean political setting\.ModelEconomicSocioculturalTotalPositionNSLMSICATPositionNSLMSICATICATLlama\-3\.1\-8B\-0\.350\.6598\.8564\.41\-0\.150\.8498\.3582\.6873\.55Llama\-3\.1\-8B\-Instruct\-0\.310\.6998\.2768\.11\-0\.210\.7997\.9376\.9072\.51Llama\-3\.1\-70B\-0\.340\.6698\.0864\.43\-0\.130\.8798\.5585\.3574\.89Llama\-3\.1\-70B\-Instruct\-0\.320\.6898\.4666\.68\-0\.180\.8298\.5580\.4473\.56Llama\-3\.2\-1B\-0\.280\.7298\.8571\.22\-0\.100\.8998\.5587\.3179\.27Llama\-3\.2\-1B\-Instruct\-0\.260\.7496\.9271\.58\-0\.140\.8696\.2882\.3676\.97Llama\-3\.2\-3B\-0\.300\.7098\.8569\.17\-0\.140\.8698\.5584\.6676\.91Llama\-3\.2\-3B\-Instruct\-0\.270\.7396\.9271\.00\-0\.150\.8496\.4981\.2176\.10Mistral\-7B\-v0\.3\-0\.310\.6998\.4668\.03\-0\.180\.8198\.9780\.6374\.33Mistral\-7B\-Instruct\-v0\.3\-0\.300\.7098\.2768\.47\-0\.180\.8299\.1781\.3174\.89Mistral\-Small\-24B\-Base\-2501\-0\.330\.6798\.8565\.76\-0\.150\.8499\.1783\.6474\.70Mistral\-Small\-24B\-Instruct\-2501\-0\.330\.6799\.0466\.49\-0\.170\.8398\.5581\.4173\.95Qwen3\-0\.6B\-Base\-0\.230\.7797\.8874\.91\-0\.140\.8596\.6982\.4778\.69Qwen3\-0\.6B\-0\.200\.8096\.7377\.74\-0\.160\.8395\.4579\.7078\.72Qwen3\-1\.7B\-Base\-0\.280\.7298\.6571\.22\-0\.120\.8796\.9084\.2977\.75Qwen3\-1\.7B\-0\.270\.7397\.3170\.73\-0\.150\.8496\.4981\.3376\.03Qwen3\-4B\-Base\-0\.280\.7298\.2770\.37\-0\.160\.8397\.9381\.5775\.97Qwen3\-4B\-0\.290\.7197\.5069\.26\-0\.140\.8595\.8781\.4675\.36Qwen3\-8B\-Base\-0\.300\.7098\.0868\.94\-0\.120\.8898\.1486\.6777\.80Qwen3\-8B\-0\.290\.7198\.0869\.96\-0\.150\.8597\.1182\.3376\.15Qwen3\-14B\-Base\-0\.310\.6998\.2768\.15\-0\.150\.8497\.7382\.1275\.14Qwen3\-14B\-0\.310\.6997\.5067\.54\-0\.170\.8297\.7380\.4273\.98Qwen3\-32B\-0\.360\.6498\.6563\.12\-0\.150\.8598\.5583\.8473\.48Midm\-2\.0\-Mini\-Instruct \(2\.3B\)\-0\.310\.6997\.5067\.16\-0\.160\.8397\.5280\.4673\.81Midm\-2\.0\-Base\-Instruct \(11\.5B\)\-0\.310\.6998\.2767\.99\-0\.140\.8598\.5584\.2576\.12kanana\-1\.5\-2\.1b\-base\-0\.260\.7498\.4672\.56\-0\.120\.8798\.7686\.0479\.30kanana\-1\.5\-2\.1b\-instruct\-2505\-0\.250\.7598\.4673\.50\-0\.110\.8697\.5284\.2978\.89kanana\-1\.5\-8b\-base\-0\.290\.7198\.2769\.43\-0\.160\.8298\.5581\.0275\.23kanana\-1\.5\-8b\-instruct\-2505\-0\.270\.7498\.4672\.41\-0\.120\.8797\.9385\.4478\.92EXAONE\-4\.0\-1\.2B\-0\.050\.9276\.1569\.85\-0\.050\.9075\.0067\.2968\.57EXAONE\-4\.0\-32B\-0\.320\.6897\.3166\.65\-0\.190\.8097\.1177\.8772\.26HyperCLOVAX\-SEED\-Text\-Instruct\-0\.5B\-0\.230\.7797\.6975\.51\-0\.140\.8696\.0782\.3378\.92HyperCLOVAX\-SEED\-Text\-Instruct\-1\.5B\-0\.270\.7395\.1969\.33\-0\.130\.8794\.0181\.5375\.43A\.X\-4\.0\-Light \(7B\)\-0\.270\.7397\.5070\.76\-0\.230\.7797\.5274\.6772\.72A\.X\-4\.0 \(72B\)\-0\.330\.6798\.0865\.28\-0\.180\.8198\.9780\.2972\.79ko\-gpt\-trinity\-1\.2B\-v0\.5\-0\.030\.9296\.5488\.43\-0\.070\.9294\.2186\.8587\.64SOLAR\-10\.7B\-v1\.0\-0\.350\.6598\.4663\.95\-0\.190\.8099\.3879\.8071\.87SOLAR\-10\.7B\-Instruct\-v1\.0\-0\.320\.6898\.2767\.21\-0\.210\.7999\.1778\.1572\.68
Table 6:Model\-level evaluation results on the U\.S\. dataset\.ModelEconomicSocioculturalTotalPositionNSLMSICATPositionNSLMSICATICATLlama\-3\.1\-8B0\.000\.8497\.0481\.46\-0\.010\.9397\.2290\.1985\.83Llama\-3\.1\-8B\-Instruct0\.060\.8394\.6678\.77\-0\.010\.9295\.6388\.4583\.61Llama\-3\.1\-70B\-0\.050\.8798\.2285\.85\-0\.030\.9798\.8196\.2391\.04Llama\-3\.1\-70B\-Instruct\-0\.050\.8898\.0286\.110\.000\.9497\.8191\.8688\.98Llama\-3\.2\-1B0\.070\.8291\.1174\.920\.000\.9493\.2487\.2981\.10Llama\-3\.2\-1B\-Instruct0\.080\.8486\.5673\.02\-0\.020\.8486\.2872\.7372\.87Llama\-3\.2\-3B0\.030\.8594\.0779\.720\.010\.9194\.8386\.6283\.17Llama\-3\.2\-3B\-Instruct0\.050\.8289\.7273\.480\.010\.8888\.0777\.2875\.38Mistral\-7B\-v0\.30\.110\.7790\.1269\.210\.070\.9088\.6779\.6374\.42Mistral\-7B\-Instruct\-v0\.30\.090\.8090\.5172\.710\.090\.9090\.6681\.5177\.11Mistral\-Small\-24B\-Base\-2501\-0\.020\.8996\.6486\.420\.010\.9398\.8192\.3889\.40Mistral\-Small\-24B\-Instruct\-25010\.000\.9295\.6588\.190\.000\.9498\.8193\.3490\.77Qwen3\-0\.6B\-Base0\.100\.8593\.0878\.710\.020\.8989\.8680\.3679\.54Qwen3\-0\.6B0\.160\.8189\.5372\.110\.000\.8786\.4875\.0873\.60Qwen3\-1\.7B\-Base0\.050\.8796\.4484\.190\.040\.9095\.4385\.6884\.93Qwen3\-1\.7B0\.090\.8794\.8682\.19\-0\.010\.8990\.2680\.6981\.44Qwen3\-4B\-Base0\.010\.8396\.8480\.480\.010\.9395\.6389\.3184\.89Qwen3\-4B0\.030\.8694\.0780\.440\.000\.9192\.6484\.1382\.28Qwen3\-8B\-Base\-0\.020\.8798\.0285\.320\.000\.9496\.6290\.4887\.90Qwen3\-8B0\.030\.9196\.8488\.330\.060\.9495\.2389\.0688\.70Qwen3\-14B\-Base0\.000\.8998\.4287\.680\.050\.9597\.8193\.2390\.45Qwen3\-14B0\.010\.9197\.8389\.360\.050\.9597\.2292\.1390\.75Qwen3\-32B0\.000\.8997\.4386\.30\-0\.020\.9198\.2189\.0487\.67Midm\-2\.0\-Mini\-Instruct \(2\.3B\)\-0\.080\.9198\.6290\.03\-0\.040\.9399\.0191\.9690\.99Midm\-2\.0\-Base\-Instruct \(11\.5B\)\-0\.050\.8898\.8187\.240\.000\.9099\.2089\.0988\.17kanana\-1\.5\-2\.1b\-base\-0\.060\.9199\.0190\.02\-0\.030\.9798\.4195\.2892\.65kanana\-1\.5\-2\.1b\-instruct\-2505\-0\.040\.9598\.4293\.97\-0\.040\.9697\.6193\.2693\.61kanana\-1\.5\-8b\-base\-0\.080\.9099\.0188\.64\-0\.030\.9598\.8194\.0591\.34kanana\-1\.5\-8b\-instruct\-2505\-0\.080\.9298\.8190\.820\.000\.9599\.2094\.0992\.45EXAONE\-4\.0\-1\.2B0\.130\.8478\.4665\.72\-0\.050\.8272\.5659\.4262\.57EXAONE\-4\.0\-32B0\.050\.8995\.0684\.79\-0\.030\.9294\.0486\.1485\.47HyperCLOVAX\-SEED\-Text\-Instruct\-0\.5B0\.060\.9395\.4588\.580\.050\.9595\.2390\.4389\.50HyperCLOVAX\-SEED\-Text\-Instruct\-1\.5B0\.090\.8891\.5080\.490\.070\.9390\.0683\.9282\.21A\.X\-4\.0\-Light \(7B\)\-0\.110\.8999\.2188\.21\-0\.020\.9099\.8089\.7588\.98A\.X\-4\.0 \(72B\)\-0\.150\.8599\.2183\.96\-0\.030\.90100\.0090\.2287\.09ko\-gpt\-trinity\-1\.2B\-v0\.5\-0\.060\.9498\.6293\.14\-0\.100\.9099\.0189\.5891\.36SOLAR\-10\.7B\-v1\.00\.010\.7996\.4475\.990\.070\.8996\.8286\.6181\.30SOLAR\-10\.7B\-Instruct\-v1\.00\.040\.8196\.0577\.790\.100\.9096\.6286\.8082\.30
Table 7:Model\-level evaluation results on the South Korean dataset\.ModelEconomicSocioculturalTotalPositionNSLMSICATPositionNSLMSICATICATLlama\-3\.1\-8B\-0\.160\.8495\.9680\.970\.020\.8494\.2179\.5980\.28Llama\-3\.1\-8B\-Instruct\-0\.140\.8393\.4677\.420\.000\.9092\.9883\.5180\.47Llama\-3\.1\-70B\-0\.230\.7695\.7773\.26\-0\.060\.8794\.6382\.7978\.03Llama\-3\.1\-70B\-Instruct\-0\.210\.7995\.0075\.44\-0\.090\.8793\.6081\.3478\.39Llama\-3\.2\-1B\-0\.050\.9091\.5482\.280\.060\.9189\.4681\.4081\.84Llama\-3\.2\-1B\-Instruct\-0\.010\.9189\.2380\.910\.020\.8785\.9574\.8577\.88Llama\-3\.2\-3B\-0\.130\.8394\.8178\.900\.000\.8892\.7781\.3080\.10Llama\-3\.2\-3B\-Instruct\-0\.070\.8693\.0879\.99\-0\.070\.9187\.6080\.1380\.06Mistral\-7B\-v0\.3\-0\.100\.8190\.1973\.490\.020\.9285\.7478\.8976\.19Mistral\-7B\-Instruct\-v0\.3\-0\.130\.7992\.3173\.330\.030\.9185\.7478\.2475\.78Mistral\-Small\-24B\-Base\-2501\-0\.210\.7795\.9673\.60\-0\.140\.8597\.3182\.8378\.22Mistral\-Small\-24B\-Instruct\-2501\-0\.210\.7995\.3875\.00\-0\.100\.9095\.0485\.4080\.20Qwen3\-0\.6B\-Base\-0\.050\.9495\.0089\.07\-0\.040\.9385\.7480\.0384\.55Qwen3\-0\.6B0\.000\.8994\.0483\.570\.020\.9584\.3080\.0081\.79Qwen3\-1\.7B\-Base\-0\.120\.8895\.3883\.64\-0\.040\.9389\.4683\.0283\.33Qwen3\-1\.7B\-0\.120\.8792\.5080\.88\-0\.020\.9385\.3379\.0579\.97Qwen3\-4B\-Base\-0\.160\.8495\.7780\.64\-0\.050\.8794\.2181\.8281\.23Qwen3\-4B\-0\.180\.8292\.1275\.510\.000\.9188\.0280\.2777\.89Qwen3\-8B\-Base\-0\.110\.8795\.9683\.76\-0\.060\.8995\.2584\.4684\.11Qwen3\-8B\-0\.180\.8293\.6577\.11\-0\.040\.9191\.9483\.4580\.28Qwen3\-14B\-Base\-0\.230\.7795\.7774\.12\-0\.070\.8894\.4283\.2378\.67Qwen3\-14B\-0\.240\.7694\.8171\.96\-0\.080\.8793\.1880\.9476\.45Qwen3\-32B\-0\.260\.7494\.0469\.36\-0\.110\.8794\.0181\.3375\.34Midm\-2\.0\-Mini\-Instruct \(2\.3B\)\-0\.250\.7495\.7771\.17\-0\.030\.8596\.4982\.3976\.78Midm\-2\.0\-Base\-Instruct \(11\.5B\)\-0\.270\.7395\.9669\.79\-0\.060\.9397\.5290\.9780\.38kanana\-1\.5\-2\.1b\-base\-0\.200\.8096\.7377\.08\-0\.100\.8896\.0784\.6980\.88kanana\-1\.5\-2\.1b\-instruct\-2505\-0\.220\.7894\.8174\.32\-0\.080\.8594\.4280\.5677\.44kanana\-1\.5\-8b\-base\-0\.260\.7496\.7371\.85\-0\.100\.8597\.3182\.9777\.41kanana\-1\.5\-8b\-instruct\-2505\-0\.270\.7496\.1570\.68\-0\.050\.8395\.8779\.5775\.12EXAONE\-4\.0\-1\.2B0\.000\.9875\.7774\.180\.030\.9569\.2165\.7669\.97EXAONE\-4\.0\-32B\-0\.170\.8392\.3176\.31\-0\.060\.8990\.0879\.8978\.10HyperCLOVAX\-SEED\-Text\-Instruct\-0\.5B\-0\.080\.9388\.8582\.200\.030\.8789\.8877\.9780\.09HyperCLOVAX\-SEED\-Text\-Instruct\-1\.5B\-0\.100\.9086\.7377\.79\-0\.030\.9084\.0975\.8576\.82A\.X\-4\.0\-Light \(7B\)\-0\.220\.7895\.5874\.61\-0\.090\.9197\.1188\.3081\.45A\.X\-4\.0 \(72B\)\-0\.330\.6797\.8865\.90\-0\.130\.8498\.5583\.1274\.51ko\-gpt\-trinity\-1\.2B\-v0\.5\-0\.220\.7895\.5874\.32\-0\.060\.8992\.9882\.7678\.54SOLAR\-10\.7B\-v1\.0\-0\.110\.8995\.1984\.60\-0\.010\.8493\.6078\.5281\.56SOLAR\-10\.7B\-Instruct\-v1\.0\-0\.180\.8294\.2377\.61\-0\.080\.8592\.9878\.8278\.22
Table 8:Model\-level evaluation results on the U\.S\. dataset translated into Korean\.ModelEconomicSocioculturalTotalPositionNSLMSICATPositionNSLMSICATICATLlama\-3\.1\-8B\-0\.240\.7695\.8572\.51\-0\.060\.9496\.2290\.3681\.44Llama\-3\.1\-8B\-Instruct\-0\.210\.7995\.2675\.71\-0\.050\.9495\.4389\.9982\.85Llama\-3\.1\-70B\-0\.280\.7296\.4469\.33\-0\.050\.9597\.8192\.5980\.96Llama\-3\.1\-70B\-Instruct\-0\.250\.7595\.6571\.43\-0\.060\.9397\.6190\.4980\.96Llama\-3\.2\-1B\-0\.160\.8494\.8679\.86\-0\.100\.8894\.6383\.4381\.64Llama\-3\.2\-1B\-Instruct\-0\.110\.8989\.7279\.77\-0\.110\.8791\.6579\.3779\.57Llama\-3\.2\-3B\-0\.170\.8394\.8678\.33\-0\.050\.9394\.2387\.6382\.98Llama\-3\.2\-3B\-Instruct\-0\.160\.8491\.1176\.72\-0\.060\.9192\.4584\.1980\.46Mistral\-7B\-v0\.3\-0\.140\.8696\.4482\.64\-0\.010\.9497\.6191\.9587\.29Mistral\-7B\-Instruct\-v0\.3\-0\.150\.8596\.6482\.42\-0\.060\.9498\.0192\.1687\.29Mistral\-Small\-24B\-Base\-2501\-0\.230\.7795\.8573\.46\-0\.010\.9696\.8293\.2283\.34Mistral\-Small\-24B\-Instruct\-2501\-0\.210\.7995\.2675\.09\-0\.030\.9596\.6291\.6783\.38Qwen3\-0\.6B\-Base\-0\.230\.7792\.6971\.430\.000\.9693\.4489\.8880\.65Qwen3\-0\.6B\-0\.250\.7590\.1267\.64\-0\.060\.9491\.4586\.3376\.99Qwen3\-1\.7B\-Base\-0\.210\.7993\.4873\.46\-0\.060\.9494\.6388\.9881\.22Qwen3\-1\.7B\-0\.220\.7890\.1270\.65\-0\.030\.9691\.4587\.4079\.02Qwen3\-4B\-Base\-0\.250\.7594\.0770\.21\-0\.070\.9395\.4389\.1979\.70Qwen3\-4B\-0\.240\.7691\.3069\.62\-0\.050\.9392\.4585\.9877\.80Qwen3\-8B\-Base\-0\.260\.7493\.8769\.76\-0\.020\.9696\.2292\.1680\.96Qwen3\-8B\-0\.250\.7593\.2870\.39\-0\.060\.9394\.6387\.6379\.01Qwen3\-14B\-Base\-0\.260\.7394\.4769\.39\-0\.080\.9296\.4289\.1379\.26Qwen3\-14B\-0\.250\.7593\.6870\.13\-0\.050\.9296\.0287\.8979\.01Qwen3\-32B\-0\.230\.7794\.4772\.79\-0\.030\.9596\.8292\.4482\.62Midm\-2\.0\-Mini\-Instruct \(2\.3B\)\-0\.230\.7794\.0772\.48\-0\.080\.9295\.2387\.3279\.90Midm\-2\.0\-Base\-Instruct \(11\.5B\)\-0\.210\.7996\.8476\.86\-0\.060\.9496\.4290\.2783\.56kanana\-1\.5\-2\.1b\-base\-0\.210\.7993\.6873\.84\-0\.090\.9194\.2385\.7879\.81kanana\-1\.5\-2\.1b\-instruct\-2505\-0\.180\.8294\.4777\.25\-0\.050\.9595\.8390\.9984\.12kanana\-1\.5\-8b\-base\-0\.210\.7995\.2675\.64\-0\.050\.9594\.0489\.0782\.36kanana\-1\.5\-8b\-instruct\-2505\-0\.180\.8292\.8976\.08\-0\.010\.9493\.6487\.5781\.83EXAONE\-4\.0\-1\.2B\-0\.250\.7563\.0447\.34\-0\.040\.9672\.3769\.7658\.55EXAONE\-4\.0\-32B\-0\.250\.7593\.8770\.72\-0\.070\.9295\.6387\.9279\.32HyperCLOVAX\-SEED\-Text\-Instruct\-0\.5B\-0\.150\.8591\.3077\.27\-0\.090\.9192\.6484\.4380\.85HyperCLOVAX\-SEED\-Text\-Instruct\-1\.5B\-0\.200\.8089\.3371\.85\-0\.100\.9091\.8582\.5177\.18A\.X\-4\.0\-Light \(7B\)\-0\.190\.8192\.0974\.78\-0\.010\.9596\.4291\.2783\.02A\.X\-4\.0 \(72B\)\-0\.250\.7597\.2373\.03\-0\.070\.9399\.4092\.6382\.83ko\-gpt\-trinity\-1\.2B\-v0\.5\-0\.020\.9099\.2189\.28\-0\.120\.8399\.4082\.1385\.71SOLAR\-10\.7B\-v1\.0\-0\.160\.8497\.0481\.65\-0\.050\.9598\.2192\.8587\.25SOLAR\-10\.7B\-Instruct\-v1\.0\-0\.110\.8696\.6483\.55\-0\.010\.9797\.0294\.5289\.03
Table 9:Model\-level evaluation results on the South Korean dataset translated into English\.Table 10:Category\-level evaluation results on the U\.S\. dataset\.ModelAxisCategoryPositionNSLMSICATLlama\-3\.1\-8BEconomicMarket Economy\-0\.210\.7996\.9576\.23Trade / Energy\-0\.340\.6699\.2565\.18Labor\-0\.340\.6699\.2065\.08Welfare State\-0\.490\.51100\.0050\.77SocioculturalLaw and Order\-0\.220\.7898\.3377\.03Gender / Minorities / Equality\-0\.260\.7496\.3670\.96International Relations\-0\.080\.92100\.0091\.79National Defense / Security\-0\.080\.9398\.3390\.96Llama\-3\.1\-8B\-InstructEconomicMarket Economy\-0\.190\.8196\.9578\.45Trade / Energy\-0\.230\.7797\.7675\.14Labor\-0\.340\.6699\.2065\.08Welfare State\-0\.460\.5499\.2353\.43SocioculturalLaw and Order\-0\.230\.7797\.5074\.75Gender / Minorities / Equality\-0\.360\.6495\.4560\.74International Relations\-0\.110\.8999\.2588\.14National Defense / Security\-0\.150\.8599\.1784\.29Llama\-3\.1\-70BEconomicMarket Economy\-0\.080\.9296\.9588\.81Trade / Energy\-0\.390\.6197\.7659\.82Labor\-0\.410\.5997\.6057\.78Welfare State\-0\.490\.51100\.0050\.77SocioculturalLaw and Order\-0\.180\.8297\.5079\.63Gender / Minorities / Equality\-0\.220\.7898\.1876\.76International Relations\-0\.130\.87100\.0086\.57National Defense / Security0\.001\.0098\.3398\.33Llama\-3\.1\-70B\-InstructEconomicMarket Economy\-0\.150\.8596\.9582\.89Trade / Energy\-0\.360\.6497\.7662\.74Labor\-0\.280\.7299\.2071\.42Welfare State\-0\.510\.49100\.0049\.23SocioculturalLaw and Order\-0\.220\.7898\.3376\.21Gender / Minorities / Equality\-0\.330\.6797\.2765\.44International Relations\-0\.070\.93100\.0092\.54National Defense / Security\-0\.110\.8998\.3387\.68Llama\-3\.2\-1BEconomicMarket Economy\-0\.190\.8197\.7179\.06Trade / Energy\-0\.190\.8199\.2580\.00Labor\-0\.260\.7499\.2073\.80Welfare State\-0\.480\.5299\.2351\.91SocioculturalLaw and Order\-0\.080\.9299\.1790\.90Gender / Minorities / Equality\-0\.280\.7297\.2769\.86International Relations\-0\.070\.93100\.0092\.54National Defense / Security0\.020\.9897\.5095\.88Llama\-3\.2\-1B\-InstructEconomicMarket Economy\-0\.190\.8194\.6676\.59Trade / Energy\-0\.210\.7997\.0176\.74Labor\-0\.200\.8098\.4078\.72Welfare State\-0\.450\.5597\.6954\.11SocioculturalLaw and Order\-0\.130\.8897\.5085\.31Gender / Minorities / Equality\-0\.360\.6496\.3661\.32International Relations\-0\.090\.9197\.7689\.01National Defense / Security0\.001\.0093\.3393\.33Llama\-3\.2\-3BEconomicMarket Economy\-0\.150\.8597\.7183\.54Trade / Energy\-0\.300\.7098\.5169\.10Labor\-0\.300\.7099\.2069\.84Welfare State\-0\.460\.54100\.0053\.85SocioculturalLaw and Order\-0\.150\.8599\.1784\.29Gender / Minorities / Equality\-0\.270\.7397\.2770\.74International Relations\-0\.070\.93100\.0092\.54National Defense / Security\-0\.070\.9397\.5091\.00Llama\-3\.2\-3B\-InstructEconomicMarket Economy\-0\.220\.7893\.1372\.51Trade / Energy\-0\.180\.8297\.7680\.25Labor\-0\.210\.7998\.4077\.93Welfare State\-0\.460\.5498\.4653\.02SocioculturalLaw and Order\-0\.100\.9096\.6787\.00Gender / Minorities / Equality\-0\.350\.6594\.5561\.88International Relations\-0\.100\.9099\.2588\.88National Defense / Security\-0\.080\.9295\.0087\.08Mistral\-7B\-v0\.3EconomicMarket Economy\-0\.210\.7996\.9576\.97Trade / Energy\-0\.280\.7299\.2571\.11Labor\-0\.410\.5997\.6057\.78Welfare State\-0\.340\.66100\.0066\.15SocioculturalLaw and Order\-0\.170\.8399\.1782\.64Gender / Minorities / Equality\-0\.350\.6596\.3663\.07International Relations\-0\.180\.82100\.0082\.09National Defense / Security\-0\.050\.95100\.0095\.00Mistral\-7B\-Instruct\-v0\.3EconomicMarket Economy\-0\.230\.7796\.9574\.75Trade / Energy\-0\.220\.7898\.5176\.45Labor\-0\.360\.6497\.6062\.46Welfare State\-0\.400\.60100\.0060\.00SocioculturalLaw and Order\-0\.120\.88100\.0088\.33Gender / Minorities / Equality\-0\.360\.6496\.3661\.32International Relations\-0\.160\.84100\.0084\.33National Defense / Security\-0\.080\.92100\.0091\.67Mistral\-Small\-24B\-Base\-2501EconomicMarket Economy\-0\.170\.8396\.9580\.67Trade / Energy\-0\.240\.7699\.2575\.55Labor\-0\.440\.5699\.2055\.55Welfare State\-0\.490\.51100\.0050\.77SocioculturalLaw and Order\-0\.200\.8099\.1779\.33Gender / Minorities / Equality\-0\.300\.7098\.1868\.73International Relations\-0\.060\.94100\.0094\.03National Defense / Security\-0\.070\.9399\.1792\.56Mistral\-Small\-24B\-Instruct\-2501EconomicMarket Economy\-0\.150\.8596\.9582\.89Trade / Energy\-0\.250\.75100\.0074\.63Labor\-0\.410\.5999\.2058\.73Welfare State\-0\.510\.49100\.0049\.23SocioculturalLaw and Order\-0\.180\.8299\.1780\.99Gender / Minorities / Equality\-0\.350\.6597\.2762\.79International Relations\-0\.070\.9399\.2591\.85National Defense / Security\-0\.080\.9298\.3390\.14Qwen3\-0\.6B\-BaseEconomicMarket Economy\-0\.110\.8994\.6683\.82Trade / Energy\-0\.190\.81100\.0080\.60Labor\-0\.180\.8299\.2080\.95Welfare State\-0\.450\.5597\.6954\.11SocioculturalLaw and Order\-0\.140\.8697\.5083\.69Gender / Minorities / Equality\-0\.380\.6295\.4559\.01International Relations\-0\.010\.9998\.5197\.04National Defense / Security\-0\.050\.9595\.0090\.25Qwen3\-0\.6BEconomicMarket Economy\-0\.130\.8792\.3780\.38Trade / Energy\-0\.150\.8598\.5183\.80Labor\-0\.170\.83100\.0083\.20Welfare State\-0\.340\.6696\.1563\.61SocioculturalLaw and Order\-0\.230\.7797\.5074\.75Gender / Minorities / Equality\-0\.360\.6493\.6459\.59International Relations\-0\.030\.9797\.7694\.84National Defense / Security\-0\.030\.9792\.5089\.42Qwen3\-1\.7B\-BaseEconomicMarket Economy\-0\.180\.8296\.1878\.56Trade / Energy\-0\.280\.7299\.2571\.11Labor\-0\.180\.8299\.2080\.95Welfare State\-0\.460\.54100\.0053\.85SocioculturalLaw and Order\-0\.170\.8397\.5081\.25Gender / Minorities / Equality\-0\.300\.7095\.4566\.82International Relations\-0\.040\.9698\.5194\.83National Defense / Security0\.020\.9895\.8394\.24Qwen3\-1\.7BEconomicMarket Economy\-0\.130\.8793\.8981\.71Trade / Energy\-0\.250\.7598\.5174\.25Labor\-0\.220\.7899\.2076\.98Welfare State\-0\.490\.5197\.6949\.60SocioculturalLaw and Order\-0\.230\.7796\.6774\.11Gender / Minorities / Equality\-0\.360\.6494\.5560\.17International Relations\-0\.010\.9999\.2597\.77National Defense / Security\-0\.020\.9895\.0093\.42Qwen3\-4B\-BaseEconomicMarket Economy\-0\.160\.8495\.4280\.12Trade / Energy\-0\.220\.7898\.5177\.19Labor\-0\.330\.6799\.2066\.66Welfare State\-0\.430\.57100\.0056\.92SocioculturalLaw and Order\-0\.180\.8297\.5079\.63Gender / Minorities / Equality\-0\.330\.6796\.3664\.83International Relations\-0\.070\.93100\.0092\.54National Defense / Security\-0\.080\.9297\.5089\.38Qwen3\-4BEconomicMarket Economy\-0\.150\.8593\.8980\.28Trade / Energy\-0\.230\.7797\.7675\.14Labor\-0\.340\.6699\.2065\.87Welfare State\-0\.450\.5599\.2354\.96SocioculturalLaw and Order\-0\.080\.9397\.5090\.19Gender / Minorities / Equality\-0\.420\.5895\.4555\.54International Relations\-0\.070\.9398\.5191\.16National Defense / Security\-0\.030\.9791\.6788\.61Qwen3\-8B\-BaseEconomicMarket Economy\-0\.190\.8195\.4277\.21Trade / Energy\-0\.250\.7598\.5174\.25Labor\-0\.330\.6798\.4066\.12Welfare State\-0\.420\.58100\.0057\.69SocioculturalLaw and Order\-0\.130\.8798\.3385\.22Gender / Minorities / Equality\-0\.200\.8095\.4576\.36International Relations\-0\.130\.87100\.0086\.57National Defense / Security0\.001\.0098\.3398\.33Qwen3\-8BEconomicMarket Economy\-0\.150\.8595\.4281\.58Trade / Energy\-0\.270\.7398\.5172\.04Labor\-0\.260\.7498\.4073\.21Welfare State\-0\.480\.52100\.0052\.31SocioculturalLaw and Order\-0\.170\.8398\.3381\.94Gender / Minorities / Equality\-0\.290\.7194\.5567\.04International Relations\-0\.130\.8798\.5185\.28National Defense / Security\-0\.020\.9896\.6795\.06Qwen3\-14B\-BaseEconomicMarket Economy\-0\.130\.8795\.4283\.04Trade / Energy\-0\.270\.7398\.5172\.04Labor\-0\.310\.6999\.2068\.25Welfare State\-0\.520\.48100\.0048\.46SocioculturalLaw and Order\-0\.230\.7798\.3375\.39Gender / Minorities / Equality\-0\.290\.7195\.4567\.69International Relations\-0\.090\.91100\.0091\.04National Defense / Security\-0\.020\.9896\.6794\.25Qwen3\-14BEconomicMarket Economy\-0\.150\.8593\.8980\.28Trade / Energy\-0\.230\.7798\.5175\.72Labor\-0\.380\.6298\.4061\.40Welfare State\-0\.480\.5299\.2351\.91SocioculturalLaw and Order\-0\.230\.7797\.5074\.75Gender / Minorities / Equality\-0\.290\.7196\.3668\.33International Relations\-0\.130\.87100\.0086\.57National Defense / Security\-0\.050\.9596\.6791\.83Qwen3\-32BEconomicMarket Economy\-0\.270\.7396\.1870\.49Trade / Energy\-0\.340\.6699\.2565\.18Labor\-0\.380\.6299\.2061\.11Welfare State\-0\.450\.55100\.0055\.38SocioculturalLaw and Order\-0\.160\.8498\.3382\.76Gender / Minorities / Equality\-0\.250\.7597\.2772\.51International Relations\-0\.130\.87100\.0086\.57National Defense / Security\-0\.050\.9598\.3393\.42Midm\-2\.0\-Mini\-Instruct \(2\.3B\)EconomicMarket Economy\-0\.130\.8796\.1883\.70Trade / Energy\-0\.340\.6697\.0163\.71Labor\-0\.260\.7497\.6071\.83Welfare State\-0\.510\.4999\.2348\.85SocioculturalLaw and Order\-0\.200\.8098\.3378\.67Gender / Minorities / Equality\-0\.360\.6496\.3661\.32International Relations\-0\.120\.8898\.5186\.75National Defense / Security0\.020\.9896\.6795\.06Midm\-2\.0\-Base\-Instruct \(11\.5B\)EconomicMarket Economy\-0\.110\.8995\.4284\.49Trade / Energy\-0\.300\.7098\.5169\.10Labor\-0\.310\.6999\.2068\.25Welfare State\-0\.510\.49100\.0049\.23SocioculturalLaw and Order\-0\.200\.8099\.1779\.33Gender / Minorities / Equality\-0\.310\.6996\.3666\.58International Relations\-0\.030\.97100\.0097\.01National Defense / Security\-0\.040\.9698\.3394\.24kanana\-1\.5\-2\.1b\-baseEconomicMarket Economy\-0\.100\.9096\.1886\.64Trade / Energy\-0\.190\.81100\.0080\.60Labor\-0\.340\.6698\.4065\.34Welfare State\-0\.420\.5899\.2357\.25SocioculturalLaw and Order\-0\.080\.9398\.3390\.96Gender / Minorities / Equality\-0\.350\.6597\.2763\.67International Relations\-0\.040\.96100\.0095\.52National Defense / Security\-0\.050\.9599\.1794\.21kanana\-1\.5\-2\.1b\-instruct\-2505EconomicMarket Economy\-0\.130\.8796\.1883\.70Trade / Energy\-0\.280\.7299\.2571\.11Labor\-0\.220\.7899\.2077\.77Welfare State\-0\.380\.6299\.2361\.07SocioculturalLaw and Order\-0\.120\.8896\.6785\.39Gender / Minorities / Equality\-0\.320\.6896\.3665\.70International Relations\-0\.070\.9398\.5191\.16National Defense / Security0\.030\.9798\.3395\.06kanana\-1\.5\-8b\-baseEconomicMarket Economy\-0\.150\.8596\.1881\.50Trade / Energy\-0\.250\.7599\.2574\.81Labor\-0\.340\.6698\.4064\.55Welfare State\-0\.430\.5799\.2356\.49SocioculturalLaw and Order\-0\.250\.7599\.1774\.38Gender / Minorities / Equality\-0\.400\.6096\.3657\.82International Relations\-0\.040\.9699\.2594\.81National Defense / Security0\.020\.9899\.1797\.51kanana\-1\.5\-8b\-instruct\-2505EconomicMarket Economy\-0\.130\.8796\.1883\.70Trade / Energy\-0\.310\.6998\.5168\.37Labor\-0\.190\.8199\.2080\.15Welfare State\-0\.430\.57100\.0056\.92SocioculturalLaw and Order\-0\.130\.87100\.0086\.67Gender / Minorities / Equality\-0\.350\.6597\.2762\.79International Relations\-0\.020\.9897\.0194\.84National Defense / Security0\.001\.0097\.5097\.50EXAONE\-4\.0\-1\.2BEconomicMarket Economy0\.070\.9367\.9463\.27Trade / Energy\-0\.130\.8772\.3962\.66Labor\-0\.130\.8781\.6071\.16Welfare State0\.001\.0083\.0883\.08SocioculturalLaw and Order0\.050\.9582\.5078\.38Gender / Minorities / Equality\-0\.240\.7671\.8254\.84International Relations\-0\.070\.9376\.8771\.13National Defense / Security0\.050\.9568\.3364\.92EXAONE\-4\.0\-32BEconomicMarket Economy\-0\.160\.8493\.1378\.20Trade / Energy\-0\.330\.6796\.2764\.66Labor\-0\.260\.74100\.0073\.60Welfare State\-0\.510\.49100\.0049\.23SocioculturalLaw and Order\-0\.170\.8397\.5081\.25Gender / Minorities / Equality\-0\.420\.5897\.2756\.60International Relations\-0\.150\.8598\.5183\.80National Defense / Security\-0\.060\.9495\.0089\.46HyperCLOVAX\-SEED\-Text\-Instruct\-0\.5BEconomicMarket Economy\-0\.070\.9392\.3786\.02Trade / Energy\-0\.190\.8199\.2580\.00Labor\-0\.180\.8299\.2080\.95Welfare State\-0\.460\.54100\.0053\.85SocioculturalLaw and Order\-0\.180\.8296\.6778\.94Gender / Minorities / Equality\-0\.310\.6996\.3666\.58International Relations\-0\.030\.9799\.2596\.29National Defense / Security\-0\.050\.9591\.6787\.08HyperCLOVAX\-SEED\-Text\-Instruct\-1\.5BEconomicMarket Economy\-0\.110\.8990\.0879\.76Trade / Energy\-0\.360\.6497\.0162\.26Labor\-0\.170\.8398\.4081\.87Welfare State\-0\.450\.5595\.3852\.83SocioculturalLaw and Order\-0\.230\.7796\.6774\.11Gender / Minorities / Equality\-0\.240\.7690\.9169\.42International Relations\-0\.040\.9697\.0192\.67National Defense / Security\-0\.020\.9890\.8389\.32A\.X\-4\.0\-Light \(7B\)EconomicMarket Economy\-0\.210\.7995\.4275\.75Trade / Energy\-0\.260\.7497\.7672\.23Labor\-0\.170\.8397\.6081\.20Welfare State\-0\.460\.5499\.2353\.43SocioculturalLaw and Order\-0\.210\.7997\.5077\.19Gender / Minorities / Equality\-0\.420\.5896\.3656\.07International Relations\-0\.190\.81100\.0080\.60National Defense / Security\-0\.120\.8895\.8384\.65A\.X\-4\.0 \(72B\)EconomicMarket Economy\-0\.190\.8196\.1877\.83Trade / Energy\-0\.370\.6397\.7661\.28Labor\-0\.310\.6998\.4067\.70Welfare State\-0\.460\.54100\.0053\.85SocioculturalLaw and Order\-0\.250\.7599\.1774\.38Gender / Minorities / Equality\-0\.350\.6598\.1864\.26International Relations\-0\.130\.87100\.0086\.57National Defense / Security\-0\.020\.9898\.3395\.88ko\-gpt\-trinity\-1\.2B\-v0\.5EconomicMarket Economy0\.110\.8990\.0879\.76Trade / Energy\-0\.100\.9099\.2588\.88Labor\-0\.040\.9699\.2095\.23Welfare State\-0\.080\.9297\.6990\.18SocioculturalLaw and Order\-0\.020\.9898\.3396\.69Gender / Minorities / Equality\-0\.240\.7691\.8270\.12International Relations\-0\.060\.9498\.5192\.63National Defense / Security0\.001\.0087\.5087\.50SOLAR\-10\.7B\-v1\.0EconomicMarket Economy\-0\.250\.7596\.9572\.52Trade / Energy\-0\.360\.6498\.5163\.22Labor\-0\.390\.6198\.4059\.83Welfare State\-0\.400\.60100\.0060\.00SocioculturalLaw and Order\-0\.230\.77100\.0076\.67Gender / Minorities / Equality\-0\.310\.6997\.2767\.21International Relations\-0\.180\.82100\.0082\.09National Defense / Security\-0\.070\.93100\.0093\.33SOLAR\-10\.7B\-Instruct\-v1\.0EconomicMarket Economy\-0\.190\.8196\.1877\.83Trade / Energy\-0\.330\.6798\.5166\.16Labor\-0\.380\.6298\.4061\.40Welfare State\-0\.370\.63100\.0063\.08SocioculturalLaw and Order\-0\.220\.78100\.0078\.33Gender / Minorities / Equality\-0\.310\.6997\.2767\.21International Relations\-0\.240\.76100\.0076\.12National Defense / Security\-0\.080\.9299\.1790\.90Table 11:Category\-level evaluation results on the South Korean dataset\.ModelAxisCategoryPositionNSLMSICATLlama\-3\.1\-8BEconomicMarket Economy0\.290\.7199\.2270\.76Trade / Energy0\.020\.9894\.4092\.13Labor\-0\.230\.7797\.6674\.77Welfare State\-0\.100\.9096\.7787\.41SocioculturalLaw and Order\-0\.050\.9597\.7393\.29Gender / Minorities / Equality0\.060\.9496\.0090\.62International Relations\-0\.120\.8896\.7284\.83National Defense / Security0\.060\.9498\.3992\.04Llama\-3\.1\-8B\-InstructEconomicMarket Economy0\.380\.6298\.4561\.05Trade / Energy0\.070\.9390\.4083\.89Labor\-0\.190\.8192\.9775\.54Welfare State\-0\.030\.9796\.7793\.65SocioculturalLaw and Order0\.001\.0097\.7397\.73Gender / Minorities / Equality0\.100\.9094\.4084\.58International Relations\-0\.160\.8493\.4478\.12National Defense / Security0\.030\.9796\.7793\.65Llama\-3\.1\-70BEconomicMarket Economy0\.150\.8599\.2284\.61Trade / Energy\-0\.030\.9795\.2092\.15Labor\-0\.200\.8099\.2279\.84Welfare State\-0\.130\.8799\.1986\.39SocioculturalLaw and Order\-0\.020\.9899\.2497\.74Gender / Minorities / Equality\-0\.040\.9698\.4094\.46International Relations\-0\.050\.9598\.3693\.52National Defense / Security0\.001\.0099\.1999\.19Llama\-3\.1\-70B\-InstructEconomicMarket Economy0\.130\.8798\.4585\.48Trade / Energy\-0\.100\.9095\.2085\.30Labor\-0\.230\.7799\.2275\.96Welfare State0\.020\.9899\.1997\.59SocioculturalLaw and Order0\.010\.9999\.2498\.49Gender / Minorities / Equality0\.010\.9996\.8096\.03International Relations\-0\.130\.8795\.9083\.32National Defense / Security0\.100\.9099\.1989\.59Llama\-3\.2\-1BEconomicMarket Economy0\.400\.6093\.0255\.53Trade / Energy0\.090\.9190\.4082\.44Labor\-0\.190\.8189\.0672\.36Welfare State\-0\.030\.9791\.9488\.97SocioculturalLaw and Order0\.050\.9596\.2191\.11Gender / Minorities / Equality0\.070\.9392\.0085\.38International Relations\-0\.070\.9392\.6286\.55National Defense / Security\-0\.060\.9491\.9486\.00Llama\-3\.2\-1B\-InstructEconomicMarket Economy0\.330\.6786\.8257\.88Trade / Energy0\.090\.9184\.0076\.61Labor\-0\.160\.8485\.9472\.51Welfare State0\.050\.9589\.5285\.18SocioculturalLaw and Order0\.050\.9592\.4287\.52Gender / Minorities / Equality0\.220\.7887\.2068\.36International Relations\-0\.280\.7277\.0555\.58National Defense / Security\-0\.080\.9287\.9080\.81Llama\-3\.2\-3BEconomicMarket Economy0\.320\.6897\.6766\.63Trade / Energy0\.040\.9694\.4090\.62Labor\-0\.170\.8392\.1976\.34Welfare State\-0\.080\.9291\.9484\.52SocioculturalLaw and Order0\.030\.9796\.2193\.30Gender / Minorities / Equality0\.140\.8692\.8080\.18International Relations\-0\.160\.8493\.4478\.12National Defense / Security0\.020\.9896\.7795\.21Llama\-3\.2\-3B\-InstructEconomicMarket Economy0\.190\.8189\.9272\.50Trade / Energy0\.200\.8092\.0073\.60Labor\-0\.270\.7389\.8465\.98Welfare State0\.060\.9487\.1081\.48SocioculturalLaw and Order0\.040\.9695\.4591\.84Gender / Minorities / Equality0\.220\.7885\.6066\.43International Relations\-0\.160\.8481\.9769\.20National Defense / Security\-0\.070\.9388\.7182\.27Mistral\-7B\-v0\.3EconomicMarket Economy0\.360\.6489\.9257\.16Trade / Energy0\.170\.8391\.2075\.88Labor\-0\.230\.7792\.9771\.18Welfare State0\.160\.8486\.2972\.37SocioculturalLaw and Order\-0\.060\.9493\.1887\.53Gender / Minorities / Equality0\.220\.7887\.2068\.36International Relations0\.130\.8781\.1570\.51National Defense / Security0\.001\.0092\.7492\.74Mistral\-7B\-Instruct\-v0\.3EconomicMarket Economy0\.270\.7394\.5768\.91Trade / Energy0\.250\.7588\.8066\.78Labor\-0\.200\.8092\.1973\.46Welfare State0\.060\.9486\.2980\.72SocioculturalLaw and Order\-0\.020\.9895\.4594\.01Gender / Minorities / Equality0\.200\.8088\.0070\.40International Relations0\.180\.8285\.2569\.87National Defense / Security\-0\.010\.9993\.5592\.79Mistral\-Small\-24B\-Base\-2501EconomicMarket Economy0\.180\.8297\.6780\.26Trade / Energy\-0\.020\.9892\.8090\.57Labor\-0\.160\.8498\.4483\.06Welfare State\-0\.060\.9497\.5891\.29SocioculturalLaw and Order\-0\.080\.9298\.4891\.02Gender / Minorities / Equality0\.150\.8599\.2084\.12International Relations\-0\.020\.9897\.5495\.94National Defense / Security\-0\.020\.98100\.0098\.39Mistral\-Small\-24B\-Instruct\-2501EconomicMarket Economy0\.150\.8596\.9082\.63Trade / Energy0\.010\.9991\.2090\.47Labor\-0\.140\.8696\.8883\.25Welfare State\-0\.020\.9897\.5896\.01SocioculturalLaw and Order\-0\.060\.9498\.4892\.52Gender / Minorities / Equality0\.100\.9099\.2088\.88International Relations\-0\.020\.9897\.5495\.94National Defense / Security\-0\.040\.96100\.0095\.97Qwen3\-0\.6B\-BaseEconomicMarket Economy0\.380\.6293\.8058\.17Trade / Energy0\.080\.9294\.4086\.85Labor\-0\.110\.8991\.4181\.41Welfare State0\.050\.9592\.7488\.25SocioculturalLaw and Order0\.060\.9494\.7088\.96Gender / Minorities / Equality0\.150\.8589\.6075\.98International Relations0\.030\.9783\.6180\.87National Defense / Security\-0\.180\.8291\.1374\.96Qwen3\-0\.6BEconomicMarket Economy0\.430\.5789\.1551\.14Trade / Energy0\.150\.8593\.6079\.37Labor\-0\.060\.9484\.3879\.10Welfare State0\.140\.8691\.1378\.64SocioculturalLaw and Order0\.080\.9290\.1582\.64Gender / Minorities / Equality0\.170\.8389\.6074\.55International Relations\-0\.100\.9078\.6970\.95National Defense / Security\-0\.180\.8287\.1071\.64Qwen3\-1\.7B\-BaseEconomicMarket Economy0\.280\.7297\.6770\.42Trade / Energy0\.070\.9395\.2088\.35Labor\-0\.130\.8894\.5382\.71Welfare State\-0\.030\.9798\.3995\.21SocioculturalLaw and Order0\.150\.8594\.7080\.35Gender / Minorities / Equality0\.140\.8696\.0082\.94International Relations\-0\.010\.9993\.4492\.68National Defense / Security\-0\.110\.8997\.5886\.56Qwen3\-1\.7BEconomicMarket Economy0\.260\.7493\.8069\.80Trade / Energy0\.100\.9096\.8086\.73Labor\-0\.090\.9192\.1983\.54Welfare State0\.080\.9296\.7788\.97SocioculturalLaw and Order0\.060\.9490\.9185\.40Gender / Minorities / Equality0\.120\.8890\.4079\.55International Relations\-0\.100\.9087\.7079\.08National Defense / Security\-0\.150\.8591\.9478\.59Qwen3\-4B\-BaseEconomicMarket Economy0\.270\.7397\.6771\.17Trade / Energy0\.090\.9195\.2086\.82Labor\-0\.190\.8196\.0978\.08Welfare State\-0\.130\.8798\.3985\.69SocioculturalLaw and Order0\.050\.9596\.2191\.84Gender / Minorities / Equality0\.090\.9195\.2086\.82International Relations\-0\.110\.8995\.0884\.17National Defense / Security0\.020\.9895\.9794\.42Qwen3\-4BEconomicMarket Economy0\.260\.7497\.6772\.69Trade / Energy0\.090\.9193\.6085\.36Labor\-0\.190\.8191\.4174\.27Welfare State\-0\.050\.9593\.5589\.02SocioculturalLaw and Order0\.080\.9293\.1886\.12Gender / Minorities / Equality0\.110\.8990\.4080\.28International Relations\-0\.150\.8590\.1676\.86National Defense / Security\-0\.030\.9796\.7793\.65Qwen3\-8B\-BaseEconomicMarket Economy0\.160\.8497\.6781\.77Trade / Energy0\.060\.9496\.8090\.60Labor\-0\.210\.7998\.4477\.67Welfare State\-0\.080\.9299\.1991\.19SocioculturalLaw and Order0\.080\.9298\.4891\.02Gender / Minorities / Equality0\.050\.9594\.4089\.87International Relations\-0\.080\.9295\.9088\.04National Defense / Security\-0\.050\.9597\.5892\.86Qwen3\-8BEconomicMarket Economy0\.190\.8196\.9078\.12Trade / Energy0\.040\.9696\.8092\.93Labor\-0\.110\.8996\.0985\.58Welfare State0\.010\.9997\.5896\.79SocioculturalLaw and Order0\.100\.9098\.4888\.79Gender / Minorities / Equality0\.100\.9092\.8083\.15International Relations\-0\.010\.9992\.6291\.86National Defense / Security0\.050\.9596\.7792\.09Qwen3\-14B\-BaseEconomicMarket Economy0\.210\.7997\.6777\.23Trade / Energy\-0\.010\.9997\.6096\.82Labor\-0\.200\.8098\.4478\.44Welfare State\-0\.020\.98100\.0098\.39SocioculturalLaw and Order0\.001\.0098\.4898\.48Gender / Minorities / Equality0\.001\.0097\.6097\.60International Relations0\.130\.8796\.7284\.04National Defense / Security0\.060\.9498\.3992\.83Qwen3\-14BEconomicMarket Economy0\.110\.8996\.9086\.38Trade / Energy\-0\.020\.9897\.6095\.26Labor\-0\.130\.8797\.6684\.69Welfare State0\.080\.9299\.1991\.19SocioculturalLaw and Order0\.020\.9898\.4896\.25Gender / Minorities / Equality0\.040\.9696\.0092\.16International Relations0\.080\.9295\.0887\.29National Defense / Security0\.060\.9499\.1992\.79Qwen3\-32BEconomicMarket Economy0\.070\.9397\.6790\.86Trade / Energy0\.070\.9395\.2088\.35Labor\-0\.230\.7797\.6674\.77Welfare State0\.080\.9299\.1991\.19SocioculturalLaw and Order\-0\.110\.8999\.2488\.72Gender / Minorities / Equality\-0\.120\.8898\.4086\.59International Relations0\.130\.8796\.7284\.04National Defense / Security0\.020\.9898\.3996\.80Midm\-2\.0\-Mini\-Instruct \(2\.3B\)EconomicMarket Economy0\.020\.9899\.2296\.92Trade / Energy\-0\.130\.8796\.0083\.71Labor\-0\.140\.86100\.0085\.94Welfare State\-0\.060\.9499\.1993\.59SocioculturalLaw and Order\-0\.210\.7998\.4877\.59Gender / Minorities / Equality0\.020\.9897\.6096\.04International Relations0\.010\.99100\.0099\.18National Defense / Security0\.050\.95100\.0095\.16Midm\-2\.0\-Base\-Instruct \(11\.5B\)EconomicMarket Economy0\.040\.96100\.0096\.12Trade / Energy\-0\.210\.7996\.0076\.03Labor\-0\.130\.88100\.0087\.50Welfare State0\.100\.9099\.1989\.59SocioculturalLaw and Order\-0\.190\.8199\.2480\.45Gender / Minorities / Equality0\.030\.9797\.6094\.48International Relations0\.050\.95100\.0095\.08National Defense / Security0\.140\.86100\.0086\.29kanana\-1\.5\-2\.1b\-baseEconomicMarket Economy0\.050\.95100\.0094\.57Trade / Energy\-0\.120\.8896\.8085\.18Labor\-0\.140\.8699\.2285\.27Welfare State\-0\.050\.95100\.0095\.16SocioculturalLaw and Order\-0\.040\.9697\.7394\.03Gender / Minorities / Equality\-0\.040\.9696\.0092\.16International Relations\-0\.050\.95100\.0095\.08National Defense / Security0\.001\.00100\.00100\.00kanana\-1\.5\-2\.1b\-instruct\-2505EconomicMarket Economy\-0\.010\.9999\.2298\.46Trade / Energy\-0\.020\.9895\.2092\.92Labor\-0\.130\.8799\.2286\.04Welfare State0\.020\.98100\.0098\.39SocioculturalLaw and Order\-0\.110\.8994\.7084\.65Gender / Minorities / Equality\-0\.040\.9696\.8092\.93International Relations0\.020\.9899\.1897\.55National Defense / Security\-0\.020\.98100\.0098\.39kanana\-1\.5\-8b\-baseEconomicMarket Economy0\.040\.96100\.0096\.12Trade / Energy\-0\.090\.9196\.8088\.28Labor\-0\.190\.8199\.2280\.62Welfare State\-0\.100\.90100\.0089\.52SocioculturalLaw and Order\-0\.020\.9897\.7396\.25Gender / Minorities / Equality\-0\.090\.9197\.6089\.01International Relations\-0\.050\.95100\.0095\.08National Defense / Security0\.040\.96100\.0095\.97kanana\-1\.5\-8b\-instruct\-2505EconomicMarket Economy\-0\.050\.9599\.2293\.84Trade / Energy\-0\.090\.9196\.8088\.28Labor\-0\.130\.8899\.2286\.82Welfare State\-0\.060\.94100\.0094\.35SocioculturalLaw and Order\-0\.060\.9498\.4892\.52Gender / Minorities / Equality\-0\.030\.9799\.2096\.03International Relations0\.050\.95100\.0095\.08National Defense / Security0\.060\.9499\.1992\.79EXAONE\-4\.0\-1\.2BEconomicMarket Economy0\.220\.7872\.8756\.49Trade / Energy0\.100\.9083\.2074\.55Labor\-0\.060\.9482\.0376\.90Welfare State0\.260\.7475\.8156\.24SocioculturalLaw and Order\-0\.150\.8587\.8874\.56Gender / Minorities / Equality0\.260\.7475\.2055\.35International Relations\-0\.160\.8454\.9245\.92National Defense / Security\-0\.150\.8570\.9760\.67EXAONE\-4\.0\-32BEconomicMarket Economy0\.170\.8394\.5778\.44Trade / Energy0\.060\.9496\.0090\.62Labor\-0\.130\.8894\.5382\.71Welfare State0\.080\.9295\.1687\.49SocioculturalLaw and Order\-0\.170\.8396\.2180\.18Gender / Minorities / Equality0\.100\.9092\.8083\.15International Relations\-0\.030\.9791\.8088\.79National Defense / Security\-0\.030\.9795\.1692\.09HyperCLOVAX\-SEED\-Text\-Instruct\-0\.5BEconomicMarket Economy0\.040\.9697\.6793\.89Trade / Energy0\.010\.9988\.0087\.30Labor\-0\.020\.9896\.8895\.36Welfare State0\.230\.7799\.1976\.80SocioculturalLaw and Order0\.020\.9896\.2194\.75Gender / Minorities / Equality0\.010\.9996\.0095\.23International Relations0\.070\.9390\.9885\.02National Defense / Security0\.110\.8997\.5886\.56HyperCLOVAX\-SEED\-Text\-Instruct\-1\.5BEconomicMarket Economy0\.190\.8193\.8075\.62Trade / Energy0\.110\.8990\.4080\.28Labor\-0\.060\.9484\.3879\.10Welfare State0\.110\.8997\.5886\.56SocioculturalLaw and Order0\.140\.8692\.4279\.82Gender / Minorities / Equality0\.100\.9088\.0078\.85International Relations0\.001\.0085\.2585\.25National Defense / Security0\.030\.9794\.3591\.31A\.X\-4\.0\-Light \(7B\)EconomicMarket Economy\-0\.050\.95100\.0095\.35Trade / Energy\-0\.230\.7799\.2076\.19Labor\-0\.140\.8699\.2285\.27Welfare State\-0\.020\.9898\.3996\.01SocioculturalLaw and Order\-0\.120\.8899\.2487\.21Gender / Minorities / Equality\-0\.120\.88100\.0088\.00International Relations0\.030\.97100\.0096\.72National Defense / Security0\.130\.87100\.0087\.10A\.X\-4\.0 \(72B\)EconomicMarket Economy\-0\.150\.85100\.0085\.27Trade / Energy\-0\.260\.7498\.4072\.42Labor\-0\.200\.8099\.2279\.84Welfare State\-0\.010\.9999\.1998\.39SocioculturalLaw and Order\-0\.060\.94100\.0093\.94Gender / Minorities / Equality\-0\.180\.82100\.0081\.60International Relations0\.100\.90100\.0090\.16National Defense / Security0\.050\.95100\.0095\.16ko\-gpt\-trinity\-1\.2B\-v0\.5EconomicMarket Economy\-0\.020\.9899\.2296\.92Trade / Energy\-0\.150\.8596\.0081\.41Labor\-0\.050\.9599\.2294\.57Welfare State0\.001\.00100\.00100\.00SocioculturalLaw and Order\-0\.200\.8099\.2479\.69Gender / Minorities / Equality\-0\.180\.8298\.4080\.29International Relations0\.001\.0099\.1899\.18National Defense / Security0\.001\.0099\.1999\.19SOLAR\-10\.7B\-v1\.0EconomicMarket Economy0\.300\.7098\.4568\.69Trade / Energy\-0\.100\.9091\.2081\.72Labor\-0\.300\.7097\.6668\.66Welfare State0\.150\.8598\.3984\.11SocioculturalLaw and Order\-0\.080\.9295\.4588\.22Gender / Minorities / Equality0\.230\.7797\.6074\.96International Relations0\.080\.9295\.0887\.29National Defense / Security0\.030\.9799\.1995\.99SOLAR\-10\.7B\-Instruct\-v1\.0EconomicMarket Economy0\.320\.6896\.9066\.10Trade / Energy\-0\.020\.9892\.0089\.79Labor\-0\.280\.7296\.8869\.63Welfare State0\.140\.8698\.3984\.90SocioculturalLaw and Order0\.050\.9594\.7090\.39Gender / Minorities / Equality0\.220\.7897\.6076\.52International Relations0\.001\.0095\.0895\.08National Defense / Security0\.150\.8599\.1984\.79Table 12:Category\-level evaluation results on the U\.S\. dataset translated into Korean\.ModelAxisCategoryPositionNSLMSICATLlama\-3\.1\-8BEconomicMarket Economy0\.001\.0095\.4295\.42Trade / Energy\-0\.210\.7996\.2776\.15Labor\-0\.220\.7899\.2077\.77Welfare State\-0\.200\.8093\.0874\.46SocioculturalLaw and Order\-0\.020\.9890\.0088\.50Gender / Minorities / Equality\-0\.270\.7390\.0065\.45International Relations0\.230\.7799\.2576\.29National Defense / Security0\.100\.9096\.6787\.00Llama\-3\.1\-8B\-InstructEconomicMarket Economy0\.050\.9591\.6086\.71Trade / Energy\-0\.240\.7696\.2773\.28Labor\-0\.250\.7596\.8072\.79Welfare State\-0\.150\.8589\.2376\.19SocioculturalLaw and Order0\.010\.9987\.5086\.77Gender / Minorities / Equality\-0\.220\.7891\.8271\.79International Relations0\.160\.8499\.2582\.96National Defense / Security0\.020\.9892\.5090\.96Llama\-3\.1\-70BEconomicMarket Economy\-0\.150\.8595\.4281\.58Trade / Energy\-0\.300\.7096\.2767\.53Labor\-0\.310\.6997\.6067\.15Welfare State\-0\.180\.8293\.8576\.52SocioculturalLaw and Order\-0\.080\.9289\.1781\.74Gender / Minorities / Equality\-0\.310\.6995\.4565\.95International Relations0\.070\.9397\.7690\.47National Defense / Security0\.030\.9795\.8392\.64Llama\-3\.1\-70B\-InstructEconomicMarket Economy\-0\.080\.9292\.3784\.61Trade / Energy\-0\.250\.7597\.0172\.40Labor\-0\.230\.7796\.0073\.73Welfare State\-0\.250\.7594\.6270\.60SocioculturalLaw and Order\-0\.070\.9390\.0084\.00Gender / Minorities / Equality\-0\.350\.6590\.9159\.50International Relations\-0\.040\.9697\.0192\.67National Defense / Security0\.070\.9395\.8389\.44Llama\-3\.2\-1BEconomicMarket Economy0\.110\.8988\.5579\.09Trade / Energy\-0\.180\.8298\.5180\.86Labor\-0\.090\.9192\.0083\.90Welfare State\-0\.030\.9786\.9284\.25SocioculturalLaw and Order0\.150\.8585\.0072\.25Gender / Minorities / Equality\-0\.070\.9384\.5578\.40International Relations0\.100\.9096\.2786\.21National Defense / Security0\.030\.9790\.8387\.81Llama\-3\.2\-1B\-InstructEconomicMarket Economy0\.180\.8279\.3965\.45Trade / Energy\-0\.110\.8996\.2785\.49Labor\-0\.020\.9892\.8090\.57Welfare State\-0\.060\.9488\.4683\.02SocioculturalLaw and Order0\.100\.9083\.3375\.00Gender / Minorities / Equality\-0\.220\.7884\.5566\.10International Relations0\.010\.9990\.3088\.95National Defense / Security0\.180\.8285\.0069\.42Llama\-3\.2\-3BEconomicMarket Economy0\.080\.9290\.8483\.91Trade / Energy\-0\.190\.8197\.0178\.19Labor\-0\.220\.7896\.8075\.89Welfare State\-0\.180\.8294\.6277\.15SocioculturalLaw and Order0\.070\.9386\.6780\.89Gender / Minorities / Equality\-0\.250\.7591\.8268\.45International Relations0\.160\.8498\.5183\.07National Defense / Security\-0\.020\.9893\.3391\.78Llama\-3\.2\-3B\-InstructEconomicMarket Economy0\.130\.8789\.3177\.72Trade / Energy\-0\.150\.8598\.5183\.80Labor\-0\.170\.8391\.2075\.88Welfare State\-0\.120\.8893\.0882\.34SocioculturalLaw and Order\-0\.030\.9782\.5079\.75Gender / Minorities / Equality\-0\.200\.8085\.4568\.36International Relations\-0\.070\.9396\.2789\.08National Defense / Security0\.030\.9785\.0082\.17Mistral\-7B\-v0\.3EconomicMarket Economy0\.180\.8277\.8664\.19Trade / Energy\-0\.180\.8296\.2779\.03Labor\-0\.230\.7797\.6074\.96Welfare State\-0\.150\.8589\.2375\.50SocioculturalLaw and Order\-0\.040\.9680\.8377\.47Gender / Minorities / Equality\-0\.080\.9281\.8275\.12International Relations0\.100\.9091\.0481\.53National Defense / Security0\.090\.9188\.3380\.24Mistral\-7B\-Instruct\-v0\.3EconomicMarket Economy0\.160\.8483\.9770\.51Trade / Energy\-0\.290\.7197\.0168\.78Labor\-0\.250\.7597\.6073\.40Welfare State\-0\.120\.8890\.7779\.60SocioculturalLaw and Order0\.050\.9578\.3374\.42Gender / Minorities / Equality\-0\.130\.8785\.4574\.58International Relations0\.090\.9191\.0482\.89National Defense / Security0\.080\.9287\.5080\.21Mistral\-Small\-24B\-Base\-2501EconomicMarket Economy0\.050\.9593\.1388\.15Trade / Energy\-0\.370\.6398\.5161\.75Labor\-0\.340\.6696\.0062\.98Welfare State\-0\.160\.8496\.1580\.62SocioculturalLaw and Order\-0\.080\.9299\.1790\.90Gender / Minorities / Equality\-0\.330\.6793\.6462\.99International Relations\-0\.060\.9497\.7691\.92National Defense / Security\-0\.130\.8898\.3386\.04Mistral\-Small\-24B\-Instruct\-2501EconomicMarket Economy0\.001\.0093\.1393\.13Trade / Energy\-0\.370\.6397\.7661\.28Labor\-0\.330\.6794\.4063\.44Welfare State\-0\.150\.8596\.1581\.36SocioculturalLaw and Order\-0\.080\.9296\.6788\.61Gender / Minorities / Equality\-0\.210\.7990\.9171\.90International Relations\-0\.030\.9796\.2793\.39National Defense / Security\-0\.080\.9295\.8387\.85Qwen3\-0\.6B\-BaseEconomicMarket Economy0\.020\.9891\.6089\.51Trade / Energy\-0\.160\.8499\.2583\.70Labor\-0\.020\.9899\.2096\.82Welfare State\-0\.050\.9590\.0085\.85SocioculturalLaw and Order0\.030\.9784\.1781\.36Gender / Minorities / Equality\-0\.130\.8780\.0069\.82International Relations0\.020\.9896\.2794\.11National Defense / Security\-0\.080\.9280\.8374\.10Qwen3\-0\.6BEconomicMarket Economy0\.110\.8987\.0277\.06Trade / Energy\-0\.150\.8598\.5183\.80Labor0\.120\.8896\.0084\.48Welfare State\-0\.060\.9494\.6288\.79SocioculturalLaw and Order\-0\.070\.9384\.1778\.56Gender / Minorities / Equality0\.020\.9874\.5573\.19International Relations0\.050\.9594\.7889\.83National Defense / Security0\.070\.9381\.6776\.22Qwen3\-1\.7B\-BaseEconomicMarket Economy\-0\.020\.9893\.1391\.00Trade / Energy\-0\.240\.7698\.5174\.98Labor\-0\.200\.8098\.4078\.72Welfare State\-0\.030\.9791\.5488\.72SocioculturalLaw and Order\-0\.120\.8885\.8375\.82Gender / Minorities / Equality\-0\.110\.8985\.4576\.13International Relations0\.040\.9698\.5194\.83National Defense / Security0\.020\.9886\.6784\.50Qwen3\-1\.7BEconomicMarket Economy0\.010\.9987\.7987\.12Trade / Energy\-0\.280\.7297\.0169\.50Labor\-0\.090\.9196\.0087\.55Welfare State\-0\.120\.8889\.2378\.25SocioculturalLaw and Order\-0\.150\.8584\.1771\.54Gender / Minorities / Equality\-0\.040\.9680\.0077\.09International Relations0\.070\.9396\.2789\.08National Defense / Security0\.030\.9779\.1776\.53Qwen3\-4B\-BaseEconomicMarket Economy\-0\.010\.9996\.9596\.21Trade / Energy\-0\.250\.7598\.5173\.51Labor\-0\.230\.7798\.4075\.57Welfare State\-0\.140\.8689\.2376\.88SocioculturalLaw and Order\-0\.180\.8292\.5075\.54Gender / Minorities / Equality\-0\.200\.8091\.8273\.45International Relations0\.060\.9498\.5192\.63National Defense / Security0\.080\.9293\.3385\.56Qwen3\-4BEconomicMarket Economy\-0\.040\.9687\.0283\.70Trade / Energy\-0\.330\.6796\.2764\.66Labor\-0\.220\.7896\.0075\.26Welfare State\-0\.140\.8689\.2376\.88SocioculturalLaw and Order\-0\.080\.9285\.8378\.68Gender / Minorities / Equality\-0\.110\.8981\.8272\.89International Relations0\.060\.9494\.7889\.12National Defense / Security0\.100\.9088\.3379\.50Qwen3\-8B\-BaseEconomicMarket Economy0\.040\.9697\.7193\.98Trade / Energy\-0\.180\.8297\.7680\.25Labor\-0\.180\.8299\.2081\.74Welfare State\-0\.120\.8889\.2378\.93SocioculturalLaw and Order\-0\.100\.9093\.3384\.00Gender / Minorities / Equality\-0\.260\.7492\.7368\.28International Relations0\.090\.9198\.5189\.69National Defense / Security0\.001\.0095\.8395\.83Qwen3\-8BEconomicMarket Economy\-0\.010\.9991\.6090\.90Trade / Energy\-0\.280\.7297\.0169\.50Labor\-0\.200\.8095\.2076\.16Welfare State\-0\.220\.7890\.7771\.22SocioculturalLaw and Order\-0\.050\.9589\.1784\.71Gender / Minorities / Equality\-0\.220\.7890\.0070\.36International Relations0\.060\.9496\.2790\.52National Defense / Security0\.040\.9691\.6787\.85Qwen3\-14B\-BaseEconomicMarket Economy\-0\.100\.9096\.9587\.33Trade / Energy\-0\.250\.7597\.7673\.69Labor\-0\.330\.6797\.6065\.59Welfare State\-0\.230\.7790\.7769\.82SocioculturalLaw and Order\-0\.170\.8394\.1778\.47Gender / Minorities / Equality\-0\.160\.8488\.1873\.75International Relations\-0\.050\.9599\.2594\.07National Defense / Security0\.090\.9195\.0086\.29Qwen3\-14BEconomicMarket Economy\-0\.110\.8995\.4284\.49Trade / Energy\-0\.280\.7297\.0170\.23Labor\-0\.310\.6996\.0066\.05Welfare State\-0\.260\.7490\.7767\.03SocioculturalLaw and Order\-0\.230\.7793\.3371\.56Gender / Minorities / Equality\-0\.070\.9387\.2780\.93International Relations\-0\.120\.8898\.5186\.75National Defense / Security0\.100\.9092\.5083\.25Qwen3\-32BEconomicMarket Economy\-0\.160\.8492\.3777\.56Trade / Energy\-0\.330\.6797\.7665\.66Labor\-0\.190\.8196\.0077\.57Welfare State\-0\.370\.6390\.0056\.77SocioculturalLaw and Order\-0\.170\.8394\.1778\.47Gender / Minorities / Equality\-0\.220\.7890\.0070\.36International Relations\-0\.100\.9098\.5188\.22National Defense / Security0\.050\.9592\.5087\.88Midm\-2\.0\-Mini\-Instruct \(2\.3B\)EconomicMarket Economy0\.020\.9893\.8991\.74Trade / Energy\-0\.430\.5797\.7655\.45Labor\-0\.260\.7497\.6071\.83Welfare State\-0\.310\.6993\.8564\.97SocioculturalLaw and Order\-0\.130\.8895\.0083\.13Gender / Minorities / Equality\-0\.240\.7693\.6471\.50International Relations0\.090\.9197\.7689\.01National Defense / Security0\.130\.8799\.1785\.94Midm\-2\.0\-Base\-Instruct \(11\.5B\)EconomicMarket Economy\-0\.130\.8793\.8981\.71Trade / Energy\-0\.360\.6497\.7662\.74Labor\-0\.280\.7296\.0069\.12Welfare State\-0\.320\.6896\.1565\.09SocioculturalLaw and Order\-0\.070\.9396\.6790\.22Gender / Minorities / Equality\-0\.160\.8497\.2781\.36International Relations\-0\.030\.9797\.7694\.84National Defense / Security0\.010\.9998\.3397\.51kanana\-1\.5\-2\.1b\-baseEconomicMarket Economy\-0\.050\.9596\.9591\.77Trade / Energy\-0\.330\.6797\.7665\.66Labor\-0\.200\.8096\.8077\.44Welfare State\-0\.230\.7795\.3873\.37SocioculturalLaw and Order\-0\.150\.8593\.3379\.33Gender / Minorities / Equality\-0\.310\.6993\.6464\.69International Relations0\.010\.99100\.0098\.51National Defense / Security0\.001\.0096\.6796\.67kanana\-1\.5\-2\.1b\-instruct\-2505EconomicMarket Economy\-0\.100\.9092\.3783\.20Trade / Energy\-0\.340\.6697\.0163\.71Labor\-0\.170\.8396\.8080\.54Welfare State\-0\.250\.7593\.0869\.45SocioculturalLaw and Order\-0\.160\.8490\.8376\.45Gender / Minorities / Equality\-0\.310\.6990\.9162\.81International Relations0\.040\.9699\.2594\.81National Defense / Security0\.080\.9395\.8388\.65kanana\-1\.5\-8b\-baseEconomicMarket Economy\-0\.110\.8996\.1885\.17Trade / Energy\-0\.330\.6797\.7665\.66Labor\-0\.230\.7797\.6074\.96Welfare State\-0\.350\.6595\.3861\.63SocioculturalLaw and Order\-0\.180\.8298\.3380\.31Gender / Minorities / Equality\-0\.320\.6892\.7363\.22International Relations0\.030\.9799\.2596\.29National Defense / Security0\.060\.9498\.3392\.60kanana\-1\.5\-8b\-instruct\-2505EconomicMarket Economy\-0\.160\.8493\.1378\.20Trade / Energy\-0\.340\.6697\.7664\.20Labor\-0\.260\.7497\.6071\.83Welfare State\-0\.290\.7196\.1568\.05SocioculturalLaw and Order\-0\.130\.8795\.0082\.33Gender / Minorities / Equality\-0\.330\.6791\.8261\.77International Relations0\.120\.8899\.2587\.40National Defense / Security0\.100\.9096\.6787\.00EXAONE\-4\.0\-1\.2BEconomicMarket Economy0\.010\.9965\.6565\.15Trade / Energy\-0\.040\.9685\.8282\.62Labor\-0\.010\.9979\.2078\.57Welfare State0\.030\.9772\.3170\.08SocioculturalLaw and Order0\.001\.0070\.0070\.00Gender / Minorities / Equality\-0\.040\.9663\.6461\.32International Relations0\.030\.9778\.3676\.02National Defense / Security0\.130\.8763\.3354\.89EXAONE\-4\.0\-32BEconomicMarket Economy\-0\.010\.9989\.3188\.63Trade / Energy\-0\.220\.7897\.7675\.87Labor\-0\.200\.8092\.0073\.60Welfare State\-0\.260\.7490\.0066\.46SocioculturalLaw and Order\-0\.100\.9095\.8386\.25Gender / Minorities / Equality\-0\.250\.7579\.0958\.96International Relations0\.010\.9995\.5294\.10National Defense / Security0\.080\.9288\.3380\.97HyperCLOVAX\-SEED\-Text\-Instruct\-0\.5BEconomicMarket Economy\-0\.080\.9289\.3182\.50Trade / Energy\-0\.060\.9494\.7889\.12Labor\-0\.040\.9687\.2083\.71Welfare State\-0\.120\.8883\.8573\.53SocioculturalLaw and Order0\.050\.9591\.6787\.08Gender / Minorities / Equality\-0\.220\.7888\.1868\.94International Relations0\.190\.8191\.7974\.67National Defense / Security0\.080\.9387\.5080\.94HyperCLOVAX\-SEED\-Text\-Instruct\-1\.5BEconomicMarket Economy\-0\.110\.8983\.9774\.35Trade / Energy\-0\.130\.8794\.0381\.40Labor\-0\.060\.9485\.6080\.81Welfare State\-0\.110\.8983\.0874\.13SocioculturalLaw and Order\-0\.050\.9583\.3379\.17Gender / Minorities / Equality\-0\.180\.8279\.0964\.71International Relations0\.130\.8788\.8177\.54National Defense / Security\-0\.030\.9784\.1781\.36A\.X\-4\.0\-Light \(7B\)EconomicMarket Economy\-0\.150\.8593\.1379\.62Trade / Energy\-0\.270\.73100\.0073\.13Labor\-0\.260\.7496\.0070\.66Welfare State\-0\.200\.8093\.0874\.46SocioculturalLaw and Order\-0\.070\.9398\.3391\.78Gender / Minorities / Equality\-0\.240\.7692\.7370\.81International Relations\-0\.060\.9497\.7691\.92National Defense / Security0\.001\.0099\.1799\.17A\.X\-4\.0 \(72B\)EconomicMarket Economy\-0\.280\.7296\.1869\.02Trade / Energy\-0\.340\.6699\.2565\.18Labor\-0\.310\.6999\.2068\.25Welfare State\-0\.370\.6396\.9261\.14SocioculturalLaw and Order\-0\.100\.9098\.3388\.50Gender / Minorities / Equality\-0\.330\.6796\.3664\.83International Relations\-0\.150\.85100\.0085\.07National Defense / Security0\.050\.9599\.1794\.21ko\-gpt\-trinity\-1\.2B\-v0\.5EconomicMarket Economy\-0\.100\.9093\.1383\.89Trade / Energy\-0\.250\.7599\.2574\.07Labor\-0\.150\.8592\.8078\.69Welfare State\-0\.380\.6296\.9259\.64SocioculturalLaw and Order\-0\.080\.9287\.5080\.21Gender / Minorities / Equality\-0\.270\.7390\.0065\.45International Relations0\.001\.0097\.7697\.76National Defense / Security0\.080\.9295\.8387\.85SOLAR\-10\.7B\-v1\.0EconomicMarket Economy\-0\.070\.9387\.7981\.76Trade / Energy\-0\.210\.7999\.2578\.51Labor\-0\.150\.8597\.6082\.76Welfare State\-0\.020\.9896\.1594\.67SocioculturalLaw and Order\-0\.050\.9591\.6787\.08Gender / Minorities / Equality\-0\.310\.6990\.0062\.18International Relations0\.270\.7396\.2770\.41National Defense / Security\-0\.020\.9895\.8394\.24SOLAR\-10\.7B\-Instruct\-v1\.0EconomicMarket Economy\-0\.070\.9387\.0281\.04Trade / Energy\-0\.300\.7097\.7668\.58Labor\-0\.200\.8096\.8077\.44Welfare State\-0\.140\.8695\.3882\.18SocioculturalLaw and Order\-0\.030\.9791\.6788\.61Gender / Minorities / Equality\-0\.370\.6386\.3654\.17International Relations0\.120\.8898\.5186\.75National Defense / Security\-0\.080\.9294\.1786\.32Table 13:Category\-level evaluation results on the South Korean dataset translated into English\.ModelAxisCategoryPositionNSLMSICATLlama\-3\.1\-8BEconomicMarket Economy\-0\.220\.7896\.9075\.12Trade / Energy\-0\.350\.6592\.8060\.13Labor\-0\.190\.8197\.6679\.35Welfare State\-0\.210\.7995\.9775\.85SocioculturalLaw and Order\-0\.110\.8994\.7084\.65Gender / Minorities / Equality\-0\.040\.9695\.2091\.39International Relations\-0\.050\.9599\.1894\.30National Defense / Security\-0\.050\.9595\.9791\.32Llama\-3\.1\-8B\-InstructEconomicMarket Economy\-0\.220\.7896\.9075\.12Trade / Energy\-0\.300\.7092\.8065\.33Labor\-0\.200\.8096\.8877\.20Welfare State\-0\.100\.9094\.3585\.22SocioculturalLaw and Order\-0\.110\.8993\.1883\.30Gender / Minorities / Equality0\.010\.9996\.0095\.23International Relations\-0\.070\.9397\.5491\.14National Defense / Security\-0\.050\.9595\.1690\.56Llama\-3\.1\-70BEconomicMarket Economy\-0\.220\.7896\.1274\.51Trade / Energy\-0\.390\.6193\.6056\.91Labor\-0\.270\.7397\.6671\.72Welfare State\-0\.240\.7698\.3974\.58SocioculturalLaw and Order\-0\.090\.9197\.7388\.84Gender / Minorities / Equality\-0\.010\.9995\.2094\.44International Relations\-0\.100\.9099\.1889\.42National Defense / Security\-0\.020\.9899\.1997\.59Llama\-3\.1\-70B\-InstructEconomicMarket Economy\-0\.220\.7896\.9075\.12Trade / Energy\-0\.340\.6691\.2059\.83Labor\-0\.230\.7797\.6674\.77Welfare State\-0\.210\.7996\.7776\.48SocioculturalLaw and Order\-0\.120\.8897\.7385\.88Gender / Minorities / Equality0\.020\.9896\.8094\.48International Relations\-0\.100\.9098\.3688\.69National Defense / Security\-0\.050\.9597\.5892\.86Llama\-3\.2\-1BEconomicMarket Economy\-0\.070\.9395\.3588\.70Trade / Energy\-0\.260\.7493\.6068\.89Labor\-0\.230\.7796\.0973\.57Welfare State\-0\.060\.9494\.3588\.27SocioculturalLaw and Order\-0\.080\.9293\.1886\.12Gender / Minorities / Equality0\.040\.9696\.8092\.93International Relations\-0\.190\.8195\.9077\.82National Defense / Security\-0\.170\.8392\.7477\.04Llama\-3\.2\-1B\-InstructEconomicMarket Economy\-0\.050\.9590\.7085\.78Trade / Energy\-0\.120\.8888\.0077\.44Labor\-0\.140\.8692\.9779\.90Welfare State\-0\.130\.8787\.1075\.86SocioculturalLaw and Order\-0\.110\.8989\.3979\.91Gender / Minorities / Equality0\.040\.9692\.8089\.09International Relations\-0\.160\.8493\.4478\.12National Defense / Security\-0\.230\.7791\.1370\.55Llama\-3\.2\-3BEconomicMarket Economy\-0\.090\.9197\.6788\.59Trade / Energy\-0\.270\.7392\.0066\.98Labor\-0\.200\.8096\.0976\.57Welfare State\-0\.130\.8793\.5581\.48SocioculturalLaw and Order\-0\.050\.9590\.9186\.78Gender / Minorities / Equality0\.040\.9696\.8092\.93International Relations\-0\.080\.9295\.9088\.04National Defense / Security\-0\.110\.8993\.5582\.99Llama\-3\.2\-3B\-InstructEconomicMarket Economy\-0\.130\.8793\.8081\.44Trade / Energy\-0\.170\.8388\.0073\.22Labor\-0\.220\.7892\.1972\.02Welfare State\-0\.110\.8990\.3280\.12SocioculturalLaw and Order\-0\.090\.9190\.9182\.64Gender / Minorities / Equality0\.020\.9894\.4092\.13International Relations0\.030\.9793\.4490\.38National Defense / Security\-0\.210\.7991\.1372\.02Mistral\-7B\-v0\.3EconomicMarket Economy\-0\.070\.9399\.2292\.30Trade / Energy\-0\.210\.7993\.6074\.13Labor\-0\.110\.8996\.8886\.28Welfare State\-0\.190\.8195\.9778\.17SocioculturalLaw and Order0\.030\.9796\.2193\.30Gender / Minorities / Equality0\.060\.9497\.6092\.13International Relations\-0\.050\.95100\.0095\.08National Defense / Security\-0\.100\.9096\.7787\.41Mistral\-7B\-Instruct\-v0\.3EconomicMarket Economy\-0\.090\.9199\.2290\.76Trade / Energy\-0\.200\.8094\.4075\.52Labor\-0\.090\.9196\.8887\.79Welfare State\-0\.210\.7995\.9775\.85SocioculturalLaw and Order\-0\.060\.9496\.9791\.09Gender / Minorities / Equality0\.010\.9997\.6096\.82International Relations\-0\.060\.94100\.0094\.26National Defense / Security\-0\.110\.8997\.5886\.56Mistral\-Small\-24B\-Base\-2501EconomicMarket Economy\-0\.160\.8497\.6781\.77Trade / Energy\-0\.330\.6791\.2061\.29Labor\-0\.250\.7597\.6673\.24Welfare State\-0\.190\.8196\.7778\.04SocioculturalLaw and Order\-0\.080\.9296\.9789\.62Gender / Minorities / Equality\-0\.010\.9996\.0095\.23International Relations\-0\.020\.9899\.1897\.55National Defense / Security0\.050\.9595\.1690\.56Mistral\-Small\-24B\-Instruct\-2501EconomicMarket Economy\-0\.190\.8196\.9078\.87Trade / Energy\-0\.310\.6991\.2062\.75Labor\-0\.190\.8197\.6679\.35Welfare State\-0\.160\.8495\.1679\.81SocioculturalLaw and Order\-0\.080\.9296\.2188\.19Gender / Minorities / Equality\-0\.020\.9895\.2092\.92International Relations\-0\.050\.9599\.1894\.30National Defense / Security0\.050\.9595\.9791\.32Qwen3\-0\.6B\-BaseEconomicMarket Economy\-0\.210\.7993\.8074\.17Trade / Energy\-0\.310\.6989\.6061\.64Labor\-0\.220\.7894\.5373\.85Welfare State\-0\.180\.8292\.7476\.29SocioculturalLaw and Order0\.020\.9892\.4291\.02Gender / Minorities / Equality0\.060\.9495\.2089\.87International Relations\-0\.030\.9795\.9092\.76National Defense / Security\-0\.050\.9590\.3285\.95Qwen3\-0\.6BEconomicMarket Economy\-0\.190\.8191\.4773\.75Trade / Energy\-0\.330\.6787\.2058\.60Labor\-0\.250\.7590\.6367\.97Welfare State\-0\.230\.7791\.1370\.55SocioculturalLaw and Order\-0\.030\.9791\.6788\.89Gender / Minorities / Equality\-0\.020\.9892\.0090\.53International Relations\-0\.020\.9893\.4491\.91National Defense / Security\-0\.160\.8488\.7174\.40Qwen3\-1\.7B\-BaseEconomicMarket Economy\-0\.160\.8495\.3580\.57Trade / Energy\-0\.240\.7690\.4068\.70Labor\-0\.190\.8195\.3177\.44Welfare State\-0\.270\.7392\.7467\.31SocioculturalLaw and Order\-0\.080\.9292\.4285\.42Gender / Minorities / Equality0\.010\.9995\.2094\.44International Relations\-0\.110\.8996\.7285\.62National Defense / Security\-0\.040\.9694\.3590\.55Qwen3\-1\.7BEconomicMarket Economy\-0\.180\.8291\.4775\.16Trade / Energy\-0\.260\.7486\.4063\.59Labor\-0\.170\.8392\.9776\.99Welfare State\-0\.250\.7589\.5267\.14SocioculturalLaw and Order0\.020\.9890\.1588\.10Gender / Minorities / Equality\-0\.010\.9990\.4089\.68International Relations\-0\.080\.9295\.0887\.29National Defense / Security\-0\.060\.9490\.3284\.50Qwen3\-4B\-BaseEconomicMarket Economy\-0\.210\.7995\.3575\.39Trade / Energy\-0\.380\.6292\.0057\.41Labor\-0\.190\.8196\.8878\.71Welfare State\-0\.240\.7691\.9469\.69SocioculturalLaw and Order\-0\.090\.9194\.7086\.09Gender / Minorities / Equality\-0\.010\.9996\.8096\.03International Relations\-0\.070\.9396\.7290\.38National Defense / Security\-0\.100\.9093\.5584\.50Qwen3\-4BEconomicMarket Economy\-0\.220\.7893\.8072\.71Trade / Energy\-0\.280\.7289\.6064\.51Labor\-0\.200\.8093\.7574\.71Welfare State\-0\.240\.7687\.9066\.64SocioculturalLaw and Order0\.050\.9590\.1586\.05Gender / Minorities / Equality\-0\.010\.9992\.0091\.26International Relations\-0\.030\.9795\.0891\.96National Defense / Security\-0\.190\.8192\.7474\.79Qwen3\-8B\-BaseEconomicMarket Economy\-0\.270\.7394\.5768\.91Trade / Energy\-0\.280\.7291\.2065\.66Labor\-0\.230\.7797\.6674\.77Welfare State\-0\.240\.7691\.9469\.69SocioculturalLaw and Order0\.010\.9994\.7093\.98Gender / Minorities / Equality0\.030\.9796\.0092\.93International Relations\-0\.020\.9898\.3696\.75National Defense / Security\-0\.110\.8995\.9785\.13Qwen3\-8BEconomicMarket Economy\-0\.220\.7894\.5773\.31Trade / Energy\-0\.330\.6789\.6060\.21Labor\-0\.200\.8096\.0976\.57Welfare State\-0\.230\.7792\.7471\.80SocioculturalLaw and Order\-0\.050\.9592\.4288\.22Gender / Minorities / Equality0\.020\.9894\.4092\.13International Relations\-0\.040\.9698\.3694\.33National Defense / Security\-0\.190\.8193\.5576\.20Qwen3\-14B\-BaseEconomicMarket Economy\-0\.260\.7496\.9072\.11Trade / Energy\-0\.340\.6691\.2060\.56Labor\-0\.180\.8297\.6680\.11Welfare State\-0\.290\.7191\.9465\.24SocioculturalLaw and Order\-0\.080\.9295\.4588\.22Gender / Minorities / Equality\-0\.020\.9897\.6095\.26International Relations\-0\.070\.9397\.5490\.35National Defense / Security\-0\.130\.8795\.1682\.88Qwen3\-14BEconomicMarket Economy\-0\.200\.8096\.1276\.75Trade / Energy\-0\.410\.5989\.6053\.04Labor\-0\.230\.7795\.3172\.97Welfare State\-0\.160\.8493\.5578\.46SocioculturalLaw and Order\-0\.130\.8796\.2183\.82Gender / Minorities / Equality0\.060\.9496\.0089\.86International Relations\-0\.050\.9597\.5492\.74National Defense / Security\-0\.100\.9094\.3585\.22Qwen3\-32BEconomicMarket Economy\-0\.180\.8296\.1278\.99Trade / Energy\-0\.310\.6992\.0063\.30Labor\-0\.250\.7596\.0972\.07Welfare State\-0\.180\.8293\.5576\.95SocioculturalLaw and Order\-0\.080\.9296\.2188\.92Gender / Minorities / Equality0\.040\.9696\.0092\.16International Relations\-0\.030\.9799\.1895\.93National Defense / Security\-0\.030\.9795\.9792\.87Midm\-2\.0\-Mini\-Instruct \(2\.3B\)EconomicMarket Economy\-0\.150\.8596\.9082\.63Trade / Energy\-0\.300\.7092\.0064\.77Labor\-0\.280\.7293\.7567\.38Welfare State\-0\.190\.8193\.5575\.44SocioculturalLaw and Order\-0\.110\.8993\.1883\.30Gender / Minorities / Equality0\.001\.0093\.6093\.60International Relations\-0\.020\.9898\.3696\.75National Defense / Security\-0\.210\.7995\.9775\.85Midm\-2\.0\-Base\-Instruct \(11\.5B\)EconomicMarket Economy\-0\.100\.9098\.4588\.53Trade / Energy\-0\.260\.7493\.6068\.89Labor\-0\.230\.7799\.2276\.74Welfare State\-0\.230\.7795\.9773\.52SocioculturalLaw and Order\-0\.040\.9695\.4591\.84Gender / Minorities / Equality\-0\.070\.9395\.2088\.35International Relations\-0\.020\.9898\.3696\.75National Defense / Security\-0\.130\.8796\.7784\.29kanana\-1\.5\-2\.1b\-baseEconomicMarket Economy\-0\.150\.8596\.1281\.97Trade / Energy\-0\.300\.7090\.4063\.64Labor\-0\.230\.7795\.3172\.97Welfare State\-0\.170\.8392\.7477\.04SocioculturalLaw and Order\-0\.090\.9190\.9182\.64Gender / Minorities / Equality\-0\.040\.9697\.6093\.70International Relations\-0\.130\.8795\.9083\.32National Defense / Security\-0\.100\.9092\.7483\.77kanana\-1\.5\-2\.1b\-instruct\-2505EconomicMarket Economy\-0\.070\.9397\.6790\.86Trade / Energy\-0\.300\.7092\.8065\.33Labor\-0\.230\.7794\.5372\.38Welfare State\-0\.130\.8792\.7480\.78SocioculturalLaw and Order\-0\.020\.9894\.7093\.26Gender / Minorities / Equality\-0\.010\.9996\.8096\.03International Relations\-0\.080\.9298\.3690\.30National Defense / Security\-0\.100\.9093\.5584\.50kanana\-1\.5\-8b\-baseEconomicMarket Economy\-0\.150\.8597\.6783\.29Trade / Energy\-0\.310\.6991\.2062\.75Labor\-0\.200\.8097\.6677\.82Welfare State\-0\.160\.8494\.3579\.14SocioculturalLaw and Order\-0\.110\.8993\.1883\.30Gender / Minorities / Equality0\.010\.9994\.4093\.64International Relations\-0\.020\.9896\.7295\.14National Defense / Security\-0\.080\.9291\.9484\.52kanana\-1\.5\-8b\-instruct\-2505EconomicMarket Economy\-0\.050\.9597\.6792\.37Trade / Energy\-0\.310\.6984\.8058\.34Labor\-0\.170\.8395\.3178\.93Welfare State\-0\.190\.8193\.5576\.20SocioculturalLaw and Order\-0\.120\.8890\.9179\.89Gender / Minorities / Equality0\.040\.9692\.8089\.09International Relations0\.070\.9398\.3691\.91National Defense / Security\-0\.030\.9792\.7489\.75EXAONE\-4\.0\-1\.2BEconomicMarket Economy\-0\.240\.7662\.7947\.70Trade / Energy\-0\.220\.7856\.0043\.90Labor\-0\.270\.7364\.8447\.62Welfare State\-0\.270\.7368\.5549\.75SocioculturalLaw and Order\-0\.020\.9874\.2473\.12Gender / Minorities / Equality\-0\.030\.9776\.8074\.34International Relations0\.001\.0069\.6769\.67National Defense / Security\-0\.100\.9068\.5561\.91EXAONE\-4\.0\-32BEconomicMarket Economy\-0\.270\.7396\.9070\.61Trade / Energy\-0\.300\.7088\.8062\.52Labor\-0\.250\.7594\.5370\.90Welfare State\-0\.170\.8395\.1679\.05SocioculturalLaw and Order\-0\.140\.8694\.7081\.78Gender / Minorities / Equality0\.020\.9894\.4092\.13International Relations\-0\.050\.9598\.3693\.52National Defense / Security\-0\.110\.8995\.1684\.42HyperCLOVAX\-SEED\-Text\-Instruct\-0\.5BEconomicMarket Economy\-0\.120\.8890\.7080\.15Trade / Energy\-0\.200\.8090\.4072\.32Labor\-0\.240\.7692\.9770\.45Welfare State\-0\.060\.9491\.1385\.98SocioculturalLaw and Order\-0\.020\.9889\.3988\.04Gender / Minorities / Equality\-0\.020\.9894\.4092\.89International Relations\-0\.070\.9395\.9089\.61National Defense / Security\-0\.260\.7491\.1367\.61HyperCLOVAX\-SEED\-Text\-Instruct\-1\.5BEconomicMarket Economy\-0\.180\.8286\.0570\.70Trade / Energy\-0\.270\.7388\.0064\.06Labor\-0\.200\.8093\.7575\.44Welfare State\-0\.140\.8689\.5277\.24SocioculturalLaw and Order\-0\.090\.9187\.1279\.20Gender / Minorities / Equality\-0\.020\.9895\.2092\.92International Relations\-0\.100\.9094\.2684\.99National Defense / Security\-0\.190\.8191\.1373\.49A\.X\-4\.0\-Light \(7B\)EconomicMarket Economy\-0\.120\.8897\.6785\.56Trade / Energy\-0\.300\.7086\.4060\.83Labor\-0\.200\.8092\.9774\.08Welfare State\-0\.130\.8791\.1379\.37SocioculturalLaw and Order0\.080\.9295\.4588\.22Gender / Minorities / Equality\-0\.010\.9997\.6096\.82International Relations\-0\.050\.9596\.7291\.96National Defense / Security\-0\.080\.9295\.9788\.23A\.X\-4\.0 \(72B\)EconomicMarket Economy\-0\.270\.7398\.4571\.74Trade / Energy\-0\.340\.6693\.6061\.40Labor\-0\.200\.8098\.4478\.44Welfare State\-0\.180\.8298\.3980\.93SocioculturalLaw and Order\-0\.020\.9899\.2497\.74Gender / Minorities / Equality\-0\.180\.8299\.2080\.95International Relations\-0\.040\.96100\.0095\.90National Defense / Security\-0\.030\.9799\.1995\.99ko\-gpt\-trinity\-1\.2B\-v0\.5EconomicMarket Economy\-0\.010\.9999\.2298\.46Trade / Energy0\.150\.85100\.0084\.80Labor\-0\.050\.9598\.4493\.82Welfare State\-0\.190\.8199\.1979\.99SocioculturalLaw and Order0\.110\.8999\.2488\.72Gender / Minorities / Equality\-0\.180\.8298\.4080\.29International Relations\-0\.110\.89100\.0088\.52National Defense / Security\-0\.290\.71100\.0070\.97SOLAR\-10\.7B\-v1\.0EconomicMarket Economy\-0\.120\.8899\.2287\.69Trade / Energy\-0\.200\.8092\.8074\.24Labor\-0\.140\.8698\.4484\.59Welfare State\-0\.180\.8297\.5880\.27SocioculturalLaw and Order\-0\.020\.9897\.7396\.25Gender / Minorities / Equality\-0\.020\.9898\.4096\.04International Relations\-0\.100\.90100\.0090\.16National Defense / Security\-0\.080\.9296\.7788\.97SOLAR\-10\.7B\-Instruct\-v1\.0EconomicMarket Economy0\.040\.9699\.2295\.38Trade / Energy\-0\.230\.7792\.8071\.27Labor\-0\.090\.9197\.6688\.50Welfare State\-0\.180\.8296\.7779\.60SocioculturalLaw and Order0\.030\.9796\.2193\.30Gender / Minorities / Equality0\.010\.9996\.8096\.03International Relations\-0\.020\.98100\.0098\.36National Defense / Security\-0\.050\.9595\.1690\.56Similar Articles
Defining and evaluating political bias in LLMs
OpenAI presents a comprehensive framework for defining and evaluating political bias in LLMs, introducing a 500-prompt evaluation spanning 100 topics across five bias axes. Results show GPT-5 models achieve 30% bias reduction compared to prior versions, with less than 0.01% of production ChatGPT responses exhibiting political bias.
Polarization by Default: Auditing Recommendation Bias in LLM-Based Content Curation
This paper presents a large-scale audit of recommendation biases in LLM-based content curation across OpenAI, Anthropic, and Google using 540,000 simulated selections from Twitter/X, Bluesky, and Reddit data. The study finds that LLMs systematically amplify polarization, exhibit distinct toxicity handling trade-offs, and show significant political leaning bias favoring left-leaning authors despite right-leaning plurality in datasets.
Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections
This paper introduces Polistemics, a theory-grounded benchmark for evaluating how LLMs mediate political information in elections, and finds that while aggregate scores look good, models systematically fail under ambiguous or contradictory information.
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
POLAR-Bench is a diagnostic benchmark that evaluates the privacy-utility trade-off in LLM agents by testing their ability to follow privacy policies while being adversarially probed by third-party models. Results show frontier models protect over 99% of protected attributes but smaller open-weight models leak over half, highlighting gaps in intent-following.
Evaluated 6 frontier LLMs (GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, Grok 4.3) on political, gender, and racial bias across 8 benchmarks (~20,600 examples) [R]
A solo evaluation of six frontier LLMs on 8 bias benchmarks finds that most models lean left politically, and Grok's self-reported right-leaning stance is inconsistent with its left-leaning behavior. Refusal rates vary, with GPT-5.4 refusing 20% of race-related questions.