Comparing and Modeling Argumentation in German Political Communication across Arenas
Summary
This paper presents a 17k-sentence corpus with annotations for argumentative passages across three German political arenas during COVID-19, and a pilot study on automatically identifying such passages, finding that boundaries are hard to pin down and models exhibit confirmation bias.
View Cached Full Text
Cached at: 08/04/26, 07:41 AM
# Comparing and Modeling Argumentation in German Political Communication across Arenas
Source: [https://arxiv.org/html/2608.00288](https://arxiv.org/html/2608.00288)
Nina Vikhrova1and Johannes Kühling2and Sebastian Haunss2and Sebastian Padó1 1IMS, Universität Stuttgart,2SOCIUM, Universität Bremen \{nina\.vikhrova\|pado\}@ims\.uni\-stuttgart\.de,\{kuehling\|haunss\}@uni\-bremen\.de
\(March 2026\)
###### Abstract
Deliberation, involving the formulation and exchange of arguments, forms an integral part of political decision making in democracies\. Argumentation patterns however differ substantially across different political arenas, such as plenary speeches and committee meetings\. However, despite a lot of interest in argumentation, there is comparatively little computational work on analyzing differences in patterns of political argumentation between arenas\. Our work addresses this research gap\. First, we present a 17k\-sentence corpus with annotation for argumentative passages \(argument and their justifications, both their boundaries and their categories\) across three German political arenas \(plenary speeches, committee meetings, and press conferences\), keeping the topic \(COVID\-19\) constant\. Our analysis of the corpus finds that contrary to expectations, justification by domain\-specific expertise is more frequent in press conferences than in committee meetings\. Second, we present a pilot study on automatically identifying such argumentative passages\. The results show that boundaries are hard to pin down, and models predictions additionally suffer from confirmation bias\.
Comparing and Modeling Argumentation in German Political Communication across Arenas
Nina Vikhrova1and Johannes Kühling2and Sebastian Haunss2and Sebastian Padó11IMS, Universität Stuttgart,2SOCIUM, Universität Bremen\{nina\.vikhrova\|pado\}@ims\.uni\-stuttgart\.de,\{kuehling\|haunss\}@uni\-bremen\.de
## 1Introduction
Even if existing democracies are often far from fulfilling the ideal of deliberative democraciesBächtigeret al\.\([2018](https://arxiv.org/html/2608.00288#bib.bib22)\), political decision\-making always contains elements of deliberation in parliaments, committees, and other arenas\. In these fora, actors make claims and argue to support their positions and to attack opposing political actors\. To advance their goal, they use various rhetorical and argumentation strategies, one of which is what Reyes has called “voices of expertise”\(Reyes,[2011](https://arxiv.org/html/2608.00288#bib.bib25), 786\), that is, the referral to experts or to expert knowledge in order to legitimise one’s claim\.
One should expect that argumentation strategies, in particular with regard to the appeal to expertise, are context dependent: Argumentation in plenary speeches should differ from argumentation in committees because the former address a general public, while the latter address mainly other MPs who are likely experts or at least experienced in the respective field\. Consequently, one could expect that expert argumentation should be more common in committees than in plenary speeches\. But studies so far have generally been limited to specific debates in one specific forum at a time\. Comparative studies on argumentation strategies are very rare\(e\.g\. Karlssonet al\.,[2024](https://arxiv.org/html/2608.00288#bib.bib26)\)because of the required amount of annotation\.
In our study, we aim to make one step forward in the study of political debate across arenas by annotating, evaluating, and modeling arguments from three different corpora of political speech in Germany: plenary speeches, health committee meetings, and press conferences\. We focus our analysis on the political discourse in these three arenas during the COVID pandemic in Germany \(2020–2022\)\. We choose this specific period and topic since during the COVID pandemic the reference to \(medical\) expert knowledge was omnipresent and we can therefore expect to see the use of expert arguments by many political actors, independent of whether they are themselves medical experts or not\. Focusing on one domain also reduces the impact of topic shifts on our analysis\.
Our paper makes two main contributions:
1. 1\.We present and analyse a manually annotated dataset of political discourse from three different German federal\-level political arenas – to our knowledge, the first such corpus\. It makes explicit argumentative passages, their justifications, and the type of justification used\. Contrary to expectations, expert arguments play a bigger role in press conferences than in committee meetings\.
2. 2\.We present a pilot study on the ability of LLMs to identify argumentative passages in political arenas within and across fora\. This would be a first step towards reducing manual effort for such analyses\. We find, however, that argumentative passages are hard to identify: Many models are reluctant to conclude that documents are non\-argumentative\.
## 2Related Work
#### Argument Mining for Political Texts\.
In computational linguistics, the automatic extraction of arguments \(“argument mining”\) has grown rapidly over the last ten years, but has focused mainly on text types where ample data was available, including social media and student texts, seeCabrio and Villata \([2018](https://arxiv.org/html/2608.00288#bib.bib13)\)for a review\. Generalizing argument mining methods across text types is often challenging\(Daxenbergeret al\.,[2017](https://arxiv.org/html/2608.00288#bib.bib12); Schaeferet al\.,[2022](https://arxiv.org/html/2608.00288#bib.bib2)\)\. When it comes to political texts, most studies focus on uniform datasets of current interest, typically election campaigns\(Lippi and Torroni,[2016](https://arxiv.org/html/2608.00288#bib.bib34); Visseret al\.,[2020](https://arxiv.org/html/2608.00288#bib.bib35)\)\. An exception isPoiaganova and Stede \([2025](https://arxiv.org/html/2608.00288#bib.bib5)\)who evaluate various argument mining\-related tasks within and across two political arenas \(US presidential debates and UN security council speeches\) and find reasonable generalization for most tasks across the two arenas\. Although argument mining in parliamentary debates is already an established field of research, our approach is timely because it applies a zero\-shot large language model to the task\. Moreover, our annotation scheme extends existing political argumentation corpora by distinguishing different types of justification, including domain\-specific justifications tailored to the analysis of expert argumentation\.
#### Political Argument Quality\.
From a political science perspective, Steenbergen et al\.’s Discourse Quality Index \(DCI\) is a normative evalutation of argumentation and in essence attempts to measure to which degree a debate fulfils the Habermasian ideal of deliberative politicsSteenbergenet al\.\([2003](https://arxiv.org/html/2608.00288#bib.bib24)\); Steffensmeier and Schenck\-Hamlin \([2008](https://arxiv.org/html/2608.00288#bib.bib23)\)\. Research from a critical discourse analysis perspective puts a stronger focus on the use of various legitimization strategies, namely legitimization through emotions, through references to the future, through rationality, expertise, and altruismReyes \([2011](https://arxiv.org/html/2608.00288#bib.bib25)\)\. In the context of argument mining, it has been found that argument quality is best considered as a multidimensional phenomenon\.Wachsmuthet al\.\([2017](https://arxiv.org/html/2608.00288#bib.bib27)\)have proposed a taxonomy for argumentation quality based on the cogency, reasonableness, and effectiveness of an argument\.
#### Political Arguments across Arenas\.
Regarding arena specific argumentation strategies, political science research has highlighted differences between “debating” \(e\.g\. UK\) and “working” \(e\.g\. Denmark, Sweden\) parliaments, and between parliamentary “frontstage” and “committee backstage”Karlssonet al\.\([2024](https://arxiv.org/html/2608.00288#bib.bib26)\)and pointed to more pragmatic argumentation in committeesAndone \([2016](https://arxiv.org/html/2608.00288#bib.bib29)\)\. On the computational side, there are few studies that compare properties of arguments across arenas\.Reiniget al\.\([2024](https://arxiv.org/html/2608.00288#bib.bib28)\)characterize political speech in terms of speech acts, which are however distributed fairly complementarily to arguments\. As mentioned above,Poiaganova and Stede \([2025](https://arxiv.org/html/2608.00288#bib.bib5)\)consider two political arenas\. Their mostly good generalization results imply that the arenas contain arguments with similar properties, but a deeper comparative analysis \(e\.g\., with regard to expertise\) was beyond the scope of their study\. We included government press conferences as the third arena because they were important during the COVID\-19 pandemic\. They became a central venue for announcing new measures, with the health minister regularly appearing alongside leading scientific experts\. framing the severity of the crisis and influencing subsequent media coverageHayek \([2024](https://arxiv.org/html/2608.00288#bib.bib42)\)\. While press conferences gained relevance during the pandemic, we expect a moderate share of expert argumentation\.
## 3Dataset
### 3\.1Corpus Selection
Our dataset of political discourse in the German political sphere consists of plenary speeches \(Bundestagsreden\), health committee protocols \(Gesundheitsausschussprotokolle\), and government press conferences \(Bundespressekonferenzen\)\. In this way, we aim to capture the full spectrum of parliamentary discursive arenas for one topical domain\. Plenary speeches represent politicized front\-stage deliberation and political competition\. The health committee is a parliamentary back stage\. And government press conferences are an interface between the general public and the parliament\.
While plenary speeches are held on the agenda set in parliament and are delivered in monologue form, health committee and press conferences are dialogical formats, mostly organised as Q&A sessions\. Press conferences usually begin with statements by the participants, after which the government spokesperson responds to questions\. The health committee is more or less scripted, as each party can invite \(scientific\) experts and question them according to the set agenda\.
Our sampling strategy was determined in advance, as we sought to capture all COVID\-related parliamentary discourse\. Or period of observation starts at the beginning of 2020 with the onset of the pandemic and ends in April 2023 after all pandemic measures had ended and Health Minister Karl Lauterbach officially declared the pandemic over\. We used the GermaParl CorpusBlaette and Leonhardt \([2023](https://arxiv.org/html/2608.00288#bib.bib31)\)as data source for plenary speeches\. The protocols of all public meetings of the health committee were downloaded from the archived websites of each legislative period \([https://www\.bundestag\.de/ausschuesse/gesundheit](https://www.bundestag.de/ausschuesse/gesundheit)\) and the written transcripts of the Bundespressekonferenzen were obtained from[jungundnaiv\.de](https://arxiv.org/html/2608.00288v1/jungundnaiv.de), a German journalist platform that provides complete coverage of these press conferencesJung & Naiv \([undated](https://arxiv.org/html/2608.00288#bib.bib37)\)\.
### 3\.2Preprocessing
The resulting dataset contains 118,745 speeches, 205 protocols and 512 press conferences\. We then filtered the dataset for COVID\-19\-related content using a dictionary of 1,958 keywords, which was using terms from the Leibniz Institute for the German Language relating to the COVID\-19 pandemicLeibniz IDS \([undated](https://arxiv.org/html/2608.00288#bib.bib36)\)\. We refined our approach with regular expression patterns to capture all relevant text segments and removed all terms not related to the medical domain\. These are mostly terms from socially triggered discourse without medical import, such as ‘click and collect’, ‘Drosten\-Ultra’ or ‘Wiesn\-Party’\. This resulted in a total of 9,737 COVID\-related texts \(9,299 speeches, 159 protocols, 279 press conferences\)\. We added metadata such as the speaker’s first and last name, date, party, and session\.
### 3\.3Annotation Schema
FollowingNordin and Schiappa \([2024](https://arxiv.org/html/2608.00288#bib.bib39), p\. 8\)we define our core concept, anargument, as a text segment consisting of one or more sentences and comprising a claim and its justification\. Claims are statements that suggest an action, such as a demand, plan, position, or recommendation\. Our definition thus represents a variant of the definition of arguments as claims supported by reasons\.
Rather than adopting annotation schemes from the argument mining literature, our approach is grounded in the framework of practical argumentation and informal logic, particularly drawing on the Toulmin model\. Initial attempts to operationalize speech act classification proved insufficient, as our first annotation sample did not align with the speech act categories\. We therefore adopted a claim\-based understanding of argumentation, showing how political claims are justified, rather than identifying argumentative units\. While argument mining primarily focuses on the identification and the reconstruction of arguments, our aim is to capture modes of practical argumentation in political discourse\. Specifically, we aim at classifying different modes of justification for political claims\.
To further distinguish these different types of justifications, we used a manual annotation of a subset of texts as an inductive baseline to assess whether stable categories could be derived\. Building on this foundation, we developed a classification scheme aimed at capturing systematic differences in argumentative styles\. We sought to differentiate between expert\-driven and layperson’s argumentation\. Therefore we defined three basic types:evidence\-based \(disciplinary\) justifications, which rely on data, studies, models, or measurable evidence;normative and pragmatic justifications, which are based on values, norms, feasibility, or utility; andideological and analogical justifications, which are based on opinions, rhetoric, or anecdotal and selective reasoning without verifiable evidenceToulmin \([2003](https://arxiv.org/html/2608.00288#bib.bib40)\); Reyes \([2011](https://arxiv.org/html/2608.00288#bib.bib25)\); van Dijk \([2000](https://arxiv.org/html/2608.00288#bib.bib41)\)\. To capture expert argumentation, we assumed that only evidence\-based \(disciplinary\) arguments qualify as expert arguments and added a binary indicator for whether the justification falls within the subject area\. The annotation scheme was further extended to distinguish between medical and non\-medical evidence\-based justifications\.
We sought to distinguish the COVID\-19 discourse thematically in order to examine whether specific subtopics of the debate offer increased expert argumentation\. We therefore classified each argument according to its subject\. We used the manual annotation as a baseline and further refined it into six main categories:healthcare\(treatment, access, personnel\),containment measures\(e\.g\. contact reduction, testing\),social issues\(socio\-economic issues, education, rights\),research\(methods, data, expertise\), international issues \(cross\-border issues\), andlogistics\(implementation, administration, communication, funding\)\.
Table 1:Agumentative spans and annotated labels\.Table[1](https://arxiv.org/html/2608.00288#S3.T1)provides English translations of two argumentative passages found in the corpus together with their annotation\. The examples show how the annotation scheme distinguishes between similarities and differences in argumentative reasoning\. In both passages, a clear claim is supported by a justification, but the justifications differ\. The first example represents an evidence\-based argument, where the justification refers to an empirically testable mechanism\. The second example, by contrast, exemplifies a normative argument, relying on a general appeal to urgency\.
### 3\.4Annotation Procedure
The annotation guidelines were presented to two annotators, both political science students, and first tested on a small sample\. This formed part of the inductive annotation procedure, in which new documents were selected and uploaded each week\. To ensure a sufficient number of positive examples in our data, we pre\-screen data for annotation using multiple factors\. We sampled by date and political party, to cover different phases of the COVID\-19 pandemic over time\. We validated the selection by quantifying the presence of claims with an automatic classifier for detecting claims within German newspaper articles and manifestosBlokkeret al\.\([2020](https://arxiv.org/html/2608.00288#bib.bib43)\)\. This classifier provides a lower bound for the presence of claims, due to the shift in text type which tends to decrease recall\. Our analysis showed that approximately 65% of the documents picked contained high number of claims, which we consider sufficient for our purposes\.
The comparatively larger number of speeches is primarily due to structural differences\. Speeches are substantially shorter and contain contributions by only a single speaker, whereas press conferences and committee protocols include multiple speakers and interaction sequences\. The annotation effort per speech was approximately one quarter of the time required for the other documents\.
The annotators coded the documents independently, while each document was annotated by both annotators, providing a reliable way to reduce subjectivity in the annotation process\. Difficult cases were discussed after each annotation round to further refine the guidelines and assess their robustness\. We used INCEpTION\(Klieet al\.,[2018](https://arxiv.org/html/2608.00288#bib.bib38)\)as annotation tool and configured the platform in accordance with the annotation guidelines by defining the annotation layers \(ARG/JUST\), features \(topic for ARG; justification type and domain for JUST\), and corresponding tagsets \(topics for ARG; types and domains for JUST\)\. As a result, the guidelines were accessible during annotation, and also structurally reflected in the annotation environment\. The estimated annotation speed was roughly 23–30 arguments per hour, when we consider 15\-25 min per speech and one hour for the other documents due to their length\. The full \(German\) guidelines are available in Appendix[A](https://arxiv.org/html/2608.00288#A1)\.
After the first annotation round, we analyzed the inter\-coder agreement \(ICA\) with the Gamma statistic\(Mathetet al\.,[2015](https://arxiv.org/html/2608.00288#bib.bib4)\), a chance\-corrected measure of agreement between sequence annotations that incorporates both labeling agreement \(like Cohen’sκ\\kappa\) and segmentation choice \(whichκ\\kappadoes not take into account\)\. Like forκ\\kappa, its range is\[−1,1\]\[\-1,1\], where 1 indicates perfect agreement and 0 agreement at chance level\.
As the second\-to\-last row of Table[2](https://arxiv.org/html/2608.00288#S3.T2)shows, ICA after initial annotation was mediocre, mainly due to differences in the identification of argument boundaries\. We therefore created a gold standard in iterative steps, similar to other semantic annotation projects \(e\.g\., OntoNotes,Hovyet al\.[2006](https://arxiv.org/html/2608.00288#bib.bib33)\)\. First, matching annotations were merged in the INCEpTION tool\. Second, arguments and justifications without matching codes were manually reviewed by a third annotator with advanced expertise in political science and refined in accordance with the codebook\. Thereafter, the ICA between the created gold standard and the initial annotations was calculated \(final row of Table[2](https://arxiv.org/html/2608.00288#S3.T2)\)\. This resulted inγ\\gammascores around 0\.7, reflecting the success of the adjudication procedure and demonstrating that the gold standard indeed represents a consensus between the two annotations\. Still, we take away that it is hard to get annotators to agree perfectly on argument spans\.
We do not carry out more detailed analysis of category correspondences, since these are hard to align in a combined segmentation\-and\-categorization task\.
Table 2:Descriptive statistics by arena: BT \(Bundestag speeches\), GA \(Gesundheitsausschuss protocols\), BPK \(Bundespressekonferenz press conference\)
### 3\.5Corpus Findings
Our corpus consists of 9,299 parliamentary speeches, 159 committee protocols, and 279 press conferences, for a total of 9,737 texts\. Note that the two latter types collapse multiple contributions in one document and are therefore much longer\. Of these, we fully annotated 87 documents, with an average of 19\.47 argument spans per documents\.
Table[2](https://arxiv.org/html/2608.00288#S3.T2)shows statistics for the resulting annotated corpus, both in total and broken down by arena\. We identified 1,689 argument spans \(ARG\) and 1,918 justification spans \(JUST\), amounting to 3,607 spans overall\. On average, this corresponds to 18\.35 ARG spans per 100 sentences and 20\.65 JUST spans per 100 sentences\. The ARG spans were differentiated by topic, while the JUST spans were differentiated by justification type\.
Looking at the topic distribution, logistics accounts for 60\.1%, followed by healthcare at 15\.1%\. Social issues amount to 10\.6%, and containment measures to 8\.3%, while international issues account for 2\.8% and research for 2\.8%\.
Regarding justification types, evidence\-based justifications make up only 7\.1% of all JUST spans, while normative and pragmatic justifications account for 88\.4% and ideological/analogical justifications for 4\.4%\. Among the evidence\-based JUST spans, 61\.5% were classified as medical and 38\.5% as non\-medical\. The sources also differ in their argument density\. Speeches show the highest average, with 23\.02 ARG and JUST spans per text, followed by protocols \(12\.43\)\. Press conferences contain only 3\.54 ARG and JUST spans per text\.
Overall, most JUST spans appear to be normatively driven justifications\. However, when looking at the share of evidence\-based justifications, BPK shows the highest proportion at 16\.7%\. Speeches account for 6\.6% and protocols for 5\.7%\.
The data reveal a discipline\-specific concentration of evidence\-based arguments, which is in line with our assumption that expert argumentation primarily takes place in evidence\-based arguments\. The observed argument density also corresponds to our expectations: Although plenary speeches are a central site of argumentation, the Q&A format of press conferences appears to constrain argument density, with committee protocols falling between the other two sources\. The press conferences appear to be the primary venue for expert argumentation, while plenary speeches are driven by normative justification\. Contrary to expectations, the health committee is only third\.
These findings align directly with our expectation that appeals to expertise are context dependent but contrast with the expectation that such forms would be concentrated in committees, where argument density is high and expert knowledge might be expected most strongly\. Although previous research focuses on public “frontstage” debate versus expertise\-oriented “backstage” committee deliberation, we observe that expert argument appears more often in government press conferences, rather than in committee deliberationKarlssonet al\.\([2024](https://arxiv.org/html/2608.00288#bib.bib26)\); Andone \([2016](https://arxiv.org/html/2608.00288#bib.bib29)\)\. Although committees show a dense argumentative interaction, expertise\-based justification appears most prominently in press conferences\. In plenary speeches, we observe more normative reasoning: Justifications rest on general appeals to collective protection rather than on empirical evidence\. For example, one MP argues that restrictions on fundamental rights are necessary to enable containment measures aimed at protecting the population\.111Example: ”Ja, zum Schutz der Bürger vor Corona sind Grundrechtseingriffe erforderlich, etwa durch Hygienekonzepte oder Abstandsmaßnahmen\.”
The prominence of expert arguments in government press conferences may reflect their institutional logic: Arguments in press conferences are used to communicate governmental decisions in a formalized setting, where speakers usually first deliver a relatively short statement and then respond to the journalists’ questions\. This setting seems to trigger expert\-like justifications to provide reasonable and deliberative answers\. Press conferences thus tend to feature dense and evidence\-based reasoning, where statements e\.g\. about increasing incidence rates, the spread of the B\.1\.1\.7 variant, and increasing ICU occupancy are used to justify concrete government decisions, reflecting a stronger reliance on empirically grounded argumentation\.222Example: ”Es gibt das Auftreten neuer Virusvarianten, das die epidemiologische Lage verändert hat\. Es gibt insbesondere die deutlich ansteckendere Virusvariante B\.1\.1\.7\. Es gibt stetig steigende Inzidenzen\. Wir liegen heute bei einem Wert von 140,9\. Wir haben eben auch eine wieder stetig steigende Zahl von Menschen, die auf den Intensivstationen behandelt werden müssen\. Deswegen hat das Kabinett heute eine Ergänzung des Infektionsschutzgesetzes beschlossen, und zwar, wie ich gesagt habe, als Formulierungshilfe für die Fraktionen von CDU/CSU und SPD\.”
In the health committee, experts provide descriptive input, e\.g\., when discussing staff vaccination rates\. This leads to a high number of normative justifications\. We attribute this to the interaction between politicians and invited experts, where justifications are more normative than evidence\-based\.333Example: ”Ich habe im Moment bei den Mitarbeitern eine Impfquote von 98 Prozent\. Zu keinem Zeitpunkt hatten wir größere Probleme, außer den üblichen, den Dienstplan zu gestalten\. In Moment habe ich, wenn ich die gesamte Gruppe anschaue, sechs Stellen zu besetzen\. Das sind also auch hier keine Horrorszenarien, wie sie immer wieder kommuniziert werden\.”
Thus, we can see overall that expertise is not simply tied to institutional proximity to expert knowledge, but also to communicative function and audience\. At the same time, the more strongly normative plenary speech aligns with theories of deliberative and representative politics that emphasize public justification, value contestation, and political positioning in plenary arenasSteenbergenet al\.\([2003](https://arxiv.org/html/2608.00288#bib.bib24)\)\.
### 3\.6Text Preparation
For modeling use, we split the datasets individually into three parts: training, development, and test\. We decided to make these parts equal\-sized \(33% each\), to accommodate the various possible approaches to modeling \(traditional training/evaluation; pre\-training and fine\-tuning; zero\-shot/few\-shot LLM use\) while keeping the test set large enough for robust evaluation\.444The dataset is available at Zenodo:[https://doi\.org/10\.5281/zenodo\.21718040](https://doi.org/10.5281/zenodo.21718040)
## 4Pilot Modeling Study
We now present the results of a pilot study limited to the automatic recognition of argumentative spans for the different arenas in our annotated corpus\. From our perspective, this task is the first step in a future more comprehensive modeling endeavour that will also consider other aspects of the annotation that we have creating, including claim and justification spans, claim domains and justification types\. We restrict ourselves to the argument recognition subtask since \(a\) it is the logically first step when facing unanalyzed political text: all other steps presuppose recognized arguments; \(b\) a reliable model for just this step would already play an enabling role in the analysis of argumentation in political texts across arenas \(cf\. Section[1](https://arxiv.org/html/2608.00288#S1)\)\.
Figure 1:Return rates of unfaithful quotations for combinations of models and prompt templates\.### 4\.1Model Choice
Since the models need to process German data, we require LLMs that have a good German proficiency\. Monolingual German models are rare, and those that exist such as LLäMmlein\(Pfisteret al\.,[2025](https://arxiv.org/html/2608.00288#bib.bib6)\)have undergone no or only basic instruction tuning\. Given that our task is complex, we focus on multilingual instruction\-tuned models\. For reasons of practicality and reproducibility, we restricted ourselves to open\-weights models in the range between 4B and 12B parameters\(Liesenfeldet al\.,[2023](https://arxiv.org/html/2608.00288#bib.bib16)\)\. Specifically, we investigated Phi4\-mini \(4B,Aboueleninet al\.[2025](https://arxiv.org/html/2608.00288#bib.bib11)\), Llama3\.1 \(8B,Grattafioriet al\.[2024](https://arxiv.org/html/2608.00288#bib.bib9)\), Gemma3 \(12B,Kamathet al\.[2025](https://arxiv.org/html/2608.00288#bib.bib7)\) and Qwen3 \(8B,Yanget al\.[2025](https://arxiv.org/html/2608.00288#bib.bib8)\)\.
### 4\.2Experimental design
We follow the current standard usage mode for zero\-shot text processing with large language models: we craft a prompt for our task, present the prompt with each input to LLMs, and parse the outputs\.
#### Prompt selection\.
The model was instructed to return argumentative passages related to COVID from a given text\. This is a task with a fairly complex output structure, namely a list of argumentative passages\. Models can misquote the input, justify their output \(contrary to instructions\), or add arbitrary comments\. Since we considered it unlikely that different LLMs would all work well with the same single prompt, we formulated a set of 8 prompt candidates, listed in Appendix[B](https://arxiv.org/html/2608.00288#A2), that varied in length and detail of definition, closeness to guidelines, instructions on output form as well as language of instruction\. These templates were evaluated on the train set\. We then selected the prompt that led to the lowest number of unfaithful quotations from the input, across arenas\. The results are shown in Figure[1](https://arxiv.org/html/2608.00288#S4.F1)\. Note that this evaluation does not require gold standard annotation, only input text\. The results show that one prompt \(\#4\) fails for all models and that models indeed prefer different prompts \(e\.g\., Phi4mini prompt \#1 vs\. prompt \#8 for Qwen3\-8B\)\. For the respective best prompt, the unfaithful quotation rate lies between 6% \(Gemma3\) and 16% \(Phi4mini\)\.
#### Answer interpretation\.
Our answer interpretation component accommodates variability in LLM output by trying to identify three types: \(a\), a list of arguments with quotation marks; \(b\), any other kind of list of arguments; \(c\) an empty list, typically accompanied with a justification\. For \(a\), the text in\-between quotation marks is extracted\. For \(b\), we match a list with a regular expression\. When some bullet point is recognized, everything located on the line is considered an argument\. For \(c\), we look for keyphrases related to an empty return, such as"leere Liste"or"keine argumentativen Textstellen"\. Outputs that remain unrecognized are treated as an empty list for the purposes of evaluation; however we keep track of the number of such ’invalid response’ cases and report them separately\.
### 4\.3Evaluation Metrics
We assess model predictions with a couple of evaluation metrics\. First, we consider three variants of sequence\-level F1score\. All of these are computed at the level of annotation spans \(not individual tokens\) and differ in what they consider true positives \(TPs\): In the strict match condition, a model\-predicted span is a TP only if it exactly matches a gold standard TP\. In the inclusion condition, a model\-predicted span is a TP if it is included in a gold standard span\. In the partial match condition, each predicted span is aligned with the best\-aligned gold span, and the TP count is the fraction of their overlap divided by predicted span length\. Thus, strict match is a harsh criterion while partial match rewards models already for mostly wrong predictions\. Inclusion attempts to strike a balance by rewarding models for not overpredicting spans\.
The other metric we consider is Gamma \(γ\\gamma\), which is usually used as a chance\-corrected metric for inter\-coder agreement \(ICA\) in sequence annotation \(cf\. Section[3\.4](https://arxiv.org/html/2608.00288#S3.SS4)\)\. Agreement between model\-predicted spans and gold\-standard spans withγ\\gammacan thus be compared to ICA numbers from above\.
Finally, we report the percentage of invalid model responses, the number of non\-argumentative contributions according to models and gold standard, and predicted argument density, measured as number of arguments per 100 sentences\.
### 4\.4Baselines
As points of comparison for the LLM\-based approach, we consider three simple baselines\. Baseline 1 does not predict any arguments\. Since this baseline has a recall of 0, its performance is always zero\. Baseline 2 predicts that the whole text is a single argument\. Since this baseline has a very low precision, its F1 score is also very low and we do not report detailed results\. Baseline 3 predicts that every sentence is \(its own\) argument\. This is the main baseline worth comparing against\.
Table 3:Performance of LLMs on argumentative passage extraction in three arenas \(BT, Bundestag; GA, Gesundheitsausschuss; BPK, Bundespressekonferenz\)\. Highest results for each metric in each arena boldfaced\. Baseline 3: Every sentence is its own argument\.
### 4\.5Modeling Results
The results for extracting argumentative passages from contributions in the three arenas are shown in Table[3](https://arxiv.org/html/2608.00288#S4.T3)\. We use the best prompt for each model \(cf\. Figure[1](https://arxiv.org/html/2608.00288#S4.F1)\)\. Our main observations are as follows\.
#### Exact match F1\.
Since exact match is a harsh success criterion for the models, results are in the single digits for almost all models\. This is not surprising, given the difficulty of human annotators to agree on exact boundaries \(cf\. Section[3\.4](https://arxiv.org/html/2608.00288#S3.SS4)\) but underlines the difficulty to precisely delineate argumentative passages in our data\. It is Gemma3, the largest model, which obtains the highest numbers on Bundestag and Gesundheitsausschuss\.
#### Partial / inclusion F1\.
According to these more lenient metrics, the identification of argumentative passages is not impossible, but still a seriously challenging task\. Phi4mini is the overall winner, showing robust performance across all three arenas\. This is striking, given that it is the smallest among our LLMs \(4B parameters\)\. Gemma3 and Llama3\.1 follow in second place for Bundestag and Gesundheitsausschuss, while QWen shows rather good performance on Bundespressekonferenz – see below for an explanation\.
#### Gamma scores\.
The Gamma scores indicate that some of the models perform close to or even at chance level, notably Llama3\.1 and Gemma3 on Bundestag\. This is different for the two other arenas, where Gamma scores reach 0\.35 \(Gesundheitsausschuss\) and 0\.44 \(Bundespressekonferenz\)\.This is a level comparable to the agreement among human annotators \(cf\. Table[2](https://arxiv.org/html/2608.00288#S3.T2)\)\. According this metric, QWen3 is the model with the overall most robust performance\.
#### Argument density\.
The results above can be explained, to an extent, by the differences in argument density among arenas\. According to Table[2](https://arxiv.org/html/2608.00288#S3.T2), the gold standard argument densities range between 22 arguments per 100 sentences \(Bundestag\), 12 arguments per 100 sentences \(Gesundheitsausschuss\) and 3 arguments per 100 sentences \(Bundespressekonferenz\)\. We observe striking differences among models regarding the predicted argument density\. All models underpredict arguments on the Bundestag data \(8–14 arg\. / 100 sent\.\) but most mildly overpredict arguments on the Gesundheitausschuss and extremely overpredict on the Bundespressekonferenz \(except QWen3\)\. The most extreme case is Gemma3 which predicts 26 arguments per 100 sentences on BPK, almost 10 times the correct density\. This naturally results in low precision\. In comparison, Phi4mini estimates the argument density most accurately on Bundestag and Gesundheitsausschuss, which accounts for its good performance on these corpora\. Only on Bundespressekonferenz, its density estimate is off; here, it is QWen3 which comes closest\.
The differences between assumed argument density across arenas is particularly striking since all of these models are zero\-shot\. The only contact that the models have had with the evaluation data is the prompt selection step \(cf\. Section[4\.2](https://arxiv.org/html/2608.00288#S4.SS2)\)\.
Across all models and arenas, the closer a model manages to get to the argument density in the gold standard, the better its predictions: F1 scores \(partial match\) and density accuracy \(difference between gold and predicted density, divided by gold density\) are inversely correlated \(Spearman’sρ\\rho=\-0\.56,nn=12,pp=0\.06\)\. Arguably, the lower argument density of Bundespressekonferenz stems from the fact that about 80 percent \(406/490\) of the contributions are purely descriptive, with no argumentation present\. As the numbers on non\-argumentative contributions according to models and gold\-standards show, this is a major stumbling block for some of the models\.
#### Baseline\.
The baseline which assumes that every sentence is a separate argument predicts, by definition, an argument density of 100, and never returns an empty argument list\. It performs surprisingly strongly\. In terms of F1 scores, the baseline outperforms the best LLMs on the Bundestag and Gesundheitsausschuss arenas and does about as well as the median LLM on Bundespressekonferenz\. However, the much smaller Gamma scores on all arenas \(even negative on BT and GA\) provide evidence for the essentially random nature of the predictions\. Again, the differences among the arenas reflect their differences in argument density: a baseline which assumes that every sentence is an argument performs better for arenas with denser argumentation\.
## 5Conclusions and Future Work
This paper investigated the differences in argumentation patterns among different arenas in German political discourse during the COVID\-19 pandemic\. In the absence of existing corpora to ground our analysis, we make two contributions: \(a\) manual annotation of a 17k\-sentence corpus comprising three important political arenas with argumentative passages, justifications, and corresponding categories; and \(b\) a pilot study to automatically recognize argument boundaries as a first step towards automating the information that we annotated\.
Our main insights are as follows\. First, identifying the boundaries of argumentative passages in political discourse is difficult both for human annotators and for models\. We also found lower agreement compared to previous studies\(Haddadanet al\.,[2019](https://arxiv.org/html/2608.00288#bib.bib32); Poiaganova and Stede,[2025](https://arxiv.org/html/2608.00288#bib.bib5)\)\. The extent to which this is due to our formulation of the task as sequence identification \(and not sentence\-level classification\) is a matter of future work\. An alternative route would be to embrace the differences among annotators in the spirit of perspectivism\(Falket al\.,[2024](https://arxiv.org/html/2608.00288#bib.bib30)\)\.
Second, we found major differences in argumentation patterns among our three arenas: the Bundespressekonferenz, while having the lowest argument density, shows the highest proportion of domain expertise\-related justifications, while in plenary speeches normative justifications dominate\.
Third, a major contributing reason for the difficulty that zero\-shot LLMs have in identifying argumentative passages is that they struggle with estimating the argument density in input texts\. This is likely because the models are trained to accept the presuppositions of instructions \(in our case, the presence of arguments in the documents\) and thus are reluctant to dismiss these presuppositions\. As a consequence, our LLMs found it hard to beat a simple ’everything is an argument’ baseline on the Bundestag and Gesundheitsausschuss texts\. In future work, we plan to explore the role of simple embedding\-based classifiers to ’pre\-screen’ documents for argumentative content\(Ruiz\-Dolz and Lawrence,[2023](https://arxiv.org/html/2608.00288#bib.bib1)\)\.
Other modeling options that were out of scope for this first study is the use of few\-shot setups which could provide LLMs with priors about argument density\(Schaeferet al\.,[2022](https://arxiv.org/html/2608.00288#bib.bib2)\); replacing LLMs with ’more traditional’ embedding\-based classifiers that have shown good performance for argument recognition in previous work\(Poiaganova and Stede,[2025](https://arxiv.org/html/2608.00288#bib.bib5)\); and modeling the remaining aspects of our annotation \(justification spans as well as argument and justification categories\)\.
## Limitations
At the corpus level, our annotation effort is focused on a single domain \(COVID\-19\)\. While we believe this to be a reasonable choice \(cf\. Section[3\.1](https://arxiv.org/html/2608.00288#S3.SS1)\), we acknowledge that our findings may not generalize straightforwardly to other domains\. Related to this, we only consider a single committee \(the health committee – Gesundheitsausschuss\); investigations of other domains should consider other committees\.
At the model level, we have considered five ’consumer\-sizes’ open\-weights LLMs but considered neither simpler solutions \(embedding\-based models\) nor very large properietary LLMs which might provide considerably better performance\. Prompts were also selected from a small number of templates\.
## References
- A\. Abouelenin, A\. Ashfaq, A\. Atkinson, H\. Awadalla, N\. Bach, J\. Bao, A\. Benhaim, M\. Cai, V\. Chaudhary, C\. Chen, D\. Chen, D\. Chen, J\. Chen, W\. Chen, Y\. Chen, Y\. Chen, Q\. Dai, X\. Dai, R\. Fan, M\. Gao, M\. Gao, A\. Garg, A\. Goswami, J\. Hao, A\. Hendy, Y\. Hu, X\. Jin, M\. Khademi, D\. Kim, Y\. J\. Kim, G\. Lee, J\. Li, Y\. Li, C\. Liang, X\. Lin, Z\. Lin, M\. Liu, Y\. Liu, G\. Lopez, C\. Luo, P\. Madan, V\. Mazalov, A\. Mitra, A\. Mousavi, A\. Nguyen, J\. Pan, D\. Perez\-Becker, J\. Platin, T\. Portet, K\. Qiu, B\. Ren, L\. Ren, S\. Roy, N\. Shang, Y\. Shen, S\. Singhal, S\. Som, X\. Song, T\. Sych, P\. Vaddamanu, S\. Wang, Y\. Wang, Z\. Wang, H\. Wu, H\. Xu, W\. Xu, Y\. Yang, Z\. Yang, D\. Yu, I\. Zabir, J\. Zhang, L\. L\. Zhang, Y\. Zhang, and X\. Zhou \(2025\)Phi\-4\-Mini technical report: compact yet powerful multimodal language models via mixture\-of\-LoRAs\.External Links:2503\.01743,[Link](https://arxiv.org/abs/2503.01743)Cited by:[§4\.1](https://arxiv.org/html/2608.00288#S4.SS1.p1.1)\.
- Argumentative patterns in the political domain: the case of European parliamentary committees of inquiry\.Argumentation30\(1\),pp\. 45–60\.External Links:ISSN 1572\-8374,[Document](https://dx.doi.org/10.1007/s10503-015-9372-4)Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px3.p1.1),[§3\.5](https://arxiv.org/html/2608.00288#S3.SS5.p7.1)\.
- A\. Bächtiger, J\. S\. Dryzek, J\. Mansbridge, and M\. E\. Warren \(2018\)Deliberative democracy\. an introduction\.InThe Oxford Handbook of Deliberative Democracy,Oxford Handbooks,pp\. 1–31\.External Links:ISBN 978\-0\-19\-106456\-2Cited by:[§1](https://arxiv.org/html/2608.00288#S1.p1.1)\.
- A\. Blaette and C\. Leonhardt \(2023\)GermaParl corpus of plenary protocols\.Technical reportZenodo,zenodo\.Note:\[Data set\]External Links:[Document](https://dx.doi.org/10.5281/zenodo.10416536),[Link](https://doi.org/10.5281/zenodo.10416536)Cited by:[§3\.1](https://arxiv.org/html/2608.00288#S3.SS1.p3.1)\.
- N\. Blokker, E\. Dayanik, G\. Lapesa, and S\. Padó \(2020\)Swimming with the tide? positional claim detection across political text types\.InProceedings of the Fourth Workshop on Natural Language Processing and Computational Social Science,D\. Bamman, D\. Hovy, D\. Jurgens, B\. O’Connor, and S\. Volkova \(Eds\.\),Online,pp\. 24–34\.External Links:[Link](https://aclanthology.org/2020.nlpcss-1.3/),[Document](https://dx.doi.org/10.18653/v1/2020.nlpcss-1.3)Cited by:[§3\.4](https://arxiv.org/html/2608.00288#S3.SS4.p1.1)\.
- E\. Cabrio and S\. Villata \(2018\)Five years of argument mining: a data\-driven analysis\.InProceedings of the Twenty\-Seventh International Joint Conference on Artificial Intelligence \(IJCAI 2018\),pp\. 5427–5433\.External Links:[Document](https://dx.doi.org/10.24963/ijcai.2018/764)Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Daxenberger, S\. Eger, I\. Habernal, C\. Stab, and I\. Gurevych \(2017\)What is the essence of a claim? Cross\-domain claim identification\.InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing,M\. Palmer, R\. Hwa, and S\. Riedel \(Eds\.\),Copenhagen, Denmark,pp\. 2055–2066\.External Links:[Link](https://aclanthology.org/D17-1218/),[Document](https://dx.doi.org/10.18653/v1/D17-1218)Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px1.p1.1)\.
- N\. Falk, A\. Waldis, and I\. Gurevych \(2024\)Overview of PerspectiveArg2024 the first shared task on perspective argument retrieval\.InProceedings of the 11th Workshop on Argument Mining \(ArgMining 2024\),Y\. Ajjour, R\. Bar\-Haim, R\. El Baff, Z\. Liu, and G\. Skitalinskaya \(Eds\.\),Bangkok, Thailand,pp\. 130–149\.External Links:[Link](https://aclanthology.org/2024.argmining-1.14/),[Document](https://dx.doi.org/10.18653/v1/2024.argmining-1.14)Cited by:[§5](https://arxiv.org/html/2608.00288#S5.p2.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma \(2024\)The Llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§4\.1](https://arxiv.org/html/2608.00288#S4.SS1.p1.1)\.
- S\. Haddadan, E\. Cabrio, and S\. Villata \(2019\)Yes, we can\! mining arguments in 50 years of US presidential campaign debates\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 4684–4690\.External Links:[Link](https://aclanthology.org/P19-1463/),[Document](https://dx.doi.org/10.18653/v1/P19-1463)Cited by:[§5](https://arxiv.org/html/2608.00288#S5.p2.1)\.
- L\. Hayek \(2024\)Media framing of government crisis communication during Covid\-19\.Media and Communication12,pp\. 7774\.External Links:ISSN 2183\-2439,[Link](https://www.cogitatiopress.com/mediaandcommunication/article/view/7774),[Document](https://dx.doi.org/10.17645/mac.7774)Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px3.p1.1)\.
- E\. Hovy, M\. Marcus, M\. Palmer, L\. Ramshaw, and R\. Weischedel \(2006\)OntoNotes: the 90% solution\.InProceedings of the Human Language Technology Conference of the NAACL, Companion Volume: Short Papers,R\. C\. Moore, J\. Bilmes, J\. Chu\-Carroll, and M\. Sanderson \(Eds\.\),New York City, USA,pp\. 57–60\.External Links:[Link](https://aclanthology.org/N06-2015/)Cited by:[§3\.4](https://arxiv.org/html/2608.00288#S3.SS4.p5.1)\.
- Jung & Naiv \(undated\)Bundespressekonferenz transcripts\.Note:Accessed 25\.06\.2025External Links:[Link](https://www.jungundnaiv.de/)Cited by:[§3\.1](https://arxiv.org/html/2608.00288#S3.SS1.p3.1)\.
- K\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. Hussenot \(2025\)Gemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§4\.1](https://arxiv.org/html/2608.00288#S4.SS1.p1.1)\.
- C\. Karlsson, T\. Persson, and M\. Mårtensson \(2024\)Do members of parliament express more opposition in the plenary than in the committee? Comparing frontstage and backstage behaviour in five national parliaments\.Parliamentary Affairs77\(1\),pp\. 173–195\.External Links:ISSN 0031\-2290,[Document](https://dx.doi.org/10.1093/pa/gsac016)Cited by:[§1](https://arxiv.org/html/2608.00288#S1.p2.1),[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px3.p1.1),[§3\.5](https://arxiv.org/html/2608.00288#S3.SS5.p7.1)\.
- J\. Klie, M\. Bugert, B\. Boullosa, R\. Eckart de Castilho, and I\. Gurevych \(2018\)The inception platform: machine\-assisted and knowledge\-oriented interactive annotation\.InProceedings of System Demonstrations of the 27th International Conference on Computational Linguistics \(COLING 2018\),Santa Fe, New Mexico, USA\.Cited by:[§3\.4](https://arxiv.org/html/2608.00288#S3.SS4.p3.1)\.
- Leibniz IDS \(undated\)Word list for the COVID\-19 pandemic\.Note:Accessed 25\.06\.2025External Links:[Link](https://www.owid.de/docs/neo/listen/corona.jsp)Cited by:[§3\.2](https://arxiv.org/html/2608.00288#S3.SS2.p1.1)\.
- A\. Liesenfeld, A\. Lopez, and M\. Dingemanse \(2023\)Opening up ChatGPT: tracking openness, transparency, and accountability in instruction\-tuned text generators\.InProceedings of the 5th International Conference on Conversational User Interfaces,pp\. 1–6\.External Links:[Document](https://dx.doi.org/10.1145/3571884.3604316)Cited by:[§4\.1](https://arxiv.org/html/2608.00288#S4.SS1.p1.1)\.
- M\. Lippi and P\. Torroni \(2016\)Argument mining from speech: detecting claims in political debates\.Proceedings of the AAAI Conference on Artificial Intelligence30\(1\)\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/10384),[Document](https://dx.doi.org/10.1609/aaai.v30i1.10384)Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Mathet, A\. Widlöcher, and J\. Métivier \(2015\)The unified and holistic method gamma \(γ\\gamma\) for inter\-annotator agreement measure and alignment\.Computational Linguistics41\(3\),pp\. 437–479\.External Links:ISSN 0891\-2017,[Document](https://dx.doi.org/10.1162/COLI%5Fa%5F00227),[Link](https://doi.org/10.1162/COLI_a_00227),https://direct\.mit\.edu/coli/article\-pdf/41/3/437/1806629/coli\_a\_00227\.pdfCited by:[§3\.4](https://arxiv.org/html/2608.00288#S3.SS4.p4.4)\.
- J\. P\. Nordin and E\. Schiappa \(2024\)Argumentation: keeping faith with reason\.2 edition,Routledge,New York\(en\)\.External Links:ISBN 978\-1\-003\-41526\-8,[Link](https://www.taylorfrancis.com/books/9781003415268),[Document](https://dx.doi.org/10.4324/9781003415268)Cited by:[§3\.3](https://arxiv.org/html/2608.00288#S3.SS3.p1.1)\.
- J\. Pfister, J\. Wunderle, and A\. Hotho \(2025\)LLäMmlein: transparent, compact and competitive German\-only language models from scratch\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 2227–2246\.External Links:[Link](https://aclanthology.org/2025.acl-long.111/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.111),ISBN 979\-8\-89176\-251\-0Cited by:[§4\.1](https://arxiv.org/html/2608.00288#S4.SS1.p1.1)\.
- M\. Poiaganova and M\. Stede \(2025\)From debates to diplomacy: argument mining across political registers\.InProceedings of the 12th Argument mining Workshop,E\. Chistova, P\. Cimiano, S\. Haddadan, G\. Lapesa, and R\. Ruiz\-Dolz \(Eds\.\),Vienna, Austria,pp\. 205–216\.External Links:[Link](https://aclanthology.org/2025.argmining-1.20/),[Document](https://dx.doi.org/10.18653/v1/2025.argmining-1.20),ISBN 979\-8\-89176\-258\-9Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2608.00288#S5.p2.1),[§5](https://arxiv.org/html/2608.00288#S5.p5.1)\.
- I\. Reinig, I\. Rehbein, and S\. P\. Ponzetto \(2024\)How to do politics with words: investigating speech acts in parliamentary debates\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),Torino, Italia,pp\. 8287–8300\.External Links:[Link](https://aclanthology.org/2024.lrec-main.727/)Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px3.p1.1)\.
- A\. Reyes \(2011\)Strategies of legitimization in political discourse: from words to actions\.Discourse & Society22\(6\),pp\. 781–807\.External Links:ISSN 0957\-9265,[Document](https://dx.doi.org/10.1177/0957926511419927)Cited by:[§1](https://arxiv.org/html/2608.00288#S1.p1.1),[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2608.00288#S3.SS3.p3.1)\.
- R\. Ruiz\-Dolz and J\. Lawrence \(2023\)Detecting argumentative fallacies in the wild: problems and limitations of large language models\.InProceedings of the 10th Workshop on Argument Mining,M\. Alshomary, C\. Chen, S\. Muresan, J\. Park, and J\. Romberg \(Eds\.\),Singapore,pp\. 1–10\.External Links:[Link](https://aclanthology.org/2023.argmining-1.1/),[Document](https://dx.doi.org/10.18653/v1/2023.argmining-1.1)Cited by:[§5](https://arxiv.org/html/2608.00288#S5.p4.1)\.
- R\. Schaefer, R\. Knaebel, and M\. Stede \(2022\)On selecting training corpora for cross\-domain claim detection\.InProceedings of the 9th Workshop on Argument Mining,G\. Lapesa, J\. Schneider, Y\. Jo, and S\. Saha \(Eds\.\),Online and in Gyeongju, Republic of Korea,pp\. 181–186\.External Links:[Link](https://aclanthology.org/2022.argmining-1.17/)Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2608.00288#S5.p5.1)\.
- M\. R\. Steenbergen, A\. Bächtiger, M\. Spörndli, and J\. Steiner \(2003\)Measuring political deliberation: a discourse quality index\.Comparative European Politics1\(1\),pp\. 21–48\.External Links:[Document](https://dx.doi.org/10.1057/palgrave.cep.6110002)Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px2.p1.1),[§3\.5](https://arxiv.org/html/2608.00288#S3.SS5.p10.1)\.
- T\. Steffensmeier and W\. Schenck\-Hamlin \(2008\)Argument quality in public deliberations\.Argumentation and Advocacy45\(1\),pp\. 21–36\.External Links:ISSN 1051\-1431,[Document](https://dx.doi.org/10.1080/00028533.2008.11821693)Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px2.p1.1)\.
- S\. E\. Toulmin \(2003\)The uses of argument\.2 edition,Cambridge University Press\.Cited by:[§3\.3](https://arxiv.org/html/2608.00288#S3.SS3.p3.1)\.
- T\. van Dijk \(2000\)Ideology: a multidisciplinary approach\.SAGE Publications Ltd,London\.External Links:[Link](https://sk.sagepub.com/dict/mono/ideology/toc),[Document](https://dx.doi.org/10.4135/9781446217856)Cited by:[§3\.3](https://arxiv.org/html/2608.00288#S3.SS3.p3.1)\.
- J\. Visser, B\. Konat, R\. Duthie, M\. Koszowy, K\. Budzynska, and C\. Reed \(2020\)Argumentation in the 2016 us presidential elections: annotated corpora of television debates and social media reaction\.Language Resources and Evaluation54,pp\. 123–154\.External Links:[Document](https://dx.doi.org/10.1007/s10579-019-09446-8),[Link](https://doi.org/10.1007/s10579-019-09446-8)Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Wachsmuth, N\. Naderi, Y\. Hou, Y\. Bilu, V\. Prabhakaran, T\. A\. Thijm, G\. Hirst, and B\. Stein \(2017\)Computational argumentation quality assessment in natural language\.InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers,M\. Lapata, P\. Blunsom, and A\. Koller \(Eds\.\),Valencia, Spain,pp\. 176–187\.External Links:[Link](https://aclanthology.org/E17-1017/)Cited by:[§2](https://arxiv.org/html/2608.00288#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu \(2025\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§4\.1](https://arxiv.org/html/2608.00288#S4.SS1.p1.1)\.
## Appendix AFull Annotation Guidelines
\(follow on next complete page\)
## Appendix BPrompt CandidatesSimilar Articles
Ideology Prediction of German Political Texts
This paper presents a transformer-based model that projects political orientation of German texts onto a continuous left-to-right spectrum, achieving high accuracy across multiple corpora including Bundestag plenary notes, Wahl-O-Mat, newspapers, and tweets.
Ideology Prediction of German Political Texts
The paper proposes a transformer-based model to predict political ideology of German political texts on a continuous left-to-right spectrum. The study compares 13 models and finds DeBERTa-large and Gemma2-2B perform best on different tasks.
A Corpus of Persuasion Techniques in Slavic Languages
A new annotated corpus of persuasion techniques in Bulgarian, Polish, and Russian, covering parliamentary debates and social media, with 25 fine-grained techniques and baseline models for detection and classification.
A Community-Based Approach for Stance Distribution and Argument Organization
Researchers from the University of British Columbia propose an unsupervised graph-based system for organizing arguments from online debates by constructing interaction graphs and applying community detection to reveal diverse viewpoint distributions. The approach requires no training data and aims to help users navigate complex argumentative landscapes and combat filter bubbles.
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
A systematic study on detecting Schwartz values in political text, comparing context lengths, model sizes, and retrieval-augmented generation methods. Results show that full-document context improves supervised models but not zero-shot LLMs, while retrieved moral knowledge consistently helps via early fusion.