Framing by Wording, Framing by Selection: A Large-Scale Two-Dimensional Audit of French News Headlines, 2022-2025
Summary
This paper introduces a two-dimensional framework for analyzing news headlines, separating salience framing from selection framing, and applies it to a large-scale audit of French headlines using LLM-based annotation, revealing disparities in media framing.
View Cached Full Text
Cached at: 09/25/26, 09:14 AM
# Framing by Wording, Framing by Selection: A Large-Scale Two-Dimensional Audit of French News Headlines, 2022–2025
Source: [https://arxiv.org/html/2609.28487](https://arxiv.org/html/2609.28487)
###### Abstract
News headlines frame public issues both by*what*they select and by*how*they word it, yet computational framing work typically collapses these operations into a single score\. We introduce a two\-dimensional framework that separates*salience framing*, measured through four wording devices \(loaded vocabulary, blame attribution, threat framing, rhetorical question\), from*selection framing*, measured through outlet\-level story\-form and high\-charge distributions\. We build a 10,000\-headline French supervision set using three LLM annotators with majority\-vote resolution and human arbitration, validate the labels against two annotator\-independent blind human studies, and apply the strongest classifier to 902,111 deduplicated headlines from 25 French outlets \(2022–2025\)\. Three main findings emerge\. First, salience and selection divergence are positively correlated yet leave nearly half of outlet\-level variance unexplained, populating interpretively distinct off\-diagonal cells in a four\-cell outlet typology\. Second, default classification thresholds systematically inflate corpus\-level salience estimates; a precision\-floor recalibration protocol corrects this distortion\. Third, group\-mention analysis reveals sharply unequal salience contexts: headlines mentioning Jews, the Far\-right, and Muslims carry the highest detected salience rates, which broad event\-context composition does not fully explain \(residuals are descriptive, not same\-event causal estimates; per\-group lexicon precision is reported alongside\)\. To our knowledge, this is the largest framing\-focused French*headline*audit to date; we release the supervision set, lexicons, and analysis code\.
## 1Introduction
Headlines are high\-reach editorial artifacts: a majority of shared news URLs are never clicked\([Gabielkov et al\. 2016](https://arxiv.org/html/2609.28487#bib.bib14)\), and experimental work shows that headlines can shape memory and inference even when the article is later read\([Ecker et al\. 2014](https://arxiv.org/html/2609.28487#bib.bib9)\)\. Their compactness makes editorial choice unusually visible\. A headline can frame by wording an event as invasion, betrayal, danger, scandal, or responsibility; it can also frame by repeatedly selecting particular event types for particular groups\. When a French outlet writes*“Les migrants envahissent”*\(“Migrants are invading”\) rather than*“Des migrants arrivent”*\(“Migrants are arriving”\), the word choice is not accidental; it encodes a threat frame, an intergroup opposition, and a causal attribution in three words\. When a far\-right aggregator fills four in five headlines with crime stories about identifiable social groups, that selection pattern is not random; it systematically concentrates group\-threat associations in editorial output\. These are two distinct editorial operations\.
Computational framing measures often blur them\. Sentiment scores and bias classifiers capture how language is used but cannot distinguish neutral wording on a concentrated crime agenda from charged wording on a broader one; topic\-distribution measures capture what is selected but not how it is packaged\. Auditing harmful or polarizing media ecosystems, as the EU Digital Services Act’s independent\-audit regime now requires at platform scale, means knowing whether an outlet’s effect is driven by selection, salience, or their interaction\. We study this problem in French news headlines, a setting with rich ideological variation but less computational framing infrastructure than English\. We ask:RQ1how salience framing varies across French outlets;RQ2whether salience and selection are empirically dissociable, taking as the null that they are one construct measured twice, which would show as near\-collinear divergences \(r≥0\.90r\\geq 0\.90\) with unpopulated off\-diagonal cells; andRQ3which social and political groups appear in the highest\-salience headline contexts\. Our contributions include the 10,000\-headline supervision set, a precision\-floor recalibration protocol, and a 902,111\-headline audit showing substantial overlap between salience and selection without collapsing them into one axis\.
## 2Framework and Related Work
[Entman 1993](https://arxiv.org/html/2609.28487#bib.bib11)’s definition says framing involves both*selection*and*salience*: selecting aspects of reality and making them more noticeable or meaningful\. Entman treats these as coupled within a single framing act; we separate them not as a claim about the act but as a measurement decision at the outlet level, where the two need not coincide and, as §[4\.1](https://arxiv.org/html/2609.28487#S4.SS1)shows, only partially do\. We operationalize this as two measurable streams\.Salience framingis the wording stream: four primary devices \(loaded vocabulary\([Stevenson 1944](https://arxiv.org/html/2609.28487#bib.bib39)\), blame attribution, threat framing, and rhetorical question\) plus a supplementary indicator of ingroup/outgroup construction \(us\-vs\-them\)\. These devices are grounded in problem definition, causal interpretation, moral evaluation, and treatment implication, and in framing theory, critical discourse analysis, and argumentation theory\([Entman 1993](https://arxiv.org/html/2609.28487#bib.bib11);[Gamson and Modigliani 1989](https://arxiv.org/html/2609.28487#bib.bib16);[van Dijk 1991](https://arxiv.org/html/2609.28487#bib.bib43);[Reisigl and Wodak 2001](https://arxiv.org/html/2609.28487#bib.bib31);[Walton 1996](https://arxiv.org/html/2609.28487#bib.bib44);[Tankard 2001](https://arxiv.org/html/2609.28487#bib.bib40)\)\. The four primary devices are selected because they are theoretically grounded, observable in headline\-length text, and stable enough for supervised measurement at corpus scale\. The second criterion is the binding one and excludes devices the same traditions treat as central: personalisation and metaphor, for instance, are recoverable from article context but rarely decidable from a headline alone\.Selection framingis the distributional stream: which story forms an outlet repeatedly chooses to cover, and how often those story forms are high\-charge events such as crime, conflict, scandal, and crisis\. This follows agenda\-setting and gatekeeping traditions, where influence operates through coverage allocation and accessibility\([McCombs and Shaw 1972](https://arxiv.org/html/2609.28487#bib.bib24);[Scheufele and Tewksbury 2007](https://arxiv.org/html/2609.28487#bib.bib35);[Shoemaker and Vos 2009](https://arxiv.org/html/2609.28487#bib.bib36)\)\. Our selection measure is conditional on coverage: it observes story\-form distributions*among headlines an outlet chose to publish*, closer to second\-level attribute agenda\-setting\([McCombs 2005](https://arxiv.org/html/2609.28487#bib.bib23)\)than to first\-level gatekeeping\. Outlets with identical story\-form distributions could differ substantially in which events from the world they select to cover at all: a dimension the published\-headline corpus cannot observe\.
Table[1](https://arxiv.org/html/2609.28487#S2.T1)summarizes the measurement schema\. The four primary devices are grounded in[Entman 1993](https://arxiv.org/html/2609.28487#bib.bib11)’s four frame functions: blame attribution \(causal interpretation\), threat framing \(problem definition and treatment implication\), loaded vocabulary and rhetorical question \(moral evaluation and problem definition\), with many\-to\-many correspondence per headline; us\-vs\-them marks ingroup/outgroup opposition\. Prior work either assigns frame labels from a single typology without decomposing wording from selection\([Card et al\. 2015](https://arxiv.org/html/2609.28487#bib.bib7);[Liu et al\. 2019](https://arxiv.org/html/2609.28487#bib.bib22);[Mendelsohn et al\. 2021](https://arxiv.org/html/2609.28487#bib.bib25)\), captures a single bias dimension \(informational and lexical bias\([Fan et al\. 2019](https://arxiv.org/html/2609.28487#bib.bib13);[Spinde et al\. 2021](https://arxiv.org/html/2609.28487#bib.bib38)\)or persuasion techniques\([Piskorski et al\. 2023](https://arxiv.org/html/2609.28487#bib.bib29)\)\), or separates issue filtering from slant rather than wording from selection\([Budak et al\. 2016](https://arxiv.org/html/2609.28487#bib.bib5)\), and so cannot serve as quantitative baselines for a two\-dimensional decomposition\([Otmakhova et al\. 2024](https://arxiv.org/html/2609.28487#bib.bib27)\)\. English\-language auditing resources including NELA\-GT\([Gruppi et al\. 2022](https://arxiv.org/html/2609.28487#bib.bib18)\)and BERT\-based political leaning classifiers\([Baly et al\. 2020](https://arxiv.org/html/2609.28487#bib.bib2)\)address related questions at scale but do not separate wording from selection mechanisms\.
French\-language media studies motivate the setting\.[Benson and Wood 2015](https://arxiv.org/html/2609.28487#bib.bib4)documents that immigration news in France is dominated by governmental and political sources, with limited non\-governmental voices\(cf\.[Benson 2013](https://arxiv.org/html/2609.28487#bib.bib3), for cross\-national comparison\);[Dalibert 2015](https://arxiv.org/html/2609.28487#bib.bib8)shows how minority\-movement access to the French public sphere is regulated by “francité,” constraining visibility and legitimacy\. Work on migrant representation records oscillation between victim, threat, and voiceless framings, echoing our threat and us\-vs\-them findings\. This literature motivates the 12\-group inventory analysed in §[4\.2](https://arxiv.org/html/2609.28487#S4.SS2), which includes migrants, Muslims, Jews, and the far right\. Related francophone work detects opinion/information genre at scale in Québécois and Belgian media\([Escouflaire et al\. 2024](https://arxiv.org/html/2609.28487#bib.bib12)\)and links ownership structures to speaking\-time\-based political slant in French broadcasting\([Cagé et al\. 2022](https://arxiv.org/html/2609.28487#bib.bib6)\)\. More recent French\-media framing and discourse studies remain either small and issue\-specific, such as the 120\-article immigration corpus of[Song et al\. 2024](https://arxiv.org/html/2609.28487#bib.bib37), or topic\-bound at moderate scale, such as the 13,795\-article AI press corpus of[Tsimpoukis 2025](https://arxiv.org/html/2609.28487#bib.bib42)\. Larger French news corpora, such as[Jehle and Le Gallo 2025](https://arxiv.org/html/2609.28487#bib.bib20)’s 400,000\-article EU\-sentiment study, analyze tone or agenda dynamics rather than explicit framing measurement\. Our unit is narrower, focusing on headlines only, but covers more outlets and explicitly separates selection from salience; to our knowledge, that makes the present corpus unusually large for framing\-focused work in the French setting\.
Table 1:Measurement schema\. Salience devices are wording\-level and non\-mutually exclusive; story form and high\-charge are selection\-level\.The boundary matters most for attribute agenda\-setting\([McCombs 2005](https://arxiv.org/html/2609.28487#bib.bib23)\): repeated CRIME coverage operates through salience transfer and accessibility, making considerations available in memory\([McCombs and Shaw 1972](https://arxiv.org/html/2609.28487#bib.bib24);[Scheufele 1999](https://arxiv.org/html/2609.28487#bib.bib34)\), while charged wording operates through applicability, shaping which interpretive schemas audiences invoke\([Scheufele 1999](https://arxiv.org/html/2609.28487#bib.bib34);[Scheufele and Tewksbury 2007](https://arxiv.org/html/2609.28487#bib.bib35)\)\. A single “framing intensity” score cannot preserve that distinction\.
## 3Data and Method
#### Corpora\.
We use two datasets with different roles\. The supervised development set contains 10,000 France\-based news headlines from 25 outlets, designed as a structurally comprehensive cross\-section of the French media ecosystem across the political spectrum\([Newman et al\. 2023](https://arxiv.org/html/2609.28487#bib.bib26);[ACPM 2026](https://arxiv.org/html/2609.28487#bib.bib1)\)\. It includes three of France’s leading broadcast and continuous\-news brands \(Franceinfo,BFMTV,TF1 INFO\) and seven major national dailies \(Le Monde,Le Figaro,Le Parisien,Les Echos,Libération,La Croix,L’Humanité\)\. Selection criteria were reach or circulation rank; format diversity; ideological breadth, explicitly anchoring the range withLa Croix\(Catholic daily\),L’Humanité\(communist historical title\),Fdesouche\(far\-right aggregator\),Valeurs actuelles\(conservative\-nationalist weekly\), andCauseur\(conservative intellectual magazine\); and restriction to outlets covering three shared sections, with headlines jointly balanced across outlets and sections\.
Each headline is labeled by three schema\-constrained LLM annotators from different model families, following evidence that LLMs can match or outperform crowdworkers on structured annotation tasks\([Gilardi et al\. 2023](https://arxiv.org/html/2609.28487#bib.bib17);[Törnberg 2025](https://arxiv.org/html/2609.28487#bib.bib41);[Ziems et al\. 2024](https://arxiv.org/html/2609.28487#bib.bib45)\)—a task\-specific comparison that remains contested, and one our own validation qualifies \(Table[4](https://arxiv.org/html/2609.28487#S3.T4)\)\. The annotators,openai/gpt\-oss\-120b\(released August 2025\),google/gemma\-4\-31B\(released April 2026\), andmeta\-llama/Llama\-3\.3\-70B\-Instruct, produced full\-field labels under a shared schema\-constrained instruction set; Appendix Table[24](https://arxiv.org/html/2609.28487#A0.T24)summarizes the released schema; the full annotation prompt is included in the release\. Three candidate second annotators were run against the primary annotator at full scale and two were retained; pairwise agreement statistics for all three are released\. Final binary labels are resolved by 2/3 majority vote\. The 642 three\-way story\-form conflicts \(6\.4%\) are assigned to a single human arbitrator; post\-hoc blind inter\-annotator reliability on these cases is reported in Appendix Table[25](https://arxiv.org/html/2609.28487#A0.T25)\. Binary conflicts are not arbitrated: the panel is three models with overlapping pretraining rather than three independent coders, so majority vote can propagate shared error \(Appendix Table[9](https://arxiv.org/html/2609.28487#A0.T9)\), and the blind human studies rather than arbitration are the control for the binary heads; the weakest head, us\-vs\-them, is reported only as a supplementary indicator\. The final split is 6,999 train, 1,500 validation, and 1,501 test headlines\.
The corpus\-level analysis set approximates real editorial output\. We drew a proportionally stratified sample targeting one million headlines from the 25\-outlet panel across outlet, section, and year for 2022–2025, then removed exact duplicate URLs and within\-outlet duplicate headlines, yielding 902,111 unique headlines\. Eligibility was restricted to the Politics, Economy, and Society sections used throughout corpus construction\. Per\-outlet counts range from 646 \(Blast\) to 108,307 \(Le Figaro\); full counts appear in Appendix Table[11](https://arxiv.org/html/2609.28487#A0.T11)\. Identical headlines across outlets are retained \(syndication is part of the observable editorial field\); classification is restricted to headline text, the standalone framing unit\. The unit also strips context that disambiguates attribution: quotation, irony, and blame voiced by a cited source rather than by the outlet are often unrecoverable, which bounds the blame\-attribution and threat\-framing devices in particular\.
#### Labels\.
Salience is multi\-label: loaded vocabulary, blame attribution, threat framing, rhetorical question, and us\-vs\-them construction\. Selection is modeled through a binary high\-charge label\([Galtung and Ruge 1965](https://arxiv.org/html/2609.28487#bib.bib15);[Harcup and O’Neill 2017](https://arxiv.org/html/2609.28487#bib.bib19)\)and a 10\-class story\-form taxonomy \(Table[1](https://arxiv.org/html/2609.28487#S2.T1)\), developed iteratively during pilot annotation on a 500\-headline exploratory sample and guided by French press section conventions and prior news\-values classifications\([Harcup and O’Neill 2017](https://arxiv.org/html/2609.28487#bib.bib19)\); unlike topic\-oriented systems such as IPTC Media Topics, story form classifies the journalistic action a headline reports rather than just the subject domain, so one immigration headline may be POLICY and another CRIME \(Appendix Table[10](https://arxiv.org/html/2609.28487#A0.T10)\)\. Target\-group analysis is handled separately through a hand\-curated lexicon of 219 French surface\-form terms spanning 12 groups \(Appendix Table[7](https://arxiv.org/html/2609.28487#A0.T7)reports the fullv6label distribution\)\.
#### Models\.
We compare three supervised measurement stacks on the finalv6split: a TF\-IDF \+ Logistic Regression baseline,camembert\-base, andxlm\-roberta\-large\(Table[3](https://arxiv.org/html/2609.28487#S3.T3)\)\. For transformer models, a multi\-task encoder predicts the five salience devices and charge band through six independent binary heads on a shared headline representation, while a separate encoder predicts story form through a single 10\-class head\. We usexlm\-roberta\-largefor corpus inference because it delivers the strongest overall held\-out performance: mean salience\+charge F1 is \.742 \(XLM\-R\), \.726 \(CamemBERT\), \.476 \(TF\-IDF\); XLM\-R also leads on story\-form macro\-F1 \(\.731\)\. CamemBERT scores higher on rhetorical questions alone \(\.900 vs\. \.841\), likely reflecting French\-specific interrogative morphosyntax; for corpus inference we therefore use XLM\-R for all heads except rhetorical question, where CamemBERT is substituted\. The ensemble was fixed on validation\-set metrics before corpus inference; no post\-hoc adjustments were made\. Exhaustive encoder benchmarking is not the contribution; the substantive cross\-outlet findings are bounded against encoder choice by the TF\-IDF replication in §[4\.3](https://arxiv.org/html/2609.28487#S4.SS3)\(r=0\.667r=0\.667vs\.0\.7360\.736\)\.
Both models use a 48\-token maximum sequence length \(<<2% of headlines exceed this; details in Appendix Table[23](https://arxiv.org/html/2609.28487#A0.T23)\)\. Three\-seed evaluation confirms robust model selection \(Appendix Table[23](https://arxiv.org/html/2609.28487#A0.T23)\); results in Table[3](https://arxiv.org/html/2609.28487#S3.T3)are from seed 42\. Thresholds are selected on validation only via a precision\-floor\-constrained recalibration policy: for each binary head, we search a 0\.10–0\.99 grid and retain the threshold maximizing validation F1 subject to pre\-specified per\-head precision floors, assigned from annotator agreement tiers before validation metrics are observed: rhetorical question \(κ=\.770\\kappa\{=\}\.770, highest\) receives a \.85 floor; the mid\-agreement heads \(loaded vocabulary, blame attribution, charge band\) receive \.70; and the lowest\-agreement heads \(threat framing, us\-vs\-them\) receive \.60 \(per\-head floors and thresholds in Table[2](https://arxiv.org/html/2609.28487#S3.T2)\)\.
Table 2:Finalv6thresholds \(validation set only\) for the mixed\-encoder ensemble; see Table[3](https://arxiv.org/html/2609.28487#S3.T3)and Models\.‡Us\-vs\-them is a supplementary indicator \(below the0\.600\.60precision floor; Appendix Table[22](https://arxiv.org/html/2609.28487#A0.T22)\), not a primary prevalence estimate\.Table 3:Held\-out test\-set F1 by model \(finalv6split, seed 42, production thresholds\); per\-model means and three\-seed robustness are in Models and Appendix Table[23](https://arxiv.org/html/2609.28487#A0.T23)\.†Production model; CamemBERT substituted for rhetorical question only \(see Models\)\.‡Us\-vs\-them is a supplementary indicator \(below the0\.600\.60precision floor; Appendix Table[22](https://arxiv.org/html/2609.28487#A0.T22)\), not a primary prevalence estimate\.Agreement patterns \(Appendix Table[9](https://arxiv.org/html/2609.28487#A0.T9)\) are consistent with thev6merge design: binary fields show high unanimous agreement while story form is harder, matching the need for arbitration; us\-vs\-them is weakest \(κ=\.408\\kappa=\.408\)\.
#### Independent blind validation\.
Two annotators with no affiliation to this work independently re\-annotated a stratified 499\-headline held\-out sample \(primary validation\) using a bilingual annotation guide derived from theoretical construct definitions, with no access to LLM labels or expected label distributions \(Table[4](https://arxiv.org/html/2609.28487#S3.T4)\)\. The sample is temporally representative of the held\-out test split \(39\.5% post\-October 7, 2023, vs\. 41\.3% of date\-resolved test headlines\), directly covering the period most exposed to possible LLM pretraining overlap; within this subset, any\-salience human–LLM agreement declines only modestly \(κ=\.424\\kappa=\.424vs\.\.463\.463full\-sample\)\. Majority\-vote human–LLMκ\\kapparanges\.608\.608–\.869\.869across the four primary devices; recall is consistently high \(≥\.65\\geq\.65\) while precision is lower \(\.37\.37–\.82\.82\), confirming that the merged labels are systematically liberal \(see Limitations\)\. Human–humanκ=\.680\\kappa=\.680on the any\-salience composite confirms that annotators agree on overall salience judgements even where device attributions diverge\. A corroborating independent study on a separateN=350N\{=\}350sample yields consistent results \(Appendix Table[17](https://arxiv.org/html/2609.28487#A0.T17)\)\. Appendix Table[16](https://arxiv.org/html/2609.28487#A0.T16)provides a qualitative error analysis illustrating principal failure modes\.
Table 4:Independent blind re\-annotation of the 499\-headline validation sample \(N=499N=499\) by two unaffiliated annotators\. Maj–LLM = majority\-vote human labels vs\. LLM consensus; conflict rows excluded, effectiveNNper head 383–488\.κ\\kappainterpretation bands follow[Landis and Koch 1977](https://arxiv.org/html/2609.28487#bib.bib21)\.†Us\-vs\-them is a supplementary indicator \(below the0\.600\.60precision floor; see Appendix Table[22](https://arxiv.org/html/2609.28487#A0.T22)\), not a primary prevalence estimate\.‡Derived composite\. See Appendix Table[17](https://arxiv.org/html/2609.28487#A0.T17)for the corroboratingN=350N=350study\.Group analysis uses a lexicon of explicit anchor terms mapping French surface forms to 12 canonical groups\. The lexicon contains 219 terms and detects 134,046 headlines \(14\.9% of the corpus\), covering explicit mentions only; indirect or paraphrastic references are missed by design\. Groups were selected by theoretical relevance \(targeting identities whose framing asymmetries are documented in French and European media research\([Benson and Wood 2015](https://arxiv.org/html/2609.28487#bib.bib4);[Dalibert 2015](https://arxiv.org/html/2609.28487#bib.bib8);[van Dijk 1991](https://arxiv.org/html/2609.28487#bib.bib43)\)\) and by a minimum corpus\-frequency criterion of at least 1,000 detectable headline mentions in a pilot pass\. A blinded two\-human audit of a 200\-headline sample confirms high agreement:κ=\.94\\kappa=\.94–\.98\.98on the representative tier \(n=140n\{=\}140\) andκ=\.83\\kappa=\.83–\.85\.85on a harder stress\-test tier \(n=60n\{=\}60\)\. The lexicon shows precision≥\.83\\geq\.83for 9 of 12 groups; LFI \(\.69\), Jews \(\.70\), and Unions \(\.82, marginal\) are lower\-precision and should be treated as indicative \(Appendix Table[8](https://arxiv.org/html/2609.28487#A0.T8)\)\.
#### Analysis plan\.
We summarize salience and selection distinctiveness with Jensen–Shannon divergence from the corpus baseline, following the use of Jensen–Shannon measures for media\-agenda comparison\([Pinto et al\. 2019](https://arxiv.org/html/2609.28487#bib.bib28)\), and test dissociability with Pearson and Spearman correlations against the conventionalr<0\.90r<0\.90criterion \(RQ2\)\. We build the four\-cell typology from median splits on salience and selection divergence as a cartographic convenience rather than a claim of categorical differences; continuous outlet positions are visible in Figure[1](https://arxiv.org/html/2609.28487#S4.F1)\. Robustness checks compare focal outlets against all other outlets within the same story form usingχ2\\chi^\{2\}andϕ\\phias a coarse descriptive stress test; story form is partly editorially mediated and therefore is not an exogenous same\-event control\. Salience distinctiveness measures divergence in the distribution of rhetorical devices relative to the corpus baseline rather than raw device frequency, so an outlet may exhibit frequent charged language while remaining low in Sal\.JS if its device mixture mirrors the broader media distribution\. Because Sal\.JS is computed over a 5\-bin device distribution and Sel\.JS over a 10\-bin story\-form distribution, the two divergences are not directly comparable in magnitude; the typology uses within\-axis rankings rather than cross\-axis absolute values\.
## 4Results
Across the full corpus, 34\.6% of headlines contain at least one salience device \(loaded vocabulary most frequent, us\-vs\-them rarest; Appendix Table[7](https://arxiv.org/html/2609.28487#A0.T7)\)\. High\-charge story forms account for 31\.3%\. These rates are lower than default\-threshold estimates would produce, reflecting the precision\-floor recalibration; the 34\.6% figure is best interpreted relative to the French corpus baseline rather than as a universal prevalence benchmark\.
Figure 1:Selection vs\. salience divergence across French news outlets\.Note that*Ouest\-France*sits high in \(b\) because it is far from the baseline overall while its headlines deploy framing devices unusually*rarely*, not because they deploy them aggressively\. \(a\) Each outlet’s distance from the corpus baseline decomposed into selection divergence \(DselJSD\_\{\\mathrm\{sel\}\}^\{\\mathrm\{JS\}\}, what stories are emphasised\) and salience divergence \(DsalJSD\_\{\\mathrm\{sal\}\}^\{\\mathrm\{JS\}\}, which framing devices are deployed in the headline\)\. Both quantities are Jensen–Shannon divergences against the cross\-outlet mixture distribution\. Marker colour encodes the four\-class typology defined in §[4\.1](https://arxiv.org/html/2609.28487#S4.SS1)\. \(b\) Same outlets re\-expressed as a salience share \(DsalJS/\(DsalJS\+DselJS\)D\_\{\\mathrm\{sal\}\}^\{\\mathrm\{JS\}\}/\(D\_\{\\mathrm\{sal\}\}^\{\\mathrm\{JS\}\}\+D\_\{\\mathrm\{sel\}\}^\{\\mathrm\{JS\}\}\)\) against summed divergence\. The dashed line at0\.50\.5separates outlets whose share falls on the selection side from those on the salience side\. Vertical position in \(b\) is summed distance from the baseline, not salience intensity; outlets close to the bottom edge are close to the corpus baseline on both dimensions\. Both divergences share the same bounded scale, so the share is well defined, but their expected magnitudes differ with support size \(5 vs\. 10 bins\); both coordinates in \(b\) are therefore read ordinally — as which axis dominates within an outlet, and as rank distance from the baseline — not as an equivalence of magnitudes\.### 4\.1Outlet Profiles and Dissociability
Figure[1](https://arxiv.org/html/2609.28487#S4.F1)maps all 25 outlets in the two\-dimensional divergence space, Table[5](https://arxiv.org/html/2609.28487#S4.T5)summarizes the four\-cell membership, and Appendix Table[11](https://arxiv.org/html/2609.28487#A0.T11)reports the full outlet\-level values\. Salience divergence from the corpus baseline varies sharply: Slate\.fr is the most salience\-distinctive outlet, Fdesouche carries the clearest threat/blame profile, Causeur is distinctive through loaded and interrogative wording, and Ouest\-France is double\-distinctive through unusually low device and high\-charge rates \(7\.2% any\-salience, 4\.6% high\-charge, both lowest in the panel\)\. The Double\-distinctive cell spans 7\.2%–72\.6% any\-salience: cell membership reflects position relative to thresholds, not editorial similarity within the cell\. Selection divergence produces a different ranking, with Fdesouche, Causeur, and Blast among the clearest agenda outliers \(Causeur’s ELITE\-heavy selection profile should be read with caution because ELITE\-category conflict cases show lower arbitration agreement; Appendix Table[25](https://arxiv.org/html/2609.28487#A0.T25)\) and high\-charge coverage concentrated especially in Fdesouche, Blast, and Valeurs actuelles\.
Device composition reveals three outlet signatures:threat\-blame, where threat and blame co\-elevate \(Fdesouche most clearly, with Valeurs actuelles at lower intensity\);interrogative\-evaluative, where rhetorical questions and loaded vocabulary dominate without the broad threat/blame profile \(Slate\.fr, Causeur, L’Express; rhetorical\-question rates index headline form, see Limitations\); andinstitutional blame, where accountability language rises without a strong threat register \(Mediapart, L’Humanité, Le Parisien; blame rates are similarly directional\)\. \(Loaded vocabulary: est\. precision0\.6430\.643post\-prior\-shift,−5\.7\-5\.7pp below the0\.700\.70floor; Appendix Table[28](https://arxiv.org/html/2609.28487#A0.T28)\.\) Similar any\-salience rates can therefore reflect distinct interpretive mechanisms\.
Rate estimates widen materially only for the smallest outlets; Blast \(646 headlines\) should be read as an illustrative case profile \(Appendix Table[12](https://arxiv.org/html/2609.28487#A0.T12)\)\.
Table 5:Four\-cell typology\. Sal\.JS/Sel\.JS = salience/selection Jensen–Shannon divergence from the corpus baseline \(§[3](https://arxiv.org/html/2609.28487#S3)\); cells assigned by median splits; Sal\.% = mean any\-salience prevalence \(selection\-dominant outlets show the highest Sal\.% yet the lowest Sal\.JS; see Results\)\. Outlet assignments in Fig\.[1](https://arxiv.org/html/2609.28487#S4.F1)and Appendix Table[11](https://arxiv.org/html/2609.28487#A0.T11)\.Dissociability rests directly on the populated off\-diagonal cells of the bivariate distribution \(Figure[1](https://arxiv.org/html/2609.28487#S4.F1); four\-cell summary in Table[5](https://arxiv.org/html/2609.28487#S4.T5)\): the cells are perturbation\-stable \(only seven near\-boundary outlets shift under±\\pm20% median changes, Appendix Table[26](https://arxiv.org/html/2609.28487#A0.T26)\), and concrete cases anchor the corners \(JDD high on selection yet low on salience, TF1 INFO the reverse\), holding independently of any correlation threshold\. The correlation only quantifies the overlap: salience distinctiveness explains 54\.2% of outlet\-level selection variance \(r=0\.736r\{=\}0\.736111Split\-half reliability:rxx=0\.979r\_\{xx\}\{=\}0\.979,ryy=0\.988r\_\{yy\}\{=\}0\.988; Spearman\-corrected latentr=0\.748r\{=\}0\.748\(see Limitations for the leave\-one\-out baseline correction;r=0\.721r\{=\}0\.721\)\., LOO\-correctedr=0\.721r\{=\}0\.721; 95% bootstrap CI \[\.503, \.936\],N=25N\{=\}25; reliably non\-zero, permutationp=\.0001p\{=\}\.0001\), leaving 45\.8% unexplained\. Both estimates fall within the conventionalr<0\.90r<0\.90criterion; withN=25N\{=\}25the interval is wide and its upper bound \(\.936\.936\) exceeds it, so the claim does not rest on the threshold\. This residual maps onto interpretable off\-diagonal profiles: selection\-dominant outlets exhibit thehighestraw salience prevalence \(any\-salience=0\.510=0\.510\) despite thelowestrhetorical divergence \(Sal\.JS=0\.008=0\.008\), while double\-distinctive outlets combine high salience \(any\-salience=0\.484=0\.484\) with substantially greater divergence \(Sal\.JS=0\.027=0\.027\)\. Frequent charged language therefore does not by itself imply rhetorical distinctiveness\.
To test whether this association is ecological aggregation, we estimate a longitudinal outlet×\\timesmonth panel \(1,184 observations; Appendix Table[29](https://arxiv.org/html/2609.28487#A0.T29)\)\. Under two\-way outlet\+month fixed effects with cluster\-robust inference \(outlet clusters,G=25G\{=\}25\), the within\-outlet coupling remains positive \(β=0\.758\\beta=0\.758, 95% CI \[−\-\.11, 1\.63\];p=\.084p=\.084\); we treat this as confirmatory rather than a stand\-alone test\. Excluding Fdesouche lowers the cross\-sectional correlation tor=0\.688r=0\.688\(ρ=0\.711\\rho=0\.711\), and the unexplained variance \(1−r2=0\.531\{\-\}r^\{2\}=0\.53\) remains sufficient to populate distinct off\-diagonal cells\. Because three high\-volume low\-divergence outlets \(Le Figaro, Le Parisien, Franceinfo\) contribute≈\{\\approx\}32% of the corpus baseline, their low divergence is partly self\-referential; a leave\-one\-out correction yieldsr=0\.721r\{=\}0\.721,ρ=0\.727\\rho\{=\}0\.727, preserving the dissociability criterion \(full LOO analysis in Limitations\)\. A tone\-only audit would therefore miss selection\-dominant outlets, while a topic\-only audit would miss salience\-dominant ones\.
### 4\.2Group\-Mention Salience Contexts
[van Dijk 1991](https://arxiv.org/html/2609.28487#bib.bib43)establishes that group framing operates through both event\-context selection and adversarial wording devices; our two\-dimensional framework keeps those pathways analytically separate\. Explicit group mentions occur in highly asymmetric model\-estimated salience contexts \(Table[6](https://arxiv.org/html/2609.28487#S4.T6)\)\. Group salience rates reflect headline rhetorical intensity, not editorial sentiment or targeting; interpretive caveats in §Ethics apply throughout\. Headlines mentioning Jews have the highest observed any\-salience rate, driven predominantly by antisemitism reporting and Israel/Gaza security coverage \(lexicon precision for this group is \.70, among the lowest of the twelve, though recall is \.97, so the detector is liberal rather than blind\): 78\.7% carry at least one detected salience device and 79\.8% are high\-charge\. A year×\\timesstory\-form reweighting estimates that 49\.6% salience would be expected from their coarse event\-context mix alone, leaving a descriptive \+29\.1 pp residual; this residual is evidence that broad story\-form composition does not fully explain the pattern, not evidence of same\-event causal editorial targeting\. Far\-right mentions show a similar raw/residual profile \(77\.7% raw; \+27\.7 pp residual\), reflecting electoral, parliamentary, and conflict coverage of a political movement rather than targeting of its members\. Muslims \(66\.4%; \+19\.8 pp residual\) and Migrants \(58\.4%; \+17\.7 pp\) concentrate in policy and security event contexts; Police \(57\.6%; \+11\.1 pp\) coverage is dominated by blame attribution in accountability\-focused reporting; and Unions \(53\.8%; \+3\.6 pp\) reflect the conflict\-heavy 2023–2024 pension\-reform mobilization cycle\. Seniors remain below baseline, and Workers fall below the salience rate expected from their year×\\timesstory\-form mix\. The 2024–2025 salience rise for Jews \(77\.2% to 84\.1%; Appendix Table[21](https://arxiv.org/html/2609.28487#A0.T21)\) is consistent with post\-October\-2023 crisis coverage; high salience reflects event\-selection and wording contexts, not a uniform targeting pattern \(caveats in §Ethics\)\.
Table 6:Group\-mention salience contexts \(top 8 of 12 groups\)\. Sal% = observed any\-salience; Exp% = expected rate under year×\\timesstory\-form reweighting \(focal group excluded from cell baselines to avoid circular self\-inclusion\); Res\. = observed−\-expected \(pp\)\. Residuals are descriptive, not same\-event causal estimates\. Full 12\-group table in Appendix Table[13](https://arxiv.org/html/2609.28487#A0.T13)\.All differences from baseline are significant under BH\-FDR\-corrected two\-proportionzz\-tests \(12 hypotheses;p<0\.001p<0\.001; intervals do not propagate classifier uncertainty; see Limitations for causal\-identification disclosures\)\.
Device profiles differ by group \(Appendix Table[14](https://arxiv.org/html/2609.28487#A0.T14)\): Jews show the strongest multi\-device signal \(blame, loaded vocabulary, threat\); Police are blame\-dominated; and the Far\-right shows broad elevation across all three adversarial devices\. Selection adds a second layer: Police and Migrants concentrate in CRIME; Workers and Unions concentrate in CONFLICT \(56\.9% and 56\.8%, reflecting the 2023–2024 pension\-reform cycle\); and RN concentrates in ELECTIONS and PARLIAMENT\.
Outlet\-by\-group patterns \(Appendix Table[14](https://arxiv.org/html/2609.28487#A0.T14)\) add an important caution: Fdesouche’s elevation spans multiple groups, indicating a broad high\-intensity editorial register rather than targeted group framing, while Mediapart and L’Humanité peak on Police and the Far\-right through an accountability lens\. Removing Fdesouche confirms most group effects are robust \(median exclusion effect−\-0\.5 pp; Appendix Table[20](https://arxiv.org/html/2609.28487#A0.T20)\)\. Results for the supplementary us\-vs\-them indicator are reported in Appendix Table[22](https://arxiv.org/html/2609.28487#A0.T22)\.
### 4\.3Robustness
A dual\-ablation study confirms that predictions are not driven by entity memorization: named\-entity masking \(spaCy\) yields negligible confidence drops \(\|Δp\|≤0\.085\|\\Delta p\|\\leq 0\.085\) and low label\-flip rates \(5\.3%–11\.4%\), while masking the single most salient semantic token \(Integrated Gradients\) produces substantially larger effects \(e\.g\.,\|Δp\|=0\.266\|\\Delta p\|=0\.266, 31\.6% flip rate for rhetorical questions\)\. Within\-story\-form comparisons show that salience differences persist inside broad story\-form bins \(Appendix Table[15](https://arxiv.org/html/2609.28487#A0.T15)\): Fdesouche and Valeurs actuelles remain elevated within CRIME, and Slate\.fr retains its interrogative\-evaluative profile within SOCIAL; these are descriptive stress tests rather than causal controls because story\-form assignment is itself partly shaped by editorial judgment\. A within\-outlet headline bootstrap confirms near\-perfect rank stability \(ρ=\.996/\.998\\rho\{=\}\.996/\.998for salience/selection JS\), and selection\-divergence rankings are robust to taxonomy granularity \(Appendix Table[27](https://arxiv.org/html/2609.28487#A0.T27)\)\. The top group pattern is temporally persistent: Jews, the Far\-right, and Muslims remain the three highest\-salience groups in every year from 2022 to 2025, with annual lifts in the×\\times1\.78–×\\times2\.35 range \(Appendix Table[21](https://arxiv.org/html/2609.28487#A0.T21)\)\.
An encoder\-independence check recomputes Sal\.JS/Sel\.JS with a TF\-IDF story\-form model \(macro\-F1=0\.543\):rrdrops from 0\.736 to 0\.667 \(still<0\.90<0\.90\);ρ\\rhois stable \(0\.741 vs\. 0\.747\), bounding the shared\-encoder confound atΔr≈0\.07\\Delta r\\approx 0\.07\. Cook’sDDanalysis shows that excluding three high\-leverage outlets raisesrrto 0\.865, still leaving∼\\sim25% unexplained variance; RQ\-head exclusion flips four near\-boundary outlets \(Appendix Table[27](https://arxiv.org/html/2609.28487#A0.T27)\)\. Broadcast\-cluster format artifacts are null \(R2=0\.0003R^\{2\}\{=\}0\.0003\), and XLM\-R confidence gaps across ideological clusters are negligible \(≤0\.015\{\\leq\}0\.015\)\.
#### Temporal trend\.
Any\-salience rises monotonically, 31\.7% \(2022\) to 38\.1% \(2025\)\. A rate–composition decomposition assigns\+4\.86\+4\.86pp of the\+6\.34\+6\.34pp change to within\-story\-form wording and\+1\.20\+1\.20pp to composition\. It holds in nine of ten story forms, under joint outlet×\\timessection×\\timesstory\-form standardisation \(31\.8%→\\to37\.4%\), and with the October–December 2023 and June–July 2024 windows excluded; 23 of 25 outlets rise\. One classifier scores all years, so drift is excluded by construction; standardisation controls event category, not intensity\.
## 5Discussion
The two\-dimensional framework resolves a structural ambiguity in single\-axis audits: outlets similarly scored may differ sharply in mechanism, as the outlet profiles illustrate\. TF1 INFO combines moderate raw salience \(38\.3%\) with elevated rhetorical divergence \(Sal\.JS=0\.029\) but near\-baseline selection divergence \(Sel\.JS=0\.011\), consistent with salience\-led broadcast packaging\. JDD shows the inverse pattern: substantial raw salience \(41\.9%\) but almost no rhetorical divergence \(Sal\.JS=0\.002\) alongside clearer selection divergence \(Sel\.JS=0\.037\), indicating agenda concentration without comparable wording distinctiveness\. Fdesouche’s dual\-distinctiveness \(72\.6% any\-salience; 82\.1% high\-charge; Sal\.JS=0\.038, Sel\.JS=0\.147\) therefore reflects not mere bias but a structural editorial register where accessibility and applicability amplify together\([Scheufele and Tewksbury 2007](https://arxiv.org/html/2609.28487#bib.bib35);[McCombs and Shaw 1972](https://arxiv.org/html/2609.28487#bib.bib24);[Scheufele 1999](https://arxiv.org/html/2609.28487#bib.bib34);[van Dijk 1991](https://arxiv.org/html/2609.28487#bib.bib43)\)\. The group\-mention findings corroborate prior unequal\-salience research\([van Dijk 1991](https://arxiv.org/html/2609.28487#bib.bib43)\); separating wording\-device load from story\-form concentration adds resolution unavailable to single\-axis designs\. Whether the pattern reflects stable editorial stance, event\-context concentration, or their interaction requires richer annotation; the present corpus enables such follow\-on work\([Benson and Wood 2015](https://arxiv.org/html/2609.28487#bib.bib4);[Dalibert 2015](https://arxiv.org/html/2609.28487#bib.bib8)\)\. The precision\-floor recalibration protocol, with its fully released threshold log, extends to any longitudinal multi\-label framing audit; whether default\-threshold inflation would accumulate to distort findings remains an empirical question\.
## Limitations
#### Scope of the audit\.
The audit measures editorial output patterns; whether those patterns produce accessibility or applicability effects in audiences requires experimental work\([Price and Tewksbury 1997](https://arxiv.org/html/2609.28487#bib.bib30);[Scheufele 2004](https://arxiv.org/html/2609.28487#bib.bib33)\)\. The study measures headlines only and does not capture article\-body framing or cross\-platform dynamics\. The selection framing measure is further constrained to published\-headline story\-form concentration and cannot detect first\-level event\-selection gatekeeping; outlets indistinguishable on Sel\.JS may differ in which real\-world events they choose to cover\. The outlet typology is two\-dimensional by construction; finer\-grained axes could further separate mobilisation\-register from adversarial\-register outlets within the double\-distinctive cell\.
#### Device performance and ensemble characterisation\.
The rhetorical\-question head indexes headline form, not rhetorical intent: the coding guide marks any interrogative headline, the annotator prompt excludes purely informational ones, and both fire on nearly all question\-marked headlines \(of 56 in the 499\-item sample, coders mark 96\.4% and 94\.6%, annotators 85\.7%\)\. Itsκ=\.869\\kappa=\.869therefore reflects a near\-syntactic judgement, bounding the interrogative\-evaluative signature in §[4\.1](https://arxiv.org/html/2609.28487#S4.SS1); its loaded\-vocabulary component is unaffected\. Across the three remaining primary heads, blame attribution achieves the strongest independent validation \(Maj–LLMκ=\.718\\kappa=\.718; Table[4](https://arxiv.org/html/2609.28487#S3.T4)\), while loaded vocabulary shows the lowest human–human agreement \(κHH=\.542\\kappa\_\{\\text\{HH\}\}=\.542\), indicating genuine task ambiguity rather than model failure\. Threat framing and us\-vs\-them reach moderate\-to\-substantial Maj–LLM agreement \(κ=\.668\\kappa=\.668and\.596\.596; Table[4](https://arxiv.org/html/2609.28487#S3.T4)\), with a corroborating independent study on a separate 350\-headline sample yielding consistent results \(Appendix Table[17](https://arxiv.org/html/2609.28487#A0.T17)\)\. Isolated French headlines remain difficult for implicit\-causality and rhetorical\-register inference: even where annotators agree on device presence, thev6ensemble operates as a recall\-oriented liberal annotator \(recall≥\.65\\geq\.65across all heads\), so corpus\-level rates should be interpreted as upper\-bound prevalence estimates\. Cross\-cluster confidence gaps are negligible \(≤0\.015\{\\leq\}0\.015; §[4\.3](https://arxiv.org/html/2609.28487#S4.SS3)\), suggesting that the recall\-biased training signal does not produce outlet\-differential annotation error sufficient to distort distributional rankings\. A French\-specific alternative \(mistral\-small3\.2\), one of the candidate annotators not retained, underperformed the ensemble on all five devices \(Δκ=−0\.03\\Delta\\kappa=\{\-\}0\.03to−0\.25\{\-\}0\.25; Appendix Table[18](https://arxiv.org/html/2609.28487#A0.T18)\), suggesting the pragmatic ceiling is task\-inherent rather than an English\-pretraining artefact, consistent with broader evidence that pragmatic understanding in LLMs remains highly sensitive to training strategy\([Ruis et al\. 2023](https://arxiv.org/html/2609.28487#bib.bib32)\)\.
#### Story\-form classification and outlet support\.
Story\-form classification is imperfect around the SOCIAL/OTHER/ELITE boundary\. The supervised development set is uniformly balanced across outlets while the production corpus is proportionally stratified, introducing an estimated precision drift of−\-4\.8 pp for the XLM\-R any\-salience baseline under prior shift \(full per\-device estimates in Appendix Table[28](https://arxiv.org/html/2609.28487#A0.T28); the final ensemble corpus any\-salience rate is slightly higher because the rhetorical\-question head is patched from CamemBERT\)\. Outlet\-level estimates are descriptive; smaller outlets, especially Blast \(646 headlines\), should be read as case profiles rather than equally precise population estimates\.
#### Group detection and event\-context confounds\.
Group detection is lexicon\-based and covers explicit mentions only\. We chose explicit surface\-form matching over NER or LLM\-based entity extraction for reproducibility and to eliminate hallucination risk; the lexicon is a transparent lower bound, not an exhaustive account of group references\. Group salience rates are descriptive: all reported differences from the corpus baseline are significant after Benjamini\-Hochberg correction, but extreme baseline volume differences \(e\.g\., Jews 0\.52% vs\. Farmers 2\.45%\) mean that high salience for rare groups is heavily driven by narrow, high\-intensity news events, with the yearly breakdown \(Appendix Table[21](https://arxiv.org/html/2609.28487#A0.T21)\) reducing sub\-annual clustering concern\. The year×\\timesstory\-form reweighting in Appendix Table[13](https://arxiv.org/html/2609.28487#A0.T13)is a coarse event\-context normalization, not a same\-event causal design: story forms are useful bins, but assigning a headline to CRIME, POLICY, CONFLICT, or SOCIAL is itself partly editorially mediated\. A true causal separation of event intensity from outlet framing would require matched same\-event corpora, article\-level event clustering, or parallel cross\-national coverage\. A blinded two\-human audit confirms high agreement with the released group labels \(κ=\.83\\kappa=\.83–\.98\.98\), though residual ambiguity concentrates in boundary\-heavy cases \(workers, broader far\-right references\)\.
#### Annotator release\-date and temporal contamination\.
Two LLM annotators \(openai/gpt\-oss\-120b, released August 2025;google/gemma\-4\-31B, released April 2026\) have training data that may overlap with the 2024–2025 portion of the annotation period; their pre\-training could introduce systematic annotation priors for high\-salience events \(Israel/Gaza coverage, French elections\) beyond pure device detection\. We directly tested this concern with a period\-stratified agreement analysis comparing 2022–2023 vs\. 2024–2025 annotations\. The analysis finds no evidence of systematic contamination\-inflated annotations: any\-salienceκ\\kappadeltas across the three annotator pairs range from\+0\.003\+0\.003to\+0\.032\+0\.032, and several binary heads show equal or lower cross\-model agreement in the later period, inconsistent with the contamination hypothesis\. Theκ\\kappadelta test is an aggregate check and cannot rule out systematic annotation priors on specific high\-salience events where LLM pretraining coverage is densest\. The two highest\-risk events, Israel/Gaza coverage from October 2023 onward and the June 2024 French legislative elections, are exactly the events most likely to carry model\-specific framing priors beyond the annotation schema\. The independent human validation sample \(N=499N\{=\}499\) is 39\.5% post\-October 7, 2023, providing annotation\-independent ground truth directly for the contamination\-exposed period; within this subset, any\-salience human–LLM agreement declines only modestly \(κ=\.424\\kappa\{=\}\.424vs\.\.463\.463full\-sample\), bounding the contamination effect at the aggregate level\. A residual risk remains for specific event clusters within this window; downstream users who require contamination\-free labels for Israel/Gaza or election coverage should use the human\-validatedN=499N\{=\}499subset only\. Contamination is also distinct from homogeneity: priors shared across the three annotators would not register in a period\-stratified comparison\. Appendix Table[9](https://arxiv.org/html/2609.28487#A0.T9)tests this separately and finds annotator–annotator agreement materially exceeding annotator–human agreement for loaded vocabulary and threat framing, so corpus rates for those two heads are the least independent of the five\.
#### Typology sensitivity\.
Causeur’s Double\-distinctive typology assignment rests partly on its ELITE\-heavy selection profile, where ELITE\-category arbitration agreement is lower \(κ=\.276\\kappa=\.276–\.417\.417\); a sensitivity analysis merging ELITE into OTHER confirms Causeur remains Double\-distinctive under this taxonomy perturbation \(BFMTV and Marianne shift cells; no focal outlet affected\)\.
#### Baseline self\-reference and precision\-floor shortfalls\.
The Sal\.JS and Sel\.JS corpus baseline is partially self\-referential for three high\-volume low\-divergence outlets \(Le Figaro, Le Parisien, Franceinfo; collectively≈\\approx32% of corpus\): their low observed divergence partially reflects their own weight in the baseline rather than absolute editorial similarity to the panel median\. A leave\-one\-out \(LOO\) baseline confirms rankings are robust: LOO\-correctedr=0\.721r=0\.721,ρ=0\.727\\rho=0\.727\(vs\. standardr=0\.736r=0\.736,ρ=0\.741\\rho=0\.741\), with only two near\-boundary outlets changing typology cells \(Valeurs actuelles: double→\\toselection\-dominant; L’Humanité: selection→\\todouble\-distinctive\); all focal outlets discussed in the text are stable, and the dissociability criterion \(r<0\.90r<0\.90\) holds comfortably under LOO correction\. Two device heads fall below their stated production\-precision floors after prior\-shift correction \(Elkan 2001 method, Appendix Table[28](https://arxiv.org/html/2609.28487#A0.T28)\): loaded vocabulary \(est\. 0\.643, floor 0\.70,−\-5\.7 pp\) and us\-vs\-them \(est\. 0\.550, floor 0\.60,−\-5\.0 pp\); blame attribution is marginally above floor \(est\. 0\.701\); threat framing \(est\. 0\.715\) and rhetorical question \(est\. 0\.929\) remain comfortably above their floors\. Corpus\-level loaded\-vocabulary and us\-vs\-them rates are therefore lower\-precision estimates relative to the other heads\. See also the inline note in §[4\.1](https://arxiv.org/html/2609.28487#S4.SS1)for loaded vocabulary\.
## Ethics, Reproducibility, and Data Availability
This study audits public editorial output and does not infer protected attributes of private individuals\. The study period covers the first enforcement year of the EU Digital Services Act \(DSA, February 2024 onward\), which introduced requirements for independent audits of very large online platforms’ algorithmic systems; our methodology is a descriptive research audit and is not a regulatory compliance instrument, but the transparent lexicon, public dataset release, and reproducible threshold protocol are designed to be consistent with independent audit principles\. Group labels refer to explicit lexical mentions in headlines and should not be interpreted as sentiment, endorsement, or hostility labels\. Outlet\-level and group\-level results describe aggregate measurement patterns and should not be used for individual profiling or content moderation decisions\.
The group salience rates carry a specific dual\-use risk: high salience for Jews\-mention headlines reflects antisemitism reporting and post\-October\-2023 Israel/Gaza security coverage, not editorial targeting; Far\-right salience reflects electoral, parliamentary, and conflict coverage of a political movement; elevated Muslim rates concentrate in policy and security contexts\. These numbers describe the rhetorical and event contexts in which groups appear in headlines \(the story types and wording conventions surrounding their mention\), not editorial sentiment, intent, or uniform targeting\. The Fdesouche exclusion check \(Appendix Table[20](https://arxiv.org/html/2609.28487#A0.T20)\) confirms elevated salience for Jews, Far\-right, and Muslims is distributed across the broader outlet panel rather than concentrated in a single far\-right source \(Migrants salience drops 10\.3 pp without Fdesouche, remaining 13\.5 pp above baseline\)\.
Specific political misuse scenarios warrant explicit counter\-framing: \(1\) antisemitic actors may cite “Jews: highest salience lift” as evidence of Jewish over\-representation in media discourse, inverting the finding, which measures the rhetorical intensity of event\-driven coverage contexts surrounding group mentions, not claims about the communities themselves; \(2\) immigration\-restriction actors may cite Migrants and Muslims figures as evidence of legitimizing media concern, ignoring that these rates are event\-driven and concentrated in security and policy story forms; \(3\) Far\-right actors may frame their high salience as evidence of media “persecution” rather than coverage density of a contested political movement in an election\-intensive period\. In all three scenarios, the interpretive key is the same: high salience in this audit measures the charge level of news contexts, driven by event type and outlet editorial register, not editorial hostility or targeting of the named groups\. LLM\-assisted annotations are not treated as ground truth: the final supervision set uses majority\-vote resolution and selective human arbitration\.
The finalv6supervision set, comprising 10,000 French headlines with binary salience labels, story\-form labels, split assignments, and outlet/section/date metadata, is released as a manifest\-only artifact: headline text is distributed as \(outlet, date, section, URL\) pointers to respect publisher copyright, so each record is sufficient for authorized reconstruction but does not redistribute verbatim content\. The human\-validation samples are the one disclosed exception: 516 headlines are released verbatim, because a blind double\-annotation study cannot be re\-run without the text that was annotated\. The group\-mention lexicons \(219 terms, 12 groups\) and the full annotation schema are included in full\. The 902,111\-headline corpus inference layer is released as a predictions\-only artifact: headline identifier, outlet, section, date, predicted labels, and raw classifier confidence scores at six\-decimal precision, with no headline text\. The annotator\-panel materials are included in full: pairwise agreement tables for all three candidate second annotators against the primary annotator, and per\-annotator device labels for the human\-validated 499\-headline subset, which together make the panel homogeneity check in Appendix Table[9](https://arxiv.org/html/2609.28487#A0.T9)independently reproducible\. Analysis and visualization scripts are released, except those requiring headline text\. Model training and threshold\-calibration code is not released: it is bound to the licensed 902,111\-headline corpus and to the trained checkpoints, neither of which can be redistributed\. All materials are available at[https://github\.com/lefrenchnewslab/framing\-wording\-selection](https://github.com/lefrenchnewslab/framing-wording-selection)\.
## References
- ACPM \(2026\)ACPM\. 2026\.[Classement audience OneNext — presse quotidienne nationale 2026 S1](https://www.acpm.fr/classements/onenext-pqn)\.Alliance pour les Chiffres de la Presse et des Médias \(ACPM\)\.
- Baly et al\. \(2020\)Ramy Baly, Giovanni Da San Martino, James Glass, and Preslav Nakov\. 2020\.[We can detect your bias: Predicting the political ideology of news articles](https://doi.org/10.18653/v1/2020.emnlp-main.404)\.In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pages 4982–4991, Online\. Association for Computational Linguistics\.
- Benson \(2013\)Rodney Benson\. 2013\.*Shaping Immigration News: A French\-American Comparison*\.Cambridge University Press, Cambridge\.
- Benson and Wood \(2015\)Rodney Benson and Timothy Wood\. 2015\.[Who says what or nothing at all? speakers, frames, and frameless quotes in unauthorized immigration news in the united states, norway, and france](https://doi.org/10.1177/0002764215573257)\.*American Behavioral Scientist*, 59\(7\):802–821\.
- Budak et al\. \(2016\)Ceren Budak, Sharad Goel, and Justin M\. Rao\. 2016\.[Fair and balanced? quantifying media bias through crowdsourced content analysis](https://doi.org/10.1093/poq/nfw007)\.*Public Opinion Quarterly*, 80\(S1\):250–271\.
- Cagé et al\. \(2022\)Julia Cagé, Moritz Hengel, Nicolas Hervé, and Camille Urvoy\. 2022\.[Hosting media bias: Evidence from the universe of french broadcasts, 2002–2020](https://doi.org/10.2139/ssrn.4036211)\.Technical Report hal\-03878119, Sciences Po / HAL\.
- Card et al\. \(2015\)Dallas Card, Amber E\. Boydstun, Justin H\. Gross, Philip Resnik, and Noah A\. Smith\. 2015\.[The media frames corpus: Annotations of frames across issues](https://doi.org/10.3115/v1/P15-2072)\.In*Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing \(Volume 2: Short Papers\)*, pages 438–444, Beijing, China\. Association for Computational Linguistics\.
- Dalibert \(2015\)Marion Dalibert\. 2015\.[Médias et mouvements sociaux minoritaires : un accès à la sphère publique régulé par la “francité” ?](https://doi.org/10.4000/sds.2406)*Sciences de la société*, 94:15–29\.
- Ecker et al\. \(2014\)Ullrich K\. H\. Ecker, Stephan Lewandowsky, Ee Pin Chang, and Rekha Pillai\. 2014\.[The effects of subtle misinformation in news headlines](https://doi.org/10.1037/xap0000028)\.*Journal of Experimental Psychology: Applied*, 20\(4\):323–335\.
- Elkan \(2001\)Charles Elkan\. 2001\.The foundations of cost\-sensitive learning\.In*Proceedings of the 17th International Joint Conference on Artificial Intelligence \(IJCAI\)*, pages 973–978\.
- Entman \(1993\)Robert M\. Entman\. 1993\.[Framing: Toward clarification of a fractured paradigm](https://doi.org/10.1111/j.1460-2466.1993.tb01304.x)\.*Journal of Communication*, 43\(4\):51–58\.
- Escouflaire et al\. \(2024\)Louis Escouflaire, Antonin Descampe, Antoine Venant, and Cédrick Fairon\. 2024\.La subjectivité dans le journalisme québécois et belge : transfert de connaissance inter\-médias et inter\-cultures\.In*Actes de JEP\-TALN\-RECITAL 2024, 31ème Conférence sur le Traitement Automatique des Langues Naturelles, volume 2 : traductions d’articles publiés*, pages 12–13\.Translation of article originally published at JADT 2024 \(17th International Conference on Statistical Analysis of Textual Data\)\. Data is francophone \(Québec \+ Belgium\), not France\.
- Fan et al\. \(2019\)Lisa Fan, Marshall White, Eva Sharma, Ruisi Su, Prafulla Kumar Choubey, Ruihong Huang, and Lu Wang\. 2019\.[In plain sight: Media bias through the lens of factual reporting](https://doi.org/10.18653/v1/D19-1664)\.In*Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\)*, pages 6343–6349, Hong Kong, China\. Association for Computational Linguistics\.
- Gabielkov et al\. \(2016\)Maksym Gabielkov, Arthi Ramachandran, Augustin Chaintreau, and Arnaud Legout\. 2016\.Social clicks: What and who gets read on twitter?In*ACM SIGMETRICS / IFIP Performance 2016*, Antibes Juan\-les\-Pins, France\.
- Galtung and Ruge \(1965\)Johan Galtung and Mari Holmboe Ruge\. 1965\.The structure of foreign news\.*Journal of Peace Research*, 2\(1\):64–91\.
- Gamson and Modigliani \(1989\)William A\. Gamson and Andre Modigliani\. 1989\.Media discourse and public opinion on nuclear power\.*American Journal of Sociology*, 95\(1\):1–37\.
- Gilardi et al\. \(2023\)Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli\. 2023\.[ChatGPT outperforms crowd workers for text\-annotation tasks](https://doi.org/10.1073/pnas.2305016120)\.*Proceedings of the National Academy of Sciences*, 120\(30\):e2305016120\.
- Gruppi et al\. \(2022\)Mauricio Gruppi, Benjamin D\. Horne, and Sibel Adali\. 2022\.[NELA\-GT\-2022: A large multi\-labelled news dataset for the study of misinformation in news articles](https://doi.org/10.48550/arXiv.2203.05659)\.*Preprint*, arXiv:2203\.05659\.
- Harcup and O’Neill \(2017\)Tony Harcup and Deirdre O’Neill\. 2017\.[What is news? news values revisited \(again\)](https://doi.org/10.1080/1461670X.2016.1150193)\.*Journalism Studies*, 18\(12\):1470–1488\.
- Jehle and Le Gallo \(2025\)Camille Jehle and Florian Le Gallo\. 2025\.[Europe in the headlines: What two decades of french news reveal about EU sentiment](https://www.banque-france.fr/en/publications-and-statistics/publications/europe-headlines-what-two-decades-french-news-reveal-about-eu-sentiment)\.Technical Report Working Paper 1008, Banque de France\.
- Landis and Koch \(1977\)J\. Richard Landis and Gary G\. Koch\. 1977\.[The measurement of observer agreement for categorical data](https://doi.org/10.2307/2529310)\.*Biometrics*, 33\(1\):159–174\.
- Liu et al\. \(2019\)Siyi Liu, Lei Guo, Kate Mays, Margrit Betke, and Derry Tanti Wijaya\. 2019\.[Detecting frames in news headlines and its application to analyzing news framing trends surrounding u\.s\. gun violence](https://doi.org/10.18653/v1/K19-1047)\.In*Proceedings of the 23rd Conference on Computational Natural Language Learning \(CoNLL\)*, pages 504–514, Hong Kong\. Association for Computational Linguistics\.
- McCombs \(2005\)Maxwell E\. McCombs\. 2005\.[A look at agenda\-setting: Past, present and future](https://doi.org/10.1080/14616700500250438)\.*Journalism Studies*, 6\(4\):543–557\.
- McCombs and Shaw \(1972\)Maxwell E\. McCombs and Donald L\. Shaw\. 1972\.[The agenda\-setting function of mass media](https://doi.org/10.1086/267990)\.*Public Opinion Quarterly*, 36\(2\):176–187\.
- Mendelsohn et al\. \(2021\)Julia Mendelsohn, Ceren Budak, and David Jurgens\. 2021\.[Modeling framing in immigration discourse on social media](https://doi.org/10.18653/v1/2021.naacl-main.179)\.In*Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, pages 2219–2263\.
- Newman et al\. \(2023\)Nic Newman, Richard Fletcher, Craig T\. Robertson, Kirsten Eddy, and Rasmus Kleis Nielsen\. 2023\.[Reuters Institute Digital News Report 2023](https://reutersinstitute.politics.ox.ac.uk/digital-news-report/2023)\.Technical report, Reuters Institute for the Study of Journalism, University of Oxford\.
- Otmakhova et al\. \(2024\)Yulia Otmakhova, Shima Khanehzar, and Lea Frermann\. 2024\.[Media framing: A typology and survey of computational approaches across disciplines](https://doi.org/10.18653/v1/2024.acl-long.822)\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 15407–15428, Bangkok, Thailand\. Association for Computational Linguistics\.Outstanding Paper Award, ACL 2024\.
- Pinto et al\. \(2019\)Sebastián Pinto, Federico Albanese, Claudio O\. Dorso, and Pablo Balenzuela\. 2019\.[Quantifying time\-dependent media agenda and public opinion by topic modeling](https://doi.org/10.1016/j.physa.2019.04.108)\.*Physica A: Statistical Mechanics and its Applications*, 524:614–624\.
- Piskorski et al\. \(2023\)Jakub Piskorski, Nicolas Stefanovitch, Giovanni Da San Martino, and Preslav Nakov\. 2023\.[SemEval\-2023 task 3: Detecting the category, the framing, and the persuasion techniques in online news in a multi\-lingual setup](https://doi.org/10.18653/v1/2023.semeval-1.317)\.In*Proceedings of the 17th International Workshop on Semantic Evaluation \(SemEval\-2023\)*, pages 2343–2361, Toronto, Canada\. Association for Computational Linguistics\.
- Price and Tewksbury \(1997\)Vincent Price and David Tewksbury\. 1997\.News values and public opinion: A theoretical account of media priming and framing\.In G\. A\. Barnett and F\. J\. Boster, editors,*Progress in the Communication Sciences*, volume 13, pages 173–212\. Ablex, New York\.
- Reisigl and Wodak \(2001\)Martin Reisigl and Ruth Wodak\. 2001\.*Discourse and Discrimination: Rhetorics of Racism and Antisemitism*\.Routledge, London\.
- Ruis et al\. \(2023\)Laura Ruis, Akbir Khan, Stella Biderman, Sara Hooker, Tim Rocktäschel, and Edward Grefenstette\. 2023\.[The Goldilocks of pragmatic understanding: Fine\-tuning strategy matters for implicature resolution by LLMs](https://proceedings.neurips.cc/paper_files/paper/2023/hash/4241fec6e94221526b0a9b24828bb774-Abstract-Conference.html)\.In*Advances in Neural Information Processing Systems*, volume 36\.
- Scheufele \(2004\)Bertram Scheufele\. 2004\.[Framing\-effects approach: A theoretical and methodological critique](https://doi.org/10.1515/comm.2004.29.4.401)\.*Communications*, 29\(4\):401–428\.
- Scheufele \(1999\)Dietram A\. Scheufele\. 1999\.[Framing as a theory of media effects](https://doi.org/10.1111/j.1460-2466.1999.tb02784.x)\.*Journal of Communication*, 49\(1\):103–122\.
- Scheufele and Tewksbury \(2007\)Dietram A\. Scheufele and David Tewksbury\. 2007\.[Framing, agenda setting, and priming: The evolution of three media effects models](https://doi.org/10.1111/j.0021-9916.2007.00326.x)\.*Journal of Communication*, 57\(1\):9–20\.
- Shoemaker and Vos \(2009\)Pamela J\. Shoemaker and Tim P\. Vos\. 2009\.*Gatekeeping Theory*\.Routledge, New York\.
- Song et al\. \(2024\)Wenlong Song, Bo Pang, Zihan Wang, Yilang Xu, Zetong Liu, and Murong Tan\. 2024\.[Differentiated secularism: discourse shaping of immigrants from different religious backgrounds in the french media](https://doi.org/10.1057/s41599-024-03390-x)\.*Humanities and Social Sciences Communications*, 11:893\.
- Spinde et al\. \(2021\)Timo Spinde, Manuel Plank, Jan\-David Krieger, Terry Ruas, Bela Gipp, and Akiko Aizawa\. 2021\.[Neural media bias detection using distant supervision with babe \- bias annotations by experts](https://doi.org/10.18653/v1/2021.findings-emnlp.101)\.In*Findings of the Association for Computational Linguistics: EMNLP 2021*, pages 1166–1177, Punta Cana, Dominican Republic\. Association for Computational Linguistics\.
- Stevenson \(1944\)Charles L\. Stevenson\. 1944\.*Ethics and Language*\.Yale University Press, New Haven\.
- Tankard \(2001\)James W\. Tankard\. 2001\.The empirical approach to the study of media framing\.In Stephen D\. Reese, Oscar H\. Gandy, and August E\. Grant, editors,*Framing Public Life*, pages 95–106\. Lawrence Erlbaum, Mahwah, NJ\.
- Törnberg \(2025\)Petter Törnberg\. 2025\.[Large language models outperform expert coders and supervised classifiers at annotating political social media messages](https://doi.org/10.1177/08944393241286471)\.*Social Science Computer Review*, 43\(6\):1181–1195\.
- Tsimpoukis \(2025\)Panos Tsimpoukis\. 2025\.[Contesting dominant AI narratives on an industry\-shaped ground: public discourse and actors around AI in the french press and social media \(2012–2022\)](https://doi.org/10.22323/2.24020210)\.*Journal of Science Communication*, 24\(2\):A10\.
- van Dijk \(1991\)Teun A\. van Dijk\. 1991\.*Racism and the Press*\.Routledge, London\.
- Walton \(1996\)Douglas N\. Walton\. 1996\.*Argumentation Schemes for Presumptive Reasoning*\.Lawrence Erlbaum, Mahwah, NJ\.
- Ziems et al\. \(2024\)Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang\. 2024\.[Can large language models transform computational social science?](https://doi.org/10.1162/coli_a_00502)*Computational Linguistics*, 50\(1\):237–291\.
Table 7:Finalv6label distribution in the 10,000\-headline development set after majority\-vote merging and story\-form arbitration\.Table 8:Held\-out validation of group\-mention lexicons\.\(b\) Homogeneity check \(N=499N=499\)
Table 9:\(a\) Compact reliability summary for the finalv6annotation stack, computed on the 10,000 rows shared by all three composite annotators after UVT harmonization\. \(b\) Homogeneity check on the 499\-headline validation sample; all cells are pairwise between individuals and therefore comparable\. LLM–LLM = mean pairwiseκ\\kappaamong the three annotators; LLM–hum\. = each annotator against each human coder \(6 pairs\); H–H = the two coders\. Shared blind spots would show as LLM–LLM exceeding*both*LLM–hum\. and H–H, which holds for loaded vocabulary and threat framing only\. On unrounded values, the LLM–LLM minus LLM–hum\. gap is\+\.177\+\.177and\+\.139\+\.139for those two, against\+\.067\+\.067for us\-vs\-them,\+\.012\+\.012for blame attribution and\+\.002\+\.002for rhetorical question\.Table 10:One example per primary salience device from thev6supervision set, each carrying the stated device label by unanimous agreement of all three LLM annotators; the released headline identifier and split are given so every row can be located in the distributed data\. The final row illustrates the story\-form/topic distinction\.Table 11:Full four\-cell outlet typology with per\-outlet headline counts\. Median split thresholds: Sal\.JS=0\.0110, Sel\.JS=0\.0249\. Displayed JS divergence values are rounded to three decimal places; cell assignments use full\-precision values\. L’Humanité’s full\-precision Sal\.JS \(0\.0110\) equals the median; strict\>\>convention places it in the low\-salience half\.†Blast \(N=646N\{=\}646\): case profile, wide CIs \(Table[12](https://arxiv.org/html/2609.28487#A0.T12)\)\.Table 12:Representative 95% Wilson intervals for outlet\-level and group\-level headline rates, recomputed from the canonical paper artifacts\. These are descriptive intervals over observed predicted headline rates only; they do not propagate classifier or lexicon uncertainty\.†Blast:N=646N\{=\}646headlines; estimates carry wide bootstrap CIs \(displayed\) and should be read as a case profile rather than a population estimate\.Table 13:Group\-level salience contexts across 902,111 headlines, with coarse event\-context normalization\. Y×\\timesStory exp\.% is the expected any\-salience rate after exact reweighting to each group’s year×\\timesstory\-form composition, subtracting focal\-group rows from each cell baseline\. Resid\. pp is observed minus expected; it is descriptive, not a same\-event causal estimate\.
Table 14:Condensed group\-level mechanism table\. It separates rhetorical device lift from selected story\-form context\.
Table 15:Within\-story\-form checks\. Values are percentages; “Others” pools non\-focal outlets within the same story form\. All contrasts are significant atp<\.001p<\.001\. Device\-levelχ2\\chi^\{2\}andϕ\\phifor Fdesouche CRIME blame and threat are included because these rows are directly cited in the text\.
Table 16:Qualitative error analysis: ten representative false positives \(FP\) and false negatives \(FN\) from the 1,501\-headline test set, two per salience head, drawn from highest\-confidence misclassifications\. Headlines are shown verbatim \(truncated where necessary\); error\-source annotations are manual\.Table 17:Corroborating independent blind study on a separate sample: two unaffiliated annotators \(N=350N\{=\}350\)\.Note\.Stratified 350\-headline sample drawn exclusively from thev6test split \(all HIGH\-confidence; seed 42; no overlap with the primary 499\-headline sample in Table[4](https://arxiv.org/html/2609.28487#S3.T4)\)\. Annotation protocol matches the primary study \(two unaffiliated annotators, blind to LLM labels and expected distributions\)\. H–H = human–human reliability; Maj–LLM = majority\-vote human labels vs\. LLM consensus \(conflict rows, where A1≠\\neqA2, excluded; conflict counts: loaded 53, blame 31, threat 72, rhetorical question 3, us\-vs\-them 60, any salience 48\)\. F1/Prec\./Rec\.: human majority vote as gold standard\. Interpretation thresholds follow[Landis and Koch 1977](https://arxiv.org/html/2609.28487#bib.bib21)\.
FieldMistral–humanκ\\kappaEnsemble–humanκ\\kappaΔ\\DeltaLoaded vocabulary\.553\.684−\-\.131Blame attribution\.351\.379−\-\.028Threat framing\.517\.614−\-\.097Rhetorical question\.324\.400−\-\.076Us\-vs\-them\.298\.546−\-\.248Any salience\.435\.559−\-\.124Story form\.449\.546−\-\.097Table 18:French\-specific Mistral comparison on the 499\-headline human validation sample\. Mistral =mistral\-small3\.2; ensemble = releasedv6consensus labels\.Δ\\Deltais Mistral–humanκ\\kappaminus ensemble–humanκ\\kappa; negative values indicate lower agreement than the released ensemble\.Table 19:Threshold recalibration sensitivity by outlet\.Note\.“Default” uses a uniform 0\.50 threshold; “Recal\.” uses the precision\-floor production thresholds from Table[2](https://arxiv.org/html/2609.28487#S3.T2)\. Default thresholds inflate corpus\-wide any\-salience by 6\.9 pp and high\-charge by 3\.3 pp, but the outlet ranking is preserved\.†Blast \(N=646N\{=\}646\): case profile, wide CIs \(Table[12](https://arxiv.org/html/2609.28487#A0.T12)\)\.
Table 20:Fdesouche exclusion sensitivity\.Note\.Median exclusion effect across all 12 groups is−\-0\.5 pp\. Migrants is the most sensitive case: removing Fdesouche lowers salience by 10\.3 pp \(58\.4% to 48\.1%\), but the group still sits 13\.5 pp above the corpus baseline\.
Table 21:Year\-by\-year any\-salience rates for the five highest\-salience groups in the full\-corpus group audit\. Parenthetical values are lifts relative to the corpus\-wide any\-salience baseline for the same calendar year, which rises from 31\.7% \(2022\) to 38\.1% \(2025\)\. The same three groups, Jews, the Far\-right, and Muslims, rank highest in every year, indicating that the main RQ3 pattern is not reducible to a single late\-period news cycle\.Table 22:Supplementary us\-vs\-them analysis\. Because the us\-vs\-them head achieved only moderate human–human \(κ=\.510\\kappa=\.510\) and human–model \(κ=\.596\\kappa=\.596\) agreement, it is retained as a supplementary indicator rather than a primary prevalence metric\. The values above are therefore best read as qualitative profile signals rather than exact point estimates\.Table 23:Training hyperparameters for CamemBERT\-base and XLM\-RoBERTa\-large on the finalv6supervision set\. Both models train on the same 6,999\-headline training split; thresholds are selected on the 1,500\-headline validation split only\. Results in Table[3](https://arxiv.org/html/2609.28487#S3.T3)are from the seed\-42 run; three\-seed evaluation \(seeds 42, 123, 456\) confirms seed\-robust model selection\. Full per\-seed results:v6\_model/colab/multiseed\_variance\_eval\.ipynb\.
Table 24:Condensed annotation schema for the released supervision criteria\.†Us\-vs\-them requires both groups to hold explicit social, political, ethnic, national, class, or institutional identity; borderline guidance covers politician\-vs\.\-politician \(classified as CONFLICT\) and commercial competition as canonical exclusions\. The full annotation prompt, user template, lexicons, and analysis scripts are available at[https://github\.com/lefrenchnewslab/framing\-wording\-selection](https://github.com/lefrenchnewslab/framing-wording-selection)\.MetricValueNCohen’sκ\\kappa\(nominal\)\.654642Cohen’sκ\\kappa\(linear\)\.666642Krippendorff’sα\\alpha\.653642% raw agreement69\.9%642By primary annotator confidenceHIGH\.729318MEDIUM\.586314LOW\.15710
Table 25:Post\-hoc human–human inter\-annotator agreement on the 642 three\-way story\-form conflict cases resolved by single\-annotator arbitration\. A second independent annotator applied the same codebook without access to the primary annotator’s decisions\. The original gold labels were not modified\. Agreement is higher for cases the primary annotator rated HIGH\-confidence \(κ=\.729\\kappa=\.729\) than MEDIUM\-confidence \(κ=\.586\\kappa=\.586\), confirming well\-calibrated self\-assessed uncertainty\. Lower agreement on conflict trios involving ELITE \(κ=\.276\\kappa=\.276–\.417\) reflects the residual\-category nature of ELITE in the schema\. Interpretation thresholds follow[Landis and Koch 1977](https://arxiv.org/html/2609.28487#bib.bib21)\.Table 26:Outlets whose four\-cell assignments change under±\\pm20% perturbations of the salience and selection median split points\. All remaining outlets keep the same cell assignment under every perturbation, indicating that instability concentrates in near\-boundary cases, not distinctive outliers\.Table 27:Rhetorical\-question\-excluded any\-salience robustness\. Recomputing any\-salience as loaded vocabulary OR blame attribution OR threat framing \(dropping the rhetorical\-question head\) reduces the corpus\-level detected rate from 34\.6% to 30\.6% \(4\.0 pp\)\. The Sal\.JS/Sel\.JS Pearson correlation drops fromr=\.736r=\.736tor=\.561r=\.561\(r=\.650r=\.650excluding Fdesouche\)\. Four near\-median outlets flip typology cells: Mediapart and Valeurs actuelles \(double\-distinctive→\\rightarrowselection\-dominant\) and Marianne and L’Humanité \(selection\-dominant→\\rightarrowdouble\-distinctive\); all other 21 outlets are stable\.†Blast \(N=646N\{=\}646\): case profile, wide CIs \(Table[12](https://arxiv.org/html/2609.28487#A0.T12)\)\.Table 28:Estimated precision drift under class\-prior shift\. Expected production precision is computed from the validation\-set TPR and FPR at each head’s chosen threshold, applying the production corpus prevalence via[Elkan 2001](https://arxiv.org/html/2609.28487#bib.bib10)\. The any\-salience OR row is an XLM\-R baseline diagnostic computed before the final CamemBERT rhetorical\-question substitution, so its production prevalence \(33\.7%\) is slightly lower than the final ensemble corpus any\-salience rate in the main text \(34\.6%\)\.\(a\) Sparse\-cell exclusion sensitivity
\(b\) Cluster\-robust inference \(G=25G\{=\}25\)
Table 29:Temporal panel robustness\. \(a\) Sensitivity to sparse\-cell exclusion thresholds \(two\-way outlet\+\+month fixed effects, classical SEs\): the within\-outlet coupling is stable across thresholds\. \(b\) Cluster\-robust standard errors \(clustered by outlet,G=25G\{=\}25, small\-sample correction\): positive under both specifications but not individually significant, consistent with treating the panel as confirmatory rather than a stand\-alone test\.Similar Articles
Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, 2022-2025
This paper analyzes 28,592 French news headlines about La France insoumise and Rassemblement National using an LLM annotation pipeline, finding asymmetric role framing where LFI is more often framed as aggressors and RN as strategic actors.
Uncovering Temporal Framing in the News
This paper proposes a taxonomy of eight temporal frames for news discourse, presents a multilingual dataset with expert annotations, and evaluates supervised and zero-shot classification for detecting temporal framing.
Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions
This paper investigates how contextual framing affects LLM responses in mental health interactions, finding systematic behavioral variation and demonstrating that internal representations encode framing information throughout transformer layers.
GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models
This paper introduces GPF-LiveNews, a streaming evaluation protocol for auditing how large language models frame live news events differently for various demographic groups, using semantic sensitivity and sentiment disparity measures across 42 identity labels and seven prompt families.
Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
This paper empirically evaluates how well LLMs align with human emotional perception of news framing, using a YouGov survey of 3,011 UK adults and seven LLMs assessing sympathy in headlines. It finds that alignment varies across models and demographic subgroups, highlighting the importance of differential alignment for AI development.