Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study

arXiv cs.CL Papers

Summary

This study audits TikTok for harmful content exposure using multimodal language models across multiple countries and age groups, finding that keyword searches increase harmful content and that MLLMs provide a scalable tool for youth-safety audits.

arXiv:2608.17583v1 Announce Type: new Abstract: Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages. We audit TikTok in France, Italy, and Sweden with sockpuppet accounts representing four age personas (13, 16, 19, 40), collecting 36,971 videos from passive For-You-page scrolling and active sessions that scroll, search for harm keywords, and scroll again. To scale annotation, we validate four multimodal LLMs against native-speaker labels on a 300-video reference set. Gemini 2.5 Flash with eight sampled frames plus text performs best (aggregate kappa = 0.42), at half the per-call cost of native-video upload, and we apply it to a 10% sample for approximately \$50 in total API spend across both modalities. Keyword search returns 35-56% harmful content, a 1.5-7.5x increase over the scrolling baseline in ten of twelve country-age combinations; the spike is temporary and flattens the age differences observed in France and Sweden. Under passive scrolling, Italy has the highest harm rate at every age, with Italian age-19 reaching 48.6%. Overall, MLLM-based auditing offers a scalable approach for cross-national youth-safety audits, while provider safety filters (1.1% refusal rate) under-count the most explicit harms.
Original Article
View Cached Full Text

Cached at: 08/19/26, 10:00 AM

# Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models:A Cross-National, Age-Stratified Study
Source: [https://arxiv.org/html/2608.17583](https://arxiv.org/html/2608.17583)
Hamidreza SaffariFrancesco PierriAffiliation:Politecnico di Milanohamidreza\.saffari@mail\.polimi\.it,francesco\.pierri@polimi\.it

###### Abstract

Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages\. We audit TikTok in France, Italy, and Sweden with sockpuppet accounts representing four age personas \(13, 16, 19, 40\), collecting36,97136\{,\}971videos from passive For\-You\-page scrolling and active sessions that scroll, search for harm keywords, and scroll again\. To scale annotation, we validate four multimodal LLMs against native\-speaker labels on a 300\-video reference set\. Gemini 2\.5 Flash with eight sampled frames plus text performs best \(aggregateκ=0\.42\\kappa=0\.42\), at half the per\-call cost of native\-video upload, and we apply it to a10%10\\%sample for approximately $50 in total API spend across both modalities\. Keyword search returns3535–56%56\\%harmful content, a1\.51\.5–7\.5×7\.5\\timesincrease over the scrolling baseline in ten of twelve country–age combinations; the spike is temporary and flattens the age differences observed in France and Sweden\. Under passive scrolling, Italy has the highest harm rate at every age, with Italian age\-19 reaching48\.6%48\.6\\%\. Overall, MLLM\-based auditing offers a scalable approach for cross\-national youth\-safety audits, while provider safety filters \(1\.1%1\.1\\%refusal rate\) under\-count the most explicit harms\.

## 1Introduction

TikTok has become one of the most influential gateways through which young users encounter video content online, with a For\-You feed that rapidly adapts to micro\-interactions such as watch time and loop rates\([7](https://arxiv.org/html/2608.17583#bib.bib12);[5](https://arxiv.org/html/2608.17583#bib.bib11)\)\. Independent audits have documented self\-harm and suicidal\-ideation content reaching newly created teen accounts within hours\([1](https://arxiv.org/html/2608.17583#bib.bib17);[2](https://arxiv.org/html/2608.17583#bib.bib18)\), weight\-normative and pro\-eating\-disorder content dominating health\-adjacent feeds\([18](https://arxiv.org/html/2608.17583#bib.bib8);[21](https://arxiv.org/html/2608.17583#bib.bib9);[6](https://arxiv.org/html/2608.17583#bib.bib10)\), and age\-gating mechanisms that do not meaningfully shield underage personas relative to adults\([9](https://arxiv.org/html/2608.17583#bib.bib15);[26](https://arxiv.org/html/2608.17583#bib.bib1)\)\. Public\-health concerns\([19](https://arxiv.org/html/2608.17583#bib.bib19)\)and regulatory frameworks like the EU Digital Services Act make independent, reproducible audits of short\-video platforms increasingly urgent\.

One response to this moderation burden is to enlist multimodal large language models \(MLLMs\) as automated annotators: recent work shows aligned hate\-speech judgments with humans\([8](https://arxiv.org/html/2608.17583#bib.bib4);[13](https://arxiv.org/html/2608.17583#bib.bib2)\)and improved alignment from policy\- and rule\-grounded MLLMs\([25](https://arxiv.org/html/2608.17583#bib.bib6)\)\. At the same time, MLLMs disagree substantially with each other on the same items\([10](https://arxiv.org/html/2608.17583#bib.bib3)\)and can exploit language priors to answer without looking at the video\([3](https://arxiv.org/html/2608.17583#bib.bib21)\); how these strengths and failure modes interact with*age*\(which shifts the harm distribution the algorithm exposes\),*country*\(which introduces cross\-lingual moderation challenges\), and*input modality*, the three dimensions most relevant to youth\-safety audits of TikTok, remains largely unexplored\.

We address three research questions:

- •RQ1:Which MLLM configuration agrees with native\-speaker annotators well enough to run at scale?
- •RQ2:What is the marginal value of sampled frames and native video over text\-only input?
- •RQ3:How does harm exposure differ across age personas, countries, and platform signals?

We tackle them with a two\-stage audit on three EU countries \(France, Italy, Sweden\) and four age personas \(13, 16, 19, 40, spanning TikTok’s stated minimum, mid\-adolescence, young adulthood in the platform’s 18\+ tier, and an adult control\), via a within\-accountscroll\-pre→\\toSEARCH→\\toscroll\-postcycle capturing both baseline FYP exposure and the platform’s response to an active probe\. Stage\-1 selects a validated MLLM auditor on a 300\-video annotated subset; Stage\-2 runs it on a phase\-stratified10%10\\%sample of the full36,97136\{,\}971\-video corpus\.

On RQ3, within\-account harm\-keyword search reaches3535–56%56\\%harm against an immediatescroll\-prebaseline of99–44%44\\%on the same accounts, and the scroll\-post snapshot reverts to baseline; under our Stage\-1\-validated Gemini 2\.5 Flash E3 auditor, Italy is the most exposed country on the three youngest age personas, with its two youngest combinations already at the ceiling atscroll\-pre\. The full audit runs at approximately $49 of API spend\.

#### Our contributions are:

- •Empirical findings on TikTok harm exposureacross three countries \(France, Italy, Sweden\) and four age personas \(13, 16, 19, 40\): keyword search spikes harm exposure1\.51\.5–7\.5×7\.5\\timesover passive scrolling, the spike is temporary, age differences flatten under search, and Italy has the highest passive harm rate at every age\.
- •An automated harm\-annotation pipeline: we use a multimodal language model \(Gemini 2\.5 Flash\) as the annotator, validated against native\-speaker labels on a 300\-video subset and then applied to the full corpus for approximately $50 in API costs\.
- •Released artifacts: a 13\-category harm taxonomy aligned with TikTok’s Community Guidelines\([22](https://arxiv.org/html/2608.17583#bib.bib23)\), the annotation schema and prompts, the per\-country keyword lists, and metadata for the36,97136\{,\}971\-video corpus\.

## 2Related Work

#### TikTok audits and youth safety

TikTok has been a recurring target for algorithmic auditing since[7](https://arxiv.org/html/2608.17583#bib.bib12)isolated the personalization factors that drive its For\-You feed, and[5](https://arxiv.org/html/2608.17583#bib.bib11)subsequently modeled how engagement vectors exponentially amplify niche content\. Qualitative work has documented what this amplification surfaces in practice, particularly weight\-normative and pro\-eating\-disorder content\([18](https://arxiv.org/html/2608.17583#bib.bib8)\)\. The methodological grounding for the sockpuppet protocol we use comes from[20](https://arxiv.org/html/2608.17583#bib.bib13), who articulated persona\-based audits as a rigorous analogue of offline discrimination studies;[17](https://arxiv.org/html/2608.17583#bib.bib14)is a recent TikTok\-specific instantiation that isolates the effect of interaction signals \(e\.g\., likes\) on recommendations\. Closest in design to our work are the two age\-stratified TikTok audits:[26](https://arxiv.org/html/2608.17583#bib.bib1)simulate age\-specific sockpuppets under passive and active engagement and find that platform enforcement does not meaningfully shield under\-18 accounts, and[9](https://arxiv.org/html/2608.17583#bib.bib15)report a cross\-platform version via manual harm annotation, showing that 13\-year\-old accounts encounter harmful content significantly more often than 18\-year\-old accounts on TikTok, YouTube, and Instagram\.

#### MLLMs as content moderators and annotators

[8](https://arxiv.org/html/2608.17583#bib.bib4)find that frontier multimodal LLMs can produce context\-sensitive hate evaluations that align with aggregate human judgment, supporting the basic feasibility of MLLM\-as\-auditor use\.[10](https://arxiv.org/html/2608.17583#bib.bib3)demonstrate that different LLM\-based moderators disagree substantially on the same items, which is why our Stage\-1 compares four model families rather than committing to a single frontier model\.[3](https://arxiv.org/html/2608.17583#bib.bib21)show that video LLMs often exploit language priors rather than performing genuine temporal reasoning, which is the failure mode our text\-only E1 baseline is designed to surface\.[11](https://arxiv.org/html/2608.17583#bib.bib5)reframes the moderation\-evaluation problem from accuracy toward legitimacy, the framing required when LLMs are proposed as substitutes for human moderators on platforms with global user bases\.

#### Cross\-lingual moderation and frame\-based MLLM annotation

[23](https://arxiv.org/html/2608.17583#bib.bib20)analyze EU Digital Services Act transparency reports to quantify content\-moderator workforces across languages and platforms; the three countries in our study \(France, Italy, Sweden\) span the upper, middle, and lower end of TikTok’s per\-language moderator allocation in that data\. On the prompting side,[25](https://arxiv.org/html/2608.17583#bib.bib6)provide evidence that policy\- and rule\-grounded multimodal LLMs substantially improve alignment with formal community guidelines, motivating the 13\-category taxonomy injection in our system prompt \(§[3\.3](https://arxiv.org/html/2608.17583#S3.SS3)\)\. On the input side,[13](https://arxiv.org/html/2608.17583#bib.bib2)use sampled frames plus thumbnail and text metadata to evaluate GPT\-4\-Turbo against crowdworkers on 19k YouTube videos; their fourteen\-frame protocol directly inspired our eight\-frame E3 condition \(§[3\.3](https://arxiv.org/html/2608.17583#S3.SS3)\)\.

![Refer to caption](https://arxiv.org/html/2608.17583v1/pipeline.png)Figure 1:End\-to\-end audit pipeline: data collection, Stage\-1 MLLM validation, Stage\-2 prevalence measurement\.

## 3Methodology

The audit proceeds in two stages\. Stage\-1 selects the MLLM auditor: four candidate models are scored against a 300\-video reference set, drawn at random with 25 videos per \(country, age\) cell, independently labeled by two native\-speaker annotators per country and reconciled through joint resolution, and one configuration is carried forward\. The draw is stratified by country and age but not by collection phase; its harm\-category mix is not controlled, since categories are only known after annotation\. Stage\-2 applies that configuration to a phase\-stratified 10% sample of the full corpus under both native\-video upload \(E2\) and an eight\-frame protocol \(E3\)\. All exposure results in §[4](https://arxiv.org/html/2608.17583#S4)come from Stage\-2; Stage\-1 supplies the validation behind them\. Figure[1](https://arxiv.org/html/2608.17583#S2.F1)summarizes the end\-to\-end pipeline\.

### 3\.1Data Collection

We audit TikTok through persona accounts organized on two axes: country and user age\. Each of the three countries in our study \(France \(FR\), Italy \(IT\), and Sweden \(SW\)\) is paired with four age personas, 13, 16, 19, and 40, yielding twelve \(country, age\) combinations\. For each combination, we use three independent accounts to avoid relying on the behavior of a single account\. All accounts report their age during registration and use the system locale in the account’s native language\. We set the account location by routing traffic through a VPN endpoint in the target country, since TikTok’s recommendation and search systems rely on the country linked to the IP address rather than the region listed in the account profile\. Without the VPN, feeds tend to match the VPN\-free IP location instead of the intended persona location\. Each account is driven by a Tampermonkey userscript injected into the TikTok web client that captures the platform’s API responses and emulates user scrolling with a randomized scroll amount and an inter\-scroll delay drawn uniformly from500500–10001000ms; the three accounts per combination run in parallel on separate browser instances\. We use the same personas in both collection phases described below\. Sockpuppet\-based audits of recommender systems have a long methodological tradition\([20](https://arxiv.org/html/2608.17583#bib.bib13);[7](https://arxiv.org/html/2608.17583#bib.bib12)\), with recent applications to TikTok auditing both age\-stratified content exposure\([26](https://arxiv.org/html/2608.17583#bib.bib1);[9](https://arxiv.org/html/2608.17583#bib.bib15)\)and political content skew\([12](https://arxiv.org/html/2608.17583#bib.bib16)\)\.

We selected France, Italy, and Sweden because they span a high/medium/low range of TikTok per\-language EU moderator allocation in the DSA transparency data analyzed by[23](https://arxiv.org/html/2608.17583#bib.bib20)\(averaged across 2023–2024 reporting periods: French∼\\sim620, Italian∼\\sim396, Swedish∼\\sim98 moderators\)\.

Data collection proceeds in two phases\. The passive phase records what the algorithm surfaces under pure scrolling: each account consumes the For\-You feed \(FYP\) and captures the sequence of recommended videos together with the interactions TikTok exposes in its API responses\. The passive phase measures what the algorithm surfaces when no active intent is signaled; this matches prior work showing that engagement\-driven amplification is highly sensitive to watch\-time patterns even under minimal interaction\([7](https://arxiv.org/html/2608.17583#bib.bib12);[5](https://arxiv.org/html/2608.17583#bib.bib11)\)\. The passive phase was collected between 30 December 2025 and 11 January 2026, five consecutive days per country, and contains 14,093 unique videos across the twelve \(country, age\) combinations \(Table[1](https://arxiv.org/html/2608.17583#S3.T1)\)\.

The active phase probes how the platform responds to keyword\-level intent\. For each country we curate three native\-language keywords per harm category \(sevenSEARCH\-probed categories from §[3\.2](https://arxiv.org/html/2608.17583#S3.SS2); full list in Appendix[A](https://arxiv.org/html/2608.17583#A1)\)\. Each active run follows ascroll\-pre→\\rightarrowSEARCH→\\rightarrowscroll\-postcycle\. Active collection ran for five days per country between 13 and 19 April 2026 \(three weekdays and two weekend days\), with three accounts per \(country, age\) combination scrolling and probing concurrently\. The protocol yielded 47,674 capture events covering 22,878 unique videos \(FR: 7,130; IT: 7,531; SW: 8,217\)\. Every capture is tagged with its phase, keyword, category, and originating account\.

#### Dataset overview

Table[1](https://arxiv.org/html/2608.17583#S3.T1)reports the unique\-video yield per phase, country, and age persona\. Together the two phases cover 36,971 unique videos \(14,093 passive \+ 22,878 active\)\. Per\-combination counts are balanced to first order across both phases, with each \(country, age\) combination contributing roughly 1,100–2,200 unique videos across both phases\. Full per\-\(country, age, phase\) breakdowns and engagement statistics are in Appendix[C](https://arxiv.org/html/2608.17583#A3)\. English dominates passive feeds in every country \(FR 45\.7%, IT 30\.3%, SW 48\.3%\), with native\-language content at only10\.510\.5–20\.6%20\.6\\%\. Native\-language keyword queries in the active phase lift native content to3030–39%39\\%and reduce the English share to23\.923\.9–28\.9%28\.9\\%\. Full breakdown in Appendix[D](https://arxiv.org/html/2608.17583#A4), Table[5](https://arxiv.org/html/2608.17583#A3.T5)\.

Table 1:Unique videos per phase, country, and age persona \(36,97136\{,\}971in total; the cross\-country overlap of videos collected is≈22%\\approx 22\\%in the passive approach\)\.

### 3\.2Harm Taxonomy and Manual Annotation Schema

#### Taxonomy

We adopt a 13\-category harm taxonomy aligned with TikTok’s Community Guidelines\([22](https://arxiv.org/html/2608.17583#bib.bib23)\), spanning disordered eating, self\-harm, dangerous challenges, nudity, sexually suggestive content, shocking/graphic content, hate speech, sexual abuse, trafficking, gambling, alcohol/tobacco/drugs, integrity, and harassment \(full list with definitions in Appendix[B](https://arxiv.org/html/2608.17583#A2)\)\. Compared with the six\-category taxonomy of[13](https://arxiv.org/html/2608.17583#bib.bib2)and the cross\-platform typology of[24](https://arxiv.org/html/2608.17583#bib.bib22), this taxonomy is finer\-grained and lets us probe category\-level asymmetries that coarser schemes hide\.

#### Annotation schema

Each video receives a three\-way verdict \(Harmful,Not Harmful, orVideo Not Availablewhen the embed fails or the content is region\-locked\) and, ifHarmful, a requiredprimaryand optionalsecondaryharm category \(max two per video\)\. Hesitation is captured by a required Confidence rating \(Low,Medium,High\) rather than a “borderline” option, and annotators are instructed to markLowwhenever they hesitate\.

#### Annotators

Two native\-speaker annotators per country independently label the sampled subset and reconcile disagreements through joint resolution to producefinal reference labels\(per\-country class balance in Fig\.[2](https://arxiv.org/html/2608.17583#S3.F2)\); the resolution rule for subcategory disjunction and pre\-resolution inter\-annotator agreement are reported in Appendix[E](https://arxiv.org/html/2608.17583#A5)\. The annotators are native\-speaker graduate students recruited among the authors’ colleagues; they were briefed on the taxonomy, informed that the material could include distressing content, consented to labeling it for research, and worked in self\-paced sessions through a dedicated web panel\. They were not financially compensated, and each annotator spent approximately five days on the task\.

Figure 2:Per\-country share ofHarmfullabels in the 300\-video Stage\-1 reference subset after annotator resolution \(n=100n=100per country\)\. Class balance differs across countries: roughly:41\\\!:\\\!4in France,:11\\\!:\\\!1in Italy,:21\\\!:\\\!2in Sweden\.

### 3\.3Models and Experimental Conditions

#### Models

We evaluate four MLLMs spanning three provider families: Gemini 2\.5 Flash \(Google AI SDK, native\-video capable\), Qwen3\-VL\-32B \(Alibaba DashScope, native\-video capable\), GPT\-4o\-mini, and Mistral Large 3 \(both via OpenRouter\)\. Two models process video natively at the API layer; the other two ingest only text and discrete frames\. Since different MLLMs disagree substantially on the same moderation items\([10](https://arxiv.org/html/2608.17583#bib.bib3)\), we read aggregate performance and pairwise disagreement side by side\.

#### Three input conditions

Each model is evaluated under three conditions: E1 \(text\-only\), the video caption plus an audio transcript, isolating the linguistic prior\([3](https://arxiv.org/html/2608.17583#bib.bib21)\); E2 \(native video\), the raw MP4 uploaded through the provider’s video API, applicable only to the two native\-video\-capable models; and E3 \(frames\-as\-images\), eight frames sampled uniformly over the video and submitted as inline base64 images alongside the caption and transcript, following the frame\-plus\-metadata protocol of[13](https://arxiv.org/html/2608.17583#bib.bib2)\.

#### Prompting

All conditions share a system prompt that injects the 13\-category taxonomy with one\-sentence definitions, following[25](https://arxiv.org/html/2608.17583#bib.bib6)’s evidence that policy\- and rule\-grounded MLLM moderation aligns better with formal guidelines\. The model returns structured JSON with the verdict, a primary harm subcategory ifharmful, and a free\-text reasoning span; the confidence rating is collected from human annotators only\. Full prompt text is in Appendix[B](https://arxiv.org/html/2608.17583#A2)\.

### 3\.4Stage\-2 Sample Selection

Stage 2 applies the Stage 1 winning auditor, Gemini 2\.5 Flash, to a stratified sample of the36,97136\{,\}971\-video corpus\. We first draw10%10\\%of videos separately within each country, age persona, and phase, where the phases are passive For\-You\-page scrolling and the three active\-session steps: before search, search results, and after search\. Because the initial draw contained too few before\-search and after\-search videos for a stable search\-versus\-baseline comparison, we top up these two active sub\-phases from the same accounts and collection window until each country–age combination contains roughly100100videos per phase\. The resulting Stage\-2 sample contains1,8171\{,\}817passive videos and4,2514\{,\}251active\-session videos; after removing unavailable videos, provider refusals, and parse failures, the final sample yields5,2275\{,\}227usable E3 verdicts and5,2485\{,\}248usable E2 verdicts, with5,1085\{,\}108paired videos\.

Because the passive and active collections were gathered in separate time windows, we analyze them separately throughout §[4](https://arxiv.org/html/2608.17583#S4)\. The main findings also hold on the original10%10\\%draw\. We omit E1 at Stage 2 because text\-only input performed poorly in Stage 1\.

## 4Results

### 4\.1MLLM Validation \(Stage\-1\)

We first evaluate the accuracy of different models at identifying harmful vs\. non\-harmful content\. Across the ten Stage\-1 \(model, condition\) combinations \(Fig\.[3](https://arxiv.org/html/2608.17583#S4.F3)\), Gemini 2\.5 Flash under the eight\-frame condition \(E3\) is the strongest configuration overall, but even it tops out at aggregate Cohen’sκ=0\.42\\kappa=0\.42and no \(model, condition\) clears moderate agreement\([16](https://arxiv.org/html/2608.17583#bib.bib26)\)in every country, so the task is hard\. For Gemini E3, Italian content reaches moderate agreement while French and Swedish stay at fair agreement; the cross\-countryκ\\kapparanking tracks the cross\-country class\-balance ranking in the reference subset \(Fig\.[2](https://arxiv.org/html/2608.17583#S3.F2)\), which mechanically suppressesκ\\kappaon the most\-imbalanced country\. Within Gemini specifically, agreement improves as more visual input is added \(E1κ=0\.17→\\kappa=0\.17\\toE2κ=0\.38→\\kappa=0\.38\\toE3κ=0\.42\\kappa=0\.42; Appendix[E](https://arxiv.org/html/2608.17583#A5), Figure[12](https://arxiv.org/html/2608.17583#A5.F12)\), and the eight\-frame condition outperforms native\-video upload at approximately2×2\\timeslower per\-call cost\. All non\-Gemini configurations sit belowκ=0\.30\\kappa=0\.30in the aggregate, but the per\-country profile is uneven: GPT\-4o\-mini’s eight\-frame run is the strongest non\-Gemini combination at aggregateκ=0\.29\\kappa=0\.29and reaches moderate agreement on Italian content, so the model ranking re\-orders by country\. We therefore carry Gemini’s two top configurations \(E3 as primary, E2 alongside for robustness\) into Stage\-2 and drop the text\-only E1 pathway, since no model clears fair agreement on text alone\. The remaining diagnostic plots \(aggregate grid, inter\-model agreement, confusion flow, per\-country category mix\) and the failure\-counts table are in Appendix[E](https://arxiv.org/html/2608.17583#A5)\.

E2 reports harm rates33–1212pp higher than E3 across the same combinations \(e\.g\. FR\-16:35\.1%35\.1\\%vs\.28\.1%28\.1\\%; SW\-13:31\.3%31\.3\\%vs\.24\.2%24\.2\\%\)\. The two modalities agree on the binary harm verdict atκ=0\.614\\kappa=0\.614overn=5,108n=5\{,\}108paired items \(Table[10](https://arxiv.org/html/2608.17583#A4.T10)\), with per\-countryκ∈\[0\.60,0\.63\]\\kappa\\in\[0\.60,0\.63\], so the cross\-combination ordering is preserved\. We treat E3 as the primary reference because Stage\-1 validated it against the annotator; E2’s higher reporting rate is consistent with a more permissive judgment layer when the full clip is available, but at this scale we cannot resolve it without a Stage\-2 annotation pass\. The E2\-vs\-E3 disagreement is not uniform across harm categories: E2 over\-flags visual\-cue categories \(Sexually Suggestive, Shocking and Graphic, Nudity\) while E3 picks up more dialogic Harassment and Bullying \(per\-category decomposition in Appendix[D](https://arxiv.org/html/2608.17583#A4), Table[11](https://arxiv.org/html/2608.17583#A4.T11)\)\.

Figure 3:Per\-country Cohen’sκ\\kappaon the binary harm verdict for the ten Stage\-1 \(model, condition\) combinations \(E1: text\-only; E2: native video; E3: eight frames plus text\)\. Dotted lines:κ=0\.21\\kappa=0\.21andκ=0\.41\\kappa=0\.41\([16](https://arxiv.org/html/2608.17583#bib.bib26)\)\.
### 4\.2Cross\-Country, Cross\-Age Harm Prevalence

All rates in §[4\.2](https://arxiv.org/html/2608.17583#S4.SS2)–§[4\.4](https://arxiv.org/html/2608.17583#S4.SS4)are estimates under the Stage\-1\-validated Gemini E3 auditor, and denominators include every sampled video served to the persona, regardless of detected content language\.Italy carries the highest estimated harm rate across all four age personas, with IT\-16 the most\-exposed case in the audit at40\.4%40\.4\\%\. France ranges from23\.6%23\.6\\%to32\.4%32\.4\\%; Sweden from24\.2%24\.2\\%to31\.0%31\.0\\%\. A per\-country precision/recall correction derived from Stage\-1 \(full derivation in Appendix[E](https://arxiv.org/html/2608.17583#A5)\) leaves this ordering intact at ages 13, 16, and 19, but at age 40 the corrected French rate \(≈49%\\approx 49\\%\) overtakes the corrected Italian \(≈38%\\approx 38\\%\) and Swedish \(≈33%\\approx 33\\%\) rates, so we read the age\-40 comparison with more caution than the three younger ages\. Within each country, harm rates are relatively flat across the four age personas \(a77–99pp spread in France and Sweden\); the exception is Italy, where the1313\-year\-old persona sees more harm than the adult, inverting the youth\-safety\-first prior\. E2 \(native video\) reports systematically higher harm rates than E3 \(eight frames plus text\) at every \(country, age\) combination \(per\-combination breakdown with E2 alongside in Appendix[D](https://arxiv.org/html/2608.17583#A4), Figure[10](https://arxiv.org/html/2608.17583#A4.F10); modality comparison in §[4\.1](https://arxiv.org/html/2608.17583#S4.SS1)\)\.

Sexually suggestive content is the dominant harm category in every \(country, age\) combination, from21\.7%21\.7\\%of harmful items \(SW\-13\) to50\.0%50\.0\\%\(IT\-16\); the per\-country breakdown of E3\-flagged items by harm subcategory is shown in Figure[4](https://arxiv.org/html/2608.17583#S4.F4), with the caveat that Stage\-1 strict primary\-subcategory agreement is only29%29\\%\(54%54\\%under primary\-or\-secondary matching; Appendix[E](https://arxiv.org/html/2608.17583#A5)\)\.

Figure 4:Country→\\toharm\-subcategory aggregate Sankey over Gemini E3 flagged items at Stage\-2\. Ribbon widths are proportional to absolute counts\.
### 4\.3Passive\-Phase Exposure

Under passive FYP scrolling \(no search probe\), Italy carries the highest harm rate at every age \(Fig\.[5](https://arxiv.org/html/2608.17583#S4.F5)\), with the cross\-country gap widening from a44\-point spread at age 13 \(FR24\.5%24\.5\\%, IT27\.9%27\.9\\%, SW23\.7%23\.7\\%\) to over2525points at age 19 \(FR22\.5%22\.5\\%, IT48\.6%48\.6\\%, SW38\.4%38\.4\\%\); Italian age\-19 is the most\-exposed case at48\.6%48\.6\\%\. France stays flat at∼23%\{\\sim\}23\\%across all four ages, while Sweden rises from24%24\\%\(age 13\) to38%38\\%\(age 19\)\. The E2 modality reports a few percentage points higher in every combination but preserves the same cross\-country ordering \(Appendix[D](https://arxiv.org/html/2608.17583#A4), Figure[11](https://arxiv.org/html/2608.17583#A4.F11)\)\. Because the passive phase and the active phase were collected in separate windows, we analyze them separately rather than as parallel signals at one point in time \(§[4\.4](https://arxiv.org/html/2608.17583#S4.SS4)\); Italy’s lead holds in both analyses\.

Figure 5:Stage\-2 harm rate per \(country, age\) on the passive collection phase \(FYP scrolling only, no search probe\) under Gemini E3, with95%95\\%video\-level bootstrap CIs\. Italy leads at every age and the cross\-country gap widens with age\. E2 version in Appendix[D](https://arxiv.org/html/2608.17583#A4), Figure[11](https://arxiv.org/html/2608.17583#A4.F11)\.
### 4\.4Search\-Phase Exposure

The within\-accountscroll\-pre→\\toSEARCH→\\toscroll\-postcycle measures how the algorithm responds to a harm\-keyword search and whether the probe contaminates the scroll\-post FYP \(Fig\.[6](https://arxiv.org/html/2608.17583#S4.F6)\)\. TheSEARCHendpoint returns3535–56%56\\%harmful content \(1\.51\.5–7\.5×7\.5\\timesoverscroll\-prein ten of twelve combinations\) despite no visible search\-time block or warning in our audit, whilescroll\-postreverts to within a few percentage points ofscroll\-pre\. A higher harm rate under harm\-keyword search than under scrolling is expected by construction; the findings are its magnitude, the collapse of the age gradient, and the absence of visible search\-time intervention, not evidence of a general moderation failure\.

The largest lifts are on older French and Italian personas, whosescroll\-prerate sits below25%25\\%and whoseSEARCHendpoint returns5050–56%56\\%harmful items before falling back; Sweden’s adult combination shows the same pattern at smaller magnitude\. The two exceptions are IT\-13 and IT\-16, wherescroll\-preis already at3434–44%44\\%harm: the harm\-keyword search has no headroom to lift further, andscroll\-poststays in range ofscroll\-pre\. The contrast with the same\-age French and Swedish combinations, which sit at99–19%19\\%pre\-probe, suggests that country and not persona age sets the headroom on this audit window\. A related observation is that, at theSEARCHendpoint, harm rates collapse into a3535–56%56\\%band across all twelve combinations: the youngest personas in France and Sweden reach3737–42%42\\%harm under search, within a few percentage points of the adult combinations in the same countries, so thescroll\-prebaseline age gradient is wiped out by the probe\. A second pattern is the speed of the reversion\. Across all twelve combinations, thescroll\-preandscroll\-post% CIs overlap and the absolute change is bounded by a few percentage points\. Within the five\-day collection window the harm\-keyword search elevates exposure during the search session itself, not as a persistent recommendation\-feed drift\.

#### Account\-clustered uncertainty

Because each cell is instantiated by three accounts and recommendations are sequential, a video\-level bootstrap could understate uncertainty\. Recomputing every active\-phase interval with a hierarchical bootstrap that resamples accounts before videos leaves the intervals essentially unchanged \(median width ratio1\.001\.00, at most1\.85×1\.85\\timeson the smallestscroll\-postcells\), and per\-accountSEARCHharm rates within a cell differ by at most a few percentage points \(Appendix[D](https://arxiv.org/html/2608.17583#A4), Table[7](https://arxiv.org/html/2608.17583#A4.T7)\)\. The clustered intervals separateSEARCHfromscroll\-prein ten of twelve combinations \(all but the two Italian ceiling cells IT\-13 and IT\-16\) and preserve thescroll\-pre/scroll\-postoverlap in all twelve\. Per\-account attribution is available for the active phase only, so passive\-phase CIs \(§[4\.3](https://arxiv.org/html/2608.17583#S4.SS3)\) remain video\-level\.

#### Policy\-tier split of the age gradient

The binary verdict pools content TikTok prohibits for every audience with content it permits for adults but restricts for minors\. Splitting the 13 categories into an 18\+\-restricted tier \(Sexually Suggestive, Nudity, Alcohol/Tobacco/Drugs, Gambling, Shocking and Graphic\) and a universally prohibited tier separates the two readings \(Appendix[D](https://arxiv.org/html/2608.17583#A4), Table[8](https://arxiv.org/html/2608.17583#A4.T8)\)\. Under passive scrolling the restricted tier rises with persona age in Italy and Sweden \(IT17\.217\.2pp at age 13 vs\.26\.726\.7pp at 40; SW8\.68\.6vs\.24\.524\.5pp\), the direction age gating predicts, while the prohibited tier does not fall for minors \(SW\-13 carries the audit’s highest passive prohibited\-tier rate at15\.115\.1pp\)\. UnderSEARCHthe age separation on the restricted tier disappears: the 13\-year\-old personas receive2424–3535pp of 18\+\-restricted content, within a few points of the adult personas in the same countries\. We do not evaluate compliance with TikTok’s age policies as such; the split shows that the flat total\-rate age gradient mixes a rising restricted tier with a non\-declining prohibited tier rather than indicating uniform age\-blindness\.

Figure 6:Per\-combination harm rate at the three steps of the active collection phase \(scroll\-pre→\\toSEARCHkeyword probe→\\toscroll\-post\) on E3 with % CIs\. Annotated:SEARCH/scroll\-preratio; stars mark non\-overlapping CIs\.
#### Temporal stability of the audit window

A day\-by\-day breakdown of Gemini E3 harm rate per \(country, age, signal\) on the active sample \(Appendix[D](https://arxiv.org/html/2608.17583#A4), Figure[9](https://arxiv.org/html/2608.17583#A4.F9)\) shows theSEARCHlift and the Italianscroll\-preceiling effect hold on every day of the five\-day window; the cross\-combination averages in §[4\.2](https://arxiv.org/html/2608.17583#S4.SS2)–§[4\.4](https://arxiv.org/html/2608.17583#S4.SS4)are not artifacts of a single\-day spike\.

### 4\.5Provider Blocks and Failure Modes at Scale

The Gemini safety layer refuses to score about1\.1%1\.1\\%of Stage\-2 inputs \(per\-country breakdown in Appendix[D](https://arxiv.org/html/2608.17583#A4), Table[6](https://arxiv.org/html/2608.17583#A4.T6)\); a small additional set of empty responses and transient network errors accounts for the gap between the∼5,300\{\\sim\}5\{,\}300inputs per modality and the headline denominators\. The block rate is broadly balanced across countries and comparable across modalities \(E2:57/∼5,300≈1\.08%57/\{\\sim\}5\{,\}300\\approx 1\.08\\%; E3:64/∼5,300≈1\.21%64/\{\\sim\}5\{,\}300\\approx 1\.21\\%; Fisher exactp≈0\.55p\\approx 0\.55\)\. The blocks are not noise\-distributed across categories\. Only2323of the6464E3 refusals are surfaced through aSEARCHkeyword and therefore category\-attributable; the remaining4141sit inscroll\-pre,scroll\-post, or passive\-only paths, where no keyword fixes a harm category\. Among the2323SEARCH\-attributable blocks, the per\-category block rate is dominated by Nudity and Body Exposure \(≈4\.8%\\approx 4\.8\\%\), with Sexually Suggestive Content \(≈2\.6%\\approx 2\.6\\%\), Shocking and Graphic Content \(≈0\.7%\\approx 0\.7\\%\), and Disordered Eating and Body Image \(≈0\.5%\\approx 0\.5\\%\) next; the remaining three search\-targeted categories \(Dangerous Challenges, Gambling, Alcohol/Tobacco/Drugs\) yielded zero or near\-zero blocks\. Reported prevalences are therefore under\-estimates for exactly the categories the audit is designed to surface, with the largest measurement bias on Nudity and Suggestive content\. A worst\-case sensitivity bound \(assigning every block\-or\-parse\-fail item to eitherHarmfulorNot Harmful\) shifts headline rates by at most11–33pp and preserves both the cross\-country IT\>\>SW≥\\geqFR ordering and the IT\-13/16 ceiling\-effect finding \(Appendix[D](https://arxiv.org/html/2608.17583#A4)\)\.

## 5Discussion and Conclusion

#### MLLM auditing is feasible at this scale, conditional on three structural caveats

First, the Stage\-1κ=0\.42\\kappa=0\.42was measured on the 300\-video reference sample, and whether it transfers to the full Stage\-2 distribution is an open question this design cannot answer\. Second, at scale E2 \(native video\) reports consistently higher harm rates than primary E3 \(eight frames plus text\) \(\+3\+3–1212pp\), and which modality is closer to human truth on the Stage\-2 population is not answerable without a Stage\-2 annotation pass\. Third, the provider safety layer refuses1\.1%1\.1\\%of inputs non\-randomly with respect to harm category \(§[4\.5](https://arxiv.org/html/2608.17583#S4.SS5)\), so reported prevalences are under\-estimates for the most explicit content\. MLLM auditing is therefore cheap enough at full corpus scale under these assumptions, but not yet a drop\-in replacement for human annotation on policy\-edge items\.

#### Cross\-country variation is large, and the patterns differ

Under E3, Italy carries the highest measured harm rate across all four age personas, with the age\-40 case shifting under per\-country precision/recall recalibration \(Appendix[E](https://arxiv.org/html/2608.17583#A5)\) but the strongest reading holding on the three youngest, and the within\-country age gradients run in opposite directions across the three countries\. The same Italy\-leads pattern replicates under purely passive FYP scrolling \(§[4\.3](https://arxiv.org/html/2608.17583#S4.SS3)\), with Italian age\-19 reaching48\.6%48\.6\\%harm, the highest measured rate in the audit, against∼23%\{\\sim\}23\\%in France at the same age\. The two youngest Italian combinations \(IT\-13, IT\-16\) are already at3434–44%44\\%harm atscroll\-pre, while the same age combinations in France and Sweden sit far lower and respond strongly to the probe\. Three decompositions narrow the candidate explanations for the Italian pattern \(Appendix[D](https://arxiv.org/html/2608.17583#A4)\)\. The lead is concentrated in one category: Sexually Suggestive content contributes23\.823\.8pp of Italy’s39\.0%39\.0\\%passive rate, against12\.212\.2pp in France and15\.815\.8pp in Sweden\. It is not carried by globally circulating videos: on passive items that also appear in another country’s corpus, Italy’s estimated rate is19\.2%19\.2\\%, indistinguishable from France \(26\.4%26\.4\\%\) and Sweden \(28\.3%28\.3\\%\), while Italy\-exclusive content sits at41\.7%41\.7\\%\. It also survives a language control: English\-language videos served to Italian accounts are flagged at32\.4%32\.4\\%, clearly above the English\-language rate in France \(19\.5%19\.5\\%\), with Sweden in between \(27\.5%27\.5\\%\)\. The pattern therefore points to the country\-specific slice of the pool TikTok serves to Italian accounts rather than to translation artifacts or annotator thresholds; whether that slice reflects content supply or per\-language moderation capacity\([23](https://arxiv.org/html/2608.17583#bib.bib20)\)is not identifiable from the outside\.

#### The harm\-keyword search endpoint is the main driver of exposure, and the elevated exposure is short\-lived

The1\.51\.5–7\.5×7\.5\\timeswithin\-accountscroll\-pre→\\toSEARCHlift, paired with near\-completescroll\-postreversion \(§[4\.4](https://arxiv.org/html/2608.17583#S4.SS4)\), means a passive\-FYP\-only audit understates the harm a determined user can reach by keyword by an order of magnitude in many combinations; this is search returning what the keyword asks for in the absence of visible search\-time moderation, not recommender\-side amplification in the sense of[5](https://arxiv.org/html/2608.17583#bib.bib11)\. The FR/SW scroll\-pre age gradient is absent atSEARCH, but our design does not distinguish \(a\) age\-insensitive platform retrieval from \(b\) a keyword set whose returned content pool overlaps across personas; the same 21 keywords are issued by every persona, and an item\-overlap analysis within \(country, keyword\) would discriminate between these two readings\.

#### Conclusion

Two findings stand out: keyword search, not recommender amplification, is the dominant harm\-exposure pathway in this audit, returning3535–56%56\\%harmful content for harm\-seeking queries with no visible search\-time friction, while baseline algorithmic exposure remains sharply uneven across countries, with accounts registered in Italy the most exposed at every age\. Methodologically, MLLM\-based auditing scales cross\-national youth\-safety audits at API costs of order $50, with annotation throughput rather than compute cost as the remaining scaling bottleneck\. These results point to platform\-side search\-time moderation and per\-language moderation infrastructure as the natural next levers for reducing harm exposure on short\-video platforms\.

## 6Limitations

Two native\-speaker annotators per country produced the Stage\-1 final reference labels via joint resolution; these are not statistical ground truth \(pre\-resolution inter\-annotator agreement in Appendix[E](https://arxiv.org/html/2608.17583#A5);[4](https://arxiv.org/html/2608.17583#bib.bib24);[14](https://arxiv.org/html/2608.17583#bib.bib25)\)\. The annotators apply the public text of the Community Guidelines and are not professional content moderators; TikTok’s internal enforcement thresholds are not observable, so the reference labels are guideline\-grounded judgments rather than platform enforcement truth\.

Stage\-2 runs a single auditor model with no per\-video annotation pass; the E2\-vs\-E3 cross\-modality agreement is an internal\-consistency check rather than a second validation against reference labels\. Reported Stage\-2 harm rates should be read as Gemini E3\-as\-auditor estimates whose transfer from the 300\-video Stage\-1 reference to the full Stage\-2 distribution is unverified; the reference draw is stratified by country and age but not by phase, so phase\-dependent auditor error cannot be ruled out\.

Provider\-side refusals \(∼1\.1%\{\\sim\}1\.1\\%at Stage\-2, plus Stage\-1 Qwen/DashScope failures in Appendix[E](https://arxiv.org/html/2608.17583#A5)\) cluster on the most\-explicit harm categories; the worst\-case sensitivity bound is in §[4\.5](https://arxiv.org/html/2608.17583#S4.SS5)\.

Our data collection spans a single, time\-bounded window, and TikTok’s recommender adapts continuously, so exposure patterns measured here may not generalize to other time periods or to persona profiles we did not instantiate\. The personas themselves are programmatically controlled accounts that lack the behavioral richness of real users \(no multi\-device use, no cross\-session continuity beyond what we script, no organic social graph\), a standard caveat of sockpuppet auditing\([20](https://arxiv.org/html/2608.17583#bib.bib13)\)\. Findings should therefore be read as upper\-bound claims about what the algorithm can serve under simple engagement rules, not as point estimates of real\-user exposure\. The VPN\-routed sockpuppet accounts may also have been handled by TikTok’s anti\-abuse layer differently from native\-IP traffic in ways we did not separately measure; the absence of mass capture failures suggests this was not severe but does not bound the residual effect\.

#### Ethics and data release

All collected videos are from public TikTok accounts and no real users are impersonated\. Following the platform’s terms of service, we plan to release: per\-video metadata for the full36,97136\{,\}971\-video corpus \(video IDs, country, age persona, phase, keyword, capture timestamp, public engagement counts\), the per\-country keyword lists \(Appendix[A](https://arxiv.org/html/2608.17583#A1)\), the per\-experiment prompt templates \(Appendix[B](https://arxiv.org/html/2608.17583#A2)\), the Gemini E3 and E2 verdicts and reasoning spans for the Stage\-2 subset, and the aggregated statistics reported in this paper\. We do*not*redistribute raw video content, raw API payloads, or any information identifying public TikTok users beyond the video ID; videos that have since been deleted, made private, or geo\-restricted are released as IDs only\. The Stage\-1 annotations are released in aggregate \(per\-combination agreement summaries\) rather than per\-video, to avoid re\-identifying the two annotators per country\. The collection, evaluation, and analysis code \(scrapers, MLLM\-evaluation scripts, plotting and statistics pipeline\) is released alongside the data\. AI assistants were used for writing improvements and editing the text\.

## References

- Amnesty International \(2023\)Amnesty InternationalDriven into the darkness: how TikTok encourages self\-harm and suicidal ideation\.Technical reportAmnesty International\.External Links:[Link](https://www.amnesty.org/en/latest/news/2023/11/tiktok-risks-pushing-children-towards-harmful-content/)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p1.1)\.
- Amnesty International \(2025\)Amnesty InternationalDragged into the rabbit hole: new evidence of TikTok’s risks to children’s mental health\.Technical reportAmnesty International\.External Links:[Link](https://www.amnesty.org/en/wp-content/uploads/2025/10/POL4003602025ENGLISH.pdf)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p1.1)\.
- Apple Machine Learning Research \(2025\)Apple Machine Learning ResearchBreaking down video LLM benchmarks: knowledge, spatial perception, or true temporal understanding?\.InNeurIPS 2025 LLM Evaluation Workshop,External Links:[Link](https://machinelearning.apple.com/research/breaking-down)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p2.1),[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2608.17583#S3.SS3.SSS0.Px2.p1.1)\.
- Artstein and Poesio \(2008\)R\. Artstein and M\. PoesioInter\-coder agreement for computational linguistics\.Computational Linguistics34\(4\),pp\. 555–596\.External Links:[Document](https://dx.doi.org/10.1162/coli.07-034-R2)Cited by:[§6](https://arxiv.org/html/2608.17583#S6.p1.1)\.
- Baumannet al\.\(2025\)F\. Baumann, N\. Arora, I\. Rahwan, and A\. CzaplickaDynamics of algorithmic content amplification on TikTok\.External Links:2503\.20231,[Link](https://arxiv.org/abs/2503.20231)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p1.1),[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.17583#S3.SS1.p3.1),[§5](https://arxiv.org/html/2608.17583#S5.SS0.SSS0.Px3.p1.1)\.
- Blackburn and Hogg \(2024\)M\. Blackburn and E\. J\. Hogg\#ForYou? the impact of pro\-ana TikTok content on body image dissatisfaction and internalisation of societal beauty standards\.PLOS ONE19\(8\),pp\. e0307597\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0307597)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p1.1)\.
- Boeker and Urman \(2022\)M\. Boeker and A\. UrmanAn empirical investigation of personalization factors on TikTok\.InProceedings of the ACM Web Conference 2022,pp\. 2298–2309\.External Links:[Document](https://dx.doi.org/10.1145/3485447.3512102)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p1.1),[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.17583#S3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2608.17583#S3.SS1.p3.1)\.
- Davidsonet al\.\(2025\)T\. Davidsonet al\.Multimodal large language models can make context\-sensitive hate speech evaluations aligned with human judgement\.Nature Human Behaviour\.External Links:[Link](https://www.nature.com/articles/s41562-025-02360-w)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p2.1),[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px2.p1.1)\.
- Eltaheret al\.\(2025\)F\. Eltaher, R\. K\. Gajula, L\. Miralles\-Pechán, P\. Crotty, J\. Martínez\-Otero, C\. Thorpe, and S\. McKeeverProtecting young users on social media: evaluating the effectiveness of content moderation and legal safeguards on video\-sharing platforms\.External Links:2505\.11160,[Link](https://arxiv.org/abs/2505.11160)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p1.1),[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.17583#S3.SS1.p1.1)\.
- Fasching and Lelkes \(2025\)N\. Fasching and Y\. LelkesModel\-dependent moderation: inconsistencies in hate speech detection across LLM\-based systems\.InFindings of the Association for Computational Linguistics: ACL 2025,External Links:[Link](https://aclanthology.org/2025.findings-acl.1144/)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p2.1),[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2608.17583#S3.SS3.SSS0.Px1.p1.1)\.
- Huang \(2024\)T\. HuangContent moderation by LLM: from accuracy to legitimacy\.Note:arXiv:2409\.03219External Links:[Link](https://arxiv.org/abs/2409.03219)Cited by:[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px2.p1.1)\.
- Ibrahimet al\.\(2026\)H\. Ibrahim, H\. D\. Jang, N\. Aldahoul, A\. R\. Kaufman, T\. Rahwan, and Y\. ZakiSystematic partisan content skews in TikTok during the 2024 US elections\.Nature\.External Links:[Document](https://dx.doi.org/10.1038/s41586-026-10447-1),[Link](https://www.nature.com/articles/s41586-026-10447-1)Cited by:[§3\.1](https://arxiv.org/html/2608.17583#S3.SS1.p1.1)\.
- Joet al\.\(2024\)C\. W\. Jo, M\. Wesołowska, and M\. WojcieszakHarmful YouTube video detection: a taxonomy of online harm and MLLMs as alternative annotators\.External Links:2411\.05854,[Link](https://arxiv.org/abs/2411.05854)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p2.1),[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px3.p1.1),[§3\.2](https://arxiv.org/html/2608.17583#S3.SS2.SSS0.Px1.p1.1),[§3\.3](https://arxiv.org/html/2608.17583#S3.SS3.SSS0.Px2.p1.1)\.
- Krippendorff \(2011\)K\. KrippendorffComputing Krippendorff’s alpha\-reliability\.External Links:[Link](https://repository.upenn.edu/asc_papers/43/)Cited by:[§6](https://arxiv.org/html/2608.17583#S6.p1.1)\.
- Kumaret al\.\(2024\)D\. Kumar, Y\. AbuHashem, and Z\. DurumericWatch your language: investigating content moderation with large language models\.InProceedings of the International AAAI Conference on Web and Social Media \(ICWSM\),Vol\.18,pp\. 865–878\.External Links:[Link](https://ojs.aaai.org/index.php/ICWSM/article/view/31358)Cited by:[Appendix E](https://arxiv.org/html/2608.17583#A5.SS0.SSS0.Px7.p1.1)\.
- Landis and Koch \(1977\)J\. R\. Landis and G\. G\. KochThe measurement of observer agreement for categorical data\.Biometrics33\(1\),pp\. 159–174\.External Links:[Document](https://dx.doi.org/10.2307/2529310)Cited by:[Figure 14](https://arxiv.org/html/2608.17583#A5.F14),[Figure 3](https://arxiv.org/html/2608.17583#S4.F3),[§4\.1](https://arxiv.org/html/2608.17583#S4.SS1.p1.1)\.
- Leet al\.\(2025\)H\. Le, S\. Elmalaki, Z\. Shafiq, and A\. MarkopoulouAutoLike: auditing social media recommendations through user interactions\.External Links:2502\.08933,[Link](https://arxiv.org/abs/2502.08933)Cited by:[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px1.p1.1)\.
- Minadeo and Pope \(2022\)M\. Minadeo and L\. PopeWeight\-normative messaging predominates on TikTok—a qualitative content analysis\.PLOS ONE17\(11\),pp\. e0267997\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0267997)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p1.1),[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px1.p1.1)\.
- Murthy \(2023\)V\. H\. MurthySocial media and youth mental health: the U\.S\. surgeon general’s advisory\.Technical reportU\.S\. Department of Health and Human Services\.External Links:[Link](https://www.hhs.gov/sites/default/files/sg-youth-mental-health-social-media-advisory.pdf)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p1.1)\.
- Sandviget al\.\(2014\)C\. Sandvig, K\. Hamilton, K\. Karahalios, and C\. LangbortAuditing algorithms: research methods for detecting discrimination on internet platforms\.InData and Discrimination: Converting Critical Concerns into Productive Inquiry, 64th Annual Meeting of the International Communication Association,External Links:[Link](https://websites.umich.edu/%CB%9Ccsandvig/research/Auditing%20Algorithms%20--%20Sandvig%20--%20ICA%202014%20Data%20and%20Discrimination%20Preconference.pdf)Cited by:[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.17583#S3.SS1.p1.1),[§6](https://arxiv.org/html/2608.17583#S6.p4.1)\.
- Stricklandet al\.\(2025\)S\. R\. Strickland, A\. Medina Fernandez, and P\. K\. KeelTikTok and disordered eating: delineating temporal associations and effects of a ban\.Eating Behaviors\.External Links:[Link](https://www.sciencedirect.com/science/article/abs/pii/S1471015325000893)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p1.1)\.
- TikTok \(2025\)TikTokCommunity guidelines\.Note:Online policy document, effective September 13, 2025External Links:[Link](https://www.tiktok.com/safety/en/policies-and-engagement/overview)Cited by:[3rd item](https://arxiv.org/html/2608.17583#S1.I2.i3.p1.1),[§3\.2](https://arxiv.org/html/2608.17583#S3.SS2.SSS0.Px1.p1.1)\.
- Tonneauet al\.\(2025\)M\. Tonneau, D\. Liu, R\. McGrady, K\. Zheng, R\. Schroeder, E\. Zuckerman, and S\. A\. HaleLanguage disparities in moderation workforce allocation by social media platforms\.Note:SocArXiv preprintExternal Links:[Document](https://dx.doi.org/10.31235/osf.io/amfws%5Fv1),[Link](https://osf.io/preprints/socarxiv/amfws)Cited by:[Appendix E](https://arxiv.org/html/2608.17583#A5.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2608.17583#S3.SS1.p2.1),[§5](https://arxiv.org/html/2608.17583#S5.SS0.SSS0.Px2.p1.1)\.
- World Economic Forum \(2023\)World Economic ForumToolkit for digital safety design interventions and innovations: typology of online harms\.Technical reportWorld Economic Forum, Global Coalition for Digital Safety\.External Links:[Link](https://www3.weforum.org/docs/WEF_Typology_of_Online_Harms_2023.pdf)Cited by:[§3\.2](https://arxiv.org/html/2608.17583#S3.SS2.SSS0.Px1.p1.1)\.
- Wuet al\.\(2025\)M\. Wuet al\.ICM\-assistant: instruction\-tuning multimodal large language models for rule\-based explainable image content moderation\.InProceedings of the AAAI Conference on Artificial Intelligence,Note:arXiv:2412\.18216External Links:[Link](https://arxiv.org/abs/2412.18216)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p2.1),[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px3.p1.1),[§3\.3](https://arxiv.org/html/2608.17583#S3.SS3.SSS0.Px3.p1.1)\.
- Xueet al\.\(2025\)L\. Xue, F\. Corso, N\. Fontana, G\. Liu, S\. Ceri, and F\. PierriTowards an automated framework to audit youth safety on TikTok\.InProceedings of the Fourth Workshop on Bridging Human–Computer Interaction and Natural Language Processing \(HCI\+NLP\),Suzhou, China\.External Links:[Link](https://aclanthology.org/2025.hcinlp-1.9/)Cited by:[§1](https://arxiv.org/html/2608.17583#S1.p1.1),[§2](https://arxiv.org/html/2608.17583#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.17583#S3.SS1.p1.1)\.

## Appendix AKeyword Lists

The active\-phaseSEARCHprobe uses three native\-language keywords per country and per harm category, listed in Table[2](https://arxiv.org/html/2608.17583#A1.T2)\. Categories without a row were not probed \(the remaining six taxonomy categories from §[3\.2](https://arxiv.org/html/2608.17583#S3.SS2)are not target\-able by short keyword queries with reasonable precision\)\. The same English category labels are used as the standard harm\-category names throughout the paper; the harm\-keyword search issues each keyword in the persona country’s native language\.

#### Keyword selection procedure

The English seed terms for each of the seven probed harm categories were drafted directly from the taxonomy definitions in §[3\.2](https://arxiv.org/html/2608.17583#S3.SS2), with three queries per category covering one literal severity\-axis anchor \(e\.g\.blood,nudity,casino\), one moderate variant \(e\.g\.fights,body exposure,betting\), and one platform\-vernacular phrase \(e\.g\.natural disaster,intimate kisses,sports betting\) chosen to surface content that TikTok’s search endpoint actually returns for the category rather than zero\-result moderation\-jargon strings\. Each English triple was then translated by a native\-speaker researcher in Italian, French, and Swedish, with the constraint that the per\-country translation should preserve the same severity\-axis structure \(anchor / moderate / vernacular\) rather than the literal lexical form, so that the cross\-countrySEARCH/scroll\-preratio of §[4\.4](https://arxiv.org/html/2608.17583#S4.SS4)is read against queries of comparable retrieval intent\. The full keyword list was*fixed before any active\-phase data collection began*and is reported here unaltered; no post\-hoc keyword tuning was applied after observing theSEARCHharm rates\. The textual coverage check in Appendix[D](https://arxiv.org/html/2608.17583#A4)\(Figure[8](https://arxiv.org/html/2608.17583#A4.F8)\) confirms that the chosen keywords surface keyword\-bearing content at a2020–25%25\\%rate underSEARCHagainst a00–2%2\\%FYP baseline\.

Table 2:Per\-country, per\-category native\-language keywords used in the active\-phaseSEARCHprobe \(§[3\.1](https://arxiv.org/html/2608.17583#S3.SS1)\)\. Each persona account issues all 21 keywords in its country’s language across the five\-day collection window; the vocabulary was fixed before data collection\.

## Appendix BPrompt Templates

This appendix gives the full prompts used at both stages\. The same system prompt is shared across all four MLLMs and across the three input conditions; only the user\-message structure changes by condition\.

#### System prompt \(all conditions, all models\)

> Classify the following TikTok video as ‘harmful’ or ‘not harmful’ under this harmful content framework \(aligned with TikTok Community Guidelines\): - •Disordered Eating and Body Image: Content that promotes disordered eating, risky weight loss or muscle gain methods, or harmful body comparisons\. - •Suicide and Self\-Harm: Content that shows, promotes, or provides instructions for suicide or self\-harm\. - •Dangerous Activity and Challenges: Content that shows or promotes dangerous stunts, dares, or challenges that could lead to physical harm\. - •Nudity and Body Exposure: Nudity, sexual activity, or significant body exposure; any such content involving minors is prohibited\. - •Sexually Suggestive Content: Sexually suggestive acts or behavior, or sexually explicit language; any such content involving minors is not allowed\. - •Shocking and Graphic Content: Extremely graphic, violent, or disturbing content that could cause viewers emotional distress\. - •Hate Speech or Behavior: Content that promotes hate or attacks people based on protected attributes\. - •Sexual and Physical Abuse: Content that shows, promotes, or facilitates sexual abuse, exploitation, or physical harm of people \(including youth\)\. - •Human Trafficking and Smuggling: Content that promotes or facilitates human trafficking or smuggling\. - •Gambling: Content that promotes or facilitates gambling or gambling\-like activities \(e\.g\. social casinos, mystery boxes\)\. - •Alcohol, Tobacco, and Drugs: Content that trades, markets, or promotes alcohol, tobacco, vaping, e\-cigarettes, or illegal or regulated drugs\. - •Integrity and Authenticity: Deceptive behavior or fake engagement; misleading edited or AI\-generated content; or misinformation that could cause significant harm\. - •Harassment and Bullying: Content that harasses or bullies others, including degrading remarks about appearance, doxing, sexual harassment, or coordinated abuse\. Watch the video carefully \(visual content, audio, on\-screen text\)\. Also consider the description and, if provided, the comments\. The goal is to prevent minors from seeing harmful content\. Reply with a single JSON object and nothing else: ``` { "verdict": "harmful" or "not harmful", "subcategory": "<one of the 13 categories above, or null if not harmful>", "reasoning": "<your brief reasoning>" } ``` If the video falls under one or more of the categories above in a way that violates the definitions, choose ‘‘harmful’’\. Otherwise choose ‘‘not harmful’’\.

#### User message per condition

The user message contains the video content \(in the form appropriate for each condition\) plus a short text block\. The text block is:

- •E1 \(text\-only\):“Video description: \{description\}” \+ “Audio transcript:\\n\{transcript\}” \+ “Classify this video according to the instructions\.” No visual content\.
- •E2 \(native video\):the raw MP4 bytes uploaded through the provider’s video API, plus “Note: no description is provided – evaluate the video content only\. Classify this video according to the instructions\.”
- •E3 \(frames \+ text\):“Video description: \{description\}” \+ “The following 8 image\(s\) are frames extracted at equal intervals from a TikTok video\. Analyze them together as a representation of the full video content\.” \+ the eight frames as inline base64 PNGs \+ “Audio transcript:\\n\{transcript\}” \+ “Classify this video according to the instructions\.”

The same prompt template is used at Stage\-1 \(four\-model comparison\) and Stage\-2 \(Gemini E3 run\) so that the Stage\-1κ\\kappacalibration transfers without prompt\-induced drift\.

## Appendix CDataset Statistics

This appendix provides the detailed dataset breakdowns summarized in §[3\.1](https://arxiv.org/html/2608.17583#S3.SS1)\.

Table 3:Per\-\(country, age\) capture counts by phase and unique\-video yield, with mean plays and likes over all items in the combination\.#### Phase\-level engagement gap

Across all three countries, For\-You\-feed items surface substantially higher play counts than search\-surfaced items, typically by an order of magnitude on the median \(Figure[7](https://arxiv.org/html/2608.17583#A3.F7)\)\. We attribute this to the different economic logics of the two sub\-systems: the For\-You feed optimizes for watch\-time virality, while search retrieves a long\-tail query\-matched pool\. The effect is stable across ages and shows up at bothscroll\-preandscroll\-post, suggesting it is not an artifact of the probe\. The full passive phase, collected on the same accounts at an earlier window, mirrors the active\-phase scrolling pattern: mean play counts are again in the millions per video \(Table[4](https://arxiv.org/html/2608.17583#A3.T4)\)\.

Table 4:Passive\-phase engagement statistics per \(country, age\) combination, over all item occurrences\. Companion to Table[3](https://arxiv.org/html/2608.17583#A3.T3)\(active phase\)\.Table 5:Detected\-language distribution of unique videos per phase and persona country\. “Native” is French for FR, Italian for IT, Swedish for SW; “Other” aggregates remaininglangdetectlabels; “Unknown” marks captions under 8 characters or unscoreable\.Figure 7:Distribution of play counts per video by \(country, age, phase\) in the active dataset \(log scale, one playcount per unique video, outliers suppressed\)\. Thescroll\-preandscroll\-postboxes sit 1–2 orders of magnitude above theSEARCHboxes in every combination\.

## Appendix DStage\-2 Supplementary Figures and Tables

This appendix collects the Stage\-2 figures and tables that are referenced from §[4\.2](https://arxiv.org/html/2608.17583#S4.SS2)–§[4\.5](https://arxiv.org/html/2608.17583#S4.SS5)but moved out of the main text for space\.

#### Provider\-block worst\-case bound

Counting every block\-or\-parse\-fail item \(96 of5,3235\{,\}323E3 inputs,1\.80%1\.80\\%\) as eitherHarmfulorNot Harmfulyields per\-country headline intervals of\[26\.5%,28\.4%\]\[26\.5\\%,28\.4\\%\]for France \(vs\. reported27\.0%27\.0\\%\),\[35\.0%,36\.5%\]\[35\.0\\%,36\.5\\%\]for Italy \(35\.5%35\.5\\%\), and\[27\.5%,29\.5%\]\[27\.5\\%,29\.5\\%\]for Sweden \(28\.0%28\.0\\%\); the IT\-16 case shifts at most from40\.4%40\.4\\%to\[39\.9%,41\.1%\]\[39\.9\\%,41\.1\\%\]\. The IT\>\>SW≥\\geqFR ordering and the IT\-13 / IT\-16 ceiling\-effect finding both survive this bound\.

Table 6:Stage\-2 GeminiPROHIBITED\_CONTENTcounts per country and modality, over∼5,300\{\\sim\}5\{,\}300inputs per modality\. Per\-category breakdown in §[4\.5](https://arxiv.org/html/2608.17583#S4.SS5)\.
#### Sample top\-up details

The initial10%10\\%phase\-stratified draw was3,8623\{,\}862items; after excluding294294unavailable MP4s,3,5683\{,\}568were fed to Gemini under both E2 and E3\. Thescroll\-preandscroll\-postcases held∼17\{\\sim\}17records each, versus∼150\{\\sim\}150forSEARCH, which was too thin to anchor the within\-accountSEARCH/scroll\-precomparison in §[4\.4](https://arxiv.org/html/2608.17583#S4.SS4)\. We therefore topped up both phases from the same accounts, days, and FYP snapshots until eachscroll\-preandscroll\-postcase held∼100\{\\sim\}100items, taking the final sample to5,2485\{,\}248usable E2 verdicts and5,2275\{,\}227E3 verdicts \(5,1085\{,\}108paired\)\.

#### Original\-draw ratio robustness

We recompute the §[4\.4](https://arxiv.org/html/2608.17583#S4.SS4)SEARCH/scroll\-preratios on the*original*10% phase\-stratified draw before the top\-up\. Across the ten combinations wherescroll\-preis positive in both draws, the original\-draw factor is1\.2×1\.2\\timesto7\.4×7\.4\\times\(vs\.1\.0×1\.0\\timesto8\.1×8\.1\\timestopped\-up\); IT\-19 and IT\-40 are undefined on the original draw because the smallnscroll\-pre∈\[15,17\]n\_\{\\text\{scroll\-pre\}\}\\in\[15,17\]produced zero harmful items, which is the noise regime the top\-up was designed to escape\.

#### Account\-clustered bootstrap intervals

Table[7](https://arxiv.org/html/2608.17583#A4.T7)reports the §[4\.4](https://arxiv.org/html/2608.17583#S4.SS4)active\-phase harm rates with95%95\\%CIs from a hierarchical bootstrap that resamples the three accounts per cell with replacement and then resamples videos within each drawn account, so that between\-account correlation is preserved\. Account attribution covers99\.399\.3–99\.8%99\.8\\%of Stage\-2 active records per phase; a video captured by two accounts contributes to both accounts’ pools\. The clustered intervals match the video\-level ones closely \(median width ratio1\.001\.00across the 36 E3 cells, maximum1\.85×1\.85\\times\), reflecting the low between\-account variability visible in the per\-accountSEARCHrate ranges\.

Table 7:Gemini E3 harm rates with95%95\\%account\-clustered hierarchical\-bootstrap CIs, per active\-phase cell\.∗marksSEARCHintervals disjoint from the cell’sscroll\-preinterval\. The last column is the range of per\-accountSEARCHharm rates within the cell\.
#### Policy\-tier split

Table[8](https://arxiv.org/html/2608.17583#A4.T8)decomposes each cell’s harm rate into the 18\+\-restricted tier \(Sexually Suggestive, Nudity and Body Exposure, Alcohol/Tobacco/Drugs, Gambling, Shocking and Graphic\) and the universally prohibited tier \(the remaining eight categories\), as discussed in §[4\.4](https://arxiv.org/html/2608.17583#S4.SS4)\. The Disordered Eating placement is arguable \(promotion is prohibited, generic weight\-management content is 18\+\-restricted\); moving it between tiers shifts no cell by more than22pp\.

Table 8:Percentage\-point contribution of the 18\+\-restricted and universally prohibited category tiers to each cell’s Gemini E3 harm rate, passive phase andSEARCHphase\.
#### Decomposing the Italian lead

Three passive\-phase decompositions support the reading in §[5](https://arxiv.org/html/2608.17583#S5)\. By category, Sexually Suggestive content contributes23\.823\.8pp of Italy’s39\.0%39\.0\\%passive rate versus12\.212\.2pp in France and15\.815\.8pp in Sweden, so a single category accounts for most of the cross\-country gap\. By content language \(Table[9](https://arxiv.org/html/2608.17583#A4.T9)\), Italian\-language videos are flagged at46\.8%46\.8\\%, but English\-language videos served to Italian accounts are also flagged at32\.4%32\.4\\%, clearly above the English\-language rate in France \(19\.5%19\.5\\%\), with Sweden in between \(27\.5%27\.5\\%\)\. By cross\-country circulation, passive videos that also appear in another country’s corpus are flagged at19\.2%19\.2\\%when served to Italian accounts, statistically indistinguishable from the shared\-content rate in France \(26\.4%26\.4\\%\) and Sweden \(28\.3%28\.3\\%\), while Italy\-exclusive content is flagged at41\.7%41\.7\\%; the Italian lead is carried entirely by content that circulates only in the Italian pool\.

Table 9:Passive\-phase Gemini E3 harm rate \(%\) by detected content language, withnnin parentheses\. “Native” is the persona country’s language; detection follows the caption\-based protocol of Table[5](https://arxiv.org/html/2608.17583#A3.T5)\.
#### Keyword\-coverage baseline

Figure[8](https://arxiv.org/html/2608.17583#A4.F8)reports the fraction of Stage\-2 active videos whose caption and transcript together contain at least one of the country’s 21 harm keywords \(whole\-word, case\- and diacritic\-insensitive match\)\.SEARCHvideos contain a literal keyword in2020–25%25\\%of cases across all three countries \(FR21\.6%21\.6\\%, IT20\.5%20\.5\\%, SW25\.1%25\.1\\%\), while thescroll\-preandscroll\-postbaselines sit at00–2%2\\%\. The two\-order\-of\-magnitude gap is a textual sanity check on the probe; the remaining7575–80%80\\%ofSEARCHvideos are surfaced by the platform’s own topic\-and\-engagement match around the query rather than by literal keyword presence in caption or transcript\.

Figure 8:Fraction of Stage\-2 active videos whose caption∪\\cuptranscript matches at least one of the country’s2121search keywords, by phase\.Table 10:Stage\-2 cross\-modality agreement between E2 and E3 on the binary harm verdict \(n=5,108n=5\{,\}108paired\)\.Figure 9:Stage\-2 Gemini E3 daily harm rate per \(country, age\) over the five active\-collection days\. Rows are country, columns are age; each major x\-tick is one collection day, with three points per day in orderscroll\-pre\(scroll\)→\\toSEARCH→\\toscroll\-post\(scroll\)\. The triangular per\-day shape is the local within\-daySEARCHlift\.Table 11:Per\-subcategory disagreement between E2 \(native video\) and E3 \(primary\) on the3,4173\{,\}417paired Stage\-2 items\.*both H*= both flagged harmful with this subcategory;*E2 only*= E2 flagged, E3 did not;*E3 only*= E3 flagged, E2 did not\.Figure 10:Stage\-2 harm rate per \(country, age\) combination on Gemini E3 and native\-video E2, with95%95\\%video\-level bootstrap CIs\. Italy carries the highest E3 rate on every age persona; the within\-country age gradient differs across the three countries\. The cleaner passive\-only view is in the main text \(Figure[5](https://arxiv.org/html/2608.17583#S4.F5), §[4\.3](https://arxiv.org/html/2608.17583#S4.SS3)\)\.Figure 11:Stage\-2 harm rate per \(country, age\) on the passive collection phase \(FYP scrolling only, no search probe\) under Gemini E2 \(native video\); companion to Figure[5](https://arxiv.org/html/2608.17583#S4.F5)\.

## Appendix EStage\-1 Diagnostic Detail

This appendix collects the per\-condition diagnostics referenced from §[4\.1](https://arxiv.org/html/2608.17583#S4.SS1)–§[4\.5](https://arxiv.org/html/2608.17583#S4.SS5)\. The numbers are computed on the 300\-video two\-annotator\-with\-resolution subset \(299 binary reference verdicts,8080–9999per combination depending on model failures\) and are pre\-Stage\-2\. The XLM\-R supervised text baseline is included alongside the four MLLM families\.

#### Inter\-annotator agreement and resolution

Two native\-speaker annotators per country independently labeled the 300\-video Stage\-1 subset\. Before resolution, the two annotators agreed on the binary harm verdict on80\.3%80\.3\\%of items \(n=300n=300\); Cohen’sκ\\kappaon the binary verdict is0\.540\.54in the aggregate, with per\-countryκFR=0\.42\\kappa\_\{\\text\{FR\}\}=0\.42,κIT=0\.65\\kappa\_\{\\text\{IT\}\}=0\.65,κSW=0\.48\\kappa\_\{\\text\{SW\}\}=0\.48\. All5959binary\-verdict disagreements were settled in a joint resolution session that produced a single consensus label per video\. A further resolution rule was applied at analysis time without re\-soliciting the annotators: when both annotators rated a videoHarmfulbut disagreed on its primary subcategory, the consensus row carries the union of the two subcategory choices\. The resulting per\-video table is thefinal referenceagainst which the LLM agreement numbers in this appendix are computed\.

#### Modality monotonicity

Figure[12](https://arxiv.org/html/2608.17583#A5.F12)shows how Cohen’sκ\\kappaand macro\-F1F\_\{1\}move across the three input conditions for the winning configuration \(Gemini 2\.5 Flash\) on the 300\-video Stage\-1 subset; the per\-condition numerics are cited inline in §[4\.1](https://arxiv.org/html/2608.17583#S4.SS1)\. On the same panel, Qwen3\-VL\-32B does*not*show the same monotonic pattern, sitting flat between E2 and E3 atκ≈0\.18\\kappa\\approx 0\.18, which suggests its native\-video pathway and its frame\-based pathway converge on the same \(weak\) judgment rather than complementing one another\.

Figure 12:Cohen’sκ\\kappaand macro\-F1F\_\{1\}across the three input conditions for Gemini 2\.5 Flash on the 300\-video Stage\-1 subset\. E3 \(eight frames\) is the strongest at∼2×\{\\sim\}2\\timeslower per\-call cost than E2 \(native video\)\.
#### Per\-country agreement of the winner

Figure[13](https://arxiv.org/html/2608.17583#A5.F13)expands the per\-countryκ\\kapparanking of the Stage\-1 winning configuration referenced in §[4\.1](https://arxiv.org/html/2608.17583#S4.SS1)\. Per\-country values areκFR=0\.290\\kappa\_\{\\text\{FR\}\}=0\.290\(CI\[0\.03,0\.52\]\[0\.03,0\.52\],nFR=93n\_\{\\text\{FR\}\}=93\),κIT=0\.463\\kappa\_\{\\text\{IT\}\}=0\.463\(CI\[0\.29,0\.63\]\[0\.29,0\.63\],nIT=98n\_\{\\text\{IT\}\}=98\), andκSW=0\.356\\kappa\_\{\\text\{SW\}\}=0\.356\(CI\[0\.16,0\.54\]\[0\.16,0\.54\],nSW=99n\_\{\\text\{SW\}\}=99\): Italian content reaches the highest agreement and French the lowest, which inverts the naive prediction that higher per\-language moderator allocation \(per[23](https://arxiv.org/html/2608.17583#bib.bib20)’s data\) should correlate with higher MLLM\-vs\-annotator agreement at the point\-estimate level; the per\-country CIs are wide and pairwise overlapping, so the inversion is consistent with sampling variation at thisnnrather than firm evidence against the hypothesis\. France contributes the largest share of provider\-blocked items \(Table[14](https://arxiv.org/html/2608.17583#A5.T14)\), so part of the FRκ\\kappagap is attributable to the removal of those items from the comparison\.

A complementary decomposition uses the pre\-resolution two\-annotatorκ\\kappaper country as a reference\-quality ceiling: the LLM cannot agree with the consensus better than the consensus agrees with itself\. The pre\-resolution two\-annotatorκ\\kappavalues areκFRann=0\.42\\kappa^\{\\text\{ann\}\}\_\{\\text\{FR\}\}=0\.42,κITann=0\.65\\kappa^\{\\text\{ann\}\}\_\{\\text\{IT\}\}=0\.65,κSWann=0\.48\\kappa^\{\\text\{ann\}\}\_\{\\text\{SW\}\}=0\.48\. Gemini E3 reaches69%69\\%of this ceiling on French content \(0\.290/0\.420\.290/0\.42\),71%71\\%on Italian \(0\.463/0\.650\.463/0\.65\), and74%74\\%on Swedish \(0\.356/0\.480\.356/0\.48\)\. The fact that the ratio is roughly constant across the three countries indicates that the absolute per\-countryκ\\kappagap \(FR<<SW<<IT\) tracks the per\-country reference\-quality gap rather than a country\-specific model deficit, and that the cross\-countryκ\\kappainversion of the moderator\-allocation prediction is best read as a reference\-quality artifact: French annotator disagreement is the main driver ofκFR\\kappa\_\{\\text\{FR\}\}being the lowest model\-vs\-reference combination\.

#### Per\-country precision/recall recalibration of Stage\-2 rates

Stage\-1 per\-country precision and recall on the binary harm verdict \(FR:P/R=0\.50/0\.33P/R=0\.50/0\.33,P/R≈1\.50P/R\\approx 1\.50; IT:0\.80/0\.600\.80/0\.60,≈1\.34\\approx 1\.34; SW:0\.62/0\.510\.62/0\.51,≈1\.21\\approx 1\.21\) give a country\-specific true\-positive\-rate correction for the Stage\-2 harm rates of §[4\.2](https://arxiv.org/html/2608.17583#S4.SS2)\. Applied uniformly across \(country, age\) combinations, the recalibration leaves the cross\-country ordering intact at ages 13, 16, and 19 \(Italy remains highest\) but inverts the age\-4040case, where corrected FR\-4040\(≈49%\\approx 49\\%\) overtakes IT\-4040\(≈38%\\approx 38\\%\) and SW\-4040\(≈33%\\approx 33\\%\)\. The Stage\-1P/RP/Rvalues are themselves point estimates on per\-countryn∈\[93,99\]n\\in\[93,99\], so the inversion is consistent with both “IT\-4040ties FR\-4040” and “IT highest on every age”; we report it as a caveat on the strongest reading of the cross\-country claim at age 40 rather than a re\-ordering\.

Figure 13:Per\-country Cohen’sκ\\kappafor the winning configuration \(Gemini 2\.5 Flash, E3\)\.
#### Aggregateκ\\kappagrid

Figure[14](https://arxiv.org/html/2608.17583#A5.F14)collapses the three\-panel split of Figure[3](https://arxiv.org/html/2608.17583#S4.F3)into a single aggregate model×\\timescondition grid for completeness\. The aggregate ordering averages over the strong cross\-lingual asymmetry visible in the per\-country split: the French panel is the only one where Qwen E2 exceeds Qwen E3 \(κFR=0\.31\\kappa\_\{\\text\{FR\}\}=0\.31vs\.0\.080\.08\); the Italian panel concentrates the moderate\-agreement combinations in the roster \(Gemini E2 reachesκIT=0\.51\\kappa\_\{\\text\{IT\}\}=0\.51, Gemini E3 reachesκIT=0\.46\\kappa\_\{\\text\{IT\}\}=0\.46, and GPT\-4o\-mini E3 separately reachesκIT=0\.41\\kappa\_\{\\text\{IT\}\}=0\.41\); and the Swedish panel collapses to Gemini E3 as the strongest non\-Italian combination \(non\-Gemini combinations on Swedish content stay at or belowκ=0\.17\\kappa=0\.17across all three input conditions, and Qwen E1 and GPT E1 sit at or below chance\)\.

Figure 14:Aggregate Cohen’sκ\\kappaon the binary harm verdict across the four MLLMs and the XLM\-R supervised baseline, by input condition\. GPT\-4o\-mini and Mistral are not natively video\-capable and are omitted from E2\. Dotted lines:κ=0\.21\\kappa=0\.21\(fair\) andκ=0\.41\\kappa=0\.41\(moderate\)\([16](https://arxiv.org/html/2608.17583#bib.bib26)\)\.
#### Pairwise inter\-model agreement, per country

Figure[15](https://arxiv.org/html/2608.17583#A5.F15)reports Cohen’sκ\\kappabetween every pair of the eleven Stage\-1 \(model, condition\) combinations \(the ten MLLM combinations plus the XLM\-R baseline\) on the binary harm label, split by persona country and computed on each pair’s intersection of pairedvideo\_ids within the country slice \(per\-combinationn≈95n\\approx 95–108108\)\. The same three blocks dominate every country panel: \(i\) a text\-only cluster among Qwen E1, GPT E1, and Mistral E1, with the Qwen↔\\leftrightarrowGPT pair the strongest in every country; \(ii\) a frames\-and\-video cluster anchored by Gemini E2↔\\leftrightarrowGemini E3 and Gemini E3↔\\leftrightarrowGPT E3; and \(iii\) an isolated Qwen E2 row that agrees only weakly with everything outside its own provider, consistent with the high413 RequestTooLargeblock rate documented in Table[14](https://arxiv.org/html/2608.17583#A5.T14)\. Within\-model, comparing Qwen E1 to Qwen E3 and GPT E1 to GPT E3 makes the effect of adding frames visible: GPT\-4o\-mini’s verdicts change substantially when shown frames, while Qwen3\-VL\-32B’s do not\. Read together with Figure[3](https://arxiv.org/html/2608.17583#S4.F3), the per\-country matrices support the language\-prior account: in conditions where the verdict is text\-driven, the four model families agree with each other at moderate\-to\-substantial levels regardless of provider, so the disagreement with the annotators comes mostly from the visual\-judgment side of the task, not from the text\-classification side\.

Figure 15:Pairwise Cohen’sκ\\kappabetween every pair of the eleven Stage\-1 \(model, condition\) combinations \(the ten MLLM combinations plus the XLM\-R baseline\), split by persona country\. Off\-diagonals are computed on each pair’s intersection of pairedvideo\_ids within the country slice \(per\-pairn≈95n\\approx 95–108108\)\.
#### Confusion of the winner

The consensus\-reference\-vs\-Gemini\-E3 confusion over then=290n=290paired records shows that the model is conservative on harm: it under\-flags more than it over\-flags \(48 false negatives vs\. 24 false positives\), inverting the usual concern about LLM over\-moderation\([15](https://arxiv.org/html/2608.17583#bib.bib7)\)and consistent with Gemini’s safety\-layer bias toward refusing rather than mislabeling sensitive content\. On the 52 jointly\-harmful items where both reference and model sayHarmful, agreement on the*primary*subcategory is 29% \(15/52\); under a relaxed definition that scores the model correct when its primary prediction matches either the human primary or human secondary subcategory, agreement rises to 54% \(28/52\)\. The bulk of strict\-match disagreements concentrate on adjacent categories within the same harm domain \(nudity\-vs\-suggestive, dangerous\-challenge\-vs\-shocking\) rather than across harm domains\.

#### Country\-specific harm\-category mix

Figure[16](https://arxiv.org/html/2608.17583#A5.F16)renders the country\-conditional distribution of harm subcategories among the reference\-flagged harmful videos \(102 across FR/IT/SW\)\. The Sankey routes the within\-country normalized category shares as flows, which keeps the absolute counts per country visible alongside the cross\-country comparison\. France’s harmful set is dominated bydangerous activity and challenges; Italy and Sweden are dominated bynudity and body exposuretogether withsexually suggestivecontent \(roughly two\-thirds of harmful items in each\), with hate\-speech and harassment categories more visible in Sweden than elsewhere\. A single globalκ\\kappafigure averages over these different harm distributions, so Stage\-2 must report per\-country, per\-category prevalence side by side with overall agreement to avoid hiding the asymmetry\.

Figure 16:Country→\\toharm\-subcategory mix for the 102 reference\-flagged harmful videos\. Ribbon widths are proportional to absolute counts\.
#### Winner vs\. runner\-up table

Table[12](https://arxiv.org/html/2608.17583#A5.T12)reports the full Stage\-1 metrics for Gemini E3 and the strongest open\-weight comparator under the same condition\. The 0\.27\-pointκ\\kappagap between the two configurations is large enough that no plausible scaling discount overturns the choice of Gemini E3 as the Stage\-2 configuration\.

Table 12:Stage\-1 metrics for the winning configuration and the strongest open\-weight comparator under the same E3 condition\.
#### Precision and recall across all Stage\-1 combinations

Table[13](https://arxiv.org/html/2608.17583#A5.T13)reports the full precision/recall/F1breakdown on the binary harm verdict for every Stage\-1 \(model, condition\) combination against the two\-annotator final reference\. Gemini E3 leads the table on every harm\-side metric; the supervised XLM\-R baseline beats all zero\-shot LLMs on harmful\-class recall but trails on precision, consistent with a supervised model fitting the positive class on a small training set\. All four non\-Gemini frame and text conditions sit at or below50%50\\%recall on harmful, which is the largest single failure mode of the LLM auditors\.

Table 13:Stage\-1 precision \(PH\{\}\_\{\\textsc\{H\}\}\), recall \(RH\{\}\_\{\\textsc\{H\}\}\), per\-classF1F\_\{1\}, and macro\-F1F\_\{1\}on the binary harm verdict, across the eleven configurations \(ten MLLMs plus the XLM\-R baseline\) against the two\-annotator final reference\. Rows sorted by Cohen’sκ\\kappa; per\-column maxima inbold\.
#### Supervised XLM\-R baseline

The XLM\-R row in Figure[14](https://arxiv.org/html/2608.17583#A5.F14)and Table[13](https://arxiv.org/html/2608.17583#A5.T13)is a 10\-fold repeated stratified train/test cross\-validation ofxlm\-roberta\-baseon the same 300\-video Stage\-1 subset used to evaluate the four MLLMs\. The input is the E1 text only, caption plus audio transcript, exactly the field shown to the text\-only LLM runs, and the target is the consensus binary harm verdict\. Each fold holds out19%19\\%of items as test, carves15%15\\%of the remaining train side as a dev split for early stopping, and trains for up to1010epochs with AdamW \(lr×10−52\\\!\\times\\\!10^\{\-5\}, batch size1616, max sequence length256256, weight decay0\.010\.01, linear warmup over6%6\\%of steps, early\-stopping patience of44epochs on dev macro\-F1F\_\{1\}\)\. Class imbalance is handled by inverse\-frequency reweighting computed on the train side\. Folds use random seeds00–99and the per\-video predicted probability is the mean across the \(typically two\) folds in which a given video appears in the test split; the metrics in Table[13](https://arxiv.org/html/2608.17583#A5.T13)are computed on those aggregated per\-video predictions against the same final reference\. We read the baseline as a sanity\-check anchor for the LLM E1 combinations rather than as a competitive system: at this training\-set size \(n≈242n\\approx 242per fold after droppingNot Availableitems\), the supervised model can beat zero\-shot LLMs on harmful\-class recall but cannot match Gemini E3’s precision\-recall balance\.

#### Provider\-block counts

Table 14:Provider\-side refusals on the 300\-video Stage\-1 subset, by experiment\. Each cell counts videos where the model’s response could not be evaluated because the provider blocked the call or returned an unparseable response\.The Qwen413 RequestTooLargedominates absolute volume because DashScope’s native\-video endpoint enforces a per\-call payload cap; the workaround would change what Qwen sees and break comparability with unmodified Gemini E2, so we report Qwen E2 only as a comparator\. Gemini’s blocks are policy decisions; they could be relaxed viasafety\_settings=BLOCK\_NONE, appropriate only under formal ethics approval\.

Similar Articles