Tag
This paper conducts a multi-method computational analysis of Telegram discourse on the Israel-Palestine conflict, revealing that pro-Israel channels use a neutral, report-style tone while pro-Palestine channels exhibit more negative sentiment and framing.
This paper analyzes 28,592 French news headlines about La France insoumise and Rassemblement National using an LLM annotation pipeline, finding asymmetric role framing where LFI is more often framed as aggressors and RN as strategic actors.
This paper compares RoBERTa-based sentiment analysis with an LLM-based multi-dimensional framing analysis on political news articles, finding that traditional SA suffers from 'neutral collapse' and that LLM-based approaches better capture bias, sensationalism, and framing for social science research.
This paper presents a 17k-sentence corpus with annotations for argumentative passages across three German political arenas during COVID-19, and a pilot study on automatically identifying such passages, finding that boundaries are hard to pin down and models exhibit confirmation bias.
This Data Descriptor presents a large-scale corpus of transcribed religious radio broadcasts captured from live webstreams over one month in July 2025, comprising over 700,000 recordings and 60 million transcript lines, annotated using LLMs for program format and topic. It enables descriptive study of religious broadcasting and analysis of social/political issues in religious media.
Japanese researchers developed CitySim, an urban simulator powered by LLMs that populates a digital twin of Tokyo with up to 1 million autonomous AI agents, accurately predicting real-world patterns like commuting and shopping behavior.
This paper proposes a diagnostic framework to separate preprocessing pipeline instability from measurement method instability in LLM-based stance analysis of public discourse, finding that cross-method disagreement is larger and more systematic than pipeline effects, and that aggregate metrics can mask these instabilities.
This paper develops a codebook for self-stigma among people who use drugs and analyzes 72,115 Reddit posts to examine prevalence, co-occurrence, and temporal patterns of cognitive, affective, and behavioral stigma indicators, finding that self-stigma is expressed as an integrated phenomenon with behavioral indicators often preceding core indicators.
This paper presents the largest computational analysis of Canadian news coverage of police-involved deaths over 25 years, introducing a novel model (PerspectiveGap) that quantifies the dominance of state bureaucrat perspectives compared to civilian voices in media narratives.
This paper introduces conditional hypothesis generation, a framework that incorporates researcher-specified covariates to steer LLM-based text analysis toward discovering meaningful subgroup differences while addressing confounds like stratum imbalance and sign reversal.
This paper discusses the need for multilingual LLMs that are epistemically grounded and responsible for applications in computational social science and humanities.
This paper proposes a label-light measurement diagnostic to evaluate whether popular text analysis methods (dictionaries, topic models, embeddings, LLMs) capture substantive stance versus symbolic rhetoric in entrepreneurial-discourse measurement, using a corpus of 80 Chinese SOE speeches and a natural experiment with same-company different-speaker pairs. The authors find that zero-shot LLMs show higher sensitivity but a significant portion of the effect may be due to speaker idiolect rather than substantive stance.
This paper presents the Arabic Women and Society Corpus, a ten-year collection of over 250,000 Arabic Facebook posts related to women's empowerment and social wellbeing, with engagement metrics for analyzing gender discourse and sentiment.
This paper critiques the 'Proxy Presumption' in NLP, where geometric embedding properties are incorrectly equated with social constructs. It introduces the Construct Validity Protocol and Counterfactual Neutralization methods to ensure rigorous validation of social measures derived from semantic embeddings.