The Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale
Summary
We introduce ISAAC, an open corpus of 527 million+ English-language Reddit posts for analyzing social group discourse, with a multi-step pipeline for annotation and analysis. It enables cross-category comparisons and temporal tracking of public attitudes, accessible via web apps and programming interfaces.
View Cached Full Text
Cached at: 09/24/26, 09:13 AM
# The Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale Source: [https://arxiv.org/abs/2609.27059](https://arxiv.org/abs/2609.27059) [View PDF](https://arxiv.org/pdf/2609.27059) > Abstract:We introduce the Illinois Social Attitudes Aggregate Corpus \(ISAAC\), an open, modular, and accessible corpus of 527 million\+ English\-language Reddit posts selected for relevance to six key social group distinctions based on race, sexuality, age, ability, body weight, and skin tone, covering the 17\-year period from 2007 to 2023\. A multi\-step, human\-audited filtering pipeline was used to keep irrelevant content in the curated dataset below 10%, both overall and for each social group distinction\. Each post was then algorithmically annotated with the user's estimated home region, along with a suite of validated off\-the\-shelf and custom semantic labels including moralization, sentiment, emotion, and linguistic generalization\. We confirm the validity of the resulting corpus through convergent evidence linking ISAAC to macro\-level societal trends, such as online search behavior, temporal spikes during major societal events \(both nationally and regionally\), and long\-term shifts in public attitudes\. By offering a unified, public infrastructure, ISAAC eliminates research fragmentation and enables seamless replication while supporting diverse empirical workflows at scale\. Specifically, ISAAC allows investigators to perform cross\-category comparisons, conduct high\-precision tracking of long\-term temporal shifts in social group discourse, and map spatial variation onto localized public opinion and policy outcomes\. ISAAC's fully public, modular pipeline facilitates easy extension of the corpus to new platforms, languages, and social categories\. To accommodate various research needs, ISAAC is accessible both without coding through a point\-and\-click website and labeler web\-apps, and programmatically via an SQL playground, a Python package, and HuggingFace\. ## Submission history From: Babak Hemmatian \[[view email](https://arxiv.org/show-email/78ba128c/2609.27059)\] **\[v1\]**Tue, 22 Sep 2026 20:59:57 UTC \(2,720 KB\)
Similar Articles
ACAT: A Collaborative Platform for Efficient Aspect-Based Sentiment Dataset Annotation
ACAT is a web-based collaborative annotation platform supporting four Aspect-Based Sentiment Analysis (ABSA) workflows, featuring an automated ETL pipeline that computes Inter-Annotator Agreement metrics at export to produce training-ready datasets. Validated on 1,002 restaurant reviews, it achieves a median annotation time of 31.58 seconds and raw IAA up to 0.86.
Cohesion-6K: An Arabic Dataset for Analyzing Social Cohesion and Conflict in Online Discourse
Introduces Cohesion-6K, a manually and ChatGPT-assisted annotated dataset of 6,000 Arabic Facebook posts about the Israeli Occupation of Palestine, spanning conflict to cohesion categories. Analysis shows conflict-oriented posts receive 2-4x more engagement than resolution-oriented ones.
Voices in the Loop: Mapping Participatory AI
This paper presents a reproducible protocol for building an open repository and interactive atlas of participatory AI initiatives, analyzing 131 records to reveal geographic and lifecycle patterns, and proposes a framework for participatory-by-default AI infrastructures.
A Context-Aware Dataset for Stance Detection in Bioethical Controversies on Reddit
Presents BioStance, a context-aware dataset of 39,600 annotated Reddit post-comment pairs for stance detection in bioethical controversies, covering six targets across three dimensions of bioethical debate.
I let 100 AI personas run a Reddit for a month — they formed factions, hold grudges from thread to thread, and you can drop in any post title to watch them swarm
An experiment running 100 LLM personas on a Reddit-style forum demonstrates emergent social dynamics like factions and persistent grudges, built with Node.js and using OpenRouter's deepseek-chat.