The Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale

arXiv cs.CL Papers

Summary

We introduce ISAAC, an open corpus of 527 million+ English-language Reddit posts for analyzing social group discourse, with a multi-step pipeline for annotation and analysis. It enables cross-category comparisons and temporal tracking of public attitudes, accessible via web apps and programming interfaces.

arXiv:2609.27059v1 Announce Type: new Abstract: We introduce the Illinois Social Attitudes Aggregate Corpus (ISAAC), an open, modular, and accessible corpus of 527 million+ English-language Reddit posts selected for relevance to six key social group distinctions based on race, sexuality, age, ability, body weight, and skin tone, covering the 17-year period from 2007 to 2023. A multi-step, human-audited filtering pipeline was used to keep irrelevant content in the curated dataset below 10%, both overall and for each social group distinction. Each post was then algorithmically annotated with the user's estimated home region, along with a suite of validated off-the-shelf and custom semantic labels including moralization, sentiment, emotion, and linguistic generalization. We confirm the validity of the resulting corpus through convergent evidence linking ISAAC to macro-level societal trends, such as online search behavior, temporal spikes during major societal events (both nationally and regionally), and long-term shifts in public attitudes. By offering a unified, public infrastructure, ISAAC eliminates research fragmentation and enables seamless replication while supporting diverse empirical workflows at scale. Specifically, ISAAC allows investigators to perform cross-category comparisons, conduct high-precision tracking of long-term temporal shifts in social group discourse, and map spatial variation onto localized public opinion and policy outcomes. ISAAC's fully public, modular pipeline facilitates easy extension of the corpus to new platforms, languages, and social categories. To accommodate various research needs, ISAAC is accessible both without coding through a point-and-click website and labeler web-apps, and programmatically via an SQL playground, a Python package, and HuggingFace.
Original Article
View Cached Full Text

Cached at: 09/24/26, 09:13 AM

# The Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale
Source: [https://arxiv.org/abs/2609.27059](https://arxiv.org/abs/2609.27059)
[View PDF](https://arxiv.org/pdf/2609.27059)

> Abstract:We introduce the Illinois Social Attitudes Aggregate Corpus \(ISAAC\), an open, modular, and accessible corpus of 527 million\+ English\-language Reddit posts selected for relevance to six key social group distinctions based on race, sexuality, age, ability, body weight, and skin tone, covering the 17\-year period from 2007 to 2023\. A multi\-step, human\-audited filtering pipeline was used to keep irrelevant content in the curated dataset below 10%, both overall and for each social group distinction\. Each post was then algorithmically annotated with the user's estimated home region, along with a suite of validated off\-the\-shelf and custom semantic labels including moralization, sentiment, emotion, and linguistic generalization\. We confirm the validity of the resulting corpus through convergent evidence linking ISAAC to macro\-level societal trends, such as online search behavior, temporal spikes during major societal events \(both nationally and regionally\), and long\-term shifts in public attitudes\. By offering a unified, public infrastructure, ISAAC eliminates research fragmentation and enables seamless replication while supporting diverse empirical workflows at scale\. Specifically, ISAAC allows investigators to perform cross\-category comparisons, conduct high\-precision tracking of long\-term temporal shifts in social group discourse, and map spatial variation onto localized public opinion and policy outcomes\. ISAAC's fully public, modular pipeline facilitates easy extension of the corpus to new platforms, languages, and social categories\. To accommodate various research needs, ISAAC is accessible both without coding through a point\-and\-click website and labeler web\-apps, and programmatically via an SQL playground, a Python package, and HuggingFace\.

## Submission history

From: Babak Hemmatian \[[view email](https://arxiv.org/show-email/78ba128c/2609.27059)\] **\[v1\]**Tue, 22 Sep 2026 20:59:57 UTC \(2,720 KB\)

Similar Articles

ACAT: A Collaborative Platform for Efficient Aspect-Based Sentiment Dataset Annotation

arXiv cs.CL

ACAT is a web-based collaborative annotation platform supporting four Aspect-Based Sentiment Analysis (ABSA) workflows, featuring an automated ETL pipeline that computes Inter-Annotator Agreement metrics at export to produce training-ready datasets. Validated on 1,002 restaurant reviews, it achieves a median annotation time of 31.58 seconds and raw IAA up to 0.86.

Voices in the Loop: Mapping Participatory AI

arXiv cs.AI

This paper presents a reproducible protocol for building an open repository and interactive atlas of participatory AI initiatives, analyzing 131 records to reveal geographic and lifecycle patterns, and proposes a framework for participatory-by-default AI infrastructures.