Assessing socio-economic climate impacts from text data
Summary
This paper reviews recent advances in using natural language processing and large language models to extract socio-economic impact data from text sources for climate hazards, identifies key challenges, and provides recommendations for robust dataset construction.
View Cached Full Text
Cached at: 05/21/26, 06:35 AM
# Assessing socio-economic climate impacts from text data Source: [https://arxiv.org/abs/2605.20793](https://arxiv.org/abs/2605.20793) Authors:[Mariana Madruga de Brito](https://arxiv.org/search/cs?searchtype=author&query=de+Brito,+M+M),[Brielen Madureira](https://arxiv.org/search/cs?searchtype=author&query=Madureira,+B),[Taís Maria Nunes Carvalho](https://arxiv.org/search/cs?searchtype=author&query=Carvalho,+T+M+N),[Damien Delforge](https://arxiv.org/search/cs?searchtype=author&query=Delforge,+D),[Aglaé Jézéquel](https://arxiv.org/search/cs?searchtype=author&query=J%C3%A9z%C3%A9quel,+A),[Murathan Kurfalı](https://arxiv.org/search/cs?searchtype=author&query=Kurfal%C4%B1,+M),[Ni Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+N),[Gabriele Messori](https://arxiv.org/search/cs?searchtype=author&query=Messori,+G),[Joakim Nivre](https://arxiv.org/search/cs?searchtype=author&query=Nivre,+J),[Barbara Pernici](https://arxiv.org/search/cs?searchtype=author&query=Pernici,+B),[Niko Speybroeck](https://arxiv.org/search/cs?searchtype=author&query=Speybroeck,+N),[Stefano Terzi](https://arxiv.org/search/cs?searchtype=author&query=Terzi,+S),[Wim Thiery](https://arxiv.org/search/cs?searchtype=author&query=Thiery,+W),[Bram Valkenborg](https://arxiv.org/search/cs?searchtype=author&query=Valkenborg,+B),[Jingxian Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+J),[Shorouq Zahra](https://arxiv.org/search/cs?searchtype=author&query=Zahra,+S),[Jakob Zscheischler](https://arxiv.org/search/cs?searchtype=author&query=Zscheischler,+J),[Jan Sodoge](https://arxiv.org/search/cs?searchtype=author&query=Sodoge,+J) [View PDF](https://arxiv.org/pdf/2605.20793) > Abstract:Recent advances in natural language processing \(NLP\) and large language models \(LLMs\) have enabled the systematic use of large\-scale textual data from news, social media, and reports to create datasets with socio\-economic impacts of climate hazards such as floods, droughts, storms, and multi\-hazard events\. As the field of text\-as\-data for impact assessment expands, so does its methodological complexity\. Yet research remains fragmented, with no clear guidelines for defining what constitutes an impact, handling temporal and spatial biases, and selecting appropriate modeling and post\-processing strategies\. This lack of coherence limits transparency and comparability across studies\. Here, we address this gap by synthesising common practices, describing key challenges specific to the use of text\-as\-data methods for analyzing socio\-economic impact data, and proposing recommendations to address them\. By providing guidance on best practices, we aim to support the construction of robust text\-derived socio\-economic impact datasets that can more accurately inform disaster risk management and attribution studies\. ## Submission history From: Brielen Madureira \[[view email](https://arxiv.org/show-email/e55a8a4f/2605.20793)\] **\[v1\]**Wed, 20 May 2026 06:40:00 UTC \(1,245 KB\)
Similar Articles
Large Language Models for Causal Relations Extraction in Social Media: A Validation Framework for Disaster Intelligence
This paper proposes a validation framework for using Large Language Models to extract causal relations from social media posts during disasters. It evaluates the effectiveness of LLMs in identifying cause-effect relationships and compares them against expert-grounded reference graphs to assess reliability and risks.
Opportunities and Challenges of Large Language Models for Low-Resource Languages in Humanities Research
This paper systematically evaluates the applications of large language models in low-resource language research, analyzing opportunities and challenges across linguistic variation, historical documentation, cultural expressions, and literary analysis. The study emphasizes interdisciplinary collaboration and customized model development to preserve linguistic and cultural heritage while addressing issues of data accessibility, model adaptability, and cultural sensitivity.
Automated Detection and Structuring of Social Tipping Point Evidence in Climate related Documents: A Modular AI Framework
This paper introduces an open modular transformer-based framework for detecting and structuring social tipping point evidence in climate documents, utilizing models like DistilBERT, RoBERTa, Mistral 7B, and LLaMA 3.2 3B, with evaluation on expert-labeled benchmarks.
Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses
This paper presents a five-stage framework integrating large language models into survey research, addressing declining response rates, sample bias, and fraudulent completions. Using 2024 Hurricane Milton survey data, the authors propose a theory-informed LLM (A-TLM) that outperforms classical imputation methods in missing-data scenarios and demonstrates manageable hallucination risk through grounded refusal.
AI-integrated models for assessing agricultural resilience
This paper presents an AI-powered tool that integrates economic (GTAP) and biophysical (APSIM) models to analyze agricultural supply chain shocks, enabling natural language queries for cross-disciplinary impact assessment.