Assessing socio-economic climate impacts from text data

arXiv cs.CL Papers

Summary

This paper reviews recent advances in using natural language processing and large language models to extract socio-economic impact data from text sources for climate hazards, identifies key challenges, and provides recommendations for robust dataset construction.

arXiv:2605.20793v1 Announce Type: new Abstract: Recent advances in natural language processing (NLP) and large language models (LLMs) have enabled the systematic use of large-scale textual data from news, social media, and reports to create datasets with socio-economic impacts of climate hazards such as floods, droughts, storms, and multi-hazard events. As the field of text-as-data for impact assessment expands, so does its methodological complexity. Yet research remains fragmented, with no clear guidelines for defining what constitutes an impact, handling temporal and spatial biases, and selecting appropriate modeling and post-processing strategies. This lack of coherence limits transparency and comparability across studies. Here, we address this gap by synthesising common practices, describing key challenges specific to the use of text-as-data methods for analyzing socio-economic impact data, and proposing recommendations to address them. By providing guidance on best practices, we aim to support the construction of robust text-derived socio-economic impact datasets that can more accurately inform disaster risk management and attribution studies.
Original Article
View Cached Full Text

Cached at: 05/21/26, 06:35 AM

# Assessing socio-economic climate impacts from text data
Source: [https://arxiv.org/abs/2605.20793](https://arxiv.org/abs/2605.20793)
Authors:[Mariana Madruga de Brito](https://arxiv.org/search/cs?searchtype=author&query=de+Brito,+M+M),[Brielen Madureira](https://arxiv.org/search/cs?searchtype=author&query=Madureira,+B),[Taís Maria Nunes Carvalho](https://arxiv.org/search/cs?searchtype=author&query=Carvalho,+T+M+N),[Damien Delforge](https://arxiv.org/search/cs?searchtype=author&query=Delforge,+D),[Aglaé Jézéquel](https://arxiv.org/search/cs?searchtype=author&query=J%C3%A9z%C3%A9quel,+A),[Murathan Kurfalı](https://arxiv.org/search/cs?searchtype=author&query=Kurfal%C4%B1,+M),[Ni Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+N),[Gabriele Messori](https://arxiv.org/search/cs?searchtype=author&query=Messori,+G),[Joakim Nivre](https://arxiv.org/search/cs?searchtype=author&query=Nivre,+J),[Barbara Pernici](https://arxiv.org/search/cs?searchtype=author&query=Pernici,+B),[Niko Speybroeck](https://arxiv.org/search/cs?searchtype=author&query=Speybroeck,+N),[Stefano Terzi](https://arxiv.org/search/cs?searchtype=author&query=Terzi,+S),[Wim Thiery](https://arxiv.org/search/cs?searchtype=author&query=Thiery,+W),[Bram Valkenborg](https://arxiv.org/search/cs?searchtype=author&query=Valkenborg,+B),[Jingxian Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+J),[Shorouq Zahra](https://arxiv.org/search/cs?searchtype=author&query=Zahra,+S),[Jakob Zscheischler](https://arxiv.org/search/cs?searchtype=author&query=Zscheischler,+J),[Jan Sodoge](https://arxiv.org/search/cs?searchtype=author&query=Sodoge,+J)

[View PDF](https://arxiv.org/pdf/2605.20793)

> Abstract:Recent advances in natural language processing \(NLP\) and large language models \(LLMs\) have enabled the systematic use of large\-scale textual data from news, social media, and reports to create datasets with socio\-economic impacts of climate hazards such as floods, droughts, storms, and multi\-hazard events\. As the field of text\-as\-data for impact assessment expands, so does its methodological complexity\. Yet research remains fragmented, with no clear guidelines for defining what constitutes an impact, handling temporal and spatial biases, and selecting appropriate modeling and post\-processing strategies\. This lack of coherence limits transparency and comparability across studies\. Here, we address this gap by synthesising common practices, describing key challenges specific to the use of text\-as\-data methods for analyzing socio\-economic impact data, and proposing recommendations to address them\. By providing guidance on best practices, we aim to support the construction of robust text\-derived socio\-economic impact datasets that can more accurately inform disaster risk management and attribution studies\.

## Submission history

From: Brielen Madureira \[[view email](https://arxiv.org/show-email/e55a8a4f/2605.20793)\] **\[v1\]**Wed, 20 May 2026 06:40:00 UTC \(1,245 KB\)

Similar Articles

Opportunities and Challenges of Large Language Models for Low-Resource Languages in Humanities Research

arXiv cs.CL

This paper systematically evaluates the applications of large language models in low-resource language research, analyzing opportunities and challenges across linguistic variation, historical documentation, cultural expressions, and literary analysis. The study emphasizes interdisciplinary collaboration and customized model development to preserve linguistic and cultural heritage while addressing issues of data accessibility, model adaptability, and cultural sensitivity.

Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses

arXiv cs.AI

This paper presents a five-stage framework integrating large language models into survey research, addressing declining response rates, sample bias, and fraudulent completions. Using 2024 Hurricane Milton survey data, the authors propose a theory-informed LLM (A-TLM) that outperforms classical imputation methods in missing-data scenarios and demonstrates manageable hallucination risk through grounded refusal.

AI-integrated models for assessing agricultural resilience

arXiv cs.AI

This paper presents an AI-powered tool that integrates economic (GTAP) and biophysical (APSIM) models to analyze agricultural supply chain shocks, enabling natural language queries for cross-disciplinary impact assessment.