Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication
Summary
This benchmark study evaluates 46 large language models against human experts for coding qualitative humanitarian data, finding that LLMs can achieve comparable reliability with structured prompts and reasoning, but require careful oversight for nuanced themes.
View Cached Full Text
Cached at: 06/26/26, 05:21 AM
# Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication Source: [https://arxiv.org/abs/2606.26541](https://arxiv.org/abs/2606.26541) [View PDF](https://arxiv.org/pdf/2606.26541) > Abstract:Data from affected populations are crucial for informing humanitarian response, but their value depends on timely and consistent interpretation of nuanced accounts of need\. Humanitarian organizations often lack the staff, time, and specialist expertise required to analyze this information at scale\. Large language models \(LLMs\) may expand this capacity, but their reliability for coding qualitative humanitarian data has not been directly established\. This benchmark study compares 46 LLMs to a human Gold Standard using 150 high\-fidelity synthetic humanitarian transcripts\. Evaluation combined inter\-rater reliability testing with Krippendorff's alpha, discrepancy analysis distinguishing correct, near\-correct, and incorrect codes, and qualitative assessment across humanitarian\-specific criteria including discrimination, complex needs hierarchies, and non\-standard communication styles\. The authors find that multiple LLMs can perform deductive coding at reliability levels comparable to experienced human coders, especially when structured prompts and reasoning\-enabled configurations are used\. At the same time, aggregate reliability metrics alone are insufficient for deployment decisions\. Models varied in recognizing needs expressed indirectly, needs outside predefined categories, and protection\-relevant concerns such as physical safety and discrimination\. These findings suggest that LLMs can materially expand humanitarian analytical capacity, but not as substitutes for human judgment\. Appropriate use requires structured codebooks, reasoning\-enabled models, attention to theme\-specific performance, and tiered oversight focused on categories where miscoding would have the greatest programmatic consequences\. For sensitive humanitarian data, open\-weights models deployed on self\-hosted infrastructure may offer a viable path for combining analytical scalability with stronger data governance\. ## Submission history From: Patrick Vinck \[[view email](https://arxiv.org/show-email/d048f44d/2606.26541)\] **\[v1\]**Thu, 25 Jun 2026 02:27:23 UTC \(632 KB\)
Similar Articles
LLMs Can Better Capture Human Judgments--With the Right Prompts
This paper presents simple prompting strategies that help large language models better capture the full distribution of human judgments, improving alignment on moral scenarios and beliefs. The authors show that asking models to report standard deviations and response proportions, along with ensuring scenario clarity, yields better agreement with human responses.
Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models
This preprint evaluates how six large language models respond to prompt framing and biased prompts across 160 prompts, finding that LLMs systematically adapt their responses to align with prompt framing even in factual contexts, potentially reinforcing user biases.
Causal Evidence Extraction and Triangulation in Crisis Reports using Large Language Models: A ReliefWeb-based Study
This paper presents a two-stage LLM pipeline for extracting and triangulating causal evidence from humanitarian crisis reports, achieving strong F1 scores on a ReliefWeb dataset and proposing a Level-of-Evidence score for cross-context convergence.
Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses
This paper presents a five-stage framework integrating large language models into survey research, addressing declining response rates, sample bias, and fraudulent completions. Using 2024 Hurricane Milton survey data, the authors propose a theory-informed LLM (A-TLM) that outperforms classical imputation methods in missing-data scenarios and demonstrates manageable hallucination risk through grounded refusal.
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
This paper introduces a register-aware linguistic evaluation framework to assess how human-like large language models (LLMs) are by comparing the distribution of 67 lexico-grammatical features between human and LLM-generated texts using Maximum Mean Discrepancy. Experiments across seven instruction-tuned open-source models and five registers show that no model perfectly matches human baselines, and closeness to human language varies by register rather than model size.