Challenges in annotations by humans and LLMs: A case study of evaluative language
Summary
This paper compares human linguists in training, a trained linguist, and LLMs on annotating evaluative language using Appraisal theory, finding that LLMs achieve strong performance and can assist in complex annotation tasks.
View Cached Full Text
Cached at: 07/31/26, 10:03 AM
# Challenges in annotations by humans and LLMs: A case study of evaluative language Source: [https://arxiv.org/abs/2607.28119](https://arxiv.org/abs/2607.28119) [View PDF](https://arxiv.org/pdf/2607.28119) > Abstract:In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models \(LLMs\) to find out if they struggle with complex linguistic phenomena in a similar way\. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts\. We focus on the Appraisal theory and its Attitude subsystem, including the categories \(classes\) of Affect, Judgement, and Appreciation\. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges\. First, we assess human annotations on a sentence level in specific scientific domains\. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes\. We assess the performance of three LLMs using the best\-performing prompt and finetune the model, reaching an F1\-score of 0\.77\. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores\. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies\. ## Submission history From: Aenne Cecilia Kristine Knierim \[[view email](https://arxiv.org/show-email/006c83dd/2607.28119)\] **\[v1\]**Thu, 30 Jul 2026 12:28:54 UTC \(1,308 KB\)
Similar Articles
LLMs for automatic annotation of Mandarin narrative transcripts
This paper evaluates LLMs for automatically annotating narrative macrostructure in spoken Mandarin, finding that the best model achieves near-human reliability while reducing annotation time by 65%, though performance degrades on semantically complex or lexically diverse narratives.
Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA
This paper tests the assumption that LLMs judge better than they generate in in-context QA, finding generation accuracy exceeds self-evaluation on most benchmarks, with evaluation attending less to context. The findings challenge core assumptions in self-evaluation pipelines.
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
This paper proposes an Interpretive Audit Pipeline that leverages multi-model disagreement to detect interpretive complexity in LLM-based public comment analysis, arguing that disagreement-based evaluation is a necessary complement to standard accuracy metrics.
Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages
This paper analyzes the use of LLM-as-a-Judge in multilingual and low-resource settings, finding inconsistent evaluation outcomes and overtrust in LLM judgments, and provides recommendations for better practices.
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
This paper proposes a two-stage sampling design where LLM evaluations are used to augment, rather than replace, human ratings, and provides guidance on determining sample sizes for human and LLM reviews using a doubly robust estimator from missing data literature.