Challenges in annotations by humans and LLMs: A case study of evaluative language

arXiv cs.CL Papers

Summary

This paper compares human linguists in training, a trained linguist, and LLMs on annotating evaluative language using Appraisal theory, finding that LLMs achieve strong performance and can assist in complex annotation tasks.

arXiv:2607.28119v1 Announce Type: new Abstract: In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.
Original Article
View Cached Full Text

Cached at: 07/31/26, 10:03 AM

# Challenges in annotations by humans and LLMs: A case study of evaluative language
Source: [https://arxiv.org/abs/2607.28119](https://arxiv.org/abs/2607.28119)
[View PDF](https://arxiv.org/pdf/2607.28119)

> Abstract:In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models \(LLMs\) to find out if they struggle with complex linguistic phenomena in a similar way\. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts\. We focus on the Appraisal theory and its Attitude subsystem, including the categories \(classes\) of Affect, Judgement, and Appreciation\. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges\. First, we assess human annotations on a sentence level in specific scientific domains\. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes\. We assess the performance of three LLMs using the best\-performing prompt and finetune the model, reaching an F1\-score of 0\.77\. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores\. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies\.

## Submission history

From: Aenne Cecilia Kristine Knierim \[[view email](https://arxiv.org/show-email/006c83dd/2607.28119)\] **\[v1\]**Thu, 30 Jul 2026 12:28:54 UTC \(1,308 KB\)

Similar Articles

LLMs for automatic annotation of Mandarin narrative transcripts

arXiv cs.CL

This paper evaluates LLMs for automatically annotating narrative macrostructure in spoken Mandarin, finding that the best model achieves near-human reliability while reducing annotation time by 65%, though performance degrades on semantically complex or lexically diverse narratives.