Tag
The study examines how large language models generate and detect fake news under different scenarios, revealing variations in performance and that refined prompts do not always improve detection.
An interview with Professor Beth Singler discusses how AI discourse often employs religious terminology while dismissing religion, resulting in a lack of self-reflection and valuable insights from religious perspectives.
This paper presents a study examining how anthropomorphic language in AI discourse affects public perceptions, finding that while overall views can shift, the specific effect of anthropomorphic framing is modest in controlled settings.
This study investigates whether instruction-tuned LLMs (Llama-3.1-8B, Qwen2.5-7B, Mistral-7B, Phi-3-mini) can reliably classify Correct Information Units in aphasic discourse transcripts. Few-shot prompting yields competitive F1 scores (0.776–0.817) for three models, but performance varies by severity and human agreement remains insufficient for fully autonomous use.
This paper evaluates LLMs for automatically annotating narrative macrostructure in spoken Mandarin, finding that the best model achieves near-human reliability while reducing annotation time by 65%, though performance degrades on semantically complex or lexically diverse narratives.
This paper introduces a benchmark for semantic segmentation in low-resource dialectal Arabic and proposes a model that improves performance on conversational speech compared to standard baselines.