LLMs论辩行为的基准测试:针对人身攻击的防御策略研究

arXiv cs.CL 论文

摘要

本文对大语言模型(LLMs)在论辩对话中处理人身攻击的能力进行了基准测试,揭示了由于安全限制,LLMs优先选择逻辑防御,而不像人类辩手。

arXiv:2609.28673v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on character attacks (ad hominem arguments), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content. Specifically, we investigate whether modern LLMs can replicate human competence to strategically use and respond to such attacks. We analyse a corpus of natural language political dialogues to identify defensive strategies human interlocutors naturally employ in ethos-centred debates and structure them into a dialogue game. Empirically, we benchmark LLM-generated dialogues against the ElecDeb60to16-fallacy corpus of U.S. presidential debates, contrasting human debaters' repertoire of defensive strategies with those of artificial agents. Results reveal a substantial difference: most LLMs rigidly prioritise logical defences, failing to exploit ethotic counterattacks as valid moves in political discourse. We argue that current safety fine-tuning constraints the strategic action space of these LLMs, making them unable to fully engage in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy.
查看原文
查看缓存全文

缓存时间: 2026/09/25 09:14

# Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks
Source: [https://arxiv.org/abs/2609.28673](https://arxiv.org/abs/2609.28673)
[View PDF](https://arxiv.org/pdf/2609.28673)

> Abstract:Large Language Models \(LLMs\) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors\. In this study, we focus on character attacks \(ad hominem arguments\), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content\. Specifically, we investigate whether modern LLMs can replicate human competence to strategically use and respond to such attacks\. We analyse a corpus of natural language political dialogues to identify defensive strategies human interlocutors naturally employ in ethos\-centred debates and structure them into a dialogue game\. Empirically, we benchmark LLM\-generated dialogues against the ElecDeb60to16\-fallacy corpus of U\.S\. presidential debates, contrasting human debaters' repertoire of defensive strategies with those of artificial agents\. Results reveal a substantial difference: most LLMs rigidly prioritise logical defences, failing to exploit ethotic counterattacks as valid moves in political discourse\. We argue that current safety fine\-tuning constraints the strategic action space of these LLMs, making them unable to fully engage in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy\.

## Submission history

From: Ewelina Gajewska \[[view email](https://arxiv.org/show-email/4a9da5a9/2609.28673)\] **\[v1\]**Wed, 23 Sep 2026 18:14:32 UTC \(103 KB\)

相似文章

鲁棒批评者:防御LLMs免受多轮攻击

arXiv cs.AI

本文提出对话批评者引导采样(DCGS)框架,通过从对话历史推断用户意图,并利用基于价值/遗憾的批评者对回复进行评分,在不进行微调的情况下提高鲁棒性,从而防御LLMs免受多轮对抗性攻击。

LLM基准测试

Reddit r/AI_Agents

一项关于大语言模型基准测试的研究或报告,可能比较了各种任务上的表现。