LLMs论辩行为的基准测试:针对人身攻击的防御策略研究
摘要
本文对大语言模型(LLMs)在论辩对话中处理人身攻击的能力进行了基准测试,揭示了由于安全限制,LLMs优先选择逻辑防御,而不像人类辩手。
arXiv:2609.28673v1 Announce Type: new
Abstract: Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on character attacks (ad hominem arguments), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content. Specifically, we investigate whether modern LLMs can replicate human competence to strategically use and respond to such attacks. We analyse a corpus of natural language political dialogues to identify defensive strategies human interlocutors naturally employ in ethos-centred debates and structure them into a dialogue game. Empirically, we benchmark LLM-generated dialogues against the ElecDeb60to16-fallacy corpus of U.S. presidential debates, contrasting human debaters' repertoire of defensive strategies with those of artificial agents. Results reveal a substantial difference: most LLMs rigidly prioritise logical defences, failing to exploit ethotic counterattacks as valid moves in political discourse. We argue that current safety fine-tuning constraints the strategic action space of these LLMs, making them unable to fully engage in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy.
查看缓存全文
缓存时间: 2026/09/25 09:14
# Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks Source: [https://arxiv.org/abs/2609.28673](https://arxiv.org/abs/2609.28673) [View PDF](https://arxiv.org/pdf/2609.28673) > Abstract:Large Language Models \(LLMs\) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors\. In this study, we focus on character attacks \(ad hominem arguments\), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content\. Specifically, we investigate whether modern LLMs can replicate human competence to strategically use and respond to such attacks\. We analyse a corpus of natural language political dialogues to identify defensive strategies human interlocutors naturally employ in ethos\-centred debates and structure them into a dialogue game\. Empirically, we benchmark LLM\-generated dialogues against the ElecDeb60to16\-fallacy corpus of U\.S\. presidential debates, contrasting human debaters' repertoire of defensive strategies with those of artificial agents\. Results reveal a substantial difference: most LLMs rigidly prioritise logical defences, failing to exploit ethotic counterattacks as valid moves in political discourse\. We argue that current safety fine\-tuning constraints the strategic action space of these LLMs, making them unable to fully engage in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy\. ## Submission history From: Ewelina Gajewska \[[view email](https://arxiv.org/show-email/4a9da5a9/2609.28673)\] **\[v1\]**Wed, 23 Sep 2026 18:14:32 UTC \(103 KB\)
相似文章
鲁棒批评者:防御LLMs免受多轮攻击
本文提出对话批评者引导采样(DCGS)框架,通过从对话历史推断用户意图,并利用基于价值/遗憾的批评者对回复进行评分,在不进行微调的情况下提高鲁棒性,从而防御LLMs免受多轮对抗性攻击。
通过多轮对话说服评估大语言模型的事实鲁棒性
本文提出SAST-IR框架,用于评估大语言模型在面对说服性攻击时的事实鲁棒性,揭示了高攻击成功率以及防御策略中的复杂性悖论。
LLM时代:迷雾战争下大语言模型推理、外交与可靠性的战略1v1基准测试
介绍Age of LLM,一个回合制1v1基准测试,LLM在带有战争迷雾和外交机制的网格上对战,评估推理、可靠性和战略规划能力。结果显示核速攻战术占主导,且可靠性与获胜之间存在弱关联。
LLM基准测试
一项关于大语言模型基准测试的研究或报告,可能比较了各种任务上的表现。
超越检测:在实时逐轮交互中评估防御性LLM对抗AI生成的社会工程攻击
本文研究防御性大语言模型能否识别AI生成的社会工程中的结构性风险来源,引入了信任链定位和包含300个案例的语料库。在实时逐轮和静态场景下评估五种模型后发现,仅依赖看似安全的行为是不够的;干预率差异显著,且结构性定位往往与保护性行动脱钩。