标签
This paper introduces a reproducible auditing framework for detecting systematic political preferences in LLMs, demonstrated through an Italian case study evaluating parties and leaders across nine criteria.
一项独立分析对 100 多个大语言模型进行了 117 个政治问题的测试,以绘制其意识形态倾向图谱,结果显示 DeepSeek 和 Grok 偏向左翼,而大多数其他模型则聚集在中间或右翼。