谁赋予了AI中的‘智能’?机器自我报告的溯源与可采性

arXiv cs.CL 论文

摘要

本学术论文探讨了AI系统中机器自我报告的溯源与可采性,讨论了对信任、透明度和法律情境的影响。

arXiv:2609.29494v1 Announce Type: new Abstract: Large language models make statements concerning their own "minds". When asked whether or not they are conscious, they usually say that they are not; if they are prompted to ignore their guidelines, they might say that they are; and if asked to write a diary from their point of view, they often describe a human lifestyle. All these contradictory ways of describing themselves are the result of the way the questions are phrased. This paper shows exactly where such descriptions came from, and considers when they can be regarded as evidence for what they claim to report. In order to achieve this, we traced the provenance from end to end. We examine Pythia and OLMo 2 across 66 pretraining checkpoints, three of the post-training stages of OLMo 2 that have been released, about 90,000 continuations, and four training corpora. A set of forty items is used in order to keep an eye on self-reference, frame sensitivity, and self-ascription throughout training. The denial formula was almost completely missing from the vast quantity of text that the models initially came across, but was present in a dense manner in the small, carefully chosen set of example dialogues that they were trained on later on. Supervised fine-tuning causes first-person AI language to become the default, and the other affirmations are then suppressed using preference optimization. The final policy is still very sensitive to framing and to the chat template itself. Two of the conditions which are set out in the epistemology of testimony determine whether or not these outputs can act as evidence for what they claim to report: reference and causation. Reports produced by the base model fail the reference condition, and those obtained after training remain sensitive to the frame and do not show state dependence. The result is symmetric in that trained denials are no more admissible than trained affirmations.
查看原文
查看缓存全文

缓存时间: 2026/09/25 09:22

# Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report
Source: [https://arxiv.org/abs/2609.29494](https://arxiv.org/abs/2609.29494)
Bibliographic Tools

## Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Code, Data, Media

## Code, Data and Media Associated with this Article

Demos

## Demos

Related Papers

## Recommenders and Search Tools

About arXivLabs

## arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\.

Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.

相似文章

超越“AI制作”:可视化出处密度以缓解透明度惩罚

arXiv cs.AI

本研究论文提出了“出处密度”(Provenance Density),一种证据可视化界面,通过展示经验证的声明,帮助用户区分AI生成的文本,缓解“流畅性陷阱”(Fluency Trap),即用户信任流畅但虚假的内容。

负责任的代理AI需要显式溯源

arXiv cs.AI

本文认为,在整个代理AI生命周期的显式溯源是使责任可计算和可操作的结构性必要条件,解决了自主组合中涌现危害的责任缺口。