谁赋予了AI中的‘智能’?机器自我报告的溯源与可采性
摘要
本学术论文探讨了AI系统中机器自我报告的溯源与可采性,讨论了对信任、透明度和法律情境的影响。
arXiv:2609.29494v1 Announce Type: new
Abstract: Large language models make statements concerning their own "minds". When asked whether or not they are conscious, they usually say that they are not; if they are prompted to ignore their guidelines, they might say that they are; and if asked to write a diary from their point of view, they often describe a human lifestyle. All these contradictory ways of describing themselves are the result of the way the questions are phrased. This paper shows exactly where such descriptions came from, and considers when they can be regarded as evidence for what they claim to report.
In order to achieve this, we traced the provenance from end to end. We examine Pythia and OLMo 2 across 66 pretraining checkpoints, three of the post-training stages of OLMo 2 that have been released, about 90,000 continuations, and four training corpora. A set of forty items is used in order to keep an eye on self-reference, frame sensitivity, and self-ascription throughout training. The denial formula was almost completely missing from the vast quantity of text that the models initially came across, but was present in a dense manner in the small, carefully chosen set of example dialogues that they were trained on later on. Supervised fine-tuning causes first-person AI language to become the default, and the other affirmations are then suppressed using preference optimization. The final policy is still very sensitive to framing and to the chat template itself.
Two of the conditions which are set out in the epistemology of testimony determine whether or not these outputs can act as evidence for what they claim to report: reference and causation. Reports produced by the base model fail the reference condition, and those obtained after training remain sensitive to the frame and do not show state dependence. The result is symmetric in that trained denials are no more admissible than trained affirmations.
查看缓存全文
缓存时间: 2026/09/25 09:22
# Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report Source: [https://arxiv.org/abs/2609.29494](https://arxiv.org/abs/2609.29494) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
相似文章
真正让你信任AI的是什么?不是“听起来正确”,而是像信任一个人或一个机构那样信任它?
一场讨论,探讨哪些具体条件(透明度、可验证的记录、持久的身份、可问责性)能让人们像信任人类或机构一样信任AI系统,而不仅仅是将其视为工具。
AI代理应如何证明其代表身份?
本文探讨了AI代理验证身份并证明其所代表对象的方法,解决了自主系统中的关键信任和安全挑战。
最近写了一篇文章,主张AI必须表明自己非人类的身份。希望得到社区的反馈!
主张在法律上要求AI在与人类互动时披露其非人属性,引用了科罗拉多州的新法律,以及更广泛的AI决策透明度的必要性。
超越“AI制作”:可视化出处密度以缓解透明度惩罚
本研究论文提出了“出处密度”(Provenance Density),一种证据可视化界面,通过展示经验证的声明,帮助用户区分AI生成的文本,缓解“流畅性陷阱”(Fluency Trap),即用户信任流畅但虚假的内容。
负责任的代理AI需要显式溯源
本文认为,在整个代理AI生命周期的显式溯源是使责任可计算和可操作的结构性必要条件,解决了自主组合中涌现危害的责任缺口。