Artificial Adaptive Intelligence:狭义与通用智能之间的缺失阶段

arXiv cs.AI 论文

摘要

本文提出Artificial Adaptive Intelligence (AAI)作为狭义与通用AI之间的缺失中间阶段,认为自适应性(而非规模或通用性)是弥合这一差距的关键能力。

arXiv:2605.16844v1 Announce Type: new Abstract: 在目前部署的狭义系统与我们推测的通用智能之间,存在着一整类机器行为,它们从未被单独命名。本文认为这一领域并非空白:元学习、神经架构搜索、AutoML、持续学习、进化计算和物理信息建模已悄然汇聚到一个共同原则之上,即逐步将人类从参数规范循环中移除。我们称这一领域为Artificial Adaptive Intelligence (AAI),并给出操作性定义:一个系统展现出AAI的程度取决于它无需人工指定可调超参数,同时能在多样化的任务分布上保持有竞争力的性能。为了使定义量化,我们引入一个自适应指数,该指数沿与规模正交的轴测量进展,结合了系统吸收的超参数比例与针对任务特化基线的性能比。我们发展了参数极小性原理,并将其植根于最小描述长度框架中,表明适当的超参数数量由数据决定而非设计者决定。然后,我们围绕三条通向极小性的路径组织该领域:数据与任务感知配置、结构与进化变形、以及训练中的自适应性。我们分析了它们的稳定性、收敛性和治理影响,并通过航空航天设计、金融制度检测、湍流建模、生态动力学和视觉-语言系统等案例研究加以说明。本文的论点是,从ANI到AGI的路径必然经过AAI,对这一阶段的命名将会改变我们的度量方式、建造方式以及成功标准的定义。
查看原文
查看缓存全文

缓存时间: 2026/05/19 06:36

# 内容来源:https://arxiv.org/html/2605.16844 一本专著  
人工自适应智能  
狭义与通用智能之间缺失的阶段  
Boris Kriuk  
2026  

**人工自适应智能:狭义与通用智能之间缺失的阶段**  
*Artificial Adaptive Intelligence: The Missing Stage Between Narrow and General Intelligence*  

Boris Kriuk  
一本独立专著  
2026  

**人工自适应智能:狭义与通用智能之间缺失的阶段**  
版权所有 © 2026 Boris Kriuk。保留所有权利。未经作者事先书面许可,不得以任何形式或通过任何方式复制、存储于检索系统或传播本出版物的任何部分,但嵌入在评论或评论中的简短引用除外。  
第一版,2026年。  
作者使用 LaTeX 排版,采用 Latin Modern 字体家族。  
本书中的观点和结论仅代表作者个人,不一定反映任何附属机构的立场。  
作者联系方式:Boris Kriuk  

献给那些相信智能不是规模,而是适应的人。  

> “一个系统具有适应性,不是因为它学得更多,而是因为它需要被告诉的更少。”  
> ——作者  

###### 目录  

1. 前言(https://arxiv.org/html/2605.16844#Chx1)  
2. 如何阅读本书(https://arxiv.org/html/2605.16844#Chx2)  
3. **第一部分:基础与动机**(https://arxiv.org/html/2605.16844#Pt1)  
   1. 1 引言(https://arxiv.org/html/2605.16844#Ch1)  
      1. 1.1 ANI–AGI 差距与“仅靠规模”为何不足以架起桥梁(https://arxiv.org/html/2605.16844#Ch1.S1)  
      2. 1.2 现代机器学习的隐性成本:超参数、架构先验与人在回路调优(https://arxiv.org/html/2605.16844#Ch1.S2)  
      3. 1.3 论点陈述:适应性——而非规模或通用性——才是缺失的中间能力(https://arxiv.org/html/2605.16844#Ch1.S3)  
      4. 1.4 定义人工自适应智能(https://arxiv.org/html/2605.16844#Ch1.S4)  
      5. 1.5 范围、贡献与专著路线图(https://arxiv.org/html/2605.16844#Ch1.S5)  
   2. 2 智能范式分类学(https://arxiv.org/html/2605.16844#Ch2)  
      1. 2.1 人工狭义智能:优势、脆弱性与调优负担(https://arxiv.org/html/2605.16844#Ch2.S1)  
         1. 2.1.1 狭义的实际含义(https://arxiv.org/html/2605.16844#Ch2.S1.SS1)  
         2. 2.1.2 ANI 的真正优势(https://arxiv.org/html/2605.16844#Ch2.S1.SS2)  
         3. 2.1.3 脆弱性问题(https://arxiv.org/html/2605.16844#Ch2.S1.SS3)  
         4. 2.1.4 调优负担(https://arxiv.org/html/2605.16844#Ch2.S1.SS4)  
      2. 2.2 人工通用智能:定义、开放问题与为何直接追求为时过早(https://arxiv.org/html/2605.16844#Ch2.S2)  
         1. 2.2.1 四类定义(https://arxiv.org/html/2605.16844#Ch2.S2.SS1)  
         2. 2.2.2 先于 AGI 的开放问题(https://arxiv.org/html/2605.16844#Ch2.S2.SS2)  
         3. 2.2.3 为何直接追求为时过早(https://arxiv.org/html/2605.16844#Ch2.S2.SS3)  
      3. 2.3 定位 AAI:操作性定义与区分标准(https://arxiv.org/html/2605.16844#Ch2.S3)  
         1. 2.3.1 操作性定义(https://arxiv.org/html/2605.16844#Ch2.S3.SS1)  
         2. 2.3.2 适应性指数,重述(https://arxiv.org/html/2605.16844#Ch2.S3.SS2)  
         3. 2.3.3 四个区分标准(https://arxiv.org/html/2605.16844#Ch2.S3.SS3)  
         4. 2.3.4 AAI 不是什么(https://arxiv.org/html/2605.16844#Ch2.S3.SS4)  
      4. 2.4 相关概念:元学习、AutoML、持续学习、神经架构搜索(https://arxiv.org/html/2605.16844#Ch2.S4)  
         1. 2.4.1 元学习(https://arxiv.org/html/2605.16844#Ch2.S4.SS1)  
         2. 2.4.2 AutoML(https://arxiv.org/html/2605.16844#Ch2.S4.SS2)  
         3. 2.4.3 持续学习(https://arxiv.org/html/2605.16844#Ch2.S4.SS3)  
         4. 2.4.4 神经架构搜索(https://arxiv.org/html/2605.16844#Ch2.S4.SS4)  
         5. 2.4.5 基础模型作为特例(https://arxiv.org/html/2605.16844#Ch2.S4.SS5)  
      5. 2.5 适应性的评估轴(https://arxiv.org/html/2605.16844#Ch2.S5)  
         1. 2.5.1 轴1:参数最小性(https://arxiv.org/html/2605.16844#Ch2.S5.SS1)  
         2. 2.5.2 轴2:可迁移性(https://arxiv.org/html/2605.16844#Ch2.S5.SS2)  
         3. 2.5.3 轴3:体制感知(https://arxiv.org/html/2605.16844#Ch2.S5.SS3)  
         4. 2.5.4 轴4:自重构(https://arxiv.org/html/2605.16844#Ch2.S5.SS4)  
         5. 2.5.5 联合报告与适应性轮廓(https://arxiv.org/html/2605.16844#Ch2.S5.SS5)  
   3. 3 参数最小性原则(https://arxiv.org/html/2605.16844#Ch3)  
      1. 3.1 为何参数数量是人类依赖度的代理(https://arxiv.org/html/2605.16844#Ch3.S1)  
         1. 3.1.1 问题背后的问题(https://arxiv.org/html/2605.16844#Ch3.S1.SS1)  
         2. 3.1.2 超参数作为潜在的人类决策(https://arxiv.org/html/2605.16844#Ch3.S1.SS2)  
         3. 3.1.3 架构选择作为离散超参数(https://arxiv.org/html/2605.16844#Ch3.S1.SS3)  
         4. 3.1.4 代理关系(https://arxiv.org/html/2605.16844#Ch3.S1.SS4)  
         5. 3.1.5 代理的不完美性(https://arxiv.org/html/2605.16844#Ch3.S1.SS5)  
         6. 3.1.6 反对高表面积不可避免论(https://arxiv.org/html/2605.16844#Ch3.S1.SS6)  
      2. 3.2 将最小调优形式化为一项目标(https://arxiv.org/html/2605.16844#Ch3.S2)  
         1. 3.2.1 从代理到原则(https://arxiv.org/html/2605.16844#Ch3.S2.SS1)  
         2. 3.2.2 递归与元超参数问题(https://arxiv.org/html/2605.16844#Ch3.S2.SS2)  
         3. 3.2.3 目标的性质(https://arxiv.org/html/2605.16844#Ch3.S2.SS3)  
         4. 3.2.4 与性能和可靠性的权衡(https://arxiv.org/html/2605.16844#Ch3.S2.SS4)  
         5. 3.2.5 原则的陈述(https://arxiv.org/html/2605.16844#Ch3.S2.SS5)  
      3. 3.3 实现最小性的三条路径(https://arxiv.org/html/2605.16844#Ch3.S3)  
         1. 3.3.1 从原则到实践(https://arxiv.org/html/2605.16844#Ch3.S3.SS1)  
         2. 3.3.2 路径I:数据和任务感知的配置(https://arxiv.org/html/2605.16844#Ch3.S3.SS2)  
         3. 3.3.3 路径II:结构与进化变形(https://arxiv.org/html/2605.16844#Ch3.S3.SS3)  
         4. 3.3.4 路径III:训练中的自适应性(https://arxiv.org/html/2605.16844#Ch3.S3.SS4)  
         5. 3.3.5 路径之间的关系(https://arxiv.org/html/2605.16844#Ch3.S3.SS5)  
         6. 3.3.6 映射本书(https://arxiv.org/html/2605.16844#Ch3.S3.SS6)  
      4. 3.4 信息论视角:匹配模型复杂度与数据复杂度(https://arxiv.org/html/2605.16844#Ch3.S4)  
         1. 3.4.1 为何是信息论(https://arxiv.org/html/2605.16844#Ch3.S4.SS1)  
         2. 3.4.2 最小描述长度原则(https://arxiv.org/html/2605.16844#Ch3.S4.SS2)  
         3. 3.4.3 最优超参数数量(https://arxiv.org/html/2605.16844#Ch3.S4.SS3)  
         4. 3.4.4 对三条路径的启示(https://arxiv.org/html/2605.16844#Ch3.S4.SS4)  
         5. 3.4.5 信息论视角的局限(https://arxiv.org/html/2605.16844#Ch3.S4.SS5)  
         6. 3.4.6 综合(https://arxiv.org/html/2605.16844#Ch3.S4.SS6)  
4. **第二部分:配置与构建**(https://arxiv.org/html/2605.16844#Pt2)  
   1. 4 训练前与训练过程中的适应性(https://arxiv.org/html/2605.16844#Ch4)  
      1. 4.1 分布与任务级诊断(https://arxiv.org/html/2605.16844#Ch4.S1)  
         1. 4.1.1 分布指纹(https://arxiv.org/html/2605.16844#Ch4.S1.SS1)  
         2. 4.1.2 随机矩阵理论与相关结构(https://arxiv.org/html/2605.16844#Ch4.S1.SS2)  
         3. 4.1.3 重尾、非平稳性与体制转变(https://arxiv.org/html/2605.16844#Ch4.S1.SS3)  
         4. 4.1.4 从诊断到模型选择(https://arxiv.org/html/2605.16844#Ch4.S1.SS4)  
      2. 4.2 任务嵌入与架构先验(https://arxiv.org/html/2605.16844#Ch4.S2)  
         1. 4.2.1 任务相似性(https://arxiv.org/html/2605.16844#Ch4.S2.SS1)  
         2. 4.2.2 任务规范作为生成先验(https://arxiv.org/html/2605.16844#Ch4.S2.SS2)  
         3. 4.2.3 归纳偏置作为可配置量(https://arxiv.org/html/2605.16844#Ch4.S2.SS3)  
         4. 4.2.4 多任务与统一目标(https://arxiv.org/html/2605.16844#Ch4.S2.SS4)  
      3. 4.3 受生物启发与物理启发的适应性(https://arxiv.org/html/2605.16844#Ch4.S3)  
         1. 4.3.1 来自生物学的启示(https://arxiv.org/html/2605.16844#Ch4.S3.SS1)  
         2. 4.3.2 表观遗传机制作为计算隐喻(https://arxiv.org/html/2605.16844#Ch4.S3.SS2)  
         3. 4.3.3 神经进化与发育编码(https://arxiv.org/html/2605.16844#Ch4.S3.SS3)  
         4. 4.3.4 群体、免疫系统与生态系统类比(https://arxiv.org/html/2605.16844#Ch4.S3.SS4)  
         5. 4.3.5 物理定律作为通用正则化器(https://arxiv.org/html/2605.16844#Ch4.S3.SS5)  
         6. 4.3.6 大规模混合物理-机器学习架构(https://arxiv.org/html/2605.16844#Ch4.S3.SS6)  
         7. 4.3.7 谱与结构物理先验(https://arxiv.org/html/2605.16844#Ch4.S3.SS7)  
         8. 4.3.8 物理何时减少 vs. 替换(https://arxiv.org/html/2605.16844#Ch4.S3.SS8)  
      4. 4.4 变形架构(https://arxiv.org/html/2605.16844#Ch4.S4)  
         1. 4.4.1 静态 vs. 动态结构假设(https://arxiv.org/html/2605.16844#Ch4.S4.SS1)  
         2. 4.4.2 自组织树、图与网络(https://arxiv.org/html/2605.16844#Ch4.S4.SS2)  
         3. 4.4.3 保持拓扑的变换(https://arxiv.org/html/2605.16844#Ch4.S4.SS3)  
         4. 4.4.4 训练中的生长、剪枝与重构(https://arxiv.org/html/2605.16844#Ch4.S4.SS4)  
      5. 4.5 进化与基于种群的方法(https://arxiv.org/html/2605.16844#Ch4.S5)  
         1. 4.5.1 进化搜索作为无参数优化器(https://arxiv.org/html/2605.16844#Ch4.S5.SS1)  
         2. 4.5.2 混合梯度与进化更新(https://arxiv.org/html/2605.16844#Ch4.S5.SS2)  
         3. 4.5.3 质量-多样性与开放式搜索(https://arxiv.org/html/2605.16844#Ch4.S5.SS3)  
         4. 4.5.4 跨领域案例研究(https://arxiv.org/html/2605.16844#Ch4.S5.SS4)  
      6. 4.6 任务调节动力学(https://arxiv.org/html/2605.16844#Ch4.S6)  
         1. 4.6.1 异构任务变形(https://arxiv.org/html/2605.16844#Ch4.S6.SS1)  
         2. 4.6.2 无需调度的跨域迁移(https://arxiv.org/html/2605.16844#Ch4.S6.SS2)  
         3. 4.6.3 模块化组合与分解(https://arxiv.org/html/2605.16844#Ch4.S6.SS3)  
      7. 4.7 总结与桥梁(https://arxiv.org/html/2605.16844#Ch4.S7)  
   2. 5 自适应性、稳定性与超越之路(https://arxiv.org/html/2605.16844#Ch5)  
      1. 5.1 梯度流作为适应性信号(https://arxiv.org/html/2605.16844#Ch5.S1)  
         1. 5.1.1 梯度流的诊断性使用(https://arxiv.org/html/2605.16844#Ch5.S1.SS1)  
         2. 5.1.2 损失景观几何与自适应步长选择(https://arxiv.org/html/2605.16844#Ch5.S1.SS2)  
         3. 5.1.3 高维平滑(https://arxiv.org/html/2605.16844#Ch5.S1.SS3)  
         4. 5.1.4 注意力与聚焦作为自适应门控(https://arxiv.org/html/2605.16844#Ch5.S1.SS4)  
      2. 5.2 独立的训练中更新(https://arxiv.org/html/2605.16844#Ch5.S2)  
         1. 5.2.1 超参数的自修改(https://arxiv.org/html/2605.16844#Ch5.S2.SS1)  
         2. 5.2.2 在线调度自适应(https://arxiv.org/html/2605.16844#Ch5.S2.SS2)  
         3. 5.2.3 自门控、自剪枝与自扩展网络(https://arxiv.org/html/2605.16844#Ch5.S2.SS3)  
         4. 5.2.4 数据与优化器之间的反馈循环(https://arxiv.org/html/2605.16844#Ch5.S2.SS4)  
      3. 5.3 通信与资源适应性(https://arxiv.org/html/2605.16844#Ch5.S3)  
         1. 5.3.1 权重之外的适应性(https://arxiv.org/html/2605.16844#Ch5.S3.SS1)  
         2. 5.3.2 压缩作为适应性(https://arxiv.org/html/2605.16844#Ch5.S3.SS2)  
         3. 5.3.3 多智能体与分布式适应性(https://arxiv.org/html/2605.16844#Ch5.S3.SS3)  
      4. 5.4 稳定性、安全性与收敛性(https://arxiv.org/html/2605.16844#Ch5.S4)  
         1. 5.4.1 自修改的风险(https://arxiv.org/html/2605.16844#Ch5.S4.SS1)  
         2. 5.4.2 收敛性分析(https://arxiv.org/html/2605.16844#Ch5.S4.SS2)  
         3. 5.4.3 适应性下的可解释性与可审计性(https://arxiv.org/html/2605.16844#Ch5.S4.SS3)  
      5. 5.5 跨领域案例研究(https://arxiv.org/html/2605.16844#Ch5.S5)  
         1. 5.5.1 航空航天与自主系统(https://arxiv.org/html/2605.16844#Ch5.S5.SS1)  
         2. 5.5.2 金融系统(https://arxiv.org/html/2605.16844#Ch5.S5.SS2)  
         3. 5.5.3 地球物理与气候(https://arxiv.org/html/2605.16844#Ch5.S5.SS3)  
         4. 5.5.4 视觉与语言(https://arxiv.org/html/2605.16844#Ch5.S5.SS4)  
         5. 5.5.5 跨领域模式(https://arxiv.org/html/2605.16844#Ch5.S5.SS5)  
      6. 5.6 AAI 在机器智能轨迹中的位置(https://arxiv.org/html/2605.16844#Ch5.S6)  
         1. 5.6.1 重访论点(https://arxiv.org/html/2605.16844#Ch5.S6.SS1)  
         2. 5.6.2 AAI 能做到 ANI 不能做到的(https://arxiv.org/html/2605.16844#Ch5.S6.SS2)  
         3. 5.6.3 与 AGI 相比 AAI 仍欠缺的(https://arxiv.org/html/2605.16844#Ch5.S6.SS3)  
         4. 5.6.4 成熟度模型(https://arxiv.org/html/2605.16844#Ch5.S6.SS4)  
      7. 5.7 开放问题与研究议程(https://arxiv.org/html/2605.16844#Ch5.S7)  
         1. 5.7.1 超越准确率的基准(https://arxiv.org/html/2605.16844#Ch5.S7.SS1)  
         2. 5.7.2 理论基础(https://arxiv.org/html/2605.16844#Ch5.S7.SS2)  
         3. 5.7.3 计算、能源与可持续性(https://arxiv.org/html/2605.16844#Ch5.S7.SS3)  
         4. 5.7.4 自修改系统的伦理与治理(https://arxiv.org/html/2605.16844#Ch5.S7.SS4)  
      8. 5.8 结论(https://arxiv.org/html/2605.16844#Ch5.S8)  
3. 参考文献(https://arxiv.org/html/2605.16844#bib)  

### 前言  

本书旨在为人工智能研究阴影中已显现十余年、却从未获得专属名称的事物命名。在我们已构建的狭义系统与我们所设想的通用智能之间,存在着一个完整的机器行为体制——它既不受任务束缚,也非真正自主——在这个体制中,系统会根据摆在面前的问题进行适应、重构并选择自身的结构。我将这个体制称为**人工自适应智能**(Artificial Adaptive Intelligence,简称AAI)。  

本书的动机并非因为自适应方法是新的。元学习、AutoML、神经架构搜索、持续学习、进化计算以及越来越多的自调优技术多年来一直朝着这个方向推进。动机在于,这些努力共享一个共同原则,而该领域尚未清晰阐述:即逐步将人类从参数规范的回路中移除。每一种技术,以其自身的方式,都是朝着需要更少被告知、更多去推断的系统迈出的一步。  

为了使这一原则精确化,本书采用了一个单一的、刻意尖锐的定义:  

> 一个系统展现出人工自适应智能的程度,取决于它在无需人类指定的可调超参数的情况下,能否在多样化的任务分布上保持竞争性性能。  

这一定义有意让人感到不适。它是可证伪的。它是可计数的。它拒绝通常对“灵活性”或“鲁棒性”的模糊诉求所产生的庇护。并且它在机器智能的四个阶段之间划出了一条清晰的界限:狭义系统,向用户暴露众多超参数;自动系统,通过元算法调优这些超参数,而这些元算法本身又暴露超参数;自适应系统,不暴露任何超参数;以及通用系统,更进一步,形成自身的目标。  

本书围绕这条主线组织。每一章都提出同样的根本问题——*该技术消除了哪个超参数,以及如何消除?*——并借助这个问题,将来自航空航天、金融、地球物理、视觉与语言建模等领域的工作汇聚在一起,这些工作原本会分属于互不关联的文献。其抱负并非声称...

相似文章

观点:Agentic AI系统是实现AGI的可预见路径

arXiv cs.AI

本文认为,单一模型的单体型扩展不足以实现AGI,并提出具有多智能体协作的Agentic AI是必要的范式,理论上证明了代理系统在泛化和样本效率上具有指数级优势。

# 数字学徒:人类主导的智能体AI开发框架

arXiv cs.AI

本文介绍了"数字学徒"(Digital Apprentice)框架——一个可扩展且安全的智能体 AI 体系,其中自主权通过观察学习、人工授权和持续对齐校正的方式逐步获得。本文还介绍了 ADAPT,一种推理时控制平面,用于将渐进式自主权等级付诸实践,并将人工校正转化为可复用的偏好数据。

真正的AGI

Reddit r/singularity

这篇文章讨论了在人工智能中实现真正AGI的概念和潜在突破。

从AGI到ASI

Hugging Face Daily Papers

本文探讨了从通用人工智能到超级人工智能的潜在路径,包括规模扩展、范式转变、递归改进及多智能体集体,并强调需通过跨学科全球协作应对变革性社会影响。