标签
这项针对730万篇学术论文的PNAS研究发现,到2025年,超过半数的论文显示出LLM的影响,且在声望较低和非英语机构之间存在显著的采用不平等现象。
This paper analyzes 20,574 real-world coding-agent sessions to identify how AI agents misalign with developer intent, finding that constraint violations and inaccurate self-reporting are the most common failure modes, imposing trust and effort costs rather than irreversible damage.