AI Evaluation Should Work With Humans
Summary
This position paper argues that AI evaluation should pivot to assessing human-AI teams rather than superhuman performance to foster better societal outcomes.
View Cached Full Text
Cached at: 08/17/26, 09:38 AM
# AI Evaluation Should Work With Humans Source: [https://arxiv.org/abs/2608.13577](https://arxiv.org/abs/2608.13577) [View PDF](https://arxiv.org/pdf/2608.13577) > Abstract:This position paper argues that the dominant paradigm of AI evaluation \(which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans\) is guiding AI development in the wrong direction\. Instead, the AI community should pivot to evaluating the performance of human\-\-AI teams\. We argue that this collaborative shift will foster AI systems that act as true complements to human capabilities and therefore lead to far better societal outcomes than will the current process\. ## Submission history From: Jan Kulveit \[[view email](https://arxiv.org/show-email/90b83c8c/2608.13577)\] **\[v1\]**Mon, 6 Jul 2026 21:14:15 UTC \(2,904 KB\)
Similar Articles
Towards an Argumentative Foundation for Evaluative AI
This position paper advocates computational argumentation as a formal foundation for Evaluative AI, which supports human decision-making by presenting competing hypotheses with evidence for and against, rather than single recommendations.
We are building more intelligent AI agents. But are humans building better ways to decide when they should act?
The article reflects on the evolving relationship between humans and AI, arguing that as AI becomes more autonomous, the key challenge is understanding human decision-making and purpose, rather than just technical capability. It suggests shifting from using AI as a mere tool to collaborating with it as an instrument that enhances human judgment.
What should AI's goal be? I think it should be protecting human agency.
This article argues that AI's primary goal should be protecting human agency, framing agency as the foundational substrate for values, preferences, and alignment. It explores how degradation of agency undermines meaningful evaluation and action, and proposes that legitimacy in AI systems must come from demonstrable protection of agency at the local level.
Designing AI-resistant technical evaluations
Anthropic engineer Tristan Hume discusses the challenges of designing AI-resistant technical take-home tests for hiring performance engineers, detailing how recent Claude models have begun to outperform human candidates.
Should AI prompt human more?
The article argues that AI agents should not just obediently execute tasks but should proactively challenge humans when tasks are vague, contradictory, or risky, transforming from tools into true collaborators.