AI Evaluation Should Work With Humans
Summary
This position paper argues that AI evaluation should pivot to assessing human-AI teams rather than superhuman performance to foster better societal outcomes.
View Cached Full Text
Cached at: 08/17/26, 09:38 AM
# AI Evaluation Should Work With Humans Source: [https://arxiv.org/abs/2608.13577](https://arxiv.org/abs/2608.13577) [View PDF](https://arxiv.org/pdf/2608.13577) > Abstract:This position paper argues that the dominant paradigm of AI evaluation \(which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans\) is guiding AI development in the wrong direction\. Instead, the AI community should pivot to evaluating the performance of human\-\-AI teams\. We argue that this collaborative shift will foster AI systems that act as true complements to human capabilities and therefore lead to far better societal outcomes than will the current process\. ## Submission history From: Jan Kulveit \[[view email](https://arxiv.org/show-email/90b83c8c/2608.13577)\] **\[v1\]**Mon, 6 Jul 2026 21:14:15 UTC \(2,904 KB\)
Similar Articles
Why Do We Measure AI Progress by What It Can Replace, Rather Than What It Enables Humans to Do?
This article critiques the dominant focus on AI replacing human roles and proposes evaluating AI progress by how it augments human capabilities, enabling people to achieve tasks previously beyond their reach.
Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era
This paper proposes TEAM-Design, a budgeted rule for allocating replay tasks to evaluate human-AI workflow effectiveness compared to human-only or agent-only alternatives, with applications in clinical and coding settings.
Towards an Argumentative Foundation for Evaluative AI
This position paper advocates computational argumentation as a formal foundation for Evaluative AI, which supports human decision-making by presenting competing hypotheses with evidence for and against, rather than single recommendations.
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
This position paper argues that AI agents in scientific teams should be studied as human-agent systems to enhance collaboration and mitigate risks such as reduced diversity in scientific inquiry.
We are building more intelligent AI agents. But are humans building better ways to decide when they should act?
The article reflects on the evolving relationship between humans and AI, arguing that as AI becomes more autonomous, the key challenge is understanding human decision-making and purpose, rather than just technical capability. It suggests shifting from using AI as a mere tool to collaborating with it as an instrument that enhances human judgment.