Tag
The tweet critiques traditional AI benchmarks and introduces TRACES, a new benchmark that evaluates AI's discovery process by focusing on how models reach answers, including tool usage, error correction, and evidence tracing.
This paper introduces Apodex Discovery, a framework and reality benchmark (TRACES) for evaluating and building 'discoverative AI'—AI systems that investigate open-ended real-world problems. The proposed heavy-duty solvers outperform published state-of-the-art on AAV capsid design and improve drug repurposing prediction scores.