Tag
The author criticizes the broken system of measuring AI models against test harnesses due to biases and lack of standardization, while speculating that models like Claude may soon generate harnesses dynamically.