Tag
The paper shows that incomplete reference sets in open-ended Theory-of-Mind tracking can reverse calibration rankings, and proposes methods to correct evaluations using human pilots.