Tag
This paper proposes NxN E-valuation, an e-value-based hypothesis certification algorithm that uses a large training set to let samples serve as null hypotheses for one another, enabling conditional randomization tests to certify LLM-proposed hypotheses without bespoke statistical procedures.