Tag
This paper audits temporal leakage in financial news NLP benchmarks across multiple models, finding that random splits inflate performance metrics and identifies M&A events as a category with a localized positive signal under chronological evaluation.
This paper shows that the standard pre/post training-cutoff check for temporal leakage in LLM backtesting is uninformative, as recency effects mimic leakage. It proposes new estimators using known cutoffs and matched clean controls to measure leakage and compute adjusted scores, validated on frontier models.