Tag
ReLaG is a scalable, modality-agnostic framework that infers groups of related samples via proximity graphs and community detection to produce independent train–test splits, addressing the over-optimistic generalization estimates caused by random splits on data with latent relations. It scales better than existing relation-aware methods and offers a label-free procedure to adapt splitting resolution to production settings, available as pip-installable open-source software.
This paper systematically evaluates five train-test splitting strategies for AutoML, showing that geometry-based methods are less effective than random/stratified splitting in preserving distributional similarity, and proposes an Optimised-Distribution method that achieves 89% similarity.