Tag
This paper investigates whether temporally drifting data streams can be partitioned into discrete regimes by fitting a hidden Markov model to the trajectory of neural network weights trained on successive time windows, showing that recovered latent states correlate with transfer performance across two datasets.
This paper presents an empirical study comparing how different neural architectures (MLPs, CNNs, RNNs, pretrained transformers) degrade under temporal distribution shift across image and text domains, finding that models exploiting localized features degrade fastest while pretrained encoders drift more gradually.