Tag
This paper demonstrates that careful preprocessing—especially context length selection, normalization, and regularization—can make simple linear models like Ridge regression competitive with or superior to large Transformer, MLP, and CNN models on time-series forecasting benchmarks.
This paper argues that measurement noise, not model inadequacy, explains why nonlinear models often fail to outperform linear regression in biomedical prediction, as noise attenuates nonlinear structure faster than linear structure, a limitation that cannot be overcome by more data or model complexity.