Tag
This paper proposes Joint Affine Spectral Shaping (JRI), which extends weight-only spectral optimizers like Muon to jointly update weight and bias in affine layers, showing small but consistent accuracy improvements on a BERT-mini IMDb classification task.
This paper compares multiple machine learning and transformer models for sentiment classification on movie reviews, finding RoBERTa achieves 93.02% accuracy, and a soft voting ensemble improves performance.