Tag
Research indicates that discarding singleton proteins in metagenomic datasets during the training of protein language models is likely incorrect, based on findings from a collaborative study.
This paper proposes a fine-grained taxonomy for curriculum learning in NLP, separating difficulty evaluation from training scheduling to enable systematic analysis and comparison of CL strategies. It identifies an incomparability problem in prior work and provides a framework for designing and evaluating CL approaches.