Tag
The paper curates and releases open datasets and a model for Armenian, demonstrating that continued pretraining with news and STEM data improves performance and addresses data scarcity in low-resource NLP.