Language Models for Portuguese: A Systematic Mapping Study

arXiv cs.CL Papers

Summary

This survey systematically maps 46 language models for Portuguese, analyzing their development, characteristics, evolution, and identifying research gaps and future directions in the field.

arXiv:2608.18138v1 Announce Type: new Abstract: In recent years, the rapid development of language models has transformed the field of Natural Language Processing through a wide range of applications. However, the development of language models has not progressed uniformly across all languages. In the case of the Portuguese language, there has recently been a growing effort by academia and companies to develop language models and create data resources for Portuguese. These efforts have resulted in the rise of an increasingly diverse ecosystem of language models for Portuguese. However, information on these models remains dispersed in scientific publications, technical reports, model repositories, and project documentation. This survey presents a systematic mapping study of language models developed for Portuguese, providing a comprehensive overview of the current state of the field. We map a total of 46 models, characterizing them by various aspects, including base model, architecture, computational resources, training datasets, licensing, code availability, data, and model weights. Furthermore, we analyzed the evolution and relationships among these models through a phylogenetic perspective, identified current research gaps and opportunities, and discussed future directions for the development of language models for Portuguese.
Original Article
View Cached Full Text

Cached at: 08/20/26, 10:05 AM

# Language Models for Portuguese: A Systematic Mapping Study
Source: [https://arxiv.org/abs/2608.18138](https://arxiv.org/abs/2608.18138)
[View PDF](https://arxiv.org/pdf/2608.18138)

> Abstract:In recent years, the rapid development of language models has transformed the field of Natural Language Processing through a wide range of applications\. However, the development of language models has not progressed uniformly across all languages\. In the case of the Portuguese language, there has recently been a growing effort by academia and companies to develop language models and create data resources for Portuguese\. These efforts have resulted in the rise of an increasingly diverse ecosystem of language models for Portuguese\. However, information on these models remains dispersed in scientific publications, technical reports, model repositories, and project documentation\. This survey presents a systematic mapping study of language models developed for Portuguese, providing a comprehensive overview of the current state of the field\. We map a total of 46 models, characterizing them by various aspects, including base model, architecture, computational resources, training datasets, licensing, code availability, data, and model weights\. Furthermore, we analyzed the evolution and relationships among these models through a phylogenetic perspective, identified current research gaps and opportunities, and discussed future directions for the development of language models for Portuguese\.

## Submission history

From: Sandra Avila \[[view email](https://arxiv.org/show-email/91232f7e/2608.18138)\] **\[v1\]**Mon, 3 Aug 2026 20:51:31 UTC \(430 KB\)

Similar Articles

Amália and the Future of European Portuguese LLMs

Hacker News Top

The Portuguese government invested €5.5M in AMÁLIA, an open-source LLM for European Portuguese based on EuroLLM, but the model's data, weights, and benchmarks are not yet publicly available.

Opportunities and Challenges of Large Language Models for Low-Resource Languages in Humanities Research

arXiv cs.CL

This paper systematically evaluates the applications of large language models in low-resource language research, analyzing opportunities and challenges across linguistic variation, historical documentation, cultural expressions, and literary analysis. The study emphasizes interdisciplinary collaboration and customized model development to preserve linguistic and cultural heritage while addressing issues of data accessibility, model adaptability, and cultural sensitivity.

What language do language models speak?

Reddit r/artificial

The article investigates whether large language models have a 'mother tongue' (likely English) and how they process multilingual inputs via a shared concept space or 'semantic hub', using logit lenses to probe internal representations.