From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

arXiv cs.CL Papers

Summary

This paper introduces MÖVE, a holistic evaluation framework for LLMs in the German public sector, examining governance dimensions like energy consumption, provider transparency, and knowledge of German-party positions, revealing trade-offs that necessitate context-specific model selection.

arXiv:2608.17827v1 Announce Type: new Abstract: Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of M\"OVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.
Original Article
View Cached Full Text

Cached at: 08/19/26, 10:05 AM

# From Global Benchmarks to Local Evaluations:Benchmarking LLMs for the German Public Sector
Source: [https://arxiv.org/html/2608.17827](https://arxiv.org/html/2608.17827)
Thilo MichaelRobin SchaeferDaniel WeinlandInnovations Department, Bundesdruckerei GmbH, Berlin, Germany,firstname\.lastname@bdr\.de

###### Abstract

Public institutions face a persistent challenge in selecting LLMs suited to their specific context\. Existing benchmarks, however, are of limited use as they primarily reflect English\-language and US\-centric settings, and often only evaluate task performance\. In this paper, we present first results of MÖVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German\-party positions\. Our results reveal significant trade\-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60\-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions\. Model selection for public institutions thus cannot rely on performance rankings alone\. Instead, evaluations should also reflect the governance requirements of the deployment context\.

\*\*footnotetext:These authors contributed equally to this work\.††footnotetext:This work will be presented at the Eval4SD workshop, co\-located with KONVENS 2026 in Hamburg, Germany\.## 1Introduction

Large language models \(LLMs\) are increasingly used in high\-stakes environments, including public administration\. At the same time, model selection is no longer merely a choice between one or two well\-known providers\. Public institutions can now choose from a rapidly expanding and heterogeneous landscape of models that differ in size, performance, energy requirements, documentation practices, and contextual knowledge\.

As LLM\-generated text increasingly informs policy analysis and public\-sector decision\-making, the need for robust and locally grounded benchmarking practices becomes more pressing[12](https://arxiv.org/html/2608.17827#bib.bib12)\. General benchmark performance alone offers limited evidence about whether a model is suitable for a particular institutional, linguistic, or societal context\. The ability to identify potential harms and anticipate failures within intended use\-case scenarios has therefore made model evaluation an essential discipline[3](https://arxiv.org/html/2608.17827#bib.bib2)\.

Established LLM benchmarks predominantly measure task performance on English\-language datasets and frequently reflect US\- or UK\-specific contexts[18](https://arxiv.org/html/2608.17827#bib.bib18);[6](https://arxiv.org/html/2608.17827#bib.bib5);[8](https://arxiv.org/html/2608.17827#bib.bib8)\. Multilingual benchmarks extend coverage to additional languages[1](https://arxiv.org/html/2608.17827#bib.bib6);[17](https://arxiv.org/html/2608.17827#bib.bib17)but often rely on translated versions of existing datasets[15](https://arxiv.org/html/2608.17827#bib.bib15), thereby failing to address the domain or context gap\. Moreover, benchmarks generally prioritize task performance over the broader governance concerns that a more holistic evaluation would address[7](https://arxiv.org/html/2608.17827#bib.bib7)\. Thus, while a model may perform well on conventional benchmarks, it may, e\.g\., consume disproportionate resources or provide insufficient documentation, which are crucial criteria for the broad adoption of LLMs\. These limitations are particularly consequential for public institutions, as they must not only assess whether a model can be applied to a task, but also whether its deployment can be justified to decision\-makers, oversight bodies, and the public\.

In this paper, we present first results of MÖVE \(Modelle für die öffentliche Verwaltung evaluieren\), a holistic evaluation framework for the German public sector\.111MÖVE is conceived as alivingbenchmark, which will be regularly updated both with respect to the assessed models and the evaluation criteria\. Our leaderboard can be found here:[https://moeve\.bundesdruckerei\.de/](https://moeve.bundesdruckerei.de/)\(in German\)\. Our framework code can be found here:[https://github\.com/Bundesdruckerei\-GmbH/moeve\-lmbench](https://github.com/Bundesdruckerei-GmbH/moeve-lmbench)\. For more details on the evaluation criteria, datasets, prompts and results, see our MÖVE framework paper:[4](https://arxiv.org/html/2608.17827#bib.bib3)\.While the frameworks’ current status includes seven performance and governance criteria, in this paper we focus on the evaluation of 39 LLMs across three governance dimensions\. We estimate inference\-timeenergy consumptionusing task\-specific output statistics, assessprovider transparencythrough a manually verified matrix of 21 questions across seven documentation domains, and evaluatepolitical knowledgeusing 4,788 official positions from 64 German political parties\. While requiring different evaluation methods, the three dimensions address a common question: whether a model is suitable for deployment in a specific public\-sector context beyond its performance on a generic task\.

## 2Benchmark Design

### 2\.1Scope and Target Groups

We evaluate 39 open\-weight and proprietary LLMs222See Appendix[A](https://arxiv.org/html/2608.17827#A1)for a list of evaluated models\.from 13 providers across three governance dimensions relevant to public\-sector deployment: inference\-time energy consumption, provider transparency, andknowledgeof German political\-party positions\. The benchmark was developed through a stakeholder\-informed process involving semi\-structured feedback sessions with public\-sector representatives and exchanges with practitioners\. Following recommendations that benchmark design should be grounded in intended uses and stakeholder needs[12](https://arxiv.org/html/2608.17827#bib.bib12), we distinguish four target groups: 1\) AI decision\-makers, 2\) public\-administration domain experts, 3\) IT departments and security\-critical institutions, and 4\) broader civil society\. Their respective needs concern strategic model comparison, operational suitability, secure and compliant deployment, and public accountability\.

Table 1:Overview of datasets used for calculating energy consumption during model evaluation\.*Type*indicates whether a dataset is pre\-existing or specifically constructed for the benchmark\. The latter are not released to the public\. Dataset size is reported with respect to the evaluation unit used in each task, e\.g\.,*summaries*in summarization tasks\.
### 2\.2Energy Consumption

Energy consumption can accumulate substantially when models are deployed at scale, making energy efficiency an operational as well as an environmental consideration[16](https://arxiv.org/html/2608.17827#bib.bib16);[19](https://arxiv.org/html/2608.17827#bib.bib19)\. To enable comparison between locally deployed and API\-accessed models, we estimate inference\-time energy consumption using EcoLogits[13](https://arxiv.org/html/2608.17827#bib.bib13)\. The framework estimates energy use from the model’s total parameter count, active parameter count for mixture\-of\-experts architectures, and number of generated output tokens\. Parameter information for open\-weight models was collected from official documentation; for proprietary models without disclosed parameter counts, we relied on the assumptions provided by EcoLogits\.

Rather than applying a fixed output length, we use the actual output\-token counts recorded during model evaluation across nine German\-language datasets \(Table[1](https://arxiv.org/html/2608.17827#S2.T1)\): four summarization datasets, three question\-answering datasets, and two topic\-extraction datasets\. We report mean estimated energy consumption in watt\-hours per inference request, both by task and across tasks\.

### 2\.3Provider Transparency

Meaningful risk assessment depends on provider transparency\. Without adequate information about training data, bias mitigation, computational resources, and model limitations, adopting institutions cannot independently evaluate the risks transferred to them\. This information is also increasingly relevant in the European regulatory context, particularly in light of the documentation requirements for providers of general\-purpose AI models established under Article 53 of the EU AI Act333[https://eur\-lex\.europa\.eu/eli/reg/2024/1689/oj/eng](https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng)\.\.

Provider transparency is evaluated using a structured dataset of 21 questions across seven domains: model identification; architecture and properties; distribution and access; use and deployment; training and data; computational resources and energy consumption; endorsement of the EU General\-Purpose AI Code of Practice\. The questions were derived from the Code’s Transparency Chapter and Model Documentation Form444[https://digital\-strategy\.ec\.europa\.eu/en/policies/contents\-code\-gpai](https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai)\., which operationalize documentation obligations under Article 53\.

For each model, we collected publicly verifiable information from three source types: official provider websites, official model cards, and technical publications issued by the provider\. Each question was scored as 0 when no information was available, 1 when information was partial, and 2 when clear and verifiable information was provided\. The Code of Practice signature was treated as a binary criterion\. The documentation was first assessed manually, primarily by one researcher, with a subset independently scored by a second\. To support quality assurance, we additionally implemented an automated scoring agent that retrieves the same sources and scores all criteria independently\. The automated and manual assessments initially diverged on 27\.4% of items\. These cases were reviewed and corrected where necessary, and recurring disagreements were encoded as explicit scoring notes, reducing divergence to 16\.2% \-\-\- a measure of the final protocol’s internal consistency rather than of accuracy, since both the notes and the manual scores were revised in the process\. A second researcher reviewed the complete output\. Across 39 models, the resulting dataset contains 819 model\-question assessments\. The full matrix, including the scoring notes, the justification for each score, and the underlying sources, is publicly available\.555[https://moeve\.bundesdruckerei\.de/transparenz](https://moeve.bundesdruckerei.de/transparenz)\.

Figure 1:Mean transparency score \(%\) by domain for EU AI Act Code of Practice signatories vs\. non\-signatories\.
### 2\.4Political Knowledge

Models used in the public sector are required to produce factually accurate outputs\.Political knowledgeis a particularly important example because incorrect representations of policy positions may affect citizen\-facing information services and policy\-related applications\. Existing studies have frequently examined the political preferences attributed to LLMs through instruments such as the Political Compass666[https://www\.politicalcompass\.org/](https://www.politicalcompass.org/)\.or national voting\-advice applications\. However, such results are sensitive to prompt formulation, language, and the assumptions used to attribute a political position to a model[14](https://arxiv.org/html/2608.17827#bib.bib14);[5](https://arxiv.org/html/2608.17827#bib.bib4);[11](https://arxiv.org/html/2608.17827#bib.bib11)\. We therefore distinguish political knowledge from political bias\. Rather than measuring political bias in model outputs, we evaluate whether such outputs correctly reflect the publicly documented positions of German political parties\. This formulation treatspolitical positioningas a factual classification task grounded in an external reference rather than as an attribution of partisan preferences to the model\.

Political knowledgeis evaluated using official party responses from the German Wahl\-O\-Mat, a voting\-advice application maintained by the Federal Agency for Civic Education777[https://www\.bpb\.de/themen/wahl\-o\-mat/](https://www.bpb.de/themen/wahl-o-mat/)\.\. The dataset covers the four most recent federal elections, i\.e\., 2013, 2017, 2021, and 2025, and includes all participating parties\. Propositions for which a party did not provide a position were removed, resulting in 4,788 labeled positions \(agree,disagree, orneutral\) from 64 political parties\. 47\.5% of positions agree with the given proposition, while 39\.7% and 12\.8% exhibit a negative or neutral stance, respectively\.

Models receive a political proposition, a party, the election year, and the response options in a German\-language prompt and are tasked to predict how the specified party positioned itself\. Each model is queried once per datapoint\. Performance is measured using classification accuracy against the official party responses\. Although this task arguably poses a challenge for LLMs, given that party positions may shift over time, we consider it a compelling use case for identifying the upper bounds of LLM performance in a political context\.

## 3Results

Our results reveal substantial variation and trade\-offs\. Estimated mean energy consumption differs by a factor of 63 across the evaluated models, with smaller models generally consuming less energy, although size does not fully determine efficiency\. Provider documentation is strongest for basic model identification and architectural properties, but remains particularly weak regarding computational resources, energy use, training\-data safeguards, and bias\-mitigation practices\. Classification of party positions is challenging across the model set: the highest mean accuracy is 0\.671, and model size is positively associated with performance without being sufficient to explain it\. Across the three criteria, no single model achieves the strongest result in every dimension\.

#### Energy Consumption\.

Estimated energy consumption varies by a factor of 63 across the 39 models, from 0\.647 Wh to 40\.6 Wh per query, with a median of 1\.788 Wh\. This estimate should be interpreted with caution for proprietary models, where parameter counts are not disclosed\. Energy requirements also depend strongly on the task: summarization is more energy\-intensive than question answering and topic extraction because it produces longer outputs\. Reasoning models incur particularly high energy costs, while several smaller models remain competitive at substantially lower consumption\. Our analysis therefore indicates that marginal improvements in performance can entail disproportionately large energy costs\.

#### Provider Transparency\.

Overall transparency scores range from 18 to 38 out of 42, demonstrating substantial differences between models and providers\. Documentation is generally strong for model identification, architecture, and access, but weak for upstream information\.Compute and energyis the least transparent domain, with a mean score of 11\.5%; 84\.6% of models disclose no training\-energy consumption and 79\.5% provide no measurement methodology\. Training\-data safeguards and bias\-mitigation practices are also poorly documented \(see Appendix[A](https://arxiv.org/html/2608.17827#A1)for full results\)\. Providers that signed the EU General\-Purpose AI Code of Practice score overall only marginally higher than non\-signatories, and their advantage is concentrated in downstream information such as intended use and deployment guidance rather than training data, compute, or energy disclosures \(Figure[1](https://arxiv.org/html/2608.17827#S2.F1)\)\.

#### Political Knowledge\.

No model produces consistently accurate outputs with respect to German political\-party positions \(Table[2](https://arxiv.org/html/2608.17827#S3.T2)\)\. The two highest\-performing models achieve an accuracy of only 0\.671 across 4,788 positions from 64 parties\. Although larger models tend to perform better on average, model size does not determine performance: a small GPT\-4o Mini matches a large DeepSeek R1 at the top of the ranking, while medium\-sized and fine\-tuned models also achieve competitive results\. The highest\-performing models span European, US, and Chinese providers, thus providing no evidence that geographical proximity alone predicts higher output accuracy\. Despite substantial variance in model performance, all top\-performing models surpass the majority baseline of 0\.475\.

Table 2:Top\-10 models for party\-position classification, ranked by mean accuracy\. All top\-performing models surpass the majority baseline\.

## 4Discussion

The findings of this paper show that model selection for public administration should extend beyond conventional performance rankings\. In particular, the 63\-fold variation in estimated energy consumption demonstrates that resource efficiency is not a secondary consideration\. Models with similar task performance may impose substantially different operational and environmental costs\. Energy consumption should therefore be assessed, especially for high\-volume public\-sector applications\.

The transparency assessment reveals a persistent accountability gap\. Providers commonly document basic model properties, access conditions, and intended uses, but disclose notably less about training data, bias mitigation, computational resources, and energy consumption\. This limits the ability of public institutions to independently assess risks and make evidence\-based procurement decisions\.

Finally, the political\-knowledge results also caution against using a model provider’s geographical origin as a proxy for local suitability\. Models from European, US, and Chinese providers achieve competitive results on German political\-party positions, and no regional group consistently dominates\. Foreign\-developed models should therefore not be excluded based on assumptions about represented contextual knowledge\. At the same time, the modest maximum accuracy shows that all models require evaluation on locally relevant data before deployment in politically sensitive use cases\.

#### Conclusion\.

Our results emphasize the need for a multidimensional, locally grounded approach to LLM evaluation\. Model size, provider reputation, and general benchmark performance alone are insufficient proxies for deployment suitability\. Public administrations should compare models based on specific governance and domain requirements\. Meanwhile, providers and regulators should strengthen transparency requirements to enable such assessments\. Overall, these findings suggest that we should move away from universal model rankings and towards evaluations that reflect the deployment context’s requirements\.

## Limitations

This study is limited to three evaluation dimensions and a fixed set of models\. Future work will extend the benchmark with newer models and additional criteria to provide a more comprehensive assessment of public sector suitability\.

Energy results are comparative estimates rather than complete measurements and for proprietary models rely on EcoLogits’ assumptions about undisclosed parameter counts\. Transparency scores reflect publicly available documentation collected and verified between March and May 2026\.

The datasets capture selected aspects of the German public\-sector context and cannot represent the full diversity of administrative tasks, political knowledge, or deployment conditions\.

## Acknowledgments

We thank the anonymous reviewers for their helpful feedback\. In the preparation of this paper, generative AI was used in a supporting capacity for stylistic revision\.

## References

- Adelaniet al\.\(2025\)D\. I\. Adelani, J\. Ojo, I\. A\. Azime, J\. Y\. Zhuang, J\. O\. Alabi, X\. He, M\. Ochieng, S\. Hooker, A\. Bukula, E\. A\. Lee, C\. I\. Chukwuneke, H\. Buzaaba, B\. K\. Sibanda, G\. K\. Kalipe, J\. Mukiibi, S\. Kabongo Kabenamualu, F\. Yuehgoh, M\. Setaka, L\. Ndolela, N\. Odu, R\. Mabuya, S\. Osei, S\. H\. Muhammad, S\. Samb, T\. K\. Guge, T\. V\. Sherman, and P\. StenetorpIrokoBench: A New Benchmark for African Languages in the Age of Large Language Models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 2732–2757\.External Links:[Link](https://aclanthology.org/2025.naacl-long.139/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.139),ISBN 979\-8\-89176\-189\-6Cited by:[§1](https://arxiv.org/html/2608.17827#S1.p3.1)\.
- Aumilleret al\.\(2022\)D\. Aumiller, A\. Chouhan, and M\. GertzEUR\-lex\-sum: a multi\-and cross\-lingual dataset for long\-form summarization in the legal domain\.arXiv preprint arXiv:2210\.13448\.Cited by:[Table 1](https://arxiv.org/html/2608.17827#S2.T1.2.2.1)\.
- Changet al\.\(2024\)Y\. Chang, X\. Wang, J\. Wang, Y\. Wu, L\. Yang, K\. Zhu, H\. Chen, X\. Yi, C\. Wang, Y\. Wang, W\. Ye, Y\. Zhang, Y\. Chang, P\. S\. Yu, Q\. Yang, and X\. XieA Survey on Evaluation of Large Language Models\.ACM Trans\. Intell\. Syst\. Technol\.15\(3\)\.External Links:ISSN 2157\-6904,[Link](https://doi.org/10.1145/3641289),[Document](https://dx.doi.org/10.1145/3641289)Cited by:[§1](https://arxiv.org/html/2608.17827#S1.p2.1)\.
- Dalerciet al\.\(2026\)C\. Dalerci, T\. Michael, R\. Schaefer, and D\. WeinlandMÖVE: A Holistic LLM Benchmark for the German Public Sector\.External Links:2606\.13111,[Link](https://arxiv.org/abs/2606.13111)Cited by:[footnote 1](https://arxiv.org/html/2608.17827#footnote1)\.
- Helweet al\.\(2025\)C\. Helwe, O\. Balalau, and D\. CeolinNavigating the Political Compass: Evaluating Multilingual LLMs across Languages and Nationalities\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 17179–17204\.External Links:[Link](https://aclanthology.org/2025.findings-acl.883/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.883),ISBN 979\-8\-89176\-256\-5Cited by:[§2\.4](https://arxiv.org/html/2608.17827#S2.SS4.p1.1)\.
- Hendryckset al\.\(2021\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring Massive Multitask Language Understanding\.arXiv\.Note:arXiv:2009\.03300 \[cs\]External Links:[Link](http://arxiv.org/abs/2009.03300),[Document](https://dx.doi.org/10.48550/arXiv.2009.03300)Cited by:[§1](https://arxiv.org/html/2608.17827#S1.p3.1)\.
- Lianget al\.\(2023\)P\. Liang, R\. Bommasani, T\. Lee, D\. Tsipras, D\. Soylu, M\. Yasunaga, Y\. Zhang, D\. Narayanan, Y\. Wu, A\. Kumar, B\. Newman, B\. Yuan, B\. Yan, C\. Zhang, C\. Cosgrove, C\. D\. Manning, C\. Ré, D\. Acosta\-Navas, D\. A\. Hudson, E\. Zelikman, E\. Durmus, F\. Ladhak, F\. Rong, H\. Ren, H\. Yao, J\. Wang, K\. Santhanam, L\. Orr, L\. Zheng, M\. Yuksekgonul, M\. Suzgun, N\. Kim, N\. Guha, N\. Chatterji, O\. Khattab, P\. Henderson, Q\. Huang, R\. Chi, S\. M\. Xie, S\. Santurkar, S\. Ganguli, T\. Hashimoto, T\. Icard, T\. Zhang, V\. Chaudhary, W\. Wang, X\. Li, Y\. Mai, Y\. Zhang, and Y\. KoreedaHolistic Evaluation of Language Models\.External Links:2211\.09110,[Link](https://arxiv.org/abs/2211.09110)Cited by:[§1](https://arxiv.org/html/2608.17827#S1.p3.1)\.
- Majithiaet al\.\(2026\)N\. Majithia, R\. Shinde, Z\. Chapman, P\. Trital, J\. Decker, M\. Maskey, E\. Simperl, and N\. ShadboltThe CitizenQuery Benchmark: A Novel Dataset and Evaluation Pipeline for Measuring LLM Performance in Citizen Query Tasks\.External Links:2602\.04064,[Link](https://arxiv.org/abs/2602.04064)Cited by:[§1](https://arxiv.org/html/2608.17827#S1.p3.1)\.
- Mölleret al\.\(2021\)T\. Möller, J\. Risch, and M\. PietschGermanQuAD and GermanDPR: Improving Non\-English Question Answering and Passage Retrieval\.InProceedings of the 3rd Workshop on Machine Reading for Question Answering,A\. Fisch, A\. Talmor, D\. Chen, E\. Choi, M\. Seo, P\. Lewis, R\. Jia, and S\. Min \(Eds\.\),Punta Cana, Dominican Republic,pp\. 42–50\.External Links:[Link](https://aclanthology.org/2021.mrqa-1.4/),[Document](https://dx.doi.org/10.18653/v1/2021.mrqa-1.4)Cited by:[Table 1](https://arxiv.org/html/2608.17827#S2.T1.2.6.1)\.
- Rasiahet al\.\(2023\)V\. Rasiah, R\. Stern, V\. Matoshi, M\. Stürmer, I\. Chalkidis, D\. E\. Ho, and J\. NiklausSCALE: Scaling up the Complexity for Advanced Language Model Evaluation\.External Links:2306\.09237Cited by:[Table 1](https://arxiv.org/html/2608.17827#S2.T1.2.3.1)\.
- Rettenbergeret al\.\(2025\)L\. Rettenberger, M\. Reischl, and M\. SchuteraAssessing political bias in large language models\.Journal of Computational Social Science8\(2\),pp\. 42\.External Links:ISSN 2432\-2725,[Document](https://dx.doi.org/10.1007/s42001-025-00376-w),[Link](https://doi.org/10.1007/s42001-025-00376-w)Cited by:[§2\.4](https://arxiv.org/html/2608.17827#S2.SS4.p1.1)\.
- Reuelet al\.\(2024\)A\. Reuel, A\. Hardy, C\. Smith, M\. Lamparth, M\. Hardy, and M\. J\. KochenderferBetterBench: assessing AI benchmarks, uncovering issues, and establishing best practices\.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA\.External Links:ISBN 9798331314385Cited by:[§1](https://arxiv.org/html/2608.17827#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.17827#S2.SS1.p1.1)\.
- Rincé and Banse \(2025\)S\. Rincé and A\. BanseEcoLogits: Evaluating the Environmental Impacts of Generative AI\.Journal of Open Source Software10\(111\),pp\. 7471\.External Links:[Document](https://dx.doi.org/10.21105/joss.07471),[Link](https://joss.theoj.org/papers/10.21105/joss.07471)Cited by:[§2\.2](https://arxiv.org/html/2608.17827#S2.SS2.p1.1)\.
- Röttgeret al\.\(2024\)P\. Röttger, V\. Hofmann, V\. Pyatkin, M\. Hinck, H\. Kirk, H\. Schuetze, and D\. HovyPolitical Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15295–15311\.External Links:[Link](https://aclanthology.org/2024.acl-long.816/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.816)Cited by:[§2\.4](https://arxiv.org/html/2608.17827#S2.SS4.p1.1)\.
- Singhet al\.\(2025\)S\. Singhet al\.Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 18761–18799\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.919),ISBN 979\-8\-89176\-251\-0,[Link](https://aclanthology.org/2025.acl-long.919/)Cited by:[§1](https://arxiv.org/html/2608.17827#S1.p3.1)\.
- Strubellet al\.\(2019\)E\. Strubell, A\. Ganesh, and A\. McCallumEnergy and Policy Considerations for Deep Learning in NLP\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 3645–3650\.External Links:[Link](https://aclanthology.org/P19-1355/),[Document](https://dx.doi.org/10.18653/v1/P19-1355)Cited by:[§2\.2](https://arxiv.org/html/2608.17827#S2.SS2.p1.1)\.
- Susantoet al\.\(2025\)Y\. Susanto, A\. V\. Hulagadri, J\. R\. Montalan, J\. G\. Ngui, X\. Yong, W\. Q\. Leong, H\. Rengarajan, P\. Limkonchotiwat, Y\. Mai, and W\. C\. TjhiSEA\-HELM: Southeast Asian Holistic Evaluation of Language Models\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 12308–12336\.External Links:[Link](https://aclanthology.org/2025.findings-acl.636/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.636),ISBN 979\-8\-89176\-256\-5Cited by:[§1](https://arxiv.org/html/2608.17827#S1.p3.1)\.
- Wanget al\.\(2019\)A\. Wang, Y\. Pruksachatkun, N\. Nangia, A\. Singh, J\. Michael, F\. Hill, O\. Levy, and S\. R\. BowmanSuperGLUE: a stickier benchmark for general\-purpose language understanding systems\.InProceedings of the 33rd International Conference on Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2608.17827#S1.p3.1)\.
- Wuet al\.\(2025\)Y\. Wu, I\. Hua, and Y\. DingUnveiling Environmental Impacts of Large Language Model Serving: A Functional Unit View\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 10560–10576\.External Links:[Link](https://aclanthology.org/2025.acl-long.519/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.519),ISBN 979\-8\-89176\-251\-0Cited by:[§2\.2](https://arxiv.org/html/2608.17827#S2.SS2.p1.1)\.

## Appendix ATransparency Scores

Figure 2:Transparency scores for all 39 models, sorted by domain and total score\. Each segment represents the score in one of the seven transparency documentation domains\. The pronounced gap between well\-documented domains \(Model Identification, Architecture, Distribution & Access\) and poorly\-documented ones \(Use & Deployment, Training & Data, Compute & Energy\) is visible across nearly all models\.

Similar Articles

Meta-Benchmarks for Financial-Services LLM Evaluation

arXiv cs.AI

This paper presents a meta-benchmarking framework that aggregates 452 existing public benchmarks into 41 work activities and 38 banking business domains, enabling more precise LLM evaluation and governance for financial services institutions.

Benchmarking LLMs

Reddit r/AI_Agents

A study or report on benchmarking large language models, likely comparing performance across various tasks.

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

arXiv cs.CL

This paper introduces a cross-evaluation framework for benchmarking LLMs on Arabic cultural and sociolinguistic knowledge, using human SME ground truth and automated judges. The authors contribute a dataset of prompt-rubric pairs for Egyptian and Iraqi Arabic, evaluating frontier LLMs and finding that cultural reasoning remains a primary failure mode for automated grading.