On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora?
Summary
This study evaluates four proprietary LLMs (GPT-4o, GPT-5.2, Claude Sonnet 4.5, DeepSeek) for specialized terminology translation from English to French across two domains, comparing prompting strategies. Results show Claude Sonnet 4.5 performs best, but LLMs cannot yet replace specialized corpora.
View Cached Full Text
Cached at: 07/29/26, 09:52 AM
# On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora? Source: [https://arxiv.org/abs/2607.24784](https://arxiv.org/abs/2607.24784) [View PDF](https://arxiv.org/pdf/2607.24784) > Abstract:Specialised translation relies on the use of documentary and terminological resources, including corpora\. These resources are particularly useful for terminology\. However, their compilation and exploitation have several limitations: they require time, technical skills and access to data that can be difficult to collect\. This study examines the extent to which LLMs can assist specialised translators in finding equivalents from English to French\. We evaluate four proprietary models, GPT\-4o, GPT\-5\.2, Claude Sonnet 4\.5 and DeepSeek, in two specialised domains, Earth, Environmental and Planetary Sciences \(EEPS\) and Natural Language Processing \(NLP\)\. The experiment is based on 80 terms per domain and compares two prompting strategies: a terminology and a translation mode\. The results highlight clear differences between models, prompting strategies and, to a lesser extent, domains\. Claude Sonnet 4\.5 achieves the best results in the most favourable configuration, while DeepSeek stands out for its greater stability\. Analysis of confidence estimates also shows that they are only a partial indicator of terminological accuracy\. Overall, the findings suggest that LLMs can be useful tools for specialised translators, but cannot, at this stage, replace specialised corpora\. This research therefore paves the way for future work on the real practical usefulness of LLMs for specialised translators in work and educational contexts\. ## Submission history From: Joachim Minder \[[view email](https://arxiv.org/show-email/45af6ef9/2607.24784)\] \[via CCSD proxy\] **\[v1\]**Mon, 22 Jun 2026 09:03:13 UTC \(931 KB\)
Similar Articles
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
This paper presents an LLM-driven pipeline using GPT-5, GPT-4o, and Claude Sonnet 4 to automatically design neural network architectures for cross-lingual handwritten OCR, achieving over 93% accuracy across Arabic, English, and Persian scripts without human intervention.
Small LLMs for Biomedical Claim Verification: Cost-Effective Fine-Tuning, Structural Dataset Shortcuts, and Cross-Domain Generalization
Fine-tuning small LLMs (3B-7B) with QLoRA on biomedical claim verification achieves higher F1 than GPT-4o and GPT-5 at 44.5x lower cost, and reveals a structural artifact in SciFact. The study demonstrates robust cross-domain transfer when training on structurally sound data.
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
DeepSeek LLM is an open-source language model project that develops a large dataset and employs SFT and DPO to achieve performance surpassing LLaMA-2 70B and GPT-3.5 in various benchmarks and open-ended evaluations.
Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs
This study compares probing techniques for identifying latent language in multilingual LLMs, finding that different methods yield inconsistent results, indicating they expose distinct aspects of multilingual processing rather than a single internal lingua franca.
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
A comprehensive dual-aspect evaluation framework for large language models on Vietnamese legal text simplification, combining quantitative benchmarking (Accuracy, Readability, Consistency) with qualitative error analysis across GPT-4o, Claude 3 Opus, Gemini 1.5 Pro, and Grok-1.