Workload-Driven Optimization for On-Device Real-Time Subtitle Translation
Summary
This paper presents a workload-driven optimization for on-device real-time English-to-Traditional-Chinese subtitle translation by reducing vocabulary size from 151k to 64k, adapting the model via embedding migration and fine-tuning, achieving a 1.63× speedup on Apple M2 with competitive translation quality.
View Cached Full Text
Cached at: 07/14/26, 04:21 AM
# Workload-Driven Optimization for On-Device Real-Time Subtitle Translation
Source: [https://arxiv.org/html/2607.09957](https://arxiv.org/html/2607.09957)
###### Abstract
This report studies on\-device English\-to\-Traditional\-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch\-size\-one inference, low latency, and privacy constraints\. These conditions limit the value of optimizations designed for long\-context or high\-throughput language\-model serving\.
Starting from LMT\-60\-0\.6B\[[2](https://arxiv.org/html/2607.09957#bib.bib2)\], preliminary profiling suggests that vocabulary projection becomes a more important decode\-time cost after GGUF quantization reduces the relative cost of Transformer blocks\. We replace the original 151k\-token vocabulary with a 64k\-token subtitle\-domain tokenizer, migrate the embedding space, and adapt the model through embedding calibration followed by full supervised fine\-tuning\.
On a fixed 500\-example subset of the OpenSubtitles2024 test set, the LocalSubs achieves a 59\.2% tie\-excluded win rate against Google Translate under GPT\-4o pairwise judging\. Performance is strongest on short cues and declines as cue length increases\. Preliminary Apple M2 Metal measurements on a 64k\-vocabulary model show a 1\.63×\\timesspeedup over a 151k\-vocabulary profiling baseline\. The raw benchmark configuration is incomplete, so the latency result is treated as preliminary\.
## 1Introduction
Real\-time subtitle translation imposes a different set of constraints from conventional machine translation and general\-purpose language\-model serving\. Each request typically contains one current subtitle cue and only a small amount of preceding context\. Both the input and output are short, inference is effectively performed at batch size one, and the translation must be produced before the subtitle display window expires\. For privacy\-sensitive applications, the complete pipeline must also run on local consumer hardware while generating natural Traditional Chinese as used in Taiwan\.
These characteristics change the inference bottleneck\. Techniques designed for long contexts, repeated prefixes, or large batches offer limited benefit when prompts are short and requests arrive sequentially\. In this setting, fixed per\-request overhead, decode\-time computation, vocabulary projection, and output token count can have a greater effect on end\-to\-end latency\. A large multilingual tokenizer may further increase the output\-projection cost while providing inefficient tokenization for domain\-specific Chinese subtitle text\.
This work investigates a workload\-driven optimization path based on LMT\-60\-0\.6B\[[2](https://arxiv.org/html/2607.09957#bib.bib2)\]\. The original 151k vocabulary is replaced with a 64k subtitle\-domain tokenizer trained on English and Taiwan Traditional Chinese subtitle text\. Because tokenizer replacement changes token identities and invalidates the original embedding matrix, the model is adapted through embedding migration, an embedding\-calibration stage, and full supervised fine\-tuning\. The resulting system is evaluated using pairwise preference judgments against a fixed Google Translate anchor, while deployment performance is examined through preliminary Apple M2 Metal measurements\. The main contributions are:
- •A 64k\-vocabulary subtitle tokenizer that reduces the output\-projection dimension and improves Chinese token density\.
- •An embedding migration and two\-stage adaptation procedure for replacing the tokenizer of a pretrained translation model\.
- •A translation\-quality evaluation based on pairwise preference judgments, including context ablation and cue\-length analysis\.
## 2Task and Design Goals
### 2\.1Cue\-level context\-aware translation
The model receives up to three previous English subtitle cues as context and one current cue as the translation target\. It must output only the Taiwan Traditional Chinese translation of the current cue\. Context is used for disambiguation and tone continuity, not as additional translation content\.
The task is
yt=f\(xt−k:t−1,xt\),0≤k≤3,y\_\{t\}=f\(x\_\{t\-k:t\-1\},x\_\{t\}\),\\qquad 0\\leq k\\leq 3,wherextx\_\{t\}is the current English cue andyty\_\{t\}is its translation\.
The main quality requirements are semantic correctness, current\-cue\-only output, natural Taiwan usage, concise subtitle style, and stable handling of short ambiguous utterances\.
### 2\.2On\-device constraints
The target environment has four practical constraints:
- •Low latency:each cue should complete within the subtitle display window\.
- •Low batch:interactive playback is effectively batch size one\.
- •Limited hardware:the model must run on consumer devices such as Apple M\-series systems\.
- •Privacy:subtitle content should remain on the local device\.
## 3Method
### 3\.1Workload profile and decode cost
This workload uses short prompts, short outputs, and little reusable prefix\. Optimizations aimed mainly at long attention sequences or large\-batch inference therefore offer limited benefit\.
Preliminary profiling suggests that, after GGUF Q5\_K\_M quantization reduces the relative cost of Transformer blocks, the output projection becomes a more important part of decode time\. Each decode step computes
logits=hWvocab⊤,\\mathrm\{logits\}=hW\_\{\\mathrm\{vocab\}\}^\{\\top\},with cost proportional to\|𝒱\|dmodel\|\\mathcal\{V\}\|d\_\{\\mathrm\{model\}\}\. The original vocabulary contains 151,936 tokens with hidden size 1024, so every decode step projects to more than 150,000 logits\.
This observation motivates a combined strategy:
1. 1\.quantization reduces per\-step Transformer cost; and
2. 2\.tokenizer redesign reduces the projection dimension and may reduce the number of Chinese decode tokens\.
Because the operation\-level profiling log is not yet complete, this report treats vocabulary projection as a supported hypothesis rather than a fully isolated bottleneck\.
### 3\.264k\-vocabulary subtitle tokenizer
We train a ByteLevel BPE tokenizer on English and Taiwan Traditional Chinese subtitle text\. BPE provides an open\-vocabulary subword representation while allowing the vocabulary to be tuned to the target domain\[[1](https://arxiv.org/html/2607.09957#bib.bib1)\]\. The base tokenizer contains 64,020 tokens; four structural tokens increase the final vocabulary to 64,024\.
Table 1:Effect of tokenizer replacementThe smaller vocabulary reduces the output\-projection dimension by approximately 57\.9%\. On the same Chinese text, the subtitle tokenizer also increases character density, which reduces the number of tokens required to represent the output\. This fixed\-text measurement isolates tokenizer efficiency from differences in model generation behavior\.
### 3\.3Embedding migration and two\-stage adaptation
Tokenizer replacement changes token identities and prevents direct reuse of the original embedding matrix\. We initialize the new embedding space using string\-level correspondence:
- •copy the original embedding when the token string already exists;
- •otherwise, tokenize the new token with the original tokenizer and average the corresponding embeddings; and
- •use a mean fallback only when decomposition is impossible\.
All 64k base\-vocabulary tokens were initialized by direct copy or sub\-token averaging\. The four structural tokens were initialized separately from their original string decompositions\. No mean fallback was required\.
The model is then adapted in two stages:
embedding migration→embedding calibration→full SFT\.\\text\{embedding migration\}\\rightarrow\\text\{embedding calibration\}\\rightarrow\\text\{full SFT\}\.
During embedding calibration, Transformer layers are frozen while embeddings, the output head, and RMSNorm parameters are updated\. Full supervised fine\-tuning then updates the complete model on the final subtitle dataset\.
## 4Experimental Setup
### 4\.1Model and training data
The base model is NiuTrans/LMT\-60\-0\.6B, a compact multilingual translation model\[[2](https://arxiv.org/html/2607.09957#bib.bib2),[3](https://arxiv.org/html/2607.09957#bib.bib3)\]\. The LocalSubs is trained on a quality\-filtered, length\-rebalanced subtitle SFT dataset\.
Table 2:Core experimental configurationThe final SFT dataset combines rule\-based cleaning, Traditional Chinese filtering, LLM\-assisted subtitle\-pair quality filtering, shuffling, and increased coverage of medium\-to\-long cues\. The filtering model belongs to the Qwen3 family\[[5](https://arxiv.org/html/2607.09957#bib.bib5)\]\. Detailed construction rules are provided in Appendix[B](https://arxiv.org/html/2607.09957#A2)\.
### 4\.2Evaluation data and protocol
The OpenSubtitles2024 benchmark contains bilingual subtitle alignments held out for machine\-translation development and evaluation\[[4](https://arxiv.org/html/2607.09957#bib.bib4)\]\. The primary quality evaluation uses a fixed 500\-example subset\.
The main protocol is pairwise preference against a fixed Google Translate anchor\. GPT\-4o judges the candidate and anchor translations with randomized A/B order and temperature zero\. The metric excludes ties:
winrate=winswins\+losses\.\\mathrm\{win\\ rate\}=\\frac\{\\mathrm\{wins\}\}\{\\mathrm\{wins\}\+\\mathrm\{losses\}\}\.
LLM\-based pairwise judging is scalable but can exhibit position and other biases, so randomized candidate order and a future human audit are important\[[6](https://arxiv.org/html/2607.09957#bib.bib6)\]\. All results are reported with sample size and win/loss/tie counts\. Character\-level F1, simplified\-Chinese rate, and English echo rate are used only as surface\-form diagnostics\.
The exact GPT\-4o API snapshot was not retained; the experiment record identifies only the model alias\. This is a reproducibility limitation\. GPT\-4o mini is used as a cloud baseline\[[8](https://arxiv.org/html/2607.09957#bib.bib8),[9](https://arxiv.org/html/2607.09957#bib.bib9)\]\.
### 4\.3Deployment benchmark
The deployment path uses llama\.cpp, GGUF Q5\_K\_M quantization, and the Apple Metal backend\. llama\.cpp provides local inference, integer quantization, and optimized Apple Silicon support\[[7](https://arxiv.org/html/2607.09957#bib.bib7)\]\. The recorded latency summary uses a separate 64k\-vocabulary profiling model, not the final model evaluated for translation quality\.
The exact raw log, llama\.cpp commit, converted\-model checksum, warm\-up count, and run count are incomplete\. The latency measurements are therefore profiling evidence rather than final deployment claims\.
## 5Results
### 5\.1Pairwise translation quality
Table 3:Pairwise preference results against Google TranslateThe LocalSubs achieves a 59\.2% tie\-excluded win rate against Google Translate, compared with 64\.6% for GPT\-4o mini under the same anchor\-relative evaluation protocol\. The 5\.4\-percentage\-point difference suggests that the local 0\.6B model approaches the translation quality of a commercial cloud model despite operating under substantially stricter constraints, including on\-device execution, batch\-size\-one inference, low\-latency requirements, and no reliance on a remote API\.
This result is particularly relevant to the target application\. The objective is not to maximize translation quality without deployment constraints, but to achieve competitive subtitle translation while preserving privacy and enabling real\-time local inference\. The final model therefore represents a practical quality–efficiency trade\-off rather than a direct replacement for a larger cloud model\. Because both systems were evaluated independently against Google Translate, the reported win rates are anchor\-relative and do not constitute a direct pairwise comparison between the final model and GPT\-4o mini\. Surface\-form diagnostics further show a simplified\-Chinese rate of 0\.3% and an English echo rate of 0\.0%\. Character\-level F1 is reported only as a secondary diagnostic and is not used as the main measure of translation quality\.
### 5\.2Context ablation
Table 4:Context ablation against the same Google Translate anchor with short cue subsetContext changes the full\-set point estimate by 0\.6 percentage points\. On the independently defined 260\-example short\-response subset, the with\-context estimate is 4\.5 points higher\. The intervals in Table[4](https://arxiv.org/html/2607.09957#S5.T4)are marginal bootstrap intervals for each condition rather than a paired interval for the difference\. The 4\.5\-point difference is therefore descriptive and is not presented as statistically significant\.
### 5\.3Performance by cue length
Table 5:Final\-model win rate by source\-cue lengthThe model is strongest on short cues and falls below the anchor for cues of eight words or more\. This pattern is consistent with qualitative observations that longer cues are more vulnerable to omission and incomplete meaning preservation\. The 16\+ word result is descriptive only because the subset contains eight examples\.
### 5\.4Embedding calibration comparison
Table 6:Embedding calibration comparison on 105 aligned examplesThe comparison supports embedding calibration before full fine\-tuning\. However, char\-F1 is only a surface metric, and the archived experiment record does not establish that calibration was the only difference between the two training runs\. The result should therefore be treated as supporting evidence rather than a fully isolated ablation\.
### 5\.5Preliminary latency results
Table 7:Preliminary Apple M2 Metal profiling resultsThe latency difference is consistent with a smaller output projection and denser Chinese tokenization\. However, the generated\-output token counts compare two systems and may reflect both tokenizer efficiency and differences in translation length or omission\. The fixed\-text characters\-per\-token result in Table[1](https://arxiv.org/html/2607.09957#S3.T1)is the cleaner tokenizer\-efficiency measurement\.
End\-to\-end speedup is smaller than the decode\-only estimate because prefill, tokenization, runtime overhead, and a higher English prompt\-token count remain\. These measurements require a reproducible rerun before final publication\.
## 6Discussion
### 6\.1Tokenizer redesign as a serving decision
Quantization and tokenizer redesign address different parts of the latency problem\. Quantization reduces Transformer\-block cost\. A smaller vocabulary then reduces output\-projection work, while subtitle\-domain tokenization represents Chinese text with fewer tokens\. In this workload, tokenizer design is therefore part of the serving architecture rather than only a preprocessing choice\.
This interpretation remains provisional because the operation\-level profiler record is incomplete\. A controlled 151k/64k/32k comparison under the same model, data, and deployment settings is needed to isolate vocabulary size from output\-token count and other changes\.
### 6\.2Role of embedding calibration
Embedding averaging avoids random initialization but does not guarantee alignment with the pretrained semantic space\. Direct full fine\-tuning asks the model to adapt to both a new embedding distribution and a new task format\. The calibration stage separates these problems and is supported by the archived comparison\.
An early run also showed that calibration must cover enough of the data distribution\. Ordered subtitle data and a very short calibration stage produced a train/evaluation loss inversion when training reached later, less familiar content\. Detailed diagnostics are provided in Appendix[E](https://arxiv.org/html/2607.09957#A5)\.
### 6\.3Quality trade\-offs
The observed advantage is concentrated in subtitle style and short\-cue translation rather than broad performance across cue lengths\. The final model often produces concise, natural Taiwan subtitle phrasing on short cues\. Its main weakness is over\-compression on longer cues, where missing details and semantic errors become more frequent\.
Qualitative inspection also suggests unstable handling of names, transliterations, and some region\-specific lexical choices\. These error categories have not yet been systematically counted\.
## 7Limitations and Future Work
This study has several limitations\. First, the Apple M2 latency experiment does not yet have a complete reproducibility record and was conducted with a profiling model rather than the final evaluated model\. The vocabulary\-projection interpretation is also based on incomplete profiling evidence because an operation\-level trace was not retained\. Second, the pairwise evaluation still requires human auditing, and the exact GPT\-4o snapshot used for judging was not recorded\. In addition, the final model and GPT\-4o mini were evaluated independently against the same Google Translate anchor rather than compared directly across the full evaluation set\. The automatically scenario\-tagged benchmark has not yet been evaluated with a complete aligned scenario\-stratified protocol, and important error categories, including omission, named\-entity translation, and Taiwan\-specific lexical choice, have not been manually quantified\.
Future work should therefore prioritize rerunning the latency and profiling experiments with complete logs and the final evaluated model, conducting controlled comparisons across tokenizer sizes, and validating the pairwise judgments through human review\. It should also include a full scenario\-stratified evaluation of the tagged test set and the construction of a manually verified error taxonomy\.
## 8Conclusion
This work presents a workload\-driven optimization path for on\-device real\-time subtitle translation\. Short inputs, short outputs, and batch\-size\-one inference make fixed decode costs more relevant than optimizations designed for long\-context or high\-throughput serving\. Preliminary profiling suggests that vocabulary projection becomes increasingly important after quantization reduces Transformer\-block cost\.
A 64k\-vocabulary subtitle tokenizer reduces the output\-projection dimension, increases Chinese token density, and decreases model size\. Embedding migration and calibration provide a practical adaptation path before full supervised fine\-tuning\.
The LocalSubs achieves a 59\.2% tie\-excluded win rate against Google Translate on 500 aligned examples\. It performs best on short cues and remains weaker on longer inputs\. A separate reduced\-vocabulary profiling model shows a preliminary 1\.63×\\timesaverage\-latency speedup on Apple M2 Metal, but this result requires a controlled and reproducible rerun\.
## Appendix ATask Format and Structural Tokens
The training and inference input format is:
```
CTX:
<0--3 previous English subtitle cues>
CUR:
<current English subtitle cue>
```
The output contains only the translation ofCUR\. Context must not be copied or translated into the output\.
The tokenizer includes four structural tokens corresponding to the user role, assistant role,`\\nCTX:\\n`, and`\\nCUR:\\n`\. The latter two delimit context and current\-cue content\. Their embeddings are initialized from the average embeddings of their original sub\-token decompositions\.
## Appendix BDataset Construction
The source pipeline begins with 14,237,823 English–Chinese subtitle pairs\. The final SFT dataset contains 233,088 examples\.
The filtering pipeline includes:
1. 1\.removal of empty records, extreme lengths, abnormal length ratios, OCR corruption, release\-group advertisements, and lyric metadata;
2. 2\.Traditional Chinese filtering, including rejection of targets containing simplified\-only characters;
3. 3\.LLM\-assisted subtitle\-pair quality filtering for semantic correctness, natural Taiwan usage, and subtitle readability; and
4. 4\.length rebalancing followed by shuffling before the training/evaluation split\.
The observed acceptance rate of the LLM\-assisted filtering stage was approximately 30%\.
Table 8:Cue\-length distribution of the final SFT datasetThe initial cleaned subtitle SFT dataset is retained only for analysis of data ordering, evaluation splitting, and training instability\. The final model uses the quality\-filtered, length\-rebalanced subtitle SFT dataset\.
## Appendix CAutomatic Scenario Annotation
The 4,246\-example OpenSubtitles2024 test set was automatically assigned multi\-label scenario tags for future stratified evaluation\. Annotation does not constitute completed model evaluation\.
Table 9:Automatic scenario\-tag distributionThe context\-ablation subsetshort\_responseis defined independently from the scenario tagS1\_short\_response\. They must not be merged without row\-level verification\.
## Appendix DTokenizer and Embedding Migration Details
Table 10:Tokenizer configurationTable 11:Embedding initialization coverage for the 64k\-token base vocabularyThe four structural tokens are initialized separately\. The parameter reduction is concentrated in the tied embedding and output\-projection matrix\.
Table 12:Model size before and after vocabulary replacement
## Appendix ETraining Configuration and Diagnostic Run
### E\.1Embedding calibration
Table 13:Embedding calibration configuration
### E\.2Early loss inversion
The initial cleaned dataset and a short 0\.07\-epoch calibration stage produced increasing training loss while evaluation loss decreased\.
Table 14:Loss inversion in the early training runA likely cause is ordered subtitle data combined with insufficient calibration coverage\. Later portions of the data contained less familiar names, transliterations, and domain terms\. The calibration stage was extended to 0\.5 epoch, and the final dataset was shuffled before splitting\.
## Appendix FEvaluation Details
The judge prioritizes semantic correctness, Taiwan Traditional Chinese usage, naturalness, subtitle concision, and current\-cue\-only output\. Simplified Chinese is treated as a hard failure\. Candidate order is randomized to reduce position bias\.
Google Translate is a fixed comparison anchor, not ground truth\. The OpenSubtitles translation remains the dataset reference and is used only for reference\-based diagnostics\.
Character\-level F1 is useful for internal tracking but unreliable as a standalone translation\-quality metric\. Valid paraphrases may have low overlap, while an incorrect translation may share many characters with a noisy reference\.
The evaluation artifacts preserve sample\-level judgments and A/B assignments\. The exact GPT\-4o snapshot, API date, and A/B randomization seed were not retained and should be recorded in future reruns\.
### F\.1Bootstrap confidence intervals
Win\-rate confidence intervals are computed byeval/summarize\_winrate\.py\. The implementation uses 10,000 bootstrap resamples and random seed 13 by default\. The resampling unit is one judged row\. For each replicate,nnrows are sampled with replacement from the originalnnrows, and win/loss/tie counts are recomputed from the sampled rows\. Rows labeledmodel\_wins,anchor\_wins, andtiemap to W, L, and T, respectively\. Ties remain in the resampled data, while the statistic is recomputed as
p^win=WW\+L\.\\widehat\{p\}\_\{\\mathrm\{win\}\}=\\frac\{W\}\{W\+L\}\.
The implementation sorts the 10,000 bootstrap statistics and selects the entries at indices⌊0\.025B⌋\\lfloor 0\.025B\\rfloorand⌊0\.975B⌋\\lfloor 0\.975B\\rfloor, whereB=10,000B=10\{,\}000\. This is a custom order\-statistic implementation of a percentile interval; it does not use SciPy quantile interpolation\. The reported intervals describe each condition separately\. A paired confidence interval for the difference between with\-context and without\-context conditions would require resampling aligned example pairs and recomputing the difference within each replicate\.
“‘latex
## Appendix GLatency Benchmark Details
Table 15:Full preliminary latency summaryTable 16:Preliminary throughput summaryTable 17:Average token counts in the recorded profiling runThe 64k tokenizer produces 7\.7% fewer tokens for the English prompts in the recorded profiling set\. The 64k\-vocabulary model also generates 40\.0% fewer Chinese output tokens than the 151k\-vocabulary baseline\. The prompt\-token comparison uses the same source text and therefore provides direct evidence of improved input tokenization efficiency\. In contrast, the generated\-output comparison may reflect both tokenizer compression and differences in translation content or generation behavior\.
The latency improvement is supported by three directly measured observations: the 64k\-vocabulary model achieves 1\.52×\\timeshigher decode throughput, uses fewer prompt tokens, and generates fewer output tokens in the recorded run\. Together, these factors are consistent with the measured 1\.63×\\timesimprovement in average end\-to\-end latency\. A precise attribution of the gain to prefill, decoding, tokenization, and runtime overhead would require stage\-level timing measurements\.
A publication\-quality report should record the llama\.cpp commit, GGUF filenames and checksums, Apple M2 model and memory configuration, benchmark command, input set, decoding parameters, warm\-up count, measured run count, and raw per\-run log\.
## References
- \[1\]Rico Sennrich, Barry Haddow, and Alexandra Birch\.Neural Machine Translation of Rare Words with Subword Units\.In*Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 1715–1725, 2016\.[https://aclanthology\.org/P16\-1162/](https://aclanthology.org/P16-1162/)\.[https://doi\.org/10\.18653/v1/P16\-1162](https://doi.org/10.18653/v1/P16-1162)\.
- \[2\]Yingfeng Luo, Ziqiang Xu, Yuxuan Ouyang, Murun Yang, Dingyang Lin, Kaiyan Chang, Tong Zheng, Bei Li, Peinan Feng, Quan Du, Tong Xiao, and Jingbo Zhu\.NiuTrans\.LMT: Toward Inclusive and Scalable Multilingual Machine Translation with LLMs\.In*Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 25151–25179, 2026\.[https://aclanthology\.org/2026\.acl\-long\.1153/](https://aclanthology.org/2026.acl-long.1153/)\.[https://doi\.org/10\.18653/v1/2026\.acl\-long\.1153](https://doi.org/10.18653/v1/2026.acl-long.1153)\.
- \[3\]NiuTrans\.NiuTrans/LMT\-60\-0\.6B\.Hugging Face model card\.[https://huggingface\.co/NiuTrans/LMT\-60\-0\.6B](https://huggingface.co/NiuTrans/LMT-60-0.6B), accessed July 11, 2026\.
- \[4\]Joerg Tiedemann and Hengyu Luo\.OpenSubtitles2024: A Massively Parallel Dataset of Movie Subtitles for MT Development and Evaluation\.In*Proceedings of the Fifteenth Language Resources and Evaluation Conference \(LREC 2026\)*, pages 8897–8907, 2026\.[https://doi\.org/10\.63317/4ivg578ub2ob](https://doi.org/10.63317/4ivg578ub2ob)\.Dataset available at[https://huggingface\.co/datasets/Helsinki\-NLP/OpenSubtitles2024](https://huggingface.co/datasets/Helsinki-NLP/OpenSubtitles2024), accessed July 11, 2026\.
- \[5\]An Yang et al\.Qwen3 Technical Report\.arXiv preprint arXiv:2505\.09388, 2025\.[https://arxiv\.org/abs/2505\.09388](https://arxiv.org/abs/2505.09388)\.[https://doi\.org/10\.48550/arXiv\.2505\.09388](https://doi.org/10.48550/arXiv.2505.09388)\.
- \[6\]Lianmin Zheng et al\.Judging LLM\-as\-a\-Judge with MT\-Bench and Chatbot Arena\.In*Advances in Neural Information Processing Systems*, volume 36, pages 46595–46623, 2023\.[https://proceedings\.neurips\.cc/paper\_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832\-Abstract\-Datasets\_and\_Benchmarks\.html](https://proceedings.neurips.cc/paper_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets_and_Benchmarks.html)\.
- \[7\]ggml\-org\.llama\.cpp: LLM inference in C/C\+\+\.GitHub repository\.[https://github\.com/ggml\-org/llama\.cpp](https://github.com/ggml-org/llama.cpp), accessed July 11, 2026\.
- \[8\]OpenAI\.GPT\-4o Model\.OpenAI API documentation\.[https://developers\.openai\.com/api/docs/models/gpt\-4o](https://developers.openai.com/api/docs/models/gpt-4o), accessed July 11, 2026\.
- \[9\]OpenAI\.GPT\-4o mini Model\.OpenAI API documentation\.[https://developers\.openai\.com/api/docs/models/gpt\-4o\-mini](https://developers.openai.com/api/docs/models/gpt-4o-mini), accessed July 11, 2026\.Similar Articles
tencent/Hy-MT2-7B
Tencent open-sourced the Hy-MT2 family of fast-thinking multilingual translation models (1.8B, 7B, 30B-A3B) supporting 33 languages, along with extreme quantization for on-device deployment and a new instruction-following benchmark IFMTBench.
Systematic Optimization of Real-Time Diffusion Model Inference on Apple M3 Ultra
This paper presents a systematic optimization study of real-time diffusion model inference on the Apple M3 Ultra, achieving 22.7 FPS at 512x512 resolution using CoreML conversion and a distillation model, revealing that CUDA-optimized techniques do not directly transfer to Apple's unified memory architecture.
Towards Visually-Guided Movie Subtitle Translation for Indic Languages
This paper presents a case study on visually-guided movie subtitle translation for low-resource Indic languages, demonstrating that selective visual grounding improves translation quality while addressing temporal misalignment challenges.
Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild
Hy-MT2 is a family of fast, efficient multilingual translation models from Tencent, available in 1.8B, 7B, and 30B-A3B sizes, supporting 33 languages and outperforming previous open-source and commercial models.
Translation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource Languages
This paper investigates the feasibility of using translation-based fine-tuning as a resource-efficient alternative to native-language BERT models for low-resource languages, finding it comparable or superior in 53.3% of cases across six NLP tasks.