TeochewBench: A Human-Reviewed Benchmark for Teochew Hanzi Translation
Summary
TeochewBench is a human-reviewed benchmark designed to evaluate translation systems for Teochew Hanzi, providing a resource for advancing research in dialect-specific natural language processing.
View Cached Full Text
Cached at: 09/17/26, 09:09 AM
# TeochewBench: A Human-Reviewed Benchmark for Teochew Hanzi Translation Source: [https://arxiv.org/abs/2609.18156](https://arxiv.org/abs/2609.18156) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
Researchers introduce HoWToBench, a large-scale Chinese writing benchmark with 1,302 instructions across 12 genres, and Tree-of-Writing (ToW), a tree-structured evaluation method that achieves 0.93 Pearson correlation with human judgments while mitigating biases in LLM writing assessment.
OpenSTBench: Beyond Semantic Evaluation for Speech Translation
OpenSTBench is a unified multidimensional evaluation framework for speech translation systems that jointly assesses translation quality, speech quality, speaker preservation, emotion fidelity, and latency across both S2TT and S2ST systems in offline and streaming settings. The framework addresses the gap left by fragmented evaluation protocols and provides a reproducible benchmark for comparing heterogeneous speech translation systems.
PaliBench: A Multi-Reference Blueprint for Classical Language Translation Benchmarks
Introduces PaliBench, a multi-reference benchmark for Pali-to-English translation using independent translations from multiple scholars, and a reusable methodology for creating similar benchmarks for classical languages.
TW-LegalBench: Measuring Taiwanese Legal Understanding
TW-LegalBench is a benchmark for evaluating large language models on Taiwanese legal understanding, including over 16,000 multiple-choice questions, 117 essay questions, and 14,000 legal judgment prediction instances. Results show top models exceed the passing threshold for lawyers but fall short for judges, highlighting challenges in reliable legal text generation.
SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages
Introduces SEATauBench, the first agent-focused evaluation framework for Southeast Asian languages, adapting τ²-Bench to Mandarin, Vietnamese, Thai, Indonesian, and Filipino, and reveals a significant capability gap when moving from English to localized settings.