Tag
WaseiGo is a product that teaches over 1,000 Japanese wasei-eigo words through illustrated dialogues, audio, a dictionary, and quizzes, focusing on English-looking words with reinvented meanings.
AIVTuber Shizuku-chan is planning a secret summer event on August 15th and 16th, heightening fans' expectations.
JOR-Bench is a collection of five Japanese-language benchmarks for evaluating large language models on operations research problem formulation, translated from existing English benchmarks. Evaluation shows overall language-neutral performance with minor cross-lingual differences.
This paper investigates the computational cost of reasoning in non-English languages, using Japanese as a case study, to highlight inefficiencies in current AI systems.
YOMI-Bench is a benchmark designed to evaluate large language models on kanji reading and phonological understanding in Japanese. It consists of four tasks and reveals that even Japanese-specific models perform poorly on generation tasks requiring kanji reading knowledge.
The MIT textbook 'Mathematics for Computer Science' has been translated into Japanese and is available at inzkyk.xyz.
Toku Reader is a language learning app that lets users read and listen to native Japanese and Chinese content, with tap-to-translate functionality.
Dango is a 1.8B-parameter LLM trained strictly on Japanese (L1) then fine-tuned on English (L2) to study language transfer effects in second language acquisition. The model filters English contamination from the pretraining corpus and shows human-like L2 production patterns.
This paper applies the likelihood ratio framework for forensic authorship attribution to Japanese texts, fusing stylometric features with embedding-based systems to improve discrimination and calibration.
A blog post exploring a 1994 CD-ROM that archives the XD FirstClass Network BBS from Japan's Kansai region, including its client software and community posts, offering a rare glimpse into a pre-internet online community.
HakushoBench is a Japanese chart and table VQA benchmark built from governmental white papers to evaluate vision-language models' understanding of complex visual data, challenging open-weight models with a 58.6% accuracy and a 34.9-point gap to proprietary models.
This paper investigates how character-level transformer models generalize to irregular verb subtypes in Japanese past-tense inflection. Controlled experiments show that including irregular examples can improve generalization, challenging the assumption that regularity simplifies learning.