开发用于从韩文发票中提取信息的OCR模型

arXiv cs.CL 论文

摘要

研究人员提出了一种结合深度学习与图像预处理技术的OCR模型,用于自动提取韩文发票中的关键信息,在自建数据集上达到了87%的F1分数,且处理耗时极低。

arXiv:2609.35796v1 Announce Type: new Abstract: Invoices are commercial documents that contain various pieces of information, including the purchased items, time, and total money. Making the extraction of important information crucial. The stored information serves different purposes. Korean language is the native language of about 80 million people, playing an important role in not only South and North Korea but also in many other countries such as Vietnam, Philippine where a large number of Korean companies are located. In this context, to automatically extract proper information from the invoices with Korean language, we propose an efficient Optical Character Recognition (OCR) model in which a deep learning model is combined with some image preprocessing techniques. The proposed OCR model is assessed in a rich set of collected invoices showing that 87% F1-score can be achieved with negligible time processing.
查看原文
查看缓存全文

缓存时间: 2026/09/30 09:47

# Developing an OCR model for Extracting Information from Invoices with Korean Language
Source: [https://arxiv.org/abs/2609.35796](https://arxiv.org/abs/2609.35796)
[View PDF](https://arxiv.org/pdf/2609.35796)

> Abstract:Invoices are commercial documents that contain various pieces of information, including the purchased items, time, and total money\. Making the extraction of important information crucial\. The stored information serves different purposes\. Korean language is the native language of about 80 million people, playing an important role in not only South and North Korea but also in many other countries such as Vietnam, Philippine where a large number of Korean companies are located\. In this context, to automatically extract proper information from the invoices with Korean language, we propose an efficient Optical Character Recognition \(OCR\) model in which a deep learning model is combined with some image preprocessing techniques\. The proposed OCR model is assessed in a rich set of collected invoices showing that 87% F1\-score can be achieved with negligible time processing\.

## Submission history

From: Bao\-Minh Dinh \[[view email](https://arxiv.org/show-email/56081d74/2609.35796)\] **\[v1\]**Thu, 17 Sep 2026 15:13:03 UTC \(624 KB\)

相似文章

@DataScienceDojo: 一家中国公司刚刚开源了一个𝐎𝐂𝐑,它修复了大多数AI驱动OCR工具默默挣扎的一个问题:……

X AI KOLs Timeline

Unlimited-OCR,一家中国公司新发布的开源OCR模型,解决了AI OCR工具中常见的内存增长问题——无论文档多长,内存使用都保持平稳,从而能够在32K上下文长度下一次性读取数十页。它采用MIT许可,拥有30亿参数,支持多语言,并且在GitHub上已经广受欢迎。

@berryxia: 卧槽,这一波直接把DeepSeek的“墙角挖到了啊”! 昨晚看到HuggingFace刷到这个有意思的OCR开源模型和原来背后有趣的故事。 这个OCR模型直接与传统的OCR模型完全不同! 光着速度和精准度真的就无敌了~~ 先说说背景,熟悉…

X AI KOLs Timeline

百度开源了Unlimited OCR模型,采用R-SWA注意力机制,可一次性处理数百页文档,无需分页,KV Cache恒定。该模型创新性地借鉴了人类抄书时的注意力模式,并与DeepSeek OCR有技术渊源,引发了对人才流动的关注。