基于CRNN模型的中文场景文字识别
2021-07-01辜双佳栗智
辜双佳 栗智
摘 要:中文場景文字识别(STR)是光学字符识别(OCR)技术的重要研究方向,在拍照翻译、无人驾驶等领域广泛应用。但是,中文场景下的文字面临着字体和字符种类多、文字背景复杂等问题。本文着眼于“中国街景”图像,基于CRNN模型提出了一种免分割、端到端的中文场景文字识别方法。首先CNN提取图像卷积特征,然后RNN进行序列特征预测,其中Bi-GRU有效抑制梯度消失或梯度爆炸,Dropout可以防止过拟合,最后引入CTC作为损失函数解决训练时字符无法对齐的问题。本文用Python实现了算法,以较好的效果完成了实验。
关键词:中文OCR;CRNN;免分割;端到端;中国街景
Chinese Scene Text Recognition based on CRNN model
Gu Shuangjia1 Li Zhi21.School of Computer Science and Engineering, Chongqing University of Technology Chongqing400054 2.School of Computer Science, Chongqing UniversityChongqing 400044
Abstract: Chinese scene character recognition (STR) is an important research direction of optical character recognition (OCR) technology, which is widely used in the fields of photo translation and unmanned driving. However, the characters in Chinese scene are faced with many problems, such as many types of fonts and characters, complex text background and so on. This paper focuses on the "Chinese street view" image, and proposes a segmentation free, end-to-end Chinese scene text recognition method based on crnn model. Firstly, CNN extracts image convolution features, and then RNN performs sequence feature prediction. Bi Gru can effectively suppress gradient disappearance or gradient explosion, dropout can prevent over fitting. Finally, CTC is introduced as a loss function to solve the problem that characters cannot be aligned during training. In this paper, Python is used to implement the algorithm, and the experiment is completed with good effect.
Keywords: Chinese OCR; CRNN; No split; End-to-End; Chinese street scene
绪论
背景及意义
图像和视频中的文字包含了丰富而精确的高层语言描述,准确有效地提取这些文字信息在多媒体检索、人机交互、机器人导航和工业自动化等领域具有重要的应用[1]。中国是一个世界性的大国,中文字符种类繁多,有篆书、楷书、行书等多种字体,如图1所示。
目前CVPR、ICCV、ECCV 等国际顶级会议,已将场景文字检测与识别列为其重要主题之一,场景文字检测与识别技术广泛运用在图片搜索和无人驾驶[2]等方面,是当前研究的一个前沿课题。场景文字识别(STR, Scene Text Recognition)是在各种复杂情况下将图像输入翻译为自然语言输出,需要包括文字检测和文字识别两个步骤,文字检测即发现文字的位置和范围,文字识别即将文字区域转化为字符信息。如图2所示。
在实际应用中, 场景文字的检测和识别往往串联在一起,能同时检测到文字位置并对其进行识别的方法被称为“端到端”文字识别方法[3]。……
