基于学习的OCR字符识别
2018-09-17肖坚
肖坚
摘 要: OCR(Optical CharacterRecognition)是通过检测字符每个像素亮度的模式确定其形状,然后用字符识别方法将形状翻译成计算机文字的过程。文章利用Java语言实现OCR步骤,包括像素二值化,图像分割,训练识别和输出等。测试开发是在web验证码识别场景中进行的,web验证码是将一串随机产生的符号,生成为图片,再加上一些干扰线,使之能有效防止恶意注册和灌水。通过测试表明,该方法可行、有效;拒识率、误识率低;识别速度快,具有一定的实用意义。
关键词: OCR; 验证码; 文字识别; 干扰线; 拒识率; 误识率
中图分类号:TP3 文献标志码:A 文章编号:1006-8228(2018)07-48-04
Abstract: OCR (Optical Character Recognition) is the process of translating the shape, which is determined by detecting the pattern of the brightness of each pixel of the character, into computer text by character recognition method. In this paper, OCR procedure is implemented in Java language, including pixel binarization, image segmentation, recognition training and output. The test development is carried out in the Web verification code identification scene. The Web verification code is a string of randomly generated symbols, generated as a picture, and a number of interference lines added on it, so that it can effectively prevent malicious registration and irrigation. The test shows that the method is feasible and effective, with low rejection rate and error rate, fast recognition speed, and has practical significance.
Key words: OCR; verification code; character recognition; interfering line; rejection rate; error rate
0 引言
識别原理及实现方法:OCR采用光学的方式,将纸质或图片文档中的文字转换成为黑白点阵的图像文件,并通过识别软件将图像中的文字转换成文本格式,供文字处理软件进一步编辑加工的技术。
识别过程一般分以下几个步骤:
首先训练学习过程,分成图像生成,预处理,图像分割三个步骤,图像分割不是简单的将图片等份分割,常常需要程序员像素级微调,才能最终生成合适的样本。
其次才是识别过程,识别前3个步骤和训练学习是一致的,而且各个步骤处理的参数必须和训练完全一样,否则获取的单字符图片完全没有可比性,识别步骤是把单字符图片和样本数据一一比较,获得最为接近的作为结果。识别流程如图1所示。
以下是各个步骤详细内容及代码实例。
1 图像预处理
预处理过程就是用阈值分割法把图片上每个像素二值化,像素红绿蓝在一定范围内置成白色,反之黑色。……
