基于密度特征与KNN算法的最优特征维数选择
2018-08-21孙国栋梅术正汤汉兵周振
孙国栋 梅术正 汤汉兵 周振
摘 要: 为了保证基于同步触发双相机的仪表复杂字符识别中误识率为0,采用K最近邻算法对仪表字符特征进行训练分类,结合字符自身特点,提出最优特征提取与高宽维度选择方法,并设计实验获取1~4 096维密度特征的误识率与运行时间。实验结果表明,图像的密度特征总维度在230~260,高宽维度比接近1.4时,误识率为0的概率最大。该规律对采用KNN算法进行分类识别时最优密度特征维数选择具有一定指导意义。
关键词: 复杂仪表; 特征维数; 误识率; KNN算法; 密度特征; 最优特征
中图分类号: TN911.73?34; TP23 文献标识码: A 文章编号: 1004?373X(2018)16?0080?04
Abstract: The K?Nearest Neighbor (KNN) algorithm is adopted to train and classify the character features of the instrument to guarantee the zero error recognition rate during the complex character recognition of the instrument based on the double cameras with synchronous trigger. In combination with the features of the character, a method of extracting the optimum feature and selecting the width and height dimensions is proposed. An experiment was designed to obtain the error recognition rate and running time of 1~4096 dimensions density features. The experimental results show that when the total number of dimensions of the image density feature is 230~260 and the dimension ratio of height to width is close to 1.4, the probability of zero error recognition rate reaches the maximum. This rule has a certain guiding significance to the selection of the optimal density feature dimension when the KNN algorithm is used for classification and recognition.
Keywords: complex instrument; feature dimension; error recognition rate; KNN algorithm; density feature; optimum feature
国家计量单位需定期检测电力仪表以保证其测量精度,对标准表与被试表的读数进行直接比较是最常用的校对方法[1],而两表的同步触发尤为关键。特别是对不含通信接口的仪表同步相对较难,因此通过同时触发双相机分别识别两表的读数成为了一种新的解决方案。为了保证仪表上复杂字符误识率[2]为0,有必要研究其字符特征提取与维度选择以保证100%的识别率与最低的运行时间。目前常用的字符特征有结构特征[3]和统计特征[4]。其中,结构特征包括圈、断点、交叉点、笔画、轮廓等;统计特征包括点的密度测量、矩、特征区域等。结构特征能描述字符的结构,误识率较低,但相对复杂[5]。统计特征算法简单,鲁棒性强,统计特征的分类器易于训练,且误识率低,但要提取能反映模式精细结构的特征[6]较难。K最近邻(K?Nearest Neighbor,KNN)分类算法[7?8]根据样本邻近性找出最近的K个训练样本作为其K近邻,并根据K近邻采用投票策略预测待分类样本类型[9]。……
