基于迁移学习优化的DCNN语音识别技术
2020-09-21张安安邓芳明
张安安 邓芳明



摘 要: 针对现有语音识别技术识别精准度低的问题,提出一种基于深度卷积神经网络算法与迁移学习相结合的语音识别技术。由于深度卷积神经网络应用范围有限,当输入输出参数发生变化时,需要重新开始构建,体系结构训练时间过长,因此,采用迁移学习方法有利于降低数据集规模。仿真实验结果表明,迁移学习不仅适用于源数据集与迁移问题的目标数据集比较,而且也适用于两种不同数据集情况,小数据集应用不仅有利于降低数据集生成时间和费用,而且有利于降低模型培训时间和对计算能力的要求。
关键词: 语音识别; 深度卷积神经网络; 迁移学习; 数据集规模; 识别精度; 培训时间
中图分类号: TN912.34?34; TN925 文献标识码: A 文章编号: 1004?373X(2020)17?0069?03
Abstract: Since the recognition accuracy of existing speech recognition technology is low, a speech recognition technology based on deep convolution neural network algorithm is proposed. Due to the limited application scope of deep convolutional neural network (DCNN), when the input and output parameters change, the deep convolution neural network needs to be rebuilt and the training duration of architecture is time?consuming. Therefore, the migration learning method is adopted, which is beneficial to the reduction of the data set scale. The results of simulation experiments show that the migration learning is not only suitable for comparing the source data set with the target data set of migration problem, but also suitable for situations of two different data sets. The application of small data sets is favorable to the reduction of not only the time and cost of data set generation, but also the training duration and computational ability requirement of the model.
Keywords: speech recognition; deep convolution neural network; transfer learning; data set scale; recognition precision; training duration
0 引 言
语音识别是机器的听觉系统,能够实现人与机器的交流[1]。一般来说,语音识别的方法通常分为以下3种:基于声道模型和语音知识方法、模板匹配方案以及利用人工神经网络方法[2]。人工神经网络方法模拟了人类神经活动,相比于传统的语音识别法,在建模能力以及语音识别准确率上都有了很大的提升。
深度学习的概念源于人工神经网络[3],2009年深度学习首次被应用于语音识别任务中[4]。根据目前语音识别技术的发展现状,基于深度学习的语音识别技术算法主要分为长短时记忆(Long Short?term Memory,LSTM)网络[5]、深层神经网络(Deep Neural Network,DNN)[4]、卷积神经网络(Convolutional Neural Network,CNN)[6]。CNN通过采用局部滤波和最大池化技术可以获得更好的鲁棒性,因此,CNN近年来在图像、视频及语音识别领域得到了广泛的关注[7?8]。……
