APP下载

基于改进BFDNN的远距离语音识别方法

2018-07-28王旭东王冬霞周城旭

电脑知识与技术 2018年15期

王旭东 王冬霞 周城旭

摘要:针对复杂环境下远距离智能语音识别的问题,提出了一种基于深度神经网络(DNN)的波束形成与声学模型联合训练的改进方法。即首先提取麦克风阵列信号之间的多通道互相关系数(MCCC)来估计频域波束形成器权重,进而对阵列信号进行滤波得到增强信号,然后对增强信号提取梅尔滤波器组(Fbank)特征送入声学模型进行训练识别,最后再将识别信息反馈回波束形成网络(BFDNN)来更新网络参数。实验通过Theano与Kaldi工具箱结合搭建大词汇量远距离语音识别系统进行。仿真结果表明了该方法的有效性。

关键词: 远距离语音识别; 波束形成; 麦克风阵列; MCCC; BFDNN

中图分类号:TN912.34 文献标识码:A 文章编号:1009-3044(2018)15-0182-04

Far-field Speech Recognition Method based on Improved BFDNN

WANG Xu-dong,WANG Dong-xia,ZHOU Cheng-xu

(School of Electronic and Information Engineering Liaoning University of Technology, Jinzhou 121001, China)

Abstract: For the speech recognition in far field scenes, an improved method is introduced which trains jointly beamforming based on Deep Neural Networks (DNN) and acoustic model. Specifically, the parameters of a frequency-domain beamformer are first estimated by multichannel cross-correlation coefficient (MCCC) extracted from the microphone channels, and then the array signals filtered by the parameters to form an enhanced signal, Mel FilterBank (Fbank) features are thus extracted from this signal and passed to acoustic model for training and recognition. Finally the output information of beamforming DNN (BFDNN) is used to update the whole network parameters. A far-field large vocabulary speech recognizer is proposed to implement by Theano coupled with the Kaldi toolkit. The simulation results show that the proposed system performance has improved.

Key words: far-field speech recognizer; beamforming; microphone channels; MCCC; BFDNN

近距離场景下的语音识别取得了令人满意的结果,实现了较高的识别准确率。但是由于噪声和混响等因素的影响,远距离场景下的语音识别仍然具有很大的挑战性[1-4],有待进一步改进和完善。而在实际的生产应用中,更多时候恰恰处于远距离场景。因此,对远距离语音识别的研究显得尤为重要。随着深度神经网络方法的引进,DNN-HMM框架下的语音识别准确率和之前相比有了显著的提高[5-7],所以,基于DNN的语音识别成为现阶段人们的研究热点。

在提高多通道远场下自动语音识别系统的鲁棒性方面,波束形成是一种重要的处理方法。文献[8]提出了一种基于学习的深度波束形成网络,提取麦克风阵列之间的GCC信息,使用BFDNN来估计波束形成参数,并与声学模型部分进行联合训练,从而提高了语音识别的鲁棒性。……

登录APP查看全文