APP下载

基于字符的递归神经网络在中文语言模型中的研究与实现

2018-10-21伍逸凡朱龙娇石俊萍

现代信息科技 2018年8期

伍逸凡 朱龙娇 石俊萍

摘 要:本文通过对基于字符的长短记忆递归神经网络的研究与实现,探究了其在自然语言模型中的应用,并选用了小说《挪威的森林》对递归神经网络进行了训练与文本生成,总结了不足之处,探讨了未来应该解决的问题与研究方向。研究结果表明递归神经网络仅能学会字与字或词与词之间在表面的连接或变化关系,而自然语言不仅仅是文字表面的异同,更多的是字里行间中情感或思维上的变化,这些是一组序列数据所不能表达的。因此,未来自然语言模型应更加注重对于文字间情感和思维的学习,构建更接近自然语言的模型。

关键词:长短记忆单元;递归神经网络;自然语言处理;字词嵌入

中图分类号:TP391.1;TP183 文献标识码:A 文章编号:2096-4706(2018)08-0012-03

Abstract:Through the research and implementation of character-based recursive neural networks of long and short memory,this essay explored its application in natural language models,and selected the novel Forest in Norway to train recurrent neural networks and generate the corresponding text. Summed up the shortcomings,discussed the problems and research directions that should be solved in the future. The research results show that the recurrent neural network can only learn the connection or change relations between word and words or words on the surface,and the natural language is not only the similarities and differences between the surface of the words,but also more changes in emotions or thoughts between lines. These are a group of sequence data far from being able to express,so in the future natural language models should pay more attention to the study of sentiment and thinking between words to build a model that is closer to natural language.

Keywords:long short term memory unit;recursive neural network;natural language processing;word embedding

0 引 言

自然語言是人类智慧的结晶,而自然语言处理(Nature Language Processing)是尝试通过计算机技术结合概率论与数理统计等数学方法,让计算机理解或生成自然语言的技术。近年来,自然语言处理技术随着时代的进步逐渐兴起,并迅速发展,让计算机正确有效地理解和处理人类自然语言,并进一步实现与人类的对话,已成为当今具有巨大挑战性的难题。

随着时代的变迁与技术的发展,在自然语言处理中,词汇的表征由最先的One-hot编码发展为如今的词嵌入编码,词嵌入将词汇嵌入到一个低纬而紧凑的向量空间中,大大加强了词汇间的联系;文本的处理由最先的N-Grams模型发展为如今的递归神经网络模型,递归神经网络通过神经元在时序上的连接,成功捕获了文本长短期的顺序依赖关系;……

登录APP查看全文