融入多特征的汉韩双语自动句对齐方法
2021-07-11刘晨阳唐慧丰
刘晨阳 唐慧丰



摘 要:为解决汉韩双语平行语料库资源匮乏以及传统句对齐算法面向跨语系语言准确率较低的问题,提出了融合特征的汉韩双语句对齐方法。首先将Bi-LSTM融入孪生神经网络构建句对齐模型,用以分别提取汉语和韩语句子的特征并进行对齐。之后基于语料的特点提取句对齐特征融入输入层。通过与传统Bi-LSTM和不同特征组合的孪生Bi-LSTM的对比实验证明,融入特征的孪生Bi-LSTM方法在句对齐任务中具有更优越的性能。
关键词: 汉语-韩语;自动句对齐;孪生Bi-LSTM;多特征;平行语料库
文章编号: 2095-2163(2021)01-0028-04 中图分类号:TP391 文献标志码:A
【Abstract】In order to solve the problem of the lack of resources in the Chinese-Korean bilingual parallel corpus and the low accuracy of traditional sentence alignment algorithms for cross-lingual languages, a method of Chinese-Korean bilingual sentence alignment with multiple-features is proposed. Firstly, Bi-LSTM is integrated into Siamese Neural Network to construct sentence alignment model, which is used to extract the features of Chinese and Korean sentences and align them. After that, sentence alignment features are extracted based on the features of corpus and integrated into the input layer. Compared with the traditional Bi-LSTM and the Siamese Bi-LSTM with different feature combinations, it is proved that the Siamese Bi-LSTM method with features has better performance in sentence alignment task.
【Key words】the Chinese-Korean; automatic sentence alignment; Siamese Bi-LSTM; multi-features; parallel corpus
0 引 言
雙语句对齐是指将语料中的双语互译句对进行匹配,该技术可以将篇章、段落级别对齐的语料进一步细化为句对齐语料,从而构建高质量的双语平行语料库。目前,汉韩双语平行语料库资源较为匮乏,构建方式也多以人工为主,因此通过自动句对齐技术高效地构建汉韩双语平行语料库,对以机器翻译为代表的汉韩自然语言处理任务有着重大意义。
1 相关研究
研究可知,双语句对齐的方法主要有基于长度的方法、基于词汇的方法和基于长度与词汇相结合的方法。对此拟展开研究论述如下。
1.1 基于长度的句对齐方法
该方法最初由Gale&Church提出,其依据是源语言与译文文本长度具有关联性,据此区分对齐与非对齐句对[1]。传统的对齐算法多以字节、字符或词数作为长度计量单位,此后的研究者利用其他元素计算句子长度,如张霞等人[2]将句子所含的动词、名词、形容词等词语作为句长计量单位,在英汉句对齐上取得了良好的效果。……
