基于n元词组表示的去噪 方法及其在跨语言映射中的应用
2016-05-03于墨赵铁军
于墨 赵铁军



摘 要:具有结构化输出的学习任务(结构化学习)在自然语言处理领域广泛存在。近年来研究人员们从理论上证明了数据标记的噪声对于结构化学习的巨大影响,因此为适应结构化学习任务的去噪算法提出了需求。受到近年来表示学习发展的启发,本文提出将自然语言的子结构低维表示引入结构化学习任务的样本去噪算法中。这一新的去噪算法通过n元词组的表示为序列标注问题中每个节点寻找近邻,并根据节点标记与其近邻标记的一致性实现去噪。本文在命名实体识别和词性标注任务的跨语言映射上对上述去噪方法进行了验证,证明了这一方法的有效性。
关键词:表示学习;半监督学习;去噪算法;自然语言处理;跨语言映射
中图分类号:TP181 文献标识号:A 文章编号:2095-2163(2015)06-
Noise Removing based on N-gram Representations and its Applications to Cross-Lingual Projection
YU Mo, ZHAO Tiejun
(School of Computer Science and Technology, Harbin Institute of Technology, Harbin 150001, China)
Abstract: Problems with structured predictions (structured learning) widely exist in natural language processing. Recent research found that compared to classification problems, structured learning problems were affected more seriously by label noises, suggesting the importance of noise removing algorithms for these problems. Inspired by the development of representation learning methods, the paper proposes a noise-removing algorithm for structured learning based on low-dimensional representations of sub-structures. The algorithm finds neighbors of each node in a sequential labeling task based on its associated n-gram representation, and then performs noise removing on the label of a node according to its consistency with the labels of its neighbors. Therefore the paper proves the effectiveness of the proposed algorithm on the cross-lingual projection of named entity recognition and POS tagging tasks.
Keywords: Representation Learning; Semi-supervised Learning; Noise Removing; Natural Language Processing; Cross-lingual Projection
0引 言
很多自然语言处理(NLP)技术依赖于有监督学习方法训练的模型,而有监督学习方法的性能不仅依赖于标记样本的数量,也依赖于标记样本的质量。当标记样本中发生了错误,即标记存在着噪声时,学习得到的分类器的推广能力会受到影响(相应的理论分析被称为噪声可学习性理论,见文献[1-2])。由于具有结构化输出的学习任务(结构化学习)在NLP领域的广泛存在性和重要性,于墨等人[3]将上述噪声可学习性理论推广到结构化学习问题中,证明了对同样的噪声率,学习的难度随结构的复杂性的增加而有所提升。这一理论分析结果说明了当NLP的结构化学习任务包含噪声时,一个好的去噪算法对模型性能的改进完善极为重要。……
