基于fastText的可视化作者归属模型
2021-07-11李逍顾长贵杨雷鑫陆祺灵
李逍 顾长贵 杨雷鑫 陆祺灵



摘 要:基于滑动窗口的方法,结合机器学习分类技术,可以判定文本的作者归属。但是此类方法需要精心挑选对应的文本特征,不同的文本特征选取可能会影响判定结果。针对以上问题,提出了一种基于快速文本分类(fastText)的文本作者归属判定模型。该模型融合滑动窗口的思想,引入词(字)向量、数据增强技术,从而充分利用文本信息、自动提取文本特征,并且以可视化的方式将结果呈现出来。使用该模型来检测《红楼梦》、《Roman de la Rose》的作者归属,实验结果表明《红楼梦》的前八十回与后四十回为不同作者所著、《Roman de la Rose》开篇4 058行(约50 000字)与后面17 724行(约218 000字)为不同作者所著。证明了Rolling-fastText模型判定文本作者归属的有效性。
关键词: 滑动窗口;作者归属;快速文本分类器;数据增强技术;可视化
文章编号: 2095-2163(2021)01-0014-06 中图分类号:TP391 文献标志码:A
【Abstract】Some methods are based on sliding window and machine learning, which can determine the authorship attribution of text. However, these methods require careful selection of text features, and different text features may affect the outcome of the authorship attribution. In response to the above problems, this paper proposes a model based on fastText classification to determine authorship attribution. The model incorporates the idea of the sliding window, introduces word (character) vectors and data enhancement technology, so as to make full use of text information and extract text features automatically, and presents the results in a manner of visualization. Finally, this paper uses the model to detect the authorship attribution of 《A Dream of Red Mansions》 and 《Roman de la Rose》. The experimental results show that the first 80 chapters and the last 40 chapters of 《A Dream of Red Mansions》 are written by different authors, the opening 4 058 lines (approximately 50 000 words) and the following 17 724 lines (approximately 218 000 words) of 《Roman de la Rose》 are written by different authors. It is proved that this model is effective to determine the authorship attribution.
【Key words】sliding window; authorship attribution; fast text classifier; data enhancement technology; visualization
0 引 言
在文體学中,文本作者归属判定是指从有争议的文本中提取作者风格信息,然后在一组“候选者”中确定最佳匹配的过程;题材识别则是找到一组在风格上相似的作品。在研究这些问题时,许多经典方法的目标都是计算语料库中文本之间的相似度,以便发现隐藏的模式或规律[1]。
早在20世纪上半叶就有学者开始使用统计学方法对文本进行作者归属研究[2],在研究《联邦党人文集》中12篇存在作者争议的文章时,Mosteller等人[3]提出了一种以贝叶斯公式为核心的分类算法,基于文本中的文体特征,如句子长度、词长或高频虚词的分布,推断出12篇文章的作者。……
