基于评论与转发的微博联合主题挖掘
2016-03-02赵臣升吴国文胡福玲
赵臣升 吴国文 胡福玲
摘要:微博文本简短、信息量少且语法随意,传统主题分类并不理想。Labeled LDA在LDA主题模型上附加类别标签协同计算隐含主题分配量使文本分类效果有所改进,但标签在处理隐式微博或主题频率相近的分类上,存在一定的模糊分配。本文提出的Union Labeled LDA模型通过引入评论转发信息丰富Label标签,进一步提升标签监督下的主题词频强度,一定程度上显化隐式微博、优化同频分配,采用吉布斯采样的方法求解模型。在真实数据集上的实验表明,Union Labeled LDA模型能更有效地对微博进行主题挖掘。
关键字:微博;主题挖掘;LDA;Union Labeled LDA;词频
中图分类号: TP391.1 文献标识码: A文章编号:2095-2163(2016)01-
Abstract:Microblog is brief and short, with a little information and irregular grammar, cause traditional method of topic classification effect is not satisfying. The Labeled LDA topic model attach classification label to original LDA model to help cooperative computing the implicit topics, but still exist some vague allocate when handling microblog whose topic frequency are neck and neck. This paper proposes to use the Union Labeled LDA model with comments and retransmissions which enrich the information of labels to enhance the supervision of topic frequency strength by themselves. The experimental results on actual dataset show that the Union Labeled LDA model can effectively mining the topics of Microblog.
Keywords:Microblog; Topic Mining; LDA; Union Labeled LDA; Word Frequence
0 引言
随着Web技术的日益完善和大数据时代的悄然来临,微博已经成为人们思想汇聚和信息交流的重要媒介,从海量数据中挖掘出有效的主题信息,分析其内在语义关联则正日显其现实突出的技术主导作用。微博本身文本简短、数据稀疏、语法随意和网络词汇大量出现,这些特点给传统文本挖掘算法带来了挑战[1-2]。
LDA(latent dirichlet allocation)主题模型是近年来文本挖掘领域热门研究方向,模型具有优秀的建模能力、文本分析降维能力和良好的概率模型扩展性,挖掘出的主题能帮助人们理解大数据文本背后的语义。LDA模型假设各主题权重在Dirichlet分布上相同,因此在处理隐性主题划分时存在部分主题强制分配的现象。Labeled LDA主题模型通过引入Label标签,单独对各类主题计算分布,在一定程度上克服了LDA的不足[3]。
本文在研究LDA和Labeled LDA模型的基础上,引入微博评论与转发数据信息,进一步丰富Labeled LDA模型的Label标签信息。……
