APP下载

基于LDA模型和聚类算法的城市热点推荐与应用

2018-09-05王诗童刘美玲孙立研

智能计算机与应用 2018年3期

王诗童 刘美玲 孙立研

文章编号: 2095-2163(2018)03-0136-04中图分类号: 文献标志码: A

摘要: 關键词: application of city hot sites

(College of Information and Computer Engineering, Northeast Forestry University, Harbin 150040, China)

Abstract: According to the functions of short text posting and sign-in to elicit the details post by the users. Cutting the vast short texts and geography positions to the phrases by LDA(Latent Dirichlet Allocation) Model, in order to count up the frequency of every phrase, and then obtain the hot geography positions, as well as label them on the map. With the Spatial Distance Clustering Algorithm, optimizing the recommendation function when the users offer their situations and restrict the searching conditions. And the system shows the details of some active sites, such as shopping malls, hot sites and restaurants to recommend to the users.

Key words:

基金项目: 国家自然科学基金(61702091);省自然科学基金(F2015037); 东北林业大学大学生创新训练计划项目(201610225196)。

作者简介: 王诗童(1996-),女,本科生,主要研究方向:数据分析; 刘美玲(1981-),女,博士,讲师,CFF高级会员,IEEE CS会员,ACM会员,主要研究方向:自然语言处理、数据挖掘、数据分析;孙立研(1994—),男,硕士研究生,主要研究方向:林业信息工程、空间数据挖掘。

通讯作者: 收稿日期: 引言

随着计算机技术的进步和Web2.0的日益完善,社交媒体在不断向前发展。在这其中,新浪微博是较为广泛应用和流行的社交媒体软件。与其他社交软件相比,新浪微博具有信息发布方式多,信息传播速度快,交互性强等特点。因此,利用新浪微博上用户发布的文本进行数据分析和挖掘亦可以获取大量潜在的且有价值的信息。

本文利用新浪微博开放平台获取的用户数据,采用LDA模型和多距离空间聚类算法,收集微博数据,挖掘出其中的地理位置信息和相应的用户评价,获取用户感兴趣的内容,在地图中形成定位点并标注,并向用户进行推荐。

1相关工作

1.1文本主题聚类的方法

基于文本主题的聚类,顾名思义,就是以文本为主题,即描述对象的标准,将数据聚集成不同的类[1]。Ivan Titov等[2]人提出一种情感总结的文本和方面评分的联合模型来挖掘文本中相关联的主题,提高情感分析结果的准确性和高效性。Chao Shen等[3]人提出基于参与者的事件提取方法zooms-in 来侦测和捕捉与参与者相关的突发性和连续性的重要子事件。……

登录APP查看全文