APP下载

面向数字资源的自动标签模型

2020-08-26雷智文黄玲

哈尔滨理工大学学报 2020年3期

雷智文 黄玲

摘 要:针对数字资源标签数量不足,获取困难的问题,提出了一种新的自动标签方法,对于收集的公共文化资源数据集和其它公开数据集,能够有效的进行标签扩展。提出过程依据神经网络理论和生成学习理论,采用隐含狄利克雷分布(latent dirichlet allocation, LDA)和Word2Vec方法分别对资源和初始标签进行处理,生成资源和初始标签的表示向量,然后以此两种向量作为深度结构语义模型的输入,建立面向数字资源的自动标签模型。从结果来看,该方法的标签扩展效果在精确度、平均排序倒数、平均准确率等指标上表现上总体优于文中提到的其它对比方法,能够解决某些情况下资源标签不足的问题,提高资源的利用率。

关键词:标签扩展;隐含狄利克雷分布;Word2Vec

DOI:10.15938/j.jhust.2020.03.022

中图分类号: TP181

文献标志码: A

文章编号: 1007-2683(2020)03-0144-07

Abstract:In this paper, we proposed a novel automatic tagging system which aimed at the lack of tags about digital resources and the difficulty of extending tags. This tagging system can effectively extend tags for public cultural resources we collected and other public data sets. The algorithm of tagging system based on neural network and generative learning. We use Latent Dirichlet Allocation (LDA) and Word2Vec to process resources and initial tags, generating the representation vectors of resources and initial tags, then use these two kinds of vector to build this automatic tagging system focused on digital resources. From the results, the Precision, MRR, MAP and other indexes of this method is better than other comparison tagging methods mentioned in this paper, and it can solve the lack of tags in some cases. Increasing utilization of resources.

Keywords:automatic tagging; latent dirichlet allocation; Word2Vec

0 引言

在互联网应用中,对象和标签的结合方法是一种非常有用的技术,标签能够大幅度提高信息检索的效率,高质量的标签还能够帮助对资源进行分类和整合,使得资源的利用变得更加有效。对图像、视频及文本等资源进行自动标注的方法通常有两类,一类是关键词提取方法,另一类是近年来逐渐兴起的关键词生成方法,关键词提取只依赖于文本本身的信息,不能生成新的信息,标签提取的效果已经到了瓶颈。因此,能够生成新信息的标签提取方法近年来越来越受到人们的重视,这种新的标签提取方法和传统基于关键词提取的方法最主要的不同点就是它往往拥有更加优化的词库和非线性結构,从而能够取得更好的标签提取效果。……

登录APP查看全文