基于无监督学习的中文电子病历分词
2014-04-29张立邦关毅杨锦峰
张立邦 关毅 杨锦峰
摘 要:电子病历中包含大量有用的医疗知识,抽取这些知识对于构建临床决策支持系统和个性化医疗健康信息服务具有重要意义。自动分词是分析和挖掘中文电子病历的关键基础。为了克服获取标注语料的困难,提出了一种基于无监督学习的中文电子病历分词方法。首先,使用通用领域的词典对电子病历进行初步的切分,为了更好地解决歧义问题,引入概率模型,并通过EM算法从生语料中估计词的出现概率。然后,利用字串的左右分支信息熵构建良度,将未登录词识别转化为最优化问题,并使用动态规划算法进行求解。最后,在3000来自神经内科的中文电子病历上进行实验,证明了该方法的有效性。
摘关键词:中文电子病历;无监督分词;EM算法;分支信息熵;动态规划
中图分类号: TP391 文献标识码: A 文章编号:2095-2163(2014)02-
An Unsupervised Approach to Word Segmentation in Chinese EMRs
ZHANG Libang,GUAN Yi,YANG Jinfeng
( School of Computer Science and Technology,Harbin Institute of Technology,Harbin 150001,China)
Abstract: Electronic medical records(EMR) contain a lot of useful medical knowledge. Extracting these knowledge are important for building clinical decision support system and personalized healthcare information service. Automatic word segmentation is a key precursor for analysis and mining of Chinese EMRs. In order to overcome the difficulties of obtaining labeled corpus, the paper proposes an unsupervised approach to word segmentation in Chinese EMRs. First, the paper uses a lexicon of general domain to generate an initial segmentation. To deal with the ambiguity problem, the paper also builds a probabilistic model. The probabilities of words are estimated by an EM procedure. Then the paper uses the left and right branching entropy to build goodness measure and regards the recognition of unknown words as an optimization problem which can be solved by dynamic programming. Finally, to prove the effectiveness of our approach, experiments are conducted on 3,000 copies of Chinese EMRs from the Department of Neurology.
Key words: Chinese EMRs; Unsupervised Segmentation; EM Algorithm; Branching Entropy; Dynamic Programming
0 引 言
电子病历是指医务人员在医疗活动过程中,使用医疗机构信息系统生成的面向患者个体的数字化医疗记录[1]。近年来,随着医院信息化建设的发展,电子病历的使用在临床中已经逐渐普及。电子病历包含了关于病人个体健康信息的全面、详实、专业、即时、准确的描述,是一种非常宝贵的知识资源。通过分析和挖掘电子病历,可以从中获得大量的医疗知识[2],而这些知识可应用于临床决策支持[3]和个性化医疗健康信息服务[4]等方面。电子病历由结构化数据和非结构化数据组成。其中,自由文本形式的非结构化数据是电子病历中最为重要的部分,包括主诉、现病史、病程记录、病历小结。……
