基于SentencePiece的中医学分词模型建模研究
2021-07-09刘双巧周璐李彩艳袁慧敏张异卓李昱达刘锦钢郑丰杰孙燕李宇航
刘双巧 周璐 李彩艳 袁慧敏 张异卓 李昱达 刘锦钢 郑丰杰 孙燕 李宇航



摘要 目的:探索构建适用于中医学领域的分词模型。方法:采用基于SentencePiece的无监督学习分词方法,提出利用出版教材、名家著作及中医临床病历这3种不同类型的文献构建中医学分词模型;选择中医临床病历、名医医案作为测试集进行模型测试。结果:中医学分词模型在测试集中的Kappa系数为0.79(一致性程度很高),准确率为0.84,宏观精确率为0.84,宏观召回率为0.83,宏观f1得分为0.83。结论:所构建的分词模型对于中医学专业术语有着较好的切分效果,表明该方法可运用于中医学领域的分词模型的构建,可为进一步地研究中医学分词提供方法学参考。
关键词 分词;中文分词;分词模型;无监督学习;无监督分词;SentencePiece
Research on Modeling of Traditional Chinese Medicine Word Segmentation Model Based on SentencePiece
LIU Shuangqiao,ZHOU Lu,LI Caiyan,YUAN Huimin,ZHANG Yizhuo,LI Yuda,LIU Jingang,ZHENG Fengjie,SUN Yan,LI Yuhang
(School of Traditional Chinese Medicine,Beijing University of Chinese Medicine,Beijing 100029,China)
Abstract Objective:To explore the construction of word segmentation model suitable for the field of traditional Chinese medicine (TCM).Methods:Using the unsupervised learning word segmentation method based on SentencePiece,we proposed to use 3 different types of documents,such as published textbooks,famous works and clinical medical records of TCM,to construct a word segmentation model of TCM; choosed the clinical records of TCM and medical records of famous doctors as the test set for model testing.Results:The Kappa coefficient of the word segmentation model of TCM established in this study was 0.79 (with substantial consistency),the accuracy rate was 0.84,the macro precision rate was 0.84,the macro recall rate was 0.83,and the macro f1 score was 0.83.Conclusion:The word segmentation model constructed by this study has a good segmentation effect on the terminology of TCM,indicating that this method can be applied to the construction of the word segmentation model in the field of TCM,and can provide a methodological reference for further study of TCM word segmentation.
Keywords Word segmentation; Chinese word segmentation; Word segmentation model; Unsupervised learning; Unsupervised word segmentation; Sentence piece
中圖分类号:R2-03文献标识码:Adoi:10.3969/j.issn.1673-7202.2021.06.024
中医学发展历程中产生了众多的医学文献,这些文献中蕴含着丰富的医药知识及临证经验,如何快速有效地从这些文献中提取信息并加以利用,是中医现代化研究过程中面临的一大难题。中文分词是信息处理过程中的基础与关键[1],词是最小的能够独立活动的有意义的语言成分[2],中文分词即是将没有天然分隔符号(如英文的空格)的汉字序列切分成词序列,如将“患者发热头痛三天”利用分词工具切分为“患者”“发热”“头痛”“三天”“。”,即提取句子中的词汇,以便于进一步实现LDA主题挖掘[3]、命名实体识别[4]、信息提取[5]、文本分类[6]等研究。因此,在中医学文献挖掘研究的过程中,对其文本作分词处理,可以为下一步研究工作打下基础。……
