APP下载

基于有限状态自动机的藏文音节组织研究

2021-06-08更桑吉安见才让

计算机时代 2021年1期

更桑吉 安见才让

摘  要: 通过对藏文的字形特征、拼写规律,以及文法规则的分析和研究,实现藏文词语的实时检错。借助形式语言有限状态自动机的方法,对藏文字结构中的基字、前加字、上加字、下加字、后加字、再后加字之间的搭配规则设计了状态图和邻接矩阵。该方法提高了藏文文本质量,使原本复杂的书面语法规则变得简单直观,从而使符合现代藏文音节组织结构的词语能实时检错。该研究为实现藏文的自动校对提供了基础。

关键词: 藏文; 文法规则; 有限状态自动机; 校对

中图分类号:TP391.1          文献标识码:A     文章编号:1006-8228(2021)01-65-03

Research on Tibetan syllable organization using finite state automata

Geng Sangji, Anjian Cairang

(School of computer, Qinghai University for Nationalities, Xining, Qinghai 810007, China)

Abstract: By analyzing and studying the characteristics of Tibetan character, the spelling rule and grammar rule, the real-time error detection of Tibetan words is realized. With the help of finite state automata of formal language, this paper designs the state diagram and adjacency matrix for the matching rules among the basic characters, prefix letters, superfixed letters, subjoined letters, suffixed letters and up-adding characters in the Tibetan character structure. This method improves the quality of Tibetan text, makes the complex original written grammar rules simple and intuitive, so that the words in line with the modern Tibetan syllable organization structure can be error detected in real time. This research provides a basis for the realization of Tibetan automatic proofreading.

Key words: Tibetan; grammar rules; finite state automata; proofreading

0 引言

隨着藏区人民对信息数字化需求的提高,学习和利用信息数字化的技术手段来记载和传承民族文字显得非常重要,而人工智能领域对藏语信息研究发展有着不可忽略的重要性。通过研究藏文音节和字形结构[1-2],判断基字所在位置、特殊音节的处理等步骤解决藏文构件元素的识别[3];基于规则和CNN模型、基字定位等方法实现检错[4-6],这些方法都各有利弊,因此本研究提出基于有限状态自动机的藏文音节组织结构的研究方法处理检错。

研究藏文或文本校对的主要对象是语言单位,在藏语言中最小的语言单位是字母,其次是音节,音节由字母组成。而字形是字的形状和结构,藏文字形以一个辅音字母为核心其余字母以此为基础前后附加和上下叠加组合成一个字的结构,因此人们都说藏文是由字母组合而成的一种拼音文字。……

登录APP查看全文