基于样本熵的DNA序列相似性分析
2016-03-02万力超周小安
万力超 周小安
摘 要:针对传统方法在分析DNA序列相似性方面的不足,提出了一种基于样本熵的DNA序列相似性分析方法。以五种东亚钳蝎神经毒素的基因序列作为分析对象,首先通过DNA序列的图形表示把DNA序列转换为时间序列,然后运用样本熵算法计算出时间序列的样本熵值,将样本熵的互值大小作为分析序列之间相似性的依据,最后将样本熵方法与DTW(Dynamic Time Warping,动态时间弯曲)方法的实验结果进行比较。实验结果表明,样本熵分析方法能有效分析序列之间的相似性,与DTW分析方法相比较,显示出更强的相似性和区别度,可将其进一步应用于生物序列的分析。
关 键 词:样本熵;DNA序列;序列相似性;DTW距离
中图分类号: TP391文献标识码: A文章编号:2095-2163(2016)01-
Abstract:This paper studies the application of sample entropy for similarity analysis of DNA sequences. The gene sequences of five kinds of Buthus martensi Karsch neurotoxins are analyzed. The graphical representation of DNA sequences are converted into digital sequences, and their sample entropy are calculated based on sample entropy method. The mutual value between different sample entropy is used to analysis sequence similarity. Analysis result is compared with the method of DTW distance. The analysis result of the proposed method provides good analysis efficiency and higher sensitivity and distinction than the results of DTW distance method. The method of sample entropy can be used for further biological sequences analysis.
Key words: DNA sequence; similarity analysis; sample entropy; DTW distance
0 引 言
随着生物序列测序技术的不断进步,人们已经获得了海量的生物序列信息,对于如何提取挖掘生物序列中的有用内容,解读DNA序列中的遗传信息和功能信息,DNA序列的相似性分析即已成为研究关注热点和实施应用亮点。DNA序列的相似性是指两条DNA序列的相似程度,相似程度越高表明两物种“同源”的可能性越大,反之,两物种的结构和功能差别越大。每当得到一个新物种的DNA序列,人们总是想通过比较该物种与其他已知序列的相似性,由此来分析其基因的功能,如果两个基因序列相似程度越高,新物种的结构和功能就与已知物种越相似,对于预测新物种基因信息就越有利,如此将会大大降低基因检测与测序的工程量,这在庞大的基因序列面前即显得尤为重要。……
