Recognition of Human G-protein Coupled Receptors Using Compressed Amino Acid Alphabets
2013-07-02GuanCuiping
Guan Cui-ping
College of Life Science, Ningxia University, Yinchuan 750021, China
Introduction
The superfamily of G-protein coupled receptors(GPCRs)is one of the largest and most diverse families of proteins in mammals (Bockaert and Pin, 1999).It is estimated that the human genome might encode as many as 2 000 different GPCRs. They are in the group of proteins that draw the most attention in the pharmaeutical industry, by the fact that about one quarter of all current prescription drugs act as ligands that bind to this huge superfamily of receptors.
A number of bioinformatics strategies have been used to search for novel GPCRs. Early strategies involved similarity searches using known GPCRs and 7tm patterns. However, generating reliable alignments of sequences lacked significant similarity are practically not possible, and the resulting models are easily affected by sampling bias (Lapinsh et al., 2002; Lu et al., 2009). In order to overcome these limitations,alignment-free protein classification methods are developed. Instead of alignments, various descriptors are extracted from each sequence (e.g., amino acid composition (Chou, 2005; Elrod and Chou, 2002;Chou and Elrod, 2002; Wang et al., 2005), dipeptide frequencies (Wang et al., 2005; Bhasin and Raghava,2004, 2005; Huang et al., 2004), physico-chemical properties (Lapinsh et al., 2002), pseudo amino acid composition (Rehman and Khan, 2012 ), genetic feature ensemble (Naveed and Khan, 2012; Peng et al., 2010), and pattern recognition or multivariate statistical methods, such as k-nearest neighbor, support vector machine, Naives Bayes model (Papasaikas et al., 2003; Guo et al., 2006; Gupta et al., 2008;Karchin et al., 2002; Worth et al., 2011), are trained to discriminate GPCRs from non-GPCRs. The same features are also used to classify GPCRs at the family,sub-family and sub-subfamily levels.
In this study, the support vector machine (SVM)method combined with the feature extraction of sequence evolutionary distance was applied to recognize human GPCRs. Firstly, N and C termini of GPCRs sequences, which did not contain clear evolutionary conservation from each other, were removed from the receptor sequences. The remaining part was called NCcutSeq. Secondly, NCcutSeqs were converted to CAseqs (compressed sequences)using the compressed alphabets. Then, the feature frequency in the form of amino acid composition was analyzed on CAseqs. Finally, classifiers were constructed to discriminate GPCRs from non-GPCRs using above statistical features and SVM. The five-fold crossvalidation showed that the overall accuracy of the proposed method could achieve up to 97.86%.Comparison with Bhasin's method also showed that this method could obtain better prediction accuracy.
Materials and Methods
Dataset
The dataset used to train this system contained 729 human GPCRs, which were collected from GPCRDB,and 1 834 human non-GPCRs selected from UniProt(uniref 90.fasta)randomly. This dataset didn't contain the sequences which labeled with "probable","putative" or "fragment" in GPCRDB and UniProt.Then all the samples were pretreated as the following steps:
Firstly, amino acids located in N and C termini of GPCRs sequences were identif i ed by TMHMM (Krogh et al., 2001)searches. The length distributions of N and C termini were analyzed on 729 human GPCRs,the length of N terminus was from 6 to 1 013 amino acids, most of them were distributed around 20 amino acids. The length of C terminus was from 4 to 578 amino acids and most were distributed around 12 amino acids.
Secondly, the front 20 letters and the last 12 letters of the sequence were removed. By reason that the termini parts were various in GPCRs, in contrast, the7tm middle part was more conservative, which could better display the characteristics of GPCRs. So we removed N and C termini parts of each sequence based on the analysis of length distribution in step one. The rest part of the sequence was called NCcutSeq, which was shorter than the original one.

Table 1 Compressed alphabets produced by different methods
We did these treatments to both the positive and negative samples in order to keep a unif i ed standard.Finally, we got two datasets as GPCR- NccutSeq (729)and nonGPCR- NccutSeq (1834)to build a GPCR prediction system.
GPCR prediction system
Coding of NCcutSeqs using compressed amino acid
Compressed alphabets are always used for sequence alignment, which can maximize the local similarities between sequences. It is described as that a compressed alphabet 'C' of size N is a partition of the standard 20-letter amino acid alphabet 'A' into N disjoint subsets (classes)containing similar amino acids. Several methods for constructing such alphabets have been proposed (Edgar, 2004), as listed in Table 1.Using the compressed amino acid to code NCcutSeqs could extend the identity in conserved region, especially to GPCR- NCcutSeqs, which could amplify the similarities in the 7tm region of GPCRs.
In this study, we compared the performance of 11 different compressed methods. Referred to Table 1,sequences in GPCR-NCcutSeq and nonGPCRNCcutSeq were compressed successively, e.g. if used SE-B (14)method to compress a NCcutSeq, the resulting compressed sequence would constitute of 14 amino acid classes, and so on. Finally, a NCcutSeq was transformed into 11 compressed sequences, which were called CAseqs. Then, the feature frequency vectors encoded by CAseq were calculated, respectively.
Frequency feature of amino acid composition
CAseqs were classified into 11 classes (SE-B (14),SE-B (10), SE-V (10), Li-A (10), Li-B (10), Solis-D(10), Solis-G (10), Murphy (10), SE-B (8), SE-B (6),and Dayhoff (6))according to the up step. Frequency features of amino acid compositions, including the single amino acid and dipeptide compositions,were extracted based on the compressed amino acid compositions in each class. Feature vector of each CAseq was calculated using Equations 1 and 2:

Where, Fiwas the occurrence frequency of amino acid i in CAseq; Aiwas the total number of amino acid i in CAseq; n was the total number of all amino acids in CAseq; Fijwas the occurrence frequency of dipeptide ij in CAseq; depijwas the total number of dipeptide ij in CAseq; m was the total number of all possible dipeptides in CAseq; and N was the class of compressed alphabets (Table 1).
Finally, a CAseq was converted into a vector with N+N2dimensions as the input of SVM classifier,where N had different values according to different compressed methods.
SVM
In this study, SVM prediction models were developed using the freely available software LIBSVM(Chang and Lin, 2011)and 11 SVM-based classif i ers were developed. The linear kernel function was chosen as the kernel, and the related parameter C was set from 10 to 100, when C=20, the classifiers could get the best prediction results.
Results
For the recognition of GPCRs, the performances of different compressed methods were compared. Total 11 SVM-based classifiers were developed using statistical features of single amino acid and dipeptide compositions on CAseqs. To evaluate their prediction performances, five-fold cross-validation process was used, the training and test sets were segregated at random and the process was repeated 10 times. At the same time, three measures including sensitivity (Sn),specif i city (Sp)and accuracy (Acc)were also utilized to evaluate the prediction performance. Let TP (true positive)and TN (true negative)be the numbers of correctly predicted positive and negative samples, FP(false positive)and FN (false negative)be the number of incorrectly predicted positive and negative samples,respectively. Then, Sn, Sp and Acc were def i ned as:

The results of discriminating 729 human GPCRs from 1 834 non-GPCRs with 11 compressed methods were compared in Table 2. For the purpose of further comparison, the performance of NCcutSeq and fulllength sequences, which were both compressed with Li-B (10)method, was compared in Table 3, and the Bhasin's method that exploited dipeptide composition features of full-length sequence was also compared in Table 3, when tested with the same data and evaluated with the same measures.
Table 2 showed that not all the compressed methods were suitable to classify GPCRs, the compressed sequences produced by the method of Li-B (10)could get better results than others. Table 3 showed that the sequences which N and C termini were removed by our method also could get better results than the sequences with full-length, and our system encoded by compressed alphabets was better than that of Bhasin's method. It proved that suitable compressed method combined with removing N and C termini with low conservation of GPCRs sequences did intensify the features of 7tm regions in GPCRs family.

Table 2 Predicting results based on NCcutSeq

Table 3 Performance comparison of NCcutSeq, full-length sequences and Bhasin's method
Discussion
The aim of this study was to improve the accuracy of GPCRs recognition, for that, we did some treatments to the samples by such measures as removed N and C termini which had low conservation of GPCRs sequences, coded sequences with compressed amino acids and integrated the compositional features of single amino acid and dipeptide. By comparing the results listed in Table 3, it was distinct that the prediction accuracy based on the pretreated sequences was better than the full-length ones, meaning that the features of GPCRs were intensified by removing N and C termini and compressing the middle part with compressed amino acids.
It was worth noticing that the prediction accuracy of our study was mainly affected by the choices of compressed method and the amino acid composition.The previous researches have proved that dipeptide composition is a better feature for recognizing GPCRs from non-GPCRs, because related sequences tend to have more information about the fraction of amino acids as well as their local order in common than expected by chance. So it can well detect similarities among proteins in a family. Whereas sequences diverge, the number of single amino acid and dipeptide in 20 amino acids would on average be reduced, ultimately reached a limit comparable to the sequences with low similarities in GPCRs families. If a compressed alphabet was used, it might expect this limit to be reached at a greater evolutionary distance.The results listed in Table 2 justif i ed that the suitable choice of compressed method could give better recognition of GPCRs. In this study, Li-B (10)proposed by Li et al. (2003)was better than others, they used a heuristic search procedure inspired by the Monte Carlo algorithm to compress the sequences, and this kind of residues grouping could well minimize the information loss in GPCRs families and maximize the GPCRs features which were different from non-GPCRs.
Conclusions
In this study, we pretreated the dataset by removing N and C termini parts of the sequences, and the remaining parts were compressed using different methods. Then frequency features in the form of single amino acid and dipeptide compositions were calculated based on the compressed sequences. The testing results demonstrated that the computational model encoded by Li-B (10)compressed method could obtain a better prediction accuracy of 97.86%. This is an efficient method to recognize human GPCRs.Future work of this study will focus on exploring and integrating more features of GPCRs to further improve the accuracy of GPCR prediction and apply it to the human genomic sequence data to fi nd novel GPCRs as well as their family classif i cation.
Bhasin M, Raghava G P S. 2004. GPCRpred: an SVM-based method for prediction of families and subfamilies of G-protein coupled receptors.Nucleic Acids Res, 32(2): 383-389.
Bhasin M, Raghava G P S. 2005. GPCRsclass: a web tool for the classif i cation of amine type of G-protein-coupled receptors. Nucleic Acids Res, 33(2): 143-147.
Bockaert J, Pin J P. 1999. Molecular tinkering of G protein-coupled receptors: an evolutionary success. EMBO J, 18(7): 1723-1729.
Chou K C. 2005. Prediction of G-protein-coupled receptor classes. J Proteome Res, 4(4): 1413-1418.
Chang C C, Lin C J. 2011. LIBSVM: a library for support vector machines. ACM TIST, 2(3): 1-27.
Chou K C, Elrod D W. 2002. Bioinformatical analysis of G-proteincoupled receptors. J Proteome Res, 1(5): 429-433.
Edgar R C. 2004. Local homology recognition and distance measures in linear time using compressed amino acid alphabets. Nucleic Acids Research, 32(1): 380-385.
Elrod D W, Chou K C. 2002. A study on the correlation of G-proteincoupled receptor types with amino acid composition. Protein Eng Des Sel, 15(9): 713-715.
Gupta R, Mittal A, Singh K. 2008. A novel and efficient technique for identification and classification of GPCRs. IEEE Trans Inform Technol Biomed, 12(4): 541-548.
Guo Y Z, Li M, Lu M, et al. 2006. Classifying G protein coupled receptors and nuclear receptors on the basis of protein power spectrum from fast fourier transform. Amino Acids, 30(4): 397-402.
Huang Y, Cai J, Ji L, et al. 2004. Classifying G-protein coupled receptors with bagging classifition tree. Comput Biol Chem, 28(4):275-280.
Karchin R, Kevin K, David H. 2002. Classifying G-protein coupled receptors with support vector machines. Bioinformatics, 18(1):147-159.
Krogh A, Larsson B, Von H G, et al. 2001. Predicting transmembrane protein topology with a hidden markov model: application to complete genomes. J MOL BIOL, 305(3): 567-580.
Lapinsh M, Gutcaits A, Prusis P, et al. 2002. Classif i cation of G-protein coupled receptors by alignment-independent extraction of principal chemical properties of primary amino acid sequences. Protein Sci,11(4): 795-805.
Li T, Fan K, Wang J, et al. 2003. Reduction of protein sequence complexity by residue grouping. Protein Eng, 16(5): 323-330.
Lu G Q, Wang Z F, Jones A M, et al. 2009. 7TMRmine: a web server for hierarchical mining of 7TMR proteins. BMC Genomics, 10: 275.
Naveed M, Khan A U. 2012. GPCR-MPredictor: multi-level prediction of G protein-coupled receptors using genetic ensemble. Amino Acids,42(5): 1809-1823.
Peng Z L, Yang J Y, Chen X. 2010. An improved classification of G-protein-coupled receptors using sequence-derived features. BMC Bioinformatics, 11: 420.
Papasaikas P K, Bagos P G, Litou Z I, et al. 2003. A novel method for GPCR recognition and family classification from sequence alone using signatures derived from profile hidden Markov models. SAR QSAR Environ Res, 14(5-6): 413-420.
Rehman Z U, Khan A. 2012. Identifying GPCRs and their types with Chou's pseudo amino acid composition: an approach from multi-scale energy representation and position specific scoring matrix. Protein Pept Lett, 19(8): 890-903.
Wang Y F, Chen H, Zhou Y H. 2005. Prediction and classification of human G-protein coupled receptors based on support vector machines. Geno Prot Bioinfo, 3(4): 242-246.
Worth C L, Kreuchwig A, Kleinau G, et al. 2011. GPCR-SSFE: A comprehensive database of G-protein-coupled receptor template predictions and homology models. BMC Bioinformatics, 12: 185.
杂志排行
Journal of Northeast Agricultural University(English Edition)的其它文章
- Strategies for Promoting Rice Self-suff i ciency in Sierra Leone
- Constraints on Problems of Rural Surplus Manpower Capital Transfer in China
- Seasonal Variation of Moisture Availability at Water-wind Erosion Crisscross Region in Northern Loess Plateau China
- Isolation of Nile Tilapia (Oreochromis niloticus) β-actin Promoter and Assay of Its Transcription Activity
