WO2018120128A1 - 一种膜蛋白残基的作用关系的预测方法和装置 - Google Patents

一种膜蛋白残基的作用关系的预测方法和装置 Download PDF

Info

Publication number
WO2018120128A1
WO2018120128A1 PCT/CN2016/113754 CN2016113754W WO2018120128A1 WO 2018120128 A1 WO2018120128 A1 WO 2018120128A1 CN 2016113754 W CN2016113754 W CN 2016113754W WO 2018120128 A1 WO2018120128 A1 WO 2018120128A1
Authority
WO
WIPO (PCT)
Prior art keywords
residue
membrane protein
feature
features
pair
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2016/113754
Other languages
English (en)
French (fr)
Inventor
张慧玲
魏彦杰
郭宁
贝振东
朱昱寰
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Shenzhen Institute of Advanced Technology of CAS
Original Assignee
Shenzhen Institute of Advanced Technology of CAS
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Shenzhen Institute of Advanced Technology of CAS filed Critical Shenzhen Institute of Advanced Technology of CAS
Priority to PCT/CN2016/113754 priority Critical patent/WO2018120128A1/zh
Publication of WO2018120128A1 publication Critical patent/WO2018120128A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations

Definitions

  • the present invention belongs to the field of data mining, machine learning and computer biology, and in particular relates to a method and apparatus for predicting the relationship between membrane protein residues.
  • membrane proteins account for about 60%. Due to the difficulty of experimental analysis of membrane protein structure, in the protein database (Protein Data Bank-PDB), more than 90,000 known protein structures, known membrane protein structures account for only 1% of the known protein structure. .
  • the amino group and the carboxyl group between the amino acids are dehydrated into a bond, and the amino acid participates in the formation of the peptide bond due to a part of the amino acid.
  • the structural part is called an amino acid residue.
  • the term "residual relationship" refers to those pairs of residues that are not adjacent in the primary sequence of the protein but are adjacent in the tertiary structure.
  • the object of the present invention is to provide a method for predicting the relationship between membrane protein residues, so as to solve the problem that the prediction method in the prior art leads to a large amount of useful information loss, affecting the accuracy and coverage of the prediction.
  • an embodiment of the present invention provides a method for predicting a relationship of a membrane protein residue, the method comprising:
  • the extracted unbalanced classification features are trained by the smote-boost algorithm to obtain a predicted model after training;
  • the extracting the residual and non-interacting residues in the membrane protein of the resolved protein structure include: a position-specific score matrix PSSM feature, a relative distance feature in the spiral, a sequence interval feature, a residue type feature, and a number of spirals One or more of a feature, a sequence length feature.
  • each of the residues in the location-specific scoring matrix PSSM is represented by a 20-dimensional vector
  • the location specific score matrix PSSM features include:
  • a residue pair includes two amino acids, the residue type characteristic comprising an acidic amino acid, a base Ten combinations of any two of amino acids, polar amino acids, and non-polar amino acids.
  • the interacting residue pair is a residue pair having a CB-CB atom distance of less than 8 angstroms on the alpha helix of the membrane protein .
  • an embodiment of the present invention provides a device for predicting a relationship between membrane protein residues, and the device includes: [0018] a training set acquisition unit, configured to acquire a membrane protein of the resolved protein structure as a training set;
  • a feature extraction unit configured to extract a feature of the unbalanced classification of the pair of residues and the pair of non-interactive residues in the membrane protein of the resolved protein structure
  • a training unit configured to train the predicted model by using the smote-boost algorithm to obtain the trained predictive model
  • a prediction unit configured to predict a relationship of membrane protein residues of an unknown protein structure according to the predicted model after training.
  • the feature of the unbalanced classification includes: a location-specific score matrix PSSM feature, and a residue in the spiral One or more of a relative distance feature, a sequence interval feature, a residue type feature, a spiral number feature, and a sequence length feature.
  • each of the residues in the location-specific scoring matrix PSSM is represented by a 20-dimensional vector.
  • the location specific score matrix PSSM features include:
  • a residue pair includes two amino acids, the residue type characteristic comprising an acidic amino acid, a base Ten combinations of any two of amino acids, polar amino acids, and non-polar amino acids.
  • the interacting residue pair is a residue pair having a CB-CB atom distance of less than 8 angstroms on the alpha helix of the membrane protein .
  • a membrane protein of a resolved protein structure is obtained as a training set, and a residue pair and a non-interacting residue pair for distinguishing interactions in the membrane protein of the resolved protein structure are extracted.
  • the characteristics of the unbalanced classification, the extracted features are trained by the smooth-boost algorithm to obtain the training model.
  • the post-predictive model and based on the post-training prediction model, predicts the relationship of membrane protein residues of unknown protein structures. Because the non-equilibrium classification features are used to train the prediction model, the post-training prediction model can avoid the loss of useful information, which is beneficial to improve the accuracy and coverage of the prediction.
  • FIG. 1 is a flow chart for realizing a prediction method for the action relationship of membrane protein residues provided by an embodiment of the present invention
  • FIG. 2 is a prediction device for the action relationship of membrane protein residues provided by an embodiment of the present invention
  • the purpose of the embodiments of the present invention is to provide a method for predicting the relationship between membrane protein residues, in order to solve the prior art prediction process for the relationship between membrane protein residues of unknown structures, generally from equilibrium classification.
  • the angle of the interacting residue pair or non-interacting residue pair is trained in a 1:1 ratio, and in fact, the ratio of interacting or non-interacting residues is much greater than 1: 1, according to equilibrium.
  • a peer-to-peer proportional training model can result in the loss of a large amount of useful information, which leads to the problem of the accuracy and coverage of the predicted membrane protein residues.
  • FIG. 1 shows an implementation flow of a prediction method for the action relationship of membrane protein residues provided by the first embodiment of the present invention, which is described in detail as follows:
  • step S101 a membrane protein of the resolved protein structure is obtained as a training set.
  • the membrane protein of the resolved protein structure should have a determined relationship of membrane protein residues.
  • a membrane protein which was analyzed before February 2012 in PDBTM (English full name: protein data bank of transmembrane proteins) can be used as a training set.
  • step S102 extracting features of the membrane protein of the resolved protein structure for distinguishing the unbalanced classification of the interacting residue pair and the non-interacting residue pair;
  • the feature for distinguishing the unbalanced classification of the interacting residue pair and the non-interacting residue pair in the embodiment of the present invention may include a position-specific scoring matrix PSSM feature, and the residue is One or more of the relative distance characteristics, sequence spacing characteristics, residue type characteristics, spiral number characteristics, and sequence length characteristics in the spiral.
  • PSSM Position-Specific Scoring Matrix
  • PSI-BLAST Alternative Basic Local Alignment Search Tool
  • Chinese full name Position-specific iterative search algorithm
  • the database that can be used to run PSI-BLAST ⁇ is the UNIREF90 database.
  • the number of iterations for running ⁇ can be 2, and the E-value truncation value is le-10 (expressed as 1*10 -10 powers).
  • each residue in the position-specific scoring matrix PSSM is represented by a 20-dimensional vector representing the frequency at which 20 amino acids occur at corresponding positions in the PSSM.
  • Feature Extraction ⁇ , Location-Specific Scoring Matrix PSSM features are divided into two categories, namely:
  • the method may be:
  • Position-specific scoring matrix PSSM features
  • the relative distance characteristic of the residue in the helix is specifically: assuming that p is a residue in the pair of residues is long The relative position on the helix of degree 1, then the relative distance characteristic of the residue in the a helix is defined as p/l, and for each residue pair including two residues, the residue corresponding to the residue can be extracted separately In the "relative distance feature in the spiral, a total of 2 residues in the spiral relative distance characteristics.
  • the sequence spacing feature may be partitioned according to the position of the residue pair in the primary sequence. For example, a specific interval division method can be divided into the following multiple intervals:
  • the corresponding sequence interval feature code 000000000 can be set to 0 or set to 1 (0 means not in the interval, and vice versa is 1) for expressing the sequence interval feature.
  • one of the nine sequence interval features may be corresponding to the interval division method described above.
  • amino acids constituting the protein considering 20 kinds of amino acids constituting the protein, according to the polar nature of the amino acid R group, it can be divided into acidic amino acids (glutamic acid and aspartic acid) and basic amino acids (Lai And arginine, and neutral amino acids, which can be divided into polar amino acids (glycine, serine, cysteine, threonine, tyrosine, asparagine and valley) (aminoamide) and non-polar amino acids (alanine, leucine, isoleucine, phenylalanine, methionine, tryptophan, valine and proline).
  • acidic amino acids glutmic acid and aspartic acid
  • basic amino acids Lai And arginine
  • neutral amino acids which can be divided into polar amino acids (glycine, serine, cysteine, threonine, tyrosine, asparagine and valley) (aminoamide) and non-polar amino acids (alanine, leucine, isoleu
  • one residue pair (corresponding to two amino acids) can produce 10 different combinations, which can be binary code 0000000000 respectively. Set to 0 or set to 1 to represent different combinations. It can include 10 residue type features.
  • the "spiral number characteristic” may be divided into sections according to the number of alpha helices included in the membrane protein. For example, it can be divided into 4 intervals of 2-4, 5-7, 8-10, and greater than 10.
  • the "number of spirals” is represented by the binary vector 0000 being set to 0 or set to 1 (0 means not in the interval, and vice versa). This class of features is consistent for all pairs of residues in a membrane protein. Each residue pair of feature vectors contains four such features.
  • the sequence length characteristic can be divided into four intervals of ⁇ 100, 100-400, 400-800, >800 according to the length of the primary sequence of the membrane protein, and the binary vector 0000 is set to 0 or set to 1. Indicates this feature (0 means not in the interval, and vice versa). Such features are consistent for all residue pairs in the same membrane protein. Each residue pair of feature vectors contains four such features.
  • the present invention can use 340 position-specific scoring matrix PSSM features, 2 alpha helices Medium relative distance feature, 9 sequence interval features and 10 residue type features, 4 alpha spiral number features
  • the proportion of the interacting residue pair and the non-interacting residue pair in the embodiment of the present invention may be 1 to 50 to 1 to 80, and a preferred embodiment may be set to 1 ratio. 67.
  • CB-C which will be located on the alpha helix of the membrane protein.
  • ⁇ Atom distance is less than 8 ⁇ (Angstrom)
  • Residue pairs are defined as pairs of interacting residues.
  • CA, CB are atom types in gromacs, groma cs molecular dynamics software.
  • step S103 the extracted non-equalized classification features are trained by the smote-boost algorithm to obtain a predicted model after training;
  • the features may be substituted into a prediction model for training.
  • the prediction model may be a vector machine training model or the like.
  • the training algorithm smote-boost is a new training method combining smote technology and boost technology, wherein: the boost method increases the weight of the sample without correct classification in each iteration, and reduces the weight of the correctly classified sample. , pay more attention to the sample of the wrong classification. Because a few samples are more susceptible to misclassification, this approach improves the predictive performance of a few classes.
  • SMOTE commonly known as synthetic minority over-sampling rechnique
  • SMOTE technology commonly known as synthetic minority over-sampling rechnique
  • step S104 the action relationship of the membrane protein residues of the unknown protein structure is predicted based on the predicted model after training.
  • the present invention extracts the membrane protein of the resolved protein structure as a training set, and extracts the non-reactive residue pair and the non-interacting residue pair in the membrane protein of the resolved protein structure.
  • the characteristics of the equilibrium classification are obtained, and the extracted features are trained by the smooth-boost algorithm to obtain a predicted model after training, and the membrane protein residue of the unknown protein structure is predicted according to the trained prediction model.
  • the role of the base is trained by using the features of unbalanced classification, so that the post-training prediction model can avoid the loss of useful information, which is beneficial to improve the accuracy and coverage of prediction.
  • FIG. 2 is a schematic view showing the structure of a prediction apparatus for a membrane protein residue relationship according to an embodiment of the present invention, which is described in detail as follows:
  • the apparatus for predicting the action relationship of the membrane protein residues in the embodiments of the present invention comprises:
  • a training set obtaining unit 201 configured to acquire a membrane protein of the resolved protein structure as a training set
  • a feature extraction unit 202 configured to extract a feature of the unbalanced classification of the pair of residual and non-interactive residues in the membrane protein of the resolved protein structure
  • the training unit 203 is configured to train the predicted feature of the extracted unbalanced classification by using a smote-boost algorithm to obtain a predicted model after training;
  • the prediction unit 204 is configured to predict a relationship of membrane protein residues of an unknown protein structure according to the predicted model after training.
  • the features of the unbalanced classification include: a position-specific score matrix, a PSSM feature, a relative distance feature in a spiral, a sequence interval feature, and a residue type feature.
  • One or more of the spiral number feature and the sequence length feature are selected from the spiral number feature and the sequence length feature.
  • each residue in the position-specific scoring matrix PSSM is represented by a 20-dimensional vector
  • the location-specific scoring matrix PSSM features include:
  • a residue pair includes two amino acids, and the residue type features include 10 combinations of any one of an acidic amino acid, a basic amino acid, a polar amino acid, and a non-polar amino acid.
  • the interacting residue pair is a residue pair having a CB-CB atom distance of less than 8 angstroms on the alpha helix of the membrane protein.
  • FIG. 2 is a prediction device for the action relationship of the membrane protein residue, and the membrane protein residue of the first embodiment is used. Corresponding to the prediction method of the relationship, the details are not repeated here.
  • the disclosed apparatus and method can be implemented in other manners.
  • the device embodiments described above are merely illustrative.
  • the division of the unit is only a logical function division, and the actual implementation may have another division manner, for example, multiple units or components may be combined or Can be integrated into another system, or some features can be ignored, or not executed.
  • the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or unit, and may be electrical, mechanical or otherwise.
  • the unit described as a separate component may or may not be physically distributed, and the component displayed as a unit may or may not be a physical unit, that is, may be located in one place, or may be distributed to multiple On the network unit. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of the embodiment.
  • each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
  • the above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
  • the integrated unit if implemented in the form of a software functional unit and sold or used as a standalone product, may be stored in a computer readable storage medium.
  • the technical solution of the present invention may contribute to the prior art or all or part of the technical solution may be embodied in the form of a software product stored in a storage medium.
  • a number of instructions are included to cause a computer device (which may be a personal computer, server, or network device, etc.) to perform all or part of the methods described in various embodiments of the present invention.
  • the foregoing storage medium includes: a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and the like, which can store program codes. .

Landscapes

  • Bioinformatics & Cheminformatics (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Genetics & Genomics (AREA)
  • Biotechnology (AREA)
  • Biophysics (AREA)
  • Chemical & Material Sciences (AREA)
  • Molecular Biology (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Analytical Chemistry (AREA)
  • Evolutionary Biology (AREA)
  • General Health & Medical Sciences (AREA)
  • Medical Informatics (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Theoretical Computer Science (AREA)
  • Investigating Or Analysing Biological Materials (AREA)

Abstract

一种膜蛋白残基的作用关系的预测方法包括:获取已解析蛋白质结构的膜蛋白作为训练集;提取所述已解析蛋白质结构的膜蛋白中用于区分相互作用的残基对和非相互作用的残基对的非均衡分类的特征;将所提取的非均衡分类的特征通过smote-boost算法训练预测模型;根据训练后的预测模型,预测未知蛋白质结构的膜蛋白残基的作用关系。由于使用非均衡分类的特征进行预测模型的训练,从而使得训练后的预测模型能够避免有用信息的流失,有利于提高预测的精准度和覆盖度。

Description

说明书 发明名称:一种膜蛋白残基的作用关系的预测方法和装置 技术领域
[0001] 本发明属于数据挖掘、 机器学习和计算机生物学的交叉领域, 尤其涉及一种膜 蛋白残基作用关系的预测方法和装置。
背景技术
[0002] 在目前已知的药物靶点中, 膜蛋白约占 60%。 由于膜蛋白结构的实验解析难度 较大, 在蛋白质数据库 (Protein Data Bank-PDB) 中, 超过 9万个的已知蛋白质 结构里, 已知的膜蛋白结构仅占已知的蛋白质结构的 1%。
[0003] 现有的解析蛋白质三维结构的生物学实验方法主要包括 X-RAY和 NMR法。 这 些生物学实验方法不仅操作过程较为复杂, 耗吋, 而且实验花费的成本也较高 。 正是由于实验解析法的这些不足, 使得计算机计算方法的发展成为必然。 目 前用于蛋白质三维结构预测的计算方法主要有同源模建法、 折叠识别法和从头 预测法。 并且通常从均衡分类的角度, 将相互作用的残基对或非相互作用的残 基对按照 1 : 1的比例训练模型。 其中, 残基是指由 20种不同的氨基酸连接形成 的多聚体, 在形成蛋白质后, 这些氨基酸之间的氨基和羧基脱水成键, 氨基酸 由于其部分基团参与了肽键的形成, 剩余的结构部分称为氨基酸残基。 所谓残 基作用关系是指那些在蛋白质的一级序列中不相邻而在三级结构中邻近的残基 对。
[0004] 由于相互作用的残基对与非相互作用的残基对的比例一般会远远大于 1 : 1, 从 而使得现有的预测方法会导致大量有用的信息流失, 影响预测的准确度和覆盖 度。
技术问题
[0005] 本发明的目的在于提供一种膜蛋白残基的作用关系的预测方法, 以解决现有技 术中的预测方法会导致大量有用的信息流失, 影响预测的准确度和覆盖度的问 技术解决方案
[0006] 第一方面, 本发明实施例提供了一种膜蛋白残基的作用关系的预测方法, 所述 方法包括:
[0007] 获取已解析蛋白质结构的膜蛋白作为训练集;
[0008] 提取所述已解析蛋白质结构的膜蛋白中用于区分相互作用的残基对和非相互作 用的残基对的非均衡分类的特征;
[0009] 将所提取的非均衡分类的特征通过 smote-boost算法训练预测模型, 得到训练后 的预测模型;
[0010] 根据训练后的预测模型, 预测未知蛋白质结构的膜蛋白残基的作用关系。
[0011] 结合第一方面, 在第一方面的第一种可能实现方式中, 所述提取所述已解析蛋 白质结构的膜蛋白中用于区分相互作用的残基对和非相互作用的残基对的非均 衡分类的特征步骤中, 所述非均衡分类的特征包括: 位置特异性得分矩阵 PSSM 特征、 残基在《螺旋中相对距离特征、 序列间隔特征、 残基类型特征、 《螺旋个 数特征、 序列长度特征中的一种或者多种。
[0012] 结合第一方面的第一种可能实现方式, 在第一方面的第二种可能实现方式中, 所述位置特异性得分矩阵 PSSM中的每个残基由一个 20维的向量表示, 所述位置 特异性得分矩阵 PSSM特征包括:
[0013] 以残基对 (i,j) 中的残基 i和残基 j分别为中心取一个大小为 a的滑动容器, 每个 残基对得到 40a个位置特异性得分矩阵 PSSM特征;
[0014] 以残基对 (i,j)的中间位置 (i+j) /2为中心取一个大小为 b的滑动窗口, 获得 20*b个 位置特异性得分矩阵 PSSM特征。
[0015] 结合第一方面的第一种可能实现方式, 在第一方面的第三种可能实现方式中, 一个残基作用对包括两个氨基酸, 所述残基类型特征包括由酸性氨基酸、 碱性 氨基酸、 极性氨基酸、 非极性氨基酸中的任意两种所产生的 10种组合。
[0016] 结合第一方面, 在第一方面的第四种可能实现方式中, 所述相互作用的残基对 为位于膜蛋白的 α螺旋上的 CB-CB原子距离小于 8埃的残基对。
[0017] 第二方面, 本发明实施例提供了一种膜蛋白残基的作用关系的预测装置, 所述 装置包括: [0018] 训练集获取单元, 用于获取已解析蛋白质结构的膜蛋白作为训练集;
[0019] 特征提取单元, 用于提取所述已解析蛋白质结构的膜蛋白中用于区分相互作用 的残基对和非相互作用的残基对的非均衡分类的特征;
[0020] 训练单元, 用于将所提取的非均衡分类的特征通过 smote-boost算法训练预测模 型, 得到训练后的预测模型;
[0021] 预测单元, 用于根据训练后的预测模型, 预测未知蛋白质结构的膜蛋白残基的 作用关系。
[0022] 结合第二方面, 在第二方面的第一种可能实现方式中, 所述特征提取单元中, 所述非均衡分类的特征包括: 位置特异性得分矩阵 PSSM特征、 残基在《螺旋中 相对距离特征、 序列间隔特征、 残基类型特征、 《螺旋个数特征、 序列长度特征 中的一种或者多种。
[0023] 结合第二方面的第一种可能实现方式, 在第二方面的第二种可能实现方式中, 所述位置特异性得分矩阵 PSSM中的每个残基由一个 20维的向量表示, 所述位置 特异性得分矩阵 PSSM特征包括:
[0024] 以残基对 (i,j) 中的残基 i和残基 j分别为中心取一个大小为 a的滑动容器, 每个 残基对得到 40a个位置特异性得分矩阵 PSSM特征;
[0025] 以残基对 (i,j)的中间位置 (i+j) /2为中心取一个大小为 b的滑动窗口, 获得 20*b个 位置特异性得分矩阵 PSSM特征。
[0026] 结合第二方面的第一种可能实现方式, 在第二方面的第三种可能实现方式中, 一个残基作用对包括两个氨基酸, 所述残基类型特征包括由酸性氨基酸、 碱性 氨基酸、 极性氨基酸、 非极性氨基酸中的任意两种所产生的 10种组合。
[0027] 结合第二方面, 在第二方面的第四种可能实现方式中, 所述相互作用的残基对 为位于膜蛋白的 α螺旋上的 CB-CB原子距离小于 8埃的残基对。
发明的有益效果
有益效果
[0028] 在本发明中, 获取已解析的蛋白质结构的膜蛋白作为训练集, 提取所述已解析 的蛋白质结构的膜蛋白中用于区分相互作用的残基对和非相互作用的残基对的 非均衡分类的特征, 将提取的特征通过 smote-boost算法训练预测模型, 得到训练 后的预测模型, 并根据所述训练后的预测模型, 预测未知蛋白质结构的膜蛋白 残基的作用关系。 由于使用非均衡分类的特征进行预测模型的训练, 从而使得 训练后的预测模型能够避免有用信息的流失, 有利于提高预测的精准度和覆盖 度。
对附图的简要说明
附图说明
[0029] 图 1是本发明实施例提供的膜蛋白残基的作用关系的预测方法的实现流程图; [0030] 图 2是本发明实施例提供的膜蛋白残基的作用关系的预测装置的结构示意图。
本发明的实施方式
[0031] 为了使本发明的目的、 技术方案及优点更加清楚明白, 以下结合附图及实施例 , 对本发明进行进一步详细说明。 应当理解, 此处所描述的具体实施例仅仅用 以解释本发明, 并不用于限定本发明。
[0032] 本发明实施例的目的在于提供一种膜蛋白残基的作用关系的预测方法, 以解决 现有技术中对于未知结构的膜蛋白残基的作用关系的预测过程中, 一般从均衡 分类的角度将相互作用的残基对或非相互作用的残基对按照 1: 1的比例训练模 型, 而实际上, 相互作用或非相互作用的残基对比例远远大于 1 : 1, 按照均衡 对等的比例训练模型会造成大量有用信息的流失, 从而会导致预测的膜蛋白残 基的作用关系的精准度和覆盖度不高的问题。 下面结合附图对本发明作进一步 的说明。
[0033] 图 1示出了本发明第一实施例提供的膜蛋白残基的作用关系的预测方法的实现 流程, 详述如下:
[0034] 在步骤 S101中, 获取已解析蛋白质结构的膜蛋白作为训练集。
[0035] 具体的, 所述已解析蛋白质结构的膜蛋白, 应当已确定的膜蛋白残基的作用关 系。 优选的一种实施方式, 可以使用 PDBTM (英文全称为: protein data bank of transmembrane proteins, 中文全称为: 跨膜蛋白的蛋白质数据库)中 2012年 2月以 前解析的膜蛋白作为训练集。
[0036] 当然, 上述膜蛋白数据库的训练集的选取只是其中一种优选的实施方式, 随着 解析和识别技术的发展, 越来越多的膜蛋白结构被解析, 能够得到确定的膜蛋 白残基的作用关系, 因而所述训练集中的样本数据也会越来越丰富, 因而也会 更加有利于提高预测模型的训练的准确度。
[0037] 在步骤 S102中, 提取所述已解析蛋白质结构的膜蛋白中用于区分相互作用的残 基对和非相互作用的残基对的非均衡分类的特征;
[0038] 具体的, 本发明实施例中所述用于区分相互作用的残基对和非相互作用的残基 对的非均衡分类的特征, 可以包括位置特异性得分矩阵 PSSM特征、 残基在《螺 旋中相对距离特征、 序列间隔特征、 残基类型特征、 《螺旋个数特征、 序列长度 特征中的一种或者多种。
[0039] 其中, 所述位置特异性得分矩阵 PSSM (英文全称为: Position-Specific Scoring Matrix)特征, 可以通过运行 PSI-BLAST (英文全称为: Position-Specific Iterative Basic Local Alignment Search Tool, 中文全称为: 位置特异性迭代搜索算法)的方 式获取。 其中, 运行 PSI-BLAST吋可以采用的数据库是 UNIREF90数据库, 运行 吋的迭代次数可以为 2, E-value截断值为 le-10 (表示为 1*10的 -10次方) 。
[0040] 在本发明实施例中, 所述位置特异性得分矩阵 PSSM中的每个残基都由一个 20 维的向量表示, 表示 20种氨基酸在 PSSM相应位置出现的频率。 特征提取吋, 位 置特异性得分矩阵 PSSM特征分为两类, 分别为:
[0041] 以残基对 (i,j) 中的残基 i和残基 j分别为中心取一个大小为 a的滑动容器, 每个 残基对得到 40a个位置特异性得分矩阵 PSSM特征;
[0042] 以残基对 (i,j)的中间位置 (i+j) /2为中心取一个大小为 b的滑动窗口, 获得 20*b 位置特异性得分矩阵 PSSM特征。
[0043] 比如, 具体的一种实施方式中, 可以为:
[0044] 第一类是以残基对 (i, j) 中的残基 i和残基 j分别为中心取一个大小为 7的滑动 窗口, 即对每个残基对可得到 2x7x20=280个位置特异性得分矩阵 PSSM特征;
[0045] 第二类是以残基对 (i, j) 的中间位置 (i+j)/2为中心取一个大小为 3的滑动窗口 , 即可获得 3x20=60个位置特异性得分矩阵 PSSM特征。
[0046] 两类位置特异性得分矩阵 PSSM特征的总数为 280+60=340个。
[0047] 所述残基在《螺旋中相对距离特征具体为: 假设 p为残基对中的一个残基在长 度为 1的螺旋上的相对位置, 那么残基在 a螺旋中相对距离特征就定义为 p/l, 对于 每个残基对中包括两个残基, 可以分别提取残基所对应的残基在《螺旋中相对距 离特征, 一共包括 2个残基在 螺旋中相对距离特征。
[0048] 所述序列间隔特征可以根据残基对在一级序列中的位置进行划分。 比如, 一种 具体的间隔划分方式可以划分为以下多个区间:
[0049] <25、 25-50、 50-75、 75-100、 100-125、 125-150、 150-175、 175-200和 >200这 九个区间。
[0050] 可将使用相应的序列间隔特征码 000000000置 0或置 1 (0表示不在该区间, 反之 为 1) 用于表述序列间隔特征。 对于每个残基对而言, 按照上述区间划分方式, 可以对应 9个序列间隔特征中的一个。
[0051] 对于所述残基类型特征, 考虑到组成蛋白质的氨基酸共 20种, 根据氨基酸 R基 的极性性质可分为酸性氨基酸 (谷氨酸及天冬氨酸) 、 碱性氨基酸 (赖氨酸、 精氨酸及组氨酸) 和中性氨基酸, 其中中性氨基酸又可分为极性氨基酸 (甘氨酸 、 丝氨酸、 半胱氨酸、 苏氨酸、 酪氨酸、 天冬酰胺及谷氨酰胺)和非极性氨基酸 (丙氨酸、 亮氨酸、 异亮氨酸、 苯丙氨酸、 甲硫氨酸、 色氨酸、 缬氨酸及脯氨 酸) 。 根据这 4种不同的氨基酸类型 (酸性氨基酸、 碱性氨基酸、 极性氨基酸和 非极性氨基酸) , 一个残基作用对 (对应两个氨基酸) 可以产生 10种不同的组 合, 可以二进制码 0000000000分别置 0或置 1来代表不同的组合类型。 可以包括 1 0个残基类型特征。
[0052] 所述《螺旋个数特征可以根据膜蛋白所包含的 α螺旋个数进行区间划分。 比如 , 可以划分为 2-4、 5-7、 8-10、 以及大于 10这 4个区间。 通过二进制向量 0000置 0 或置 1来表示该《螺旋个数特征 (0表示不在该区间, 反之为 1) 。 该类特征对某 一膜蛋白中所有残基对具有一致性。 每个残基对特征向量包含 4个该类特征。
[0053] 所述序列长度特征, 可以根据膜蛋白所一级序列的长度可分为 <100, 100-400 , 400-800, >800这 4个区间, 以二进制向量 0000置 0或置 1来表示该特征 (0表示 不在该区间, 反之为 1) 。 这类特征对同一个膜蛋白中的所有残基对均一致。 每 个残基对特征向量包含 4个该类特征。
[0054] 综上所述, 本发明可以使用 340个位置特异性得分矩阵 PSSM特征, 2个 α螺旋 中相对距离特征, 9个序列间隔特征以及 10个残基类型特征, 4个 α螺旋个数特征
, 4个序列长度特征, 共计 369个特征。
[0055] 另外, 本发明实施例中所述相互作用的残基对和非相互作用的残基对的比例, 可以为 1比 50至 1比 80, 优选的一种实施方式可以设置为 1比 67。
[0056] 具体的, 蛋白质残基作用对的定义有多种, 例如基于原子的范德华距离的定义
, 基于 CA-CA原子距离的定义以及基于 CB-CB原子距离的定义。 本发明关于残 基作用对的定义将沿用一个被广泛采用的定义: 将位于膜蛋白的 α螺旋上的 CB-C
Β原子距离小于 8Α (埃)
的残基对定义为相互作用的残基对。 CA、 CB是 gromacs里面的原子类型, groma cs分子动力学软件。
[0057] 在步骤 S103中, 将所提取的非均衡分类的特征通过 smote-boost算法训练预测模 型, 得到训练后的预测模型;
[0058] 在提到到所述非均衡分类的特征后, 可以将所述特征代入到预测模型中进行训 练。 所述预测模型可以为向量机训练模型等。
[0059] 所述训练算法 smote-boost, 是将 smote技术和 boost技术结合的新型训练方法, 其中: boost方法在每次迭代中, 增加没有正确分类样本的权值, 减少正确分类 样本的权值, 更加关注于分类错误的样本。 因为少数样本更容易被错误分类, 所以该方法能够改进对少数类的预测性能。 SMOTE (英文全称为 synthetic minority over-sampling rechnique)技术是非均衡数据集学习的一种新办法, 通过 对少数样本的人工合成提高少数类样本的比例, 降低数据的过度偏斜。 SMOTE 技术与 BOOST技术相结合, 可以有效避免由于赋予少数样本更大权值可能产生 的过度拟合。
[0060] 在步骤 S104中, 根据训练后的预测模型, 预测未知蛋白质结构的膜蛋白残基的 作用关系。
[0061] 本发明通过获取已解析的蛋白质结构的膜蛋白作为训练集, 提取所述已解析的 蛋白质结构的膜蛋白中用于区分相互作用的残基对和非相互作用的残基对的非 均衡分类的特征, 将提取的特征通过 smote-boost算法训练预测模型, 得到训练后 的预测模型, 并根据所述训练后的预测模型, 预测未知蛋白质结构的膜蛋白残 基的作用关系。 由于使用非均衡分类的特征进行预测模型的训练, 从而使得训 练后的预测模型能够避免有用信息的流失, 有利于提高预测的精准度和覆盖度
[0062] 图 2示出了本发明实施例提供的一种膜蛋白残基的作用关系的预测装置的结构 示意图, 详述如下:
[0063] 本发明实施例所述膜蛋白残基的作用关系的预测装置, 包括:
[0064] 训练集获取单元 201, 用于获取已解析蛋白质结构的膜蛋白作为训练集;
[0065] 特征提取单元 202, 用于提取所述已解析蛋白质结构的膜蛋白中用于区分相互 作用的残基对和非相互作用的残基对的非均衡分类的特征;
[0066] 训练单元 203, 用于将所提取的非均衡分类的特征通过 smote-boost算法训练预 测模型, 得到训练后的预测模型;
[0067] 预测单元 204, 用于根据训练后的预测模型, 预测未知蛋白质结构的膜蛋白残 基的作用关系。
[0068] 优选的, 所述特征提取单元中, 所述非均衡分类的特征包括: 位置特异性得分 矩阵 PSSM特征、 残基在《螺旋中相对距离特征、 序列间隔特征、 残基类型特征
、 《螺旋个数特征、 序列长度特征中的一种或者多种。
[0069] 优选的, 所述位置特异性得分矩阵 PSSM中的每个残基由一个 20维的向量表示
, 所述位置特异性得分矩阵 PSSM特征包括:
[0070] 以残基对 (i,j) 中的残基 i和残基 j分别为中心取一个大小为 a的滑动容器, 每个 残基对得到 40a个位置特异性得分矩阵 PSSM特征;
[0071] 以残基对 (i,j)的中间位置 (i+j) /2为中心取一个大小为 b的滑动窗口, 获得 20*b个 位置特异性得分矩阵 PSSM特征。
[0072] 优选的, 一个残基作用对包括两个氨基酸, 所述残基类型特征包括由酸性氨基 酸、 碱性氨基酸、 极性氨基酸、 非极性氨基酸中的任意两种所产生的 10种组合
[0073] 优选的, 所述相互作用的残基对为位于膜蛋白的 α螺旋上的 CB-CB原子距离小 于 8埃的残基对。
[0074] 图 2所述膜蛋白残基的作用关系的预测装置, 与实施例一所述膜蛋白残基的作 用关系的预测方法对应, 在此不作重复赘述。
[0075] 在本发明所提供的几个实施例中, 应该理解到, 所揭露的装置和方法, 可以通 过其它的方式实现。 例如, 以上所描述的装置实施例仅仅是示意性的, 例如, 所述单元的划分, 仅仅为一种逻辑功能划分, 实际实现吋可以有另外的划分方 式, 例如多个单元或组件可以结合或者可以集成到另一个系统, 或一些特征可 以忽略, 或不执行。 另一点, 所显示或讨论的相互之间的耦合或直接耦合或通 信连接可以是通过一些接口, 装置或单元的间接耦合或通信连接, 可以是电性 , 机械或其它的形式。
[0076] 所述作为分离部件说明的单元可以是或者也可以不是物理上分幵的, 作为单元 显示的部件可以是或者也可以不是物理单元, 即可以位于一个地方, 或者也可 以分布到多个网络单元上。 可以根据实际的需要选择其中的部分或者全部单元 来实现本实施例方案的目的。
[0077] 另外, 在本发明各个实施例中的各功能单元可以集成在一个处理单元中, 也可 以是各个单元单独物理存在, 也可以两个或两个以上单元集成在一个单元中。 上述集成的单元既可以采用硬件的形式实现, 也可以采用软件功能单元的形式 实现。
[0078] 所述集成的单元如果以软件功能单元的形式实现并作为独立的产品销售或使用 吋, 可以存储在一个计算机可读取存储介质中。 基于这样的理解, 本发明的技 术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的全部或部分 可以以软件产品的形式体现出来, 该计算机软件产品存储在一个存储介质中, 包括若干指令用以使得一台计算机设备 (可以是个人计算机, 服务器, 或者网 络设备等) 执行本发明各个实施例所述方法的全部或部分。 而前述的存储介质 包括: U盘、 移动硬盘、 只读存储器 (ROM , Read-Only Memory) . 随机存取存储 器 (RAM, Random Access Memory) 、 磁碟或者光盘等各种可以存储程序代码 的介质。
[0079] 以上所述仅为本发明的较佳实施例而已, 并不用以限制本发明, 凡在本发明的 精神和原则之内所作的任何修改、 等同替换和改进等, 均应包含在本发明的保 护范围之内。

Claims

权利要求书
一种膜蛋白残基的作用关系的预测方法, 其特征在于, 所述方法包括 获取已解析蛋白质结构的膜蛋白作为训练集;
提取所述已解析蛋白质结构的膜蛋白中用于区分相互作用的残基对和 非相互作用的残基对的非均衡分类的特征;
将所提取的非均衡分类的特征通过 smote-boost算法训练预测模型, 得 到训练后的预测模型;
根据训练后的预测模型, 预测未知蛋白质结构的膜蛋白残基的作用关 系。
根据权利要求 1所述方法, 其特征在于, 所述提取所述已解析蛋白质 结构的膜蛋白中用于区分相互作用的残基对和非相互作用的残基对的 非均衡分类的特征步骤中, 所述非均衡分类的特征包括: 位置特异性 得分矩阵 PSSM特征、 残基在《螺旋中相对距离特征、 序列间隔特征 、 残基类型特征、 《螺旋个数特征、 序列长度特征中的一种或者多种 根据权利要求 2所述方法, 其特征在于, 所述位置特异性得分矩阵 PS SM中的每个残基由一个 20维的向量表示, 所述位置特异性得分矩阵 P SSM特征包括:
以残基对 (i,j) 中的残基 i和残基 j分别为中心取一个大小为 a的滑动容 器, 每个残基对得到 40a个位置特异性得分矩阵 PSSM特征; 以残基对 (i,j)的中间位置 (i+j) /2为中心取一个大小为 b的滑动窗口, 获得 20*b个位置特异性得分矩阵 PSSM特征。
根据权利要求 2所述方法, 其特征在于, 一个残基作用对包括两个氨 基酸, 所述残基类型特征包括由酸性氨基酸、 碱性氨基酸、 极性氨基 酸、 非极性氨基酸中的任意两种所产生的 10种组合。
根据权利要求 1所述方法, 其特征在于, 所述相互作用的残基对为位 于膜蛋白的 α螺旋上的 CB-CB原子距离小于 8埃的残基对。 一种膜蛋白残基的作用关系的预测装置, 其特征在于, 所述装置包括 训练集获取单元, 用于获取已解析蛋白质结构的膜蛋白作为训练集; 特征提取单元, 用于提取所述已解析蛋白质结构的膜蛋白中用于区分 相互作用的残基对和非相互作用的残基对的非均衡分类的特征; 训练单元, 用于将所提取的非均衡分类的特征通过 smote-boost算法训 练预测模型, 得到训练后的预测模型;
预测单元, 用于根据训练后的预测模型, 预测未知蛋白质结构的膜蛋 白残基的作用关系。
根据权利要求 6所述装置, 其特征在于, 所述特征提取单元中, 所述 非均衡分类的特征包括: 位置特异性得分矩阵 PSSM特征、 残基在《 螺旋中相对距离特征、 序列间隔特征、 残基类型特征、 《螺旋个数特 征、 序列长度特征中的一种或者多种。
根据权利要求 7所述装置, 其特征在于, 所述位置特异性得分矩阵 PS SM中的每个残基由一个 20维的向量表示, 所述位置特异性得分矩阵 P SSM特征包括:
以残基对 (i,j) 中的残基 i和残基 j分别为中心取一个大小为 a的滑动容 器, 每个残基对得到 40a个位置特异性得分矩阵 PSSM特征; 以残基对 (i,j)的中间位置 (i+j) /2为中心取一个大小为 b的滑动窗口, 获得 20*b个位置特异性得分矩阵 PSSM特征。
根据权利要求 7所述装置, 其特征在于, 一个残基作用对包括两个氨 基酸, 所述残基类型特征包括由酸性氨基酸、 碱性氨基酸、 极性氨基 酸、 非极性氨基酸中的任意两种所产生的 10种组合。
根据权利要求 6所述装置, 其特征在于, 所述相互作用的残基对为位 于膜蛋白的 α螺旋上的 CB-CB原子距离小于 8埃的残基对。
PCT/CN2016/113754 2016-12-30 2016-12-30 一种膜蛋白残基的作用关系的预测方法和装置 Ceased WO2018120128A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/CN2016/113754 WO2018120128A1 (zh) 2016-12-30 2016-12-30 一种膜蛋白残基的作用关系的预测方法和装置

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2016/113754 WO2018120128A1 (zh) 2016-12-30 2016-12-30 一种膜蛋白残基的作用关系的预测方法和装置

Publications (1)

Publication Number Publication Date
WO2018120128A1 true WO2018120128A1 (zh) 2018-07-05

Family

ID=62706805

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2016/113754 Ceased WO2018120128A1 (zh) 2016-12-30 2016-12-30 一种膜蛋白残基的作用关系的预测方法和装置

Country Status (1)

Country Link
WO (1) WO2018120128A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN117593783A (zh) * 2023-11-20 2024-02-23 广州视景医疗软件有限公司 基于自适应smote的视觉训练方案生成方法及装置

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104252581A (zh) * 2013-06-26 2014-12-31 中国科学院深圳先进技术研究院 一种基于支持向量机的跨膜蛋白残基作用关系预测方法
CN104504299A (zh) * 2014-12-29 2015-04-08 中国科学院深圳先进技术研究院 预测膜蛋白的残基间的作用关系的方法
CN104615910A (zh) * 2014-12-30 2015-05-13 中国科学院深圳先进技术研究院 基于随机森林预测α跨膜蛋白的螺旋相互作用关系的方法
CN106650309A (zh) * 2016-12-30 2017-05-10 中国科学院深圳先进技术研究院 一种膜蛋白残基的作用关系的预测方法和装置

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104252581A (zh) * 2013-06-26 2014-12-31 中国科学院深圳先进技术研究院 一种基于支持向量机的跨膜蛋白残基作用关系预测方法
CN104504299A (zh) * 2014-12-29 2015-04-08 中国科学院深圳先进技术研究院 预测膜蛋白的残基间的作用关系的方法
CN104615910A (zh) * 2014-12-30 2015-05-13 中国科学院深圳先进技术研究院 基于随机森林预测α跨膜蛋白的螺旋相互作用关系的方法
CN106650309A (zh) * 2016-12-30 2017-05-10 中国科学院深圳先进技术研究院 一种膜蛋白残基的作用关系的预测方法和装置

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
ARUN KUMAR, M.N. ET AL.: "On the Classification of Imbalanced Datasets", INTERNATIONAL JOURNAL OF COMPUTER APPLICATIONS, vol. 44, no. 8, 30 April 2012 (2012-04-30), pages 1 - 7, XP055509822 *
FATTAHI, S.: "NEW APPROACH FOR IMBALANCED BIOLOGICAL DATASET CLASSIFICATION", JOURNAL OF THEORETICAL AND APPLIED INFORMATION TECHNOLOGY, vol. 72, no. 1, 10 February 2015 (2015-02-10), pages 40 - 57, XP055509829 *

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN117593783A (zh) * 2023-11-20 2024-02-23 广州视景医疗软件有限公司 基于自适应smote的视觉训练方案生成方法及装置
CN117593783B (zh) * 2023-11-20 2024-04-05 广州视景医疗软件有限公司 基于自适应smote的视觉训练方案生成方法及装置

Similar Documents

Publication Publication Date Title
Tsukiyama et al. LSTM-PHV: prediction of human-virus protein–protein interactions by LSTM with word2vec
Ju et al. Prediction of lysine crotonylation sites by incorporating the composition of k-spaced amino acid pairs into Chou’s general PseAAC
You et al. An improved sequence-based prediction protocol for protein-protein interactions using amino acids substitution matrix and rotation forest ensemble classifiers
Li et al. Protein contact map prediction based on ResNet and DenseNet
Jia et al. S-SulfPred: A sensitive predictor to capture S-sulfenylation sites based on a resampling one-sided selection undersampling-synthetic minority oversampling technique
CN109817275B (zh) 蛋白质功能预测模型生成、蛋白质功能预测方法及装置
CN112837747A (zh) 基于注意力孪生网络的蛋白质结合位点预测方法
Li et al. Using weighted extreme learning machine combined with scale-invariant feature transform to predict protein-protein interactions from protein evolutionary information
Anishchenko et al. Structural templates for comparative protein docking
Zhang et al. Adaptive compressive learning for prediction of protein–protein interactions from primary sequence
Shao et al. DeepSec: a deep learning framework for secreted protein discovery in human body fluids
Lopez et al. C-iSUMO: a sumoylation site predictor that incorporates intrinsic characteristics of amino acid sequences
CN106650309A (zh) 一种膜蛋白残基的作用关系的预测方法和装置
Zhang et al. SPIN-CGNN: Improved fixed backbone protein design with contact map-based graph construction and contact graph neural network
Spadaro et al. Predicting lysine methylation sites using a convolutional neural network
Wang et al. SPDesign: protein sequence designer based on structural sequence profile using ultrafast shape recognition
Zhao et al. PGlcS: prediction of protein O-GlcNAcylation sites with multiple features and analysis
WO2018120128A1 (zh) 一种膜蛋白残基的作用关系的预测方法和装置
Du et al. Improving protein domain classification for third-generation sequencing reads using deep learning
Torrisi et al. Protein structure annotations
Yang et al. Large scale video data analysis based on spark
CN112417163B (zh) 基于实体线索片段的候选实体对齐方法及装置
CN104504299B (zh) 预测膜蛋白的残基间的作用关系的方法
Zuo et al. CarSitePred: an integrated algorithm for identifying carbonylated sites based on KNDUA-LNDOT resampling technique
Rahmani et al. An extension of Wang’s protein design model using Blosum62 substitution matrix

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16925593

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16925593

Country of ref document: EP

Kind code of ref document: A1

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 02/10/2019)

122 Ep: pct application non-entry in european phase

Ref document number: 16925593

Country of ref document: EP

Kind code of ref document: A1