WO2018120128A1 - 一种膜蛋白残基的作用关系的预测方法和装置 - Google Patents
一种膜蛋白残基的作用关系的预测方法和装置 Download PDFInfo
- Publication number
- WO2018120128A1 WO2018120128A1 PCT/CN2016/113754 CN2016113754W WO2018120128A1 WO 2018120128 A1 WO2018120128 A1 WO 2018120128A1 CN 2016113754 W CN2016113754 W CN 2016113754W WO 2018120128 A1 WO2018120128 A1 WO 2018120128A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- residue
- membrane protein
- feature
- features
- pair
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
Definitions
- the present invention belongs to the field of data mining, machine learning and computer biology, and in particular relates to a method and apparatus for predicting the relationship between membrane protein residues.
- membrane proteins account for about 60%. Due to the difficulty of experimental analysis of membrane protein structure, in the protein database (Protein Data Bank-PDB), more than 90,000 known protein structures, known membrane protein structures account for only 1% of the known protein structure. .
- the amino group and the carboxyl group between the amino acids are dehydrated into a bond, and the amino acid participates in the formation of the peptide bond due to a part of the amino acid.
- the structural part is called an amino acid residue.
- the term "residual relationship" refers to those pairs of residues that are not adjacent in the primary sequence of the protein but are adjacent in the tertiary structure.
- the object of the present invention is to provide a method for predicting the relationship between membrane protein residues, so as to solve the problem that the prediction method in the prior art leads to a large amount of useful information loss, affecting the accuracy and coverage of the prediction.
- an embodiment of the present invention provides a method for predicting a relationship of a membrane protein residue, the method comprising:
- the extracted unbalanced classification features are trained by the smote-boost algorithm to obtain a predicted model after training;
- the extracting the residual and non-interacting residues in the membrane protein of the resolved protein structure include: a position-specific score matrix PSSM feature, a relative distance feature in the spiral, a sequence interval feature, a residue type feature, and a number of spirals One or more of a feature, a sequence length feature.
- each of the residues in the location-specific scoring matrix PSSM is represented by a 20-dimensional vector
- the location specific score matrix PSSM features include:
- a residue pair includes two amino acids, the residue type characteristic comprising an acidic amino acid, a base Ten combinations of any two of amino acids, polar amino acids, and non-polar amino acids.
- the interacting residue pair is a residue pair having a CB-CB atom distance of less than 8 angstroms on the alpha helix of the membrane protein .
- an embodiment of the present invention provides a device for predicting a relationship between membrane protein residues, and the device includes: [0018] a training set acquisition unit, configured to acquire a membrane protein of the resolved protein structure as a training set;
- a feature extraction unit configured to extract a feature of the unbalanced classification of the pair of residues and the pair of non-interactive residues in the membrane protein of the resolved protein structure
- a training unit configured to train the predicted model by using the smote-boost algorithm to obtain the trained predictive model
- a prediction unit configured to predict a relationship of membrane protein residues of an unknown protein structure according to the predicted model after training.
- the feature of the unbalanced classification includes: a location-specific score matrix PSSM feature, and a residue in the spiral One or more of a relative distance feature, a sequence interval feature, a residue type feature, a spiral number feature, and a sequence length feature.
- each of the residues in the location-specific scoring matrix PSSM is represented by a 20-dimensional vector.
- the location specific score matrix PSSM features include:
- a residue pair includes two amino acids, the residue type characteristic comprising an acidic amino acid, a base Ten combinations of any two of amino acids, polar amino acids, and non-polar amino acids.
- the interacting residue pair is a residue pair having a CB-CB atom distance of less than 8 angstroms on the alpha helix of the membrane protein .
- a membrane protein of a resolved protein structure is obtained as a training set, and a residue pair and a non-interacting residue pair for distinguishing interactions in the membrane protein of the resolved protein structure are extracted.
- the characteristics of the unbalanced classification, the extracted features are trained by the smooth-boost algorithm to obtain the training model.
- the post-predictive model and based on the post-training prediction model, predicts the relationship of membrane protein residues of unknown protein structures. Because the non-equilibrium classification features are used to train the prediction model, the post-training prediction model can avoid the loss of useful information, which is beneficial to improve the accuracy and coverage of the prediction.
- FIG. 1 is a flow chart for realizing a prediction method for the action relationship of membrane protein residues provided by an embodiment of the present invention
- FIG. 2 is a prediction device for the action relationship of membrane protein residues provided by an embodiment of the present invention
- the purpose of the embodiments of the present invention is to provide a method for predicting the relationship between membrane protein residues, in order to solve the prior art prediction process for the relationship between membrane protein residues of unknown structures, generally from equilibrium classification.
- the angle of the interacting residue pair or non-interacting residue pair is trained in a 1:1 ratio, and in fact, the ratio of interacting or non-interacting residues is much greater than 1: 1, according to equilibrium.
- a peer-to-peer proportional training model can result in the loss of a large amount of useful information, which leads to the problem of the accuracy and coverage of the predicted membrane protein residues.
- FIG. 1 shows an implementation flow of a prediction method for the action relationship of membrane protein residues provided by the first embodiment of the present invention, which is described in detail as follows:
- step S101 a membrane protein of the resolved protein structure is obtained as a training set.
- the membrane protein of the resolved protein structure should have a determined relationship of membrane protein residues.
- a membrane protein which was analyzed before February 2012 in PDBTM (English full name: protein data bank of transmembrane proteins) can be used as a training set.
- step S102 extracting features of the membrane protein of the resolved protein structure for distinguishing the unbalanced classification of the interacting residue pair and the non-interacting residue pair;
- the feature for distinguishing the unbalanced classification of the interacting residue pair and the non-interacting residue pair in the embodiment of the present invention may include a position-specific scoring matrix PSSM feature, and the residue is One or more of the relative distance characteristics, sequence spacing characteristics, residue type characteristics, spiral number characteristics, and sequence length characteristics in the spiral.
- PSSM Position-Specific Scoring Matrix
- PSI-BLAST Alternative Basic Local Alignment Search Tool
- Chinese full name Position-specific iterative search algorithm
- the database that can be used to run PSI-BLAST ⁇ is the UNIREF90 database.
- the number of iterations for running ⁇ can be 2, and the E-value truncation value is le-10 (expressed as 1*10 -10 powers).
- each residue in the position-specific scoring matrix PSSM is represented by a 20-dimensional vector representing the frequency at which 20 amino acids occur at corresponding positions in the PSSM.
- Feature Extraction ⁇ , Location-Specific Scoring Matrix PSSM features are divided into two categories, namely:
- the method may be:
- Position-specific scoring matrix PSSM features
- the relative distance characteristic of the residue in the helix is specifically: assuming that p is a residue in the pair of residues is long The relative position on the helix of degree 1, then the relative distance characteristic of the residue in the a helix is defined as p/l, and for each residue pair including two residues, the residue corresponding to the residue can be extracted separately In the "relative distance feature in the spiral, a total of 2 residues in the spiral relative distance characteristics.
- the sequence spacing feature may be partitioned according to the position of the residue pair in the primary sequence. For example, a specific interval division method can be divided into the following multiple intervals:
- the corresponding sequence interval feature code 000000000 can be set to 0 or set to 1 (0 means not in the interval, and vice versa is 1) for expressing the sequence interval feature.
- one of the nine sequence interval features may be corresponding to the interval division method described above.
- amino acids constituting the protein considering 20 kinds of amino acids constituting the protein, according to the polar nature of the amino acid R group, it can be divided into acidic amino acids (glutamic acid and aspartic acid) and basic amino acids (Lai And arginine, and neutral amino acids, which can be divided into polar amino acids (glycine, serine, cysteine, threonine, tyrosine, asparagine and valley) (aminoamide) and non-polar amino acids (alanine, leucine, isoleucine, phenylalanine, methionine, tryptophan, valine and proline).
- acidic amino acids glutmic acid and aspartic acid
- basic amino acids Lai And arginine
- neutral amino acids which can be divided into polar amino acids (glycine, serine, cysteine, threonine, tyrosine, asparagine and valley) (aminoamide) and non-polar amino acids (alanine, leucine, isoleu
- one residue pair (corresponding to two amino acids) can produce 10 different combinations, which can be binary code 0000000000 respectively. Set to 0 or set to 1 to represent different combinations. It can include 10 residue type features.
- the "spiral number characteristic” may be divided into sections according to the number of alpha helices included in the membrane protein. For example, it can be divided into 4 intervals of 2-4, 5-7, 8-10, and greater than 10.
- the "number of spirals” is represented by the binary vector 0000 being set to 0 or set to 1 (0 means not in the interval, and vice versa). This class of features is consistent for all pairs of residues in a membrane protein. Each residue pair of feature vectors contains four such features.
- the sequence length characteristic can be divided into four intervals of ⁇ 100, 100-400, 400-800, >800 according to the length of the primary sequence of the membrane protein, and the binary vector 0000 is set to 0 or set to 1. Indicates this feature (0 means not in the interval, and vice versa). Such features are consistent for all residue pairs in the same membrane protein. Each residue pair of feature vectors contains four such features.
- the present invention can use 340 position-specific scoring matrix PSSM features, 2 alpha helices Medium relative distance feature, 9 sequence interval features and 10 residue type features, 4 alpha spiral number features
- the proportion of the interacting residue pair and the non-interacting residue pair in the embodiment of the present invention may be 1 to 50 to 1 to 80, and a preferred embodiment may be set to 1 ratio. 67.
- CB-C which will be located on the alpha helix of the membrane protein.
- ⁇ Atom distance is less than 8 ⁇ (Angstrom)
- Residue pairs are defined as pairs of interacting residues.
- CA, CB are atom types in gromacs, groma cs molecular dynamics software.
- step S103 the extracted non-equalized classification features are trained by the smote-boost algorithm to obtain a predicted model after training;
- the features may be substituted into a prediction model for training.
- the prediction model may be a vector machine training model or the like.
- the training algorithm smote-boost is a new training method combining smote technology and boost technology, wherein: the boost method increases the weight of the sample without correct classification in each iteration, and reduces the weight of the correctly classified sample. , pay more attention to the sample of the wrong classification. Because a few samples are more susceptible to misclassification, this approach improves the predictive performance of a few classes.
- SMOTE commonly known as synthetic minority over-sampling rechnique
- SMOTE technology commonly known as synthetic minority over-sampling rechnique
- step S104 the action relationship of the membrane protein residues of the unknown protein structure is predicted based on the predicted model after training.
- the present invention extracts the membrane protein of the resolved protein structure as a training set, and extracts the non-reactive residue pair and the non-interacting residue pair in the membrane protein of the resolved protein structure.
- the characteristics of the equilibrium classification are obtained, and the extracted features are trained by the smooth-boost algorithm to obtain a predicted model after training, and the membrane protein residue of the unknown protein structure is predicted according to the trained prediction model.
- the role of the base is trained by using the features of unbalanced classification, so that the post-training prediction model can avoid the loss of useful information, which is beneficial to improve the accuracy and coverage of prediction.
- FIG. 2 is a schematic view showing the structure of a prediction apparatus for a membrane protein residue relationship according to an embodiment of the present invention, which is described in detail as follows:
- the apparatus for predicting the action relationship of the membrane protein residues in the embodiments of the present invention comprises:
- a training set obtaining unit 201 configured to acquire a membrane protein of the resolved protein structure as a training set
- a feature extraction unit 202 configured to extract a feature of the unbalanced classification of the pair of residual and non-interactive residues in the membrane protein of the resolved protein structure
- the training unit 203 is configured to train the predicted feature of the extracted unbalanced classification by using a smote-boost algorithm to obtain a predicted model after training;
- the prediction unit 204 is configured to predict a relationship of membrane protein residues of an unknown protein structure according to the predicted model after training.
- the features of the unbalanced classification include: a position-specific score matrix, a PSSM feature, a relative distance feature in a spiral, a sequence interval feature, and a residue type feature.
- One or more of the spiral number feature and the sequence length feature are selected from the spiral number feature and the sequence length feature.
- each residue in the position-specific scoring matrix PSSM is represented by a 20-dimensional vector
- the location-specific scoring matrix PSSM features include:
- a residue pair includes two amino acids, and the residue type features include 10 combinations of any one of an acidic amino acid, a basic amino acid, a polar amino acid, and a non-polar amino acid.
- the interacting residue pair is a residue pair having a CB-CB atom distance of less than 8 angstroms on the alpha helix of the membrane protein.
- FIG. 2 is a prediction device for the action relationship of the membrane protein residue, and the membrane protein residue of the first embodiment is used. Corresponding to the prediction method of the relationship, the details are not repeated here.
- the disclosed apparatus and method can be implemented in other manners.
- the device embodiments described above are merely illustrative.
- the division of the unit is only a logical function division, and the actual implementation may have another division manner, for example, multiple units or components may be combined or Can be integrated into another system, or some features can be ignored, or not executed.
- the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or unit, and may be electrical, mechanical or otherwise.
- the unit described as a separate component may or may not be physically distributed, and the component displayed as a unit may or may not be a physical unit, that is, may be located in one place, or may be distributed to multiple On the network unit. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of the embodiment.
- each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
- the above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
- the integrated unit if implemented in the form of a software functional unit and sold or used as a standalone product, may be stored in a computer readable storage medium.
- the technical solution of the present invention may contribute to the prior art or all or part of the technical solution may be embodied in the form of a software product stored in a storage medium.
- a number of instructions are included to cause a computer device (which may be a personal computer, server, or network device, etc.) to perform all or part of the methods described in various embodiments of the present invention.
- the foregoing storage medium includes: a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and the like, which can store program codes. .
Landscapes
- Bioinformatics & Cheminformatics (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Genetics & Genomics (AREA)
- Biotechnology (AREA)
- Biophysics (AREA)
- Chemical & Material Sciences (AREA)
- Molecular Biology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Bioinformatics & Computational Biology (AREA)
- Analytical Chemistry (AREA)
- Evolutionary Biology (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Theoretical Computer Science (AREA)
- Investigating Or Analysing Biological Materials (AREA)
Abstract
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2016/113754 WO2018120128A1 (zh) | 2016-12-30 | 2016-12-30 | 一种膜蛋白残基的作用关系的预测方法和装置 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2016/113754 WO2018120128A1 (zh) | 2016-12-30 | 2016-12-30 | 一种膜蛋白残基的作用关系的预测方法和装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018120128A1 true WO2018120128A1 (zh) | 2018-07-05 |
Family
ID=62706805
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2016/113754 Ceased WO2018120128A1 (zh) | 2016-12-30 | 2016-12-30 | 一种膜蛋白残基的作用关系的预测方法和装置 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2018120128A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117593783A (zh) * | 2023-11-20 | 2024-02-23 | 广州视景医疗软件有限公司 | 基于自适应smote的视觉训练方案生成方法及装置 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104252581A (zh) * | 2013-06-26 | 2014-12-31 | 中国科学院深圳先进技术研究院 | 一种基于支持向量机的跨膜蛋白残基作用关系预测方法 |
| CN104504299A (zh) * | 2014-12-29 | 2015-04-08 | 中国科学院深圳先进技术研究院 | 预测膜蛋白的残基间的作用关系的方法 |
| CN104615910A (zh) * | 2014-12-30 | 2015-05-13 | 中国科学院深圳先进技术研究院 | 基于随机森林预测α跨膜蛋白的螺旋相互作用关系的方法 |
| CN106650309A (zh) * | 2016-12-30 | 2017-05-10 | 中国科学院深圳先进技术研究院 | 一种膜蛋白残基的作用关系的预测方法和装置 |
-
2016
- 2016-12-30 WO PCT/CN2016/113754 patent/WO2018120128A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104252581A (zh) * | 2013-06-26 | 2014-12-31 | 中国科学院深圳先进技术研究院 | 一种基于支持向量机的跨膜蛋白残基作用关系预测方法 |
| CN104504299A (zh) * | 2014-12-29 | 2015-04-08 | 中国科学院深圳先进技术研究院 | 预测膜蛋白的残基间的作用关系的方法 |
| CN104615910A (zh) * | 2014-12-30 | 2015-05-13 | 中国科学院深圳先进技术研究院 | 基于随机森林预测α跨膜蛋白的螺旋相互作用关系的方法 |
| CN106650309A (zh) * | 2016-12-30 | 2017-05-10 | 中国科学院深圳先进技术研究院 | 一种膜蛋白残基的作用关系的预测方法和装置 |
Non-Patent Citations (2)
| Title |
|---|
| ARUN KUMAR, M.N. ET AL.: "On the Classification of Imbalanced Datasets", INTERNATIONAL JOURNAL OF COMPUTER APPLICATIONS, vol. 44, no. 8, 30 April 2012 (2012-04-30), pages 1 - 7, XP055509822 * |
| FATTAHI, S.: "NEW APPROACH FOR IMBALANCED BIOLOGICAL DATASET CLASSIFICATION", JOURNAL OF THEORETICAL AND APPLIED INFORMATION TECHNOLOGY, vol. 72, no. 1, 10 February 2015 (2015-02-10), pages 40 - 57, XP055509829 * |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117593783A (zh) * | 2023-11-20 | 2024-02-23 | 广州视景医疗软件有限公司 | 基于自适应smote的视觉训练方案生成方法及装置 |
| CN117593783B (zh) * | 2023-11-20 | 2024-04-05 | 广州视景医疗软件有限公司 | 基于自适应smote的视觉训练方案生成方法及装置 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Tsukiyama et al. | LSTM-PHV: prediction of human-virus protein–protein interactions by LSTM with word2vec | |
| Ju et al. | Prediction of lysine crotonylation sites by incorporating the composition of k-spaced amino acid pairs into Chou’s general PseAAC | |
| You et al. | An improved sequence-based prediction protocol for protein-protein interactions using amino acids substitution matrix and rotation forest ensemble classifiers | |
| Li et al. | Protein contact map prediction based on ResNet and DenseNet | |
| Jia et al. | S-SulfPred: A sensitive predictor to capture S-sulfenylation sites based on a resampling one-sided selection undersampling-synthetic minority oversampling technique | |
| CN109817275B (zh) | 蛋白质功能预测模型生成、蛋白质功能预测方法及装置 | |
| CN112837747A (zh) | 基于注意力孪生网络的蛋白质结合位点预测方法 | |
| Li et al. | Using weighted extreme learning machine combined with scale-invariant feature transform to predict protein-protein interactions from protein evolutionary information | |
| Anishchenko et al. | Structural templates for comparative protein docking | |
| Zhang et al. | Adaptive compressive learning for prediction of protein–protein interactions from primary sequence | |
| Shao et al. | DeepSec: a deep learning framework for secreted protein discovery in human body fluids | |
| Lopez et al. | C-iSUMO: a sumoylation site predictor that incorporates intrinsic characteristics of amino acid sequences | |
| CN106650309A (zh) | 一种膜蛋白残基的作用关系的预测方法和装置 | |
| Zhang et al. | SPIN-CGNN: Improved fixed backbone protein design with contact map-based graph construction and contact graph neural network | |
| Spadaro et al. | Predicting lysine methylation sites using a convolutional neural network | |
| Wang et al. | SPDesign: protein sequence designer based on structural sequence profile using ultrafast shape recognition | |
| Zhao et al. | PGlcS: prediction of protein O-GlcNAcylation sites with multiple features and analysis | |
| WO2018120128A1 (zh) | 一种膜蛋白残基的作用关系的预测方法和装置 | |
| Du et al. | Improving protein domain classification for third-generation sequencing reads using deep learning | |
| Torrisi et al. | Protein structure annotations | |
| Yang et al. | Large scale video data analysis based on spark | |
| CN112417163B (zh) | 基于实体线索片段的候选实体对齐方法及装置 | |
| CN104504299B (zh) | 预测膜蛋白的残基间的作用关系的方法 | |
| Zuo et al. | CarSitePred: an integrated algorithm for identifying carbonylated sites based on KNDUA-LNDOT resampling technique | |
| Rahmani et al. | An extension of Wang’s protein design model using Blosum62 substitution matrix |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 16925593 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 16925593 Country of ref document: EP Kind code of ref document: A1 |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 02/10/2019) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 16925593 Country of ref document: EP Kind code of ref document: A1 |