WO2009113752A1 - System comprising scoring algorithm and method for identifying alternative splicing isoforms using peptide mass fingerprinting, and recording media having program therefor - Google Patents
System comprising scoring algorithm and method for identifying alternative splicing isoforms using peptide mass fingerprinting, and recording media having program therefor Download PDFInfo
- Publication number
- WO2009113752A1 WO2009113752A1 PCT/KR2008/002390 KR2008002390W WO2009113752A1 WO 2009113752 A1 WO2009113752 A1 WO 2009113752A1 KR 2008002390 W KR2008002390 W KR 2008002390W WO 2009113752 A1 WO2009113752 A1 WO 2009113752A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- mass
- database
- isoforms
- alternative splicing
- algorithm
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N33/00—Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
- G01N33/48—Biological material, e.g. blood, urine; Haemocytometers
- G01N33/50—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
- G01N33/68—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
- G01N33/6803—General methods of protein analysis not limited to specific proteins or families of proteins
- G01N33/6848—Methods of protein analysis involving mass spectrometry
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
Definitions
- the present invention relates to a system comprising a scoring algorithm and a method for identifying alternative splicing isoforms by using peptide mass fingerprinting, and recording media having a computer readable program to implement the method.
- Peptide mass fingerprinting (herein after, sometimes referred to as "PMF") is a technique used to determine the mass distribution of peptide fragments that are obtained from protein hydrolysis. Similar to a line spectrum which serves as a fingerprint for an element, mass of peptide fragments is a fingerprint of a protein. In this case, it is unnecessary to measure the molecular weight of all peptides obtained from the protein hydrolysis. Instead, an accurate measurement of the molecular weight of several peptides is usually enough and it is because for two separate proteins a probability of obtaining a combination of peptides having exactly the same molecular weight is extremely low.
- Alternative splicing is a mechanism of producing two or more various and separate proteins from a single gene in a complex organism. With such alternative splicing, a great variety of proteins can be produced from a limited number of genes at different time and location. It was also reported that alternative splicing is related to the occurrence of some diseases. [5] For an identification of an alternative splicing isoform, usually proteins and transcripts are experimentally determined.
- a microarray system As a method for determining a transcript, a microarray system is used. There is a method which designs and uses exon probes and exon-exon junction probes (Nagao, K. (2005) Human Molecular Genetics, 14, 3379-3388; Shoemaker, D. (2001) Nature, 409, 922-927). However, because not all of those determined at transcription level are translated into a protein, it is advantageous that the determination is carried out at protein level.
- Korean Patent Registration No. 10-0789430 discloses a method for determining isotopic clusters and monoisotopic masses of polypeptides on mass spectra of complex polypeptide mixtures and computer-readable medium thereof.
- Korean Patent Registration No. 10-0757040 discloses a system and method for analysis of interfaces based on protein domain and recording medium therefor.
- PCT publication WO 01/57519 discloses a method for identifying and/or characterizing a (poly)peptide. Disclosure of Invention Technical Problem
- the present intention which is devised in view of said necessity, relates to a development of a system and a method for identifying alternative splicing isoforms that have been known or not known from prior art, and an identification of alternative splicing isoforms using such system and method.
- the present invention provides a system and a method for identifying alternative splicing isoforms using peptide mass fingerprinting.
- the present invention provides recording media having a computer readable program that can be used to implement said method.
- Figure 1 is a schematic diagram representing the system for identifying alternative splicing isoforms according to the present invention.
- Figure 2 is a diagram describing the algorithm of the present invention.
- Figure 3 shows a screen for input section.
- Figure 4 shows a screen for output section.
- the present invention provides a system encompassing entire processes for identifying alternative splicing isoforms by using peptide mass fingerprinting, which comprises,
- transcripts of splicing isoforms are collected for a species having an alternative splicing, and then they can be translated into a protein sequence to construct the database.
- said algorithm can be an algorithm which determines priority of each candidate with a process that mass data measured from an experiment is searched against the established database and then the results are calculated for masses that are in match with respective isoforms by using the following formulae 1, 2 and 3.
- P s value is a sum of E Mj for the mass values in match with respective isoforms.
- P s is a sum of E Mj that are generated from applying k number of matching mass values to formulae 1 and 2. Smaller the P s value is, a more significant candidate it becomes. This means that, smaller the possibility of an accidental match with measured mass values, a more significant candidate it becomes. Because alternative isoforms originating from a single gene share a common peptide, they will have a similar score and thus can be concentrated.
- the present invention further provides a method for identifying alternative splicing isoforms by using peptide mass fingerprinting.
- said method comprises following steps:
- transcripts from the species having an alternative splicing event are virtually translated into proteins and the resulting protein sequences are also virtually digested with a proteolytic enzyme as if they are digested with it in real. Consequently, generated peptides are stored in the database of the present invention.
- the algorithm can be an algorithm which determines priority of each candidate with a process that mass data measured from an experiment is searched against the established database and then the results are calculated for masses that are in match with respective isoforms by using the following formulae 1, 2 and 3.
- the implementation of said algorithm can be achieved by processes including, searching within error range of mass the mass data that has been supplied from the input section and the mass considering the error of mass against the database, sorting the results for each isoforms, scoring them by using said algorithm, and sorting them in an ascending order to decide the priority among candidates.
- the present invention still further provides recording media having a computer readable program to implement said method for identifying alternative splicing isoforms by using peptide mass fingerprinting.
- Figure 1 is a schematic diagram representing the system for identifying alternative splicing isoforms according to the present invention.
- the system for identifying alternative splicing isoforms by using peptide mass fingerprinting includes an input section; a database; a search section; an algorithm; and an output section.
- the input section is where the mass data input of the peptide fragments obtained from protein hydrolysis occurs.
- Figure 3 shows a screen for an input section. Essential mass values, optional protein mass values and isopotential values are fed to the input section.
- the above-described search section is where a search against the database is carried out.
- the output section is where the candidate isoforms are sorted by their score.
- Figure 4 shows a screen for output section. Closer to the top of the output screen, it is more likely that it is a correct isoform.
- the present invention furthermore provides a method for identifying alternative splicing isoforms by using peptide mass fingerprinting.
- Said method for identifying alternative splicing isoforms by using peptide mass fingerprinting according to the present invention can provide a database which can be used for program search of mass data that is generated from mass spectrometry. That is, after searching mass data generated from mass spectrometry against said database, the isoforms having a mass with match in the database are scored by using said algorithm, and the candidate isoforms are aligned in an ascending order with their scores.
- FIG. 2 mass value (201) obtained from a mass spectrometer is searched against database (202).
- A, B, C, D, L, M, and N in 202 are imaginary circles having a size proportional to the number of peptides that are in match after the search with said measured mass values. Bigger sized-circle indicates that there are more peptides in match. On the other hand, smaller the circle is, probability of accidental occurrence is lower.
- isoforms having a matching mass value is presented as scores as shown in 203. Each isoforms are aligned in an ascending order of score such as from Cl, C2, ..., to C7.
- the isoforms originating from a single gene tend to share the same peptide fragments, thus they end up having a similar score.
- Cl, C2, C3 isoforms originating from a single gene show an order of 1, 2 and 3 from the top. Then, an order of significance of each candidate is decided based on matched mass values of L, M, and N that are different to each other. In this case, because the size of L is smallest, Cl becomes the most significant candidate.
- ⁇ m 1 1 °g
- Mass data (201) measured by a mass spectrometer is searched against the mass value of the database (202) considering an error value.
- the number determined by search with all of the measured mass values corresponds to the size of dotted-line circle of Figure 2.
- Each mass value can be searched against the database.
- Circles A, B, C, D, L, M and N of (202) are represented by f ⁇ and the formula 2 can be solved using the values obtained before.
- E Mj values of formula 2 which correspond to the mass values in match with each isoform are all added up to yield a specific score.
- the present invention still further provides recording media having a computer readable program to implement the method for identifying alternative splicing isoforms by using peptide mass fingerprinting.
- Recording media having a computer readable program can be any media that can be directly read and accessed through a computer.
- Examples of such recording medic includes; magnetic recording media such as a floppy disk, a hard disk and a magnetic tape, etc., optical recording media such as CD-ROM, CD-R, CD, RW, DVD-ROM, DVD-RAM, and DVD-RW, etc., and electric recording media such as RAM and ROM, etc., and combinations thereof (e.g., magnetic/optical recording media such as MO, etc.), but are not limited thereto.
- Selection of an instrument that is used for recording and storing data into the above- described recording media and an instrument or a device that is used for reading the data information from recording media is determined by the kinds of recording media and an access method being used. Further, various kinds of data processor program, software, comparator, and format are used for recording in corresponding media the program that is required for the implementation of the method of the present invention. Information can be presented in a form of, for example, a binary file formatted with commercially available software, text file or ASCII file.
Landscapes
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Molecular Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Chemical & Material Sciences (AREA)
- Biophysics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Biotechnology (AREA)
- Analytical Chemistry (AREA)
- General Health & Medical Sciences (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Theoretical Computer Science (AREA)
- Medical Informatics (AREA)
- Evolutionary Biology (AREA)
- Genetics & Genomics (AREA)
- Immunology (AREA)
- Urology & Nephrology (AREA)
- Hematology (AREA)
- Biomedical Technology (AREA)
- Medicinal Chemistry (AREA)
- Food Science & Technology (AREA)
- Biochemistry (AREA)
- General Physics & Mathematics (AREA)
- Pathology (AREA)
- Microbiology (AREA)
- Cell Biology (AREA)
- Other Investigation Or Analysis Of Materials By Electrical Means (AREA)
Abstract
The present invention relates to a system comprising a scoring algorithm and a method for identifying alternative splicing isoforms by using peptide mass fingerprinting, and recording media having a computer readable program to implement the method. By using a database including information of tissue-specific gene and an algorithm which can efficiently distinguish alternative splicing isoforms of a gene from another, it is possible according to the present invention that alternative splicing isoforms can be efficiently characterized and identified and tissue-specific isoforms can be determined.
Description
Description
SYSTEM COMPRISING SCORING ALGORITHM AND
METHOD FOR IDENTIFYING ALTERNATIVE SPLICING
ISOFORMS USING PEPTIDE MASS FINGERPRINTING, AND
RECORDING MEDIA HAVING PROGRAM THEREFOR
Technical Field
[1] The present invention relates to a system comprising a scoring algorithm and a method for identifying alternative splicing isoforms by using peptide mass fingerprinting, and recording media having a computer readable program to implement the method. Background Art
[2] Inside cells hundreds or thousands of different proteins are constantly being expressed in various amounts. As such, it is very challenging to determine the kinds and the amounts of such proteins. Even if a protein of interest is isolated in pure state, molecular weight determination alone is not enough for a precise identification of protein since there are many different kinds of proteins having similar molecular weight present in a cell. Meanwhile, if a protein is digested with an enzyme which can hydrolyze the protein at a specific site and molecular weights of the resulting peptide fragments are determined, the volume of information that can be used for protein identification will dramatically increase and it would lead a way to characterize the original protein.
[3] Peptide mass fingerprinting (herein after, sometimes referred to as "PMF") is a technique used to determine the mass distribution of peptide fragments that are obtained from protein hydrolysis. Similar to a line spectrum which serves as a fingerprint for an element, mass of peptide fragments is a fingerprint of a protein. In this case, it is unnecessary to measure the molecular weight of all peptides obtained from the protein hydrolysis. Instead, an accurate measurement of the molecular weight of several peptides is usually enough and it is because for two separate proteins a probability of obtaining a combination of peptides having exactly the same molecular weight is extremely low.
[4] Alternative splicing is a mechanism of producing two or more various and separate proteins from a single gene in a complex organism. With such alternative splicing, a great variety of proteins can be produced from a limited number of genes at different time and location. It was also reported that alternative splicing is related to the occurrence of some diseases.
[5] For an identification of an alternative splicing isoform, usually proteins and transcripts are experimentally determined.
[6] As a method for determining a transcript, a microarray system is used. There is a method which designs and uses exon probes and exon-exon junction probes (Nagao, K. (2005) Human Molecular Genetics, 14, 3379-3388; Shoemaker, D. (2001) Nature, 409, 922-927). However, because not all of those determined at transcription level are translated into a protein, it is advantageous that the determination is carried out at protein level.
[7] For a protein level identification, mass analysis is used. For a determination of splicing sites, there is a method based on protein sequencing with a combined mass spectrometry and a mass spectrometry (Tanner, S. (2007) Genome research, 17, 231-239). In addition, there is a method by which a genome is translated into six frames of proteins and then splicing site is identified (Giddings, M. (2003) Proc Natl Acad Sci USA, 100, 20-25). Although being suitable for locating a splicing site in a gene, these methods are not useful for obtaining an isoform. The reason is that, not all of the fragment mass of the peptides that are produced by proteolytic enzymes are confirmed during mass analysis. In addition, if an alternative splicing isoform is not taken into consideration, a wrong type of isoform can be identified from peptide mass fingerprinting, because there can be many isoforms present that have not yet been verified by an experiment.
[8] Meanwhile, in order to determine an alternative splicing isoform by using PMF technique, a database wherein the presence of such alternative splicing is considered should be constructed. A gene having an event of alternative splicing in such database should have a common exon so that it may have a protein sequence that is partially same to each other. In addition, there can be more than one hundred isoforms which can be generated from a single gene. Due to such characteristics, a new algorithm that can distinguish isoforms different to each other is in need.
[9] Korean Patent Registration No. 10-0789430 discloses a method for determining isotopic clusters and monoisotopic masses of polypeptides on mass spectra of complex polypeptide mixtures and computer-readable medium thereof. Korean Patent Registration No. 10-0757040 discloses a system and method for analysis of interfaces based on protein domain and recording medium therefor. In addition, PCT publication WO 01/57519 discloses a method for identifying and/or characterizing a (poly)peptide. Disclosure of Invention Technical Problem
[10] The present intention, which is devised in view of said necessity, relates to a development of a system and a method for identifying alternative splicing isoforms that
have been known or not known from prior art, and an identification of alternative splicing isoforms using such system and method.
[11] According to the present invention, by using an algorithm which can efficiently distinguish alternative splicing forms of a gene from another and a database which includes tissue-specific genetic information, the alternative splicing isoforms can be efficiently characterized and identified and tissue- specific isoforms can be confirmed. Technical Solution
[12] In order to solve the problems described above, the present invention provides a system and a method for identifying alternative splicing isoforms using peptide mass fingerprinting. In addition, the present invention provides recording media having a computer readable program that can be used to implement said method.
Advantageous Effects
[13] According to the present invention, alternative splicing isoforms can be identified from the mass distribution of peptides that are generated by mass spectrometry. As a result, isoforms that have been previously unknown and isoforms that are related to a certain disease or specific for certain tissues can be identified. Brief Description of the Drawings
[14] Figure 1 is a schematic diagram representing the system for identifying alternative splicing isoforms according to the present invention.
[15] Figure 2 is a diagram describing the algorithm of the present invention.
[16] Figure 3 shows a screen for input section.
[17] Figure 4 shows a screen for output section.
Mode for the Invention
[18] In order to achieve said object of the invention, the present invention provides a system encompassing entire processes for identifying alternative splicing isoforms by using peptide mass fingerprinting, which comprises,
[19] - an input section to which mass data measured for peptide fragments that are obtained from protein hydrolysis are input;
[20] - a database comprising mass and isopotential value of a protein, site information of an amino acid which might be involved in protein modification and sequence information of a peptide that is produced after the hydrolysis of the protein;
[21] - a search section by which the peptide mass data is searched against an established database;
[22] - an algorithm for scoring a potential of candidate alternative splicing isoforms having a mass with match found in database; and
[23] - an output section by which the candidate isoforms are aligned according to their score.
[24] According to a system of one embodiment of the present invention, transcripts of splicing isoforms are collected for a species having an alternative splicing, and then they can be translated into a protein sequence to construct the database.
[25] According to a system of one embodiment of the present invention, said algorithm can be an algorithm which determines priority of each candidate with a process that mass data measured from an experiment is searched against the established database and then the results are calculated for masses that are in match with respective isoforms by using the following formulae 1, 2 and 3.
[26] r m. = J fm. I T Fo.mula l
En,, = lθg|θ (c, ) Fθl im'lθ:
Formula 3
'Mj /-1
[27] In the above-described formulae 1, 2 and 3, symbols used are as follows. Specifically, when measured mass value is represented by m = {ml, m2,..., mi} and a mass value with match in the database is represented by M = {M1,M2, ..., Mk}, mi indicates mass value of m while Mk indicates mass value of M. Mk is a single value that is in match with one candidate isoform among mi. When a search is carried out using a measured data and total number of peptides that are searched from the database is T and the number of peptides that are searched by using each of the mass values from the database is fm, rm can be obtained from the above formula 1. In this case, since the values can be too small to handle, rm can be converted to Em based on the above-described logarithmic formula 2. Ps value is a sum of EMj for the mass values in match with respective isoforms. In other words, if there are k mass values in match with a certain isoform, Ps is a sum of EMj that are generated from applying k number of matching mass values to formulae 1 and 2. Smaller the Ps value is, a more significant candidate it becomes. This means that, smaller the possibility of an accidental match with measured mass values, a more significant candidate it becomes. Because alternative isoforms originating from a single gene share a common peptide, they will have a similar score and thus can be concentrated.
[28] The present invention further provides a method for identifying alternative splicing isoforms by using peptide mass fingerprinting.
[29] More specifically, said method comprises following steps:
[30] - constructing a database for searching alternative splicing isoforms;
[31] - searching the mass data of the peptide fragments obtained from protein hydrolysis against said database;
[32] - scoring by algorithm which can score the potential of each candidate that have been
searched for the alternative splicing isoforms having a mass with match in the database; and,
[33] - identifying the isoforms by sorting them out by the score.
[34] In the present invention, transcripts from the species having an alternative splicing event are virtually translated into proteins and the resulting protein sequences are also virtually digested with a proteolytic enzyme as if they are digested with it in real. Consequently, generated peptides are stored in the database of the present invention.
[35] According to a method of one embodiment of the present invention, the database is constructed by processes including collecting transcripts for the species having an alternative splicing event and translating them into protein sequences, and the database may comprise mass and isopotential value of protein, site information of amino acid which might be involved in protein modification and sequence information of a peptide that is produced after the hydrolysis of the protein.
[36] According to a method of one embodiment of the present invention, the algorithm can be an algorithm which determines priority of each candidate with a process that mass data measured from an experiment is searched against the established database and then the results are calculated for masses that are in match with respective isoforms by using the following formulae 1, 2 and 3.
[37] r = f I T Foimula l
E m, = logic lO Foπnula2
[38] Definition of the symbols described in said formulae 1 to 3 is the same as described before.
[39] According to a method of one embodiment of the present invention, the implementation of said algorithm can be achieved by processes including, searching within error range of mass the mass data that has been supplied from the input section and the mass considering the error of mass against the database, sorting the results for each isoforms, scoring them by using said algorithm, and sorting them in an ascending order to decide the priority among candidates.
[40] The present invention still further provides recording media having a computer readable program to implement said method for identifying alternative splicing isoforms by using peptide mass fingerprinting.
[41] Herein after, preferred examples of the present invention will be explained in greater detail in view of the attached drawings.
[42] Figure 1 is a schematic diagram representing the system for identifying alternative
splicing isoforms according to the present invention.
[43] The system for identifying alternative splicing isoforms by using peptide mass fingerprinting according to the present invention includes an input section; a database; a search section; an algorithm; and an output section.
[44] The input section is where the mass data input of the peptide fragments obtained from protein hydrolysis occurs. Figure 3 shows a screen for an input section. Essential mass values, optional protein mass values and isopotential values are fed to the input section.
[45] The above-described database includes information regarding mass and isopotential value of protein sequence, site of amino acid that might be involved in protein modification, and peptide sequences that are produced after the enzyme treatment, etc. To construct the database, after collecting transcripts for the species having an alternative splicing, they are translated into protein sequences and redundant protein sequences are removed.
[46] The above-described search section is where a search against the database is carried out.
[47] The above-described algorithm scores potential of candidates that are found during search step in relation with the alternative splicing isoforms which have masses with match in the database constructed as above. More specifically, said algorithm can be an algorithm which determines priority of each candidate with a process that mass data measured from an experiment is searched against the established database and then the results are calculated for masses that are in match with respective isoforms by using the following formulae 1, 2 and 3.
[48] r in. = J f m. I T Formula !
E m, = lθ8ιo (/'«,. ) Formula2
[49] Definition of the symbols described in said formulae 1 to 3 is the same as described before.
[50] The output section is where the candidate isoforms are sorted by their score. Figure 4 shows a screen for output section. Closer to the top of the output screen, it is more likely that it is a correct isoform.
[51] The present invention furthermore provides a method for identifying alternative splicing isoforms by using peptide mass fingerprinting. Said method for identifying alternative splicing isoforms by using peptide mass fingerprinting according to the present invention can provide a database which can be used for program search of
mass data that is generated from mass spectrometry. That is, after searching mass data generated from mass spectrometry against said database, the isoforms having a mass with match in the database are scored by using said algorithm, and the candidate isoforms are aligned in an ascending order with their scores.
[52] In order to identify alternative splicing isoforms by using peptide mass fingerprinting, a database for searching mass data obtained from mass spectrometry, wherein an alternative splicing is taken into consideration, is required. In addition, an algorithm guaranteeing precise identification is necessary. Using a protein sequence comprising alternative splicing isoforms, a database which can be used for analyzing the results obtained from mass spectrometry is constructed based on the program for producing the search database according to the present invention. This database includes information regarding site of amino acid that might be involved in protein modification, isopotential value of each isoforms and mass of each isoforms, etc. The amino acid site at which a protein sequence can be modified indicates a site wherein the chemical structure of a certain amino acid can be modified after the translation of corresponding transcript.
[53] Considering that an alternative splicing may occur, there can be several to hundred or more of isoforms can be present for a single gene. However, they tend to have a peptide sequence that is partially same to each other. In this regard, a new algorithm which can take an advantage of such characteristic is in need to efficiently distinguish the isoforms.
[54] To identify the isoforms from the masses that are generated from an experiment, a search using a web-based software can be carried out. In this case, mass values that are essential for data input, optional protein mass values, and isopotential values are added (Figure 3). The added data is searched against the database described above, scored by said algorithm, and aligned in an ascending order with their scores (Figure 4). Closer to the top, it is more likely that it corresponds to a correct isoform. In other words, more significant candidate will have a lower score.
[55] Detailed description of the algorithm of the present invention is given as a diagram of
Figure 2. As it can be seen from Figure 2, mass value (201) obtained from a mass spectrometer is searched against database (202). A, B, C, D, L, M, and N in 202 are imaginary circles having a size proportional to the number of peptides that are in match after the search with said measured mass values. Bigger sized-circle indicates that there are more peptides in match. On the other hand, smaller the circle is, probability of accidental occurrence is lower. With the search of the measured mass values against the database, isoforms having a matching mass value is presented as scores as shown in 203. Each isoforms are aligned in an ascending order of score such as from Cl, C2, ..., to C7. The isoforms originating from a single gene tend to share the same peptide
fragments, thus they end up having a similar score. As it is shown in 204, Cl, C2, C3 isoforms originating from a single gene show an order of 1, 2 and 3 from the top. Then, an order of significance of each candidate is decided based on matched mass values of L, M, and N that are different to each other. In this case, because the size of L is smallest, Cl becomes the most significant candidate.
[56] r m. — J fm. I T Formula 1
^m1 = 1°g|θ(0 FθmUlla2 k
P s = / > i K ' ΛM1.j Formula 3
[57] Definition of the symbols described in said formulae 1 to 3 is the same as described before.
[58] Mass data (201) measured by a mass spectrometer is searched against the mass value of the database (202) considering an error value. The number determined by search with all of the measured mass values corresponds to the size of dotted-line circle of Figure 2. Each mass value can be searched against the database. Circles A, B, C, D, L, M and N of (202) are represented by f^ and the formula 2 can be solved using the values obtained before. In addition, for the isoforms that have been searched using all of the measured mass values, EMj values of formula 2 which correspond to the mass values in match with each isoform are all added up to yield a specific score.
[59] The present invention still further provides recording media having a computer readable program to implement the method for identifying alternative splicing isoforms by using peptide mass fingerprinting.
[60] Recording media having a computer readable program can be any media that can be directly read and accessed through a computer. Examples of such recording medic includes; magnetic recording media such as a floppy disk, a hard disk and a magnetic tape, etc., optical recording media such as CD-ROM, CD-R, CD, RW, DVD-ROM, DVD-RAM, and DVD-RW, etc., and electric recording media such as RAM and ROM, etc., and combinations thereof (e.g., magnetic/optical recording media such as MO, etc.), but are not limited thereto.
[61] Selection of an instrument that is used for recording and storing data into the above- described recording media and an instrument or a device that is used for reading the data information from recording media is determined by the kinds of recording media and an access method being used. Further, various kinds of data processor program, software, comparator, and format are used for recording in corresponding media the program that is required for the implementation of the method of the present invention. Information can be presented in a form of, for example, a binary file formatted with
commercially available software, text file or ASCII file.
[62] A skilled person in the pertinent art would understand that the present invention can be implemented in another specific forms without changing the technical spirit or essential characteristics of the invention. Therefore, it must be understood that the examples described above are illustrative only and no way limit the present invention. The scope of the present invention is described by the claims that are given below, rather than by the detailed description of the invention. In addition, it should be understood that meaning and scope of the claims and variants or varied forms that can be deduced from the claims and their equivalents are all within the scope of the present invention.
Claims
[1] A system encompassing entire processes for identifying alternative splicing isoforms using peptide mass fingerprinting, which comprises;
- an input section to which mass data measured for peptide fragments that are obtained from protein hydrolysis are input;
- a database comprising mass and isopotential value of a protein, site information of an amino acid which might be involved in protein modification and sequence information of a peptide that is produced after the hydrolysis of the protein;
- a search section wherein a search is carried out against the database;
- an algorithm for scoring a potential of candidate alternative splicing isoforms having a mass with match found in database; and
- an output section by which the candidate isoforms are aligned according to their score.
[2] The system according to Claim 1, characterized in that said database is constructed by processes including collecting transcripts for a species having an alternative splicing event, translating them into protein sequences, and removing the overlapped protein sequences.
[3] The system according to Claim 1, characterized in that said algorithm is an algorithm which determines priority of each candidate with a process that mass data measured from an experiment is searched against the database and then the results are calculated for masses that are in match with respective isoforms by using the following formulae 1, 2 and 3: r = f I T Formula 1
Em, = 1°8lθ('' m, ) Fθ"1"1"'2
Formula 1
'M]
(wherein, when measured mass value is represented by m = {ml, m2,..., mi} and a mass value with match in the database is represented by M = {M1,M2, ..., Mk}, mi indicates mass value of m, Mk indicates mass value of M and Mk is a single value that is in match with one candidate isoform among mi, T represents total number of peptides that are searched from the database when the search is carried out using measured data, fm represents the number of peptides that are searched with each of the mass values from the database, Em is a logarithmic value of rm, and Ps (score) is a sum of EMj for the mass values in match with respective isoforms).
[4] A method for identifying alternative splicing isoforms by using peptide mass fingerprinting, which comprises;
- constructing a database for searching alternative splicing isoforms;
- searching the mass data of the peptide fragments obtained from protein hydrolysis against said database;
- scoring by algorithm which can score the potential of each candidate that has been searched for the alternative splicing isoforms having a mass with match in the database; and
- identifying the isoforms by sorting them out by the score.
[5] The method according to Claim 4, characterized in said algorithm is an algorithm which determines priority of each candidate with a process that mass data measured from an experiment is searched against the database and then the results are calculated for masses that are in match with respective isoforms by using the following formulae 1, 2 and 3: r = f I T Formula 1 ) Foπnula2
(wherein, definition of the symbols described in said formulae 1 to 3 is the same as described in Claim 3).
[6] The method according to Claim 5, characterized in that the implementation of said algorithm is achieved by processes including, searching within error range of mass the mass data that is supplied from the input section and the mass considering the error of mass against the database, sorting the results for each isoforms, scoring them by using said algorithm, and sorting them in an ascending order to decide the priority among the candidates.
[7] Recording media having a computer readable program to implement the method described in any one of Claims 4 to 6.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR1020080023411A KR100856526B1 (en) | 2008-03-13 | 2008-03-13 | Systems and methods including scoring algorithms for identifying alternating splicing isoforms using peptide mass fingerprint tracking and recording media having programs for performing the methods |
| KR10-2008-0023411 | 2008-03-13 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2009113752A1 true WO2009113752A1 (en) | 2009-09-17 |
Family
ID=40022402
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2008/002390 Ceased WO2009113752A1 (en) | 2008-03-13 | 2008-04-28 | System comprising scoring algorithm and method for identifying alternative splicing isoforms using peptide mass fingerprinting, and recording media having program therefor |
Country Status (2)
| Country | Link |
|---|---|
| KR (1) | KR100856526B1 (en) |
| WO (1) | WO2009113752A1 (en) |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR101168371B1 (en) | 2009-12-17 | 2012-07-24 | 경상대학교산학협력단 | Mass Spectrometry based Pathogen Diagnosis and Biomarker analysis |
| WO2012086859A1 (en) * | 2010-12-22 | 2012-06-28 | 경상대학교 산학협력단 | Pathogen diagnosis and biomarker analysis using mass spectroscope |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR100531207B1 (en) * | 2005-06-04 | 2005-11-29 | 씨비에스소프트주식회사 | Protein identification system |
| JP2006058111A (en) * | 2004-08-19 | 2006-03-02 | Shimadzu Corp | Protein identification processing method and apparatus |
| KR20070017676A (en) * | 2005-08-08 | 2007-02-13 | 한국기초과학지원연구원 | How to Analyze Sequence and Modification Information of Modified Polypeptides |
| KR100757040B1 (en) * | 2005-12-12 | 2007-09-07 | 오브젝트인터랙션테크놀로지스(주) | System and method for analyzing surface of interaction between proteins or proteins and compounds based on protein domains, and recording media therefor |
| US20070224704A1 (en) * | 2006-03-23 | 2007-09-27 | Epitome Biosystems, Inc. | Protein splice variant / isoform discrimination and quantitative measurements thereof |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP1254373B1 (en) | 2000-02-07 | 2007-04-11 | Max-Planck-Gesellschaft zur Förderung der Wissenschaften e.V. | Method for identifying and/or characterizing a (poly)peptide |
-
2008
- 2008-03-13 KR KR1020080023411A patent/KR100856526B1/en not_active Expired - Fee Related
- 2008-04-28 WO PCT/KR2008/002390 patent/WO2009113752A1/en not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006058111A (en) * | 2004-08-19 | 2006-03-02 | Shimadzu Corp | Protein identification processing method and apparatus |
| KR100531207B1 (en) * | 2005-06-04 | 2005-11-29 | 씨비에스소프트주식회사 | Protein identification system |
| KR20070017676A (en) * | 2005-08-08 | 2007-02-13 | 한국기초과학지원연구원 | How to Analyze Sequence and Modification Information of Modified Polypeptides |
| KR100757040B1 (en) * | 2005-12-12 | 2007-09-07 | 오브젝트인터랙션테크놀로지스(주) | System and method for analyzing surface of interaction between proteins or proteins and compounds based on protein domains, and recording media therefor |
| US20070224704A1 (en) * | 2006-03-23 | 2007-09-27 | Epitome Biosystems, Inc. | Protein splice variant / isoform discrimination and quantitative measurements thereof |
Also Published As
| Publication number | Publication date |
|---|---|
| KR100856526B1 (en) | 2008-09-04 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Nesvizhskii et al. | Analysis, statistical validation and dissemination of large-scale proteomics datasets generated by tandem MS | |
| US6393367B1 (en) | Method for evaluating the quality of comparisons between experimental and theoretical mass data | |
| CN103890164B (en) | Cell recognition device and program | |
| KR101542529B1 (en) | Examination methods of the bio-marker of allele | |
| KR20150024232A (en) | Examination methods of the origin marker of resistance from drug resistance gene about disease | |
| CN113393903B (en) | Construction method of reference protein database, storage medium and electronic equipment | |
| Wacholder et al. | Detection of human unannotated microproteins by mass spectrometry-based proteomics: a community assessment | |
| CN103488913A (en) | A computational method for mapping peptides to proteins using sequencing data | |
| US20070059842A1 (en) | Mass analysis method and mass analysis apparatus | |
| EP4428864A1 (en) | Method for diagnosing cancer by using sequence frequency and size at each position of cell-free nucleic acid fragment | |
| CN112415208A (en) | Method for evaluating quality of proteomics mass spectrum data | |
| WO2009113752A1 (en) | System comprising scoring algorithm and method for identifying alternative splicing isoforms using peptide mass fingerprinting, and recording media having program therefor | |
| CN119673284A (en) | Third-generation sequencing read analysis methods, applications and devices | |
| CN110438235B (en) | A method for population origin inference based on hair shaft proteomic nsSNPs | |
| JP2012235723A (en) | Large-scale base sequence analysis method, program, and apparatus | |
| US8296300B2 (en) | Method for reconstructing protein database and a method for screening proteins by using the same method | |
| EP1820133B1 (en) | Method and system for identifying polypeptides | |
| Mueller et al. | Analysis of the experimental detection of central nervous system‐related genes in human brain and cerebrospinal fluid datasets | |
| JP4286075B2 (en) | Protein identification processing method | |
| US20050042682A1 (en) | System and method for scoring peptide mass fingerprinting | |
| Halligan et al. | Peptide identification using peptide amino acid attribute vectors | |
| KR20250056440A (en) | Method for idntifying of de novo peptides and feature vector generation method for identification of de novo peptides | |
| Song et al. | Bioinformatics methods for protein identification using peptide mass fingerprinting | |
| Jarrahi et al. | A Multi-Objective Scoring (MOS) Framework for Detecting Cross-Modal Spatial Similarity: Conceptual and Direct Formulations | |
| Gundry | Bioinformatics for mass spectrometry-based proteomics |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 08753199 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 08753199 Country of ref document: EP Kind code of ref document: A1 |

