EP1550069A1 - Verfahren zur analyse von transkriptionsvariationen in einer genmenge - Google Patents
Verfahren zur analyse von transkriptionsvariationen in einer genmengeInfo
- Publication number
- EP1550069A1 EP1550069A1 EP03756043A EP03756043A EP1550069A1 EP 1550069 A1 EP1550069 A1 EP 1550069A1 EP 03756043 A EP03756043 A EP 03756043A EP 03756043 A EP03756043 A EP 03756043A EP 1550069 A1 EP1550069 A1 EP 1550069A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- gene
- genes
- value
- variation
- calibration
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6809—Methods for determination or identification of nucleic acids involving differential detection
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B25/00—ICT specially adapted for hybridisation; ICT specially adapted for gene or protein expression
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B25/00—ICT specially adapted for hybridisation; ICT specially adapted for gene or protein expression
- G16B25/10—Gene or protein expression profiling; Expression-ratio estimation or normalisation
Definitions
- the present invention relates to the analysis of variations in m-RNA concentrations of a set of genes carried out using DNA chips.
- Each DNA molecule is made up of two complementary polynucleotide strands, an "antisense” strand (-) and a “sense” strand (+).
- Each polynucleotide strand consists of a polymeric chain of nucleotides.
- Each nucleotide consists of a phosphate, a sugar (deoxyribose) and a base, the bases possibly being a guanine (G), an adenine (A), a cytosine (C) and a thyine (T) .
- each gene When a cell is active and living, each gene synthesizes RNA-messenger molecules, or mRNA, which are base-to-base copies of the sense (+) strand of the gene. This phenomenon is called transcription or expression of the gene. More precisely, the transcription of a gene is carried out only for certain groups of consecutive bases, or sequences, of the strand of the gene which is expressed, the sense strand (+). L 1 mRNA produced by a gene is in fact a grouping of copies of sequences. Depending on the cell, not all genes are expressed in the same proportions. Thus, • the concentration of mRNA relative to a given gene can be zero, or vary between 1 and 10,000 per cell.
- a known method for measuring the concentration of mRNA is to use DNA chips.
- Cells are taken from a culture or from a human body by biopsy. The transcription activity of these cells is then stopped, for example by freezing.
- a sample is then prepared containing the mRNA extracted from a certain number of cells in solution.
- a DNA chip is also prepared, an example of which is illustrated in FIG. 1 in order to analyze a set of genes.
- each gene is analyzed by means of two sets of around twenty hybridization units.
- a hybridization unit groups together a set of identical DNA strands called probes. These DNA strands are complementary strands of a gene sequence which is found in the mRNA of the cells analyzed. These DNA strands have sequences identical to those of the antisense (-) strand of the gene.
- a first set of hybridization units, called perfect (UP) contains probes which correspond to different sequences of a gene.
- a second set of hybridization units contains probes which differ from the probes of the first set for at least one of the bases, each perfect hybridization unit being associated with an imperfect hybridization unit.
- a perfect hybridization unit 2 contains probes 3, 4, 5, 6 and 7.
- the perfect hybridization unit 2 is associated with an imperfect hybridization unit 10 which contains probes 11, 12, 13, 14 and 15 which differ by a base (A, G) compared to probes 3 to 7.
- the messenger RNAs of the previously prepared sample are "labeled", for example rendered fluorescent.
- the fluorescence of the strands is represented by a cross in a circle attached to the fluorescent strand.
- the tagged RNA-messengers are called targets.
- a washing step possibly makes it possible to dissociate the strands which are not very complementary and thus limit the number of false appearances.
- a photograph is then taken of each of the hybridization units of the DNA chip in order to determine for each hybridization unit a fluorescence intensity. After measuring the fluorescence intensities, two fluorescence intensity values iy and ⁇ J are obtained for each pair of perfect and imperfect hybridization units corresponding to a gene sequence. A fluorescence intensity is calculated for each gene sequence equal to the difference between the fluorescence intensity values i-gp and iui- This method of measuring the fluorescence intensity of each sequence makes it possible to obtain a better signal ratio on noise.
- the reference cells could be, for example, healthy liver cells and the test cells, diseased liver cells.
- the same DNA chip models are used, and in both cases the sequence of operations described above is carried out.
- the study of variations in the concentration of m-RNA for each gene makes it possible to identify which genes have the concentration of m-RNA changed, following a modification of the transcription activity, or a change in the lifespan of mRNAs.
- the lifespan of mRNA fluctuates among other things as a function of more or less significant protein synthesis activity.
- the analysis of variations in mRNA concentrations for each of the genes is carried out by calculating the ratio of the mRNA concentrations of the same gene.
- This method is known as the "fold change" method.
- the change in m-RNA concentration is considered to be significant when the ratio of RN-m concentrations is above a predetermined threshold. This threshold is identical for all of the genes and this method therefore does not allow the specificity of each of them to be taken into account.
- the processes of creation and destruction of m-RNA are interrupted randomly during the collection of cells and the concentration of m-RNA may fluctuate slightly from one cell to another. In the case where a gene produces on average 10 mRNA in each cell, a difference of only one
- MRNA between two cells leads to a ratio of 1.1, or 10% difference, and the gene in question will be considered to have a significant difference in mRNA concentration.
- a difference of 10 mRNA leads to a ratio of 1.01, or 1% difference, and this will go unnoticed when it can be completely abnormal.
- the concentration of m-RNA relative to a gene can naturally vary in its own proportions. With a simple fold change analysis, it is impossible to know to what extent the variation in the concentration of m-RNA relative to a gene remains or not within acceptable proportions.
- One way of knowing the range of natural variation of the mRNA concentration relative to a gene, or more precisely the cumulative distribution of frequencies, would be to carry out a large number of mRNA concentration measurements, for each gene. from identical reference cells. In the case where 100 measurements have been made for each gene, it is possible to define threshold values corresponding to probabilities in increments of 0.01 so that the same discomfort associated with identical cells has a higher concentration of mRNA at these threshold values.
- Another object of the present invention is to provide such a method which makes it possible to define a threshold value very precisely.
- the present invention provides a method for analyzing variations in concentrations of messenger RNAs obtained by transcription of a set of genes comprising the following steps: a) measuring the concentration of messenger RNAs for each of the genes in so-called reference cells and report the results on a reference list (L re f); b) measure the concentration of messenger RNA for each of the genes in so-called test cells and report the results on a test list (L ⁇ est) • 'c) calculate for each gene a variation value (Var j ) , k being an integer between 1 and n, which is a measure of the difference between the mRNA concentrations of said gene between the reference list (L re f) and the test list
- the step of identifying the genes consists in selecting the genes whose normalized variation value is greater than a determined threshold value (Z se ⁇ j _] _).
- the determination of the threshold value (Z seu j_) comprises the following steps: h) measuring the concentration of m-RNA for each of the genes of two identical groups of so-called calibration cells and report the respective results on first CL> etal l) and second (Iié al 2 ⁇ calibration lists; i) calculate for each gene a variation value (Vargtal k) according to the method of steps c ) to e) from the first (L e t a li) and second ⁇ I * stall 2) calibration lists; j) calculating for each gene a normalized calibration variation value (Z re fj according to the method of step f); k) construct the cumulative frequency distribution, called calibration, of the normalized calibration variation values associating with any calibration variation value normalized
- n (number of genes for which Z> Z threshold) where n is the number of genes considered.
- the step of identifying the genes consists in selecting the genes whose normalized variation value is greater than a first threshold value for the genes of the first group and greater than a second threshold value for the genes of the second group.
- the determination of the first and second threshold values consists in choosing first and second probabilities of selection error desired respectively for the first and second groups and in defining the first and second corresponding threshold values using the cumulative distribution of calibration frequencies.
- the choice of the first and second threshold values consists in carrying out the method of claim 4 successively for the first and the second group.
- the variation value Var ⁇ of a gene is equal to the difference between the concentrations of m-RNA of said gene for different cells.
- the value of variation Var ⁇ - of a gene is equal to the ratio of the concentrations of m-RNA of said gene for different cells.
- the method comprises for each list the following steps:
- the variation value of a gene is equal to the difference between the ranks of the gene for the two lists analyzed.
- the normalized variation value Z of each gene is obtained according to the following formula: Var - ⁇ (g)
- the normalized variation value is calculated according to the following steps: - assign a unique rank value r to each gene equal to the rank value of the reference list for the genes of the first group and equal to the rank value of the test list for genes of the second group.
- the method aims to analyze the variations in m-RNA concentrations of a set of genes from m identical groups of so-called reference cells (GR ⁇ to G ⁇ and q groups identical to so-called test cells (GT ] _ to GTg), the method comprising the following steps:
- the first and second calibration groups (GRétal i and Ggtal _X) are identical whatever the combination of groups considered.
- the determination of the threshold grouping value (Rseuil) comprises the following steps:
- the step of selecting a probability of selection of grouping error comprises the steps of: - defining the maximum acceptable rate of false positive for one identification genes;
- the grouping method comprises the following steps:
- the method aims to analyze the variations in mRNA concentrations of a set of genes from m identical groups of so-called reference cells (GR ⁇ _ to GR j ⁇ ) and q identical groups of so-called test cells (GT ] _ to GTg), the method comprising the following steps:
- one or more reference, test or calibration lists are obtained according to a method of creating an artificial data set comprising the following steps:
- FIG. 1 represents a chip DNA
- FIG. 2 is a representation of variation values of m-RNA concentration relating to a set of genes used according to a first step of the invention
- FIG. 3 is a representation of normalized mRNA concentration variation values relating to a set of genes used according to a second step of the invention
- FIG. 1 represents a chip DNA
- FIG. 2 is a representation of variation values of m-RNA concentration relating to a set of genes used according to a first step of the invention
- FIG. 3 is a representation of normalized mRNA concentration variation values relating to a set of genes used according to a second step of the invention
- FIG. 1 represents a chip DNA
- FIG. 2 is a representation of variation values of m-RNA concentration relating to a set of genes used according to a first step of the invention
- FIG. 3 is a representation of normalized mRNA concentration variation values relating to a set of genes used according to a second step of the invention
- FIG. 1 represents a chip DNA
- FIG. 4A represents a cumulative frequency distribution of RN-m concentration variation values for a first set of genes
- Figure 4B shows a cumulative frequency distribution of mRNA concentration variation values for a second set of genes
- FIG. 4C is a "quantile versus quantile" curve of the variation values of m-RNA concentrations of the first and second sets of genes
- FIG. 5A represents a set of "quantile against quantile” curves of non-normalized variation values obtained according to a "fold change”method
- FIG. 5B represents a set of "quantile against quantile” curves of non-normalized variation values obtained according to a row shift method
- FIG. 6A represents a set of curves
- FIG. 6B represents a set of "quantile against quantile” curves of normalized variation values obtained according to a row shift method.
- the method of analysis of the present invention provides for using DNA chips to analyze a set of n genes and to study the variations in m-RNA concentrations between reference cells and test cells.
- an analysis of the variations between a group of test cells and a group of reference cells will be described.
- the method according to the invention will be generalized to the analysis of several groups of test and reference cells.
- the method of analysis of the present invention provides for using DNA chips to analyze a set of n genes and to study the variations in m-RNA concentrations between a group of reference cells and a group of test cells.
- concentration of mRNA Ck relative to each gk gene is measured beforehand and the values are reported on reference lists L re f and test £ eS .
- the method of analysis begins with the calculation for each of the genes of a value of variation of mRNA concentration, or value of variation Var, which can be equal to the difference of the concentrations of mRNA of each gene between the reference and test groups ref or c k test and Ck ref respectively the mRNA concentrations of the gk gene on the test and reference lists) or also equal to the ratio of the mRNA concentrations (Va ⁇ ⁇ Ck test / c k ref) ⁇ • which corresponds to the method "fold change" described above.
- the genes are classified in ascending order of their mRNA concentrations for each of the reference and test lists.
- a value of zero rank is then assigned to all the genes whose mRNA concentration is equal to zero or more broadly to all the genes whose mRNA concentration is less than a threshold concentration value corresponding to a estimation of measurement noise.
- Each of the ni other genes is then assigned a unique rank value, the rank value being between 1 and ni.
- the set of rank values forms a continuous series of integers between 0 and ni. The higher the rank of a gene, the higher its mRNA concentration.
- variations in the method of measuring the concentration of mRNA from DNA chips results in a greater or lesser variation in the values of RNA concentration.
- Two identical groups of cells can have concentration values varying between 10 and 10,000 for the first group and between 50 and 11,000 for the second group.
- Vark The variation value, Vark, of each gk gene is calculated as follows: Var k ⁇ r test , k - r re f, k ( D where r ⁇ st k and r ref k are respectively the ranks of the gk gene from the lists of test and reference.
- FIG. 2 represents a set of positive Vrk variation values calculated according to the "row shift” method.
- the rows are indicated on the abscissa.
- the variations are indicated on the ordinate.
- Each variation value of a gene is represented by a cross whose abscissa corresponds to the rank of this gene for the reference list. Although this is not visible in Figure 2 'because of the large number of genes considered, each value of x-axis (row) corresponds to a single gene, and thus to a single value of variation.
- the present invention provides for defining a threshold variation value which is a function of the rank of the discomfort. More particularly, the analysis method of the present invention includes a normalization method. Genes are classified into two groups. The genes whose variation value indicates an increase in their mRNA concentrations between the reference list and the test list are placed in a first group. The others " are put in a second group and a new variation value is calculated for these genes by inverting the test and reference lists.
- the genes of the second group are the n ne g genes whose variation is strictly negative ( ⁇ is k ⁇ r ref k For a g gene).
- V ⁇ - the variation value V ⁇ - equal to the opposite of the initial value. All variation values are now positive.
- the variation values of the genes exhibiting a decrease in their concentration (value less than 1) between the reference group and the test group are replaced by 1 inverse of the initial values.
- variation values are therefore all greater than 1.
- a set of neighboring rows, or else "window" of rows is selected for each gene gk of rank ⁇ .
- a normalized variation value Zk is calculated for each of the gk genes according to the following formula: z Vark ⁇ ⁇ (9k) ⁇ ( g)
- the normalization process is carried out separately for each of the first and second groups of genes.
- the values ⁇ (gk) and ⁇ (g) are calculated for each group from the variation values of a set of genes from the same group.
- FIG. 3 represents the set of normalized variation values Z obtained for each of the variation values Vark ⁇ e l in FIG. 2.
- the abscissa designates the rows and a value of abscissa corresponds to a single value of normalized variation.
- the curves 30 and 31 correspond respectively to the local means and to the local standard deviations, not smoothed, calculated from the Z values in the same way as that had been done previously from the Vark values, and described above. Curves 30 and 31 show that the local means and the local standard deviations are now substantially constant whatever the rank, which means that genes with different mean mRNA concentrations have normalized variation values that follow the same cumulative frequency distribution.
- any normalization method can be used such that the cumulative frequency distribution of a subset of normalized variation values corresponding to genes in the same row window is substantially identical regardless of the subset. considered.
- a threshold value seu j_ ⁇ possibly different for the first and the second group of genes, and selecting the genes whose standardized variation value exceeds the threshold value.
- this threshold value is identical for all the genes and the selection criterion is homogeneous whatever the rank of the genes analyzed, that is to say regardless of their concentration of RNA- m average.
- An advantage of the analysis method according to the present invention is that it makes it possible to identify genes exhibiting a significant variation in their mRNA concentrations from a limited number of measurements.
- the present invention also proposes to define a threshold value according to the method below.
- a calibration step is carried out which consists in determining the variations in the normal RN-m concentrations of each of the genes by studying two groups of identical cells called calibration, the concentration of m-RNA of each gene being plotted on two calibration lists Lg al 1 and L êtal 2 •
- a calculation of normalized calibration variation values is carried out according to the row offset method and the normalization method previously described.
- One of the two calibration lists Lg ⁇ al 1 and L étal 2 is considered as a test list and the other as a reference list.
- local averages are smoothed used for the calculation of Zetal k- ® n obtains two calibration curves representing the mean ⁇ etal ( r ) and the standard deviation r ⁇ tal ⁇ ) of the variations in calibration as a function of rank, any reference to a given gene being deleted .
- the normalized variation values Z are calculated from these calibration curves according to the formula:
- the groups of calibration cells can be reference cells, test cells or other cells deemed suitable.
- the choice of cells used is dictated by the effect of the ⁇ êt values (r) and ⁇ stall (r) are normalized on variation values Z - These are even smaller than the mean values and standard type are great.
- the values ⁇ etal ( r ) and ⁇ etal ( r ) depend on the one hand on the reproducibility of the experimental conditions (DNA chips not perfectly identical) and on the other hand on the stability of the biological system of the chosen cells.
- the experimental conditions are assumed reproducible biological system ⁇ étal present values (r) and ⁇ stall (r) all the greater that it is unstable.
- the calibration curves are constructed independently for each of the pairs, which leads to two pairs of calibration curves ( ⁇ test ' ⁇ test) and ⁇ ref' ⁇ ref) • 0n then evaluates which of the two systems is more unstable ( ⁇ or / and ⁇ higher).
- This assessment can be done in different ways.
- the results of the analysis method of the present invention are better if the calibration curves constructed from the most unstable system are used.
- a cumulative distribution of calibration frequencies is constructed from all the normalized variation values. Normalized variation values for all genes, regardless of their rank, follow this cumulative distribution of calibration frequencies. Indeed, as will be established more precisely in relation to FIG. 6B, any subset of normalized calibration variation values corresponding to genes of the same row window follows the same cumulative distribution of frequencies and it is therefore possible to construct a single cumulative distribution of frequencies from all the normalized calibration variation values. Given the large number of genes studied and therefore the large number of normalized calibration variation values obtained, the cumulative distribution of resulting calibration frequencies is very precise. From this cumulative distribution calibration frequencies, is associated with all normalized calibration variation value z Stall, k Probability, called p selection error probability is uil k 'for that there are values of normalized calibration variation naturally greater than the latter.
- the probability of error can now be defined using the cumulative distribution of calibration frequencies.
- p is uil selection corresponding to the probability that it exists naturally standard variation values greater than the threshold value seu ⁇ ; L chosen to select genes.
- An advantage of the analysis method according to the present invention is that it makes it possible to associate a probability of selection error with any threshold value Z seu j_ ⁇ _ chosen.
- Another advantage of the analysis method according to the present invention is that it allows to choose a threshold value seu j_. very precise with a limited number of measurements.
- a first and a second false positive rate are defined.
- n the number of genes of the first group np OS or of the second group n ne g, the threshold Pseuil / z values possibly being different for each group of genes.
- the cumulative frequency distribution of the normalized variation values Zk obtained during the comparison between test and reference cells is constructed beforehand. From this distribution, it is possible to associate with any normalized variation value k a probability, called probability of observation Pobs k 'so that normalized variation values greater than the latter are observed.
- the false positive rate can be defined as being equal to Pseuil k / Pobs k-
- Pseuil / Z threshold 'l has sensitivity, equal to (Pobs k "Pseuil k) / F 'makes it possible to know if among the selected genes, the number of genes actually showing significant variations is representative of the number of genes whose variation values have increased (Vark> ariai) •
- An advantage of the analysis method according to the present invention is that it allows to associate a false positive rate and a sensitivity value of any threshold value seu i] _ and therefore to any Pseuil selection error probability value chosen.
- FIGS. 4A to 4C illustrate the construction of a "quantile against quantile" curve.
- FIG. 4A represents a cumulative distribution of frequencies C ⁇ of a first subset of variation values taken from the set of variation values (Var) obtained during a comparative study. The variation values are plotted on the abscissa. We indicate on the ordinate the probability (proba) so that there are variation values lower than the variation value on the abscissa.
- FIG. 4B is another cumulative distribution of frequencies C2 of a second set of variation values taken from the set of variation values of the comparative study.
- FIG. 4C is a "quantile against quantile" curve C3 obtained from curves C1 and C2 in FIGS. 4A and 4B.
- the variation values of the first studied set are represented on the ordinate, and the variation values of the second studied set are represented on the abscissa.
- “quantile against quantile” is obtained by taking for each probability value (between 0 and 1) the corresponding variation values on the curves C1 and C2 and by defining a point having these two values respectively for ordinate and abscissa.
- the point 40 of the curve C3 has the abscissa VI 'and the ordinate VI, VI and VI' being respectively the values of variation of the curves Cl and C2 corresponding to the probability 0.1.
- the points 41 and 42 of the curve C3 have the respective abscissa V2 'and V3' and for the ordinate V2 and V3, the variation values V2, V3 of the curve C ⁇ _ and
- V2 ', V3' of curve C2 having respective probabilities 0, 5 and 0.9.
- a “quantile against quantile” curve is thus obtained for two subsets of variation values.
- the curve C3 is relatively far from the diagonal drawn in dotted lines, which means that the first and second subsets of variation values have different distribution functions.
- FIG. 5A represents a set of "quantile against quantile" curves obtained by studying different subsets of variation values calculated according to a Fold Change method. The most flattened curves are obtained by taking subsets of variation values whose respective ranks are very far apart. This demonstrates that genes with different ranks have variation values that follow different distribution functions.
- FIG. 5B likewise represents a set of "quantile against quantile" curves obtained by studying different subsets of non-normalized variation values calculated according to a row shift function. We can also observe a difference between the distribution functions for genes with very distant ranks.
- FIG. 6A represents a set of "quantile against quantile" curves obtained by studying different subsets of normalized variation values calculated according to the Fold Change function and the normalization method of the present invention.
- the curves approach the diagonal which means that genes with different ranks have normalized variation values which follow relatively similar distribution functions. However, there are relatively large divergences for the values corresponding to high probabilities.
- FIG. 6B represents a set of "quantile against quantile" curves obtained by studying different subsets of normalized variation values calculated according to the row shift method and the normalization method of the present invention.
- the curves are all very close to the diagonal, which means that the set of normalized variation values follows the same cumulative frequency distribution. This demonstrates that, by combining a calculation of the variation values according to the row shift method of the invention and a normalization of these values according to the normalization method of the invention, a set of normalized variation values is obtained which follow the same cumulative distribution of reference frequencies.
- a method of multiple analysis aims to identify more precisely which genes exhibit the most significant variations in mRNA concentrations.
- the multiple analysis method includes multiple analyzes of variation between reference and test lists. For all or "part of the combinations C i j comprising a reference group GR ⁇ and a test group GT is calculated for each gene gk, an amount of change Var ⁇ j ⁇ according to the offset method of ranks and an amount of change normalized Zj_ jk according to the normalization process of the invention.
- a calibration step identical to that described above is carried out. After selecting two GR ⁇ calibration groups have] _ and ⁇ al GR 2 among the m reference groups, is calculated for each gene g a normalized calibration variation value Zgtal k by means of the offset method of rows and of the standardization process of the invention. A cumulative distribution of calibration frequencies is constructed from all the variation values calibration standards. It is thus possible to associate with a normalized value of variation of calibration Zg ⁇ al k a probability, called probability of error of calibration Pétai k 'so that there exist values of normalized variation naturally higher than this last.
- a cumulative distribution of grouping frequencies is constructed for each combination C- ⁇ j chosen from two reference groups one of which is the GRi group or two test groups whose one of them is the group GT-j of the combination Cj_ considered.
- a probability is defined for each gene gk, called the probability of error Pi k / corresponding to the normalized variation value Z j k of said gene.
- the error probabilities Pi j, k are all equal.
- some of the probabilities Pi -1 k correspond to positive variations and other values Pk i correspond to negative variations.
- the product Prodpp OS of the values Pi, j, k corresponding to positive variations is compared to the product Prodp n gg of the values Pi, j, k corresponding to negative values.
- Prodp OS is less than Prod n £ g we consider that the variation of the gene is positive and all the probabilities Pi, k corresponding to negative variations take the value 1 (conversely if Prodp OS > Prod n gg, the variation of the discomfort is considered negative and all the probabilities Pi H k take the value 1).
- the result is homogeneous, i.e. the variation of the k gene is considered to be positive (or negative) for all combinations. If for a minority of sets the assignment procedure has resulted in giving the gk gene a sense of opposite variation, this is explained by the presence of an abnormal variation called artefactual which is easily detectable. These values are eliminated, which leads to a correct reassignment of the direction of variation.
- a grouping value Rk is calculated for each gene g from the gene error probabilities according to a grouping method.
- a grouping value Rk is calculated for each gene gk worth RETAL calibration combination, k using the calibration petai error probabilities, i, j, k corresponding to the normalized variation values Zêtal, i, k each gene obtained from the cumulative frequency distributions previously calculated.
- the combinations chosen are distributed in different sets. We could for example constitute sets of independent combinations, two combinations Ci ⁇ ji and C 2 r j2 being independent if the groups GR ⁇ and GR2 are different and if the groups GTji and
- G j2 are different.
- computed for each 'discomfort gk Rk a grouping value by taking the average of the intermediate values of each set.
- a threshold grouping value R S euil is defined in order to select the genes having grouping values greater than the latter.
- grouping frequencies a cumulative distribution of frequencies, called grouping frequencies, from all the calibration grouping values.
- Pthéo k 'so a probability of group selection error
- P2seuil any threshold grouping value Rgeuil chosen.
- R S and Pthr euil be chosen according to the false positive rate and the desired sensitivity.
- the method of multiple analysis by analysis of means consists in constructing for the groups G] _ to GR j n and GT ⁇ _ to GTg a single group GR and GT.
- the concentration values of mRNA-m of the groups GR ⁇ to GR j n and GT X to GTq are expressed in the form of rank values, normalized on a scale of 0 to 100, as described in chapter 1.
- the cumulative distribution of frequencies of the variations of transcription signal normalized for a biological system makes it possible to construct artificial data sets, in the form of an artificial list ar t associating with each gene a concentration value, the data set having the same statistical characteristics as the actual data used for the calibration. From two identical groups of Gl cells and
- rj eUf k consists in successively calculating, starting from the value immediately below rk, the absolute value of ⁇ r for any value r game, k less than r and taking the rank r game for new rank.
- the new set of values thus obtained can be easily transformed into mRNA concentration values by the reverse transformation of that which gives the rank.
- concentration of mRNA of each gene being reported on the artificial list L ar t.
- a multiple method according to the present invention plans to identify more precisely the genes exhibiting the most significant transcription variations.
- the groups GC1 to GCn can represent measurements carried out on the same biological system but at different and increasing times (kinetics experiment), or subjected to a stimulus of strictly increasing or decreasing intensity (dose / response experiments).
- the common characteristic of these two types of experiment is that it is sought for each gk gene whether there has been a significant variation in transcription signal over the entire interval of the independent variable VI (time in kinetics or dose of a product in the case of a dose / response).
- one of the analyzes will relate to the GCO and GC1 groups, another to the GC1 and GC2 groups, and the last will relate to the GCn-1 and GCn groups.
- the Pthéor k is determined (° u l es Pthr k s' ⁇ there is only one group) and p s k OD 0n selects genes that have undergone an RNA concentration variation -m significant using the selection parameters such as the probability of grouping selection error, the false positive rate or the sensitivity.
- the list s sel k is completed as follows: if a significant variation has been detected between the values i and i + 2 of VI, and if the positions i and i + 1 were at zero in the previous step, then we changes positions i and i + 1 to one. If one of the positions were already at one, the new result is not considered significant with regard to the second position.
- the new suite for k could be s Sel k ⁇ 1 '1' 0 '1' 1 '1 -' 0 '0 -
- the present invention is susceptible of various variants and modifications which will appear to one skilled in the art.
- the method of the present invention can be applied to the analysis of variations in the number of different proteins present in living cells.
- the analysis method of the present invention can be implemented from the concentrations of m-RNA noted for each of the gene sequences studied corresponding to a hybridization unit of the DNA chip used. We will therefore not study the variations in the concentration of mRNA relating to a gene but that relating to a given sequence.
- a different definition of variation values can be used.
- other normalization methods can be provided which satisfy the requirement of uniformity of the cumulative frequency distributions of any subset of normalized variation values.
- those skilled in the art will be able to define the optimal grouping process making it possible to identify the genes having the most significant values of variation in mRNA concentrations.
Landscapes
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Physics & Mathematics (AREA)
- Chemical & Material Sciences (AREA)
- Genetics & Genomics (AREA)
- Engineering & Computer Science (AREA)
- Biophysics (AREA)
- Molecular Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Organic Chemistry (AREA)
- Biotechnology (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Zoology (AREA)
- Theoretical Computer Science (AREA)
- Evolutionary Biology (AREA)
- Analytical Chemistry (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Wood Science & Technology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Microbiology (AREA)
- Immunology (AREA)
- Biochemistry (AREA)
- General Engineering & Computer Science (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FR0206749 | 2002-05-31 | ||
| FR0206749A FR2840323B1 (fr) | 2002-05-31 | 2002-05-31 | Methode d'analyse des variations de transcription d'un ensemble de genes |
| PCT/FR2003/001655 WO2003102849A1 (fr) | 2002-05-31 | 2003-06-02 | Methode d'analyse des variations de transcription d'un ensemble de genes |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP1550069A1 true EP1550069A1 (de) | 2005-07-06 |
Family
ID=29558893
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP03756043A Withdrawn EP1550069A1 (de) | 2002-05-31 | 2003-06-02 | Verfahren zur analyse von transkriptionsvariationen in einer genmenge |
Country Status (5)
| Country | Link |
|---|---|
| US (1) | US20050255471A1 (de) |
| EP (1) | EP1550069A1 (de) |
| AU (1) | AU2003255623A1 (de) |
| FR (1) | FR2840323B1 (de) |
| WO (1) | WO2003102849A1 (de) |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP0880598A4 (de) * | 1996-01-23 | 2005-02-23 | Affymetrix Inc | Verfahren zur analyse von nukleinsäure |
| KR20010052341A (ko) * | 1998-05-12 | 2001-06-25 | 로제타 인파마틱스 인코포레이티드 | 유전자 발현 분석을 위한 정량 방법, 시스템, 장치 |
-
2002
- 2002-05-31 FR FR0206749A patent/FR2840323B1/fr not_active Expired - Fee Related
-
2003
- 2003-06-02 WO PCT/FR2003/001655 patent/WO2003102849A1/fr not_active Ceased
- 2003-06-02 US US10/516,278 patent/US20050255471A1/en not_active Abandoned
- 2003-06-02 AU AU2003255623A patent/AU2003255623A1/en not_active Abandoned
- 2003-06-02 EP EP03756043A patent/EP1550069A1/de not_active Withdrawn
Non-Patent Citations (1)
| Title |
|---|
| See references of WO03102849A1 * |
Also Published As
| Publication number | Publication date |
|---|---|
| FR2840323B1 (fr) | 2006-07-07 |
| FR2840323A1 (fr) | 2003-12-05 |
| WO2003102849A1 (fr) | 2003-12-11 |
| US20050255471A1 (en) | 2005-11-17 |
| WO2003102849A9 (fr) | 2004-04-22 |
| AU2003255623A1 (en) | 2003-12-19 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| FR3116064A1 (fr) | Procede de conception de sequences d’adn base sur l’algorithme d’optimisation chaotique des baleines | |
| US20230259588A1 (en) | Inter-cluster intensity variation correction and base calling | |
| EP0552575B1 (de) | Mehrfachteilungsegmentierungsverfahren | |
| Garry et al. | Bayesian counting of photobleaching steps with physical priors | |
| Pigani et al. | Classification of red wines by chemometric analysis of voltammetric signals from PEDOT-modified electrodes | |
| WO2004001673A2 (fr) | Procede d'analyse d'image pour la mesure du signal sur des biopuces | |
| EP1550069A1 (de) | Verfahren zur analyse von transkriptionsvariationen in einer genmenge | |
| JP2010512777A (ja) | 差分解析により取得されるトランスクリプトーム実験の結果を処理するための補正方法 | |
| EP4002223A1 (de) | Verfahren zur aktualisierung eines künstlichen neuronalen netzes | |
| EP3405899B1 (de) | Verfahren zur klassifizierung von einer biologischen probe | |
| FR3143174A1 (fr) | Procede d’identification de biomarqueurs d’arn extracellulaires dans des echantillons de fluide corporel en vue de detecter des cancers et d’autres pathologies telles que l’endometriose | |
| EP2952888B1 (de) | Grössenmarkierer und verfahren zur steuerung der auflösung eines elektropherogramms | |
| EP3227813B1 (de) | Verfahren zur kalkulation der sonden-ziel-affinität von einem dna-chip sowie verfahren zur herstellung eines dna-chips | |
| FR2861406A1 (fr) | Methode d'analyse d'un ensemble de genes | |
| FR3155300A1 (fr) | Procédé de calcul de paramètres physiques d’une chaussée à partir de mesures de déflexion | |
| EP4300129B1 (de) | Verfahren zur gruppierung von wellenformbeschreibungen | |
| WO2020242603A1 (en) | Methods and usage for quantitative evaluation of clonal amplified products and sequencing qualities | |
| JP2006170670A (ja) | 遺伝子発現量規格化方法、プログラム、並びにシステム | |
| EP4721065A1 (de) | Verfahren und system zur probabilistischen typisierung mikrobieller stämme | |
| FR3152885A1 (fr) | Procédé de génération de rapports d'analyse protéomique comparatifs à partir de multiples échantillons biologiques mis en œuvre par ordinateur | |
| EP4471785A1 (de) | Verfahren und system zur probabilistischen typisierung mikrobieller stämme | |
| EP4405666A1 (de) | Verfahren zur bestimmung der qualität einer bei der hefeherstellung verwendeten melasse | |
| Lun et al. | Package ‘scran’ | |
| HK40090027B (en) | Systems and methods for per-cluster intensity correction and base calling | |
| FR3149977A1 (fr) | Procédé d’analyse métabolomique d’échantillons biologiques par intégration des paramètres dynamiques de la RMN |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 20041201 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LI LU MC NL PT RO SE SI SK TR |
|
| AX | Request for extension of the european patent |
Extension state: AL LT LV MK |
|
| DAX | Request for extension of the european patent (deleted) | ||
| 17Q | First examination report despatched |
Effective date: 20061019 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN |
|
| 18D | Application deemed to be withdrawn |
Effective date: 20100105 |