EP1141878A2 - Normalization, scaling, and difference finding among data sets - Google Patents
Normalization, scaling, and difference finding among data setsInfo
- Publication number
- EP1141878A2 EP1141878A2 EP00903107A EP00903107A EP1141878A2 EP 1141878 A2 EP1141878 A2 EP 1141878A2 EP 00903107 A EP00903107 A EP 00903107A EP 00903107 A EP00903107 A EP 00903107A EP 1141878 A2 EP1141878 A2 EP 1141878A2
- Authority
- EP
- European Patent Office
- Prior art keywords
- data set
- calculation
- scaling
- representation
- display means
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
- 238000010606 normalization Methods 0.000 title claims abstract description 29
- 238000004364 calculation method Methods 0.000 claims abstract description 113
- 238000000034 method Methods 0.000 claims abstract description 99
- 230000006870 function Effects 0.000 claims description 45
- 238000012935 Averaging Methods 0.000 claims description 42
- 230000009466 transformation Effects 0.000 claims description 34
- 150000007523 nucleic acids Chemical class 0.000 claims description 27
- 108020004707 nucleic acids Proteins 0.000 claims description 26
- 102000039446 nucleic acids Human genes 0.000 claims description 26
- 238000005284 basis set Methods 0.000 claims description 24
- 238000004422 calculation algorithm Methods 0.000 claims description 21
- 238000004458 analytical method Methods 0.000 claims description 19
- 238000009396 hybridization Methods 0.000 claims description 16
- 238000001962 electrophoresis Methods 0.000 claims description 13
- 239000003153 chemical reaction reagent Substances 0.000 claims description 11
- 230000000873 masking effect Effects 0.000 claims description 7
- 238000005457 optimization Methods 0.000 claims description 7
- 230000008569 process Effects 0.000 claims description 7
- 239000000203 mixture Substances 0.000 claims description 5
- 238000002474 experimental method Methods 0.000 description 28
- 230000014509 gene expression Effects 0.000 description 23
- 239000012634 fragment Substances 0.000 description 16
- 239000000523 sample Substances 0.000 description 15
- 238000001514 detection method Methods 0.000 description 12
- 230000000694 effects Effects 0.000 description 11
- 108090000623 proteins and genes Proteins 0.000 description 10
- XLYOFNOQVPJJNP-UHFFFAOYSA-N water Chemical compound O XLYOFNOQVPJJNP-UHFFFAOYSA-N 0.000 description 7
- 238000005259 measurement Methods 0.000 description 6
- 239000008223 sterile water Substances 0.000 description 6
- 241001465754 Metazoa Species 0.000 description 5
- 241000700159 Rattus Species 0.000 description 5
- 239000002299 complementary DNA Substances 0.000 description 5
- DDBREPKUVSBGFI-UHFFFAOYSA-N phenobarbital Chemical compound C=1C=CC=CC=1C1(CC)C(=O)NC(=O)NC1=O DDBREPKUVSBGFI-UHFFFAOYSA-N 0.000 description 5
- 238000000692 Student's t-test Methods 0.000 description 4
- 108020004999 messenger RNA Proteins 0.000 description 4
- 239000002773 nucleotide Substances 0.000 description 4
- 125000003729 nucleotide group Chemical group 0.000 description 4
- 229960002695 phenobarbital Drugs 0.000 description 4
- 238000012353 t test Methods 0.000 description 4
- 108090000790 Enzymes Proteins 0.000 description 3
- 102000004190 Enzymes Human genes 0.000 description 3
- 238000001134 F-test Methods 0.000 description 3
- 238000003491 array Methods 0.000 description 3
- 239000003814 drug Substances 0.000 description 3
- 238000003860 storage Methods 0.000 description 3
- 238000011282 treatment Methods 0.000 description 3
- 210000004556 brain Anatomy 0.000 description 2
- 238000012937 correction Methods 0.000 description 2
- 238000010586 diagram Methods 0.000 description 2
- 230000009274 differential gene expression Effects 0.000 description 2
- 229940079593 drug Drugs 0.000 description 2
- 108091008146 restriction endonucleases Proteins 0.000 description 2
- 230000000717 retained effect Effects 0.000 description 2
- 229920006395 saturated elastomer Polymers 0.000 description 2
- 238000001228 spectrum Methods 0.000 description 2
- 108091093088 Amplicon Proteins 0.000 description 1
- 238000007476 Maximum Likelihood Methods 0.000 description 1
- 230000004075 alteration Effects 0.000 description 1
- 230000003321 amplification Effects 0.000 description 1
- 238000013459 approach Methods 0.000 description 1
- 230000008901 benefit Effects 0.000 description 1
- 239000012472 biological sample Substances 0.000 description 1
- 230000015572 biosynthetic process Effects 0.000 description 1
- 239000012496 blank sample Substances 0.000 description 1
- 239000003054 catalyst Substances 0.000 description 1
- 230000033077 cellular process Effects 0.000 description 1
- 230000008859 change Effects 0.000 description 1
- 238000006243 chemical reaction Methods 0.000 description 1
- 238000004587 chromatography analysis Methods 0.000 description 1
- 239000013068 control sample Substances 0.000 description 1
- 238000007405 data analysis Methods 0.000 description 1
- 230000003247 decreasing effect Effects 0.000 description 1
- 230000001627 detrimental effect Effects 0.000 description 1
- 238000006073 displacement reaction Methods 0.000 description 1
- 238000009826 distribution Methods 0.000 description 1
- 239000003596 drug target Substances 0.000 description 1
- 238000010828 elution Methods 0.000 description 1
- 230000009088 enzymatic function Effects 0.000 description 1
- 238000010195 expression analysis Methods 0.000 description 1
- 238000003304 gavage Methods 0.000 description 1
- 230000007614 genetic variation Effects 0.000 description 1
- 238000011835 investigation Methods 0.000 description 1
- 238000012417 linear regression Methods 0.000 description 1
- 239000000463 material Substances 0.000 description 1
- 230000037323 metabolic rate Effects 0.000 description 1
- 238000012986 modification Methods 0.000 description 1
- 230000004048 modification Effects 0.000 description 1
- 230000007935 neutral effect Effects 0.000 description 1
- 238000003199 nucleic acid amplification method Methods 0.000 description 1
- 230000010363 phase shift Effects 0.000 description 1
- 238000002360 preparation method Methods 0.000 description 1
- 238000012545 processing Methods 0.000 description 1
- 102000004169 proteins and genes Human genes 0.000 description 1
- 230000001105 regulatory effect Effects 0.000 description 1
- 230000003252 repetitive effect Effects 0.000 description 1
- 230000004044 response Effects 0.000 description 1
- 238000005070 sampling Methods 0.000 description 1
- 238000010187 selection method Methods 0.000 description 1
- 241000894007 species Species 0.000 description 1
- 238000013222 sprague-dawley male rat Methods 0.000 description 1
- 238000012453 sprague-dawley rat model Methods 0.000 description 1
- 238000007619 statistical method Methods 0.000 description 1
- 239000000126 substance Substances 0.000 description 1
- 238000006467 substitution reaction Methods 0.000 description 1
- 238000003786 synthesis reaction Methods 0.000 description 1
- 230000001225 therapeutic effect Effects 0.000 description 1
- 238000000844 transformation Methods 0.000 description 1
- 230000001131 transforming effect Effects 0.000 description 1
- 230000000007 visual effect Effects 0.000 description 1
- 238000011179 visual inspection Methods 0.000 description 1
- 230000001755 vocal effect Effects 0.000 description 1
Definitions
- This invention relates to statistical analysis of differences between at least two data sets.
- Measurements of the expression levels of individual genes within a cell provide a wealth of information about cellular processes. This is done by extracting messenger R A molecules (mRNA) from a cell, possibly converting these to more stable cDNA molecules, and measuring the concentrations of each individual species by methods such as differential display or hybridization.
- mRNA messenger R A molecules
- a typical analysis strategy is to identify genes whose expression levels differ between particular biological states.
- One difficulty in performing such an analysis is that experimental measurements of expression levels include variation due to noise. Distinguishing the true differences from the false differences (those due simply to noise) has presented a challenge for gene expression analysis.
- the relevance of this problem is that genes that are differentially regulated can be converted to commercial products, including protein therapeutics, antibody targets, therapeutic markers, as well as conventional drug targets.
- One of the major sources of noise in such experiments is that the amount of material analyzed, such as mRNA or cDNA, can differ from experiment to experiment, or among the replicates of a single experiment.
- Data analysis strategies typically account for this overall variation by performing a global scaling of all the measurements from such experiments. For example, if sample A has twice the overall cDNA concentration of sample B, then the expression level for a gene in sample B must be doubled before comparison with sample A. Often, however, such an overall scaling is not sufficient to discriminate between true differences and those that can be attributed to noise.
- One source of difficulty is identifying the particular features in the data set or sets that can be used as scaling landmarks. It is not always possible to identify a priori such unchanging features ahead of time
- the amount of gene expression is related to the amount of PCR product generated in an amplification reaction.
- the amount of product can depend on the activity of the polymerase enzymes as well as the length of a fragment being replicated. If the enzyme functions effectively, the amount of PCR product is uniformly high from small fragments to long fragments. If the enzyme activity is less effective, however, the amount of PCR product can be relatively less for long fragments than for short fragments.
- An overall scaling does not account for the non-uniform tapering of the signal with the size of the amplicon.
- the present invention discloses a method of identifying a difference between at least two groups, wherein each group comprises a data set containing ordered elements.
- the method includes the steps of: (a) providing a first group having one or more elements in a first data set; (b) applying at least one transformation to said first data set to provide a transformed data set, wherein said transformation is a calculation selected from a normalizing calculation, an averaging calculation and a scaling calculation; and (c) distinguishing differences, if present, between elements of said first transformed data set and a second groups having one or more elements in a second data set; thereby identifying a difference between the data sets.
- the method corrects the effects of the noise prior to distinguishing the differences.
- regions in a data set that do not contain useful data may be masked, and regions that have a higher information content may be highlighted. These include regions where the signal intensity in the data set is either too low (noise) or too high (saturation) for accurate measurement, or is at locations of local peaks.
- the noise includes low frequency noise.
- the noise includes jiggle. In the latter embodiment, the jiggle includes positional shifts of elements between different data sets and signal alignment within a data set. In a further embodiment, the jiggle is corrected. In another embodiment, correction of jiggle may be considered as signal alignment between two or more data sets.
- the elements of a data set may represent, for example, a trace, such as a trace arising in an electrophoretogram or a chromatogram.
- an element of a data set represents a position in a reagent array.
- the position in the reagent array determines one extent of matching between a reagent affixed to the array at the position and a sample contacting the array position.
- the reagent may be a first nucleic acid and the sample may include a second nucleic acid.
- a data set is obtained in an experiment related to identifying significant differences in gene expression.
- a group includes more than one individual.
- a data set of the method may be subjected to a masking operation.
- the data set for each group is obtained by applying at least one calculation, chosen from a normalizing calculation, an averaging calculation and a scaling calculation, to the data sets from each individual.
- each individual provides at least one replicate sample that is employed to provide a trace.
- the traces for each individual are advantageously transformed by applying at least one of a normalizing calculation, an averaging calculation and a scaling calculation to the trace from each replicate; and in additional such embodiments, the traces for each replicate are discretized prior to the normalizing, the averaging and/or the scaling.
- the normalization includes adjusting each data set such that a subset of elements in each data set has similar or identical values.
- the averaging includes calculating an average for a location or for a discretized position across a collection of data sets. The average may be an unweighted average or a weighted average.
- the scaling includes a calculation that causes a first data set to resemble a second data set except that an element in the scaled first data set whose intensity differs significantly from the intensity of the element in the second data set at the same location or the same position contributes to identifying the difference between the data sets.
- the scaling includes calculating a distance between the data sets, or calculating a similarity between the data sets.
- the scaling calculation employs a scaling function; and in other embodiments, the scaling function is a basis set expansion, such as a piecewise linear basis set or a direct product of basis functions.
- successive iterations of a cycle that includes at least one of a normalization calculation, an averaging calculation and a scaling calculation are carried out until a specified termination condition has been satisfied.
- the termination condition is that the transformed data set has converged.
- the termination condition is that a predetermined number of iterations has been reached.
- the distinguishing of differences among the elements of the transformed data sets includes application of a difference finding algorithm.
- the invention also discloses a display means that displays a representation of a difference between data sets, and also discloses the representation itself, wherein the representation is obtained in general by applying methods disclosed herein to the data sets.
- FIG. 1 is a graphic representation of jiggle arising between two traces.
- FIG. 2 is a schematic diagram illustrating the flow from different groups to the transformed data sets for those groups.
- FIG. 3 is a schematic estimation of ⁇ , the experimental noise in a data set.
- FIG. 4. is a schematic representation of averages of three replicate raw traces for each individual animal in Example 1 , prior to normalization or scaling.
- FIG. 5. is a schematic representation of averages of 3 normalized replicate traces for each individual animal in Example 1.
- FIG. 6. is a schematic representation of scaling factors employed to scale the phenobarbital-treated individual average to the sterile-water-treated individual average in Example 1, obtained as a result of iterative scaling.
- FIG. 7. is a schematic representation of iteratively scaled traces for each individual animal in Example 1.
- the invention discloses methods for normalizing, scaling, and difference finding that may be used in any experimental study, including gene expression data, in which noise and other uncontrolled variations exist between data sets.
- these methods have been adapted for differential display experiments, in which gene expression levels are represented by fragment intensities in, for example, an electrophoresis trace.
- the methods are also applicable to, for example, hybridization experiments, such as those employed with nucleic acid microchip arrays, as well as other experiments not related to gene expression.
- the present invention discloses a method of identifying a difference between at least two groups, wherein each group comprises a data set containing ordered elements.
- the method includes the steps of: (a) providing a first group having one or more elements in a first data set; (b) applying at least one transformation to said first data set to provide a transformed data set, wherein said transformation is a calculation selected from a normalizing calculation, an averaging calculation and a scaling calculation; and (c) distinguishing differences, if present, between elements of said first transformed data set and a second groups having one or more elements in a second data set; thereby identifying a difference between the data sets.
- Each data set can be represented as a set of discretized intensity values, or elements in the data set.
- the intensity of an element may include effects of noise, as described below, and the method operates to correct the differences for the effects of the noise.
- the noise includes low frequency noise.
- the noise includes differences in jiggle.
- the jiggle includes positional phase shifts of elements between different data sets.
- the invention also discloses a display means that displays a representation of a difference between data sets, and also discloses the representation itself, wherein the representation is obtained in general by applying methods disclosed herein to the data sets.
- representation relates to any graphical, visual, or equivalent non-verbal display that provides an image of the results, such as differences between data sets, obtained according to the methods of the present invention. More specifically, a "representation" of the invention is obtained by transforming the quantitative results gathered by experiments underlying the invention. Examples of such data include, by way of non-limiting example, traces from differential gene expression, and intensities from an array, and/or equivalent types of experimental parameter.
- a representation of the invention is generated by algorithms executed in a computer and is suitable for display on a display means, such as a display screen or monitor, employed in the operation of the computer.
- the representation is also suitable for storing in a storage module or data archive of such a computer. It is still further suitable for printing from the computer onto a medium such as paper or equivalent physical medium, and for recording it onto a portable storage medium, including, for example, magnetic media, CD ROMs and equivalent storage media.
- display means includes any of the objects and media identified above in this paragraph, as well as equivalent apparatuses and objects suitable for displaying the results of computational processes for visual inspection.
- normalization is defined herein as a means for standardizing or correcting elements in a data set, for example, but not by way of limitation, for correcting overall signal strength within a given data set.
- Features of given elements to be normalized are first identified within a data set. For example, one such feature may be the median peak height of signals within a data set.
- a summary statistic for the given feature is generated for a data set, and used to normalize the elements, as described below, to allow comparisons across data sets.
- Algorithms that are designed to either mask or highlight chosen features identified among the elements of a data set may be applied prior to normalization, averaging or scaling. Such features include low intensity signal regions that comprise noise, high intensity signal regions that comprise saturation zones, and local maxima that comprise peaks.
- Average is defined as combining multiple data sets to generate one average representative data set. Averages are combined into the representative data set in such a way that noise from any one data set so combined does not affect any other data sets used to generate the average.
- scaling is defined as a correction for low frequency difference is signal strength across data sets.
- the data sets may arise in any of a number of ways. Any experiment or study in which one group is compared with another may provide the data sets employed in the invention. Such groups may be distinguished by the experimental conditions experienced by the respective groups, or by the experimental state characterizing the respective groups. Experimental subjects may be animate or inanimate, or may be inanimate samples derived from animate subjects.
- display means and representations the data sets arise from experiments conducted in investigations in which identification of the differential expression of a gene or genes between data sets from at least one experimental group and at least one control group is sought.
- such differential expression arises in GeneCallingT experiments. See, e.g., United States Patent No. 5,871,697; Shimkets et al, Nat. Biotechnol. 17: 798-803 (1999).
- such differential expression is evaluated using nucleic acid microchip arrays in order to detect the presence, absence or extent of expression of a gene or gene fragment. Any alternative, equivalent differential expression formats and methods of analysis are encompassed within the scope of the present invention as well.
- noise may arise during the course of gathering the data elements comprising the data sets.
- Nonlimiting examples of noise include intensity noise and extension noise leading to longitudinal differences.
- Intensity noise includes relatively high frequency noise such as that commonly associated with short-time fluctuations in the electronic and/or mechanical components of an experimental system.
- High frequency noise may be defined as having a frequency greater than about 1 Hz. Examples of high frequency noise include shot noise in photodetectors and comparable electronic noise arising in the various electronic components and circuits of an experimental instrument employed in gathering the data elements of a data set.
- Low frequency noise has a frequency less than about 1 Hz, and may have frequencies less than about 0.1 Hz, or less than about 0.01 Hz, or even less than about 0.001 Hz or lower.
- Such low frequency noise may arise during an experiment, for example, by decay of activity of a reagent, catalyst or enzyme during the course of preparing a sample that is applied to generate a data set.
- an uncompensated low frequency change in response of an electronic instrument may arise during the time in which a data set is being gathered.
- an array is being used to generate the data set, uncompensated variations in detection across the various positions and/or dimensions of the array may arise that behave as low frequency noise (i.e., they may be considered as low frequency noise even though an array may be subjected to simultaneous detection of all the sample points on the array, since positional variations behave as if they have a long wavelength across the array.)
- Equivalent sources of low frequency noise are also encompassed in this definition. Normalization and scaling algorithms employed are particularly effective in minimizing or eliminating the effects of low frequency noise.
- jiggle An additional detrimental effect that may arise in identifying differences between data sets is termed "jiggle".
- Longitudinal displacement relates to variation in the location or discretized position of a particular feature in a trace even though the feature appears in the traces of more than one group.
- uncompensated variation in the location or discretized position of the feature may occur due to variations in physical or chemical conditions during the process of accumulating the data elements of the various data sets being considered.
- Such a variation, or jiggle may be considered to be low frequency noise in the longitudinal, or positional, direction. Jiggle is illustrated in FIG. 1.
- A(n) and B(n) two discretized, normalized data sets, A(n) and B(n), are shown.
- A(n) and B(n) should be thought of as each representing the same feature. Nevertheless, they are displayed with a jiggle of 1.75 units on the n axis.
- the normalization and scaling algorithms employed are particularly effective in compensating for and/or overcoming the effects of low frequency longitudinal noise. Such procedures, as employed in the methods of the present invention, largely or completely eliminate the jiggle and, referring to FIG. 1, restore the overlap of the points for A(n) and B(n). Compensation for jiggle as shown in FIG. 1 is also termed "signal alignment." Groups, Individuals, Replicates and Transformed Data
- a hierarchy of notation is used herein to indicate data elements and/or the data sets. These notations are discussed below and furthermore are illustrated in the flow diagram presented in FIG. 2.
- Raw, i.e. untreated or untransformed, data arise from the carrying out the experiments on actual samples obtained from experimental groups.
- a "group” represents a particular experimental state or condition.
- the groups are denoted herein in capital letters A, B, ... without any indices or delimiters.
- FIG. 2 at least two groups comprise the subject matter on which the methods, display means and representations of the present invention are based. Each group gives rise to data elements comprising data sets after samples from the groups have been subjected to a given experimental method of detection or analysis.
- a group may be initially composed of one or more individuals. The number of individuals is not fixed or constant, but may vary. Each individual in the group is subjected to the same experimental conditions or experimental state. For the case of animate groups, each individual may represent an individual animal, a plant (such as a seedling), or a set of cells grown in cell or tissue culture.
- each individual may represent, again by way of nonlimiting example, a separate execution of a particular experimental protocol such as a synthetic or preparative procedure, or the implementation of a particular set of physical conditions on separate samples or objects.
- Equivalent ways of designating individuals of a group are encompassed within the scope of the present invention.
- the data sets obtained from the individuals of a group may be transformed by any one or more of the normalization, averaging and scaling calculations of this invention in arriving at the differences determined by the present methods.
- each individual of a group may furnish one or more replicate samples for detection or analysis according to the experimental method employed. Such replicates also represent raw, or untreated, data.
- Each replicate of an individual is designated with a second index or delimiter j, shown, for example, by ajj, bjj, ... (see FIG. 2).
- the number of replicates may vary due to experimental circumstances. Commonly replicates are obtained by repetitive sampling from the same individual.
- the data sets obtained from the replicates of a particular individual may be operated upon by any one or more of the normalization, averaging and scaling calculations of the present invention in arriving at the differences determined by the present methods.
- the normalization, averaging and/or scaling calculations that are applied to the replicates may be applied prior to, or simultaneously with, the similar calculations applied to the individuals and discussed in the preceding paragraph.
- each data set may be considered to be comprised of an infinite number of data elements designated using a further delimiter, ajj(x), where x denotes the continuous longitudinal dimension of the analytical method.
- ajj(x) a further delimiter
- Such data sets therefore, in general need not carry the additional delimiter x.
- the delimiter n replaces the delimiter x when the intensity trace has been discretized; i.e., ajj(x) becomes ajj(n) (see FIG. 2).
- any data sets that have been transformed using the calculations disclosed herein are designated in upper case letters including an index and/or a delimiter.
- the transformations may include at least one operation chosen from among normalization, averaging and scaling.
- a transformation that operates to combine the replicates of an individual while leaving the individuals of a group intact results in a transformed data set indicated by one index and a delimiter, Aj(n), B j (n), ... (see FIG. 2).
- Further transformation that operates to combine the individuals of a group to provide a single data set for an entire group is designated by a delimiter only, as shown, for example, by A(n), B(n), ... (see FIG. 2).
- the Aj(n), B ⁇ (n), ... are obtained directly without discretization. They may still arise from replicate samples, however.
- the detection method is based on an array, one or more positions in the array may represent the results of one or more replicates, respectively.
- a particular embodiment of a data set envisioned in the present invention is differential display.
- mRNA is extracted from sample, converted to cDNA, and digested with restriction enzymes into fragments (United States Patent No. 5,871,697; Shimkets et al, Nat. Biotechnol. 17: 798-803 (1999)). The fragments are then separated according to length using electrophoresis.
- nucleic acids consist of an integer number of nucleotides, their electrophoretic transport properties also depend on the nucleotide composition.
- Electrophoresis experiments currently in use measure the length of a nucleic acid fragment determined electrophoretically to a precision of 0.1 nt, and the actual value of the electrophoretic length is usually within 1-2 nt of the actual number of nucleotides in the fragment.
- the intensity signal a(x) of the electrophoresis trace for sample A represents the amount of fragments of electrophoretic length x generated from the sample.
- the intensity a(x) also depends on the particular restriction enzymes used to generate fragments; for simplicity, this dependence is suppressed in the notation.
- the intensity at length x corresponds to a single fragment; sometimes multiple fragments have the same length and their signals are combined; sometimes no fragments are present and a(x) is a baseline signal.
- a(x) is a measured intensity, it should be a positive quantity.
- the intensity a(x) from samples in group A is compared with the intensity b(x) generated using an identical protocol from samples in group B.
- Some inadvertent differences between A and B can be attributed to underlying genetic variation in the individuals chosen.
- samples A and B may include organisms or individuals having an allelic variation between them that generates a difference in the measured expression levels but has no biological relevance in the context of the particular experimental study.
- a neutral single nucleotide polymorphism (SNP) can add or remove a band. For this reason it is preferable to include multiple organisms or individuals for samples A and B to control for these types of individual differences.
- the expression profile of the i m individual of group A is denoted aj(x), and similarly bj(x) is the expression profile for individual i of group B.
- each organism or individual may have multiple experimental replicates of the expression profiles for each organism or individual.
- the j m expression profile, or replicate, of the i m individual of group A is denoted as ajj(x), and similarly for group B.
- each group may have one or more individuals, and each individual may have one or more replicates. More elaborate hierarchies are also possible and may be analyzed directly with the methods outlined below.
- An alternative embodiment of an experimental system relates to hybridization.
- ajj(x) represent the intensity from the j m experimental replicate of the i m organism in group A measured at position x on a hybridization array or chip.
- x is a two- dimensional coordinate that identifies the location of a particular spot on the hybridization surface.
- trace is used herein to denote a one-dimensional data set
- array or hybridization data is used herein to represent a two-dimensional data set. Terms such as data set, signal, and intensity may represent one- or two-dimensional data sets.
- repeated experiments such as hybridization experiments conducted on a series of biological samples collected over time, may have additional dimensions.
- each time point in a time course study generates a two-dimensional plane of data, and the time coordinate adds a third dimension.
- the methods disclosed herein are applicable to data sets such as these, as well as to those of the preceding paragraphs. In full generality the methods disclosed herein are generally applicable to any experimental study that generates multi-dimensional data sets. In particular cases, attention may be restricted to a particular dimensionality of data sets, as the specific character of the study may provide.
- the terms "signal” and "intensity" may be considered interchangeable references to either measured, normalized, scaled, or averaged data. Data sets can have many representations in the memory of a computer.
- each data set can be represented as a set of discretized intensity values, or elements in the data set.
- ⁇ x is preferably close to the reproducibility of the instrument. With currently available instruments, a value of 0.1 nt is preferable. For raw hybridization images, a value corresponding to a single pixel in an image is preferable. For processed hybridization images, it is preferable that each discretization point represents an individual spot with a different probe.
- the invention allows for the identification of a difference between data sets, wherein each data set contains data elements as described above.
- the differences are identified by operating on at least two transformed data sets, A(n) and B(n), to discern particular discrete positions n at which differences that exceed a lower limit of distinction are found.
- the methods are that they use algorithms that automatically identify scaling landmarks that may exist in the data sets being evaluated.
- the noise mask m no j se (n) depends on a noise level I n oise tnat characterized the experimental uncertainty in the measured intensity. This uncertainty may be estimated, for example, as the standard deviation of the background signal obtained for a blank or control sample. I no ise ma Y a ⁇ so De preferably assigned a value that is a small multiple (0.2X to 5X) of the low end of the dynamic range of the detection instrument.
- the noise mask is calculated as follows:
- the saturation mask m sa t(n) marks data collected near the upper limit of the detection range of an instrument.
- the mask depends on a saturation level, as follows: • For each position n
- the threshold I sa ⁇ is preferably close to the high end of the dynamic range of the detection instrument (0.95X or to IX).
- points that are marked as not saturated are checked for saturation in a second pass that depends on a saturation width w sa t as follows:
- n o m sa t(n) 1 if a(n) is part of a plateau of constant value over a range of ⁇ w sa ⁇ in each dimension.
- m sa t(n) 1 for each of these points.
- m sa ⁇ (n) 1 only for the center point; the remaining points require their own saturation checks.
- o m sa t(n) 0 otherwise.
- a parameter pg ⁇ defines the half-width for a peak as follows:
- pg ⁇ n 1 if a(n + ⁇ n') ⁇ a(n + ⁇ n) for all ⁇ n and ⁇ n' such that ⁇ n' is farther than ⁇ n from n, ⁇ n' is no farther than Wp ⁇ from n, and
- ⁇ n and ⁇ n' are identical except for a single dimension in which they differ by 1.
- distances may be calculated by any of a number of methods, including the Euclidean metric, the Manhattan metric, or the maximum absolute difference in any dimension.
- ⁇ aj ⁇ n) 1 if a ⁇ -Wpg ⁇ ) ⁇ ... ⁇ a(n-2) ⁇ a(n-l) ⁇ a(n) > a(n+l) > a(n+2) > ... > a(n+w pea ' c ).
- a summary peak intensity Ip ea k is calculated from the individual values. If desired, the values can be rank-ordered, and Ip ea k can be defined as the 75 m percentile value (75% of peaks are smaller in value; 25% of peaks are larger). Other methods include using a different percentile, for example the median, or calculating an average value. Rank-order selection methods are more robust than averages.
- the data set is rescaled by multiplying each point a(n) by the factor (Inorm ⁇ peak)' nere Inorm ls identical for each data set and sets a convenient scale.
- I norm is arbitrary, a value such as 100 is convenient.
- the noise threshold I no ise mav a ⁇ so ⁇ > e subject to the same normalization.
- Inoise mav De set to a fixed value.
- Averaging is an operation that is applied to a collection of data sets.
- the average A(n) of a collection of r data sets a](n), a2(n), ... , a r (n) is calculated as
- weighting functions may be used, and may account for characteristics of a data set such as those discussed in the following.
- SD ⁇ (n) it is preferable to calculate a standard deviation SD ⁇ (n) to describe the distribution of data points leading to the average A(n).
- a preferable formula for SD ⁇ (n) is
- This extent of agreement can be measured, by way of nonlimiting example, by the distance Dist[A,B] or the similarity Sim[A,B] between two data sets A(n) and B(n).
- Dist[A,B] ⁇ n w[A(n),B(n)] dist[A(n),B(n)] and
- Dist[A,B] ⁇ n w[A(n),B(n)] dist[A(n),B(n)] / ⁇ n w[A(n),B(n)].
- dist[A(n),B(n)] is a function that measures the distance between two values
- the term w[A(n),B(n)] is a mask that determines whether the data points at location n should be included in the calculation.
- the second formula is preferable.
- a and b must be regularized to prevent values close to 0 from causing a divergence. This can be accomplished, for example, by replacing a or b by a minimum value I mm if either is smaller than I mm , or by adding a positive constant to raise all values A(n) and B(n) above 0.
- and D max 3.
- the weight w ⁇ n ) is preferably [1 - n A,noise( n )][I - niA,sat( n )]> anc s i m il ar ly f° r w B(n)-
- a similarity Sim[A,B] between two data sets A and B may be defined as
- Sim[A,B] ⁇ n sim[A(n),B(n)] where the similarity sim(a,b) between two numbers a and b is larger when the quantities are larger and also larger when a and b are closer in value.
- ⁇ /2 , d
- sim(a,b) F(p) G(d) where F(p) is an increasing function of p and G(d) is a decreasing function of d.
- This algorithm is related to one of the literature as a method of decomposing spectra of multicomponent mixtures into separate spectra for each of the pure components.
- Scaling is an operation that is applied to a subordinate data set a(n) to bring it in closer agreement with a master data set A(n).
- a scaling algorithm optimizes a scaling function s(n) to minimize the distance or maximize the similarity between the scaled slave data set s(n)a(n) and the master data set A(n).
- the scaling function s(n) may have various mathematical representations.
- a basis set expansion may be a cosine series, a sine series, or more generally, a Fourier series. For a one-dimensional data set, a preferred choice is a piecewise linear basis.
- the p m basis function is zero outside the interval np.i to n p , with X Q taken as the left-most point of the data set and np as the right-most point.
- s(n) c p _ ⁇ + (c p - c p _ ⁇ ) (n - n p _ ⁇ )/(n p - n p _ ⁇ ).
- s(n) ⁇ p c p ⁇ p (n), where here n and p are both d-dimensional and ⁇ p (n) can be expressed as
- ⁇ p(n) ⁇ pl(n ⁇ ) ⁇ p 2(n2) ... ⁇ pd d)
- n; and p; are the components of n and p in dimension j and ⁇ pj(ni) is a one- dimensional basis set in dimension j .
- a preferred choice for the one-dimensional basis sets in the multidimensional direct product is an orthogonal basis.
- a preferred choice for an orthogonal basis is a trigonometric basis,
- ⁇ pj(n) cos[(pj-l) ⁇ (n-nj 0 )/(n-nji)] , where njo and nu are the left-most and right-most points in dimension j .
- the coefficients Cp are selected to minimize the distance Dist[A(n),s(n)a(n)] or maximize the similarity Sim[A(n),s(n)a(n)]. Methods to perform this optimization are well-known in the art. Preferable methods are conjugate direction minimization or conjugate gradient minimization, which use linear algebra to optimize the P basis set coefficients simultaneously. See, e.g., Press et al., NUMERICAL RECIPES IN C, THE ART OF SCIENTIFIC COMPUTING, Second Edition, Cambridge Univ. Press, Cambridge UK, 1992, Chapter 10.
- a preferable approximation that is faster computationally than a full minimization is to obtain Cp from an interval surrounding np, preferably the interval from n p _ ⁇ to n p + ⁇ , by minimizing the distance Dist[A(n),c p a(n)] or maximizing the similarity Sim [ A(n) ,c p a(n)] .
- the number of piecewise linear basis functions is selected so that the low-frequency noise in the data occurs on a length scale of L/(0.3P) or longer.
- L 400 nt and a noise length scale approximately 100 nt, P « 13 is preferable (interpolation points spaced every 35 nt).
- a group of data sets ⁇ aj(n) ⁇ can be brought into closer agreement with each other by first normalizing each data set, then generating an average A(n), then scaling each data set aj(n) to the average A(n), then repeating these steps. If desired, the average A(n) can be re-normalized after every iteration.
- a preferable termination condition is that A(n) has converged. This means that the distance Dist[A(n),A'(n)] between the value of A(n) after an iteration and its value A'(n) after the next iteration is smaller than some threshold value.
- the scaling functions sj(n) for each of the slave data sets aj(n) can be checked for convergence.
- a second preferable termination condition is that a predetermined number of iterations has been reached.
- the square distance measure essentially calculates the standard deviation of the data sets.
- minimizing the square distance is essentially identical with performing scaling that minimizes the standard deviation of the scaled traces.
- Iterative scaling may occur at a hierarchy of levels including experimental replicates, independent individuals, and groups. Recall that the data set corresponding to experimental replicate j of organism i of group A is aij(n). Similarly, the data sets bji(n) are obtained for group B, data sets cjj(n) for group C, and so forth for each of the groupings.
- data sets are normalized, scaled, and averaged within each organism, then within each group, and then between groups.
- One process is as follows:
- the scaling terms must be back-propagated to compare averages other than the final, scaled group averages. For example, if scaled group averages are required, s(n)A(n) is used. If scaled individual averages are required, then s(n)s j (n)A ⁇ (n) is used. If scaled data sets are required, then s(n)si(n)sjj(n)ajj(n) is used.
- each data set from individual i is preferably given a weight proportional to 1 /(number of replicates from individual i) . This gives each individual equal weight and prevents an individual with many replicates from dominating the average.
- Other preferable methods include weighting each data set equally and weighting each data sets to give each group equal weight. If each group has equal weight, one method is to weight each replicate equally. Thus, each data set from group A is given a weight proportional to 1 /(number of replicates from all the individuals belonging to group A).
- An alternate method is to weight each data set variably to give each individual within a group equal weight. Thus, each data set from individual i of group A is given a weight proportional to 1 /[(number of individuals in group A)(number of replicates in individual i)].
- a preferable threshold for differential display data is 2 iterations.
- Jiggling One aspect of difference finding is comparing the heights of peaks in two data sets.
- the same peak may occur at different positions in different data sets. For example, a peak in one replicate of data set may occur at position n, while in a second data set it may occur at position n+1 or n-1 due to experimental variability. (See FIG. 1)
- a jiggling algorithm identifies the peak height in a data set a(n) that corresponds to a given location n'.
- a preferred jiggling algorithm requires a parameter wjjggi e , which describes the width of the jiggling window.
- the preferred algorithm starts at position n' and searches for the peak in a(n) closest to n' and within distance wjig g ⁇ e .
- the height of a(n) at this peak position is termed the jiggled height of a(n) at n'. If two peaks are within equal distance, the higher value is preferably taken as the jiggled height. If there is no peak within distance wjigg' e , then the height a(n') is the jiggled height.
- the data range for the jiggling peak search is n'-wjiggi e through n'+wjiggi e . If n' is a peak in a(n), then the value a(n') is the jiggled height of a(n) at n'. Otherwise the positions n' ⁇ l, n' ⁇ 2, ..., n' ⁇ W ⁇ ggi e are tested in turn for peaks; if a peak is found at location n", then a(n") is the jiggled height of a(n) at n'. If no peak is found, then a(n') is the jiggled height.
- Difference Finding Difference finding identifies locations where at least one of the groups has a peak and its value is significantly different from the other groups.
- the group averages and individual averages produced by iterative scaling serve as inputs to difference finding.
- a preferable method employs a parameter Wrjjff that defines the minimum distance between differences.
- Wrjjff a parameter that defines the minimum distance between differences.
- Wpg ⁇ a parameter that defines the minimum distance between differences.
- W fj jff ⁇ x 1.1 nt.
- a preferred algorithm is as follows: • Generate a master peak mask Mp ea k(n) using one of the following alternatives: o Option 1 : For each group A and individual Aj, calculate a peak mask m pea j (n) from the individual average Aj(n). Then, for each position n, M pea i c (n) is 1 if at least one of the individual peak masks is 1 and is 0 otherwise. o Option 2: Calculate the peak mask Mp ea ] f (n) directly from the grand mean M(n) of all the groups. o Option 3: If there are only two groups A and B, generate a peak mask from the difference A(n) — B(n).
- a pooled variance t-test may be employed instead of an F-test.
- a less preferred algorithm for a comparison between two groups, A and B, and one- dimensional data set is as follows: • Perform the final step of iterative scaling by scaling A(n) to B(n).
- LASTPOSITION n
- LASTDIRECTION direction(n)
- LASTPVALUE p-value(n). o Otherwise if p-value(n) ⁇ LASTPVALUE then
- mice Male Sprague-Dawley rats (Harlan Sprague Dawley, Inc., Indianapolis, Indiana) of 10-14 weeks of age were gavage-fed and dosed with phenobarbitol once a day for three days at a dose of 3.81 mg/kg/day.
- the drug was dissolved in sterile water prior to treatment. This dosage corresponds to the ED 100 (the upper limit of the effective dose for humans) adjusted for the difference in metabolic rate between rats and humans.
- Three rats were used for the drug treatment group, and an additional three rats were treated with sterile water to serve as the control group.
- Rats were sacrificed 24 hours after the final dose and their brains were harvested. Collection of mRNA from the harvested brains, synthesis of the corresponding cDNA, and differential display protocols were done as has been described elsewhere. See, U. S. Patent No. 5,871,697; Shimkets et al. Nat. Biotechnol. 17: 798-803 (1999).
- the basis set used was piecewise linear with 13 scaling points located every 35 nt beginning at 30 nt and ending at 450 nt.
- the distance function was [ln(a(n)/A(n))] 2 , and points for which
- the first step in the scaling procedure was that the 3 normalized traces for each individual were averaged, scaled to the average, re-averaged, then re-scaled to the average for 2 rounds of iterative scaling.
- the phenobarbital-treated individual average and the sterile-water-treated individual average were themselves averaged, the individual averages scaled to the grand average, then the process repeated for 2 rounds of iterative scaling.
- the phenobarbital- treated individual average was scaled to the sterile-water-treated individual average.
- the final scaling factors are shown in FIG. 6 for the two individuals.
- the final individual averages are shown in FIG. 7.
Landscapes
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
- Apparatus Associated With Microorganisms And Enzymes (AREA)
- Complex Calculations (AREA)
Abstract
This invention relates to a method of identifying a difference between at least two data sets made up of ordered elements utilizing internal features within the data sets for calculations relating to normalization, scaling, and difference finding.
Description
NORMALIZATION, SCALING, AND DIFFERENCE FINDING
AMONG DATA SETS
FIELD OF THE INVENTION
This invention relates to statistical analysis of differences between at least two data sets.
RELATED APPLICATIONS
The present patent application claims priority to the United States provisional patent application U.S.S.N. 60/114,806, entitled "Scaling and Normalization" filed January 5, 1999, which is incorporated herein by reference in its entirety.
BACKGROUND OF THE INVENTION
Measurements of the expression levels of individual genes within a cell provide a wealth of information about cellular processes. This is done by extracting messenger R A molecules (mRNA) from a cell, possibly converting these to more stable cDNA molecules, and measuring the concentrations of each individual species by methods such as differential display or hybridization. A typical analysis strategy is to identify genes whose expression levels differ between particular biological states. One difficulty in performing such an analysis is that experimental measurements of expression levels include variation due to noise. Distinguishing the true differences from the false differences (those due simply to noise) has presented a challenge for gene expression analysis. The relevance of this problem is that genes that are differentially regulated can be converted to commercial products, including protein therapeutics, antibody targets, therapeutic markers, as well as conventional drug targets.
More broadly, similar data sets may arise in any of a number of ways. Any experiment or study in which one group is compared with another may provide data sets whose elements include the effects of noise. Noise and other uncontrolled variations within and among data sets arising in such experiments make comparisons between the data sets more difficult, and present challenges in evaluating the results appropriately.
One of the major sources of noise in such experiments, including gene expression experiments, is that the amount of material analyzed, such as mRNA or cDNA, can differ from experiment to experiment, or among the replicates of a single experiment. Data analysis
strategies typically account for this overall variation by performing a global scaling of all the measurements from such experiments. For example, if sample A has twice the overall cDNA concentration of sample B, then the expression level for a gene in sample B must be doubled before comparison with sample A. Often, however, such an overall scaling is not sufficient to discriminate between true differences and those that can be attributed to noise. One source of difficulty is identifying the particular features in the data set or sets that can be used as scaling landmarks. It is not always possible to identify a priori such unchanging features ahead of time
An additional source of noise generally present in experimental studies is noise from analytical instruments and methods. In differential display experiments, for example, the amount of gene expression is related to the amount of PCR product generated in an amplification reaction. The amount of product can depend on the activity of the polymerase enzymes as well as the length of a fragment being replicated. If the enzyme functions effectively, the amount of PCR product is uniformly high from small fragments to long fragments. If the enzyme activity is less effective, however, the amount of PCR product can be relatively less for long fragments than for short fragments. An overall scaling does not account for the non-uniform tapering of the signal with the size of the amplicon.
There thus remains a strong need for counteracting and overcoming the effects of noise in comparing data sets in an experimental study. There is a lack of adequate means for identifying constant or unvarying components in a data set, which may serve as reference markers in normalizing, scaling and distinguishing differences among data sets in such a study. There further is a significant need for a means to identify scaling landmarks automatically in data sets being compared to one another. In a particular framework addressed in this invention, there is a need for robust methods that normalize scale and find differences in experiments related to the differential expression of genes in cells and tissues subjected to specific experimental treatments. These and comparable needs are addressed by the present invention.
SUMMARY OF THE INVENTION
The present invention discloses a method of identifying a difference between at least two groups, wherein each group comprises a data set containing ordered elements. The method includes the steps of: (a) providing a first group having one or more elements in a first data set;
(b) applying at least one transformation to said first data set to provide a transformed data set, wherein said transformation is a calculation selected from a normalizing calculation, an averaging calculation and a scaling calculation; and (c) distinguishing differences, if present, between elements of said first transformed data set and a second groups having one or more elements in a second data set; thereby identifying a difference between the data sets.
In some embodiments, the method corrects the effects of the noise prior to distinguishing the differences. Prior to normalization, averaging, and/or scaling calculations, selected regions in a data set that do not contain useful data may be masked, and regions that have a higher information content may be highlighted. These include regions where the signal intensity in the data set is either too low (noise) or too high (saturation) for accurate measurement, or is at locations of local peaks. In one embodiment, the noise includes low frequency noise. In another embodiment, the noise includes jiggle. In the latter embodiment, the jiggle includes positional shifts of elements between different data sets and signal alignment within a data set. In a further embodiment, the jiggle is corrected. In another embodiment, correction of jiggle may be considered as signal alignment between two or more data sets.
In some embodiments, the elements of a data set may represent, for example, a trace, such as a trace arising in an electrophoretogram or a chromatogram. In an alternative embodiment, an element of a data set represents a position in a reagent array. In a further significant embodiment the position in the reagent array determines one extent of matching between a reagent affixed to the array at the position and a sample contacting the array position. For example, the reagent may be a first nucleic acid and the sample may include a second nucleic acid.
In a further embodiment, a data set is obtained in an experiment related to identifying significant differences in gene expression. In some embodiments, a group includes more than one individual. Additionally, a data set of the method may be subjected to a masking operation. In further embodiments, the data set for each group is obtained by applying at least one calculation, chosen from a normalizing calculation, an averaging calculation and a scaling calculation, to the data sets from each individual. In still further embodiments, each individual provides at least one replicate sample that is employed to provide a trace. In such embodiments, the traces for each individual are advantageously transformed by applying at least one of a normalizing calculation, an averaging calculation and a scaling calculation to the trace from each
replicate; and in additional such embodiments, the traces for each replicate are discretized prior to the normalizing, the averaging and/or the scaling.
In another embodiment, the normalization includes adjusting each data set such that a subset of elements in each data set has similar or identical values. In further embodiments, the averaging includes calculating an average for a location or for a discretized position across a collection of data sets. The average may be an unweighted average or a weighted average.
In additional embodiments, the scaling includes a calculation that causes a first data set to resemble a second data set except that an element in the scaled first data set whose intensity differs significantly from the intensity of the element in the second data set at the same location or the same position contributes to identifying the difference between the data sets. In further embodiments, the scaling includes calculating a distance between the data sets, or calculating a similarity between the data sets. In particularly embodiments, the scaling calculation employs a scaling function; and in other embodiments, the scaling function is a basis set expansion, such as a piecewise linear basis set or a direct product of basis functions.
In yet additional embodiments, successive iterations of a cycle that includes at least one of a normalization calculation, an averaging calculation and a scaling calculation are carried out until a specified termination condition has been satisfied. In particularly embodiments, the termination condition is that the transformed data set has converged. Alternatively, the termination condition is that a predetermined number of iterations has been reached.
In a further embodiment, the distinguishing of differences among the elements of the transformed data sets includes application of a difference finding algorithm.
The invention also discloses a display means that displays a representation of a difference between data sets, and also discloses the representation itself, wherein the representation is obtained in general by applying methods disclosed herein to the data sets.
BRIEF DESCRIPTION OF THE DRAWING
FIG. 1 is a graphic representation of jiggle arising between two traces.
FIG. 2 is a schematic diagram illustrating the flow from different groups to the transformed data sets for those groups.
FIG. 3 is a schematic estimation of σ, the experimental noise in a data set.
FIG. 4. is a schematic representation of averages of three replicate raw traces for each individual animal in Example 1 , prior to normalization or scaling.
FIG. 5. is a schematic representation of averages of 3 normalized replicate traces for each individual animal in Example 1.
FIG. 6. is a schematic representation of scaling factors employed to scale the phenobarbital-treated individual average to the sterile-water-treated individual average in Example 1, obtained as a result of iterative scaling.
FIG. 7. is a schematic representation of iteratively scaled traces for each individual animal in Example 1.
DETAILED DESCRIPTION OF THE INVENTION
The invention discloses methods for normalizing, scaling, and difference finding that may be used in any experimental study, including gene expression data, in which noise and other uncontrolled variations exist between data sets. In particular embodiments, these methods have been adapted for differential display experiments, in which gene expression levels are represented by fragment intensities in, for example, an electrophoresis trace. The methods are also applicable to, for example, hybridization experiments, such as those employed with nucleic acid microchip arrays, as well as other experiments not related to gene expression.
The present invention discloses a method of identifying a difference between at least two groups, wherein each group comprises a data set containing ordered elements. The method includes the steps of: (a) providing a first group having one or more elements in a first data set; (b) applying at least one transformation to said first data set to provide a transformed data set, wherein said transformation is a calculation selected from a normalizing calculation, an averaging calculation and a scaling calculation; and (c) distinguishing differences, if present, between elements of said first transformed data set and a second groups having one or more elements in a second data set; thereby identifying a difference between the data sets.
Each data set can be represented as a set of discretized intensity values, or elements in the data set. The intensity of an element may include effects of noise, as described below, and the method operates to correct the differences for the effects of the noise. In one embodiment, the
noise includes low frequency noise. In another embodiment, the noise includes differences in jiggle. In the latter embodiment, the jiggle includes positional phase shifts of elements between different data sets. The invention also discloses a display means that displays a representation of a difference between data sets, and also discloses the representation itself, wherein the representation is obtained in general by applying methods disclosed herein to the data sets.
As used herein, "representation" relates to any graphical, visual, or equivalent non-verbal display that provides an image of the results, such as differences between data sets, obtained according to the methods of the present invention. More specifically, a "representation" of the invention is obtained by transforming the quantitative results gathered by experiments underlying the invention. Examples of such data include, by way of non-limiting example, traces from differential gene expression, and intensities from an array, and/or equivalent types of experimental parameter.
In some embodiments, a representation of the invention is generated by algorithms executed in a computer and is suitable for display on a display means, such as a display screen or monitor, employed in the operation of the computer. The representation is also suitable for storing in a storage module or data archive of such a computer. It is still further suitable for printing from the computer onto a medium such as paper or equivalent physical medium, and for recording it onto a portable storage medium, including, for example, magnetic media, CD ROMs and equivalent storage media. As used herein, "display means" includes any of the objects and media identified above in this paragraph, as well as equivalent apparatuses and objects suitable for displaying the results of computational processes for visual inspection.
In addition, "normalization" is defined herein as a means for standardizing or correcting elements in a data set, for example, but not by way of limitation, for correcting overall signal strength within a given data set. Features of given elements to be normalized are first identified within a data set. For example, one such feature may be the median peak height of signals within a data set. A summary statistic for the given feature is generated for a data set, and used to normalize the elements, as described below, to allow comparisons across data sets. Algorithms that are designed to either mask or highlight chosen features identified among the elements of a data set may be applied prior to normalization, averaging or scaling. Such features include low intensity signal regions that comprise noise, high intensity signal regions that comprise saturation zones, and local maxima that comprise peaks. "Averaging" is defined as combining multiple
data sets to generate one average representative data set. Averages are combined into the representative data set in such a way that noise from any one data set so combined does not affect any other data sets used to generate the average. "Scaling" is defined as a correction for low frequency difference is signal strength across data sets. The data sets may arise in any of a number of ways. Any experiment or study in which one group is compared with another may provide the data sets employed in the invention. Such groups may be distinguished by the experimental conditions experienced by the respective groups, or by the experimental state characterizing the respective groups. Experimental subjects may be animate or inanimate, or may be inanimate samples derived from animate subjects. In some embodiments, display means and representations, the data sets arise from experiments conducted in investigations in which identification of the differential expression of a gene or genes between data sets from at least one experimental group and at least one control group is sought. In certain embodiments of the invention, such differential expression arises in GeneCallingT experiments. See, e.g., United States Patent No. 5,871,697; Shimkets et al, Nat. Biotechnol. 17: 798-803 (1999). In other embodiments of the invention, such differential expression is evaluated using nucleic acid microchip arrays in order to detect the presence, absence or extent of expression of a gene or gene fragment. Any alternative, equivalent differential expression formats and methods of analysis are encompassed within the scope of the present invention as well. Various types of noise may arise during the course of gathering the data elements comprising the data sets. Nonlimiting examples of noise include intensity noise and extension noise leading to longitudinal differences. Intensity noise includes relatively high frequency noise such as that commonly associated with short-time fluctuations in the electronic and/or mechanical components of an experimental system. High frequency noise may be defined as having a frequency greater than about 1 Hz. Examples of high frequency noise include shot noise in photodetectors and comparable electronic noise arising in the various electronic components and circuits of an experimental instrument employed in gathering the data elements of a data set.
The methods of the present invention can minimize or eliminate low frequency noise. Low frequency noise has a frequency less than about 1 Hz, and may have frequencies less than about 0.1 Hz, or less than about 0.01 Hz, or even less than about 0.001 Hz or lower. Such low
frequency noise may arise during an experiment, for example, by decay of activity of a reagent, catalyst or enzyme during the course of preparing a sample that is applied to generate a data set.
Alternatively, an uncompensated low frequency change in response of an electronic instrument may arise during the time in which a data set is being gathered. Additionally, if an array is being used to generate the data set, uncompensated variations in detection across the various positions and/or dimensions of the array may arise that behave as low frequency noise (i.e., they may be considered as low frequency noise even though an array may be subjected to simultaneous detection of all the sample points on the array, since positional variations behave as if they have a long wavelength across the array.) Equivalent sources of low frequency noise are also encompassed in this definition. Normalization and scaling algorithms employed are particularly effective in minimizing or eliminating the effects of low frequency noise.
An additional detrimental effect that may arise in identifying differences between data sets is termed "jiggle". By this term is meant that the elements of one data set are offset in a longitudinal direction in comparison with the elements of a second data set with which the first data set is being compared. Longitudinal displacement relates to variation in the location or discretized position of a particular feature in a trace even though the feature appears in the traces of more than one group. By way of nonlimiting example, uncompensated variation in the location or discretized position of the feature may occur due to variations in physical or chemical conditions during the process of accumulating the data elements of the various data sets being considered. Such a variation, or jiggle, may be considered to be low frequency noise in the longitudinal, or positional, direction. Jiggle is illustrated in FIG. 1. In this figure, two discretized, normalized data sets, A(n) and B(n), are shown. A(n) and B(n) should be thought of as each representing the same feature. Nevertheless, they are displayed with a jiggle of 1.75 units on the n axis. It is an additional aspect of the present invention that the normalization and scaling algorithms employed are particularly effective in compensating for and/or overcoming the effects of low frequency longitudinal noise. Such procedures, as employed in the methods of the present invention, largely or completely eliminate the jiggle and, referring to FIG. 1, restore the overlap of the points for A(n) and B(n). Compensation for jiggle as shown in FIG. 1 is also termed "signal alignment."
Groups, Individuals, Replicates and Transformed Data
A hierarchy of notation is used herein to indicate data elements and/or the data sets. These notations are discussed below and furthermore are illustrated in the flow diagram presented in FIG. 2. Raw, i.e. untreated or untransformed, data arise from the carrying out the experiments on actual samples obtained from experimental groups. A "group" represents a particular experimental state or condition. The groups are denoted herein in capital letters A, B, ... without any indices or delimiters. As shown in FIG. 2, at least two groups comprise the subject matter on which the methods, display means and representations of the present invention are based. Each group gives rise to data elements comprising data sets after samples from the groups have been subjected to a given experimental method of detection or analysis. Experimental data not transformed by any calculations of the methods disclosed herein are designated using lower case letters together with at least one index or delimiter i, shown, for example, by a\, bj, ... (see FIG. 2). A group may be initially composed of one or more individuals. The number of individuals is not fixed or constant, but may vary. Each individual in the group is subjected to the same experimental conditions or experimental state. For the case of animate groups, each individual may represent an individual animal, a plant (such as a seedling), or a set of cells grown in cell or tissue culture. Correspondingly, for inanimate groups, each individual may represent, again by way of nonlimiting example, a separate execution of a particular experimental protocol such as a synthetic or preparative procedure, or the implementation of a particular set of physical conditions on separate samples or objects. Equivalent ways of designating individuals of a group are encompassed within the scope of the present invention. In general, the data sets obtained from the individuals of a group may be transformed by any one or more of the normalization, averaging and scaling calculations of this invention in arriving at the differences determined by the present methods.
As a further hierarchical subclassification, each individual of a group may furnish one or more replicate samples for detection or analysis according to the experimental method employed. Such replicates also represent raw, or untreated, data. Each replicate of an individual is designated with a second index or delimiter j, shown, for example, by ajj, bjj, ... (see FIG. 2). As shown for illustration in FIG. 2, the number of replicates may vary due to experimental circumstances. Commonly replicates are obtained by repetitive sampling from the same
individual. In general, the data sets obtained from the replicates of a particular individual may be operated upon by any one or more of the normalization, averaging and scaling calculations of the present invention in arriving at the differences determined by the present methods. The normalization, averaging and/or scaling calculations that are applied to the replicates may be applied prior to, or simultaneously with, the similar calculations applied to the individuals and discussed in the preceding paragraph.
In many of the detection or analytical methods employed in the experiments underlying the gathering of the presently disclosed data sets, continuous traces of an experimental intensity as a function of a longitudinal variable such as time, elution volume or distance are obtained. Such traces arise, for example, in the use of chromatographic or electrophoretic methods of detection or analysis. Since the traces are continuous, each data set may be considered to be comprised of an infinite number of data elements designated using a further delimiter, ajj(x), where x denotes the continuous longitudinal dimension of the analytical method. (It may be noted that use of alternative analytical or detection methods, for example use of arrays with discrete positions on them, does not generate a continuous trace. Such data sets, therefore, in general need not carry the additional delimiter x.) It is convenient for virtually all calculations carried out as disclosed herein, using computers with discrete memory locations for separate data elements, to discretize a continuous trace into discrete intensities at specified locations, or discrete positions n, on the trace. As used herein, the delimiter n replaces the delimiter x when the intensity trace has been discretized; i.e., ajj(x) becomes ajj(n) (see FIG. 2).
As used herein, any data sets that have been transformed using the calculations disclosed herein are designated in upper case letters including an index and/or a delimiter. As noted, the transformations may include at least one operation chosen from among normalization, averaging and scaling. A transformation that operates to combine the replicates of an individual while leaving the individuals of a group intact results in a transformed data set indicated by one index and a delimiter, Aj(n), Bj(n), ... (see FIG. 2). Further transformation that operates to combine the individuals of a group to provide a single data set for an entire group is designated by a delimiter only, as shown, for example, by A(n), B(n), ... (see FIG. 2). Conversely, in experimental methods of analysis or detection that do not rely on developing traces, the Aj(n), B}(n), ... are obtained directly without discretization. They may still arise from replicate samples, however.
For example, if the detection method is based on an array, one or more positions in the array may represent the results of one or more replicates, respectively.
A particular embodiment of a data set envisioned in the present invention is differential display. In a differential display experiment involving gene expression, mRNA is extracted from sample, converted to cDNA, and digested with restriction enzymes into fragments (United States Patent No. 5,871,697; Shimkets et al, Nat. Biotechnol. 17: 798-803 (1999)). The fragments are then separated according to length using electrophoresis. Although nucleic acids consist of an integer number of nucleotides, their electrophoretic transport properties also depend on the nucleotide composition. Electrophoresis experiments currently in use measure the length of a nucleic acid fragment determined electrophoretically to a precision of 0.1 nt, and the actual value of the electrophoretic length is usually within 1-2 nt of the actual number of nucleotides in the fragment. The intensity signal a(x) of the electrophoresis trace for sample A represents the amount of fragments of electrophoretic length x generated from the sample. Of course, the intensity a(x) also depends on the particular restriction enzymes used to generate fragments; for simplicity, this dependence is suppressed in the notation. Sometimes the intensity at length x corresponds to a single fragment; sometimes multiple fragments have the same length and their signals are combined; sometimes no fragments are present and a(x) is a baseline signal. Because a(x) is a measured intensity, it should be a positive quantity. Mathematical operations used during signal processing, such as the subtraction of a baseline, might result in negative values at certain locations a(x). If negative values exist, their magnitude should be preferably of the same order as the measurement error in the data set.
To detect differences, the intensity a(x) from samples in group A is compared with the intensity b(x) generated using an identical protocol from samples in group B. Some inadvertent differences between A and B can be attributed to underlying genetic variation in the individuals chosen. For example, samples A and B may include organisms or individuals having an allelic variation between them that generates a difference in the measured expression levels but has no biological relevance in the context of the particular experimental study. In differential display, for example, a neutral single nucleotide polymorphism (SNP) can add or remove a band. For this reason it is preferable to include multiple organisms or individuals for samples A and B to control for these types of individual differences. In general, as noted earlier, the expression
profile of the im individual of group A is denoted aj(x), and similarly bj(x) is the expression profile for individual i of group B.
Furthermore, it is preferable to have multiple experimental replicates of the expression profiles for each organism or individual. The jm expression profile, or replicate, of the im individual of group A is denoted as ajj(x), and similarly for group B. Thus, each group may have one or more individuals, and each individual may have one or more replicates. More elaborate hierarchies are also possible and may be analyzed directly with the methods outlined below.
An alternative embodiment of an experimental system relates to hybridization. In this embodiment, let ajj(x) represent the intensity from the jm experimental replicate of the im organism in group A measured at position x on a hybridization array or chip. Here x is a two- dimensional coordinate that identifies the location of a particular spot on the hybridization surface. The term trace is used herein to denote a one-dimensional data set, and the term array or hybridization data is used herein to represent a two-dimensional data set. Terms such as data set, signal, and intensity may represent one- or two-dimensional data sets. Furthermore, repeated experiments, such as hybridization experiments conducted on a series of biological samples collected over time, may have additional dimensions. By way of nonlimiting example, each time point in a time course study generates a two-dimensional plane of data, and the time coordinate adds a third dimension. The methods disclosed herein are applicable to data sets such as these, as well as to those of the preceding paragraphs. In full generality the methods disclosed herein are generally applicable to any experimental study that generates multi-dimensional data sets. In particular cases, attention may be restricted to a particular dimensionality of data sets, as the specific character of the study may provide. Furthermore, as used herein the terms "signal" and "intensity" may be considered interchangeable references to either measured, normalized, scaled, or averaged data. Data sets can have many representations in the memory of a computer. Here it is assumed that each data set can be represented as a set of discretized intensity values, or elements in the data set. Although it is not necessary to use a regular grid to store the intensity, it is convenient to do so. Using a regular grid in one dimension, an intensity a(x) is stored at locations x = 0, Δx, 2Δx, ... , lΔx (where lΔx = L, and L represents the full length of a trace in the longitudinal direction). In two dimensions, a(x) is stored at locations (xj,x2) where x\ - 0, Δx, 2Δx, ... , lΔx, and X2 = 0, Δy, 2Δy, ... , mΔy (mΔy = M). In d-dimensions, a(x) is stored at
locations (x\, X2, ... , x<j) where, for k = 1, 2, ... , d, x^ takes on the values 0, Δx^, 2Δxjς, ... , ljζΔxj . For convenience, we also introduce the notation a(n) where n is a d-tuple of integers (n\, n2, ... , nj) and a(n) corresponds to the intensity a(x) where x^ = nj^Δx^.
For electrophoresis data, such as that generated by differential-display data, Δx is preferably close to the reproducibility of the instrument. With currently available instruments, a value of 0.1 nt is preferable. For raw hybridization images, a value corresponding to a single pixel in an image is preferable. For processed hybridization images, it is preferable that each discretization point represents an individual spot with a different probe.
Methods of Calculation, Algorithms The invention allows for the identification of a difference between data sets, wherein each data set contains data elements as described above. The differences are identified by operating on at least two transformed data sets, A(n) and B(n), to discern particular discrete positions n at which differences that exceed a lower limit of distinction are found. The methods are that they use algorithms that automatically identify scaling landmarks that may exist in the data sets being evaluated.
Details of the ways of identifying statistically significant differences are presented in the following sections and in the EXAMPLE.
Masks for Noise, Saturation, and Peaks
It is useful to mask out regions in the data set that do not contain useful data and to highlight regions that have a higher information content. These include regions in which the intensity is too low or too high for an accurate measurement and locations of peaks. A series of masks m(n) records this information for a data set a(n).
The noise mask mnojse(n) depends on a noise level Inoise tnat characterized the experimental uncertainty in the measured intensity. This uncertainty may be estimated, for example, as the standard deviation of the background signal obtained for a blank or control sample. Inoise maY a^so De preferably assigned a value that is a small multiple (0.2X to 5X) of the low end of the dynamic range of the detection instrument. The noise mask is calculated as follows:
• For each position n
° ™noise(n) = 1 if a(n) < Inoise o mnojse(n) = 0 otherwise.
The saturation mask msat(n) marks data collected near the upper limit of the detection range of an instrument. The mask depends on a saturation level, as follows: • For each position n
° msat(n) = 1 if a(n) > Isat. o msat(n) = 0 otherwise.
The threshold Isa^ is preferably close to the high end of the dynamic range of the detection instrument (0.95X or to IX). Preferably, points that are marked as not saturated are checked for saturation in a second pass that depends on a saturation width wsat as follows:
• For each position n o msat(n) = 1 if a(n) is part of a plateau of constant value over a range of ±wsa^ in each dimension. Thus, in one dimension, if a(n) = a(n+l) = a(n+2) = ... = a(n+wsat) and a(n) = a(n-l) = a(n-2) = ... = a(n-wsal), then msat(n) = 1 for each of these points. Less preferably, msa^(n) = 1 only for the center point; the remaining points require their own saturation checks. o msat(n) = 0 otherwise.
For differential display data, wsatΔx = 0.2 nt is preferable, and Isat corresponds to the camera intensity at saturation.
Next points that are local maxima are identified as peaks. A parameter pg^ defines the half-width for a peak as follows:
• For each position n o In multi-dimensions, pg^n) = 1 if a(n + Δn') < a(n + Δn) for all Δn and Δn' such that Δn' is farther than Δn from n, Δn' is no farther than Wp^ from n, and
Δn and Δn' are identical except for a single dimension in which they differ by 1. For this purpose, distances may be calculated by any of a number of methods,
including the Euclidean metric, the Manhattan metric, or the maximum absolute difference in any dimension. In one dimension, π aj^n) = 1 if a^-Wpg^) < ... < a(n-2) < a(n-l) < a(n) > a(n+l) > a(n+2) > ... > a(n+wpea'c). Preferably, a(n+wpeak) > Inoise mά a(n' peak) > ^oise as well. o
= 0 otherwise.
For one-dimensional differential display data, Wpea] Δx = 0.3 nt is preferable. For hybridization, if an image has already been processed such that each point represents a different probe on the surface, then each point is regarded as a peak.
Normalization A data set is normalized by first determining the peak intensity values a(n) at locations where the peak mask π ai^n) = 1. A summary peak intensity Ipeak is calculated from the individual values. If desired, the values can be rank-ordered, and Ipeak can be defined as the 75m percentile value (75% of peaks are smaller in value; 25% of peaks are larger). Other methods include using a different percentile, for example the median, or calculating an average value. Rank-order selection methods are more robust than averages.
After a peak intensity has been calculated, the data set is rescaled by multiplying each point a(n) by the factor (Inorm^peak)' nere Inorm ls identical for each data set and sets a convenient scale. Although the precise choice for Inorm is arbitrary, a value such as 100 is convenient. The noise threshold Inoise mav a^so ^>e subject to the same normalization. Alternatively,
Inoise mav De set to a fixed value. A preferable fixed value for differential display data is Inoise = 10.
Averaging
Averaging is an operation that is applied to a collection of data sets. The average A(n) of a collection of r data sets a](n), a2(n), ... , ar(n) is calculated as
A(n) = ∑(i=l..r) wj atfn) / [ ∑(i=l..r) wj ] where WJ is a weighting applied to data set i. If ∑(i=l ... r) wj = 0, then another method must be used. It is preferable to use local information from the closest points where the summed weight does not vanish to estimate A(n). Most preferably, the value A(n) can be set equal to the
value at the closest point where the weight does not vanish. Alternatively, the unweighted values aj(n) may be used.
Any of a variety of weighting functions may be used, and may account for characteristics of a data set such as those discussed in the following. A preferred weighting function is wj = [1 - msat j(n)] where msat j(n) is the saturation mask for data set i. Weights may also incorporate error estimates from the data sets. For example, suppose that the data set aj(n) is known with statistical error ej(n). Then a maximum likelihood estimate for A(n) is obtained by minimizing a chi-square statistic χ2 = Σ(i=l ..r) [A(n) - ai(n)]2 / ei(n)2
with respect to the final average A(n) to obtain WJ = l/ej(n)2, or, if desired, wj = [1- msat i(n)]/ei(n)^- If ai(n) is itself derived from an average of other data sets, then the standard error of the mean is an appropriate choice for ej(n). If aj(n) is an unnormalized data set, then an appropriate choice for ej(n) is the background noise level Inoise defined previously. If aj(n) is a normalized data set, then it is appropriate to scale the noise level as well and use Inoise^peak I°r ej(n).
It is preferable to calculate a standard deviation SD^(n) to describe the distribution of data points leading to the average A(n). A preferable formula for SD^(n) is
SD(n) = [ ∑(i=l..r) [A(n) - aj(n)]2 / (r-l)]1 2 .
The standard error for the average is preferably calculated as E(n) = SD(n)/r1 2.
Similarity and Difference
In preparation for describing the scaling operation below, it is necessary to determine the extent of agreement between two data sets. This extent of agreement can be measured, by way of nonlimiting example, by the distance Dist[A,B] or the similarity Sim[A,B] between two data sets A(n) and B(n).
Two possible formulas for the difference Dist[A,B] are
Dist[A,B] = ∑n w[A(n),B(n)] dist[A(n),B(n)] and
Dist[A,B] = ∑n w[A(n),B(n)] dist[A(n),B(n)] / ∑n w[A(n),B(n)].
The term dist[A(n),B(n)] is a function that measures the distance between two values
A(n) and B(n). The term w[A(n),B(n)] is a mask that determines whether the data points at location n should be included in the calculation. The second formula is preferable.
A preferable formula for the distance dist(a,b) for two numbers a and b is dist(a,b) = [ln(a/b)]2, where ln() is the natural logarithm. Here a and b must be regularized to prevent values close to 0 from causing a divergence. This can be accomplished, for example, by replacing a or b by a minimum value Imm if either is smaller than Imm , or by adding a positive constant to raise all values A(n) and B(n) above 0.
Other acceptable formulas are as follows: the absolute difference, dist(a,b) = | a - b | ; the Euclidean distance, dist(a,b) = [ (a - b)2 j^2 ; the square distance, dist(a,b) = (a-b)2 ; or any non-negative function F(a,b) that is 0 when a = b and increases with increasing |a-b| or increasing |ln(a/b)|. A preferable formula for the weight w[A(n),B(n)] is w[A(n),B(n)] = wA(n)wB(n) if dist'(a,b) < Dmax and w[A(n),B(n)] = 0 otherwise,
where w^tn) and B(n) are weights for the individual data sets, dist'(a,b) is a distance measure and Dmax is some maximum distance. Possible choices for the distance measure dist'(a,b) are the same as the choices for dist, but the same distance measure need not be used for both. A preferable choice is dist'(a,b) = |ln(a b)| and Dmax = 3.
The weight w^n) is preferably [1 - n A,noise(n)][I - niA,sat(n)]> anc similarly f°r wB(n)- Other acceptable alternatives are to use either the noise mask or the saturation mask, or to use no mask and set w^(n) = wβ(n) = 1. It is also acceptable to use w[A(n),B(n)] = 1.
A similarity Sim[A,B] between two data sets A and B may be defined as
Sim[A,B] = ∑n sim[A(n),B(n)]
where the similarity sim(a,b) between two numbers a and b is larger when the quantities are larger and also larger when a and b are closer in value. A preferred method for calculating sim(a,b) is as follows. Plot the point (a,b) and measure the length r of its projection onto the line a=b and its distance d from the same line, as shown in the figure below. Then define sim(a,b) = p exp(-d2/2σ2)/[2πσ2]I'2 where p = |a + b|Λ/2 , d = |a - b|/V2 , and σ characterizes the experimental noise in the data sets.
As depicted in FIG. 3, one method to estimate an appropriate value for σ is to calculate the slope m of the best linear regression line b = m a, then calculate σ as the root mean square residual of the points (A(n),B(n)) from the line b = ma, σ = [ ∑n 2 [ (mA(n) - B(n)) / (m+1) ]2 / (r-1) ]1/2 where r-1 is the number of degrees of freedom in the fit.
Other preferable formulas for sim(a,b) are sim(a,b) = F(p) G(d) where F(p) is an increasing function of p and G(d) is a decreasing function of d.
This algorithm is related to one of the literature as a method of decomposing spectra of multicomponent mixtures into separate spectra for each of the pure components.
Equivalent procedures for evaluating the similarity and difference between data sets is encompassed within the scope of the present invention.
Scaling
Scaling is an operation that is applied to a subordinate data set a(n) to bring it in closer agreement with a master data set A(n). A scaling algorithm optimizes a scaling function s(n) to minimize the distance or maximize the similarity between the scaled slave data set s(n)a(n) and the master data set A(n).
The scaling function s(n) may have various mathematical representations. One representation is a basis set expansion
s(n) = ∑p cpφp(n) where cp is an expansion coefficient, φp(n) is the value of the pm basis function at position n, and p ranges over the P basis functions numbered p = 1 to P. A basis set expansion may be a cosine series, a sine series, or more generally, a Fourier series. For a one-dimensional data set, a preferred choice is a piecewise linear basis. The pm basis function is zero outside the interval np.i to np, with X Q taken as the left-most point of the data set and np as the right-most point. Within the interval, s(n) = cp_ι + (cp - cp_ι) (n - np_ι )/(np - np_ι ).
For a one-dimensional or multi-dimensional data set, a preferred choice is a direct product of basis functions, s(n) = ∑p cp φp(n), where here n and p are both d-dimensional and φp(n) can be expressed as
Φp(n) = φpl(nι ) φp2(n2) ... Φpd d) where n; and p; are the components of n and p in dimension j and φpj(ni) is a one- dimensional basis set in dimension j .
A preferred choice for the one-dimensional basis sets in the multidimensional direct product is an orthogonal basis. A preferred choice for an orthogonal basis is a trigonometric basis,
Φpj(n) = cos[(pj-l)π(n-nj0)/(n-nji)] , where njo and nu are the left-most and right-most points in dimension j . In one dimension, for example, with points n = 0 to 1 corresponding to distances 0 to L, the pm basis function is φp(x) = cos[(p-l)πx/L].
An advantage of this basis set is that the low-order basis functions describe low- frequency variations. Typically, the low-frequency variation in the scaling is larger in amplitude than the high-frequency variation. Therefore truncating a trigonometric basis at low order still provides a good approximation of the scaling function from a complete basis (P approaches infinity). A preferable choice is to choose a value of P such that the low-frequency noise in the
data occurs on a length scale of L/P or longer. Preferably for differential display, L = 400 nt and the noise length scale is approximately 100 nt, so P « 4 is preferable. It is this feature of a basis set such as the presently described basis set that contributes significantly to overcoming or eliminating the effects of low frequency noise. Other acceptable basis sets include, for example, polynomials (φp(n) = nP), special functions, and wavelets, and are well-known in the art See, e.g., Press et al., NUMERICAL RECIPES IN C, THE ART OF SCIENTIFIC COMPUTING, Second Edition, Cambridge Univ. Press, Cambridge UK, 1992, Chapters 5, 12 and 13.
The coefficients Cp are selected to minimize the distance Dist[A(n),s(n)a(n)] or maximize the similarity Sim[A(n),s(n)a(n)]. Methods to perform this optimization are well-known in the art. Preferable methods are conjugate direction minimization or conjugate gradient minimization, which use linear algebra to optimize the P basis set coefficients simultaneously. See, e.g., Press et al., NUMERICAL RECIPES IN C, THE ART OF SCIENTIFIC COMPUTING, Second Edition, Cambridge Univ. Press, Cambridge UK, 1992, Chapter 10. For a piecewise linear basis, a preferable approximation that is faster computationally than a full minimization is to obtain Cp from an interval surrounding np, preferably the interval from np_ι to np+ι, by minimizing the distance Dist[A(n),cpa(n)] or maximizing the similarity Sim [ A(n) ,cpa(n)] .
Preferably for differential display, the number of piecewise linear basis functions is selected so that the low-frequency noise in the data occurs on a length scale of L/(0.3P) or longer. With L = 400 nt and a noise length scale approximately 100 nt, P « 13 is preferable (interpolation points spaced every 35 nt).
Iterative Scaling
A group of data sets {aj(n)} can be brought into closer agreement with each other by first normalizing each data set, then generating an average A(n), then scaling each data set aj(n) to the average A(n), then repeating these steps. If desired, the average A(n) can be re-normalized after every iteration.
Iterations continue until a termination condition has been satisfied. A preferable termination condition is that A(n) has converged. This means that the distance Dist[A(n),A'(n)] between the value of A(n) after an iteration and its value A'(n) after the next iteration is smaller
than some threshold value. Alternatively, the scaling functions sj(n) for each of the slave data sets aj(n) can be checked for convergence. A second preferable termination condition is that a predetermined number of iterations has been reached.
It is possible to allow multiple termination conditions, with iterations ending after just one condition is satisfied.
Note that the square distance measure essentially calculates the standard deviation of the data sets. Thus, minimizing the square distance is essentially identical with performing scaling that minimizes the standard deviation of the scaled traces.
Iterative scaling may occur at a hierarchy of levels including experimental replicates, independent individuals, and groups. Recall that the data set corresponding to experimental replicate j of organism i of group A is aij(n). Similarly, the data sets bji(n) are obtained for group B, data sets cjj(n) for group C, and so forth for each of the groupings.
In one implementation of iterative scaling, data sets are normalized, scaled, and averaged within each organism, then within each group, and then between groups. One process is as follows:
• For each individual i in each group A, compute the average A}(n) by iterative scaling of the data sets ajj(n) as follows: o Each data set in ajj(n) is normalized.
o Initialize Aj(n) as the average of the experimental replicates.
o Repeat the following steps until Aj(n) has converged or the number of iterations has reached a threshold:
• Optimize the scaling function sjj(n) to bring ajj(n) into best agreement with Aj(n).
• Compute the new average Aj(n) from the scaled sets sjj(n)aij(n).
• Optionally normalize the new Aj(n). o Calculate the standard deviation SDj(n) as a measure of the experimental variance between scaled data sets.
• For each group A, compute the average A(n) by iterative scaling of the individual averages Aj(n) as follows: o Initialize A(n) as the average of the individual averages Aj(n).
o Repeat the following steps until A(n) has converged or the number of iterations has reached a threshold:
• Optimize the scaling function sj(n) to bring each Aj(n) into best agreement with the group average A(n).
• Compute the new average A(n) from the scaled individual averages si(n)Aj(n). • Optionally, normalize the new A(n). o Calculate the standard deviation SDA(Π) as a measure of the variance between scaled individual averages.
• For the final scaling of groups to each other, perform one of the following two operations: o Option 1 : scale by computing the grand mean M(n) from all the group averages
A(n), B(n), ... , as follows:
• Initialize the grand mean M(n) as the average of A(n), B(n), ... .
• Repeat the following steps until M(n) has converged or the number of iterations has reached a threshold: • Optimize the scaling functions SA(Π), sg(n), ... , that bring A(n),
B(n), ... , into best agreement with M(n).
• Compute the new M(n) from the scaled group averages SA.(n)A(n),
• Optionally, normalize the new M(n)
• Calculate the standard deviation SD(n) as a measure of the variance between scaled group averages.
o Option 2: select one of the groups R(n) as a reference and scale the remaining groups to R(n). Calculate the standard deviation SD(n) as a measure of the variance between the scaled group averages.
Using this process, the scaling terms must be back-propagated to compare averages other than the final, scaled group averages. For example, if scaled group averages are required, s(n)A(n) is used. If scaled individual averages are required, then s(n)sj(n)A}(n) is used. If scaled data sets are required, then s(n)si(n)sjj(n)ajj(n) is used.
In a second implementation, intermediate averages are not required. This implementation requires that a weighting method be selected. With a preferred weighting method, each data set from individual i is preferably given a weight proportional to 1 /(number of replicates from individual i) . This gives each individual equal weight and prevents an individual with many replicates from dominating the average. Other preferable methods include weighting each data set equally and weighting each data sets to give each group equal weight. If each group has equal weight, one method is to weight each replicate equally. Thus, each data set from group A is given a weight proportional to 1 /(number of replicates from all the individuals belonging to group A). An alternate method is to weight each data set variably to give each individual within a group equal weight. Thus, each data set from individual i of group A is given a weight proportional to 1 /[(number of individuals in group A)(number of replicates in individual i)].
After selecting a weighting method, apply the following algorithm: • Initialize the grand mean M(n) by averaging all the data sets ajj(n) according to the selected weighting method..
• Repeat the following steps until M(n) has converged or the number of iterations reaches a threshold: o Optimize the scaling functions sji(n) to bring each ajj(n) into best agreement with M(n). o Compute the new M(n) from the scaled data sets Sij(n)a}j(n) using the selected weighting method. o Optionally, normalize the new M(n).
• Calculate the individual averages Aj(n) by averaging the scaled replicates sij(n)ajj(n) belonging to individual i with weights according to the selected weighting method. Calculate the standard deviation SDj(n) within each individual.
• Calculate the group averages A(n), B(n), ... , by averaging the individual averages Aj(n) according to the selected weighting method. Calculate the standard deviation SDA(Π),
SDg(n), ... , within each group.
For each of the iterative scaling steps, a preferable threshold for differential display data is 2 iterations.
Jiggling One aspect of difference finding is comparing the heights of peaks in two data sets. In many data sets, the same peak may occur at different positions in different data sets. For example, a peak in one replicate of data set may occur at position n, while in a second data set it may occur at position n+1 or n-1 due to experimental variability. (See FIG. 1)
A jiggling algorithm identifies the peak height in a data set a(n) that corresponds to a given location n'. A preferred jiggling algorithm requires a parameter wjjggie, which describes the width of the jiggling window.
The preferred algorithm starts at position n' and searches for the peak in a(n) closest to n' and within distance wjiggιe. The height of a(n) at this peak position is termed the jiggled height of a(n) at n'. If two peaks are within equal distance, the higher value is preferably taken as the jiggled height. If there is no peak within distance wjigg'e, then the height a(n') is the jiggled height.
For a one-dimensional data set, for example, the data range for the jiggling peak search is n'-wjiggie through n'+wjiggie. If n' is a peak in a(n), then the value a(n') is the jiggled height of a(n) at n'. Otherwise the positions n'±l, n' ±2, ..., n'±Wμggie are tested in turn for peaks; if a peak is found at location n", then a(n") is the jiggled height of a(n) at n'. If no peak is found, then a(n') is the jiggled height.
Less preferably, all of the peaks within distance w;jggie of n' are examined and the maximum value is taken as the jiggled height of a(n) at n'. For a one-dimensional data set, all of the peaks in the window n'-wjjggie, ... , n'-l, n', n'+l, ... , n'+w;jgg'e are examined and the
maximum value is taken as the jiggled height of a(n) at n'. If there is no peak in the interval, then a(n') is the jiggled height.
For differential display data, a preferable value is wjjggieΔx = 0.4 nt.
Difference Finding Difference finding identifies locations where at least one of the groups has a peak and its value is significantly different from the other groups. The group averages and individual averages produced by iterative scaling serve as inputs to difference finding.
It is preferable to employ an algorithm that uses jiggling to identify corresponding peaks in different data sets. This avoids spurious differences due to slight offsets in peak positions. It is also preferable to employ an algorithm that identifies at most one difference from peaks that correspond. A preferable method employs a parameter Wrjjff that defines the minimum distance between differences. A preferable choice is wjjff > Wpg^. For differential display data, a preferable value is WfjjffΔx = 1.1 nt.
A preferred algorithm is as follows: • Generate a master peak mask Mpeak(n) using one of the following alternatives: o Option 1 : For each group A and individual Aj, calculate a peak mask mpeaj (n) from the individual average Aj(n). Then, for each position n, Mpeaic(n) is 1 if at least one of the individual peak masks is 1 and is 0 otherwise. o Option 2: Calculate the peak mask Mpea]f(n) directly from the grand mean M(n) of all the groups. o Option 3: If there are only two groups A and B, generate a peak mask from the difference A(n) — B(n).
• For each position n that appears in the peak mask Mpeaic(n), calculate the significance of a difference as follows: o For each individual i in each group A, find the jiggled height of Ai at n.
o Calculate the group averages based on the jiggled heights.
o Perform an F-test that compares the variance between group averages with the variance within groups.
o Associate the p-value of the F-test with the position n. Also record the number of samples that had a peak at position n.
• Generate a list of peak positions sorted from lowest p-value to highest p-value. If two positions have the same p-value, break the tie by listing first the position with more sample peaks. Break any remaining ties by listing first the lower position.
• Repeat the following steps until the list is empty: o Remove the first element from the list and record its position n as a difference.
o Strike out any remaining elements in the list that are within distance w^ff of n. For a one-dimensional data set, for example, strike out any differences at positions n±l, n±2, ... , n±w^if .
For difference finding with two groups A and B, a pooled variance t-test may be employed instead of an F-test.
A less preferred algorithm for a comparison between two groups, A and B, and one- dimensional data set, is as follows: • Perform the final step of iterative scaling by scaling A(n) to B(n).
• Calculate the peak mask Mpg^n) from the difference A(n) — B(n).
• Generate a list of peak positions sorted from smallest n to largest n.
• Initialize a variable LASTPOSITION as 0 and a variable LASTDIRECTION as 0.
• Repeat the following steps until the list of peaks is empty: o Remove the first element n from the list and calculate p-value(n) from a t-test comparing the jiggled heights of individuals from group A to the jiggled heights of individuals from group B at position n. If the average of the group A heights is larger than the average of the group B heights, then direction(n) = +1; otherwise, direction(n) = -l . o If direction(n) is not equal to LASTDIRECTION, or if (n — LASTPOSITION) > wdiff> then
• If LASTPOSITION is not 0, save LASTPOSITION as a difference with p- value equal to LASTPVALUE
• Update LASTPOSITION = n, LASTDIRECTION = direction(n), and
LASTPVALUE = p-value(n). o Otherwise if p-value(n) < LASTPVALUE then
• Update LASTPOSITION = n, LASTDIRECTION = direction(n), and LASTPVALUE - p-value(n). o Otherwise
• Continue with the next peak from the list.
• If LASTDIRECTION is not 0, then save the final difference at position LASTPOSITION with p-value equal to LASTPVALUE.
EXAMPLE
Example 1. Differential Gene Expression in Phenobarbitol-Treated Rats
Male Sprague-Dawley rats (Harlan Sprague Dawley, Inc., Indianapolis, Indiana) of 10-14 weeks of age were gavage-fed and dosed with phenobarbitol once a day for three days at a dose of 3.81 mg/kg/day. The drug was dissolved in sterile water prior to treatment. This dosage corresponds to the ED 100 (the upper limit of the effective dose for humans) adjusted for the difference in metabolic rate between rats and humans. Three rats were used for the drug treatment group, and an additional three rats were treated with sterile water to serve as the control group.
Rats were sacrificed 24 hours after the final dose and their brains were harvested. Collection of mRNA from the harvested brains, synthesis of the corresponding cDNA, and differential display protocols were done as has been described elsewhere. See, U. S. Patent No. 5,871,697; Shimkets et al. Nat. Biotechnol. 17: 798-803 (1999).
Three experimental replicate raw traces were collected from each of two individual animals, one treated with phenobarbital and the other treated with sterile water. Data points were collected for fragments from 30 nt to 450 nt in length, and the data sets were discretized for a point every 0.1 nt. The averages of the 3 replicate raw traces for each individual, prior to normalization or scaling, are shown in FIG. 4. Each trace was weighted equally, and no data were masked for either noise or saturation.
Next, noise was masked using Inoise = 500 and wsat = 5, and peaks were identified in each replicate using wpeak = 3 with the condition that all 7 points contributing to a peak be above the noise level and not saturated. The peaks were sorted in increasing order of height, and each replicate was normalized to give the 75th percentile peak an intensity of 100. The 3 normalized traces were averaged for each individual, and the individual averages are displayed in FIG. 5.
For all the scaling operations that follow, the basis set used was piecewise linear with 13 scaling points located every 35 nt beginning at 30 nt and ending at 450 nt. The distance function was [ln(a(n)/A(n))]2, and points for which |ln(a(n)/A(n)]| > 3 were masked out.
The first step in the scaling procedure was that the 3 normalized traces for each individual were averaged, scaled to the average, re-averaged, then re-scaled to the average for 2 rounds of iterative scaling. Next, the phenobarbital-treated individual average and the sterile-water-treated individual average were themselves averaged, the individual averages scaled to the grand average, then the process repeated for 2 rounds of iterative scaling. Finally, the phenobarbital- treated individual average was scaled to the sterile-water-treated individual average. The final scaling factors are shown in FIG. 6 for the two individuals. The final individual averages are shown in FIG. 7.
With both the normalized traces and the normalized and scaled traces, difference finding was performed using wpeak = 2, wjigg]e = 4, and wdiff = 11. Significance levels were calculated as 1 - p-value from a t-test based on the scaled replicate traces. See Table 1, below. Only differences with a significance greater than 0.9 and a ratio |ln(Phenobarbital/ Water) | > ln(l .5) were retained. The differences identified are listed in the table below (significances greater that 0.99 are reported as 1). The normalized traces generated 35 differences, whereas the scaled traces generated 18 differences, of which 12 were in common with the normalized traces. The differences in the normalized traces that are removed by scaling tend to have lower significance than the differences that are retained, indicating that scaling helps identify the differences with greater support in the data..
Table 1
Length t-test (1 - p-value)
Normalized Normalized and Only Scaled
51.9 1 1
53.6 1 1
67.3 0.94
81.8 0.99 0.97
87.1 - 1
103.2 - 1
120.8 1 0.95
127.5 1
154.6 - 1
158.9 - 1 165.0 0.98
170.0 1 0.99
190.1 0.97
205.7 1 1 218.6 1 0.91 228.1 1
233.0 0.98
263.1 0.94
274.1 1 1
280.0 0.91
303.2 - 0.99
331.8 1 1
340.6 0.98 0.9 352.2 - 1
353.7 0.99
354.9 0.96
388.8 1 1 394.5 1
395.7 1 0.99
402.4 0.96
404.7 1
406.4 1
437.7 0.99
443.1 0.98
447.9 1
EQUIVALENTS
From the foregoing detailed description of the specific embodiments of the invention, it should be apparent that unique methods for identifying differences between at least two data sets have been described. Although particular embodiments have been disclosed herein in detail, this has been done by way of example for purposes of illustration only, and is not intended to be limiting with respect to the scope of the appended claims which follow. In particular, it is
contemplated by the inventor that various substitutions, alterations, and modifications may be made to the invention without departing from the spirit and scope of the invention as defined by the claims. For instance, the choice of algorithms used to transform data sets, such as normalization calculations, averaging calculations or scaling calculations, or the choice of data sets to be analyzed is believed to be a matter of routine for a person of ordinary skill in the art with knowledge of the embodiments described herein.
Claims
1. A method of identifying a difference between at least two groups, the method comprising: a) providing a first group having one or more elements in a first data set; b) applying at least one transformation to said first data set to provide a transformed data set, wherein said transformation is a calculation selected from a normalizing calculation, an averaging calculation and a scaling calculation; and c) distinguishing differences, if present, between elements of said first transformed data set and a second groups having one or more elements in a second data set; thereby identifying a difference between the groups.
2. The method of claim 1 , wherein said second data set is a transformed data set.
3. The method of claim 1 , wherein said transformation minimizes noise associated with one or more signals associated with elements of at least one data set.
4. The method of claim 3, wherein said noise comprises low frequency noise.
5. The method of claim 3, wherein said noise comprises jiggle.
6. The method of claim 5, wherein said jiggle includes positional shifts of corresponding elements between said first and second data sets.
7. The method of claim 1 , wherein elements of at least one data set comprise a trace.
8. The method of claim 7, wherein said trace represents the result of an electrophoretogram or a chromatogram.
9. The method of claim 1, wherein elements of at least one data set comprise one or more positions in an array.
10. The method of claim 9, wherein any one position in the array is used to determine an extent of matching between a reagent affixed to the array at said position and a sample contacting the array position.
11. The method of claim 10, wherein the reagent is a first nucleic acid and the sample comprises a second nucleic acid, wherein the second nucleic acid is in a hybridization solution mixture, and the matching constitutes hybridization of a sample nucleic acid to the affixed nucleic acid.
12. The method of claim 1, wherein at least one data set comprises elements derived from an analysis of one or more differentially expressed nucleic acids.
13. The method of claim 1, wherein at least one data set is derived from an analysis of one or more individuals in a group.
14. The method of claim 13, wherein at least one data set is subjected to a masking operation.
15. The method of claim 13, wherein the elements in the data sets comprise values derived from an analysis of differentially expressed nucleic acids in the plurality of individuals in a group.
16. The method of claim 15, wherein a data set obtained for each group is obtained by applying at least one transformation to the data set from each individual, wherein said transformation is selected from the group consisting of a normalizing calculation, an averaging calculation and a scaling calculation.
17. The method of claim 15, wherein each individual provides at least one replicate sample.
18. The method of claim 15, wherein the data set from any one individual or replicate comprises a data set derived from a trace of an electropherogram or a chromatogram.
19. The method of claim 18, wherein said data set is transformed by applying at least one of a normalizing calculation, an averaging calculation and a scaling calculation to the data set from each individual or replicate.
20. The method of claim 18, wherein said data set is discretized prior to the normalizing, the scaling and/or the averaging calculation.
21. The method of claim 1 , wherein said transformation comprises a normalization calculation.
22. The method of claim 21 , wherein said normalization calculation comprises adjusting each data set such that a subset of elements in each data set has similar or identical values.
23. The method of claim 21 , wherein at least one data set is subjected to a masking operation.
24. The method of claim 1, wherein said transformation comprises an averaging calculation.
25. The method of claim 24, wherein said averaging calculation comprises calculating an average for a location or for a discretized position of said first and second data sets, and wherein the average may be an unweighted average or a weighted average.
26. The method of claim 1, wherein said transformation comprises a scaling calculation.
27. The method of claim 26, wherein said scaling calculation comprises a calculation that causes a first data set to resemble a second data set, provided that an element in the scaled first data set whose intensity differs significantly from the intensity of the element in the second data set at the same location or the same position contributes to identifying the difference between the data sets.
28. The method of claim 27, wherein scaling comprises calculating a scaling function based on optimization of a distance between the data sets.
29. The method of claim 27, wherein scaling comprises calculating a scaling function based on optimization of a similarity between the data sets.
30. The method of claim 27, wherein the scaling calculation employs a scaling function.
31. The method of claim 30, wherein the scaling function is a basis set expansion.
32. The method of claim 31 , wherein the basis set is a piecewise linear basis set.
33. The method of claim 31 , wherein the basis set is a Fourier series.
34. The method of claim 30, wherein the scaling function is a direct product of basis functions.
35. The method of claim 1, wherein the transformation comprises successive iterations of a cycle of calculations that are carried out until a specified termination condition has been satisfied, wherein the calculations comprise at least one of a normalization calculation, an averaging calculation and a scaling calculation.
36. The method of claim 35, wherein the termination condition is met when a transformed data set has converged.
37. The method of claim 35, wherein the termination condition is met when a predetermined number of iterations of a cycle of calculations has been reached.
38. The method of claim 1 , wherein elements within any two or more data sets have been corrected for signal alignment.
39. The method of claim 1, wherein the distinguishing of differences between the elements of the resulting data sets comprises an application of at least one difference finding algorithm.
40. A display means displaying a representation of a difference between two or more transformed data sets, wherein each data set comprises ordered elements; and the data sets are transformed by at least one calculation selected from the group consisting of a normalizing calculation, an averaging calculation and a scaling calculation.
41. The display means of claim 40, wherein the representation is obtained by a process comprising the steps of: a) providing a first group having one or more elements in a first data set; b) applying at least one transformation to said first data set to provide a transformed data set, wherein said transformation is a calculation selected from a normalizing calculation, an averaging calculation and a scaling calculation; and c) distinguishing differences, if present, between elements of said first transformed data set and a second groups having one or more elements in a second data set; thereby identifying a difference between the groups.
42. The display means of claim 41, wherein said second data set is a transformed data set.
43. The display means of claim 41 , wherein said transformation minimizes noise associated with one or more signals associated with elements of at least one data set.
44. The display means of claim 43, wherein said noise comprises low frequency noise.
45. The display means of claim 43, wherein said noise comprises jiggle.
46. The display means of claim 45, wherein said jiggle includes positional shifts of corresponding elements between said first and second data sets.
47. The display means of claim 41, wherein elements of at least one data set comprise a trace.
48. The display means of claim 47, wherein said trace represents the result of an electrophoretogram or a chromatogram.
49. The display means of claim 41 , wherein elements of at least one data set comprise one or more positions in an array.
50. The display means of claim 49, wherein any one position in the array is used to determine an extent of matching between a reagent affixed to the array at said position and a sample contacting the array position.
51. The display means of claim 50, wherein the reagent is a first nucleic acid and the sample comprises a second nucleic acid, wherein the second nucleic acid is in a hybridization solution mixture, and the matching constitutes hybridization of a sample nucleic acid to the affixed nucleic acid.
52. The display means of claim 41 , wherein at least one data set comprises elements derived from an analysis of one or more differentially expressed nucleic acids.
53. The display means of claim 41, wherein at least one data set is derived from an analysis of one or more individuals in a group.
54. The display means of claim 53, wherein at least one data set is subjected to a masking operation.
55. The display means of claim 53, wherein the elements in the data sets comprise values derived from an analysis of differentially expressed nucleic acids in the plurality of individuals in a group.
56. The display means of claim 55, wherein a data set obtained for each group is obtained by applying at least one transformation to the data set from each individual, wherein said transformation is selected from the group consisting of a normalizing calculation, an averaging calculation and a scaling calculation.
57. The display means of claim 55, wherein each individual provides at least one replicate sample.
58. The display means of claim 55, wherein the data set from any one individual or replicate comprises a data set derived from a trace of an electropherogram or a chromatogram.
59. The display means of claim 58, wherein said data set is transformed by applying at least one of a normalizing calculation, an averaging calculation and a scaling calculation to the data set from each individual or replicate.
60. The display means of claim 58, wherein said data set is discretized prior to the normalizing, the scaling and/or the averaging calculation.
61. The display means of claim 41 , wherein said transformation comprises a normalization calculation.
62. The display means of claim 61 , wherein said normalization calculation comprises adjusting each data set such that a subset of elements in each data set has similar or identical values.
63. The display means of claim 61 , wherein at least one data set is subjected to a masking operation.
64. The display means of claim 41 , wherein said transformation comprises an averaging calculation.
65. The display means of claim 64, wherein said averaging calculation comprises calculating an average for a location or for a discretized position of said first and second data sets, and wherein the average may be an unweighted average or a weighted average.
66. The display means of claim 41 , wherein said transformation comprises a scaling calculation.
67. The display means of claim 66, wherein said scaling calculation comprises a calculation that causes a first data set to resemble a second data set, provided that an element in the scaled first data set whose intensity differs significantly from the intensity of the element in the second data set at the same location or the same position contributes to identifying the difference between the data sets.
68. The display means of claim 67, wherein scaling comprises calculating a scaling function based on optimization of a distance between the data sets.
69. The display means of claim 67, wherein scaling comprises calculating a scaling function based on optimization of a similarity between the data sets.
70. The display means of claim 67, wherein the scaling calculation employs a scaling function.
71. The display means of claim 70, wherein the scaling function is a basis set expansion.
72. The display means of claim 71, wherein the basis set is a piecewise linear basis set.
73. The display means of claim 71 , wherein the basis set is a Fourier series.
74. The display means of claim 70, wherein the scaling function is a direct product of basis functions.
75. The display means of claim 41, wherein the transformation comprises successive iterations of a cycle of calculations that are carried out until a specified termination condition has been satisfied, wherein the calculations comprise at least one of a normalization calculation, an averaging calculation and a scaling calculation.
76. The display means of claim 75, wherein the termination condition is met when a transformed data set has converged.
77. The display means of claim 75, wherein the termination condition is met when a predetermined number of iterations of a cycle of calculations has been reached.
78. The display means of claim 41 , wherein elements within any two or more data sets have been corrected for signal alignment.
79. The display means of claim 41 , wherein the distinguishing of differences between the elements of the resulting data sets comprises an application of at least one difference finding algorithm.
80. A representation of a difference between normalized, averaged and scaled data sets, wherein each data set comprises ordered elements.
81. The representation of claim 80 wherein the representation is obtained by a process comprising the steps of a) providing a first group having one or more elements in a first data set; b) applying at least one transformation to said first data set to provide a transformed data set, wherein said transformation is a calculation selected from a normalizing calculation, an averaging calculation and a scaling calculation; and c) distinguishing differences, if present, between elements of said first transformed data set and a second groups having one or more elements in a second data set; thereby identifying a difference between the groups.
82. The representation of claim 81 , wherein said second data set is a transformed data set.
83. The representation of claim 81 , wherein said transformation minimizes noise associated with one or more signals associated with elements of at least one data set.
84. The representation of claim 83, wherein said noise comprises low frequency noise.
85. The representation of claim 83, wherein said noise comprises jiggle.
86. The representation of claim 85, wherein said jiggle includes positional shifts of corresponding elements between said first and second data sets.
87. The representation of claim 81 , wherein elements of at least one data set comprise a trace.
88. The representation of claim 87, wherein said trace represents the result of an electrophoretogram or a chromatogram.
89. The representation of claim 81 , wherein elements of at least one data set comprise one or more positions in an array.
90. The representation of claim 89, wherein any one position in the array is used to determine an extent of matching between a reagent affixed to the array at said position and a sample contacting the array position.
91. The representation of claim 90, wherein the reagent is a first nucleic acid and the sample comprises a second nucleic acid, wherein the second nucleic acid is in a hybridization solution mixture, and the matching constitutes hybridization of a sample nucleic acid to the affixed nucleic acid.
92. The representation of claim 81 , wherein at least one data set comprises elements derived from an analysis of one or more differentially expressed nucleic acids.
93. The representation of claim 81 , wherein at least one data set is derived from an analysis of one or more individuals in a group.
94. The representation of claim 93, wherein at least one data set is subjected to a masking operation.
95. The representation of claim 93, wherein the elements in the data sets comprise values derived from an analysis of differentially expressed nucleic acids in the plurality of individuals in a group.
96. The representation of claim 95, wherein a data set obtained for each group is obtained by applying at least one transformation to the data set from each individual, wherein said transformation is selected from the group consisting of a normalizing calculation, an averaging calculation and a scaling calculation.
97. The representation of claim 95, wherein each individual provides at least one replicate sample.
98. The representation of claim 95, wherein the data set from any one individual or replicate comprises a data set derived from a trace of an electropherogram or a chromatogram.
99. The representation of claim 98, wherein said data set is transformed by applying at least one of a normalizing calculation, an averaging calculation and a scaling calculation to the data set from each individual or replicate.
100. The representation of claim 98, wherein said data set is discretized prior to the normalizing, the scaling and/or the averaging calculation.
101. The representation of claim 81 , wherein said transformation comprises a normalization calculation.
102. The representation of claim 101, wherein said normalization calculation comprises adjusting each data set such that a subset of elements in each data set has similar or identical values.
103. The representation of claim 101, wherein at least one data set is subjected to a masking operation.
104. The representation of claim 81 , wherein said transformation comprises an averaging calculation.
105. The representation of claim 104, wherein said averaging calculation comprises calculating an average for a location or for a discretized position of said first and second data sets, and wherein the average may be an unweighted average or a weighted average.
106. The representation of claim 81, wherein said transformation comprises a scaling calculation.
107. The representation of claim 106, wherein said scaling calculation comprises a calculation that causes a first data set to resemble a second data set, provided that an element in the scaled first data set whose intensity differs significantly from the intensity of the element in the second data set at the same location or the same position contributes to identifying the difference between the data sets.
108. The representation of claim 107, wherein scaling comprises calculating a scaling function based on optimization of a distance between the data sets.
109. The representation of claim 107, wherein scaling comprises calculating a scaling function based on optimization of a similarity between the data sets.
110. The representation of claim 107, wherein the scaling calculation employs a scaling function.
111. The representation of claim 110, wherein the scaling function is a basis set expansion.
112. The representation of claim 111, wherein the basis set is a piecewise linear basis set.
113. The display means of claim 111, wherein the basis set is a Fourier series.
114. The representation of claim 110, wherein the scaling function is a direct product of basis functions.
115. The representation of claim 81 , wherein the transformation comprises successive iterations of a cycle of calculations that are carried out until a specified termination condition has been satisfied, wherein the calculations comprise at least one of a normalization calculation, an averaging calculation and a scaling calculation.
116. The representation of claim 115, wherein the termination condition is met when a transformed data set has converged.
117. The representation of claim 115, wherein the termination condition is met when a predetermined number of iterations of a cycle of calculations has been reached.
118. The representation of claim 81 , wherein elements within any two or more data sets have been corrected for signal alignment.
119. The representation of claim 81 , wherein the distinguishing of differences between the elements of the resulting data sets comprises an application of at least one difference finding algorithm.
Applications Claiming Priority (5)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US477273 | 1995-06-07 | ||
| US11480699P | 1999-01-05 | 1999-01-05 | |
| US114806P | 1999-01-05 | ||
| US47727300A | 2000-01-04 | 2000-01-04 | |
| PCT/US2000/000167 WO2000041122A2 (en) | 1999-01-05 | 2000-01-05 | Normalization, scaling, and difference finding among data sets |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP1141878A2 true EP1141878A2 (en) | 2001-10-10 |
Family
ID=26812554
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP00903107A Withdrawn EP1141878A2 (en) | 1999-01-05 | 2000-01-05 | Normalization, scaling, and difference finding among data sets |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP1141878A2 (en) |
| JP (1) | JP2002539768A (en) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| IL259466B (en) * | 2015-11-20 | 2022-09-01 | Seegene Inc | A method for calibrating a target analyte data set |
-
2000
- 2000-01-05 EP EP00903107A patent/EP1141878A2/en not_active Withdrawn
- 2000-01-05 JP JP2000592779A patent/JP2002539768A/en not_active Withdrawn
Non-Patent Citations (1)
| Title |
|---|
| See references of WO0041122A3 * |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2002539768A (en) | 2002-11-26 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Gottardo et al. | Bayesian robust inference for differential gene expression in microarrays with multiple samples | |
| Lai et al. | Comparative analysis of algorithms for identifying amplifications and deletions in array CGH data | |
| Franks et al. | Feature specific quantile normalization enables cross-platform classification of molecular subtypes using gene expression data | |
| Mitteroecker et al. | Multivariate analysis of genotype–phenotype association | |
| Calza et al. | Normalization of oligonucleotide arrays based on the least-variant set of genes | |
| Frost et al. | Principal component gene set enrichment (PCGSE) | |
| Borisov et al. | Shambhala: a platform-agnostic data harmonizer for gene expression data | |
| Bolstad et al. | Preprocessing high-density oligonucleotide arrays | |
| CN116364180B (en) | A Cell Type Unbiased Localization Method and System Based on Spatial Transcriptomics | |
| Ma et al. | Belayer: Modeling discrete and continuous spatial variation in gene expression from spatially resolved transcriptomics | |
| Giai Gianetto | Statistical analysis of post-translational modifications quantified by label-free proteomics across multiple biological conditions with R: illustration from SARS-CoV-2 infected cells | |
| WO2000041122A2 (en) | Normalization, scaling, and difference finding among data sets | |
| Lees et al. | Novel methods for secondary structure determination using low wavelength (VUV) circular dichroism spectroscopic data | |
| Malik et al. | Restricted maximum-likelihood method for learning latent variance components in gene expression data with known and unknown confounders | |
| Guerrero Montero et al. | Self-contained Beta-with-Spikes approximation for inference under a Wright–Fisher model | |
| JP2002539768A (en) | Normalize, scale and find differences between datasets | |
| Hellicar et al. | Machine learning approach for pooled DNA sample calibration | |
| Nieto-Reyes et al. | Statistical depth based normalization and outlier detection of gene expression data | |
| Anglada-Girotto et al. | robustica: customizable robust independent component analysis | |
| Rensink et al. | Statistical issues in microarray data analysis | |
| Lesne et al. | Probability landscapes for integrative genomics | |
| Trussart et al. | Fast, accurate and scalable normalization for RNA sequencing data with the RUVprps software package | |
| Marimuktu | Review on gene expression meta-analysis: techniques and implementations | |
| Sadygov et al. | Exact Integral Formulas for False Discovery Rate and the Variance of False Discovery Proportion | |
| Song et al. | scAEQN: A Batch Correction Joint Dimension Reduction Method on scRNA-seq Data |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 20010706 |
|
| AK | Designated contracting states |
Kind code of ref document: A2 Designated state(s): AT BE CH CY DE DK ES FI FR GB GR IE IT LI LU MC NL PT SE |
|
| AX | Request for extension of the european patent |
Free format text: AL;LT;LV;MK;RO;SI |
|
| 17Q | First examination report despatched |
Effective date: 20011001 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN |
|
| 18D | Application deemed to be withdrawn |
Effective date: 20040214 |