WO2007134167A2 - Computational analysis of the synergy among multiple interacting factors - Google Patents
Computational analysis of the synergy among multiple interacting factors Download PDFInfo
- Publication number
- WO2007134167A2 WO2007134167A2 PCT/US2007/068666 US2007068666W WO2007134167A2 WO 2007134167 A2 WO2007134167 A2 WO 2007134167A2 US 2007068666 W US2007068666 W US 2007068666W WO 2007134167 A2 WO2007134167 A2 WO 2007134167A2
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- genes
- factors
- module
- gene expression
- data
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6876—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
- C12Q1/6883—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2600/00—Oligonucleotides characterized by their use
- C12Q2600/158—Expression markers
Definitions
- the disclosed subject matter relates generally to systems and methods for factor selection, including factors useful in gene expression analysis.
- Gene selection techniques based on microarray analysis often involve individual gene ranking depending on a numerical score measuring the correlation of each gene with particular disease types.
- the expression levels of the highest-ranked genes tend to be either consistently higher in the presence of disease and lower in the absence of disease, or vice versa.
- Such genes usually have the property that their joint expression levels corresponding to diseased tissues and the joint expression levels corresponding to healthy tissues can be cleanly separated into two distinct clusters. These techniques are therefore convenient for classification purposes between disease and health, or between different disease types. However, they do not identify cooperative relationships or the synergy among multiple interacting genes.
- the disclosed subject matter provides techniques for the analysis of the cooperative interactions or synergy among multiple interacting factors.
- the factors can be features, elements or outcomes that are cooperatively associated with one or more factors or outcomes by their joint presence or absence.
- methods for selecting factors from a data set of measurements are provided.
- the measurements can include values of factors and/or outcomes.
- Two or more factors that are jointly associated with one or more outcomes from the data set are identified.
- Each of the two or more factors are analyzed to determine at least one cooperative interaction among the factors with respect to an outcome or factor.
- the factors can be a module of factors or a sub-module of factors.
- the interaction can be a structure of interactions.
- the factors can include two or more genes.
- the data set can include gene expression data including expression levels for each of the genes.
- the outcomes includes presence or absence of a disease, and the genes can be a module of genes, preferably a smallest cooperative module of genes with joint expression levels that can be used for a prediction of the presence of a disease with high accuracy.
- the gene expression data can include expression levels for each of the two or more genes.
- the method includes providing gene expression data for two or more genes, discretizing the gene expression data, identifying the two or more genes with a high synergy, and identifying a cooperative interaction that connects the expression levels in the two or more genes with presence or absence of a disease.
- the gene expression data can be derived from at least one microarray of gene expression data.
- the genes can be a module of genes.
- the cooperative interaction can be modeled using a Boolean function, a parsimonious Boolean function, or the most parsimonious Boolean function.
- the high synergy can be a maximum synergy.
- the gene expression data includes expression levels for each of the two or more genes.
- the system includes at least one processor, a computer readable medium coupled to the processor including instructions which when executed cause the processor to provide gene expression data for the genes.
- the instructions also cause the processor to discretize the gene expression data, choose a single threshold for each of the two or more genes, identify the two or more genes with a high synergy, and identify a cooperative interaction that connects the expression levels in the two or more genes with presence of a disease.
- the gene expression data can be derived from a microarray of gene expression data.
- the two or more genes can be a module of genes.
- the high synergy can be a maximum synergy.
- systems for selecting factors from a data set of measurements are provided. Each measurement can include values of the factors and outcomes.
- the system includes at least one processor and a computer readable medium coupled to the at least one processor.
- the computer readable medium includes instructions which when executed cause the processor to identify two or more factors that are jointly associated with one or more outcomes or factors from the data, and, analyze each of the two or more factors to determine at least one cooperative interaction among the factors with respect to an outcome or factor.
- the factors can be a module of factors.
- the module of factors can include at least one sub-module of factors.
- the cooperative interaction can include a structure of interactions, such as a logic function.
- the two or more factors can be two or more genes
- the data can include gene expression data including expression levels
- the outcomes can be the presence or absence of a disease.
- the genes can be a module of genes.
- the module of genes can include at least one sub-module of genes.
- the module of genes can be a smallest module of genes with joint expression levels that can be used for a prediction of the presence of disease.
- Fig. 1 is a diagram according to some embodiments of the disclosed subject matter.
- Fig. 2 is a diagram according to some embodiments of the disclosed subject matter.
- Fig. 3 is a diagram according to some embodiments of the disclosed subject matter.
- Fig. 4 is a schematic diagram according to some embodiments of the disclosed subject matter.
- a method for selecting factors from a data set of measurements includes identifying factors cooperatively associated with an outcome from a data set for two or more factors, and analyzing each of the factors or modules of factors to determine cooperative or synergistic interactions among the factors or outcomes with respect to the outcome.
- the data set can be a set of measurements that includes values of the factors and the outcomes.
- the factors are genes and the inference of the cooperative relationship or synergy among multiple interacting genes is desired.
- disease data it can be more generally applicable to other data sets.
- Other applicable data sets include other biological data such as how cells are influenced by stimuli jointly, financial data, internet traffic data, scheduling data for industries, marketing data, and manufacturing data, for example.
- Table 1, below, specifies other data sets relevant to various objectives, including factors and outcomes, to which the disclosed subject matter can also be applied: TABLE 1
- gene expression data for two or more genes is provided 101 in the form of, for example, a microarray of expression data.
- the gene expression data is then discretized 102.
- the data can be binarized into two levels.
- Other levels of discretization, such as trinarization, can also be used.
- each gene can be assumed to be either expressed or not expressed in a particular tissue. It can also be assumed that there are two types of tissues, either healthy ones or tissues suffering from a particular disease.
- Table 2 shows results from hypothetical microarray measurements of three genes, Gj, Gi, and G 3 in both the presence and absence of a particular cancer C. N 0 and Ni represent the presence and absence of cancer, respectively.
- the latter assumption can also be generalized to include more than two types of tissues, or modified to be used for classification among several types of cancer.
- two or more genes with a high synergy are identified 103, as detailed below.
- Cooperative or synergistic interactions between multiple genes can then be determined 104, as shown in Figure 1. Because systems biology is based on a holistic view of biological systems, synergy is important.
- a gene set of n genes can have expression levels -.,G n and a particular outcome C can be any phenotype, such as the presence of a particular disease or the differentiation of stem cells into a particular cell type when analyzing expression data of human tissues.
- the synergy Syn(Gi,G 2 ,...,G n ;C) of the gene set with respect to the phenotype C is defined by equation (1):
- Equation (2) is symmetric with respect to the three random variables and equal to the opposite of the mutual information /(Gi;G 2 ;C) common to the three variables Gi,G 2 ,C. Contrary to the mutual information common to two variables, the mutual information common to three variables is not necessarily a nonnegative quantity. This allows for a strictly positive synergy.
- Positive synergy implies some form of direct or indirect interaction of the genes, as a system. This definition of synergy also allows for insight into the structure of potential pathways by making iterative use of the "synergistic partition," defined earlier, to generate a hierarchical decomposition of the gene set into smaller modules.
- a rooted and not necessarily binary tree with n leaves, each of which represents one of the genes.
- Each node of the tree represents a subset or sub-module of genes, which contains the genes represented by the leaves of the clade formed by the node. Therefore the root represents the whole gene set.
- the synergistic partition of the whole gene set can then be represented by the branching of the root, so that the nodes that are neighboring to the root represent the gene subsets or sub-modules defined by the synergistic partition 105, 106 as shown in Figure 1.
- Some of these nodes can be leaves, representing a single gene. If they are not leaves, then they represent a subset of genes, which has its own synergistic partition, defined and evaluated as above, with respect to the phenotype. This methodology can be repeated for all gene subsets, until the full tree is formed 107 as outlined in Figure 1.
- G ⁇ and G 2 can be a subset of genes or a module or sub-module of genes that together have a nonnegative synergy 202 and a third independent gene G 3 with a negative synergy 201 between itself and the subset of Gi and G 2 .
- the synergy between Gi and G 2 is +0.0582 and the synergy between the subset of Gi and G 2 and G 3 is -0.1971.
- the synergy refers to the combined cooperative participation of all n genes. If, for example, the expression of one of these genes is independent of all the other genes including the phenotype, then the synergy of the ft-gene set will be zero, even if the set contains synergistic subsets or sub-modules. Therefore, for a thorough synergistic analysis of a gene set, it can be desirable to also identify the most synergistic subsets of size n — 1 , n - 2, ... , 2, which may not necessarily appear in the tree of synergy.
- the disclosed subject matter identifies a cooperative or synergistic interaction that connects the expression levels in two or more genes with the presence or absence of disease.
- the interaction can be modeled using, for example, using a most parsimonious Boolean function as disclosed in International Application No. PCT/US06/61749.
- n such as 3 or 2
- search of all gene sets become computationally expensive and therefore heuristic methods can be employed to find a collection of gene sets with positive synergy, as shown in Figure 3.
- an initial random gene set is first chosen 302 and the synergy of the gene set is calculated 304. If the synergy is greater than a predetermined threshold 305 the gene set is added to the list of gene sets with synergy.
- the predetermined threshold could be zero, for example.
- the gene set is then modified using a heuristic algorithm such as simulated annealing 303 to identify a new gene set whose synergy is evaluated 303. If the synergy of the new gene set is positive, it is added to the list of existing gene sets 306. This process is repeated until the desired number of gene sets is reached. The gene sets are then selected based on the statistical significance of their synergy values 307.
- a heuristic algorithm such as simulated annealing 303 to identify a new gene set whose synergy is evaluated 303. If the synergy of the new gene set is positive, it is added to the list of existing gene sets 306. This process is repeated until the desired number of gene sets is reached. The gene sets are then selected based on the statistical significance of their synergy values 307.
- the techniques of the disclosed subject matter can be implemented by way of off-the-shelf software such as MATLAB, JAVA, C++, or other software. Machine language or other low level languages can also be utilized. Multiple processors working in parallel can also be utilized.
- a system in accordance with the disclosed subject matter can include a processor or multiple processors 404 and a computer readable medium 401 coupled to the processor or processors 404.
- the computer readable medium includes data such as factors and outcomes 402 and can also include programs for synergy analysis and/or EMBP analysis 403.
- the system leads to the identification of the synergy, synergies or cooperative interaction(s) among multiple interacting factors 405.
- Multiple processors 404 working in parallel can also be utilized.
Landscapes
- Chemical & Material Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Organic Chemistry (AREA)
- Genetics & Genomics (AREA)
- Zoology (AREA)
- Analytical Chemistry (AREA)
- Wood Science & Technology (AREA)
- Engineering & Computer Science (AREA)
- Microbiology (AREA)
- Biochemistry (AREA)
- Biotechnology (AREA)
- Molecular Biology (AREA)
- Biophysics (AREA)
- Physics & Mathematics (AREA)
- Pathology (AREA)
- Immunology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
- Apparatus Associated With Microorganisms And Enzymes (AREA)
- Magnetic Resonance Imaging Apparatus (AREA)
- Measuring Volume Flow (AREA)
- Complex Calculations (AREA)
Abstract
Systems and methods for selecting factors from a data set of measurements are provided. The measurements include values of factors and/or outcomes. Two or more factors that are jointly associated with one or more outcomes from the data set are identified. Each of the two or more factors are analyzed to determine at least one cooperative interaction among the factors with respect to an outcome. The two or more factors can be a module of factors.
Description
COMPUTATIONAL ANALYSIS OF THE SYNERGY AMONG MULTIPLE INTERACTING FACTORS
PATENT APPLICATION SPECIFICATION
CROSS-REFERENCE TO RELATED APPLICATIONS
The present application claims priority to U.S. Provisional Application Serial No. 60/799,271 filed May 10, 2006 and U.S. Provisional Application Serial No. 60/885,349 filed January 17, 2007, the entire contents of which are incorporated by reference herein. BACKGROUND
The disclosed subject matter relates generally to systems and methods for factor selection, including factors useful in gene expression analysis.
The expression levels of thousands of genes, measured simultaneously using DNA microarrays, can provide information useful for medical diagnosis and prognosis. However, gene expression measurements have not provided significant insight into the development of therapeutic approaches. This can be partly attributed to the fact that while traditional gene selection techniques typically produce a "list of genes" that are correlated with disease, they do not reflect interrelationships.
Gene selection techniques based on microarray analysis often involve individual gene ranking depending on a numerical score measuring the correlation of each gene with particular disease types. The expression levels of the highest-ranked genes tend to be either consistently higher in the presence of disease and lower in the absence of disease, or vice versa. Such genes usually have the property that their joint expression levels corresponding to diseased tissues and the joint expression levels corresponding to healthy tissues can be cleanly separated into two distinct clusters. These techniques are therefore convenient for classification purposes between disease and health, or between different disease types. However, they do not identify cooperative relationships or the synergy among multiple interacting genes.
There is therefore a need for the ability to analyze genes in terms of the cooperative, as opposed to independent, nature of their contributions towards a phenotype.
SUMMARY
The disclosed subject matter provides techniques for the analysis of the cooperative interactions or synergy among multiple interacting factors. The factors can be features, elements or outcomes that are cooperatively associated with one or more factors or outcomes by their joint presence or absence.
In some embodiments of the disclosed subject matter, methods for selecting factors from a data set of measurements are provided. The measurements can include values of factors and/or outcomes. Two or more factors that are jointly associated with one or more outcomes from the data set are identified. Each of the two or more factors are analyzed to determine at least one cooperative interaction among the factors with respect to an outcome or factor. The factors can be a module of factors or a sub-module of factors. The interaction can be a structure of interactions.
The factors can include two or more genes. The data set can include gene expression data including expression levels for each of the genes. The outcomes includes presence or absence of a disease, and the genes can be a module of genes, preferably a smallest cooperative module of genes with joint expression levels that can be used for a prediction of the presence of a disease with high accuracy.
In other embodiments of the disclosed subject matter, methods for selecting two or more genes from gene expression data are provided. The gene expression data can include expression levels for each of the two or more genes. The method includes providing gene expression data for two or more genes, discretizing the gene expression data, identifying the two or more genes with a high synergy, and identifying a cooperative interaction that connects the expression levels in the two or more genes with presence or absence of a disease. The gene expression data can be derived from at least one microarray of gene expression data. The genes can be a module of genes. The cooperative interaction can be modeled using a Boolean function, a parsimonious Boolean function, or the most parsimonious Boolean function. The high synergy can be a maximum synergy. In other embodiments of the disclosed subject matter, systems for selecting two or more genes from gene expression data are provided. The gene expression data includes expression levels for each of the two or more genes. The system includes at least one processor, a computer readable medium coupled to the processor including instructions which when executed cause the processor to provide
gene expression data for the genes. The instructions also cause the processor to discretize the gene expression data, choose a single threshold for each of the two or more genes, identify the two or more genes with a high synergy, and identify a cooperative interaction that connects the expression levels in the two or more genes with presence of a disease. The gene expression data can be derived from a microarray of gene expression data. The two or more genes can be a module of genes. The high synergy can be a maximum synergy.
In other embodiments of the disclosed subject matter, systems for selecting factors from a data set of measurements are provided. Each measurement can include values of the factors and outcomes. The system includes at least one processor and a computer readable medium coupled to the at least one processor. The computer readable medium includes instructions which when executed cause the processor to identify two or more factors that are jointly associated with one or more outcomes or factors from the data, and, analyze each of the two or more factors to determine at least one cooperative interaction among the factors with respect to an outcome or factor. The factors can be a module of factors. The module of factors can include at least one sub-module of factors. The cooperative interaction can include a structure of interactions, such as a logic function.
The two or more factors can be two or more genes, the data can include gene expression data including expression levels, and the outcomes can be the presence or absence of a disease. The genes can be a module of genes. The module of genes can include at least one sub-module of genes. The module of genes can be a smallest module of genes with joint expression levels that can be used for a prediction of the presence of disease.
BRIEF DESCRIPTION QF THE DRAWINGS
Fig. 1 is a diagram according to some embodiments of the disclosed subject matter.
Fig. 2 is a diagram according to some embodiments of the disclosed subject matter.
Fig. 3 is a diagram according to some embodiments of the disclosed subject matter.
Fig. 4 is a schematic diagram according to some embodiments of the disclosed subject matter.
DETAILED DESCRIPTION
According to some aspects of the disclosed subject matter, a method for selecting factors from a data set of measurements is provided. The method includes identifying factors cooperatively associated with an outcome from a data set for two or more factors, and analyzing each of the factors or modules of factors to determine cooperative or synergistic interactions among the factors or outcomes with respect to the outcome. The data set can be a set of measurements that includes values of the factors and the outcomes.
One application of the disclosed subject matter is where the factors are genes and the inference of the cooperative relationship or synergy among multiple interacting genes is desired. Although the following will describe disease data it can be more generally applicable to other data sets. Other applicable data sets include other biological data such as how cells are influenced by stimuli jointly, financial data, internet traffic data, scheduling data for industries, marketing data, and manufacturing data, for example. Table 1, below, specifies other data sets relevant to various objectives, including factors and outcomes, to which the disclosed subject matter can also be applied: TABLE 1
Referring to Figure 1 , gene expression data for two or more genes is provided 101 in the form of, for example, a microarray of expression data. The gene expression data is then discretized 102. For example, the data can be binarized into two levels. Other levels of discretization, such as trinarization, can also be used.
Rather than independently binarizing each gene's expression level, which can be more appropriate for an individual gene ranking approach, single thresholds are used for the genes. This approach is consistent with the fact that finding global interrelationships among genes is desirable and that the microarray data have already been normalized across the tissues and genes. Therefore, a choice of high threshold will identify the genes that are "strongly" expressed, while a choice of a low threshold will identify the genes that are expressed even "weakly." Entropy Minimization and Boolean Parsimony ("EMBP") analyses can be performed across several thresholds and to determine the threshold levels that provide optimized performance, as disclosed in International Application No. PCT/US06/61749, the disclosure of which is incorporated herein by reference.
Following binarization, each gene can be assumed to be either expressed or not expressed in a particular tissue. It can also be assumed that there are two types of tissues, either healthy ones or tissues suffering from a particular disease. For example, Table 2, below shows results from hypothetical microarray measurements of three genes, Gj, Gi, and G3 in both the presence and absence of a particular cancer C. N0 and Ni represent the presence and absence of cancer, respectively.
TABLE 2
The latter assumption can also be generalized to include more than two types of tissues, or modified to be used for classification among several types of cancer. According to one aspect of the disclosed subject matter, two or more genes with a high synergy are identified 103, as detailed below.
Cooperative or synergistic interactions between multiple genes can then be determined 104, as shown in Figure 1. Because systems biology is based on a holistic view of biological systems, synergy is important.
A gene set of n genes can have expression levels
-.,Gn and a particular outcome C can be any phenotype, such as the presence of a particular disease or the differentiation of stem cells into a particular cell type when analyzing expression data of human tissues. The synergy Syn(Gi,G2,...,Gn;C) of the gene set with respect to the phenotype C is defined by equation (1):
1(G11G2,...,G111 C) - max Y1I(Sn-Q (l) all partitions ^T^ v J
[S1 } such that '
U 5,={G, ,...,G. } and f] S,=0 The partition of the gene set that is chosen in equation (1) above is the one that maximizes the sum of the amounts of mutual information connecting the subsets of that partition with the phenotype 104 as shown in Figure 1, and it is referred to as the "synergistic partition" of the gene set (G]5G2,...,Gn) with respect to the phenotype C. The definition is consistent with the intuitive concept that synergy is the additional amount of contribution for a particular task provided by an integrated "whole" compared with what can best be achieved, after breaking the whole into "parts," by the sum of the contributions of these parts. The above quantity can be divided by the entropy H(C) from EMBP analysis, in which case the maximum possible thus normalized synergy will be +1. For the special case of n - 2, the synergy is defined as shown in equation (2):
Syn(Gl 5G2;O = I(GUG2;Q - U(GuQ + I(G2,Q]. (2)
Equation (2) is symmetric with respect to the three random variables and equal to the opposite of the mutual information /(Gi;G2;C) common to the three variables Gi,G2,C. Contrary to the mutual information common to two variables, the
mutual information common to three variables is not necessarily a nonnegative quantity. This allows for a strictly positive synergy.
In one example, it can be assumed that each of the genes Gi and G2 is equally (50% of the time) expressed when C = I and C = O. In that case, it would appear that the two genes are uncorrelated with the phenotype C, because
/(Gi ;Q = I((J2,C) = 0, and the genes would not be found high up in any typical "gene ranking" computational method. However, C can to be determined with absolute certainty from the joint state of the two genes, for example when C = 1 if G\ = G2, and C = 0 if Gi ≠ G2, in which case /(Gi,G2;C) = 1 , and the synergy is positive and equal to +1. On the other hand, if Gj = G2 = C then the synergy is negative and equal to -1. More generally, if Gi = G2 = ... = Gn = C (ultimate redundancy) then the multivariate synergy can become even more negative and equal to -in - 1).
Since H(C | GuG2,...,Gn) = H(Q - /(G15G2,...,Gn;C), EMBP analysis naturally tends to find high-synergy results, although not necessarily the most synergistic ones.
In another example, as shown in Table 2, equation (3) can be used for n = 3 and is simplified by omitting the phenotypes.
Syn123 = I123 - max (Ii + 123, 12 + 1]3, 13 + 1]2, Ii + 12 + 13) (3)
The values for mutual information between gene subsets and the phenotype can then be found in this example as follows: Ij23 = 0.7963; I]2 = 0.7152; Ii3 = 0.5094; I23 = 0.5163; Ii = 0.3235; I2 = 0.3335; and I3 = 0.2782.
Positive synergy implies some form of direct or indirect interaction of the genes, as a system. This definition of synergy also allows for insight into the structure of potential pathways by making iterative use of the "synergistic partition," defined earlier, to generate a hierarchical decomposition of the gene set into smaller modules. In particular, consider a rooted and not necessarily binary tree with n leaves, each of which represents one of the genes. Each node of the tree represents a subset or sub-module of genes, which contains the genes represented by the leaves of the clade formed by the node. Therefore the root represents the whole gene set. The synergistic partition of the whole gene set, as defined above, can then be represented by the branching of the root, so that the nodes that are neighboring to the
root represent the gene subsets or sub-modules defined by the synergistic partition 105, 106 as shown in Figure 1. Some of these nodes can be leaves, representing a single gene. If they are not leaves, then they represent a subset of genes, which has its own synergistic partition, defined and evaluated as above, with respect to the phenotype. This methodology can be repeated for all gene subsets, until the full tree is formed 107 as outlined in Figure 1. This is referred to as the tree of synergy of the gene set (Cr15G2,...,Gn] with respect to the phenotype C. Each intermediate node of the tree of synergy identifies a gene subset or sub-module with nonnegative synergy. As shown in Figure 2, G\ and G2 can be a subset of genes or a module or sub-module of genes that together have a nonnegative synergy 202 and a third independent gene G3 with a negative synergy 201 between itself and the subset of Gi and G2. In the example of Table 2, the synergy between Gi and G2 is +0.0582 and the synergy between the subset of Gi and G2 and G3 is -0.1971.
The synergy, as defined above, refers to the combined cooperative participation of all n genes. If, for example, the expression of one of these genes is independent of all the other genes including the phenotype, then the synergy of the ft-gene set will be zero, even if the set contains synergistic subsets or sub-modules. Therefore, for a thorough synergistic analysis of a gene set, it can be desirable to also identify the most synergistic subsets of size n — 1 , n - 2, ... , 2, which may not necessarily appear in the tree of synergy. For /2 = 3, however, it can be proved that the most synergistic subset of size 2, if it has positive synergy, is defined by an existing clade of the full tree of synergy, In this way, the disclosed subject matter identifies a cooperative or synergistic interaction that connects the expression levels in two or more genes with the presence or absence of disease. The interaction can be modeled using, for example, using a most parsimonious Boolean function as disclosed in International Application No. PCT/US06/61749.
For small gene set sizes n, such as 3 or 2, it is possible to identify all the synergistic sets of genes by doing a search of possible gene sets of size n. However, as the size increases, searches of all gene sets become computationally expensive and therefore heuristic methods can be employed to find a collection of gene sets with positive synergy, as shown in Figure 3. Starting with the binarized gene expression data 301 an initial random gene set is first chosen 302 and the synergy of the gene set is calculated 304. If the synergy is greater than a predetermined threshold 305 the gene set is added to the list of gene sets with
synergy. The predetermined threshold could be zero, for example. The gene set is then modified using a heuristic algorithm such as simulated annealing 303 to identify a new gene set whose synergy is evaluated 303. If the synergy of the new gene set is positive, it is added to the list of existing gene sets 306. This process is repeated until the desired number of gene sets is reached. The gene sets are then selected based on the statistical significance of their synergy values 307.
For small sizes of gene sets, synergistic analysis can be done with algorithms that list all the partitions of a particular set of genes. The total number of partitions of a set with n elements is given by the "Bell number." As the values of n increase, however, the increased computational complexity makes the problems more complex and heuristic solutions can be used to do the analysis.
The techniques of the disclosed subject matter can be implemented by way of off-the-shelf software such as MATLAB, JAVA, C++, or other software. Machine language or other low level languages can also be utilized. Multiple processors working in parallel can also be utilized.
As illustrated in the embodiment depicted in Figure 4, a system in accordance with the disclosed subject matter can include a processor or multiple processors 404 and a computer readable medium 401 coupled to the processor or processors 404. The computer readable medium includes data such as factors and outcomes 402 and can also include programs for synergy analysis and/or EMBP analysis 403. The system leads to the identification of the synergy, synergies or cooperative interaction(s) among multiple interacting factors 405. Multiple processors 404 working in parallel can also be utilized.
The foregoing merely illustrates the principles of the disclosed subject matter. Various modifications and alterations to the described embodiments will be apparent to those skilled in the art in view of the teachings herein. It will thus be appreciated that those skilled in the art will be able to devise numerous techniques which, although not explicitly described herein, embody the principles of the disclosed subject matter and are thus within the spirit and scope of the disclosed subject matter.
Claims
1. A method for selecting factors from a data set of measurements, the measurements including values of the factors and outcomes, comprising: identifying two or more factors that are jointly associated with one or more outcomes from the data set; and analyzing each of the two or more factors to determine at least one cooperative interaction among the factors with respect to an outcome.
2. The method of claim 1 , wherein the two or more factors comprise a module of factors.
3. The method of claim 1 , wherein the two or more factors comprise a sub- module of factors.
4. The method of claim 1, wherein the at least one interaction comprises a structure of interactions.
5. The method of claim 1, wherein the two or more factors comprise two or more genes, the data includes gene expression data comprising expression levels for each of the two or more genes, and the one or more outcomes includes presence or absence of a disease.
6. The method of claim 5, wherein the two or more genes comprise a module of genes.
7. The method of claim 6, wherein the module of genes comprise a smallest cooperative module of genes with joint expression levels that can be used for a prediction of the presence of disease.
8. A method for selecting two or more genes from gene expression data, comprising: providing gene expression data for two or more genes, the gene expression data comprising expression levels for each of the two or more genes; discretizing the gene expression data; identifying the two or more genes with a high synergy; and identifying a cooperative interaction that connects the expression levels in the two or more genes with presence or absence of a disease.
9. The method of claim 8 wherein the gene expression data is derived from at least one microarray of gene expression data.
10. The method of claim 8 wherein the two or more genes comprise a module of genes.
11. The method of claim 8, wherein the cooperative interaction is modeled using a most parsimonious Boolean function.
12. The method of claim 8, wherein the high synergy comprises a maximum synergy.
13. A system for selecting two or more genes from gene expression data, comprising: at least one processor, and a computer readable medium, coupled to the at least one processor, having stored thereon instructions which when executed cause the processor to: provide gene expression data for the two or more genes, the gene expression data includes expression levels for each of the two or more genes; discretize the gene expression data; choose a single threshold for each of the two or more genes; identify the two or more genes with a high synergy; and identify a cooperative interaction that connects the expression levels in the two or more genes with presence of a disease.
14. The system of claim 13 wherein the gene expression data is derived from a microarray of gene expression data.
15. The method of claim 13 wherein the two or more genes comprise a module of genes.
16. The method of claim 13, wherein the high synergy comprises a maximum synergy.
17. A system for selecting factors from a data set of measurements, each measurement comprising values of the factors and outcomes comprising: at least one processor, and a computer readable medium coupled to the at least one processor, having stored thereon instructions which when executed cause the at least one processor to: identify two or more factors that are jointly associated with one or more outcomes or factors from the data; and analyze each of the two or more factors to determine at least one cooperative interaction among the factors with respect to an outcome or factor.
18. The system of claim 17, wherein the two or more factors comprise a module of factors.
19. The system of claim 18, wherein the module of factors comprises at least one sub-module of factors.
20. The system of claim 18, wherein the at least one cooperative interaction comprises a structure of interactions.
21. The system of claim 20, wherein the at least one cooperative interaction comprises a logic function.
22. The system of claim 21, wherein the two or more factors comprise two or more genes, the data comprises gene expression data comprising expression levels for each of the two or more, and the one or more outcomes comprise presence or absence of a disease.
23. The system of claim 22, wherein the two or more genes comprise a module of genes.
24. The system of claim 23, wherein the module of genes comprises at least one sub-module of genes.
25. The system of claim 23, wherein the module of genes comprises a smallest module of genes with joint expression levels that can be used for a prediction of the presence of disease with high accuracy.
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US12/307,694 US8234077B2 (en) | 2006-05-10 | 2007-05-10 | Method of selecting genes from gene expression data based on synergistic interactions among the genes |
| US13/534,578 US20120290218A1 (en) | 2006-05-10 | 2012-06-27 | Computational Analysis of the Synergy Among Multiple Interacting Factors |
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US79927106P | 2006-05-10 | 2006-05-10 | |
| US60/799,271 | 2006-05-10 | ||
| US88534907P | 2007-01-17 | 2007-01-17 | |
| US60/885,349 | 2007-01-17 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US13/534,578 Division US20120290218A1 (en) | 2006-05-10 | 2012-06-27 | Computational Analysis of the Synergy Among Multiple Interacting Factors |
Publications (3)
| Publication Number | Publication Date |
|---|---|
| WO2007134167A2 true WO2007134167A2 (en) | 2007-11-22 |
| WO2007134167A3 WO2007134167A3 (en) | 2008-10-16 |
| WO2007134167A9 WO2007134167A9 (en) | 2008-11-27 |
Family
ID=38694700
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2007/068666 Ceased WO2007134167A2 (en) | 2006-05-10 | 2007-05-10 | Computational analysis of the synergy among multiple interacting factors |
Country Status (2)
| Country | Link |
|---|---|
| US (2) | US8234077B2 (en) |
| WO (1) | WO2007134167A2 (en) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8086409B2 (en) | 2007-01-30 | 2011-12-27 | The Trustees Of Columbia University In The City Of New York | Method of selecting genes from continuous gene expression data based on synergistic interactions among genes |
| US8234077B2 (en) | 2006-05-10 | 2012-07-31 | The Trustees Of Columbia University In The City Of New York | Method of selecting genes from gene expression data based on synergistic interactions among the genes |
| US8290715B2 (en) * | 2005-12-07 | 2012-10-16 | The Trustees Of Columbia University In The City Of New York | System and method for multiple-factor selection |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8809037B2 (en) | 2008-10-24 | 2014-08-19 | Bioprocessh20 Llc | Systems, apparatuses and methods for treating wastewater |
Family Cites Families (18)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| NL277716A (en) | 1962-04-27 | |||
| US3354297A (en) | 1964-02-06 | 1967-11-21 | Ford Motor Co | Apparatus for measuring dynamic characteristics of systems by crosscorrelation |
| US3535084A (en) | 1965-05-29 | 1970-10-20 | Katsuhisa Furuta | Automatic and continuous analysis of multiple components |
| US3449553A (en) | 1965-08-23 | 1969-06-10 | United Geophysical Corp | Computer for determining the correlation function of variable signals |
| US5697369A (en) | 1988-12-22 | 1997-12-16 | Biofield Corp. | Method and apparatus for disease, injury and bodily condition screening or sensing |
| US5597719A (en) | 1994-07-14 | 1997-01-28 | Onyx Pharmaceuticals, Inc. | Interaction of RAF-1 and 14-3-3 proteins |
| US6110109A (en) | 1999-03-26 | 2000-08-29 | Biosignia, Inc. | System and method for predicting disease onset |
| US20040253637A1 (en) | 2001-04-13 | 2004-12-16 | Biosite Incorporated | Markers for differential diagnosis and methods of use thereof |
| DE10159262B4 (en) | 2001-12-03 | 2007-12-13 | Siemens Ag | Identify pharmaceutical targets |
| US20030215866A1 (en) | 2002-05-01 | 2003-11-20 | Liebovitch Larry S. | Models of genetic interactions and methods of use |
| US9898578B2 (en) | 2003-04-04 | 2018-02-20 | Agilent Technologies, Inc. | Visualizing expression data on chromosomal graphic schemes |
| CN1316419C (en) | 2002-08-22 | 2007-05-16 | 新加坡科技研究局 | Make predictions from common likelihoods that form models |
| US6996476B2 (en) * | 2003-11-07 | 2006-02-07 | University Of North Carolina At Charlotte | Methods and systems for gene expression array analysis |
| US20070299645A1 (en) | 2004-04-27 | 2007-12-27 | Yeda Research And Development Co., Ltd. | Autonomous Molecular Computer Diagnoses Molecular Disease Markers and Administers Requisite Drug in Vitro |
| WO2007012052A1 (en) | 2005-07-20 | 2007-01-25 | Medical College Of Georgia Research Institute | Use of protein profiles in disease diagnosis and treatment |
| WO2007067956A2 (en) | 2005-12-07 | 2007-06-14 | The Trustees Of Columbia University In The City Of New York | System and method for multiple-factor selection |
| US8234077B2 (en) | 2006-05-10 | 2012-07-31 | The Trustees Of Columbia University In The City Of New York | Method of selecting genes from gene expression data based on synergistic interactions among the genes |
| US8086409B2 (en) | 2007-01-30 | 2011-12-27 | The Trustees Of Columbia University In The City Of New York | Method of selecting genes from continuous gene expression data based on synergistic interactions among genes |
-
2007
- 2007-05-10 US US12/307,694 patent/US8234077B2/en active Active
- 2007-05-10 WO PCT/US2007/068666 patent/WO2007134167A2/en not_active Ceased
-
2012
- 2012-06-27 US US13/534,578 patent/US20120290218A1/en not_active Abandoned
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8290715B2 (en) * | 2005-12-07 | 2012-10-16 | The Trustees Of Columbia University In The City Of New York | System and method for multiple-factor selection |
| US8234077B2 (en) | 2006-05-10 | 2012-07-31 | The Trustees Of Columbia University In The City Of New York | Method of selecting genes from gene expression data based on synergistic interactions among the genes |
| US8086409B2 (en) | 2007-01-30 | 2011-12-27 | The Trustees Of Columbia University In The City Of New York | Method of selecting genes from continuous gene expression data based on synergistic interactions among genes |
Also Published As
| Publication number | Publication date |
|---|---|
| US20120290218A1 (en) | 2012-11-15 |
| WO2007134167A3 (en) | 2008-10-16 |
| WO2007134167A9 (en) | 2008-11-27 |
| US8234077B2 (en) | 2012-07-31 |
| US20090299643A1 (en) | 2009-12-03 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7684287B2 (en) | Single-cell RNA-SEQ data processing | |
| Ansari et al. | Performance evaluation of machine learning techniques (MLT) for heart disease prediction | |
| US20170193157A1 (en) | Testing of Medicinal Drugs and Drug Combinations | |
| Liao et al. | Appropriate medical data categorization for data mining classification techniques | |
| Shandilya et al. | Survey on recent cancer classification systems for cancer diagnosis | |
| CA3092303A1 (en) | A computer-implemented method of analysing genetic data about an organism | |
| US20120253686A1 (en) | System and method for multiple-factor selection | |
| WO2007134167A2 (en) | Computational analysis of the synergy among multiple interacting factors | |
| Zawar et al. | Variants of uncertain significance: At the crux of diagnostic odyssey | |
| US20240273359A1 (en) | Apparatus and method for discovering biomarkers of health outcomes using machine learning | |
| Stephen et al. | Feature selection/dimensionality reduction | |
| Hermans et al. | Elongate dendritic phytoliths as indicators for cereal identification and domestication: exploring a 3D morphometric approach | |
| Bugrim | Identification of disease mechanisms and novel disease genes using clinical concept embeddings learned from massive amounts of biomedical data | |
| US11348662B2 (en) | Biomarkers based on sets of molecular signatures | |
| Shaikh et al. | Quasi Opposition-based Learning in Beluga Whale Optimization Feature Selection Approach for Diabetes Prediction in IoT System. | |
| Qureshi et al. | Bioinformatics and Genomics in Rare Disease Diagnosis: Leveraging AI and Next-Generation Sequencing for Identifying Novel Genetic Disorders | |
| Pandey et al. | Machine Learning and Statistic-Based | |
| Mandal et al. | Identification of genetic pathway for cervical cancer development using rough and Bayesian theory | |
| La Cava et al. | Application of concise machine learning to construct accurate and interpretable EHR computable phenotypes | |
| Aziz et al. | SMOTE-ENN-LR: LEVERAGING MACHINE LEARNING FOR BREAST CANCER CLASSIFICATION IN MICROARRAY GENE EXPRESSION WITH EXPLAINABLE AI | |
| Choobdar et al. | Discovering weighted motifs in gene co-expression networks | |
| Jeon et al. | Denoiseit: denoising gene expression data using rank based isolation trees | |
| Balogh et al. | Contrastive learning of adverse events to provide effective and interpretable vector representations for machine-assisted pharmacovigilance | |
| Duvall et al. | Analyzing Co-expression Networks with Network Skeleton Extraction | |
| Sha | Genetic Data Analysis and Interpretation Via Feature Selection and Network Science |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 07762093 Country of ref document: EP Kind code of ref document: A2 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 12307694 Country of ref document: US |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 07762093 Country of ref document: EP Kind code of ref document: A2 |

