WO2006059770A1 - 確率算出方法 - Google Patents
確率算出方法 Download PDFInfo
- Publication number
- WO2006059770A1 WO2006059770A1 PCT/JP2005/022317 JP2005022317W WO2006059770A1 WO 2006059770 A1 WO2006059770 A1 WO 2006059770A1 JP 2005022317 W JP2005022317 W JP 2005022317W WO 2006059770 A1 WO2006059770 A1 WO 2006059770A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sample
- haplotype
- frequency
- diplotype
- locus
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/40—Population genetics; Linkage disequilibrium
Definitions
- the present invention relates to a probability calculation method for calculating the probability (significance level) of the first type of error in an independence test using case-contro 1 correlation analysis using SNP (Single Nucleotide Polymorphism). .
- case-control correlation analysis using SNP as risk factors, (1) allele frequency model (AZa), (2) dominant genetic model (AA + Aa / aa), (3) recessive genetic model (AA / Aa + aa), (4) Genotype model (AAZAaZaa) is assumed.
- A represents a major allele
- a represents a minor allele.
- a 2 X 2 or 2 X 3 contingency table is created and an independence test is performed.
- Bonferroni's correction method the value obtained by multiplying the p-value obtained by the independence test by the number of independence tests (4 times in the case of the four models described above) is regarded as the multiple test. This excludes the effects of multiple testing. This method interprets the test results in the most conservative (strict) way.
- the method described in Non-Patent Document 1 is known.
- the linkage disequilibrium index value ( ⁇ ) between each two SNPs is calculated for a small number of SNPs, and approximation is performed by performing eigenvalue decomposition of the matrix containing the index values as components. Corrections are made.
- the method described in Non-Patent Document 1 is a method for correcting the p-value in a multiple test focusing on linkage disequilibrium.
- Patent Document 1 Dale R. Nyholt, A simple correction for multiple test ing for single- nucleotide polymorphisms in linkage disequilibrium with each other, Am. J. Hum. Genet., Vol. 74, pp. 765- 769 Disclosure of the Invention
- the Bonferroni correction method has a problem that there is a risk of overlooking a disease-related SNP because the corrected p-value is likely to be larger than the true p-value.
- the p-value of the entire multiple test is set to a commonly used value such as 0.01 or 0.05, in order to allow the set p-value. Since the p-value of each SNP is very small, that is, it is a very severe condition, there is also a problem that there is a risk of overlooking the disease-related SNP as well. For example, if 50000 SNP data are targeted and the p-value of the entire multiple test is set to 0.01, the p-value of each SNP is always 0.0000002 (0.0.1 + 50000) / J, Value.
- Non-Patent Document 1 has a problem that haplotype frequency is not taken into account because it focuses on pairwise linkage disequilibrium. Furthermore, there was a problem that the direct p-value was not calculated.
- the present invention has been made in view of the above problems.
- the P value in the independence test performed by case-control correlation analysis using SNP that is, the probability of the first type of error is considered, and the haplotype frequency is considered. It is an object of the present invention to provide a probability calculation method that can be directly calculated and, as a result, can prevent oversight of disease-related SNPs.
- the probability calculation method calculates the probability of the first type error in the independence test performed by case-control correlation analysis using SNP.
- a sample acquisition step for acquiring a first sample and a second sample including a plurality of haplotype information related to haplotypes composed of a plurality of SNPs, and a first sample acquired in the sample acquisition step.
- a haplotype frequency calculation step for calculating the haplotype frequency for each haplotype that can exist, and each locus of each haplotype is a minor allele.
- a matrix definition step for defining a matrix having the predetermined value as a component, a haplotype frequency of each haplotype in the first sample, and a haplotype frequency of each haplotype in the second sample are set.
- a haplotype frequency setting step that gives a combination of each haplotype frequency of the first sample and each haplotype frequency of the second sample, a plurality of haplotype frequencies corresponding to the first sample set in the haplotype frequency setting step, and the matrix definition step
- the allele frequency at each locus in the first sample is calculated based on the matrix defined in, and the multiple haplotype frequencies corresponding to the set second sample and the defined matrix at each locus in the second sample are calculated.
- a contingency table creation step for creating a contingency table corresponding to each locus based on the allele frequency at each locus in the first sample and the allele frequency at each locus in the second sample; Either the independence test execution step for executing the independence test for each locus or the test result of the independence test executed in the independence test execution step based on the contingency table created in the generation step is If significant, based on the number of samples in the first sample, the number of samples in the second sample, the total number of haplotypes that may be present, the haplotype frequency in the first sample, the haplotype frequency in the second sample, and the no-protype frequency The probability corresponding to the combination of each haplotype frequency of the first sample and each haplotype frequency of the second sample set in the haplotype frequency setting step is a haplo of the first sample and the second sample set in advance.
- the probability calculation method according to claim 2 is a probability calculation method for calculating the probability of the first type of error in the independence test performed in case-contro relation analysis using SNP.
- the contingency table creation step for creating the contingency table corresponding to each locus and the contingency table created in the contingency table creation step for each locus
- one of the independence test execution step that executes the independence test and the test result of the independence test executed in the independence test execution step is significant
- the probability corresponding to the combination of each diplotype form frequency of the first sample set in the diplotype form frequency setting step and each diplotype form frequency of the second sample is set between the first sample and the second sample set in advance.
- the probability calculation method according to claim 3, which is useful for the present invention, is a probability calculation method for calculating the probability of the first type of error in the independence test performed in case-contro relation analysis using SNP.
- the haplotype frequency calculation step for calculating the haplotype frequency for each haplotype that can exist, and whether each locus of each haplotype has a minor allele or major allele is predetermined.
- a matrix definition step for defining a matrix having the predetermined value as a component and the first sample
- the diplotype form frequency setting step that gives the combination, the plurality of diplotype form frequencies corresponding to the first sample set in the diplotype form frequency setting step, and the matrix defined in the matrix definition step
- Genotypes that calculate genotype frequencies at each locus and calculate genotype frequencies at each locus in the second sample based on multiple diplotype frequencies corresponding to the set second sample and a defined matrix The genotype frequency at each locus in the first sample calculated in the frequency calculation step and the genotype frequency calculation step.
- the contingency table creation step for creating the contingency table corresponding to each locus, and the contingency table created in the contingency table creation step ,
- the independence test execution step that executes the independence test for each locus and the test result of the independence test executed in the independence test execution step!
- the first sample Based on the number of samples, the number of samples in the second sample, the total number of possible haplotypes, the number of diplotypes in the first sample, the number of diplotypes in the second sample, and the frequency of the diplotype, Probability corresponding to the combination of each diplotype power of the first sample and each diplotype power of the second sample set in the setting step is related to the diplotype power of the first and second samples set in advance.
- It includes a probability calculation step of calculating using the interleaf distribution, and the diplotype power setting step, the genotype of The number calculation step, the contingency table creation step, the independence test execution step, and the probability calculation step are performed for all combinations of each diplotype form frequency of the first sample and each of the divotype form frequencies of the second sample.
- the probability of the first type error is finally calculated by adding the probabilities calculated each time repeatedly.
- the probability calculation method according to claim 4 which is useful for the present invention, approximates the probability of the first type of error in the independence test performed by case-control correlation analysis using SNP using the MCMC method.
- the first haplotype selection step for selecting the first haplotype and the haplotype power of the first haplotype in the sample selected in the sample selection step
- the second haplotype selection step for selecting a second haplotype different from the first haplotype, and the haplotype of the first haplotype based on the preset haplotype frequency of the first haplotype and the haplotype frequency of the second haplotype.
- a haplotype frequency update step for updating the haplotype frequency and the haplotype frequency of the second haplotype, a plurality of haplotype frequencies corresponding to the first sample updated in the haplotype frequency update step, and a matrix defined in the matrix definition step.
- An allele frequency calculating step to calculate, and the allele Create a contingency table for each locus based on the allele frequency at each locus in the first sample and the allele frequency at each locus in the second sample calculated in the number calculation step For the step and the contingency table created in the contingency table creation step, the independence test execution step for performing the independence test for each locus and the independence test test result performed in the independence test execution step.
- a cumulative number updating step of updating the cumulative number of times in the case where any one of the two is significant the sample selection step, the first haptic A mouth type selection step, the second haplotype selection step, the haplotype frequency update step, the allele frequency calculation step, the contingency table creation step, the independence detection execution step, and the cumulative number update step are repeated a predetermined number of times, The ratio of the final cumulative number of times to the predetermined number of times is calculated as the first type error probability.
- the probability calculation method includes: (1) acquiring a first sample and a second sample including a plurality of haplotype information related to a haplotype composed of a plurality of SNPs; and (2) acquiring the acquired first Based on the haplotype information contained in the sample and the haplotype information contained in the acquired second sample, the haplotype frequency is calculated for each haplotype that can exist. (3) Each locus of each haplotype is a minor allele (with the same locus). By setting a predetermined value according to whether it has a less frequent allele or a major allele (the more frequent allele at the same sitting position), a matrix having the predetermined value as a component is defined.
- each haplotype of the first sample is set.
- the haplotype frequencies of the second sample are combined, and (5) the haplotype frequencies corresponding to the set first sample and the matrix frequencies defined in the first sample based on the defined matrix and the defined matrix.
- the contingency table corresponding to each locus (note that the contingency table is contingency t able, translated into contingency table in Japanese) (7) Independence test is performed for each locus on the generated contingency table.
- Sample number of the first sample, second sample The number of haplotypes in the first sample and the number of haplotypes in the second sample, based on the number of haplotypes in the first sample, the haplotype frequency in the first sample, the haplotype frequency in the second sample, and the frequency of the haplotypes The probability corresponding to the combination of and is calculated using a preset distribution of the haplotype frequencies of the first and second samples. And (4), ( 5), (6), (7), and (8) are repeated for all combinations of each haplotype frequency of the first sample and each haplotype frequency of the second sample. By adding the probabilities, the probability of type 1 error is finally calculated. As a result, it is possible to directly calculate the probability of the first type error in consideration of the haplotype frequency, and as a result, it is possible to prevent the oversight of the disease-related SNP.
- the probability calculation method includes: (1) acquiring a first sample and a second sample including a plurality of haplotype information related to a haplotype composed of a plurality of SNPs, and (2) acquiring the acquired first Based on the haplotype information contained in the specimen and the haplotype information contained in the acquired second specimen, the haplotype frequency is calculated for each haplotype that can exist, and (3) each locus of each haplotype is identified as a minor or major allele.
- a matrix having the predetermined value as a component is defined.
- the diplotype frequency of each diplotype in the first sample and each matrix in the second sample By setting the diplotype power of the diplotype, the set of each diplotype power of the first sample and each diplotype power of the second sample (5) Calculate the allele frequency at each locus in the first sample based on the defined diplotype frequencies corresponding to the set first sample and the defined matrix, and correspond to the set second sample Calculate the allele frequency at each locus in the second sample based on the multiple diplotype frequencies and the defined matrix, and (6) the calculated allele frequency at each locus in the first sample and the second sample Based on the allele frequency at each sitting position, create an even table corresponding to each sitting position. (7) Perform an independence test for each sitting position.
- the number of samples in the first sample, the number of samples in the second sample, the total number of haplotypes that can exist, the diplotype form frequency in the first sample, the number of samples in the second sample Diplotype power and Noplota Based on the frequency, the probability corresponding to the combination of each diplotype form frequency of the first sample and each diplotype form frequency of the second sample is determined between the preset first sample and second sample. Calculated using the bond distribution for the diplotype form frequency.
- the probability calculation method includes (1) acquiring a first sample and a second sample including a plurality of haplotype information related to haplotypes configured by a plurality of SNPs, and (2) acquiring Based on the haplotype information included in the first sample and the haplotype information included in the acquired second sample, the haplotype frequency is calculated for each haplotype that can exist.
- Each locus of each haplotype is a minor allele or major By setting a predetermined value according to whether alleles are present, a matrix with the predetermined value as a component is defined, and (4) the diplotype frequency of each diplotype in the first sample and the second sample By setting the diplotype power of each diplotype in, each diplotype power of the first sample and each diplotype power of the second sample (5) Based on the multiple diplotype frequencies corresponding to the set first sample and the defined matrix, genotype frequencies at each locus in the first sample are calculated, and the set second sample is calculated. The genotype frequency at each locus in the second sample is calculated based on the corresponding multiple diplotype frequencies and the defined matrix.
- the genotype frequency and the Based on the genotype frequency at each locus in 2 samples create an even table corresponding to each locus, (7) Run an independent test for each locus on the created even table, (8) If one of the results of the independence test performed is significant, the number of samples in the first sample, the number of samples in the second sample, the total number of haplotypes that can exist, the diplotype form factor of the first sample, The diplotype form factor of the second sample.
- the probability corresponding to the combination of each diplotype form frequency of the set first sample and each diplotype form frequency of the second sample is determined based on the frequency of the first and second samples. Calculated using the joint distribution for the diplotype frequency with the second sample.
- the probability calculation method includes (1) setting a predetermined value that is predetermined according to whether each locus of each haplotype has a minor allele or a major allele. Define a matrix whose components are predetermined values, (2) generate a combination of haplotype frequencies for the first and second samples set in advance, and (3) either the first sample or the second sample (4) Select the first haplotype, (5) If the haplotype frequency of the first haplotype in the selected sample is not 0, select a second haplotype that is different from the first haplotype, (6 ) Based on the haplotype frequency of the first haplotype and the haplotype frequency of the second haplotype set in advance, the haplotype frequency of the first haplotype and the haplotype frequency of the second haplotype (7) Calculate the allele frequencies at each locus in the first sample based on multiple haplotype frequencies corresponding to the updated first sample and the defined matrix, and update the second sample The allele frequencies at each locus in the second sample are
- FIG. 1 is a flowchart showing an example of a probability calculation method for accurately calculating the probability of the first type of error using an allele frequency model.
- Figure 2 shows the probability calculation that accurately calculates the probability of type 1 error in the dominant / recessive genetic model. It is a flowchart which shows an example of the exit method.
- FIG. 3 is a flowchart showing an example of a probability calculation method for accurately calculating the probability of the first type error in the genotype model.
- FIG. 4 is a flowchart showing an example of a probability calculation method for accurately calculating the probability of the first type of error when an accurate haplotype frequency is not known.
- FIG. 5 is a flowchart showing an example of a probability calculation method for approximately calculating the probability of the first type of error using the MCMC method in the allele frequency model.
- FIG. 6 is a flowchart showing an example of a probability calculation method for calculating the probability of the first type of error using the MCMC method in a dominant / recessive genetic model.
- FIG. 7 is a flowchart showing an example of a probability calculation method for approximately calculating the probability of the first type of error using the MCMC method in the genotype model.
- FIG. 8 is a diagram illustrating an example of information included in the first sample and the second sample.
- FIG. 9 is a diagram showing an example of information on all haplotypes that can exist.
- FIG. 10 is a diagram showing an example of haplotype frequencies for all haplotypes that can exist.
- FIG. 11 is a diagram illustrating an example of a defined matrix.
- FIG. 12 is a diagram showing an example of information related to haplotype frequency.
- FIG. 13 is a diagram showing an example of an even table created in the case of an allele frequency model.
- FIG. 14 is a diagram showing the relationship between diplotype model numbers and haplotype numbers.
- FIG. 15 is a diagram showing an example of information related to diplotype form frequencies.
- FIG. 16 is a diagram showing an example of a contingency table created in the case of a recessive genetic model.
- FIG. 17 is a diagram showing an example of a contingency table created in the case of a dominant genetic model.
- FIG. 18 is a diagram showing an example of a contingency table created in the case of a genotype model.
- FIG. 19 is a diagram showing an example of an even table created in the case of an allele frequency model.
- FIG. 20 is a diagram showing an example of an even table created in the case of the allele frequency model.
- FIG. 21 is a diagram showing an example of an even table created in the case of an allele frequency model.
- FIG. 22 is a diagram showing an example of an even table created in the case of an allele frequency model.
- FIG. 23 is a diagram showing an example of an even table created in the case of an allele frequency model.
- FIG. 24 is a diagram showing an example of an even table created in the case of an allele frequency model.
- FIG. 25 is a diagram showing an example of an even table created in the case of an allele frequency model.
- the probability calculation method for the present invention which accurately calculates the probability of the first type of error, is as follows: (1 1) allele frequency model, (1 2) dominant. Recessive genetic model, (1 3) genotype The model, (1-4) If the exact haplotype frequency is weak, explain in detail with reference to the figure.
- FIG. 1 is a flowchart showing an example of a probability calculation method for accurately calculating the probability of the first type of error using the allele frequency model.
- a first sample and a second sample containing a plurality of haplotype information related to a haplotype composed of a plurality of SNPs are acquired (step SA-1). Specifically, as shown in Fig. 8, N pieces of haplotype information consisting of L (L is a positive integer) chained SNPs are included.
- the haplotype frequency is calculated for each haplotype that may exist (Ste SA—2).
- the total number of haplotypes that can exist is specifically 2 or less and M may be described as shown in FIG. is there.
- the haplotype frequency h (where j is a number that identifies the haplotype and satisfies l ⁇ j ⁇ M)
- the haplotype frequency h satisfies the following formula 1.
- a predetermined value is set according to the force at which each locus of each haplotype has a minor allele or major allele, thereby defining a matrix having the predetermined value as a component (step SA-3). Specifically, if the i-th position in the haplotype of the haplotype number j (i is an integer that identifies the locus and satisfies l ⁇ i ⁇ L) is the minor allele (a), 1 is set as the major allele. In the case of (A), 0 is set as a predetermined value.
- the matrix to be defined can be expressed as shown in Fig. 11, and a value of 0 or 1 for each combination of the locus number i (i is an integer satisfying l ⁇ i ⁇ L) and the haplotype number j. Is stored.
- the haplotype frequency X of haplotype number j is a random variable.
- the actual value X of the haplotype frequency X of the haplotype number j in the first sample is set for all haplotype numbers j, and the second
- Haplotype frequency of sample with haplotype number j Realized value of X
- step SA-4 based on the multiple haplotype frequencies corresponding to the first sample set in step SA-4 and the matrix defined in step SA-3, the allele frequency at each locus in the first sample is calculated.
- the key at each locus in the second sample is Calculate the real frequency (step SA—5).
- the minor allele at the locus number i in the first sample and the minor allele at the locus number i in the second sample is calculated.
- the allele frequency Y is a random variable, and is expressed by the following Equation 3. And below
- ⁇ is the matrix defined in step SA-3.
- Step SA-5 based on the allele frequency at each locus in the first sample and the allele frequency at each locus in the second sample calculated in step SA-5, the occurrence corresponding to each locus.
- Step SA—6 Specifically, 2 x 2 contingency tables shown in Fig. 13 are created for the number of sitting positions (L).
- step SA-7 an independence test is performed for each locus on the contingency table created in step SA-6 (step SA-7).
- the independence test will be briefly described.
- X defined by Equation 4 below is calculated for an even table, and the calculated X value has a degree of freedom (1, 1) and a significance level (for example, 0.01 or 0.05).
- Equation 4 i is a number that identifies the specimen.
- j is a number for identifying the allele.
- rij is the allele frequency.
- n is the minor allele frequency (n) and measure for sample number i
- n is the allele frequency of the case group corresponding to allele number j (
- n is the number of cases and lj 2j
- n ⁇ 2 ⁇ 2 n .
- step SA—8 Yes
- the number of samples in the first sample, the second sample Sample haplotypes, the total number of haplotypes that can be present, the haplotype frequency of the first sample, the haplotype frequency of the second sample, and the frequency of the negative protypes, and then each haplo of the first sample set in step SA-4.
- the probability corresponding to the combination of the type frequency and each haplotype frequency of the second sample is calculated using the preset distribution of the haplotype frequencies of the first and second samples (step SA-9). Specifically, the number of samples of the first sample, N, the number of samples of the second sample, N,
- Equation 5 k is a number that identifies the sample.
- step SA—10 Yes
- step SA—4 to Step SA—9
- All the probabilities that have been issued are added, and the added probability is finally set as the first type error probability (step SA-11).
- step SA-10 No
- step SA-4 executes step SA-4 to step SA-9 and repeat until all combinations are completed.
- the probability calculation method (1) obtains a first sample and a second sample including a plurality of haplotype information related to a haplotype composed of a plurality of SNPs, and (2) obtains Based on the haplotype information contained in the first sample and the haplotype information contained in the acquired second sample, calculate the haplotype frequency for each haplotype that may exist. (3) Each locus of each haplotype is minor By setting a predetermined value according to whether the allele or major allele is present, a matrix with the predetermined value as a component is defined.
- the haplotype frequency of each haplotype in the first sample and the second sample By setting the haplotype frequency of each haplotype, a combination of each haplotype frequency of the first sample and each haplotype frequency of the second sample is given, (5) The allele frequencies at each locus in the first sample are calculated based on the defined haplotype frequencies corresponding to the first sample and the defined matrix, and the haplotype frequencies and the corresponding haplotype frequencies corresponding to the set second sample are calculated. Based on the defined matrix, the allele frequency at each locus in the second sample is calculated.
- step SA-5 the allele frequency Y of the minor allele at the locus number i in the first sample shown in Equation 3 and the minor allele at the locus number i in the second sample are shown.
- the allele frequency Y of haplota is related to the allele frequency p of the minor allele at locus number i.
- step SA-9 the combined distribution related to the haplotype frequency between the first sample and the second sample shown in Equation 5 is the combined distribution related to the haplotype frequency of the first sample shown in Equation 7 below and Equation 8 below. It is given by the product of the second sample shown and the joint distribution of haplotype frequencies.
- two samples are obtained from the same population as the null hypothesis (H).
- haplotype frequencies of haplotype number j in these two samples are random variables X and X, respectively, the two samplings are
- the 2j probabilities are expressed by the following formula 7 and the following formula 8, respectively. Note that both samples are assumed to be from a population with the same haplotype frequency (s) (null hypothesis).
- X and X lj 2j are real values of the random variables X and X, respectively.
- the difference between the concept of haplotype copy and haplotype in the experiment of extracting haplotypes is as follows. For example, if one individual has two identical wild-types (homozygotes), this individual is considered to possess one haplotype and two haplotype copies. Also, the first species calculated in step SA-11 The probability of error can be expressed by Equation 9 below. The probability calculated by Equation 9 below is the probability under the null hypothesis, and means the probability that any one of the test results of the independence test performed in Step SA-7 will be significant.
- Equation 9 Z is 1 if one of the test results of the independence test performed in Step SA-7 is significant, and 0 if it is not. It is a random variable defined to take The variable Z is a function of the random variable Y described above, and
- the random variable Y is a function of the random variable X described above
- the variable Z is a function f (x, ⁇ ⁇ ki kj 11
- ⁇ a ⁇ is an s XL matrix.
- the function Is 21 22) is defined as follows. ⁇ dl,
- b is the proportion of js + l of s-Lth haplotype haplotype copies of the jth sample that is a minor allele at the kth position. This is a value that varies depending on the composition of the X actual haplotypes of the s + 1st haplotype of the result, and the run js + l
- the independence test is performed for each sitting position on the created contingency table. Then define a function f that is 1 if any of the results of the independence test is significant, and 0 otherwise.
- a and b are non-negative integers, and m ⁇ a and n ⁇ b.
- Equation 12 is negative when a ⁇ bmZn and positive when a> bmZn. It should be noted that m ⁇ a and n ⁇ b in Equation 12 and Equation 13 are not 0 at the same time. Therefore, T is monotonically decreasing at a bm / n and monotonically increasing at a ⁇ bmZn. When b is fixed, in the range of a ⁇ a ⁇ a
- Equation 13 is negative when b ⁇ anZm and positive when b> anZm.
- T is monotonically decreasing when b ⁇ anZm, and monotonically increasing when b> anZm. Therefore, T is b ⁇ b ⁇ b when a is fixed
- Equation 11 is expressed as (a 2 a, b 2 b), (a
- the conservative probability calculation method for the major haplotypes is described in detail again.
- the allele at the k-th position of the s + 1st haplotype is all minor alleles or all in the first and second samples.
- the test is performed as a major allele, and if any test is significant, the k-position is determined to be significant.
- s (the cumulative frequency of major haplotypes) should be as large as possible. Then the confidence in the estimated frequency of major haplotypes will be less reliable.
- the reliability of the haplotype frequency at the s + l, ..., L locus may be low.
- the minor allele frequency at the kth locus in a population is often known. If the minor allele frequency at the k locus in this population is q, then k
- the minor allele frequency of the kth locus in other haplotypes can be considered as q— ⁇ s «h.
- the number of copies of other haplotypes in the jth sample is X
- the calculation method may be a method such as linear regression or curve regression.
- FIG. 2 is a flowchart showing an example of a probability calculation method for accurately calculating the probability of the first type of error in the dominant / recessive genetic model.
- a first sample and a second sample containing a plurality of haplotype information related to a haplotype composed of a plurality of SNPs are acquired (step SB-1). Note that the explanation regarding the acquisition of the first specimen and the second specimen is the same as the above-mentioned (11) allele frequency model.
- the haplotype frequency is calculated for each haplotype that may exist. (Step SB—2).
- the haplotype frequency is the same as the (1-1) allele frequency model described above.
- a predetermined value is set according to the force at which each locus of each haplotype has a minor allele or major allele, thereby defining a matrix having the predetermined value as a component (step SB-3).
- the explanation about the definition of the matrix is the same as the above (11) array frequency model.
- each diplotype power and the first diplotype power in the second sample are set.
- each diplotype form factor of the two samples step SB—4.
- Mouth type form frequency X is a random variable, and satisfies the following formula 14 respectively.
- the actual value X of the diplopivo frequency of diplotype model number jj 'in the first sample is set for all diplotype model numbers jj', and the second sample
- the allele frequencies at each locus in the first sample And calculate the allele frequencies at each position in the second sample based on the multiple probabilistic frequencies corresponding to the second sample set in step SB-4 and the matrix defined in step SB-3. (Step SB—5).
- the allele frequency Y of the minor allele at sitting position i in the first sample and the minor allele at sitting position i in the second sample is calculated.
- the allele frequency Y is a random variable, and is represented by the following formula 15. And below Based on Formula 15, the actual value y of the minor allele allele frequency at locus i in the first sample and the minor allele frequency y at locus i in the second sample y li
- the allele frequency of the minor allele (aa + Aa) is calculated using Formula 15, but in the case of the dominant genetic model, the minor allele (aa) is calculated using Formula 16 below. Calculate the allele frequency.
- the contingency table corresponding to each locus is calculated. Create (Step SB—6). Specifically, in the case of the recessive genetic model, the 2 x 2 contingency table shown in Figure 16 for the allele frequency calculated using Equation 15 is created for the number of loci (L), and the dominant genetic model Then, the 2 x 2 contingency table for the allele frequency calculated using Equation 16 is created for the number of sitting positions (S) shown in Fig. 17.
- step SB-7 an independence test is performed for each sitting position on the contingency table created in step SB-6 (step SB-7).
- step SB—8 Yes
- the number of samples in the first sample, the second sample Each diplo of the first sample set in step SB-4 based on the number of haplotypes, the total number of haplotypes that can be present, the diplotype form frequency of the first sample, the diploma form frequency of the second sample, and the haplotype frequency.
- the probability is calculated using the preset distribution of the diplotype frequency of the first and second samples (step SB-9). Specifically, the number of samples of the first sample N, the second
- each diplotype power of the first sample and each diplotype power of the second sample (X, ⁇ ⁇ , X, X, ⁇ ⁇ ⁇ ⁇ , X)
- the corresponding probability P is calculated using the joint distribution of the diplotype frequency between the first and second samples shown in Equation 17 below.
- Equation 17 k is a number that identifies the sample.
- step SB-10 it is determined whether or not all combinations of the diplotype powers of the first sample and the diplotype powers of the second sample have been completed, and if all the combinations have been completed (step SB-10: Yes), each time step SB-4 to step SB-9 are repeated, the probability calculated in step SB-9 is added, and the added probability is finally set as the probability of type 1 error ( Step SB—11). On the other hand, if all combinations have not been completed (step SB-10: No), execute steps SB-4 to SB-9 and repeat until all combinations are completed.
- the probability calculation method (1) obtains a first sample and a second sample including a plurality of haplotype information related to haplotypes configured by a plurality of SNPs, and (2 ) Calculate the haplotype frequency for each haplotype that can exist based on the haplotype information contained in the acquired first sample and the haplotype information contained in the acquired second sample, and (3) each haplotype Sitting position minor allele or major allele By defining a predetermined value according to whether or not it has, a matrix with the predetermined value as a component is defined.
- the diplotype form frequency of each diplotype in the first sample and each diplotype in the second sample By setting the type diplotype power, the combination of each diplotype power of the first sample and each diplotype power of the second sample is given, and (5) corresponds to the set first sample.
- the allele frequency at each locus in the first sample is calculated based on the multiple diplotype powers and the defined matrix, and the second diplotype powers corresponding to the set second sample and the defined matrix are calculated. Calculate the allele frequencies at each locus in the two samples, and (6) calculate the allele frequencies at each locus in the first sample and the allele frequencies at each locus in the second sample.
- the predetermined diplotype of the first sample and the second sample of the probability corresponding to the combination of each diplotype form frequency of the first sample and each diplotype form factor of the second sample Calculation is performed using a joint distribution relating to the form frequency.
- step SB-9 the joint distribution of the first and second samples in terms of the diplotype frequency shown in Equation 17 is the joint distribution of X for all j and j ', that is,
- the dominant genetic model test is a major allele.
- the test method that compares the ratio of individuals with those that do not, and the ratio of individuals in two groups.
- the test with the recessive genetic model is a test that compares the ratio of individuals with minor alleles and those that do not in two groups. Let's go.
- j j 'may be sufficient. If the total number of haplotypes is M, the total number of permutations diplotype is M 2. If Hardy-Weinberg equilibrium is established for haplotypes, the frequency of D in the population is hh. Now, N people as the first specimen and N people as the second specimen
- Equation 18 It is a multinomial distribution, and the probability is given by Equation 18 and Equation 19 below for each of the first and second samples.
- the probability of type 1 error calculated in step SB-11 can be expressed by the following equation 20.
- the probability calculated by Equation 20 below means the probability that any one of the results of the independence test performed in Step SB-7 will be significant.
- the probability of the first error calculated by this test coincides with the probability of the first error calculated by the above-described (1 1) allele frequency model.
- the probability of the first type of error due to the method of performing the three tests of the dominant genetic model, the recessive genetic model, and the allele frequency model and determining that it is significant if any of them is significant should be accurately calculated. Is possible. Specifically, for the contingency tables in Fig. 13, Fig. 16 and Fig. 17, an independence test is performed at each locus, and if the deviation is significant, the function f is set to 1, and Otherwise, it can be 0.
- FIG. 3 is a flowchart showing an example of a probability calculation method for accurately calculating the probability of the first type of error in the genotype model.
- a first sample and a second sample containing a plurality of haplotype information related to haplotypes composed of multiple SNPs are acquired (step SC-1).
- the explanation for obtaining the first specimen and the second specimen is the same as the (1 1) allele frequency model and (1 2) dominant 'recessive genetic model described above.
- the haplotype frequency is calculated for each haplotype that may exist ( Step SC—2).
- the explanation for the calculation of the haplotype frequency is the same as the (1 1) allele frequency model and (1 2) dominant / recessive genetic model described above.
- a predetermined value is set according to the force at which each locus of each haplotype has a minor allele or major allele, thereby defining a matrix having the predetermined value as a component (step SC-3).
- the description of the matrix definition is the same as the (1-1) array frequency model and (1-2) dominant / recessive genetic model described above.
- each diplotype power and the first diplotype power in the second sample are set.
- a combination of two specimens with each diplotype power is given (step SC—4).
- the explanation regarding the setting of the diplotype form frequency is the same as the above-mentioned (12) dominant / recessive genetic model.
- the genes at each locus in the first sample based on the multiple diplotype diopter corresponding to the first sample set in step SC-4 and the matrix defined in step SC-3.
- the genotype at each locus in the second sample is calculated based on the multiple dive mouth type frequencies corresponding to the second sample set in step SC-4 and the matrix defined in step SC-3. Calculate the frequency (step SC—5).
- the major homozygote (AA) genotype frequency Y at locus i in the first sample and at locus i in the second sample Major homozygous (AA) genotype frequency Y, minor homozygous at locus i in the first sample (a)
- the genotype frequency Y "of hetero (Aa) at locus i in 2i li and the second sample is
- step SC-5 based on the genotype frequency at each locus in the first sample and the genotype frequency at each locus in the second sample calculated in step SC-5, the even number corresponding to each locus is calculated.
- Create a current table step SC—6). Specifically, the 2 ⁇ 3 contingency table shown in FIG. 18 relating to the genotype frequency calculated using Equation 22 is created for the number of loci (L).
- step SC-7 an independence test is performed for each locus on the contingency table created in step SC-6 (step SC-7).
- step SC—8 Yes
- the number of samples in the first sample, the second sample Each diplo of the first sample set in step SC-4 based on the number of samples, the total number of haplotypes that can be present, the diplotype form frequency of the first sample, the diploma form frequency of the second sample, and the haplotype frequency.
- the probability corresponding to the combination of the type frequency and each diplotype frequency of the second sample is calculated using the preset distribution of the diplotype frequency of the first and second samples (step SC). — 9). Note that the explanation regarding the calculation of the probability is described above. (1 2) Same as dominant / recessive genetic model.
- step SC-10 it is determined whether or not all combinations of each diplotype power of the first sample and each diplotype power of the second sample have been completed, and if all the combinations have been completed (step SC-10: Yes), each time step SC-4 to step SC-9 are repeated, the probability calculated in step SC-9 is added, and the added probability is finally set as the probability of type 1 error ( Step SC—11). On the other hand, if all combinations have not been completed (step SC-10: No), execute steps SC-4 to SC-9 and repeat until all combinations are completed.
- the probability calculation method (1) obtains a first sample and a second sample including a plurality of haplotype information related to haplotypes composed of a plurality of SNPs, and (2 ) Calculate the haplotype frequency for each haplotype that can exist based on the haplotype information contained in the acquired first sample and the haplotype information contained in the acquired second sample, and (3) each haplotype By setting a predetermined value according to whether the sitting position has a minor allele or a major allele, a matrix with the predetermined value as a component is defined.
- each diplotype power for the first sample and each diplotype for the second sample (5) Calculate and set genotypic frequencies at each locus in the first sample based on multiple diplotype shape frequencies corresponding to the set first sample and the defined matrix The genotype frequency at each locus in the second sample is calculated based on a plurality of diplotype frequencies corresponding to the second sample and the defined matrix. (6) Genes at each locus in the calculated first sample Based on the type frequency and the genotype frequency at each locus in the second sample, create an even table corresponding to each locus, and (7) perform an independence test for each locus on the created even table.
- the number of samples in the first sample, the number of samples in the second sample, the total number of haplotypes that can exist, and the diplotype form of the first sample Frequency, diplotype shape of the second sample Based on the protype frequency, the probability corresponding to the combination of each diplotype form frequency of the first sample and each dive mouth type form frequency of the second sample is set to the preset first and second samples. With specimen Calculation is made using a bond distribution relating to diplotype power.
- FIG. 4 is a flowchart showing an example of a probability calculation method for accurately calculating the probability of the first type of error when the exact haplotype frequency is inefficient.
- step SD-1 obtain the first and second samples that contain multiple pieces of haplotype information related to haplotypes composed of multiple SNPs.
- the explanation for obtaining the first specimen and the second specimen is the same as the above-mentioned (1 1) allele frequency model, (1 2) dominant / recessive genetic model, and (1-3) genotype model.
- a predetermined value is set according to the force at which each locus of each haplotype has a minor allele or major allele, thereby defining a matrix having the predetermined value as a component (step SD-2).
- the explanation for the definition of the matrix is the same as the above (1 1) allelic frequency model, (1 2) dominant / recessive genetic model, and (1 3) genotype model.
- the haplotype frequency X of haplotype number j is a random variable.
- the actual value X of the haplotype frequency X of the haplotype number j in the first sample is set for all haplotype numbers j.
- haplotype frequency of haplotype number j in 2 samples Realized value of X Set for type number j.
- step SD -Four the explanation regarding the calculation of the allele frequency is the same as the above-described (1-1) allele frequency model.
- step SD-4 based on the allele frequency at each locus in the first sample and the allele frequency at each locus in the second sample calculated in step SD-4, the contingency table corresponding to each locus is calculated. Create (Step SD—5).
- step SD-6 the independence test is performed for each locus based on the contingency table created in step SD-5 (step SD-6).
- step SD—7 Yes
- the number of samples in the first sample, the number of samples in the second sample Based on the total number of haplotypes that can exist, the haplotype frequency of the first sample, and the haplotype frequency of the second sample, each haplotype frequency of the first sample and each haplotype frequency of the second sample set in step SD—3
- the probability corresponding to the combination is calculated using the preset distribution of the haplotype frequencies of the first and second samples (step SD-8). Specifically, sample number N of the first sample, sample number N of the second sample, haplotypes that may exist
- each haplo of the first sample set in step SD-3 Based on the actual value X of the type frequency, each haplo of the first sample set in step SD-3
- the probability P corresponding to 11 1 21 2 is calculated using the joint distribution related to the haplotype frequency between the first and second samples shown in Equation 24 below. Note that the first and second samples shown in Equation 24 below
- the joint distribution regarding the haplotype frequency is hypergeometric distribution.
- step SD—9 it is determined whether all combinations of each haplotype frequency of the first sample and each haplotype frequency of the second sample have been completed, and all combinations have been completed.
- step SD—9 Yes
- step SD-3 to step SD-8 all the probabilities calculated in step SD-8 are added, and the added probability is finally set as the first type error probability (step SD - Ten).
- Step SD-9 No
- Step SD-3 to Step SD-8 execute Step SD-3 to Step SD-8 and repeat until all combinations are completed.
- the probability calculation method (1) obtains a first sample and a second sample including a plurality of haplotype information related to a haplotype composed of a plurality of SNPs, and (2 ) By defining a predetermined value according to whether each locus of each haplotype has a minor allele or a major allele, a matrix whose component is the predetermined value is defined.
- the probability corresponding to the combination of each haplotype frequency of the set first sample and each haplotype frequency of the second sample is set in advance. Calculated using the joint distribution of the haplotype frequency between the first and second samples. Then, repeat (3), (4), (5), (6), and (7) for all combinations of each haplotype frequency in the first sample and each haplotype frequency in the second sample. By adding the probabilities calculated in (7), we finally calculate the probability of type 1 error. As a result, it is possible to directly calculate the probability of type 1 error, and as a result, it is possible to prevent oversight of disease-related SNPs.
- Equation 25 The probability of type 1 error calculated in step SD-10 can be expressed by Equation 25 below.
- the probability calculated by Equation 25 below means the probability that any one of the results of the independence test performed in step SD-6 will be significant.
- Equation 2 5 Equation 25
- Z is 1 if one of the test results of the independence test performed in step SD-6 is significant, and 0 if it is not. It is a random variable defined to take The variable Z is a function of the random variable Y described above, and
- the random variable Y is a function of the random variable X described above. Also, the random variable X
- Equation 23 the variable Z can be expressed by the function f (x, ..., X)
- the probability calculation method according to the present invention for approximately calculating the probability of the first type of error using the MCMC method is as follows: (2-1) Allele frequency model, (2-2) (2) 3) Genotype model, (2) 4)
- Allele frequency model (2-2) (2) 3) Genotype model, (2) 4)
- FIG. 5 is a flowchart showing an example of a probability calculation method for approximately calculating the first type error probability using the MCMC method V in the allele frequency model. It is assumed that all haplotype frequencies are known.
- a predetermined value is set according to whether each locus of each haplotype has a minor allele or a major allele, thereby defining a matrix having the predetermined value as a component (step SE-1).
- the description regarding the definition of the matrix is the same as the (1-1) allele frequency model in the first embodiment described above.
- a combination of haplotype frequencies is generated for the first and second samples set in advance (step SE-2). Specifically, the realized value X of the haplotype frequency X of haplotype number j in the first sample and the haplotype j of haplotype number j in the second sample
- one of the samples of the first sample and the second sample is selected (step SE-3).
- step SE—6 If the haplotype frequency of the haplotype number j in the selected sample number k is X ⁇ 0 (or X> 0), in step SE-3
- the haplotype frequency of the first haplotype and the haplotype frequency of the second haplotype are updated based on the haplotype frequency of the first haplotype and the haplotype frequency of the second haplotype set in advance (step SE-7).
- the value of c defined by Equation 26 below is calculated.
- h and h are the haplotype frequency of node protype number j and the haplotype frequency of haplotype number j ′, respectively, and their values are preset.
- kj Keep the values of kj and X.
- step SE-7 based on a plurality of haplotype frequencies corresponding to the first sample updated in step SE-7 and the matrix defined in step SE-1, the key at each locus in the first sample is determined.
- the allele at each locus in the second sample is calculated based on the multiple haptic type frequencies corresponding to the second sample after the real frequency is calculated and updated in step SE-7 and the matrix defined in step SE-1 Calculate the frequency (step SE—8).
- the explanation for the calculation of the allele frequency is the same as the (1-1) allele frequency model of the first embodiment described above.
- Step SE-9 based on the allele frequencies at each locus in the first sample and the allele frequencies at each locus in the second sample calculated in step SE-8, the contingency table corresponding to each locus is created. (Step SE-9).
- Step SE-10 an independence test is performed for each locus based on the contingency table created in step SE-9 (step SE-10).
- Step SE-11 Yes
- the cumulative number in that case is updated (Step SE). — 12). Specifically, 1 is added to the cumulative number.
- step SE-13 it is determined whether or not the force has reached the predetermined number of times, and when the predetermined number of times has ended (step SE-13: Yes), the ratio of the final cumulative number to the predetermined number of times (cumulative number ⁇ predetermined number of times) Is calculated as the probability of type 1 error (step SE-14). On the other hand, if the predetermined number of times is over (Step SE-13: No), execute Step SE-3 to Step SE-12 and repeat until the predetermined number of times is completed.
- the probability calculation method that works for the present invention generates haplotype samples according to the multinomial distribution expected for each sample by the Metropolis-Hastings method, and assigns each sample to each sample. ! /, Perform a test at each position, and find the percentage that is significant at any position.
- the state of the Markov chain is represented by a contingency table of 2 X M (total number of haplotypes that can exist) based on the actual value of the haplotype frequency.
- the test is performed by sampling a conservative test method or other haplotype copy number.
- the probability of type 1 error can be calculated by calculating the proportion of steps determined to be significant.
- the first and second samples are independently generated by the Monte Carlo sampler that follows the multinomial distribution, and the above-described calculation of the function f is performed. A method is conceivable.
- Figure 6 shows a dominant / recessive genetic model that approximates the probability of type 1 error using the MCMC method. It is a flowchart which shows an example of the probability calculation method calculated in step. It is assumed that all haplotype frequencies are known.
- a predetermined value is set according to whether each locus of each haplotype has a minor allele or a major allele, thereby defining a matrix having the predetermined value as a component (step SF-1).
- the description of the matrix definition is the same as the (1-2) dominant / recessive genetic model in the first embodiment described above.
- a combination of diplotype powers is generated for the first and second samples set in advance (step SF-2). Specifically, the actual value X of the diplotype form frequency of diplotype form number ⁇ 'in the first sample is set to ⁇ ⁇ ' for all diplotype form numbers ⁇ '.
- the realization value X of the diplotype form frequency of the diplotype form number jj 'in the second sample is generated so as to satisfy the above-mentioned equation 14. (X, ..., X, X
- step SF—3 either the first sample or the second sample is selected.
- Step SF—4 select the first diplotype. Specifically, the diplotype model number '(1 ⁇ ordered natural numbers i and j ⁇ M ⁇ S 1 ⁇ ); L is the number of sitting positions) is selected.
- the diplotype frequency of the first diplotype form in the specimen selected in step SF-3 is 0 (SF-5: Yes)
- the first diplotype form is different from the first diplotype form. 2
- Select the diplotype type (Step SF—6). Specifically, the diplotype shape number j j 'of the selected sample number k sample x half 0 (or X
- the value of the diplotype power X is maintained as kjljl 'kjljl'.
- the diplotype of the first diplotype Degree and second diplotype Update the diplotype form factor for (step SF—7). Specifically, first, the value of c defined by Equation 27 below is calculated. In Equation 27 below, h, h, h and h
- jl jl '] 2 is the haplotype frequency of haplotype number j and the haplotype of j'
- each locus in the first sample is Based on multiple diplotype powers corresponding to the second sample after calculation of the allele frequency and updated in step SF—7 and the matrix defined in step SF— !, and each locus in the second sample Calculate the allele frequency at (step SF—8).
- the explanation for calculating the allele frequency is the same as the (12) dominant 'recessive genetic model of the first embodiment described above.
- Step SF-9 based on the allele frequencies at each locus in the first sample and the allele frequencies at each locus in the second sample calculated in step SF-8, the contingency table corresponding to each locus is created. (Step SF-9).
- step SF-10 the independence test is performed for each locus based on the contingency table created in step SF-9 (step SF-10).
- step SF—11 Yes
- step SF— 12 update the cumulative number in that case. Specifically, 1 is added to the cumulative number.
- Step SF—13 Yes
- Step SF-13 No
- Step SF-3 Step SF-12 and repeat until the predetermined number of times ends.
- the probability calculation method according to the present invention generates a diplotype sample according to the multinomial distribution expected for each sample by the Metropolis-Hastings method, and each locus for each sample. Perform a test at, and find the proportion that is significant at any of these loci.
- FIG. 7 is a flowchart showing an example of a probability calculation method for approximately calculating the first type error probability using the MCMC method in the genotype model. It is assumed that all haplotype frequencies are known.
- a matrix having the predetermined value as a component is defined by setting a predetermined value according to whether each locus of each haplotype has a minor allele or a major allele (step SG-1).
- the description regarding the definition of the matrix is the same as the (1-3) genotype model in the first embodiment described above.
- step SG-2 a combination of diplotype form frequencies is generated for the first and second samples set in advance (step SG-2).
- the explanation for the generation of the diplotype form frequency is the same as in (2-2) Dominant / recessive genetic model described above.
- step SG-3 either the first sample or the second sample is selected.
- step SG-3 if the diplotype frequency of the first diplotype form in the sample selected in step SG-3 is 0 (SG-5: Yes), the second diplotype form is different from the first diplotype form.
- step SG—6 The explanation for the selection of the second diplotype is the same as in (2-2) Dominant / recessive genetic model described above.
- the diplotype of the first diplotype is updated (step SG7 ).
- the explanation for updating the diplotype form frequency is the same as the (2-2) dominant / recessive genetic model described above.
- genotypes at each locus in the first sample based on the multiple diplotypes corresponding to the first sample after updating in step SG-7 and the matrix defined in step SG-1
- the genotype at each locus in the second sample based on the multiple diplotype frequencies corresponding to the second sample after updating the frequency in step SG-7 and the matrix defined in step SG-1 Calculate the frequency (step SG—8).
- the description regarding the calculation of the genotype frequency is the same as the (1-3) genotype model of the first embodiment described above.
- step SG-8 based on the genotype frequency at each locus in the first sample and the genotype frequency at each locus in the second sample calculated in step SG-8, the contingency table corresponding to each locus. (Step SG—9).
- step SG-10 the independence test is performed for each sitting position using the contingency table created in step SG-9 (step SG-10).
- step SG—11 Yes
- step SG— 12 the cumulative number in that case is updated. Specifically, 1 is added to the cumulative number.
- step SG—13: Yes the ratio of the final cumulative number to the predetermined number (cumulative number ⁇ predetermined number of times) Is calculated as the probability of type 1 error.
- step SG-13: No the ratio of the final cumulative number to the predetermined number (cumulative number ⁇ predetermined number of times) Is calculated as the probability of type 1 error.
- the probability calculation method that works on the present invention is Metropolis- Hastings.
- a diplotype sample that follows the expected multinomial distribution for each specimen is generated by the method, and a test is performed at each locus for each sample, and a ratio that is significant at any one of the loci is obtained.
- the true haplotype frequency is often unknown. Therefore, even if the haplotype frequency is not weak, if the diplotype shape in the specimen is weak, it can be calculated as follows.
- Permutation haplotype copy numbers follow a multinomial distribution given the haplotype frequency.
- probability under specific observational data rather than the probability under certain conditions of the sample space. For example, consider the probability under the condition of a set of results that satisfy the restriction as shown in Equation 28 or 29 below.
- the above-described dominant genetic model, recessive genetic model, allele frequency model, and genotype model can all be tested if the combined diplotype copy number is known.
- ⁇ j ⁇ i follows hypergeometric distribution.
- the probabilities are expressed by the following formula 30 and the following formula 31, respectively.
- Equation 31 (A-2) Generate a sample according to the hypergeometric distribution shown in Equation 31 using the MCMC method. However, ⁇ y ⁇ satisfies Equation 32 below.
- Equation 32 (A-2)
- A-3 The generated sample is tested at each locus using a dominant genetic model, a recessive genetic model, an allele frequency model, and a genotype model.
- the number of sitting positions is given under certain constraints.
- the state space of the Markov chain (Markov-chain) is a plurality of element y forces that take different values under the predetermined constraint.
- N 1 A given (given) fixed non-negative integer value.
- N is included in sample number k
- the diplotype model number uv is determined by selecting two ordered integers (u, V) with equal probability.
- U is ;
- L satisfies the number of sitting positions, and
- V satisfies "l ⁇ v ⁇ u”.
- the diplotype shape number uv of the selected sample number k sample uv has two orders different from the selected (u, V) when y
- the diplotype model number WS is further determined by further selecting (W, S), which is an integer.
- the power described for the independence test using one of the three models (dominant genetic model, recessive genetic model, and genotype model).
- the probability calculation method according to the present invention includes three methods. It can be easily extended to an independence test using the entire model. Therefore, as described in (B-6) above, the independence test is performed at each sitting position using three different contingency tables. Independence test If the constant test result is significant, the test result of the overall (all loci) independence test is significant.
- the probability calculation method according to the present invention can directly calculate the probability of the first type of error in consideration of the haplotype frequency, and as a result, the disease-related SNP can be calculated. Oversight can be prevented. Therefore, the probability calculation method according to the present invention is extremely useful in fields such as medicine and drug discovery.
Landscapes
- Bioinformatics & Cheminformatics (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Genetics & Genomics (AREA)
- Biotechnology (AREA)
- Biophysics (AREA)
- Chemical & Material Sciences (AREA)
- Molecular Biology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Bioinformatics & Computational Biology (AREA)
- Analytical Chemistry (AREA)
- Evolutionary Biology (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Theoretical Computer Science (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2006546685A JP4755601B2 (ja) | 2004-12-03 | 2005-12-05 | 確率算出方法 |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2004-351839 | 2004-12-03 | ||
| JP2004351839 | 2004-12-03 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2006059770A1 true WO2006059770A1 (ja) | 2006-06-08 |
Family
ID=36565195
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2005/022317 Ceased WO2006059770A1 (ja) | 2004-12-03 | 2005-12-05 | 確率算出方法 |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JP4755601B2 (ja) |
| WO (1) | WO2006059770A1 (ja) |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2004192018A (ja) * | 2002-10-16 | 2004-07-08 | Japan Biological Informatics Consortium | Dnaプールによるハプロタイプ頻度推定方法 |
| JP2004229511A (ja) * | 2003-01-28 | 2004-08-19 | Ntt Data Corp | ハプロタイプ解析装置、および、ハプロタイプ解析方法をコンピュータに実行させることを特徴とするプログラム |
-
2005
- 2005-12-05 JP JP2006546685A patent/JP4755601B2/ja not_active Expired - Fee Related
- 2005-12-05 WO PCT/JP2005/022317 patent/WO2006059770A1/ja not_active Ceased
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2004192018A (ja) * | 2002-10-16 | 2004-07-08 | Japan Biological Informatics Consortium | Dnaプールによるハプロタイプ頻度推定方法 |
| JP2004229511A (ja) * | 2003-01-28 | 2004-08-19 | Ntt Data Corp | ハプロタイプ解析装置、および、ハプロタイプ解析方法をコンピュータに実行させることを特徴とするプログラム |
Non-Patent Citations (4)
| Title |
|---|
| KAMATANI N. ET AL: "Haplotype o Kiso ni shita Human Genome Takei to Hyogengata no Kanren no Kaiseki Shuho", JAPANESE FEDERATION OF STATISTICAL SCIENCE ASSOCIATIONS, vol. 2004, 6 September 2004 (2004-09-06), pages 188 - 189, XP002999646 * |
| NYHOLT D.R.: "A simple correction for multiple testing for single-nucleotide polymorphisms in linkage disequilibrium with each other", AM.J.HUM.GENET., vol. 74, 2004, pages 765 - 769, XP002999648 * |
| TOMITA M. ET AL: "Rensa Fuheiko Kaiseki ni okeru Tazai no Kankei", JAPANESE FEDERATION OF STATISTICAL SCIENCE ASSOCIATIONS, vol. 2003, 5 September 2003 (2003-09-05), pages 227 - 228, XP002999647 * |
| XIONG M. ET AL: "The haplotype linkage disequilibrium test for genome-wide screens: its power and study design", PACIFIC SYMPOSIUM ON BIOCOMPUTING, 2000, pages 672 - 681, XP002999649 * |
Also Published As
| Publication number | Publication date |
|---|---|
| JP4755601B2 (ja) | 2011-08-24 |
| JPWO2006059770A1 (ja) | 2008-06-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Mathers et al. | Chromosome-scale genome assemblies of aphids reveal extensively rearranged autosomes and long-term conservation of the X chromosome | |
| Di Rienzo et al. | Heterogeneity of microsatellite mutations within and between loci, and implications for human demographic histories | |
| Liu et al. | Ancient and modern genomes unravel the evolutionary history of the rhinoceros family | |
| Burgess et al. | Estimation of hominoid ancestral population sizes under Bayesian coalescent models incorporating mutation rate variation and sequencing errors | |
| Vakhrusheva et al. | Genomic signatures of recombination in a natural population of the bdelloid rotifer Adineta vaga | |
| Fernández-Mazuecos et al. | Resolving recent plant radiations: power and robustness of genotyping-by-sequencing | |
| Thomas et al. | Coding single-nucleotide polymorphisms associated with complex vs. Mendelian disease: evolutionary evidence for differences in molecular effects | |
| Tsai et al. | Population genomics of the wild yeast Saccharomyces paradoxus: quantifying the life cycle | |
| Schaffner et al. | Calibrating a coalescent simulation of human genome sequence variation | |
| Blumenstiel et al. | An age-of-allele test of neutrality for transposable element insertions | |
| Morgan et al. | Informatics resources for the Collaborative Cross and related mouse populations | |
| Ding et al. | Genome structure-based Juglandaceae phylogenies contradict alignment-based phylogenies and substitution rates vary with DNA repair genes | |
| Whelan et al. | Estimating the frequency of events that cause multiple-nucleotide changes | |
| Long et al. | Low base-substitution mutation rate in the germline genome of the ciliate Tetrahymena thermophila | |
| Kim | Allele frequency distribution under recurrent selective sweeps | |
| Collet et al. | Rapid evolution of the intersexual genetic correlation for fitness in Drosophila melanogaster | |
| Menardo et al. | Reconstructing the evolutionary history of powdery mildew lineages (Blumeria graminis) at different evolutionary time scales with NGS data | |
| Barton et al. | New methods for inferring the distribution of fitness effects for INDELs and SNPs | |
| Mugal et al. | Polymorphism data assist estimation of the nonsynonymous over synonymous fixation rate ratio ω for closely related species | |
| Riebler et al. | Bayesian variable selection for detecting adaptive genomic differences among populations | |
| Ebler et al. | Pangenome-based genome inference | |
| Maruki et al. | Evolutionary genomics of a subdivided species | |
| WO2015006668A1 (en) | Methods for identification of individuals | |
| Fazalova et al. | Low spontaneous mutation rate and pleistocene radiation of pea aphids | |
| Cohen et al. | The social supergene dates back to the speciation time of two Solenopsis fire ant species |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| AK | Designated states |
Kind code of ref document: A1 Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS JP KE KG KM KN KP KR KZ LC LK LR LS LT LU LV LY MA MD MG MK MN MW MX MZ NA NG NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SM SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW |
|
| AL | Designated countries for regional patents |
Kind code of ref document: A1 Designated state(s): GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IS IT LT LU LV MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application | ||
| WWE | Wipo information: entry into national phase |
Ref document number: 2006546685 Country of ref document: JP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 05811804 Country of ref document: EP Kind code of ref document: A1 |
|
| WWW | Wipo information: withdrawn in national office |
Ref document number: 5811804 Country of ref document: EP |





















