WO2013102441A1 - Cyp450基因型别数据库及基因分型、酶活性鉴定方法 - Google Patents

Cyp450基因型别数据库及基因分型、酶活性鉴定方法 Download PDF

Info

Publication number
WO2013102441A1
WO2013102441A1 PCT/CN2013/070080 CN2013070080W WO2013102441A1 WO 2013102441 A1 WO2013102441 A1 WO 2013102441A1 CN 2013070080 W CN2013070080 W CN 2013070080W WO 2013102441 A1 WO2013102441 A1 WO 2013102441A1
Authority
WO
WIPO (PCT)
Prior art keywords
cyp450
sequence
gene
sample
genotype
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2013/070080
Other languages
English (en)
French (fr)
Inventor
刘晓
徐怀前
张伟
苏政
王冠
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
BGI Shenzhen Co Ltd
Original Assignee
BGI Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by BGI Shenzhen Co Ltd filed Critical BGI Shenzhen Co Ltd
Publication of WO2013102441A1 publication Critical patent/WO2013102441A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6883Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • G16B20/20Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6883Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
    • C12Q1/6886Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material for cancer
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • G16B20/10Ploidy or copy number detection
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • G16B30/10Sequence alignment; Homology search
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B50/00ICT programming tools or database systems specially adapted for bioinformatics
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B50/00ICT programming tools or database systems specially adapted for bioinformatics
    • G16B50/30Data warehousing; Computing architectures
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/156Polymorphic or mutational markers
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids

Definitions

  • the invention relates to the field of gene detection, in particular to a standard genotype database of CYP450 and a construction method thereof, and a method for CYP450 genotyping and enzyme activity identification. Background technique
  • the cytochrome P450 cartridge is called CYP450.
  • CYP450 cytochrome P450 oxidases
  • the participating drug metabolism is mainly concentrated in the CYP1, CYP2 and CYP3 families, and is involved in the metabolism of more than 90% of the drugs.
  • the important role of CYP450 in drug metabolism leads to the CYP450 gene polymorphism becoming one of the most important factors affecting individual drug differences. Polymorphisms include point mutations, insertions or deletions, deletion or replication of whole genes, and ultimately enzyme activity. Increased, weakened or completely missing.
  • the CYP450 polymorphism can result in severe side effects or inactivity in individuals at standard drug doses.
  • warfarin as the current first-line oral anticoagulant, its effect of preventing and treating thromboembolism and the convenience and economy of oral administration are obvious, but the warfarin treatment window is narrow, and The differences in the individual administration and the ethnic differences are very large. To achieve the same effect, the high and low doses can differ by more than 10 times. Studies have shown that this is closely related to the genetic polymorphism of CYP2C9 between individuals.
  • tamoxifen which is widely used in the treatment of breast and ovarian cancer, requires a series of CYP450 enzymes to form a therapeutically effective product, and the genotype of CYP2D6 affects him.
  • the main limiting factor for the efficacy of cixifloxacin, its defective form can lead to shorter disease-free survival, while the CYP2D6 strong metabolite requires only a smaller dose of tamoxifen to achieve the same effect.
  • CYP450 gene is closely related to the occurrence of many diseases.
  • CYP21B gene defects account for 90%-95% of the cause, and CYP11B1 gene mutation accounts for about 5%.
  • Defective mutations of CYP17AK CYP1 1A1 can also cause this disease.
  • CYP450 The important role of CYP450 in metabolism and disease occurrence makes the detection of CYP450 gene polymorphism very important social significance and research value.
  • the current techniques for detecting CYP450 mutations mainly focus on known mutations of single or several P450 genes; and, the mutation sites in each study are relative to their specific sequences, and there is no uniform measurement between them. Standard.
  • the existing CYP450 detection method is not efficient for detecting unknown mutation sites still present in the CYP450 gene; it takes a long time to detect a large number of samples, which does not meet the needs of research and practical applications.
  • the present invention aims to solve at least one of the technical problems existing in the prior art. To this end, it is an object of the present invention to provide a standard database of CYP450 genotypes and a method for constructing the same, and a method for rapidly detecting CYP450 genotypes based on the database and a method for identifying CYP450 enzyme activity.
  • the invention provides a method of constructing a CYP450 gene standard type database.
  • the method comprises the steps of: aligning a specific sequence corresponding to the mutation information of the CYP450 genotype with a human whole genome standard sequence, obtaining a CYP450 specific sequence and a human whole genome standard sequence at each base Correspondence in position; According to the obtained correspondence, the CYP450 genotype was converted into a genotype with reference to the human whole genome standard sequence, and the normalized genotype of the CYP450 gene was obtained.
  • the inventors have surprisingly found that the method for constructing a standard type database of the C YP450 gene of the present invention can efficiently construct a CYP450 gene standard type database, and the database can provide a unified standard for each CYP450 genotyping, thereby
  • the CYP450 gene standard type database for the known genotype of CYP450 gene, can quickly and accurately determine the genotype of the CYP450 gene of the sample to be tested, and further, can diagnose or treat a disease involving the CYP450 gene for the sample to be tested. Treatment, etc. provide a more accurate basis for judgment.
  • the present invention also provides a CYP450 gene standard type database constructed by the method of the present invention for constructing a CYP450 gene standard type database as described above.
  • the inventors found that the CYP450 gene standard type database can provide a unified standard for each CYP450 genotyping, thereby enabling rapid and accurate determination of the CYP450 gene of known genotypes based on the CYP450 gene standard type database.
  • the genotype of the CYP450 gene of the sample further provides a more accurate basis for disease diagnosis or drug treatment involving the C YP450 gene for the sample to be tested.
  • the present invention also provides a method of genotyping CYP450.
  • the method comprises: obtaining an exon sequence of a sample CYP450 gene to be tested, sequencing with a high-throughput sequencing platform, and performing data analysis, and analyzing the result with the CYP450 genotype of the present invention described above Compare with other databases to get the genotype of the sample to be tested.
  • the CYP450 gene of the known genotype can quickly and accurately determine the genotype of the CYP450 gene of the sample to be tested, and further, it can provide more for the diagnosis or drug treatment of the CYP450 gene for the sample to be tested. Accurate judgment basis.
  • the present invention also provides a method for identifying CYP450 enzyme activity.
  • the method comprises: obtaining an exon sequence of a sample CYP450 gene to be tested, sequencing with a high-throughput sequencing platform, and performing data analysis, and combining the analysis result with the CYP450 provided by the present invention as described above
  • the CYP450 gene standard type database of the enzyme activity information is compared, the genotype of the sample to be tested is obtained, and the CYP450 enzyme activity result of the sample to be tested is obtained according to the enzyme activity information corresponding to the genotype.
  • the CYP450 gene of known genotype can quickly and accurately determine the genotype of CYP450 gene and the activity of CYP450 enzyme activity of the sample to be tested, and further, can provide more accurate judgment for diseases or drugs involving CYP450. in accordance with.
  • FIG. 1 is a schematic flow chart showing the steps of obtaining an exon sequence of a sample CYP450 gene in a CYP450 genotyping and enzymatic activity identification method according to an embodiment of the present invention
  • FIG. 2 is a schematic flow chart showing the steps of data analysis in the CYP450 genotyping and enzyme activity identification method of the present invention according to an embodiment of the present invention. Detailed description of the invention
  • the present invention provides a method of constructing a CYP450 gene standard type database.
  • the method comprises the steps of: aligning a specific sequence corresponding to the mutation information of the CYP450 genotype with a human whole genome standard sequence, obtaining a CYP450 specific sequence and a human whole genome standard sequence at each base Correspondence in position; According to the obtained correspondence, the CYP450 genotype was converted into a genotype with reference to the human whole genome standard sequence, and the normalized genotype of the CYP450 gene was obtained.
  • the CYP450 gene standard type database can be efficiently constructed by the method for constructing a CYP450 gene standard type database of the present invention, and the database can provide a unified standard for each CYP450 genotyping, thereby
  • the CYP450 gene standard type database for the known genotype of CYP450 gene, can quickly and accurately determine the genotype of the CYP450 gene of the sample to be tested, and further, can diagnose or treat the disease involving the CYP450 gene for the sample to be tested. Provide a more accurate basis for judgment.
  • the CYP450 gene comprises a member selected from the group consisting of CYP11A1, CYP11B1, CYP1 1B2, CYP17AK CYP1A1, CYP1A2, CYP1B1, CYP20A1, CYP21A2, CYP24A CYP26AK CYP26B1, CYP26C CYP27A1, CYP27B1, CYP27C1, CYP2A13, CYP2A6, CYP2A7, CYP2B6, CYP2C18, CYP2C19. CYP2C8, CYP2C9. CYP2D6.
  • the above method further comprises localizing the CYP450 gene to a human full base
  • the specific sequence corresponding to the CYP450 genotype mutation information was obtained by determining the starting position and the termination position of the CYP450 gene coding sequence in the standard sequence of the group.
  • the specific sequence corresponding to the CYP450 genotype mutation information comprises a DNA fragment of 5000 bp upstream from the start position of the coding sequence of the CYP450 gene to a region of 500 bp downstream of the stop position of the coding sequence.
  • the human whole genome standard sequence is hgl9.
  • the present invention also provides a CYP450 gene standard type database constructed by the method of the present invention for constructing a CYP450 gene standard type database as described above.
  • the inventors found that the CYP450 gene standard type database can provide a unified standard for each CYP450 genotyping, thereby enabling rapid and accurate determination of the CYP450 gene of known genotypes based on the CYP450 gene standard type database.
  • the genotype of the CYP450 gene in the sample further, can be used for the sample to be tested
  • the diagnosis or drug treatment of the C YP450 gene provides a more accurate basis for judgment.
  • the normalized genotypes of the CYP450 gene correspond to enzyme activity information.
  • the CYP450 gene comprises 58 humans as described in Table 1 below
  • At least one of the CYP450 genes At least one of the CYP450 genes.
  • the present invention also provides a method of genotyping CYP450.
  • the method comprises: obtaining an exon sequence of a sample CYP450 gene to be tested, sequencing by a high-throughput sequencing platform, and performing data analysis, and analyzing the result with the CYP450 genotype of the invention described above. Database The comparison is made to obtain the genotype of the sample to be tested.
  • the CYP450 gene of the known genotype can quickly and accurately determine the genotype of the CYP450 gene of the sample to be tested, and further, it can provide more for the diagnosis or drug treatment of the CYP450 gene for the sample to be tested. Accurate judgment basis.
  • the present invention also provides a method for identifying CYP450 enzyme activity.
  • the method comprises: obtaining an exon sequence of a sample CYP450 gene to be tested, sequencing using a high-throughput sequencing platform, and performing data analysis, and analyzing the result with the CYP450 enzyme provided by the present invention as described above
  • the CYP450 gene standard type database of the activity information is compared, the genotype of the sample to be tested is obtained, and the CYP450 enzyme activity result of the sample to be tested is obtained according to the enzyme activity information corresponding to the genotype.
  • the CYP450 gene of known genotype can quickly and accurately determine the genotype of CYP450 gene and the activity of CYP450 enzyme activity of the sample to be tested, and further, can provide more accurate judgment for diseases or drugs involving CYP450. in accordance with.
  • obtaining the exon sequence of the sample CYP450 gene to be tested is achieved by the following steps:
  • step B The sequence capture library prepared in step B is hybridized with the chip of step A to obtain an exon library of the CYP450 gene of the sample to be tested.
  • the chip in the above step A, contains an oligonucleotide probe which is complementary to each exon sequence of 58 human CYP450 genes, respectively, and the length of the oligonucleotide probe is 55-105bp.
  • the genomic DNA of the sample to be tested is interrupted into fragments of 200 to 300 bp in size.
  • the terminal treatment comprises performing a terminal repair to form a blunt-end phosphorylated DNA fragment, and adding a base "A" at the 3' end of the blunt-ended phosphorylated DNA fragment " and further connect the tags.
  • the sequence capture libraries from the plurality of different samples to be tested are mixed and then hybridized with the chip of the step A, each library having a different The base sequence of the tag is different from each other, and the length of the base sequence of the tag is preferably 6 to 8 bp.
  • the data analysis further comprises:
  • step i using the human whole genome standard sequence as a reference sequence, comparing the sequence obtained in step i with the comparison software, preferably using SOAP or BWA;
  • Iii selecting a sequence aligned to the target region, wherein the target region refers to the region where the exon sequence of the CYP450 gene is located;
  • Iv Perform variation analysis after passing the data quality control, and the variation analysis includes detecting at least one of the following: single nucleotide polymorphism, insertion and deletion, structural variation, copy number variation.
  • the method for constructing a CYP450 gene standard type database of the present invention can provide a unified standard database for each CYP450 genotyping. Based on this database, the CYP450 gene of the known genotype can quickly and accurately give the CYP450 gene genotype of the sample to be tested, thereby providing a more accurate basis for the disease or drug involved in CYP450. Further, by adding information on the enzyme activity corresponding to the genotype in the CYP450 gene standard type database, it is possible to directly obtain the CYP450 enzyme activity result while genotyping the sample to be tested.
  • the CYP450 genotyping method of the present invention after obtaining the exon sequences of all CYP450 genes, and then performing sequencing analysis, and comparing with the CYP450 gene standard type database, can effectively identify unknown mutations in the CYP450 gene. Point to detect.
  • the CYP450 genotyping and enzyme activity identification method of the invention utilizes the high-flux property of the chip capture, and can simultaneously detect up to hundreds of samples in one experiment, which not only improves the number of detection samples, but also greatly reduces the number of samples. The cost of testing each sample.
  • the CYP450 genotyping and enzyme activity identification method of the present invention comprises 57 CYP450 oxidase genes and one P450 reductase gene which have been identified in human beings, and has wide coverage, which is greatly convenient for CYP450. Gene research.
  • the present invention also provides an overall technical scheme for constructing a CYP450 genotype standardized database, CYP450 genotyping and enzymatic activity identification, which is based on high-throughput sequencing of target regions after sequence capture, specifically, The steps can be included:
  • CYP450 functional genes including 57 CYP450 oxidase genes and 1 CYP450 reductase gene (see Table 1 above), by BLAST (http://blast.ncbi.nlm.nih.gov/Blast.cgi)
  • BLAST http://blast.ncbi.nlm.nih.gov/Blast.cgi
  • all genotype sequences of 58 C YP450 genes were aligned with the hg 19 reference sequence, and the mutation site information relative to hgl9 was obtained based on the alignment result.
  • All genotypes of the CYP450 gene are converted to a uniform format and standard. The genotype is converted to a hgl9-based type based on the annotation information of the gene on the whole genome. Specifically, the following steps are included:
  • the information on the mutation information and type and enzyme activity of all genotypes of the existing 58 CYP450 genes were collected.
  • the information mainly includes the name of the genotype, the protein shield number corresponding to the genotype, the mutation information of the genotype and the specific sequence, the activity of the enzyme corresponding to the genotype in the living body, and the enzyme corresponding to the genotype in vitro test. Activity in etc.
  • the "specific sequence” as used in the present application refers to a DNA sequence fragment or a cDNA sequence used as a reference in the study.
  • the inventors first identified the position of the CYP450 gene on hgl9, and then from the 5000 bp upstream of the CDS start position to 500 bp downstream of the CDS termination position as the region of the CYP450 gene, but some genotypes were farther away from the CDS region. Exceeding the above range, for these genes, the inventor will set the region of this gene to be longer, in order to include the above-mentioned mutation sites.
  • a specific sequence was BLAST aligned with hgl9. If the specific sequence is cDNA, the inventors used BLAT for alignment.
  • the specific sequence may be aligned with multiple positions on hgl9, selecting the alignment of the best position, analyzing the bases at each position, and obtaining a specific sequence with hgl9 at each Base correspondence at one position. It should be noted that if it is aligned to the negative strand on the stain, it needs to be converted to tt on the positive strand.
  • CYP450 genotypes were converted to hgl9-based mutation site information based on the alignment of specific sequences with hgl9.
  • the CDS starting position and the defined gene region of the above gene are required.
  • the mutation site information on some genotypes is negatively chained, and the negative strand information needs to be converted into a positive strand during the conversion.
  • the file format is organized and information on the genotype enzyme activity is also added. Specific examples are listed in Table 2. Then check the correctness of the results.
  • CYP2C 6701979 Del Maekawa et al
  • the exon sequence of the CYP450 gene of the sample to be tested is determined, specifically:
  • the human genome hgl9 was used as the reference sequence, and all the exon regions of the 58 genes were selected as the target sequences, and the total length of the target sequences was about 276 kb.
  • an oligonucleotide capture probe of approximately 55-105 bp in length complementary to the exon sequence was designed.
  • a high density of immobilized capture probes was immobilized on the chip to form a capture chip containing all of the exon capture probes of 58 CYP450 genes.
  • the designed probe was produced by Roche-Nimblegen and assembled and fixed on the capture chip.
  • probe sequences are designed with reference to hg 19, and because of the differences in genomic sequences among different species, the probe is preferentially suitable for human genomic DNA capture, and other genomes of species with higher homology to the human genome can be applied. However, the capture effect may not be as good as the human genome.
  • Different species can design probes similar to the present invention based on their reference sequences for capture in target regions of different species.
  • the purified and fragmented DNA is recovered by the action of an enzyme such as T4 DNA polymerase, Klenow fragment and T4 polynucleotide kinase using dNTP as a substrate to form a blunt-ended terminal phosphorylated DNA fragment, which is then purified.
  • an enzyme such as T4 DNA polymerase, Klenow fragment and T4 polynucleotide kinase using dNTP as a substrate to form a blunt-ended terminal phosphorylated DNA fragment, which is then purified.
  • the Kendow fragment (3,-5,exo-) polymerase and dATP were used to add the base "A" to the 3' end of the purified terminal phosphorylated DNA fragment, followed by purification.
  • the purified DNA fragment "A” was ligated to the tag linker using T4 DNA ligase, and the adaptor ligation product was purified using a kit.
  • the PCR product, the linker blocking sequence, and the Cotl DNA obtained above were mixed to constitute an exon capture library.
  • the libraries of the plurality of samples to be tested are mixed, in order to distinguish the libraries from different samples in the sequencing, the DNA of each library contains a different 6 bp or 8 bp tag base in the linker when the tag linker is ligated.
  • the base sequence, the amount of DNA mixed per library can be mixed in equal amounts or in a certain ratio as needed. It should be noted that when the amount of sequencing data is the same for each sample, the amount of mixed DNA in each library is the same; some studies may have different amounts of sequencing data, and the amount of library used will be different.
  • the mixing ratio is in accordance with the field. The specific research purpose or design requirements of the technician are determined.
  • the probe hybridizes to the target area
  • the hybrid library of the above plurality of samples was hybridized with the chip according to the mblegen solid phase chip hybridization standard operating instructions.
  • the hybridized DNA is eluted, purified, and then amplified using a linker sequence as a primer to obtain an amplified product.
  • the amplification products obtained above were subjected to quality control using Agilent 2100 and Q-PCR, and were ready for quality control.
  • Each of the amplified products constitutes a sequencing library.
  • the sequencing library obtained above was sequenced using a sequencing method that was synthesized by sequencing. Then, using the human genome hgl9 (UCSC) as a reference sequence, the obtained sequencing data was analyzed.
  • the data analysis step can include the following steps:
  • Step one Filter Remove sequences with low mass values and contamination with sequencing primers
  • each base in the sequence corresponds to a sequencing quality value
  • the average quality value of the sequence is calculated, if the average quality value of the sequence is lower than The conventional empirical threshold, this sequence will be filtered out; on the other hand, the sequencing sequence may be contaminated by the connector on the machine, and the sequence contained in this part will also be filtered out.
  • the sequence filtered by step 1 is aligned with alignment software (such as SOAP, BWA). These alignment software is able to select an optimal alignment position for a sequence. For multiple repeat sequences in the alignment position, the software selects a position output and adds a label.
  • alignment software such as SOAP, BWA
  • Step 3 Select the sequence that is aligned to the target area. After the hybridization of the chip, the sequence of some non-target regions will be captured. In step 2, the whole genome sequence of hgl9 is used as the reference sequence, and the sequence of the non-target region will be compared to the corresponding position according to the best matching principle, and the comparison will not be performed. target area. The sequence aligned to the target area is selected for subsequent analysis, ensuring that the selected sequences are all target region sequences.
  • Data control includes multiple aspects, such as the percentage of aligned sequences, the percentage of unique reads (only one optimal alignment position when the sequence is aligned with the reference sequence), the ratio of duplication (same sequence), sequencing Depth, coverage of the target area, etc.
  • quality controls are subject to conventional empirical thresholds for further analysis. For example, the depth of the sample is consistent with expectations, and the single base depth overlay is subject to the Poisson distribution.
  • the mutation analysis can be performed, including detection of SNP (single nucleotide polymorphism), INDEL (insertion and deletion), SV (structural variation), and CNV (copy number variation). Each variation detection can be implemented in different ways as needed.
  • the mutation site information in each gene was sorted and compared with the corresponding genotypes in the previously compiled CYP450 standard database to obtain the genotype of each sample. Since humans are diploid organisms, there are at most two types of each gene type. The final CYP450 gene typing result is a homozygous or heterozygous type. Some genotypes have information on the corresponding enzyme activity, so after the sample genotyping, the reaction of the sample to the enzyme activity can also be obtained.
  • the CYP450 gene standard type database constructed by the present invention comprises all functional 57 CYP450 oxidase genes and one CYP450 reductase gene which have been identified in humans, and has a wide range of contents, and can be detected by using the database. All known and unknown polymorphic sites of these CYP450 genes in the sample to be tested greatly facilitated the study of the CYP450 gene.
  • a CYP450 gene standard type database with hgl9 as a reference sequence can be established.
  • genotype CYPP450 gene the corresponding genotype information can be quickly and accurately given, which provides more information for diseases or drugs involving CYP450. Accurate judgment basis.
  • This example constructs a genotype standardized database according to the method for constructing a CYP450 genotype standardized database of the present invention, specifically: The inventors collected all functional genes of CYP450, including 57 CYP450 oxidase genes and 1 CYP450 reductase gene. (See Table 1 above), using the BLAST (http://blast.ncbi.nlm.mh.
  • gov Blast.cgi comparison software, with the human genome-wide standard sequence hgl9 as the reference sequence, all genes of 58 CYP450 genes
  • the type sequence is aligned with the hgl9 reference sequence, and the mutation site information relative to hgl9 is obtained according to the alignment result, and all genotypes of the CYP450 gene are converted into a uniform format and standard.
  • the genotype is converted to a hgl9-based type based on the annotation information of the gene on the whole genome. Specifically, the following steps are included:
  • the information mainly includes the name of the genotype, the protein number corresponding to the genotype, the mutation information of the genotype and the specific sequence, the activity of the enzyme corresponding to the genotype in the living body, and the enzyme corresponding to the genotype in an in vitro test. Activity.
  • the region is set to be longer, in order to include the above mutation sites as a principle.
  • the alignment result of the best position is selected, and the bases at each position are analyzed to obtain a specific sequence with hgl9 at each Base correspondence at one position. It should be noted that if it is aligned to the negative strand on the stain, it needs to be converted to tt on the positive strand.
  • CYP450 genotypes were converted to hgl9-based mutation site information based on the alignment of specific sequences with hgl9.
  • the CDS starting position and the defined gene region of the above gene are required.
  • the mutation site information on some genotypes is negatively chained, and the negative strand information needs to be converted into a positive strand during the conversion.
  • the experimental flow portion of this example is described as a chip hybridization of 50 samples including Yanhuang.
  • the number of samples in this example is used to explain the present invention, rather than limiting the number of samples that each chip can hybridize. 1.
  • reagents in this example are shown in Table 3. Other reagents, consumables, and equipment are not indicated in Table 3, and are all general-purpose products that can be purchased through the market.
  • the disrupted fragment was tested by electrophoresis (the main band was concentrated between 200 bp and 300 bp), it was recovered by QIAquick PCR Purification Kit, and the sample was dissolved in 75 L of elution buffer.
  • the DNA fragment obtained after the interruption was recovered and purified, and the end-repair reaction system was prepared in a 1.5 mL centrifuge tube to form a flattened terminal phosphorylated DNA fragment.
  • reaction mixture was lightly mixed for 4 Torr, and then purified by a QIAquick PCR purification kit in a Thermomixer (Eppendorf) at 20 ° C for 30 min, and finally the DNA was sufficiently dissolved in 32 ⁇ L of ddH 2 0 .
  • reaction mixture was lightly shaken and mixed uniformly, placed in a Thermomixer (Eppendorf) at 20 °C for 15 min after transient centrifugation, and then purified by MiniElute PCR Purification Kit after final reaction. Finally, the sample was dissolved in 25 ⁇ . Flush.
  • the DNA obtained in the above step (4) is used as a template, and the primers containing the linker sequence are amplified, and the amplification system and conditions are as follows:
  • the PCR program was 94 °C for 2 min; 4 cycles of 94 °C for 15 s, 62 V for 30 s, 72 °C for 30 s; and 72 °C for 5 min.
  • the PCR product was purified using a QIAquick PCR purification kit with an elution volume of 30 ⁇ .
  • the construction of the exon library comprises hybridization of the prepared sequence capture library with the capture chip, enriching all exons of 58 CYP450 genes onto the capture chip, eluting the hybridized capture chip, and eluting the product as an exon Sequence, exon sequence amplification treatment to obtain an exon library, as follows:
  • the sample was loaded with 35 ⁇ l and hybridized at 42 °C for 64-72 hr. After hybridization and post-hybridization of the chip, the sequence enriched on the chip was eluted with 900 ⁇ l of 160 mM NaOH. The eluted product was purified by MinElute PCR purification kit. Finally eluted with 80 ⁇ l elution buffer.
  • PCR amplification was performed using the sequence eluted from the capture chip as a template, the system was Phusion Mix 150 ⁇ 1, the upstream and downstream primers were each 4.2 ⁇ l (Multixing sequencing primer and Phix Control kit), and the above 80 ⁇ l elution sample was added with 85 ⁇ l ⁇ 2 0, after mixing, 6 tubes were used for PCR.
  • PCR reaction conditions 94 ° C, lmin; 16 cycles of 94 ° C 30s, 58 ° C 30s, 72 ° C 30s; 72 "C 5min.
  • 6 tubes were mixed and purified by QIAquick PCR purification kit magnetic beads Fragments of 300-450 bp size with an elution volume of 50 ⁇ l.
  • the data obtained by sequencing is filtered in two aspects. First, the quality value of the sequencing is performed, and the base quality value is calculated for the entire sequence. When the average mass value of the entire sequence is less than 10, the filter is filtered off; The Adapter connector is contaminated. If the sequence contains an Adapter sequence, it is also filtered out.
  • the data filtered sequences were compared using BWA (Burrows-Wheeler Aligner) comparison software.
  • BWA Backrows-Wheeler Aligner
  • each sequence can allow up to 5 mismatches, and open gap (allowing insertion and deletion when comparing).
  • open gap allowing insertion and deletion when comparing.
  • a sequence has multiple optimal alignment positions, randomly select a position output, but There are tags. In the test of this example, the sequence on the sample alignment accounted for approximately 97% of all aligned sequences.
  • the sequences of the target areas on the reference sequence are reserved for the next analysis.
  • the data quality control includes the data volume of the sample, the amount of data filtered, the ratio of the sequence alignment to the upper sequence, whether the average depth of the sample is in line with expectations, whether the single base depth coverage map conforms to the Poisson distribution, and the target area of the sample. Coverage, etc.
  • data quality control includes two aspects. On the one hand, it is to see whether the samples are relatively consistent. If the data between the samples is similar, the requirements are met. If there are individual samples, the other samples are quite different. Explain that this sample is likely to have problems; on the other hand, each quality control data of each sample, those skilled in the art can determine a rough range based on experience, and different sequencing areas may have some changes, specifically , "Remaining amount after data filtering" is generally above 90%, the ratio of the aligned sequence (%) is more than 90%, the amount of remaining data after deduplication is more than 60%, and the proportion of unique reads is related to the specific sequencing target area and 90 Above %, the average depth meets the expected experimental design requirements, and the coverage is over 95%, which is acceptable.
  • the SNP is obtained by using samtools. After selecting the sequence to the target area, using samtools to convert the format and sorting, use the mpileup command to perform SNP Callings.
  • the original SNP also performs some filtering, including bits. Point depth, quality value, etc. Usually, the depth is in accordance with the requirements of 4-400, and the quality value is calculated by statistically calculating the significance of the quality value.
  • the mutation site information of each gene is extracted based on the region of each gene on the whole genome. Based on these mutation site information, the genotype information and enzyme activity information of the sample were determined by comparison with the CYP450 gene standard type database constructed in Example 1. The test results of some samples are shown in Table 7.
  • the method for constructing the CYP450 genotype standardized database, CYP450 genotyping and enzyme activity identification can effectively be used for CYP450 genotyping and enzyme activity identification, and saves time and labor, low cost and accurate results.

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Chemical & Material Sciences (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • General Health & Medical Sciences (AREA)
  • Biotechnology (AREA)
  • Theoretical Computer Science (AREA)
  • Biophysics (AREA)
  • Analytical Chemistry (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Evolutionary Biology (AREA)
  • Medical Informatics (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Genetics & Genomics (AREA)
  • Organic Chemistry (AREA)
  • Molecular Biology (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • Bioethics (AREA)
  • Pathology (AREA)
  • Immunology (AREA)
  • Databases & Information Systems (AREA)
  • Microbiology (AREA)
  • Biochemistry (AREA)
  • General Engineering & Computer Science (AREA)
  • Oncology (AREA)
  • Hospice & Palliative Care (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Description

CYP450基因型别数据库及基因分型、 酶活性鉴定方法 优先权信息
本申请请求 2012 年 01 月 06 日向中国国家知识产权局提交的、 专利申请号为 201210002976.3的专利申请的优先权和权益, 并且通过参照将其全文并入此处。 技术领域
本发明涉及基因检测领域, 特别涉及 CYP450 的标准基因型别数据库及其构建方 法, 以及 CYP450基因分型和酶活性鉴定方法。 背景技术
细胞色素 P450筒称 CYP450, 目前已经从人体中鉴定出了 57种 CYP450氧化酶, 其中参与药物代谢主要集中于 CYP1、 CYP2及 CYP3家族, 参与代谢目前 90%以上的 药物。 CYP450在药物代谢中的重要作用导致 CYP450基因多态性成为影响药物个体差 异的最重要因素之一, 其多态性包括点突变、 插入或缺失、 整个基因的缺失或复制, 最 终导致酶的活性增强、 减弱或完全缺失。 CYP450多态性可导致在标准的药物剂量下可 能产生个体严重的副反应或不起作用。 比如, 华法林(warfarin )作为目前的一线口服 抗凝药, 其预防和治疗血栓栓塞的作用效果和口服给药的便利性和经济性是显而易见 的, 但是华法林治疗窗口很窄, 且给药各个体差异和种族差异很大, 要达到同样的作用 效果, 高低剂量可相差 10倍以上。 研究表明, 这与个体间的 CYP2C9的基因多态性密 切相关。 再如, 广泛用于治疗乳腺癌和卵巢癌的他莫昔芬 (tamoxifen ) 需要经过一系 列的 CYP450酶类代谢, 才最终形成有活性的产物发挥治疗效果,其中 CYP2D6的基因 型是影响他莫昔芬疗效的主要限制因素, 其缺陷型可导致无病生存期变短, 而 CYP2D6 强代谢型只需要较小剂量的他莫昔芬即可达到同样的效果。
越来越多的研究显示, CYP450基因除了影响药物代谢外, 其基因多态性与许多疾 病的发生紧密相关。在先天性肾上腺皮质增生症中, CYP21B基因缺陷占病因 90%-95% 左右, CYP11B1基因突变约占 5%, CYP17AK CYP1 1A1的缺陷型突变也可导致此疾 病的发生。
CYP450在代谢以及疾病发生中的重要作用, 使得 CYP450基因多态性的检测具有 非常重要的社会意义和研究价值。
但是, 目前对于 CYP450突变检测的技术主要集中于单个或几个 P450基因的已知 突变; 并且, 各研究中的突变位点都是相对各自特定的序列而言的, 彼此之间缺乏统一 衡量比较的标准。 另外, 现有的 CYP450检测方法对 CYP450基因中仍然存在的未知突 变位点检测效率不高; 需要对大量样本进行检测时, 耗时长, 不能很好的满足研究和实 践应用的需要。
因而, 目前的 CYP450突变检测及基因分型的方法仍有待改进。 发明内容
本发明旨在至少解决现有技术中存在的技术问题之一。 为此,本发明的一个目的是 提供一种 CYP450基因型别的标准数据库及其构建方法,以及基于该数据库的快速检测 CYP450基因型的方法和 CYP450酶活性鉴定方法。
因而,根据本发明的一个方面, 本发明提供了一种构建 CYP450基因标准型别数据 库的方法。 根据本发明的实施例, 该方法包括以下步驟: 将 CYP450基因型别的突变信 息对应的特定序列与人类全基因组标准序列进行比对,获得 CYP450特定序列与人类全 基因组标准序列在每个碱基位置上的对应关系; 根据所获得的对应关系, 将 CYP450基 因型转换成以人类全基因组标准序列为参考序列的基因型,获得 CYP450基因的标准化 基因型别。发明人惊奇地发现,利用本发明的构建 C YP450基因标准型别数据库的方法, 能够有效地构建 CYP450基因标准型别数据库,并且该数据库能够为各 CYP450基因分 型提供一个统一标准, 从而, 基于该 CYP450基因标准型别数据库, 对于已知基因型的 CYP450基因, 能够快速准确地确定待测样本的 CYP450基因的基因型别, 进一步, 能 够为针对待测样本的涉及 CYP450基因的疾病诊断或药物治疗等提供更精确的判断依 据。
根据本发明的另一方面, 本发明还提供了一种 CYP450基因标准型别数据库, 其是 采用前面所述的本发明的构建 CYP450基因标准型别数据库的方法构建的。 发明人发 现, 该 CYP450基因标准型别数据库能够为各 CYP450基因分型提供一个统一标准, 从 而, 基于该 CYP450基因标准型别数据库, 对于已知基因型的 CYP450基因, 能够快速 准确地确定待测样本的 CYP450基因的基因型别, 进一步, 能够为针对待测样本的涉及 C YP450基因的疾病诊断或药物治疗等提供更精确的判断依据。
根据本发明的另一方面, 本发明还提供了一种 CYP450基因分型的方法。根据本发 明的实施例, 该方法包括: 获取待测样本 CYP450基因的外显子序列, 釆用高通量测序 平台测序并进行数据分析,将分析结果与前面所述的本发明的 CYP450基因型别数据库 进行比较, 从而得到待测样本的基因型别。 利用该方法, 对于已知基因型的 CYP450基 因, 能够快速准确地确定待测样本的 CYP450基因的基因型别, 进一步, 能够为针对待 测样本的涉及 CYP450基因的疾病诊断或药物治疗等提供更精确的判断依据。
根据本发明的再一方面, 本发明还提供了一种 CYP450酶活性鉴定方法。根据本发 明的实施例, 该方法包括: 获取待测样本 CYP450基因的外显子序列, 釆用高通量测序 平台测序并进行数据分析,将分析结果与前面所述的本发明提供的含有 CYP450酶活性 信息的 CYP450基因标准型别数据库进行比较,得到待测样本的基因型别, 并根据基因 型别对应的酶活信息获得待测样本的 CYP450酶活性结果。 利用该方法, 对于已知基因 型的 CYP450 基因, 能够快速准确地确定待测样本的 CYP450基因的基因型别以及 CYP450酶活性结果, 进一步, 能够为涉及 CYP450的疾病或药物等提供更精确的判断 依据。 本发明的附加方面和优点将在下面的描述中部分给出,部分将从下面的描述中变得 明显, 或通过本发明的实践了解到。 附图说明
本发明的上述和 /或附加的方面和优点从结合下面附图对实施例的描述中将变得明 显和容易理解, 其中:
图 1为根据本发明一个实施例, 本发明的 CYP450基因分型及酶活性鉴定方法中, 获取待测样本 CYP450基因的外显子序列的步骤的具体流程示意图;
图 2为根据本发明一个实施例, 本发明的 CYP450基因分型及酶活性鉴定方法中, 数据分析的步骤的具体流程示意图。 发明详细描述
下面详细描述本发明的实施例。下面通过参考附图描述的实施例是示例性的,仅用 于解释本发明, 而不能理解为对本发明的限制。在本发明的描述中,除非另有说明, "多 个" 的含义是两个或两个以上。
根据本发明的一个方面,本发明公开提供了一种构建 CYP450基因标准型别数据库 的方法。 根据本发明的实施例, 该方法包括以下步驟: 将 CYP450基因型别的突变信息 对应的特定序列与人类全基因组标准序列进行比对,获得 CYP450特定序列与人类全基 因组标准序列在每个碱基位置上的对应关系; 根据所获得的对应关系, 将 CYP450基因 型转换成以人类全基因组标准序列为参考序列的基因型,获得 CYP450基因的标准化基 因型别。 发明人惊奇地发现, 利用本发明的构建 CYP450基因标准型别数据库的方法, 能够有效地构建 CYP450基因标准型别数据库,并且该数据库能够为各 CYP450基因分 型提供一个统一标准, 从而, 基于该 CYP450基因标准型别数据库, 对于已知基因型的 CYP450基因, 能够快速准确地确定待测样本的 CYP450基因的基因型别, 进一步, 能 够为针对待测样本的涉及 CYP450基因的疾病诊断或药物治疗等提供更精确的判断依 据。
根据本发明的一个实施例, CYP450基因包括选自 CYP11A1、 CYP11B1、 CYP1 1B2、 CYP17AK CYP1A1、 CYP1A2、 CYP1B1、 CYP20A1、 CYP21A2、 CYP24A CYP26AK CYP26B1、 CYP26C CYP27A1、 CYP27B1 , CYP27C1 , CYP2A13 , CYP2A6、 CYP2A7、 CYP2B6、 CYP2C18、 CYP2C19. CYP2C8、 CYP2C9. CYP2D6. CYP2E CYP2FK CYP2J2、 CYP2R1、 CYP2S1、 CYP2U1、 CYP2W CYP39A1 , CYP3A4、 CYP3A43 , CYP3A5、 CYP3A7、 CYP46A CYP4A1 CYP4A22、 CYP4B1、 CYP4F11、 CYP4F12、 CYP4F2、 CYP4F22、 CYP4F3、 CYP4F8、 CYP4V2、 CYP4X1、 CYP4Z1、 CYP51A CYP5A1、 CYP7A1、 CYP7B1、 CYP8A1、 CYP8BK POR等 58个人类 CYP450基因的 至少一种。
根据本发明的另一个实施例,上述方法进一步包括将 CYP450基因定位于人类全基 因组标准序列上, 确定 CYP450基因编码序列的起始位置和终止位置, 获得 CYP450基 因型别突变信息对应的特定序列。
根据本发明的一个实施例, 所述 CYP450基因型别突变信息对应的特定序列包含 CYP450基因编码序列的起始位置上游 5000bp至编码序列终止位置下游 500bp区域的 DNA片段。
根据本发明的一个优选实施例, 所述人类全基因组标准序列为 hgl9。
根据本发明的另一方面, 本发明还提供了一种 CYP450基因标准型别数据库, 其是 采用前面所述的本发明的构建 CYP450基因标准型别数据库的方法构建的。 发明人发 现, 该 CYP450基因标准型别数据库能够为各 CYP450基因分型提供一个统一标准, 从 而, 基于该 CYP450基因标准型别数据库, 对于已知基因型的 CYP450基因, 能够快速 准确地确定待测样本的 CYP450基因的基因型别, 进一步, 能够为针对待测样本的涉及
C YP450基因的疾病诊断或药物治疗等提供更精确的判断依据。
根据本发明的一个优选实施例, 在本发明的 CYP450 基因标准型别数据库中,
CYP450基因的各标准化基因型别对应有酶活性信息。
根据本发明的另一个实施例, 该 CYP450 基因包括下表 1 中所述的 58 个人类
CYP450基因的至少一种。
表 1 人类 CYP450基因
Figure imgf000005_0001
根据本发明的另一方面, 本发明还提供了一种 CYP450基因分型的方法。根据本发 明的实施例, 该方法包括: 获取待测样本 CYP450基因的外显子序列, 采用高通量测序 平台测序并进行数据分析,将分析结果与前面所述的本发明的 CYP450基因型别数据库 进行比较, 从而得到待测样本的基因型别。 利用该方法, 对于已知基因型的 CYP450基 因, 能够快速准确地确定待测样本的 CYP450基因的基因型别, 进一步, 能够为针对待 测样本的涉及 CYP450基因的疾病诊断或药物治疗等提供更精确的判断依据。
根据本发明的再一方面, 本发明还提供了一种 CYP450酶活性鉴定方法。根据本发 明的实施例, 该方法包括: 获取待测样本 CYP450基因的外显子序列, 采用高通量测序 平台测序并进行数据分析,将分析结果与前面所述的本发明提供的含有 CYP450酶活性 信息的 CYP450基因标准型别数据库进行比较,得到待测样本的基因型别, 并根据基因 型别对应的酶活信息获得待测样本的 CYP450酶活性结果。 利用该方法, 对于已知基因 型的 CYP450 基因, 能够快速准确地确定待测样本的 CYP450基因的基因型别以及 CYP450酶活性结果, 进一步, 能够为涉及 CYP450的疾病或药物等提供更精确的判断 依据。
根据本发明的一个实施例, 获取待测样本 CYP450基因的外显子序列是通过以下步 骤实现的:
A、 制备能够捕获 CYP450基因外显子序列的芯片, 所述芯片上含有与 CYP450基 因外显子序列反向互补的寡核苷酸探针;
B、 用待测样本的基因组 DNA制备序列捕获文库, 包括将待测样本基因组 DNA打 断为 200~500bp大小的片段, 进行末端处理后扩增得到序列捕获文库;
C、将步骤 B制备得到的序列捕获文库与步骤 A的芯片杂交,从而获取得到待测样 本的 CYP450基因外显子文库。
根据本发明的一个实施例,在上述步驟 A中,芯片含有能分别与 58个人类 CYP450 基因的所有外显子序列反向互补的寡核苷酸探针, 寡核苷酸探针的长度为 55-105bp。
根据本发明的另一个实施例, 在上述步骤 B 中, 将待测样本基因组 DNA打断为 200~300bp大小的片段。
根据本发明的再一个实施例,在上述步骤 B中,末端处理包括进行末端修复形成平 末端磷酸化的 DNA片段, 并在该平末端磷酸化的 DNA片段的 3,末端加上碱基 "A" , 并进一步连接标签。
进一步, 根据本发明的一个实施例, 在上述步骤 C中, 在进行杂交之前, 将来自多 个不同待测样本的序列捕获文库混合后再同时与步骤 A的芯片杂交, 每个文库带有不 同的标签( Index )碱基序列而相互区别, 且该标签碱基序列长度优选为 6~8bp。
根据本发明的一个实施例, 在本发明的 CYP450基因分型和酶活性鉴定方法中, 数 据分析进一步包括:
1、 过滤去掉影响信息分析的低质量测序序列;
ii、 以人类全基因组标准序列为参考序列, 将步骤 i得到的序列用比对软件进行比 对, 比对软件优选用 SOAP或 BWA;
iii、 选取比对到目标区域的序列进行后续分析, 所述目标区域是指 CYP450基因外 显子序列所在区域; iv、 数据质控合格后进行变异分析, 所述变异分析包括检测以下中的至少一种: 单 核苷酸多态性、 插入和删除、 结构性变异、 拷贝数变异。
需要说明的是, 由于采用了以上技术方案, 使本发明至少具备以下有益效果: 1、 本发明的构建 CYP450基因标准型别数据库的方法, 能够为各 CYP450基因分 型提供一个统一标准的数据库, 基于该数据库, 对于已知基因型的 CYP450基因, 能够 快速准确地给出待测样本的 CYP450基因基因型别,从而能够对涉及 CYP450的疾病或 药物等提供更精确的判断依据。 进一步, 通过在 CYP450基因标准型别数据库中增加与 基因型别对应的酶活性信息, 能够在对待测样本进行基因分型的同时, 直接得到其 CYP450酶活结果。
2、 本发明的 CYP450基因分型方法, 先获取所有 CYP450基因的外显子序列后, 再进行测序分析, 与 CYP450基因标准型别数据库进行比对后, 能够有效地对 CYP450 基因中未知突变位点进行检测。
3、 本发明的 CYP450基因分型及酶活鉴定方法, 利用芯片捕获具有高通量的性质, 可以实现一次实验同时检测多达上百个样本, 不仅提高了检测样本的数量, 同时也大大 降低了每个样本的检测费用。
4、 本发明的 CYP450基因分型及酶活鉴定方法, 包含了目前人类中已经鉴定出的 57个 CYP450氧化酶基因及 1个 P450还原酶基因, 覆盖范围全面广泛, 极大地方便了 专门针对 CYP450基因的研究。
进一步, 本发明还提供了构建 CYP450基因型别标准化数据库、 CYP450基因分型及酶 活性鉴定的整体技术方案, 其是以目标区域经序列捕获后的高通量测序为基础进行的, 具 体地, 可以包括以下步骤:
一、 CYP450基因型别标准化数据库的构建
收集 CYP450全部功能基因 , 其中包含 57个 CYP450氧化酶基因和 1个 CYP450还原 酶基因 (见上述表 1 ), 通过 BLAST ( http://blast.ncbi.nlm.nih. gov/Blast.cgi ) 比对软件, 以 人类全基因组标准序列 hg 19为参考序列,将 58个 C YP450基因的所有基因型别序列与 hg 19 参考序列比对, 根据比对结果得到相对于 hgl9的突变位点信息, 将 CYP450基因的所有基 因型别转换成统一的格式和标准。 根据基因在全基因组上的注释信息, 将基因型别转换为 以 hgl9为标准的型别。 具体包括以下步骤:
1. 收集 CYP450基因型别相关突变及酶活性信息
收集现有的 58个 CYP450基因的所有基因型别的突变信息和型别与酶活性相关信息。 这些信息主要包括基因型别的名称、 基因型别对应的蛋白盾编号、 基因型别与特定序列的 突变信息、 基因型别对应的酶在活体中的活性、 基因型别对应的酶在体外试验中的活性等。 需要说明的是, 本申请中所述 "特定序列", 是指研究中所采用的作为参考的 DNA序列片 段或者一段 cDNA序列。 对收集的资料分析发现, 每个型别的突变信息都是相对于其中一 个特定序列给出的; 也就是说, 不同的研究资料中, 58个 CYP450基因其基因型别的参考 对象不同, 而针对不同的参考对象, 同一个基因的不同基因型别也存在差异。 对于不同资 料上格式的不一致, 需要改成统一的格式, 以便后续的整理。
2. 收集基因在特定序列上的 CDS区域, 及基因在 hgl9上的位置
在收集的资料中, 很多基因型别突变信息是相对于给定的特定序列的, 并且, 突变位 点信息是以 1998公布的基因突变命名规则 ( Recommendations for a nomenclature system for human gene mutations. Nomenclature Working Group )为才示准的, 以基因的 CDS (编码序歹1 J ) 起始位置为十 1 的标准来给出突变位置的。 所以为了后续的分析, 需要找出所有基因在特定 序列上的 CDS起始位置。 叉因为特定序列非常的长, 有些序列上包括了多个基因, 所以要 确定哪一段区域是发明人需要的 CYP450基因。 发明人是先找出 CYP450基因在 hgl9上的 位置, 然后从 CDS起始位置上游的 5000bp到 CDS终止位置下游的 500bp作为 CYP450基 因的区域, 但有些基因型别的突变位点离 CDS区比较远, 超出了上述的范围, 对于这些基 因, 发明人会把这个基因的区域定得更长一些, 以嚢括上述突变位点为原则。
3. BLAST比对
将特定序列与 hgl9进行 BLAST比对。 如果特定序列是 cDNA, 发明人用 BLAT进行 比对。
4. 确定特定序列与 hgl9的突变信息
在比对结果中, 特定序列可能会比对上 hgl9的多个位置, 选择比对最好的一个位置的 比对结果, 对每一个位置上的碱基进行分析, 得到特定序列与 hgl9在每一个位置上的碱基 对应关系。 需要注意的是, 如果比对到染色上的负链上, 需要将 转换成正链上的 tt。
5. 转换所有 CYP450基因型别
根据特定序列与 hgl9的比对情况,将所有 CYP450基因型别转换为以 hgl9为标准的突 变位点信息。 在进行坐标转换时, 需要用到上面基因的 CDS起始位置和定义的基因区域。 有些基因型别上的突变位点信息都是负链的, 在转换时需要将负链信息转换为正链。
6. 整理文件格式及检查
整理文件格式, 将基因型别酶活性的信息也加入进来, 具体例子如表格 2所列。 之后 再检查结果的正确性。
CYP450基因的标准化基因分型数据库信息 (部分)
酶在体外
基因 染色 SNP突变 INDEL突变 其他突 酶在体
基因 实验中活 参考文献 型别 体 位点 位点 变信息 内活性
CYP2
CYP2C 96748777 Blaisdell et al,
C9*l chrlO ― - ― 降低
9 :C>T 2004
2
CYP2 Si et al, 2004; Guo
CYP2C 96701715
C9*l chrlO ― 降低 降低 et al, 2005a; Guo 9 : T>C
3 et al, 2005b CYP2 Zhao et al, 2004;
CYP2C 96701991
C9*l chrlO ― - ― 降低 Delozier et al, 9 : G>A
4 2005
CYP2 Zhao et al, 2004;
CYP2C 96707539
C9*l chrlO ― - ― 无活性 Delozier et al, 9 : C>A
5 2005
96701970-9
CYP2
CYP2C 6701979:Del Maekawa et al,
C9*2 chrlO ― - ― 无活性
9 AGAAATG 2006 5
GAA 二、 CYP450基因型别检测
首先, 参照图 1所示的获取待测样本 CYP450基因的外显子序列的步骤的具体流程 示意图, 确定待测样本 CYP450基因的外显子序列, 具体地:
1、 P450外显子探针设计
发明人才艮据 57个 CYP450氧化酶基因及一个 CYP450还原酶基因, 以人类基因组 hgl9 为参考序列,选取这 58个基因的全部外显子区域作为靶序列,靶序列长度之总和约 276 kb。 针对每一个外显子序列, 设计与外显子序列反向互补的长度约为 55-105bp的寡核苷酸捕获 探针。 将设计的捕获探针高密度的固定合成在芯片上, 形成包含 58个 CYP450基因所有外 显子捕获探针的捕获芯片。 设计好的探针由 Roche-Nimblegen生产并合成固定在捕获芯片 上。
上述探针序列是参照 hg 19设计的, 由于不同物种间基因组序列存在一定的差异, 因此 该探针优先适用于人源基因组 DNA捕获,其它跟人类基因组同源性较高的物种的基因组可 以适用, 但捕获效果可能不如人源基因组理想。 不同物种可以根据其参考序列设计跟本发 明类似的探针, 应用于不同物种靶区域的捕获。
2、 基因组打断、 纯化
以没有 R A、 蛋白质污染且没有降解的人基因组 DNA作为实验材料, 利用物理或化 学的方法将 DNA打断成 200~300bp大小的片段, 使用相关回收试剂盒回收 DNA片段。
3、 末端修复、 纯化
回收纯化的片段化 DNA通过 T4 DNA聚合酶、 Klenow片段和 T4多聚核苷酸激酶等酶 的作用以 dNTP为作用底物进行末端修复,形成补平的末端磷酸化的 DNA片段,然后纯化。
4、 3'加 "A"、 纯化
利用 Klenow片段(3,-5,exo-)聚合酶及 dATP,将经过纯化的末端磷酸化的 DNA片段的 3,末端加上碱基 "A" , 然后纯化。
5、 接头连接、 纯化 利用 T4 DNA连接酶, 将经过纯化的末端加 "A" 的 DNA片段与标签接头连接, 并用 试剂盒纯化接头连接产物。
6、 连接产物 PCR、 定量
以标签接头序列引物对加接头后的 DNA文库进行扩增, 扩增产物经纯化后经 Agilent 2100和 NanoDrop定量、 质控合格后备用。
7、 PCR产物、 接头封闭序列、 Cotl DNA混合
将上述获得的 PCR产物、 接头封闭序列、 Cotl DNA混合以构成外显子捕获文库。 然后, 将建好的多个待测样本的文库混合, 为了在测序中区别来自不同样本的文库, 每个文库的 DNA在连接标签接头时, 其接头中都含有不同的 6bp或 8bp的标签碱基序列, 每个文库 DNA混合量可根据需要等量或按照一定比例混合。 需要说明的是, 等量即在需要 每个样本测序数据量相同时, 每个文库混合 DNA量一致; 有的研究不同样本测序数据量可 能不同, 文库使用量也就不同 , 混合比例按照本领域技术人员具体的研究目的或设计要求 来确定。
8、 探针与目标区域杂交
按照 mblegen固相芯片杂交标准操作说明,将上述多个样本的混合文库与芯片进行杂 交。
9、 探针洗脱、 纯化
将杂交后的 DNA进行洗脱、 纯化, 然后以接头序列为引物进行扩增, 以便获得扩增产 物。
10、 质控
利用 Agilent 2100和 Q-PCR将上述获得的扩增产物进行质控, 经质控合格后备用。 其 中各扩增产物构成测序文库。
11、 上机测序
使用 Hiseq2000平台, 釆用边合成边测序的测序方法对上述获得的测序文库进行测序。 然后, 以人类基因组 hgl9 ( UCSC )为参考序列, 对获得的测序数据进行数据分析。 参 照图 2 , 该数据分析步驟可以包括以下步驟:
步骤一 过滤: 去掉质量值较低和有测序接头污染的序列
首先去掉影响信息分析的低质量测序序列: 序列中每个碱基分别对应一个测序质量值, 对于测序结果的一段序列, 计算这段序列的平均质量值, 若这条序列的平均质量值低于常 规的经验阈值, 这条序列会被过滤掉; 另一方面, 测序序列可能会被机器上的接头污染, 这部分含有的序列也会被过滤掉。
步驟二 与参考序列 (hgl9 )进行序列比对
以 hgl9( UCSC )为参考序列,将经过步骤一过滤后的序列用比对软件(如 SOAP , BWA ) 进行序列比对。 这些比对软件对于一段序列, 能够选择一个最佳的比对位置。 对于比对位 置有多个的重复序列, 软件会选择一个位置输出, 并添加一个标签。
步驟三 选取比对到目标区域的序列 芯片杂交后会捕获到部分非目标区域的序列, 步骤二中以 hgl9全基因组序列作为参考 序列, 非目标区域的序列就会根据最佳匹配原则比对到相应的位置, 而不会比对的目标区 域。 选取比对到目标区域的序列用于后续分析, 保证选取的序列都是目标区域序列。
步驟四 数据控制
数据控制 (质控) 包括多个方面, 如比对上序列的百分比, unique reads (序列与参考 序列比对时只有一个最佳比对位置) 的百分比, duplication (相同的序列)的比例, 测序深 度, 目标区域的覆盖度等。 这些质控要符合常规的经验阈值才能进行下一步的分析, 如测 序深度与预期一致, 单碱基深度覆盖图服从泊松分布。
步驟五 变异检测
数据质控合格后, 才能进行变异分析, 包括检测 SNP (单核苷酸多态性), INDEL (插 入和删除), SV (结构性变异), CNV (拷贝数变异)等。 每种变异检测可根据需要使用不 同的方式来实现。
步驟六 CYP450基因分型
当变异检测分析完之后, 整理每个基因中的突变位点信息, 与之前整理好的 CYP450 标准数据库中的相应基因型别比较, 得到每个样本的基因型别。 由于人是二倍体生物, 每 个基因的型别最多只有两种型别,最后 CYP450基因的分型结果是一种纯合型别或者杂合型 别。 一些基因型别有相应酶活性信息, 所以通过样本基因分型之后, 同时也可以得到样本 对酶活性的反应情况。
此外,还需要说明的是,现阶段,对于 CYP450突变检测的技术主要集中于单个或几个 CYP450基因的已知突变,对于未知突变或者大量样本检测存在耗时长、费用高等限制因素。 而本发明的上述整体技术方案, 明显具有以下几个优点:
一、本发明构建的 CYP450基因标准型别数据库, 包含了目前人类中已经鉴定出的所有 有功能的 57个 CYP450氧化酶基因及 1个 CYP450还原酶基因, 包含范围全面广泛, 利用 该数据库可以检测出待测样本这些 CYP450基因所有已知和未知的多态性位点,极大的方便 了专门针对 CYP450基因的研究。
二、 能够建立一个以 hgl9为参考序列的 CYP450基因标准型别数据库, 对于已知基因 型的 CYPP450基因, 能够快速准确的给出相应基因型别信息, 这对于涉及 CYP450的疾病 或药物等提供更精确的判断依据。
三、 利用芯片捕获具有高通量的性质, 一次实验同时检测多达上百个样本, 不仅提高 了检测样本的数量, 同时也大大降低了每个样本的检测费用。 下面将结合实施例对本发明的方案进行解释。 本领域技术人员将会理解, 下面的实施 例仅用于说明本发明, 而不应视为限定本发明的范围。 实施例中未注明具体技术或条件的, 按照本领域内的文献所描述的技术或条件(例如参考 J.萨姆布鲁克等著, 黄培堂等译的《分 子克隆实验指南》, 第三版, 科学出版社)或者按照产品说明书进行。 所用试剂或仪器未注 明生产厂商者, 均为可以通过市购获得的常规产品, 例如可以釆购自 Illumina公司。 实施例 1
本实施例根据本发明的构建 CYP450基因型别标准化数据库的方法构建基因型别标准 化数据库, 具体地: 发明人收集了 CYP450全部功能基因, 其中包含 57个 CYP450氧化酶 基因和 1 个 CYP450 还原酶基因 (见上述表 1 ), 通过 BLAST ( http://blast.ncbi.nlm.mh. gov Blast.cgi ) 比对软件, 以人类全基因组标准序列 hgl9为参考序列, 将 58个 CYP450基 因的所有基因型别序列与 hgl9参考序列比对, 才艮据比对结果得到相对于 hgl9的突变位点 信息,将 CYP450基因的所有基因型别转换成统一的格式和标准。根据基因在全基因组上的 注释信息, 将基因型别转换为以 hgl9为标准的型别。 具体包括以下步骤:
1. 收集 CYP450基因型别相关突变及酶活性信息
收集现有的 58个 CYP450基因的所有基因型别的突变信息和型别与酶活性相关信息。 这些信息主要包括基因型别的名称、 基因型别对应的蛋白质编号、 基因型别与特定序列的 突变信息、 基因型别对应的酶在活体中的活性、 基因型别对应的酶在体外试验中的活性。
2. 收集基因在特定序列上的 CDS区域, 及基因在 hgl9上的位置
找出收集的资料中所有基因在特定序列上的 CDS起始位置, 并确定哪一段区域是本发 明需要的 CYP450基因。 具体地, 先找出 CYP450基因在 hgl9上的位置, 然后从 CDS起始 位置上游的 5000bp到 CDS终止位置下游的 500bp作为 CYP450基因的区域,其中有些基因 型别的突变位点离 CDS区比较远, 超出了上述的范围, 对于这些基因, 将其区域定得更长 一些, 以嚢括上述突变位点为原则。
3. BLAST比对
将特定序列与 hgl9进行 BLAST比对。
4. 确定特定序列与 hgl9的突变信息
在比对结果中, 对于比对上 hgl9的多个位置的特定序列, 选择比对最好的一个位置的 比对结果, 对每一个位置上的碱基进行分析, 得到特定序列与 hgl9在每一个位置上的碱基 对应关系。 需要注意的是, 如果比对到染色上的负链上, 需要将 转换成正链上的 tt。
5. 转换所有 CYP450基因型别
根据特定序列与 hgl9的比对情况,将所有 CYP450基因型别转换为以 hgl9为标准的突 变位点信息。 在进行坐标转换时, 需要用到上面基因的 CDS起始位置和定义的基因区域。 有些基因型别上的突变位点信息都是负链的, 在转换时需要将负链信息转换为正链。
6. 整理文件格式及检查
整理文件格式, 将基因型别酶活性的信息也加入进来, 之后再检查结果的正确性。 由此, 获得类似表 2所示的包含 58个 CYP450基因性别及酶活信息的 CYP450基因标 准型别数据库。 实施例 1
本实施例实验流程部分描述为包括炎黄在内的 50个样本建库杂交一张芯片, 本实施例 中的样本数用以解释本发明, 而不是限制每张芯片可以杂交的样本数。 1、 实验材料
本实施例中的试剂见表 3, 其它试剂、耗材和仪器设备未在表 3中注明者, 均为可通过 市场购买的通用产品。
表 3 本实施例所用试剂
2、 序列捕获文库制备
( 1 )基因组 DNA片段化 以 3 g无蛋白质、 R A污染且没有降解的炎黄基因组 DNA为材料, 使用 Covaris-S2
Figure imgf000014_0001
打断后的片段经电泳检测合格(主带集中在 200bp-300bp之间)后,使用 QIAquick PCR 纯化试剂盒回收纯化, 样本溶于 75 L 洗脱緩冲液中。
( 2 ) DNA片段末端修复
将打断后回收纯化得到的 DNA 片段按下表在 1.5mL 的离心管中配制末端修复反应体 系, 形成补平的末端磷酸化的 DNA片段。
Figure imgf000014_0002
将上述 100 μL反应混合物轻 4敎混匀后,在 Thermomixer( Eppendorf )中 20°C温浴 30 min 后用 QIAquick PCR纯化试剂盒纯化, DNA最后于 32 μL ddH20中充分溶解。
( 3 ) 3,末端加 "A"碱基修饰
在末端补平修复后的 DNA片段 3,末端加上 "A" 碱基, 以便于下一步标签接头(Index Adapter ) 连接。 末端加 "A" 碱基反应体系如下表。
DNA 32μL
10x blue緩冲液 5μL
dATP(lmM) 10μΕ
Klenow (3 '-5' exo-) 3μL
总体积 50μL 将上述 50μL反应混合物轻微混匀后, 在 Thermomixer ( Eppendorf ) 中 37°C温浴 30min 后用 QIAquick PC 纯化试剂盒纯化, DNA最后于 15 μL ddH20中充分溶解。
( 4 ) Index Adapter接头连接
末端加 "A"后的 DNA片段纯化后在 T4 DNA Ligase作用下与标签接头连接。在 1.5 ml 的离心管中配制标签接头连接反应体系:
Figure imgf000015_0001
上述 5(^L反应混合物轻 t振荡混合均匀, 瞬时离心后置于 Thermomixer ( Eppendorf ) 中 20°C温浴 15min,反应完后用 MiniElute PCR纯化试剂盒进行纯化,最后将样品溶于 25μί 洗脱緩冲液。
( 5 ) 杂交前 PCR扩增
以上述步 (4 )得到的 DNA为模板, 以含有接头序列的引物进行扩增, 扩增体系和 条件如下:
Figure imgf000015_0002
PCR程序为 94 °C 2min; 4个循环的 94 °C 15s, 62V 30s, 72 °C 30s; 72 °C 5min。 PCR 产物用 QIAquick PCR纯化试剂盒純化, 洗脱体积为 30μί。
( 7 )样本文库混合
按照上述 DNA打断、 末端修复、 加标签接头 ( Index Adapter )、 杂交前 PCR等步骤, 构建其它 49个样本文库,包括炎黄基因组 DNA样本文库共计 50个文库(包含 4个 HapMap 样本、 1个炎黄样本和 45个正常人样本, 其中 45正常人样本用于测试一张芯片可以杂交的 样本数目), 从这 50个文库中取等量的 DNA均匀混合。 为了在测序中区别来自不同样本的 文库, 在加 Index Adapter时, 每个文库的 DNA末端都含有不同的 6bp或 8bp的 Index碱基 序列。 需要说明的是, Index Adapter包括两部分,分别为用于区分各文库的标签碱基序列和 接头序列。
4、 外显子文库构建
外显子文库的构建包括采用制备的序列捕获文库与捕获芯片杂交, 将 58个 CYP450基 因的全部外显子富集到捕获芯片上, 洗脱杂交后的捕获芯片, 洗脱产物即外显子序列, 对 外显子序列扩增处理得到外显子文库, 具体如下:
( 1 )芯片杂交
A )在 1.5mL离心管中加入 45(^g的 Cot-1 DNA、 3 g来自混合文库的 DNA、 lnmol Index-adpaterl -block和 Index-adpater2-block ( Multiplexing Sample Preparation Oligonucleotide Kit, Illumina ), 混合物置于 SpeedVac ( Thermo ) 中蒸干, 温度设置为 60°C。
B )在蒸干的离心管中加入 n.2μL纯水, 充分溶解 DNA后加入 18.5 L的 2 <SC 杂交 緩冲液和 7.3 L的 SC Hybridiation, 充分混匀后将混合物转移至杂交仪 ( Nimblegen )上 95 °C千浴器中 10分钟使 DNA变性。
C )将样品取出震荡后置于离心机上全速离心 30秒, 置于杂交仪 ( imblegen )上 42 °C位置, 与外显子捕获芯片杂交。
D )杂交方法参照 MmbleGen公司芯片杂交方法( MmbleGen Arrays User's Guide, Version
3.1 , 7 Jul 2009, Roche NimbleGen, Inc. , 通过参照将其全文并入本文)。 样品上样量 35μ1, 42 °C杂交 64-72hr, 杂交完成并经过芯片的杂交后处理后, 用 900μ1 160mM NaOH洗脱富集在 芯片上的序列, 洗脱产物用 MinElute PCR纯化试剂盒纯化, 最终用 80μ1洗脱緩冲液洗脱。
( 2 )捕获后 PCR扩增
以从捕获芯片上洗脱下来的序列为模板进行 PCR扩增, 体系为 Phusion Mix 150μ1, 上 下游引物各 4.2μ1 ( Multiplexing测序引物和 Phix Control试剂盒), 上述的 80μ1洗脱样品加 85μ1 άάΗ20 , 混合后分 6管进行 PCR。 PCR反应条件 94°C , lmin; 16个循环的 94°C 30s , 58 °C 30s, 72 °C 30s; 72 "C 5min。 PCR反应后把 6管混合并用 QIAquick PCR纯化试剂盒 磁珠纯化回收 300-450bp大小的片段, 洗脱体积为 50μ1。
( 3 )文库检测:
采用 Bioanalyzer analysis system (Agilent, Santa Clara, USA)检测文库插入片段大小及含 量; Q-PCR精确定量文库的浓度。
5、 序列测定
对上述经过纯化和质量检测合格的 PCR扩增产物进行测序, 测序方法参照 Illumina公 司 HiSeq2000操作方法( HiSeq 2000 User Guide. Catalog # SY-940-1001 Part # 15011190 Rev B , Illumina )„ 6、 数据分析
( 1 )测序数据过滤
对测序获得的数据进行两方面的过滤, 一是测序质量值, 对整条序列, 计算其碱基质 量值, 当整条序列的平均质量值低于 10时, 将其过滤掉; 二是检测 Adapter接头污染, 如 果序列中含有 Adapter序列, 也将其过滤掉。
测序数据过滤结果显示, 被过滤掉的序列约占 7%, 其余 93%用于下一步的分析。
( 2 )序列比对
以 hgl9为参考序列, 用 BWA ( Burrows-Wheeler Aligner ) 比对软件对经过数据过滤的 序列进行比对。 比对时每条序列最多允许 5个错配, 开 gap (比对时允许有插入和删除)的 比对, 当一条序列有多个最佳比对位置时, 随机选择一个位置输出, 但会有标记。 在本实 施例的测试中, 样本比对上的序列占所有进行比对的序列的约 97%。
( 3 )选取比对到目标区域的序列
比对完之后, 首先, 根据比对的结果, 去掉非 unique reads, 只保留那些唯一比对到全 基因组中的序列; 再去 duplication, 对于比对到参考序列上同一位置的配对 reads, 去重复 任意保留其中一对 reads, 因为比对到同一位置的配对序列很可能是 PCR过程引起的。
上面处理完后,根据 CYP450芯片设计的目标区域,保留那些比对到参考序列上的目标 区域的序列, 进行下一步的分析。
( 4 )数据盾控
数据质控包括样本的数据量, 过滤的数据量大小, 序列比对时比对上序列的比例, 样 本的平均深度是否符合预期, 单碱基深度覆盖图是否符合泊松分布, 样本的目标区域覆盖 度等。
统计分析结果显示, 本实施例的 50个样本均符合质控要求, 部分结果见表 4。
具体地, 数据质控包括两方面, 一方面是看各样本之间是不是比较一致, 如果各样本 之间的数据都差不多, 表示符合要求, 如果有个别样本的数据其他大多数样本相差很多, 说明这个样本很可能有问题; 另一方面是每个样本的各质控数据, 这些标准本领域技术人 员都可根据经验来确定一个大概的范围, 不同的测序区域可能会有些变化, 具体来说, "数 据过滤后剩余量"一般在 90%以上, 比对序列的比例 (%) 90%以上, 去重复后剩余数据量 60%以上, unique reads占的比例与具体的测序目标区域相关且 90%以上, 平均深度符合预 期的实验设计要求, 覆盖度要 95%以上, 都是可以接受的。
表 4 部分样本的数据质控结果
去重复 unique
数据过滤 比对序 目标区 样本名 原始数 后剩余 reads占 平均深 覆盖度 后剩余量 列的比 域数据 称 据量 数据量 的比例 度(% ) ( % )
( % ) 例(% ) 里
( % ) ( % )
样本 1 83.47M 93.18 96.85 68.47 93.12 19.26M 70.44 99.04 样本 2 81.73M 92.66 96.77 68.04 93.01 18.25M 66.61 98.95 样本 3 72.84M 92.91 96.77 73.52 93.60 17.80M 65.06 98.75 样本 4 71.07M 93.11 96.74 71.23 93.28 16.73M 61.16 99.05 样本 5 72.71M 92.84 96.71 73.67 93.24 17.07M 62.33 98.87 样本 6 80.67M 92.69 96.64 69.05 93.23 17.51M 63.92 98.61 样本 7 63.65M 93.10 96.71 74.63 93.66 15.63M 57.27 98.77 样本 8 77.48M 92.97 96.79 70.90 93.22 18.45M 67.46 98.89 样本 9 66.75M 93.49 96.79 74.93 93.98 16.71M 61.01 98.76 样本 10 73.58M 92.95 96.79 70.25 93.32 16.97M 62.17 98.86
( 5 ) SNP分析
本实施例中, SNP是用 samtools得到的, 当选取比对到目标区域的序列后, 用 samtools 转换格式、 排序之后 , 用其中的 mpileup命令进行 SNP Callings 原始的 SNP还会进行一些 过滤, 包括位点的深度、 质量值等。 通常, 深度在 4-400符合要求, 质量值则是通过用统计 的方法计算质量值的显著性, 对显著性过滤。
在本实施例的样本中, 包括 4个 HapMap样本 、 b、 c、 d )和 1个炎黄样本(这 5 个样本已经有公布的基因组及分型数据),其中炎黄样本测了两次,对这五个样本的 SNP进 行了评价。 4个 HapMap样本与已有的 HapMap数据进行比较, 炎黄样本的 SNP与已有的 炎黄样本 Genotyping位点进行了比较, 结果见下表 5和表 6。
表 5 HapMap样本的 SNP分析结果
Figure imgf000018_0001
表 6 炎黄样本的 SNP分析结果
炎黄 假阴性
SNP真阳 假阳性(假
样本名称 Genotyping位 (假阴性
性个数 阳性率 )
点数目 率)
单独测炎黄样本 397 89 0(0.00%) 0(0.00%)
Pooling中炎黄样
397 89 0(0.00%) 0(0.00%) 本 ( 6 ) CYP450基因分型
做完变异检测后, 根据每个基因在全基因组上的区域, 提取出每个基因的突变位点信 息。 根据这些突变位点信息与实施例 1中构建好的 CYP450基因标准型别数据库进行比较, 确定样本的基因型别信息和酶活性信息。 部分样本的检测结果如表 7所示。
Figure imgf000019_0001
分型结果显示, 采用本发明的方法得到的基因型别信息及酶活性信息与现有参考文献 记载一致。 工业实用性
本发明的构建 CYP450基因型别标准化数据库、 CYP450基因分型及酶活性鉴定的方 法, 能够有效地用于 CYP450基因分型和酶活性鉴定, 并且省时省工、 成本低、 结果准确。 尽管本发明的具体实施方式已经得到详细的描述, 本领域技术人员将会理解。 根据已 经公开的所有教导, 可以对那些细节进行各种修改和替换, 这些改变均在本发明的保护范 围之内。 本发明的全部范围由所附权利要求及其任何等同物给出。
在本说明书的描述中, 参考术语 "一个实施例"、 "一些实施例"、 "示意性实施例"、 "示 例"、 "具体示例"、 或 "一些示例" 等的描述意指结合该实施例或示例描述的具体特征、 结 构、 材料或者特点包含于本发明的至少一个实施例或示例中。 在本说明书中, 对上述术语 的示意性表述不一定指的是相同的实施例或示例。 而且, 描述的具体特征、 结构、 材料或 者特点可以在任何的一个或多个实施例或示例中以合适的方式结合。

Claims

权利要求书
1、 一种构建 CYP450基因标准型别数据库的方法, 其特征在于, 包括以下步骤: 将 CYP450基因型别突变信息对应的特定序列与人类全基因组标准序列进行比对,获得 CYP450特定序列与所述人类全基因组标准序列在每个碱基位置上的对应关系;
根据所述的对应关系,将所述 CYP450基因型别转换成以所述人类全基因组标准序列为 参考序列的基因型 , 获得所述 CYP450基因的标准化基因型别。
2、 根据权利要求 1所述的方法, 其特征在于, 所述 CYP450基因包括选自 CYP11A1、 CYP11BK CYP11B2、 CYP17A1、 CYP1A1、 CYP1A2, CYP1B1、 CYP20A1、 CYP21A2、 CYP24A1、 CYP26A1、 CYP26BK CYP26C1 , CYP27AK CYP27B1、 CYP27CK CYP2A13、 CYP2A6、 CYP2A7、 CYP2B6、 CYP2C18、 CYP2C19、 CYP2C8、 CYP2C9、 CYP2D6、 CYP2E1、 CYP2F1、 CYP2J2、 CYP2 CYP2S CYP2U1、 CYP2WK CYP39A CYP3A4, CYP3A43, CYP3A5. CYP3A7、 CYP46A1、 CYP4A11、 CYP4A22. CYP4B CYP4F1 CYP4F12、 CYP4F2, CYP4F22, CYP4F3 , CYP4F8、 CYP4V2、 CYP4X1、 CYP4Z1、 CYP51A1、 CYP5A1、 CYP7A1、 CYP7B1、 CYP8A1、 CYP8B1、 POR等 58个人类 CYP450基因的至少一种。
3、 根据权利要求 1所述的方法, 其特征在于, 进一步包括:
定位所述 CYP450基因于所述人类全基因组标准序列上,确定所述 CYP450基因编码序 列的起始位置和终止位置, 获得所述 CYP450基因型别突变信息对应的特定序列。
4、 根据权利要求 1所述的方法, 其特征在于, 所述 CYP450基因型别突变信息对应的 特定序列包含所述 CYP450基因编码序列的起始位置上游 5000bp至终止位置下游的 500bp 区域。
5、 根据权利要求 1所述的方法, 其特征在于, 所述人类全基因组标准序列为 hgl9。
6、 一种 CYP450基因标准型别数据库, 其特征在于, 所述数据库釆用权利要求 1-5任 一项所述的方法构建。
7、 根据权利要求 6所述的 CYP450基因标准型别数据库, 其特征在于, 所述 CYP450 基因标准型别数据库中, CYP450基因的各标准化基因型别对应有酶活性信息。
8、 根据权利要求 6所述的 CYP450基因标准型别数据库, 其特征在于, 所述 CYP450 基因包括选自 CYP11A1、 CYP11BK CYP11B2, CYP17AK CYP1A1、 CYP1A2、 CYP1B1、 CYP20A CYP21A2, CYP24AK CYP26A1 , CYP26BK CYP26C CYP27A1、 CYP27B CYP27C CYP2A13. CYP2A6. CYP2A7、 CYP2B6、 CYP2C18、 CYP2C19. CYP2C8、 CYP2C9、 CYP2D6、 CYP2E1、 CYP2FK CYP2J2, CYP2R1、 CYP2S1、 CYP2U1、 CYP2W1、 CYP39AK CYP3A4, CYP3A43、 CYP3A5、 CYP3A7、 CYP46A1、 CYP4A1 CYP4A22, CYP4B CYP4F11、 CYP4F12、 CYP4F2、 CYP4F22、 CYP4F3、 CYP4F8、 CYP4V2、 CYP4X1、 CYP4Z1、 CYP51A1、 CYP5A1、 CYP7AK CYP7B1、 CYP8A1、 CYP8BK POR基因的至 少一种。
9、 一种 CYP450基因分型的方法, 其特征在于, 包括: 获取待测样本 CYP450基因的外显子序列 , 釆用高通量测序平台测序并进行数据分析, 将分析结果与权利要求 6-8任一项所述的 CYP450基因标准型别数据库进行比较,从而得到 所述待测样本的基因型别。
10、 一种 CYP450酶活性鉴定方法, 其特征在于, 包括:
获取待测样本 CYP450基因的外显子序列, 采用高通量测序平台测序并进行数据分析 , 将分析结果与权利要求 7所述的 CYP450基因标准型别数据库进行比较,得到待测样本的基 因型别,并根据所述待测样本的基因型别对应的酶活信息获得所述待测样本的 CYP450酶活 性结果。
11、 根据权利要求 9或 10所述的方法, 其特征在于, 所述获取待测样本 CYP450基因 的外显子序列是通过以下步骤实现的:
A、 制备能够捕获 CYP450基因外显子序列的芯片, 所述芯片上含有与所述 CYP450基 因外显子序列反向互补的寡核苷酸探针;
B、 用待测样本的基因组 DNA制备序列捕获文库, 包括将所述待测样本基因组 DNA 打断为 200~500bp大小的片段, 进行末端处理后扩增得到所述序列捕获文库;
C、 将步驟 B制备得到的序列捕获文库与步骤 A的芯片杂交, 从而获取得到所述待测 样本的 CYP450基因外显子文库。
12、 根据权利要求 11所述的方法, 其特征在于, 在所述步驟 A中, 所述芯片含有能分 别与 58个人类 CYP450基因的所有外显子序列反向互补的寡核苷酸探针, 所述寡核苷酸探 针的长度为 55-105bp。
13、 根据权利要求 11所述的方法, 其特征在于, 在所述步骤 B中, 将所述待测样本基 因组 DNA打断为 200~300bp大小的片段。
14、 根据权利要求 11所述的方法, 其特征在于, 在所述步骤 B中, 所述末端处理包括 进行末端修复形成平末端磷酸化的 DNA片段, 并在所述平末端磷酸化的 DNA片段的 3'末 端加上碱基 "A" , 并进一步连接标签。
15、 根据权利要求 11所述的方法, 其特征在于, 在所述步骤 C中, 在进行所述杂交之 前, 将来自多个不同待测样本的序列捕获文库混合后再同时与步驟 A的芯片杂交, 每个文 库带有不同的标签碱基序列而相互区别。
16、 根据权利要求 11所述的方法, 其特征在于, 所述标签碱基序列长度为 6~8bp。
17、 根据权利要求 9或 10所述的方法, 其特征在于, 所述数据分析进一步包括: i、 过滤去掉影响信息分析的低质量测序序列;
ii、 以人类全基因组标准序列为参考序列, 将步骤 i得到的序列用比对软件进行比对; iii、选取比对到目标区域的序列进行后续分析, 所述目标区域是指 CYP450基因外显子 序列所在区域;
iv、数据质控合格后进行变异分析, 所述变异分析包括检测以下中的至少一种: 单核苷 酸多态性、 插入和删除、 结构性变异、 拷贝数变异。
18、 根据权利要求 17所述的方法, 其特征在于, 所述比对软件为选自 SOAP和 BWA 的至少一种。
PCT/CN2013/070080 2012-01-06 2013-01-05 Cyp450基因型别数据库及基因分型、酶活性鉴定方法 Ceased WO2013102441A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201210002976.3 2012-01-06
CN201210002976.3A CN103198236B (zh) 2012-01-06 2012-01-06 Cyp450基因型别数据库及基因分型、酶活性鉴定方法

Publications (1)

Publication Number Publication Date
WO2013102441A1 true WO2013102441A1 (zh) 2013-07-11

Family

ID=48720790

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2013/070080 Ceased WO2013102441A1 (zh) 2012-01-06 2013-01-05 Cyp450基因型别数据库及基因分型、酶活性鉴定方法

Country Status (2)

Country Link
CN (1) CN103198236B (zh)
WO (1) WO2013102441A1 (zh)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10395759B2 (en) 2015-05-18 2019-08-27 Regeneron Pharmaceuticals, Inc. Methods and systems for copy number variant detection
CN114231606A (zh) * 2021-11-29 2022-03-25 北京艾迪康医学检验实验室有限公司 一种快速分析cyp2c9基因型的方法
CN118335195A (zh) * 2024-06-13 2024-07-12 浙江省标准化研究院(金砖国家标准化(浙江)研究中心、浙江省物品编码中心) 一种基于高通量测序数据的str分型方法
US12071669B2 (en) 2016-02-12 2024-08-27 Regeneron Pharmaceuticals, Inc. Methods and systems for detection of abnormal karyotypes

Families Citing this family (20)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103436545B (zh) * 2013-09-05 2015-07-22 蔡剑平 包括4094c>a突变的cyp2d6基因片段、所编码的蛋白质片段及其应用
CN103436546B (zh) * 2013-09-05 2015-10-28 蔡剑平 包括4110c>g突变的cyp2d6基因片段、所编码的蛋白质片段及其应用
CN103436544B (zh) * 2013-09-05 2016-05-18 蔡剑平 包括1745c>g突变的cyp2d6基因片段、所编码的蛋白质片段及其应用
CN103468820B (zh) * 2013-09-27 2015-06-10 中国人民解放军第三军医大学 检测cyp4v2基因常见突变的试剂盒
CN104745592B (zh) * 2013-12-31 2020-02-21 第三军医大学第一附属医院 Cyp4v2基因突变体及其应用
CN104232753A (zh) * 2014-07-22 2014-12-24 百世诺(北京)医疗科技有限公司 一种检测17α-羟化酶缺乏症相关基因突变的试剂盒
CN104232754A (zh) * 2014-07-22 2014-12-24 百世诺(北京)医疗科技有限公司 一种检测11β-羟化酶缺乏症相关基因突变的试剂盒
CN106086192A (zh) * 2016-06-27 2016-11-09 上海泽因生物科技有限公司 他克莫司个性化用药相关基因的分型检测试剂盒
CN106529211A (zh) * 2016-11-04 2017-03-22 成都鑫云解码科技有限公司 变异位点的获取方法及装置
CN107944224B (zh) * 2017-12-06 2021-04-13 懿奈(上海)生物科技有限公司 构建皮肤相关基因标准型别数据库的方法及应用
CN107974490B (zh) * 2017-12-08 2019-05-14 东莞博奥木华基因科技有限公司 基于半导体测序的pku致病基因突变检测方法及装置
CN108647494B (zh) * 2018-05-21 2021-05-11 深圳华大基因科技服务有限公司 一种评估组装基因序列的准确性的方法
CN108707658B (zh) * 2018-06-12 2019-09-17 东莞博奥木华基因科技有限公司 一种个体化用药基因检测试剂盒及应用
CN109136266B (zh) * 2018-08-10 2022-02-18 深圳泓熙生物科技发展有限公司 用于治疗或预防结晶样视网膜色素变性的基因载体及其用途
CN110942806A (zh) * 2018-09-25 2020-03-31 深圳华大法医科技有限公司 一种血型基因分型方法和装置及存储介质
CN109825566B (zh) * 2018-12-26 2021-09-21 阅尔基因技术(苏州)有限公司 基因cyp11b2外显子的pcr引物组、试剂盒、扩增体系和检测方法
CN112992277B (zh) * 2021-03-18 2021-10-26 南京先声医学检验实验室有限公司 一种微生物基因组数据库构建方法及其应用
CN114023391A (zh) * 2021-10-28 2022-02-08 中国科学院西北高原生物研究所 一种呼伦贝尔羊snp数据库的构建方法
CN115976004B (zh) * 2022-12-27 2024-07-09 天津科技大学 一种黄体酮17α-羟化酶突变体及其应用
CN117106917B (zh) * 2023-09-05 2026-04-14 上海交通大学医学院附属仁济医院 Cyp2a6基因型作为预测乳腺癌新辅助化疗疗效标记物的用途

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101429559A (zh) * 2008-12-12 2009-05-13 深圳华大基因研究院 一种环境微生物检测方法和系统

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
JIANG, TAO ET AL.: "High-performance single-chip exon capture allows accurate whole exome sequencing using the Illumina Genome Analyzer.", SCIENTIA SINICA VITAE, vol. 41, no. 9, 2011, pages 714 - 721, XP019969844 *
MARSH, S. ET AL.: "SNP databases and pharmacogenetics: great start, but a long way to go", HUMAN MUTATION, vol. 20, 2002, pages 174 - 179, XP055077435 *
SIM, S.C. ET AL.: "The human cytochrome P450 (CYP) allele nomenclature website: a peer-reviewed database of CYP variants and their associated effects", HUMAN GENOMICS, vol. 4, no. 4, April 2010 (2010-04-01), pages 278 - 281, XP021126954 *

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10395759B2 (en) 2015-05-18 2019-08-27 Regeneron Pharmaceuticals, Inc. Methods and systems for copy number variant detection
US11568957B2 (en) 2015-05-18 2023-01-31 Regeneron Pharmaceuticals Inc. Methods and systems for copy number variant detection
US12071669B2 (en) 2016-02-12 2024-08-27 Regeneron Pharmaceuticals, Inc. Methods and systems for detection of abnormal karyotypes
CN114231606A (zh) * 2021-11-29 2022-03-25 北京艾迪康医学检验实验室有限公司 一种快速分析cyp2c9基因型的方法
CN118335195A (zh) * 2024-06-13 2024-07-12 浙江省标准化研究院(金砖国家标准化(浙江)研究中心、浙江省物品编码中心) 一种基于高通量测序数据的str分型方法

Also Published As

Publication number Publication date
CN103198236B (zh) 2017-02-15
CN103198236A (zh) 2013-07-10

Similar Documents

Publication Publication Date Title
CN103198236B (zh) Cyp450基因型别数据库及基因分型、酶活性鉴定方法
CN103198238B (zh) 构建药物反应相关基因标准型别数据库的方法及其应用
JP7810455B2 (ja) 血漿dnaの単分子配列決定
ES3037438T3 (en) Methods for analysis of circulating cells
CN102770558B (zh) 由母本生物样品进行胎儿基因组的分析
CN106715711B (zh) 确定探针序列的方法和基因组结构变异的检测方法
CN107177670B (zh) 一种高通量检测帕金森病致病基因突变的方法
JP6073461B2 (ja) 標的大規模並列配列決定法を使用した対立遺伝子比分析による胎児トリソミーの非侵襲的出生前診断
JP2025148581A (ja) メチル化分配アッセイにおいて無細胞dnaを解析するための組成物および方法
CN102409047B (zh) 一种构建杂交测序文库的方法
CN109971846A (zh) 使用双等位基因snp靶向下一代测序的非侵入性产前测定非整倍体的方法
CN110023509A (zh) 基因型分型测定中的非独特条形码
JP2020521216A (ja) 挿入および欠失を検出するための方法およびシステム
WO2015042980A1 (zh) 确定染色体预定区域中snp信息的方法、系统和计算机可读介质
WO2013091276A1 (zh) 一种检测dmd基因外显子缺失和/或重复的方法
CN102127819A (zh) Mhc区域核酸文库的构建方法及用途
CN103748234B (zh) 用于膀胱癌诊断的序列及其使用方法和应用
CN105950709A (zh) 试剂盒、建库方法以及检测目标区域变异的方法及系统
US20190316112A1 (en) Methods of capturing a nucleic acid including a target oligonucleotide sequence and uses thereof
CN105838720A (zh) Ptprq基因突变体及其应用
JP7626430B2 (ja) 薬物動態関連遺伝子の網羅的配列解析法とそれに使用されるプライマーセット
Jiang et al. High-performance single-chip exon capture allows accurate whole exome sequencing using the Illumina Genome Analyzer
US20190316195A1 (en) Methods of capturing a nucleic acid including a target oligonucleotide sequence and uses thereof
Żmieńko et al. PERSPECTIVES Transcriptome sequencing: next generation approach to RNA functional analysis
US11512346B2 (en) Method for sequencing a direct repeat

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13733772

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205 DATED 26/11/2015)

122 Ep: pct application non-entry in european phase

Ref document number: 13733772

Country of ref document: EP

Kind code of ref document: A1