WO2014023076A1 - 一种地中海贫血的分型方法及其应用 - Google Patents

一种地中海贫血的分型方法及其应用 Download PDF

Info

Publication number
WO2014023076A1
WO2014023076A1 PCT/CN2012/087645 CN2012087645W WO2014023076A1 WO 2014023076 A1 WO2014023076 A1 WO 2014023076A1 CN 2012087645 W CN2012087645 W CN 2012087645W WO 2014023076 A1 WO2014023076 A1 WO 2014023076A1
Authority
WO
WIPO (PCT)
Prior art keywords
gene
sample
tested
sequence
snp
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2012/087645
Other languages
English (en)
French (fr)
Inventor
曹红志
张涛
柴相花
陈仕平
李剑
王俊
汪建
杨焕明
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
BGI Shenzhen Co Ltd
Original Assignee
BGI Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by BGI Shenzhen Co Ltd filed Critical BGI Shenzhen Co Ltd
Publication of WO2014023076A1 publication Critical patent/WO2014023076A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6883Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/156Polymorphic or mutational markers

Definitions

  • the present invention belongs to the field of genomic bioinformatics analysis, and in particular, to a method for typing thalassemia and its use. Background technique
  • Thalassemia is a common autosomal recessive disorder. It is caused by globin gene variation, impaired globin biosynthesis, and insufficient or absent production. Thalassemia is extremely high in all provinces south of the Yangtze River in China, and it is one of the most influential genetic diseases in China.
  • peptide chains constituting globin There are four kinds of peptide chains constituting globin, namely ⁇ , , and 5 chains, which are respectively encoded by their corresponding genes. Deletion or point mutation of these genes can cause structural disorder of the peptide chain, resulting in changes in the composition of hemoglobin.
  • the two most common types are: a thalassemia due to alpha chain variation and beta thalassemia due to beta chain variation.
  • the ⁇ -globin gene is located in the ⁇ -globin gene cluster at the end of the short arm of chromosome 16, which consists of two repeated ⁇ genes ( ⁇ 2 and al), one embryonic a gene ( ⁇ 2), and three pseudogenes. ( ⁇ ⁇ , ⁇ 2 and ⁇ ) and a functionally unknown gene ( ⁇ 1), the order of which is: 5 '-2 ⁇ 22 ⁇ 12 ⁇ 22 ⁇ 12 ⁇ 22 ⁇ 12 ⁇ 12-3 ', and the total length is about 30kb.
  • the ⁇ -globin gene is located in the short arm of the chromosome 1 (1 1 ⁇ 15), and the gene sequence is 5'-2 ⁇ 20 ⁇ 2 ⁇ 2 ⁇ 2 ⁇ 2 ⁇ 2-3', and the length is about 60kb.
  • the classification of thalassemia is the detection of mutations in ⁇ -chain and ⁇ -chain related genes, including point mutations, insertion deletions and copy number variations.
  • There are many methods for detecting thalassemia including reverse dot blot hybridization, PCR2 single-strand conformation polymorphism analysis, and gene chip method.
  • Reverse dot blotting is a new technique for detecting point mutations. It replaces the immobilized target DNA with a fixed probe on the membrane, thus changing the traditional hybridization method to detect only one mutation at a time. the way. The membrane strips from which a plurality of specific probes are immobilized are hybridized with the amplified target sequence, so that one hybridization can be simultaneously detected and detected. However, reverse dot hybridization can only detect partial point mutations and is not universal.
  • PC 2 single strand conformation polymerase is based on the principle that the DNA single strand can be electrophoresed on a neutral polyacrylamide gel even if only one base changes. The mobility has changed.
  • the process is as follows: First, the target DNA fragment is amplified by PCR, and then the PCR product is denatured at 95 ° C and then subjected to neutral polyacrylamide gel. Electrophoresis, DNA electrophoresis bands can be displayed by methods such as radionuclide and silver staining. Because silver staining is simple and easy, and there is no radionuclide contamination, it has become the preferred method for displaying SSCP electrophoresis bands. However, this method is slow in typing and high in cost, and is not suitable for popularization and application.
  • Affymetrix began to use DNA chip technology to diagnose thalassemia.
  • the basic principle is to simultaneously hybridize with the same cDNA chip with two different fluorescent dye-labeled target sequences, and analyze the fluorescence signal intensity of different colors. It can reflect changes in gene expression. Since this technology can immobilize an extremely large number of probes on the support at the same time, a large number of biomolecules can be detected and analyzed at one time, thereby solving the single defect of the detection items such as conventional PCR and nucleic acid blot hybridization, but the method exists. Low throughput, low cost and other shortcomings.
  • Another object of the present invention is to provide a technique for detecting SNP and Indel mutations and an application thereof.
  • Another object of the present invention is to provide a method for detecting copy number variation and an application thereof.
  • Another object of the present invention is to provide a thalassemia typing technique and its use.
  • a method of detecting a copy number variation of a HBB gene of a sample to be tested comprising the steps of:
  • step (1) Correcting and/or counting the reading sequence obtained in step (1), obtaining the following absolute data values: reading number of the HBB gene of the sample to be tested, reading number of the control gene of the sample to be tested, reading of the normal sample HBB gene Ordinal number, the number of readings of normal sample control genes;
  • the HBB gene of the sample to be tested is sequenced using a high throughput sequencing technique to obtain a gene sequence.
  • the length of the HBB gene sequence of the sample to be tested obtained in (1) is 70-100 bp. It is preferably 80-90 bp, more preferably 87 bp.
  • the HBB gene sequence of the sample to be tested corresponds to the sample and/or gene using the index and the specificity of the primer.
  • the reading order obtained by the step (1) is corrected based on the SNP and/or InDel.
  • a method of detecting a copy number variation of a HBA gene of a sample to be tested comprising the steps of:
  • step (1) Correcting and/or counting the reading order obtained in step (1), obtaining the following absolute data values: the reading number of the normal sample HBA1 gene, the reading number of the normal sample HBA2 gene, the reading number of the normal sample control gene, The number of readings of the HBA1 gene to be tested, the number of readings of the HBA2 gene to be tested, and the number of readings of the control gene to be tested;
  • the reading number of the normal sample HBA1 gene/the reading number of the normal sample control gene The number of readings of the HBA2 gene to be tested The number of readings of the HBA2 gene/the number of readings of the test sample to be tested.
  • the HBA1 gene or HBA2 gene sequence of the sample to be tested obtained in (1) has a sequence length of 100-150 bp, preferably 120-140 bp, more preferably 128 bp.
  • the HBA1 gene or HBA2 gene of the sample to be tested is sequenced using a high-throughput sequencing technique to obtain a gene sequence.
  • the HBA1 gene or HBA2 gene of the sample to be tested is associated with the sample and the gene using the index and the specificity of the primer.
  • the reading order obtained by the step (1) is corrected according to the SNP and/or InDel.
  • the (4) includes the steps of:
  • a detecting unit for detecting a copy number variation of a HBB gene of a sample to be tested, the unit comprising a module:
  • a sequence acquisition module configured to obtain a HBB gene sequence of the sample to be tested, the sequence corresponding to the sample and/or the gene;
  • a read order correction and/or statistics module for correcting and/or counting the read order obtained by the module (1);
  • An output module for outputting information on the copy number variation of the HBB gene of the sample to be tested.
  • a detecting unit for detecting a copy number variation of a HBA gene of a sample to be tested, the unit comprising a module:
  • a sequence acquisition module configured to obtain an HBA1 gene sequence and an HBA2 gene sequence of the sample to be tested, the sequence corresponding to the sample and/or the gene;
  • An output module for outputting HBA gene copy number variation information of the sample to be tested.
  • a typing system for detecting thalassemia comprising:
  • system further comprises: (3) a detection unit of the SNP and/or InDel.
  • a method for detecting SNP and Indel variation of a sample to be tested comprising the steps of:
  • sequencing data in (1) is aligned with a reference sequence using an alignment software selected from the group consisting of BWA, SOAP, BLAST, MAQ, and the like.
  • the sample to be tested is a sample of a person.
  • the reference sequence is from hgl9.
  • a method of Mediterranean typing of a sample to be tested comprising the steps of:
  • step (2) Combine all the variation data of step (1) to obtain the complete variation of the sample to be tested, and then perform the Mediterranean classification of the sample to be tested.
  • sample SNP and InDel variation data to be tested are obtained by the method described in the sixth aspect.
  • the HBB gene copy number variation of the sample to be tested is detected by the method described in the first aspect.
  • the HBA gene copy number variation of the sample to be tested is detected by the method described in the second aspect.
  • a method of constructing a thalassemia type SNP and InDel relational database comprising the steps of:
  • a SNP and Indel database constructed using the method of the eighth aspect. It is to be understood that within the scope of the present invention, the various technical features of the present invention and the technical features specifically described hereinafter (as in the embodiments) may be combined with each other to constitute a new or preferred technical solution. Due to space limitations, we will not repeat them here. DRAWINGS
  • Figure 1 shows the SNP and InDel information analysis process for thalassemia.
  • FIG. 2 shows the HBB gene copy number detection information flow.
  • FIG. 3 shows the flow of HBA gene copy number information analysis.
  • Figure 4 shows the design of the primer.
  • Figure 5 shows a single site depth profile.
  • the inventors have conducted extensive and intensive research to provide a method for detecting the copy number variation of the HBB gene and/or HBA gene of a sample to be tested, the method comprising the steps of: obtaining a sequence of a gene of interest, the sequence and the sample and/or Corresponding to the gene; correcting and/or counting the reading sequence to obtain absolute data values of related genes; calculating A value or calculating R1 value, R2 value, R3 value; according to A value or R1 value, R2 value, R3 value Obtain information on the copy number variation of the sample HBB and/or HBA gene to be tested.
  • the present invention also provides a typing system and analytical method for detecting thalassemia. The present invention has been completed on this basis. the term
  • SNP single nucleotide polymorphism
  • Polymorphisms expressed by SNPs involve only a single base variation, which can be caused by a single base transition or transversion, or by the insertion or deletion of a base.
  • InDel and "insertion deletion mutation” are used interchangeably and refer to mutations involving nucleoside insertions and/or deletions.
  • CNV copy number variation
  • SNV DNA fragments. Segment size ranges from kb to Mb submicroscopic mutations. As a variation within a chromosome, CNV means that there are too many or insufficient DNA fragments in a certain cell. Double-end sequencing
  • the gene fragments (including DNA and cDNA) are sequenced, and the sequenced objects are a piece of physically continuous base sequence called an insert, the length of which is called the insert size.
  • double-end sequencing is the sequencing of the base sequences on both sides of the fragment from edge to interior.
  • the sequence measured is called read and the length is called read-length.
  • the read order measured on both sides is from the same insert, and the distance between the ends is insertsize, so the pairing relationship of the read order on both sides is determined. These two readings are called Pair-end reads.
  • High-throughput sequencing of the genome enables humans to detect abnormal changes in disease-associated genes as early as possible, and to facilitate in-depth research into the diagnosis and treatment of individual diseases.
  • Those skilled in the art can generally perform high throughput sequencing using conventional methods including, but not limited to, 454 FLX (Roche), Solexa Genome Analyzer (Illumina), Applied Biosystems, SOLID, and the like.
  • the common feature of these platforms is the extremely high sequencing throughput.
  • high-throughput sequencing can read 400,000 to 4 million sequences in one experiment. The read length is from 25bp depending on the platform. Up to 450 bp, so different sequencing platforms can read base numbers ranging from 1G to 14G in one experiment.
  • Solexa high-throughput sequencing includes two steps: DNA cluster formation and on-machine sequencing: a mixture of PCR amplification products is hybridized with a fixed sequencing probe immobilized on a solid phase carrier, and subjected to solid phase bridge PCR amplification to form a sequencing cluster; The sequencing cluster is sequenced by "edge synthesis-edge sequencing” to obtain a sequence of nucleic acid molecules in the sample.
  • the DNA cluster is formed by using a flow cell with a single-stranded primer attached to the surface, and the DNA fragment of the single-stranded state is immobilized on the chip by the principle that the linker sequence and the primer on the surface of the chip are complementary to each other by base complementation.
  • the fixed single-stranded DNA becomes double-stranded DNA
  • the double strand is denatured into a single strand, one end of which is anchored on the sequencing chip, and the other end is randomly and adjacent to another primer to be anchored, Forming a "bridge"
  • on the sequencing chip there are tens of millions of DNA single molecules simultaneously reacting; forming a single-stranded bridge, using the surrounding primers as amplification primers, and amplifying again on the surface of the amplification chip to form a double
  • the chain the double strand is denatured into a single strand, and becomes a bridge again.
  • the template called the next round of amplification continues to expand; after repeated rounds of 30 rounds of amplification, each single molecule is amplified 1000 times, called a single gram. Long DNA cluster.
  • DNA clusters were sequenced on a Solexa sequencer. During the sequencing reaction, the four bases were labeled with different fluorescence, and each base was blocked by a protected base. Only one base could be added to a single reaction. After reading the color of the reaction, the protection group is removed, and the next reaction can be continued. Thus, the exact sequence of the base is obtained. In the Solexa Multiplexed Sequencing process, Index is used to distinguish the samples, and after the conventional sequencing is completed, the Index part is additionally sequenced. By index identification, a plurality of different sequencing channels can be distinguished. sample. Thalassemia type SNP&InDel database
  • the invention provides a thalassemia type SNP&InDel database and a construction method thereof.
  • the contents of the database are as follows:
  • the first column is the relative position of the mutation
  • the second column is the specific mutation information: as the first row HBA1 represents the gene, A>G indicates that the reference sequence is abrupt from G to G, and the last column represents the type, indicating HBA1
  • the 71 bp position of the gene was changed from A to G, and the thalassemia type was Initiation codon (A -> G).
  • the type of the database can be supplemented.
  • the reference sequence of position 542 of HBA1 has G to T, which is a synonymous mutation, and the last type gives Normal.
  • the present invention also provides a method for constructing a thalassemia type-SNP&InDel relational database.
  • the method comprises the steps of: a) selecting a sequence of a plurality of normal disease-free genes as a reference sequence (such as a sequence corresponding to coordinates in hgl9) Extracted as a thalassemia reference sequence); Thalassemia-related genes include HBB, HBA 1, HBA2; b) Align the known types of HB A 1 gene, HBA2 gene and HBB gene in the existing thalassemia database with reference sequences The difference sites of the reference sequence, namely the SNP locus and the InDel locus, were obtained, and the SNP and InDel relationship of each type with respect to the reference sequence were obtained, and the thalassemia type SNP-InDel database was obtained.
  • the invention provides a method for detecting SNP and Indel mutations, comprising the steps of: sequencing data
  • the SNP&InDel linkage relationship with respect to the reference sequence is obtained by comparison with the corresponding reference sequence, and the type is determined by the SNP and InDel linkage relationship.
  • the sample sequence is compared with the reference sequence by using comparison software (preferably BWA, SOAP, BLAST, MAQ, etc.); then the comparison result file is format converted (preferably samtools software tool is converted), and the consistency is obtained. Sexual sequence file; Finally, compared with the SNP&InDel linkage database, the type results of thalassemia were obtained.
  • comparison software preferably BWA, SOAP, BLAST, MAQ, etc.
  • the SNP&InDel mutation detection method comprises the steps of:
  • the obtained SNP and InDel are annotated; then the SNP and InDel are compared to the Mediterranean type database to obtain the Mediterranean type, and the supplementary database does not exist but is obtained by the sample test. Type.
  • Figure 1 shows the SNP and InDel information analysis process for thalassemia. Detection of copy number variation
  • the present invention provides a method for detecting copy number variation, the method comprising detecting a copy number variation portion of the HBB gene and detecting a copy number variation portion of the HBA gene.
  • the method includes the following steps:
  • the CNV part is divided into the HBB gene detection part and the HBA gene detection part.
  • the number of reads of the control gene is calculated; in the HBA gene portion, the number of reads of the HBA 1 gene, the HBA2 gene, and the control gene are detected.
  • Detect sample to be read HBB /reads ( control): normal sample reads ( HBB ) /reads ( control ).
  • Sample to be tested reads HBAl /reads(control): normal sample reads ( HBAl ) /reads(control);
  • Sample to be tested reads HBA2 /reads(control): normal sample reads ( HBA2 ) /reads(control);
  • Sample to be tested reads HBAl /reads(HBA2): normal sample reads ( HBAl ) /reads(HBA2).
  • the overall variance is ( 19*0.1+69*0.03 ) /88.
  • the resulting mean bias and variance are used as fixed errors due to experimental errors, and the theoretical values for each class are regained.
  • the theoretical value of the 0.5 category is 0.5-( 0.02*20+0.03*70 ) /90.
  • the distance setting can be further away from the class 1 when the distance is limited.
  • the distance from 0.8 to 1 is calculated to be 0.2, then the distance from 0.7 to 0.5 is also 0.2, and then 0.15 is set as the limit distance. It is considered that the distance other than 0.15 is not reliable, but if 0.8 is modified to 0.85, then 0.8 belongs to class 1. The distance is 1.5, within the allowable range, and 0.7 does not belong to category 0.5.
  • the HBB gene copy number detection information flow and the HBA gene copy number information analysis flow are shown in Fig. 2 and Fig. 3, respectively.
  • type of thalassemia refers to the detection of variants of alpha chain and beta chain related genes, including point mutations, insertion deletions and copy number variations, and the like.
  • the present invention provides a method for typing thalassemia. Specifically, the following steps are included:
  • copy HBB gene copy number variation and data set corresponds to the HBA gene copy number variation and the data set type; at the same time consider the depth of the reading order, mark the label with too few readings; summarize the three parts of the results according to the unified label, get each The complete variation of a sample.
  • the main advantages of the invention include:
  • the method and system of the present invention have high accuracy, and the data to be tested is relatively low;
  • Detection of small fragment insertion and point mutations includes the HBB gene and the HBA gene. After primer-specific capture of genes, different samples were labeled with different indices, and 200-500 bp gel fragments were randomly interrupted and sequenced by high-throughput sequencing. Finally, the reads are divided into different samples and genes using the specificity of index+primer.
  • This embodiment uses the BWA comparison software and the samtools software, and the GATK software package as an example.
  • BWA comparison software Through the BWA comparison software, these sequences are sequence-aligned with a specific reference sequence, and the comparison result *.sampe file is obtained through two steps of aln and sampe.
  • the GATK software package is used to recalibrate the bam file.
  • the bwa comparison is not the same for the different reads of the same site.
  • the GATK software package is used to recalibrate the low quality SNP near InDel, which is more favorable and accurate. And SNP.
  • the original consistency sequence is obtained by the samtools pileup, and the quality value and depth are re-filtered by the program written in the PERL language to obtain the trusted SNP and InDeL.
  • the SNP near InDel is corrected, for example, HBB97 C G , which indicates the position of HBB gene 97, and the mutation of ref C is G.
  • HBB 97 C -T indicating the position of HBB gene 98, missing a T.
  • Such mutations can be replaced by delins, ie insertions after deletion.
  • the correction is CT/G, which means that G replaces the CT base of ref.
  • SNP and InDel formats are organized, with "gene-location-variation" as the recognition unit, and the type-to-type database is given, and the type of thalassemia is given.
  • the variant is not in the database and is marked as "Unknown”.
  • the soft chipping position generated in the local coupling algorithm soft chipping is due to the pair end reads in the comparison due to a reads can be Uniq, and the second appears a lot of mismatch, the second reads It will not be discarded, but will be retained and marked s.
  • the cleavage sites can be found by using S. These cleavage sites may be caused by large fragment insertion or by deletion. Complex fragments such as insertion, deletion and insertion after deletion can be found by collecting S sites. .
  • HBbp gene fragment of 87 bp in length was captured, and each sample was labeled with an index.
  • the HBB gene was sequenced using high-throughput sequencing technology.
  • the data is indexed to the specificity of the index+primer, and the reads are divided into each sample and each different gene, and the detection is performed simultaneously.
  • the recognition sequence is corrected, and the recognition sequence is corrected according to the SNP condition in the same sample due to the mutation of the recognition sequence region.
  • an error message is given. Because there are linkages between multiple SNPs, it is not certain that InDel will seriously affect the sequence, so both cases are output.
  • the recognition sequence is used to re-divide the reads into each sample to obtain 4 absolute data values, gp: the number of reads of the HBB gene of the sample to be tested, the number of reads of the control sample of the sample to be tested, the number of reads of the normal sample HBB gene, and the normal sample. Control the number of gene reads.
  • Ratio sample to be tested HBB gene reads / test sample control gene reads: normal sample HBB gene reads / normal sample control gene reads number.
  • HBA1 and HBA2 are very similar in the HBA gene, so only one pair of primers is designed to capture both HBA1 and HBA2.
  • the HBA gene fragment of 128 bp in length was captured, and the HBA gene was sequenced by hiseq sequencing technique after index-specific labeling of each sample at both ends. After the machine is down, the data is divided into each sample and each different gene by using the specificity of index+primer, and the detection is performed simultaneously.
  • the modified recognition sequence is partially identical to the HBB gene.
  • the HBA is divided into HBA1 and HBA2 using the difference sequence of HBA1 and HBA2 from the HBA.
  • it is necessary to count 6 original reads, which are normal sample HBA1 gene reads, normal sample HBA2 gene reads, The normal sample control gene reads, the number of HBA1 gene reads in the sample to be tested, the number of HBA2 gene reads in the sample to be tested, and the number of control gene reads in the sample to be tested.
  • the HBA part is more complicated, and three ratios need to be calculated, defined as Rl, R2, and R3.
  • R 1 sample to be tested HB A 1 gene reads / sample to be tested control gene reads: normal sample HBA 1 gene reads / normal sample control gene reads number.
  • test sample control gene reads number normal sample HBA2 gene reads number / normal sample control gene reads number.
  • R3 sample to be tested HB A2 gene reads number I sample to be tested HB A 1 gene reads: normal sample HBA2 gene reads / normal sample HBA 1 gene reads number.
  • R1 and R2 are used to obtain the absolute value of the HBA1 gene copy number and the absolute value of the HBA2 gene copy number, and finally use R3 for detection.

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Chemical & Material Sciences (AREA)
  • Physics & Mathematics (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Health & Medical Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Analytical Chemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • Biophysics (AREA)
  • Organic Chemistry (AREA)
  • Biotechnology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Genetics & Genomics (AREA)
  • Wood Science & Technology (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Theoretical Computer Science (AREA)
  • Evolutionary Biology (AREA)
  • Zoology (AREA)
  • Medical Informatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Molecular Biology (AREA)
  • Microbiology (AREA)
  • Immunology (AREA)
  • Pathology (AREA)
  • Biochemistry (AREA)
  • General Engineering & Computer Science (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Description

一种地中海贫血的分型方法及其应用
技术领域
本发明属于基因组生物信息分析领域, 具体地, 本发明涉及一种地中海贫 血的分型方法及其应用。 背景技术
地中海贫血是一种常见的常染色体隐性遗传病。 它是由于珠蛋白基因变 异, 使珠蛋白生物合成受阻、 产量不足或缺失所致。 地中海贫血在我国长江以 南各省发病率极高, 是我国影响最大的遗传病之一。
组成珠蛋白的肽链有 4种, 即 α、 、 和 5链, 分别由其相应的基因编码。 这些基因的缺失或点突变可造成肽链的合成障碍, 致使血红蛋白的组分改变。 最常见有两种类型: 由于 α链变异导致的 a地中海贫血和由于 β链变异导致的 β地中海贫血。
α珠蛋白基因位于第 16号染色体短臂末端的 α珠蛋白基因簇中, 该基因 簇包括二个重复的 α基因 (α2和 al)、 一个胚胎期类 a基因 (ζ2 )、 三个假基因 (Ψζ ΐ、 Ψα2 和 Ψαΐ)和一个功能未知基因 (Θ 1), 其排列顺序为: 5 '-2ζ22Ψζ 12Ψα22Ψα 12α 22α 12Θ 12-3 ' , 全长约 30kb。 β珠蛋白基因位于 1 1号染 色体短臂(1 1ρ 15), 基因排列顺序为 5'-2ε20γ2Αγ2Ψβ2δ2β2-3 ', 长度约 60kb。
地中海贫血的分型即检测 α链和 β链相关基因的变异, 包括点突变, 插入 缺失和拷贝数变异。地中海贫血检测方法有很多种, 包括反向点杂交法, PCR2 单链构象多态性分析方法, 基因芯片法等。
反向点杂交法: 反向点杂交是检测点突变的一种新技术, 它与是将膜上固 定探针取代了固定靶 DNA方式, 从而改变了传统杂交法一次只能检测一种突 变的方式。 用固化了多种特异性探针的膜条与扩增靶序列杂交, 使一次杂交即 可同时筛查出被检测出来。 但是反向点杂交法只能检测部分点突变, 不具有普 遍性。
PC 2单链构象多态性分析( single strand conformation polymerase , SSCP) 其原理是 DNA单链上即使只有一个碱基发生改变, 就可使该 DNA单链在中 性聚丙烯酰胺凝胶上的电泳迁移率发生改变。 其过程是: 先用 PCR方法扩增 靶 DNA片段, 然后将 PCR产物经 95 °C变性后在中性聚丙烯酰胺凝胶上进行 电泳, DNA 电泳带可用放射性核素及银染色等方法显示。 由于银染法简便易 行, 且无放射性核素污染, 目前已成为首选方法用于显示 SSCP电泳带。 但是 这种方法分型慢, 费用高, 不适宜推广应用。
基因芯片法: 1997年 Affymetrix公司开始将 DNA芯片技术用于诊断地中 海贫血, 其基本原理是用 2 种不同荧光染料标记的靶序列同时与同 1 个 cDNA芯片杂交, 通过不同颜色的荧光信号强度分析即可反映出基因表达的变 化。 由于用该技术可以将极其大量的探针同时固定于支持物上, 所以一次可以 对大量的生物分子进行检测分析, 从而解决了传统 PCR和核酸印迹杂交等检 测项目单一的缺陷, 但是该方法存在低通量、 费用昂贵不足等缺点。
目前本领域还没有一种准确简便的一次性检测 SNP、 InDeK CNV的方法, 因此迫切需要开发适合于实际应用的地中海贫血的分型方法。 发明内容
本发明的目的是提供一种地中海贫血型别关系数据库及其建立方法。
本发明的另一目的是提供一种检测 SNP和 Indel变异的技术及其应用。 本发明的另一目的是提供一种拷贝数变异的检测方法及其应用。
本发明的另一目的是提供一种地中海贫血分型技术及其应用。 在本发明的第一方面,提供了一种检测待测样本 HBB基因拷贝数变异的方 法, 包括步骤:
(1) 获得待测样本的 HBB基因序列, 所述序列与样品和 /或基因相对应;
(2) 对步骤 (1)获得的读序进行修正和 /或统计, 获得下述绝对数据值: 待测样本 HBB基因的读序数、待测样本对照基因的读序数、 正常样本 HBB 基因的读序数、 正常样本对照基因的读序数;
(3) 计算 A值,
待测样本 HBB基因的读序数 /待测样本对照基因的读序数
Ait= ;
正常样本 ΗΒΒ基因的读序数 /正常样本对照基因的读序数
(4) 根据 Α值, 获得待测样本 HBB基因拷贝数变异的检测结果。
在另一优选例中, 用高通量测序技术对待测样本的 HBB基因测序, 获得基 因序列。
在另一优选例中, (1)中获得的待测样本的 HBB基因序列长度为 70-100bp, 较佳地为 80-90bp, 更佳地为 87bp。
在另一优选例中, 利用标签 (index) 以及引物的特异性将待测样本的 HBB 基因序列与样品和 /或基因相对应。
在另一优选例中,在 (2)中,根据 SNP和 /或 InDel对骤 (1)获得的读序进行修正。 在本发明的第二方面,提供了一种检测待测样本 HBA基因拷贝数变异的方 法, 包括步骤:
(1) 获得待测样本的 HBA1基因序列和 HBA2基因序列, 所述的序列与样品 和 /或基因相对应;
(2) 对步骤 (1)获得的读序进行修正和 /或统计, 获得下述绝对数据值: 正常样本 HBA1基因的读序数, 正常样本 HBA2基因的读序数, 正常样本对 照基因的读序数,待测样本 HBA1基因的读序数,待测样本 HBA2基因的读序数, 待测样本对照基因的读序数;
(3) 计算下述的比值:
待测样本 HBA1基因的读序数 /待测样本对照基因的读序数
R1值 =
正常样本 HBA1基因的读序数 /正常样本对照基因的读序数 待测样本 HBA2基因的读序数 /待测样本对照基因的读序数 正常样本 HBA2基因的读序数 /正常样本对照基因的读序数 待测样本 HBA2基因的读序数 /待测样本 HBA1基因的读序数
1 3值= ; 正常样本 HBA2基因的读序数 /正常样本 HBA1基因的读序数 (4) 根据 R1值、 R2值和 R3值, 获得待测样本 HBA基因拷贝数变异检测结 果。
在另一优选例中, (1)中获得的待测样本的 HBA1基因或 HBA2基因序列长度 为 100-150bp, 较佳地为 120-140bp, 更佳地为 128bp。
在另一优选例中, 用高通量测序技术对待测样本的 HBA1基因或 HBA2基因 测序, 获得基因序列。
在另一优选例中, 利用标签(index) 以及引物的特异性将待测样本的 HBA1 基因或 HBA2基因与样品和基因相对应。 在另一优选例中,在 (2)中,根据 SNP和 /或 InDel对骤 (1)获得的读序进行修正。 在另一优选例中, 所述的 (4)包括步骤:
(i) 对一个文库下所有样品的 Rl值、 R2值、 R3值进行大概归类;
(ϋ) 计算每个类别中的方差和偏离度 (偏离度=均值一类别理论值) ;
(iii) 综合考虑每个类别个数大于 15个值的方差与偏离度, 构建整个文库 的方差和偏离度;
(iv) 计算每个值与每类别的理论值之间的距离;
(v) 选取每个标签 (index) 的 R1 值、 R2值、 R3值的最短距离作为最佳 选项。
在本发明的第三方面, 提供了一种检测待测样本 HBB基因拷贝数变异的 检测单元, 所述单元包括模块:
(1) 序列获取模块, 用于获得待测样本的 HBB基因序列, 所述序列与样品和 /或基因相对应;
(2) 读序修正和 /或统计模块, 用于对模块 (1)获得的读序进行修正和 /或统计;
(3) 比值计算模块, 用于计算 A值; 和
(4) 输出模块, 用于输出待测样本 HBB基因拷贝数变异的信息。
在本发明的第四方面,提供了一种检测待测样本 HBA基因拷贝数变异的检 测单元, 所述单元包括模块:
(1) 序列获取模块, 用于获得待测样本的 HBA1基因序列和 HBA2基因序列, 所述的序列与样品和 /或基因相对应;
(2) 读序修正和 /或统计模块, 用于对模块 (1)获得的序列读序进行修正和 /或 统计;
(3) 比值计算模块, 用于计算 R1值、 R2值、 R3值;
(4) 型别判断模块, 用于判断 HBA基因拷贝数变异的型别; 和
(5) 输出模块, 用于输出待测样本 HBA基因拷贝数变异信息。
在本发明的第五方面, 提供了一种检测地中海贫血的分型系统, 包括:
(1) 第三方面所述的检测待测样本 HBB基因拷贝数变异的检测单元; 和
(2) 第四方面所述的检测待测样本 HBA基因拷贝数变异的检测单元。
在另一优选例中, 所述的系统还包括: (3) SNP和 /或 InDel的检测单元。 在本发明的第六方面, 提供了一种检测待测样本 SNP和 Indel变异的方法, 包括步骤:
(1) 获得待测样本的基因测序数据; (2) 将 (1)中的测序数据与参考序列进行比对, 获得相对于参考序列的 SNP 和 InDel连锁关系数据;
(3) 根据 (2)中所述的连锁关系数据确定待测样本 SNP和 Indel变异信息。 在另一优选例中,用选自下组的比对软件对 (1)中的测序数据与参考序列进 行比对: BWA、 SOAP, BLAST, MAQ等。
在另一优选例中, 所述的待测样本为人的样本。
在另一优选例中, 所述的参考序列来自 hgl9。
在本发明的第七方面, 提供了一种对待测样本进行地中海分型的方法, 包 括步骤:
(1) 获得待测样本 SNP和 InDel变异数据, 以及 HBB基因、 HBA基因拷贝数 变异数据;
(2) 结合步骤 (1)的所有变异数据, 获得待测样本的完整变异情况, 从而对待 测样本进行地中海分型。
在另一优选例中, 用第六方面所述的方法获得待测样本 SNP和 InDel变异数 据。
在另一优选例中, 用第一方面所述的方法对待测样本 HBB基因拷贝数变异 进行检测。
在另一优选例中,用第二方面所述的方法对待测样本 HBA基因拷贝数变异 进行检测。
在本发明的第八方面, 提供了一种构建地中海贫血型别 SNP和 InDel关系数 据库的方法, 包括步骤:
(1) 将现有地中海贫血数据库中已知型别的 HBA1基因、 HBA2基因禾 ΠΗΒΒ 基因与作为参考序列的已知正常无病基因序列进行比对, 获得差异位点, 所述 的差异位点为 SNP和 InDel位点;
(2) 获得每个型别相对于参考序列的 SNP和 InDel关系, 得到地中海贫血型 别 SNP-InDel数据库。
在本发明的第九方面, 提供了一种 SNP和 Indel数据库, 是用第八方面所述 的方法构建的。 应理解, 在本发明范围内中, 本发明的上述各技术特征和在下文 (如实施例) 中具体描述的各技术特征可以互相组合, 从而构成新的或优选的技术方案。 限于 篇幅, 在此不再一一累述。 附图说明
下列附图用于说明本发明的具体实施方案, 而不用于限定由权利要求书所 界定的本发明范围。
图 1显示了地中海贫血 SNP和 InDel信息分析流程。
图 2显示了 HBB基因拷贝数检测信息流程。
图 3显示了 HBA基因拷贝数信息分析流程。
图 4显示了引物 (primer) 的设计思路。
图 5显示了单个位点深度分布图。 具体实施方式
本发明人经过广泛而深入的研究, 提供了一种检测待测样本 HBB基因和 / 或 HBA基因拷贝数变异的方法, 所述方法包括步骤: 获得目的基因序列, 所 述序列与样品和 /或基因相对应; 对所述读序进行修正和 /或统计, 获得相关基因 的绝对数据值; 计算 A值或计算 R1值、 R2值、 R3值; 根据 A值或 R1值、 R2 值、 R3值获得待测样本 HBB和 /或 HBA基因拷贝数变异信息。本发明还提供了 检测地中海贫血的分型系统和分析方法。 在此基础上完成了本发明。 术语
SNP
如本文所用, 术语" SNP"或"单核苷酸多态性" 可以互换使用, 是指在基因 组水平上由单个核苷酸的变异所引起的 DNA序列的多态性。 SNP是可遗传变异 中最常见的一种, 占所有已知多态性的绝大多数。
SNP 所表现的多态性只涉及到单个碱基的变异, 这种变异可由单个碱基 的转换 (transition)或颠换 (transver sion)所引起, 也可由碱基的插入或缺失所致。
InDel
如本文所用, 术语" InDel"和"插入缺失突变 "可以互换使用, 是指涉及核苷 酸插入和 /或缺失的突变。
CNV
如本文所用, 术语"拷贝数变异"和" CNV"可以互换使用, 都是指 DNA 片 段大小范围从 kb到 Mb的亚微观突变, 作为一种染色体内的变异, CNV意味 着某个细胞内的某些 DNA片段过多或者不足。 双末端测序
对基因片段 (包括 DNA和 cDNA)进行测序, 其测序对象都是一段物理连续 的碱基序列片段, 该片段称为插入片段, 其长度称为插入片段长度 (insertsize )。
如本文所用, 术语"双末端测序"是对该片段的两侧碱基序列从边缘向内部 的测序, 测得的序列称为读序 (read) , 长度称为读长 (read-length)。 两侧测得的 读序是来自于同一个插入片段, 并且其末端距离为 insertsize, 故两侧读序的配 对关系确定。 这两个读序被称为配对读序 (Pair-end reads)。
高通量测序
基因组的高通量测序使得人类能够尽早地发现与疾病相关基因的异常变 化, 有助于对个体疾病的诊断和治疗进行深入的研究。 本领域技术人员通常可 以采用常规方法进行高通量测序, 包括但不限于: 454FLX(Roche公司)、 Solexa Genome Analyzer(Illumina公司)禾卩 Applied Biosystems 公司的 SOLID等。这些平 台共同的特点是极高的测序通量, 相对于传统测序的 96道毛细管测序, 高通量 测序一次实验可以读取 40万到 400万条序列,根据平台的不同,读取长度从 25bp 到 450bp不等, 因此不同的测序平台在一次实验中, 可以读取 1G到 14G不等的 碱基数。
Solexa 高通量测序包括 DNA簇形成和上机测序两个步骤: PCR扩增产物的 混合物与固相载体上固定的测序探针进行杂交, 并进行固相桥式 PCR扩增, 形成 测序簇; 对所述测序簇用"边合成 -边测序法"进行测序, 从而得到样本中核酸分子 的序列。
DNA簇的形成是使用表面连有一层单链引物 (primer)的测序芯片 (flow cell), 单链状态的 DNA片段通过接头序列与芯片表面的引物通过碱基互补配对的原 理被固定在芯片的表面, 通过扩增反应, 固定的单链 DNA变为双链 DNA, 双 链再次变性成为单链, 其一端锚定在测序芯片上, 另一端随机和附近的另一个 引物互补从而被锚定, 形成"桥"; 在测序芯片上同时有上千万个 DNA单分子发 生以上的反应; 形成的单链桥, 以周围的引物为扩增引物, 在扩增芯片的表面 再次扩增, 形成双链, 双链经变性成单链, 再次成为桥, 称为下一轮扩增的模 板继续扩增; 反复进行了 30轮扩增后, 每个单分子得到 1000倍扩增, 称为单克 隆的 DNA簇。
DNA簇在 Solexa测序仪上进行边合成边测序, 测序反应中, 四种碱基分别 标记不同的荧光,每个碱基末端被保护碱基封闭,单次反应只能加入一个碱基, 经过扫描, 读取该次反应的颜色后, 该保护集团被除去, 下一个反应可以继续 进行, 如此反复, 即得到碱基的精确序列。 在 Solexa多重测序 (Multiplexed Sequencing)过程中会使用 Index(标签)来区分样品, 并在常规测序完成后, 针对 Index部分额外进行测序, 通过 Index的识别, 可以在 1条测序甬道中区分多种不 同的样品。 地中海贫血型别 SNP&InDel数据库
本发明提供了一种地中海贫血型别 SNP&InDel数据库及其构建方法。所述 的数据库的内容如下所示:
Figure imgf000009_0001
如上标所示, 第一列是突变的相对位置, 第二列是具体突变信息: 如第一 行 HBA1表示基因, A>G表示参考序列由 A突变为 G, 最后一列代表型别, 指 出 HBA1 基因 71bp位置由 A突变为 G, 地中海贫血型别为 Initiation codon (A -〉 G)。
在不断的测试中, 可对数据库的型别进行补充, 如第二行, 在测试样本中 发现 HBA1第 542位置参考序列有 G变为 T, 属于同义突变, 最后的型别给出 Normal。
本发明还提供了一种构建地中海贫血型别 -SNP&InDel关系数据库的方法, 在一个优选例中,包括步骤: a) 选择多种正常无病基因的序列作为参考序列(如 hgl9中对应坐标的序列提取出来作为地中海贫血参考序列) ; 地中海贫血相关 基因包括 HBB, HBA 1, HBA2; b) 将现有地中海贫血数据库中已知型别的 HB A 1 基因、 HBA2基因和 HBB基因与参考序列比对, 获得参考序列的差异位点, 即 SNP位点和 InDel位点, 得到每个型别相对于参考序列的 SNP和 InDel关系, 得到地中海贫血型别 SNP-InDel数据库。
SNP和 InDel的检测
本发明提供了一种检测 SNP和 Indel变异的方法, 包括步骤: 将测序数据 与对应的参考序列做比对, 得到相对于参考序列的 SNP&InDel连锁关系, 通过 所述的 SNP和 InDel连锁关系确定出型别。
具体地, 用比对软件 (优选 BWA、 SOAP, BLAST, MAQ等工具) 将样 本序列与参考序列进行比对;再将比对的结果文件进行格式转换(优选 samtools 软件工具进行转换) , 得到一致性序列文件; 最后与 SNP&InDel连锁关系数据 库比较, 获得地中海贫血的型别结果。
在一个优选例中, 所述的 SNP&InDel变异检测方法包括步骤:
a) 寻找 SNP和 InDel位点: 对样本进行高通量测序, 获得测序数据; 将所 述的测序数据与参考序列进行比对, 得到 SNP位点和 InDel位点, 获得深度分 布 (图 5 ) 结果; 对比对结果进行质量控制, 过滤 SNP, InDel支持数中质量值 低于 20的位点;
b) 校正 SNP位点和 InDel位点: 由于利用 bwa比对的罚分策略, InDel附 近如果有一样的重复序列, 会导致 InDel calling出现偏差: 比如 ACTCT缺失 为 CT, 无论是前面 CT缺失还是后面 CT缺失, bwa都展示为前面的 CT缺失, 导致得到的 InDel跟数据库比对不上, 对于这些重复序列, 都会基于数据库将 缺失位置调整过来, 然后利用校正好的 InDel对 SNP再进行校正。 例如 97位 置缺失 C, 98位置参考序列由 T突变为 G, 这时将调整为 CT被替换 G。
c) 寻找异常变异: 利用 bwa比对中的 smith-waterman算法, 将 bwa比对 中出现的 S来记录复杂变异; 利用 pairend关系, 一条链比对很好, 第二条链 的一部分不能很好比对到参考序列, 这时就会出现 S, 例如 28S72M表示前面 28bp碱基没有比对到这个位置, 利用更多的 reads在同一位置出现的 S来定义 异常变异。本领域的普通技术人员可以使用常规方法,如 Shin Suzuki, Tomohiro Yasuda, Yuichi Shiraishi, Satoru Miyano, Masao Nagasaki , ClipCrop等人在文献 BMC Bioinformatics201 1 , 12(Suppl 14):S7: A tool for detecting structural variationswith single-base resolution using soft-clipping information中描述的 H 样, 来寻找大片段的插入缺失和插入后缺失等复杂变异。 将得到的变异位点作 为错误输出, 即得到复杂变异的信号数据, 利用 Sanger测序法重新测序直接得 到正确的变异位点。
d) 为了判定变异是否是同义突变, 将得到的 SNP和 InDel做注释; 然后 将 SNP和 InDel比对到地中海型别数据库后得到地中海型别, 同时补充数据库 中不存在但通过样本测试得到的型别。
图 1显示了地中海贫血 SNP和 InDel信息分析流程。 拷贝数变异的检测
本发明提供了一种检测拷贝数变异的方法, 所述方法包括检测 HBB 基因 的拷贝数变异部分和检测 HBA基因的拷贝数变异部分。
在本发明的一个优选例中, 包括如下步骤:
a) 计算参考基因与待测基因的读序 (reads ) 数: CNV部分分为 HBB基 因检测部分和 HBA基因检测部分。 首先利用 index+primer将读序 (reads ) 分 到每个样品中。因为 HBA1和 HBA2基因同源性很高,设计时只设计一对引物, 对两个基因进行捕获, 利用 HBA1和 HBA2中的差异部分对实验得到的数据再 次进行过滤得到 HBA1和 HBA2的 reads数。 在 HBB基因部分中, 计算 HBB 基因, control基因的 reads数; 在 HBA基因部分中, 检测 HBA 1基因, HBA2 基因和 control基因的 reads数。
b )计算相对比例: HBA基因检测部分和 HBB基因检测部分引入正常样本, 排除标签 (index) 的扩增问题。
在 HBB基因检测部分中,
检测待测样本 reads ( HBB ) /reads ( control): 正常样本 reads ( HBB ) /reads ( control ) 。
在 HBA基因检测部分中,
待测样本 reads ( HBAl ) /reads(control): 正常样本 reads ( HBAl ) /reads(control);
待测样本 reads ( HBA2 ) /reads(control): 正常样本 reads ( HBA2 ) /reads(control);
待测样本 reads ( HBAl ) /reads(HBA2): 正常样本 reads ( HBAl ) /reads(HBA2)。
因为 HBA和 HBB的 CNV检测很相似, 但 HBA更为复杂, 以 HBA为例 描述 CNV检测: HBA定义 Rl=待测样本 reads ( HBAl ) /reads(control): 正常 样本 reads ( HBAl ) /reads(control); R2=待测样本 reads ( HBA2 ) /reads(control): 正常样本 reads ( HBA2 ) /reads(control) , R3=待测样本 reads ( HBAl ) /reads(HBA2): 正常样本 reads ( HBAl ) /reads(HBA2)。
而 HBB定义为 Rl=检测待测样本 reads ( HBB ) /reads ( control) : 正常样 本 reads ( HBB ) /reads ( control) 。 因为一个文库有数个 (一般 96个) 样品, 将这 96个样品结合在一起进行分型, 以排除实验对整个文库的影响。 c) 对一个文库下所有样品的 Rl, R2, R3进行大概归类, 因为 Rl, R2, R3的理论值是 0、 0.5、 1、 1.5等, 将距离最近的值归于这一类, 例如 0.6归于 0.5类, 表示只有一个拷贝, 1.1归于 1类, 表示拷贝数正常, 对于中间值, 例 如 0.75, 可以分到 0.5类和 1类中, 表示这个值暂时无法确认属于哪一类。
d) 计算每个类别中的方差和偏离度 (偏离度=均值一类别理论值) , 计算 均值与理论值之间的差距。例如理论值是 0.5, 而属于这类别的均值是 0.48, 那 么这一类的偏离度为 0.02。
e)综合考虑每个类别个数大于 15个值的方差与偏离度, 构建整个文库的 方差和偏离度。 例如 0.5类别中有 20个样品, 偏离度为 0.02, 方差为 0.1, 类 别 1 中有 70个样品, 偏离度为 0.03, 方差为 0.03。 那么整体的均值偏离度为
( 0.02*20+0.03*70 ) /90, 整体的方差为 ( 19*0.1+69*0.03 ) /88。 得到的均值偏 离度和方差作为由于实验误差导致的固定误差, 重新得到每类别的理论值。 例 如 0.5类别的理论值为 0.5- ( 0.02*20+0.03*70 ) /90。
f ) 计算每个值与每类别的理论值之间的距离, 计算公式: (Xi - ave(X))/sqrt(var(x)), 其中 ave (X) 表示理论值, sqrt (var ( x) ) 表示方差。 计算每个值的最短距离作为这个值应该属于的类别。例如最原始的值 0.75无法 分开, 但 1类实际值为 0.85,0.5类值的实际值为 0.35, 这时, 0.75判定距离是, 就距离 1类更近了。
g) 选取每个 index 的 Rl, R2, R3的最短距离作为最佳选项。 同时设定 距离限制, 防止部分波动。 同时利用贝叶斯思想, 即正常样品的概率很大, 所 以, 对于距离限制时设置可以距离 1类更远些。 例如 0.8到 1的距离算出来为 0.2, 然后 0.7到 0.5的距离也为 0.2, 然后设置 0.15为限制距离, 认为 0.15以 外的距离都不可靠,但如果 0.8修改为 0.85,这时 0.8属于 1类的距离就是 1.5, 在允许之内, 而 0.7就不属于 0.5类。
HBB基因拷贝数检测信息流程和 HBA基因拷贝数信息分析流程分别见图 2和图 3。 地中海贫血分型
如本文所用, 术语"地中海贫血的分型"是指检测 α链和 β链相关基因的变 异, 包括点突变, 插入缺失和拷贝数变异等。
本发明提供了一种地中海贫血分型方法。 具体包括如下步骤:
将 SNP, InDel与型别数据库对应起来, 将 HBB基因拷贝数变异与数据集 型别对应起来, 将 HBA 基因拷贝数变异与数据集型别对应起来; 同时考虑读 序的深度问题, 标记读序太少的标签; 将三部分变异结果根据统一的标签总结 到一起, 获得每一个样品的完整变异情况。 本发明的主要优点包括:
(1) 本发明方法和系统准确性高, 对待测的数据要求比较低;
(2) 本发明方法和系统对于信息的分析过程快速, 操作简单。 下面结合具体实施例, 进一步阐述本发明。 应理解, 这些实施例仅用于说 明本发明而不用于限制本发明的范围。 下列实施例中未注明具体条件的实验方 法, 通常按照常规条件如 Sambrook等人, 分子克隆: 实验室手册 (New York: Cold Spring Harbor Laboratory Press, 1989)中所述的条件, 或按照制造厂商所建 议的条件。 实施例 1
960个样品的地中海贫血分型
本实施例的目的: 通过捕获地中海贫血相关基因 HBA基因和 HBB基因, 用 高通量测序技术对 SNP和 InDel与 CNV部分序列进行测序, 利用重测序技术和 CNV分析技术对变异进行检测, 最后得出地中海贫血型别。 图 4显示了测序中 引物设计的思路。
实验流程:
1. SNP&InDel检测
1.1 构建地中海贫血型别 -SNP&InDel关系数据库
构建方法:
a) 从 hgl9 的对应坐标中提取出 HBA1、 HBA2、 HBB 的序列作为参考序 列。
b) 将现有地中海贫血数据库中已知型别的 HBA1基因、 HBA2基因和 HBB 基因与参考序列比对, 获得参考序列的差异位点, 即 SNP位点和 InDel位点, 得 到每个型别相对于参考序列的 SNP和 InDel关系, 得到地中海贫血型别 SNP-InDel数据库, 地中海贫血已知的型别是从 EBI (EBI: http://www.ebi.ac.uk) 数据库下载的。 同时将数据库中目前并不存在的但发明人在此过程中得到的型 别添加进地中海贫血数据库中。 1.2 得到下机数据
小片段的插入缺失和点突变的检测包括 HBB基因和 HBA基因。 利用引物 (primer) 特异性捕获基因后, 在不同的样品上加上不同的标签 (index) 进行 区分, 随机打断回收 200-500bp的胶片段, 用高通量测序技术对其进行测序。 最 后用 index+primer的特异性将 reads分到不同的样品和基因中。
1.3 SNP和 InDel初步检测
本实施例使用 BWA比对软件和 samtools软件, GATK软件包为例来说明。 通过 BWA比对软件, 将这些序列与特定的参考序列进行序列比对, 经过 aln和 sampe两步, 得到比对结果 *.sampe文件。
为了加快后续程序的运行速度, 使用 samtools软件工具的 view将文件转化 为 *.bam格式, 分别经过 sort、 index得到建立好 index的 bam文件。
然后利用 GATK软件包对 bam文件进行重新校正, bwa比对后同一位点的不 同 reads的比对并不相同, 利用 GATK软件包对 InDel附近的低质量的 SNP重新进 行校正, 更有利准确的 InDel和 SNP。
最后经过 samtools的 pileup得到原始的一致性序列, 利用 PERL语言编写的 程序重新对质量值, 深度进行过滤, 得到可信的 SNP和 InDeL
1.4 SNP和 InDel校正
由于全基因组存在很多重复序列, 例如 AGCCGCC , 如果缺失为最后 3bp 的 GCC, 最后 consensus文件会展现为 A后面缺失 GCC, 所以要对这些重复区域 附近的 InDel重新纠正。
同时对 InDel附近的 SNP进行纠正, 例如 HBB97 C G , 表示 HBB基因 97位 置, ref的 C突变为 G。 HBB 97 C -T, 表示 HBB基因 98位置, 缺失了一个 T。 这 样的变异可以替换成 delins, 即缺失后插入变异。对于这种变异情况要进行修正 为 CT/G, 表示 G替换了 ref的 CT碱基。
1.5 比对型别数据库
整理 SNP和 InDel格式, 以"基因-位置-变异"为识别单位, 比对到型别数据 库, 给出地中海贫血型别, 变异型别不在数据库中的, 标记为 "Unknown"。
1.6 找复杂变异
性别数据库中存在一些复杂的变异情况, 大片段的插入缺失, 大片段的缺 失后插入, bwa比对软件不能更好的展现这些变异情况。 利用 bwa比对算法, 局 部偶联算法中产生的 soft chipping位置, soft chipping是由于 pair end reads在比对 时由于一条 reads可以 Uniq比对, 而第二条出现了很多的 mismatch, 第二条 reads 不会被丢弃, 而是被保留下来, 并标记 s。
利用 S可以找到断裂位点, 这些断裂位点的产生可能是由于大片段插入, 也可以是由于缺失导致的, 可以通过收集 S位点找到大片段的插入, 缺失和缺 失后插入等复杂的变异。
2. HBB基因拷贝数检测
2.1 HBB基因测序
基于引物特异性扩增, 捕获长度为 87bp的 HBB基因片段, 在两端加上标签 ( index) 特异性标记每个样品后, 利用高通量测序技术对 HBB基因测序。
对下机后数据利用 index+primer的特异性将读序 (reads) 分到每个样品和 每个不同的基因中, 分别同时进行检测。
2.2 修正识别序列并计算读序 (reads) 数
根据 SNP和 InDel对识别序列进行修正, 由于存在识别序列区域发生突变, 所以根据同一个样品中 SNP情况对识别序列进行修正。 对于多个 SNP和 InDel 的情况, 给出错误信息, 因为多个 SNP存在连锁关系, 不能确定, InDel会严重 影响序列, 所以这两种情况都输出。 利用识别序列对分到每个样品中的 reads 重新进行计数, 得到 4个绝对数据值, gp : 待测样本 HBB基因 reads数, 待测样 本控制基因 reads数, 正常样本 HBB基因 reads数, 正常样本控制基因 reads数。
2.3 计算比值
在检测 HBB基因中加入正常样本排除了 index对于不同的样品的影响,加入 control基因作为一个绝对参考, 排除不同的样品取样的绝对数量的不同。 比值 =待测样本 HBB基因 reads数 /待测样本控制基因 reads数: 正常样本 HBB基因 reads数 /正常样本控制基因 reads数。
3. HBA基因拷贝数变异检测
3.1 HBA1基因和 HBA2基因测序
HBA基因中 HBA1和 HBA2非常相似, 所以只设计一对引物将 HBA1, HBA2 同时捕获。 捕获长度为 128bp的 HBA基因片段, 在两端加上 index特异性标记每 个样品后利用 hiseq测序技术对 HBA基因测序。 对下机后数据利用 index+primer 的特异性将 reads分到每个样品和每个不同的基因中, 分别同时进行检测。
3.2 修正识别序列并统计 reads数
修正识别序列与 HBB基因部分一致。从 HBA中利用 HBA1与 HBA2的差异序 列将 HBA分为 HBA1和 HBA2。 按照本发明 HBA基因拷贝数检测的流程, 需要统 计 6个原始 reads数,分别是正常样本 HBA1基因 reads,正常样本 HBA2基因 reads, 正常样本 control基因 reads , 待测样本 HBA1基因 reads数, 待测样本 HBA2基因 reads数, 待测样本 control基因 reads数。
3.3 计算比值
HBA部分比较复杂, 需要计算 3个比值, 定义为 Rl, R2, R3。
R 1 =待测样本 HB A 1基因 reads数 /待测样本 control基因 reads数: 正常样本 HBA 1基因 reads数 /正常样本 control基因 reads数。
2=待测样本 HB A2基因 reads数 I待测样本 control基因 reads数: 正常样本 HBA2基因 reads数 /正常样本 control基因 reads数。
R3=待测样本 HB A2基因 reads数 I待测样本 HB A 1基因 reads数: 正常样本 HBA2基因 reads数 /正常样本 HBA 1基因 reads数。
R1和 R2用来得到 HBA1基因拷贝数绝对数值和 HBA2基因拷贝数绝对数值, 最后利用 R3进行检测。
3.4 判定 CNV型别
首先利用类似聚类的方法, 对一个文库下所有的比值即 Rl, R2, R3进行 大概聚类, 判定规则为距离哪个值最近就归于哪类。 然后求出每个类别中的方 差和偏离度。 偏离度表针的是由于实验误差导致的理论值的一个偏差。 然后综 合考虑每个类别有大于 15个样品, 用于构建出实验对整个文库的影响, 包括方 差和偏离度。 计算公式 (Xi - ave(X))/sqrt(var(X))也称为数据值的标准化。 选取 每个 index 的 Rl, R2, R3的最短距离作为最佳选项。 同时利用贝叶斯思想, 对 属于正常样本的距离进行了减缩。
4. 整理所有变异
将 SNP, InDel , HBB基因拷贝数变异, HBA拷贝数变异情况以 index作 为唯一标记, 将他们整合到一起, 同时整合一些质控信息, 例如 CNV primer 区域的变异, reads绝对数据值过少等信息, 最后给出完整的每个样品的变异信 息。
在本发明提及的所有文献都在本申请中引用作为参考, 就如同每一篇文献 被单独引用作为参考那样。此外应理解,在阅读了本发明的上述讲授内容之后, 本领域技术人员可以对本发明作各种改动或修改, 这些等价形式同样落于本申 请所附权利要求书所限定的范围。

Claims

权 利 要 求
1. 一种检测待测样本 HBB基因拷贝数变异的方法,其特征在于,包括步骤:
(1) 获得待测样本的 HBB基因序列, 所述序列与样品和 /或基因相对应;
(2) 对步骤 (1)获得的读序进行修正和 /或统计, 获得下述绝对数据值: 待测样本 HBB基因的读序数、待测样本对照基因的读序数、 正常样本 HBB 基因的读序数、 正常样本对照基因的读序数;
(3) 计算 A值,
待测样本 HBB基因的读序数 /待测样本对照基因的读序数
Ait= ;
正常样本 ΗΒΒ基因的读序数 /正常样本对照基因的读序数
(4) 根据 Α值, 获得待测样本 HBB基因拷贝数变异的检测结果。
2. 一种检测待测样本 HBA基因拷贝数变异的方法, 其特征在于, 包括步 骤:
(1) 获得待测样本的 HBA1基因序列和 HBA2基因序列, 所述的序列与样品 和 /或基因相对应;
(2) 对步骤 (1)获得的读序进行修正和 /或统计, 获得下述绝对数据值: 正常样 本 HBA1基因的读序数, 正常样本 HBA2基因的读序数, 正常样本对照基因的读 序数, 待测样本 HBA1基因的读序数, 待测样本 HBA2基因的读序数, 待测样本 对照基因的读序数;
(3) 计算下述的比值:
待测样本 HBA1基因的读序数 /待测样本对照基因的读序数
R1值 =
正常样本 HBA1基因的读序数 /正常样本对照基因的读序数 待测样本 HBA2基因的读序数 /待测样本对照基因的读序数 正常样本 HBA2基因的读序数 /正常样本对照基因的读序数 待测样本 HBA2基因的读序数 /待测样本 HBA1基因的读序数
1 3值= ; 正常样本 HBA2基因的读序数 /正常样本 HBA1基因的读序数 (4) 根据 Rl值、 R2值和 R3值, 获得待测样本 HBA基因拷贝数变异检测结 果。
3. 一种检测待测样本 HBB基因拷贝数变异的检测单元, 其特征在于, 所 述单元包括模块:
(1) 序列获取模块, 用于获得待测样本的 HBB基因序列, 所述序列与样品和 /或基因相对应;
(2) 读序修正和 /或统计模块, 用于对模块 (1)获得的读序进行修正和 /或统计;
(3) 比值计算模块, 用于计算 A值; 和
(4) 输出模块, 用于输出待测样本 HBB基因拷贝数变异的信息。
4. 一种检测待测样本 HBA基因拷贝数变异的检测单元, 其特征在于, 所 述单元包括模块:
(1) 序列获取模块, 用于获得待测样本的 HBA1基因序列和 HBA2基因序列, 所述的序列与样品和 /或基因相对应;
(2) 读序修正和 /或统计模块, 用于对模块 (1)获得的序列读序进行修正和 /或 统计;
(3) 比值计算模块, 用于计算 R1值、 R2值、 R3值;
(4) 型别判断模块, 用于判断 HBA基因拷贝数变异的型别; 和
(5) 输出模块, 用于输出待测样本 HBA基因拷贝数变异信息。
5. 一种检测地中海贫血的分型系统, 其特征在于, 包括:
(1) 权利要求 3所述的检测待测样本 HBB基因拷贝数变异的检测单元; 和
(2) 权利要求 4所述的检测待测样本 HBA基因拷贝数变异的检测单元。
6. 如权利要求 5所述的系统, 其特征在于, 还包括:
(3) SNP和 /或 InDel的检测单元。
7. 一种检测待测样本 SNP和 Indel变异的方法, 其特征在于, 包括步骤:
(1) 获得待测样本的基因测序数据;
(2) 将 (1)中的测序数据与参考序列进行比对, 获得相对于参考序列的 SNP 和 InDel连锁关系数据;
(3) 根据 (2)中所述的连锁关系数据确定待测样本 SNP和 Indel变异信息。
8. 一种对待测样本进行地中海分型的方法, 其特征在于, 包括步骤:
(1) 获得待测样本 SNP和 InDel变异数据, 以及 HBB基因、 HBA基因拷贝数 变异数据;
(2) 结合步骤 (1)的所有变异数据, 获得待测样本的完整变异情况, 从而对待 测样本进行地中海分型。
9. 一种构建地中海贫血型别 SNP和 InDel关系数据库的方法, 其特征在于, 包括步骤:
(1) 将现有地中海贫血数据库中已知型别的 HBA1基因、 HBA2基因和 HBB 基因与作为参考序列的已知正常无病基因序列进行比对, 获得差异位点, 所述 的差异位点为 SNP和 InDel位点; 和
(2) 获得每个型别相对于参考序列的 SNP和 InDel关系, 得到地中海贫血型 别 SNP-InDel数据库。
10. 一种 SNP和 Indel数据库, 其特征在于, 是用权利要求 9所述的方法构建 的。
PCT/CN2012/087645 2012-08-10 2012-12-27 一种地中海贫血的分型方法及其应用 Ceased WO2014023076A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201210284823.2 2012-08-10
CN201210284823 2012-08-10

Publications (1)

Publication Number Publication Date
WO2014023076A1 true WO2014023076A1 (zh) 2014-02-13

Family

ID=50067415

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2012/087645 Ceased WO2014023076A1 (zh) 2012-08-10 2012-12-27 一种地中海贫血的分型方法及其应用

Country Status (1)

Country Link
WO (1) WO2014023076A1 (zh)

Cited By (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106591441A (zh) * 2016-12-02 2017-04-26 深圳市易基因科技有限公司 基于全基因捕获测序的α和/或β‑地中海贫血突变的检测探针、方法、芯片及应用
CN106636344A (zh) * 2016-10-28 2017-05-10 上海阅尔基因技术有限公司 一种基于二代高通量测序技术的地中海贫血症的基因检测试剂盒
CN106755329A (zh) * 2016-11-28 2017-05-31 广西壮族自治区妇幼保健院 基于二代测序技术检测α和β地中海贫血点突变的方法及试剂盒
CN109280701A (zh) * 2017-07-21 2019-01-29 深圳华大基因股份有限公司 用于地中海贫血检测的探针、基因芯片及制备方法和应用
CN110564843A (zh) * 2019-10-12 2019-12-13 广西安仁欣生物科技有限公司 一种用于地中海贫血突变型与缺失型基因检测的引物组、试剂盒及其应用
CN110942806A (zh) * 2018-09-25 2020-03-31 深圳华大法医科技有限公司 一种血型基因分型方法和装置及存储介质
CN111326211A (zh) * 2020-01-07 2020-06-23 深圳市早知道科技有限公司 一种检测地中海贫血基因变异的方法及检测装置
CN112176051A (zh) * 2019-07-05 2021-01-05 右江民族医学院 缺失型alpha-地中海贫血三重实时荧光定量PCR检测引物、探针和试剂盒
CN115831222A (zh) * 2022-12-20 2023-03-21 北京希望组生物科技有限公司 一种基于三代测序的全基因组结构变异鉴定方法
CN117012274A (zh) * 2023-10-07 2023-11-07 北京智因东方转化医学研究中心有限公司 基于高通量测序识别基因缺失的装置
WO2024010812A3 (en) * 2022-07-07 2024-02-15 Illumina Software, Inc. Methods and systems for determining copy number variant genotypes

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
KITTS, A. ET AL.: "The Single Nucleotide Polymorphism Database (dbSNP) of Nucleotide Sequence Variation.", THE NCBI HANDBOOK, 2 February 2011 (2011-02-02) *
SATHIRAPONGSASUTI, J. F. ET AL.: "Exome sequencing-based copy-number variation and loss of heterozygosity detection: ExomeCNV.", BIOINFORMATICS., vol. 27, no. 19, 9 August 2011 (2011-08-09), pages 2648 - 2654 *
YANG, XU ET AL.: "New-generation high-throughput technologies based 'omics' research strategy in human disease.", HEREDITAS., vol. 33, no. 8, August 2011 (2011-08-01), pages 829 - 846 *

Cited By (15)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106636344A (zh) * 2016-10-28 2017-05-10 上海阅尔基因技术有限公司 一种基于二代高通量测序技术的地中海贫血症的基因检测试剂盒
CN106755329A (zh) * 2016-11-28 2017-05-31 广西壮族自治区妇幼保健院 基于二代测序技术检测α和β地中海贫血点突变的方法及试剂盒
CN106591441A (zh) * 2016-12-02 2017-04-26 深圳市易基因科技有限公司 基于全基因捕获测序的α和/或β‑地中海贫血突变的检测探针、方法、芯片及应用
CN106591441B (zh) * 2016-12-02 2021-07-09 王君文 基于全基因捕获测序的α和/或β-地中海贫血突变的检测探针、方法、芯片及应用
CN109280701A (zh) * 2017-07-21 2019-01-29 深圳华大基因股份有限公司 用于地中海贫血检测的探针、基因芯片及制备方法和应用
CN110942806A (zh) * 2018-09-25 2020-03-31 深圳华大法医科技有限公司 一种血型基因分型方法和装置及存储介质
CN112176051A (zh) * 2019-07-05 2021-01-05 右江民族医学院 缺失型alpha-地中海贫血三重实时荧光定量PCR检测引物、探针和试剂盒
CN110564843B (zh) * 2019-10-12 2023-09-19 广西安仁欣生物科技有限公司 一种用于地中海贫血突变型与缺失型基因检测的引物组、试剂盒及其应用
CN110564843A (zh) * 2019-10-12 2019-12-13 广西安仁欣生物科技有限公司 一种用于地中海贫血突变型与缺失型基因检测的引物组、试剂盒及其应用
CN111326211A (zh) * 2020-01-07 2020-06-23 深圳市早知道科技有限公司 一种检测地中海贫血基因变异的方法及检测装置
CN111326211B (zh) * 2020-01-07 2023-12-19 深圳市早知道科技有限公司 一种检测地中海贫血基因变异的方法及检测装置
WO2024010812A3 (en) * 2022-07-07 2024-02-15 Illumina Software, Inc. Methods and systems for determining copy number variant genotypes
CN115831222A (zh) * 2022-12-20 2023-03-21 北京希望组生物科技有限公司 一种基于三代测序的全基因组结构变异鉴定方法
CN117012274A (zh) * 2023-10-07 2023-11-07 北京智因东方转化医学研究中心有限公司 基于高通量测序识别基因缺失的装置
CN117012274B (zh) * 2023-10-07 2024-01-16 北京智因东方转化医学研究中心有限公司 基于高通量测序识别基因缺失的装置

Similar Documents

Publication Publication Date Title
Yin et al. Challenges in the application of NGS in the clinical laboratory
WO2014023076A1 (zh) 一种地中海贫血的分型方法及其应用
US20240304280A1 (en) Validation methods and systems for sequence variant calls
Cipriani et al. A set of microsatellite markers with long core repeat optimized for grape (Vitis spp.) genotyping
ES2357549T3 (es) Estrategias para la identificación y detección de alto rendimiento de polimorfismos.
CN106591441B (zh) 基于全基因捕获测序的α和/或β-地中海贫血突变的检测探针、方法、芯片及应用
Delahunty et al. Testing the feasibility of DNA typing for human identification by PCR and an oligonucleotide ligation assay
Ding et al. Five-color-based high-information-content fingerprinting of bacterial artificial chromosome clones using type IIS restriction endonucleases
US20130337447A1 (en) Methods and compositions for evaluating genetic markers
WO2014116729A2 (en) Haplotying of hla loci with ultra-deep shotgun sequencing
US11293067B2 (en) Method for genotyping Mycobacterium tuberculosis
CN105441432A (zh) 组合物及其在序列测定和变异检测中的用途
CN110846429B (zh) 一种玉米全基因组InDel芯片及其应用
Claes et al. Dealing with pseudogenes in molecular diagnostics in the next generation sequencing era
Jian et al. A narrative review of single-nucleotide polymorphism detection methods and their application in studies of Staphylococcus aureus
CN102618549A (zh) Ncstn突变型基因、其鉴定方法和工具
US12562237B2 (en) Methods and systems for detection and phasing of complex genetic variants
CN110904220A (zh) 检测cyp2d6基因多态性和拷贝数的组合物、试剂盒及方法
Zascavage et al. Deep-sequencing technologies and potential applications in forensic DNA testing
Berg et al. Pyrosequencing™ technology and the need for versatile solutions in molecular clinical research
CN105154543A (zh) 一种用于生物样本核酸检测的质控方法
CN104769129B (zh) 一种主要组织相容性复合体mhc分型方法及其应用
US20130143746A1 (en) Method for detecting gene region features based on inter-alu polymerase chain reaction
US20240294982A1 (en) Method, kit and cartridge for detecting nucleic acid molecule
WO2003025198A2 (en) Regulatory single nucleotide polymorphisms and methods therefor

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 12882628

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 15/06/2015)

122 Ep: pct application non-entry in european phase

Ref document number: 12882628

Country of ref document: EP

Kind code of ref document: A1