WO2019014936A1 - 用于检测低频突变的靶向富集方法和试剂盒 - Google Patents
用于检测低频突变的靶向富集方法和试剂盒 Download PDFInfo
- Publication number
- WO2019014936A1 WO2019014936A1 PCT/CN2017/093914 CN2017093914W WO2019014936A1 WO 2019014936 A1 WO2019014936 A1 WO 2019014936A1 CN 2017093914 W CN2017093914 W CN 2017093914W WO 2019014936 A1 WO2019014936 A1 WO 2019014936A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- primer
- sequencing
- universal primer
- pcr amplification
- universal
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
Definitions
- the invention relates to the technical field of molecular biology, in particular to a targeted enrichment method and a kit for detecting low frequency mutations.
- NGS second-generation sequencing
- target region targeted enrichment sequencing technology is a technique for enriching the target gene of interest and combining with the second generation sequencing technology to obtain the base information of the target region, which has achieved the purpose of detecting the disease, and is compared with the whole genome sequencing.
- Targeted enrichment sequencing technology can reduce the cost of sequencing, simplify the information analysis process, improve the sequencing depth of the target area, and improve the sensitivity and accuracy of detection.
- probe-based capture enrichment technology using the principle of complementary hybridization of nucleic acid molecules, designing a reverse-complementary oligonucleotide probe according to the target region, then interrupting the genomic DNA, plus the linker for sequencing The probe is hybridized, the unhybridized DNA is eluted, the target DNA fragment is recovered, and the library preparation is performed for DNA sequencing.
- This technique requires a high amount of starting DNA of the sample (generally required to reach micrograms), long experimental operation time, and cumbersome experiment, which is not conducive to automated database construction, low data utilization rate, and high cost.
- the multiplex PCR-based enrichment technique is based on designing primers according to the target region, and then enriching the target region by multiplex PCR, and then performing PCR library sequencing on the PCR product.
- this technique is short in experiment time, it requires complicated primer design work in the early stage, and a lot of cumbersome primer optimization work is needed in the later stage.
- the above two techniques have strict requirements on the amount and integrity of the template, and are incapable of testing samples such as cell free DNA, highly degraded DNA, and paraffin-embedded formaldehyde. Therefore, it is particularly important to develop a simple targeted enrichment method for efficient enrichment of short fragment DNA.
- the invention provides a targeted enrichment method and a kit for detecting low frequency mutations, which can effectively enrich short fragments, single strands, double strands and lost DNA, and can be detected in DNA by combining second generation sequencing technology. Low frequency mutation.
- an embodiment provides a targeted enrichment method for detecting low frequency mutations, comprising:
- First PCR amplification is performed by using an upstream specific primer and a first universal primer for the region where the mutation target site is located, wherein the upstream specific primer is anchored with a target sequence complementary thereto, and the first universal primer is as described above.
- the single nucleotide tail is an anchor anchor
- the first universal primer includes a contiguous single base complementary to the above-described single nucleotide tail at the 1' and 3' ends of the sequencing primer sequence at the 5' end;
- the first PCR-specific primer and the second universal primer perform a second PCR amplification on the first PCR-amplified product, wherein the downstream specific primer is located downstream of the upstream specific primer and has a sequencing primer at the 5' end Sequence 2, the second universal primer described above includes the sequencing primer sequence 1 of the 5' end of the first universal primer described above.
- an embodiment provides a targeted enrichment kit for detecting low frequency mutations, comprising:
- terminal transferase and a single base nucleotide for the addition of a single base single nucleotide to the 3' end of each strand of single-stranded DNA and/or double-stranded DNA under the action of a terminal transferase tail;
- the nucleotide tail is an anchor site
- the first universal primer comprises a single single base complementary to the above-described single nucleotide tail of the sequencing primer sequence 1 and 3' of the 5' end;
- the downstream specific primer is located downstream of the upstream specific primer and has a 5' end Primer primer sequence 2
- the second universal primer described above includes the sequencing primer sequence 1 of the 5' end of the first universal primer described above.
- the method of the invention can effectively enrich short-segment single-stranded DNA, double-stranded DNA and nicked DNA with high stencil utilization; has high detection sensitivity and can detect low-frequency mutations as low as 0.1%; Efficient enrichment of multiple targeted regions and good specificity, uniformity and stability.
- FIG. 1 is a schematic diagram showing the principle of a target enrichment method for detecting low frequency mutations according to an embodiment of the present invention
- FIG. 2 is a schematic diagram showing the principle of a target enrichment method for detecting low frequency mutations according to another embodiment of the present invention
- 3 is a schematic diagram of molecular label correction in an embodiment of the present invention.
- Example 4 is a diagram showing the results of Agilent 2100 quality inspection of a target-enriched library in Example 1 of the present invention
- Figure 5 is a graph showing the results of uniformity of each amplicon of HBB in Example 1 of the present invention.
- Example 6 is a diagram showing the results of Agilent 2100 quality inspection of a targeted enrichment library in Example 2 of the present invention
- Example 8 is a depth distribution diagram of sequencing data of 10 amplicon regions after molecular label correction in Example 2 of the present invention.
- Figure 9 is a graph showing the results of the consistency of mutation detection in Example 2 of the present invention.
- FIG. 1 shows the principle of a targeted enrichment method for detecting low frequency mutations according to an embodiment of the present invention, the method comprising:
- Step I A single base single nucleotide tail is added to the 3' end of each strand of single-stranded DNA and/or double-stranded DNA under the action of a terminal transferase.
- a double-stranded DNA molecule (dsDNA) is subjected to denaturing and melting to obtain single-stranded DNA (ssDNA), and 15-30 of the single-stranded DNA is added to the 3' end of the single-stranded DNA by a terminal transferase.
- Single base C or any of A, T, G
- the composition consists of a 15-30 bp length of a single nucleotide tail. It should be noted that the double-stranded DNA molecule can be directly added to the 3' end of each strand by the terminal transferase without denaturation.
- Step II performing the first PCR amplification with the upstream specific primer and the first universal primer for the region where the mutation target site is located, wherein the upstream specific primer is anchored with the target sequence complementary thereto, and the first universal primer is The single nucleotide tail is an anchor site, and the first universal primer includes a 5' end of the sequencing primer sequence 1 and a 3' end of a contiguous single base complementary to the single nucleotide tail.
- upstream specific primers were designed 25-150 bp before the mutation target site.
- One end uses a single nucleotide tail as an anchor site, and the other end uses an upstream specific primer (USP) complementary target sequence as an anchor site for 10-20 cycles of amplification.
- a primer sequence with a single nucleotide tail as an anchor site is referred to as a "first universal primer” (ie, universal primer 1 in Figure 1) including: a 5'-end sequencing primer sequence 1 (eg, BGISEQ-500, Illumina, or Proton) Such as the sequencing primer sequence of the sequencing platform), the continuous base G in the middle (or any one of T, A, C), The number of consecutive bases C or G is 11-15, and the Tm value ranges from 54 to 70 ° C. Too many numbers are not conducive to PCR amplification; the number of consecutive bases A or T is 25-35, Tm The range of values is 54-62 ° C, too many is not conducive to PCR amplification.
- a 5'-end sequencing primer sequence 1 eg, BGISEQ-500, Illumina, or Proton
- the continuous base G in the middle or any one of T, A, C
- the number of consecutive bases C or G is 11-15, and the Tm value ranges from 54 to 70
- universal primer 1 is 13 consecutive bases G in the middle, and, preferably, a degenerate base H is added last (in other cases it may be V, B or D) ), the degenerate base at this end is used to fix the length of the 3' end product, and the template with C bases of different lengths is subjected to PCR amplification by primers with degenerate bases, and the 3' end is fixed.
- a product of base length C (which in other cases may be A, T or G).
- the product is from the 5' end to the 3' end in turn, the target upstream specific primer sequence, the target region sequence, and the continuous single nucleotide C (in other cases, A, T) Or G), and the final sequencing primer sequence.
- the method of the embodiments of the present invention is applicable to targeted enrichment of various mutation types, including single nucleotide polymorphism (SNP), insertion and deletion (INDEL), and copy number variation (CNV).
- Step III performing a second PCR amplification of the first PCR amplified product with a downstream specific primer and a second universal primer, wherein the downstream specific primer is located downstream of the upstream specific primer and has a sequencing primer at the 5' end Sequence 2, the second universal primer comprises the sequencing primer sequence 1 of the 5' end of the first universal primer.
- downstream specific primers are designed downstream of the upstream specific primer (USP) (eg, 0-10 bp downstream of USP), and downstream specific primers ( The 5' end of DSP) plus sequencing primer sequence 2 (eg, sequencing primer sequences for sequencing platforms such as BGISEQ-500, Illumina or Proton), using downstream specific primers (DSP) with sequencing primer sequence 2 and universal primer 2 ( Also known as "second universal primers") amplification is performed for 15-30 cycles.
- the second universal primer comprises a sequencing primer sequence 1 at the 5' end of the first universal primer.
- a sequencing tag primer can also be introduced, and the 3' end of the sequencing tag primer includes the sequencing primer sequence 2 at the 5' end of the downstream specific primer.
- the downstream specific primer (DSP) and universal primer 2 were amplified, and since the downstream specific primer (DSP) also had the sequencing primer sequence 2 at the 5' end, the product was obtained at both ends.
- Sequencing the primer sequence starting from the second cycle, using the sequencing primer sequence as the anchor site, two universal primers (ie, universal primer 2 and sequencing tag primer (BC)) are simultaneously amplified with downstream specific primers to obtain a target. Sequencing library.
- the advantages of introducing sequencing tag primers are: (1) reducing the specific amplification step, increasing the universal amplification step, facilitating the homogeneity of different amplicon regions; (2) reducing the specific primer-introduced linker sequences, and more The linker sequence was introduced by sequencing tag primers.
- Both single-stranded and double-stranded DNA can be efficiently enriched and have a very high template utilization rate.
- the terminal transferase can add bases to free single-stranded, double-stranded and lost DNA, and the efficiency of adding the template can reach 99% or more, that is, more than 99% of the templates can be added at the 3' end.
- the single nucleotide tail Enrichment of the targeted region by specific primers and a single nucleotide tail requires efficient enrichment by only having a specific primer binding site on the template.
- multiplex PCR amplification is performed using specific primers and anchor primers, and the specific primers may be primers for different targeting regions, and the anchor primers are fixed primer sequences, by mixing specific primers and immobilized targets. Enrichment of multiple regions was performed on the primers, and two-round nested PCR was used to increase the specificity of enrichment, and a multi-region targeted sequencing library was obtained.
- FIG. 2 illustrates the principle of a targeted enrichment method for detecting low frequency mutations in accordance with another embodiment of the present invention, the method comprising:
- the double-stranded DNA molecule is subjected to denaturing and melting to obtain single-stranded DNA, and 5-10 random bases are randomly added to the 3' end of the single-stranded DNA template by a terminal transferase as a molecular tag for labeling the original template.
- the residual dNTPs were removed by magnetic bead purification.
- terminal transferase may also be A, T, Any of G
- a single nucleotide tail is obtained, and then the remaining dCTP (or dATP, dTTP, dGTP) is removed by magnetic bead purification.
- all the free single-stranded, double-stranded and damaged DNA templates have a 5-10 bp molecular tag consisting of four bases and a single base consisting of 15-30 bp in length.
- Polynucleotide tail it should be noted that the double-stranded DNA molecule can be directly added to the 3' end of each strand with a molecular tag and a single nucleotide tail without denaturing.
- Upstream specific primers were designed 25-150 bp before the mutation target site. One end uses a single nucleotide tail as an anchor site, and the other end uses an upstream specific primer (USP) complementary target sequence as an anchor site for 10-20 cycles of amplification.
- a primer sequence with a polynucleotide tail as an anchor site is referred to as a "first universal primer” (ie, universal primer 1 in Figure 2) including: a 5'-end sequencing primer sequence 1 (eg, BGISEQ-500, Illumina, or Proton) Such as the sequencing primer sequence of the sequencing platform), the continuous base G in the middle (or any one of T, A, C), the number of consecutive bases C or G is 11-15, and the Tm value ranges from 54- At 70 ° C, too many numbers are not conducive to PCR amplification; the number of consecutive bases A or T is 25-35, and the Tm value ranges from 54-62 ° C. Too many numbers are not conducive to PCR amplification.
- a 5'-end sequencing primer sequence 1 eg, BGISEQ-500, Illumina, or Proton
- universal primer 1 is 13 consecutive bases G in the middle, and, preferably, a degenerate base H is added last (in other cases it may be V, B or D) ), the degenerate base at this end is used to fix the length of the 3' end product, and the template with C bases of different lengths is subjected to PCR amplification by primers with degenerate bases, and the 3' end is fixed.
- a product of base length C (which in other cases may be A, T or G).
- the product is, from the 5' end to the 3' end, a target upstream specific primer sequence, a target region sequence, a molecular tag consisting of 5-10 random bases, and a continuous mononucleoside.
- Acid C which in other cases can be A, T or G
- the method of the embodiments of the present invention is applicable to targeted enrichment of various mutation types, including single nucleotide polymorphism (SNP), insertion and deletion (INDEL), and copy number variation (CNV). .
- SNP single nucleotide polymorphism
- INDEL insertion and deletion
- CNV copy number variation
- Downstream specific primers are designed downstream of upstream specific primers (USP) and sequencing primers 2 are added to the 5' end of downstream specific primers (DSP) (eg sequencing platforms such as BGISEQ-500, Illumina or Proton)
- DSP downstream specific primers
- the sequencing primer sequence was subjected to 15-30 cycles of amplification using a downstream specific primer (DSP) with sequencing primer sequence 2 and universal primer 2 (also referred to as "second universal primer”).
- the second universal primer comprises a sequencing primer sequence 1 at the 5' end of the first universal primer.
- a sequencing tag primer can also be introduced, and the 3' end of the sequencing tag primer includes the sequencing primer sequence 2 at the 5' end of the downstream specific primer.
- the downstream specific primer (DSP) and universal primer 2 were amplified, and since the downstream specific primer (DSP) also had the sequencing primer sequence 2 at the 5' end, the product was obtained at both ends.
- Sequencing the primer sequence starting from the second cycle, using the sequencing primer sequence as the anchor site, two universal primers (ie, universal primer 2 and sequencing tag primer (BC)) are simultaneously amplified with downstream specific primers to obtain a target. Sequencing library.
- the advantages of introducing sequencing tag primers are: (1) reducing the specific amplification step, increasing the universal amplification step, facilitating the homogeneity of different amplicon regions; (2) reducing the specific primer-introduced linker sequences, and more The linker sequence was introduced by sequencing tag primers.
- the method shown in Fig. 2 has the following advantageous effects in addition to the advantageous effects of the method shown in Fig. 1: a high degree of detection sensitivity, and a low frequency mutation further reduced to 0.1% can be detected. Specifically, a random sequence of 5-10 bp in length consisting of four bases randomly added to the 3' end of the free DNA template by a terminal transferase, the sequence of which can be up to one million species, which can be opposed The initial template is uniquely labeled, and a low frequency mutation further reduced to 0.1% can be detected by molecular marker binding information analysis.
- the targeted sequencing library obtained in the present embodiment is subjected to double-end sequencing, and the specific primer is used for detecting the targeted enrichment region, and the other end is used for reading the molecular tag information for labeling the template.
- the detection of very low frequency mutations is achieved by combining PCR errors and sequencing errors with molecular tags in conjunction with specific data analysis algorithms.
- the plasma source is plasma free DNA of pregnant women at 12 weeks, mother carries CD41/42 ⁇ E (del CTTT) mutation, and father carries CD71/72 (Ins A) mutation.
- CD41/42 ⁇ E del CTTT
- father carries CD71/72 (Ins A) mutation.
- CD71/72 mutation.
- the cfDNA was first heat denatured at 95 ° C for 5 minutes, then rapidly inserted into the ice, followed by an enzymatic reaction.
- the oligonucleotide tail was added to the 3' end of the DNA by a terminal transferase (Terminal Transferase, NEB, USA, Cat. No. M0315S).
- the upstream specific primer pool is shown in Table 3, and the universal primer 1 is shown in Table 4.
- HBB-USP14 CCTTAAACCTGTCTTGTAACCTTGAT SEQ ID NO: 14 HBB-USP15 CAGTAACGGCAGACTTCTCCTC SEQ ID NO: 15 HBB-USP16 GTTGTGTCAGAAGCAAATGTAAGC SEQ ID NO: 16 HBB-USP17 CTGACTTTTATGCCCAGCC SEQ ID NO: 17 HBB-USP18 CTAGGGTGTGGCTCCACAG SEQ ID NO:18 HBB-USP19 CAGCCGTACCTGTCCTTGG SEQ ID NO: 19
- the upstream specific primer pool consisted of a mixture of the equimolar numbers of the primers shown in Table 3.
- the downstream specific primer pool is shown in Table 7, and the universal primer 2 and sequencing primer primers are shown in Table 4.
- the downstream specific primer pool consisted of a mixture of the equimolar numbers of the primers shown in Table 7.
- Targeted enriched libraries were detected with Agilent 2100, and the quality results are shown in Figure 4.
- the BGISEQ-500 sequencing platform was used for sequencing, single-ended 100 bp sequencing, and the obtained data was obtained. After data conversion and quality filtering, the following information analysis is used.
- the obtained data was de-joined, and single-ended sequencing results were obtained.
- the genome (reference genome hg19) was compared, and the mutation of the target site was statistically analyzed by data analysis to obtain information of the target site.
- the results are shown in Table 9-10.
- the depth of each target area is calculated according to the location of the target area, and the uniformity information of the target area is obtained. The result is shown in FIG. 5, and the uniformity is good from the figure.
- Design primers for 10 hotspot mutations related to lung cancer construct a targeted sequencing library for plasma free DNA, and combine high-throughput sequencing and specific information analysis to detect lung cancer-related hotspots.
- the plasma free DNA was Horizon cfDNA standard: 0.1% Multiplex I cfDNA Reference Standard (Cat. No. HD779), and the mutation information was as shown in Table 11, starting at 10 ng, and the experiment was carried out as follows.
- the cfDNA was first heat denatured at 95 ° C for 5 minutes, then rapidly inserted into the ice, followed by an enzymatic reaction. 5-10 random bases added to the 3' end of the DNA by terminal transferase (Terminal Transferase, NEB, USA, Cat. No. M0315S).
- the upstream specific primer pool is shown in Table 15, and the universal primer 1 is shown in Table 4.
- the upstream specific primer pool consisted of a mixture of the equimolar numbers of the primers described in Table 15.
- the amplification system is shown in Table 16 below:
- the downstream specific primer pool is shown in Table 18, and the universal primer 2 and sequencing primer primers are shown in Table 4.
- the downstream specific primer pool consisted of a mixture of equimolar numbers of primers shown in Table 18.
- Targeted enriched libraries were detected using the Agilent 2100 and the results are shown in Figure 6.
- the BGISEQ-500 sequencing platform was used for sequencing, and the double-ended 50 bp sequencing was performed. After the data was converted and mass filtered, the following information was analyzed.
- the obtained data was deligated to obtain double-end sequencing results.
- One end of the sequencing results was used to compare the genome (reference genome hg19), and the other end of the result removed the consecutive G bases, and then 10 base sequences were intercepted from the end.
- Molecular label used to mark the sequence information at the front end; perform basic parameter statistics (Table 20) to compare the proportion of data on the genome, and calculate the depth of each target area according to the target area position, and obtain the target area uniformity information (Figure 7 -8), wherein FIG. 7 shows the depth of sequencing of the original data, and FIG. 8 shows the depth of sequencing after the deduplication of the reads. The results showed good homogeneity.
- the depth of the target site and the ratio of the four bases were counted, and the molecular tag was used to remove the repetition and sequencing errors, PCR errors, and the mutation information of the target site was obtained by a specific information analysis algorithm.
- the results are shown in Table 21.
- Table 22 shows the results of the consistency of the test results.
- Figure 9 shows the mutation detection identity information indicating that the detected mutation information of the target site is consistent with the expectation.
- the mutation ratio detected by our method is 0.08%, 0.10%, 0.10%, 0.15%, 0.12%, 0.11%, 0.15%, and the detected value and the true value are within ⁇ 0.02% error range, indicating The method can accurately detect mutations as low as 0.10%. It can be seen that the molecular marker (UID) correction can remove the repetition and sequencing errors, and reduce the sequencing background from 0.60% to 0.00%. (It should be noted that the error value before the V600E correction is the largest, which means that the error rate of the method can be 0.60% dropped to 0.00%).
- UID molecular marker
Landscapes
- Chemical & Material Sciences (AREA)
- Organic Chemistry (AREA)
- Life Sciences & Earth Sciences (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Microbiology (AREA)
- Immunology (AREA)
- Physics & Mathematics (AREA)
- Molecular Biology (AREA)
- Biotechnology (AREA)
- Biophysics (AREA)
- Analytical Chemistry (AREA)
- Biochemistry (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Genetics & Genomics (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
本发明提供了一种用于检测低频突变的靶向富集方法和试剂盒。该方法包括:在末端转移酶的作用下,在单链DNA和/或双链DNA每条链的3'端加上一段单一碱基的单聚核苷酸尾巴;以上游特异性引物和第一通用引物对突变目标位点所在的区域进行第一次PCR扩增,其中第一通用引物包括5'端的测序引物序列1和3'端的与单聚核苷酸尾巴互补的连续单一碱基;以下游特异性引物和第二通用引物对第一次PCR扩增的产物进行第二次PCR扩增,其中第二通用引物包括第一通用引物的测序引物序列1。
Description
本发明涉及分子生物学技术领域,具体涉及一种用于检测低频突变的靶向富集方法和试剂盒。
二代测序(NGS)技术的发展,为现代基因组学的研究打开了新的局面,然而全基因组测序的成本和分析的复杂程度还是让科研人员倍感困难,尽管二代测序的通量越来越高,而费用越来越低,但它仍不是大多数遗传实验室的可行选择。对于复杂疾病的研究更是如此,这类研究至少需要数百个样本,以实现足够的统计能力。然而,这么多样本的全基因组测序,无论从成本考虑,还是从数据分析考虑,都是相对困难的。
目标区域靶向富集测序(Target region sequencing)技术的出现,缓解了上述问题。目标区域靶向富集测序技术,是对感兴趣的目标基因进行富集并结合二代测序技术进行测序的技术,得到目标区域的碱基信息,已达到检测疾病的目的,相对于全基因组测序,靶向富集测序技术可以降低测序成本,简化信息分析流程,提高目标区域测序深度,提高检测的灵敏性和准确性。
目前市场上针对目标区域进行富集的技术主要有两种,一种是基于探针的捕获富集技术,另一种是基于多重PCR的富集技术。基于探针的捕获富集技术,利用核酸分子碱基互补杂交的原理,根据目标区域设计反向互补的寡聚核苷酸探针,然后打断基因组DNA,加上用于测序的接头后与探针杂交,洗脱未杂交上的DNA,回收目标DNA片段,再进行文库制备进行DNA测序。这种技术需要样本起始DNA量高(一般需要达到微克),实验操作时间长,实验繁琐,不利于自动化建库,数据利用率低,成本高。基于多重PCR的富集技术,是根据目标区域设计引物,然后通过多重PCR富集目标区域,再对PCR产物进行文库制备进行DNA测序。这种技术虽然实验时间短,但前期需要进行复杂的引物设计工作,并且在后期需要进行大量繁琐的引物优化工作。此外,上述两种技术对模板的量和完整度有严格的要求,对细胞游离DNA、高度降解DNA、石蜡包埋甲醛固定的医学等样本无能为力。因此开发一种简单的针对短片段DNA能够进行有效富集的靶向富集方法显得尤为重要。
发明内容
本发明提供一种用于检测低频突变的靶向富集方法和试剂盒,能够对短片段、单链、双链及损失的DNA进行有效富集,结合二代测序技术可以检测发生在DNA中的低频突变。
根据第一方面,一种实施例中提供一种用于检测低频突变的靶向富集方法,包括:
在末端转移酶的作用下,在单链DNA和/或双链DNA每条链的3’端加上一段单一碱基的单聚核苷酸尾巴;
以上游特异性引物和第一通用引物对突变目标位点所在的区域进行第一次PCR扩增,其中上述上游特异性引物以与其互补的目标序列为锚定位点,上述第一通用引物以上述单聚核苷酸尾巴为锚定位点,且上述第一通用引物包括5’端的测序引物序列1和3’端的与上述单聚核苷酸尾巴互补的连续单一碱基;
以下游特异性引物和第二通用引物对上述第一次PCR扩增的产物进行第二次PCR扩增,其中上述下游特异性引物位于上述上游特异性引物的下游且5’端带有测序引物序列2,上述第二通用引物包括上述第一通用引物5’端的测序引物序列1。
根据第二方面,一种实施例中提供一种用于检测低频突变的靶向富集试剂盒,包括:
末端转移酶和单一碱基核苷酸,用于在末端转移酶的作用下,在单链DNA和/或双链DNA每条链的3’端加上一段单一碱基的单聚核苷酸尾巴;
上游特异性引物和第一通用引物,用于对突变目标位点所在的区域进行第一次PCR扩增,其中上述上游特异性引物以与其互补的目标序列为锚定位点,上述第一通用引物以上述单聚
核苷酸尾巴为锚定位点,且上述第一通用引物包括5’端的测序引物序列1和3’端的与上述单聚核苷酸尾巴互补的连续单一碱基;
下游特异性引物和第二通用引物,用于对上述第一次PCR扩增的产物进行第二次PCR扩增,其中上述下游特异性引物位于上述上游特异性引物的下游且5’端带有测序引物序列2,上述第二通用引物包括上述第一通用引物5’端的测序引物序列1。
本发明的方法,能够有效富集短片段单链DNA、双链DNA以及带缺口损伤的DNA且具有极高的模板利用率;具有高度的检测灵敏性,可以检测低至0.1%的低频突变;能够一次对多个靶向区域进行有效富集,并且能够保证很好的特异性、均一性和稳定性。
图1为本发明一种实施例的用于检测低频突变的靶向富集方法的原理示意图;
图2为本发明另一种实施例的用于检测低频突变的靶向富集方法的原理示意图;
图3为本发明一种实施例中的分子标签校正示意图;
图4为本发明实施例1中靶向富集文库的安捷伦2100质检结果图;
图5为本发明实施例1中HBB各扩增子均一性结果图;
图6为本发明实施例2中靶向富集文库的安捷伦2100质检结果图;
图7为本发明实施例2中10个扩增子区域测序数据深度分布图;
图8为本发明实施例2中分子标签校正后的10个扩增子区域测序数据深度分布图;
图9为本发明实施例2中突变检测一致性结果图。
下面通过具体实施方式结合附图对本发明作进一步详细说明。在以下的实施方式中,很多细节描述是为了使得本发明能被更好的理解。然而,本领域技术人员可以毫不费力的认识到,其中部分特征在不同情况下是可以省略的,或者可以由其他原材料、方法所替代。在某些情况下,本发明相关的一些操作并没有在说明书中显示或者描述,这是为了避免本发明的核心部分被过多的描述所淹没,而对于本领域技术人员而言,详细描述这些相关操作并不是必要的,他们根据说明书中的描述以及本领域的一般技术知识即可完整了解相关操作。
本文中为部件所编序号本身,例如“第一”、“第二”等,仅用于区分所描述的对象,不具有任何顺序或技术含义。
图1示出了本发明一种实施例的用于检测低频突变的靶向富集方法的原理,该方法包括:
步骤I:在末端转移酶的作用下,在单链DNA和/或双链DNA每条链的3’端加上一段单一碱基的单聚核苷酸尾巴。
具体地,在图1所示的实施例中,双链DNA分子(dsDNA)通过变性解链后得到单链DNA(ssDNA),通过末端转移酶在单链DNA的3’端加入15-30个单一碱基C(或者A、T、G中的任意一种),得到一个单聚核苷酸尾巴,此时所有单链、双链及损伤DNA模板的3’端都有一段由单一碱基组成的长度为15-30bp的单聚核苷酸尾巴。需要说明的是,双链DNA分子不经变性也可以在末端转移酶的作用下直接在每条链的3’端加入单聚核苷酸尾巴。
步骤II:以上游特异性引物和第一通用引物对突变目标位点所在的区域进行第一次PCR扩增,其中上游特异性引物以与其互补的目标序列为锚定位点,第一通用引物以单聚核苷酸尾巴为锚定位点,且第一通用引物包括5’端的测序引物序列1和3’端的与单聚核苷酸尾巴互补的连续单一碱基。
具体地,在图1所示的实施例中,在突变目标位点前25-150bp设计上游特异性引物(USP)。一端以单聚核苷酸尾巴为锚定位点,另一端以上游特异性引物(USP)互补的目标序列为锚定位点进行10-20个循环的扩增。以单聚核苷酸尾巴为锚定位点的引物序列称为“第一通用引物”(即图1中的通用引物1)包括:5’端的测序引物序列1(例如BGISEQ-500、Illumina或Proton等测序平台的测序引物序列),中间的连续碱基G(或者T、A、C中的任意一种),
连续碱基C或G的个数为11-15个,Tm值范围为54-70℃,个数太多不利于PCR扩增;连续碱基A或T的个数为25-35个,Tm值范围为54-62℃,个数太多不利于PCR扩增。在图1所示的优选的实施例中,通用引物1中间是13个连续碱基G,并且,优选地,最后加上一个简并碱基H(在其他情况下可以是V、B或D),这个末端的简并碱基用来固定3’端产物的长度,将末端为不同长度C碱基的模板经过带有简并碱基的引物进行PCR扩增以后,得到3’端为固定长度C(在其他情况下可以是A、T或G)碱基的产物。得到的PCR产物经过磁珠纯化后,产物从5’端到3’端依次是目标上游特异性引物序列、目标区域序列、连续的单聚核苷酸C(在其他情况下可以是A、T或G),以及最后的测序引物序列。需要说明的是,本发明实施例的方法,适用于各种突变类型的靶向富集,包括单核苷酸多态性(SNP)、插入和删除(INDEL)以及拷贝数变异(CNV)等。
步骤III:以下游特异性引物和第二通用引物对第一次PCR扩增的产物进行第二次PCR扩增,其中下游特异性引物位于上游特异性引物的下游且5’端带有测序引物序列2,第二通用引物包括第一通用引物5’端的测序引物序列1。
具体地,在图1所示的实施例中,在上游特异性引物(USP)的下游(例如USP的下游0-10bp的区域)设计下游特异性引物(DSP),并在下游特异性引物(DSP)的5’端加上测序引物序列2(例如BGISEQ-500、Illumina或Proton等测序平台的测序引物序列),用带有测序引物序列2的下游特异性引物(DSP)和通用引物2(也称为“第二通用引物”)进行15-30个循环的扩增。所述第二通用引物包括所述第一通用引物5’端的测序引物序列1。
作为优选技术方案,还可以引入一个测序标签引物(BC),所述测序标签引物3’端包括所述下游特异性引物5’端的测序引物序列2。在PCR的第一个循环中,下游特异性引物(DSP)和通用引物2进行扩增,由于下游特异性引物(DSP)5’端也有测序引物序列2,因此得到产物的两端都带有测序引物序列;从第二个循环开始,以测序引物序列为锚定位点,两个通用引物(即通用引物2和测序标签引物(BC))与下游特异性引物同时进行扩增,得到靶向测序文库。引入测序标签引物的好处在于:(1)减少特异性扩增步骤,增加通用扩增步骤,有利于不同扩增子区域的均一性;(2)减少特异性引物引入的接头序列,更多的接头序列由测序标签引物引入。
本实施例的方法具有以下有益效果:
(1)对单链和双链DNA都能有效富集且具有极高的模板利用率。具体而言,末端转移酶对游离单链、双链及损失DNA都能够进行碱基的添加,且对模板的添加效率能够达到99%以上,即99%以上的模板都能在3’端加上单聚核苷酸尾巴。通过特异性引物和单聚核苷酸尾巴对靶向区域进行富集,只需要模板上有特异性引物结合位点就能够产生有效富集。对于大小为165bp左右的游离DNA,引物所占范围在25bp左右,对模板的利用率可以达到(165-25bp)/165bp=85%,因此对于一些高度片段化单链、双链以及损失DNA都能够达到很好的富集效果。
(2)能够对多个靶向区域进行富集。具体而言,采用特异性引物和锚定引物进行多重PCR扩增,特异性引物可以是针对不同靶向区域的引物,锚定引物是固定的引物序列,通过混合的特异性引物和固定的靶向引物进行多个区域的富集,并且采用两轮巢式PCR可以提高富集的特异性,得到多个区域的靶向测序文库。
图2示出了本发明另一种实施例的用于检测低频突变的靶向富集方法的原理,该方法包括:
双链DNA分子通过变性解链后得到单链DNA,通过末端转移酶对单链DNA模板的3’端随机加入5-10个随机碱基,作为分子标签,用于标记原始模板。其中加上的分子标签的种类达到百万种(N=54+64+74+…104=1397760),远远超过模板的个数,因此每条模板大于99.9%概率都会加上一个唯一的由A、T、C、G四种碱基组成的5-10bp的分子标签,用以对原始模板进行标记。通过磁珠纯化去除掉残留的dNTP。
然后再通过末端转移酶在分子标签的基础上再加入15-30个单一碱基C(也可以是A、T、
G中的任意一种),得到一个单聚核苷酸尾巴,然后再用磁珠纯化去除掉残留的dCTP(或dATP、dTTP、dGTP)。此时所有的游离单链、双链及损伤DNA模板的3’端都带有一个由四种碱基组成的5-10bp的分子标签和由一种碱基组成的长度为15-30bp的单聚核苷酸尾巴。需要说明的是,双链DNA分子不经变性也可以在末端转移酶的作用下直接在每条链的3’端加入分子标签和单聚核苷酸尾巴。
在突变目标位点前25-150bp设计上游特异性引物(USP)。一端以单聚核苷酸尾巴为锚定位点,另一端以上游特异性引物(USP)互补的目标序列为锚定位点进行10-20个循环的扩增。以单聚核苷酸尾巴为锚定位点的引物序列称为“第一通用引物”(即图2中的通用引物1)包括:5’端的测序引物序列1(例如BGISEQ-500、Illumina或Proton等测序平台的测序引物序列),中间的连续碱基G(或者T、A、C中的任意一种),连续碱基C或G的个数为11-15个,Tm值范围为54-70℃,个数太多不利于PCR扩增;连续碱基A或T的个数为25-35个,Tm值范围为54-62℃,个数太多不利于PCR扩增。在图2所示的优选的实施例中,通用引物1中间是13个连续碱基G,并且,优选地,最后加上一个简并碱基H(在其他情况下可以是V、B或D),这个末端的简并碱基用来固定3’端产物的长度,将末端为不同长度C碱基的模板经过带有简并碱基的引物进行PCR扩增以后,得到3’端为固定长度C(在其他情况下可以是A、T或G)碱基的产物。得到的PCR产物经过磁珠纯化后,产物从5’端到3’端依次是目标上游特异性引物序列、目标区域序列、5-10个随机碱基组成的分子标签、连续的单聚核苷酸C(在其他情况下可以是A、T或G),以及最后的测序引物序列。需要说明的是,本发明实施例的方法,适用于各种突变类型的靶向富集,包括单核苷酸多态性(SNP)、插入和删除(INDEL)以及拷贝数变异(CNV)等。
在上游特异性引物(USP)的下游设计下游特异性引物(DSP),并在下游特异性引物(DSP)的5’端加上测序引物序列2(例如BGISEQ-500、Illumina或Proton等测序平台的测序引物序列),用带有测序引物序列2的下游特异性引物(DSP)和通用引物2(也称为“第二通用引物”)进行15-30个循环的扩增。所述第二通用引物包括所述第一通用引物5’端的测序引物序列1。
作为优选技术方案,还可以引入一个测序标签引物(BC),所述测序标签引物3’端包括所述下游特异性引物5’端的测序引物序列2。在PCR的第一个循环中,下游特异性引物(DSP)和通用引物2进行扩增,由于下游特异性引物(DSP)5’端也有测序引物序列2,因此得到产物的两端都带有测序引物序列;从第二个循环开始,以测序引物序列为锚定位点,两个通用引物(即通用引物2和测序标签引物(BC))与下游特异性引物同时进行扩增,得到靶向测序文库。引入测序标签引物的好处在于:(1)减少特异性扩增步骤,增加通用扩增步骤,有利于不同扩增子区域的均一性;(2)减少特异性引物引入的接头序列,更多的接头序列由测序标签引物引入。
图2所示的方法除具有图1所示的方法的有益效果以外,还具有以下有益效果:高度的检测灵敏性,可以检测进一步降低至0.1%的低频突变。具体而言,通过末端转移酶在游离DNA模板的3’端随机加上的由四种碱基组成的长度为5-10bp的一段随机序列,该序列的种类可达百万种,可以对起始模板进行唯一标记,通过分子标记结合信息分析方法可以检测进一步降低至0.1%的低频突变。
如图3所示,对本实施例得到的靶向测序文库进行双端测序,特异性引物那一端用来进行靶向富集区域的检测,另一端用来读取分子标签信息用以标记模板,通过结合特定的数据分析算法通过分子标签进行PCR错误和测序错误的去除,实现检测非常低频的突变。
以下通过实施例详细说明本发明的技术方案和效果,应当理解的是,实施例仅是示例性的,不能理解为对本发明保护范围的限制。
实施例1:地贫父源突变检测
针对与beta地贫相关的HBB基因设计19对引物,对常见的beta地贫突变位点进行检测,针对父母携带不同突变类型,检测孕妇血浆游离DNA中胎儿是否携带父源突变来达到排除性诊断。
血浆来源是12周孕妇血浆游离DNA,母亲携带CD41/42βE(del CTTT)突变,父亲携带CD71/72(Ins A)突变,通过对血浆游离DNA捕获建库后测序,检测胎儿是否携带父亲来源的(CD71/72)突变。
实验步骤:
1.寡聚核苷酸尾巴的添加
先将cfDNA在95℃热变性5分钟,然后迅速插冰上,然后再进行酶反应。通过末端转移酶(美国NEB公司的Terminal Transferase,货号M0315S)在DNA的3’端加上寡聚核苷酸尾巴。
反应体系如表1所示:
表1
37℃孵育30分钟,然后加入10μl浓度为0.5M的EDTA终止反应。加入1.8倍体积的Agencourt AMPure XP磁珠(美国贝克曼库尔特有限公司)108μl,按照说明书进行纯化,纯化后用20μl蒸馏水溶解DNA。
2.第一轮PCR扩增
PCR反应体系如表2所示:
表2
上游特异性引物池如表3所示,通用引物1如表4所示。
表3
| 引物名称 | 引物序列(5-3’) | 引物编号 |
| HBB-USP1 | TGAGAGATGCAGGATAAGCAA | SEQ ID NO:1 |
| HBB-USP2 | GTTGCCAATGTGCATTAGCT | SEQ ID NO:2 |
| HBB-USP3 | TCCCAAGGTTTGAACTAGCTC | SEQ ID NO:3 |
| HBB-USP4 | TTAGGGAACAAAGGAACCTTTAAT | SEQ ID NO:4 |
| HBB-USP5 | GTGGGAGGAAGATAAGAGGTATGA | SEQ ID NO:5 |
| HBB-USP6 | GCTGCTATTAGCAATATGAAACCTC | SEQ ID NO:6 |
| HBB-USP7 | TGATACATTGTATCATTATTGCCCTG | SEQ ID NO:7 |
| HBB-USP8 | TAGTAATGTACTAGGCAGACTGTGT | SEQ ID NO:8 |
| HBB-USP9 | TCATTCGTCTGTTTCCCATTC | SEQ ID NO:9 |
| HBB-USP10 | CCTTCCTATGACATGAACTTAACC | SEQ ID NO:10 |
| HBB-USP11 | GCGTCCCATAGACTCACCC | SEQ ID NO:11 |
| HBB-USP12 | CACCGAGCACTTTCTTGCC | SEQ ID NO:12 |
| HBB-USP13 | GAAAATAGACCAATAGGCAGAGAGA | SEQ ID NO:13 |
| HBB-USP14 | CCTTAAACCTGTCTTGTAACCTTGAT | SEQ ID NO:14 |
| HBB-USP15 | CAGTAACGGCAGACTTCTCCTC | SEQ ID NO:15 |
| HBB-USP16 | GTTGTGTCAGAAGCAAATGTAAGC | SEQ ID NO:16 |
| HBB-USP17 | CTGACTTTTATGCCCAGCC | SEQ ID NO:17 |
| HBB-USP18 | CTAGGGTGTGGCTCCACAG | SEQ ID NO:18 |
| HBB-USP19 | CAGCCGTACCTGTCCTTGG | SEQ ID NO:19 |
上游特异性引物池由表3所示的引物等摩尔数混合组成。
表4
扩增体系如下表5所示:
表5
加入1.8倍体积的Agencourt AMPure XP磁珠(美国贝克曼库尔特有限公司)90μl,按照说明书进行纯化,纯化后用20μl蒸馏水溶解DNA。
3.第二轮PCR扩增
PCR反应体系如下表6所示:
表6
下游特异性引物池如表7所示,通用引物2和测序标签引物如表4所示。
表7
下游特异性引物池由表7所示的引物等摩尔数混合组成。
扩增体系如下表8所示:
表8
加入1倍体积的Agencourt AMPure XP磁珠(美国贝克曼库尔特有限公司)50μl,按照说明书进行纯化,纯化后用30μl蒸馏水溶解DNA。
4.文库质检
用安捷伦2100检测靶向富集文库,其质检结果如图4所示。
5.上机测序
质检合格后,采用BGISEQ-500测序平台进行测序,单端100bp测序,得到的下机数据
经过数据转换和质量过滤后采用如下信息分析。
6.信息分析
首先将得到的数据去接头,得到单端测序结果,比对基因组(参考基因组hg19),通过数据分析对目标位点的突变进行统计,得到目标位点的信息。结果如表9-10所示。根据目标区域位置统计各目标区域的深度,得出目标区域均一性信息,结果如图5所示,从图可以看到均一性好。
表9:下机数据统计
| 样本 | 下机数据 | 比对率 | 捕获率 | 0.1X平均深度 |
| 1 | 6789141 | 93.2% | 93.5% | 100% |
表10:Beta地贫检测结果
结论:检测得到胎儿不携带父亲的致病突变,通过sanger验证得到胎儿不携带父亲的致病突变,与本实施例的检测结果相同。可以通过这种方法来达到无创检测胎儿是否beta地贫父源突变。
实施例2:血浆游离DNA低频突变检测
针对肺癌相关的10个热点突变设计引物,对血浆游离DNA构建靶向测序文库,结合高通量测序和特定的信息分析对与肺癌相关的热点区域进行检测。
血浆游离DNA采用的是Horizon公司cfDNA标准品:0.1%Multiplex I cfDNA Reference Standard(货号:HD779),突变信息如表11,起始用量10ng,按照下面进行实验。
表11:Horizon公司cfDNA标准品突变信息
| 基因 | 突变名称 | 突变类型 | 突变频率 |
| BRAF | V600E | c.1799T>A(exon15) | 0.00% |
| cKIT | D816V | c.2447A>T | 0.00% |
| EGFR | G719S | c.2155G>A | 0.00% |
| EGFR | T790M | c.2369C>T | 0.10% |
| EGFR | L858R | c.2573T>G | 0.10% |
| EGFR | ΔE746-A750 | c.2235_2249del15(Deletion) | 0.10% |
| KRAS | G12D | c.35G>A | 0.13% |
| KRAS | G13D | c.38G>A | 0.00% |
| NRAS | Q61K | c.35G>A | 0.13% |
| PIK3CA | E545K | c.35G>A | 0.13% |
| PIK3CA | H1047R | c.35G>A | 0.00% |
实验步骤:
1.分子标签的添加
先将cfDNA在95℃热变性5分钟,然后迅速插冰上,然后再进行酶反应。通过末端转移酶(美国NEB公司的Terminal Transferase,货号M0315S)在DNA的3’端加上的5-10个随机碱基。
反应体系如下表12所示:
表12
37℃孵育30分钟,然后加入10μl浓度为0.5M的EDTA终止反应。加入1.8倍体积的Agencourt AMPure XP磁珠(美国贝克曼库尔特有限公司)108μl,按照说明书进行纯化,纯化后用34μl蒸馏水溶解DNA。
2.寡聚核苷酸尾巴的添加
反应体系如下表13所示:
表13
37℃孵育30分钟,然后加入10μl浓度为0.5M的EDTA终止反应。加入1.8倍体积的Agencourt AMPure XP磁珠(美国贝克曼库尔特有限公司)108μl,按照说明书进行纯化,纯化后用20μl蒸馏水溶解DNA。
3.第一轮PCR扩增
反应体系如下表14所示:
表14
上游特异性引物池如表15所示,通用引物1如表4所示。
表15
上游特异性引物池由表15所述的引物等摩尔数混合组成。
扩增体系如下表16所示:
表16
加入1.8倍体积的Agencourt AMPure XP磁珠(美国贝克曼库尔特有限公司)90μl,按照说明书进行纯化,纯化后用20μl蒸馏水溶解DNA。
4.第二轮PCR扩增:
PCR反应体系如下表17所示:
表17
下游特异性引物池如表18所示,通用引物2和测序标签引物如表4所示。
表18
下游特异性引物池由表18所示的引物等摩尔数混合组成。
扩增体系如下表19所示:
表19
加入1倍体积的Agencourt AMPure XP磁珠(美国贝克曼库尔特有限公司)50μl,按照说明书进行纯化,纯化后用30μl蒸馏水溶解DNA。
5.文库质检
采用安捷伦2100检测靶向富集文库,其结果如图6所示。
6.上机测序
质检合格后,采用BGISEQ-500测序平台进行测序,双端50bp测序,得到的下机数据经过数据转换和质量过滤后采用如下信息分析。
7.信息分析
首先将得到的数据去接头,得到双端测序结果,一端测序结果用来比对基因组(参考基因组hg19),另一端结果去除连续的G碱基后,从去掉一端开始截取10个碱基序列作为分子标签,用来标记前面一端的序列信息;进行基本参数统计(表20)比对到基因组上的数据比例,根据目标区域位置统计各目标区域的深度,得出目标区域均一性信息(图7-8),其中,图7示出了原始数据测序深度,图8示出了读段(reads)去重后的测序深度。结果显示均一性良好。统计目标位点的深度及四碱基比例,通过分子标签去除重复和测序错误、PCR错误,通过特定的信息分析算法得到目标位点的突变信息,结果如表21所示。表22示出了检测结果一致性的结果。图9示出了突变检测一致性信息,表明目标位点的检出的突变信息和预期一致。
表20:下机数据统计
表21:检测结果
表22:检测结果一致性
| 突变位点 | 原始突变 | 校正后的突变 | 标准品突变 |
| V600E | 0.60% | 0.00% | 0.00% |
| D816V | 0.14% | 0.00% | 0.00% |
| G719S | 0.20% | 0.00% | 0.00% |
| T790M | 0.11% | 0.08% | 0.10% |
| L858R | 0.47% | 0.10% | 0.10% |
| ΔE746-A750 | 0.30% | 0.15% | 0.10% |
| G12D | 0.07% | 0.12% | 0.13% |
| G13D | 0.30% | 0.00% | 0.00% |
| Q61K | 0.14% | 0.11% | 0.13% |
| E545K | 0.20% | 0.15% | 0.13% |
| H1047R | 0.11% | 0.00% | 0.00% |
结论:在标准品中,V600E、D816V、G719S、G13D、H1047R均为阴性,突变比例为0.00%,本实施例得到的结果也是阴性,突变比例也为0.00%;在标准品中,T790M、L858R、ΔE746-A750、G12D、Q61K、E545K为阳性,突变比例分别为0.10%、0.10%、0.10%、0.13%、
0.13%、0.13%,我们方法检测到的突变比例分别为0.08%、0.10%、0.10%、0.15%、0.12%、0.11%、0.15%,检测值和真实值在±0.02%误差范围内,表明方法可以准确检测低至0.10%的突变。可以看到通过分子标记(UID)校正可以去除重复和测序错误,将测序背景从0.60%降到0.00%(此处需要说明的是V600E校正前的错误值最大,意即该方法错误率可以从0.60%降到0.00%)。
由图9可以看出,目标位点检出突变信息与预期一致,阴性突变点都未检出,阳性突变点都检出,并且检出频率和预期值相差不大,可以认为该方法在血浆游离DNA中可以检测低至0.10%的突变。
以上应用了具体个例对本发明进行阐述,只是用于帮助理解本发明,并不用以限制本发明。对于本发明所属技术领域的技术人员,依据本发明的思想,还可以做出若干简单推演、变形或替换。
Claims (24)
- 一种用于检测低频突变的靶向富集方法,其特征在于,包括:在末端转移酶的作用下,在单链DNA和/或双链DNA每条链的3’端加上一段单一碱基的单聚核苷酸尾巴;以上游特异性引物和第一通用引物对突变目标位点所在的区域进行第一次PCR扩增,其中所述上游特异性引物以与其互补的目标序列为锚定位点,所述第一通用引物以所述单聚核苷酸尾巴为锚定位点,且所述第一通用引物包括5’端的测序引物序列1和3’端的与所述单聚核苷酸尾巴互补的连续单一碱基;以下游特异性引物和第二通用引物对所述第一次PCR扩增的产物进行第二次PCR扩增,其中所述下游特异性引物位于所述上游特异性引物的下游且5’端带有测序引物序列2,所述第二通用引物包括所述第一通用引物5’端的测序引物序列1。
- 根据权利要求1所述的方法,其特征在于,所述第二次PCR扩增中还加入测序标签引物,所述测序标签引物3’端包括所述下游特异性引物5’端的测序引物序列2。
- 根据权利要求1所述的方法,其特征在于,所述方法还包括:在加所述单聚核苷酸尾巴之前,在末端转移酶的作用下,在单链DNA和/或双链DNA每条链的3’端加上一段随机碱基。
- 根据权利要求3所述的方法,其特征在于,所述随机碱基的长度是5-10个碱基。
- 根据权利要求1所述的方法,其特征在于,所述单聚核苷酸尾巴的长度是15-30个单一碱基。
- 根据权利要求1所述的方法,其特征在于,所述上游特异性引物距离所述突变目标位点25-150bp。
- 根据权利要求1所述的方法,其特征在于,所述第一次PCR扩增进行10-20个循环。
- 根据权利要求1所述的方法,其特征在于,所述连续单一碱基是11-15个连续的C或G碱基,或25-35个连续的A或T碱基。
- 根据权利要求1所述的方法,其特征在于,所述第一通用引物在所述连续单一碱基之后还包括一个简并碱基。
- 根据权利要求9所述的方法,其特征在于,所述简并碱基是H、V、B或D。
- 根据权利要求1所述的方法,其特征在于,所述第二次PCR扩增进行15-30个循环。
- 根据权利要求1所述的方法,其特征在于,所述突变目标位点包括单核苷酸多态性、插入和删除以及拷贝数变异。
- 一种用于检测低频突变的靶向富集试剂盒,其特征在于,包括:末端转移酶和单一碱基核苷酸,用于在末端转移酶的作用下,在单链DNA和/或双链DNA每条链的3’端加上一段单一碱基的单聚核苷酸尾巴;上游特异性引物和第一通用引物,用于对突变目标位点所在的区域进行第一次PCR扩增,其中所述上游特异性引物以与其互补的目标序列为锚定位点,所述第一通用引物以所述单聚核苷酸尾巴为锚定位点,且所述第一通用引物包括5’端的测序引物序列1和3’端的与所述单聚核苷酸尾巴互补的连续单一碱基;下游特异性引物和第二通用引物,用于对所述第一次PCR扩增的产物进行第二次PCR扩增,其中所述下游特异性引物位于所述上游特异性引物的下游且5’端带有测序引物序列2,所述第二通用引物包括所述第一通用引物5’端的测序引物序列1。
- 根据权利要求13所述的试剂盒,其特征在于,所述试剂盒还包括测序标签引物,用于对所述第一次PCR扩增的产物进行第二次PCR扩增,所述测序标签引物3’端包括所述下游特异性引物5’端的测序引物序列2。
- 根据权利要求13所述的试剂盒,其特征在于,所述试剂盒还包括:混合核苷酸,用于在加所述单聚核苷酸尾巴之前,在末端转移酶的作用下,在单链DNA和/或双链DNA每条链的3’端加上一段随机碱基。
- 根据权利要求15所述的试剂盒,其特征在于,所述随机碱基的长度是5-10个碱基。
- 根据权利要求13所述的试剂盒,其特征在于,所述单聚核苷酸尾巴的长度是15-30个单一碱基。
- 根据权利要求13所述的试剂盒,其特征在于,所述上游特异性引物距离所述突变目标位点25-150bp。
- 根据权利要求13所述的试剂盒,其特征在于,所述第一次PCR扩增进行10-20个循环。
- 根据权利要求13所述的试剂盒,其特征在于,所述连续单一碱基是11-15个连续的C或G碱基,或25-35个连续的A或T碱基。
- 根据权利要求13所述的试剂盒,其特征在于,所述第一通用引物在所述连续单一碱基之后还包括一个简并碱基。
- 根据权利要求21所述的试剂盒,其特征在于,所述简并碱基是H、V、B或D。
- 根据权利要求13所述的试剂盒,其特征在于,所述第二次PCR扩增进行15-30个循环。
- 根据权利要求13所述的试剂盒,其特征在于,所述突变目标位点包括单核苷酸多态性、插入和删除以及拷贝数变异。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201780091041.8A CN110651050A (zh) | 2017-07-21 | 2017-07-21 | 用于检测低频突变的靶向富集方法和试剂盒 |
| PCT/CN2017/093914 WO2019014936A1 (zh) | 2017-07-21 | 2017-07-21 | 用于检测低频突变的靶向富集方法和试剂盒 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2017/093914 WO2019014936A1 (zh) | 2017-07-21 | 2017-07-21 | 用于检测低频突变的靶向富集方法和试剂盒 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019014936A1 true WO2019014936A1 (zh) | 2019-01-24 |
Family
ID=65014966
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2017/093914 Ceased WO2019014936A1 (zh) | 2017-07-21 | 2017-07-21 | 用于检测低频突变的靶向富集方法和试剂盒 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN110651050A (zh) |
| WO (1) | WO2019014936A1 (zh) |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112359101B (zh) * | 2020-11-13 | 2023-10-03 | 苏州金唯智生物科技有限公司 | 一种质检寡核苷酸交叉污染的方法 |
| CN114317696B (zh) * | 2021-12-24 | 2024-07-09 | 深圳裕康医学检验实验室 | 一种试剂盒及其文库构建方法与污染检测方法 |
| CN117343929B (zh) * | 2023-12-06 | 2024-04-05 | 广州迈景基因医学科技有限公司 | 一种pcr随机引物及用其加强靶向富集的方法 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20160115532A1 (en) * | 2012-08-10 | 2016-04-28 | Sequenta, Inc. | High sensitivity mutation detection using sequence tags |
| CN106192018A (zh) * | 2015-05-07 | 2016-12-07 | 深圳华大基因研究院 | 一种锚定巢式多重pcr富集dna目标区域的方法和试剂盒 |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105861724B (zh) * | 2016-06-03 | 2019-07-16 | 人和未来生物科技(长沙)有限公司 | 一种kras基因超低频突变检测试剂盒 |
| CN106676182B (zh) * | 2017-02-07 | 2020-08-14 | 北京诺禾致源科技股份有限公司 | 一种低频率基因融合的检测方法及装置 |
-
2017
- 2017-07-21 CN CN201780091041.8A patent/CN110651050A/zh active Pending
- 2017-07-21 WO PCT/CN2017/093914 patent/WO2019014936A1/zh not_active Ceased
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20160115532A1 (en) * | 2012-08-10 | 2016-04-28 | Sequenta, Inc. | High sensitivity mutation detection using sequence tags |
| CN106192018A (zh) * | 2015-05-07 | 2016-12-07 | 深圳华大基因研究院 | 一种锚定巢式多重pcr富集dna目标区域的方法和试剂盒 |
Non-Patent Citations (2)
| Title |
|---|
| VARDI, O. ET AL.: "Biases in the SMART-DNA Library Preparation Method Associated with Genomic Poly dA/dT Sequences", PLOS ONE, vol. 12, no. 2, 24 February 2017 (2017-02-24), pages e0172769-1 - e0172769-14, XP055562515 * |
| ZHENG, Z.L. ET AL.: "Anchored Multiplex PCR for Targeted Next-Generation Sequencing", NATURE MEDICINE, vol. 20, no. 12, 31 December 2014 (2014-12-31), pages 1479 - 1486, XP055169023 * |
Also Published As
| Publication number | Publication date |
|---|---|
| CN110651050A (zh) | 2020-01-03 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP3608420B1 (en) | Nucleic acids and methods for detecting chromosomal abnormalities | |
| CN103198238B (zh) | 构建药物反应相关基因标准型别数据库的方法及其应用 | |
| US20230065345A1 (en) | Method for bidirectional sequencing | |
| CN102533985B (zh) | 一种检测dmd基因外显子缺失和/或重复的方法 | |
| CN104372093B (zh) | 一种基于高通量测序的snp检测方法 | |
| CN105861678B (zh) | 一种用于扩增低浓度突变靶序列的引物和探针的设计方法 | |
| CN110628891B (zh) | 一种对胚胎进行基因异常筛查的方法 | |
| CN106591441B (zh) | 基于全基因捕获测序的α和/或β-地中海贫血突变的检测探针、方法、芯片及应用 | |
| US20140051585A1 (en) | Methods and compositions for reducing genetic library contamination | |
| JP2022510723A (ja) | 遺伝子標的エリアの富化方法及びキット | |
| CN108085315A (zh) | 一种用于无创产前检测的文库构建方法及试剂盒 | |
| CN105385755A (zh) | 一种利用多重pcr技术进行snp-单体型分析的方法 | |
| JP2017176181A (ja) | 胎児の染色体異数性の診断 | |
| HK1222684A1 (zh) | 检测稀有突变和拷贝数变异的系统和方法 | |
| WO2018184495A1 (zh) | 一步法构建扩增子文库的方法 | |
| JP2019536474A (ja) | メチル化dnaの多重検出方法 | |
| CN104264231B (zh) | 构建测序文库的方法及其应用 | |
| US11261479B2 (en) | Methods and compositions for enrichment of target nucleic acids | |
| WO2018133546A1 (zh) | 无创产前胎儿α型地贫基因突变检测文库构建方法、检测方法和试剂盒 | |
| CN103571822B (zh) | 一种用于新一代测序分析的多重目的dna片段富集方法 | |
| CN106399553B (zh) | 一种基于多重pcr的人线粒体全基因组高通量测序方法 | |
| WO2019014936A1 (zh) | 用于检测低频突变的靶向富集方法和试剂盒 | |
| Eboreime et al. | Estimating exceptionally rare germline and somatic mutation frequencies via next generation sequencing | |
| CN104450872A (zh) | 一种高通量多样本多靶点单碱基分辨率的甲基化水平检测方法 | |
| WO2018133547A1 (zh) | 无创产前胎儿β型地贫基因突变检测文库构建方法、检测方法和试剂盒 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 17918129 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 17918129 Country of ref document: EP Kind code of ref document: A1 |





















