WO2016169431A1 - 一种长片段dna文库构建方法 - Google Patents

一种长片段dna文库构建方法 Download PDF

Info

Publication number
WO2016169431A1
WO2016169431A1 PCT/CN2016/079278 CN2016079278W WO2016169431A1 WO 2016169431 A1 WO2016169431 A1 WO 2016169431A1 CN 2016079278 W CN2016079278 W CN 2016079278W WO 2016169431 A1 WO2016169431 A1 WO 2016169431A1
Authority
WO
WIPO (PCT)
Prior art keywords
fragment
linker
sequencing
product
strand
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2016/079278
Other languages
English (en)
French (fr)
Inventor
王欧
常灿坤
林霖
蒋慧
章文蔚
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
BGI Shenzhen Co Ltd
Original Assignee
BGI Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by BGI Shenzhen Co Ltd filed Critical BGI Shenzhen Co Ltd
Priority to CN201680012412.4A priority Critical patent/CN107250447B/zh
Priority to US15/567,963 priority patent/US20180195060A1/en
Publication of WO2016169431A1 publication Critical patent/WO2016169431A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/10Processes for the isolation, preparation or purification of DNA or RNA
    • C12N15/1034Isolating an individual clone by screening libraries
    • C12N15/1093General methods of preparing gene libraries, not provided for in other subgroups
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/10Processes for the isolation, preparation or purification of DNA or RNA
    • CCHEMISTRY; METALLURGY
    • C40COMBINATORIAL TECHNOLOGY
    • C40BCOMBINATORIAL CHEMISTRY; LIBRARIES, e.g. CHEMICAL LIBRARIES
    • C40B40/00Libraries per se, e.g. arrays, mixtures
    • C40B40/04Libraries containing only organic compounds
    • C40B40/06Libraries containing nucleotides or polynucleotides, or derivatives thereof
    • CCHEMISTRY; METALLURGY
    • C40COMBINATORIAL TECHNOLOGY
    • C40BCOMBINATORIAL CHEMISTRY; LIBRARIES, e.g. CHEMICAL LIBRARIES
    • C40B50/00Methods of creating libraries, e.g. combinatorial synthesis
    • C40B50/06Biochemical methods, e.g. using enzymes or whole viable microorganisms

Definitions

  • the invention relates to the field of biotechnology, and in particular to a method for constructing a long segment DNA library.
  • LFR Long Fragment Read
  • 384-well plate physical separation method By separating the genomic sample by 384-well plate physical separation method, the long DNA fragment from the male parent and the female parent is separated and the different tag sequences are added to construct the library. After the sequencing is completed, the genome can be completely phased to confirm the mutation position. Whether the point is on the same parental chromosome. MDA amplification of long DNA fragments isolated into the well plates is performed during library construction, and dUTP or other dNTP analogs are incorporated during amplification.
  • CG Company uses Multiple Displace Amplification (MDA) method to input each well of 384-well plate. A long fragment of DNA is amplified.
  • MDA Multiple Displace Amplification
  • the MDA method is easy to form non-specific amplification, and the single strand that is replaced during the amplification process will combine with the new random primer to form a high-level complex structure to influence the subsequent reaction.
  • the multi-step enzymatic reaction of the linker A is complicated and cumbersome.
  • the method provided by the invention comprises the following steps:
  • sequencing linker single strand A with a different tag and the sequenced linker with a different tag are single-chain B annealed to form the sequencing linker;
  • the obtained PCR amplification product is a PCR amplification product for connecting different sequencing linkers
  • the construction library (specifically see the examples) is to sequentially digest the PCR amplification product of the ligation sequencing linker product, dUTP, double-stranded cyclization, EcoP15 digestion, terminal fill, dephosphorylation, second linker,
  • the second linker ligation product is amplified, isolated to obtain a single strand, and single-stranded, that is, a long-segment DNA library is obtained.
  • the method of step 1) comprises the following steps:
  • the above PCR amplification reaction system (without template): 2x buffer 291.2uL, Primer B (20uM) 8.48uL, Primer C (20uM) 8.48uL, dNTP (25mM each) 20.17uL, dUTP (4mM) 5.04uL, Pfu turbo Cx polymerase 7.68 uL, 20% (by volume) Triton X-100 aqueous solution 20.8 uL, 10% (by volume) Tween 20 aqueous solution 20.8 uL, adding nuclease-free water to a total volume of 560 uL.
  • the transposase is a transposase that occludes an amplification adaptor
  • the amplification linker is one or two, and the amplification linker is formed by a transposase-recognizing single-stranded DNA molecule and a single-stranded DNA molecule partially complementary thereto;
  • the fragment size between the positions of the transposases of two adjacent said amplifying adaptors is 3-10 kb.
  • the DNA matching the amplification adaptor in (2) is a primer which is a reverse complement of the transposase-recognizing single-stranded DNA molecule in an amplification adaptor ligated to both ends of the fragment after the fragmentation.
  • the stranded DNA molecule in addition to the transposase recognizing the reverse complement of the single-stranded DNA molecule, the remaining portion of the sequence forms a primer pair.
  • the amplification linker is a linker 1 and a linker 2,
  • the amplification linker is a linker 1 and a linker 2, which is composed of a transposase-recognizing single-stranded DNA molecule A and a single-stranded DNA molecule B partially complementary thereto, the linker 2 being The enzyme recognizes a single-stranded DNA molecule A and a single-stranded DNA molecule C partially complementary to its counterpart;
  • the DNA molecule that matches the amplification linker is composed of a primer B and a primer C
  • the primer B is a single-stranded DNA molecule B which is removed from the complementary portion of the single-stranded DNA molecule A of the transposase.
  • the remaining sequence; the primer C is the remaining sequence of the single-stranded DNA molecule C except that the transposase recognizes the complementary portion of the single-stranded DNA molecule A;
  • the method of removing dUTP in the first PCR amplification product comprises the following: the first PCR amplification using uracil DNA glycosylase and human depurination pyrimidine endonuclease
  • the product is subjected to a digestion reaction to obtain a digested product; the polymerized product is further polymerized by polymerase I, Taq polymerase and dATP to obtain a DNA fragment of 300-1200 bp in size.
  • polymerase I (10U/uL) 2.86uL
  • Taq polymerase (5U/uL) 5.7uL
  • dATP 100mM
  • nuclease-free water to replenish the total volume to 560uL .
  • the tag sequence is obtained by arranging and combining n bases, and the base is at least one of A, G, C, and T, and n is greater than or equal to 8.
  • the single-strand A of the sequencing linker with different tags includes a fragment A, a fragment B, a tag sequence and a fragment C in order from 5' to 3';
  • the sequencing linker single-strand B with a different tag includes, in order from the 5' to 3' direction, a fragment complementary to the fragment C, a tag sequence, a fragment, and a reverse complement to the fragment A.
  • one primer in the primer matches the fragment B in the single strand A of the sequencing adaptor (same or reverse complement), and the other primer in the primer and the fragment in the single strand B Ding match (reverse complement or the same).
  • the number of the single-strand A of the sequencing linker with different tags is 72 or less, and the sequence of each single-stranded tag is different;
  • the number of the single-stranded B of the sequencing linker with different tags is 72 or less, and the sequence of each single-stranded tag is different;
  • n in the label sequence is greater than or equal to 8 and less than 15.
  • step 2) the tagged sequencing linker single chain A and its complementary sequenced linker single chain B are separately added to the system containing the fragment in a single strand form.
  • the reaction includes the following steps:
  • the sequencing linker single-chain B reaction buffer contains T4 ligase
  • one primer in the primer is identical or reversely complementary to the fragment B in the single strand A of the sequencing adaptor;
  • the other primer in the primer is complementary or identical to the fragment of the fragment in the single-strand B of the sequencing linker.
  • the PCR amplification is carried out by mixing together the products of the different sequencing linkers in each well of a 5184-well plate;
  • step 1) of the method are all carried out in a 5184-well chip.
  • the fragment is divided into several parts, it means that the fragment is divided into the pores of the 5184-well chip, and the dispensing can be carried out in step 1) or step 2).
  • the long fragment DNA is a fragment larger than 100 Kb. Specifically, a fragment having a size of 400 kb;
  • the transposase is a Tn5 transposase; in this embodiment, the transposase encapsulating the amplifying adaptor is a product of Vazyme, and is named TruePrep mini DNA Sample Prep Kit (S302-01-B) transposase kit.
  • the optimal use concentration is 100 times diluted transposase.
  • the amount of the transposase and the long fragment DNA encapsulating the amplification adaptor is as follows.
  • the reaction system of the above reaction includes the single-stranded DNA molecule B, the single-stranded DNA molecule C, the reaction buffer, dNTP, dUTP.
  • a wafergen MSND pipetting platform is used to add various substances into the pores of the chip.
  • the method further comprises the steps of: digesting the transposase with a denaturation reagent; the denaturation reagent specifically adopts a 0.1-1% SDS solution; the digestion condition Specifically, it was allowed to stand at 25 ° C for 10 minutes.
  • the transposition enzyme recognizes the nucleotide sequence of single-stranded DNA molecule A as sequence 5 in the sequence listing;
  • the nucleotide sequence of the single-stranded DNA molecule B is the sequence 6 in the sequence listing;
  • the nucleotide sequence of the single-stranded DNA molecule C is the sequence 7 in the sequence listing;
  • the nucleotide sequence of the primer B is the sequence 12 in the sequence listing;
  • the nucleotide sequence of the primer C is the sequence 13 in the sequence listing;
  • the nucleotide sequence of the single-strand A of the sequencing linker is sequence 1 or sequence 8 in the sequence listing;
  • the nucleotide sequence of the single-strand B of the sequencing linker is sequence 2 or sequence 9 in the sequence listing;
  • the nucleotide sequence of the single-stranded primer A (one primer) is sequence 3 or sequence 10 in the sequence listing;
  • the nucleotide sequence of the single-stranded primer B (another primer) is SEQ ID NO: 4 or SEQ ID NO: 11 in the Sequence Listing.
  • the method further comprises the steps of: digesting the transposase with a denaturation reagent; the denaturation reagent specifically adopts a 0.1-1% SDS solution; the digestion condition Specifically, it was allowed to stand at 25 ° C for 10 minutes.
  • the above denaturing reagent is a solution containing SDS, specifically from 11.2 uL of 1% SDS (mass/volume percentage) and 548.8 uL of 2x buffer (MP01137, Complete Genomics).
  • the condition of the transposase cleavage reaction is 55 ° C for 5 minutes;
  • the annealing temperature of the amplification is 68 ° C, and the annealing time is 18 min;
  • the digestion reaction conditions are 37 ° C for 2 hours, and then 65 ° C for 15 minutes; the polymerization reaction conditions are 23 ° C for 1 hour, 65 ° C for 30 minutes;
  • step 2) the reaction conditions are 20 ° C for 2 hours;
  • step 3 the annealing temperature of the PCR amplification is 68 ° C, and the annealing time is 10 min.
  • the long fragment DNA library prepared by the above method is also within the scope of protection of the present invention.
  • Another object of the present invention is to provide a method for the preparation of a ligation sequencing linker for the construction of a long fragment DNA library.
  • the method provided by the invention comprises the following steps: step 1)-2) in the above method.
  • the denaturation reagent of the present invention may be an NTbuffer, SDS or other reagent which denatures the transposase and detaches from the long DNA fragment.
  • Figure 1 is a schematic diagram of a transposase disrupting genomic fragment.
  • Figure 2 is a schematic diagram of the wafergen MSND pipetting platform chip.
  • Figure 3 is a schematic diagram showing the binding of a transposase to genomic DNA.
  • Figure 4 shows the amplification of the primer sequence matched with the linker 1/2 inserted by the transposase after detachment of the transposase.
  • Figure 5 shows that dUTP was cleaved from the PCR product under the action of UDG enzyme and APE1 enzyme and left a gap of 1 bp on the product.
  • Figure 6 shows the notch translation produced by digestion with dUTP.
  • Fig. 7 is a schematic view showing the manner in which two single chains of the first linker each having a tag sequence are added.
  • Figure 8 is a schematic view showing the manner in which two single strands of the first link are connected to the insert.
  • Figure 9 is a schematic diagram of the PCR amplification reaction after the first linker is completed.
  • Figure 10 is a schematic diagram of two spray methods of wafergen MSND.
  • Figure 11 is a gel electrophoresis gel of the PCR product after the transposase fragment was interrupted.
  • Figure 12 is a product electrophoresis diagram of each step of the chip building process.
  • Figure 13 is an EcoP15 endonuclease cleavage product.
  • Figure 14 is a diagram showing the electrophoresis after single-chain cyclization.
  • Fig. 15 is a histogram of the amount of data and the corresponding frequency in each hole.
  • Figure 16 shows the depth profile of the sequencing.
  • the present invention finally realizes library construction by inserting a specific sequence using a transposase, and performing PCR on a chip having 5,184 wells by means of the specific sequence.
  • sampled or reaction liquid can be sprayed in different ways by means of wafergen's MSND pipetting platform and a chip with 5,184 holes, as shown in the following Table 1:
  • Table 1 shows different ways of spraying
  • the long fragment DNA is cleaved into a target fragment of 3-10 kb in size by a transposase.
  • Figure 1 is a schematic diagram of a transposase disrupting genomic fragment, in which a transposase encoding a linker 1 and a linker 2 is embedded, and after being mixed with genomic DNA, randomly binds to a position in the genome, and controls the amount of transposase used to control
  • the fragment size between the positions of two adjacent transposases is between about 3 and 10 kb.
  • Figure 2 is a schematic diagram of the wafergen MSND pipetting platform chip.
  • the chip has a total of 72 rows in the horizontal direction, a total of 72 columns in the longitudinal direction, and the total number of micropores is 5,184.
  • Transposase enzymes can be used to embed one or two types of linkers.
  • the transposase enzyme of the embedded amplification adaptor used in this embodiment is embedded with two kinds of adaptors, and the transposase is used.
  • a linker obtained by annealing a single-stranded DNA molecule A and a single-stranded DNA molecule B which is inversely complementary to the A moiety, a transposase-recognizing single-stranded DNA molecule A, and a single-stranded DNA molecule C which is inversely complementary to the A moiety .
  • These two kinds of linkers were mixed in an equal volume of 100 uM each, and then mixed with a transposase and embedded.
  • the transposase for embedding the amplification adaptor used in this example was a Tn5 transposase of Vazyme, and was named TruePrep mini DNA Sample Prep Kit (S302-01-B) transposase kit.
  • Transposase recognizes single-stranded DNA molecule A: 5'-CTGTCTCTTATACACATCT-3' (sequence 5)
  • the amount of Tn5 transposase embedded in the adaptor sequence is controlled to act on the genomic DNA such that two adjacent transposase sites are separated by a long distance in the same DNA fragment (about 3 10K). Due to the nature of the transposase itself, it does not immediately detach from the DNA after the completion of the action, but is still present at the site of action, and the long segment of the fragmented DNA is ligated, so that the entire length of the entire DNA remains at its original length. Not disconnected.
  • the above transposase product was diluted to a certain concentration (to allow sufficient separation between long DNA fragments), and then transferred to the chip using wafergen's micropipetting platform to achieve physical separation of long DNA fragments.
  • the chip carries 5184 microwells, which allows up to 5184 individual DNA samples to be separated.
  • the reagents added in the subsequent steps were all carried out using the pipetting platform.
  • the fragment length required for the experiment was 3-10 kb.
  • the transposases encapsulating the amplification adaptors were diluted by three gradients of 50-fold, 100-fold, and 150-fold, respectively.
  • the following reaction system was added to the PCR reaction tube: 5x buffer 2uL (S302-01, Vazyme), human genomic DNA (genomic DNA extracted from ex vivo blood cells; 7 ng, greater than 100 kb) 6 uL, transposase (different dilution) Multiples) 2uL.
  • the above reaction system with different dilution multiple transposases was mixed and placed in a PCR machine, reacted at 55 ° C for 5 minutes, then cooled to normal temperature, and the reaction mixture was added to a final concentration of 0.1%.
  • the denaturing agent SDS was mixed and allowed to stand at room temperature for 10 minutes to obtain a target fragment having a size of 3 to 10 kb.
  • the target fragment having a size of 3 to 10 kb obtained in the above 1) was subjected to PCR amplification in the following PCR reaction system to obtain an amplification product of the desired fragment.
  • the primer for amplification is a single-stranded DNA molecule which is ligated in the amplification adaptor at both ends of the fragment after the fragmentation and which is complementary to the transposase-recognized single-stranded DNA molecule, except for the transposase recognition
  • the primer pairs formed by the remaining partial sequences in addition to the reverse complementary sequence of the single-stranded DNA molecule are as follows:
  • Primer B 5'-TCG CGGCAGCGTC-3' (sequence 12)
  • the above 100 uL PCR reaction system comprises: a target fragment of 3-10 kb in size (8 pg/uL) 5 uL, 2 x buffer 50 uL, primer B (20 uM) 0.5 uL, primer C (20 uM) 0.5 uL, DNA polymerase 0.5 uL, hydration to 100uL.
  • the optimal transposase dilution factor is 100-fold dilution.
  • the genomic DNA was treated with the optimal 100-fold diluted transposase described above.
  • the PCR adaptor-embedded PCR adaptor sequence was inserted between the genomic DNA at a distance of about 3-10K, and the PCR sequence was amplified from the linker sequence as follows:
  • transposase and the human genomic DNA in which the amplification adaptor is embedded are reacted according to the following reaction system, as follows:
  • the above reaction system was as follows: 2 uL of 100-fold diluted transposase, 2 uL of 5x buffer (S302-01, Vazyme), 7 ng of human genomic DNA (length 400 kb), and water supplemented to 10 uL.
  • the reaction system can be scaled up or down.
  • reaction system is placed in a PCR machine, reacted at 55 ° C for 5 minutes, and then cooled to normal temperature to obtain a reaction product, which is diluted to a final concentration of about 8 pg/uL to obtain a diluted reaction product (inserted every 3-10 KD).
  • a reaction product is diluted to a final concentration of about 8 pg/uL to obtain a diluted reaction product (inserted every 3-10 KD).
  • the diluted fragmented product was dispensed into each well of a 5184-well plate at a dose of 35 nL/well using a wafergen MSND pipetting platform single sample dispensing procedure (the solvent of the well was 200 nl); the chip was centrifuged (4000 rpm, 5 min) to allow The sprayed liquid settled to the bottom of the chip to obtain a 5184-well plate containing the reaction product.
  • Figure 3 shows that the transposase binds to genomic DNA and inserts into the linker 1/2, and does not immediately detach but attach to its site of action. A certain concentration of SDS is added to the reaction solution to cause the transposase to fall off the DNA.
  • the denaturing reagent was added to each well of the 5184-well plate containing the reaction product obtained in the above 1) using a wafergen MSND pipetting platform single sample dispensing procedure, and the chip was centrifuged (time 4000 rpm, 5 min) to allow the sprayed liquid to settle. At the bottom of the chip; after centrifugation, it is allowed to stand at room temperature for about 10 minutes to digest the transposase, detach it from the genomic DNA and fragment the genomic DNA to obtain a 3-10KD target fragment, and the chip containing the same contains 3-10KD.
  • the 5184-well chip of the target fragment was added to each well of the 5184-well plate containing the reaction product obtained in the above 1) using a wafergen MSND pipetting platform single sample dispensing procedure, and the chip was centrifuged (time 4000 rpm, 5 min) to allow the sprayed liquid to settle. At the bottom of the chip; after centrifugation, it is allowed to stand at room temperature for about 10 minutes to
  • denaturing reagents and concentrations can be used for this procedure, not limited to SDS.
  • the above denaturing reagent consists of 11.2 uL of 1% SDS (mass/volume percentage) and 548.8 uL of 2x buffer;
  • Figure 4 shows the addition of primer B and primer C to the PCR system for amplification after completion of the detachment of the transposase.
  • a certain proportion of dUTP 4%, dUTP: dATP + dTTP + dCTP + dGTP was added to the PCR system to incorporate a certain proportion of dUTP into the amplification product.
  • PCR was performed on the product obtained in the previous step using the designed primers.
  • the following PCR reaction system was added to each well of the 5184-well chip containing the 3-10 KD target fragment obtained in step 2) using a wafergen MSND pipetting platform single sample dispensing procedure, followed by centrifugation to sediment the sprayed liquid, and then placing the chip board.
  • the PCR amplification reaction was carried out in a PCR machine to obtain a 5184-well chip containing the amplification product of the target fragment of dUTP.
  • a certain proportion of dUTP is added to the above reaction system, so that a part of the dTTP is replaced by dUTP in the amplification, so that the amplification product can be interrupted again by the removal of dUTP in the later stage.
  • the above PCR amplification reaction system 2x buffer 291.2 uL, single-stranded DNA molecule B (20 uM) 8.48 uL, single-stranded DNA molecule C (20 uM) 8.48 uL, dNTP (25 mM each) 20.17 uL, dUTP (4 mM) 5.04 uL, Pfu Turbo Cx polymerase 7.68 uL, 20% (by volume) Triton X-100 aqueous solution 20.8 uL, 10% (by volume) Tween 20 aqueous solution 20.8 uL, adding nuclease-free water to a total volume of 560 uL.
  • Table 3 shows the PCR amplification reaction procedure
  • Figure 11 is a gel electrophoresis gel of the PCR product after the transposase fragment was interrupted.
  • the chip was placed in a 37 ° C metal bath and allowed to stand for 15.5 minutes to evaporate excess liquid.
  • the 3-10KD target fragment is fragmented into 300-1200bp DNA short fragments.
  • the dUTP introduced in the digestion was digested with uracil DNA glycosylase (UDG) and human depurinated apyrimidine endonuclease (APE1).
  • UDG uracil DNA glycosylase
  • APE1 human depurinated apyrimidine endonuclease
  • the digestion reaction system was added to each well of the 5184-well chip containing the amplification product of the target fragment containing dUTP according to the single sample fractionation procedure of the wafergen MSND pipetting platform, and then the liquid was sedimented by centrifugation (4000 rpm, 5 min).
  • the chip was placed in a dedicated PCR machine and subjected to a digestion reaction to obtain a 5184-well chip containing the enzyme-cut product.
  • the above digestion conditions were 2 hours at 37 ° C, 15 minutes at 65 ° C, and then returned to normal temperature.
  • the nick translation interrupts the DNA and adds A to the end.
  • DNA polymerase I and Taq enzyme were added to the reaction system.
  • the product obtained in the previous step was disrupted by the action of DNA polymerase I (5'-3' exonuclease activity and 5'-3' polymerase activity) at 23 °C.
  • DNA polymerase I binds to the gap left after the removal of dUTP, then excises the DNA base from the 5'-3' direction, and polymerizes the DNA base from the 5'-3' direction to achieve the gap from 5'-3 'The translation of the direction.
  • the DNA duplex breaks into smaller fragments.
  • Taq polymerase was added to the system to fill in the DNA double strand at 65 ° C and add a dATP base at the 3' end. At this temperature, DNA polymerase I loses its activity due to denaturation.
  • Fig. 6 is a gap generated by digestion of dUTP, and the DNA polymerase I which is later introduced into the reaction system combines the 5'-3' direction cutting and the 5'-3' direction synthesis at the position of the gap, so that The notch translates in the 5'-3' direction.
  • the gaps in the double strands on both the positive and negative strands simultaneously translate and meet, eventually breaking the entire double strand at the encounter position.
  • the Taq enzyme added to the system exerts its activity at a temperature of 65 ° C, so that the 3' end of the product forms a prominent A base.
  • the polymerization reaction system was sprayed into the above 1) pores of the 5184-well chip containing the enzyme-cut product, and then the liquid was sedimented by centrifugation (4000 rpm, 5 min); Into a dedicated PCR machine, polymerization was carried out to obtain a 5184-well chip containing a DNA fragment of 300-1200 bp in size.
  • reaction conditions were 1 hour at 23 ° C, 30 minutes at 65 ° C, and then returned to room temperature.
  • a sequencing linker with a partial sequencing primer is ligated at both ends of the 300-1200 bp DNA fragment obtained above, and the sequencing linker is formed by annealing a single-strand A with a different tag and a single-strand B of a sequencing link with a different tag.
  • each column in the longitudinal direction of the chip (corresponding to the single-strand A of the sequencing linker) is added to the single-chain A with different tags sequencing link number 1-72.
  • Fragment B The first strand of the PCR primer shown in SEQ ID NO:3 is identical, and only the U of SEQ ID NO:3 is replaced by T;
  • NNNNNNNNNN consists of 10 bases to form a tag sequence
  • thiophosphoric acid that is, a nucleic acid of thiophosphoric acid used in the synthesis of the last base.
  • the second strand of the PCR primer shown in SEQ ID NO:4 is reverse complementary, and only the U of SEQ ID NO:4 is replaced by T;
  • Figure 7 is a schematic diagram showing the manner in which two single strands of the first sequencing link each have a tag sequence.
  • two tag sequences are used to form a combined form.
  • the two single strands of the first linker were added separately in the experiment.
  • the single strand of the 72 tag sequences of the first strand (sequence linker single strand A) is added in a longitudinal form, that is, the tag sequence of the first strand added in each column is the same.
  • the single strand of the respective 72 tag sequences of the second strand is added in a lateral form, i.e., the tag sequence of the second strand added in each row is the same. This ultimately results in a matrix of tag sequences in 72x72, with each well obtaining a unique combination of two-two tag sequences.
  • the sequencing linker with the tag sequence is added in a longitudinal fashion, and the sequencing link with the tag sequence is single-stranded B in a lateral manner. Join as follows:
  • the above-mentioned 5184-well plate to which the sequencing linker single-strand B was added was placed in a PCR apparatus and reacted at 20 ° C for 2 hours to obtain a 5184-well plate containing the product of the ligation sequencing linker.
  • Figure 8 is a schematic diagram showing the connection of two single strands and an insert of a first adaptor (sequencing linker).
  • Figure 10 is a schematic diagram of two spray methods of wafergen MSND.
  • the left picture shows the "72 sample 35nL liquid separation process" combined with the sequencing method of the single-strand A of the sequencing linker, so that the single-strand (sequence linker single-strand A) of the 72-tag sequence of the first strand is added in the longitudinal form.
  • the picture on the right is the "72 sample 50nL liquid separation process" combined with the second chain of the first link, so that the single strand of the 72 strands of the second strand (sequence linker single strand B) is added in a lateral form.
  • Figure 9 is a schematic diagram showing the PCR amplification reaction after completion of ligation of the first linker (sequencing linker).
  • the product of all the wells of the 5184-well plate containing the sequencing sequencing linker obtained from the above three is separated from The hearts were collected in a 1.5 mL centrifuge tube and purified using 1x AmpureXP magnetic beads. The purified product was recovered in 100 uL of TE solution. After purification, the sequencing linker product was ligated and subjected to PCR amplification reaction to obtain a ligation sequencing product. PCR amplification products.
  • the primers used in the above PCR amplification are as follows:
  • PCR primer first strand upstream primer: GGUCGCCAGCCCUATGGC (sequence 3)
  • PCR primer second strand (downstream primer): AGGGCUGGCGACCUTGTCAG (sequence 4)
  • the reaction system used for the above PCR amplification is as follows: ligation sequencing product 100uL, 2x PfuCx buffer 275uL, PCR primer first strand (20uM) 14uL, PCR primer second strand (20uM) 14uL, PfuCx polymerase 11uL, plus ribozyme
  • the water is replenished to 550 uL.
  • Table 4 shows the amplification reaction procedure
  • the PCR amplification product to which the sequencing linker product was ligated was purified using 1x AmpureXP magnetic beads, and the purified product was dissolved in 60 uL of TE solution.
  • the electrophoresis detection of the reaction products of the above steps is as shown in Fig. 12, 1.
  • the product is replaced by water); 5, the cleavage translation end is added with the A product (step 2 of 2) polymerization reaction product; 6, the nick translation end plus A negative product (step 2 of 2), the enzyme digestion product in the polymerization reaction is replaced with water 7; linker ligation product 1 (step 3 of the ligation sequencing linker product); 8 linker ligation product 2 (step 3 of the 2 ligation sequencing linker product); 9, PCR product 1 (step 4 PCR amplification product 10; PCR product 2 (PCR amplification product of step 4); 11, PCR negative product; M2, 100 bp molecular weight marker. Can be seen from Figure 12 The target product obtained in each step.
  • the PCR amplification product of the above-described ligation-splicing linker product needs to digest the dUTP carried in the primer to form a sticky end, thereby realizing self-cyclization in subsequent experiments.
  • the dUTP digestion reaction was as follows, 10 ⁇ Taq buffer 11 uL (RM00059, Complete Genomics), User enzyme (RM00017, Complete Genomics), and the purified PCR product was 60 uL, supplemented with 110 ⁇ L with nuclease-free water, and reacted at 37 ° C for 1 hour to obtain a reaction product.
  • the reaction product obtained in the above five was added to 100 ⁇ L of 10 ⁇ TAbuffer (RM06601, Complete Genomics), and the enzyme-free water was added to 1810 uL.
  • the mixture was divided into 4 tubes on average, and reacted in a water bath at 60 ° C for 30 minutes, and then transferred to a room temperature water bath for 20 minutes.
  • Each tube reaction product was purified by Ampure XP magnetic beads and the product was recovered in 70 uL of TE.
  • the digestion reaction was as follows. For each of the above obtained purified cyclized products, 9x PS mix (MP01154, Complete Genomics) 8.9 uL, Plasmid-Safe (RM02046, Complete Genomics) 10.4 uL, and nuclease-free water was added to 80 uL. After mixing, the reaction was carried out at 37 ° C for 1 hour. The reaction product was purified using Ampure XP magnetic beads and dissolved in 40 uL of TE solution to give a cyclized product after digestion.
  • the cyclized product, the first two-terminal linker was ligated together, and here, EcoP15 was digested, and the cyclized product was cut from the EcoP15 restriction site at both ends of the linker to about 27 bp at the inner ends of the insert. Subsequent screening of the digested product fragments, obtaining the intermediate band linker sequence, with the products of the inserts cut at both ends, as follows:
  • the digestive digestion system was prepared, and the product obtained after the above digestion was cyclized with 37 uL, 5xEcoP 15 Mix 3 (MP01149, Complete Genomics) 72 uL, EcoP 15 (RM00063, Complete Genomics) 10.8 uL, supplemented with 360 ⁇ L with nuclease-free water, and reacted at 37 ° C for 16 hours.
  • Figure 13 shows the EcoP15 endonuclease cleavage and the fragment recovered product.
  • the product fragment is around 140 bp.
  • the EcoP15 digested product was end-filled to join the subsequent second linker.
  • the end-filling reaction system was prepared, and the above-obtained digested product 44uL, 10x NEB buffer 2 (New England biolabs) 5.4 uL, 25 mM dNTP 0.8 uL, 10 mg/mL BSA 0.4 uL, T4 DNA polymerase (M0203-com, New) was prepared. England Biolabs), reacted at 12 ° C for 20 minutes to obtain a terminal fill product.
  • the reaction product was purified using PEG32 magnetic beads and dissolved in 48 uL of TE.
  • the terminal fill product obtained in the above nine was subjected to dephosphorylation treatment to cooperate with the subsequent second linker.
  • the reaction system was prepared as follows. 10x NEB buffer 2 (New England Biolabs) 5.75 uL, Fast AP (EF0651, Fermentas) 5.75 uL, and the product was purified by adding 40 ⁇ L of the terminal-purified product, and reacted at 37 ° C for 45 minutes. The reaction product was purified using 75 uL of PEG32 magnetic beads and dissolved in 42 uL of TE solution to give the dephosphorylated product.
  • the second linker is directionally connected, and after 4 steps of enzymatic reaction, the dephosphorylation product obtained by the above ten is first introduced into the 3' end linker sequence, and the terminal is filled again and the dATP is introduced into the end, and the 5' end linker is introduced. The sequence is finally replaced with a stretch of oligonucleotide sequence and ligated using a ligase.
  • a reaction mixture was prepared, 5x klex NTA mix (MP01150, Complete Genomics) 10.7 uL, klenow (RM00066, Complete Genomics) 1.1 uL, nuclease-free water 1.5 uL, and 40 uL of the above step product was added. The mixture was reacted at 37 ° C for 1 hour. The reaction was completed using 69 uL of PEG32 magnetic beads for purification and dissolved in 42 uL of TE solution.
  • reaction mixture was prepared, 3xHB 24.7 uL, nuclease-free water 1.9 uL, T4 ligase 1.9 uL, two-joint 5' end link 5.6 uL (9 uM), and 42 uL of the above step product was added. Mixture at 14 ° C After 2 hours of reaction, the reaction was completed using 63 uL of PEG32 magnetic beads for purification and dissolved in 42 uL of TE solution.
  • the enzyme reaction solution was prepared, 10 ⁇ Taq buffer 0.4 uL, T4 ligase 4.8 uL, Taq polymerase 4.8 uL.
  • a second linker is obtained.
  • Table 5 shows the amplification reaction procedure
  • the reaction was purified using 550 uL of PEG32 magnetic beads and dissolved in 85 uL of TE solution.
  • the purified product was quantitatively detected, and 600 ng was taken for the subsequent single-chain cyclization step.
  • a 1xBWB/tween mixture was prepared, and 1xBWB (MP01111, Complete Genomics), 0.5% Tween 20 20 uL was added to the centrifuge tube and mixed well.
  • Stretavidin beads (MP01162, Complete Genomics) 120 uL were placed and placed on a magnetic stand to fully adsorb the magnetic beads, and the supernatant was removed. The beads were washed twice with 1 x BBB (MP01110, Complete Genomics) 600 uL, then 120 uL of 1 x BBB was added to resuspend the beads, and then 1% by volume of 0.5% Tween 20 was added.
  • the product was subjected to single-strand separation using cleaned magnetic beads. Take 600 ng of the amplified product obtained in the above 12, add TE to 60 uL, add 4xBBB (MP01145, Complete Genomics) 20 uL, 30 uL of stretavidin beads after washing, mix thoroughly, let stand for 15 minutes, and then place on a magnetic stand to make magnetic The beads were fully adsorbed and washed twice with a prepared 1xBWB/tween mixture. After the completion of the washing, 75 uL of 0.1 M NaOH solution was added, and after mixing for 2 minutes, it was placed on a magnetic stand to fully adsorb the magnetic beads, and the supernatant was recovered. Finally, 0.3 M MOPS acid (MP01165, Complete Genomics) 37.5 uL was added to the recovered supernatant, and the mixture was uniformly mixed.
  • 4xBBB MP01145, Complete Genomics
  • Primer mixture take Bridge Oligo (20uM) 20uL, add 43uL of nuclease-free enzyme water, mix and mix for use;
  • the enzyme reaction mixture was centrifuged with 135.3 uL of water, and 10 ⁇ TA buffer (RM06601, Complete Genomics) 35 uL, 100 mM ATP 3.5 uL, ligase (600 U/uL) 1.2 uL was added, and the mixture was uniformly used.
  • 10 ⁇ TA buffer (RM06601, Complete Genomics) 35 uL, 100 mM ATP 3.5 uL, ligase (600 U/uL) 1.2 uL was added, and the mixture was uniformly used.
  • Figure 14 is a diagram showing the electrophoresis after single-chain cyclization.
  • the library preparation process is complete.
  • Example 2 library construction method (adapted to illumina sequencing platform)
  • the long fragment DNA is cleaved into a target fragment of 3-10 kb in size by a transposase.
  • the target fragment 3-10KD target fragment is secondarily fragmented into 300-1200bp DNA short fragments
  • a sequencing link with a partial sequencing primer is ligated to both ends of the one-step reaction product of the 300-1200 bp DNA fragment obtained in the above two, and the double-strand of the linker of the step has a different label sequence, respectively, for each hole on the chip.
  • Both can be distinguished in the process of sequencing, and each column in the longitudinal direction of the chip (corresponding to the first chain of the sequencing link) is added to the tag sequence numbered 1-72, each row in the horizontal direction of the chip (corresponding to the second link of the connector) Chain) Add a sequence of tags numbered 1-72. This forms a 72x72 matrix through two separate tag sequences, resulting in a unique dual tag sequence combination for each well.
  • a linker with a portion of the sequencing primer sequence was ligated at both ends of the reaction product of the previous step.
  • the Flow cell linker P5 italic
  • the 10-base tag sequence N
  • the single-stranded DNA molecule complementary to the Read 1 sequencing linker symbold, * represents thiophosphoric acid
  • the Read 2 sequencing linker (bold), the 10-base tag sequence (N), and the single-stranded DNA molecule complementary to the Flow cell linker P7 (italic, Phos, phosphate, AmMo) Amino-modified) composition; the 7-base random primers of the 72 3'-end tag primers are all different.
  • Fig. 7 is a schematic view showing the manner in which two single chains of the first linker each having a tag sequence are added.
  • two tag sequences are used to form a combined form.
  • the two single strands of the first linker were added separately in the experiment.
  • the single strand of each of the 72 tag sequences of the first strand is added in a longitudinal form, that is, the sequence of the first strand added in each column is the same.
  • the single strands of the respective 72 tag sequences of the second strand are added in a lateral form, i.e., the tag sequence of the second strand added to each transverse row is the same. This ultimately results in a matrix of tag sequences in 72x72, with each well obtaining a unique combination of two-two tag sequences.
  • the first linker with the tag sequence is added longitudinally, and the second link with the tag sequence is single-stranded.
  • Join in the horizontal mode as follows:
  • the sequencing linker single-chain A reaction solution was prepared as follows: 10 ⁇ buffer (25% (volume ratio) PEG8000, 500 mM Tris-HCl, 10 mM ATP, 100 mM MaCl 2 ) 252 uL, and the total volume was added to 957 uL by adding nuclease-free water;
  • sequencing linker single-strand B reaction solution 10xbuffer (25% (volume ratio) PEG8000, 500mM Tris-HCl, 10mM ATP, 100mM MaCl 2 ) 296.4uL, T4 ligase (600U/uL, Enzymatics) 125.97uL, plus no core
  • the enzyme water supplemented the total volume to 1186 uL.
  • the above-mentioned 5184-well plate to which the sequencing linker single-strand B was added was placed in a PCR apparatus and reacted at 20 ° C for 2 hours to obtain a 5184-well plate containing the product of the ligation sequencing linker.
  • Figure 8 is a schematic view showing the manner in which two single strands of the first link are connected to the insert.
  • Figure 10 is a schematic diagram of two spray methods of wafergen MSND.
  • the left picture shows the "72 sample 35nL liquid separation process" with the first chain of the first link, so that the single chain of the 72 different label sequences of the first chain is added in the longitudinal form.
  • the picture on the right is the "72 sample 50nL liquid separation process" with the second chain of the first link, so that the single chain of the 72 different label sequences of the second chain is added in a lateral form.
  • Figure 9 is a schematic diagram of the PCR amplification reaction after the first linker is completed.
  • the product obtained from the above three obtained all the wells of the 5184-well plate containing the sequencing-splicing product was centrifuged and collected in a 1.5 mL centrifuge tube, and purified using 1 time AmpureXP magnetic beads, and the purified product was dissolved in 100 uL of TE solution, and purified.
  • the sequencing linker product is ligated and subjected to PCR amplification reaction to obtain a PCR amplification product which is ligated to the sequencing linker product.
  • the primers used in the above PCR amplification are as follows:
  • PCR primer first strand upstream primer: AATGATAGCGCACCCACCGAGATCT (sequence 10)
  • the reaction system used for the above PCR amplification is as follows: ligation sequencing product 100uL, 2x PfuCx buffer 275uL, PCR primer first strand (20uM) 14uL, PCR primer second strand (20uM) 14uL, PfuCx polymerase 11uL, plus ribozyme
  • the water is replenished to 550 uL.
  • Table 6 is the amplification reaction procedure
  • the PCR amplification product to which the sequencing linker product was ligated was purified using 1x AmpureXP magnetic beads, and the purified product was dissolved in 60 uL of TE solution.
  • reaction products of the above steps were electrophoresed to obtain the target product.
  • the DNA long fragment library prepared in Example 2 was sequenced on an illumina platform with a sequencing depth of 40X.
  • the sequence alignment software SOAP aligner 2.20 was used (LiR, LiY, Kristiansen K, et al, SOAP: short oligonucleotide alignment program. Bioinformatics 2008, 24(5): 713-714; LiR, YuC, LiY, et al, SOAP2: an improved Ultrafast tool for short read alignment. Bioinformatics 2009, 25(15): 1966-1967;
  • the method can successfully mark all 5184 wells.
  • a long DNA fragment library was prepared using CG's Multiple Displacement Amplification (MDA) method and methods for long fragment read sequencing (United States Patent 8, 592, 150), and then on the Complete Genomincs platform. Sequencing was performed and the sequencing depth was 80X.
  • MDA Multiple Displacement Amplification
  • a long DNA fragment library was prepared using Example 1 and then sequenced on a Complete Genomincs platform with a sequencing depth of 60X.
  • the sequence alignment software SOAP aligner 2.20 was used (LiR, LiY, Kristiansen K, et al, SOAP: short oligonucleotide alignment program. Bioinformatics 2008, 24(5): 713-714; LiR, YuC, LiY, et al, SOAP2: an improved Ultrafast tool for short read alignment. Bioinformatics 2009, 25(15): 1966-1967;
  • the amplification of the method is better than that of the MDA amplification method, and is closer to the sequencing depth distribution of the conventional sequencing method.
  • the present invention replaces the MDA amplification method used for the original DNA amplification and the connection mode of the linker A, and replaces the 384-well plate with a chip of the wafergen micropipetting platform.
  • the present invention uses a transposase to insert a specific sequence and utilizes that particular Sequences PCR amplification of long DNA fragments in each single well is a conventional "denaturation-anneal-extension" approach during amplification, avoiding the problem of MDA amplification forming complex structures and non-specific amplification.
  • two independent tag sequences are creatively designed in two single chains of a double link header, and then respectively added in a single chain form, and two tag sequences are simultaneously connected to the two insert segments at the time of connection. end.
  • the primer designed by the present invention is first annealed in a normal temperature reaction system, and then joined to both ends of the insert by a ligase.
  • the 3' end of the interrupted insert is extended by a dATP prior to ligation, and the two single strands of the linker are directionally joined at both ends of the insert by pair A-T pairing in the ligation reaction.
  • the invention combines with the wafergen micro pipetting platform to separate the long DNA fragments into the chip with 5184 pores, and the number of single sample separation holes is more than 10 times that of the 384-well plate, and the separation effect is more sufficient. This results in better location and assembly of the mutation site.
  • the experiments of the present invention prove that the present invention improves the separation probability of homologous long fragments in the genome by combining with the micropore chip, and the 5184 in the chip is compared with the 384-well plate separation method of Complete Genomics. Hole separation is better. In the separation of 384-well plates, about 5% of the homologous long fragments will be separated into the same well and cannot be distinguished. However, by the separation of the 5184-well chip, the probability will be further reduced to improve the efficiency of the analysis.
  • the invention uses long fragment PCR to amplify the DNA fragments in each well after the separation, instead of using the MDA amplification in the original method, thereby avoiding the formation of multi-level complex structures in the local regions caused by MDA amplification. happening. In the data, the data generated by these complex structures is removed because they cannot be identified, limiting the resulting effective data ratio. Using long-segment PCR amplification, the obtained data is more efficient.
  • the two joint single chains are respectively carried with the label sequence and respectively added, and the joint is connected while annealing in the connecting process.
  • the joint connection-extension-connection of the other end joint is carried out, and the one-step reaction of the method can realize the connection of different joints at both ends and add a combination of different label sequences at both ends, thereby greatly reducing the operation complexity and Operating time.

Landscapes

  • Chemical & Material Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Organic Chemistry (AREA)
  • Engineering & Computer Science (AREA)
  • Biochemistry (AREA)
  • Zoology (AREA)
  • Biomedical Technology (AREA)
  • Wood Science & Technology (AREA)
  • Biotechnology (AREA)
  • General Engineering & Computer Science (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Molecular Biology (AREA)
  • Microbiology (AREA)
  • Plant Pathology (AREA)
  • Biophysics (AREA)
  • Physics & Mathematics (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • Chemical Kinetics & Catalysis (AREA)
  • General Chemical & Material Sciences (AREA)
  • Medicinal Chemistry (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

提供了一种长片段DNA文库构建方法,包括如下步骤:1)通过转座酶使长片段DNA断开为大小为3-10kb的目的片段,再扩增所述目的片段,得到含有dUTP的目的片段扩增产物;2)通过去除所述目的片段扩增产物中的dUTP,使所述目的片段二次片段化为300-1200bp DNA短片段;3)在所述DNA短片段两端分别连接测序接头单链A和测序接头单链B;得到连接测序接头产物;4)PCR扩增所述连接测序接头产物,得到扩增产物。

Description

一种长片段DNA文库构建方法 技术领域
本发明涉及生物技术领域,尤其涉及一种长片段DNA文库构建方法。
背景技术
长片段读取技术(Long Fragment Read,LFR),是美国Complete Genomics公司提出的DNA文库构建测序技术(Methods and compositions for long fragment read sequencing,United States Patent 8,592,150)。其通过对基因组样品进行384孔板物理分隔的方法,将来自父本与母本的DNA长片段进行分离并分别加入不同的标签序列构建文库,在测序完成后可以完全定相基因组,确认突变位点是否在同一亲本染色体上。文库构建过程中对分离至孔板中的DNA长片段进行MDA扩增,在扩增过程中,掺入dUTP或其它dNTP类似物。后续通过内切酶的作用去除dNTP类似物并使用缺口平移的方法将扩增产物打断成较小的片段,然后使用带方向性的接头通过“连接接头1-延伸反应-连接接头2”的多步反应实现接头A的连接操作。完成以上步骤后进行环化并按照CG建库的后续流程连接测序接头B并完成文库构建。长片段读取技术可以有效改善测序的准确性并且降低起始DNA的投入量。
为满足文库构建实验的需求,分入384孔板中的DNA片段需要先进行扩增处理,CG公司使用多重链置换扩增(Multiple Displace Amplification,MDA)的方法对投入384孔板每个单孔中的DNA长片段进行扩增。而MDA方法容易形成非特异扩增,其在扩增过程中被置换的单链会与新的随机引物结合而形成高级复杂结构影响后续反应。并且,接头A连接的多步酶反应操作复杂繁琐。
发明公开
本发明的一个目的是提供一种长片段DNA文库构建方法。
本发明提供的方法,包括如下步骤:
1)对长片段DNA依次进行转座酶断裂、引入dUTP扩增和去除dUTP,得到断裂片段;
2)将带有不同标签的测序接头单链A和与其部分互补的带有不同标签的测序接头单链B以单链的形式分别加入含有所述断裂片段的体系中反 应,使所述断裂片段两端连接测序接头,通过测序接头单链A和测序接头单链B中标签序列的排列组合,使每份所述断裂片段对应的测序接头均相互区别,得到连接不同测序接头的产物;
所述带有不同标签的测序接头单链A和所述带有不同标签的测序接头单链B退火可成所述测序接头;
3)以所述连接测序接头后的产物为模板,以与所述测序接头匹配的引物,进行PCR扩增,得到的PCR扩增产物为连接不同测序接头的PCR扩增产物;
4)用所述连接不同测序接头的PCR扩增产物构建文库,即得到长片段DNA文库。
所述构建文库(具体见实施例)为将所述连接测序接头产物的PCR扩增产物依次进行消化dUTP、双链环化、EcoP15酶切、末端补平、去磷酸化、第二接头连接、第二接头连接产物扩增、分离得到单链,单链环化、即得到长片段DNA文库。
上述方法中,所述步骤1)的方法包括如下步骤:
(1)用转座酶将所述长片段DNA进行断裂、并在断裂后的片段两端均连上扩增接头,得到连接扩增接头的产物;
(2)以所述连接扩增接头的产物为模板,以与所述扩增接头匹配的DNA为引物,进行第一次PCR扩增,得到第一次PCR扩增产物;所述第一次PCR扩增过程中引入dUTP;
上述PCR扩增的反应体系(没有模板):2x buffer 291.2uL,Primer B(20uM)8.48uL,Primer C(20uM)8.48uL,dNTP(各25mM)20.17uL,dUTP(4mM)5.04uL,Pfu turbo Cx聚合酶7.68uL,20%(体积百分比)Triton X-100水溶液20.8uL,10%(体积百分比)Tween20水溶液20.8uL,加入无核酶水至总体积560uL。
(3)去除所述第一次PCR扩增产物中的dUTP,得到缺口,进行缺口平移,并在3’末端加A,得到的产物为1)中所述断裂片段。
上述方法中,所述转座酶为包埋扩增接头的转座酶;
和/或,所述扩增接头为1种或2种,所述扩增接头由转座酶识别单链DNA分子和与其部分反向互补的单链DNA分子形成;
和/或,相邻两个所述包埋扩增接头的转座酶的作用位置之间的片段大小为3-10kb。
和/或,(2)中与所述扩增接头匹配的DNA为引物为连接在所述断裂后的片段两端的扩增接头中与所述转座酶识别单链DNA分子反向互补的单链DNA分子中,除所述转座酶识别单链DNA分子反向互补序列外,剩余部分序列形成的引物对。
上述方法中,
所述扩增接头为接头1和接头2,
所述扩增接头为接头1和接头2,所述接头1为由转座酶识别单链DNA分子A和与其部分反向互补的单链DNA分子B组成,所述接头2为由所述转座酶识别单链DNA分子A和与其部分反向互补的单链DNA分子C组成;
所述与所述扩增接头匹配DNA分子作为引物由引物B和引物C组成,所述引物B为所述单链DNA分子B中除去与所述转座酶识别单链DNA分子A互补部分外的其余序列;所述引物C为所述单链DNA分子C中除去与所述转座酶识别单链DNA分子A互补部分外的其余序列;
和/或,去除所述第一次PCR扩增产物中的dUTP的方法包括如下:用尿嘧啶DNA糖基化酶和人源脱嘌呤脱嘧啶核酸内切酶对所述第一次PCR扩增产物进行酶切反应,得到酶切产物;再用聚合酶I、Taq聚合酶和dATP对所述酶切产物进行聚合反应,得到300-1200bp大小的DNA片段。
上述酶切反应体系(不含有模板):UDG(2U/uL)14.56uL,APE 1(10U/uL)2.9uL,加无核酶水将总体积补充至560uL。
上述聚合反应体系(不含有模板):聚合酶I(10U/uL)2.86uL,Taq聚合酶(5U/uL)5.7uL,dATP(100mM)50.4uL,加无核酶水将总体积补充至560uL。
上述方法中,所述标签序列是由n个碱基经过排列组合得到,所述碱基为A、G、C、T中至少一种,n大于等于8。
上述方法中,所述带有不同标签的测序接头单链A从5’至3’方向依次包括片段甲、片段乙、标签序列和片段丙;
所述带有不同标签的测序接头单链B从5’至3’方向依次包括与所述片段丙反向互补的片段、标签序列、片段丁和与所述片段甲反向互补的 片段;
所述步骤3)中,所述引物中的一条引物与所述测序接头单链A中的片段乙匹配(相同或反向互补),所述引物中的另一条引物与单链B中的片段丁匹配(反向互补或相同)。
上述方法中,所述带有不同标签的测序接头单链A的条数为小于等于72,且各条单链标签序列不同;
所述带有不同标签的测序接头单链B的条数为小于等于72,且各条单链标签序列不同;
所述标签序列中n大于等于8且小于15。
上述方法中,步骤2)中,所述将带有标签的测序接头单链A和与其互补的带有不同标签的测序接头单链B以单链的形式分别加入含有所述断裂片段的体系中反应包括如下步骤:
2)-A,将72条所述带有不同标签序列的测序接头单链A分别加入含有所述断裂片段体系的芯片的纵向72列中;所述测序接头单链A反应缓冲液不含T4连接酶;
2)-B,将72条所述带有不同标签序列的测序接头单链B分别加入2)-A处理后的芯片的横向72行中,反应,得到连接测序接头产物。所述测序接头单链B反应缓冲液含有T4连接酶;
上述方法中,步骤3)中,所述引物中的一条引物与所述测序接头单链A中的片段乙相同或反向互补;
所述引物中的另一条引物与所述测序接头单链B中的片段丁的片段反向互补或相同。
所述PCR扩增是将5184孔板上各孔中所述连接不同测序接头的产物混合一起进行扩增;
上述方法中,所述方法的步骤1)的(2)、(3)和步骤2)均在5184孔芯片中进行。所述在所述断裂片段被分成若干份的情况下,是指将所述断裂片段分装在5184孔芯片的各孔中,分装可在步骤1)或步骤2)中进行。
上述方法中,所述长片段DNA为大于100Kb的片段。具体为大小为400kb的片段;
所述转座酶为Tn5转座酶;本实施例中,包埋扩增接头的转座酶为Vazyme公司产品,命名为TruePrep mini DNA Sample Prep Kit(S302-01-B)转座酶试剂盒,最佳使用浓度为100倍稀释转座酶。包埋扩增接头的转座酶和长片段DNA的量见如下反应体系中,上述反应的反应体系包括所述单链DNA分子B、所述单链DNA分子C、反应缓冲液、dNTP、dUTP和聚合酶;2uL 100倍稀释后包埋扩增接头的转座酶、2uL 5x buffer(S302-01,Vazyme公司)、7ng人基因组DNA(长度为大于100kb),加水补充至10uL。
上述方法中,所述方法的步骤1)的(2)、(3)和步骤2)中均采用wafergen MSND移液平台将各种物质加入芯片的孔中。
上述方法中,在所述用wafergen MSND移液平台将各种物质加入芯片的孔中后,还包括如下步骤:离心沉降各物质。
上述方法中,步骤1)的(1)和(2)中,还包括如下步骤:用变性试剂消化所述转座酶;所述变性试剂具体采用0.1-1%SDS溶液;所述消化的条件具体为25℃静置10分钟。
所述转座酶识别单链DNA分子A的核苷酸序列为序列表中序列5;
所述单链DNA分子B的核苷酸序列为序列表中序列6;
所述单链DNA分子C的核苷酸序列为序列表中序列7;
所述引物B的核苷酸序列为序列表中序列12;
所述引物C的核苷酸序列为序列表中序列13;
所述测序接头单链A的核苷酸序列为序列表中序列1或序列8;
所述测序接头单链B的核苷酸序列为序列表中序列2或序列9;
所述单链引物甲(一条引物)的核苷酸序列为序列表中序列3或序列10;
所述单链引物乙(另一条引物)的核苷酸序列为序列表中序列4或序列11。
上述方法中,步骤1)的(1)和(2)中,还包括如下步骤:用变性试剂消化所述转座酶;所述变性试剂具体采用0.1-1%SDS溶液;所述消化的条件具体为25℃静置10分钟。上述变性试剂为含有SDS溶液,具体由11.2uL 1%SDS(质量/体积百分比)和548.8uL 2x buffer(MP01137, Complete Genomics)组成。
上述方法中,步骤1)的(1)中,所述转座酶断裂反应的条件为55℃5分钟;
步骤1)的(2)中,所述扩增的退火温度为68℃,退火时间为18min;
步骤1)的(3)中,所述酶切反应条件为37℃2小时,再65℃15分钟;所述聚合反应条件为23℃1小时,65℃30分钟;
步骤2)的2)-B中,所述反应条件为20℃2小时;
步骤3)中,所述PCR扩增的退火温度为68℃,退火时间为10min。
由上述方法制备得到的长片段DNA文库也是本发明保护的范围。
本发明的另一个目的是提供一种用于长片段DNA文库构建的连接测序接头产物制备方法。
本发明提供的方法,包括如下步骤:为上述方法中的步骤1)-2)。
上述的方法在构建长片段DNA文库中的应用也是本发明保护的范围。
本发明的变性试剂可以为NTbuffer、SDS或其他可以使转座酶变性并从DNA长片段中脱落的试剂。
附图说明
图1为转座酶打断基因组片段示意图。
图2为wafergen MSND移液平台芯片示意图。
图3为转座酶在与基因组DNA结合示意图。
图4为转座酶脱落后加入与转座酶插入的接头1/2相匹配的引物序列进行扩增。
图5为dUTP在UDG酶及APE1酶的作用下被从PCR产物中切除且在产物上留下1bp的缺口。
图6为依靠dUTP的消化所产生的缺口平移。
图7为第一接头的各自带有标签序列的两条单链的加入方式示意图。
图8为第一接头的两条单链与插入片段连接方式示意图。
图9为第一接头连接完成后PCR扩增反应示意图。
图10为wafergen MSND两种喷液方式示意图。
图11为转座酶片段打断后PCR产物电泳胶图。
图12为芯片建库过程各步骤产物电泳图。
图13为EcoP15内切酶切割产物。
图14为单链环化后电泳图。
图15为各孔中的数据量及对应频数的直方图。
图16测序深度分布曲线图。
实施发明的最佳方式
下述实施例中所使用的实验方法如无特殊说明,均为常规方法。
下述实施例中所用的材料、试剂等,如无特殊说明,均可从商业途径得到。
本发明通过利用转座酶插入特异序列,借助该特异序列在带有5184孔的芯片中进行PCR,最终实现文库构建。
下述实施例中借助wafergen公司的MSND移液平台及带有5184孔的芯片,每次可以对样品或反应液进行不同方式的喷液,具体举例如下表1:
表1 为不同方式的喷液
Figure PCTCN2016079278-appb-000001
实施例1、构建长片段DNA文库(适配Complete Genomics测序平台)
一、通过转座酶使长片段DNA断开为大小为3-10kb的目的片段
图1为转座酶打断基因组片段示意图,包埋有接头1和接头2的转座酶,在与基因组DNA混合后,随机与基因组中的位置结合,通过控制转座酶的使用量,控制相邻两个转座酶作用位置之间的片段大小约为3~10kb之间。
图2为wafergen MSND移液平台芯片示意图。该芯片横向共72行,纵向共72列,总微孔数量为5184个。
转座酶包埋1种或2种接头都可以。
本实施例采用的包埋扩增接头的转座酶包埋有两种接头,为将转座酶 识别单链DNA分子A和与A部分反向互补的单链DNA分子B退火得到的接头、转座酶识别单链DNA分子A和与A部分反向互补的单链DNA分子C退火得到的接头。这两种接头按各100uM的浓度等体积混合,然后与转座酶混合并进行包埋得到。本实施例采用的包埋扩增接头的转座酶为Vazyme公司产品Tn5转座酶,命名为TruePrep mini DNA Sample Prep Kit(S302-01-B)转座酶试剂盒。
上述序列信息如下:
转座酶识别单链DNA分子A:5’-CTGTCTCTTATACACATCT-3’(序列5)
单链DNA分子B:5’-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3’(序列6)
单链DNA分子C:5’-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3’(序列7)
在离心管中,控制包埋有接头序列的Tn5转座酶的使用量对基因组DNA进行作用,使两个相邻的转座酶作用位点在同一条DNA片段相距较长距离(约3-10K)。由于转座酶本身的特性,在作用完成后并不会立刻脱离DNA,而是仍然存在于作用位点,连接住被打断的长片段DNA,故此时整条DNA长片段仍然保持原有长度而非断开状态。
将上述转座酶作用产物稀释至一定浓度(使DNA长片段间可以实现充分分离),然后利用wafergen的微量移液平台转移到芯片之中,实现DNA长片段的物理分隔。该芯片带有5184个微孔,即最多可将单个DNA样品分离成5184份。后续步骤加入的试剂均使用该移液平台进行。
1、摸索转座酶使用浓度
利用转座酶介导PCR反应,获得实验需求的片段长度3-10kb。
将包埋扩增接头的转座酶分别稀释50倍、100倍和150倍三个梯度。
在PCR反应管中加入如下反应体系:5x buffer 2uL(S302-01,Vazyme公司),人基因组DNA(从离体血液细胞中提取得到基因组DNA;7ng,大于100kb)6uL,转座酶(不同稀释倍数)2uL。
将上述加有不同稀释倍数转座酶的反应体系混合后分别放入PCR仪中,55摄氏度5分钟反应,然后冷却至常温,反应混合物加入终浓度0.1% 的变性剂SDS混合,在室温下静置10分钟,得到大小为3-10kb的目的片段。然后将大小为3-10kb的目的片段稀释至终浓度约8pg/uL。
2)PCR反应
将上述1)得到的大小为3-10kb的目的片段加入如下PCR反应体系进行PCR扩增,得到目的片段扩增产物。
用于扩增的引物为连接在所述断裂后的片段两端的扩增接头中与所述转座酶识别单链DNA分子反向互补的单链DNA分子中,除与所述转座酶识别单链DNA分子反向互补序列外,剩余部分序列形成的引物对,具体如下:
引物B:5’-TCG
Figure PCTCN2016079278-appb-000002
CGGCAGCGTC-3’(序列12)
引物C:5’-GTCTCGTGGGCTCGG-3’(序列13)
上述100uL PCR反应体系包括:大小为3-10kb的目的片段(8pg/uL)5uL、2x buffer 50uL、引物B(20uM)0.5uL、引物C(20uM)0.5uL、DNA聚合酶0.5uL,补水至100uL。
使用表2所示的PCR程序进行扩增。
表2 为PCR程序
Figure PCTCN2016079278-appb-000003
不同稀释倍数的转座酶结果如下:
100倍稀释转座酶条件下目的片段扩增产物的结果所示,可以看出,得到3-10kb的扩增产物,为所需目的片段。
50倍稀释转座酶条件下和150倍稀释转座酶条件下均未得到3-10kb的目的片段扩增产物。
可以看出,最优转座酶稀释倍数为100倍稀释。
2、最优转座酶稀释浓度片段化基因组DNA目的片段
采用上述1摸索的最优100倍稀释转座酶对基因组DNA进行处理,使 基因组DNA约3-10K左右距离间加入转座酶包埋带有的PCR接头序列,并以此接头序列为起点进行PCR扩增,具体如下:
1)、转座酶片段化DNA
将包埋扩增接头的转座酶和人基因组DNA按照如下反应体系进行反应,具体如下:
上述反应体系如下:2uL 100倍稀释后包埋扩增接头的转座酶、2uL 5x buffer(S302-01,Vazyme公司)、7ng人基因组DNA(长度为400kb),加水补充至10uL。
反应体系可按照比例扩大或缩小。
上述反应为:反应体系放入PCR仪中,55摄氏度5分钟反应,然后冷却至常温,得到反应产物,将其稀释至终浓度约8pg/uL,得到稀释后的反应产物(每3-10KD插入转座酶的基因组DNA)。
使用wafergen MSND移液平台单样品分液流程将稀释后的片段化产物按照35nL/孔分入5184孔板各孔中(孔的溶剂为200nl);再对芯片进行离心(4000rpm,5min),使所喷液体沉降至芯片底部,得到含有反应产物的5184孔板。
图3为转座酶在与基因组DNA结合并插入接头1/2后,并不会立刻脱离而是附着在其作用位点中。在反应溶液中加入一定浓度的SDS,使得转座酶从DNA中脱落。
2)、消化转座酶
使用wafergen MSND移液平台单样品分液流程将变性试剂加入上述1)得到的含有反应产物的5184孔板各孔中,再对该芯片进行离心(时间4000rpm,5min),使所喷液体沉降至芯片底部;离心完毕后室温下静置10分钟左右,用于消化转座酶,使其脱离基因组DNA并实现基因组DNA的片段化,得到3-10KD目的片段,含有其的芯片为含有3-10KD目的片段的5184孔芯片。
本操作可使用其它变性试剂及浓度,不限于SDS。
上述变性试剂由11.2uL 1%SDS(质量/体积百分比)和548.8uL 2x buffer组成;
3)、扩增目的片段得到目的片段扩增产物
图4为在完成转座酶的脱落后,在PCR体系中加入引物B和引物C进行扩增。同时,在PCR体系中投入一定比例的dUTP(4%,dUTP:dATP+dTTP+dCTP+dGTP),使扩增产物中掺入一定比例的dUTP。
使用设计好的引物,对上一步所得的产物进行PCR。在PCR体系之中掺入dUTP(dUTP:dNTP=4:100),使最终产物中部分的dTTP被dUTP代替,具体如下:
使用wafergen MSND移液平台单样品分液流程将下述PCR反应体系加入步骤2)得到的含有3-10KD目的片段的5184孔芯片各孔中,之后离心使所喷液体沉降,再将芯片板放入PCR仪中进行PCR扩增反应,得到含有dUTP的目的片段扩增产物的5184孔芯片。
在上述反应体系中加入一定比例的dUTP,使得在扩增中一部分的dTTP被dUTP替代,从而使扩增产物在后期可通过dUTP的去除而被再次打断。
上述PCR扩增反应体系:2x buffer 291.2uL,单链DNA分子B(20uM)8.48uL,单链DNA分子C(20uM)8.48uL,dNTP(各25mM)20.17uL,dUTP(4mM)5.04uL,Pfu turbo Cx聚合酶7.68uL,20%(体积百分比)Triton X-100水溶液20.8uL,10%(体积百分比)Tween20水溶液20.8uL,加入无核酶水至总体积560uL。
上述PCR扩增反应程序如表3所示:
表3 为PCR扩增反应程序
Figure PCTCN2016079278-appb-000004
图11为转座酶片段打断后PCR产物电泳胶图。
反应完毕后,将芯片放置于37℃金属浴中静置15.5分钟蒸干多余液体。
二、3-10KD目的片段二次片段化为300-1200bp DNA短片段
1)去除dUTP形成单链缺口
使用具有单独去除dUTP活性的酶(UDG、USER等),对上步PCR产物作用,去除其中的dUTP及该位置的糖骨架,得到中间带有一些1个碱基长度的缺口的DNA双链产物(图5),具体如下:
使用尿嘧啶DNA糖基化酶(UDG)及人源脱嘌呤脱嘧啶核酸内切酶(APE1)消化扩增中引入的dUTP。
具体如下:
按照wafergen MSND移液平台单样品分液流程将酶切反应体系加入上述一得到的含有dUTP的目的片段扩增产物的5184孔芯片的各孔中,之后离心使所喷液体沉降(4000rpm,5min);再将芯片放入专用PCR仪中,进行酶切反应,得到含有酶切产物的5184孔芯片。
上述酶切反应体系:UDG(2U/uL)14.56uL,APE 1(10U/uL)2.9uL,加无核酶水将总体积补充至560uL。
上述酶切反应条件为,37℃2小时,65℃15分钟,之后恢复至常温放置。
2)聚合反应完成断裂
缺口平移打断DNA,并在末端加A。在反应体系中加入DNA聚合酶I和Taq酶。在23摄氏度下,利用DNA聚合酶I的作用活性(5’-3’外切酶活性和5’-3’聚合酶活性),对上一步所得产物进行打断操作。DNA聚合酶I结合于去除dUTP后留下的缺口处,然后从5’-3’方向切除DNA碱基,并从5’-3’方向进行DNA碱基的聚合,实现缺口从5’-3’方向的平移。当DNA双链的两个缺口因为平移而相遇时,该DNA双链便会断开成更小的片段。同时在体系中加入的Taq聚合酶,在65摄氏度下对DNA双链进行补平,并在3’末端加入一个dATP碱基。而在该温度下,DNA聚合酶I因变性而失去活性。
图6为依靠dUTP的消化所产生的缺口,后来投入到反应体系中的DNA聚合酶I结合在该缺口位置发挥其5'-3'方向切割及5'-3'方向合成的作用,使得该缺口沿5'-3'方向平移。双链中位于正反两条链上的缺口同时进行平移并相遇,最终在相遇位置使整个双链断开。而同时加入到体系中的Taq酶,在65℃的温度下发挥其活性,使得产物的3’末端形成突出一个A碱基。
利用聚合酶缺口平移的特性,依靠上述去除dUTP后的缺口进行缺口平移实现DNA片段打断,同时在反应体系中加入Taq聚合酶,在反应中使缺口平移后的DNA产物的3'末端加入dATP,具体如下:
按照wafergen MSND移液平台单样品分液流程将聚合反应体系喷液至上述1)含有酶切产物的5184孔芯片各孔中,之后离心使所喷液体沉降(4000rpm,5min);再将芯片放入专用PCR仪中,进行聚合反应,得到含有300-1200bp大小的DNA片段5184孔芯片。
上述聚合反应体系:聚合酶I(10U/uL)2.86uL,Taq聚合酶(5U/uL)5.7uL,dATP(100mM)50.4uL,加无核酶水将总体积补充至560uL。
上述反应条件为,23℃1小时,65℃30分钟,之后恢复至室温。
三、连接测序接头
在上述二得到的300-1200bp大小的DNA片段两端连接带有部分测序引物的测序接头,测序接头由带有不同标签的测序接头单链A和带有不同标签的测序接头单链B退火形成,为了对芯片上每一个孔都可以在测序的过程中进行区分,设计在芯片中纵向的每一列(对应测序接头单链A)加入编号为1-72的带有不同标签测序接头单链A,在芯片中横向的每一行(对应测序接头单链B)加入编号为1-72的带有不同标签测序接头单链B,这样通过两个单独的标签序列便形成72x72的矩阵,使得每一个孔都获得特有的双标签序列组合。
测序接头单链A
Figure PCTCN2016079278-appb-000005
AAA
Figure PCTCN2016079278-appb-000006
CAGTACG
Figure PCTCN2016079278-appb-000007
T(序列1)
其中,从5’至3’方向所有的粗体如下:
片段甲
Figure PCTCN2016079278-appb-000008
测序接头单链B3’端的
Figure PCTCN2016079278-appb-000009
片段乙
Figure PCTCN2016079278-appb-000010
序列3所示的PCR引物第一链相同,且仅将序列3的U替换为T;
NNNNNNNNNN为10个碱基组成标签序列;
片段丙
Figure PCTCN2016079278-appb-000011
测序接头单链B5’端的
Figure PCTCN2016079278-appb-000012
*代表的是硫代磷酸,即在最后一个碱基合成时,使用的是硫代磷酸的核酸。
测序接头单链B
/5Phos/
Figure PCTCN2016079278-appb-000013
Figure PCTCN2016079278-appb-000014
/3’AmMO/(序列2)
其中,从5’至3’方向所有的粗体如下:
Figure PCTCN2016079278-appb-000015
测序接头单链A3’端的片段丙
Figure PCTCN2016079278-appb-000016
Figure PCTCN2016079278-appb-000017
为10个碱基组成标签序列;
Figure PCTCN2016079278-appb-000018
序列4所示的PCR引物第二链反向互补,且仅将序列4的U替换为T;
Figure PCTCN2016079278-appb-000019
测序接头单链A5’端的片段甲
Figure PCTCN2016079278-appb-000020
图7为第一测序接头的各自带有标签序列的两条单链的加入方式示意图。为了使芯片中每一个微孔均带有特有的标签序列以进行相互区分,采用两个标签序列组成一个组合的形式。在实验中第一接头的两条单链分别加入。其中第一链的分别72种标签序列的单链(测序接头单链A),以纵向形式加入,即每一纵列加入的第一链的标签序列相同。第二链的分别72种标签序列的单链(测序接头单链B),以横向形式加入,即每一横行加入的第二链的标签序列相同。这样最终形成72x72中标签序列的矩阵,每一个孔获得唯一的两两标签序列组合。
为了使用两个单独的标签序列实现72x72种标签序列组合,在连接接头的步骤中,带有标签序列的测序接头单链A以纵向方式加入,带有标签序列的测序接头单链B以横向方式加入,具体如下:
1、测序接头单链A的添加
配制测序接头单链A测序接头单链A反应液如下:10x buffer(25%(体积比)PEG8000、500mM Tris-HCl、10mM ATP、100mM MaCl2)252uL,加无核酶水将总体积补充至957uL;
1)按照wafergen MSND移液平台“72样品35nL分液流程”要求将上述测序接头单链A反应液以每孔11.96uL加入到上述二得到的含有300-1200bp大小的DNA片段5184孔芯片的纵向72列孔中;
2)再将72种不同的测序接头单链A测序接头单链A(10uM)以每孔2.14uL的量按照wafergen MSND移液平台“72样品35nL分液流程”加入上述1)处理后的5184孔板的纵向72列孔中,离心沉降液体,备用, 得到添加测序接头单链A测序接头单链A的5184孔板。
2、测序接头单链B的添加
配制测序接头单链B测序接头单链B反应液:10xbuffer(25%(体积比)PEG8000、500mM Tris-HCl、10mM ATP、100mM MaCl2)296.4uL、T4连接酶(600U/uL,Enzymatics)125.97uL,加无核酶水将总体积补充至1186uL。
1)按照wafergen MSND移液平台“72样品50nL分液流程”要求将上述测序接头单链B测序接头单链B反应液以每孔13.95uL的体积加入到上述1得到的添加测序接头单链A的5184孔板的横向72行孔中;
2)再将72种不同的测序接头单链B(10uM)以每孔1.65uL的量按照wafergen MSND移液平台“72样品50nL分液流程”加入上述1)处理后的5184孔板的横向72行孔中,离心沉降液体,备用,得到添加测序接头单链B的5184孔板。
将上述添加测序接头单链B的5184孔板置于PCR仪中反应,反应条件为20℃2小时,得到含有连接测序接头产物的5184孔板。
图8为第一接头(测序接头)的两条单链与插入片段连接方式示意图。
图10为wafergen MSND两种喷液方式示意图。左图为“72样品35nL分液流程”配合测序接头单链A的加液方式,使得第一链的分别72种标签序列的单链(测序接头单链A),实现纵向形式加入。右图为“72样品50nL分液流程”配合第一接头第二链的加液方式,使得第二链的分别72种标签序列的单链(测序接头单链B),实现横向形式加入。
四、PCR扩增
图9为第一接头(测序接头)连接完成后PCR扩增反应示意图。
将上述三获得的含有连接测序接头产物的5184孔板所有孔的产物均通过离心的方式取出并混合于一离心管之中,然后进行PCR反应扩增,扩增过程会将接头不匹配的部分重新变成匹配的双链。由于产物同一条单链两端是带有不同标签序列的两条接头单链,所以即使PCR后接头的形状进行改变,但不会影响其标签序列组合。也因此,各个孔的产物混合在一起,也不会影响最后对每一个标签序列的组合的拆分,具体如下:
将上述三获得的含有连接测序接头产物的5184孔板所有孔的产物离 心收集于1.5mL离心管中,并使用1倍AmpureXP磁珠进行纯化,回收溶于100uL TE溶液中纯化产物,得到纯化后连接测序接头产物,对其进行PCR扩增反应,得到连接测序接头产物的PCR扩增产物。
上述PCR扩增采用的引物如下:
PCR引物第一链(上游引物):GGUCGCCAGCCCUATGGC(序列3)
PCR引物第二链(下游引物):AGGGCUGGCGACCUTGTCAG(序列4)
上述PCR扩增采用的反应体系如下:连接测序接头产物100uL,2x PfuCx buffer 275uL,PCR引物第一链(20uM)14uL,PCR引物第二链(20uM)14uL,PfuCx聚合酶11uL,加无核酶水补充至550uL。
上述PCR扩增的反应程序如下表4。
表4 为扩增反应程序
Figure PCTCN2016079278-appb-000021
将连接测序接头产物的PCR扩增产物使用1倍AmpureXP磁珠进行纯化,纯化产物溶于60uL TE溶液中。
上述各步骤反应产物电泳检测如下图12所示,1、长片段PCR扩增产物(步骤一的2的3)得到的扩增产物);2、长片段PCR阴性扩增产物(步骤一的2的3)得到的扩增产物,模板用水替代基因组DNA);3、去除dUTP产物(步骤二的1));4、去除dUTP反应阴性产物(步骤二的1),含有dUTP的目的片段扩增产物替换为水);5、缺口平移末端加A产物(步骤二的2)聚合反应产物);6、缺口平移末端加A阴性产物(步骤二的2)聚合反应中的酶切产物替换为水);7、接头连接产物1(步骤三的2的连接测序接头产物);8接头连接产物2(步骤三的2的连接测序接头产物);9、PCR产物1(步骤四的PCR扩增产物);10、PCR产物2(步骤四的PCR扩增产物);11、PCR阴性产物;M2、100bp分子量标记。从图12可以看 出各步骤得到的目标产物。
五、消化扩增产物的dUTP
上述连接测序接头产物的PCR扩增产物需要经过消化扩增引物中带有的dUTP,形成粘性末端,从而在后续实验中实现自身环化。dUTP消化反应如下,10x Taq buffer 11uL(RM00059,Complete Genomics),User enzyme(RM00017,Complete Genomics),纯化后PCR产物60uL,加无核酶水补充至110uL,37℃反应1小时,得到反应产物。
六、双链环化
上述五得到的反应产物加入10xTAbuffer(RM06601,Complete Genomics)180uL,加无核酶水至1810uL,混合物平均分成4管,于水浴60℃反应30分钟后,转至常温水浴20分钟。
预先配制反应混合物,无核酶水98uL,20x Circ mix(MP01134,Complete,Genomics)100uL,连接酶(L603-HC-1500,Enzymatics)2uL,混合后平均分成4份加入上述反应产物中混合,室温反应1小时以环化,得到环化产物。
对每一管反应产物,使用AmpureXP磁珠纯化回收,产物回收溶于70uL TE。
七、消化未环化DNA
环化反应中存在相当一部分未环化DNA,为了不影响后续实验结果,需要对未环化的DNA进行消化。消化反应如下,对于每管上述六得到的纯化环化产物,加入9x PS mix(MP01154,Complete Genomics)8.9uL,Plasmid-Safe(RM02046,Complete Genomics)10.4uL,加无核酶水补充至80uL,混合后于37℃反应1小时。反应产物使用AmpureXP磁珠纯化后溶于40uL TE溶液,得到消化后环化的产物。
八、使用EcoP15酶切并筛选目的片段
环化后的产物,最初加入的两端接头连接在一起,此处使用EcoP15酶切,使环化产物由接头两端的EcoP15限制酶切位点处向插入片段内部两端各切下约27bp,后续对酶切下来的产物进行片段筛选,获得中间带接头序列,两端带有酶切下来的插入片段的产物,具体如下:
酶切消化体系配制,上述七得到的消化后环化的产物37uL,5xEcoP 15 mix 3(MP01149,Complete Genomics)72uL,EcoP 15(RM00063,Complete Genomics)10.8uL,加无核酶水补充至360uL,于37℃反应16小时。
配制PEG32磁珠(按体积比混合,AmpureXP磁珠:0.5%Tween=100:1,下同)
对于上述360uL酶反应产物,加入252uL PEG32磁珠结合,结合产物保留上清,并加入PEG32磁珠184uL,混合结合后回收磁珠并溶于52uL TE回收液(带有0.001%Tween20,体积比),得到酶切产物。
图13为EcoP15内切酶切割,片段回收后产物。产物片段在140bp左右。
九、末端补平
对EcoP15酶切产物进行末端补平,以连接后续第二接头。
配制末端补平反应体系,取上述八得到的酶切产物44uL,10x NEB buffer 2(New England biolabs)5.4uL,25mM dNTP 0.8uL,10mg/mL BSA 0.4uL,T4 DNA polymerase(M0203-com,New England Biolabs),12℃反应20分钟,得到末端补平产物。
反应产物使用PEG32磁珠进行纯化并溶于48uL TE。
十、去磷酸化
对上述九得到的末端补平产物进行去磷酸化处理,以配合后续第二接头连接。配制反应体系如下,10x NEB buffer 2(New England Biolabs)5.75uL,Fast AP(EF0651,Fermentas)5.75uL,加入末端补平纯化后产物46uL,于37℃反应45分钟。反应产物使用75uL PEG32磁珠进行纯化,并溶于42uL TE溶液,得到去磷酸化产物。
十一、第二接头连接
第二接头进行定向连接,需要经过4步酶反应,先将上述十得到的去磷酸化产物进行3’末端接头序列的引入,再次进行末端补平并末端引入dATP后,连接引入5’末端接头序列,最后使用一段寡核苷酸序列替代其中一段序列并使用连接酶连接。
接头序列
3’末端接头
/phos/GTCTCCAGTCGAAGCCCGACG/3ddC/
/3ddC/AGAGGTCAGCTTCG
5’末端接头
             TTGGCCTCCGACT/3dT-Q/
/ddC/TGCTGGCGAACCGGAGGCTGA/5phos/
序列替换的两条序列
寡核苷酸序列1
/52bio/TCCTAAGACCGCTTGGCCTCCGACT
寡核苷酸序列2
/5phos/AGACAAGCTCGAGCTCGAGCGATCGGGCTTCGACTGGAGAC
PCR引物序列
PCR引物1
/52bio/TCCTAAGACCGCTTGGCCTCCGACT
PCR引物2
/5phos/AGACAAGCTCGAGCTCGAGCGATCGGGCTTCGACTGGAGAC
第二接头3’末端序列引入
配制反应混合液,3xHB(MP01139,Complete Genomics)24.7uL,无核酶水1.9uL,T4连接酶(600U/uL,Enzymatics)1.9uL,二接头3’端接头5.6uL(9uM),加入去磷酸化的产物40uL。混合物于14℃反应2小时,反应完成使用63uL PEG32磁珠进行纯化,并溶于42uL TE溶液。
末端补平加A
配制反应混合液,5x klex NTA mix(MP01150,Complete Genomics)10.7uL,klenow(RM00066,Complete Genomics)1.1uL,无核酶水1.5uL,加入上步骤产物40uL。混合物于37℃反应1小时。反应完成使用69uL PEG32磁珠进行纯化,并溶于42uL TE溶液。
第二接头3’末端序列引入
配制反应混合液,3xHB 24.7uL,无核酶水1.9uL,T4连接酶1.9uL,二接头5’端接头5.6uL(9uM),加入上步骤产物42uL。混合物于14℃ 反应2小时,反应完成使用63uL PEG32磁珠进行纯化,并溶于42uL TE溶液。
序列替换
配制Oligo反应液,10x Taq buffer(RM00059,Complete Genomics)8uL,100mM ATP(MP01164,Complete Genomics)0.8uL,25mM dNTP 0.32uL,寡核苷酸序列1(20mM)1uL,寡核苷酸序列2(20mM)1uL,加入无核酶水至32uL。
配制酶反应液,10x Taq buffer 0.4uL,T4连接酶4.8uL,Taq聚合酶4.8uL。
实验中将32uL Oligo反应液加入到40uL上述纯化后DNA产物中,然后加入8uL酶反应液,混合后于37℃中孵育20分钟。反应完成使用80uL PEG32磁珠进行纯化,并溶于47uL TE溶液。纯化产物进行定量检测。
得到第二接头连接产物。
十二、第二接头连接产物扩增
往一1.5mL离心管中,加入2xPfuCx mix3 275uL,PfuCx聚合酶11uL,第二接头PCR引物1(20uM)13.75uL,第二接头PCR引物2(20uM)13.75uL,定量后的上述十一制备的第二接头连接产物60ng,加无核酶水补充至550uL。反应物混合后平均分至4个PCR管中进行PCR扩增反应,得到扩增产物。
扩增反应程序如下表5。
表5 为扩增反应程序
Figure PCTCN2016079278-appb-000022
反应完成使用550uL PEG32磁珠进行纯化,并溶于85uL TE溶液。对纯化产物进行定量检测,取600ng进行后续单链环化步骤。
十三、单链环化
配制1xBWB/tween混合液,往离心管中加入1xBWB(MP01111,Complete Genomics),0.5%Tween20 20uL,充分混合。
清洗磁珠。取stretavidin beads(MP01162,Complete Genomics)120uL,放置于在磁力架上使磁珠充分吸附后,去除上清液。使用1xBBB(MP01110,Complete Genomics)600uL清洗两次,然后加入120uL的1xBBB使磁珠重新悬浮,然后加入1%体积的0.5%Tween20。
使用清洗后的磁珠对产物进行单链分离。取上述十二得到的扩增产物600ng,加入TE补充至60uL,加入4xBBB(MP01145,Complete Genomics)20uL,30uL清洗后的stretavidin beads,充分混合后静置15分钟,然后放置于磁力架上使磁珠充分吸附,使用配制好的1xBWB/tween混合液清洗2次。清洗完成后加入75uL 0.1M NaOH溶液,混合放置2分钟后,放置于磁力架上使磁珠充分吸附,回收上清液。最后往回收的上清液中加入0.3M MOPS acid(MP01165,Complete Genomics)37.5uL,混合均匀。
上述步骤结束后获得112.5uL单链产物,产物进入单链环化步骤。预先配制以下反应液:
引物混合物,取Bridge Oligo(20uM)20uL,加入43uL无核酶水,混合均匀待用;
酶反应混合物,取无核酶水135.3uL,加入10xTA buffer(RM06601,Complete Genomics)35uL,100mM ATP 3.5uL,连接酶(600U/uL)1.2uL,混合均匀待用。
往上述112.5uL产物中,加入引物混合物63uL,酶反应混合物175uL,混合均匀后,置于37℃孵育1.5小时,得到单链环化产物。
图14为单链环化后电泳图。
十四、消化未环化单链产物
往上述十三得到的单链环化产物中,加入无核酶水1.5uL,10xTA buffer 3.7uL,Exo I(M0293L,New England Biolabs)11.1uL,Exo III(M0206L,New England Biolabs)3.7uL,充分混合后,置于37℃反应30分钟。反应完成后加入500mm EDTA 15.4uL,充分混合以终止反应。反应混合物加入500uL PEG32磁珠进行纯化,最终溶解于16uL TE溶液,得到长 片段文库。
文库制备过程结束。
实施例2、文库构建方法(适配illumina测序平台)
一、通过转座酶使长片段DNA断开为大小为3-10kb的目的片段
与实施例1的步骤相同。
二、目的片段3-10KD目的片段二次片段化为300-1200bp DNA短片段
与实施例1的步骤相同
三、连接测序接头
在上述二得到的300-1200bp大小的DNA片段上一步反应产物的两端连接带有部分测序引物的测序接头,该步骤的接头双链分别带有不同的标签序列,为了对芯片上每一个孔都可以在测序的过程中进行区分,设计在芯片中纵向的每一列(对应测序接头第一条链)加入编号为1-72的标签序列,在芯片中横向的每一行(对应接头第二条链)加入编号为1-72的标签序列。这样通过两个单独的标签序列便形成72x72的矩阵,使得每一个孔都获得特有的双标签序列组合。在上一步反应产物两端连接带有部分测序引物序列的接头。
测序接头单链A
Figure PCTCN2016079278-appb-000023
NNNNNNNN
Figure PCTCN2016079278-appb-000024
Figure PCTCN2016079278-appb-000025
(序列8)
从5’至3’方向依次由Flow cell接头P5(斜体)、10个碱基组成标签序列(N)和与Read 1测序接头互补的单链DNA分子(加粗,*代表的是硫代磷酸,即在最后一个碱基合成时,使用的是硫代磷酸的核酸)组成。
测序接头单链B
/5Phos/
Figure PCTCN2016079278-appb-000026
NNNNNNNN
Figure PCTCN2016079278-appb-000027
Figure PCTCN2016079278-appb-000028
/3’AmMO/(序列9)
从5’至3’方向依次由Read 2测序接头(加粗)、10个碱基组成标签序列(N)和与Flow cell接头P7互补的单链DNA分子(斜体,Phos即磷酸修饰,AmMo即氨基修饰)组成;72条3’端标签引物的10个碱基组成随机片段均不相同。
图7为第一接头的各自带有标签序列的两条单链的加入方式示意图。 为了使芯片中每一个微孔均带有特有的标签序列以进行相互区分,采用两个标签序列组成一个组合的形式。在实验中第一接头的两条单链分别加入。其中第一链的分别72种标签序列的单链,以纵向形式加入,即每一纵列加入的第一链的标签序列相同。第二链的分别72种标签序列的单链,以横向形式加入,即每一横行加入的第二链的标签序列相同。这样最终形成72x72中标签序列的矩阵,每一个孔获得唯一的两两标签序列组合。
为了使用两个单独的标签序列实现72x72种标签序列组合,在连接接头的步骤中,第一个带有标签序列的接头单链以纵向方式加入,第二个带有标签序列的接头单链以横向方式加入,具体如下:
1、测序接头单链A的添加
配制测序接头单链A反应液如下:10x buffer(25%(体积比)PEG8000、500mM Tris-HCl、10mM ATP、100mM MaCl2)252uL,加无核酶水将总体积补充至957uL;
1)按照wafergen MSND移液平台“72样品35nL分液流程”要求将上述测序接头单链A反应液以每孔11.96uL加入到上述二得到的含有300-1200bp大小的DNA片段5184孔芯片的纵向72列孔中;
2)再将72种不同的测序接头单链A(10uM)以每孔2.14uL的量按照wafergen MSND移液平台“72样品35nL分液流程”加入上述1)处理后的5184孔板的纵向72列孔中,离心沉降液体,备用,得到添加测序接头单链A的5184孔板。
2、测序接头单链B的添加
配制测序接头单链B反应液:10xbuffer(25%(体积比)PEG8000、500mM Tris-HCl、10mM ATP、100mM MaCl2)296.4uL、T4连接酶(600U/uL,Enzymatics)125.97uL,加无核酶水将总体积补充至1186uL。
1)按照wafergen MSND移液平台“72样品50nL分液流程”要求将上述测序接头单链B反应液以每孔13.95uL的体积加入到上述1得到的添加测序接头单链A的5184孔板的横向72行孔中;
2)再将72种不同的测序接头单链B(10uM)以每孔1.65uL的量按照wafergen MSND移液平台“72样品50nL分液流程”加入上述1)处理后的5184孔板的横向72行孔中,离心沉降液体,备用,得到添加测序接 头单链B的5184孔板。
将上述添加测序接头单链B的5184孔板置于PCR仪中反应,反应条件为20℃2小时,得到含有连接测序接头产物的5184孔板。
图8为第一接头的两条单链与插入片段连接方式示意图。
图10为wafergen MSND两种喷液方式示意图。左图为“72样品35nL分液流程”配合第一接头第一链的加液方式,使得第一链的分别72种标签序列的单链,实现纵向形式加入。右图为“72样品50nL分液流程”配合第一接头第二链的加液方式,使得第二链的分别72种标签序列的单链,实现横向形式加入。
四、PCR扩增
图9为第一接头连接完成后PCR扩增反应示意图。
将上述三获得的含有连接测序接头产物的5184孔板所有孔的产物均通过离心的方式取出并混合于一离心管之中,然后进行PCR反应扩增,扩增过程会将接头不匹配的部分重新变成匹配的双链。由于产物同一条单链两端是带有不同的两条接头单链/标签序列的,所以即使PCR后接头的形状进行改变,但不会影响其标签序列组合。也因此,各个孔的产物混合在一起,也不会影响最后对每一个标签序列的组合的拆分,具体如下:
将上述三获得的含有连接测序接头产物的5184孔板所有孔的产物离心收集于1.5mL离心管中,并使用1倍AmpureXP磁珠进行纯化,回收溶于100uL TE溶液中纯化产物,得到纯化后连接测序接头产物,对其进行PCR扩增反应,得到连接测序接头产物的PCR扩增产物。
上述PCR扩增采用的引物如下:
PCR引物第一链(上游引物):AATGATACGGCGACCACCGAGATCT(序列10)
PCR引物第二链(下游引物):CAAGCAGAAGACGGCATACGAGAT(序列11)
上述PCR扩增采用的反应体系如下:连接测序接头产物100uL,2x PfuCx buffer 275uL,PCR引物第一链(20uM)14uL,PCR引物第二链(20uM)14uL,PfuCx聚合酶11uL,加无核酶水补充至550uL。
上述PCR扩增的反应程序如下表6。
表6 为扩增反应程序
Figure PCTCN2016079278-appb-000029
将连接测序接头产物的PCR扩增产物使用1倍AmpureXP磁珠进行纯化,纯化产物溶于60uL TE溶液中。
上述各步骤反应产物电泳检测,得到目标产物。
至此文库制备结束。
实施例3、文库测序结果分析
一、测序
使用实施例2制备的DNA长片段文库在illumina平台上进行测序,测序深度40X。
二、比对
使用序列比对软件SOAP aligner 2.20(LiR,LiY,Kristiansen K,et al,SOAP:short oligonucleotide alignment program.Bioinformatics2008,24(5):713-714;LiR,YuC,LiY,et al,SOAP2:an improved ultrafast tool for short read alignment.Bioinformatics 2009,25(15):1966-1967;
http://soap.genomics.org.cn/soapaligner.html)把读取片段(reads)比对人类参考基因组Reference hg19(http://hgdownload.cse.ucsc.edu/goldenPath/hg19/bigZips/),比对存在多重结果时只选取一个比对结果(-r 1)。
三、按照标签组合信息,确定每个孔对应的reads。
四、统计各孔中的数据量及对应频数,绘制直方图(图15)。
数据统计如下:
单位:Mb
均值:36.9000标准差:13.55647
最小值:0.3168最大值:123.2000中位数:36.7100
由上可知,本方法可以将所有5184个孔进行成功标记。
实施例4:不同文库构建方法的测序结果对比
一、文库构建和测序
使用CG公司的多重链置换扩增(Multiple Displace Amplification,MDA)的长片段DNA文库构建方法(Methods and compositions for long fragment read sequencing,United States Patent 8,592,150)制备DNA长片段文库,然后在Complete Genomincs平台上进行测序,测序深度80X。
使用实施例1制备DNA长片段文库,然后在Complete Genomincs平台上进行测序,测序深度60X。
二、比对
使用序列比对软件SOAP aligner 2.20(LiR,LiY,Kristiansen K,et al,SOAP:short oligonucleotide alignment program.Bioinformatics2008,24(5):713-714;LiR,YuC,LiY,et al,SOAP2:an improved ultrafast tool for short read alignment.Bioinformatics2009,25(15):1966-1967;
http://soap.genomics.org.cn/soapaligner.html)把测序后得到读取片段(reads)分别比对到人类参考基因组Reference hg19(http://hgdownload.cse.ucsc.edu/goldenPath/hg19/bigZips/),比对存在多重比对时只选取一个比对结果(-r 1)。
三、使用SOAP aligner 2.20中的soap.coverage计算Reads在Reference上的单点覆盖度(-cvg)。
四、绘制测序深度分布曲线图(图16)。
由图16可知,本方法的扩增较MDA扩增方法的均一性更好,更接近于常规测序方法的测序深度分布。
工业应用
本发明在LFR文库构建的原理基础上,取代原DNA扩增使用的MDA扩增方式以及接头A的连接方式,同时使用wafergen微量移液平台的芯片取代384孔板进行试验。本发明使用转座酶插入特定序列,并利用该特定 序列对每个单孔中的DNA长片段进行PCR扩增,扩增过程中为常规的“变性-退火-延伸”的方式,避免了MDA扩增形成复杂结构以及非特异扩增的问题。本发明中创造性地将两个独立的标签序列设计于双链接头的两个单链之中,然后以单链的形式分别加入,在连接的时候同时将两个标签序列连接到插入片段的两端。由于连接反应在常温下进行,本发明设计的引物在常温反应体系中会先进行退火,然后在连接酶的作用下,与插入片段的两个末端进行连接。在连接反应之前先将打断的插入片段3’末端延伸一个dATP,在连接反应中通过末端A-T配对的方式使接头的两个单链在插入片段的两端进行定向连接。
同时,本发明与wafergen微量移液平台相结合,将DNA长片段分离进入带有5184个细孔的芯片之中,单样品分离孔的数量为384孔板的10倍以上,分离效果更充分,由此获得更好的突变位点定位及组装效果。
本发明的实验证明,本发明通过与微孔芯片相结合的方式,提高了基因组中同源的长片段的分离概率,与Complete Genomics公司的384孔板分离方式相比,芯片中带有的5184孔分离效果更好。384孔板的分离,约有5%的同源长片段会被分入同一孔中而无法区分,而通过5184孔芯片的分离,该概率将进一步下降从而提升分析效率。
本发明使用长片段PCR的方式,对分离后每一个孔中的DNA片段进行扩增,取代了原方法中使用MDA扩增,因此避免了因MDA扩增导致的局部区域形成多级复杂结构的情况。在数据中,这些复杂结构所产生的数据因为无法被识别而去除,限制了所得的有效数据比例。使用长片段PCR扩增的方式,获得的数据比对效率更好。
接头连接过程中采用两接头单链分别带有标签序列并分别加入的方式,在连接过程中边进行退火边实现接头连接。相比较原方法中先进行接头连接-延伸-连接另一端接头的做法,本方法一步反应即可实现两端不同接头的连接并加入两端不同标签序列的组合,大大减少了操作的复杂度及操作时间。

Claims (18)

  1. 一种长片段DNA文库构建方法,包括如下步骤:
    1)对长片段DNA依次进行转座酶断裂、引入dUTP扩增和去除dUTP,得到断裂片段;
    2)将带有不同标签的测序接头单链A和与其部分互补的带有不同标签的测序接头单链B以单链的形式分别加入含有所述断裂片段的体系中反应,使所述断裂片段两端连接测序接头,通过测序接头单链A和测序接头单链B中标签序列的排列组合,使每份所述断裂片段对应的测序接头均相互区别,得到连接不同测序接头的产物;
    所述带有不同标签的测序接头单链A和所述带有不同标签的测序接头单链B退火可成所述测序接头;
    3)以所述连接测序接头后的产物为模板,以与所述测序接头匹配的DNA为引物,进行PCR扩增,得到的PCR扩增产物为连接不同测序接头的PCR扩增产物;
    4)用所述连接不同测序接头的PCR扩增产物构建文库,即得到长片段DNA文库。
  2. 根据权利要求1所述的方法,其特征在于:
    所述步骤1)的方法包括如下步骤:
    (1)用转座酶将所述长片段DNA进行断裂、并在断裂后的片段两端连上扩增接头,得到连接扩增接头的产物;(2)以所述连接扩增接头的产物为模板,以与所述扩增接头匹配的DNA为引物,进行第一次PCR扩增,得到第一次PCR扩增产物;所述第一次PCR扩增过程中引入dUTP;
    (3)去除所述第一次PCR扩增产物中的dUTP,得到缺口,进行缺口平移,并在3’末端加A,得到的产物为1)中所述断裂片段。
  3. 根据权利要求1或2所述的方法,其特征在于:
    步骤1)中,所述转座酶为包埋扩增接头的转座酶;
    和/或,所述扩增接头为1种或2种,所述扩增接头由转座酶识别单链DNA分子和与其部分反向互补的单链DNA分子形成;
    和/或,相邻两个所述包埋扩增接头的转座酶的作用位置之间的片段 大小为3-10kb。
  4. 根据权利要求2或3所述的方法,其特征在于:
    所述扩增接头为接头1和接头2,所述接头1为由转座酶识别单链DNA分子A和与其部分反向互补的单链DNA分子B组成,所述接头2为由所述转座酶识别单链DNA分子A和与其部分反向互补的单链DNA分子C组成;
    (2)中与所述扩增接头匹配DNA作为的引物由引物B和引物C组成,所述引物B为所述单链DNA分子B中除去与所述转座酶识别单链DNA分子A互补部分外的其余序列;所述引物C为所述单链DNA分子C中除去与所述转座酶识别单链DNA分子A互补部分外的其余序列;
    和/或,去除所述第一次PCR扩增产物中的dUTP的方法包括如下:用尿嘧啶DNA糖基化酶和人源脱嘌呤脱嘧啶核酸内切酶对所述第一次PCR扩增产物进行酶切反应,得到酶切产物;再用聚合酶I、Taq聚合酶和dATP对所述酶切产物进行聚合反应,得到300-1200bp大小的DNA片段。
  5. 根据权利要求1-4中任一所述的方法,其特征在于:
    所述标签序列是由n个碱基经过排列组合得到,所述碱基为A、G、C、T中至少一种,n大于等于8。
  6. 根据权利要求1-4中任一所述的方法,其特征在于:
    步骤2)中,所述带有不同标签的测序接头单链A从5’至3’方向依次包括片段甲、片段乙、标签序列和片段丙;
    所述带有不同标签的测序接头单链B从5’至3’方向依次包括与所述片段丙反向互补的片段、标签序列、片段丁和与所述片段甲反向互补的片段;
    所述步骤3)中,所述引物中的一条引物与所述测序接头单链A中的片段乙匹配,所述引物中的另一条引物与单链B中的片段丁匹配。
  7. 根据权利要求1-6中任一所述的方法,其特征在于:
    所述带有不同标签的测序接头单链A的条数为小于等于72,且各条单链标签序列不同;
    所述带有不同标签的测序接头单链B的条数为小于等于72,且各条单链标签序列不同;
    所述标签序列中n大于等于8且小于15。
  8. 根据权利要求1-7中任一所述的方法,其特征在于:
    步骤2)中,所述将带有标签的测序接头单链A和与其部分互补的带有不同标签的测序接头单链B以单链的形式分别加入含有所述断裂片段的体系中反应包括如下步骤:
    2)-A,将72条所述带有不同标签序列的测序接头单链A分别加入含有所述断裂片段体系的芯片的纵向72列中;
    2)-B,将72条所述带有不同标签序列的测序接头单链B分别加入2)-A处理后的芯片的横向72行中,反应,得到连接测序接头产物。
  9. 根据权利要求1-8中任一所述的方法,其特征在于:
    步骤3)中,所述引物中的一条引物与所述测序接头单链A中的片段乙相同或反向互补;
    所述引物中的另一条引物与所述测序接头单链B中的片段丁的片段互补反向互补或相同。
  10. 根据权利要求2-9中任一所述的方法,其特征在于:
    所述方法的步骤1)的(2)、(3)和步骤2)均在5184孔芯片中进行。
  11. 根据权利要求1-10中任一所述的方法,其特征在于:
    所述长片段DNA为大于100Kb的片段,具体为大小为400kb的片段;
    所述转座酶为Tn5转座酶。
  12. 根据权利要求2-11中任一所述的方法,其特征在于:
    所述方法的步骤1)的(2)、(3)和步骤2)中均采用wafergen MSND移液平台将各种物质加入芯片的孔中。
  13. 根据权利要求12所述的方法,其特征在于:在所述用wafergenMSND移液平台将各种物质加入芯片的孔中后,还包括如下步骤:离心沉降各物质。
  14. 根据权利要求2-11中任一所述的方法,其特征在于:
    步骤1)的(1)和(2)中,还包括如下步骤:用变性试剂消化所述转座酶;所述变性试剂具体采用0.1-1%SDS溶液;所述消化的条件具体为25℃静置10分钟。
  15. 根据权利要求4-14中任一所述的方法,其特征在于:
    步骤1)的(1)中,所述转座酶断裂反应的条件为55℃5分钟;
    步骤1)的(2)中,所述扩增的退火温度为68℃,退火时间为18min;
    步骤1)的(3)中,所述酶切反应条件为37℃2小时,再65℃15分钟;所述聚合反应条件为23℃1小时,65℃30分钟;
    步骤2)的2)-B中,所述反应条件为20℃2小时;
    步骤3)中,所述PCR扩增的退火温度为68℃,退火时间为10min。
  16. 由权利要求1-15所述方法制备得到的长片段DNA文库。
  17. 一种用于长片段DNA文库构建的连接测序接头产物制备方法,包括如下步骤:为权利要求1-16中任一所述方法中的步骤1)-2)。
  18. 权利要求17所述的方法在构建长片段DNA文库中的应用。
PCT/CN2016/079278 2015-04-20 2016-04-14 一种长片段dna文库构建方法 Ceased WO2016169431A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
CN201680012412.4A CN107250447B (zh) 2015-04-20 2016-04-14 一种长片段dna文库构建方法
US15/567,963 US20180195060A1 (en) 2015-04-20 2016-04-14 Method for constructing long fragment dna library

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201510187820.0 2015-04-20
CN201510187820 2015-04-20

Publications (1)

Publication Number Publication Date
WO2016169431A1 true WO2016169431A1 (zh) 2016-10-27

Family

ID=57142917

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2016/079278 Ceased WO2016169431A1 (zh) 2015-04-20 2016-04-14 一种长片段dna文库构建方法

Country Status (3)

Country Link
US (1) US20180195060A1 (zh)
CN (1) CN107250447B (zh)
WO (1) WO2016169431A1 (zh)

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107586835A (zh) * 2017-10-19 2018-01-16 东南大学 一种基于单链接头的下一代测序文库的构建方法及其应用
CN107904668A (zh) * 2018-01-02 2018-04-13 上海美吉生物医药科技有限公司 一种微生物多样性文库构建方法及其应用
CN108315387A (zh) * 2018-02-07 2018-07-24 北京大学 微量细胞ChIP法
CN113088567A (zh) * 2021-03-29 2021-07-09 中国农业科学院农业基因组研究所 一种长片段dna分子结构柔性的表征方法
CN116254320A (zh) * 2022-12-15 2023-06-13 纳昂达(南京)生物科技有限公司 平末端双链接头元件、试剂盒及平末端建库方法

Families Citing this family (31)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10323279B2 (en) 2012-08-14 2019-06-18 10X Genomics, Inc. Methods and systems for processing polynucleotides
US10584381B2 (en) 2012-08-14 2020-03-10 10X Genomics, Inc. Methods and systems for processing polynucleotides
US9701998B2 (en) 2012-12-14 2017-07-11 10X Genomics, Inc. Methods and systems for processing polynucleotides
US11591637B2 (en) 2012-08-14 2023-02-28 10X Genomics, Inc. Compositions and methods for sample processing
US10752949B2 (en) 2012-08-14 2020-08-25 10X Genomics, Inc. Methods and systems for processing polynucleotides
US9951386B2 (en) 2014-06-26 2018-04-24 10X Genomics, Inc. Methods and systems for processing polynucleotides
AU2013302756C1 (en) 2012-08-14 2018-05-17 10X Genomics, Inc. Microcapsule compositions and methods
US10533221B2 (en) 2012-12-14 2020-01-14 10X Genomics, Inc. Methods and systems for processing polynucleotides
CA2894694C (en) 2012-12-14 2023-04-25 10X Genomics, Inc. Methods and systems for processing polynucleotides
KR20200140929A (ko) 2013-02-08 2020-12-16 10엑스 제노믹스, 인크. 폴리뉴클레오티드 바코드 생성
BR112016023625A2 (pt) 2014-04-10 2018-06-26 10X Genomics, Inc. dispositivos fluídicos, sistemas e métodos para encapsular e particionar reagentes, e aplicações dos mesmos
WO2015200893A2 (en) 2014-06-26 2015-12-30 10X Genomics, Inc. Methods of analyzing nucleic acids from individual cells or cell populations
CA2953469A1 (en) 2014-06-26 2015-12-30 10X Genomics, Inc. Analysis of nucleic acid sequences
US12312640B2 (en) 2014-06-26 2025-05-27 10X Genomics, Inc. Analysis of nucleic acid sequences
JP6769969B2 (ja) 2015-01-12 2020-10-14 10エックス ジェノミクス, インコーポレイテッド 核酸配列決定ライブラリーを作製するためのプロセス及びシステム、並びにこれらを使用して作製したライブラリー
EP3262407B1 (en) 2015-02-24 2023-08-30 10X Genomics, Inc. Partition processing methods and systems
US12264411B2 (en) 2017-01-30 2025-04-01 10X Genomics, Inc. Methods and systems for analysis
EP4029939B1 (en) 2017-01-30 2023-06-28 10X Genomics, Inc. Methods and systems for droplet-based single cell barcoding
US10995333B2 (en) 2017-02-06 2021-05-04 10X Genomics, Inc. Systems and methods for nucleic acid preparation
US10400235B2 (en) 2017-05-26 2019-09-03 10X Genomics, Inc. Single cell analysis of transposase accessible chromatin
EP3445876B1 (en) 2017-05-26 2023-07-05 10X Genomics, Inc. Single cell analysis of transposase accessible chromatin
CN109680343B (zh) * 2017-10-18 2022-02-18 深圳华大生命科学研究院 一种外泌体微量dna的建库方法
SG11201913654QA (en) 2017-11-15 2020-01-30 10X Genomics Inc Functionalized gel beads
US10829815B2 (en) 2017-11-17 2020-11-10 10X Genomics, Inc. Methods and systems for associating physical and genetic properties of biological particles
CN109930206A (zh) * 2017-12-18 2019-06-25 深圳华大生命科学研究院 基于bgiseq-500的微量血小板rna测序检测试剂盒
CN108384835A (zh) * 2018-01-23 2018-08-10 安徽微分基因科技有限公司 一种基于Wafergen Smartchip系统的高效极微量基因表达定量技术
CN108411374A (zh) * 2018-01-24 2018-08-17 安徽微分基因科技有限公司 一种基于Wafergen Smartchip系统的高效极微量TE建库方法
CN112262218B (zh) 2018-04-06 2024-11-08 10X基因组学有限公司 用于单细胞处理中的质量控制的系统和方法
CN111524552B (zh) * 2020-04-24 2021-05-11 深圳市儒翰基因科技有限公司 简化基因组测序文库构建分析方法、检测设备及存储介质
CN115852495B (zh) * 2022-12-30 2023-09-15 苏州泓迅生物科技股份有限公司 一种基因突变文库的合成方法及其应用
WO2025176158A1 (en) * 2024-02-20 2025-08-28 Mgi Tech Co., Ltd. Long contiguous DNA sequence reads from longer amplified inserts by maintaining phase and improving detection

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102264914A (zh) * 2008-10-24 2011-11-30 阿霹震中科技公司 用于修饰核酸的转座子末端组合物和方法
CN102703426A (zh) * 2012-05-21 2012-10-03 吴江汇杰生物科技有限公司 构建核酸库的方法、试剂及试剂盒
CN103443338A (zh) * 2011-02-02 2013-12-11 华盛顿大学商业化中心 大规模平行邻接作图

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP2464738A4 (en) * 2009-08-12 2013-05-01 Nugen Technologies Inc METHODS, COMPOSITIONS, AND KITS FOR GENERATING NUCLEIC ACID PRODUCTS SUBSTANTIALLY FREE OF MATRIX NUCLEIC ACID
US8829171B2 (en) * 2011-02-10 2014-09-09 Illumina, Inc. Linking sequence reads using paired code tags
WO2013059746A1 (en) * 2011-10-19 2013-04-25 Nugen Technologies, Inc. Compositions and methods for directional nucleic acid amplification and sequencing
EP2807292B1 (en) * 2012-01-26 2019-05-22 Tecan Genomics, Inc. Compositions and methods for targeted nucleic acid sequence enrichment and high efficiency library generation
EP4148142B1 (en) * 2012-05-21 2025-08-13 The Scripps Research Institute Methods of sample preparation
CN104711250A (zh) * 2015-01-26 2015-06-17 北京百迈客生物科技有限公司 一种长片段核酸文库的构建方法

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102264914A (zh) * 2008-10-24 2011-11-30 阿霹震中科技公司 用于修饰核酸的转座子末端组合物和方法
CN103443338A (zh) * 2011-02-02 2013-12-11 华盛顿大学商业化中心 大规模平行邻接作图
CN102703426A (zh) * 2012-05-21 2012-10-03 吴江汇杰生物科技有限公司 构建核酸库的方法、试剂及试剂盒

Cited By (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107586835A (zh) * 2017-10-19 2018-01-16 东南大学 一种基于单链接头的下一代测序文库的构建方法及其应用
CN107586835B (zh) * 2017-10-19 2020-11-03 东南大学 一种基于单链接头的下一代测序文库的构建方法及其应用
CN107904668A (zh) * 2018-01-02 2018-04-13 上海美吉生物医药科技有限公司 一种微生物多样性文库构建方法及其应用
CN108315387A (zh) * 2018-02-07 2018-07-24 北京大学 微量细胞ChIP法
WO2019153852A1 (zh) * 2018-02-07 2019-08-15 北京大学 微量细胞ChIP法
CN113088567A (zh) * 2021-03-29 2021-07-09 中国农业科学院农业基因组研究所 一种长片段dna分子结构柔性的表征方法
CN113088567B (zh) * 2021-03-29 2023-06-13 中国农业科学院农业基因组研究所 一种长片段dna分子结构柔性的表征方法
CN116254320A (zh) * 2022-12-15 2023-06-13 纳昂达(南京)生物科技有限公司 平末端双链接头元件、试剂盒及平末端建库方法

Also Published As

Publication number Publication date
CN107250447B (zh) 2020-05-05
CN107250447A (zh) 2017-10-13
US20180195060A1 (en) 2018-07-12

Similar Documents

Publication Publication Date Title
WO2016169431A1 (zh) 一种长片段dna文库构建方法
JP2024060054A (ja) ヌクレアーゼ、リガーゼ、ポリメラーゼ、及び配列決定反応の組み合わせを用いた、核酸配列、発現、コピー、またはdnaのメチル化変化の識別及び計数方法
CN108060191B (zh) 一种双链核酸片段加接头的方法、文库构建方法和试剂盒
US20220372548A1 (en) Vitro isolation and enrichment of nucleic acids using site-specific nucleases
CN106715713B (zh) 试剂盒及其在核酸测序中的用途
CN107075512B (zh) 一种接头元件和使用其构建测序文库的方法
EP3208336B1 (en) Linker element and method of using same to construct sequencing library
CN107124888B (zh) 鼓泡状接头元件和使用其构建测序文库的方法
IL283866B (en) Method of nucleic acid enrichment using site-specific nucleases followed by capture
WO2016037394A1 (zh) 一种核酸单链环状文库的构建方法和试剂
CN107002292A (zh) 一种核酸的双接头单链环状文库的构建方法和试剂
WO2016082130A1 (zh) 一种核酸的双接头单链环状文库的构建方法和试剂
WO2012037883A1 (zh) 核酸标签及其应用
WO2016078096A1 (zh) 使用鼓泡状接头元件构建测序文库的方法
JP2023510515A (ja) 二本鎖核酸分子およびそれを用いたdnaライブラリー内の遊離アダプターの除去法
CN107002290B9 (zh) 样品制备方法
HK1240289B (zh) 一种长片段dna文库构建方法
CN104673785A (zh) 一种链特异性全基因组范围内研究apa的方法
HK1240289A1 (zh) 一种长片段dna文库构建方法
JP2022547699A (ja) ライブラリーのスクリーニング方法
HK1248767A1 (zh) 一种双链核酸片段加接头的方法、文库构建方法和试剂盒
HK1240286B (zh) 一种核酸的双接头单链环状文库的构建方法和试剂
HK1240203B (zh) 一种核酸的双接头单链环状文库的构建方法和试剂
HK1240615B (zh) 一种接头元件和使用其构建测序文库的方法
HK1240268B (zh) 鼓泡状接头元件和使用其构建测序文库的方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16782578

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16782578

Country of ref document: EP

Kind code of ref document: A1