WO2025128786A1 - Methods for simultaneous characterization of genome and transcriptome of a single cell - Google Patents

Methods for simultaneous characterization of genome and transcriptome of a single cell Download PDF

Info

Publication number
WO2025128786A1
WO2025128786A1 PCT/US2024/059713 US2024059713W WO2025128786A1 WO 2025128786 A1 WO2025128786 A1 WO 2025128786A1 US 2024059713 W US2024059713 W US 2024059713W WO 2025128786 A1 WO2025128786 A1 WO 2025128786A1
Authority
WO
WIPO (PCT)
Prior art keywords
adapter
sequence
cells
dna
genomic dna
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/US2024/059713
Other languages
French (fr)
Inventor
Chenghang ZONG
Muchun NIU
Yichi NIU
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Baylor College of Medicine
Original Assignee
Baylor College of Medicine
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Baylor College of Medicine filed Critical Baylor College of Medicine
Publication of WO2025128786A1 publication Critical patent/WO2025128786A1/en
Anticipated expiration legal-status Critical
Pending legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/10Processes for the isolation, preparation or purification of DNA or RNA
    • C12N15/1096Processes for the isolation, preparation or purification of DNA or RNA cDNA Synthesis; Subtracted cDNA library construction, e.g. RT, RT-PCR
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/10Processes for the isolation, preparation or purification of DNA or RNA
    • C12N15/1034Isolating an individual clone by screening libraries
    • C12N15/1093General methods of preparing gene libraries, not provided for in other subgroups
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12PFERMENTATION OR ENZYME-USING PROCESSES TO SYNTHESISE A DESIRED CHEMICAL COMPOUND OR COMPOSITION OR TO SEPARATE OPTICAL ISOMERS FROM A RACEMIC MIXTURE
    • C12P19/00Preparation of compounds containing saccharide radicals
    • C12P19/26Preparation of nitrogen-containing carbohydrates
    • C12P19/28N-glycosides
    • C12P19/30Nucleotides
    • C12P19/34Polynucleotides, e.g. nucleic acids, oligoribonucleotides
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12YENZYMES
    • C12Y207/00Transferases transferring phosphorus-containing groups (2.7)
    • C12Y207/07Nucleotidyltransferases (2.7.7)
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12YENZYMES
    • C12Y301/00Hydrolases acting on ester bonds (3.1)
    • C12Y301/21Endodeoxyribonucleases producing 5'-phosphomonoesters (3.1.21)
    • C12Y301/21004Type II site-specific deoxyribonuclease (3.1.21.4)
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/14Hydrolases (3)
    • C12N9/16Hydrolases (3) acting on ester bonds (3.1)
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/14Hydrolases (3)
    • C12N9/16Hydrolases (3) acting on ester bonds (3.1)
    • C12N9/22Ribonucleases [RNase]; Deoxyribonucleases [DNase]

Definitions

  • Embodiments of the present disclosure relate to methods for characterization of the genomic and transcriptomic nucleic acids in single cellular or subcellular particles, or single aggregations of biological constituents.
  • duplex-sequencing provides the solution to minimize technical artifacts and therefore results in accurate detection of somatic mutations 8 .
  • the duplex-sequencing only when the base changes are detected in both Watson and Crick strands from the same double-stranded DNA fragments, will the variants be called true mutations (illustrated in FIG. 1).
  • the duplex-sequencing approach has not been adapted for single cell level characterization. The present disclosure satisfies a need for implementation of duplex-sequencing into single cell chemistry.
  • genomic information such as the somatic mutations, DNA damage, small insertions and deletions (INDEL), copy number variations (CNV), and structural variations (SV) is resolved for a single cell
  • INDEL small insertions and deletions
  • CNV copy number variations
  • SV structural variations
  • the present disclosure also satisfies a long-felt need in the art for simultaneous characterization of the genome and transcriptome of a single cell, a single subcellular particle such as a nucleus, and a single aggregate of biological constituents.
  • Embodiments of the present disclosure relate in general to methods and compositions for producing DNA libraries representative of DNA and/or RNA sequences of any kind, including at least genomic DNA, extrachromosomal DNA, mitochondrial DNA, cell-free DNA, synthetic DNA, mRNA, nascent RNA, long non-coding RNA, and synthetic RNA.
  • Embodiments of the disclosure include methods for amplifying genomes and/or transcriptomes, including from one or more cells and/or one or more subcellular structures from one or more cells.
  • the methods of the disclosure allow for increased accuracy over known methods in the art.
  • the disclosure concerns amplifying double-stranded nucleic acids in a mode where the original strand information is resolved in part or in full.
  • the strand information could be resolved by paired-read direction in next generation sequencing or third generation sequencing, or a strand barcode introduced during library preparation steps, as examples.
  • the double-stranded nucleic acids include at least doublestranded DNA, DNA:RNA hybrid, double-stranded RNA, and the stem part of a nucleic acid forming a stem-loop structure, in some embodiments.
  • this design may allow distinguishing of mutations and nucleic variants that only display on one strand, such as a single base DNA damage, or an epimutation.
  • this design also may allow distinguishing of indels and a DNA bulge, where one strand carries one or more nucleotides than the other strand.
  • Embodiments of methods are provided herein for performing simultaneous or substantially simultaneous amplification of the genome and/or the transcriptome for single cells with high uniformity and high sensitivity across the genome and transcriptome, which allows accurate detection of genomic information and quantification of gene expression by standard high throughput sequencing platforms, in some embodiments.
  • the genomic information includes at least mutations, epimutations, single base DNA damage, small insertions and deletions (INDEL), DNA bulges, copy number variations (CNV), and/or structural variations (SV), in specific embodiments.
  • Embodiments of the disclosure include methods of producing a library representing both DNA and RNA related to a single cell and/or subcellular structure(s) from a single cell, comprising fragmenting genomic DNA using a Tn5 transposase.
  • the transposase sequence comprises one or more restriction sites, including in some embodiments Type II S restriction enzyme sites. Such activity may produce genomic DNA fragments comprising transposase sequence generally at both ends.
  • the restriction enzyme digestion allows for subsequent adapter ligation to the digested ends, and in specific embodiments the adapters have both unique and common sequences. The unique sequence allows for identification of a single nucleic acid molecule, and the common sequence allows for subsequent amplification of multiple nucleic acid molecules, such as generally at the same time.
  • the transposase sequence comprises one or more U bases.
  • Such one or more U bases may be excised to introduce one or more protruding ends to genomic DNA fragments.
  • Such one or more U bases may be excised by using Uracil.
  • the protruding ends allow for subsequent adapter ligation to the protruding ends, and in specific embodiments the adapters have both unique and common sequences.
  • the unique sequence allows for identification of a single nucleic acid molecule, and the common sequence allows for subsequent amplification of multiple nucleic acid molecules, such as generally at the same time.
  • Embodiments of the disclosure include methods of producing a library representing both DNA and RNA related to a single cell and/or subcellular structure(s) from a single cell, comprising the steps of: (a) a relatively weak cell lysis procedure is performed to lyse the whole transcriptome of a single cell with or without disturbing the genomic DNA; (b) a reverse transcription step using the cellular RNA as a template, and results in a complementary DNA (cDNA) containing common handle sequences on both 5’ and 3’ end that can be amplifiable; (c) a relatively strong cell lysis procedure to efficiently release the genomic DNA (gDNA) from the cell; (d) fragmentation of gDNA to the fragments of desired lengths; (e) introduction of unique identifiers to both strands of the DNA fragments, which can be a strand-specific barcode, or a strand-specific adapter orientation, or both; (f) simultaneous amplification of both RNA-derived products and DNA-derived products.
  • step (b) can be accomplished by one or more reverse transcriptases with template switching activity on the presence of at least one type of template switching oligo.
  • step (b) can be accomplished by one or more reverse transcriptases without template switching activity, followed by a homopolymer tailing reaction, such as the polyadenylation tailing catalyzed by terminal transferase.
  • step (b) can be accomplished by one or more reverse transcriptases without template switching activity, followed by a ligation reaction to ligate a nucleic acid sequence to the end of cDNA.
  • step (b) can be accomplished by one or more reverse transcriptases without template switching activity, followed by a second strand synthesis reaction that utilizes the cDNA as the template.
  • step (b) can be obviated in a reverse transcriptase-independent manner.
  • step (b) can be obviated by hybridization of the RNA with one or multiple probe nucleic acids.
  • step (b) can be obviated by direct ligation of adapters containing common handle sequences to the original cellular RNA.
  • the gDNA fragmentation of step (d) can be accomplished by mechanical force (such as sonication), one or more enzymes (such as fragmentase, transposase, nuclease, or a combination thereof), heating, pH changes, chemicals, or their combinations.
  • mechanical force such as sonication
  • one or more enzymes such as fragmentase, transposase, nuclease, or a combination thereof
  • heating pH changes, chemicals, or their combinations.
  • the introduction of strand-specific identifiers in step (e) can be accomplished by ligation of Y-shape adapters or stem-looped adapters. Such adapters may or may not carry unique barcodes, because the orientation of the adapters during the ligation step can already distinguish both strands of the original double-stranded DNA fragments.
  • the introduction of strand-specific identifiers in step (e) can be accomplished by a polymerase-mediated DNA synthesis reaction so that both strands of the original double-stranded DNA fragments are tagged with different barcodes or the same barcode but different barcode orientations.
  • step (d) and step (e) can be accomplished by a one-step reaction by using transposases loaded with Y-shaped adapters or stem -looped adapters as the transposons.
  • the adapter sequences are automatically pasted to the gDNA fragments, with both strands carrying the adapters in different orientations.
  • the nucleic acid amplification in step (f) can be linearly, quasi- linearly, or non-linearly produced using the same original template and primers with or without unique molecular identifiers (UMI).
  • UMI unique molecular identifiers
  • the linear amplification process makes amplification errors independent from each other. Such feature allows efficient filtering of amplification errors, and thereby resulting in accurate calling of somatic mutations and somatic variants only on a single strand, such as epimutations or single base DNA damage.
  • a chemical conversion step such as bisulfite conversion, nitrite conversion
  • the gDNA lysis in step (c) may be omitted, or may be utilized with milder treatment.
  • step (c) can be skipped.
  • the methods in the present disclosure resulted in simultaneous characterization of the transcriptome and genomic information of ATAC fragments (instead of the whole genome) of the same single cell.
  • the methods in the present disclosure can be performed to characterize a single cell, a single subcellular particle such as a nucleus, a single aggregate of biological constituents, a small number of biological constituents, as well as bulk samples.
  • the disclosure concerns amplifying transcriptomic and genomic sequences in situ, such as the transcriptome and the ATAC fragments in fixed subcellular structures or particles.
  • the methods in the present disclosure can be performed in a single-omics version, for example, transcriptome only, genome only, or ATAC fragments only.
  • the methods in the present disclosure can be performed with bulk tissue samples instead of single particles, single subcellular structures, etc.
  • double stranded nucleic acids have their strands separated and are analyzed by methods that identify a change on the opposing strands, such as compared to a known standard.
  • a change on one strand is detected and a corresponding change on the other strand is detected.
  • the nucleic acids are derived from a cell or cells from a sample from an individual.
  • the sample may be of any kind, but in specific embodiments the sample comprises a biopsy, blood (including cord blood), urine, cheek scraping, plasma, saliva, stool, sputum, semen, cerebrospinal fluid, arterial sampling, amniocentesis, or a combination thereof.
  • Aspect 1 is a method for producing a library of nucleic acids representing nucleic acid of one or more cells and/or one or more subcellular structures.
  • the method comprises: fragmenting genomic DNA using a Tn5 transposase assembled with a sequence comprised of a restriction enzyme site, to produce genomic DNA fragments comprising transposase sequence comprising the site at both ends; repairing the ends of the genomic DNA fragments; subjecting the genomic DNA fragments to the restriction enzyme to produce digested molecules; and ligating adapters to both ends of the digested molecules to produce adapter-ligated molecules, wherein each adapter at both ends of the adapter- ligated molecule comprises a unique sequence to label the DNA fragment that it was ligated to and a common sequence, optionally for amplification.
  • Aspect 2 includes all the limitations of aspect 1 and further comprises generating RNA-representing double stranded DNA molecules representing part or all of a transcriptome of the one or more cells, wherein the molecules comprise common, amplifiable handle sequences on both 5’ and 3’ ends.
  • Aspect 3 includes all the limitations of aspect 2, wherein the generating step comprises reverse transcription with template switching activity.
  • Aspect 4 includes all the limitations of aspect 2, wherein the generating step comprises reverse transcription without template switching activity.
  • Aspect 5 includes all the limitations of aspect 4, wherein the method further comprises a homopolymer tailing reaction.
  • Aspect 6 includes all the limitations of aspect 4, wherein the method further comprises a ligation reaction to ligate a nucleic acid sequence to the end of cDNA.
  • Aspect 7 includes all the limitations of aspect 4, wherein the method further comprises a second strand synthesis reaction that utilizes cDNA as a template.
  • Aspect 8 includes all the limitations of aspect 2, wherein the generating step lacks reverse transcription.
  • Aspect 9 includes all the limitations of aspect 8, wherein the generating comprises hybridization of the RNA with one or multiple probe nucleic acids.
  • Aspect 10 includes all the limitations of aspect 8, wherein the generating comprises direct ligation of adapters containing common handle sequences to the original cellular RNA.
  • Aspect 11 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the restriction enzyme is a Type II S restriction enzyme.
  • Aspect 12 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the fragmenting of the genomic DNA is by mechanical force, by using one or more enzymes, by heating, by pH changes, by chemicals, or a combination thereof.
  • Aspect 13 includes all the limitations of aspect 12, wherein the mechanical force comprises sonication and/or hydrodynamic shearing.
  • Aspect 14 includes all the limitations of aspect 12, wherein the enzymes comprise fragmentase, transposase, nuclease, or a combination thereof.
  • Aspect 15 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the unique sequence comprises a strand-specific barcode, a strand-specific adapter orientation, or both.
  • Aspect 16 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the repairing occurs using a DNA polymerase.
  • Aspect 17 includes all the limitations of aspect 16, wherein the DNA polymerase is incapable of nick translation activity, strand displacement activity, and/or 5 ’->3’ exonuclease activity.
  • Aspect 18 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapters are asymmetric.
  • Aspect 19 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapters are Y-shaped or looped.
  • Aspect 20 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapter-ligated molecules are amplified.
  • Aspect 21 includes all the limitations of aspect 20, wherein the adapter-ligated molecules are amplified by linear amplification.
  • Aspect 22 includes all the limitations of aspect 20, wherein the adapter-ligated molecules are amplified by non-linear amplification.
  • Aspect 23 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the subcellular structures comprise organelles.
  • Aspect 24 includes all the limitations of aspect 23, wherein the organelle is a mitochondria, nucleus, endoplasmic reticulum, chloroplast, Golgi apparatus, ribosome, or combination thereof.
  • Aspect 25 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the one or more cells and/or one or more subcellular structures are obtained from storage.
  • Aspect 26 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the one or more cells and/or one or more subcellular structures are obtained from a sample from one or more individuals.
  • Aspect 27 includes all the limitations of any one of the preceding aspects, respectively or collectively, further comprises the step of obtaining the one or more cells and/or one or more subcellular structures from a sample from one or more individuals.
  • Aspect 28 includes all the limitations of aspect 27, wherein the sample comprises solid matter.
  • Aspect 29 includes all the limitations of aspect 27, wherein the solid matter comprises tissue.
  • Aspect 30 includes all the limitations of aspect 27, wherein the tissue is homogenized to produce single cells.
  • Aspect 31 is a method for producing a library of nucleic acids representing nucleic acid of one or more cells and/or one or more subcellular structures.
  • the method comprises: fragmenting genomic DNA using a Tn5 transposase assembled with a sequence comprised of one or more U bases and/or a first adaptor, to produce genomic DNA fragments comprising the sequence at both ends; repairing the ends of the genomic DNA fragments; introducing one or more protruding ends to the genomic DNA fragments by excision of the one or more U bases; and ligating a second adapter to the one or more protruding ends to produce adapter-ligated molecules, wherein each of the first adaptor and second adaptor comprises a unique sequence to label the DNA fragment that it was ligated to and a common sequence, optionally for amplification.
  • Aspect 32 includes all the limitations of aspect 31, wherein the step of introducing one or more protruding ends to the genomic DNA fragments by excision of the one or more U bases comprises using Uracil for the excision.
  • Aspect 33 includes all the limitations of either aspect 31 or aspect 32, respectively or collectively, and further comprises generating RNA-representing double stranded DNA molecules representing part or all of a transcriptome of the one or more cells, wherein the molecules comprise common, amplifiable handle sequences on both 5’ and 3’ ends.
  • Aspect 34 includes all the limitations of aspect 33, wherein the generating step comprises reverse transcription with template switching activity.
  • Aspect 35 includes all the limitations of aspect 33, wherein the generating step comprises reverse transcription without template switching activity.
  • Aspect 36 includes all the limitations of aspect 35, wherein the method further comprises a homopolymer tailing reaction.
  • Aspect 37 includes all the limitations of aspect 35, wherein the method further comprises a ligation reaction to ligate a nucleic acid sequence to the end of cDNA.
  • Aspect 38 includes all the limitations of aspect 35, wherein the method further comprises a second strand synthesis reaction that utilizes cDNA as a template.
  • Aspect 39 includes all the limitations of aspect 33, wherein the generating step lacks reverse transcription.
  • Aspect 40 includes all the limitations of aspect 39, wherein the generating comprises hybridization of the RNA with one or multiple probe nucleic acids.
  • Aspect 41 includes all the limitations of aspect 39, wherein the generating comprises direct ligation of adapters containing common handle sequences to the original cellular RNA.
  • Aspect 42 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the fragmenting of the genomic DNA is by mechanical force, by using one or more enzymes, by heating, by pH changes, by chemicals, or a combination thereof.
  • Aspect 43 includes all the limitations of aspect 42, wherein the mechanical force comprises sonication and/or hydrodynamic shearing.
  • Aspect 44 includes all the limitations of aspect 42, wherein the enzymes comprise fragmentase, transposase, nuclease, or a combination thereof.
  • Aspect 45 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the unique sequence comprises a strand-specific barcode, a strand-specific adapter orientation, or both.
  • Aspect 46 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the repairing occurs using a DNA polymerase.
  • Aspect 47 includes all the limitations of aspect 46, wherein the DNA polymerase is incapable of nick translation activity, strand displacement activity, and/or 5’->3’ exonuclease activity.
  • Aspect 48 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapters are asymmetric.
  • Aspect 49 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapters are Y-shaped or looped.
  • Aspect 50 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapter-ligated molecules are amplified.
  • Aspect 51 includes all the limitations of aspect 50, wherein the adapter-ligated molecules are amplified by linear amplification.
  • Aspect 52 includes all the limitations of aspect 50, wherein the adapter-ligated molecules are amplified by non-linear amplification.
  • Aspect 53 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the subcellular structures comprise organelles.
  • Aspect 54 includes all the limitations of aspect 53, wherein the organelle is a mitochondria, nucleus, endoplasmic reticulum, chloroplast, Golgi apparatus, ribosome, or combination thereof.
  • Aspect 55 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the one or more cells and/or one or more subcellular structures are obtained from storage.
  • Aspect 56 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the one or more cells and/or one or more subcellular structures are obtained from a sample from one or more individuals.
  • Aspect 57 includes all the limitations of any one of the preceding aspects, respectively or collectively, further comprises the step of obtaining the one or more cells and/or one or more subcellular structures from a sample from one or more individuals.
  • Aspect 58 includes all the limitations of aspect 57, wherein the sample comprises solid matter.
  • Aspect 59 includes all the limitations of aspect 57, wherein the solid matter comprises tissue.
  • Aspect 60 includes all the limitations of aspect 57, wherein the tissue is homogenized to produce single cells.
  • FIG. 1 An embodiment of variant calling strategy in duplex-sequencing.
  • FIG. 2 An example workflow of the chemistry for simultaneous capture of transcriptome and genome in single cells.
  • FIG. 3 One embodiment of the design for duplex-sequencing of the genome in single cells in Example I.
  • FIG. 4 One embodiment of a design for duplex-sequencing of the genome in single cells in Example II.
  • FIG. 5 One embodiment of a design for duplex-sequencing of the genome in single cells in Example III.
  • FIG. 6 One embodiment of variant calling strategy in duplex-sequencing with linear amplification from Example IV.
  • x, y, and/or z can refer to “x” alone, “y” alone, “z” alone, “x, y, and z,” “(x and y) or z,” “x or (y and z),” or “x or y or z.” It is specifically contemplated that x, y, or z may be specifically excluded from an embodiment.
  • RNA and DNA sequences with strand information resolved, as representative sequences in a DNA library, when the corresponding RNA and DNA molecules are a part of or otherwise associated with single cellular or single subcellular structures.
  • the RNA of single cells is released by one or more stimuli, including but not limited to heating, pH changes, osmotic pressure, salt, detergent (such as Tween-20, Triton XI 00, NP40, digitonin, etc.), enzymes, or a combination thereof.
  • stimuli including but not limited to heating, pH changes, osmotic pressure, salt, detergent (such as Tween-20, Triton XI 00, NP40, digitonin, etc.), enzymes, or a combination thereof.
  • cellular RNA in the reaction mixture is converted to cDNA by at least one reverse transcriptase with template switching activity (such as superscript II reverse transcriptase, superscript IV reverse transcriptase, and/or Maxima H- reverse transcriptase) on the presence of template switching oligos.
  • template switching activity such as superscript II reverse transcriptase, superscript IV reverse transcriptase, and/or Maxima H- reverse transcriptase
  • cellular RNA in the reaction mixture is converted to cDNA by at least one reverse transcriptase without template switching activity, such as superscript III reverse transcriptase.
  • cDNA is then tailed with a common handle sequence that can be primed for further amplification.
  • the common handle sequence can be a homopolymer sequence introduced by a terminal transferase, or the common handle sequence can be a sequence ligated by a nucleic acid ligase, in particular embodiments.
  • the common handle can be introduced on the second strand by a second strand synthesis reaction using the cDNA as template and catalyzed by at least one DNA polymerase.
  • cellular RNA in the reaction mixture is hybridized to one or multiple nucleic acid probes that share reverse complementary sequences to the cellular RNA.
  • Methods to generate such probe nucleic acids include but are not limited to in vitro transcription and chemical synthesis.
  • Such probes may contain common handles on both 5’ and 3’ sides that can be primed for linear amplification or nonlinear amplification such as PCR.
  • both 5’ and 3’ ends of the cellular RNA in the reaction mixture can be ligated with adapters comprising common handle sequences that can be primed for linear amplification or nonlinear amplification, such as PCR.
  • the gDNA is physically (such as heating), chemically (such as sodium dodecyl sulfate, SDS), and/or enzymatically (such as protease, proteinase K) lysed from the cellular content.
  • SDS sodium dodecyl sulfate
  • enzymatically such as protease, proteinase K
  • the gDNA lysis step could be omitted.
  • direct fragmentation of gDNA in the intact single nuclei or partially lysed single nuclei will result in the capture of a subset of the genome rather than the whole genome.
  • the subset can be accessible genomic regions, which is comparable to transposase-accessible chromatin in ATAC-seq, and DNase hypersensitive sites.
  • the subset can be a pre-defined list of genomic loci, which can be obtained by a Cas9 nuclease with one or more guide RNAs, in certain embodiments.
  • gDNA is fragmented into the length of desire, typically 30-2,000 base pairs in length for next generation sequencing, or 30 - 1,000,000 base pairs in length for third generation sequencing.
  • the range of fragment length is subject to change in order to accommodate start-of-the-art application of currently available next generation sequencing and third generation sequencing platforms, as well as the potential development of other sequencing platforms.
  • Methods to fragment the gDNA can be mechanical methods (such as sonication), chemical methods (such as heating, pH changes, and chemicals), and/or enzymatical methods (such as fragmentase, transposase, nuclease).
  • enzymatical methods include but are not limited to fragmentase, MNase, Tn5 transposase, restriction endonuclease, Cas9 nuclease, Casl2a nuclease, a combination thereof, etc.
  • strand-specific identifiers can be introduced to both strands of the double-stranded DNA fragments by adapter ligation reaction.
  • the double-stranded DNA fragments can undergo end repair, and then ligated to nucleic acid adapters (ligation adapters).
  • the nucleic acid adapters can be Y-shaped adapters in some embodiments, in which two oligonucleotide sequences with reverse complimentary sequences are annealed to each other.
  • Such ligation adapters can be annealed before ligation, or the two different oligonucleotide sequence components can also be simultaneously or subsequently added to the ligation reaction, without pre-annealing, in particular embodiments.
  • the nucleic acid adapter can also be an oligonucleotide sequence that contains reverse complimentary sequences so that a stem- loop structure is formed.
  • Such ligation adapters can contain an optional barcode sequence, that allows more accurate identification of both strands of double-stranded DNA.
  • strand-specific identifiers can be introduced to both strands of the double-stranded DNA fragments by non-ligation reaction, such as a polymerase-mediated DNA synthesis reaction.
  • non-ligation reaction such as a polymerase-mediated DNA synthesis reaction.
  • the transposon sequences introduced by Tn5 fragmented fragments one can melt the double-stranded gDNA fragments into two single strands, and use the single strands as template and the primer with a new common sequence A on the 5’ end and the transposon sequence (or reverse complimentary sequence) on the 3’ end.
  • Such a polymerase-mediated DNA synthesis reaction will result in a semi-amplicon produced from the two strands of original DNA fragments, in particular embodiments.
  • the 5’ end of semi-amplicons carry a new common sequence A.
  • the product is melted again into single strands, and one can use the single strands as template and the primer with a new common sequence B on the 5’ end and the transposon sequence (or reverse complimentary sequence) on the 3’ end.
  • This polymerase-mediated DNA synthesis reaction will result in a full amplicon, where the 5’ end of full-amplicons carry a new common sequence B and the 3’ end of full-amplicons carry a new common sequence A. Since the two original strands have inverse 573’ orientations, based on the orientation of common sequence A and common sequence B, one can readily distinguish the product originated from both template gDNA fragment strands.
  • a Tn5 transposase may be assembled with a transposon sequence without a restriction enzyme recognition site.
  • a Tn5 transposase may be assembled with a transposon sequence comprised of one or more U bases and/or a first adaptor and/or primer.
  • Such Tn5 transposase may be used to fragment genomic DNA to produce genomic DNA fragments comprising transposase sequence comprising the one or more U bases at both ends.
  • a DNA polymerase may be used for gap filling of the transposed DNA sequence.
  • the one or more U bases may be excised (e.g., by Uracil excision) to introduce protruding ends of the genomic DNA fragments.
  • a second adapter and/or primer can then be hybridized next to the recessive ends of the genomic DNA fragments.
  • the second adapter and/or primer and the recessive ends of the genomic DNA fragments may then be ligated, catalyzed by a DNA ligase, such as E. Coli DNA ligase, 9N DNA ligase, T4 DNA ligase, etc.
  • any DNA synthesis steps before melting the two strands of original gDNA fragments may decrease the detection accuracy because they break the independency between the two strands.
  • a tailing step that blocks DNA synthesis is applied before melting the two strands of original gDNA fragments to prevent it from ligation.
  • the tailing steps include but are not limited to (a) using at least one types of deoxy -ribonucleoside triphosphate (dNTP) with a amide modification on the 3’ position of the sugar ring rather than regular dNTP, (b) using at least one type of dideoxyribonucleoside triphosphate (ddNTP) rather than dNTP.
  • amide-modified dNTP or ddNTP, or other blockers of polymerase extension can be added by an DNA polymerase, a terminal transferase, a DNA ligase etc.
  • examples of the types of dNTP are: dATP, dCTP, dGTP, and dTTP.
  • the DNA synthesis steps can be performed with one or more DNA polymerases with minimal 5’ exonuclease activity, minimal 3’ exonuclease activity, minimal strand displacement activity, or minimal nick translation activity, etc.
  • Reaction buffer conditions, and temperature conditions can be adjusted to minimize such activities to yield optimal results.
  • templated synthesis breaks the independency between the two strands, the newly synthesized strand cannot extend to the end, thereby incapable of ligation.
  • Examples of such DNA polymerases include but are not limited to Q5 DNA polymerase, DeepVent (exo-) DNA polymerase, Sulfolobus DNA Polymerase IV, etc.
  • a UMI may be a unique or random nucleotide sequence that includes a number of bases, such as at least, at most, a range of, or exactly 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22,
  • the strand-tagged nucleic acids can undergo an optional chemical conversion step, such as bisulfite conversion or nitrite, for the purpose to characterize epigenetic marks such as 5-methylcytosine (5mC), and 6-methyladenine (6mA).
  • an optional chemical conversion step such as bisulfite conversion or nitrite
  • epigenetic marks such as 5-methylcytosine (5mC), and 6-methyladenine (6mA).
  • the transcriptome sequences and strand-tagged genome sequences are simultaneously amplified in the same reaction, for example a PCR reaction.
  • the transcriptome sequences and strand-tagged genome sequences can be further split from each other by PCR amplification with specific primers, or by nucleic acid probes.
  • the chemistry presented for tagging both strands of double-stranded DNA are also applicable to other double-stranded nucleic acids, or stem-looped nucleic acids, such as cDNA:RNA hybrid, double-stranded RNA, a RNA hairpin, a DNA hairpin, etc.
  • the chemistry presented here can be applied to single tubes, microwell-plates, nanowells, plate-based liquid handlers, microfluidic systems, droplet platforms, etc.
  • Histological tissue slides were prepared using commercial cryostat. Sections of 10-20 pm thickness of paraformaldehyde-fixed solid tissue were prepared. The cells of interest were cut from the section slide by laser microdissection microscopes (Leica, MMI etc.). The individual cells were dissociated from the micro-dissected tissue section using standard cell dissociation protocols. The complete single cells were collected into individual tubes.
  • frozen tumor tissue was chopped into small pieces with a blade.
  • the tissue was resuspended in homogenization buffer for homogenization.
  • the homogenate was passed through a cell strainer and centrifuged. The supernatant was discarded, and the pellet containing nuclei was resuspended in phosphate-buff ered saline (PBS) containing DNA staining dyes such as Hoechst 33342 (Invitrogen).
  • PBS phosphate-buff ered saline
  • DNA staining dyes such as Hoechst 33342 (Invitrogen).
  • the samples were subject to fluorescence-activated nucleus sorting (e.g., FACS), and single nuclei were collected into individual tubes.
  • the single cells were lysed in a mild lysis buffer.
  • the captured cells were thermally lysed and then stored in -80 °C.
  • a reverse transcription and template switching step was performed to generate the cDNA amplicons that have common adapters on both ends (e.g., R1 adapter and R2 adapter in FIG. 2) that allow primer binding in further amplifications, and genome library and transcriptome library are generated after amplification (FIG. 2).
  • a protease-based genomic DNA lysis mixture was added into the PCR tube to release genomic DNA. Next, the protease is inactivated by heating.
  • RNA digestion step is performed to remove RNA from the RNA:cDNA hybrid to prevent the cDNA amplicons from being fragmented in the next step.
  • This RNA digestion step can be achieved by incubation with at least one RNase, such as RNase A, RNase H, and/or RNase If.
  • a custom Tn5 transposase was assembled with a transposon sequence comprised of the mosaic end sequence (referred to as “ME,” a sequence which may be recognized by transposases, shown in FIG. 3) and a specific restriction enzyme recognition site (referred to as “site RE,” shown in FIG. 3).
  • ME mosaic end sequence
  • site RE specific restriction enzyme recognition site
  • the genomic DNA is randomly fragmented by the customized Tn5 transposase, resulting in the short genomic DNA fragments with transposon sequences attached to both ends (referred to as “tagging”).
  • the transposition reaction may be quenched by EDTA and heating to release the transposase.
  • a restriction enzyme e.g., type II S restriction enzyme shown in FIG. 3
  • site RE a restriction enzyme that recognizes the “site RE”
  • a sticky end shown in FIG. 3 after the step of type II S restriction enzyme digestion
  • the fragmented DNA is then ligated to an asymmetric Y-shaped or a looped adapter, catalyzed by a DNA ligase, such as T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, ampligase thermostable DNA ligase, etc.
  • a DNA ligase such as T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, ampligase thermostable DNA ligase, etc.
  • the digestion by the designed restriction enzyme generated the designed sticky end, which will not only maximize the downstream ligation efficiency, but also reduce the length of adapters in the sequencing library, therefore facilitating the cost-effectiveness of sequencing.
  • the DNA fragments can be ligated to a Y-shape adapter, or a looped adapter.
  • the asymmetric feature of the adapters allow distinguishing of both strands in the sequencing data, and thus genome strand-aware tagging is achieved.
  • each adaptor may comprise a UMI.
  • the marked base in the UMI stands for one base (e.g., A base) that is added to the 3’ ends of DNA fragments during the end repair process.
  • the custom Tn5 transposase tagmentation reaction (shown in FIG. 3) could be performed with genomic DNA lysis.
  • the custom Tn5 transposase tagmentation reaction (shown in FIG. 3) could be performed without genomic DNA lysis, such as by removing the genomic DNA lysis step.
  • This design captures the open chromatin only, rather than the whole genome.
  • the captured sequences essentially corresponded to ATAC-seq (Assay for Transposase- Accessible Chromatin using sequencing), in particular embodiments.
  • the end repair step could be performed in the absence of certain types of dNTP (dATP, dCTP, dGTP, and/or dTTP).
  • dNTP dATP, dCTP, dGTP, and/or dTTP
  • modified nucleotides that prevent further extension such as ddNTP, or amide-modified dNTP could be used here to replace one or more of the dNTPs (referred to as “ddNTP sealing,” shown in FIG. 3).
  • the end repair step could be performed using a DNA polymerase without nick translation activity, strand displacement activity, or 5 ’->3’ exonuclease activity.

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Chemical & Material Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Organic Chemistry (AREA)
  • Engineering & Computer Science (AREA)
  • Wood Science & Technology (AREA)
  • Zoology (AREA)
  • General Engineering & Computer Science (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Biochemistry (AREA)
  • Biotechnology (AREA)
  • Molecular Biology (AREA)
  • Biomedical Technology (AREA)
  • Microbiology (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • Physics & Mathematics (AREA)
  • Biophysics (AREA)
  • Plant Pathology (AREA)
  • Chemical Kinetics & Catalysis (AREA)
  • General Chemical & Material Sciences (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

Embodiments of the disclosure include methods of producing libraries of nucleic acids, including genomes and/or transcriptomes, including for single cells and/or subcellular structures. In specific embodiments, a double stranded DNA library is produced upon fragmentation by a Tn5 transposase having restriction enzyme site(s), digestion by the restriction enzyme, and ligation of adaptors to the ends.

Description

METHODS FOR SIMULTANEOUS CHARACTERIZATION OF GENOME AND TRANSCRIPTOME OF A SINGLE CELL
[001] This application claims priority to U.S. Provisional Patent Application Serial No. 63/608,947, filed December 12, 2023, which is incorporated by reference herein in its entirety.
STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[002] This invention was made with government support under 1DP2EB020399 and 1UG3NS132132 awarded by National Institutes of Health. The government has certain rights in the invention.
TECHNICAL FIELD
[003] Embodiments of the present disclosure relate to methods for characterization of the genomic and transcriptomic nucleic acids in single cellular or subcellular particles, or single aggregations of biological constituents.
BACKGROUND
[004] Multiple recent studies including regional dissection-based sequencing and single-cell expansion-based sequencing have shown that somatic mutations accumulate with aging and the accumulated somatic mutation burden in individual human cells on average can reach thousands of mutations per cell. However, the regional dissection-based approach cannot effectively detect the variations of somatic mutations in individual cells. While the single-cell expansion-based approach can detect single-cell variations, this approach is limited to special cell types like stem cells from which single-cell expansion may be established. In order to decipher the somatic mutations unbiasedly for single cells or single nuclei that are isolated from various types of human tissues, it is greatly desired to profile the somatic mutations of single cells by single-cell whole genome amplification and sequencing.
[005] However, despite the rapid development of various single-cell whole genome amplification (scWGA) methods 1-7, the accuracy in the detection of somatic mutations in single cells is still under intensive debate largely due to various sources of technical artifacts 6,7 The recent development of duplex-sequencing based approach provides the solution to minimize technical artifacts and therefore results in accurate detection of somatic mutations8. In the duplex-sequencing, only when the base changes are detected in both Watson and Crick strands from the same double-stranded DNA fragments, will the variants be called true mutations (illustrated in FIG. 1). However, for lack of detection sensitivity, to date the duplex-sequencing approach has not been adapted for single cell level characterization. The present disclosure satisfies a need for implementation of duplex-sequencing into single cell chemistry.
[006] On the other hand, once the genomic information such as the somatic mutations, DNA damage, small insertions and deletions (INDEL), copy number variations (CNV), and structural variations (SV) is resolved for a single cell, it is greatly desired to connect such genomic information to phenotypic readouts, such as the transcriptome. The present disclosure also satisfies a long-felt need in the art for simultaneous characterization of the genome and transcriptome of a single cell, a single subcellular particle such as a nucleus, and a single aggregate of biological constituents.
SUMMARY
[007] Embodiments of the present disclosure relate in general to methods and compositions for producing DNA libraries representative of DNA and/or RNA sequences of any kind, including at least genomic DNA, extrachromosomal DNA, mitochondrial DNA, cell-free DNA, synthetic DNA, mRNA, nascent RNA, long non-coding RNA, and synthetic RNA.
[008] Embodiments of the disclosure include methods for amplifying genomes and/or transcriptomes, including from one or more cells and/or one or more subcellular structures from one or more cells. The methods of the disclosure allow for increased accuracy over known methods in the art.
[009] In particular embodiments, the disclosure concerns amplifying double-stranded nucleic acids in a mode where the original strand information is resolved in part or in full. The strand information could be resolved by paired-read direction in next generation sequencing or third generation sequencing, or a strand barcode introduced during library preparation steps, as examples. The double-stranded nucleic acids include at least doublestranded DNA, DNA:RNA hybrid, double-stranded RNA, and the stem part of a nucleic acid forming a stem-loop structure, in some embodiments. In some embodiments, this design may allow distinguishing of mutations and nucleic variants that only display on one strand, such as a single base DNA damage, or an epimutation. In some embodiments, this design also may allow distinguishing of indels and a DNA bulge, where one strand carries one or more nucleotides than the other strand.
[010] Embodiments of methods are provided herein for performing simultaneous or substantially simultaneous amplification of the genome and/or the transcriptome for single cells with high uniformity and high sensitivity across the genome and transcriptome, which allows accurate detection of genomic information and quantification of gene expression by standard high throughput sequencing platforms, in some embodiments. Here, the genomic information includes at least mutations, epimutations, single base DNA damage, small insertions and deletions (INDEL), DNA bulges, copy number variations (CNV), and/or structural variations (SV), in specific embodiments.
[OH] Embodiments of the disclosure include methods of producing a library representing both DNA and RNA related to a single cell and/or subcellular structure(s) from a single cell, comprising fragmenting genomic DNA using a Tn5 transposase. In specific embodiments, the transposase sequence comprises one or more restriction sites, including in some embodiments Type II S restriction enzyme sites. Such activity may produce genomic DNA fragments comprising transposase sequence generally at both ends. The restriction enzyme digestion allows for subsequent adapter ligation to the digested ends, and in specific embodiments the adapters have both unique and common sequences. The unique sequence allows for identification of a single nucleic acid molecule, and the common sequence allows for subsequent amplification of multiple nucleic acid molecules, such as generally at the same time. Alternatively and/or additionally, in specific embodiments, the transposase sequence comprises one or more U bases. Such one or more U bases may be excised to introduce one or more protruding ends to genomic DNA fragments. Such one or more U bases may be excised by using Uracil. The protruding ends allow for subsequent adapter ligation to the protruding ends, and in specific embodiments the adapters have both unique and common sequences. The unique sequence allows for identification of a single nucleic acid molecule, and the common sequence allows for subsequent amplification of multiple nucleic acid molecules, such as generally at the same time.
[012] Embodiments of the disclosure include methods of producing a library representing both DNA and RNA related to a single cell and/or subcellular structure(s) from a single cell, comprising the steps of: (a) a relatively weak cell lysis procedure is performed to lyse the whole transcriptome of a single cell with or without disturbing the genomic DNA; (b) a reverse transcription step using the cellular RNA as a template, and results in a complementary DNA (cDNA) containing common handle sequences on both 5’ and 3’ end that can be amplifiable; (c) a relatively strong cell lysis procedure to efficiently release the genomic DNA (gDNA) from the cell; (d) fragmentation of gDNA to the fragments of desired lengths; (e) introduction of unique identifiers to both strands of the DNA fragments, which can be a strand-specific barcode, or a strand-specific adapter orientation, or both; (f) simultaneous amplification of both RNA-derived products and DNA-derived products.
[013] In specific embodiments, step (b) can be accomplished by one or more reverse transcriptases with template switching activity on the presence of at least one type of template switching oligo. In specific embodiments, step (b) can be accomplished by one or more reverse transcriptases without template switching activity, followed by a homopolymer tailing reaction, such as the polyadenylation tailing catalyzed by terminal transferase. In specific embodiments, step (b) can be accomplished by one or more reverse transcriptases without template switching activity, followed by a ligation reaction to ligate a nucleic acid sequence to the end of cDNA. In specific embodiments, step (b) can be accomplished by one or more reverse transcriptases without template switching activity, followed by a second strand synthesis reaction that utilizes the cDNA as the template.
[014] In some embodiments, step (b) can be obviated in a reverse transcriptase-independent manner. For example, step (b) can be obviated by hybridization of the RNA with one or multiple probe nucleic acids. For another example, step (b) can be obviated by direct ligation of adapters containing common handle sequences to the original cellular RNA.
[015] In particular embodiments, the gDNA fragmentation of step (d) can be accomplished by mechanical force (such as sonication), one or more enzymes (such as fragmentase, transposase, nuclease, or a combination thereof), heating, pH changes, chemicals, or their combinations.
[016] In specific embodiments, the introduction of strand-specific identifiers in step (e) can be accomplished by ligation of Y-shape adapters or stem-looped adapters. Such adapters may or may not carry unique barcodes, because the orientation of the adapters during the ligation step can already distinguish both strands of the original double-stranded DNA fragments. In specific embodiments, the introduction of strand-specific identifiers in step (e) can be accomplished by a polymerase-mediated DNA synthesis reaction so that both strands of the original double-stranded DNA fragments are tagged with different barcodes or the same barcode but different barcode orientations.
[017] In particular embodiments, step (d) and step (e) can be accomplished by a one-step reaction by using transposases loaded with Y-shaped adapters or stem -looped adapters as the transposons. Under this scenario, the adapter sequences are automatically pasted to the gDNA fragments, with both strands carrying the adapters in different orientations.
[018] In some embodiments, the nucleic acid amplification in step (f) can be linearly, quasi- linearly, or non-linearly produced using the same original template and primers with or without unique molecular identifiers (UMI). The linear amplification process makes amplification errors independent from each other. Such feature allows efficient filtering of amplification errors, and thereby resulting in accurate calling of somatic mutations and somatic variants only on a single strand, such as epimutations or single base DNA damage.
[019] In some embodiments, one can characterize epigenetic marks including but not limited to 5-methylcytosine (5mC), and 6-methyladenine (6mA), and a chemical conversion step (such as bisulfite conversion, nitrite conversion) can be integrated in the procedure.
[020] In some embodiments that relate to characterization of a subset of a genome, the gDNA lysis in step (c) may be omitted, or may be utilized with milder treatment. For example, in an application to profile the genomic information from accessible genome regions (such as the assay for transposase-accessible chromatin with sequencing, ATAC-seq), step (c) can be skipped. In this example, the methods in the present disclosure resulted in simultaneous characterization of the transcriptome and genomic information of ATAC fragments (instead of the whole genome) of the same single cell.
[021] In particular embodiments, the methods in the present disclosure can be performed to characterize a single cell, a single subcellular particle such as a nucleus, a single aggregate of biological constituents, a small number of biological constituents, as well as bulk samples.
[022] In specific embodiments, the disclosure concerns amplifying transcriptomic and genomic sequences in situ, such as the transcriptome and the ATAC fragments in fixed subcellular structures or particles. [023] In particular embodiments, the methods in the present disclosure can be performed in a single-omics version, for example, transcriptome only, genome only, or ATAC fragments only.
[024] In particular embodiments, the methods in the present disclosure can be performed with bulk tissue samples instead of single particles, single subcellular structures, etc.
[025] In particular embodiments, double stranded nucleic acids have their strands separated and are analyzed by methods that identify a change on the opposing strands, such as compared to a known standard. In particular embodiments, a change on one strand is detected and a corresponding change on the other strand is detected.
[026] In specific embodiments, the nucleic acids are derived from a cell or cells from a sample from an individual. The sample may be of any kind, but in specific embodiments the sample comprises a biopsy, blood (including cord blood), urine, cheek scraping, plasma, saliva, stool, sputum, semen, cerebrospinal fluid, arterial sampling, amniocentesis, or a combination thereof.
[027] Also disclosed in the context of the present disclosure are aspects 1 to 60. Aspect 1 is a method for producing a library of nucleic acids representing nucleic acid of one or more cells and/or one or more subcellular structures. The method comprises: fragmenting genomic DNA using a Tn5 transposase assembled with a sequence comprised of a restriction enzyme site, to produce genomic DNA fragments comprising transposase sequence comprising the site at both ends; repairing the ends of the genomic DNA fragments; subjecting the genomic DNA fragments to the restriction enzyme to produce digested molecules; and ligating adapters to both ends of the digested molecules to produce adapter-ligated molecules, wherein each adapter at both ends of the adapter- ligated molecule comprises a unique sequence to label the DNA fragment that it was ligated to and a common sequence, optionally for amplification. Aspect 2 includes all the limitations of aspect 1 and further comprises generating RNA-representing double stranded DNA molecules representing part or all of a transcriptome of the one or more cells, wherein the molecules comprise common, amplifiable handle sequences on both 5’ and 3’ ends. Aspect 3 includes all the limitations of aspect 2, wherein the generating step comprises reverse transcription with template switching activity. Aspect 4 includes all the limitations of aspect 2, wherein the generating step comprises reverse transcription without template switching activity. Aspect 5 includes all the limitations of aspect 4, wherein the method further comprises a homopolymer tailing reaction. Aspect 6 includes all the limitations of aspect 4, wherein the method further comprises a ligation reaction to ligate a nucleic acid sequence to the end of cDNA. Aspect 7 includes all the limitations of aspect 4, wherein the method further comprises a second strand synthesis reaction that utilizes cDNA as a template. Aspect 8 includes all the limitations of aspect 2, wherein the generating step lacks reverse transcription. Aspect 9 includes all the limitations of aspect 8, wherein the generating comprises hybridization of the RNA with one or multiple probe nucleic acids. Aspect 10 includes all the limitations of aspect 8, wherein the generating comprises direct ligation of adapters containing common handle sequences to the original cellular RNA. Aspect 11 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the restriction enzyme is a Type II S restriction enzyme. Aspect 12 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the fragmenting of the genomic DNA is by mechanical force, by using one or more enzymes, by heating, by pH changes, by chemicals, or a combination thereof. Aspect 13 includes all the limitations of aspect 12, wherein the mechanical force comprises sonication and/or hydrodynamic shearing. Aspect 14 includes all the limitations of aspect 12, wherein the enzymes comprise fragmentase, transposase, nuclease, or a combination thereof. Aspect 15 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the unique sequence comprises a strand-specific barcode, a strand-specific adapter orientation, or both. Aspect 16 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the repairing occurs using a DNA polymerase. Aspect 17 includes all the limitations of aspect 16, wherein the DNA polymerase is incapable of nick translation activity, strand displacement activity, and/or 5 ’->3’ exonuclease activity. Aspect 18 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapters are asymmetric. Aspect 19 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapters are Y-shaped or looped. Aspect 20 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapter-ligated molecules are amplified. Aspect 21 includes all the limitations of aspect 20, wherein the adapter-ligated molecules are amplified by linear amplification. Aspect 22 includes all the limitations of aspect 20, wherein the adapter-ligated molecules are amplified by non-linear amplification. Aspect 23 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the subcellular structures comprise organelles. Aspect 24 includes all the limitations of aspect 23, wherein the organelle is a mitochondria, nucleus, endoplasmic reticulum, chloroplast, Golgi apparatus, ribosome, or combination thereof. Aspect 25 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the one or more cells and/or one or more subcellular structures are obtained from storage. Aspect 26 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the one or more cells and/or one or more subcellular structures are obtained from a sample from one or more individuals. Aspect 27 includes all the limitations of any one of the preceding aspects, respectively or collectively, further comprises the step of obtaining the one or more cells and/or one or more subcellular structures from a sample from one or more individuals. Aspect 28 includes all the limitations of aspect 27, wherein the sample comprises solid matter. Aspect 29 includes all the limitations of aspect 27, wherein the solid matter comprises tissue. Aspect 30 includes all the limitations of aspect 27, wherein the tissue is homogenized to produce single cells.
[028] Aspect 31 is a method for producing a library of nucleic acids representing nucleic acid of one or more cells and/or one or more subcellular structures. The method comprises: fragmenting genomic DNA using a Tn5 transposase assembled with a sequence comprised of one or more U bases and/or a first adaptor, to produce genomic DNA fragments comprising the sequence at both ends; repairing the ends of the genomic DNA fragments; introducing one or more protruding ends to the genomic DNA fragments by excision of the one or more U bases; and ligating a second adapter to the one or more protruding ends to produce adapter-ligated molecules, wherein each of the first adaptor and second adaptor comprises a unique sequence to label the DNA fragment that it was ligated to and a common sequence, optionally for amplification. Aspect 32 includes all the limitations of aspect 31, wherein the step of introducing one or more protruding ends to the genomic DNA fragments by excision of the one or more U bases comprises using Uracil for the excision. Aspect 33, includes all the limitations of either aspect 31 or aspect 32, respectively or collectively, and further comprises generating RNA-representing double stranded DNA molecules representing part or all of a transcriptome of the one or more cells, wherein the molecules comprise common, amplifiable handle sequences on both 5’ and 3’ ends. Aspect 34 includes all the limitations of aspect 33, wherein the generating step comprises reverse transcription with template switching activity. Aspect 35 includes all the limitations of aspect 33, wherein the generating step comprises reverse transcription without template switching activity. Aspect 36 includes all the limitations of aspect 35, wherein the method further comprises a homopolymer tailing reaction. Aspect 37 includes all the limitations of aspect 35, wherein the method further comprises a ligation reaction to ligate a nucleic acid sequence to the end of cDNA. Aspect 38 includes all the limitations of aspect 35, wherein the method further comprises a second strand synthesis reaction that utilizes cDNA as a template. Aspect 39 includes all the limitations of aspect 33, wherein the generating step lacks reverse transcription. Aspect 40 includes all the limitations of aspect 39, wherein the generating comprises hybridization of the RNA with one or multiple probe nucleic acids. Aspect 41 includes all the limitations of aspect 39, wherein the generating comprises direct ligation of adapters containing common handle sequences to the original cellular RNA. Aspect 42 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the fragmenting of the genomic DNA is by mechanical force, by using one or more enzymes, by heating, by pH changes, by chemicals, or a combination thereof. Aspect 43 includes all the limitations of aspect 42, wherein the mechanical force comprises sonication and/or hydrodynamic shearing. Aspect 44 includes all the limitations of aspect 42, wherein the enzymes comprise fragmentase, transposase, nuclease, or a combination thereof. Aspect 45 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the unique sequence comprises a strand-specific barcode, a strand-specific adapter orientation, or both. Aspect 46 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the repairing occurs using a DNA polymerase. Aspect 47 includes all the limitations of aspect 46, wherein the DNA polymerase is incapable of nick translation activity, strand displacement activity, and/or 5’->3’ exonuclease activity. Aspect 48 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapters are asymmetric. Aspect 49 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapters are Y-shaped or looped. Aspect 50 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the adapter-ligated molecules are amplified. Aspect 51 includes all the limitations of aspect 50, wherein the adapter-ligated molecules are amplified by linear amplification. Aspect 52 includes all the limitations of aspect 50, wherein the adapter-ligated molecules are amplified by non-linear amplification. Aspect 53 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the subcellular structures comprise organelles. Aspect 54 includes all the limitations of aspect 53, wherein the organelle is a mitochondria, nucleus, endoplasmic reticulum, chloroplast, Golgi apparatus, ribosome, or combination thereof. Aspect 55 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the one or more cells and/or one or more subcellular structures are obtained from storage. Aspect 56 includes all the limitations of any one of the preceding aspects, respectively or collectively, wherein the one or more cells and/or one or more subcellular structures are obtained from a sample from one or more individuals. Aspect 57 includes all the limitations of any one of the preceding aspects, respectively or collectively, further comprises the step of obtaining the one or more cells and/or one or more subcellular structures from a sample from one or more individuals. Aspect 58 includes all the limitations of aspect 57, wherein the sample comprises solid matter. Aspect 59 includes all the limitations of aspect 57, wherein the solid matter comprises tissue. Aspect 60 includes all the limitations of aspect 57, wherein the tissue is homogenized to produce single cells.
BRIEF DESCRIPTION OF THE FIGURES
[029] FIG. 1. An embodiment of variant calling strategy in duplex-sequencing.
[030] FIG. 2. An example workflow of the chemistry for simultaneous capture of transcriptome and genome in single cells.
[031] FIG. 3. One embodiment of the design for duplex-sequencing of the genome in single cells in Example I.
[032] FIG. 4. One embodiment of a design for duplex-sequencing of the genome in single cells in Example II.
[033] FIG. 5. One embodiment of a design for duplex-sequencing of the genome in single cells in Example III.
[034] FIG. 6. One embodiment of variant calling strategy in duplex-sequencing with linear amplification from Example IV.
DETAILED DESCRIPTION
[035] In keeping with long-standing patent law convention, the words “a” and “an” when used in the present specification in concert with the word comprising, including the claims, denote “one or more.” Some embodiments of the disclosure may consist of or consist essentially of one or more elements, method steps, and/or methods of the disclosure. It is contemplated that any method or composition described herein can be implemented with respect to any other method or composition described herein and that different embodiments may be combined.
[036] Throughout this specification, unless the context requires otherwise, the words “comprise”, “comprises” and “comprising” will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements. By “consisting of’ is meant including, and limited to, whatever follows the phrase “consisting of.” Thus, the phrase “consisting of’ indicates that the listed elements are required or mandatory, and that no other elements may be present. By “consisting essentially of’ is meant including any elements listed after the phrase, and limited to other elements that do not interfere with or contribute to the activity or action specified in the disclosure for the listed elements. Thus, the phrase “consisting essentially of’ indicates that the listed elements are required or mandatory, but that no other elements are optional and may or may not be present depending upon whether or not they affect the activity or action of the listed elements.
[037] Reference throughout this specification to “one embodiment,” “an embodiment,” “a particular embodiment,” “a related embodiment,” “a certain embodiment,” “an additional embodiment,” or “a further embodiment” or combinations thereof means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the foregoing phrases in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[038] As used herein, the terms “or” and “and/or” are utilized to describe multiple components in combination or exclusive of one another. For example, “x, y, and/or z” can refer to “x” alone, “y” alone, “z” alone, “x, y, and z,” “(x and y) or z,” “x or (y and z),” or “x or y or z.” It is specifically contemplated that x, y, or z may be specifically excluded from an embodiment.
[039] Throughout this application, the term “about” is used according to its plain and ordinary meaning in the area of cell and molecular biology to indicate that a value includes the standard deviation of error for the device or method being employed to determine the value. [040] The present disclosure concerns methods of amplifying RNA and DNA sequences with strand information resolved, as representative sequences in a DNA library, when the corresponding RNA and DNA molecules are a part of or otherwise associated with single cellular or single subcellular structures.
[041] According to certain aspects of the disclosure, the RNA of single cells is released by one or more stimuli, including but not limited to heating, pH changes, osmotic pressure, salt, detergent (such as Tween-20, Triton XI 00, NP40, digitonin, etc.), enzymes, or a combination thereof.
[042] According to one aspect of the disclosure, cellular RNA in the reaction mixture is converted to cDNA by at least one reverse transcriptase with template switching activity (such as superscript II reverse transcriptase, superscript IV reverse transcriptase, and/or Maxima H- reverse transcriptase) on the presence of template switching oligos.
[043] According to some aspects of the disclosure, cellular RNA in the reaction mixture is converted to cDNA by at least one reverse transcriptase without template switching activity, such as superscript III reverse transcriptase. cDNA is then tailed with a common handle sequence that can be primed for further amplification. The common handle sequence can be a homopolymer sequence introduced by a terminal transferase, or the common handle sequence can be a sequence ligated by a nucleic acid ligase, in particular embodiments. In some embodiments, the common handle can be introduced on the second strand by a second strand synthesis reaction using the cDNA as template and catalyzed by at least one DNA polymerase.
[044] According to one aspect of the disclosure, cellular RNA in the reaction mixture is hybridized to one or multiple nucleic acid probes that share reverse complementary sequences to the cellular RNA. Methods to generate such probe nucleic acids include but are not limited to in vitro transcription and chemical synthesis. Such probes may contain common handles on both 5’ and 3’ sides that can be primed for linear amplification or nonlinear amplification such as PCR.
[045] According to one aspect of the disclosure, both 5’ and 3’ ends of the cellular RNA in the reaction mixture can be ligated with adapters comprising common handle sequences that can be primed for linear amplification or nonlinear amplification, such as PCR. [046] According to certain aspects of the disclosure, the gDNA is physically (such as heating), chemically (such as sodium dodecyl sulfate, SDS), and/or enzymatically (such as protease, proteinase K) lysed from the cellular content. A combination of different lysis methods may be applicable.
[047] According to certain aspects of the disclosure, the gDNA lysis step could be omitted. In this scenario, direct fragmentation of gDNA in the intact single nuclei or partially lysed single nuclei will result in the capture of a subset of the genome rather than the whole genome. For example, the subset can be accessible genomic regions, which is comparable to transposase-accessible chromatin in ATAC-seq, and DNase hypersensitive sites. For another example, the subset can be a pre-defined list of genomic loci, which can be obtained by a Cas9 nuclease with one or more guide RNAs, in certain embodiments.
[048] According to certain aspects of the disclosure, gDNA is fragmented into the length of desire, typically 30-2,000 base pairs in length for next generation sequencing, or 30 - 1,000,000 base pairs in length for third generation sequencing. The range of fragment length is subject to change in order to accommodate start-of-the-art application of currently available next generation sequencing and third generation sequencing platforms, as well as the potential development of other sequencing platforms. Methods to fragment the gDNA can be mechanical methods (such as sonication), chemical methods (such as heating, pH changes, and chemicals), and/or enzymatical methods (such as fragmentase, transposase, nuclease). Detailed examples of enzymatical methods include but are not limited to fragmentase, MNase, Tn5 transposase, restriction endonuclease, Cas9 nuclease, Casl2a nuclease, a combination thereof, etc.
[049] According to certain aspects of the disclosure, strand-specific identifiers can be introduced to both strands of the double-stranded DNA fragments by adapter ligation reaction. For example, in some embodiments the double-stranded DNA fragments can undergo end repair, and then ligated to nucleic acid adapters (ligation adapters). The nucleic acid adapters can be Y-shaped adapters in some embodiments, in which two oligonucleotide sequences with reverse complimentary sequences are annealed to each other. Such ligation adapters can be annealed before ligation, or the two different oligonucleotide sequence components can also be simultaneously or subsequently added to the ligation reaction, without pre-annealing, in particular embodiments. The nucleic acid adapter can also be an oligonucleotide sequence that contains reverse complimentary sequences so that a stem- loop structure is formed. Such ligation adapters can contain an optional barcode sequence, that allows more accurate identification of both strands of double-stranded DNA.
[050] According to one aspect of the disclosure, strand-specific identifiers can be introduced to both strands of the double-stranded DNA fragments by non-ligation reaction, such as a polymerase-mediated DNA synthesis reaction. For example, in the DNA fragments that already share common sequences on both ends, such as the transposon sequences introduced by Tn5 fragmented fragments, one can melt the double-stranded gDNA fragments into two single strands, and use the single strands as template and the primer with a new common sequence A on the 5’ end and the transposon sequence (or reverse complimentary sequence) on the 3’ end. Such a polymerase-mediated DNA synthesis reaction will result in a semi-amplicon produced from the two strands of original DNA fragments, in particular embodiments. The 5’ end of semi-amplicons carry a new common sequence A. Next, the product is melted again into single strands, and one can use the single strands as template and the primer with a new common sequence B on the 5’ end and the transposon sequence (or reverse complimentary sequence) on the 3’ end. This polymerase-mediated DNA synthesis reaction will result in a full amplicon, where the 5’ end of full-amplicons carry a new common sequence B and the 3’ end of full-amplicons carry a new common sequence A. Since the two original strands have inverse 573’ orientations, based on the orientation of common sequence A and common sequence B, one can readily distinguish the product originated from both template gDNA fragment strands.
[051] According to one aspect of the disclosure, the steps of gDNA fragmentation and strandspecific identifier introduction can be merged into one step. For example, Y-shape or stem- looped adapters (similar to the ligation adapters) can be pre-loaded as the transposons in the Tn5 transposase. As a result, direct Tn5 transposase-based gDNA fragmentation will automatically paste adapter sequences to the gDNA fragments, with both strands carrying the adapters in different orientations.
[052] According to one aspect of the disclosure, a Tn5 transposase may be assembled with a transposon sequence comprised of a restriction enzyme recognition site. Such Tn5 transposase may be used to fragment genomic DNA to produce genomic DNA fragments comprising transposase sequence comprising the restriction enzyme recognition site at both ends. Next, a DNA polymerase may be used for gap filling of the transposed DNA sequence. Followed by the gap filling, a restriction enzyme that recognizes the restriction enzyme recognition site may be used to generate a sticky or a blunt end of the genomic DNA fragments. The fragmented DNA with either a sticky or a blunt end may then be ligated to an asymmetric Y-shaped or a looped adapter, catalyzed by a DNA ligase, such as T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, ampligase thermostable DNA ligase, etc.
[053] According to another aspect of the disclosure, a Tn5 transposase may be assembled with a transposon sequence without a restriction enzyme recognition site. In one aspect, a Tn5 transposase may be assembled with a transposon sequence comprised of one or more U bases and/or a first adaptor and/or primer. Such Tn5 transposase may be used to fragment genomic DNA to produce genomic DNA fragments comprising transposase sequence comprising the one or more U bases at both ends. Next, a DNA polymerase may be used for gap filling of the transposed DNA sequence. Followed by the gap filling, the one or more U bases may be excised (e.g., by Uracil excision) to introduce protruding ends of the genomic DNA fragments. A second adapter and/or primer can then be hybridized next to the recessive ends of the genomic DNA fragments. The second adapter and/or primer and the recessive ends of the genomic DNA fragments may then be ligated, catalyzed by a DNA ligase, such as E. Coli DNA ligase, 9N DNA ligase, T4 DNA ligase, etc.
[054] According to certain aspects of the disclosure, any DNA synthesis steps before melting the two strands of original gDNA fragments may decrease the detection accuracy because they break the independency between the two strands.
[055] According to one aspect of the disclosure, a tailing step that blocks DNA synthesis is applied before melting the two strands of original gDNA fragments to prevent it from ligation. Examples for the tailing steps include but are not limited to (a) using at least one types of deoxy -ribonucleoside triphosphate (dNTP) with a amide modification on the 3’ position of the sugar ring rather than regular dNTP, (b) using at least one type of dideoxyribonucleoside triphosphate (ddNTP) rather than dNTP. These amide-modified dNTP or ddNTP, or other blockers of polymerase extension can be added by an DNA polymerase, a terminal transferase, a DNA ligase etc. Herein, examples of the types of dNTP are: dATP, dCTP, dGTP, and dTTP.
[056] According to one aspect of the disclosure, the DNA synthesis steps can be performed with one or more DNA polymerases with minimal 5’ exonuclease activity, minimal 3’ exonuclease activity, minimal strand displacement activity, or minimal nick translation activity, etc. Reaction buffer conditions, and temperature conditions can be adjusted to minimize such activities to yield optimal results. In these scenarios, although templated synthesis breaks the independency between the two strands, the newly synthesized strand cannot extend to the end, thereby incapable of ligation. Examples of such DNA polymerases include but are not limited to Q5 DNA polymerase, DeepVent (exo-) DNA polymerase, Sulfolobus DNA Polymerase IV, etc.
[057] According to some aspects of the disclosure, the strand-tagged nucleic acid can be linearly amplified using the tagged DNA strands as the template. The primer used for linear amplification may or may not carry unique molecular identifiers (UMI). Because the errors in linear amplification are independent, by requiring the genomic variants to be shared between multiple independent linear copies, one can efficiently filter out amplification errors and therefore retaining accurate somatic variants, including somatic mutations, epimutations, single base DNA damage, single base mismatch, indels, DNA bulges, etc.
[058] According to some aspects of the disclosure, for example, a UMI may be a unique or random nucleotide sequence that includes a number of bases, such as at least, at most, a range of, or exactly 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22,
23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46,
47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70,
71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94,
95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131,
132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149,
150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167,
168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185,
186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, or more bases.
[059] According to one aspect of the disclosure, the strand-tagged nucleic acids can undergo an optional chemical conversion step, such as bisulfite conversion or nitrite, for the purpose to characterize epigenetic marks such as 5-methylcytosine (5mC), and 6-methyladenine (6mA).
[060] According to certain aspects of the disclosure, the transcriptome sequences and strand- tagged genome sequences are simultaneously amplified in the same reaction, for example a PCR reaction. The transcriptome sequences and strand-tagged genome sequences can be further split from each other by PCR amplification with specific primers, or by nucleic acid probes.
[061] According to certain aspects of the disclosure, the chemistry presented for tagging both strands of double-stranded DNA are also applicable to other double-stranded nucleic acids, or stem-looped nucleic acids, such as cDNA:RNA hybrid, double-stranded RNA, a RNA hairpin, a DNA hairpin, etc.
[062] According to certain aspects of the disclosure, the chemistry presented here can be applied to single tubes, microwell-plates, nanowells, plate-based liquid handlers, microfluidic systems, droplet platforms, etc.
EXAMPLES
[063] The following examples are included to demonstrate embodiments of the disclosure. It should be appreciated by those of skill in the art that the techniques disclosed in the examples that follow represent techniques discovered by the inventor to function well in the practice of the subject matter of the disclosure, and thus can be considered to constitute particular modes for its practice. However, those of skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific embodiments that are disclosed and still obtain a like or similar result without departing from the spirit and scope of the subject matter of the disclosure.
EXAMPLE I
Isolation of Complete Single Cells from Tissue Sample
[064] Histological tissue slides were prepared using commercial cryostat. Sections of 10-20 pm thickness of paraformaldehyde-fixed solid tissue were prepared. The cells of interest were cut from the section slide by laser microdissection microscopes (Leica, MMI etc.). The individual cells were dissociated from the micro-dissected tissue section using standard cell dissociation protocols. The complete single cells were collected into individual tubes.
Isolation of Complete Single Nuclei from Tissue Sample
[065] As shown in FIG. 2, frozen tumor tissue was chopped into small pieces with a blade. The tissue was resuspended in homogenization buffer for homogenization. The homogenate was passed through a cell strainer and centrifuged. The supernatant was discarded, and the pellet containing nuclei was resuspended in phosphate-buff ered saline (PBS) containing DNA staining dyes such as Hoechst 33342 (Invitrogen). The samples were subject to fluorescence-activated nucleus sorting (e.g., FACS), and single nuclei were collected into individual tubes.
Capture of Single Cell Transcriptome
[066] The single cells were lysed in a mild lysis buffer. The captured cells were thermally lysed and then stored in -80 °C. Next, a reverse transcription and template switching step was performed to generate the cDNA amplicons that have common adapters on both ends (e.g., R1 adapter and R2 adapter in FIG. 2) that allow primer binding in further amplifications, and genome library and transcriptome library are generated after amplification (FIG. 2).
Lysis of Genomic DNA
[067] A protease-based genomic DNA lysis mixture was added into the PCR tube to release genomic DNA. Next, the protease is inactivated by heating.
[068] An optional RNA digestion step is performed to remove RNA from the RNA:cDNA hybrid to prevent the cDNA amplicons from being fragmented in the next step. This RNA digestion step can be achieved by incubation with at least one RNase, such as RNase A, RNase H, and/or RNase If.
Transposase-based Genome Fragmentation and Strand-aware Tagging
[069] For the gDNA fragmentation step in FIG. 2, a custom Tn5 transposase was assembled with a transposon sequence comprised of the mosaic end sequence (referred to as “ME,” a sequence which may be recognized by transposases, shown in FIG. 3) and a specific restriction enzyme recognition site (referred to as “site RE,” shown in FIG. 3). A range of restriction enzymes can be used for this design. For example, in FIG. 3, the structure of the transposon sequence could be exemplified as: Annealing sequence 1 : 5’- /Phos/CTGTCTCTTATACACATCT -3’ (SEQ ID NO:1); Annealing sequence 2: 5’- [Handle_sequence][“site_RE”]AGATGTGTATAAGAGACAG-3’ (SEQ ID NO:2) and here the digestion sequence of a type II S restriction enzyme is used.
[070] As shown in FIG. 3, the genomic DNA is randomly fragmented by the customized Tn5 transposase, resulting in the short genomic DNA fragments with transposon sequences attached to both ends (referred to as “tagging”). The transposition reaction may be quenched by EDTA and heating to release the transposase.
[071] Next, for the gap filling/end repair step in FIG. 2 (or referred to as sealing in FIG. 3), a DNA polymerase was used for gap filling of the transposed DNA sequence. To minimize false positives, polymerases incapable of nick translation activity and strand displacement activity are utilized.
[072] Followed by the gap filling, a restriction enzyme (e.g., type II S restriction enzyme shown in FIG. 3) that recognizes the “site RE” is added to the product to generate a sticky or a blunt end. Whenever a sticky end is used (shown in FIG. 3 after the step of type II S restriction enzyme digestion), it allows efficient ligation. For the adapter ligation step in FIG. 3, the fragmented DNA is then ligated to an asymmetric Y-shaped or a looped adapter, catalyzed by a DNA ligase, such as T4 DNA ligase, T7 DNA ligase, Taq DNA ligase, ampligase thermostable DNA ligase, etc.
[073] The digestion by the designed restriction enzyme generated the designed sticky end, which will not only maximize the downstream ligation efficiency, but also reduce the length of adapters in the sequencing library, therefore facilitating the cost-effectiveness of sequencing. After end repair, the DNA fragments can be ligated to a Y-shape adapter, or a looped adapter. The asymmetric feature of the adapters allow distinguishing of both strands in the sequencing data, and thus genome strand-aware tagging is achieved.
[074] As shown in FIG. 2, each adaptor may comprise a UMI. In some embodiments, the marked base in the UMI (pointed to by an arrow in FIG. 2) stands for one base (e.g., A base) that is added to the 3’ ends of DNA fragments during the end repair process.
[075] The custom Tn5 transposase tagmentation reaction (shown in FIG. 3) could be performed with genomic DNA lysis.
[076] Alternatively, the custom Tn5 transposase tagmentation reaction (shown in FIG. 3) could be performed without genomic DNA lysis, such as by removing the genomic DNA lysis step. This design captures the open chromatin only, rather than the whole genome. The captured sequences essentially corresponded to ATAC-seq (Assay for Transposase- Accessible Chromatin using sequencing), in particular embodiments.
[077] It is worth noting that any template-based DNA synthesis could make the two DNA strands no longer independent, thereby introducing false positive variant callings. To minimize this effect, in some embodiments the end repair step could be performed in the absence of certain types of dNTP (dATP, dCTP, dGTP, and/or dTTP). Alternatively, modified nucleotides that prevent further extension, such as ddNTP, or amide-modified dNTP could be used here to replace one or more of the dNTPs (referred to as “ddNTP sealing,” shown in FIG. 3). Alternatively, the end repair step could be performed using a DNA polymerase without nick translation activity, strand displacement activity, or 5 ’->3’ exonuclease activity.
Simultaneous Amplification of Transcriptome and Genome Libraries
[078] Multiple oligos with random sequences (for example, N based oligos (6N-20N), where N refers to equal mix of A, T, G, C bases (25% each)) are added in the lysis buffer to allow the hybridization to transcripts entangled into the genomic DNA (e.g., in the form of DNA spaghetti). RNA digestion enzyme such as RNaseH is added to solution to cut RNA into smaller sized fragments, which then can be efficiently released from genomic DNA (e.g., in the form of DNA spaghetti). Centrifugation is performed to enrich the RNA fragments into the supernatant and genomic DNA (e.g., in the form of DNA spaghetti) into the bottom of tubes. Next, the supernatant is pipetted out for transcriptome using standard totalRNA based single-cell RNA-seq methods such as MATQ-seq (Multiple Annealing and dC-Tailing-based Quantitative single-cell RNA-seq) or Cel-seq (Cell Expression by Linear amplification and Sequencing). The genomic DNA (e.g., in the form of DNA spaghetti) at the bottom of tubes is used for genome profiling as described above.
[079] Additionally and/or alternatively, a reaction mix is added into each PCR tube for PCR amplification. This PCR mix contains two pairs of primers: one pair for transcriptome amplification, and the other pair for genome amplification. The transcriptome and genome libraries are amplified simultaneously for a few cycles, and then split into two for further amplification.
Single Cell Transcriptome Data Analysis
[080] Read 1 was mapped to human genome assembly (GRCh37) by using STAR (v2.5.3a). The uniquely mapped reads were then mapped to the gene annotations of GENCODE by using htseq-count with ‘intersection-strict’ mode and with option stranded=no’. To discern the amplicons from exons and introns, the ‘transcript’ feature and the ‘exon’ feature were used respectively. The UMI sequences (the first five bases of read 2) of the reads mapped to the gene regions (either the ‘transcript’ feature or the ‘exon’ feature) were extracted. The reads were then grouped by the UMI sequence and the gene that they were mapped to, as these reads were derived from the same original cDNA amplicon. If all the reads within the group were mapped to the exon regions of the corresponding gene, the original amplicon was classified as an exonic amplicon. Otherwise, the original amplicon was classified as an intronic amplicon. The number of exonic amplicons and intronic amplicons were then counted for each gene, and the exonic UMI count matrix and intronic UMI count matrix were generated.
Single Cell Genome Data Analysis
[081] The 3 -base UMI at the beginning of both reads was extracted using a Python script. Next, 3’ Truseq adapter sequences were trimmed using cutadapt. The reads were mapped to the hgl9 reference genome using bwa-mem v0.7.13-rl 126. Appropriately mapped read pairs with mapping quality no less than 50 were kept, and the resulting bam file was split by UMI. For each split, raw variant calling was performed by samtools mpileup followed by bcftools call -mv. All called variants were retained regardless of the variant-calling quality score. Tandem regions, centromere regions, and homopolymer regions were filtered out, resulting in a list of raw variants. A custom Python script was used to traverse each raw variant, a variant would be called a mutation if it fulfilled the three criteria: 1) covered in both paired-read directions, 2) at least two reads for each direction with sequencing quality no less than 30, and 3) all of the q>=30 bases support the variant base. By dividing the number of called germline heterozygous mutations (without deduplication of the same mutations called from different original DNA fragments) by the total number of germline heterozygous mutations, we can determine the mutation detection efficiency. A mutation would be called a somatic mutation if 1) the locus was covered by at least 10 reads in the matched normal bulk WGS with sequencing quality >=30, but zero read supports the variant call, and 2) the variant is not annotated in the dbSNP database. Somatic mutation load was calculated by dividing the number of somatic mutations detected by the mutation detection efficiency.
[082] Starting with the single-cell bam files, we marked the PCR duplicated reads using Picard v3.0.0 MarkDuplicates. Duplicated reads were then removed by samtools rmdup, and the deduplicated bam files were converted to bed files with bedtools v.2.26.0 bamtobed 9 and gzip compressed. The bed files were passed to the local version of Ginkgo master-build d7c7790 10 for single-cell CNV calling at the 500k bin-size resolution. An estimated ploidy file based on FANS was provided for global CNV correction in Ginkgo. All the remaining parameters were set as default. To cluster the single cells based on the CNV profile, we generated a pairwise distance matrix, where the distance between two cells was calculated as (1 - spearman correlation of copy number profiles). The distance matrix was then used for hierarchical clustering with ward.D2 linkage. To infer the time order, single cells were sorted by between-leaf or between-branch similarity based on the optimal leaf ordering method 11 implemented in R package seriation v.1.4.2. Fraction of genome altered is defined by the fraction of regions whose copy numbers deviated from the major copy number of the cell (chrY is excluded).
EXAMPLE II
Looped transposase-based Genome Fragmentation and Strand Tagging
[083] In some embodiments, the single cell gDNA lysate generated in EXAMPLE I may be transposed using the custom-assembled Tn5 transposase with looped or Y-shape transposons.
[084] As shown in FIG. 4, the steps of gDNA fragmentation and strand-aware tagging of gDNA fragments in EXAMPLE I can be streamlined into a one-step reaction. In this example, a custom Tn5 transposase was assembled with the following transposon sequence (5’- /Phos/CTGTCTCTTATACACATCT CCGAGCCCACGAGAC U
TCGTCGGCAGCGTC AGATGTGT AT AAGAGAC AG-3 ’ (SEQ ID NO:3, used for looped adaptor), or annealing the two sequences below: Annealing sequence 1 : 5’- /Phos/CTGTCTCTTATACACATCT CCGAGCCCACGAGAC-3’ (SEQ ID NO:4, used in the Y-shape adaptor); Annealing sequence 2: 5 ’-TCGTCGGCAGCGTC AGATGTGTATAAGAGACAG-3’; SEQ ID NO:5, used in the Y-shape adaptor). This resulted in a transposon with either a looped structure, or a Y-shape structure.
[085] Next, the genomic DNA is randomly fragmented by the Tn5 transposase, resulting in the short genomic DNA fragments with transposon sequences attached to both ends (referred to as “tagging”). The transposition reaction may be quenched by EDTA and heating to release the transposase.
[086] Truseq Adapter A and Truseq Adapter B may be connected by one or more U bases and form a loop and stem structure, which has a protruding end that can hybridize (or anneal) to the Tx-tagged end (possibly including ME) and Ty-tagged end (possibly including ME) of a genome DNA fragment, respectively (adapter hybridization step in FIG. 4). [087] In some embodiments. A DNA polymerase was used to fill the 9 base gaps generated by the Tn5 transposase, resulting in a DNA fragment with one nick on each strand (sealing step in FIG. 4). A DNA ligase was added to the reaction to seal the nick, such as an E. Coll DNA ligase, 9N DNA ligase, T4 DNA ligase, etc. Further amplification of transcriptome and genome libraries could follow EXAMPLE I (ligation step in FIG. 4).
[088] For the looped adapters, Uracil excision is performed by USER enzymes to remove U bases, as the result, convert the lopped adapters to Y shape adapters (USER digestion step in FIG. 4).
EXAMPLE III
Transposase-based Genome Fragmentation Using Oligos with U Bases and Strand Tagging
[089] As shown in FIG. 5, the steps of gDNA fragmentation and strand-aware tagging of gDNA fragments in EXAMPLE I can be achieved without using restriction enzyme. In this example, a custom Tn5 transposase was assembled with the following oligo duplex sequences with one or more U bases: one strand carries Tl-ME with one or more U bases (TCG/ideoxyU/CGGCAGCG/ideoxyU/CAGATGTGTA/ideoxyU/AAGAGACAG; SEQ ID NO: 6); the other strand carries the complementary sequencing Tl* (/5Phos/CTGTCTCTTATACACATCT; SEQ ID NO:7). Tl and Tl* hybridize and form double stranded parts of oligos, which will be bound by Tn5 transposase to produce Tn5 transposomes.
[090] Next, the genomic DNA is randomly fragmented by the Tn5 transposase, resulting in the short genomic DNA fragments with transposon sequences attached to both ends (referred to as “tagging”). The transposition reaction may be quenched by EDTA and heating to release the transposase.
[091] After tagging, a DNA polymerase was used for gap filling of the transposed DNA sequence. To minimize false positives, polymerases incapable of nick translation activity and strand displacement activity are utilized.
[092] The protruding ends can then be introduced by Uracil excision of the U bases. The primers with T2 sequence (GTCTCGTGGGCTCGGAGATGTGTAT; SEQ ID NO: 8) can be hybridized (or annealed) next to the recessive ends. A DNA ligase, such as an E. Colt DNA ligase, 9N DNA ligase, T4 DNA ligase, etc was added to the reaction to seal the nick to achieve efficient ligation between oligos and the recessive ends.
[093] Further, the independent construction of transcriptome libraries could follow EXAMPLE I.
EXAMPLE IV
Accurate Detection of Single Strand Variants by Additional Step with Linear Amplification
[094] The ligation product in EXAMPLES I, II, and/or III could be used as the starting materials for the methods encompassed in this example. Before PCR amplification, the product in this example underwent cycles of linear amplification. In this reaction, a primer was used with a unique molecular identifier (UMI). The structure of the primer could be exemplified as: (5’- [Illumina sequencing adapter]_[8-base sample index]_[8-base random UMI]_[Common adapter sequence introduced by ligation] -3’). This linear amplification resulted in a bunch of amplicons produced from the same original template, but with different UMIs. These linear amplicons were further amplified by PCR following the remaining steps in EXAMPLES I, II and/or III.
[095] In the bioinformatic analysis part, as shown in FIG. 6, reads from different linear copies of the same original template strand can be identified based on the UMI (each UMI with a different pattern represents a different UMI in FIG. 6). When a variant is required to be covered by at least two different UMI reads with sequencing quality exceeding a threshold (e.g., a Phred quality score or Q score of no less than 30 (q>=30)) for one strand or both strands of Watson and Crick strands in following three scenarios: 1) all of the bases that exceed the threshold (e.g., q>=30) support the variant bases existing in both Watson and Crick Strands (Left panel in FIG. 6), and one can accurately call these variants as the true mutations; 2) all of the bases that exceed the threshold (e.g., q>=30) support the variant base in a given strand (Watson strand or Crick strand, but not both), one can accurately call these variants that are unique to only one strand of DNA as a single base DNA damage, a DNA mismatch, a DNA bulge, an epimutation, a combination thereof, etc (Middle panel in FIG. 6); and 3) all of the bases that exceed the threshold (e.g., q>=30) support the variant base in a portion of reads of a given strand and with the same UMI, one can categorize these variants as amplification errors (Right panel in FIG. 6, and the top two UMIs are the same UMI). REFERENCES
[096] Citation or identification of any document in this application is not an admission that such document is available as prior art to the subject matter of the present disclosure.
[097] 1 Leung, K. et al. Robust high-performance nanoliter-volume single-cell multiple displacement amplification on planar substrates. Proc Natl Acad Sci U S A 113, 8484- 8489, doi: 10.1073/pnas.1520964113 (2016).
[098] 2 Dong, X. et al. Accurate identification of single-nucleotide variants in whole- genome-amplified single cells. Nature methods 14, 491-493, doi: 10.1038/nmeth.4227 (2017).
[099] 3 Xing, D., Tan, L., Chang, C.-H., Li, H. & Xie, X. S. Accurate SNV detection in single cells by transposon-based whole-genome amplification of complementary strands. Proceedings of the National Academy of Sciences 118, e2013106118, doi : doi : 10.1073/pnas.2013106118 (2021 ).
[0100] 4 Dean, F. B., Nelson, J. R., Giesler, T. L. & Lasken, R. S. Rapid amplification of plasmid and phage DNA using Phi 29 DNA polymerase and multiply-primed rolling circle amplification. Genome Res 11, 1095-1099, doi : 10.1101/gr.180501 (2001).
[0101] 5 Zong, C., Lu, S., Chapman, A. R. & Xie, X. S. Genome-wide detection of single nucleotide and copy number variations of a single human. Science 338, 1622-1626 (2012).
[0102] 6 Zhu, Q., Niu, Y., Gundry, M. & Zong, C. Single-Cell Damagenome Profiling Unveils Vulnerable Genes and Functional Pathways in Human Genome towards DNA damage. Science Advances in press (2021).
[0103] 7 Chen, C. et al. Single-cell whole-genome analyses by Linear Amplification via Transposon Insertion (LIANTI). Science 356, 189-194, doi: 10.1126/science.aak9787 (2017).
[0104] 8 Abascal, F. et al. Somatic mutation landscapes at single-molecule resolution. Nature 593, 405-410, doi : 10.1038/s41586-021-03477-4 (2021).
[0105] 9 Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841-842, doi: 10.1093/bioinformatics/btq033 (2010).
[0106] 10 Garvin, T. et al. Interactive analysis and assessment of single-cell copy -number variations. Nat Methods 12, 1058-1060, doi: 10.1038/nmeth.3578 (2015). [0107] 11 Bar-Joseph, Z., Gifford, D. K. & Jaakkola, T. S. Fast optimal leaf ordering for hierarchical clustering. Bioinformatics 17 Suppl 1, S22-29, doi : 10.1093/bioinformatics/ 17. suppl l . s22 (2001 ).

Claims

What is claimed:
1. A method for producing a library of nucleic acids representing nucleic acid of one or more cells and/or one or more subcellular structures, comprising: fragmenting genomic DNA using a Tn5 transposase assembled with a sequence comprised of a restriction enzyme site, to produce genomic DNA fragments comprising transposase sequence comprising the site at both ends; repairing the ends of the genomic DNA fragments; subjecting the genomic DNA fragments to the restriction enzyme to produce digested molecules; and ligating adapters to both ends of the digested molecules to produce adapter-ligated molecules, wherein each adapter at both ends of the adapter-ligated molecule comprises a unique sequence to label the DNA fragment that it was ligated to and a common sequence, optionally for amplification.
2. The method of claim 1, further comprising generating RNA-representing double stranded DNA molecules representing part or all of a transcriptome of the one or more cells, wherein the molecules comprise common, amplifiable handle sequences on both 5’ and 3’ ends.
3. The method of claim 2, wherein the generating step comprises reverse transcription with template switching activity.
4. The method of claim 2, wherein the generating step comprises reverse transcription without template switching activity.
5. The method of claim 4, wherein the method further comprises a homopolymer tailing reaction.
6. The method of claim 4, wherein the method further comprises a ligation reaction to ligate a nucleic acid sequence to the end of cDNA.
7. The method of claim 4, wherein the method further comprises a second strand synthesis reaction that utilizes cDNA as a template.
8. The method of claim 2, wherein the generating step lacks reverse transcription.
9. The method of claim 8, wherein the generating comprises hybridization of the RNA with one or multiple probe nucleic acids.
10. The method of claim 8, wherein the generating comprises direct ligation of adapters containing common handle sequences to the original cellular RNA.
11. The method of any one of the preceding claims, wherein the restriction enzyme is a Type II S restriction enzyme.
12. The method of any one of the preceding claims, wherein the fragmenting of the genomic DNA is by mechanical force, by using one or more enzymes, by heating, by pH changes, by chemicals, or a combination thereof.
13. The method of claim 12, wherein the mechanical force comprises sonication and/or hydrodynamic shearing.
14. The method of claim 12, wherein the enzymes comprise fragmentase, transposase, nuclease, or a combination thereof.
15. The method of any one of the preceding claims, wherein the unique sequence comprises a strand-specific barcode, a strand-specific adapter orientation, or both.
16. The method of any one of the preceding claims, wherein the repairing occurs using a DNA polymerase.
17. The method of claim 16, wherein the DNA polymerase is incapable of nick translation activity, strand displacement activity, and/or 5 ’->3’ exonuclease activity.
18. The method of any one of the preceding claims, wherein the adapters are asymmetric.
19. The method of any one of the preceding claims, wherein the adapters are Y-shaped or looped.
20. The method of any one of the preceding claims, wherein the adapter-ligated molecules are amplified.
21. The method of claim 20, wherein the adapter-ligated molecules are amplified by linear amplification.
22. The method of claim 20, wherein the adapter-ligated molecules are amplified by nonlinear amplification.
23. The method of any one of the preceding claims, wherein the subcellular structures comprise organelles.
24. The method of claim 23, wherein the organelle is a mitochondria, nucleus, endoplasmic reticulum, chloroplast, Golgi apparatus, ribosome, or combination thereof.
25. The method of any one of the preceding claims, wherein the one or more cells and/or one or more subcellular structures are obtained from storage.
26. The method of any one of the preceding claims, wherein the one or more cells and/or one or more subcellular structures are obtained from a sample from one or more individuals.
27. The method of any one of the preceding claims, further comprising the step of obtaining the one or more cells and/or one or more subcellular structures from a sample from one or more individuals.
28. The method of claim 27, wherein the sample comprises solid matter.
29. The method of claim 27, wherein the solid matter comprises tissue.
30. The method of claim 27, wherein the tissue is homogenized to produce single cells.
31. A method for producing a library of nucleic acids representing nucleic acid of one or more cells and/or one or more subcellular structures, comprising: fragmenting genomic DNA using a Tn5 transposase assembled with a sequence comprised of one or more U bases and/or a first adaptor, to produce genomic DNA fragments comprising the sequence at both ends; repairing the ends of the genomic DNA fragments; introducing one or more protruding ends to the genomic DNA fragments by excision of the one or more U bases; and ligating a second adapter to the one or more protruding ends to produce adapter-ligated molecules, wherein each of the first adaptor and second adaptor comprises a unique sequence to label the DNA fragment that it was ligated to and a common sequence, optionally for amplification.
32. The method of claim 31, wherein the step of introducing one or more protruding ends to the genomic DNA fragments by excision of the one or more U bases comprises using Uracil for the excision.
33. The method of claim 31 or 32, further comprising generating RNA-representing double stranded DNA molecules representing part or all of a transcriptome of the one or more cells, wherein the molecules comprise common, amplifiable handle sequences on both 5’ and 3’ ends.
34. The method of claim 33, wherein the generating step comprises reverse transcription with template switching activity.
35. The method of claim 33, wherein the generating step comprises reverse transcription without template switching activity.
36. The method of claim 35, wherein the method further comprises a homopolymer tailing reaction.
37. The method of claim 35, wherein the method further comprises a ligation reaction to ligate a nucleic acid sequence to the end of cDNA.
38. The method of claim 35, wherein the method further comprises a second strand synthesis reaction that utilizes cDNA as a template.
39. The method of claim 33, wherein the generating step lacks reverse transcription.
40. The method of claim 39, wherein the generating comprises hybridization of the RNA with one or multiple probe nucleic acids.
41. The method of claim 39, wherein the generating comprises direct ligation of adapters containing common handle sequences to the original cellular RNA.
42. The method of any one of the preceding claims, wherein the fragmenting of the genomic DNA is by mechanical force, by using one or more enzymes, by heating, by pH changes, by chemicals, or a combination thereof.
43. The method of claim 42, wherein the mechanical force comprises sonication and/or hydrodynamic shearing.
44. The method of claim 42, wherein the enzymes comprise fragmentase, transposase, nuclease, or a combination thereof.
45. The method of any one of the preceding claims, wherein the unique sequence comprises a strand-specific barcode, a strand-specific adapter orientation, or both.
46. The method of any one of the preceding claims, wherein the repairing occurs using a DNA polymerase.
47. The method of claim 46, wherein the DNA polymerase is incapable of nick translation activity, strand displacement activity, and/or 5 ’->3’ exonuclease activity.
48. The method of any one of the preceding claims, wherein the adapters are asymmetric.
49. The method of any one of the preceding claims, wherein the adapters are Y-shaped or looped.
50. The method of any one of the preceding claims, wherein the adapter-ligated molecules are amplified.
51. The method of claim 50, wherein the adapter-ligated molecules are amplified by linear amplification.
52. The method of claim 50, wherein the adapter-ligated molecules are amplified by nonlinear amplification.
53. The method of any one of the preceding claims, wherein the subcellular structures comprise organelles.
54. The method of claim 53, wherein the organelle is a mitochondria, nucleus, endoplasmic reticulum, chloroplast, Golgi apparatus, ribosome, or combination thereof.
55. The method of any one of the preceding claims, wherein the one or more cells and/or one or more subcellular structures are obtained from storage.
56. The method of any one of the preceding claims, wherein the one or more cells and/or one or more subcellular structures are obtained from a sample from one or more individuals.
57. The method of any one of the preceding claims, further comprising the step of obtaining the one or more cells and/or one or more subcellular structures from a sample from one or more individuals.
58. The method of claim 57, wherein the sample comprises solid matter.
59. The method of claim 57, wherein the solid matter comprises tissue.
60. The method of claim 57, wherein the tissue is homogenized to produce single cells.
PCT/US2024/059713 2023-12-12 2024-12-12 Methods for simultaneous characterization of genome and transcriptome of a single cell Pending WO2025128786A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202363608947P 2023-12-12 2023-12-12
US63/608,947 2023-12-12

Publications (1)

Publication Number Publication Date
WO2025128786A1 true WO2025128786A1 (en) 2025-06-19

Family

ID=96058428

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2024/059713 Pending WO2025128786A1 (en) 2023-12-12 2024-12-12 Methods for simultaneous characterization of genome and transcriptome of a single cell

Country Status (1)

Country Link
WO (1) WO2025128786A1 (en)

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20180155709A1 (en) * 2015-05-28 2018-06-07 Illumina Cambridge Limited Surface-based tagmentation
US20180245069A1 (en) * 2017-02-21 2018-08-30 Illumina, Inc. Tagmentation using immobilized transposomes with linkers
US20190127792A1 (en) * 2017-11-02 2019-05-02 Bio-Rad Laboratories, Inc. Transposase-based genomic analysis
US20210079386A1 (en) * 2014-04-29 2021-03-18 Illumina, Inc. Multiplexed Single Cell Gene Expression Analysis Using Template Switch and Tagmentation
WO2023034913A2 (en) * 2021-09-02 2023-03-09 Baylor College Of Medicine Methods of in situ total rna-based transcriptome profiling for large-scale subcellular structure profiling
WO2023225519A1 (en) * 2022-05-17 2023-11-23 10X Genomics, Inc. Modified transposons, compositions and uses thereof

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20210079386A1 (en) * 2014-04-29 2021-03-18 Illumina, Inc. Multiplexed Single Cell Gene Expression Analysis Using Template Switch and Tagmentation
US20180155709A1 (en) * 2015-05-28 2018-06-07 Illumina Cambridge Limited Surface-based tagmentation
US20180245069A1 (en) * 2017-02-21 2018-08-30 Illumina, Inc. Tagmentation using immobilized transposomes with linkers
US20190127792A1 (en) * 2017-11-02 2019-05-02 Bio-Rad Laboratories, Inc. Transposase-based genomic analysis
WO2023034913A2 (en) * 2021-09-02 2023-03-09 Baylor College Of Medicine Methods of in situ total rna-based transcriptome profiling for large-scale subcellular structure profiling
WO2023225519A1 (en) * 2022-05-17 2023-11-23 10X Genomics, Inc. Modified transposons, compositions and uses thereof

Similar Documents

Publication Publication Date Title
US11584959B2 (en) Compositions and methods for selection of nucleic acids
JP7223790B2 (en) Methods for targeted genomic analysis
AU2021204166B2 (en) Reagents, kits and methods for molecular barcoding
EP4239067A2 (en) Transposase compositions for reduction of insertion bias
CN108026575A (en) Methods for Amplifying Nucleic Acid Sequences
EP3615683B1 (en) Methods for linking polynucleotides
US20240368695A1 (en) Embryonic nucleic acid analysis
CA3213037A1 (en) Blocking oligonucleotides for the selective depletion of non-desirable fragments from amplified libraries
CN112585279A (en) RNA library building method and kit
WO2005090599A2 (en) Methods and adaptors for analyzing specific nucleic acid populations
JP2015509724A (en) Method for identifying VDJ recombination products
US20260002203A1 (en) Primary template-directed amplification and methods thereof
US20180100180A1 (en) Methods of single dna/rna molecule counting
WO2025085821A1 (en) Methods, systems, and compositions for cell storage and analysis
WO2025128786A1 (en) Methods for simultaneous characterization of genome and transcriptome of a single cell
EP4594728A2 (en) Methods and compositions for fixed sample analysis
JP7688417B2 (en) Methods for generating a population of polynucleotide molecules
CN119095979A (en) Target enrichment
WO2018081666A1 (en) Methods of single dna/rna molecule counting
CN118086457B (en) Construction and application of DNA library
HK40098903A (en) Transposase compositions for reduction of insertion bias
JP2018126078A (en) Methods for analyzing rna

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24904854

Country of ref document: EP

Kind code of ref document: A1