EP4627111A2 - Cost-effective long read metagenomic sequencing - Google Patents
Cost-effective long read metagenomic sequencingInfo
- Publication number
- EP4627111A2 EP4627111A2 EP24725039.2A EP24725039A EP4627111A2 EP 4627111 A2 EP4627111 A2 EP 4627111A2 EP 24725039 A EP24725039 A EP 24725039A EP 4627111 A2 EP4627111 A2 EP 4627111A2
- Authority
- EP
- European Patent Office
- Prior art keywords
- methylation
- seq
- motif
- methylated
- strand oligo
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6806—Preparing nucleic acids for analysis, e.g. for polymerase chain reaction [PCR] assay
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2600/00—Oligonucleotides characterized by their use
- C12Q2600/154—Methylation markers
Definitions
- the present inventive concept is directed to methods and compositions for selectively amplifying a nucleic acid of one or more prokaryotic organisms of interest in a microbiome sample.
- Metagenomics have enabled the comprehensive and culture-independent study of microbiomes. However, for many applications, only certain taxa are of interest in a microbiome sample, for example, pathogenic bacteria, beneficial microbes, or taxa of low relative abundance. It is highly desirable that a method can effectively enrich bacterial taxa of interest directly from the microbiome.
- the present disclosure is based on, in part, the suppressing discovery' that natural self vs. non-self genome differentiation as provided by bacterial DNA methylation can be exploited to rationally choose methylation-sensitive restriction enzymes (REs), individually or in combination, to deplete host DNA and most background microbial DNA while enriching bacterial taxa of interest. Accordingly, the present disclosure herein provides for a new method of selectively enriching any bacterial target of interest in a sample of mixed bacterial taxa.
- REs methylation-sensitive restriction enzymes
- a method of selectively amplifying a nucleic acid of one or more prokaryotic organisms of interest (target organisms) in a microbiome sample comprises: (a) obtaining or having obtained a microbiome sample comprising a plurality of prokaryotic-organisms, wherein the microbiome sample comprises genomic DNA (gDNA) from one or more target organisms (target gDNA) and gDNA from one or more prokaryotic organisms that are not of interest (background organism gDNA); (b) isolating gDNA from the microbiome sample to obtain target gDNA fragments and background organism gDNA fragments; (c) selecting one or more methylation-sensitive restriction enzymes based on the DNA methylome (a collection of DNA methylation sequence motifs) of the one or more target organisms wherein the one or more methylation restriction enzymes target motifs that are primarily methylated in the target gDNA (methylated
- the methods provided herein further comprise an end repair and/or a dA-tailing step prior to the linking of the universal PCR adapters.
- the universal PCR adaptor can comprise an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary' and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, and wherein the hybridized region does not comprise a nonmethylated motif targeted by the one or more restriction enzyme.
- the universal PCR adaptor further comprises a *T at the 3’ end of the upper strand oligo and/or is phosphorylated at the 5’ end of the lower strand oligo.
- the upper strand oligo may have a nucleotide sequence comprising any one of SEQ ID NOs: 1 and 3-18 and the lower strand oligo may have a nucleotide sequence comprising any one of SEQ ID NOs: 2 and 19-34.
- the upper strand oligo may have a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 36; or the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 38.
- each upper-strand oligo of the universal primers herein can comprise a primer binding site complementary’ to a forward barcoded primer, and each lower- strand oligo comprises a primer binding site complementary’ to a reverse barcoded primer.
- the methods may further comprise deconvolving genomes of prokaryotic organisms in the microbiome sample to identify a methylation profile of the target prokaryotic organism, wherein deconvolving genomes comprises: A) obtaining a microbiome sample comprising a plurality of prokaryotic-organisms; B) sequencing nucleic acids of the prokaryotic organisms using single-molecule long reads sequencing technology, wherein the sequencing comprises the step of identifying methylated nucleotides, and at least one of the steps of: i. sequencing single molecule reads of nucleic acids; ii.
- step (D) determining nucleic acid methylation profiles of the assembled contigs or the single molecule reads in the microbiome sample based on motifs identified in step (D); F) separating the assembled contigs and/or the single molecule reads into bins corresponding to distinct prokaryotic organisms based on the methylation profiles of step (E); and G) assembling the bins of step (F), thereby obtaining assembled genomes of the distinct bacterial organisms in the microbiome sample, thereby deconvolving genomes of the prokaryotic organisms in the microbiome sample.
- the method further comprises the step of combining the methylation profiles of step (E) with other sequence features of the nucleic acids of the prokaryotic organisms in the microbiome sample prior to separating the assembled contigs and/or the single molecule reads into bins.
- the target prokaryotic organism(s) comprises one or more of an E. coll strain.
- Akkermansia muciniphila. Alistipes fmegoldii and Bacteroides ovatus. Bifidobacterium longum, Ruminococcus bromii, Eubacterium rectale, Bacteroides uniformis D7, Dorea longicatena and Barnesiella intestinihominis, Eubacterium sp. CAG:248, Alistipes onderdonkii, Enterocloster clostridioformis, Bacteroides fragilis, Gordonibacter pamelaea,
- Adlercreutzia sp. 8CFCBH1 Adlercreutzia equolifaciens, or any combination thereof.
- the target prokaryotic organism(s) are selected from specific pathogenic bacteria, beneficial bacteria, probiotics, or prokaryotic organisms with low relative abundance in a microbiome sample as determined by standard metagenomic sequencing.
- the target prokaryotic organism(s) comprising a methylated CTS6mAG motif may comprise one or more species of Alistipes finegoldii and/or Bacteroides ovatus
- the target prokaryotic organism(s) comprising a methylated RG6mATCY motif may comprise one or more species of Bifidobacterium longum, Ruminococcus bromii. Eubacterium rectale and Bacteroides uniformis
- the target prokaryotic organism(s) comprising a methylated G6mATGC motif may comprise one or more species of Dorea longicatena.
- the target prokary otic organism(s) comprising a methylated G6mATGC motif and methylated CTNN6mAG motif may comprise one or more species of Barnesiella intestinihominis.
- any of the foregoing or related methods may further comprise sequencing the nucleic acid of the amplified target prokary otic organism(s).
- the sequencing is performed using a single-molecule long read real time (SMRT) technology, nanopore sequencing technology 7 or other short read sequencing technologies including but not limited to Illumina sequencing platforms.
- SMRT single-molecule long read real time
- nanopore sequencing technology 7 or other short read sequencing technologies including but not limited to Illumina sequencing platforms.
- the microbiome sample may 7 be obtained from soil, air, water, sediment, oil, and combinations thereof.
- the microbiome sample is obtained from water selected from marine water, fresh water, and rainwater.
- the microbiome sample is obtained from a subject selected from a protozoa, an animal, or a plant.
- the subject is a mammal.
- the subject is human.
- the subject is an infant.
- the adaptor comprises an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, wherein the hybridized region does not comprise a motif targeted by a methylation sensitive restriction enzyme or does comprise a methylated nucleotide in a motif targeted by the methylation sensitive restriction enzyme.
- the universal adaptor provided herein further comprises a *T at the 3’ end of the upper strand oligo and/or a phosphorylated base at the 5’ end of the lower strand oligo.
- a subgroup of bacteria of interest can be enriched (circled by the red dashed line), by depleting the rest of the bacteria (the gray circle).
- mEnrich-seq allow researchers to examine different subgroups of bacteria in the microbiome sample (circled by individual yellow or blue dashed lines). mEnrich-seq can also be tailored with the use of multiple REs to more specifically enrich bacterial taxa with multiple methylation motifs (as illustrated by both the yellow and blue scissors, enriching the bacterium shared between blue and yellow dashed line).
- FIG. 2D depicts motif frequency on the mEnrich-seq reads that were mapped to N. gonorrhoeae. In replicates 1 and 2, 94.66% and 92.26% of reads that mapped to N. gonorrhoeae do not have any GATC sites (blue), respectively.
- FIG. 2F depicts scatter plots for three urine samples (UTI-1 to -3) showing the percentage of reads mapped to E. coli, human and other taxa from the input (x-axis) vs. mEnrich-seq (y-axis). Across all the three samples, E. coli reads are significantly enriched in mEnrich-seq.
- FIG. 3A-3D depict the enrichment of A. muciniphila from fecal samples by mEnrich-seq
- FIG. 3A shows scatter plots for three fecal samples (GUT-1 to -3) showing the percentage of reads mapped to A. muciniphila and other taxa from the input (x- axis) and mEnnch-seq (y-axis).
- top 50 abundant species, in input or mEnrich-seq are shown.
- the size of the dots represents the fold change by mEnrich-seq over the input.
- Significant increases in read % are observed for A.
- FIG. 3B shows fold enrichment of reads that map to A. muciniphila from mEnrich-seq over the input for GUT-1 to -3.
- FIG. 3C shows A. muciniphila genome coverage by sequencing reads of GUT-1 to -3. Reads from mEnrich-seq (orange) and the input standard metagenomic sequencing (blue) were mapped to the genome of corresponding A. muciniphila isolates (sequenced and assembled beforehand for evaluation purpose), respectively.
- FIG. 3D shows analysis of RAATTY motif frequency (per kilobase) for the taxa enriched by mEnrich-seq (XapI as RE) in GUT-1 to -3.
- X-axis RAATTY motif frequency on the reference genomes of enriched taxa;
- Y-axis read-level RAATTY frequency in mEnrich-seq.
- the green dashed line circles A. muciniphila. which is enriched due to its methylated RA6mATTY motif sites.
- the red dashed line circles the bacterial taxa enriched due to the sparse RAATTY frequency on their genomes.
- FIG. 4A-4J depict enrichment sequencing of bacterial taxa based on de novo discovered methylation motifs, wherein: FIG. 4A depicts a workflow for rational design of mEnrich-seq based on de novo motif discovery.
- An initial metagenomic sequencing was conducted to estimate the relative abundance of different taxa. The goal is to design mEnrich- seq to enrich taxa with relatively low abundance. To do so, de novo methylation motif discovery' was performed from the initial metagenomic sequencing data. Importantly, motif discovery' does not need fully assembled genomes, because a partially assembled contig can already inform us about the methylation motifs from a strain.
- FIG. 4B depicts fold enrichment of six bacterial strains with relatively low abundance; enrichment by mEnrich-seq based on de novo discovered methylation motifs.
- One of the strain B. intestinihominis was further ennched using multiple REs simultaneously to provide greater enrichment (85-fold) than a single RE (30-fold), illustrating the flexible design of mEnrich-seq when multiple methylation motifs are discovered on the strain of interest.
- FIG. 4C depicts scatter plots showing the percentage of reads mapped to bacteria from the input (x-axis) and mEnrich-seq using the RE (BspCNI) targeting the CTCAG motif (y-axis).
- the size of the dots represents the fold change by mEnrich-seq over the input.
- FIG. 4D depicts analysis of the CTCAG motif frequency (per kilobase) for the taxa enriched by mEnrich-seq (BspCNI as RE). X-axis.
- FIG. 4E depicts data as described for FIG. 4C but for the RE targeting the RGATCY moti.
- FIG. 4F depicts data as described for FIG. 4D but for the RE targeting the RGATCY motif.
- FIG. 4G depicts data as described for FIG. 4C but for the RE targeting the GATGC motif.
- FIG. 4H depicts data as described for FIG.
- FIG. 41 depicts data as described for FIG. 4C but for the multiple REs targeting the GATGC motif and CTNNAG motif simultaneously.
- FIG. 4J shows that mEnrich-seq improved genome assembly compared to standard metagenomic sequencing. The left shows the completeness (%) of genome assembly; blue bars show the input and orange bars show the mEnrich-seq; * completeness less than 1%. On the right, contig N50 length is shown in Mb. For each targeted species, reads from input and mEnrich-seq were normalized by random subsampling to make sure reads originated from the targeted species have the same yield for genome assembly comparison.
- FIG. 5A-5E depict experimental data assessing the broad applicability of mEnrich- seq, wherein: FIG. 5A depicts a histogram for the number of methylation motifs in each bacterial genome across the 4,601 bacterial methylomes mapped to date (as of 07/18/2022, REBASE). X-axis, number of motifs per bacterium; Y -axis, number of bacterial species having a given number of motif.
- FIG. 5B depicts a histogram for the number of non-degenerate bases (excluding 'N’s) per motif across 4683 motifs. For example, both GATC and CTNNAG have 4 non-degenerate bases, FIG.
- FIG. 5C depicts a phylogenic tree containing 4,601 bacterial strains with mapped methylomes to date.
- the bacterial strains that can be targeted (enriched by depleting the background metagenomic DNA) by at least one commercially available REs are highlighted with colors, demonstrating broad distribution across the tree.
- Number of motifs recognized by commercially available REs (and motif length) for a given bacteria are color coded as detailed in the panel at the bottom right
- FIG. 5D depicts a histogram for number of bacteria targeted (enriched by depleting the background metagenomic DNA) by each commercially available RE (please note that a single motif can be targeted by multiple restriction enzymes, namely, isoschizomers).
- FIG. 5E shows taxonomic distribution of the seven most prevalent methylation motifs.
- the taxonomic information is indicated by color-coding branches by order (top 27 orders are color coded). The distribution of top seven motifs is highlighted by colored squares around the phylogenetic tree.
- FIG. 6A-6C depict mEnrich-seq of urine samples using Nanopore Flongle flow cells and wherein: FIG. 6A shows three urine samples were sequenced as the Input: UTI-1 contained mostly bacterial DNA, while UTI-2 and UTI-3 were dominated by human DNA, FIG. 6B depicts scatter plots for six urine samples (UTI-1 to -6) showing the percentage of reads mapped to E. coli, human and other taxa from the input (x-axis) vs. mEnrich-seq (y- axis). The UTI-1 to -3 samples demonstrate consistent patterns as shown in FIG.
- FIG. 12A-12B depict agarose gels of enriched low-abundance bacterial taxa from the adult fecal sample based on de novo discovered methylation motifs, wherein: FIG. 12A depicts gels for the adult fecal sample digested by different REs. Left, the Mfll-digested the adult fecal sample DNA (targeting RGATCY motif) shows the band at -5- lOkb; the BspCNI (CTC AG as recognition site, part of the CTSAG motif)- digested sample shows a faint band at -5- lOkb (based upon the DNA ladder).
- targeting RGATCY motif shows the band at -5- lOkb
- BspCNI CTC AG as recognition site, part of the CTSAG motif
- FIG. 12B depicts gels for PCR bands amplified by the barcoded primers, using RE-digested adult fecal DNA (from Fig. 7a) as the template.
- the PCR bands of RE(s)-digested DNA are shown around 5—10 kb.
- FIG. 13 depicts a plot showing de novo methylation motif discovery from an initial shallow' sequencing run (400k PacBio Sequel II CCS reads). Many methylation motifs were discovered from SMRT sequencing data, even those from bacterial genomes with -10% completeness. The relative abundance of each bacterial genome is included in the x-axis label in red color. 4-mer motifs (those with four non-degenerate bases) and similarly 5-mer and 6- mer motifs are color coded [0050]
- FIG. 14A-14B shows via a direct comparison that the mEnrich-seq design has better fold enrichment than REMoDE.
- FIG. 14A-14B shows via a direct comparison that the mEnrich-seq design has better fold enrichment than REMoDE.
- FIG. 15A-15B show that mEnrich-seq design has better fold enrichment than REMoDE when they are compared with the same, real microbiome samples with DNA fragmentation and damages.
- FIG. 15A depicts agarose gel electrophoresis tested once on genomic DNA isolated from Mock-3 (illustrating the ideal high molecular weight gDNA from cell culture) and two infant fecal samples (representing DNA quality’ in real applications).
- FIG. 15B depicts data comparing the fold enrichment of A. muciniphila from two infant fecal samples using mEnrich-seq, REMoDE, or REMoDE with DNA repair beforehand (see Supplementary’ Information). As DNA repair is part of mEnrich-seq by default, this step is also applied to REMoDE for a fair comparison.
- FIG. 16 depicts an evaluation of the benefit of adapter ligation before vs. after RE- digestion using the Mock-3.
- the bar graph demonstrates that adapter ligation before RE- digestion indeed has better fold ennchment.
- FIG. 17 shows that an alternative adapter has comparable enrichment efficiency as the original adapter.
- the fold enrichment of A. muciniphila (using Xapl as RE) was tested on Mock-3 with either the original adapter (having an upper and loyver oligonucleotide corresponding to SEQ ID NOs: 35 and 36, respectively) or the alternative adapter (having an upper and loyver oligonucleotide corresponding to SEQ ID NOs: 37 and 38, respectively).
- FIG. 18 shows that GC content does not introduce systematic bias on read coverage across the E. coll genome analyzed by mEnrich-seq. Analysis was performed on Mock-1 and Mock-2 (E. coll relative abundance as 0.91% and 1.20%, see Supplementary Information). For a fair comparison, the same yield of data from both input and mEnrich-seq were used in data analysis and figures.
- Each dot represents a 5 kb genomic region, colored by GC content.
- X- axis number of input reads mapped to each 5 kb genomic region;
- Y-axis number of mEnrich- seq reads mapped to each 5 kb genomic region. To ease visualization, a random value was added on x-values to avoid overlapping dots.
- FIG. 19 shows that GC content does not introduce systematic bias on read coverage across the A. muciniphila genome by mEnrich-seq.
- the same type of scatter plots (as the ones above for E. coli) made with sequencing data from three human gut samples (A. muciniphila relative abundance: 0.48%, 0.52% and 1.52%, respectively) to characterize the enrichment of A. muciniphila. Consistent with the Circos plot for the same samples, read depths were capped at the 95% quantile to ease visualization.
- the present disclosure is based, at least in part, on the discovery that natural self vs. non-self genome differentiation as provided by bacterial DNA methylation can be exploited to rationally choose methylation-sensitive restriction enzymes (REs), individually or in combination, to deplete host DNA and most background microbial DNA while enriching bacterial taxa of interest. Accordingly, the present disclosure herein provides for a new method of selectively enriching any bacterial target of interest in a sample of mixed bacterial taxa.
- REs methylation-sensitive restriction enzymes
- the term “about.” can mean relative to the recited value, e.g . amount, dose, temperature, time, percentage, etc., ⁇ 10%, ⁇ 9%, ⁇ 8%, ⁇ 7%, ⁇ 6%, ⁇ 5%, ⁇ 4%, ⁇ 3%, ⁇ 2%, or ⁇ 1%.
- nucleic acid' refers to deoxyribonucleic acids (DNA) or ribonucleic acids (RNA) and polymers thereof in either single- or double-stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogues of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g.. degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences as well as the sequence explicitly indicated.
- DNA deoxyribonucleic acids
- RNA ribonucleic acids
- degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and/or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolmi et ak.Afo/. Cell. Probes 8:91-98 (1994)).
- a method of selectively amplifying a nucleic acid of one or more prokaryotic organisms of interest (target organisms) in a microbiome sample comprises: (a) obtaining a microbiome sample comprising a plurality of prokaryotic-organisms, wherein the microbiome sample comprises genomic DNA (gDNA) from one or more target organisms (target gDNA) and gDNA from one or more prokaryotic organisms that are not of interest (background organism gDNA); (b) isolating gDNA from the microbiome sample to obtain target gDNA fragments and background organism gDNA fragments; (c) selecting one or more methylationsensitive restriction enzymes based on the DNA methylome (a collection of DNA methylation sequence motifs) of the one or more target organisms wherein the one or more methylation restriction enzymes target motifs that are primarily methylated in the target gDNA (methylated motifs) and
- the methods here are, in part, directed to methods that improve upon other methods of using methylation sensitive restriction enzymes to enrich target nucleic acid populations, by ligating a set of universal adaptors to the target (and non-target) DNA before any digestion occurs.
- the universal adaptors also referred to herein as “universal PCR adaptors’'
- Suitable universal PCR adaptors are described further in Section (III) below.
- These universal adaptors may be targeted by suitable PCR primers after digestion has occurred to selectively amplify and/or sequence any un-digested DNA (i.e., any DNA fragments still containing universal adaptors at both the 5‘ and 3‘ ends).
- Suitable PCR primers can target certain primer binding motifs in the universal PCR adaptors as described further in Section (III) below.
- the methylated motif comprises a nucleic acid sequence motif having methylation on one or more nucleotides, wherein the methylation optionally comprises N4-methylcytosine (4mC), N6-methyladenine (6mA) or 5- Methylcytosine (5mC), and wherein the methylated motif optionally comprises GATC, RAATTY where R is A or G and Y is C or T, or CTSAG where S is G or C, RGATCY where R is A or G and Y is C or T, GATGC, CTNNAG and each A is N6-methyladenine (6mA).
- the methylation optionally comprises N4-methylcytosine (4mC), N6-methyladenine (6mA) or 5- Methylcytosine (5mC)
- the methylated motif optionally comprises GATC, RAATTY where R is A or G and Y is C or T, or CTSAG where S is G or C, RGATCY where R is A or G and Y
- the non-methylated motif comprises a nucleic acid sequence motif that is recognized and cleaved by a restriction enzyme, and wherein the non-methylated motif optionally comprises GATC, RAATTY where R is A or G and Y is C or T, or CTSAG where S is G or C, RGATCY where R is A or G and Y is C or T, GATGC, CTNNAG and each A is non-methylated.
- the one or more methylation sensitive restriction enzymes may comprise a class of natural and/or artificial restriction enzymes that are sensitive to DNA methylation, and optionally comprise methylation-sensitive restriction enzy mes DpnII, XapI, Apol, BspCNI, BseMII, Mfll, BstX2I, BstYI, Mfll, XhoII, SfaNI, AflII, Acul. Bpml, BpuEI, PstI, any methylation-sensitive isoschizomer thereof, or a combination of any thereof.
- the methylation-sensitive restriction enzyme comprises a restriction enzyme whose cleavage ability is blocked by the methylated motif sites.
- the method comprises identifying a methylation profile of the target prokaryotic organism and selecting the one or more restriction enzy mes based on the methylation profile of the target prokaryotic organism.
- the method further comprises deconvoluting genomes of prokaryotic organisms in the microbiome sample to identify a methylation profile of the target prokaryotic organism, wherein deconvoluting genomes comprises: A) obtaining a microbiome sample comprising a plurality of prokary otic- organisms; B) sequencing nucleic acids of the prokaryotic organisms using single-molecule long reads sequencing technology, wherein the sequencing comprises the step of identifying methylated nucleotides, and at least one of the steps of: i. sequencing single molecule reads of nucleic acids; ii.
- step (D) determining nucleic acid methylation profiles of the assembled contigs or the single molecule reads in the microbiome sample based on motifs identified in step (D); F) separating the assembled contigs and/or the single molecule reads into bins corresponding to distinct prokaryotic organisms based on the methylation profiles of step (E); G) assembling the bins of step (F), thereby ⁇ obtaining assembled genomes of the distinct bacterial organisms in the microbiome sample, thereby deconvoluting genomes of the prokaryotic organisms in the microbiome sample.
- the method of deconvoluting genomes further comprises the step of combining the methylation profiles of step (E) with other sequence features of the nucleic acids of the prokaryotic organisms in the microbiome sample prior to separating the assembled contigs and/or the single molecule reads into bins.
- Additional methods useful to deconvolve genomes to derive at ideal combinations of methylation sensitive restriction enzymes or that may be used in combination with any of the methods and compositions provided herein are described in US20200160936A1, US20220254446A1, US20220230704A1, which are each incorporated herein by reference in their entirety.
- the target prokary otic organism may comprise one or multiple of the following an E. coli strain, Akkermansia muciniphila, Alistipes fmegoldii and Bacteroides ovatus. Bifidobacterium longum.. Ruminococcus bromii, Eubacterium rectale, Bacteroides uniformis D7, Dorea longicatena and Barnesiella intestinihominis, Eubacterium sp.
- the target prokaryotic organisms may be selected from specific pathogenic bacteria, beneficial bacteria, probiotics, or prokaryotic organisms with low relative abundance in a microbiome sample as determined by standard metagenomic sequencing.
- the target prokaryotic organism, methylation motif and restriction enzyme may be selected from the following groups: (i) the target prokaryotic organism(s) comprise a methylated G6mATC motif, and the one or more methylation sensitive restriction enzyme comprises methylation sensitive DpnII and/or a methylation sensitive isoschizomer thereof; (ii) the target prokaryotic organism(s) comprise a methylated RA6mATTY motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive XapI , methylation sensitive Apol, and/or a methylation sensitive isoschizomer thereof, or any combination thereof; (iii) the target prokaryotic organism(s) comprise a methylated CTS6mAG motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive BspCNL methylation sensitive BseMII, a methylation sensitive isoschizomer thereof, or any combination thereof; (iv)
- the target prokaryotic organism(s) comprising a methylated G6mATC motif comprise Escherichia coll, Klebsiella pneumoniae, one or more species of Gammaproteobacteria class or any combination thereof; (ii) the target prokaryotic organism(s) comprising a methylated RA6mATTY motif comprise one or more species of Akkermansia muciniphila, Campylobacter spp., Acinetobacter spp., Spirochaeta spp., Treponema spp.
- the target prokaryotic organism(s) comprising a methylated CTS6mAG motif comprise one or more species ofAlistipes finegoldii and/or Bacteroides ovatus;
- the target prokaryotic organism(s) comprising a methylated RG6mATCY motif comprise one or more species of Bifidobacterium longum, Ruminococcus bromii, Eubacterium rectale and Bacteroides uniformis;
- the target prokaryotic organism(s) comprising a methylated G6mATGC motif comprise one or more species of Dorea longicatena, Barnesiella intestinihominis and Eubacterium sp.
- CAG:248 and/or the target prokaryotic organism(s) comprising a methylated G6mATGC motif and methylated CTNN6mAG motif comprise one or more species of Barnesiella intestinihominis.
- the method may further comprise a DNA shearing step, and wherein shearing the gDNA optionally comprises g-TUBEs and the gDNA fragments from target and background organisms have a length of about 10 kb.
- the method further comprises an end repair and dA-tailing step prior to the linking of said universal PCR adapter.
- the method may comprise sequencing the nucleic acid of the amplified target prokaryotic organism(s).
- the sequencing is performed using a single-molecule long read real time (SMRT) technology, nanopore sequencing technology or other short read sequencing technologies including but not limited to Illumina sequencing platforms.
- SMRT single-molecule long read real time
- the microbiome sample is obtained from soil, air, water, sediment, oil, and combinations thereof.
- the microbiome sample may be obtained from water selected from marine water, fresh water, and rainwater.
- the microbiome sample may be obtained from a subject selected from a protozoa, an animal, or a plant.
- the subject may be a mammal.
- the subject is a human.
- the subject may be an infant.
- the sample obtained from a mammal e.g., human
- compositions 111.
- a universal adaptor for use in any of the methods described herein.
- the universal adaptors can comprise about 50-100 bp double stranded short DNA fragments with an overhang at the 3’ end of the top strand, which facilitates ligation to all DNA molecules through T/A ligation.
- the universal adaptors will not include any recognition sequences in the adapter sequence. Suitable universal adaptors may be “Y- shaped” or “bell-shaped”, or may come in other structures.
- the universal PCR adaptor comprises an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hy bridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, and wherein the hybridized region does not comprise a nonmethylated motif targeted by the one or more restriction enzyme.
- the universal adaptor can further comprise a *T at the 3‘ end of the upper strand oligo, wherein *T comprises a phosphorothioated thiamine base having a phosphorothioate nucleotide modification, wherein the phosphorothioate modification renders an inter-nucleotide linkage formed by the universal adaptor resistant to nuclease degradation.
- the universal adaptor can be phosphory lated at the 5’ end, particularly on the lower strand oligonucleotide. Other modifications to prevent nuclease degradation are known in the art and are contemplated herein.
- the non-hybridized region of the universal PCR adaptor can comprise or consist of SEQ ID NO: 1 (for the upper oligo) and/or SEQ ID NO: 2 (for the lower oligo).
- Each of the upper and lower oligo may comprise of an additional 10 to 25 nucleotides that hybridize and that, when hybridized, do not comprise a non-methylated motif targeted by the one or more restriction enzyme.
- the upper and lower oligos comprise the same number of additional nucleotides that hybridize.
- the upper strand may comprise a nucleotide sequence of any one of SEQ ID NOs: 3-18 and the lower strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 19-34.
- the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 36.
- the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 38.
- SEQ ID NOs: 1 -38 are provided in the Table 1 below, where the phosphorothioate linkage is indicated by an asterisk (*) and the 5’phosphorylation modification is indicated by a /5Phos/.
- each upper-strand oligo comprises a primer binding site complementary to a forward barcoded primer
- each lower-strand oligo comprises a primer binding site complementary to reverse barcoded primer
- a set of barcoded PCR primers is provided, where the barcoded PCR primers target any portion of the upper or lower oligo of the universal adaptor.
- the barcoded PCR primer may target a nucleic acid sequence comprising at least a portion of a nucleic acid sequence comprising or consisting of SEQ ID NOs 1 or 2, which are shown again in Table 2 for ease of reference.
- the term "‘target” means that the primer hybridizes to a nucleic acid comprising the nucleic acid target sequence to enable PCR amplification.
- the primer hybridizes to the 5’ strand (e g., upper oligo of universal adaptor) comprising the target sequence or to the 3 ' strand complementary to the 5 ’ strand. In other aspects, the primer hybridizes to the 3‘ strand (e.g., lower oligo of universal adaptor) comprising the target sequence or to the 5’ strand complementary to the 3’ strand.
- the 5’ strand e g., upper oligo of universal adaptor
- the primer hybridizes to the 3‘ strand (e.g., lower oligo of universal adaptor) comprising the target sequence or to the 5’ strand complementary to the 3’ strand.
- kits comprising (a) a set of universal PCR adapters and (b) one or methylation sensitive restriction enzymes.
- the kit further comprises (c) a set of barcoded PCR primers complementary to at least a portion of the universal PCR adapters.
- the one or more methylation sensitive restriction enzymes comprise restriction enzymes blocked by the methylated motifs in one or more target prokaryotic organisms, and wherein comprise but not limited to DpnII, XapI, Apol, BspCNI, BseMII, Mill, BstX2I, BstYI, Mill, XhoII, SfaNI, AflII, Acul, Bpml, BpuEI, PstI or a combination of any thereof.
- each of the universal PCR adaptors may comprise an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, wherein the hybridized region and wherein the hybridized region does not comprise a non-methylated motif targeted by the one or more restriction enzyme.
- each universal PCR adapter in the kit further comprising a *T at the 3’ end of the upper strand oligo, wherein *T comprises a phosphorothioated thiamine base having a phosphorothioate nucleotide modification, wherein the phosphorothioate modification renders an inter-nucleotide linkage formed by the Y shaped adaptor resistant to nuclease degradation.
- the universal adaptor can be phosphorylated at the 5’ end, particularly on the lower strand oligonucleotide. Other modifications to prevent nuclease degradation are known in the art and are contemplated herein.
- the upper strand oligo can have a nucleotide sequence comprising any one of SEQ ID NOs: 3-18 and/or the lower strand oligo can have a nucleotide sequence comprising any one of SEQ ID NOs: 19-34.
- the upper strand oligo can have a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo can have a nucleotide sequence comprising SEQ ID NO: 36.
- the upper strand oligo can have a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo can have a nucleotide sequence comprising SEQ ID NO: 38.
- the set of barcoded PCR primers may target any nucleic acid of the upper or lower oligo of the universal adaptor.
- the barcoded PCR primer may target a nucleic acid sequence comprising at least a portion of a nucleic acid sequence comprising or consisting of: SEQ ID NO: 1 or 2.
- the term “target” means that the primer hybridizes to a nucleic acid comprising the nucleic acid target sequence to enable PCR amplification.
- the forward primer hybridizes to the sense strand of the target sequence comprising upper oligo of universal adaptor (i.e., any one of SEQ ID NOs: 3- 18).
- the reverse primer hybridizes to the anti-sense strand of the target sequence comprising lower oligo of universal adaptor (i.e., any one of SEQ ID NOs: 19-34).
- Kits may optionally provide additional components such as buffers and interpretive information.
- the kit includes a container and a label or package insert(s) on or associated with the container.
- the invention provides articles of manufacture comprising contents of the kits described above.
- Metagenomics have enabled the comprehensive and culture-independent study of microbiomes. However, for many applications, only certain taxa are of interest in a microbiome sample, for example, pathogenic bacteria, beneficial microbes, or taxa of low relative abundance. It is highly desirable that a method can effectively enrich bacterial taxa of interest directly from the microbiome. To address this critical need. mEnrich-seq, a method that can enrich taxa of interest from metagenomic DNA before sequencing, was developed. The core idea is to exploit the natural self vs.
- non-self genome differentiation provided by bacterial DNA methylation and rationally choose methylation-sensitive restriction enzymes (REs), individually or in combination, to deplete host DNA and most background microbial DNA while enriching bacterial taxa of interest.
- This core idea is integrated with library preparation procedures in a way that only non-digested DNA libraries are sequenced. In-depth evaluations of mEnrich-seq were performed using both synthetic and real microbiome samples. The use of mEnrich-seq is demonstrated in several applications to enrich (up to 117-fold) genomic DNA of pathogenic or beneficial bacteria from urine and fecal samples, including several species that are hard to culture or of low abundance.
- mEnrich-seq provides microbiome researchers with a versatile and cost-effective approach for selective sequencing of diverse taxa of interest directly from the microbiome.
- Host genomic DNA can be eliminated using chemicals such as saponin to selectively lyse mammalian cells, which enables the efficient sequencing of pathogens in biospecimens that are dominated by host gDNA (for example, upper respiratory tract infection).
- Host genetic material can also be filtered using DNA binding proteins that specifically recognize 5-methylcytosine (5mC), abundant in mammalian genomes.
- 5mC 5-methylcytosine
- Flow cytometry has also been used to segregate complex microbiomes into small subsets, which can be sequenced and computationally analyzed separately.
- Adaptive sampling in the Nanopore platform can examine the first few hundred base pairs, during real-time sequencing, against a collection of reference sequences to reject the DNA molecules that meet certain criteria, for example, host DNA or a highly abundant bacterium for which enough data have been already collected.
- This strategy by adaptive sampling can reduce further consumption of sequencing yield by highly abundant species; however, its enrichment efficiency for specific taxa of interest is relatively modest, and insufficient to differentiate between bacterial taxa with highly similar genomes in complex microbiome samples.
- a method is described herein that takes advantage of bacterial DNA methylation in microbiomes, namely bacterial epigenomes.
- bacterial DNA methylation in the bacterial kingdom, there are three main forms of DNA methylation: N6-methyladenine (6mA), N4-methylcytosine (4mC) and 5mC, which are catalyzed by methyltransferases (MTases) that apply methyl groups in a highly sequence-specific manner, that is. some sequence motifs of a genome are nearly 100% methylated, while most of the genome is not methylated.
- the bacterial methylome has three fundamental properties.
- DNA methylation is present in nearly all bacteria (more than 95%).
- all genetic contents chromosomes, plasmids
- the methylation motifs are highly variable between different species and different strains of the same species. Based on these three properties, bacterial DNA methylation naturally differentiates self from nonself DNA, which serves as the foundation of restrictionmodification systems, and has been exploited as natural epigenetic barcodes to group highly similar species and strains in metagenomic analyses, namely methylation binning.
- E. coli K-12 MG1655 was obtained from American Type Culture Collection (ATCC, cat# 47076, VA. USA) and cultured with Miller Luria Broth (LB) medium, and on LB agar plates (Thermo Fisher Scientific, cat# BP1426-500, and cat# BP1425-500, respectively, MA, USA) at 37°C based on ATCC recommended protocols.
- N. gonorrhoeas MS 11 was obtained from ATCC (cat# BAA- 1833) and cultured based on protocols in Dillard, 2011 (Dillard, J. P. Genetic manipulation of Neisseria gonorrhoeae. Curr Protoc Microbiol 23, 4A-2 (2011)). Briefly, cells were grown in Gonococcal Base medium liquid (GCBL) with Kellogg’s supplements I&II, and on Chocolate II Agar plates (BD, cat# 22126, MD, USA) at 37°C in 5%CO 2 for 24 h.
- GCBL Gonococcal Base medium liquid
- Kellogg Kellogg
- UTI Urinary tract infection
- DNA from the human LCL cell line GM24149 E. coli MG1655, N. gonorrhoeae MS11, and urine samples was isolated with the DNeasy Blood and Tissue Kit (Qiagen, cat# 69506, MD, USA) based upon the manufacturer’s instructions, including the optional RNase A treatment provided with the kit.
- DNeasy Blood and Tissue Kit Qiagen, cat# 69506, MD, USA
- 2xl0 6 cells were collected from cell culture by centrifugation for 5 min at 300 x g. The cell pellets were resuspended in 180 pl PBS and 20 pl Proteinase K. After adding 200 pl Buffer AL, the mixture was incubated at 56°C for 30 min. The rest of extraction steps were performed according to manufacturer’s instructions.
- each 2 ml aliquot of urine was first centrifuged at 5000 x g and 4°C for 30 min, then the pellets were collected in a new tube.
- the urine pellets were resuspended with 180 pl of the bacterial enzymatic lysis buffer (20 mM Tris pH 8.0, 2mM EDTA, 1% SDS and 20 mg/ml lysozy me) and incubated at 37°C for 30 min.
- 20 pl of Proteinase K and 200 pl Buffer AL was added to the mixture, and it was incubated at 56°C for 30 min.
- the rest of the extraction steps were followed per manufacturer instructions with one modification, i.e., the DNA was eluted with 50 pl AE in the final step to concentrate the DNA.
- Fecal DNA was extracted with a QIAamp PowerFecal Pro DNA Kit (Qiagen, 51804) according to the manufacturer instructions. Briefly, 100 mg of fecal sample, 750 pl of PowerBead Solution and 60 pl of Solution Cl were added to the Bead Tube (provided). After incubation at 65°C for 10 min. the sample was homogenized with the TissueLyser II machine (Qiagen) for 10 min at 25 Hz speed (5 min each side). The homogenized samples were centrifuged to collect the supernatant at 15,000 xg for 1 min. The subsequent column-based extraction steps were performed as described in the QIAamp PowerFecal Pro DNA Kit protocol.
- the isolates were cultured from the urine samples with confirmed E. coli cases. Briefly, 10 pl of urine was spread by disposable inoculating loops (Thermo Fisher Scientific, cat# 22170201) quantitatively onto BD BBL MacConkey agar plates (Becton, Dickinson, and Company (BD), NJ, USA) and incubated for 24 h at 37°C aerobically. For each urine sample, a representative colony of E. coli was collected for DNA purification.
- Step 1 DNA shearing. Shearing was conducted to establish a uniform size of DNA inserts for better ligation efficiency, to maintain highly consistent molanty estimations for downstream loading into flow cells, and to reduce pore clogging.
- gDNA was sheared to around lOkb with g-TUBEs (Covaris, cat.520079, MA, USA) using an Eppendorf 5424 centrifuge at 5000 rpm for 3 min at room temperature. The g-TUBE was then inverted and centrifuged for another 3 min to collect the sheared DNA. The DNA was concentrated using 0.5x volume Ampure XP beads and eluted with 48 ul nuclease-free (NF) water.
- NF nuclease-free
- Step 2 End Repair and Barcode Adapter Ligation.
- the fragmented DNA was applied to ’End-prep” and “Ligation of Barcode Adapter” steps according to the manufacturer’s instructions.
- the “End-prep” was performed with 48 pl of sheared DNA in the reaction mixture using NEBNext Ultra II End Repair/dA-Tailing Module and NEBNext FFPE DNA Repair Mix (New England Biolabs (NEB), cat# E7546, M6630, respectively, MA, USA) with incubation at 20°C for 5 min and 65°C for 5 min. After purification with lx volume of AMPure XP beads, the end-repaired DNA was eluted in 30 pl nuclease-free water.
- the end-repaired DNA was ligated with 20 pl Barcode Adapter (ONT, cat# EXP-PBC096, blue cap) and Blunt/TA Ligase Master Mix (NEB, cat# M0367) for lOmin at room temperature, which attached the universal PCR handle to all of the DNA molecules.
- the Barcode-Adapter- ligated-sample was then purified with lx volume of Ampure XP beads and eluted in 26 pl NF water.
- 20 ng of the PCR Adapter ligated DNA was split into another tube as sample “Input”, while the rest of the sample was used for the restriction enzyme digestion step of mEnrich-seq.
- Step 3 Restriction enzyme (RE) digestion.
- RE digestion for mEnrich-seq was conducted following the manufacturer’s instructions.
- 10 U of enzyme (NEB, cat # R0543) was used in a 50 pl-reaction containing 5 pl of NEBuffer r3.1 (10X, NEB #B7203), 26 pl of adapter-ligated DNA and NF water to 50 pl. The reaction system was incubated at 37 °C for 30 min.
- BspCNI NEB, cat# R0624S
- 10 U of the enzyme was used in the 50 pl -reach on containing 5 pl of rCutSmart Buffer (10X, NEB # B6004S) for 30 min- incubation at 37 °C.
- the DNA was firstly digested with 20 U SfaNI in NEBuffer r3.1 for 30 min at 37 °C, then the digestion system was purified with 1.8X AMpure beads. The eluted DNA (20 pl) was further digested for 30 min at 37 °C with a 5 pl-mixture of Aflll, Acul, Bpml, Bpuel and Pstl-HF (1 pl of each).
- Step 4 Gel purification.
- the RE-digested products were loaded into a 1.5% agarose gel (Thermo Scientific, cat# R0492, MA, USA) pre-stained with SYBR Safe DNA gel stain (Invitrogen, cat# S33102, MA USA) for gel-based size selection.
- DNA ladder (NEB, cat#N3238S) was used as a comparison for size estimation and each gel was visualized under UV light on a ChemiDoc XRS+ System (Bio-Rad, CA, USA).
- the DNA band at 5 to lOkb size (as the DNA ladder indicates) was collected for extraction with a NucleoSpin Gel and PCR Clean-up purification kit following the manufacturer's instructions (Macherey-Nagel, cat# 740609.250, Germany), and was eluted with 15 pl NF water.
- Step 5 PCR amplification.
- the gel-purified DNA was amplified with the barcoded primers from the barcoding expansion pack mentioned above (EXP-PBC096, white caps, ONT).
- EXP-PBC096, white caps, ONT for PrimeSTAR GXL Premix (2x, Takara Bio, cat# R051A, Japan), a 50 pl -reaction system was prepared as following: 25 pl of Premix, 1 pl of barcoded primer (lOuM, white tube), 15 pl purified DNA and 9 pl water.
- the PCR reactions were pre-heated at 98°C for 1 min followed by 15-18 cycles of 98°C for 15 s, annealing 62°C for 15 s and extension at 68°C for 8 min, with a final extension at 68°C for 10 min on the ABI Veriti thermal cycler.
- the PCR products were purified with 0.5x volume of Ampure XP beads. Before moving to library prep, the purified PCR products were checked with an agarose gel for a distinct band at 5-10 kb in size (FIGS. 9A-9B, 10A-10B, 11A-11B, and 12A-12B).
- Step 6 Library prep and sequencing.
- the purified barcoded PCR products were pooled together for "DNA repair’ and “end-prep” step. Briefly. 47 pl of pooled PCR products were incubated at 20°C for 15min and 65°C for 15min, in a 60-reaction system containing 3.5 pl of NEBNext FFPE DNA Repair Buffer, 2 pl of NEBNext FFPE DNA Repair Mix, Ultra II End-prep reaction buffer 3.5 pl and Ultra II End-prep enzyme mix 3 pl.
- the purified DNA was ligated to Adapter Mix (ONT, SQK-LSK109) in the presence of Ligation Buffer (ONT, SQK- LSK109) and NEBNext Quick T4 Ligase (NEB, cat# E6056) for 30 min at room temperature. After beads purification, the adapter ligated DNA is ready for sequencing on ONT flow cells.
- the barcoded PCR products from ‘Step 5’ will firstly be purified and adjusted to shorter insert size (e.g., by sonication). Then, the fragmented DNA molecules are ready for users to prepare the library 7 with suitable kits compatible with various Illumina platforms. Compatible barcoding kits are required for a good sequencing result.
- the Barcode adapter sequence from the mEnrich-seq reads can be trimmed using computational methods for further analysis.
- Barcode adapter As the Barcode adapter ligation step is prior to the RE digestion during the library preparation, the DNA from targeted taxa will not be amplified by PCR if the barcode adapter is cut during the RE digestion step. Therefore, it’s critical to avoid a chosen RE cutting the sequence of the Barcode adapter or the junction between the adapter and metagenomic gDNA. It is suggested to check the sequence of the Barcode adapter (particularly the sequences that are reverse complementary to each other) for any possible RE motif site. If there’s a recognition sequence for the chosen RE, the bases can be substituted to avoid the cut. The sequences are as follows: top sequence: SEQ ID NO: 35 and bottom sequence: SEQ ID NO: 36.
- the top and bottom sequence are added to the annealing buffer (10 mM Tris, pH 7.5, 50 mM NaCl and 1 mM EDTA) to a concentration of 10 pM and annealed with the program: 95 °C for 2 min, -0. 1 C/sec for each of 800 s ( ⁇ 13 min) and hold at 4°C.
- the barcoded adapter is now ready to use and compatible with the original protocol.
- RE can be found through the Enzyme Finder (enzymefmder.neb.com) based on the chosen target motifs. For the chosen RE, it’s important to check its methylation sensitivity result provided by the manufacturer in REBASE (rebase.neb.com/rebase/rebase), which is commonly tested with modified oligos, and even better to re-confirm the sensitivity upon purchase.
- the PCR programs were as follows: pre-incubation at 95°C for 3 min, amplification for 14 cycles at 95°C for 15 s, 62°C for 15 s, and extension at 65°C for 10 min, with a final extension at 65°C for 10 min.
- the PCR products were checked with an agarose gel for successful amplification (bands at ⁇ 5 to 10 kb, FIGS. 9A-9B, 10A-10B, 11A-11B. and 12A-12B).
- the purified barcoded PCR products were used for “repair and end-prep” step in the ONT protocol.
- the library was quantified by Qubit and ready for sequencing.
- Two pairs of primers specific to tet(A) sequence were designed using Primer- BLAST as follows: te ⁇ -Fl: 5’- GAAACCCAACAGACCCCTGA-3’ (SEQ ID NO: 39), and tet(A)-R 5’- CCACCCGTTCCACGTTGTTA-3’(SEQ ID NO: 40); tet(A)-F2: 5’- TGAAACCCAACAGACCCCTG-3’(SEQ ID NO: 41), and tet(A)-R2 5’- CCCTGACGTTCCTCATCCAC-3’ (SEQ ID NO: 42) (Integrated DNA Technologies, IA, USA; 25 nmol, standard desalting).
- the tet(A) was amplified directly from the gDNA of UTI- 1 using the Q5 High-Fidelity 2X PCR Master Mix (NEB, cat #M0492S) with the following conditions: 50 ng of gDNA per reaction; 1 cycle at 95°C for 30s; 33 cycles of denature 95°C for 15s, annealing at 60 °C for 15s, extension at 72 °C for Imin; a final extension cycle at 72 °C for 5 min.
- PCR products were checked for amplification specificity on an agarose gel. Main bands of the PCR products were gel purified with NucleoSpin Gel and PCR Clean-up purification kit.
- the purified PCR products were sequenced by the Sanger method using tet(A)- F1 and tet(A)-V2 as forward primers._The two PCR products are shown in Table 3 below. Nucleotides denoted with an “n” in SEQ ID NO: 43 and 44 were undetermined by the Sanger method and may be any base (a,t,c, or g).
- the fold enrichment of E. coli was calculated based on E. coh reads from mEnrich-seq and Input samples that were mapped to the mock reference using minimap2 (v2.24-rl 122) with -x map-ont.
- the fold enrichment is defined as: target strain abundance (%) in enrichment sequencing versus target strain abundance (%) in the input sample.
- target strain abundance (%) in enrichment sequencing versus target strain abundance (%) in the input sample For genome completeness, the coverage of the E. coli was calculated using read depths across lOObp bins across the E. coli genome. For Circos plots. Nanopore reads from paired Input and mEnrich-seq samples were normalized to the same yield by random subsampling using seqtk (vl.2-r94) with 2-pass mode.
- rRNA ribosomal RNA
- muciniphila using minimap2 (v2.24-rl 122) with -x map-ont.
- sliding windows with intervals of 2.5kb and step sizes of 1.25kb were defined across the A. muciniphila isolate reference using bedtools (v2.29.2) with -w 2500 -s 1250.
- Reads mapped to defined windows were counted using bedtools multicov.
- read depths were capped at the 95% depth quantile of the whole genome for Circos plots drawing using circos (vO.69-6).
- RAATTY motif analysis For the species used for motif analysis, all genome references except A. muciniphila were retrieved fromNCBI, including C. aerofaciens (RefSeq genome sequence GCF_002736145.1), E. lenta (GCF_021378605.1), B. bifidum (GCF 000273525.1), B. longum (GCF_000196555.1), B. breve (GCF_001281425.1). The matched A. muciniphila reference was obtained from the isolated strain as previously described. Nanopore reads from the mEnrich-seq samples were mapped to the reference of each species using minimap2 (v2.24-rl 122) with -x map-ont.
- the raw SMRT reads pre-processing was performed with SMRTlink v8.0 (Pacific Biosciences).
- Raw multiplexed files were first de-multiplexed with lima.
- Circular Consensus reads (CCS reads) were generated using CCS software (v4.0.0) with --min-rq 0.99 based on subreads.
- CCS reads were assembled using metaFlye (v2.9) with -pacbio-hifi. The assembly was then binned and refined with the metaWRAP Binning and Bin refmement module 10 , which could combine the results from three metagenomic binning software - MaxBin2, metaBAT2, and CONCOCT. The completeness and the contamination level was checked using CheckM (v.1.1.3).
- the metagenomic profile computing and graphics were performed by R (v.4.2. 1).
- the genome assemblies from bacteria isolates from a previous study5 along with de novo assembled metagenomic contigs were used as the reference database for the following nanopore reads assignments and motif analyses.
- Nanopore reads w ere first mapped to the reference database described above using minimap2 (v2.24-rl 122) with -x map-ont. To reduce the possibility of low-quality mis-mapped reads from other species, non-primary and supplementary alignments were not included for reads assignment using samtools (vl.15.1) with view -F 2308. Other reads were then classified using Kraken2 (v2.0.8-beta) with the k2_pluspf_20200919 database.
- REBASE is the database constantly updated with bacterial methylomes mapped to date. For each bacterial methylome, a list of methylation motifs was reported (rebase.neb.com/rebase/rebase.html). All 4-mer to 6-mer methylation motifs were extracted and their corresponding restriction endonucleases (REs) from 4,644 bacterial methylomes (as of 7/18/2022). Among 4,644 PacBio organisms, the 1,358 most complete genomes were selected as representatives when multiple strains from the same species are present. The protein sequences from the 1,358 species level genomes were extracted to reconstruct a phylogenetic tree using PhyloPhlAn (v3.0). Visualization and annotation of the phylogenetic tree was performed by ITOL (itol.embl.de/).
- REMoDE libraries were prepared as previously described by Enam et al., (Restriction Endonuclease-Based Modification-Dependent Enrichment (REMoDE) of DNA for Metagenomic Sequencing. Appl Environ Microbiol 89, e01670-22 (2023), which is incorporated herein by reference in its entirety ). Briefly, 100 ng of DNA was used for the reaction mixture of 40 pL containing 10X Tango buffer and 1 pL of XapI (Fisher Scientific, ER1381). The mixture was incubated at 37°C for 30 mm.
- T5 exonuclease (NEB, #M0663) was added to each reaction mixture and incubated for 5 min at 37°C, after which the reaction w as immediately quenched with 8 pL 66 mM EDTA.
- REMoDE-repair For the library preparation of REMoDE with DNA repair (“REMoDE-repair’’), the DNA was first treated with the NEBNext FFPE DNA Repair Mix at 20°C for 15 min, purified with IX AMPure beads, before following the protocol described above (starting at XapI digestion).
- NGS libraries were prepared using Nextera XT (Illumina, cat # FC-131-1024) library preparation kit. The amplified libraries were resolved on an agarose gel and quantified with Qubit HS dsDNA kit. The libraries were pooled and sequenced on an Illumina MiSeq sequencer using a MiSeq reagent kit v2 (Illumina, cat# MS-102-2002; 300 cycles, paired-end). [0142] Data analysis.
- the paired-end short-read data of Input, REMoDE and mEnrich were each mapped to the reference using bwa with default parameters, and then reads mapped to each species were counted.
- 1 ng of DNA was diluted to 3 pL and incubated with 1 pL of Fragmentation Mix at 30° C for 1 min and then at 80° C for 1 min.
- the tagmented DNA was added to PCR mix containing 20 pL nuclease-free water, 25 pL LongAmp Taq 2X master mix, 1 pL Rapid Barcode Primers (RLB01-12A, at 10 pM), followed by 18 cycles of amplification (18 cycles of 95 °C for 15 s; 56 °C for 15 s; 65 °C for 6 min; 1 cycle of extension at 65 °C for 6 min).
- PCR products were purified with 0.6X AMPure Beads and eluted with 10 pl elution buffer (10 mM Tris-HCl pH 8.0 with 50 mM NaCl). 50-100 ftnoles of library pool was ready for sequencing after ligation with 1 pL Rapid Adapter for 5 min at room temperature.
- mEnrich-Dl the original mEnrich-seq
- mEnrich-D2 the standard ONT library without enrichment
- Input standard ONT library without enrichment
- Example 3 Methylation-guided enrichment for bacterial taxa of interest
- the core idea of enriching the gDNA of bacterial taxa of interest from the microbiome is to distinguish gDNA from different taxa based on their distinct characteristics.
- a methylation-guided enrichment sequencing of bacterial taxa of interest from microbiomes is shown and named mEnrich-seq (in which ‘m’ stands for methylation and seq for sequencing, FIG. 1A).
- m stands for methylation and seq for sequencing
- the gDNA of the bacterial taxa of interest (including its chromosome and mobile genetic elements, for example, plasmids) is left intact, enriched over the background, which can be sequenced by various sequencing platforms (FIG. 1A). More specifically, g-TUBEs are first used to shear the DNA size to around 10 kb, which facilitate the second step where DNA is ligated to barcode adapters. In the third step, cognate restriction enzymes (RE) corresponding to the methylation motif in the taxon of interest are used to cut the background DNA libraries into smaller fragments while preserving the targeted DNA libraries as the long fragments.
- RE restriction enzymes
- the choice of fragmentation size in the shearing step increases the chance that most gDNA molecules have at least one nonmethylated RE site that can be cut: 5-10 kb fragments expected to have 20-40 sites of a 4-mer motif, 5-10 sites of a 5-mer motif and 1-2.5 sites of a 6-mer motif.
- gel-based size selection is performed, which largely removes the digested shorter DNA fragments.
- step 2 barcoded primers are used to anchor the PCR adapter to perform amplification of the long target DNA fragments that have been protected from RE digestion (i.e., those that are methylated at restriction sites).
- step 3 the DNA-to-adapter ligation step
- step 3 only nondigested DNA libraries can be sequenced. This design minimizes allocation of sequencing throughput to background DNA fragments that are still long after RE digestion due to sparse distribution of the RE recognition motif(s) in certain genomes or genomic regions.
- mEnrich-seq is versatile along three dimensions.
- RE(s) can be tailored to enrich bacterial taxa at different taxonomic levels, which can be helpful depending on specific research goals (FIG. IB).
- the methylation-sensitive RE DpnII (which cuts doublestranded DNA (dsDNA) at nonmethylated GATC sites, but is sensitive to methylated G6mATC) was chosen for mEnrich-seq to selectively enrich E. coli gDNA while digesting N. gonorrhoeae (only 35 of the 2,434 GATC sites are methylated owing to the overlapping with GGTG6mA methylation motif) and human gDNA with hardly detectable methylated GATC sites.
- the library was prepared for Nanopore long-read sequencing and generated 200 Mb of data (mean read length of roughly 5 kb Table 5).
- the mock samples were also sequenced without mEnrich-seq protocol, labeled as ‘input’ (standard metagenomics).
- Table 5 summarizes the Nanopore sequencing dataset for the Mock-1 and Mock-2. Data were produced by Flongle flow cell. The read-length difference is due to the gel-based size selection step in the gel-based size selection step in mEnrich-seq.
- mEnrich-seq greatly increased the proportion (%) of reads mapped to the E. coli genome from 0.91 to 70.75% and from 1 .20 to 71 .41% in the two mock replicates, respectively (FIG. 2A): roughly 70-fold enrichment (FIG. 2B).
- mEnrich-seq reads mapping showed that the E. coli genome can be covered without systematic bias for both mock samples (more than 99.96% of genome covered; FIG. 2C). Also, 3.98 and 4.12% of reads from mEnrich-seq are mapped to N. gonorrhoeae.
- GATC sites are located within genomic regions with simple sequence repeats (Example 2, above) that tend to form secondary structure, making it less accessible to DpnII digestion.
- This characterization illustrates that mEnrich-seq not only enriches bacterial taxa of interest, but also carries some by-product reads that are either due to methylation at (partially) overlapping motif sites in background taxa or depleted restriction motifs in a subset of gDNA molecules.
- the highlight of mEnrich-seq is the enrichment efficiency compared to standard metagenomic sequencing despite by-product reads.
- mEnrich- seq was further tested on three urine samples (UTI-1-3) from patients with urinary tract infection (UTI) (E. coli positive urine culture, defined as more than or equal to 100.000 colony forming units per ml; see Example 2, above).
- UTI urinary tract infection
- the three UTI samples were sequenced with mEnrich-seq (with DpnII as RE) coupled with Nanopore sequencing.
- mEnrich-seq 1.16— 1.58 GB of sequencing data was generated with an average of 233,000 reads per sample (mean read length roughly 6 kb, Tables 6A-6B, below).
- coli reads was further confirmed by mapping each read manually to the National Center for Biotechnology (NCBI) Nucleotide database, and further validated by PCR (FIG. 7B) and Sanger sequencing (as described in Example 2, above). This observation is consistent with the increasing recognition that a culture-independent approach may provide additional information that can complement standard urine culture in monitoring pathogen genomes in UTI and other infectious diseases.
- NCBI National Center for Biotechnology
- FIG. 7B Sanger sequencing
- Tables 6A-6B show the summary of the Nanopore sequencing dataset for urine samples.
- Table 6A illustrates the input, isolates and mEnrich-seq of sample UTI-1 to -3 as first sequenced with a MinlON flow cell.
- Table 6B illustrates the three urine samples (UTI- 1 to -3) together with the additional three UTI urine samples (UT1-4 to -6) as sequenced on a Flongle flow cell for validation purposes.
- the efficient enrichment of E. coll genomes by mEnrich-seq can facilitate the culture-free study of E. coli genomes from urine microbiome with better sensitivity (examining E. coli with lower relative abundance) than standard metagenomics.
- mEnrich-seq can facilitate a more cost-effective examination of urine samples in the study of E. coli genomes.
- 6mA at GATC sites are also conserved in most Gammaproteobacteria including many enteric pathogens, which makes mEnrich-seq with DpnII applicable to enrich additional enteric pathogens and commensal bacteria as are described in the following Examples.
- RA6mATTY mediated by a 6mA MTase, AmuORF1905P, is one of the five methylated motifs in A. muciniphila (ATCC BAA-835) according to REBASE (The Restriction Enzyme Database) (FIG. 8A).
- mEnrich-seq was applied with XapI to enrich the muciniphila genome from three infant fecal samples (GUT-1-3) (Table 7. below), and observed a 20.0-fold increase of reads mapped to A. muciniphila in GUT-1 (from 0.48 to 9.59%). a 27.2-fold increase in GUT-2 (from 0.52 to 14.16%) and an 18.3-fold increase in GUT-3 (from 1.52 to 27.88%) (FIG. 3A-3B).
- the RAATTY frequency is comparable between the mEnrich-seq reads and RAATTY frequency of the reference genome (FIG. 3D), which is consistent with the enrichment due to methylated RA6mATTY motif on its genome.
- the frequency of RAATTY in the genomes of B. bifidum, C. aerofaciens and E. lenta is lower, which further decreased among mEnrich-seq reads that mapped to these three species (FIG. 3D).
- These three genomes with sparse frequency of RAATTY sites were enriched as by-products because of the more efficient depletion of most of the other background genomes with dense yet nonmethylated RAATTY sites.
- mEnrich-seq with XapI as the methylation-sensitive RE can efficiently enrich 4.
- muciniphila highly conserved RA6mATTY methylation across all the isolated strains examined
- mEnrich-seq can be a useful tool to study the genomes of A. muciniphila in a culture-independent, sensitive and cost-effective way, which may facilitate larger-scale association studies with different human diseases.
- mEnrich-seq is demonstrated when the methylation motifs of target bacteria of interest are known a priori.
- mEnrich- seq is used to enrich gDNA from low-abundance bacteria from a complex microbiome sample based on methylation motifs discovered de novo.
- GUT-5 an adult fecal microbiome sample (GUT-5) was built that has been recently well characterized with comprehensive cultured isolates.
- a pilot standard metagenomic sequencing of this adult fecal sample was analyzed (Nanopore sequencing, 7.08 Gb; further details in Example 2 above).
- Table 8 provides the summary of this Nanopore sequencing data for low- abundance bacterial taxa enrichment with de novo discovered methylation motifs.
- the top portion summanzes data from an adult fecal sample sequenced with the MiNION flow cell for the initial standard metagenomic profiling.
- the bottom portion shows data for the low abundance bacterial taxa enrichment by mEnrich-seq, sequenced with a Flongle flow cell.
- ovatus D3 was also enriched (4.3-fold, from 2.97 to 12.94%), which has a CTCAG frequency similar to A. finegoldii H10 and B. ovatus C6, reflecting that B. ovatus D3 has methylation at CTCAG sites that resist BspCNI digestion (FIG. 4D).
- Bifidobacterium longum F12 and B. longum DIO have the same methylation motif RG6mATCY.
- mEnrich-seq with Mfll achieved an 8.7-fold enrichment for B. longum F 12 (from 1.17 to 10.17%) and a sevenfold enrichment for B. longum D10 (from 0.32 to 2.24%) (FIG. 4E).
- enrichment was observed for Ruminococcus bromii (4.5-fold, from 2.58 to 11.69%), Eubacterium rectale G9 (3.8-fold, from 2.15 to 8.20%) and Bacteroides uniformis D7 (2.6-fold, from 2.69 to 6.89%; FIG. 4E).
- Further examination of RGATCY motif frequency supported that R. bromii, E. rectale G9 and B uniformis D7 were all enriched owing to methylation at RG6mATCY sites on their genomes (FIG. 4F).
- B. intestinihominis owing to the methylated G6mATGC motif on their genomes (FIG. 4H).
- SfaNI targeting GATGC
- this Example shows that mEnrich-seq can be tailored based on the de novo discovered methylation motifs from adult fecal microbiome samples to enrich bacterial strains with relatively low abundance in the sample.
- the mEnrich-seq strategy reduces the overall cost of recovering genomes from species with modest-to-low abundance, whose genomes would otherwise be cost prohibitive to resolve using standard metagenomics.
- mEnrich-seq significantly reduces sequencing reads from background bacteria, it can simplify the assembly of the enriched bacterial genomes.
- methylomes were compared with a list of 584 methylation-sensitive REs targeting 209 motifs that are commercially available (see Table 16, at end of the Examples) and found that 3,130 (68.03%) of the 4,601 strains have at least one (26.62% have two or more) methylation motif(s) that can be targeted (enriched by depleting the background metagenomic DNA) by commercially available RE(s).
- 3,130 strains When projected onto a phylogenetic tree (see Example 2, above), these 3,130 strains have broad taxonomic diversity (FIG. 5C), representing 54.78% of the species examined in this analysis.
- This analysis focused on REs with recognition motifs that have four, five or six nondegenerate bases, as these REs are expected to be readily applicable considering the expected frequency of their target motifs across bacterial genomes: roughly 256 bp for 4-mers, roughly 1 kb for 5-mers and roughly 4 kb for 6-mers. As researchers are actively developing more advanced DNA extraction methods that preserve high molecular weight gDNA from microbiome samples, it is expected that mEnrich-seq may expand to REs with additional target motifs with lower frequency.
- a methylation motif can be conserved at different taxonomic levels. Some methylation motifs are conserved at the family level. For example, the G6mATC motif by the Dam DNA methyltransferase is highly conserved in many families of the Gammaproteobacteria class, which includes many of the enteric pathogens (for example, E. coli, Salmonella enterica, Vibrio cholerae and so on) as well as some nonpathogenic species across many families such as Shewanellaceae and Moraxellaceae and so on. Some methylation motifs are conserved at the genus level.
- REs have very broad utility as their target motifs are methylated in a large number of bacterial genomes (FIG. 5D-5E). Taking DpnII (GATC) and XapI (RAATTY) as examples, the previous examples show their utilities with mEnrich-seq in two specific applications enriching for E. coli and A. muciniphila, respectively.
- mEnrich-seq with DpnII can help enrich many bacteria in the Gammaproteobacteria class in which G6mATC methylation is highly conserved
- mEnrich-seq with XapI can help enrich other bacteria such as Campylobacter spp., Acinetobacter spp., Spirochaeta spp., Treponema spp. and Brachyspira spp.
- the discriminative power of methylation motifs in practical applications is with respect to the bacteria in a specific microbiome sample, not among all the genomes used in the phylogenetic analysis.
- mEnrich-seq The core idea of mEnrich-seq is essentially two-fold: (i) methylation-guided digestions/enrichment followed by (ii) size selection.
- mEnrich-seq there are multiple ways to implement mEnrich-seq.
- the design described throughout the previous Examples performs adapter ligation before digestion ('ABD method’, uses gel purification for size selection, and utilizes the adapter sequence for subsequent PCR.
- AAD method adapter ligation after digestion
- Tn5-based adapter tagmentation T5 exonuclease to preferentially deplete shorter DNA fragments after digestion as described in a recent work called REMoDE (Enam, S. U.
- This library was the “Ligation-after-RE library” and it was directly compared to a standard mEnrich-seq library (500-100 ng as input) that was generated for evaluation of the adaptor ligation before RE digestion (“Ligation-before-RE”).
- the ONT data of Ligation-before-RE or Ligation-after-RE were each mapped to the reference using minimap2 (v2.24-rl !22) with -x map-ont and reads mapped to each species were counted.
- FIG. 16 and summarized in Table 11, below, adaptor-ligation before RE resulted in greater enrichment than adaptor-ligation after RE.
- FIGS 14A-14B, 15A-15B and 16 highlight the advantages of adapter ligation before digestion, it is worth noting that a single adapter sequence is not compatible with all the REs, because some enzymes can digest the adapter.
- the adapter sequence used in the previous Examples e.g., having an upper oligonucleotide sequence of SEQ ID NO: 35 and a lower oligonucleotide sequence of SEQ ID NO: 36, as described in Example 2, above
- 556 of the 584 (95%) commercially available REs is compatible with 556 of the 584 (95%) commercially available REs.
- one additional adapter was synthesized and found to be compatible with the remaining 28 commercially available REs.
- This second “alternative” adaptor contained a top oligo having a sequence of SEQ ID NO: 37 and a bottom oligo having a sequence of SEQ ID NO: 38. So, all the 584 commercially available REs are compatible with at least one of the two adapter oligos (FIG. 17).
- Example 9 Complementarity between mEnrich-seq and adaptive sampling.
- Adaptive sampling is effective in reducing the allocation of sequencing yield to host gDNA or highly abundant bacteria.
- desired taxa are indirectly enriched by the rejection of reads from highly abundant species, thus the efficiency enrichment by adaptive sampling from latest studies is still modest and insufficient for low-abundant bacteria.
- mEnrich-seq has the advantage of more efficiently depleting the vast background gDNA before sequencing, hence achieving high fold enrichments.
- Adaptive sampling can also select for specific sequences of interest (such as a desired genome): however, it does so by requiring a pre-defined genome sequence hence it will likely miss genes unique to a specific strain yet to be determined.
- mEnrich-seq has the advantage of enriching both known and unknown genes in a strain of interest, including mobile genetic elements, because all the genetic contents within a bacterial cell share the same methylation motifs. So, mEnrich-seq is complementary to the adaptive sampling method on the Nanopore sequencing platform.
- Bacterial DNA methylation provides a natural way for differentiating diverse bacterial taxa from each other. Although it has been used for metagenomic binning, it is only applicable as a post hoc analysis, after single molecule real-time (SMRT) sequencing or Nanopore sequencing. By contrast, mEnrich-seq enables the use of bacterial DNA methylation to enrich certain taxa of interest before sequencing, which allows microbiome researchers to tailor sequencing strategy to best serve their goals.
- SMRT single molecule real-time
- mEnrich-seq is directly applicable and is expected to facilitate many applications.
- the Examples herein demonstrate the use of mEnrich-seq to enrich E. coli from urine samples and to enrich A. muciniphila from fecal samples.
- they also demonstrate the use of mEnrich-seq along with de novo methylation discovers’ to enrich several low- abundance bacteria from a complex fecal microbiome sample.
- mEnrich-seq is broadly applicable and versatile. Based on a meta-analysis of the 4,601 bacterial methylomes mapped to date, it is estimated that roughly 68% of bacterial genomes can be targeted by mEnrich-seq with at least one RE, representing 54.78% of the species examined. In practice, mEnrich-seq may be applicable to a broader diversity of bacterial genomes, because epigenomes mapped to date represent only a small fraction of the bacterial diversity. mEnrich-seq is also very versatile in that it can be used along with one or multiple RE(s) (FIG. IB).
- mEnrich-seq provides a flexible way to dissect a microbiome sample when it is coupled with different REs, the choice of which can be tailored based on taxa of interest in a particular application.
- mEnrich-seq is compatible with different long-read and short-sequencing platforms (FIGs. 14A-14B and FIGs. 15A- 15B).
- the core idea of mEnrich-seq is essentially twofold: (1) methylation-guided digestions and/or enrichment follow ed by (2) size selection. Along w ith this core idea, there are multiple ways to implement mEnrich-seq.
- mEnrich-seq also has some limitations.
- mEnrich-seq is its efficacy in the enrichment of taxa of interest over standard metagenomics.
- mEnrich-seq is primarily applicable to enrich bacteria, archaea and dsDNA phages, but not applicable to single-stranded (ssDNA) phages, RNA phages and parasites, because they are not substrates of restriction-modification systems. It is worth noting a recent work leveraged the size differences between host and bacterial cells and used mechanical stress to effectively deplete host cells, which could be integrated with mEnrich-seq to study targeted taxa in host-rich samples.
- mEnrich-seq is a versatile and cost-effective approach for enrichment sequencing of diverse microbial taxa of interest directly from the microbiome. It provides microbiome researchers an additional approach to more effectively study complex microbiomes.
- Tables 14A-14B The MTase mediating RAATTY methylation is conserved across 1 12 genomes of A. muciniphila isolates.
- the amino acid sequence of the MTase (methylating RAATTY) is used as a query to blast against the 112 reference genomes.
- the prevalence of the MTase in 112 Akkermansia isolates is 100%, when using the following filter criteria for the BLAST hits: bitscores > 100 & identity > 80% & alignment length > 300.
- Table 15 The list of methylation motifs de novo discovered from the adult fecal metagenomic sample, de novo methylation discovery from microbiome samples was performed using existing tools (see Example 2). 112 methylation motifs were found across 34 bins (from 26 species) with lower abundance (Relative abundance profiled by Nanopore is indicated below). The bold motifs were the 4-mer or 5-mer motifs selected as the targets for illustrative purposes.
Landscapes
- Chemical & Material Sciences (AREA)
- Organic Chemistry (AREA)
- Life Sciences & Earth Sciences (AREA)
- Analytical Chemistry (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Biophysics (AREA)
- Immunology (AREA)
- Microbiology (AREA)
- Molecular Biology (AREA)
- Biotechnology (AREA)
- Physics & Mathematics (AREA)
- Chemical Kinetics & Catalysis (AREA)
- Biochemistry (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Genetics & Genomics (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
Methods for selectively amplifying a nucleic acid of one or more prokaryotic organisms of interest (target organisms) in a microbiome sample are provided. The methods generally comprise tagging genomic DNA fragments in a sample and then subjecting tagged DNA fragments to restriction enzyme digestion using methylation sensitive restriction enzymes chosen to selectively enrich nucleic acids of the one or more prokaryotic organisms of interest in the microbiome sample.
Description
COST-EFFECTIVE LONG READ METAGENOMIC SEQUENCING
CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of US Application No. 63/382,616, filed November 7, 2022, the entire disclosure of all of which is incorporated herein by reference.
ACKNOWLEDGEMENT OF GOVERNMENT SUPPORT
[0002] This invention was made with government support under Grant No. R35GM139655 awarded by the National Institute of General Medical Sciences National Institutes of Health. The government has certain rights in the invention.
INCORPORATION BY REFERENCE OF SEQUENCE LISTING
[0003] This application contains a Sequence Listing that has been submi tted in XML format via Patent Center and is hereby incorporated by reference in its entirety. The XML copy, created on January 3, 2024, is named 093698- 782858 SL. xml and is 81,000 bytes in size.
BACKGROUND
[0004] 1. Field
[0005] The present inventive concept is directed to methods and compositions for selectively amplifying a nucleic acid of one or more prokaryotic organisms of interest in a microbiome sample.
[0006] 2. Discussion of Related Art
[0007] Metagenomics have enabled the comprehensive and culture-independent study of microbiomes. However, for many applications, only certain taxa are of interest in a microbiome sample, for example, pathogenic bacteria, beneficial microbes, or taxa of low relative abundance. It is highly desirable that a method can effectively enrich bacterial taxa of interest directly from the microbiome.
[0008] The present disclosure is based on, in part, the suppressing discovery' that natural self vs. non-self genome differentiation as provided by bacterial DNA methylation can be exploited to rationally choose methylation-sensitive restriction enzymes (REs), individually or in combination, to deplete host DNA and most background microbial DNA while enriching bacterial taxa of interest. Accordingly, the present disclosure herein provides for a new method
of selectively enriching any bacterial target of interest in a sample of mixed bacterial taxa.
SUMMARY OF THE INVENTION
[0009] In various aspects of the present disclosure, a method of selectively amplifying a nucleic acid of one or more prokaryotic organisms of interest (target organisms) in a microbiome sample is provided. In various aspects, the method comprises: (a) obtaining or having obtained a microbiome sample comprising a plurality of prokaryotic-organisms, wherein the microbiome sample comprises genomic DNA (gDNA) from one or more target organisms (target gDNA) and gDNA from one or more prokaryotic organisms that are not of interest (background organism gDNA); (b) isolating gDNA from the microbiome sample to obtain target gDNA fragments and background organism gDNA fragments; (c) selecting one or more methylation-sensitive restriction enzymes based on the DNA methylome (a collection of DNA methylation sequence motifs) of the one or more target organisms wherein the one or more methylation restriction enzymes target motifs that are primarily methylated in the target gDNA (methylated motifs) and primarily unmethylated in the background organism gDNA (non-methylated motifs); (d) selecting universal PCR adapters that do not comprise nonmethylated motifs of the one or more methylation-sensitive restriction enzyme(s); (e) linking the universal PCR adapters to both 5’- and 3’-ends of target gDNA fragments and the background organism gDNA fragments; (!) subjecting the target gDNA fragments and the background organism gDNA fragments linked with the universal PCR adapters to digestion by the one or more methylation sensitive restriction enzyme(s) selected in (c); (g) amplifying undigested gDNA fragments comprising universal PCR adaptors at both 5 ’ and 3 ’ ends with forward and reverse barcoded primers corresponding to the universal PCR adapters, thereby selectively amplifying the genome of target prokaryotic organism in the microbiome gDNA sample.
[0010] In various aspects, the methods provided herein further comprise an end repair and/or a dA-tailing step prior to the linking of the universal PCR adapters.
[0011] In various aspects, the universal PCR adaptor can comprise an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper
strand oligo are not complementary' and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, and wherein the hybridized region does not comprise a nonmethylated motif targeted by the one or more restriction enzyme. In some aspects, the universal PCR adaptor further comprises a *T at the 3’ end of the upper strand oligo and/or is phosphorylated at the 5’ end of the lower strand oligo.
[0012] In various aspects, the upper strand oligo may have a nucleotide sequence comprising any one of SEQ ID NOs: 1 and 3-18 and the lower strand oligo may have a nucleotide sequence comprising any one of SEQ ID NOs: 2 and 19-34. For example, in some aspects, the upper strand oligo may have a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 36; or the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 38.
[0013] In various aspects, each upper-strand oligo of the universal primers herein can comprise a primer binding site complementary’ to a forward barcoded primer, and each lower- strand oligo comprises a primer binding site complementary’ to a reverse barcoded primer.
[0014] In various aspects, the methylated motif comprises a nucleic acid sequence motif having one or more methylated nucleotides. In some aspects, the one or more methylated nucleotides comprise N4-methylcytosine (4mC), N6-methyladenine (6mA) or 5- Methylcytosine (5mC) or any combination thereof. In still further aspects, the methylated motif may comprise GATC, RAATTY yvhere R is A or G and Y is C or T, or CTSAG where S is G or C. RGATCY where R is A or G and Y is C or T. GATGC. CTNNAG and each A is N6- methyladenine (6mA).
[0015] In any of the foregoing or related aspects, the non-methylated motif comprises a nucleic acid sequence motif that is recognized and cleaved by the one or more methylation sensitive restriction enzy me. In some aspects, for example, the non-methylated motif may comprise GATC. RAATTY where R is A or G and Y is C or T. or CTSAG where S is G or C, RGATCY where R is A or G and Y is C or T, GATGC. CTNNAG and each A is non- methylated.
[0016] In various aspects, the one or more methylation sensitive restriction enzymes comprise a class of natural and/or artificial restriction enzymes that are sensitive to DNA methylation. In some aspects, the methylation-sensitive restriction enzymes comprise DpnII,
XapI, Apol, BspCNI, BseMII, MflI, BstX2I, BstYI, MflI, XhoII, SfaNI, Aflll, Acul, Bpml, BpuEI, PstI, any methylation-sensitive isoschizomer thereof, or any combination thereof. In any of these or other related aspects, the methylation-sensitive restriction enzyme comprises a restriction enzyme whose cleavage ability is blocked by the methylated motif sites.
[0017] In various aspects, the methods provided herein further comprise identifying a methylation profile of the target prokaryotic organism and selecting the one or more restriction enzy mes based on the methylation profile of the target prokary otic organism.
[0018] In some aspects, the methods may further comprise deconvolving genomes of prokaryotic organisms in the microbiome sample to identify a methylation profile of the target prokaryotic organism, wherein deconvolving genomes comprises: A) obtaining a microbiome sample comprising a plurality of prokaryotic-organisms; B) sequencing nucleic acids of the prokaryotic organisms using single-molecule long reads sequencing technology, wherein the sequencing comprises the step of identifying methylated nucleotides, and at least one of the steps of: i. sequencing single molecule reads of nucleic acids; ii. assembling contigs from single molecule reads of the nucleic acids; C) assigning a methylation score reflecting the extent of methylation for sequence motifs of the nucleic acids on the assembled contig and/or the single molecule read; D) applying motif filtering to identify’ sequence motifs with methylation scores indicating methylation on the assembled contigs and/or the single molecVe reads; E) determining nucleic acid methylation profiles of the assembled contigs or the single molecule reads in the microbiome sample based on motifs identified in step (D); F) separating the assembled contigs and/or the single molecule reads into bins corresponding to distinct prokaryotic organisms based on the methylation profiles of step (E); and G) assembling the bins of step (F), thereby obtaining assembled genomes of the distinct bacterial organisms in the microbiome sample, thereby deconvolving genomes of the prokaryotic organisms in the microbiome sample. In some aspects, the method further comprises the step of combining the methylation profiles of step (E) with other sequence features of the nucleic acids of the prokaryotic organisms in the microbiome sample prior to separating the assembled contigs and/or the single molecule reads into bins.
[0019] In various aspects, the target prokaryotic organism(s) comprises one or more of an E. coll strain. Akkermansia muciniphila. Alistipes fmegoldii and Bacteroides ovatus. Bifidobacterium longum, Ruminococcus bromii, Eubacterium rectale, Bacteroides uniformis D7, Dorea longicatena and Barnesiella intestinihominis, Eubacterium sp. CAG:248, Alistipes
onderdonkii, Enterocloster clostridioformis, Bacteroides fragilis, Gordonibacter pamelaea,
Adlercreutzia sp. 8CFCBH1 , Adlercreutzia equolifaciens, or any combination thereof.
[0020] In various aspects the target prokaryotic organism(s) are selected from specific pathogenic bacteria, beneficial bacteria, probiotics, or prokaryotic organisms with low relative abundance in a microbiome sample as determined by standard metagenomic sequencing.
[0021] In any of the foregoing or related aspects, (i) the target prokaryotic organism(s) may comprise a methylated G6mATC motif, and the one or more methylation sensitive restriction enzyme comprises methylation sensitive DpnII and/or a methylation sensitive isoschizomer thereof; (ii) the target prokary otic organism(s) may comprise a methylated RA6mATTY motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive XapI , methylation sensitive Apol, and/or a methylation sensitive isoschizomer thereof, or any combination thereof; (iii) the target prokary otic organism(s) may comprise a methylated CTS6mAG motif and the one or more methylation sensitive restriction enzy me comprises methylation sensitive BspCNI, methylation sensitive BseMII, a methylation sensitive isoschizomer thereof, or any combination thereof; (iv) the target prokaryotic organism(s) may comprise a methylated RG6mATCY motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive MUI, methylation sensitive BstX2I, methylation sensitive BstYI, methylation sensitive Mfll, methylation sensitive XhoII, a methylation sensitive isoschizomer thereof, or any combination thereof; (v) the target prokaryotic organism(s) may comprise a methylated G6mATGC motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive SfaNI and/or a methylation sensitive isoschizomer; and/or (vi) the target prokaryotic organism(s) may comprise a methylated G6mATGC motif and a methylated CTNN6mAG motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive Aflll, methylation sensitive Acul, methylation sensitive Bpml, methylation sensitive BpuEI, methylation sensitive PstI, methylation sensitive SfaNI, a methylation sensitive isoschizomer thereof, or any combination thereof.
[0022] In various aspects, (i) the target prokary otic organism(s) comprising a methylated G6mATC motif may comprise Escherichia coli, Klebsiella pneumoniae, one or more species of Gammaproteobacteria class or any combination thereof; (ii) the target prokaryotic organism(s) comprising a methylated RA6mATTY motif may comprise one or more species of Akkermansia muciniphila, Campylobacter spp., Acinetobacter spp., Spirochaeta spp.,
Treponema spp. and Brachyspira spp or any combination thereof; (iii) the target prokaryotic organism(s) comprising a methylated CTS6mAG motif may comprise one or more species of Alistipes finegoldii and/or Bacteroides ovatus (iv) the target prokaryotic organism(s) comprising a methylated RG6mATCY motif may comprise one or more species of Bifidobacterium longum, Ruminococcus bromii. Eubacterium rectale and Bacteroides uniformis; (v) the target prokaryotic organism(s) comprising a methylated G6mATGC motif may comprise one or more species of Dorea longicatena. Barnesiella intestinihominis and Eubacterium sp. CAG:248 and/or (vi) the target prokary otic organism(s) comprising a methylated G6mATGC motif and methylated CTNN6mAG motif may comprise one or more species of Barnesiella intestinihominis.
[0023] In various aspects, any of the foregoing or related methods may further comprise sequencing the nucleic acid of the amplified target prokary otic organism(s). In various aspects, the sequencing is performed using a single-molecule long read real time (SMRT) technology, nanopore sequencing technology7 or other short read sequencing technologies including but not limited to Illumina sequencing platforms.
[0024] In various aspects, the microbiome sample may7 be obtained from soil, air, water, sediment, oil, and combinations thereof. In some aspects, the microbiome sample is obtained from water selected from marine water, fresh water, and rainwater. In some aspects, the microbiome sample is obtained from a subject selected from a protozoa, an animal, or a plant. In some aspects, the subject is a mammal. In some aspects, the subject is human. In some aspects, the subject is an infant.
[0025] Further aspects of the present disclosure relate to a universal adaptor, wherein the adaptor comprises an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, wherein the hybridized region does not comprise a motif targeted by a methylation sensitive restriction enzyme or does comprise a methylated nucleotide in a motif targeted by the methylation sensitive restriction enzyme.
[0026] In various aspects, the universal adaptor provided herein further comprises a *T at the 3’ end of the upper strand oligo and/or a phosphorylated base at the 5’ end of the lower
strand oligo.
[0027] In various aspects, the the upper strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 1 and 3-18 and the lower strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 2 and 19-34. In various aspects, the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 36; or the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 38.
[0028] Further aspects of the present disclosure provide for a kit comprising (a) a set of universal PCR adapters and (b) one or more methylation sensitive restriction enzymes.
[0029] In various aspects, the one or more methylation sensitive restriction enzymes provided in the kit comprise restriction enzymes sensitive to methylated motifs in one or more target prokaryotic organisms. In some aspects, the one or more methylation sensitive restriction enzymes comprise DpnII, XapI, Apol, BspCNI, BseMII, Mfll, BstX2I, BstYI, Mfll, XhoII, SfaNI, Aflll, Acul, Bpml, BpuEI, PstI or a combination of any thereof.
[0030] In various aspects, each universal PCR adapter in (a) comprises an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, wherein the hybridized region does not comprise a motif targeted by a methylation sensitive restriction enz me or does comprise a methylated nucleotide in a motif targeted by the methylation sensitive restriction enzyme.
[0031] In various aspects, each universal PCR adapter in (a) further comprising a *T at the 3’ end of the upper strand oligo and/or a phosphorylated base at the 5’ end of the lower strand oligo.
[0032] In various aspects, the upper strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 1 and 3-18 and/or the lower strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 2 and 19-34. In some aspects, the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 36; or the upper strand oligo has a nucleotide
sequence comprising SEQ ID NO: 37 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 38.
[0033] In various aspects, the kits provided herein may further comprise (c) a set of barcoded PCR primers complementary to at least a portion of the universal PCR adapters.
[0034] The foregoing is intended to be illustrative and is not meant in a limiting sense. Many features and subcombinations of the present inventive concept may be made and will be readily evident upon a study of the following specification and accompanying drawings comprising a part thereof. These features and subcombinations may be employed without reference to other features and subcombinations.
BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0036] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure, which can be better understood by reference to the drawing in combination with the detailed description of specific embodiments presented herein. Embodiments of the present inventive concept are illustrated by way of example in which like reference numerals indicate similar elements and in which.
[0037] FIG. 1A-1B depict mEnrich-seq workflow and its application, wherein: FIG. 1A depicts the mEnrich-seq workflow. 1st, the sheared DNA molecules are end-prepared and ligated to the barcode-adapter. 2nd, the barcode-adapter-ligated DNA molecules are digested with one or more restriction enzymes (REs, green scissors) that cut the non-methylated motif sites (green circles), leaving the methylated motifs (filled green dots) intact. 3rd, gel-based purification is performed to recover undigested DNA molecules at a high molecular weight. 4th, purified DNA from the previous step is amplified with PCR. In the 5th step, amplified DNA molecules are compatible with multiple sequencing platforms after library preparation per the sequencing platforms’ instructions. FIG. IB depicts the use of mEnrich-seq with different RE(s) to enrich different bacterial taxa from the same microbiome sample. The colored circles with dashed lines highlight bacteria in a microbiome sample that share common methylated motifs, which can be targeted by different REs (scissors with colors corresponding
to dashed circles). Each bacterial taxon can have multiple methylation motifs, and thus can be shared between different dashed circles. When one RE (e.g., the red scissors represent one RE) is added to the metagenomic DNA sample, a subgroup of bacteria of interest can be enriched (circled by the red dashed line), by depleting the rest of the bacteria (the gray circle). When applying the individual REs to the microbiome sample separately (represented by the individual yellow and blue scissors), mEnrich-seq allow researchers to examine different subgroups of bacteria in the microbiome sample (circled by individual yellow or blue dashed lines). mEnrich-seq can also be tailored with the use of multiple REs to more specifically enrich bacterial taxa with multiple methylation motifs (as illustrated by both the yellow and blue scissors, enriching the bacterium shared between blue and yellow dashed line).
[0038] FIG. 2A-2H depict validation of mEnrich-seq with the mock samples and urine samples and wherein: FIG. 2A depicts the percentage of reads mapped to E. coli in two mock replicates; the blue bar represents the input (human 94%, N. gonorrhoeae 5%, E. coli 1%); the orange bar represents mEnrich-seq. FIG. 2B depicts the fold change (mEnrich-seq vs. input) of percentage reads mapped to E. coli (dark green), N. gonorrhoeae (yellow) and human genome (blue) in the two mock replicates. FIG. 2C depicts E. coli genome coverage by sequencing reads (mEnrich-seq vs. input, with matched total yield). Maximum depth of each circos plot was capped at 95% quantile to ease visualization (Methods). mEnrich-seq reads mapping showed that the E. coli genome can be covered without systematic bias recovered for both mock samples (>99.96% genome covered). FIG. 2D depicts motif frequency on the mEnrich-seq reads that were mapped to N. gonorrhoeae. In replicates 1 and 2, 94.66% and 92.26% of reads that mapped to N. gonorrhoeae do not have any GATC sites (blue), respectively. The remainder of the reads have > 1 GATC/read (gray), which are further divided into two groups: the reads with GATC sites overlapping with the GGTG6mA motif of N. gonorrhoeae (labeled green), and Other (labeled orange), which reflect SNPs or sequencing errors (see Methods). FIG. 2E depicts Motif frequency on the mEnrich-seq reads that mapped to the human genome. In replicates 1 and 2, 77.04% and 84.63% of reads that mapped to human do not have any GATC sites (blue), respectively. The remainder of the reads have > 1 GATC/read (gray), which are further divided into two groups: the reads with GATC sites within repetitive elements that may be inaccessible to RE due to secondary- structure (labeled green), and Other (labeled orange), which reflect SNPs or sequencing errors (see Methods). FIG. 2F depicts scatter plots for three urine samples (UTI-1 to -3) showing the percentage of
reads mapped to E. coli, human and other taxa from the input (x-axis) vs. mEnrich-seq (y-axis). Across all the three samples, E. coli reads are significantly enriched in mEnrich-seq. In UTI-2, and UT1-3, the proportion of reads corresponding to the human genome are significantly decreased due to deletion by mEnrich-seq. FIG. 2G depicts fold enrichment of E. coli reads from mEnrich-seq compared to the input for the three urine samples. A break in the y-axis indicated by a double line indicates a deletion and a change in scale to permit all 3 samples to be graphed in the same plot. FIG. 2H depicts E. coli genome (chromosomal and plasmids) coverage by sequencing reads. Reads from mEnrich-seq of urine metagenomic DNA (orange) and the input standard urine metagenomic sequencing (blue) were mapped to the genome of corresponding E. coli isolates (sequenced and assembled beforehand for evaluation purpose), respectively. The maximum depth of circos plots were capped at 95% quantile to ease visualization (Methods). mEnrich-seq reads mapping showed that the E. coli chromosomes and plasmids can be covered without systematic bias for the three urine samples (> 99.97% of the genome covered).
[0039] FIG. 3A-3D depict the enrichment of A. muciniphila from fecal samples by mEnrich-seq wherein: FIG. 3A shows scatter plots for three fecal samples (GUT-1 to -3) showing the percentage of reads mapped to A. muciniphila and other taxa from the input (x- axis) and mEnnch-seq (y-axis). To ease visualization, top 50 abundant species, in input or mEnrich-seq, are shown. For the enriched species, the size of the dots represents the fold change by mEnrich-seq over the input. Significant increases in read % are observed for A. muciniphila and a few other taxa, going from input to mEnrich-seq samples, while some others are depleted in such as P. dorei, which undergoes around 10- to 15-fold decrease in read %. FIG. 3B shows fold enrichment of reads that map to A. muciniphila from mEnrich-seq over the input for GUT-1 to -3. FIG. 3C shows A. muciniphila genome coverage by sequencing reads of GUT-1 to -3. Reads from mEnrich-seq (orange) and the input standard metagenomic sequencing (blue) were mapped to the genome of corresponding A. muciniphila isolates (sequenced and assembled beforehand for evaluation purpose), respectively. The maximum depth of circos plot was capped at 95% quantile to ease visualization (Methods). mEnrich-seq reads mapping showed that the A. muciniphila genome can be covered without systematic bias for the three fecal samples (>99.7% genome covered). FIG. 3D shows analysis of RAATTY motif frequency (per kilobase) for the taxa enriched by mEnrich-seq (XapI as RE) in GUT-1 to -3. X-axis, RAATTY motif frequency on the reference genomes of enriched taxa; Y-axis,
read-level RAATTY frequency in mEnrich-seq. Color is used to identify enriched species, while the shapes (square, triangle, or circle) are used to indicate sample of origin. The green dashed line circles A. muciniphila. which is enriched due to its methylated RA6mATTY motif sites. The red dashed line circles the bacterial taxa enriched due to the sparse RAATTY frequency on their genomes.
[0040] FIG. 4A-4J depict enrichment sequencing of bacterial taxa based on de novo discovered methylation motifs, wherein: FIG. 4A depicts a workflow for rational design of mEnrich-seq based on de novo motif discovery. An initial metagenomic sequencing was conducted to estimate the relative abundance of different taxa. The goal is to design mEnrich- seq to enrich taxa with relatively low abundance. To do so, de novo methylation motif discovery' was performed from the initial metagenomic sequencing data. Importantly, motif discovery' does not need fully assembled genomes, because a partially assembled contig can already inform us about the methylation motifs from a strain. Based on the de novo discovered methylation motifs, REs were selected that rationally deplete the background metagenomic DNA while enriching for bacterial strains with relatively low abundance with mEnrich-seq. FIG. 4B depicts fold enrichment of six bacterial strains with relatively low abundance; enrichment by mEnrich-seq based on de novo discovered methylation motifs. One of the strain B. intestinihominis, was further ennched using multiple REs simultaneously to provide greater enrichment (85-fold) than a single RE (30-fold), illustrating the flexible design of mEnrich-seq when multiple methylation motifs are discovered on the strain of interest. FIG. 4C depicts scatter plots showing the percentage of reads mapped to bacteria from the input (x-axis) and mEnrich-seq using the RE (BspCNI) targeting the CTCAG motif (y-axis). For enriched species, the size of the dots represents the fold change by mEnrich-seq over the input. To ease visualization, top 100 abundant species, in input or mEnrich-seq, are shown. FIG. 4D depicts analysis of the CTCAG motif frequency (per kilobase) for the taxa enriched by mEnrich-seq (BspCNI as RE). X-axis. CTCAG motif frequency in the reference genomes of enriched taxa; Y-axis, read-level CTCAG frequency in mEnrich-seq. The green dashed line circles the taxa enriched due to their methylated CTC6mAG motif sites. The red dashed line circles the bacterial taxa enriched due to the sparse CTCAG frequency on their genomes. FIG. 4E depicts data as described for FIG. 4C but for the RE targeting the RGATCY moti. FIG. 4F depicts data as described for FIG. 4D but for the RE targeting the RGATCY motif. FIG. 4G depicts data as described for FIG. 4C but for the RE targeting the GATGC motif. FIG. 4H depicts data
as described for FIG. 4D but for the RE targeting the GATGC motif. FIG. 41 depicts data as described for FIG. 4C but for the multiple REs targeting the GATGC motif and CTNNAG motif simultaneously. FIG. 4J shows that mEnrich-seq improved genome assembly compared to standard metagenomic sequencing. The left shows the completeness (%) of genome assembly; blue bars show the input and orange bars show the mEnrich-seq; * completeness less than 1%. On the right, contig N50 length is shown in Mb. For each targeted species, reads from input and mEnrich-seq were normalized by random subsampling to make sure reads originated from the targeted species have the same yield for genome assembly comparison.
[0041] FIG. 5A-5E depict experimental data assessing the broad applicability of mEnrich- seq, wherein: FIG. 5A depicts a histogram for the number of methylation motifs in each bacterial genome across the 4,601 bacterial methylomes mapped to date (as of 07/18/2022, REBASE). X-axis, number of motifs per bacterium; Y -axis, number of bacterial species having a given number of motif. FIG. 5B depicts a histogram for the number of non-degenerate bases (excluding 'N’s) per motif across 4683 motifs. For example, both GATC and CTNNAG have 4 non-degenerate bases, FIG. 5C depicts a phylogenic tree containing 4,601 bacterial strains with mapped methylomes to date. The bacterial strains that can be targeted (enriched by depleting the background metagenomic DNA) by at least one commercially available REs are highlighted with colors, demonstrating broad distribution across the tree. Number of motifs recognized by commercially available REs (and motif length) for a given bacteria are color coded as detailed in the panel at the bottom right, and FIG. 5D depicts a histogram for number of bacteria targeted (enriched by depleting the background metagenomic DNA) by each commercially available RE (please note that a single motif can be targeted by multiple restriction enzymes, namely, isoschizomers). X-axis, each column represents an individual restriction enzyme; Y-axis, number of bacterial genomes per RE target. The most abundant motifs (corresponding to the REs) are indicated by label in the plot. FIG. 5E shows taxonomic distribution of the seven most prevalent methylation motifs. The taxonomic information is indicated by color-coding branches by order (top 27 orders are color coded). The distribution of top seven motifs is highlighted by colored squares around the phylogenetic tree.
[0042] FIG. 6A-6C depict mEnrich-seq of urine samples using Nanopore Flongle flow cells and wherein: FIG. 6A shows three urine samples were sequenced as the Input: UTI-1 contained mostly bacterial DNA, while UTI-2 and UTI-3 were dominated by human DNA, FIG. 6B depicts scatter plots for six urine samples (UTI-1 to -6) showing the percentage of
reads mapped to E. coli, human and other taxa from the input (x-axis) vs. mEnrich-seq (y- axis). The UTI-1 to -3 samples demonstrate consistent patterns as shown in FIG. 2 (MinlON flowcell); the additional three UTI urine samples (UT1 -4 to -6) are included to illustrate the consistent enrichment of E. coli by mEnrich-seq. Across all six samples, E. coli reads are significantly enriched in mEnrich-seq. In UTI-2, UTI-3, UTI-4, and UTI-6, the proportion of reads corresponding to the human genome are significantly lower due to cleavage by DpnII in mEnrich-seq and FIG. 6C depicts fold enrichment of E. coli reads from mEnrich-seq compared to the input for the six urine samples.
[0043] FIG. 7A-7B depicts mEnrich-seq data of UTI-1 including reads that have high- confidence mapping to E. coli yet carry additional genes beyond the isolated strain from standard urine culture. FIG. 7A 10 reads with highly confident mapping to E. coli from mEnrich-seq of UTI-1 urine sample, consistently supporting a tet(A) resistant gene (AF534183, blue boxes). For each read, lengths of the sequences flanking the resistance gene were shown on both sides. The identities between the reads and AF534183 were shown inside the blue boxes. FIG. 7B shows validation of tet(A) in UTI-1 sample. Agarose gel electrophoresis of the tet(A) regions amplified once by two sets of primers (of lengths 1150 bp and 900 bp, respectively). The bands (in red square) were further purified for gene sequence confirmation by Sanger sequencing.
[0044] FIG. 8A-8C depict confirmation of DNA 6mA methylation at RA6mATTY motif sites using PacBio sequencing of two A. muciniphila strains, wherein: FIG. 8A depicts RA6mATTY (mediated by MTase AmuORF1905P) is one of the five methylated motifs detected in A. muciniphila (ATCC-835) by PacBio sequencing as summarized in REBASE, FIG. 8B depicts DNA 6mA methylation at RAATTY motif sites (light blue) is confirmed based on the high IPD ratios observed in PacBio sequencing data of the A. muciniphila strain isolated from the infant fecal sample GUT-3, and FIG. 8C depicts DNA 6mA methylation at RAATTY motif sites (purple) is confirmed based on the high IPD ratio observed in PacBio sequencing data of another A. muciniphila strain.
[0045] FIG. 9A-9B depict agarose gels of the two mock replicates, wherein: FIG. 9A depicts bands of ~10kb size are observed for two mock replicates digested by DpnII. E. coli DNA is used as the positive control for size reference in band collection from the gel (5- lOkb) and FIG. 9B depicts the Input and DpnII-digested DNA are amplified by the barcoded primers and visualized on an agarose gel. The PCR bands for DpnII-digested DNA are 5- 10
kb size.
[0046] FIG. 10A-10B depict agarose gels for urine samples after the DpII-digestion and barcoding PCR steps, wherein: FIG. 10A depicts RE-digested urine samples (UTI-1 to -3) are visualized on an agarose gel. Bands are observed, and cut out, for all three urine samples, around 5- lOkb in size. E. coli DNA (positive control) and lymphoblastoid cells (LCL, GM24149) DNA (negative control) are included to verify size estimations and FIG. 10B depicts a gel showing bands of the Input and DpnII-digested DNA that are amplified by the barcoded primers. PCR bands of DpnII-digested DNA are shown around 5- lOkb size.
[0047] FIG. 11A-11B depict agarose gels for fecal samples after Xapl-digestion and barcoding PCR steps, wherein: FIG. 11A depicts a gel of the infant fecal samples (GUT-1 to - 3) digested by XapI (recognition site: RAATTY). Faint (highlighted by the red square) bands are observed for the three fecal samples on an agarose gel. DNA w as collected by cutting the gel at -5- lOkb (based upon the DNA ladder) and FIG. 11B depicts gels show' the bands of the Input (left panel) and Xapl-digested DNA (right panel) that are amplified by the barcoded primers. The PCR bands of Xapl-digested DNA are shown ~5- lOkb in size.
[0048] FIG. 12A-12B depict agarose gels of enriched low-abundance bacterial taxa from the adult fecal sample based on de novo discovered methylation motifs, wherein: FIG. 12A depicts gels for the adult fecal sample digested by different REs. Left, the Mfll-digested the adult fecal sample DNA (targeting RGATCY motif) shows the band at -5- lOkb; the BspCNI (CTC AG as recognition site, part of the CTSAG motif)- digested sample shows a faint band at -5- lOkb (based upon the DNA ladder). Right, faint bands (highlighted by the red squares) are observed for the adult fecal sample digested by SfaNI (GATGC recognition site) and by the RE cocktails (targeting both GATGC and CTNNAG) and FIG. 12B depicts gels for PCR bands amplified by the barcoded primers, using RE-digested adult fecal DNA (from Fig. 7a) as the template. The PCR bands of RE(s)-digested DNA are shown around 5—10 kb.
[0049] FIG. 13 depicts a plot showing de novo methylation motif discovery from an initial shallow' sequencing run (400k PacBio Sequel II CCS reads). Many methylation motifs were discovered from SMRT sequencing data, even those from bacterial genomes with -10% completeness. The relative abundance of each bacterial genome is included in the x-axis label in red color. 4-mer motifs (those with four non-degenerate bases) and similarly 5-mer and 6- mer motifs are color coded
[0050] FIG. 14A-14B shows via a direct comparison that the mEnrich-seq design has better fold enrichment than REMoDE. FIG. 14A shows agarose gel electrophoresis of Mock- 3 (a mock with five species, see Supplementary Information) after treatment by RE and T5 exonuclease for once. FIG. 14A (left) shows Agarose gel electrophoresis of the mock sample digested by RE, Xapl. Due to the sparse RE recognition sites in the background genomes in the mock microbiome sample, relatively long DNA fragments from the background genomes remain even after RE digestion; FIG. 14A (right), The gel image of the mock sample digested by RE and T5. FIG. 14B shows a direct comparison between mEnrich-seq and REMoDE was performed using the same mock microbiome sample (Mock-3), and the same sequencing platform (Illumina). To be compatible with the short-read sequencing platform, the library of mEnrichseq was prepared according to the "Compatibility’ between mEnrich-seq and other sequencing platforms’ section in Example 2.
[0051] FIG. 15A-15B show that mEnrich-seq design has better fold enrichment than REMoDE when they are compared with the same, real microbiome samples with DNA fragmentation and damages. FIG. 15A depicts agarose gel electrophoresis tested once on genomic DNA isolated from Mock-3 (illustrating the ideal high molecular weight gDNA from cell culture) and two infant fecal samples (representing DNA quality’ in real applications). FIG. 15B depicts data comparing the fold enrichment of A. muciniphila from two infant fecal samples using mEnrich-seq, REMoDE, or REMoDE with DNA repair beforehand (see Supplementary’ Information). As DNA repair is part of mEnrich-seq by default, this step is also applied to REMoDE for a fair comparison.
[0052] FIG. 16 depicts an evaluation of the benefit of adapter ligation before vs. after RE- digestion using the Mock-3. The bar graph demonstrates that adapter ligation before RE- digestion indeed has better fold ennchment.
[0053] FIG. 17 shows that an alternative adapter has comparable enrichment efficiency as the original adapter. The fold enrichment of A. muciniphila (using Xapl as RE) was tested on Mock-3 with either the original adapter (having an upper and loyver oligonucleotide corresponding to SEQ ID NOs: 35 and 36, respectively) or the alternative adapter (having an upper and loyver oligonucleotide corresponding to SEQ ID NOs: 37 and 38, respectively).
[0054] FIG. 18 shows that GC content does not introduce systematic bias on read coverage across the E. coll genome analyzed by mEnrich-seq. Analysis was performed on Mock-1 and Mock-2 (E. coll relative abundance as 0.91% and 1.20%, see Supplementary Information). For
a fair comparison, the same yield of data from both input and mEnrich-seq were used in data analysis and figures. Each dot represents a 5 kb genomic region, colored by GC content. X- axis: number of input reads mapped to each 5 kb genomic region; Y-axis: number of mEnrich- seq reads mapped to each 5 kb genomic region. To ease visualization, a random value was added on x-values to avoid overlapping dots.
[0055] FIG. 19 shows that GC content does not introduce systematic bias on read coverage across the A. muciniphila genome by mEnrich-seq. The same type of scatter plots (as the ones above for E. coli) made with sequencing data from three human gut samples (A. muciniphila relative abundance: 0.48%, 0.52% and 1.52%, respectively) to characterize the enrichment of A. muciniphila. Consistent with the Circos plot for the same samples, read depths were capped at the 95% quantile to ease visualization.
DETAILED DESCRIPTION
[0056] The following detailed description references the accompanying drawings that illustrate various embodiments of the present inventive concept. The drawings and description are intended to describe aspects and embodiments of the present inventive concept in sufficient detail to enable those skilled in the art to practice the present inventive concept. Other components can be utilized and changes can be made without departing from the scope of the present inventive concept. The following description is, therefore, not to be taken in a limiting sense. The scope of the present inventive concept is defined only by the appended claims, along with the full scope of equivalents to which such claims are entitled.
[0057] The present disclosure is based, at least in part, on the discovery that natural self vs. non-self genome differentiation as provided by bacterial DNA methylation can be exploited to rationally choose methylation-sensitive restriction enzymes (REs), individually or in combination, to deplete host DNA and most background microbial DNA while enriching bacterial taxa of interest. Accordingly, the present disclosure herein provides for a new method of selectively enriching any bacterial target of interest in a sample of mixed bacterial taxa.
I. Terminology
[0058] The phraseology and terminology employed herein are for the purpose of description and should not be regarded as limiting. For example, the use of a singular term, such as, “a” is not intended as limiting of the number of items. Also, the use of relational terms
such as, but not limited to, “top,” “bottom,” “left,” “right,” “upper,” “lower,” “down,” “up,” and “side,” are used in the description for clarity in specific reference to the figures and are not intended to limit the scope of the present inventive concept or the appended claims.
[0059] Further, as the present inventive concept is susceptible to embodiments of many different forms, it is intended that the present disclosure be considered as an example of the principles of the present inventive concept and not intended to limit the present inventive concept to the specific embodiments shown and described. Any one of the features of the present inventive concept may be used separately or in combination with any other feature. References to the terms “embodiment.” “embodiments,” and/or the like in the description mean that the feature and/or features being referred to are included in, at least, one aspect of the description. Separate references to the terms “embodiment,” “embodiments,” and/or the like in the description do not necessarily refer to the same embodiment and are also not mutually exclusive unless so stated and/or except as will be readily apparent to those skilled in the art from the description. For example, a feature, structure, process, step, action, or the like described in one embodiment may also be included in other embodiments, but is not necessarily included. Thus, the present inventive concept may include a variety of combinations and/or integrations of the embodiments described herein. Additionally, all aspects of the present disclosure, as described herein, are not essential for its practice. Likewise, other systems, methods, features, and advantages of the present inventive concept will be, or become, apparent to one with skill in the art upon examination of the figures and the description. It is intended that all such additional systems, methods, features, and advantages be included within this description, be within the scope of the present inventive concept, and be encompassed by the claims.
[0060] As used herein, the term “about.” can mean relative to the recited value, e.g . amount, dose, temperature, time, percentage, etc., ±10%, ±9%, ±8%, ±7%, ±6%, ±5%, ±4%, ±3%, ±2%, or ±1%.
[0061] The terms "comprising," "including," “encompassing” and "having" are used interchangeably in this disclosure. The terms "comprising," "including," “encompassing” and "having" mean to include, but not necessarily be limited to the things so described.
[0062] The terms “or” and “and/or,” as used herein, are to be interpreted as inclusive or meaning any one or any combination. Therefore, “A, B or C” or “A, B and/or C” mean any of the following: “A,” “B” or “C”; “A and B”; “A and C”; “B and C”; “A, B and C.” An exception
to this definition will occur only when a combination of elements, functions, steps or acts are in some way inherently mutually exclusive.
[0063] The term “nucleic acid'’ or “polynucleotide” refers to deoxyribonucleic acids (DNA) or ribonucleic acids (RNA) and polymers thereof in either single- or double-stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogues of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g.. degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and/or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolmi et ak.Afo/. Cell. Probes 8:91-98 (1994)).
II. Methods
[0064] In various aspects of the present disclosure, a method of selectively amplifying a nucleic acid of one or more prokaryotic organisms of interest (target organisms) in a microbiome sample is provided. In various aspects, the method comprises: (a) obtaining a microbiome sample comprising a plurality of prokaryotic-organisms, wherein the microbiome sample comprises genomic DNA (gDNA) from one or more target organisms (target gDNA) and gDNA from one or more prokaryotic organisms that are not of interest (background organism gDNA); (b) isolating gDNA from the microbiome sample to obtain target gDNA fragments and background organism gDNA fragments; (c) selecting one or more methylationsensitive restriction enzymes based on the DNA methylome (a collection of DNA methylation sequence motifs) of the one or more target organisms wherein the one or more methylation restriction enzymes target motifs that are primarily methylated in the target gDNA (methylated motifs) and primarily unmethylated in the background organism gDNA (non-methylated motifs); (d) selecting universal PCR adapters that do not comprise non-methylated motifs of the one or more methylation-sensitive restriction enzyme(s); (e) linking the universal PCR adapters to both 5’- and 3’-ends of target gDNA fragments and the background organism gDNA fragments; (f) subjecting the target gDNA fragments and the background organism gDNA fragments linked with the universal PCR adapters to digestion by the one or more
methylation sensitive restriction enzyme(s) selected in (c); (g) amplifying undigested gDNA fragments comprising universal PCR adaptors at both 5’ and 3' ends with forward and reverse barcoded primers corresponding to the universal PCR adapters, thereby selectively amplifying the genome of target prokaryotic organism in the microbiome gDNA sample.
Universal Adaptors and PCR Primers
[0065] The methods here are, in part, directed to methods that improve upon other methods of using methylation sensitive restriction enzymes to enrich target nucleic acid populations, by ligating a set of universal adaptors to the target (and non-target) DNA before any digestion occurs. The universal adaptors (also referred to herein as “universal PCR adaptors’') are designed to be usable across almost all restriction enzymes currently available. Suitable universal PCR adaptors are described further in Section (III) below. These universal adaptors may be targeted by suitable PCR primers after digestion has occurred to selectively amplify and/or sequence any un-digested DNA (i.e., any DNA fragments still containing universal adaptors at both the 5‘ and 3‘ ends). Suitable PCR primers can target certain primer binding motifs in the universal PCR adaptors as described further in Section (III) below.
Methylation Motifs and Methylation-Sensitive Restriction Enzymes
[0066] In any of the above or foregoing aspects, the methylated motif comprises a nucleic acid sequence motif having methylation on one or more nucleotides, wherein the methylation optionally comprises N4-methylcytosine (4mC), N6-methyladenine (6mA) or 5- Methylcytosine (5mC), and wherein the methylated motif optionally comprises GATC, RAATTY where R is A or G and Y is C or T, or CTSAG where S is G or C, RGATCY where R is A or G and Y is C or T, GATGC, CTNNAG and each A is N6-methyladenine (6mA). Accordingly, in further aspects, the non-methylated motif comprises a nucleic acid sequence motif that is recognized and cleaved by a restriction enzyme, and wherein the non-methylated motif optionally comprises GATC, RAATTY where R is A or G and Y is C or T, or CTSAG where S is G or C, RGATCY where R is A or G and Y is C or T, GATGC, CTNNAG and each A is non-methylated.
[0067] In any of the foregoing or related aspects, the one or more methylation sensitive restriction enzymes may comprise a class of natural and/or artificial restriction enzymes that are sensitive to DNA methylation, and optionally comprise methylation-sensitive restriction
enzy mes DpnII, XapI, Apol, BspCNI, BseMII, Mfll, BstX2I, BstYI, Mfll, XhoII, SfaNI, AflII, Acul. Bpml, BpuEI, PstI, any methylation-sensitive isoschizomer thereof, or a combination of any thereof. In various aspects the methylation-sensitive restriction enzyme comprises a restriction enzyme whose cleavage ability is blocked by the methylated motif sites. In still further aspects, the method comprises identifying a methylation profile of the target prokaryotic organism and selecting the one or more restriction enzy mes based on the methylation profile of the target prokaryotic organism.
Deconvoluting genomes to identify methylation profile of target microorganisms
[0068] In further aspects of the present disclosure, the method further comprises deconvoluting genomes of prokaryotic organisms in the microbiome sample to identify a methylation profile of the target prokaryotic organism, wherein deconvoluting genomes comprises: A) obtaining a microbiome sample comprising a plurality of prokary otic- organisms; B) sequencing nucleic acids of the prokaryotic organisms using single-molecule long reads sequencing technology, wherein the sequencing comprises the step of identifying methylated nucleotides, and at least one of the steps of: i. sequencing single molecule reads of nucleic acids; ii. assembling contigs from single molecule reads of the nucleic acids; and C) assigning a methylation score reflecting the extent of methylation for sequence motifs of the nucleic acids on the assembled contig and/or the single molecule read; D) applying motif filtering to identify sequence motifs with methylation scores indicating methylation on the assembled contigs and/or the single molecule reads; E) determining nucleic acid methylation profiles of the assembled contigs or the single molecule reads in the microbiome sample based on motifs identified in step (D); F) separating the assembled contigs and/or the single molecule reads into bins corresponding to distinct prokaryotic organisms based on the methylation profiles of step (E); G) assembling the bins of step (F), thereby^ obtaining assembled genomes of the distinct bacterial organisms in the microbiome sample, thereby deconvoluting genomes of the prokaryotic organisms in the microbiome sample.
[0069] In some aspects, the method of deconvoluting genomes further comprises the step of combining the methylation profiles of step (E) with other sequence features of the nucleic acids of the prokaryotic organisms in the microbiome sample prior to separating the assembled contigs and/or the single molecule reads into bins.
[0070] Additional methods useful to deconvolve genomes to derive at ideal combinations of methylation sensitive restriction enzymes or that may be used in combination with any of the methods and compositions provided herein are described in US20200160936A1, US20220254446A1, US20220230704A1, which are each incorporated herein by reference in their entirety.
Target Prokaryotic Organisms
[0071] In accordance with various aspects, the target prokary otic organism may comprise one or multiple of the following an E. coli strain, Akkermansia muciniphila, Alistipes fmegoldii and Bacteroides ovatus. Bifidobacterium longum.. Ruminococcus bromii, Eubacterium rectale, Bacteroides uniformis D7, Dorea longicatena and Barnesiella intestinihominis, Eubacterium sp. CAG:248, Alistipes onderdonkii, Enterocloster clostridioformis, Bacteroides fragilis, Collinsella aerofaciens, Eggerthella lenta, Gordonibacter pamelaea, Adlercreutzia sp. 8CFCBH1 , and/or Adlercreutzia equolifaciens In general, the target prokaryotic organisms may be selected from specific pathogenic bacteria, beneficial bacteria, probiotics, or prokaryotic organisms with low relative abundance in a microbiome sample as determined by standard metagenomic sequencing.
[0072] Accordingly, in various aspects, the target prokaryotic organism, methylation motif and restriction enzyme may be selected from the following groups: (i) the target prokaryotic organism(s) comprise a methylated G6mATC motif, and the one or more methylation sensitive restriction enzyme comprises methylation sensitive DpnII and/or a methylation sensitive isoschizomer thereof; (ii) the target prokaryotic organism(s) comprise a methylated RA6mATTY motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive XapI , methylation sensitive Apol, and/or a methylation sensitive isoschizomer thereof, or any combination thereof; (iii) the target prokaryotic organism(s) comprise a methylated CTS6mAG motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive BspCNL methylation sensitive BseMII, a methylation sensitive isoschizomer thereof, or any combination thereof; (iv) the target prokaryotic organism(s) comprise a methylated RG6mATCY motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive Mfll, methylation sensitive BstX2I, methylation sensitive BstYI, methylation sensitive Mfll, methylation sensitive XhoII, a methylation sensitive isoschizomer thereof, or any combination thereof; (v) the target
prokaryotic organism(s) comprise a methylated G6mATGC motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive SfaNI and/or a methylation sensitive isoschizomer; and/or (vi) the target prokaryotic organism(s) comprise a methylated G6mATGC motif and a methylated CTNN6mAG motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive Aflll, methylation sensitive Acul, methylation sensitive Bpml, methylation sensitive BpuEI, methylation sensitive PstI, methylation sensitive SfaNI. a methylation sensitive isoschizomer thereof, or any combination thereof.
[0073] In some aspects, the target prokaryotic organism(s) comprising a methylated G6mATC motif comprise Escherichia coll, Klebsiella pneumoniae, one or more species of Gammaproteobacteria class or any combination thereof; (ii) the target prokaryotic organism(s) comprising a methylated RA6mATTY motif comprise one or more species of Akkermansia muciniphila, Campylobacter spp., Acinetobacter spp., Spirochaeta spp., Treponema spp. and Brachyspira spp or any combination thereof; (iii) the target prokaryotic organism(s) comprising a methylated CTS6mAG motif comprise one or more species ofAlistipes finegoldii and/or Bacteroides ovatus; (iv) the target prokaryotic organism(s) comprising a methylated RG6mATCY motif comprise one or more species of Bifidobacterium longum, Ruminococcus bromii, Eubacterium rectale and Bacteroides uniformis; (v) the target prokaryotic organism(s) comprising a methylated G6mATGC motif comprise one or more species of Dorea longicatena, Barnesiella intestinihominis and Eubacterium sp. CAG:248 and/or (vi) the target prokaryotic organism(s) comprising a methylated G6mATGC motif and methylated CTNN6mAG motif comprise one or more species of Barnesiella intestinihominis.
[0074] In various aspects, the method may further comprise a DNA shearing step, and wherein shearing the gDNA optionally comprises g-TUBEs and the gDNA fragments from target and background organisms have a length of about 10 kb. In various aspects, the method further comprises an end repair and dA-tailing step prior to the linking of said universal PCR adapter.
[0075] In still further aspects the method may comprise sequencing the nucleic acid of the amplified target prokaryotic organism(s). In some aspects, the sequencing is performed using a single-molecule long read real time (SMRT) technology, nanopore sequencing technology or other short read sequencing technologies including but not limited to Illumina sequencing
platforms.
[0076] Any suitable microbiome sample may be used in the methods described herein. In some aspects, the microbiome sample is obtained from soil, air, water, sediment, oil, and combinations thereof. For example, the microbiome sample may be obtained from water selected from marine water, fresh water, and rainwater. In other aspects, the microbiome sample may be obtained from a subject selected from a protozoa, an animal, or a plant. In some aspects, the subject may be a mammal. In preferred aspects, the subject is a human. In some aspects, the subject may be an infant. In some aspects, the sample obtained from a mammal (e.g., human) may comprise a urine sample.
111. Compositions
[0077] In various aspects of the present disclosure, a universal adaptor for use in any of the methods described herein is provided. In various aspects, the universal adaptors can comprise about 50-100 bp double stranded short DNA fragments with an overhang at the 3’ end of the top strand, which facilitates ligation to all DNA molecules through T/A ligation. In further aspects, the universal adaptors will not include any recognition sequences in the adapter sequence. Suitable universal adaptors may be “Y- shaped” or “bell-shaped”, or may come in other structures. In some aspects, the universal PCR adaptor comprises an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hy bridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, and wherein the hybridized region does not comprise a nonmethylated motif targeted by the one or more restriction enzyme.
[0078] In further aspects, the universal adaptor can further comprise a *T at the 3‘ end of the upper strand oligo, wherein *T comprises a phosphorothioated thiamine base having a phosphorothioate nucleotide modification, wherein the phosphorothioate modification renders an inter-nucleotide linkage formed by the universal adaptor resistant to nuclease degradation. In further aspects the universal adaptor can be phosphory lated at the 5’ end, particularly on the lower strand oligonucleotide. Other modifications to prevent nuclease degradation are known in the art and are contemplated herein.
[0079] In various aspects, the non-hybridized region of the universal PCR adaptor can comprise or consist of SEQ ID NO: 1 (for the upper oligo) and/or SEQ ID NO: 2 (for the lower oligo). Each of the upper and lower oligo may comprise of an additional 10 to 25 nucleotides that hybridize and that, when hybridized, do not comprise a non-methylated motif targeted by the one or more restriction enzyme. Preferably, the upper and lower oligos comprise the same number of additional nucleotides that hybridize. Accordingly, in various aspects, the upper strand may comprise a nucleotide sequence of any one of SEQ ID NOs: 3-18 and the lower strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 19-34. In still further aspects, the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 36. In still further aspects, the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 38. For ease of reference, SEQ ID NOs: 1 -38 are provided in the Table 1 below, where the phosphorothioate linkage is indicated by an asterisk (*) and the 5’phosphorylation modification is indicated by a /5Phos/.
Table 1
[0080] In various aspects, each upper-strand oligo comprises a primer binding site complementary to a forward barcoded primer, and each lower-strand oligo comprises a primer binding site complementary to reverse barcoded primer.
[0081] In further aspects, a set of barcoded PCR primers is provided, where the barcoded PCR primers target any portion of the upper or lower oligo of the universal adaptor. So, for example, the barcoded PCR primer may target a nucleic acid sequence comprising at least a portion of a nucleic acid sequence comprising or consisting of SEQ ID NOs 1 or 2, which are shown again in Table 2 for ease of reference. As used herein, the term "‘target” means that the primer hybridizes to a nucleic acid comprising the nucleic acid target sequence to enable PCR amplification. In some aspects the primer hybridizes to the 5’ strand (e g., upper oligo of universal adaptor) comprising the target sequence or to the 3 ' strand complementary to the 5 ’ strand. In other aspects, the primer hybridizes to the 3‘ strand (e.g., lower oligo of universal adaptor) comprising the target sequence or to the 5’ strand complementary to the 3’ strand.
Table 2
IV. Kits
[0082] In various aspects of the present disclosure, a kit is provided, the kit comprising (a) a set of universal PCR adapters and (b) one or methylation sensitive restriction enzymes. In some aspects, the kit further comprises (c) a set of barcoded PCR primers complementary to at least a portion of the universal PCR adapters.
[0083] In various aspects, the one or more methylation sensitive restriction enzymes comprise restriction enzymes blocked by the methylated motifs in one or more target
prokaryotic organisms, and wherein comprise but not limited to DpnII, XapI, Apol, BspCNI, BseMII, Mill, BstX2I, BstYI, Mill, XhoII, SfaNI, AflII, Acul, Bpml, BpuEI, PstI or a combination of any thereof.
[0084] In further aspects, each of the universal PCR adaptors may comprise an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, wherein the hybridized region and wherein the hybridized region does not comprise a non-methylated motif targeted by the one or more restriction enzyme.
[0085] In some aspects, each universal PCR adapter in the kit further comprising a *T at the 3’ end of the upper strand oligo, wherein *T comprises a phosphorothioated thiamine base having a phosphorothioate nucleotide modification, wherein the phosphorothioate modification renders an inter-nucleotide linkage formed by the Y shaped adaptor resistant to nuclease degradation. In further aspects the universal adaptor can be phosphorylated at the 5’ end, particularly on the lower strand oligonucleotide. Other modifications to prevent nuclease degradation are known in the art and are contemplated herein. For example, the upper strand oligo can have a nucleotide sequence comprising any one of SEQ ID NOs: 3-18 and/or the lower strand oligo can have a nucleotide sequence comprising any one of SEQ ID NOs: 19-34. As another example, the upper strand oligo can have a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo can have a nucleotide sequence comprising SEQ ID NO: 36. As another example, the upper strand oligo can have a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo can have a nucleotide sequence comprising SEQ ID NO: 38.
[0086] In various aspects, the set of barcoded PCR primers may target any nucleic acid of the upper or lower oligo of the universal adaptor. So, for example, the barcoded PCR primer may target a nucleic acid sequence comprising at least a portion of a nucleic acid sequence comprising or consisting of: SEQ ID NO: 1 or 2. As used herein, the term “target” means that the primer hybridizes to a nucleic acid comprising the nucleic acid target sequence to enable PCR amplification. In some aspects the forward primer hybridizes to the sense strand of the target sequence comprising upper oligo of universal adaptor (i.e., any one of SEQ ID NOs: 3-
18). In other aspects, the reverse primer hybridizes to the anti-sense strand of the target sequence comprising lower oligo of universal adaptor (i.e., any one of SEQ ID NOs: 19-34).
[0087] Kits may optionally provide additional components such as buffers and interpretive information. Normally, the kit includes a container and a label or package insert(s) on or associated with the container. In some embodiments, the invention provides articles of manufacture comprising contents of the kits described above.
[0088] Having described several embodiments, it will be recognized by those skilled in the art that various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the present inventive concept. Additionally, a number of well- known processes and elements have not been described in order to avoid unnecessarily obscuring the present inventive concept. Accordingly, this description should not be taken as limiting the scope of the present inventive concept.
[0089] Those skilled in the art will appreciate that the presently disclosed embodiments teach by way of example and not by limitation. Therefore, the matter contained in this description or shown in the accompanying drawings should be interpreted as illustrative and not in a limiting sense. The following claims are intended to cover all generic and specific features described herein, as well as all statements of the scope of the method and assemblies, w hich, as a matter of language, might be said to fall there between.
EXAMPLES
[0090] The following examples are included to demonstrate preferred embodiments of the disclosure. It should be appreciated by those of skill in the art that the techniques disclosed in the examples that follow' represent techniques discovered by the inventor to function well in the practice of the present disclosure, and thus can be considered to constitute preferred modes for its practice. However, those of skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific embodiments which are disclosed and still obtain a like or similar result without departing from the spirit and scope of the present disclosure.
Overview of Examples
[0091] Metagenomics have enabled the comprehensive and culture-independent study of microbiomes. However, for many applications, only certain taxa are of interest in a microbiome sample, for example, pathogenic bacteria, beneficial microbes, or taxa of low relative abundance. It is highly desirable that a method can effectively enrich bacterial taxa of interest directly from the microbiome. To address this critical need. mEnrich-seq, a method that can enrich taxa of interest from metagenomic DNA before sequencing, was developed. The core idea is to exploit the natural self vs. non-self genome differentiation provided by bacterial DNA methylation and rationally choose methylation-sensitive restriction enzymes (REs), individually or in combination, to deplete host DNA and most background microbial DNA while enriching bacterial taxa of interest. This core idea is integrated with library preparation procedures in a way that only non-digested DNA libraries are sequenced. In-depth evaluations of mEnrich-seq were performed using both synthetic and real microbiome samples. The use of mEnrich-seq is demonstrated in several applications to enrich (up to 117-fold) genomic DNA of pathogenic or beneficial bacteria from urine and fecal samples, including several species that are hard to culture or of low abundance. The broad applicability of mEnrich-seq was also evaluated and it was found that 3130 (68.03%) of the 4,601 strains with mapped methylomes to date can be targeted by at least one commercially available RE, representing 54.78% of the species examined in this analysis. mEnrich-seq provides microbiome researchers with a versatile and cost-effective approach for selective sequencing of diverse taxa of interest directly from the microbiome.
Example 1 - Introduction to Examples
[0092] Microbiomes play important roles in human health, diseases and drug responses. Many studies have reported that the dysbiosis of microbiota, or the loss of microbial diversity, is not only associated with a higher risk of proliferating pathogenic microbes but is also linked to many complex diseases and patients’ response to treatments. In many applications, researchers hope to specifically examine certain bacterial taxa of interest in a microbiome sample rather than sequencing all the species. For example, a few enteric pathogens are of particular interest in infectious diseases, while certain commensal bacteria are the target of discovery' in terms of microbe-host interaction, immune homeostasis or biosynthetic gene clusters. Although culture can be used to isolate specific bacteria, it is time- and resourceconsuming, hard to scale and sometimes very challenging for certain bacterial taxa.
[0093] Major challenges arise in the sequencing of specific bacterial taxa of interest from complex microbiome samples, mostly stemming from the presence of bacterial species with complex genomes, the highly skewed abundance distribution, mixed with viruses, fungi and host cells. The sequencing throughput is largely consumed by abundant microbes while low- abundant ones cannot be sequenced and studied in a cost-effective manner. A few technologies have been developed to address this challenge. Host genomic DNA (gDNA) can be eliminated using chemicals such as saponin to selectively lyse mammalian cells, which enables the efficient sequencing of pathogens in biospecimens that are dominated by host gDNA (for example, upper respiratory tract infection). Host genetic material can also be filtered using DNA binding proteins that specifically recognize 5-methylcytosine (5mC), abundant in mammalian genomes.
[0094] Flow cytometry has also been used to segregate complex microbiomes into small subsets, which can be sequenced and computationally analyzed separately. However, it is challenging for these approaches to enrich specific taxa of interest, taxa with low abundance or resolve bacterial species and strains with highly similar genomes. Adaptive sampling in the Nanopore platform can examine the first few hundred base pairs, during real-time sequencing, against a collection of reference sequences to reject the DNA molecules that meet certain criteria, for example, host DNA or a highly abundant bacterium for which enough data have been already collected. This strategy by adaptive sampling can reduce further consumption of sequencing yield by highly abundant species; however, its enrichment efficiency for specific taxa of interest is relatively modest, and insufficient to differentiate between bacterial taxa with highly similar genomes in complex microbiome samples.
[0095] To address the need for effective enrichment of specific bacterial taxa of interest directly from metagenomic DNA, a method is described herein that takes advantage of bacterial DNA methylation in microbiomes, namely bacterial epigenomes. In the bacterial kingdom, there are three main forms of DNA methylation: N6-methyladenine (6mA), N4-methylcytosine (4mC) and 5mC, which are catalyzed by methyltransferases (MTases) that apply methyl groups in a highly sequence-specific manner, that is. some sequence motifs of a genome are nearly 100% methylated, while most of the genome is not methylated. The bacterial methylome has three fundamental properties. First, DNA methylation is present in nearly all bacteria (more than 95%). Second, all genetic contents (chromosomes, plasmids) in a single strain share the same set of methylation motifs, and most methylation sites are highly stable over time and
conditions. Third, since the presence or absence of an MTase is mainly driven by horizontal gene transfer, the methylation motifs are highly variable between different species and different strains of the same species. Based on these three properties, bacterial DNA methylation naturally differentiates self from nonself DNA, which serves as the foundation of restrictionmodification systems, and has been exploited as natural epigenetic barcodes to group highly similar species and strains in metagenomic analyses, namely methylation binning.
[0096] In this work, method exploits bacterial DNA methylation motifs to differentiate bacterial taxa from each other before sequencing to enrich bacteria of interest and deplete background DNA. The core idea is to rationally choose individual, or combinations of, methylation-sensitive restriction enzymes (REs) that digest host genome and background microbial DNA at nonmethylated sequence recognition sites to eliminate the vast background of metagenomic DNA while leaving the gDNA from bacterial taxa of interest intact. A similar strategy was tested on mock microbiome in a recent study, but here this core idea is integrated with library preparation procedures that consider a number of practical factors to achieve efficient allocation of sequencing throughput to bacterial taxa of interest. The following examples first describe the method design and the evaluation, and then provide a few applications to demonstrate its broad utilities in microbiome research.
Example 2 - Materials and Methods used in Examples 3-9
Cell Culture
[0097] E. coli K-12 MG1655 was obtained from American Type Culture Collection (ATCC, cat# 47076, VA. USA) and cultured with Miller Luria Broth (LB) medium, and on LB agar plates (Thermo Fisher Scientific, cat# BP1426-500, and cat# BP1425-500, respectively, MA, USA) at 37°C based on ATCC recommended protocols.
[0098] N. gonorrhoeas MS 11 was obtained from ATCC (cat# BAA- 1833) and cultured based on protocols in Dillard, 2011 (Dillard, J. P. Genetic manipulation of Neisseria gonorrhoeae. Curr Protoc Microbiol 23, 4A-2 (2011)). Briefly, cells were grown in Gonococcal Base medium liquid (GCBL) with Kellogg’s supplements I&II, and on Chocolate II Agar plates (BD, cat# 22126, MD, USA) at 37°C in 5%CO2 for 24 h.
[0099] Human Lymphoblastoid Cell Line (LCL; Coriell Institute for Medical Research (Coriell), cat# GM 24149, NJ, USA) suspension cell cultures were grown in RPMI-1640 media (GIBCO cat# 11875093, MA, USA) supplemented with 15% Fetal Bovine Serum (GIBCO
cat# A3160402) at 37°C in 5% CO2. Cells were split at 1 :3 to 1:5 ratios when reaching confluency, and were monitored periodically by the My coAlert Mycoplasma detection assay (Lonza, cat# LT07-118, NJ, USA) for mycoplasma. After 6 passages, cells were harvested for DNA isolation.
Sample Collection
[0100] Deidentified Urinary tract infection (UTI) positive urine samples were obtained after routine cultivation and susceptibility testing performed at the Clinical Microbiology Laboratory' (Icahn School of Medicine at Mount Sinai, NY, USA). No patient identifiable information was collected, only the identified pathogen(s) were provided. Patients provided informed consent to Mount Sinai at collection and further informed consent was not required.
[0101] Four infant fecal samples (GUT-1-4) for mEnrich-seq with a RE targeting RAATTY (A. muciniphila) was collected from an individual subject from a previous study with Institutional Review Board approval. No personal identifiable information was used in the scope of this work. The matched A. muciniphila strains were isolated and cultured by using the established methods by Becken et al. (Genoty pic and phenotypic diversity among human isolates of Akkermansia muciniphila mBio 12, e00478-21 (2021). incorporated herein by reference in its entirety7).
[0102] The adult fecal sample (GUT-5 and GUT-6) used for mEnrich-seq with de novo discovered methylation motifs was from an individual subject from a previous study with Institutional Review Board approval. Written consent was obtained from all individuals recruited in the study using a protocol approved by the Mount Sinai Institutional Review Board (HS no. 11 -01669).
DNA extraction
[0103] DNA from the human LCL cell line GM24149. E. coli MG1655, N. gonorrhoeae MS11, and urine samples was isolated with the DNeasy Blood and Tissue Kit (Qiagen, cat# 69506, MD, USA) based upon the manufacturer’s instructions, including the optional RNase A treatment provided with the kit. For the human cell line, 2xl06 cells were collected from cell culture by centrifugation for 5 min at 300 x g. The cell pellets were resuspended in 180 pl PBS and 20 pl Proteinase K. After adding 200 pl Buffer AL, the mixture was incubated at 56°C for 30 min. The rest of extraction steps were performed according to manufacturer’s instructions. For E. coli MG1655, and A. gonorrhoeae MS 11, 109 cells were harvested by centrifuging 10
min at 5000 x g. The cell pellets were resuspended in 180 pl Buffer ATL, and then lysed with 200 pl Buffer AL and 20 pl Proteinase K for 30 min at 56°C. The rest of the extraction steps were followed per manufacturer's instructions.
[0104] For urine samples, each 2 ml aliquot of urine was first centrifuged at 5000 x g and 4°C for 30 min, then the pellets were collected in a new tube. Next, the urine pellets were resuspended with 180 pl of the bacterial enzymatic lysis buffer (20 mM Tris pH 8.0, 2mM EDTA, 1% SDS and 20 mg/ml lysozy me) and incubated at 37°C for 30 min. Following the enzymatic lysis, 20 pl of Proteinase K and 200 pl Buffer AL was added to the mixture, and it was incubated at 56°C for 30 min. The rest of the extraction steps were followed per manufacturer instructions with one modification, i.e., the DNA was eluted with 50 pl AE in the final step to concentrate the DNA.
[0105] Fecal DNA was extracted with a QIAamp PowerFecal Pro DNA Kit (Qiagen, 51804) according to the manufacturer instructions. Briefly, 100 mg of fecal sample, 750 pl of PowerBead Solution and 60 pl of Solution Cl were added to the Bead Tube (provided). After incubation at 65°C for 10 min. the sample was homogenized with the TissueLyser II machine (Qiagen) for 10 min at 25 Hz speed (5 min each side). The homogenized samples were centrifuged to collect the supernatant at 15,000 xg for 1 min. The subsequent column-based extraction steps were performed as described in the QIAamp PowerFecal Pro DNA Kit protocol.
E. coli isolates from UTI positive samples
[0106] The isolates were cultured from the urine samples with confirmed E. coli cases. Briefly, 10 pl of urine was spread by disposable inoculating loops (Thermo Fisher Scientific, cat# 22170201) quantitatively onto BD BBL MacConkey agar plates (Becton, Dickinson, and Company (BD), NJ, USA) and incubated for 24 h at 37°C aerobically. For each urine sample, a representative colony of E. coli was collected for DNA purification.
Preparation of the tw o mock samples (Mock-1 & Mock-2) for the initial evaluation of mEnrich-seq
[0107] Purified human and bacterial DNA were quantified for concentration with the Qubit dsDNA HS assay (Life Technologies, Q33231, MA, USA), and pooled by adding 94% of human DNA, 5% of N. gonorrhoeae DNA and 1% of E. coli DNA by7 mass to create 500 ng of the mock samples. The final ratio of E. coli, N. gonorrhoeae and human gDNA was quantified
by metagenomic sequencing.
The mock microbiome sample with five species (Mock-3)
[0108] Purified genomic DNA of Akkermansia muciniphila, Anaerobutyricum hallii, Bifidobacterium longum, Clostridium beijerinckii, and Clostridium butyricum was quantified for concentration with the Qubit dsDNA HS assay (Life Technologies, Q33231, MA, USA), and pooled by mixing -3% of A. muciniphila. -5% of A hallii. -10% of B. longum. -27% of C. but victim and -55% of the C. beijerinckii.
Fecal samples with low-abundance A. muciniphila (Mock-4 & Mock-5)
[0109] An adult fecal (GUT-7) with no A. muciniphila (confirmed by shotgun sequencing and culture) was used as background (from a previous study with Institutional Review Board approval4) and was whole genome amplified with REPLI-g Kits (Qiagen, cat# 150023) for 16 hours, and named as “fecal -WGA”, and was purified with 0.8 X AMPure Beads. A. muciniphila genomic DNA was mixed with fecal-WGA at 1:500 to make Mock-4, and at 1: 1000 ratio to make Mock-5.
Detailed methods for mEnrich-seq
[0110] Libraries for mEnrich-seq were prepared using the Oxford Nanopore Technologies (ONT) PCR Barcoding Expansion Pack 1-96 (Catalog No. EXP-PBC096), Ligation Sequencing Kit (Catalog No. SQK-LSK109), and 100-1,000 ng of genomic DNA as staring material.
[0111] Step 1. DNA shearing. Shearing was conducted to establish a uniform size of DNA inserts for better ligation efficiency, to maintain highly consistent molanty estimations for downstream loading into flow cells, and to reduce pore clogging. gDNA was sheared to around lOkb with g-TUBEs (Covaris, cat.520079, MA, USA) using an Eppendorf 5424 centrifuge at 5000 rpm for 3 min at room temperature. The g-TUBE was then inverted and centrifuged for another 3 min to collect the sheared DNA. The DNA was concentrated using 0.5x volume Ampure XP beads and eluted with 48 ul nuclease-free (NF) water.
[0112] Step 2. End Repair and Barcode Adapter Ligation. The fragmented DNA was applied to ’End-prep" and “Ligation of Barcode Adapter” steps according to the manufacturer’s instructions. The “End-prep” was performed with 48 pl of sheared DNA in the reaction mixture using NEBNext Ultra II End Repair/dA-Tailing Module and NEBNext FFPE
DNA Repair Mix (New England Biolabs (NEB), cat# E7546, M6630, respectively, MA, USA) with incubation at 20°C for 5 min and 65°C for 5 min. After purification with lx volume of AMPure XP beads, the end-repaired DNA was eluted in 30 pl nuclease-free water. Next, the end-repaired DNA was ligated with 20 pl Barcode Adapter (ONT, cat# EXP-PBC096, blue cap) and Blunt/TA Ligase Master Mix (NEB, cat# M0367) for lOmin at room temperature, which attached the universal PCR handle to all of the DNA molecules. The Barcode-Adapter- ligated-sample was then purified with lx volume of Ampure XP beads and eluted in 26 pl NF water. After the ‘'Ligation of Barcode Adapter” step, 20 ng of the PCR Adapter ligated DNA was split into another tube as sample “Input”, while the rest of the sample was used for the restriction enzyme digestion step of mEnrich-seq.
[0113] Step 3. Restriction enzyme (RE) digestion. RE digestion for mEnrich-seq was conducted following the manufacturer’s instructions. For DpnII digestion, 10 U of enzyme (NEB, cat # R0543) was used in a 50 pl-reaction containing 5 pl of NEBuffer r3.1 (10X, NEB #B7203), 26 pl of adapter-ligated DNA and NF water to 50 pl. The reaction system was incubated at 37 °C for 30 min. For BspCNI (NEB, cat# R0624S), 10 U of the enzyme was used in the 50 pl -reach on containing 5 pl of rCutSmart Buffer (10X, NEB # B6004S) for 30 min- incubation at 37 °C. 10 U of XapI (Fisher Scientific, ER1381) was added in 50 pl-reaction system supplied with 5 pl Buffer Tango (10X. Fisher Scientific. BY5) and incubated for 30 min at 37 °C. When using SfaNI (NEB, cat# R0172S) for digestion, 10 U of enzy me was used in 50 pl reaction system with 5 pl of NEBuffer r3.1, and incubated for 30 min at 37 °C. MUI (Takara, cat # 1070A) was used with 10 U in the 50 pl-reaction system containing buffer L, and incubated at 37 °C for 30 min.
[0114] When using multiple enzymes, it’s critical to first check whether the enzymes work in the same optimal buffer and temperature as recommended by the manufacturer. In this test. 6 enzymes were used in combination, SfaNI, Aflll (NEB, cat# R0520S), Acul (NEB, cat# R0641S), Bpml (NEB, cat# R0565S), Bpuel (NEB, cat# R0633S) and Pstl-HF (NEB, cat# R3140S). Among them, only SfaNI works in NEBuffer r3. 1 at 100% efficiency. To make sure all enzymes work optimally, the DNA was firstly digested with 20 U SfaNI in NEBuffer r3.1 for 30 min at 37 °C, then the digestion system was purified with 1.8X AMpure beads. The eluted DNA (20 pl) was further digested for 30 min at 37 °C with a 5 pl-mixture of Aflll, Acul, Bpml, Bpuel and Pstl-HF (1 pl of each).
[0115] Step 4. Gel purification. The RE-digested products were loaded into a 1.5% agarose
gel (Thermo Scientific, cat# R0492, MA, USA) pre-stained with SYBR Safe DNA gel stain (Invitrogen, cat# S33102, MA USA) for gel-based size selection. DNA ladder (NEB, cat#N3238S) was used as a comparison for size estimation and each gel was visualized under UV light on a ChemiDoc XRS+ System (Bio-Rad, CA, USA). The DNA band at 5 to lOkb size (as the DNA ladder indicates) was collected for extraction with a NucleoSpin Gel and PCR Clean-up purification kit following the manufacturer's instructions (Macherey-Nagel, cat# 740609.250, Germany), and was eluted with 15 pl NF water.
[0116] Step 5. PCR amplification. To add the barcode to each sample, the gel-purified DNA was amplified with the barcoded primers from the barcoding expansion pack mentioned above (EXP-PBC096, white caps, ONT). For PrimeSTAR GXL Premix (2x, Takara Bio, cat# R051A, Japan), a 50 pl -reaction system was prepared as following: 25 pl of Premix, 1 pl of barcoded primer (lOuM, white tube), 15 pl purified DNA and 9 pl water. The PCR reactions were pre-heated at 98°C for 1 min followed by 15-18 cycles of 98°C for 15 s, annealing 62°C for 15 s and extension at 68°C for 8 min, with a final extension at 68°C for 10 min on the ABI Veriti thermal cycler. The PCR products were purified with 0.5x volume of Ampure XP beads. Before moving to library prep, the purified PCR products were checked with an agarose gel for a distinct band at 5-10 kb in size (FIGS. 9A-9B, 10A-10B, 11A-11B, and 12A-12B).
[0117] Step 6. Library prep and sequencing. The purified barcoded PCR products were pooled together for "DNA repair’ and “end-prep” step. Briefly. 47 pl of pooled PCR products were incubated at 20°C for 15min and 65°C for 15min, in a 60-reaction system containing 3.5 pl of NEBNext FFPE DNA Repair Buffer, 2 pl of NEBNext FFPE DNA Repair Mix, Ultra II End-prep reaction buffer 3.5 pl and Ultra II End-prep enzyme mix 3 pl. The purified DNA was ligated to Adapter Mix (ONT, SQK-LSK109) in the presence of Ligation Buffer (ONT, SQK- LSK109) and NEBNext Quick T4 Ligase (NEB, cat# E6056) for 30 min at room temperature. After beads purification, the adapter ligated DNA is ready for sequencing on ONT flow cells.
Compatibility between mEnrich-seq and other sequencing platforms
[0118] Although Nanopore sequencing was used in these examples, mEnrich-seq is compatible with other sequencing platforms, because the barcoded PCR products from 'step 5’ in the detailed methods are essentially double-stranded DNA. For PacBio platform users, after purification with AMPure PB beads (Pacific Biosciences, #100-265-900). the PCR products from “Step 5” are ready for library preparation following the “Preparing multiplexed amplicon libraries using SMRTbell prep kit 3.0” (PN 102-359-000 REV 02 SEP2022). It’s recommended
to use the barcode kit by PacBio to achieve the best sequencing result. The reads from mEnrich- seq are also computational applicable for data analysis by trimming the barcode adapter sequences (see “Guidance for broad applications of mEnrich-seq”).
[0119] For the users of the Illumina platforms, as Illumina is a short-read platform, the barcoded PCR products from ‘Step 5’ will firstly be purified and adjusted to shorter insert size (e.g., by sonication). Then, the fragmented DNA molecules are ready for users to prepare the library7 with suitable kits compatible with various Illumina platforms. Compatible barcoding kits are required for a good sequencing result. The Barcode adapter sequence from the mEnrich-seq reads can be trimmed using computational methods for further analysis.
Guidance for broad applications of mEnrich-seq.
[0120] Barcode adapter. As the Barcode adapter ligation step is prior to the RE digestion during the library preparation, the DNA from targeted taxa will not be amplified by PCR if the barcode adapter is cut during the RE digestion step. Therefore, it’s critical to avoid a chosen RE cutting the sequence of the Barcode adapter or the junction between the adapter and metagenomic gDNA. It is suggested to check the sequence of the Barcode adapter (particularly the sequences that are reverse complementary to each other) for any possible RE motif site. If there’s a recognition sequence for the chosen RE, the bases can be substituted to avoid the cut. The sequences are as follows: top sequence: SEQ ID NO: 35 and bottom sequence: SEQ ID NO: 36. The top and bottom sequence are added to the annealing buffer (10 mM Tris, pH 7.5, 50 mM NaCl and 1 mM EDTA) to a concentration of 10 pM and annealed with the program: 95 °C for 2 min, -0. 1 C/sec for each of 800 s (~13 min) and hold at 4°C. The barcoded adapter is now ready to use and compatible with the original protocol.
[0121] Selection of RE. The RE can be found through the Enzyme Finder (enzymefmder.neb.com) based on the chosen target motifs. For the chosen RE, it’s important to check its methylation sensitivity result provided by the manufacturer in REBASE (rebase.neb.com/rebase/rebase), which is commonly tested with modified oligos, and even better to re-confirm the sensitivity upon purchase.
Library preparation for the standard metagenomic sequencing (Input)
[0122] For the mock, urine and the infant fecal samples, libraries of the standard metagenomic sequencing (Input) for comparison with mEnrich-seq was prepared with the 20 ng DNA split from Barcode Adapter ligation step (at the end of ‘Step 2’ in mEnrich-seq). The
Input sample was first amplified with barcoded primer. The PCR master mix consisted of 1 pl PCR barcode (10 pM, ONT), 25 pl PrimeSTAR GXL Premix (2x, Takara Bio, cat# R051A, Japan). 20 ng DNA and nuclease-free water to a total volume of 50 pl. The PCR programs were as follows: pre-incubation at 95°C for 3 min, amplification for 14 cycles at 95°C for 15 s, 62°C for 15 s, and extension at 65°C for 10 min, with a final extension at 65°C for 10 min. The PCR products were checked with an agarose gel for successful amplification (bands at ~5 to 10 kb, FIGS. 9A-9B, 10A-10B, 11A-11B. and 12A-12B). After 0.5x volume bead purification, the purified barcoded PCR products were used for “repair and end-prep” step in the ONT protocol. Following the “Adapter ligation and clean-up” step, the library was quantified by Qubit and ready for sequencing.
Library preparation for the isolates and adult fecal microbiome sample (GUT-5)
[0123] For the E. coll isolates from urine samples, the library preparation protocol from Oxford Nanopore Technologies pic (ONT), “Native barcoding genomic DNA” (with EXP- NBD104, EXP-NBD114, and SQK-LSK109, ONT, NY, USA) was followed with one modification for high molecular weight shearing. Briefly, 5 pg purified gDNA of E. coli isolates was diluted in 250 pl EB buffer (10 mM Tris-Cl, pH8.0), then the DNA was sheared with a 29 G needle for 20 passes. This led to DNA samples with an N50 of approximately 20 kb. The sheared DNA was purified with 0.5x volume of AMPure PB beads (Pacific Biosciences of California. Inc.), and then was eluted with 51 pl nuclease-free water. After quantification by Qubit, 1 pg of sheared DNA w as used for further library preparation.
[0124] The sequencing data of A. muciniphila isolates from the infant fecal samples were collected from another study. In brief, the SMRTbell libraries of A. muciniphila strains were prepared according to the manufacturer’s instructions, and then sequenced on the Sequel instrument. 6mA methylation motif discovery' from A. muciniphila isolates were performed with the build-in pipeline pb basemods from SMRTlink vlO.O.
[0125] The data of adult fecal microbiome sample were generated from another study. Briefly, 1000 ng of the DNA was used for library preparation using SQK-LSK109 kit from ONT, following the Native Barcoding genomic DNA protocol (NBE_9065_vl09_revZ_14Aug2019). Briefly, the purified DNA was sheared with g-TUBE, and then was added to “End-prep and FFPE repair” reaction system and incubated at 20°C for 5 min and 65°C for 5 min. Following the DNA purification step, the eluted DNA was added to the sequencing “Adapter ligation” reaction system, containing Ligation Buffer (LNB),
NEBNext Quick T4 DNA Ligase and Adapter Mix. After the purification using 0.5X beads, the prepared library was quantified with Qubit, and 40 finol of the library was loaded onto a MinlON cell by manufacturer’s instructions.
Validation of let (A) in urine sample UTI-1
[0126] Two pairs of primers specific to tet(A) sequence were designed using Primer- BLAST as follows: te^-Fl: 5’- GAAACCCAACAGACCCCTGA-3’ (SEQ ID NO: 39), and tet(A)-R 5’- CCACCCGTTCCACGTTGTTA-3’(SEQ ID NO: 40); tet(A)-F2: 5’- TGAAACCCAACAGACCCCTG-3’(SEQ ID NO: 41), and tet(A)-R2 5’- CCCTGACGTTCCTCATCCAC-3’ (SEQ ID NO: 42) (Integrated DNA Technologies, IA, USA; 25 nmol, standard desalting). The tet(A) was amplified directly from the gDNA of UTI- 1 using the Q5 High-Fidelity 2X PCR Master Mix (NEB, cat #M0492S) with the following conditions: 50 ng of gDNA per reaction; 1 cycle at 95°C for 30s; 33 cycles of denature 95°C for 15s, annealing at 60 °C for 15s, extension at 72 °C for Imin; a final extension cycle at 72 °C for 5 min. PCR products were checked for amplification specificity on an agarose gel. Main bands of the PCR products were gel purified with NucleoSpin Gel and PCR Clean-up purification kit. The purified PCR products were sequenced by the Sanger method using tet(A)- F1 and tet(A)-V2 as forward primers._The two PCR products are shown in Table 3 below. Nucleotides denoted with an “n” in SEQ ID NO: 43 and 44 were undetermined by the Sanger method and may be any base (a,t,c, or g).
Table 3
Evaluation of adapter ligation before and after RE digestion
[0127] For the evaluation of adapter ligation after RE digestion (“Ligation-after-RE”), 5- 10 pg of Mock-3 was used as input and digested with XapI in a 50 pL reaction system at 37°C for 30 min. Next, the DNA larger than 10 kb was selected by gel purification and eluted with 15 pL nuclease-free water. Then the DNA was further processed for library preparation following the ‘’Library preparation for the standard metagenomic sequencing (Input)” in the Methods part of the manuscript. The mEnrich-seq library' (500- 1000 ng as input) was generated for evaluation of the adapter ligation before RE digestion (“Ligation-before-RE”). The ONT data of Ligation-before-RE or Ligation-after-RE were each mapped to the reference using minimap2 (v2.24-rl 122) with -x map-ont and reads mapped to each species were counted.
Nanopore sequencing and base calling
[0128] Sequencing was primarily performed with both MinlON flow cell R9.4.1 and Flongle flow cell R9.4.1 (both from ONT), and ONT MinKNOW software (v21) was used for sequencing raw data collection. For the MinlON flow cell, 50 fmol of library was loaded onto flow cell according to manufacturer's instructions and the MinlON was run for 72 h to collect the raw data. For the Flongle flow cell, 20 fmol of library was loaded and run for up to 24 h to collect the raw data. All Nanopore sequencing reads were base called using guppy (v3.2.4) with the SUP model dna_r9.4. l_450bps_sup.cfg.
Evaluation of mEnrich-seq of the two mock samples
[0129] Fold enrichment and completeness of E. coli. The genome references of E. coll KI 2 MG1655 (RefSeq accession CP014225.1), N. gonorrhoeae MS1 1 (RefSeq accession CP003909.1), and human (GRCh38) were combined as the mock reference. Repeat annotation of human genome was obtained from RepeatMasker database (www.repeatmasker.org/genomes/hg38/RepeatMasker-rm405-db20140131/hg38.fa.out.gz).
The fold enrichment of E. coli was calculated based on E. coh reads from mEnrich-seq and Input samples that were mapped to the mock reference using minimap2 (v2.24-rl 122) with -x map-ont. The fold enrichment is defined as: target strain abundance (%) in enrichment sequencing versus target strain abundance (%) in the input sample. For genome completeness, the coverage of the E. coli was calculated using read depths across lOObp bins across the E. coli genome. For Circos plots. Nanopore reads from paired Input and mEnrich-seq samples were normalized to the same yield by random subsampling using seqtk (vl.2-r94) with 2-pass mode. Reads after normalization were mapped to the mock reference using minimap2 (v2.24- rl 122) with -x map-ont. Sliding windows with intervals of 2.5kb and step sizes of 1.25kb were defined across the E. coli genome using bedtools (v2.29.2) with -w 2500 -s 1250. Reads mapped to defined windows were counted using bedtools multicov. Circos plots were drawn using circos (vO.69-6). To ease visualization, maximum depth of each circos plot was capped at the 95% depth quantile.
[0130] Motif analyses in N. gonorrhoeae and human. Nanopore reads from the Input and mEnrich-seq samples were mapped to the mock reference using minimap2 (v2.24-rl 122) with -x map-ont. Reads of N. gonorrhoeae and human mEnrich-seq were extracted using samtools (vl . 15. 1). To reduce the possibility of low-quality mis-mapped reads from other species, nonprimary and supplementary' alignments were not included for the motif analysis using
samtools (vl.15.1) with view -F 2308. Considering the moderate Nanopore sequencing error rate, analysis of motif (GATC and GGTGATC) frequencies and overlapping of GATC motif and human genomic repeats were counted based on mapped reference segment of a read instead of the read itself.
Evaluation of mEnrich-seq of the urine samples
[0131] Fold enrichment and completeness o fE. coli. The references of E. coli isolates were assembled using Flye (v2.8.1-bl684) with -nano-raw -genome-size 5m. E. coli reads were extracted by mapping the Input and mEnrich-seq Nanopore reads to the references of the E. coli isolates using minimap2 (v2.24-rl 122) with -x map-ont. Non-A. coli reads were then classified using Kraken2 (v2.0.8-beta) with the k2_pluspf_20200919 database. For genome completeness, the coverage of the E. coli was calculated using read depths across lOObp bins across the E. coli genome assembled from the isolates. For Circos plots, Nanopore reads from paired Input and mEnrich-seq samples were normalized to the same yield by random subsampling using seqtk (vl ,2-r94) with 2-pass mode. Reads after normalization were mapped to the references of the E. coli isolates using minimap2 (v2.24-rl 122) with -x map-ont. To draw the circos plots, sliding windows with intervals of 2.5kb and step sizes of 1.25kb were defined across the E. coli references using bedtools (v2.29.2) with -w 2500 -s 1250. Reads mapped to defined windows were counted using bedtools multicov. To exclude the low-quality mapped reads from redundant genomic locations, e.g., high similar reads between multiple copies of ribosomal RNA (rRNA), read depths were capped at the 95% depth quantile for each chromosome (that is, the longest circled contig roughly 4-5 Mb) and plasmid (that is, short circularized contigs) for Circos plots drawing using circos (vO.69-6).
Evaluation of mEnrich-seq of the infant fecal samples
[0132] Fold enrichment and completeness of A. muciniphila. Reads from the Input and mEnrich-seq were assessed using Kraken2 (v2.0.8-beta) classification with the k2_pluspf_20200919 database. To check the completeness of the A. muciniphila, nanopore reads were mapped to the A. muciniphila isolate reference and calculated with reads depth in lOObp bins across the genome. Before Circos plots, Nanopore reads from paired Input and mEnrich-seq samples were normalized to the same yields by random subsampling using seqtk (vl.2-r94) with 2-pass mode. Reads after normalization were mapped to the reference of A. muciniphila using minimap2 (v2.24-rl 122) with -x map-ont. To draw the Circos plots, sliding windows with intervals of 2.5kb and step sizes of 1.25kb were defined across the A.
muciniphila isolate reference using bedtools (v2.29.2) with -w 2500 -s 1250. Reads mapped to defined windows were counted using bedtools multicov. To exclude the low-quality mapped reads from redundant genomic locations, e.g., high similar reads between multiple copies of rRNA, read depths were capped at the 95% depth quantile of the whole genome for Circos plots drawing using circos (vO.69-6).
[0133] RAATTY motif analysis. For the species used for motif analysis, all genome references except A. muciniphila were retrieved fromNCBI, including C. aerofaciens (RefSeq genome sequence GCF_002736145.1), E. lenta (GCF_021378605.1), B. bifidum (GCF 000273525.1), B. longum (GCF_000196555.1), B. breve (GCF_001281425.1). The matched A. muciniphila reference was obtained from the isolated strain as previously described. Nanopore reads from the mEnrich-seq samples were mapped to the reference of each species using minimap2 (v2.24-rl 122) with -x map-ont. To reduce the possibility of low- quality mis-mapped reads from other species, non-primary and supplementary alignments were not included in motif analysis using samtools (vl.15.1) with view -F 2308. Considering the moderate sequencing error rate of Nanopore sequencing, analysis of RAATTY motif frequencies in mEnrich-seq were counted based on mapped reference segment of a read instead of the read itself. The genome frequency for RAATTY motif on each species was counted with custom PERL script.
De novo motif discovery from the adult fecal microbiome
[0134] De novo methylation motif discover}' from the fecal microbiome DNA sample was based on PacBio sequencing data (SEQUEL II). Briefly, the DNA SMRTbell library' was performed according to the manufacturer’s instructions, then the library' was annealed to sequencing primer together with the Sequel 2. 1 polymerase, followed by loading to the 8M sequencing chip at concentration of 125-175pM. 2.5 Gb CCS data were generated with 400k ccs reads and 6,341bp as median read length.
[0135] The raw SMRT reads pre-processing was performed with SMRTlink v8.0 (Pacific Biosciences). Raw multiplexed files were first de-multiplexed with lima. Circular Consensus reads (CCS reads) were generated using CCS software (v4.0.0) with --min-rq 0.99 based on subreads. CCS reads were assembled using metaFlye (v2.9) with -pacbio-hifi. The assembly was then binned and refined with the metaWRAP Binning and Bin refmement module10, which could combine the results from three metagenomic binning software - MaxBin2, metaBAT2, and CONCOCT. The completeness and the contamination level was
checked using CheckM (v.1.1.3). The metagenomic profile computing and graphics were performed by R (v.4.2. 1). For evaluation purpose, the genome assemblies from bacteria isolates from a previous study5 along with de novo assembled metagenomic contigs were used as the reference database for the following nanopore reads assignments and motif analyses.
[0136] The de novo methylation motif discovery was performed as described in Beaulaurier, J. et al. (Metagenomic binning and association of plasmids with bacterial host genomes using DNA methylation Nat. Biotechnol. 36, 61-69 (2018), which is incorporated herein by reference in its entirety). Briefly, the bins of species/strain level with low completeness (10% ~ 80%) were selected for further analysis. SMRT subreads were mapped to metagenomes using pbalign (vO.4.1) with default parameters and separated the alignments into each species/strain based on the binning results. IPD ratio were calculated with ipdSummary (v2.4.1) from the alignments for each species/strain. The de novo motifs were detected by using motifMaker (in SMRTlink v8.0 package) with default parameters. Only the motifs with detection percentage > 60% were kept for the following analysis (see Table 15 at the end of the Examples).
Evaluation of mEnrich-seq for the enrichment of low-abundant species from the adult fecal sample
[0137] Reads assignments. Nanopore reads w ere first mapped to the reference database described above using minimap2 (v2.24-rl 122) with -x map-ont. To reduce the possibility of low-quality mis-mapped reads from other species, non-primary and supplementary alignments were not included for reads assignment using samtools (vl.15.1) with view -F 2308. Other reads were then classified using Kraken2 (v2.0.8-beta) with the k2_pluspf_20200919 database.
[0138] Motif analysis. For each targeted genome, Nanopore reads from the mEnrich-seq samples were extracted from the above read assignment results. Considering the moderate sequencing error rate of Nanopore sequencing, analysis of motif frequency in mEnrich-seq were counted based on mapped reference segment of a read instead of the read itself. The genome frequency for a motif on specific species were counted with custom PERL script.
Evaluation of mEnrich-seq on metagenome assembly
[0139] To make sure that the final read yield originated from the targeted species was the same before evaluation, total reads by mEnrich-seq and standard metagenomic sequencing (‘input') were normalized by random subsampling using seqtk (v. l.2-r94) with 2-pass mode.
Reads after normalization were assembled with Flye (v.2.8.1-bl684) with -meta-nano-raw. The assemblies were then binned and refined with the metaWRAP Binning and Bin refinement module (v. 1.2). The completeness and the contamination level were measured using CheckM (v.1.1.3).
Assessing the broad applicability of mEnrich-seq
[0140] REBASE is the database constantly updated with bacterial methylomes mapped to date. For each bacterial methylome, a list of methylation motifs was reported (rebase.neb.com/rebase/rebase.html). All 4-mer to 6-mer methylation motifs were extracted and their corresponding restriction endonucleases (REs) from 4,644 bacterial methylomes (as of 7/18/2022). Among 4,644 PacBio organisms, the 1,358 most complete genomes were selected as representatives when multiple strains from the same species are present. The protein sequences from the 1,358 species level genomes were extracted to reconstruct a phylogenetic tree using PhyloPhlAn (v3.0). Visualization and annotation of the phylogenetic tree was performed by ITOL (itol.embl.de/).
Evaluation of mEnrich-seq in comparison with REMoDE
[0141] Library preparation. The REMoDE libraries were prepared as previously described by Enam et al., (Restriction Endonuclease-Based Modification-Dependent Enrichment (REMoDE) of DNA for Metagenomic Sequencing. Appl Environ Microbiol 89, e01670-22 (2023), which is incorporated herein by reference in its entirety ). Briefly, 100 ng of DNA was used for the reaction mixture of 40 pL containing 10X Tango buffer and 1 pL of XapI (Fisher Scientific, ER1381). The mixture was incubated at 37°C for 30 mm. 0.4 U of T5 exonuclease (NEB, #M0663) was added to each reaction mixture and incubated for 5 min at 37°C, after which the reaction w as immediately quenched with 8 pL 66 mM EDTA. The DNA w as purified using the NucleoSpin Gel and PCR Clean-up purification kit (Macherey-Nagel, cat# 740609.250, Germany). The DNA was eluted with 15 pL of nuclease-free water. For the library preparation of REMoDE with DNA repair (“REMoDE-repair’’), the DNA was first treated with the NEBNext FFPE DNA Repair Mix at 20°C for 15 min, purified with IX AMPure beads, before following the protocol described above (starting at XapI digestion). NGS libraries were prepared using Nextera XT (Illumina, cat # FC-131-1024) library preparation kit. The amplified libraries were resolved on an agarose gel and quantified with Qubit HS dsDNA kit. The libraries were pooled and sequenced on an Illumina MiSeq sequencer using a MiSeq reagent kit v2 (Illumina, cat# MS-102-2002; 300 cycles, paired-end).
[0142] Data analysis. For the short-read sequencing of the Mock-3 sample, the paired-end short-read data of Input, REMoDE and mEnrich were each mapped to the reference using bwa with default parameters, and then reads mapped to each species were counted. For the human infant fecal samples (‘'GUT-1” and “GUT-4”), matched standard short-read library without enrichment (“Input”) was also sequenced as the negative control. Paired-end short reads were classified using Kraken2 (v2.0.8-beta) with the k2_pluspf_20200919 database.
Evaluation of mEnrich-seq with long-read Tn5-based adapter tagmentation
[0143] Fecal samples with low-abundance A. muclnlphilci (Mock-4 &Mock-5) were prepared as described above. The long-read Tn5 library with mEnrich-seq (“mEnrich-D2”), was prepared with 1000 ng HMW genomic DNA as input. After digestion with XapI enzyme for 30 min at 37 °C, the mixture was resolved on a gel for isolation of the DNA fragments >10 kb. The gel purified DNA was quantified with Qubit for concentration and was used for further library preparation with the Rapid PCR Barcoding Kit (Oxford Nanopore, SQK-RPB004). 1 ng of DNA was diluted to 3 pL and incubated with 1 pL of Fragmentation Mix at 30° C for 1 min and then at 80° C for 1 min. The tagmented DNA was added to PCR mix containing 20 pL nuclease-free water, 25 pL LongAmp Taq 2X master mix, 1 pL Rapid Barcode Primers (RLB01-12A, at 10 pM), followed by 18 cycles of amplification (18 cycles of 95 °C for 15 s; 56 °C for 15 s; 65 °C for 6 min; 1 cycle of extension at 65 °C for 6 min). The PCR products were purified with 0.6X AMPure Beads and eluted with 10 pl elution buffer (10 mM Tris-HCl pH 8.0 with 50 mM NaCl). 50-100 ftnoles of library pool was ready for sequencing after ligation with 1 pL Rapid Adapter for 5 min at room temperature. For each sample, the original mEnrich-seq (“mEnrich-Dl”), the mEnrich-D2 and the standard ONT library without enrichment (“Input”) were sequenced and compared on the matched sample. Reads were then classified using Kraken2 (v2.0.8-beta) with the k2_pluspf_20200919 database.
Evaluation of mEnrich-seq with alternative adapters
[0144] Standard mEnrich-seq with alternative adapters was conducted on Mock-3 using the ONT platform as described in Methods. The alternative adapters are shown in Table 4 below and were designed such that certain bases (in lower case) are substituted to avoid the digestion by the 28 REs not compatible with the adapter described in the first submission. 5' phosphorylation is labeled as /5Phos/, and the phosphorothioated DNA base is labeled as “*N” in the sequence. For data analysis, reads were mapped to the reference using minimap2 (v2.24- rl 122) with -x map-ont and reads mapped to each species were counted.
Table 4
Example 3 - Methylation-guided enrichment for bacterial taxa of interest
[0145] The core idea of enriching the gDNA of bacterial taxa of interest from the microbiome is to distinguish gDNA from different taxa based on their distinct characteristics. Here a methylation-guided enrichment sequencing of bacterial taxa of interest from microbiomes is shown and named mEnrich-seq (in which ‘m’ stands for methylation and seq for sequencing, FIG. 1A). For a bacterial taxon of interest, based on its methylation motifs, a methylation-sensitive RE with the same recognition motif is chosen to cut the nonmethylated motif sites in most of the ‘background’ DNA (FIG. 1A). By contrast, the gDNA of the bacterial taxa of interest (including its chromosome and mobile genetic elements, for example, plasmids) is left intact, enriched over the background, which can be sequenced by various sequencing platforms (FIG. 1A). More specifically, g-TUBEs are first used to shear the DNA size to around 10 kb, which facilitate the second step where DNA is ligated to barcode adapters. In the third step, cognate restriction enzymes (RE) corresponding to the methylation motif in the taxon of interest are used to cut the background DNA libraries into smaller fragments while preserving the targeted DNA libraries as the long fragments. Because bacterial DNA methylation motifs usually have 2-8 nondegenerate bases (for example, GANTC has four nondegenerate bases; N = A, C, G or T), the choice of fragmentation size in the shearing step (generating roughly 10 kb fragments) increases the chance that most gDNA molecules have at least one nonmethylated RE site that can be cut: 5-10 kb fragments expected to have 20-40 sites of a 4-mer motif, 5-10 sites of a 5-mer motif and 1-2.5 sites of a 6-mer motif. In the fourth step, gel-based size selection is performed, which largely removes the digested shorter DNA fragments. In the fifth step, barcoded primers are used to anchor the PCR adapter to perform amplification of the long target DNA fragments that have been protected from RE digestion (i.e., those that are methylated at restriction sites). In the sixth and final step, the amplified
dsDNA is sequenced and analyzed using various metagenomic analysis tools. Because the DNA-to-adapter ligation step (step 2) occurs before RE cutting (step 3), only nondigested DNA libraries can be sequenced. This design minimizes allocation of sequencing throughput to background DNA fragments that are still long after RE digestion due to sparse distribution of the RE recognition motif(s) in certain genomes or genomic regions.
[0146] mEnrich-seq is versatile along three dimensions. First, while this work mainly used Nanopore sequencing, the mEnrich-seq protocol described above is compatible with all sequencing platforms (such as PacBio and Illumina sequencing platforms) by adjusting the insert size (after step 5) and chemistry according to different platforms (See detailed methods in Example 2, above) (FIG. 1A). Second, because bacterial genomes usually have multiple methylation motifs (three on average, more than 20 in certain strains), multiple REs can be used simultaneously in the third step of mEnrich-seq described above, which can further enhance the enrichment efficiency (FIG. IB and Example 2, above). Third, some methylation motifs are conserved across different strains of the same species or across different genera (for example, G6mATC in most Gammaproteobacteria). Therefore, the choice of RE(s) can be tailored to enrich bacterial taxa at different taxonomic levels, which can be helpful depending on specific research goals (FIG. IB).
[0147] This method was first applied to a mock sample for a proof of concept and illustration. gDNA from Escherichia coli (representing bacterial taxa of interest, roughly 1.00% by mass), human lymphoblastoid cell line (representing host gDNA, roughly 94.00% by mass) and Neisseria gonorrhoeae (representing background bacterial taxa of no interest, roughly 5.00% by mass) were mixed (two replicates). Because GATC sites are methylated (6mA) in the E. coli genome (a site occurs approximately every 250 base pairs (bp)) but not in N. gonorrhoeae or the human genome, the methylation-sensitive RE DpnII (which cuts doublestranded DNA (dsDNA) at nonmethylated GATC sites, but is sensitive to methylated G6mATC) was chosen for mEnrich-seq to selectively enrich E. coli gDNA while digesting N. gonorrhoeae (only 35 of the 2,434 GATC sites are methylated owing to the overlapping with GGTG6mA methylation motif) and human gDNA with hardly detectable methylated GATC sites. After digesting the DNA mock samples with DpnII, the library was prepared for Nanopore long-read sequencing and generated 200 Mb of data (mean read length of roughly 5 kb Table 5). For comparison, the mock samples were also sequenced without mEnrich-seq protocol, labeled as ‘input’ (standard metagenomics). Table 5 summarizes the Nanopore
sequencing dataset for the Mock-1 and Mock-2. Data were produced by Flongle flow cell. The read-length difference is due to the gel-based size selection step in the gel-based size selection step in mEnrich-seq.
Table 5
[0148] Compared with the input data, mEnrich-seq greatly increased the proportion (%) of reads mapped to the E. coli genome from 0.91 to 70.75% and from 1 .20 to 71 .41% in the two mock replicates, respectively (FIG. 2A): roughly 70-fold enrichment (FIG. 2B). mEnrich-seq reads mapping showed that the E. coli genome can be covered without systematic bias for both mock samples (more than 99.96% of genome covered; FIG. 2C). Also, 3.98 and 4.12% of reads from mEnrich-seq are mapped to N. gonorrhoeae. Among them, most (more than 90%) have no GATC sites, representing gDNA molecules that can survive DpnII digestion and PCR amplification; some (less than 5%) reads have GATC site(s) that overlap with the N. gonorrhoeae methylation motif GGTG6mA, thus protected from DpnII digestion; the remaining reads with GATC sites could be due to single-nucleotide polymorphisms (SNPs) or sequencing errors (FIG. 2D and see Example 2, above). Among the reads mapped to the human reference (FIG. 2E), most have no GATC sites. Among the reads with at least one GATC site, most GATC sites are located within genomic regions with simple sequence repeats (Example 2, above) that tend to form secondary structure, making it less accessible to DpnII digestion. This characterization illustrates that mEnrich-seq not only enriches bacterial taxa of interest, but also carries some by-product reads that are either due to methylation at (partially) overlapping motif sites in background taxa or depleted restriction motifs in a subset of gDNA molecules. The highlight of mEnrich-seq is the enrichment efficiency compared to standard metagenomic sequencing despite by-product reads.
Example 4 - Selective sequencing of E. coli genomes from urine samples
[0149] Building on the efficient enrichment of E. coli genome using mock data, mEnrich-
seq was further tested on three urine samples (UTI-1-3) from patients with urinary tract infection (UTI) (E. coli positive urine culture, defined as more than or equal to 100.000 colony forming units per ml; see Example 2, above). The three UTI samples were sequenced with mEnrich-seq (with DpnII as RE) coupled with Nanopore sequencing. For mEnrich-seq, 1.16— 1.58 GB of sequencing data was generated with an average of 233,000 reads per sample (mean read length roughly 6 kb, Tables 6A-6B, below). For comparison with mEnrich-seq, the three urine samples were subjected to standard metagenomic sequencing (labeled as ‘input'), which revealed that UTI-1 contained mostly bacterial DNA, while UTI-2 and UTI-3 were dominated by human DNA (FIG. 6A). E. coli isolates cultured from UTI-1-3 were also sequenced and used the de novo assembled genomes of these isolates as the background truth for evaluation purposes (see Example 2 for details). Using mEnrich-seq, it was found that a remarkable increase of reads mapped to E. coli for all three samples: 8.7-fold for UTI-1 (from 5.88 to 51.14%), 116.8-fold for sample UTI-2 (from 0.77 to 90.00%) and 3.8-fold for UTI-3 (from 24.79 to 93.59%) (FIG. 2F, FIG. 2G). These fold enrichments were further validated using Nanopore Flongle flow cells with lower throughput and cheaper cost to demonstrate flexible utility (FIG. 6B-6C and Tables 6A-6B, below). In addition to the efficient fold enrichment, mEnrich-seq was also evaluated in terms of the completeness of E. coli genomes recovered from enrichment sequencing based on the comparison with the de novo assembled genomes of E. coli isolates. mEnrich-seq reads mapping showed that more than 99.97% of the E coli genomes are covered across all three samples (FIG. 2H). When analyzing the E. coli mEnrich- seq reads, substantial nucleotide variation with the genome assembly of the corresponding E. coli isolate was observed in the UTI-1 sample, suggesting the possibility of multiple E. coli strain(s) in this sample. Consistently, multiple mEnrich-seq reads were found with confident mapping to E. coli harboring let (A) genes, yet not included in the genome assembly of the corresponding E. coli isolate (FIG. 7A-7B). To exclude the possibility of low-quality mapped reads from other species, the source of these E. coli reads was further confirmed by mapping each read manually to the National Center for Biotechnology (NCBI) Nucleotide database, and further validated by PCR (FIG. 7B) and Sanger sequencing (as described in Example 2, above). This observation is consistent with the increasing recognition that a culture-independent approach may provide additional information that can complement standard urine culture in monitoring pathogen genomes in UTI and other infectious diseases.
[0150] Tables 6A-6B, below, show the summary of the Nanopore sequencing dataset for
urine samples. Table 6A illustrates the input, isolates and mEnrich-seq of sample UTI-1 to -3 as first sequenced with a MinlON flow cell. Table 6B illustrates the three urine samples (UTI- 1 to -3) together with the additional three UTI urine samples (UT1-4 to -6) as sequenced on a Flongle flow cell for validation purposes.
Table 6A
Table 6B
[0151] For the same cost and sequencing throughput, the efficient enrichment of E. coll genomes by mEnrich-seq can facilitate the culture-free study of E. coli genomes from urine microbiome with better sensitivity (examining E. coli with lower relative abundance) than standard metagenomics. Alternatively. mEnrich-seq can facilitate a more cost-effective examination of urine samples in the study of E. coli genomes. In addition to E. coli, 6mA at GATC sites are also conserved in most Gammaproteobacteria including many enteric pathogens, which makes mEnrich-seq with DpnII applicable to enrich additional enteric pathogens and commensal bacteria as are described in the following Examples.
Example 5 - Selectively sequencing of Akkermansia muciniphila from gut microbiome
[0152] Akkermansia muciniphila, a gram-negative bacterium colonizing the intestinal tract, has been actively studied for its association with many human diseases such as obesity and type 2 diabetes, and with patient responses to cancer immunotherapies. Because^, muciniphila is strictly anaerobic with a long doubling time, it is challenging to isolate from fecal samples. RA6mATTY (R = A or G; Y = C or T), mediated by a 6mA MTase, AmuORF1905P, is one of the five methylated motifs in A. muciniphila (ATCC BAA-835) according to REBASE (The Restriction Enzyme Database) (FIG. 8A). Using comparative genomics analysis, it was found that the MTase is conserved in all the 112 genomes of muciniphila isolates (Tables 14A- 14B, at end of Examples). This conserved methylation motif is also validated by de novo methylation analysis of two in-house isolates (FIGs. 8B-8C, and as detailed in Example 2 above). This led to applying mEnrich-seq with the RE XapI (digesting nonmethylated RAATTY sites), to deplete background DNA and enrich A. muciniphila from fecal microbiome without culturing.
[0153] mEnrich-seq was applied with XapI to enrich the muciniphila genome from three infant fecal samples (GUT-1-3) (Table 7. below), and observed a 20.0-fold increase of reads mapped to A. muciniphila in GUT-1 (from 0.48 to 9.59%). a 27.2-fold increase in GUT-2 (from 0.52 to 14.16%) and an 18.3-fold increase in GUT-3 (from 1.52 to 27.88%) (FIG. 3A-3B).
Table 7
[0154] Comparative analysis against the genomes of the 4. muciniphila isolates from the three infant fecal samples showed that mEnrich-seq was able to cover more than 99.7% of the A. muciniphila genomes (FIG. 3C and Example 2 above). In addition to 4. muciniphila, several taxa showed an elevated proportion of reads by mEnrich-seq: Bifidobacterium bifidum. Collinsella aerofaciens. Eggerthella lenta and so on (FIG. 3A). To understand the enrichment of these taxa by mEnrich-seq, the target motif frequency was calculated in the mEnrich-seq data and in the corresponding reference genomes. For 4. muciniphila, the RAATTY frequency is comparable between the mEnrich-seq reads and RAATTY frequency of the reference genome (FIG. 3D), which is consistent with the enrichment due to methylated RA6mATTY motif on its genome. By contrast, the frequency of RAATTY in the genomes of B. bifidum, C. aerofaciens and E. lenta is lower, which further decreased among mEnrich-seq reads that mapped to these three species (FIG. 3D). These three genomes with sparse frequency of RAATTY sites were enriched as by-products because of the more efficient depletion of most of the other background genomes with dense yet nonmethylated RAATTY sites.
[0155] To summarize. mEnrich-seq with XapI as the methylation-sensitive RE (targeting RAATTY) can efficiently enrich 4. muciniphila (highly conserved RA6mATTY methylation across all the isolated strains examined), a difficult-to-culture commensal bacterium, directly
from metagenomic DNA. Given the important roles of A. muciniphila in human health, diseases and drug responses, mEnrich-seq can be a useful tool to study the genomes of A. muciniphila in a culture-independent, sensitive and cost-effective way, which may facilitate larger-scale association studies with different human diseases.
Example 6 - Selective sequencing of low-abundance bacterial genomes
[0156] In the previous examples, the use of mEnrich-seq is demonstrated when the methylation motifs of target bacteria of interest are known a priori. In this Example, mEnrich- seq is used to enrich gDNA from low-abundance bacteria from a complex microbiome sample based on methylation motifs discovered de novo. To facilitate the method evaluation, an adult fecal microbiome sample (GUT-5) was built that has been recently well characterized with comprehensive cultured isolates. First, a pilot standard metagenomic sequencing of this adult fecal sample was analyzed (Nanopore sequencing, 7.08 Gb; further details in Example 2 above). Table 8 below provides the summary of this Nanopore sequencing data for low- abundance bacterial taxa enrichment with de novo discovered methylation motifs. The top portion summanzes data from an adult fecal sample sequenced with the MiNION flow cell for the initial standard metagenomic profiling. The bottom portion shows data for the low abundance bacterial taxa enrichment by mEnrich-seq, sequenced with a Flongle flow cell.
Table 8
[0157] While 36 genomes were assembled (more than 90% completeness) from high relative abundance taxa using metaFlye (v.2.9), the others with lower relative abundance could
not be well assembled. To rationally enrich less-abundant bacteria with mEnnch-scq. it was necessary’ to know their methylation motifs (FIG. 4A). Using existing tools for de novo methylation discovery from microbiome samples with an initial round of shallow-depth PacBio sequencing (see details in Example 2 above), 112 methylation motifs across 34 metagenomic bins (corresponding to 34 strains from 26 species) with lower abundance were discovered (see Table 15 at end of the Examples and FIG. 13). In the illustrations described below, the use of some 4-, 5- and 6-mer motifs was prioritized, and commercially available REs across several bacterial taxa with relative low abundance (0.32-1.62%). The initial PacBio sequencing depth was intentionally designed to be shallow to mimic realistic applications, where users can uncover methylation motifs from bacterial species with modest-to-low abundance in a cost- effective manner, yet guide the choices of methylation-sensitive REs for mEnrich-seq. Many methylation motifs can be discovered from genomes that are far from complete, hence guiding follow-up targeted sequencing to get a more complete genome and better assembly (FIG. 13).
[0158] Alistipes flnegoldii H10 and Bacteroides ovatus C6 share the same 5-mer motif CTS6mAG (S = G or C), which can be partially recognized by the methylation-sensitive BspCNI (CTCAG as recognition site). Aiming to enrich both genomes simultaneously, mEnrich-seq was applied with BspCNI as the RE to the fecal metagenomic gDNA. Encouraging enrichment by mEnrich-seq was achieved (FIGs. 4B-4C): 8.7-fold for A. finegoldii H10 (from 1.12 to 9.73%) and 3.1-fold for B. ovatus C6 (from 1.62 to 5.05%). In addition, B. ovatus D3 was also enriched (4.3-fold, from 2.97 to 12.94%), which has a CTCAG frequency similar to A. finegoldii H10 and B. ovatus C6, reflecting that B. ovatus D3 has methylation at CTCAG sites that resist BspCNI digestion (FIG. 4D).
[0159] Bifidobacterium longum F12 and B. longum DIO have the same methylation motif RG6mATCY. mEnrich-seq with Mfll (targeting RGATCY) achieved an 8.7-fold enrichment for B. longum F 12 (from 1.17 to 10.17%) and a sevenfold enrichment for B. longum D10 (from 0.32 to 2.24%) (FIG. 4E). In addition, enrichment was observed for Ruminococcus bromii (4.5-fold, from 2.58 to 11.69%), Eubacterium rectale G9 (3.8-fold, from 2.15 to 8.20%) and Bacteroides uniformis D7 (2.6-fold, from 2.69 to 6.89%; FIG. 4E). Further examination of RGATCY motif frequency supported that R. bromii, E. rectale G9 and B uniformis D7 were all enriched owing to methylation at RG6mATCY sites on their genomes (FIG. 4F).
[0160] Dorea longicatena and Barnesiella intestinihominis share the methylation motif G6mATGC (reverse strand GC6mATC). mEnrich-seq with SfaNI as the RE provided a 43.8-
and 30-fold enrichment of D. longicatena (from 0.39 to 17.10%) and B. intestinihominis (from 0.71 to 21.34%), respectively. The same data also showed a 9.4-fold enrichment of Eubacterium sp. CAG:248 (from 1.62 to 15.17%) (FIG. 4G). Further examination of the target motif frequency supported that Eubacterium sp. CAG:248 was enriched similar to D. longicatena and B. intestinihominis owing to the methylated G6mATGC motif on their genomes (FIG. 4H). B. intestinihominis has one additional methylation motif, CTNN6mAG (N = A, C, G or T), which can be targeted by a mixture of additional REs: AHII (CTTAAG), Acul (CTGAAG), Bpml (CTGGAG), BpuEI (CTTGAG) and PstI (CTGCAG). Next, the use of SfaNI (targeting GATGC) was combined with the additional five REs targeting the CTNNAG motif into a single mEnrich-seq assay (see Example 2, above), which achieved an 85-fold enrichment for B. intestinihominis (from 0.71 to 60.47%), illustrating that mEnrich-seq with combinations of REs targeting multiple motifs can greatly benefit the enrichment efficiency for a specific strain (FIG. 41).
[0161] To summarize, this Example shows that mEnrich-seq can be tailored based on the de novo discovered methylation motifs from adult fecal microbiome samples to enrich bacterial strains with relatively low abundance in the sample. For this type of application, although an initial round of pilot shallow sequencing is performed, the mEnrich-seq strategy’ reduces the overall cost of recovering genomes from species with modest-to-low abundance, whose genomes would otherwise be cost prohibitive to resolve using standard metagenomics. In addition, because mEnrich-seq significantly reduces sequencing reads from background bacteria, it can simplify the assembly of the enriched bacterial genomes. With this rationale, adult fecal samples GUT-5 and GUT-6 were further analyzed to compare the genomes of several targeted species assembled from mEnrich-seq versus standard metagenomic sequencing. Subsampling was used to make sure reads originated from the targeted species have the same yield for genome assembly comparison and it was found that mEnrich-seq significantly improved genome assembly compared to standard metagenomic sequencing (FIG. 4J) Tables 9A-9B show the summary of the GUT-5 and GUT-6 Nanopore sequencing dataset for evaluation of mEnrich-seq on metagenome assembly. Data were produced by a MinlON flow cell. The read-length difference is due to the gel-based size selection step in mEnrich-seq.
Table 9A
Table 9B
Example 7 - Broad applicability of mEnrich-seq
[0162] To assess the general applicability of mEnrich-seq, a meta-analysis of the methylomes of 4,601 bacterial strains (across 1,358 species) was performed. For each methylome, the list of discovered methylation motifs is maintained at REBASE. On average, there are three methylation motifs in each bacterial methylome (FIG. 5A). The number of nondegenerate bases in a motif range from two to nine, mostly between four and eight (FIG. 5B). These methylomes were compared with a list of 584 methylation-sensitive REs targeting 209 motifs that are commercially available (see Table 16, at end of the Examples) and found that 3,130 (68.03%) of the 4,601 strains have at least one (26.62% have two or more) methylation motif(s) that can be targeted (enriched by depleting the background metagenomic DNA) by commercially available RE(s). When projected onto a phylogenetic tree (see Example 2, above), these 3,130 strains have broad taxonomic diversity (FIG. 5C), representing 54.78% of the species examined in this analysis. This analysis focused on REs with recognition motifs that have four, five or six nondegenerate bases, as these REs are expected to be readily applicable considering the expected frequency of their target motifs across bacterial genomes:
roughly 256 bp for 4-mers, roughly 1 kb for 5-mers and roughly 4 kb for 6-mers. As researchers are actively developing more advanced DNA extraction methods that preserve high molecular weight gDNA from microbiome samples, it is expected that mEnrich-seq may expand to REs with additional target motifs with lower frequency.
[0163] Based on a systematic characterization of DNA MTases conservation in a recent survey, a methylation motif can be conserved at different taxonomic levels. Some methylation motifs are conserved at the family level. For example, the G6mATC motif by the Dam DNA methyltransferase is highly conserved in many families of the Gammaproteobacteria class, which includes many of the enteric pathogens (for example, E. coli, Salmonella enterica, Vibrio cholerae and so on) as well as some nonpathogenic species across many families such as Shewanellaceae and Moraxellaceae and so on. Some methylation motifs are conserved at the genus level. For example, RA6mATTY is conserved in genus Campylobacter, and GTWW6mAC is conserved in genus Burkholder ia. Some methylation motifs are conserved at the species level, for example, RA6mATTY motif in A. muciniphila, and C5mCWGG in E. coli. Many strains have strain-specific methylation motifs in addition to the conserved motifs shared at higher taxonomic levels as illustrated in Table 15 (at end of Examples), for example, G6mATC, CCRC6mAG. CTAG- 6mAG, CTCT6mAG in strains of B. ovatus or CG6mAYNNNNNRTTC in B. longum. The complex structure of conservation and variations of methylation motifs across different taxonomic levels provide a great opportunity to target different taxa of interest using mEnrich-seq.
[0164] Some REs have very broad utility as their target motifs are methylated in a large number of bacterial genomes (FIG. 5D-5E). Taking DpnII (GATC) and XapI (RAATTY) as examples, the previous examples show their utilities with mEnrich-seq in two specific applications enriching for E. coli and A. muciniphila, respectively. In addition, mEnrich-seq with DpnII can help enrich many bacteria in the Gammaproteobacteria class in which G6mATC methylation is highly conserved, while mEnrich-seq with XapI can help enrich other bacteria such as Campylobacter spp., Acinetobacter spp., Spirochaeta spp., Treponema spp. and Brachyspira spp. In practical applications, the discriminative power of methylation motifs in practical applications is with respect to the bacteria in a specific microbiome sample, not among all the genomes used in the phylogenetic analysis.
Example 8 - Comparison of mEnrich against other metagenomic tools
[0165] The core idea of mEnrich-seq is essentially two-fold: (i) methylation-guided
digestions/enrichment followed by (ii) size selection. Along with this core idea, there are multiple ways to implement mEnrich-seq. The design described throughout the previous Examples performs adapter ligation before digestion ('ABD method’, uses gel purification for size selection, and utilizes the adapter sequence for subsequent PCR. Alternatively, one can either perform adapter ligation after digestion (‘AAD method’), or uses Tn5-based adapter tagmentation, or use T5 exonuclease to preferentially deplete shorter DNA fragments after digestion as described in a recent work called REMoDE (Enam, S. U. et al., Restriction endonuclease-based modification-dependent enrichment (REMoDE) of DNA for metagenomic sequencing Appl. Environ. Microbiol. 89, eO 167022 (2023), which is incorporated herein by reference in its entirety). Therefore, these two alternative methods differ from mEnrich-Seq primarily because they do not perform adapter ligation before digestion. Both methods were directly compared to mEnrich-Seq as described below.
[0166] In a first set of experiments, mEnrich-seq design was directly compared to REMoDE, which relies on a T5 exonuclease to preferentially degrade short DNA fragments after digestion by a methylation restriction enzyme - thereby enriching for longer, presumably methylated, non-digested fragments of interest. Details regarding preparation of the REMoDE library and data analysis are described in Example 2, above. It was found that mEnrich-seq more successfully enriched for target DNA when directly compared to REMoDE (FIG. 14A- 14B, Tables 10A-10C) This likely is because the human microbiome is very complex, in which some background genomes not methylated at a target motif can have a sparse frequency of RE recognition sites. For these background genomes with non-methylated but sparse frequency sites, even after RE digestion, some genomic regions can still have relatively long DNA fragments. This data demonstrates that adding adapters before digestion (the ABD method) can more effectively remove background species with sparse target motifs.
[0167] In another set of experiments, mEnrich-seq w as compared REMODE using samples that had been treating with bead beating, which is a common practice for fragmenting DNA in diverse microbiome samples containing species with very different bacterial cell walls to minimize biases in DNA extraction across different bacterial species. A problem with bead beating is that it usually leads to fragmented lengths rather than the ideal high molecular weight DNA. It w as again found that mEnrich-Seq has better enrichment than REMoDE when gDNA has modest fragment lengths and DNA damages, even when a DNA repair step is added to REMoDE (FIG. 15A-15B, Tables 10A-10C)
Table 10A
Table 10B
Table IOC
[0168] To further test whether adaptor ligation should occur before or after digestion, a simple experiment was performed where 5-1 Opg of Mock-3 was used as input and digested with Xapl in a 50 pL reaction system at 37°C for 30 min. followed by size selection for DNA larger than 10 kb by gel purification and elution with nuclease-free water before further processing for library preparation using the methods described in "Library preparation for the standard metagenomic sequencing (Input)” in Example 2, above. This library was the “Ligation-after-RE library” and it was directly compared to a standard mEnrich-seq library (500-100 ng as input) that was generated for evaluation of the adaptor ligation before RE digestion (“Ligation-before-RE”). The ONT data of Ligation-before-RE or Ligation-after-RE were each mapped to the reference using minimap2 (v2.24-rl !22) with -x map-ont and reads mapped to each species were counted. As shown in FIG. 16 and summarized in Table 11,
below, adaptor-ligation before RE resulted in greater enrichment than adaptor-ligation after RE.
Table 11
[0169] While the data in FIGS 14A-14B, 15A-15B and 16 highlight the advantages of adapter ligation before digestion, it is worth noting that a single adapter sequence is not compatible with all the REs, because some enzymes can digest the adapter. The adapter sequence used in the previous Examples (e.g., having an upper oligonucleotide sequence of SEQ ID NO: 35 and a lower oligonucleotide sequence of SEQ ID NO: 36, as described in Example 2, above) is compatible with 556 of the 584 (95%) commercially available REs. In addition, one additional adapter was synthesized and found to be compatible with the remaining 28 commercially available REs. This second “alternative” adaptor contained a top oligo having a sequence of SEQ ID NO: 37 and a bottom oligo having a sequence of SEQ ID NO: 38. So, all the 584 commercially available REs are compatible with at least one of the two adapter oligos (FIG. 17).
[0170] Another set of experiments evaluated the few rounds of PCR used before Nanopore sequencing and did not observe systematic bias due to GC content, much less than the GC bias previously reported in the Illumina platform (FIGs. 18-19). However, for microbiome samples with adequate input gDNA, none or fewer rounds of PCR may be preferred.
[0171] Finally, in addition to the design described earlier in the Examples, an alternative design for methylation-guided digestion and size selection was tested comprising digestion, gel-based size selection, long-read Tn5 adapter tagmentation, and PCR. Comparable enrichment was observed between the two methods (Tables 12-13).
Table 12
Table 13
Example 9 - Complementarity between mEnrich-seq and adaptive sampling.
[0172] Adaptive sampling is effective in reducing the allocation of sequencing yield to host gDNA or highly abundant bacteria. However, with this approach, desired taxa are indirectly enriched by the rejection of reads from highly abundant species, thus the efficiency enrichment by adaptive sampling from latest studies is still modest and insufficient for low-abundant bacteria. In contrast, owing to the high processivity of REs, mEnrich-seq has the advantage of more efficiently depleting the vast background gDNA before sequencing, hence achieving high fold enrichments. Adaptive sampling can also select for specific sequences of interest (such as a desired genome): however, it does so by requiring a pre-defined genome sequence hence it will likely miss genes unique to a specific strain yet to be determined. In contrast, mEnrich-seq has the advantage of enriching both known and unknown genes in a strain of interest, including mobile genetic elements, because all the genetic contents within a bacterial cell share the same methylation motifs. So, mEnrich-seq is complementary to the adaptive sampling method on the Nanopore sequencing platform.
Example 10 - Discussion of Examples 1 -9
[0173] Bacterial DNA methylation provides a natural way for differentiating diverse bacterial taxa from each other. Although it has been used for metagenomic binning, it is only applicable as a post hoc analysis, after single molecule real-time (SMRT) sequencing or
Nanopore sequencing. By contrast, mEnrich-seq enables the use of bacterial DNA methylation to enrich certain taxa of interest before sequencing, which allows microbiome researchers to tailor sequencing strategy to best serve their goals. While more than 100-fold enrichment in was achieved in some samples, even a modest fold enrichment of fivefold can provide a roughly 80% reduction in the sequencing data needed to obtain the same number of reads from a target genome, five times higher sensitivity to investigate the genetic variations in the target genome or better scalability to examine five times more samples with multiplexing.
[0174] Thousands of bacterial methylomes have been characterized to date, revealing many species with highly conserved methylation motifs. For these well-characterized bacterial taxa, mEnrich-seq is directly applicable and is expected to facilitate many applications. For example, the Examples herein demonstrate the use of mEnrich-seq to enrich E. coli from urine samples and to enrich A. muciniphila from fecal samples. In addition, they also demonstrate the use of mEnrich-seq along with de novo methylation discovers’ to enrich several low- abundance bacteria from a complex fecal microbiome sample.
[0175] mEnrich-seq is broadly applicable and versatile. Based on a meta-analysis of the 4,601 bacterial methylomes mapped to date, it is estimated that roughly 68% of bacterial genomes can be targeted by mEnrich-seq with at least one RE, representing 54.78% of the species examined. In practice, mEnrich-seq may be applicable to a broader diversity of bacterial genomes, because epigenomes mapped to date represent only a small fraction of the bacterial diversity. mEnrich-seq is also very versatile in that it can be used along with one or multiple RE(s) (FIG. IB). Essentially, mEnrich-seq provides a flexible way to dissect a microbiome sample when it is coupled with different REs, the choice of which can be tailored based on taxa of interest in a particular application. In addition, mEnrich-seq is compatible with different long-read and short-sequencing platforms (FIGs. 14A-14B and FIGs. 15A- 15B). The core idea of mEnrich-seq is essentially twofold: (1) methylation-guided digestions and/or enrichment follow ed by (2) size selection. Along w ith this core idea, there are multiple ways to implement mEnrich-seq. The design described in these Examples performs adapter ligation before digestion, uses gel purification for size selection and uses the adapter sequence for subsequent PCR. Alternatively, one can either perform adapter ligation after digestion, using Tn5-based adapter tagmentation, or use T5 exonuclease to preferentially deplete shorter DNA fragments after digestion as described in a recent work called REMoDE. As discussed in Example 8, above, these and several additional designs were examined in depth which
highlighted a few important factors critical for practical applications such as DNA quality', diverse background taxa and abundance of targeted taxa (FIG. 14A-14B, 15A-15B, 16, 17, 18, 19 and Tables 12 and 13).
[0176] mEnrich-seq also has some limitations. First, although 584 commercial REs are readily available (Table 16), they do not cover all motifs of interest. While additional REs may be cloned and purified depending on specific needs, it is expected that broader use of mEnrich- seq in microbiome applications could motivate vendors to expand the list of commercially available REs. Second, not all the reads from mEnrich-seq are from bacterial genomes with the expected DNA methylation. Some by-product reads are expected because some genomes or specific genomic regions can have very sparse RE recognition sites, such as the rare RAATTY motifs on B. longum genome due to codon bias (FIG. 3D). Meanwhile, some bacterial reads were enriched because their methylated motifs partially overlapped with the RE recognition motif (FIG. 2D). The highlight of mEnrich-seq is its efficacy in the enrichment of taxa of interest over standard metagenomics. Finally, mEnrich-seq is primarily applicable to enrich bacteria, archaea and dsDNA phages, but not applicable to single-stranded (ssDNA) phages, RNA phages and parasites, because they are not substrates of restriction-modification systems. It is worth noting a recent work leveraged the size differences between host and bacterial cells and used mechanical stress to effectively deplete host cells, which could be integrated with mEnrich-seq to study targeted taxa in host-rich samples.
[0177] In summary, mEnrich-seq is a versatile and cost-effective approach for enrichment sequencing of diverse microbial taxa of interest directly from the microbiome. It provides microbiome researchers an additional approach to more effectively study complex microbiomes.
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
Additional Tables
Tables 14A-14B. The MTase mediating RAATTY methylation is conserved across 1 12 genomes of A. muciniphila isolates. The amino acid sequence of the MTase (methylating RAATTY) is used as a query to blast against the 112 reference genomes. The prevalence of the MTase in 112 Akkermansia isolates is 100%, when using the following filter criteria for the BLAST hits: bitscores > 100 & identity > 80% & alignment length > 300.
Table 14A
64
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
65
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
66
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
67
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
68
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
69
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
70
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
Table 14B
71
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
72
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
73
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
74
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
75
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
76
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
77
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
78
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
79
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
80
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
81
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
82
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
83
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
84
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
85
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
86
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
Table 15. The list of methylation motifs de novo discovered from the adult fecal metagenomic sample, de novo methylation discovery from microbiome samples was performed using existing tools (see Example 2). 112 methylation motifs were found across 34 bins (from 26 species) with lower abundance (Relative abundance profiled by Nanopore is indicated below). The bold motifs were the 4-mer or 5-mer motifs selected as the targets for illustrative purposes.
Table 15
87
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
88
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
Table 16. The list of commercially available REs. Enzymes are listed in order of motif, first by motif size, and then by sequence. Enzyme cut sites are designated by arrows. Isoschizomers are show n on the same row. The last colunm is for selection of the BC Adapter compatible to the REs.
Table 16
89
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
90
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
91
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
92
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
93
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
94
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
95
92636236.6
PCT PA TENTAPPLICA TION Atorney Docket No. 093698-782858 Via Patent Center
96
92636236.6
Claims
1. A method of selectively amplifying a nucleic acid of one or more prokaryotic organisms of interest (target organisms) in a microbiome sample, the method comprising:
(a) obtaining or having obtained a microbiome sample comprising a plurality of prokaryotic-organisms, wherein the microbiome sample comprises genomic DNA (gDNA) from one or more target organisms (target gDNA) and gDNA from one or more prokaryotic organisms that are not of interest (background organism gDNA);
(b) isolating gDNA from the microbiome sample to obtain target gDNA fragments and background organism gDNA fragments;
(c) selecting one or more methylation-sensitive restriction enzymes based on the DNA methylome (a collection of DNA methylation sequence motifs) of the one or more target organisms wherein the one or more methylation restriction enzy mes target motifs that are primarily methylated in the target gDNA (methylated motifs) and primarily unmethylated in the background organism gDNA (non-methylated motifs);
(d) selecting universal PCR adapters that do not comprise non-methylated motifs of the one or more methylation-sensitive restriction enzyme(s);
(e) linking the -universal PCR adapters to both 5’- and 3’-ends of target gDNA fragments and the background organism gDNA fragments;
(1) subjecting the target gDNA fragments and the background organism gDNA fragments linked with the universal PCR adapters to digestion by the one or more methylation sensitive restriction enzyme(s) selected in (c);
(g) amplifying undigested gDNA fragments comprising universal PCR adaptors at both 5’ and 3’ ends with forward and reverse barcoded primers corresponding to the universal PCR adapters, thereby selectively amplifying the genome of target prokaryotic organism in the microbiome gDNA sample.
2. The method of claim 1. further comprising an end repair and/or a dA-tailing step prior to the linking of the universal PCR adapters.
3. The method of claim 1 or 2, wherein the universal PCR adaptor comprises an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, and wherein the hybridized region does not comprise a non-methylated motif targeted by the one or more restriction enzyme.
4. The method of claim 3, wherein the universal PCR adaptor of claim 2, further comprising a *T at the 3’ end of the upper strand oligo and/or is phosphorylated at the 5’ end of the lower strand oligo.
5. The method of claim 3 or 4, wherein the upper strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 1 and 3-18 and the lower strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 2 and 19-34.
6. The method of any one of claims 3 to 5, wherein the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 36; or the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 38.
7. The method of any one of claims 3 to 6, wherein each upper-strand oligo comprises a primer binding site complementary to a forward barcoded primer, and each lower-strand oligo comprises a primer binding site complementary to a reverse barcoded primer.
8. The method of any one of claims 1 to 7, wherein the methylated motif comprises a nucleic acid sequence motif having one or more methylated nucleotides.
9. The method of claim 8, wherein the one or more methylated nucleotides comprise N4-methylcytosine (4mC). N6-methyladenine (6mA) or 5-Methylcytosine (5mC) or any combination thereof.
10. The method of claim 9, wherein the methylated motif comprises GATC, RAATTY where R is A or G and Y is C or T, or CTSAG where S is G or C, RGATCY where R is A or G and Y is C or T, GATGC, CTNNAG and each A is N6-methyladenine (6mA).
11. The method of any one of claims 1 to 10, wherein the non-methylated motif comprises a nucleic acid sequence motif that is recognized and cleaved by the one or more methylation sensitive restriction enzyme.
12. The method of claim 11, wherein the non-methylated motif comprises GATC, RAATTY where R is A or G and Y is C or T, or CTSAG where S is G or C, RGATCY where R is A or G and Y is C or T, GATGC, CTNNAG and each A is non-methylated.
13. The method of any one of claims 1 to 12, wherein the one or more methylation sensitive restriction enzy mes comprise a class of natural and/or artificial restriction enzymes that are sensitive to DNA methylation.
14. The method of claim 13, wherein the methylation-sensitive restriction enzymes comprise DpnII, XapI, Apol, BspCNI, BseMII, Mfll, BstX2I, BstYI, MflI, XhoII, SfaNI, AfUI, Acul, Bpml, BpuEI, PstI, any methylation-sensitive isoschizomer thereof, or any combination thereof.
15. The method of any one of claims 1 to 14, where the methylation-sensitive restriction enzyme comprises a restriction enzyme whose cleavage abi 1 i ty is blocked by the methylated motif sites.
16. The method of any one of claims 1 to 15, further comprising identifying a methylation profile of the target prokaryotic organism and selecting the one or more restriction enzymes based on the methylation profile of the target prokaryotic organism.
17. The method of claim 16, further comprising deconvoluting genomes of prokaryotic organisms in the microbiome sample to identify a methylation profile of the target prokaryotic organism, wherein deconvoluting genomes comprises:
A) obtaining a microbiome sample comprising a plurality of prokaryotic- organisms;
B) sequencing nucleic acids of the prokaryotic organisms using single-molecule long reads sequencing technology, wherein the sequencing comprises the step of identifying methylated nucleotides, and at least one of the steps of: i. sequencing single molecule reads of nucleic acids; ii. assembling contigs from single molecule reads of the nucleic acids;
C) assigning a methylation score reflecting the extent of methylation for sequence motifs of the nucleic acids on the assembled contig and/or the single molecule read;
D) applying motif filtering to identify sequence motifs with methylation scores indicating methylation on the assembled contigs and/or the single molecule reads;
E) determining nucleic acid methylation profiles of the assembled contigs or the single molecule reads in the microbiome sample based on motifs identified in step (D);
F) separating the assembled contigs and/or the single molecule reads into bins corresponding to distinct prokary otic organisms based on the methylation profiles of step (E); and
G) assembling the bins of step (F), thereby obtaining assembled genomes of the distinct bacterial organisms in the microbiome sample, thereby deconvoluting genomes of the prokaryotic organisms in the microbiome sample.
18. The method of claim 17, further comprising the step of combining the methylation profiles of step (E) with other sequence features of the nucleic acids of the prokaryotic organisms in the microbiome sample prior to separating the assembled contigs and/or the single molecule reads into bins.
19. The method of any one of claims 1 to 18, wherein the target prokaryotic organism(s) comprises one or more of an E. colt strain, Akkermansia muciniphila. Alistipes fimegoldii and Bacteroides ovatus, Bifidobacterium longum, Ruminococcus bromii, Eubacterium rectale, Bacteroides uniformis D7, Dorea longicatena and Barnesiella intestinihominis , Eubacterium sp. CAG:248. Alistipes onderdonkii, Enterocloster clostridioformis, Bacteroides fragilis, Gordonibacter pamelaea, Adler creutzia sp. 8CFCBH1 , Adlercreutzia equolifaciens , or any combination thereof.
20. The method of claim 19, where the target prokaryotic organism(s) are selected from specific pathogenic bacteria, beneficial bacteria, probiotics, or prokaryotic organisms with low relative abundance in a microbiome sample as determined by standard metagenomic sequencing.
21. The method of any one of claims 1 to 20, wherein
(i) the target prokaryotic organism(s) comprise a methylated G6mATC motif, and the one or more methylation sensitive restriction enzyme comprises methylation sensitive DpnII and/or a methylation sensitive isoschizomer thereof;
(ii) the target prokaryotic organism(s) comprise a methylated RA6mATTY motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive XapI , methylation sensitive Apol, and/or a methylation sensitive isoschizomer thereof, or any combination thereof;
(iii) the target prokaryotic organism(s) comprise a methylated CTS6mAG motif and the one or more methylation sensitive restriction enzy me comprises methylation sensitive BspCNI, methylation sensitive BseMII, a methylation sensitive isoschizomer thereof, or any combination thereof;
(iv) the target prokaryotic organism(s) comprise a methylated RG6mATCY motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive Mfll. methylation sensitive BstX21, methylation sensitive BstYI, methylation sensitive Mfll. methylation sensitive XhoII, a methylation sensitive isoschizomer thereof, or any combination thereof;
(v) the target prokaryotic organism(s) comprise a methylated G6mATGC motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive SfaNI and/or a methylation sensitive isoschizomer; and/or
(vi) the target prokary otic organism(s) comprise a methylated G6mATGC motif and a methylated CTNN6mAG motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive Aflll, methylation sensitive Acul, methylation sensitive Bpml, methylation sensitive BpuEI, methylation sensitive PstI, methylation sensitive SfaNI, a methylation sensitive isoschizomer thereof, or any combination thereof.
22. The method of claim 21, wherein:
(i) the target prokaryotic organism(s) comprising a methylated G6mATC motif comprise Escherichia coli, Klebsiella pneumoniae, one or more species of Gammaproteobacteria class or any combination thereof;
(ii) the target prokary otic organism(s) comprising a methylated RA6mATTY motif comprise one or more species of Akkermansia muciniphila, Campylobacter spp., Acinetobacter spp., Spirochaeta spp., Treponema spp. and Brachyspira spp or any combination thereof;
(iii) the target prokaryotic organism(s) comprising a methylated CTS6mAG motif comprise one or more species of Alistipes finegoldii and/ or Bacteroides ovatus:
(iv) the target prokaryotic organism(s) comprising a methylated RG6mATCY motif comprise one or more species of Bifidobacterium longum, Ruminococcus bromii, Eubacterium rectale and Bacteroides uniformis;
(v) the target prokary otic organism(s) comprising a methylated G6mATGC motif comprise one or more species of Dorea longicatena, Barnesiella intestinihominis and Eubacterium sp. C AG: 248
; and/or
(vi) the target prokaryotic organism(s) comprising a methylated G6mATGC motif and methylated CTNN6mAG motif comprise one or more species of Barnesiella intestinihominis .
23. The method of any one of claims 1 to 22, further comprising sequencing the nucleic acid of the amplified target prokaryotic organism(s).
24. The method of claim 23, wherein the sequencing is performed using a singlemolecule long read real time (SMRT) technology, nanopore sequencing technology’ or other short read sequencing technologies including but not limited to Illumina sequencing platforms.
25. The method of any of claims 1 to 24, wherein the microbiome sample is obtained from soil, air, water, sediment, oil, and combinations thereof.
26. The method of any of claims 1 to 25, wherein the microbiome sample is obtained from water selected from marine water, fresh water, and rainwater.
27. The method of any of claims 1 to 26, wherein the microbiome sample is obtained from a subject selected from a protozoa, an animal, or a plant.
28. The method of claim 27. wherein the subject is a mammal.
29. The method of any of claim 27 or 28, wherein the subject is human.
30. The method of any of claim 27 to 29, wherein the subject is an infant.
31. A universal adaptor, wherein the adaptor comprises an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary- to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary- and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, wherein the hybridized region does not comprise a motif targeted by a methylation sensitive restriction enzyme or does comprise a methylated nucleotide in a motif targeted by the methylation sensitive restriction enzyme.
32. The universal adaptor of claim 31, further comprising a *T at the 3‘ end of the upper strand oligo and/or a phosphorylated base at the 5’ end of the lower strand oligo.
33. The universal adaptor of claim 31 or 32, wherein the upper strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 1 and 3-18 and the lower strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 2 and 19-34.
34. The universal adaptor of any one of claims 31 to 33. wherein the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 36; or the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 38.
35. A kit comprising (a) a set of universal PCR adapters and (b) one or more methylation sensitive restriction enzymes.
36. The kit of claim 35, wherein the one or more methylation sensitive restriction enzymes comprise restriction enzymes sensitive to methylated motifs in one or more target prokaryotic organisms
37. The kit of claim 36, wherein the one or more methylation sensitive restriction enzymes comprise DpnII, XapI, Apol, BspCNI, BseMII, Mfll, BstX2I, BstYI, Mfll, XhoII, SfaNI, Aflll, Acul. Bpml, BpuEI, PstI or a combination of any thereof.
38. The kit of any one of claims 35 to 37, wherein each universal PCR adapter in (a) comprises an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, wherein the hybridized region does not comprise a motif targeted by a methylation sensitive restriction enzyme or does comprise a methylated nucleotide in a motif targeted by the methylation sensitive restriction enzy me.
39. The kit of claim 38, wherein each universal PCR adapter in (a) further comprising a *T at the 3’ end of the upper strand oligo and/or a phosphorylated base at the 5’ end of the lower strand oligo.
40. The kit of claim 38 or 39, wherein the upper strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 1 and 3-18 and/or the lower strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 2 and 19-34.
41. The kit of any one of claims 38 to 40 wherein the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 36; or the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 38.
42. The kit of any one of claims 35 to 41, further compnsing (c) a set of barcoded PCR primers complementary’ to at least a portion of the universal PCR adapters.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202263382616P | 2022-11-07 | 2022-11-07 | |
| PCT/US2024/010215 WO2024103078A2 (en) | 2022-11-07 | 2024-01-03 | Cost-effective long read metagenomic sequencing |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4627111A2 true EP4627111A2 (en) | 2025-10-08 |
Family
ID=91033510
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24725039.2A Pending EP4627111A2 (en) | 2022-11-07 | 2024-01-03 | Cost-effective long read metagenomic sequencing |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4627111A2 (en) |
| WO (1) | WO2024103078A2 (en) |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP4189495B2 (en) * | 2004-08-02 | 2008-12-03 | 国立大学法人群馬大学 | Method for detecting methylation of genomic DNA |
| CA3131514A1 (en) * | 2019-02-25 | 2020-09-03 | Twist Bioscience Corporation | Compositions and methods for next generation sequencing |
| JP2020191833A (en) * | 2019-05-29 | 2020-12-03 | 森永乳業株式会社 | A method of selectively amplifying the DNA of a target bacterium in a bacterial population |
| US20230212684A1 (en) * | 2020-05-05 | 2023-07-06 | The Board Of Trustees Of The Leland Stanford Junior University | Cell-free dna biomarkers and their use in diagnosis, monitoring response to therapy, and selection of therapy for prostate cancer |
| US20220135965A1 (en) * | 2020-10-26 | 2022-05-05 | Twist Bioscience Corporation | Libraries for next generation sequencing |
-
2024
- 2024-01-03 EP EP24725039.2A patent/EP4627111A2/en active Pending
- 2024-01-03 WO PCT/US2024/010215 patent/WO2024103078A2/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024103078A8 (en) | 2024-06-20 |
| WO2024103078A2 (en) | 2024-05-16 |
| WO2024103078A3 (en) | 2024-07-18 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20250109426A1 (en) | Compositions and methods for targeted depletion, enrichment, and partitioning of nucleic acids using crispr/cas system proteins | |
| Parkinson et al. | Preparation of high-quality next-generation sequencing libraries from picogram quantities of target DNA | |
| US9850525B2 (en) | CAS9-based isothermal method of detection of specific DNA sequence | |
| US12416003B2 (en) | Methods and compositions for enrichment of target polynucleotides | |
| JP7436493B2 (en) | Normalization control for handling low sample input in next-generation sequencing | |
| AU2018256358B2 (en) | Nucleic acid characteristics as guides for sequence assembly | |
| US20230279385A1 (en) | Sequence-Specific Targeted Transposition and Selection and Sorting of Nucleic Acids | |
| JP5116481B2 (en) | A method for simplifying microbial nucleic acids by chemical modification of cytosine | |
| WO2023148235A1 (en) | Methods of enriching nucleic acids | |
| JP2024541111A (en) | Sequencing library construction methods and uses | |
| US20230068726A1 (en) | Transposon systems for genome editing | |
| EP4627111A2 (en) | Cost-effective long read metagenomic sequencing | |
| Stocks | Transposon Mediated Genetic Modification of Gram-positive Bacteria |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250822 |
|
| AK | Designated contracting states |
Kind code of ref document: A2 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |