WO2025007145A2 - Metagenomic diagnostic systems and methods thereof - Google Patents
Metagenomic diagnostic systems and methods thereof Download PDFInfo
- Publication number
- WO2025007145A2 WO2025007145A2 PCT/US2024/036435 US2024036435W WO2025007145A2 WO 2025007145 A2 WO2025007145 A2 WO 2025007145A2 US 2024036435 W US2024036435 W US 2024036435W WO 2025007145 A2 WO2025007145 A2 WO 2025007145A2
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- reads
- samples
- sample
- pathogen
- sequencing
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B35/00—ICT specially adapted for in silico combinatorial libraries of nucleic acids, proteins or peptides
- G16B35/10—Design of libraries
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
- G16B30/10—Sequence alignment; Homology search
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6869—Methods for sequencing
Definitions
- the present invention generally relates to systems and methods for metagenomic diagnosis.
- Culturing is a conventional method to identify pathogens.
- culturing is time-consuming and many pathogens that need specific culture conditions are difficult to grow.
- Metagenomic sequencing is a target-independent approach that offers a comprehensive view of the pathogenic agents and provides a detection of common and unexpected pathogens in samples. This technology performs well in detecting rare pathogens and/or novel viral strains.
- Current approaches to metagenomic sequencing can only process a small number of samples at a time and each sample costs hundreds of dollars. There are multiple challenges to processing larger amounts of samples including difficulty to prepare the samples and cross-contamination which is when the samples are mixed during processing causing incorrect results.
- Some embodiments include a method of detecting a pathogen comprising, obtaining a plurality of nucleic acids from a sample; sequencing the plurality of nucleic acids to obtain a plurality of reads; analyzing the plurality of reads using a bioinformatic process; and identifying the pathogen from the analyzed reads.
- Some embodiments further comprise obtaining the sample using at least one of: a swab, a Q-tip, a wipe, and a cloth.
- the sample is a tissue, an organ, a bodily fluid, an object, a surface, or a container.
- the sample is at least one of: a turbinate swab, a nasal pharyngeal swab, and saliva.
- the sample is a clinical sample; wherein the clinical sample is at least one of: urine, a tissue, a skin swab, bronchoalveolar lavage (BAL), sputum, and blood.
- the clinical sample is at least one of: urine, a tissue, a skin swab, bronchoalveolar lavage (BAL), sputum, and blood.
- the sample is an upper respiratory specimen or a lower respiratory specimen from a subject.
- the obtaining step obtains the plurality of nucleic acids using at least one nucleic acid extraction kit.
- Some embodiments further comprise adding a known quantity of Bacteriophage MS2 to the sample.
- Some embodiments further comprise constructing at least one next generation sequencing library.
- the at least one next generation sequencing library is at least one of: a DNA library, and an RNA library.
- the sequencing step uses the at least one next generation sequencing library.
- the analyzing step comprises matching a plurality of sequencing reads to at least one pathogen sequence in a database.
- Some embodiments further comprise generating a plurality of simulated reads using a set of curated pathogen sequences; analyzing the plurality of simulated reads using the database to check if the database contains an error; and eliminating a plurality of false reads in the database by analyzing the plurality of simulated reads.
- Some embodiments further comprise using the plurality of reads of the pathogen and a plurality of reads of a control added prior to nucleic acid extraction or a plurality of reads of a control added after nucleic acid extraction to quantify an amount of the sample.
- control is Bacteriophage MS2.
- the pathogen is a pathogenic organism, a bacterium, a virus, or a fungus.
- Some embodiments further comprise processing a plurality of samples at a time using automation.
- Some embodiments further comprise processing a plurality of samples at a time using automation by interleaving a plurality sets of samples to increase a number that is processed on an automated instrument.
- Some embodiments further comprise processing a plurality of samples using at least one microfluidic liquid handler.
- automation eliminates cross-sample contamination when processing the plurality of samples.
- Some embodiments further comprise adding at least one molecular barcode to each of the plurality of samples, and mixing the plurality of samples.
- more than 90 samples are processed together at a time.
- more than 200 samples are processed together at a time.
- more than 1000 samples are processed together at a time.
- the preparation for sequencing and analyzing is performed on a microfluidic liquid handler.
- Some embodiments further comprise a depletion process configured to improve a performance of the mixture of barcoded plurality of samples.
- Some embodiments further comprise combining a plurality of targeted diagnostic samples with a plurality of metagenomic diagnostic samples in a same sequencing process.
- Some embodiments further comprise using a nucleic extraction process that is optimized for extracting bacteria and viruses compared to host nucleic acids.
- sample to result using short read sequencers is completed in less than 11 hours.
- Some embodiments include a method of a bioinformatic analysis comprising, generating a plurality of simulated genomic reads using a plurality of pathogen genomic sequences; eliminating a plurality of false reads in a database by analyzing the plurality of simulated genomic reads; comparing a plurality of sequencing reads from a sample with the database; and identifying a pathogen from the sample; wherein the eliminated false reads in the database improves identification accuracy.
- Some embodiments further comprise analyzing the plurality of simulated genomic reads using the database to check if the database contains an error.
- Some embodiments further comprise generating a plurality reads of a known quantity of Bacteriophage MS2 and using the plurality reads of the known quantity of Bacteriophage MS2 and the plurality of sequencing reads to quantify an amount of the sample.
- Some embodiments further comprise extracting a plurality of nucleic acids from the sample.
- the extracting step uses at least one nucleic acid extraction kit.
- Some embodiments further comprise using a next generation sequencer to generate the plurality of sequencing reads.
- the pathogen is a bacterium or a virus.
- the pathogen is reconstructed using a genome assembly technique.
- eliminating the plurality of false reads uses a manual process.
- eliminating the plurality of false reads uses a computer assisted process.
- the computer assisted process uses artificial intelligence.
- Some embodiments further comprise using a control sequence that is an organism other than Bacteriophage MS2.
- Some embodiments further comprise adding a control RNA or DNA after nucleic acid extraction to generate a plurality of simulated genomic reads.
- the database used has publicly available sequencing data that is curated prior to analysis using efficient algorithms.
- Figure 1 illustrates a process of SwabSeq metagenomics diagnostic platform workflow in accordance with an embodiment.
- Figure 2 illustrates a 96 well plate with the control positions marked in accordance with an embodiment.
- Figure 3 illustrates workflows for the metagenomic diagnostic platform multiplex assay in accordance with an embodiment.
- Figure 4 illustrates a process of bioinformatic analysis process in accordance with an embodiment.
- Figure 5 illustrates sample output for 20 COVID-19 samples analyzed using analyzed using the metagenomic diagnostic platform in accordance with an embodiment.
- Figure 6A shows True Positive positions of Parainfluenza 3 in accordance with an embodiment.
- Figure 6B shows predicted positions using manual library preparation protocol in accordance with an embodiment.
- Figure 6C shows predicted predictions using automated library preparation protocol in accordance with an embodiment.
- Figures 7A and 7B illustrate interleaving of two plates in accordance with an embodiment.
- Figure 8A through 8I show the results from automation validation for three additional pathogens in accordance with an embodiment.
- Identifying different pathogens in a sample may be challenging. These challenges stem from the fact that the number of sequencing reads that correspond to a pathogen is a very small fraction of the total amount of sequence obtained for a sample, often in the range of 1 in about 10,000 reads and in advance the pathogens present in the sample are unknown species.
- sequencing reads corresponding to a pathogen genome need to be identified, and sequencing reads that do not correspond to a pathogen genome also need to be identified.
- One approach would be to obtain a set of pathogen genomes and compare each sequencing read to each of the genomes. If there is a close match, the read can be assigned to the pathogen. Once this is completed, the counts of reads assigned to pathogen genomes can be used as an indicator of which genomes are present in the sample. However, this approach can result in false positives. Many pathogens share some genetic sequence. Other genomes such as human genomes or other organisms may be present in the sample. This can result in false positives where a read is assigned to a pathogen even though it originates from a different genome. This is especially the case when an organism that is related to the pathogen is present in the sample.
- Metagenomic agnostic diagnostic platform is a next generation sequencing (NSG) based diagnostic that can detect nucleic acids from various pathogens, and identify various pathogens present in samples.
- pathogens can include (but are not limited to) bacteria and/or viruses.
- Pathogens can be detected in upper respiratory specimens from patients suspected of infection. Assays can be optimized to improve pathogen identification in clinical samples. To prepare for identification, samples undergo nucleic acid extraction followed by NGS library preparation and sequencing. An internal control which is the Bacteriophage MS2 is added to each sample prior to extraction.
- Results of sequencing can be processed by bioinformatics analysis (or pipelines) which can perform (but not limited to) quality controls, checks for internal controls, and match reads to pathogens.
- workflows to return results can be achieved within 24 hours.
- the metagenomic diagnostic platforms can analyze 5 samples in less than about 12 hours; or 96 or 192 samples in about 24 hours.
- Some embodiments use a MiniSeq Rapid Run Kit that can analyze 5 samples in less than about 12 hours.
- Some embodiments use an Illumina® NextSeq 2000 Sequencer that can analyze about 96 to 192 samples in about 24 hours at a cost of under $50 per sample. When applied to even larger numbers of samples simultaneously analyzed, this price can approach about $10 per sample.
- Many embodiments provide that existing approaches cannot process even close to that many samples and are much more expensive for each sample.
- Metagenomic diagnostic platforms do not perform targeted amplification of pathogen sequence such as (but not limited to) PCR. Instead, the platforms obtain hundreds of thousands of sequencing reads per sample, most of them corresponding to patients and/or host nucleic acids. The platforms then use a bioinformatics pipeline to identify and classify the small fraction of the reads that correspond to pathogen sequence using nuclei acid databases such as (but not limited to) annotated databases, Refseq database, and/or GenBank database. Prior to classification, a cross-reactivity analysis can be performed using curated genome sequences to verify that the pipeline can accurately identify the target pathogens.
- Figure 1 illustrates a process of SwabSeq metagenomics diagnostic platform workflow in accordance with an embodiment.
- the process starts with extracting (101 ) nucleic acid from a sample.
- Samples can be collected using a swab, a Q-tip, a wipe, and/or a cloth, from a source such as (but not limited to) an organ, a tissue, a liquid, a bodily fluid, an organism, a living species, an animal, a human, a patient, an object, a surface, a container, and/or any of a source that may contain pathogens.
- Some embodiments use turbinate swabs; or nasal pharyngeal swabs; or saliva.
- the collected sample can be prepared (if needed) or used directly for nucleic acid extraction.
- Nucleic acid extraction can be performed using various extraction kits.
- Nucleic acid extraction for various specimen types can be performed using instruments such as (but not limited to) Thermo Fisher® KingFisher Apex Instrument (5400930) and Zymo Quick-DNA/RNA Viral Magbead (R2141 ).
- Sequencing libraries can be created with the extracted RNA using the NEBNext® Single Cell/Low Input RNA Library Prep Kit for Illumina® (E6420L).
- a control which is fixed quantity of Bacteriophage MS2 is added to each sample. This control is later used to quantify the amount of pathogen in the sample.
- an additional control of RNA or DNA is added to the extracted RNA or RNA from the sample and used as control and to further quantify the amount of pathogen in the sample as well as potentially identify issues of contamination.
- Construct (102) library with extracted nucleic acids can include the steps: fragmenting and/or sizing the target sequences to a desired length; converting target to double-stranded DNA; attaching oligonucleotide adapters to the ends of target fragments; and quantitating the final library product for sequencing.
- the library preparation can be automated or manual.
- Sequence (103) nucleic acids using the prepared library can be generated using various sequencers such as (but not limited to) Illumina® sequencers, Oxford Nanopore® sequencers, and/or Qiagen® sequencers. Some embodiments use an Illumina® Sequencer (NextSeq 2000) in Illumina® BCL format.
- samples are organized into a 96 well plate where at least 2 of the positions in the plate correspond to controls.
- One control is a no-template control containing nuclease free water.
- One control is a positive control which is a contrived sample containing a mixture of target pathogens.
- Figure 2 illustrates a 96 well plate with the control positions marked in accordance with an embodiment. The control positions are always one of four positions on the 96 well plate (A01 , A02, H07, H08). Prior to RNA extraction, about 10 6 PFU of Bacteriophage MS2 is added to each sample and control as an additional positive control. Some embodiments add a control after nucleic acid extraction.
- FIG. 3 illustrates workflows for the metagenomic diagnostic platform multiplex assay in accordance with an embodiment.
- Collecting (301 ) samples such as extracting nucleic acid from a sample.
- Samples can be swabbed in about 3 mL viral transport medium (VTM) or universal transport medium (UTM).
- the library can be prepared manually (303) or automatically (304).
- Automated library preparation can include (but not limited to) automated Biomek ⁇ 7.
- Sequence (305) the extracted samples using various sequencing techniques or platforms.
- metagenomic diagnostic platforms perform a preprocessing step that utilizes a set of curated pathogen sequences. These sequences can be used to generate simulated reads which are then analyzed using computational analysis in accordance with several embodiments. Some embodiments then identify any reads and/or pieces of reads that are incorrectly classified using the database. Certain embodiments can then mark these regions and will ignore them if they show up in an actual sample. Same analysis approaches can be applied using genomes of organisms that are likely to be present in clinical samples. After performing simulated read analysis, regions of the genome that are ambiguous and/or can lead to errors can be eliminated, such that the accuracy of the computational analysis improves.
- sequencing data can be analyzed using the metagenomic diagnostic platforms that have 3 phases: fastq generation and quality control; read classification into genome of origin; internal controls and diagnostic calling.
- Figure 4 illustrates a process of bioinformatic analysis process in accordance with an embodiment of the invention.
- the SwabSeq metagenomic diagnostic platform bioinformatics analysis process can include the following steps. Convert BCL format data to fastq format data using Illumina® Dragen software. Apply (401 ) quality control by filtering the fastq files using fastp.
- the BCL data can be converted to paired fastq format using the bcl-convert software from Illumina® and a sample sheet created with barcode information obtained from New England Biolab®.
- Quality control can be performed using the fastp software using default settings except for passing the -q 7 option which reduces the sequencing quality level slightly compared to the default setting. This optimization can be determined by experiments over contrived and clinical samples. Reads can be filtered out if:
- the read has a complexity of less than 30%
- the read has less than 15 positions remaining after removing any adapter sequence.
- the result of this process is a pair of qc filtered fastq files for each sample.
- the database can include RefSeq archaea, bacteria, viral, plasmid, human and UniVec_Core from National Center for Biotechnology Information (NCBI). Any pathogen target genomes which are not present in the standard database can be added to the database. Previously, curated pathogen genomes (such as from the ARGOS database) were utilized to validate the pipeline and identify any regions of the RefSeq genome that caused misclassified reads. Any reads that mapped to these regions are removed. Reads that map to multiple target pathogens are removed from the analysis. Line data consisting of counts of unclassified reads and classified reads for human, bacteria, viruses and counts of reads that uniquely map to each respiratory pathogen target are extracted from the kraken 2 results. This line data is evaluated using internal control checks and used to generate diagnostic calls.
- NCBI National Center for Biotechnology Information
- the software kraken 2 is run using default parameters. A set of read counts for a specific set of genomes with taxonomic identifiers in Table 2 is extracted from the results.
- Table 2 lists taxonomic Identifiers with read counts extracted in SwabSeq Metagenomic Bioinformatics Pipeline.
- the taxonomies that correspond to pathogen targets are marked in the Pathogen Target column.
- Some of the pathogen sequences are labeled with a “Pathogen Group” which are closely related sequences that are difficult to distinguish.
- the diagnostic calling algorithm reports the presence or absence of a sequence from the group (e.g., “Influenza A”) since there is not often enough information to identify the actual sequence.
- Pathogen group entries with an asterisk denote sequences where contrived sequences were available to validate the detection of those specific genomes which allow the possibility of the diagnostic algorithm predicting the specific sequence (e.g., Influenza A H1 N1 ).
- results of the previous steps are a count of the number of reads for each taxonomic identifier in Table 1. “classified” is the sum of reads that correspond to any taxonomic identifier and “unclassified” are reads which do not correspond to any taxonomy identifier. Apply (403) three internal control checks to kraken 2 results to verify that a sample is valid:
- Sequence Quality (Internal Control 1 ): Checks the sequence quality by verifying that the percentage of classified reads is greater than 0.5. Passed if the number of classified reads divided by sum of classified and unclassified reads is greater than 0.5.
- Sequence Depth (Internal Control 2): Checks the sequence quality by verifying that the number of classified reads is over 100,000. Passed if the number of classified reads is over 100,000.
- Human Reads (Internal Control 3): Checks by verifying if the number of human reads is greater than 20,000. Passed if the number of human reads is greater than 20,000.
- MS2 Reads (Internal Control 4): Checks by verifying if the number of MS2 reads is greater than 100. Passed is the number of MS2 reads is greater than 100.
- MS2 Reads Ratio to Total Reads Checks by verifying if the ratio of MS2 reads is greater than 0.01. Passed if the number of MS2 reads divided by the number of classified reads is greater than 0.01 .
- Identify (404) pathogens from reads A sample that passes all internal control checks is a valid sample. A sample that fails any one of internal control checks is an inconclusive sample. For any valid sample, a pathogen target is considered present in the sample if the number of reads for the pathogen divided by the number of MS2 reads is greater than 0.01 . For example, if the number of MS2 reads is 1000, the pathogen read threshold is 10.
- the pathogen For any pathogen target below that level, the pathogen is considered not present. For pathogen targets not part of “Pathogen Groups” (Table 2), the pathogen results as “detected” if the pathogen is considered present and “not detected” otherwise. For pathogen targets part of pathogen groups, due to sequence similarity between targets within each group, the specific pathogen present is difficult to distinguish. For these targets, if any pathogen from a group is above the threshold, the pathogen group results as “detected” and otherwise the pathogen group results “not detected”. In practice, typically multiple sequences from the pathogen group are above the threshold.
- pathogen targets listed with an asterisk in their pathogen group e.g., Influenza A H1 N1
- contrived samples were available within the pathogen group to enable identification of sequences within the group.
- the specific pathogen results “detected” if 95% of the reads within the group are from the specific genome.
- NTC Control Passed if the MS2 reads divided by the total classified reads divided by the total number of classified reads if greater than 0.1 for all of the NTC positions.
- the positive control can either originate from a known positive sample or a contrived sample which has a mixture of pathogens.
- the SwabSeq Metagenomic Diagnostic Platform is a meta-genomics based untargeted NGS diagnostic designed to detect nucleic acid from a pathogen in upper respiratory specimens from patients suspected of infection. Samples undergo nucleic acid extraction followed by NGS library preparation and sequencing. Results of sequencing are then processed by a bioinformatics pipeline which performs quality control, checks for internal controls, and matches reads to pathogen.
- the SwabSeq Metagenomic Diagnostic Platform does not perform any targeted PCR amplification of pathogen sequence. Instead, the platform obtains hundreds of thousands of sequencing reads per sample, most of them corresponding to patient/host nucleic acid.
- the platform uses a bioinformatics pipeline to identify and classify the small fraction of the reads that correspond to pathogen sequence using the RefSeq database. Prior to this classification, a cross-reactivity analysis is performed using curated genome sequences to verify that the pipeline is capable of accurately identifying the target pathogens.
- SwabSeq metagenomic diagnostic platform uses sample types such as (but not limited to) mid turbinate swabs, nasal pharyngeal swabs and saliva. Sequencing data is analyzed using SwabSeq Metagenomic Diagnostic Pipeline which has 3 phases: Fastq generation and quality control, read classification into genome of origin, and internal controls and diagnostic calling.
- Some embodiments implement Fastq generation and quality control processes.
- the Illumina® NextSeq 2000 generates the sequencing data in raw BCL format.
- the BCL data is converted to paired fastq format using the bcl-convert software from Illumina® and a sample sheet created with barcode information obtained from NEB®.
- Quality control can be carried out using the fastp software using default settings except for passing the - q 7 option which reduces the sequencing quality level slightly compared to the default setting. This optimization can be determined by experiments over contrived and clinical samples.
- Using fastp reads are filtered out if: 40% of the positions have a read quality of less than 7; the read has a complexity of less than 30%; the read has less than 15 positions after removing any adapter.
- the result of this process is a pair of qc filtered fastq files for each sample.
- An additional filtering step is performed to identify filtered reads that ambiguously map to both a pathogen target and either a common contaminant or organism that is often present in respiratory samples.
- This filtering is performed using the bbduk.sh tool from the BBMap software package, bbduk.sh is run against a “contaminant kmer list” which is a set of kmers from portions of the genomes which are shared between targets and contaminants and have been shown to classify incorrectly, bbduk.sh removes any read that contains one of these kmers.
- contaminant kmer list is a set of kmers from portions of the genomes which are shared between targets and contaminants and have been shown to classify incorrectly
- bbduk.sh removes any read that contains one of these kmers. The process to identify the kmers which may cause incorrect classification is described below.
- a curated genome can be identified. This is typically a RefSeq genome that is marked as “Complete” or from a reliable source.
- RefSeq sequence “NC_045512.2 Severe acute respiratory syndrome coronavirus 2 isolate Wuhan-Hu-1 , complete genome.”
- NC_045512.2 Severe acute respiratory syndrome coronavirus 2 isolate Wuhan-Hu-1 complete genome.
- the genomes for organisms commonly occurring in nasal samples are included in Table 3.
- Some embodiments use the wgsim to generate simulated reads from each organism and run them through the pipeline. For pathogen targets, any reads that are either incorrectly classified as another target pathogen or incorrectly classified as either human or bacteria can be identified. For common organisms, any reads which are incorrectly classified as a target pathogen can be identified. Some embodiments look for kmers of length 31 that are common among the incorrectly classified reads and any kmer that occurs at least 3 times are included in a “contaminant kmers list.” The intuition is that these commonly occurring kmers are characteristic of regions of the genomes that are inherently ambiguous and may lead to incorrect classifications. This list can be used to filter any reads which contain one of these kmers.
- Some embodiments implement the processes for read classification into genome of origin. Reads from each sample are classified into a genome of origin using the kraken 2 software using the “standard” database which consists of RefSeq archaea, bacteria, viral, plasmid, human and UniVec_Core from NCBI. Pathogen target genomes which are not present in the standard database can be added, kraken 2 is run using default parameters. A set of read counts for a specific set of genomes with taxonomic identifiers in Table 1 is extracted from the results.
- Some embodiments provide the processes for internal controls and diagnostic calling.
- the results of the previous steps are a count of the number of reads for each taxonomic identifier in the table above, “classified” is the sum of reads that correspond to any taxonomic identifier and “unclassified” are reads which do not correspond to any taxonomy identifier.
- Figure 5 shows sample output for 20 COVID-19 samples analyzed using the SwabSeq Metagenomic Diagnostic Platform in accordance with an embodiment. Each line is the output for one sample. The numbers in each column are the number of reads in each category.
- sample 1 75,441 reads are not able to be classified while 1 ,354,221 reads are classified. Of these reads, a little over 1.3M are human reads, a little over 40k are bacteria reads and 2,466 are viral reads. Of those reads 2,180 are from SARS-CoV-2.
- each of the three internal controls are passed: 94.7% of the reads are classified to pass Internal Control 1 ; more than 100,000 reads are classified to pass Internal Control 2; more than 20,000 reads are classified as human to pass Internal Control 3.
- the fraction of SARS-CoV-2 reads is greater than 0.00001 for samples 1 , 2, 11 and 20 which is concordant with PCR results (last column). Interestingly, samples 5 and 10 appear to indicate RSV infection. Sample 18 also shows a low presence of SARS-CoV-2 which is both below the threshold for calling a SARS- CoV-2 positive and maybe below the limit of detection of the PCR test.
- RNA quality can result in low quality sequencing.
- Experiment with the ThermoFisher® Extraction Kit and the Illumina® Library Preparation Kit on 10 contrived samples resulted in low quality RNA in the samples from the sequencing metrics.
- Table 4 shows the number of reads in each sample before and after sequencing. 4 samples failed and the filtering removed anywhere from 40% - 75% of the reads for all but one sample due to RNA quality issues. None is apparently different between the samples that succeeded and failed that can be observed from inspection of the physical sample which suggests that low RNA quality is an inherent property of nasal swab samples.
- RNA quality after nucleic acid extraction can cause issues with the library preparation (Illumina® TruSeq) resulting in many adapter dimers.
- Figure 6 illustrates a histogram of bases per read cycle (read position) showing the adapter sequence in a highly enriched pattern.
- RNA/DNA extraction kits experiments are tested out.
- the Zymo Research® Quick-DNA/RNA Viral MagBead showed superior performance.
- Some embodiments use a different library preparation protocol which is optimized for lower quality input. NEBNext® Single Cell/Low Input RNA Library Prep Kit for Illumina® (E6420L) is used.
- depletion technologies can potentially improve the scalability and allow for more samples being processed. In most cases, depletion technologies further damage the poor RNA quality of the samples and perform worse than not using the technology. For this reason, some embodiments proceed without using depletion technology (with an exception below for using Jumpcode on the more scalable workflow).
- a set of samples are prepared that can lead to a complete limit of detection with enough samples for both the preliminary limit of detection and the confirmatory limit of detection.
- 2 plates of contrived samples are prepared with 12 samples at each of the following concentrations: 16,000 GCE/ml, 8,000 GCE/ml 4,000 GCE/ml, 2,000 GCE/ml, 1 ,000 GCE/ml, 500 GCE/ml, 250 GCE/ml, and 125 GCE/ml and a plate of 96 samples with 0 concentration.
- Preliminary LoD and confirmatory LoD testing determines the lowest detectable concentration of SARS-CoV-2 at which approximately 95% of all (true positive) replicates tested positive. Using the contrived samples produced as described above, the preliminary LoD is defined as the lowest concentration where 3 of 4 replicates are identified correctly. The results of the preliminary LoD testing are summarized in Table 6.
- LoD confirmation testing was performed by testing twenty (20) replicates at the preliminary LoD concentration determined above (1000 GCE/ml). Acceptance criteria for confirmation of the LoD was that at least 95% of the replicates (> 19/20) test positive. The LoD was confirmed to be 1000 GCE/mL for mid-turbinate samples run. Several embodiments also evaluated twenty (20) replicates at 0 concentration, and the LoD was confirmed to be 0 GCE/mL.
- Clinical samples are remnant samples from COVID-19 testing submitted to the SwabSeq laboratory. Samples from individuals with symptoms consistent with COVID-19 submitted for COVID-19 testing and transported to the laboratory at 2 °C. After samples were accessioned but prior to samples being analyzed with the SwabSeq COVID-19 Diagnostic test, 100 pL is removed from each sample tube to be extracted using the SwabSeq Agnostic Diagnostic Test Extraction procedure. The SwabSeq COVID-19 Diagnostic test only requires no more than about 100pL of sample volume and a total of about 750 pL of sample is collected from each patient and thus the removal of the sample does not affect the results of the SwabSeq COVID-19 Diagnostic test.
- RNA extract is then stored at -80 °C which is consistent with recommendations for the NEBNext® Single Cell/Low Input RNA Library Prep Kit for Illumina® (E6420L).
- the advantages of this approach to obtain remnant samples is that the samples are extracted as fresh samples and not freeze/thawed samples, and the SwabSeq COVID-19 Diagnostic Test performs heat extraction which may damage nucleic acids.
- kmers of length 31 are extracted from the reads and any kmers that occur at least 3 times in a “contaminant kmers list” can be added.
- the intuition is that these kmers are characteristic of regions of the genome which are inherently ambiguous as they lead to multiple misclassified reads.
- the pipeline filters out any read that contains such a kmer.
- Table 8 shows this approach for 3 organisms, SARS-CoV-2, RSV and Haemophilus influenzae. For each organism, 1 ,000,000 reads were simulated and run through the pipeline. The first three columns show the results of the simulations. Entries in the table with an asterisk (*) are misclassified reads. Pathogen targets with 0 reads for these three organisms are omitted for clarity.
- the SARS-CoV-2 simulations resulted in 339 misclassified reads which then resulted in 51 kmers which were then added to the “contaminant kmer list.”
- the RSV simulations added another 204 kmers to this list as well.
- the fourth column illustrates the effect of this approach by performing the same simulation and generating 1 ,000,000 reads for SARS-CoV-2 but filtering the reads that match a kmer. The results are that 4,492 reads were filtered out and the number of human and SARS reads were reduced compared to without filtering.
- At least 50% of the simulated reads are correctly classified by the pipeline; 2).
- No more than 0.1 % reads are misclassified to either Human, Bacteria or another pathogen target.
- Table 13 shows this validation for 3 pathogen targets. Without the additional contaminant kmer filtering, the pipeline is valid for those three targets.
- Table 9 shows the validation for 4 commonly occurring organisms in nasal samples.
- Table 9 Example of simulation studies of common organisms in nasal samples. 1 ,000,000 reads are simulated for each common and classified using the pipeline. Reads that are classified as pathogen targets are considered misclassified. In these examples, no reads are misclassified. [00111] Several embodiments show approaches to clean errors in the database by utilizing the technique above. Any unexpected results where reads from the simulated genomes match in other genomes suggest the possibility of the errors. Some embodiments developed manual approaches to evaluate these errors as well as automated approaches using machine learning or artificial intelligence (Al) techniques to identify database errors.
- Al artificial intelligence
- NTC no-template control
- positive control to verify the functioning of the assay.
- additional positive controls use “additional positive controls” in each sample that can be used to verify that the assay is working correctly.
- Bacteriophage MS2 in each of the samples as an additional positive control.
- a no-template control is the assay applied to the extraction buffer without a sample present.
- the purpose of a no-template control (NTC) is to determine the amount of background nucleic acids in the equipment and reagents that are detectable in the assay.
- the NTC can identify if the assay either has very high levels of contamination or is not functioning properly. After processing the assay, the NTC position should not detect any pathogens.
- an NTC is implemented by including nuclease free water (NFW) in a well prior to extraction.
- NFW nuclease free water
- Some embodiments include at least one positive control in each run of the assay.
- positive controls There are two types of positive controls. The first is a sample which is a known positive verified by an alternative technology (such as a sample verified with the Roche ePlex or a saliva sample with a positive verified by the SwabSeq COVID-19 Diagnostic Platform (targeted PCR NGS test). For this type of control, the assay identified the correct pathogen for this sample. This positive control is used when performing validation experiments when measuring the performance of the assay against other assays.
- the second type of positive control is a contrived control which is used when applying the assay to unknown samples. A mixture of multiple Twist pathogen controls is created and included on the plate. The mixture is varied in each control so that many of the target pathogens are included in some of the positive controls. For this type of control, the assay correctly identified all the components of the mixture.
- Some embodiments include a known quantity of the Bacteriophage MS2 in each sample.
- the purpose of this control is to verify that the assay is working correctly by checking to see if enough reads from MS2 are observed as well as to quantify the amount of a pathogen if it’s detected in the sample. Since a known quantity of MS2 is added to each sample, the ratio of the number of pathogens reads vs the number of MS2 reads is a robust measure of the amount of pathogen in the sample. Specifically, 10 6 PFU per mL of MS2 is added in each sample prior to extraction.
- Some embodiments provide experiments validating inclusion of MS2 in assay. Specifically, there are three aspects of the assay related to MS2 and NTC: 1 ). How much MS2 should include in each sample; 2). What substance should be used for the NTC; 3). What threshold should be used for detection of a pathogen relative to the amount of MS2. [00117] To address the first two aspects, the MS2 control and various options for the no-template control (NTC) are optimized on a sequencing run. Two quantities of MS2 (10 6 PFU/mL and 10 3 PFU/mL) are used. Both nuclease free water and DNA/RNA Shield are used as the NTC. Fresh saliva samples where MS2 is added.
- NTC no-template control
- MS2 positions in sequencing run (Table 10) corresponding to RNA extraction. Column 7 positions correspond to saliva samples with MS2 added prior to extraction. Column 8 corresponds to NTC controls. The number in Table 16 is the quantity of MS2 (PFU/mL) added to the sample. True positions show the positions expected to observe
- MS2 is consistently detected in both clinical and NTC samples if added at 10 A 6 PFU/ml concentration and this is why 10 A 6 PFU/ml concentration of MS2 is used in the assay.
- NTC nuclease free water
- DNA/RNA shield in positions C08, D08, G08 and H08
- the third aspect of the assay related to MS2 is the threshold.
- the detection threshold for pathogens is the ratio of the number of pathogen reads to the number of MS2 reads. Since the same amount of MS2 is added to each sample, this threshold is much more robust and closer to an actual pathogen concentration than just the number of reads of the pathogen.
- the threshold of 0.01 which is high enough to filter out any contamination or other artifacts but low enough to identify all the clinical samples is used.
- Several embodiments utilize contrived samples which were created using Twist Respiratory Virus Controls. Contrived samples are prepared by first obtaining negative samples from at least three individuals using mid turbinate swabs which were collected into 750 pl of 0.9% saline.
- Samples are then pooled together to create pooled negative nasal samples.
- Contrived samples are prepared by adding a quantity of the appropriate Twist Respiratory Virus Control to achieve a desired concentration.
- Positive samples were prepared at different concentrations typically 4000 GCE/ml, 2000 GCE/ml, 1000 GCE/ml, 500 GCE/ml, 250 GCE/ml and 125 GCE/ml, but in some cases in larger ranges.
- Table 12 shows the layout of the first LoD plate.
- the LoD is already comparable to multiplex-PCR tests such as the Roche ePlex. Because the ranges of limit of detection is in the same range for all pathogens, after modifying the assay and performing new LoD experiments, only the concentrations 4000 GCE/ml, 2000 GCE/ml, 1000 GCE/ml, 500 GCE/ml, 250 GCE/ml and 125 GCE/ml are considered for the experiments and can put 4 pathogens in each plate.
- the plate is shown in Table 14 which contains Rhinovirus, Influenza H1 N1 , Mumps and Parainfluenza 1 .
- the plate contained Enterovirus D68 and Influenza B.
- the plate contained Parainfluenza 4, Human Coronavirus NL63, Human Coronavirus 229E and Measles.
- the plate contained Influenza H3N2 and Human Coronavirus OC43.
- Twist provides an approximate concentration for their materials.
- the Twist controls are synthetically generated as opposed to heat inactivated.
- the improvement is specific to the batch of controls compared to different controls.
- NGS technology enables one to directly quantify the amount of carryover from previous runs.
- the LoD plates are utilized to conduct a carryover study.
- the two LoD plates contain very high viral loads of distinct pathogens. They were both generated by the same operator using the same equipment and run 2 days apart.
- the first plate has a layout shown in Table 19.
- the second plate has a layout shown in Table 15.
- Rhinovirus, Mumps, Influenza H1 N1 and Parainfluenza 1 are present. These two plates allow the direct measurement of carryover by quantifying the number of reads corresponding to pathogens from the first plate detected on the second plate. Table 16 shows the number of reads of pathogens from the first plate on the second plate which is evidence of carryover. The carryover is not in the same wells as the samples were in the previous plate and this is expected as the carryover can occur because of contamination from any of the processes of processing the samples.
- the amount of carryover can be estimated by computing the ratio of the number of reads observed for each pathogen in the second plate divided by the number of reads observed in the first plate. Table 17 shows the results for each pathogen and an overall estimate. Carryover from over 275 million pathogen reads in the second plate only generated 122 pathogen reads in the first plate. Actual plates of clinical samples will contain far fewer pathogen reads (typically under 1 million reads) and thus the amount of carryover will be even smaller than here. If assuming 1 million pathogen samples in a previous run and the same carryover rate estimated here, the amount of carryover would be in the range of 1-2 reads. The carryover contamination rate is much smaller than other sources of errors and not likely a contributor to false positives.
- Cross-contamination is a fundamental problem for NGS Diagnostic technologies.
- the rate of cross-contamination can be estimated by utilizing samples which are known positives for pathogens. In one plate, there are 5 samples which were confirmed to be positive for Parainfluenza 3 using a Roche ePlex Respiratory Panel comparator. For the 5 known Parainfluenza 3 positives, a dramatic range of the number of pathogen reads for each positive is shown in Table 18.
- Figures 6A through 6C show the results of analysis of the plate both using the manual library preparation protocol and the automated library preparation protocol. In the library preparation manual there is a much higher level of cross contamination compared to the automated library preparation.
- Figure 6A shows 5 True Positive positions of Parainfluenza 3.
- Figure 6B shows predicted positions using manual library preparation protocol where blue are true positive predictions and red are false positives due to crosscontamination.
- Figure 6C shows predicted predictions using automated library reparation protocol. Positions are highlighted only if there are more than 10 reads present.
- Table 19 Cross-contamination analysis for true positives of Parainfluenza 3 pathogen. Each row shows the position of the true position which is the potential source of crosscontamination, the neighboring position and the number of reads for both the source and neighboring positions for manual and automated library preparation protocols. The rate is estimated by the ratio of the source reads to neighbor reads. Only positions where cross-contamination occurred are shown in the table.
- a precision study is performed by applying the assay on the same set of samples twice using a different sequencer.
- One run was sequenced on a NextSeq 2000.
- the same set of samples was repeated in the assay including new RNA extraction, library preparation and sequencing.
- the second run utilized a NovaSeq X Plus.
- Table 20 shows the concordance of positives between runs.
- Table 21 Design of Analytical Specificity Experiment.
- the plate contained 18 samples contrived samples, each with three pathogens that were created using Twist materials. 9 different combinations were used with 2 samples per combination. The exact combinations of pathogens in each sample are shown in the table.
- the experiment evaluated combinations of similar pathogens (F9, F10), combinations of different pathogens (F11 , F12, G7, G8, G11 , G12, H7, H8, H9, H10) and combinations of two similar and one different pathogen (F7, F8, G9, G10, H11 , H12).
- Table 22 shows the results of three representative samples F9, F7 and H9 of three similar pathogens, two similar and one different and three different pathogens respectively. As shown, the assay can identify the correct pathogens within the mixtures. The assay correctly identified the pathogens in each of the 18 samples.
- Table 22 Results of Analytical Specificity Samples. Results of three samples from Figure 12. As shown, the assay correctly identifies the components of the mixture of pathogens. [00140] Some embodiments perform an interfering substances study considering 5 substances that can potentially interfere with the assay: Cough Drops, Mouth Wash, Cough Syrup, Bovine Mucin, Sore Throat Spray. For each substance, 3 contrived positive samples at 2x the LoD were generated and added to the substance. The substance was added to 3 negative samples. The assay was applied to the contrived samples. Results are shown in Table 23. None of the substances interfered with the assay as expected.
- Table 23 Interfering Substances. Positives are contrived samples that were generated at 2x LoD prior to adding the substance. ‘Positive generated from Twist Influenza H3N2. “Positive generated from Twist Coronavirus OC43. The number in the table is the number of samples where the assay made a correct prediction.
- Some embodiments perform a stability study using saliva samples to evaluate how consistent the assay results are by varying the time between sample collection and processing.
- the first step of the Zymo Quick-DNA/RNA Viral Magbead is to add the RNA/DNA Shield reagent. It is well established that RNA/DNA Shield preserves RNA for long periods of time stored at ambient temperature and it is well known that RNA analysis on samples which contain RNA/DNA Shield are consistent irrespective of how long analysis was performed after the RNA/DNA Shield was added. Thus, the key question in the stability study is to understand the effect on the assay of how soon after sample collection is RNA/DNA Shield added.
- Some embodiments perform a stability study using saliva samples submitted for COVID-19 testing. Prior to COVID-19 testing, 3 aliquots were obtained from 96 samples. RNA/DNA Shield was added to one set of aliquots right away. These samples are referred to as the “Immediate” samples. RNA/DNA Shield was added to the second aliquot after 3 to 6 days stored at room temperature. These samples are referred to “Delayed” samples. A third set of aliquots were frozen and thawed prior to RNA/DNA shield being added. These samples are referred to as “Frozen” samples. All three sets of 96 samples (288 samples total) had their libraries constructed together and sequenced together in order to minimize any additional sources of variance outside the stability study. [00143] Table 24 shows the stability study results. Each entry shows a pathogen identified in the immediate samples, and the corresponding pathogen identified in the delayed samples. Both the number of reads and MS2 ratio are provided.
- the LoD plates contain the following pathogens: First plate contains Enterovirus D68 and Influenza B; second plate contains Parainfluenza 1 , Influnenza H1 N1 , Mumps and Rhinovirus; third plate contains Parainfluenza 4, Human Coronavirus NL63, Human Coronavirus 229E and Measles; fourth plate contains Influenza H3N2 and Human Coronavirus OC43.
- Some embodiments measure the accuracy of the pipeline by computing the number of misclassified reads or reads that do not match the contrived pathogens in each run. There are 0 misclassified reads for the first plate. For the second plate, the only reads that are misclassified are some Influenza H1 N1 reads which are classified as Influenza H3N2. However, the number of reads is below the 5% threshold which is used to identify pathogens within pathogen groups so will have no effect on the overall predictions.
- the only misclassified reads are the 122 reads in Table 25, or the reads classified to Parainfluenza 4, Measles, Coronavirus NL63 and Coronavirus 229E pathogens which is likely due to carryover rather than a bioinformatics pipeline error. No analysis was done on the fourth plate since part of that run contained mixtures of different pathogens and every pathogen was present in that plate. Table 25 summarizes these results.
- Table 26 Summary of sets of samples in clinical validation.
- the second-row samples utilize a comparator that can detect multiple pathogens.
- the other two sets utilize a comparator that can only detect SARS-CoV-2 but have the advantage that they are prospective studies in that they are a set of samples randomly selected from samples submitted for COVID-19 testing.
- the results of detecting pathogens other than SARS-CoV-2 in these two other sets include three different sample types (Mid-turbinate swab, nasal pharyngeal swab and saliva).
- Some embodiments perform a clinical validation.
- a total of 460 remnant samples which were selected based on the results of analysis by the Roche ePlex Respiratory Pathogen Panel 2 were collected. 210 samples were positive for a pathogen on the ePlex panel and 250 were negative. Most of the samples were nasopharyngeal swabs, but other sample types were also included in the set of samples. The sample type distribution is shown in Table 27.
- Table 28 Results of positives from samples. 210 positive samples and 250 negative samples were obtained, and the table shows results by pathogen for the 204 samples that passed QC and were analyzed. *The process of selecting samples for analysis may have included samples that are positive for either SARS-CoV-2 or Influenza A as negative samples. In analysis of the negative samples, only pathogens other than SARS-CoV-2 and Influenza A are considered.
- Saliva as a sample type has many advantages, yet until the COVID-19 pandemic, saliva was not regularly utilized for infectious disease diagnostics. Many embodiments utilize the SwabSeq Metagenomic Diagnostic Platform to demonstrate that many pathogens can be detected in the saliva sample type. About 1 ,152 samples were analyzed. Table 30 shows the pathogens other than SARS-Cov-2 identified in the samples.
- Automation can enable processing many samples simultaneously. Specifically, a plate of 96 samples or multiple plates of 96 sample can be run at the same time. Each step is designed to process one plate at a time.
- Two sequencing protocols one protocol uses the NextSeq 2000 P2 flowcell that can sequence 96 samples and one protocol uses the NextSeq 2000 P3 flowcell that can sequence 3 sets of 96 samples (288 total samples).
- Some embodiments provide accessioning and nucleic acid extraction automation. Nucleic acid extraction has been performed using the Thermo Scientific® KingFisher Apex system. The extraction kit from Zymo Research® is compatible with this system. The output of the Kingfisher is a plate of the RNA extracts of the 96 input samples.
- the racks are scanned using a barcode scanner which captures the barcodes and positions of each tube in the rack.
- 250 pL are aliquoted from each sample to a 96 well plate and the captured barcodes and positions are stored long with the plate barcode in the SwabSeq Laboratory Information Management System (LIMS).
- LIMS Laboratory Information Management System
- This plate can then be processed in KingFisher Apex system for extraction while the original samples can follow the SwabSeq COVID-19 Diagnostic test workflow. This process can extract nucleic acids from samples.
- Some embodiments provide sequencing library preparation.
- the specific library preparation protocol such as NEBNext® Single Cell/Low Input RNA Library Prep Kit can be used.
- the first step is to modify the existing Biomek i7s to support the protocol. Cold blocks, magnetic modules and integration kits are used.
- a Thermocycler is not used on the Biomek which allows for a second plate to be processed while the first plate is in the thermocycler.
- the assay has 6 major steps, and, in each step, there is a combination of user time to set up the Biomek deck and add reagents, thermocycler incubations and Biomek automation. Table 31 shows an overview of the automated protocol steps.
- the total time of the library preparation assay is approximately 9 hours for 1 plate. If a 2nd plate is processed at the same time, most of the steps can be completed while the first plate is in the thermocycler. The only exception is the cDNA cleanup which performs multiple incubations on the Biomek and repeatedly uses the magnetic module. Overall, as shown in the table, a second 96 well plate requires an additional 2 hours.
- Table 31 Summary of Biomek automation for sequencing library preparation. There are 6 steps for the protocol and (almost) each step involves user preparation of the Biomek and reagents, Biomek actions and thermocycler incubations. Most of the user actions can occur during incubations and do not add any time to the protocol. Since the thermocycler is not on the Biomek deck, a second plate can be processed while the first is in the thermocycler.
- each resulting library is quantified using a Qubit and all libraries are pooled together into a single tube at a proportion so that the resulting sequencing will have a similar number of total reads for each sample.
- This step is automated using other liquid handles.
- An alternative approach to the Biomek i7 is a manual library preparation for a small number of samples which also requires a limited amount of hands on time.
- the manual library preparation for 5 samples can be performed in approximately 6 hours. 5 samples are the number of samples that can be run on the MiniSeq Rapid kit.
- This process can also be automated on a smaller automated liquid handler such as an Opentrons OT-2.
- This process can also be automated up to 8 samples on a microfluidic liquid handler. The process can be automated for many more than 8 samples on a microfluidic liquid handler as microfluidic liquid handlers increase in their capacity.
- Table 32 shows the runtimes of relevant Illumina flow cells for the final configurations.
- the three flow cells (MiniSeq Rapid 50 bp, NextSeq 2000 P2 2 x 50 bp and NextSeq P3 1 x 50 bp) are listed.
- the MiniSeq Rapid 50 bp and NextSeq P3 1 x 50 bp flow cells only collect 50 bp per read but this is enough for our application.
- the MiniSeq Rapid flow cell can generate 100 bases instead of 50 and the NextSeq P3 2 x 50 bp is practically the same cost as the NextSeq 2000 P3 1 x 50 bp.
- the only downside of these two flow cells is that they have longer runtime. Given that collecting this additional data was effectively free, the full 100 bp from the MiniSeq Rapid and NextSeq 2000 P3 2 x 50 bp flow cells were collected.
- Table 33 Concordance of analysis of 50 bp vs 100 bp showing that only 50 bp are necessary to identify viral pathogens.
- Table 34 Length of time necessary for the bioinformatics pipeline to analyze the data generated by each configuration. For the two NextSeq configurations, after BCL conversion, sample processing is parallelized on Amazon Web Services.
- Sample to Answer Time For each step of the assay which include: nucleic acid extraction; the 6 steps of sequencing library preparation; quantification, normalization and pooling; sequencing; and bioinformatics.
- the time of each step can be decomposed into two parts, a preparation time and a sample to answer time contribution.
- the preparation step involves preparing reagents and equipment for the step and can be performed either before analysis starts or during the previous step.
- the sample to answer time depends on the number of samples being processed and is summarized in Table 35. The sample to answer for 3 different configurations: 96 samples, 192 samples and 5 samples.
- Table 35 Sample to Answer Time for the three different configurations. Each step is broken down into preparatory time which are parts of the protocol that can be performed either before the samples are analyzed or during the previous step.
- the third configuration processes 2 sets of 96 sample plates on a single Biomek i7.
- One constraint is that the cDNA cleanup step is performed on the Biomek deck and has many room temperature incubations. Since the deck is occupied throughout the step, it is impossible to overlap the two plates.
- a second constraint is that the result of the Fragmentation and End Prep step is not stable, and the Adapter Ligation and cleanup step must be started immediately afterwards.
- the simplest way to address this constraint is to perform both steps for one plate before performing both plates for the next plate.
- Figures 7A and 7B show diagrams of how two plates are analyzed by single Biomek. We note that both plates are extracted simultaneously on different KingFisher systems.
- FIGs 7A and 7B illustrate interleaving of 2 96 sample plates on 1 Biomek i7. Most steps have a long thermocycler incubation at the end of the step which allows a second plate to be processed on the Biomek during the incubation. However, cDNA cleanup (L3) has no off-deck incubation step and the plates must complete the full step before releasing the deck. In addition, Fragmentation and End Prep (L4) is not stable so much be immediately followed by Adapter Ligation and cleanup (L5).
- Figure 7A shows current automation workflow. The extra time for the second plate is a little bit of time at the beginning and then the doubling of the L3, L4 and L5 steps which are performed serially.
- Figure 7B shows automation workflow that is optimized. L4 and L5 are interleaved since L4 ends with a long thermocycler incubation and L5 for the other plate can complete during that time.
- Each row is a different pathogen: (a)-(c) Human Coronavirus, (d)-(f) Parainfluenza 1. (g)-(i) Human Metapneumovirus.
- the first column (a), (d), (g) are locations of positives confirmed by Roche ePlex.
- the second column (b), (e), (h) are locations predicted by the assay using the manual library preparation protocol.
- the third column are locations predicted by the assay using the automated library preparation protocol. Blue positions are true positive predictions and red are false positives.
- Many embodiments provide methods to scale to a much larger number of samples.
- the bottleneck for the number of samples is not the capacity of the sequencer, but the effort to construct a library. This process takes multiple hours and has many steps. Each step has to be performed for each sample.
- molecular barcodes can be added to each sample at the first step, the samples can be mixed together, and all of the remaining steps can be performed on the mixture.
- Some embodiments implement this method using the BRB-Seq library. Some embodiments repeat the clinical and analytical validations using this technique. The technique is not as efficient as the standard approach and can require a bit more sequencing than the standard. Using this approach, several embodiments can find viral pathogens in the samples.
- the advantage of this approach is that the number of samples to be processed can be substantially increased and the time it takes to run the samples can be substantially decreased. Some embodiments can combine about 96 or 384 samples into a single mixture and process them all together. If this approach is combined with the automation approach that can prepare 192 samples described above, some embodiments can then process 192 mixtures of 384 samples or over 70,000 samples in close to 24 hours.
- This mixture of samples has improved performance when depletion approaches are applied in accordance with many embodiments. Specifically, some embodiments show that the mixture of sample approach benefits from CRISPR depletion unlike the standard approach.
- Some embodiments provide that the methods can lead to full reconstruction of the virus genomes using sequencing assembly. Several embodiments demonstrate this by showing that we these methods can reconstruct viruses that were unknown to be present in the sample at the time.
- the methods can be combined with targeted next generation sequencing diagnostics in the same sequencing run.
- the Illumina MiniSeq® 5 sample run can also process several hundred targeted diagnostics.
- Example 1 A method of detecting a pathogen comprising, obtaining a plurality of nucleic acids from a sample; sequencing the plurality of nucleic acids to obtain a plurality of reads; analyzing the plurality of reads using a bioinformatic process; and identifying the pathogen from the analyzed reads.
- Example 2 The method of example 1 , further comprising obtaining the sample using at least one of: a swab, a Q-tip, a wipe, and a cloth.
- Example 3 The method of example 1 or 2, wherein the sample is a tissue, an organ, a bodily fluid, an object, a surface, or a container.
- Example 4 The method of example 1 , or 2, or 3, wherein the sample is at least one of: a turbinate swab, a nasal pharyngeal swab, and saliva.
- Example 5 The method of any one of examples 1 to 4, wherein the sample is a clinical sample; wherein the clinical sample is at least one of: urine, a tissue, a skin swab, bronchoalveolar lavage (BAL), sputum, and blood.
- the sample is a clinical sample; wherein the clinical sample is at least one of: urine, a tissue, a skin swab, bronchoalveolar lavage (BAL), sputum, and blood.
- BAL bronchoalveolar lavage
- Example 6 The method of any one of examples 1 to 5, wherein the sample is an upper respiratory specimen or a lower respiratory specimen from a subject.
- Example 7 The method of any one of examples 1 to 6, wherein the obtaining step obtains the plurality of nucleic acids using at least one nucleic acid extraction kit.
- Example 8 The method of any one of examples 1 to 7, further comprising adding a known quantity of Bacteriophage MS2 to the sample.
- Example 9 The method of any one of examples 1 to 8, further comprising constructing at least one next generation sequencing library.
- Example 10 The method of any one of examples 1 to 9, wherein the at least one next generation sequencing library is at least one of: a DNA library, and an RNA library.
- Example 11 The method of any one of examples 1 to 10, wherein the sequencing step uses the at least one next generation sequencing library.
- Example 12 The method of any one of examples 1 to 11 , wherein the analyzing step comprises matching a plurality of sequencing reads to at least one pathogen sequence in a database.
- Example 13 The method of any one of examples 1 to 12, further comprising generating a plurality of simulated reads using a set of curated pathogen sequences; analyzing the plurality of simulated reads using the database to check if the database contains an error; and eliminating a plurality of false reads in the database by analyzing the plurality of simulated reads.
- Example 14 The method of any one of examples 1 to 13, further comprising using the plurality of reads of the pathogen and a plurality of reads of a control added prior to nucleic acid extraction or a plurality of reads of a control added after nucleic acid extraction to quantify an amount of the sample.
- Example 15 The method of any one of examples 1 to 14, wherein the control is Bacteriophage MS2.
- Example 16 The method of any one of examples 1 to 15, wherein the pathogen is a pathogenic organism, a bacterium, a virus, or a fungus.
- Example 17 The method of any one of examples 1 to 16, further comprising processing a plurality of samples at a time using automation.
- Example 18 The method of any one of examples 1 to 17, further comprising processing a plurality of samples at a time using automation by interleaving a plurality sets of samples to increase a number that is processed on an automated instrument.
- Example 19 The method of any one of examples 1 to 18, further comprising processing a plurality of samples using at least one microfluidic liquid handler.
- Example 20 The method of any one of examples 1 to 19, wherein automation eliminates cross-sample contamination when processing the plurality of samples.
- Example 21 The method of any one of examples 1 to 20, further comprising adding at least one molecular barcode to each of the plurality of samples, and mixing the plurality of samples.
- Example 22 The method of any one of examples 1 to 21 , wherein more than 90 samples are processed together at a time.
- Example 23 The method of any one of examples 1 to 22, wherein more than 200 samples are processed together at a time.
- Example 24 The method of any one of examples 1 to 23, wherein more than 1000 samples are processed together at a time.
- Example 25 The method of any one of examples 1 to 24, wherein the preparation for sequencing and analyzing are performed on a microfluidic liquid handler.
- Example 26 The method of any one of examples 1 to 25, further comprising a depletion process configured to improve a performance of the mixture of barcoded plurality of samples.
- Example 27 The method of any one of examples 1 to 26, further comprising combining a plurality of targeted diagnostic samples with a plurality of metagenomic diagnostic samples in a same sequencing process.
- Example 28 The method of any one of examples 1 to 27, further comprising using a nucleic extraction process that is optimized for extracting bacteria and viruses compared to host nucleic acids.
- Example 29 The method of any one of examples 1 to 28, wherein sample to result using short read sequencers is completed in less than 11 hours.
- Example 30 A method of a bioinformatic analysis comprising, generating a plurality of simulated genomic reads using a plurality of pathogen genomic sequences; eliminating a plurality of false reads in a database by analyzing the plurality of simulated genomic reads; comparing a plurality of sequencing reads from a sample with the database; and identifying a pathogen from the sample; wherein the eliminated false reads in the database improves identification accuracy.
- Example 31 The method of example 30, further comprising analyzing the plurality of simulated genomic reads using the database to check if the database contains an error.
- Example 32 The method of example 30 or 31 , further comprising generating a plurality reads of a known quantity of Bacteriophage MS2 and using the plurality reads of the known quantity of Bacteriophage MS2 and the plurality of sequencing reads to quantify an amount of the sample.
- Example 33 The method of example 30, or 31 , or 32, further comprising extracting a plurality of nucleic acids from the sample.
- Example 34 The method of any one of examples 30 to 33, wherein the extracting step uses at least one nucleic acid extraction kit.
- Example 35 The method of any one of examples 30 to 34, further comprising using a next generation sequencer to generate the plurality of sequencing reads.
- Example 36 The method of any one of examples 30 to 35, wherein the pathogen is a bacterium or a virus.
- Example 37 The method of any one of examples 30 to 36, wherein the pathogen is reconstructed using a genome assembly technique.
- Example 38 The method of any one of examples 30 to 37, wherein eliminating the plurality of false reads uses a manual process.
- Example 39 The method of any one of examples 30 to 38, wherein eliminating the plurality of false reads uses a computer assisted process.
- Example 40 The method of any one of examples 30 to 39, wherein the computer assisted processes uses artificial intelligence.
- Example 41 The method of any one of examples 30 to 40, further comprising using a control sequence that is an organism other than Bacteriophage MS2.
- Example 42 The method of any one of examples 30 to 41 , further comprising adding a control RNA or DNA after nucleic acid extraction to generate a plurality of simulated genomic reads.
- Example 43 The method of any one of examples 30 to 42, wherein the database used has publicly available sequencing data that is curated prior to analysis using efficient algorithms.
- the terms “approximately,” and “about” are used to describe and account for small variations.
- the terms can refer to instances in which the event or circumstance occurs precisely as well as instances in which the event or circumstance occurs to a close approximation.
- the terms can refer to a range of variation of less than or equal to ⁇ 10% of that numerical value, such as less than or equal to ⁇ 5%, less than or equal to ⁇ 4%, less than or equal to ⁇ 3%, less than or equal to ⁇ 2%, less than or equal to ⁇ 1 %, less than or equal to ⁇ 0.5%, less than or equal to ⁇ 0.1 %, or less than or equal to ⁇ 0.05%.
Landscapes
- Physics & Mathematics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Chemical & Material Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Biophysics (AREA)
- Theoretical Computer Science (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Medical Informatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biotechnology (AREA)
- Evolutionary Biology (AREA)
- Analytical Chemistry (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Molecular Biology (AREA)
- Library & Information Science (AREA)
- Genetics & Genomics (AREA)
- Biochemistry (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24833155.5A EP4736168A2 (en) | 2023-06-30 | 2024-07-01 | Metagenomic diagnostic systems and methods thereof |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363511532P | 2023-06-30 | 2023-06-30 | |
| US63/511,532 | 2023-06-30 |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| WO2025007145A2 true WO2025007145A2 (en) | 2025-01-02 |
| WO2025007145A3 WO2025007145A3 (en) | 2025-04-17 |
Family
ID=93939998
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2024/036435 Ceased WO2025007145A2 (en) | 2023-06-30 | 2024-07-01 | Metagenomic diagnostic systems and methods thereof |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4736168A2 (en) |
| WO (1) | WO2025007145A2 (en) |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10208339B2 (en) * | 2015-02-19 | 2019-02-19 | Takara Bio Usa, Inc. | Systems and methods for whole genome amplification |
| WO2017053446A2 (en) * | 2015-09-21 | 2017-03-30 | The Regents Of The University Of California | Pathogen detection using next generation sequencing |
| CN110366596A (en) * | 2016-12-28 | 2019-10-22 | 埃斯库斯生物科技股份公司 | Methods, apparatus and systems for the analysis, determination of functional relationships and interactions, and identification and synthesis of biologically active modifiers based on intact microbial strains in complex heterogeneous communities |
| IL299783A (en) * | 2020-08-18 | 2023-03-01 | Illumina Inc | Sequence-specific targeted transposition and selection and sorting of nucleic acids |
-
2024
- 2024-07-01 WO PCT/US2024/036435 patent/WO2025007145A2/en not_active Ceased
- 2024-07-01 EP EP24833155.5A patent/EP4736168A2/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2025007145A3 (en) | 2025-04-17 |
| EP4736168A2 (en) | 2026-05-06 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Diao et al. | Metagenomics next-generation sequencing tests take the stage in the diagnosis of lower respiratory tract infections | |
| CN111462821B (en) | Pathogenic microorganism analysis and identification system and application | |
| Bharucha et al. | STROBE-metagenomics: a STROBE extension statement to guide the reporting of metagenomics studies | |
| CN110349629A (en) | Analysis method for detecting microorganisms by using metagenome or macrotranscriptome | |
| Greninger et al. | Metagenomics to assist in the diagnosis of bloodstream infection | |
| EP2926288B1 (en) | Accurate and fast mapping of targeted sequencing reads | |
| WO2022028624A1 (en) | Method and apparatus for determining microbial species and acquiring related information by means of sequencing, computer-readable storage medium, and electronic device | |
| US20250182850A1 (en) | Creation or use of anchor-based data structures for sample-derived characteristic determination | |
| CN111349719B (en) | A kind of specific primer and application for novel coronavirus detection | |
| Jia et al. | A streamlined clinical metagenomic sequencing protocol for rapid pathogen identification | |
| CN117690483B (en) | Drug-resistant gene detection method based on pathogenic macro gene second generation sequencing | |
| EP4139482B1 (en) | Single step sample preparation for next generation sequencing | |
| Bartlow et al. | Comparing variability in diagnosis of upper respiratory tract infections in patients using syndromic, next generation sequencing, and PCR-based methods | |
| De Wolfe et al. | Multi-factorial examination of amplicon sequencing workflows from sample preparation to bioinformatic analysis | |
| Weng et al. | Calculating fast differential genome coverages among metagenomic sources using micov | |
| Marcolungo et al. | ACoRE: Accurate SARS-CoV-2 genome reconstruction for the characterization of intra-host and inter-host viral diversity in clinical samples and for the evaluation of re-infections | |
| CN115938491B (en) | A method and system for constructing a high-quality bacterial genome database for clinical pathogen diagnosis | |
| Lisha et al. | Clinical evaluation of negative mNGS reports in sterile body fluids and tissues | |
| Yángüez et al. | HiDRA-seq: high-throughput SARS-CoV-2 detection by RNA barcoding and amplicon sequencing | |
| EP4736168A2 (en) | Metagenomic diagnostic systems and methods thereof | |
| CN108570496A (en) | A kind of molecular diagnosis method and kit of constitutional bone disease | |
| CN119811490A (en) | Nanopore sequencing-based identification and analysis system and method for unknown pathogenic microorganisms | |
| Bustos et al. | Impact of non-standardized reporting on reproducibility, usability, and integration in nasopharyngeal metagenomic research: A systematic review | |
| CN117524313A (en) | Analysis method and device for pathogen metagenome sequencing data and application thereof | |
| Gigante et al. | Orthopoxvirus genome sequencing, assembly, and analysis |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24833155 Country of ref document: EP Kind code of ref document: A2 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024833155 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2024833155 Country of ref document: EP Effective date: 20260130 |
|
| ENP | Entry into the national phase |
Ref document number: 2024833155 Country of ref document: EP Effective date: 20260130 |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24833155 Country of ref document: EP Kind code of ref document: A2 |
|
| WWP | Wipo information: published in national office |
Ref document number: 2024833155 Country of ref document: EP |






































