WO2025245286A2 - Method for detecting myeloid-lineage malignancies/myeloid malignancies - Google Patents
Method for detecting myeloid-lineage malignancies/myeloid malignanciesInfo
- Publication number
- WO2025245286A2 WO2025245286A2 PCT/US2025/030454 US2025030454W WO2025245286A2 WO 2025245286 A2 WO2025245286 A2 WO 2025245286A2 US 2025030454 W US2025030454 W US 2025030454W WO 2025245286 A2 WO2025245286 A2 WO 2025245286A2
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- watson
- crick
- cells
- sequence
- myeloid
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6876—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
- C12Q1/6883—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
- C12Q1/6886—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material for cancer
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6806—Preparing nucleic acids for analysis, e.g. for polymerase chain reaction [PCR] assay
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2600/00—Oligonucleotides characterized by their use
- C12Q2600/156—Polymorphic or mutational markers
Definitions
- Myeloid lineage malignancies such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), myelodysplastic syndrome (MDS), or myeloproliferative neoplasms (MPNs) are characterized by the accumulation of abnormal myeloid progenitor cells in the bone marrow and peripheral blood. These malignancies are driven by molecular alterations including mutations, structural variants, and aberrant methylation patterns.
- AML acute myeloid leukemia
- CML chronic myeloid leukemia
- MDS myelodysplastic syndrome
- MPD myeloproliferative neoplasms
- Sensitive detection of molecular alterations that are markers of myeloid lineage malignancies such as AML, or other myeloid conditions (CML, MDS, or MPNs) is crucial for diagnosis, prognosis, treatment selection, monitoring treatment, and detecting minimal residual disease (MRD).
- CML, MDS, or MPNs myeloid lineage malignancies
- MRD minimal residual disease
- approaches incorporating positive selection for myeloid lineage malignancies are inevitably limited to the subset of patients with myeloid lineage malignancies whose leukemic cells harbor the detected markers.
- RT-PCR RT-PCR
- ddPCR digital droplet PCR
- VAF Variant Allele Frequency
- next-generation sequencing (NGS) based methods by assessing the presence of multiple potential molecular aberrations in parallel, are applicable to the majority of patients and have reduced operational complexity, but are generally cost prohibitive at an LoD below 10' 3 owing to the linear relationship between sequencing cost and LoD, and the need for deep sequencing to accommodate error correction methods, e.g.
- This disclosure features methods to enable sensitive, on-target, and cost-efficient MRD detection of myeloid lineage malignancies or conditions.
- this disclosure features an improved method for MRD analysis comprising negative enrichment selection coupled with detecting one or more myeloid condition variants and/or molecular profiling of the enriched myeloid cells.
- the methods described herein enable cost-effective detection of AML or other myeloid conditions down to 10' 6 LoD with an operationally efficient workflow.
- the negative enrichment workflow can reduce sequencing costs, e.g., by up to 10-fold.
- this disclosure provides methods for preparing nucleic acids from myeloid lineage tumor cells.
- the methods include obtaining a sample comprising myeloid lineage cells, performing negative selection enrichment to remove lymphoid cells while retaining residual myeloid cells, and isolating nucleic acids from the residual cells of myeloid lineage.
- the methods featured herein further comprise detecting molecular alterations in the isolated myeloid nucleic acids (IMNA) or nucleic acids processed therefrom (NAPT)
- the methods of this disclosure are applied to various sample types, including whole blood, peripheral blood, bone marrow, or fractions thereof such as peripheral blood mononuclear cells (PBMCs).
- PBMCs peripheral blood mononuclear cells
- the negative selection enrichment steps may involve density gradient or differential centrifugation, e.g., such as Ficoll density gradient centrifugation, and/or immunomagnetic cell separation using antibodies, e.g., antibody-bead extraction, against lymphoid cell markers, which can include or exclude: B cells, T cells, or NK cells.
- alternative or complementary negative selection techniques can be employed, including but not limited to, fluorescence-activated cell sorting (FACS) using a panel of antibodies to negatively select for myeloid cells by excluding labeled lymphoid and other non-myeloid populations, or microfluidic chip-based separation technologies that utilize affinity-based depletion or physical property differences to remove lymphoid cells.
- FACS fluorescence-activated cell sorting
- the isolated myeloid nucleic acids are analyzed using next-generation sequencing (NGS) technologies.
- NGS next-generation sequencing
- the NGS assay may be targeted to specific genes or regions frequently mutated in myeloid malignancies.
- the NGS assay may include one or more myeloid lineage malignancies-informed or myeloid-informed markers selected by comparison of sequencing results from whole genome sequencing or whole exome sequencing of myeloid lineage cells compared to sequencing results from myeloid lineage cells from patients with AML or other myeloid conditions following diagnosis but prior to any treatment, or before any treatment that significantly reduces leukemic cell levels.
- the NGS assay may incorporate unique or non-unique molecular identifiers (Mis) and/or strand distinguishing sequences to enable error correction and detection of low-frequency variants.
- the disclosure provides methods for detecting leukemic alterations using PCR- based assays such as quantitative PCR, digital PCR, or droplet digital PCR.
- this disclosure provides methods for detecting leukemic alterations using molecular inversion probes, including barcoded molecular inversion probes.
- the disclosure relates to methods for detecting leukemic molecular alterations using an NGS assay on nucleic acids processed from IMNA or NAPT.
- the NGS assay used in the methods of this disclosure includes error correction, such as duplex sequencing or SaferSeqs sequencing.
- the NGS assay is an Ion Torrent AmpliSeq sequencing assay.
- the NGS assay includes highly multiplexed targeted amplification (e.g., up to 24,000 amplicons), sample indexing and sequencing on an Ion Torrent sequencer.
- the NGS assay includes digesting primers following multiplex amplification.
- the methods of this disclosure feature a targeted NGS assay, such as a targeted duplex sequencing assay, where the processed nucleic acids are partially double-stranded and include a double-stranded target sequence, one or more double-stranded Mis located proximal to the ends of the target sequence, and one or more strand-distinguishing sequences located distal to the Mis on each strand.
- the Mis can be unique Mis (UMIs) or nonunique Mis.
- UMIs unique Mis
- non-unique Mis are used, but every processed DNA adapter has essentially a unique combination of start and end sites from the original target sequence and MI.
- the disclosure provides a duplex sequencing method for detecting leukemic molecular alterations.
- the method involves producing IMNA-adapters or NAPT- adapters with specific configurations of Mis, e.g, non-unique Mis or UMIs and stranddistinguishing sequences on the Watson and Crick strands, converting the IMNA-adapters or NAPT-adapters to a library, e.g., by amplifying the IMNA-adapters or NAPT-adapters, optionally converting the IMNA-adapters or NAPT-adapters or amplicons thereof to enriched IMNA-adapters or enriched NAPT-adapters by enriching a plurality of all or a portion of the IMNA-adapters or NAPT-adapters (e.g., by performing one or more targeted amplification steps with at least one target specific primer, or by bait hybridization), and analyzing the amplicons by NGS to obtain sequence reads.
- Mis e.g, non-unique Mis or UMIs and
- the targeted amplification is achieved using a multiplex PCR approach, using from 10 to 100, 500, 1000, 5000, or 10,000 primer pairs to amplify specific regions of interest.
- the primers may bind to, or flank, specific regions of interest.
- target enrichment is achieved by hybridization to a panel of oligonucleotide baits (hybrid capture), where said baits are designed to selectively bind to genomic regions relevant to myeloid malignancies, including coding regions, specific exons, introns, or regulatory regions of tens to hundreds of genes.
- the target specific primer used in a second or successive targeted amplication step is nested (inner) with respect the target specific primer used in a first targeted amplification step.
- the method involves producing a plurality of IMNA adapters or NAPT adapters, each comprising a Watson strand and a Crick strand.
- the Watson strand of each adapter includes, from 5' to 3', a first Watson single-stranded stranddistinguishing sequence comprising a first universal primer binding site, a first Watson UMI sequence, a Watson strand of a target sequence, a second Watson UMI sequence, and a second Watson single-stranded strand-distinguishing sequence comprising a second, different universal primer binding site.
- the Crick strand of each adapter includes, from 5' to 3', a first Crick single-stranded strand-distinguishing sequence identical to the first Watson strand-distinguishing sequence, a first Crick UMI sequence complementary to the second Watson UMI sequence, a Crick strand of the target sequence complementary to the Watson strand of the target sequence, a second Crick UMI sequence complementary to the first Watson UMI sequence, and a second Crick single-stranded strand-distinguishing sequence identical to the second Watson strand-distinguishing sequence.
- the methods featured herein further comprise amplifying the Watson and Crick strands of the adapters using universal primers that bind to the strand distinguishing sequences to obtain Watson and Crick amplicons, optionally converting the Watson and Crick amplicons to enriched and/or targeted Watson amplicons and enriched and/or targeted Crick amplicons by, e.g., performing one or more targeted amplification steps, or bait hybridization steps on the Watson and Crick amplicons.
- the methods featured herein further comprise analyzing the amplicons by nextgeneration sequencing (NGS) to obtain sequence reads, and identifying variants not present in all sequence reads from the Watson and Crick amplicons derived from the same original molecule.
- NGS nextgeneration sequencing
- the duplex sequencing method further includes identifying variants that are not present in all sequence reads from any one target sequence or selected sequence in Watson amplicons derived from any one Watson IMNA orNAPT and its corresponding Crick amplicons derived from any one Crick IMNA or NAPT.
- Variants not present in all or substantially all are identified as a sequencing process error (from amplification or sequencing) or as a damaged Watson and/or damaged Crick target sequence.
- the disclosure relates to methods where the myeloid lineage tumor cells being analyzed are selected from megakaryocytes, platelets, eosinophils, basophils, erythrocytes, monocytes, dendritic cells, macrophages, and/or neutrophils.
- the methods involve analyzing myeloid lineage tumor cells that express AML specific surface markers.
- AML specific surface markers may include or exclude any of the following markers: CDl lc, CD13, CD14, CD16, CD31, CD33, CD36, CD56, CD64, CD68, CD115, CD116, CD123, CD124, CD135, CD163, CD203c, CD244, CD300a, CD341, CD366, CD371, CD383, CD387 and/or Myeloperoxidase (MPO).
- Ranges can be expressed herein as from “about” one particular value, and/or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and/or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another embodiment. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as “about” that particular value in addition to the value itself. For example, if the value “10” is disclosed, then “about 10” is also disclosed.
- each unit between two particular units are also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.
- a data point range is disclosed, it is understood that each unit from the lowest data point to the highest stated datapoint, including the first (lowest) and last (highest) data point is disclosed.
- data points 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20, and each unit between any two particular units in the range are also disclosed.
- any range falling between any two of the recited values is also understood to be included.
- separation includes any means of substantially purifying one component from another (e.g., by filtration, magnetic attraction, etc.).
- the term “isolation” or “isolating”, includes the extraction, purification, and separation of nucleic acids from a biological sample, wherein the nucleic acids are removed from other cellular components, contaminants, or impurities, to obtain a preparation that is suitable for downstream applications such as sequencing, amplification, or analysis.
- the term “subject” refers to any individual who is the target of administration or treatment.
- the subject can be a vertebrate, for example, a mammal.
- the subject can be human, non-human primate, bovine, equine, porcine, canine, or feline.
- the subject can also be a guinea pig, rat, hamster, rabbit, mouse, or mole.
- the subject can be a human or veterinary patient.
- the term subject in some embodiments, refers to a “patient” under the treatment of a clinician, e.g., physician.
- beneficial or desired clinical results include, but are not limited to, alleviation of symptoms, diminishment of extent of disease, stabilized (i.e., not worsening) state of disease, delay or slowing of disease progression, amelioration or palliation of the disease state, and remission (whether partial or total), whether detectable or undetectable.
- Treatment can also mean prolonging survival as compared to expected survival if not receiving treatment.
- Those in need of treatment include those already with the condition or disorder and those prone to have the condition or disorder or those in which the condition or disorder is to be prevented.
- preventing means preventing in whole or in part, or ameliorating or controlling.
- the term "therapeutically effective amount” means an amount of a compound of the present disclosure that (i) treats the particular disease, condition, or disorder, (ii) attenuates, ameliorates, or eliminates one or more symptoms of the particular disease, condition, or disorder, or (iii) prevents or delays the onset of or reduces the intensity of one or more symptoms of the particular disease, condition, or disorder described herein.
- the term “substantially the same” means an amount of expression, level, and/or activity of a gene or gene product within 90% of baseline expression, level, and/or activity as in a subject unaffected by a disease. “Substantially the same” can also mean an amount of expression, level, and/or activity of a gene or gene product within 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, to 100% of the baseline expression, level, and/or activity as in a subject unaffected by unaffected by a disease.
- any concentration range, percentage range, ratio range or integer range is to be understood to include the value of any integer within the recited range and, when appropriate, fractions thereof (including one-tenth and one-hundredth of an integer), unless otherwise indicated.
- LoD Limit of Detection
- LoD refers to the lowest quantity or frequency of an analyte or alteration that can be reliably distinguished from its absence.
- the specific units for LoD depend on the analyte and assay methodology; for example, LoD for molecular variants is commonly expressed as Variant Allele Frequency (VAF), LoD for cell-based assays like flow cytometry may be expressed as the percentage or number of target cells per total cells or unit volume, and LoD for quantitative nucleic acid assays may be expressed as copies per unit of input material (e.g., copies/pg RNA or copies/mL of sample).”
- VAF Variant Allele Frequency
- alteration or variant refers to a single nucleotide polymorphism/single nucleotide variants (SNPs/SNVs), and deletions and insertions,. In some embodiments, copy number variations, translocations, inversions, and structural variations are also detected. SNPs/SNVs include synonymous and nonsynonymous mutations.
- Nonsynonymous mutations include missense mutations, frame-shift mutations, nonsense mutations, and readthrough mutations that can contribute to the development or progression of myeloid malignancies such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), myelodysplastic syndromes (MDS), or myeloproliferative neoplasms (MPNs).
- AML acute myeloid leukemia
- CML chronic myeloid leukemia
- MDS myelodysplastic syndromes
- MPNs myeloproliferative neoplasms
- risk allele “risk variant”, and “risk gene variant” include, but are not limited to a gene variant or SNP/SNV that is associated with risk to develop a disease.
- protection allele As used herein, the terms “protective allele”, “protective variant”, and “protective gene variant” include, but are not limited to a gene variant or SNP/SNV that is associated with protection against developing a disease. In some embodiments, a subject having a protective allele or protective variant is a subject that has one protective allele.
- nucleic acid As used herein, the terms “polynucleotide”, “nucleic acid” and “nucleic acid molecule” are used interchangeably herein to refer to a polymeric form of nucleotides of any length, and may comprise ribonucleotides, deoxyribonucleotides, analogs thereof, or mixtures thereof. This term refers only to the primary structure of the molecule. Thus, the term includes triple-, double- and single-stranded deoxyribonucleic acid (“DNA”), as well as triple-, double- and single-stranded ribonucleic acid (“RNA”).
- DNA triple-, double- and single-stranded deoxyribonucleic acid
- RNA triple-, double- and single-stranded ribonucleic acid
- isolated myeloid nucleic acids refers to nucleic acids that are extracted from enriched myeloid lineage cells and separated from substantially all other components of blood or plasma.
- myeloid lineage refers to those cells derived from myeloid cell lineage, such as macrophages (and monocytes), granulocytes (including neutrophils, eosinophils, and basophils), erythrocytes, and megakaryocytes/platelets.
- myeloid progenitor cells such as myeloblasts and promyelocytes are also included.
- these myeloid malignancies encompass acute myeloid leukemia (AML) and its various subtypes (e.g., AML with recurrent genetic abnormalities such as t(8;21)(q22;q22.1); RUNX1-RUNX1T1, inv(16)(pl3.1q22) or t(16;16)(pl3.1;q22); CBFB- MYH11, t(15;17)(q22;ql2); PML-RARA, AML with KMT2A rearrangement, AML with MECOM rearrangement, AML with NUP98 rearrangement, or AML with mutated NPM1, biallelic mutations of CEBPA, mutated RUNX1, mutated ASXL1, mutated TP53, mutated SF3B1, mutated SRSF2, mutated U2AF1, mutated ZRSR2, mutated BCOR, mutated EZH2, or mutated STAG2), myelody
- the methods are also applicable to detecting molecular markers associated with clonal hematopoiesis of indeterminate potential (CHIP) within the enriched myeloid cell fraction, which may indicate a risk for progression to overt myeloid malignancy.
- CHIP indeterminate potential
- AML refers to an acute myeloid leukemia, with several subtypes, including: myeloid leukemia in cells that produce neutrophils, a white blood cell; acute monocytic leukemia (AML-M5) in cells that produce monocytes, a white blood cell; acute megakaryocytic leukemia (AMLK) in cells that produce red blood cells or platelets, and acute promyelocytic leukemia (APL) in promyelocytes (immature white blood cells).
- AML-M5 acute monocytic leukemia
- APL acute megakaryocytic leukemia
- nucleic acids processed therefrom refers to nucleic acids derived from the IMNAs by further processing steps such as fragmentation, end repair (e.g., phosphorylating or dephosphorylatiing, blunting, dA-tailing, or conversion of methyl cytosinse and/or or 5 ’hydroxymethylcytosine by e.g., bisulfite conversion.
- a method for detecting one or more myeloid leukemic molecular alterations in nucleic acids from myeloid lineage tumor cells includes the steps of obtaining a sample comprising myeloid lineage cells, performing one or more negative enrichment steps to remove lymphoid lineage cells while retaining the myeloid lineage cells, isolating nucleic acids from the enriched myeloid cells, and detecting leukemic molecular alterations in the isolated nucleic acids or nucleic acids processed therefrom.
- the sample used in the method may be obtained from various sources, including whole blood, peripheral blood or fractions thereof such as PBMCs, or bone marrow aspirate.
- the sample is collected from a patient suspected of having a myeloid malignancy, such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), myelodysplastic syndrome (MDS), or myeloproliferative neoplasms (MPNs).
- AML acute myeloid leukemia
- CML chronic myeloid leukemia
- MDS myelodysplastic syndrome
- MPNs myeloproliferative neoplasms
- the methods disclosed herein comprise negative selection enrichment or negative enrichment of myeloid lineage cells, a key step in the method that allows for sensitive detection of myeloid leukemic molecular alterations.
- the negative enrichment involves density gradient centrifugation, such as Ficoll-Paque centrifugation, to separate mononuclear cells from granulocytes and red blood cells.
- the mononuclear cell fraction which includes lymphocytes and monocytes, can then be subjected to further negative selection to remove the lymphocytes.
- the negative enrichment is performed using a combination of density gradient centrifugation and immunomagnetic cell separation. This two-step procedure can achieve higher purity of the myeloid cell population compared to either method alone.
- the enriched myeloid cells are then lysed and the nucleic acids are extracted and isolated using standard techniques such as phenol -chloroform extraction, silica-based columns, or magnetic bead-based methods.
- the lymphocytes are removed using antibodies that specifically bind to cell surface markers expressed on B cells, T cells, and/or NK cells.
- these antibodies are conjugated to magnetic beads, allowing the labeled cells to be separated from the unlabeled myeloid cells using a magnetic field.
- anti -CD 19 antibodies are used to remove B cells, anti-CD3 antibodies for T cells, and anti-KIR (Killer Immunoglobulin-like Receptor Family) antibodies for NK cells.
- the specific markers targeted for depletion may vary depending on the desired purity and recovery of the myeloid cell population.
- antibodies targeting one or more of CD2, CD3, CD4, CD5, CD7, CD8, or T-cell receptor (TCR) components can be used.
- TCR T-cell receptor
- B cell depletion antibodies targeting one or more of CD 19, CD20, CD22, CD79a, or CD79b can be employed.
- NK cell depletion antibodies targeting one or more of CD56, CD 16, or members of the KIR family or NKG2 family receptors can be utilized.
- a combination or cocktail of antibodies against multiple markers for each lymphoid lineage is used to maximize depletion efficiency.
- the ratio of different antibody- conjugated beads or depletion reagents is optimized based on the typical or expected distribution of lymphoid cell subsets in the specific sample type (e.g., peripheral blood vs. bone marrow).
- the isolated nucleic acids are then analyzed for the presence of myeloid leukemic molecular alterations.
- the alterations are detected using next-generation sequencing (NGS) technologies.
- NGS next-generation sequencing
- samples can be multiplexed when using sample indexes or sample barcodes.
- RNA is isolated from the enriched myeloid cells, this RNA can include total RNA, messenger RNA (mRNA), or fractions enriched for specific RNA species such as microRNA (miRNA) or long non-coding RNA (IncRNA).
- mRNA messenger RNA
- miRNA microRNA
- IncRNA long non-coding RNA
- Such RNA can be subsequently converted to complementary DNA (cDNA) for analysis of fusion transcripts, gene expression levels, or RNA-specific variants.
- the isolated nucleic acids can be processed, e.g., fragmented, and/or end repaired (e.g., phosphorylation or dephosphorylation of 5’ or 3’ ends, blunting of 5’ or 3’ overhangs with or without the use of enzyymes such as, e.g., T4 DNA polymerase, removal of 3' overhangs using 3'-5' exonucleases such as, e.g., Exonuclease I, removal of 5' overhangs using 5'-3' exonucleases or endonucleases such as, e.g., lambda exonuclease or mung bean nuclease, dA-tailing, etc).
- enzyymes such as, e.g., T4 DNA polymerase, removal of 3' overhangs using 3'-5' exonucleases such as, e.g., Exonuclease I, removal of 5' overhangs using 5
- the isolated nucleic acids are subjected to fragmentation prior to adapter ligation.
- fragmentation is performed using enzymatic methods, for example, utilizing a non-specific endonuclease or a transposase-based system, to generate DNA fragments of a desired size range suitable for downstream sequencing.
- mechanical shearing e.g., sonication, acoustic shearing
- chemical methods are used for fragmentation.
- following fragmentation, or if the starting nucleic acids are already suitably fragmented (e.g., NAPT) the nucleic acid ends are repaired and modified to facilitate adapter ligation.
- this includes, but is not limited to, 5’ phosphorylation, 3’ adenylation (dA-tailing) to prepare for ligation with adapters having a 3’ thymidine (dT) overhang, blunting of ends, and/or repair of nicks or gaps.
- these steps can be performed using a combination of enzymes such as polymerases (e.g., T4 DNA polymerase, Klenow fragment), kinases (e.g., T4 polynucleotide kinase), and ligases.
- the fragments may undergo dephosphorylation and blunt ending prior to ligation with a 3’ adapter, and subsequent hybridization with a 5’ adapter, following by extension and ligation of the 5’ adapter.
- the nucleic acid library undergoes a conditioning step to remove or repair damaged DNA molecules.
- this can involve enzymatic treatments to excise damaged bases (e.g., uracil DNA glycosylase for uracil removal, FPG for oxidized purines) or to repair other forms of DNA damage such as abasic sites or single-strand breaks, thereby improving the quality and accuracy of the sequencing library.
- adapters are ligated to the ends of the processed nucleic acid fragments.
- the processed nucleic acids are converted to IMNA-adapters or NAPT- adapters with specific configurations of Mis, e.g, non-unique Mis or UMIs and strand-distinguishing sequences on the Watson and Crick strands, converting the IMNA- adapters or NAPT-adapters to a library, e.g., by amplifying the IMNA-adapters or NAPT- adapters, optionally converting the IMNA-adapters or NAPT-adapters or amplicons thereof to enriched IMNA-adapters or enriched NAPT-adapters by enriching a plurality of all or a portion of the IMNA-adapters or NAPT-adapters (e.g., by performing one or more targeted amplification steps with at least one target specific primer, or by bait hybridization), and analyzing the amplicons by NGS to obtain sequence reads.
- Mis e.g, non-unique Mis or UMIs and strand-distinguishing sequences on the Watson and Crick strands
- the targeted amplification is achieved using a multiplex PCR approach, potentially involving thousands of primer pairs to amplify specific regions of interest.
- target enrichment is achieved by hybridization to a panel of oligonucleotide baits (hybrid capture), where said baits are designed to selectively bind to genomic regions relevant to myeloid malignancies, including coding regions, specific exons, introns, or regulatory regions of tens to hundreds of genes.
- the target specific primer used in a second or successive targeted amplication step is nested (inner) with respect the target specific primer used in a first targeted amplification step.
- a limited number of PCR amplification cycles are performed using primers that bind to universal sequences within the ligated adapters. This step serves to increase the quantity of adapter-ligated library material available for subsequent hybridization to baits.
- the conversion to IMNA- adapters or NAPT-adapters involves the ligation of Y-shaped adapters, which facilitate subsequent amplification and sequencing steps and can incorporate features such as molecular identifiers (Mis) and primer binding sites.
- the method involves producing a plurality of IMNA adapters or NAPT adapters, each comprising a Watson strand and a Crick strand.
- the Watson strand of each adapter includes, from 5' to 3', a first Watson single-stranded stranddistinguishing sequence comprising a first universal primer binding site, a first Watson UMI sequence, a Watson strand of a target sequence, a second Watson UMI sequence, and a second Watson single-stranded strand-distinguishing sequence comprising a second, different universal primer binding site.
- the methods featured herein further comprise amplifying the Watson and Crick strands of the adapters using universal primers that bind to the strand distinguishing sequences to obtain Watson and Crick amplicons, optionally converting the Watson and Crick amplicons to enriched and/or targeted Watson amplicons and enriched and/or targeted Crick amplicons by, e.g., performing one or more targeted amplification steps, or bait hybridization steps on the Watson and Crick amplicons.
- the methods featured herein further comprise analyzing the amplicons by next-generation sequencing (NGS) to obtain sequence reads, and identifying variants not present in all sequence reads from the Watson and Crick amplicons derived from the same original molecule.
- NGS next-generation sequencing
- the present disclosure provides methods for detecting leukemic molecular alterations using a variety of molecular assays.
- these assays may include, but are not limited to, Next-Generation Sequencing (NGS), multiplex PCR, Droplet Digital PCR (ddPCR), Quantitative PCR (qPCR), hybrid capture assays, and molecular inversion probe assays.
- NGS Next-Generation Sequencing
- ddPCR Droplet Digital PCR
- qPCR Quantitative PCR
- hybrid capture assays and molecular inversion probe assays.
- an MI e.g., UMI or non-unique MI, strand-distinguishing sequence, or primer region
- amplification primers utilized in a second or subsequent round of amplification, for example for library preparation.
- An aliquot of the first amplified sample can be removed and amplified a second time using a second set of amplification primers that are specific to the MI, e.g., UMI or non-unique MI, stranddistinguishing sequence, or primer region, e.g., a MI, e.g., UMI or non-unique MI or an sequencing primer region, of the first amplification primers which may comprise of one or more additional MI, strand- distinguishing sequence, or primer sequences, such as sequence MI, strand-distinguishing sequence, or primer specific for one or more downstream sequencing workflows, and the same or a second PCR master mix.
- a second set of amplification primers that are specific to the MI, e.g., UMI or non-unique MI, stranddistinguishing sequence, or primer region, e.g., a MI, e.g., UMI or non-unique MI or an sequencing primer region, of the first amplification primers which may comprise of one or more additional
- the processing steps include fragmentation, and end repair.
- the preparing steps include converting IMNA or NAPT to IMNA-adapters or NAPT adapters by one or more ligation steps using one or more adapters, or one or more amplicfication steps using one or more adapter primers or one or more adapter primer pairs, one or more library amplifications using universal primers, and/or conversion to enriched and/or targeted Watson amplicons and enriched and/or targeted Crick amplicons using one or more targeted amplification steps or one or more bait enrichment steps
- the NGS assay includes an error correction method to improve the accuracy of variant calling.
- One such method is duplex sequencing, which involves converting DNA molecules (e.g., IMNA or NAPT) into DNA-adapters (e.g., IMNA-adapters or NAPT-adapters) using one or more adapters comprising a strand distinguishing sequence and/or a UMI before enrichment and sequencing.
- the specific implementation of Mis can be adapted based on the requirements of the sequencing platform and the desired error-correction fidelity.
- error correction utilizing UMIs and analysis of both DNA strands is a feature of Duplex Sequencing methodologies, which are designed to achieve high accuracy in calling low-frequency variants by distinguishing true mutations from PCR-induced or sequencing errors. This allows for a limit of detection (LoD) for variants that can extend to 10' 5 , 10' 6 , or even lower, depending on sequencing depth and initial DNA input from the enriched myeloid cells.
- Duplex sequencing can substantially reduce the error rate and enable the detection of low-frequency variants with high confidence.
- the NGS assay is a targeted assay that focuses on specific genomic regions of interest.
- the NGS assay is designed to target specific genes or genomic regions known to be frequently mutated in myeloid malignancies.
- the NGS assay is designed to target one or more myeloid informed genes or genomic regions specific for a subject. Either of these targeted approaches can increase the sensitivity and specificity of mutation detection by focusing the sequencing depth on the regions of interest.
- Targeted NGS can be achieved by various methods, such as amplicon sequencing, hybrid capture, or molecular inversion probes.
- the baits are allele, SNP, and/or mutation specific.
- targeted amplicon sequencing involves PCR amplification using primers specific to the desired genomic regions. Both hybrid capture and amplicon-based approaches are utilized in various commercially available targeted sequencing assays and can be effectively coupled with the negative enrichment methods described herein to enhance the detection of low-frequency variants from the myeloid cell fraction.
- the bait panel can be designed to cover a focused set of key myeloid genes (e.g., 20-50 genes) or a more comprehensive panel (e.g., 50-500 genes or more), including genes such as, but not limited to, NPM1, IDH1, IDH2, FLT3, KIT, TP53, RUNX1, ASXL1, DNMT3A, TET2, SF3B1, SRSF2, U2AF1, JAK2, CALR, MPL, CEBPA, BCOR, EZH2, GATA2, PHF6, RAD21, SETBP1, SH2B3, STAG2, and WT1.
- Such panels can target SNVs, indels, and regions informative for CNV detection.
- the design of such bait panels also considers intronic regions flanking exons to detect splice site mutations.
- the specific panel of genes may be selected based on the type of myeloid malignancy or condition, the clinical context, or the desired level of comprehensive profiling.
- the NGS assay is a SaferSeq assay.
- the NGS assay is an Ion Torrent AmpliSeq sequencing assay.
- the NGS assay includes highly multiplexed targeted amplification (e.g., up to 24,000 amplicons), sample indexing and sequencing on an Ion Torrent sequencer.
- the NGS assay includes digesting primers following multiplex amplification. The amplicons may be selected based on population myeloid cancer markers or may be informed by a subject’s myeloid markers vs. normal tissue markers.
- the duplex sequencing method may involve several steps, including: (i) producing IMNA or NAPT adapters with specific configurations of Mis and strand-distinguishing sequences on the Watson and Crick strands, (ii) amplifying the IMNA-adapters or NAPT- adapters to obtain Watson and Crick amplicons, (iii) optionally performing targeted amplification or bait hybridization on the amplicons, and (iv) analyzing the amplicons by NGS to obtain sequence reads.
- an additional step of (v) identifying variants that are not present in all sequence reads from the Watson and Crick amplicons derived from the same original molecule is performed. This step helps to filter out errors and identify true variants based on their presence on both strands of the original DNA duplex.
- other molecular assays are used to detect leukemic molecular alterations in addition to NGS.
- Multiplex PCR allows for the simultaneous amplification of multiple target regions in a single reaction.
- primers are designed to bind to regions of interest, e.g., those containing known hotspot mutations
- multiplex PCR provides a rapid and cost-effective way to screen for alterations.
- ddPCR and qPCR are useful for quantifying specific alterations, such as gene fusions or copy number variations.
- These assays rely on the use of fluorescent probes or intercalating dyes to measure the amount of DNA amplification in real-time.
- Hybrid capture assays involve the use of biotinylated probes to selectively enrich for target regions before sequencing. This can improve the sensitivity and specificity of the assay by increasing the proportion of sequencing reads that map to the regions of interest.
- the leukemic molecular alterations detected by the duplex sequencing methods of this disclosure include a wide range of genetic and epigenetic changes relevant to myeloid malignancies.
- single nucleotide variants are single base pair changes that occur in coding or non-coding regions of the genome.
- SNVs in genes frequently mutated in AML, MDS, or MPNs such as, e.g., SNVs in any one or more of AKT1, AKT2, AKT3, ALK, AR, ARAF, AXL, BRAF, BTK, CBL, CCND1, CDK4, CDK6, CHEK2, CSF1R, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, ERBB4, ERCC2, ESRI, EZH2, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, FOXL2, GATA2, GNA11, GNAQ, GNAS, H3F3A, HIST1H3B, HNF1A, HRAS, IDH1, IDH2, JAK1, JAK2, JAK3, KDR, KIT, KNSTRN, KRAS, MAGOH, MAP2K1, MAP2K2, MAP2K4, MAPK1, MAX, MDM4, MED12, MET, MTOR, MY
- the SNVs may include or exclude SNVs in one or a combination of the foregoing genes.
- SNVs in genes frequently mutated in AML, MDS, or MPNs such as, e.g., SNVs in any one or more of FLT3, NPM1, CEBPA, IDH1/2, or TET2, are detected with high accuracy.
- Indels are small insertions or deletions of one or more base pairs.
- structural variants which include larger-scale changes such as translocations, inversions, or copy number variations, can be detected based on the presence of Watson and Crick reads spanning the breakpoints.
- Aberrant methylation refers to changes in the DNA methylation patterns that affect gene expression and contribute to leukemogenesis.
- Copy number variants are changes in the number of copies of a particular gene or genomic region. In an embodiment, these alterations are driver events that contribute to the initiation and progression of myeloid malignancies, or they are passenger events that do not directly confer a selective advantage but may still have diagnostic or prognostic value.
- the panel of genes interrogated for SNVs and indels can be comprehensive, including, but not limited to, genes involved in signal transduction pathways: (e g., FLT3, KIT, JAK2, MPL, NRAS, KRAS, HRAS, PTPN11, CBL, NF1, SH2B3, CSF3R, MET), transcription factors and regulators: (e.g., NPM1, CEBPA, RUNX1, GATA1, GATA2, ETV6, IKZF1, WT1, CUX1, PHF6, ERG, MYC, TP53, ETV6, STAT3, SOX4, TALI, LYL1), epigenetic modifiers: including those involved in DNA methylation (e.g., DNMT1, DNMT3A, DNMT3B, TET1, TET2, TET3, IDH1, IDH2), histone modification (e.g, ASXL1, ASXL2, EZH1, EZH2, SUZ12, EED, KMT
- transcription factors and regulators
- a targeted panel comprises a selection of these genes known to be recurrently mutated in myeloid malignancies.
- a panel can include, or be similar to, genes targeted by sensitive assays such as the DuplexSeq AML MRD Assay, which includes but is not limited to: ASXL1, BCOR, BCORL1, CALR, CBL, CEBPA, CSF3R, DDX41, DNMT3A, ETV6, EZH2, FLT3, GATA2, GNAS, HRAS, IDH1, IDH2, IKZF1, JAK2, KIT, KRAS, KMT2A (MLL), MPL, NF1, NPM1, NRAS, PHF6, PPM1D, PTPN11, RAD21, RUNX1, SETBP1, SF3B1, SRSF2, STAG2, STAT3, TET2, TP53, U2AF1, WT1, and ZRSR2.
- the selection of genes for a particular panel can be tailored based on
- the methods described herein are also capable of detecting alterations in non-coding regions of the genome that have functional consequences in myeloid malignancies. These include, but are not limited to, mutations in promoter regions (e.g., TERT promoter mutations), enhancer regions, silencer regions, insulator elements, 5’ and 3’ untranslated regions (UTRs), or regions encoding non-coding RNAs.
- ncRNAs non-coding RNAs isolated from the enriched myeloid cell fraction is performed.
- microRNAs e.g., miR-155 family
- IncRNAs long non-coding RNAs
- circRNAs circular RNAs
- snoRNAs small nucleolar RNAs
- piRNAs PlWI-interacting RNAs
- histone modifications e.g., specific acetylation or methylation marks on histones such as H3K4me3
- ChlP-seq chromatin immunoprecipitation followed by sequencing
- alterations in chromatin accessibility for example, using Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) on the enriched myeloid cells, can provide insights into the regulatory landscape of the malignant cells.
- the number of genomic regions harboring leukemic alterations can vary depending on the type and stage of the myeloid malignancy.
- the leukemic molecular alterations are present in two or more genomic regions. This can reflect the presence of multiple driver events or the accumulation of passenger mutations over time.
- the alterations are present in 2 genomic regions.
- the alterations are present in 2-60 genomic regions frequently impacted in myeloid malignanices, such as, e.g., any one or more of alterations selected from alterations in FLT3, NPM1, CEBPA, CBFB, IDH1, IDH2, TET2, ASXL1, RUNX1, RUNX1T1, TP53, MYH11, NRAS, KRAS, KIT, CALR, WT1, JAK2, and MPL.
- the alteraction can be one or more mutations or alterations selected from alterations in AKT1, AKT2, AKT3, ALK, AR, ARAF, AXL, BRAF, BTK, CBL, CCND1, CDK4, CDK6, CHEK2, CSF1R, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, ERBB4, ERCC2, ESRI, EZH2, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, FOXL2, GATA2, GNA11, GNAQ, GNAS, H3F3A, HIST1H3B, HNF1A, HRAS, IDH1, IDH2, JAK1, JAK2, JAK3, KDR, KIT, KNSTRN, KRAS, MAGOH, MAP2K1, MAP2K2, MAP2K4, MAPK1, MAX, MDM4, MED12, MET, MTOR, MYC, MYCN, MYD88, NFE2L2, NRAS, N
- the alterations include or exclude one or more of the foregoing or combinations thereof.
- the alteraction can be one or more mutations, alterations, or gene fusions selected from a gene fusion of one or more of AKT2, ALK, AR, AXL, BRCA1, BRCA2, BRAF, CDKN2A, EGFR, ERBB2, ERBB4, ERG, ESRI, ETV1, ETV4, ETV5, FGFR1, FGFR2, FGFR3, FGR, FLT3, JAK2, KRAS, MDM4, MET, MYB, MYBL1, NF1, N0TCH1, N0TCH4, NRG1, NTRK1, NTRK2, NTRK3, NUTM1, PDGFRA, PDGFRB, PIK3CA, PRKACA, PRKACB, PTEN, PPARG, RAD51B, RAFI, RBI, RELA, RET, ROS1, RSPO2, RSPO3, and TERT.
- the alterations include or
- the methods of the present disclosure can be applied to various types of myeloid lineage tumor cells, including megakaryocytes, platelets, eosinophils, basophils, erythrocytes, monocytes, dendritic cells, macrophages, and/or neutrophils.
- these cell types represent different stages and lineages of myeloid differentiation and give rise to distinct subtypes of myeloid malignancies.
- the myeloid lineage tumor cells express specific surface markers that are used for their identification and isolation.
- markers may include CDl lc, CD13, CD14, CD31, CD33, CD36, CD64, CD68, CD115, CD116, CD123, CD124, CD135, CD163, CD203c, CD244, CD300a, CD341, CD366, CD371, CD383, CD387 and/or Myeloperoxidase (MPO).
- MPO Myeloperoxidase
- the methods of the present disclosure are applied to myeloid myeloid lineage cells.
- the methods of the present disclosure are applied to acute myeloid leukemia (AML) cells.
- AML is a heterogeneous malignancy characterized by the clonal expansion of immature myeloid cells in the bone marrow and peripheral blood.
- AML cells express various combinations of surface markers, including CD13, CD33, CD34, CD117, and MPO.
- the genetic alterations in AML are diverse.
- the methods of the present disclosure are used to comprehensively profile the genetic and epigenetic landscape of AML cells and to monitor the response to therapy and the emergence of resistant clones.
- MDS myelodysplastic syndromes
- MDS cells show dysplastic morphology and increased apoptosis in one or more myeloid lineages.
- the genetic alterations in MDS often involve mutations in genes related to RNA splicing (e.g. SF3B1, SRSF2, U2AF1), epigenetic regulation (e.g. TET2, ASXL1, DNMT3A), and transcriptional regulation (e.g. RUNX1, TP53).
- the methods of the present disclosure are used to identify the specific genetic and epigenetic alterations in MDS cells and to guide risk stratification and treatment decisions.
- the methods of the present disclosure are applied to chronic myeloid leukemia (CML) cells.
- CML is a myeloproliferative neoplasm characterized by the BCR-ABL1 fusion gene, which results from a translocation between chromosomes 9 and 22.
- the BCR-ABL1 fusion protein has constitutive tyrosine kinase activity and drives the proliferation and survival of myeloid cells.
- the methods of the present disclosure are used to monitor the response to tyrosine kinase inhibitor therapy and to detect the emergence of resistance mutations in the BCR-ABL1 kinase domain.
- the methods of the present disclosure are used to detect minimal residual disease (MRD) in myeloid malignancies.
- MRD refers to the presence of residual leukemic cells below the threshold of morphologic detection.
- MRD is an important prognostic factor and guides decisions about the intensity and duration of therapy.
- the high sensitivity of the molecular assays used in the present disclosure such as NGS and ddPCR, enables the detection of MRD at levels as low as 0.01% or lower.
- MRD allows for early intervention and personalized management of patients based on their MRD status.
- the initial amount of nucleic acid (IMNA or NAPT) input into the molecular assay following enrichment can be varied.
- DNA input amounts may have a range that is between any two values from about 0.1 ng to about 2000 ng, or more.
- the initial amount of nucleic acid (IMNA or NAPT) input into the molecular assay following enrichment can be varied.
- DNA input amounts may have a range that is between any two values of about 0.1 ng, about 1 ng, about 5 ng, about 10 ng, about 15 ng, about 20 ng, about 25 ng, about 30 ng, about 40 ng, about 50 ng, about 75 ng, about 100 ng, about 125 ng, about 150 ng, about 175 ng, about 200 ng, about 250 ng, about 300 ng, about 400 ng, about 500 ng, about 750 ng, about 1000 ng (1 pg), about 1500 ng (1.5 pg), about 2000 ng (2 pg), or more.
- the DNA input amount is, is about, or is at least 0.1 ng, 1 ng, 10 ng, 25 ng, 50 ng, 100 ng, or 150 ng. In some embodiments, the DNA input amount is, is about, or is not more than 50 ng, 100 ng, 250 ng, 500 ng, 1000 ng, or 2000 ng. In an embodiment, the DNA input amount is within a range of about 10 ng to about 50 ng. In an embodiment, the DNA input amount is within a range of about 50 ng to about 250 ng. In an embodiment, the DNA input amount is within a range of about 200 ng to about 2000 ng. In an embodiment, the selection of a specific input amount can depend on factors such as the type of nucleic acid, the expected abundance of target sequences, the downstream molecular assay to be performed, and the desired level of sensitivity.
- the size of the nucleic acid molecules may vary greatly using the methods and compositions disclosed herein. It would be appreciated by those skilled in the art that the nucleic acid molecules amplified from a target sequence comprising a tandem repeat (e.g., STR) may have a large size, while the nucleic acid molecules amplified from a target sequence comprising a SNP may have a small size. In some embodiments, the nucleic acid molecules may comprise from less than a hundred nucleotides to hundreds or even thousands of nucleotides. In some embodiments, the size of the nucleic acid molecules may have a range that is between 3- bp to 1 kb.
- the size of the nucleic acid molecules may have a range that is between any two values of about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1 kb, or more.
- the minimal size of the nucleic acid molecules may be a length that is, is about, or is less than, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, or 100 bp.
- the maximum size of the nucleic acid molecules may be a length that is, is about, or is more than, 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, or 1 kb.
- the method involves ligating adapters to the ends of doublestranded IMNA or NAPT fragments to generate adapter-attached, double-stranded DNA fragments.
- the adapters comprise a double-stranded portion comprising a UMI and a single-stranded forked portion with a 3' sequence and a non-complementary 5' sequence to create Y-shaped ends.
- the 5' sequence can include a binding site for a first sequencing primer (Rl), while the 3' sequence can include a binding site for a second sequencing primer (R2).
- each Y-shaped adapter comprises: a 3’ single-stranded arm, a 5’ single-stranded arm, an MI sequence (e.g., UMI or non-unique MI) that alone or in combination with a start or end site, or sequence from the target sequence uniquely labels a ligated IMNA-adapter or NAPT-adapter such that each individual ligated IMNA-adapter or NAPT-adapter is distinguishable from other IMNA-adapters or NAPT- adapters; and a strand-distinguishing sequence that, following the attachment of the adapter to the IMNA or NAPT, provides a region of non-complementarity between a first strand of an individual IMNA-adapter or NAPT-adapter and a second strand of the same IMNA-adapter or NAPT-adapter.
- MI sequence e.g., UMI or non-unique MI
- the R1 and R2 primer binding sites are used in sequencing systems that generate paired-end reads from opposite ends of each fragment.
- the R1 primer is used to produce "Rl” or “Read 1" sequences from one end, while the R2 primer is used to produce "R2" or “Read 2" sequences from the other end.
- the Rl and R2 reads for each fragment can be aligned as read pairs and grouped by their unique Mis for error correction.
- the number of expected duplicate fragments e.g., distinct original molecules that happen to yield the same start and end sites after fragmentation
- the number of expected duplicate fragments in a tested sample is estimated to be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or in a range such as about 1 to about 2, about 1 to about 3, about 1 to about 4, about 1 to about 5, about 1 to about 6, about 1 to about 7, about 1 to about 8, about 1 to about 9, about 1 to about 10, about 2 to about 5, about 5 to about 10, about 1 to about 15, about 1 to about 20, or any range between any two of these recited numbers
- the diversity of the MI pool total number of different MI sequences available
- the total number of available Mis may be, for example, any number between 100, 250, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 100,000, 250,000, 500,000, 1,000,000, 5,000,000, 10,000,000, 25,000,000, 50,000,000, 100,000,000, 150,000,000, 200,000,000, 250,000,000, about 268,000,000, 300,000,000, 500,000,000, or 1,000,000,000 or , or any range between any two of the recited numbers, or more.
- the MI e.g., UMI has a length.
- the length can be about 2- 4000 nt.
- the length is from about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, to about 4000 nt, or any range between any two of these recited lengths.
- the length can be about 6-100 nt.
- the length is from about 6 nt to about 100 nt.
- the length is about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 22, 24, 25, 28, 30, 32, 35, 36, 40, 42, 45, 48, 50, 55, 60, 64, 65, 70, 72, 75, 80, 85, 90, 95, or about 100 nt, or any range between any two of these recited lengths.
- the length can be about 8-50 nt.
- the length is about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or about 50 nt, or any range between any two of these recited lengths.
- the length can be about 10-20 nt.
- the length is about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or about 20 nt, or any range between any two of these recited lengths.
- the length can be about 12-14 nt.
- the MI length is sufficient to uniquely barcode the molecules without interfering with downstream processing steps.
- Mis, e.g., UMIs are randomly generated sequences that are confirmed to be distinct from the genomic sequences of the myeloid samples.
- methods described herein involve unique molecular identifiers (UMI).
- UMI comprises a randomly selected stretch of nucleotides that can be used during sequencing to correct for PCR and sequencing errors, thereby adding an additional layer of error correction to sequencing results.
- a MI e.g, UMI or non-unique identier can be associated with a selected target sequence. Mis could be from, for example from 3-50 nucleotides long (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17. 18, 19, 20, 30, 40, or 50, or any range between any of those lengths).
- the number of Mis e.g., UMIs depends on the amount of input DNA.
- the length of the UMI is sufficient to uniquely barcode the molecules and the length/sequence of the UMI does not interfere with the downstream amplification steps.
- all PCR duplicates from the same PCR reaction would have the same UMI sequence, as such the duplicates can be compared and any errors in the sequence such as single base substitutions, deletions, insertions (i.e., stutter in PCR) can be excluded from the sequencing results bioinformatically.
- the molecular identifiers e.g., non-unique or unique molecular identifier (UMI) sequences are distinct from any sequences present in the native IMNA or NAPT fragments. This ensures that the Mis can be unambiguously distinguished from genomic sequences during data analysis.
- such unique sequences can be randomly generated, e.g., by a computer readable medium, and selected by BLASTing against known nucleotide databases such as, e.g., GenBank.
- the MI sequences may be designed to match a specific sequence in the IMNA or NAPT fragments. In these cases, the position of the UMI within the read (e.g. at a specific distance from the R1 or R2 primer) is used to discriminate it from the genomic sequence during data analysis. In an embodiment, this approach may be used to label individual strands of a double-stranded fragment with the same UMI for applications like error correction.
- a primer of the present methods can comprise one or more Mis, strand distinguishing sequences, and/or primer sequences.
- the MI, strand distinguishing sequences, and/or primer sequences can be one or more of primer sequences that are not homologous to the target sequence, but for example can be used as templates for one or more amplification reactions.
- the MI, strand distinguishing sequences, and/or primer sequence can be a capture sequence, for example a hapten sequence such as biotin that can be used to purify amplicons away from reaction components.
- the present disclosure also relates to methods for detecting one or more leukemic molecular alterations in nucleic acids from myeloid lineage tumor cells.
- the methods involve obtaining a sample comprising myeloid lineage cells, performing negative selection enrichment, or negative enrichment, to remove lymphoid cells while retaining myeloid cells, isolating nucleic acids from the residual cells of myeloid lineage, and detecting molecular alterations in the isolated myeloid nucleic acids (IMNA) or nucleic acids processed therefrom (NAPT).
- IMNA isolated myeloid nucleic acids
- NAPT nucleic acids processed therefrom
- the negative enrichment steps result in a 2-fold enrichment of myeloid (e.g., AML) cells relative to the initial number of myeloid (e.g., AML) cells present in the sample.
- the negative enrichment steps result in a 3 -fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16- fold, 17-fold, 18-fold, 19-fold, or 20-fold enrichment of AML cells.
- the degree of enrichment achieved by the negative enrichment steps may depend on various factors, such as the initial proportion of myeloid (e.g., AML) cells in the sample, the specific methods used for negative enrichment, and the efficiency of cell separation.
- the negative enrichment steps are optimized to achieve a desired level of myeloid (e.g., AML) cell enrichment based on these factors.
- the negative enrichment steps involve the use of density gradient centrifugation to separate mononuclear cells, including myeloid (e.g., AML) cells, from other cell types based on their density.
- myeloid e.g., AML
- Ficoll-Paque centrifugation can be used to isolate mononuclear cells from peripheral blood or bone marrow samples. This step can result in a 2-fold to 5-fold enrichment of myeloid (e.g., AML) cells, depending on the initial myeloid (e.g., AML) cell frequency and the efficiency of the separation.
- the negative enrichment steps involve the use of immunomagnetic beads to deplete non-myeloid (e.g., non-AML) cells from the sample.
- beads coated with antibodies specific for T cells e.g., CD3 or any one or more T cell lineage marker featured herein
- specific for B cells e.g., CD19 or any one or more B cell lineage marker featured herein
- specific for NK cells e.g., KIR or any one or more NK cell lineage markers featuered herein
- This step can result in a 5-fold to 20-fold enrichment of myeloid (e.g., AML) cells, depending on the initial myeloid (e.g., AML) cell frequency and the efficiency of the depletion.
- the one or more B cell lineage markers are selected from: CD19, CD20, CD22, CD23, CD24, CD38, CD40, CD45, CD79a, CD79b, CD138, CD200 IgKappa, IgLambda, IgM, IgD, IgGl, IgG2, IgG3, IgG4, IgE, IgAl, and/or IgA2.
- the B cell lineage markers analyzed may include or exclude one or more of the foregoing.
- the one or more T cell lineage markers are selected from: CD2, CD3, CD4, CD5, CD7, CD8, CD25, CD27, CD28, CD45RA, CD45RO, CD69, TRBC1, TRBC2, and/or CD127.
- the T cell lineage markers analyzed may include or exclude one or more of the foregoing.
- the one or more NK cell lineage markers are selected from: CD57, CD94, CD122, CD158a, CD158b, CD159b, CD161, CD314, NKG2A, NKG2C, NKG2D, NKp30, NKp44, NKp46, NKp80, and KIR Family Receptors.
- the NK cell lineage markers analyzed may include or exclude one or more of the foregoing.
- the negative enrichment steps involve a combination of density gradient centrifugation and immunomagnetic depletion.
- the effectiveness of the negative enrichment steps can be assessed by measuring the proportion of myeloid cells in the sample before and after enrichment. This can be done using flow cytometry, immunohistochemistry, or other methods that allow for the specific identification and quantification of myeloid cells.
- the enriched myeloid (e.g., AML) cell population is further characterized using molecular techniques, such as nextgeneration sequencing or polymerase chain reaction, to detect myeloid -associated genetic alterations.
- the sample from a subject is obtained or derived from the subject’ s whole blood, peripheral blood, peripheral blood mononuclear cells (PBMCs), or bone marrow, cells, or tissue.
- the sample contains an amount of nucleic acids.
- obtaining or having obtained the sample results in disrupting or lysing cells in the sample to release nucleic acids from the cells.
- the subject’s sample comprises cellular nucleic acids.
- cellular nucleic acids make up less than about 1% of the total cellular nucleic acids in the sample.
- cellular nucleic acids make up less than about 5% of the total cellular nucleic acids in the sample.
- cellular nucleic acids make up less than about 10% of the total cellular nucleic acids in the sample. In some embodiments, cellular nucleic acids make up less than about 20% of the total cellular nucleic acids in the sample. In some embodiments, cellular nucleic acids make up more than about 50% of the total cellular nucleic acids in the sample. In some embodiments, cellular nucleic acids make up less than about 90% of the total cellular nucleic acids in the sample.
- the sample used in the method may be obtained from various sources, including whole blood, peripheral blood, or fractions thereof such as PBMCs, or bone marrow aspirate.
- the sample is collected from a patient suspected of having a myeloid malignancy, such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), myelodysplastic syndrome (MDS), or myeloproliferative neoplasms (MPNs).
- AML acute myeloid leukemia
- CML chronic myeloid leukemia
- MDS myelodysplastic syndrome
- MPNs myeloproliferative neoplasms
- this disclosure provides for methods for obtaining or having obtained a sample for the methods as described herein.
- a sample may be obtained directly (e.g., a doctor takes a blood sample from a subject).
- a sample may be obtained indirectly (e.g., through shipping, by a technician from a doctor or a subject).
- the sample is a whole blood, peripheral blood, peripheral blood mononuclear cells (PBMCs), or bone marrow.
- PBMCs peripheral blood mononuclear cells
- this disclosure provides for methods comprising obtaining or having obtained samples with fragmented nucleic acids.
- the sample may have been subjected to conditions that are not conducive to preserving the integrity of nucleic acids.
- the sample may have been exposed to air, heat, light, or chemicals or enzymes that degrade nucleic acids.
- methods comprise obtaining or having obtained a tissue sample wherein the tissue sample comprises fragmented nucleic acids.
- methods comprise obtaining or having obtained a tissue sample wherein the tissue sample comprises nucleic acids and fragmenting the nucleic acids to produced fragmented nucleic acids.
- the tissue sample is a frozen sample.
- the sample is a preserved sample.
- the tissue sample is a fixed sample (e.g. formaldehyde-fixed).
- methods may comprise isolating the (fragmented) nucleic acids from the sample.
- this disclosure provides for methods comprising obtaining or having obtained a sample from a subject, wherein the sample contains at least one nucleic acid of interest.
- the nucleic acid of interest may be DNA from the whole blood, peripheral blood, peripheral blood mononuclear cells (PBMCs), or bone marrow of a subject.
- PBMCs peripheral blood mononuclear cells
- this disclosure provides for methods comprising obtaining or having obtained a sample from the subject, wherein the sample contains about 1 to about 100 nucleic acids.
- the at least one nucleic acid is represented by a sequence that is unique to a target gene disclosed herein.
- this disclosure provides for methods comprising isolating or having isolated nucleic acids from a biological sample. In some embodiments, this disclosure provides for methods comprising purifying or having purified nucleic acids from a biological sample.
- biological sample means a whole blood, peripheral blood, peripheral blood mononuclear cells (PBMCs), or bone marrow sample obtained from a subject.
- PBMCs peripheral blood mononuclear cells
- this disclosure provides for methods comprising isolating or having isolated and purifying or having purified nucleic acids from a biological sample.
- this disclosure provides for methods of isolating nucleic acids comprising removing or having removed non-nucleic acid components from a biological sample described herein.
- the nucleic acids derived therefrom can include any or all the following manipulations described herein, including any method of isolating, purifying, enriching, and one or more steps involved in preparing a library of nucleic acids for detection from the nucleic acids.
- isolating or purifying comprises removing unwanted non- nucleic acid components from a biological sample.
- isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of unwanted non-nucleic acid components from a biological sample, such as, e.g., components of blood or plasma.
- isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 95% of unwanted non-nucleic acid components from a biological sample to obtain a sample derived from the initial sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 97% of unwanted non-nucleic acid components from a biological sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 98% of unwanted non-nucleic acid components from a biological sample.
- isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 99% of unwanted non-nucleic acid components from a biological sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 95% of unwanted non- nucleic acid components from a biological sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 97% of unwanted non-nucleic acid components from a biological sample.
- isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 98% of unwanted non-nucleic acid components from a biological sample. In some embodiments, isolating or having isolated and purifying or having purified IMNA or NAPT from a biological sample comprises removing at least 99% of unwanted non-nucleic acid components from a biological sample.
- this disclosure provides for methods comprising isolating or having isolated and purifying or having purified DNA, RNA, or a subset of desired nucleic acids from a mix of nucleic acids derived from a biological sample.
- methods may comprise analyzing only DNA. In methods wherein only DNA is analyzed, RNA is unwanted and creates undesirable background noise or contamination to the DNA; therefore, in some embodiments, the removed nucleic acid components are discarded.
- this disclosure provides for methods comprising removing RNA from a biological sample. In some embodiments, this disclosure provides for methods comprising removing mRNA from a biological sample. In some embodiments, this disclosure provides for methods comprising removing microRNA from a biological sample.
- removing nucleic acid components comprises contacting the nucleic acid components with an oligonucleotide capable of hybridizing to the nucleic acid, wherein the oligonucleotide is conjugated, attached or bound to a capturing device (e.g., bead, column, matrix, nanoparticle, magnetic particle, etc.). In some embodiments, the removed nucleic acid components are discarded.
- a capturing device e.g., bead, column, matrix, nanoparticle, magnetic particle, etc.
- the negative enrichment involves density gradient centrifugation, such as Ficoll-Paque centrifugation, to separate mononuclear cells from granulocytes and red blood cells.
- the mononuclear cell fraction which includes lymphocytes and monocytes, can then be subjected to further negative selection to remove the lymphocytes.
- the lymphocytes are removed using antibodies that specifically bind to cell surface markers expressed on B cells, T cells, and/or NK cells.
- the negative enrichment is performed using a microfluidic device.
- the negative enrichment is performed using a combination of density gradient centrifugation and immunomagnetic cell separation. This two-step procedure can achieve higher purity of the myeloid cell population compared to either method alone.
- the enriched myeloid cells are then lysed and the nucleic acids are extracted using standard techniques such as phenol-chloroform extraction, silica-based columns, or magnetic beadbased methods.
- this disclosure provides for methods comprising isolating or having isolated and purifying or having purified nucleic acids from one or more non-nucleic acid components of a biological sample.
- Non-nucleic acid components may also be considered unwanted substances.
- Non-limiting examples of non-nucleic acid components include cells (e.g., blood cells), cell fragments, extracellular vesicles, lipids, proteins or a combination thereof. Additional non-nucleic acid components are described herein and throughout. It should be noted that while methods may comprise isolating and/or purifying nucleic acids, they may also comprise analyzing a non-nucleic acid component of a sample that is considered an unwanted substance in a nucleic acid purifying step. Isolating or having isolated and purifying or having purified may comprise removing components of a biological sample that would inhibit, interfere with or otherwise be detrimental to the later process steps such as nucleic acid amplification or detection.
- purifying or having purified nucleic acids does not comprise washing the nucleic acids with a wash buffer.
- purifying or having purified nucleic acids comprises capturing the nucleic acids with a nucleic acid capturing moiety to produce captured nucleic acids.
- nucleic acid capturing moieties are silica particles and paramagnetic particles.
- purifying or having purified nucleic acids comprises passing the sample comprising the captured nucleic acids through a hydrophobic phase (e.g., a liquid or wax). The hydrophobic phase retains impurities in the sample that would otherwise inhibit further manipulation (e.g., amplification, sequencing) of the nucleic acids.
- Isolating and/or purifying may occur with the use of a sample purifier.
- isolating or having isolated and purifying or having purified nucleic acids comprises removing non-nucleic acid components from a biological sample described herein.
- isolating or having isolated and purifying or having purified nucleic acids comprises discarding non-nucleic acid components from a biological sample.
- isolating or having isolated and purifying or having purified nucleic acids comprises collecting, processing and analyzing the non-nucleic acid components.
- the non-nucleic acid components may be considered biomarkers because they provide additional information about the subject.
- isolating or having isolated and purifying or having purified nucleic acids comprise lysing a cell. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids avoids lysing a cell. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids does not comprise lysing a cell. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids does not comprise an active step intended to lyse a cell. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids does not comprise intentionally lysing a cell. Intentionally lysing a cell may include mechanically disrupting a cell membrane (e.g., shearing). Intentionally lysing a cell may include contacting the cell with a lysis reagent.
- Intentionally lysing a cell may include mechanically disrupting a cell membrane (e.g., shearing
- isolating or having isolated and purifying or having purified nucleic acids comprises separating components of a biological sample disclosed herein. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids comprises centrifuging the biological sample, filtering the biological sample, contacting the sample with a solid phase support, or using solid phase extraction, in order to separate components of a biological sample, or a combination thereof. In some embodiments, methods comprise subjecting the blood to vertical filtration. In some embodiments, methods comprise subjecting the blood to a sample purifier comprising a filter matrix for receiving whole blood, the filter matrix having a pore size that is prohibitive for cells to pass through, while nucleic acids can pass through the filter matrix uninhibited. Such vertical filtration and filter matrices are described for devices disclosed herein.
- isolating or having isolated and purifying or having purified nucleic acids comprises filtering the biological sample in order to remove non-nucleic acid components from the biological sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids comprises filtering the biological sample in order to capture nucleic acids from the biological sample. In some embodiments, removing non-nucleic acid components may comprise centrifuging the biological sample. In some embodiments, removing non-nucleic acid components may comprise contacting the biological sample with a binding moiety.
- Isolating or having isolated and purifying or having purified may comprise capturing an extracellular vesicle or extracellular microparticle in the biological sample with a binding moiety.
- the extracellular vesicle contains at least one of DNA and RNA.
- this disclosure provides for methods of enriching nucleic acids isolated and purified from a biological sample, e.g. by hybridizing enrichment such as bait capture or by amplification, e.g., PCR amplification, or to provide a sample enriched in selected or target nucleic acids.
- methods comprise performing targeted sequencing on the enriched nucleic acids, which may comprise amplicons of nucleic acids in the biological sample, obtained by whole genome or targeted amplification.
- targeted amplification is performed using target-specific primers for fully nested or hemi-nested PCR.
- enriching nucleic acid components comprises separating the nucleic acid components on a gel by size.
- this disclosure provides for methods comprising removing DNA from the biological sample.
- this disclosure provides for methods comprising capturing DNA from the biological sample.
- this disclosure provides for methods comprising selecting DNA from the biological sample.
- the genomic DNA has a minimum length. In some embodiments, the minimum length is about 50 base pairs. In some embodiments, the minimum length is about 100 base pairs. In some embodiments, the minimum length is about 110 base pairs. In some embodiments, the minimum length is about 120 base pairs. In some embodiments, the minimum length is about 130, about 140, about 150, about 160, about 170 or about 180 base pairs.
- the minimum length is about 130, about 140, about 150, about 160, or about.
- the DNA has a maximum length. In some embodiments, the maximum length is about 180 base pairs. In some embodiments, the maximum length is about 200 base pairs. In some embodiments, the maximum length is about 220 base pairs. In some embodiments, the maximum length is about 240 base pairs. In some embodiments, the maximum length is about 300 base pairs. Size based separation would be useful for other categories of nucleic acids having limited size ranges (e.g., microRNAs).
- enriching methods comprise capturing a nucleosome in a biological sample and analyzing nucleic acids attached to the nucleosome.
- methods comprise capturing an exosome in a biological sample and analyzing nucleic acids attached to the exosome. Capturing nucleosomes and/or exosomes may preclude the need for a lysis step or reagent, thereby simplifying the method and reducing time from sample collection to detection.
- enriching nucleic acids comprises lysing and performing sequence specific capture of a target nucleic acid with "bait" in a solution followed by binding of the "bait" to solid supports such as magnetic beads.
- methods comprise performing sequence specific capture in the presence of a recombinase or helicase to reduce the need for heat denaturation of a nucleic acid thereby speeding up the detection step.
- enriching nucleic acids comprises subjecting a biological sample, or a fraction thereof, or a modified version thereof, to a binding moiety.
- the binding moiety may be capable of binding to a component of a biological sample and removing it to produce a modified sample depleted of cells, cell fragments, nucleic acids or proteins that are unwanted or of no interest.
- enriching purified nucleic acids comprises subjecting a biological sample to a binding moiety to reduce unwanted substances or non- nucleic acid components in a biological sample.
- enriching purified nucleic acids comprises subjecting a biological sample to a binding moiety to produce a modified sample enriched with target cell, target cell fragments, target nucleic acids or target proteins.
- the resulting cell-bound binding moieties can be captured and enriched for with antibodies or other methods, e.g., low speed centrifugation.
- this disclosure provides for methods comprising amplifying at least one nucleic acid in a sample to produce at least one amplification product.
- the at least one nucleic acid may be a genomic nucleic acid or nucleic acid obtained from a cell in a sample.
- the sample may be a biological sample disclosed herein or a fraction or portion thereof.
- methods comprise producing a copy of the nucleic acid in the sample and amplifying the copy to produce the at least one amplification product.
- methods comprise producing a reverse transcript of the nucleic acid in the sample and amplifying the reverse transcript to produce the at least one amplification product.
- methods comprise performing whole genome amplification. In some embodiments, methods do not comprise performing whole genome amplification.
- the term, "whole genome amplification” also referred to as “random amplification” may refer to amplifying all of the genomic nucleic acids in a biological sample.
- the term, “whole genome amplification” may refer to amplifying at least 90% of the genomic nucleic acids in a biological sample.
- Whole genome may refer to multiple genomes.
- Whole genome amplification may comprise amplifying genomic nucleic acids from a biological sample of a subject having an infection, wherein the biological sample comprises genomic nucleic acids from the subject and a pathogen.
- targeted sequencing of the IMNA- or NAPT-adapters is performed using bait-capture methods.
- a prepared pool of NGS templates is heat denatured and mixed with a pool of capture probe oligonucleotides (“baits”).
- the baits are designed to hybridize to the regions of interest within the target genome and are 60-200 bases in length, for example, about 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or about 200 bases in length, or any range between any two of these recited lengths.
- baits are further modified to contain a ligand that permits subsequent capture of these probes.
- such oligonucleotide baits comprise deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or analogs thereof, including oligonucleotides with modified backbones or bases to enhance hybridization specificity, stability, or reduce non-specific binding.
- the length of the baits can vary widely, for example, depending on the specific application, desired hybridization kinetics, the complexity of the target regions, and the specific enrichment strategy employed.
- a commonly used approach involves baits with lengths typically ranging from about 100 nucleotides to about 150 nucleotides, for example, around about 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110,
- bait lengths for example ranging from about 30 nucleotides to about 50 nucleotides, e.g., about 30, 35, 40, 45, or 50 nucleotides; such shorter baits can, in some embodiments, be utilized in single or multiple rounds of hybridization to achieve the desired target enrichment. While these represent common ranges, the baits can also have lengths outside of these specific examples, for instance, generally ranging from about 30 nucleotides up to about 500 nucleotides or more.
- this range includes lengths such as about 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, or about 500 nucleotides, or any range between any two recited lengths, depending on the requirements of the assay.
- the capture method incorporates a biotin group (or groups) on the baits.
- other capture ligands are used.
- capture is performed with a component having affinity for only the bait.
- streptavidin-magnetic beads are used to bind the biotin moiety of biotinylated-baits that are hybridized to the desired DNA targets from the pool of NGS templates.
- washing removes unbound nucleic acids, reducing the complexity of the retained material.
- the retained material is then eluted from the magnetic beads and introduced into automated sequencing processes, providing for ‘capture enrichment’, where the captured nucleic acids are retained as an enriched pool for subsequent study.
- the design of bait oligonucleotides accounts for potential sequence variations in the population to ensure efficient capture of target regions across different individuals.
- baits can be synthesized with modified nucleotides or backbones (e.g., locked nucleic acids - LNAs) to enhance their binding affinity and specificity for target sequences, or to improve their stability.
- the density and tiling of baits across a target region can also be optimized to ensure uniform capture and coverage.
- computational methods can be used to design bait sequences that minimize cross-hybridization to off-target genomic regions, including pseudogenes or repetitive elements.
- this disclosure provides for methods comprising amplifying a nucleic acid, wherein amplifying comprises performing an isothermal amplification of the nucleic acid.
- isothermal amplification are as follows: loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicase dependent amplification (HDA), nicking enzyme amplification reaction (NEAR), and recombinase polymerase amplification (RPA).
- LAMP loop-mediated isothermal amplification
- SDA strand displacement amplification
- HDA helicase dependent amplification
- NEAR nicking enzyme amplification reaction
- RPA recombinase polymerase amplification
- the isothermal amplification is high throughput involving parallel sample processing.
- the high throughput isothermal amplification involves amplifying a nucleic acid in 12, 24, 36, 48, 60, 72, 84, 96, 108, or more samples in parallel. In some embodiments, the high throughput isothermal amplification involves amplifying a nucleic acid in between 12-24, 24-36, 36-48, 48-60, 70-72, 72-84, 84-96, 96-108, 108-120, 120-132, 132-144, 144-156-156-168, 168-180, 180-192, 192-204, 204-216, 216-228, 228-240, 240-252, or 252-264, samples in parallel.
- the high throughput isothermal amplification involves amplifying a nucleic acid in at least 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, or 1,500 samples in parallel.
- Amplification methods can include isothermal amplification.
- amplification is isothermal with the exception of an initial heating step before isothermal amplification begins.
- isothermal amplification methods can be used (Zanoli and Spoto, 2013, “Isothermal Amplification Methods for the Detection of Nucleic Acids in Microfluidic Devices," Biosensors 3: 18-43; Fakruddin, et al., 2013, “Alternative Methods of Polymerase Chain Reaction (PCR),” Journal of Pharmacy and Bioallied Sciences 5(4): 245-252).
- any appropriate isothermic amplification method is used.
- the isothermic amplification method used is selected from: Loop Mediated Isothermal Amplification (LAMP); Nucleic Acid Sequence Based Amplification (NASBA); Multiple Displacement Amplification (MDA); Rolling Circle Amplification (RCA); Helicase Dependent Amplification (HDA); Strand Displacement Amplification (SDA); Nicking Enzyme Amplification Reaction (NEAR); Ramification Amplification Method (RAM); and Recombinase Polymerase Amplification (RPA).
- LAMP Loop Mediated Isothermal Amplification
- NASBA Nucleic Acid Sequence Based Amplification
- MDA Multiple Displacement Amplification
- RCA Rolling Circle Amplification
- HDA Helicase Dependent Amplification
- SDA Strand Displacement Amplification
- NEAR Nicking Enzyme Amplification Reaction
- RAM Ramification Amplification Method
- RPA Recombinase Polymerase Amplification
- the amplification method is Nucleic Acid Sequence Based Amplification (NASBA).
- NASBA also known as 3 SR, and transcription-mediated amplification
- 3 SR is an isothermal transcription-based RNA amplification system.
- Three enzymes avian myeloblastosis virus reverse transcriptase, RNase Hand T7 DNA dependent RNA polymerase
- NASBA can be used to amplify DNA.
- the amplification reaction is performed at 41 °C, maintaining constant temperature, typically for about 60 to about 90 minutes (see, e.g., Fakruddin, et al., 2012, "Nucleic Acid Sequence Based Amplification (NASBA) Prospects and Applications,” Int. J. of Life Science and Pharma Res. 2(l):L106-L121).
- the NASBA reaction is carried out at about 40 °C to about 42 °C. In some embodiments, the NASBA reaction is carried out at 41 °C. In some embodiments, the NASBA reaction is carried out at at most about 42 °C. In some embodiments, the NASBA reaction is carried out at about 40 °C to about 41 °C, about 40 °C to about 42 °C, or about 41 °C to about 42 °C. In some embodiments, the NASBA reaction is carried out at about 40 °C, about 41 °C, or about 42 °C.
- Each SDA cycle consists of (1) primer binding to a displaced target fragment, (2) extension of the primer/target complex by exoKlenow, (3) nicking of the resultant hemiphosphothioate Hindi site, (4) dissociation ofHindl from the nicked site and (5) extension of the nick and displacement of the downstream strand by exo-Klenow.
- this disclosure provides for methods comprising contacting DNA in a sample with a helicase.
- the amplification method is Helicase Dependent Amplification (HD A).
- HDA is an isothermal reaction because a helicase, instead of heat, is used to denature DNA.
- the amplification method is Multiple Displacement Amplification (MDA).
- MDA is an isothermal, strand-displacing method based on the use of the highly processive and strand-displacing DNA polymerase from bacteriophage 029, in conjunction with modified random primers to amplify the entire genome with high fidelity. It has been developed to amplify all DNA in a sample from a very small amount of starting material.
- MDA 029 DNA polymerase is incubated with dNTPs, random hexamers and denatured template DNA at 30°C for 16 to 18 hours and the enzyme must be inactivated at high temperature (65°C) for 10 min. No repeated recycling is required, but a short initial denaturation step, the amplification step, and a final inactivation of the enzyme are needed.
- the amplification method is Rolling Circle Amplification (RCA).
- RCA is an isothermal nucleic acid amplification method which allows amplification of the probe DNA sequences by more than 109 fold at a single temperature, typically about 30°C. Numerous rounds of isothermal enzymatic synthesis are carried out by 029 DNA polymerase, which extends a circle-hybridized primer by continuously progressing around the circular DNA probe.
- the amplification reaction is carried out using RCA, at about 28°C to about 32°C.
- RCA is used to amplify ligated molecular inversion probes, optionally ligated barcoded molecular inversion probes.
- amplifying comprises performing an exponential amplification reaction (EXP AR), which is an isothermal molecular chain reaction in that the products of one reaction catalyze further reactions that create the same products.
- EXP AR exponential amplification reaction
- amplifying occurs in the presence of an endonuclease.
- the endonuclease may be a nicking endonuclease.
- iSDA Isothermal strand displacement amplification
- this disclosure provides for methods comprising performing multiple cycles of nucleic acid amplification with a pair of primers.
- the number of amplification cycles is important because amplification may introduce a bias into the representation of regions. Not all regions amplify with the same efficiency and therefore the overall representation may not be uniform which will impact the accuracy of the analysis. Fewer amplification cycles are ideal if amplification is necessary at all.
- methods comprise performing fewer than 50, 45, 40, 35, 30, or 25 cycles of amplification. In some embodiments, methods comprise performing fewer than 25 cycles of amplification. In some embodiments, methods comprise performing fewer than 20 cycles of amplification. In some embodiments, methods comprise performing fewer than 15 cycles of amplification.
- methods comprise performing fewer than 12 cycles of amplification. In some embodiments, methods comprise performing fewer than 11 cycles of amplification. In some embodiments, methods comprise performing fewer than 10 cycles of amplification. In some embodiments, methods comprise performing at least 3 cycles of amplification. In some embodiments, methods comprise performing at least 5 cycles of amplification. In some embodiments, methods comprise performing at least 8 cycles of amplification. In some embodiments, methods comprise performing at least 10 cycles of amplification.
- the amplification reaction is carried for about 30 seconds to about 90 minutes. In some embodiments, the amplification reaction is carried out for at least about 30 minutes. In some embodiments, the amplification reaction is carried out for at most about 90 minutes.
- the amplification reaction is carried out for about 30 minutes to about 35 minutes, about 30 minutes to about 40 minutes, about 30 minutes to about 45 minutes, about 30 minutes to about 50 minutes, about 30 minutes to about 55 minutes, about 30 minutes to about 60 minutes, about 30 minutes to about 65 minutes, about 30 minutes to about 70 minutes, about 30 minutes to about 75 minutes, about 30 minutes to about 80 minutes, about 30 minutes to about 90 minutes, about 35 minutes to about 40 minutes, about 35 minutes to about 45 minutes, about 35 minutes to about 50 minutes, about 35 minutes to about 55 minutes, about 35 minutes to about 60 minutes, about 35 minutes to about 65 minutes, about 35 minutes to about 70 minutes, about 35 minutes to about 75 minutes, about 35 minutes to about 80 minutes, about 35 minutes to about 90 minutes, about 40 minutes to about 45 minutes, about 40 minutes to about 50 minutes, about 40 minutes to about 55 minutes, about 40 minutes to about 60 minutes, about 40 minutes to about 65 minutes, about 40 minutes to about 70 minutes, about 40 minutes to about 75 minutes, about 40 minutes to about 80 minutes, about 40 minutes to about 90 minutes, about 40 minutes to about
- this disclosure provides for methods comprising amplifying a nucleic acid at at least one temperature. In some embodiments, this disclosure provides for methods comprising amplifying a nucleic acid at a single temperature (e.g., isothermal amplification). In some embodiments, this disclosure provides for methods comprising amplifying a nucleic acid, wherein the amplifying occurs at not more than two temperatures. Amplifying may occur in one step or multiple steps. Non-limiting examples of amplifying steps include double strand denaturing, primer hybridization, and primer extension.
- At least one step of amplifying occurs at room temperature. In some embodiments, all steps of amplifying occur at room temperature. In some embodiments, at least one step of amplifying occurs in a temperature range. In some embodiments, all steps of amplifying occur in a temperature range. In some embodiments, the temperature range is about 0° C to about 100°C. In some embodiments, the temperature range is about 15°C to about 100°C. In some embodiments, the temperature range is about 25°C to about 100°C. In some embodiments, the temperature range is about 35°C to about 100°C. In some embodiments, the temperature range is about 55°C to about 100°C. In some embodiments, the temperature range is about 65°C to about 100°C.
- the temperature range is about 15°C to about 80°C. In some embodiments, the temperature range is about 25°C to about 80°C. In some embodiments, the temperature range is about 35°C to about 80°C. In some embodiments, the temperature range is about 55°C to about 80°C. In some embodiments, the temperature range is about 65°C to about 80°C. In some embodiments, the temperature range is about 15°C to about 60°C. In some embodiments, the temperature range is about 25°C to about 60°C. In some embodiments, the temperature range is about 35°C to about 60°C. In some embodiments, the temperature range is about 15°C to about 40°C. In some embodiments, the temperature range is about -20°C to about 100°C.
- the temperature range is about -20°C to about 90°C. In some embodiments, the temperature range is about -20°C to about 50°C. In some embodiments, the temperature range is about -20°C to about 40°C. In some embodiments, the temperature range is about -20°C to about 10°C. In some embodiments, the temperature range is about 0°C to about 100°C. In some embodiments, the temperature range is about 0°C to about 40°C. In some embodiments, the temperature range is about 0°C to about 30°C. In some embodiments, the temperature range is about 0°C to about 20°C. In some embodiments, the temperature range is about 0°C to about 10°C.
- the temperature range is about 15°C to about 100°C. In some embodiments, the temperature range is about 15°C to about 90°C. In some embodiments, the temperature range is about 15°C to about 80°C. In some embodiments, the temperature range is about is about 15°C to about 70°C. In some embodiments, the temperature range is about 15°C to about 60°C. In some embodiments, the temperature range is about 15°C to about 50 °C. In some embodiments, the temperature range is about 15°C to about 30°C. In some embodiments, the temperature range is about 10°C to about 30°C. In some embodiments, methods disclose herein are performed at room temperature, not requiring cooling, freezing or heating. In some embodiments, amplifying comprises contacting the sample with random oligonucleotide primers.
- amplifying comprises targeted amplification (including that described in US6558928).
- amplifying a nucleic acid comprises contacting a nucleic acid with at least one primer having a sequence corresponding to a target gene sequence.
- amplifying comprises contacting the nucleic acid with at least one primer having a sequence corresponding to a non-target gene sequence.
- amplifying comprises contacting the nucleic acid with not more than one pair of primers, wherein each primer of the pair of primers comprises a sequence corresponding to a sequence on a target gene disclosed herein.
- amplifying comprises contacting the nucleic acid with multiple sets of primers, wherein each of a first pair in a first set and each of a pair in a second set are all different.
- amplifying comprises contacting the sample with at least one primer having a sequence corresponding to a sequence on a target gene disclosed herein. In some embodiments, amplifying comprises contacting the sample with at least one primer having a sequence corresponding to a sequence on a non-target gene disclosed herein. In some embodiments, amplifying comprises contacting the sample with not more than one pair of primers, wherein each primer of the pair of primers comprises a sequence corresponding to a sequence on a target gene disclosed herein. In some embodiments, amplifying comprises contacting the sample with multiple sets of primers, wherein each of a first pair in a first set and each of a pair in a second set are all different.
- amplifying comprises multiplexing (nucleic acid amplification of a plurality of nucleic acids in one reaction). In some embodiments, multiplexing comprises contacting nucleic acids of the biological sample with a plurality of oligonucleotide primer pairs. In some embodiments, multiplexing comprising contacting a first nucleic acid and a second nucleic acid, wherein the first nucleic acid corresponds to a first sequence and the second nucleic acid corresponds to a second sequence. In some embodiments, the first sequence and the second sequence are the same. In some embodiments, the first sequence and the second sequence are different. In some embodiments, amplifying does not comprise multiplexing.
- amplifying does not require multiplexing. In some instance, amplifying comprises nested primer amplification. Methods may comprise multiplex PCR of multiple regions, wherein each region comprises a SNP. Multiplexing may occur in a single tube. In some embodiments, methods comprise multiplex PCR of more than 100 regions wherein each region comprises a SNP. In some embodiments, methods comprise multiplex PCR of more than 500 regions wherein each region comprises a SNP. In some embodiments, methods comprise multiplex PCR of more than 1000 regions wherein each region comprises a SNP. In some embodiments, methods comprise multiplex PCR of more than 2000 regions wherein each region comprises a SNP. In some embodiments, methods comprise multiplex PCR of more than 300 regions wherein each region comprises a SNP.
- this disclosure provides for methods comprising amplifying a nucleic acid in the sample, wherein amplifying comprises contacting the sample with at least one oligonucleotide primer, wherein the at least one oligonucleotide primer is not active or extendable until it is in contact with the sample. In some embodiments, amplifying comprises contacting the sample with at least one oligonucleotide primer, wherein the at least one oligonucleotide primer is not active or extendable until it is exposed to a selected temperature.
- amplifying comprises contacting the sample with at least one oligonucleotide primer, wherein the at least one oligonucleotide primer is not active or extendable until it is contacted with an activating reagent.
- the at least one oligonucleotide primer comprises a blocking group. Using such oligonucleotide primers may minimize primer dimers, allow recognition of unused primer, and/or avoid false results caused by unused primers.
- amplifying comprises contacting the sample with at least one oligonucleotide primer comprising a sequence corresponding to a sequence on a target gene disclosed herein.
- this disclosure provides for methods comprising detecting an amplification product, wherein the amplification product is produced by amplifying at least a portion of a target gene disclosed herein.
- detecting amplification products disclosed herein does not comprise barcoding or labeling the amplification product.
- methods detect the amplification product based on its amount. For example, the methods may detect an increase in the amount of double stranded DNA in the sample.
- detecting the amplification product is at least partially based on its size.
- the amplification product has a length of about 50 base pairs to about 500 base pairs.
- this disclosure provides for methods comprising modifying genomic nucleic acids from the biological sample to produce a library of IMNA or NAPT fragments or amplicons for detection.
- methods comprise modifying nucleic acids for nucleic acid sequencing.
- methods comprise modifying nucleic acids for detection, wherein detection does not comprise nucleic acid sequencing.
- methods comprise modifying nucleic acids for detection, to obtain nucleic acids derived from the isolated myeloid nucleic acids (IMNA) or in nucleic acids processed therefrom (NAPT) in the initial sample, wherein detection comprises counting IMNA-adapters or NAPT-adapters nucleic acids based on an occurrence of IMNA-adapters, NAPT-adapters, Mis (e.g., UMIs or non-unique Mis), primers, or primer binding sites.
- this disclosure provides for methods comprising converting in one or more stepa, IMNA or NAPT in the biological sample to a library of IMNA-adapters and NAPT-adapters, wherein the method comprises amplifying the nucleic acids.
- modifying occurs before amplifying.
- modifying occurs after amplifying.
- modifying the nucleic acids comprises repairing ends of nucleic acids that are fragments of a nucleic acid.
- repairing ends may comprise restoring a 5’ phosphate group, a 3’ hydroxy group, or a combination thereof to the nucleic acid.
- repairing comprises 5’ phosphorylation, A-tailing, gap filling, closing nick sites or a combination thereof.
- repairing may comprise removing overhangs.
- repairing may comprise filling in overhangs with complementary nucleotides.
- modifying the nucleic acids for preparing a library comprises use of an adapter.
- the adapter may also be referred to herein as a sequencing adapter.
- the adapter aids in sequencing.
- the adapter can comprise an oligonucleotide.
- the adapter may simplify other steps in the methods, such as amplifying, purification and sequencing because it is a sequence that is universal to multiple, if not all, nucleic acids in a sample after modifying.
- modifying the nucleic acids comprises ligating an adapter to the nucleic acids. Ligating may comprise blunt ligation.
- modifying the nucleic acids comprises hybridizing an adapter to the nucleic acids.
- the sequencing adapter comprises a hairpin or stem-loop adapter.
- modifying the nucleic acids comprises hybridizing a hairpin or stem-loop adapter to the nucleic acids, thereby generating a circular library product that is sequenced or analyzed.
- the sequencing adapter comprises a blocked 5’ end leaving a nick at the 3’ end. Advantages of this configuration include, but are not limited to, an increase in library efficiency and reduction of unwanted byproducts such as adapter dimers.
- the adapter has a cleavable replication stop to linearize templates.
- modifying the nucleic acids for preparing a library comprises use of an IMNA-adapters, NAPT-adapters, Mis, e.g, non-unique Mis or UMIs, primers, or primer binding sites.
- the IMNA-adapters, NAPT-adapters, Mis, e.g, non-unique Mis or UMIs, primers, or primer binding sites may also be referred to herein as a barcode.
- this disclosure provides for methods comprising modifying nucleic acids with a barcode that corresponds to a chromosomal region of interest.
- this disclosure provides for methods comprising modifying nucleic acids with a barcode that is specific to a chromosomal region that is not of interest. In some embodiments, this disclosure provides for methods comprising modifying a first portion of nucleic acids with a first barcode that corresponds to at least one chromosomal region that is of interest and a second portion of nucleic acids with a second barcode that corresponds to at least one chromosomal region that is not of interest. In some embodiments, modifying the nucleic acids comprises ligating a barcode to the nucleic acids. Ligating may comprise blunt ligation. In some embodiments, modifying the nucleic acids comprises hybridizing a barcode to the nucleic acids.
- the barcodes comprise oligonucleotides. In some embodiments, the barcodes comprise a non-oligonucleotide marker or label that can be detected by means other than nucleic acid analysis. In some embodiments, a non-oligonucleotide marker or label could comprise a fluorescent molecule, a nanoparticle, a dye, a peptide, or other detectable/quantifiable small molecule. [00156] In some embodiments, modifying the nucleic acids for preparing a library comprises use of a sample index, also simply referred to herein as an index.
- the index may comprise an oligonucleotide, a small molecule, a fluorescent molecule, a dye, or other detectable/quantifiable moiety.
- a first group of nucleic acids from a first biological sample are labeled with a first index
- a first group of nucleic acids from a first biological sample are labeled with a second index, wherein the first index and the second index are different.
- methods disclose amplifying nucleic acids wherein an oligonucleotide primer used to amplify the nucleic acids comprises an index.
- the methods disclosed herein comprise amplification and library preparation in anticipation of downstream sequencing.
- An assay can include one or two PCR master mixes, one or two thermostable polymerases, one or two primer mixes and library adaptors.
- a sample of DNA may be amplified for a number of cycles by using a first set of amplification primers that comprise target specific regions and non-target specific barcode regions and a first PCR master mix.
- the barcode region can be any sequence, such as a universal barcode region, a capture barcode region, an amplification barcode region, a sequencing barcode region, a UMI (unique molecular identifier) barcode region, and the like.
- a barcode region can be the template for amplification primers utilized in a second or subsequent round of amplification, for example for library preparation.
- the methods comprise adding single stranded-binding protein to the first amplification products.
- An aliquot of the first amplified sample can be removed and amplified a second time using a second set of amplification primers that are specific to the barcode region, e.g., a universal barcode region or an amplification barcode region, of the first amplification primers which may comprise of one or more additional barcode sequences, such as sequence barcodes specific for one or more downstream sequencing workflows, and the same or a second PCR master mix.
- a library of the original DNA sample can be ready for sequencing.
- the method of preparing gene amplicons or enriched nucleic acid samples further comprises detecting the clonally amplified gene amplicons by multiplex PCR or sequencing.
- sequencing comprises targeted sequencing. In some embodiments, sequencing comprises whole genome sequencing. In some embodiments, sequencing comprises targeted sequencing and whole genome sequencing. In some embodiments, whole genome sequencing comprises massive parallel sequencing. In some embodiments, whole genome sequencing comprises random massive parallel sequencing. In some embodiments, sequencing comprises random massive parallel sequencing of target regions captured from a whole genome library.
- the sequencing can be performed by sequencing by synthesis (“SBS).
- SBS sequencing by synthesis
- sequencing comprises a flow cell wherein nucleic acids are attached at (or through a hydrogel to) fixed locations in an array such that their relative positions do not change and wherein the array is repeatedly imaged. Examples in which images are obtained in different color channels, for example, coinciding with different labels used to distinguish one nucleotide base type from another are particularly applicable.
- the different nucleotides present in a sequencing reagent can have different labels and they can be distinguished using appropriate optics as exemplified by the sequencing methods presently commercialized by Illumina, Inc. (San Diego, CA), Ultima Genomics, Pacific Biosciences, Complete Genomics (MGI), ThermoFisher, Element Biosciences, or Singular Genomics.
- SBS involves pyrosequencing.
- Pyrosequencing detects the release of inorganic pyrophosphate (PPi) or protons as particular nucleotides are incorporated into the nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) “Real-time DNA sequencing using detection of pyrophosphate release.” Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) “Pyrosequencing sheds light on DNA sequencing.” Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) “A sequencing method based on real-time pyrophosphate.” Science 281(5375), 363; U.S. Pat. No. 6,210,891; the disclosures of which are incorporated herein by reference in their entireties).
- PPi inorganic pyrophosphate
- the pyrosequencing mechanism involves detecting released PPi by being enzymatically converted to adenosine triphosphate (ATP) by ATP sulfurylase, such that the amount of ATP generated is detected via luciferase-produced photons.
- the nucleic acids to be sequenced can be attached to features in an array comprising wells to localize the released ATP and reduce or prevent ATP crossover from well to well, and the array can be imaged to capture the chemiluminescent signals that are produced due to incorporation of a nucleotides at the features of the array.
- An image can be obtained after the array is treated with a particular nucleotide type (e.g. A, T, C or G).
- SBS methods involve detection of a proton released upon incorporation of a nucleotide into an extension product.
- sequencing based on detection of released protons can use an electrical detector and associated techniques that are commercially available from Thermo Fisher Scientific as the Ion Torrent instrument or sequencing methods and systems described in US 8262900A1, incorporated herein by reference.
- Methods set forth herein for amplifying target nucleic acids using kinetic exclusion can be readily applied to substrates used for detecting protons. More specifically, methods set forth herein can be used to produce clonal populations of amplicons that are used to detect protons.
- SBS involves cycle sequencing by the stepwise addition of reversible terminator nucleotides which comprise a cleavable or photobleachable dye label (as described in U.S. Pat. No. 7,057,026, the disclosure of which is incorporated herein by reference).
- This approach is commercialized Illumina Inc., and is also described in WO 91/06678 (also referred to as US 427,321, Tsien et al.) and U.S. Patent No. 8241573, each of which is incorporated herein by reference.
- the availability of fluorescently-labeled terminators in which both the termination can be reversed and the fluorescent label cleaved facilitates efficient cyclic reversible termination sequencing.
- Polymerases can also be engineered to efficiently incorporate and extend from these modified nucleotides. Additional exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Patent Nos. 7541444, 7566537, 7057026, 8460910, 8623628, 8951781, and 9193996, the disclosures of which are incorporated herein by reference in their entireties.
- the library fragments are immobilized on a substrate, for example a slide, which comprises homologous oligonucleotide sequences for capturing and immobilizing the DNA library fragments.
- the immobilized DNA library fragments are amplified using cluster amplification methodologies as exemplified by the disclosures of U.S. Pat. Nos. 7985565 and 7115400, the contents of each of which is incorporated herein by reference in its entirety. The incorporated materials of U.S. Pat. Nos.
- SBS methods involve detection of four different nucleotides using fewer than four different labels.
- SBS can be performed utilizing methods and systems described in the incorporated materials of U.S. Patent No. 9453258.
- a pair of nucleotide types can be detected at the same wavelength, but distinguished based on a difference in intensity for one member of the pair compared to the other, or based on a change to one member of the pair (e.g. via chemical modification, photochemical modification or physical modification) that causes apparent signal to appear or disappear compared to the signal detected for the other member of the pair.
- three of four different nucleotide types can be detected under particular conditions while a fourth nucleotide type lacks a label that is detectable under those conditions, or is minimally detected under those conditions (e.g., due to background fluorescence).
- Incorporation of the first three nucleotide types into a nucleic acid can be determined based on presence of their respective signals and incorporation of the fourth nucleotide type into the nucleic acid can be determined based on absence or minimal detection of any signal.
- one nucleotide type can include label(s) that are detected in two different channels, whereas other nucleotide types are detected in no more than one of the channels.
- SBS methods can involve a first nucleotide type that is detected in a first channel (e.g. dATP having a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g. dCTP having a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and the second channel (e.g.
- dTTP having at least one label that is detected in both channels when excited by the first and/or second excitation wavelength
- a fourth nucleotide type that lacks a label that is not, or minimally, detected in either channel (e.g. dGTP having no label).
- sequencing data can be obtained using a single channel.
- the first nucleotide type is labeled but the label is removed after the first image is generated, and the second nucleotide type is labeled only after a first image is generated.
- the third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.
- sequencing can be performed by methods capable of single molecule sequencing, with or without amplification or labeling, such as nanopore sequencing (as described in: Deamer, D. W. & Akeson, M. “Nanopores and nucleic acids: prospects for ultrarapid sequencing.” Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, “Characterization of nucleic acids by nanopore analysis”. Ace. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A Golovchenko, “DNA molecules and configurations in a solid-state nanopore microscope” Nat. Mater.
- nanopore sequencing as described in: Deamer, D. W. & Akeson, M. “Nanopores and nucleic acids: prospects for ultrarapid sequencing.” Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, “Character
- the target nucleic acid passes through a nanopore.
- the nanopore can be a synthetic pore or biological membrane protein, such as alpha-hemolysin.
- Each base-pair can be identified by measuring fluctuations in the electrical conductance of the pore as the target nucleic acid passes through the nanopore.
- sequencing can be performed by methods comprising the real-time monitoring of DNA polymerase activity.
- Nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore- bearing polymerase and gamma-phosphate-labeled nucleotides as described, for example, in U.S. Pat. No. 7329492 (which is incorporated herein by reference) or nucleotide incorporations can be detected with zero-mode waveguides as described, for example, in U.S. Pat. No. 7315019 (which is incorporated herein by reference) and using fluorescent nucleotide analogs and engineered polymerases as described, for example, in U.S. Pat. No.
- FRET fluorescence resonance energy transfer
- the illumination can be restricted to a zeptoliter-scale volume around a surface-tethered polymerase such that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M. J. et al. “Zero-mode waveguides for single-molecule analysis at high concentrations.” Science 299, 682-686 (2003); Lundquist, P. M. et al. “Parallel confocal detection of single molecules in real time.” Opt. Lett. 33, 1026-1028 (2008); the disclosures of which are incorporated herein by reference in their entireties).
- Commercial systems employing such technologies are presently commercialized by PacBio (Menlo Park, CA, USA).
- the sequencing methods described herein can be advantageously carried out in multiplex formats such that multiple different target nucleic acids are manipulated simultaneously.
- different target nucleic acids can be treated in a common reaction vessel or on a surface of a particular substrate, allowing for the convenient delivery of sequencing reagents, removal of unreacted reagents and detection of incorporation events in a multiplexed manner.
- the target nucleic acids can be in an array format. In an array format, the target nucleic acids can be typically bound to a surface in a spatially distinguishable manner.
- the target nucleic acids can be bound by direct covalent attachment, attachment to a bead or other particle or binding to a polymerase or other molecule that is attached to the surface.
- the target nucleic acids can be bound through a surface-immobilized hydrogel (as described in U.S. Patent No. 9012022, incorporated herein by reference).
- the array can include a single copy of a target nucleic acid at each site (also referred to as a feature) or multiple copies having the same sequence can be present at each site or feature (also referred to as a colony). Multiple copies can be produced by amplification methods which can include or exclude bridge amplification or emulsion PCR as described in further detail below.
- the methods of the present disclosure may utilize the Illumina, Inc. SBS technology for sequencing the DNA profile libraries created by practicing the methods described herein.
- the Novaseq 6000 platform (a sequencing instrument) was used for clustering and sequencing for the examples described herein. However, as understood by a skilled artisan, the present methods are not limited by the type of sequencing platform used.
- sequencing can be performed using long-read sequencing technologies, such as those developed by Pacific Biosciences (e.g., SMRT sequencing) or Oxford Nanopore Technologies (e.g., nanopore sequencing).
- Long-read sequencing can be advantageous for resolving complex structural variants, phasing mutations over long distances, distinguishing highly homologous gene regions or pseudogenes, and characterizing full-length RNA transcripts including isoforms and fusion genes, especially from nucleic acids derived from the enriched myeloid cell population.
- sequencing libraries are prepared prior to sequencing. Sequencing library preparation involves the production of a random collection of adapter- modified DNA fragments, which are ready to be sequenced. Sequencing libraries of polynucleotides can be prepared from DNA or RNA, including equivalents, analogs of either DNA or cDNA, that is complementary or copy DNA produced from an RNA template, for example by the action of reverse transcriptase.
- the polynucleotides may originate in doublestranded DNA (dsDNA) form (e.g. genomic DNA fragments, PCR and amplification products) or polynucleotides that may have originated in single-stranded form, as DNA or RNA, and been converted to dsDNA form.
- mRNA molecules may be copied into double-stranded cDNAs suitable for use in preparing a sequencing library.
- the precise sequence of the primary polynucleotide molecules is generally not material to the method of library preparation, and may be known or unknown.
- Preparation of sequencing libraries for some sequencing platforms may require that the polynucleotides be or a specific range of fragment sizes e.g. 0-1200 bp. Therefore, fragmentation of polynucleotides e.g. genomic DNA may be required.
- Standard protocols e.g. protocols for sequencing using, for example, the Illumina platforms, instruct users to purify the end-repaired products prior to dA-tailing, and to purify the dA-tailing products prior to the adapter-ligating steps of the library preparation.
- Purification of the end-repaired products and dA-tailed products remove enzymes, buffers, salts and the like to provide favorable reaction conditions for the subsequent enzymatic step.
- the steps of end-repairing, dA-tailing and adapter ligating exclude the purification steps.
- the method disclosed herein encompasses preparing a sequencing library that comprises the consecutive steps of end-repairing, dA-tailing and adapter-ligating.
- dA-tailing is not performed.
- adapters are added via blunt-end ligation to one or both strands of a target sequence.
- an amplification reaction is prepared. The amplification step introduces to the adapter ligated template molecules the oligonucleotide sequences required for hybridization to the flow cell.
- the contents of an amplification reaction include appropriate substrates (such as dNTPs), enzymes (e.g. a DNA polymerase) and buffer components required for an amplification reaction.
- amplification of adapter-ligated polynucleotides can be omitted.
- amplification reactions require at least two amplification primers i.e.
- primer oligonucleotides which may be identical, and include an adapter-specific portion, capable of annealing to a primer-binding sequence in the polynucleotide molecule to be amplified (or the complement thereof if the template is viewed as a single strand) during the annealing step.
- the library or templates prepared according to the methods described above can be used for solid-phase nucleic acid amplification.
- solid-phase amplification refers to any nucleic acid amplification reaction carried out on or in association with a solid support such that all or a portion of the amplified products are immobilized on the solid support as they are formed.
- solid-phase PCR solid-phase polymerase chain reaction
- solid phase isothermal amplification which are reactions analogous to standard solution phase amplification, except that one or both of the forward and reverse amplification primers is/are immobilized on the solid support.
- Solid phase PCR covers systems such as emulsions, wherein one primer is anchored to a bead and the other is in free solution, and colony formation in solid phase gel matrices wherein one primer is anchored to the surface, and one is in free solution.
- the library of template polynucleotide can be used in sequencing methods.
- library templates provide templates for whole genome amplification.
- Sequencing of the amplified libraries can be carried out using any suitable sequencing technique as described herein.
- sequencing is performed using sequencing-by-synthesis (SBS), or single molecule sequencing such as nanopore sequencing.
- SBS sequencing-by-synthesis
- nanopore sequencing single molecule sequencing
- this disclosure provides for methods which employ a sequencing technology in which clonally amplified DNA templates or single DNA molecules are sequenced within a flow cell.
- the method employs sequencing of DNA fragments using Illumina's sequencing-by-synthesis (SBS) and reversible terminator-based sequencing chemistry, as described herein.
- template DNA can be genomic DNA e.g. ctDNA.
- genomic DNA from isolated cells is used as the template, and is fragmented into lengths of several hundred base pairs.
- Illumina's sequencing technology involves the attachment of fragmented genomic DNA to a planar, optically transparent surface on which oligonucleotide anchors are bound to a surface-immobilized hydrogel polymer.
- the present disclosure provides a duplex sequencing method for detecting leukemic molecular alterations in isolated myeloid nucleic acids (IMNA) or nucleic acids processed therefrom (NAPT).
- the method involves producing a plurality of IMNA adapters or NAPT adapters, each comprising a Watson strand and a Crick strand.
- the Watson strand of each adapter includes the following elements in a 5' to 3' orientation:
- a second Watson single-stranded strand-distinguishing sequence that includes at least one universal primer binding site, which is different from the universal primer binding site in the first Watson single-stranded strand-distinguishing sequence.
- each adapter can include the following elements in a 5' to 3' orientation:
- the strand-distinguishing sequences allow for the independent amplification and sequencing of the Watson and Crick strands of each adapter.
- the UMI sequences enable the identification and consensus calling of sequence reads originating from the same starting molecule, which is essential for error correction.
- the method can involve amplifying all or a portion of the Watson and Crick strands using at least one strand-distinguishing sequence that includes at least one universal primer. This generates Watson and Crick amplicons that can be further analyzed.
- the method includes an optional step of performing either (a) one or more targeted amplification steps on the Watson and Crick amplicons or their amplified products to generate targeted amplicons, or (b) bait hybridization on the Watson and Crick amplicons to generate selected amplicons.
- Targeted amplification allows for the enrichment of specific regions of interest, while bait hybridization enables the capture of sequences that are complementary to the bait probes.
- the Watson and Crick amplicons are then analyzed by next-generation sequencing (NGS) to obtain sequence reads.
- NGS next-generation sequencing
- the sequence reads are mapped to the reference sequence of the target region and grouped based on their unique molecular identifiers. Sequence reads with the same UMI are considered to have originated from the same starting molecule and are used to generate a consensus sequence.
- a key step in the duplex sequencing method is the identification of sequence variants that are not present in all sequence reads from the Watson and Crick amplicons derived from the same original molecule. Such variants are likely to be artifacts introduced during amplification or sequencing, while true variants should be present in both the Watson and Crick strands. By comparing the sequence reads from the Watson and Crick amplicons with the same UMI, errors can be identified and removed, increasing the accuracy of variant calling.
- the method further includes identifying variants that are not present in all sequence reads from any one target sequence or selected sequence in Watson amplicons derived from any one Watson IMNA or NAPT and its corresponding Crick amplicons derived from any one Crick IMNA or NAPT. This allows for the identification of errors that may be specific to certain target sequences or selected regions.
- the duplex sequencing method described herein can be used to detect a wide range of leukemic molecular alterations, including single nucleotide variants, insertions, deletions, and structural rearrangements. The high accuracy of the method enables the detection of low- frequency variants that may be missed by conventional sequencing approaches.
- the duplex sequencing method may involve several steps, including: (i) producing IMNA or NAPT adapters with specific configurations of Mis, e.g, non-unique Mis or UMIs and strand-distinguishing sequences on the Watson and Crick strands, (ii) amplifying the adapters to obtain Watson and Crick amplicons, (iii) optionally performing targeted amplification or bait hybridization on the amplicons, and (iv) analyzing the amplicons by NGS to obtain sequence reads.
- an additional step of (v) identifying variants that are not present in all sequence reads from the Watson and Crick amplicons derived from the same original molecule is performed. This step helps to filter out errors and identify true variants based on their presence on both strands of the original DNA duplex.
- the duplex sequencing method is used to detect minimal residual disease (MRD) in patients with myeloid malignancies.
- MRD refers to the presence of residual leukemic cells below the limit of detection of conventional morphologic or cytogenetic methods.
- IMNA or NAPT By analyzing IMNA or NAPT from patient samples using duplex sequencing, MRD can be detected with high sensitivity, allowing for early intervention and personalized treatment decisions.
- the duplex sequencing method is used to monitor clonal evolution and the emergence of therapy-resistant subclones in myeloid malignancies.
- serial analyses of IMNA or NAPT from patient samples before, during, and after treatment changes in the genetic landscape of the leukemic cells can be tracked over time. This can provide insights into the mechanisms of drug resistance and guide the selection of alternative therapies.
- the duplex sequencing method is used to identify novel driver mutations or cooperating mutations in myeloid malignancies.
- novel driver mutations or cooperating mutations in myeloid malignancies By analyzing IMNA or NAPT from a large cohort of patients, recurrent mutations that are not detected by conventional sequencing methods may be identified. These mutations may represent novel therapeutic targets or prognostic biomarkers.
- the duplex sequencing method is used to characterize the clonal architecture and evolution of myeloid malignancies.
- IMNA or NAPT By analyzing IMNA or NAPT from different cell populations or at different time points, the relative abundance and genetic composition of different subclones can be determined. This can provide insights into the order of acquisition of mutations and the evolutionary trajectories of the leukemic cells.
- the strand-distinguishing sequences on the Watson and Crick strands are of different lengths.
- the strand-distinguishing sequence on the Watson strand may be longer or shorter than the strand-distinguishing sequence on the Crick strand. This length difference can provide an additional level of strand discrimination during amplification and sequencing.
- the strand-distinguishing sequences on the Watson and Crick strands have different base compositions.
- the strand-distinguishing sequence on the Watson strand may be rich in G/C bases, while the strand-distinguishing sequence on the Crick strand may be rich in A/T bases.
- This difference in base composition can affect the melting temperature and hybridization specificity of the sequences, allowing for selective amplification or sequencing of one strand over the other.
- the strand-distinguishing sequences on the Watson and Crick strands are artificial sequences that are not derived from or homologous to any naturally occurring genomic sequences. This can minimize the risk of non-specific hybridization or amplification of unintended targets.
- the strand-distinguishing sequences on the Watson and Crick strands are selected to have minimal self-complementarity or secondary structure. This can improve the efficiency and specificity of primer binding and extension during amplification and sequencing.
- the strand-distinguishing sequences on the Watson and Crick strands are designed to be compatible with specific sequencing platforms or chemistries.
- the sequences may be selected to optimize cluster generation, primer hybridization, or nucleotide incorporation on a particular sequencing instrument.
- the strand-distinguishing sequences on the Watson and Crick strands are used as binding sites for strand-specific sequencing primers.
- the sequences of the two strands can be determined independently and used for error correction or haplotype phasing.
- the strand-distinguishing sequences on the Watson and Crick strands are used as binding sites for strand-specific probes or baits. This can allow for the selective capture or enrichment of one strand over the other, which can be useful for applications such as targeted sequencing or strand-specific gene expression analysis.
- the strand-distinguishing sequences on the Watson and Crick strands are used as primer binding sites for strand-specific PCR. By using different primers for the Watson and Crick strands, the two strands can be amplified separately and used for downstream applications such as sequencing or cloning.
- the strand-distinguishing sequences on the Watson and Crick strands are used as barcodes for strand-specific labeling or detection.
- the sequences may be conjugated to different fluorescent dyes or affinity tags that allow for the visual differentiation or physical separation of the two strands.
- the methods comprise generating a library from the IMNA or NAPT in the biological sample, e.g., a library of IMNA or NAPT, adapter-IMNA or NAPT and/or enriched IMNA or NAPT or enriched adapter-IMNA or NAPT, e.g., bait enriched or amplification products with or without adapters.
- the methods comprise determining the sequences of the barcode library.
- the barcode sample is obtained or derived from a sample obtained from a human.
- each of the plurality of primers comprises one or more barcode sequences.
- the one or more barcode sequences comprise a primer barcode, a capture barcode, a sequencing barcode, a unique molecular identifier barcode, or a combination thereof.
- the one or more barcode sequences comprise a primer barcode.
- the one or more barcode sequences comprise a unique molecular identifier (UMI) barcode.
- nucleic acid sequencing may comprise sequencing at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300 or more nucleotides or base pairs of the nucleic acid molecule sequences.
- sequencing may comprise sequencing at least about 200, 300, 400, 500, 600, 700, 800, 900, 1,000 or more nucleotides or base pairs of the nucleic acid molecule sequences.
- sequencing may comprise sequencing at least about 1,500; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; 10,000; 50,000 or more nucleotides or base pairs of the nucleic acid molecule sequences.
- nucleic acid sequencing may comprise at least about 200, 300, 400, 500, 600, 700, 800, 900, 1,000 or more sequencing reads per run. In some embodiments, sequencing may comprise sequencing at least about 1,500; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; or 10,000 or more sequencing reads per run. In some embodiments, nucleic acid sequencing may comprise at least about 10,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; or 100,000 or more sequencing reads per run. In some embodiments, nucleic acid sequencing may comprise at least about 250,000; 500,000; 1,000,000; 10,000,000; 100,000,000; or 1,000,000,000 or more sequencing reads per run. In some embodiments, nucleic acid sequencing may comprise less than or equal to about 1,600,000,000 sequencing reads per run. In some embodiments, nucleic acid sequencing may comprise less than or equal to about 200,000,000 reads per run.
- the number of times a single nucleotide or polynucleotide is identified or “read” is defined as the sequencing depth or read depth, or fold coverage, optionally describing a percentage of bases.
- Read depth (sequencing depth, or sampling) represents the total number of times a sequenced nucleic acid fragment (a “read”) is obtained for a sequence.
- Theoretical read depth is defined as the expected number of times the same nucleotide is read, assuming reads are perfectly distributed throughout an idealized genome. Read depth is expressed as function of percentage coverage (or coverage breadth). For example, 10 million reads of a 1 million base genome, perfectly distributed, theoretically results in 10X read depth of 100% of the sequences. In practice, a greater number of reads (higher theoretical read depth, or oversampling) may be needed to obtain the desired read depth for a percentage of the target sequences.
- sequencing is peformed at a read depth of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550. 600, 650, 700, 750, 800, 850, 900, 950, or at least 1000, at least 10,000, at least 20,000, at least 30,000, at least 50,0000, at least 75,000, or least 100,000 unique reads per base.
- sequencing may comprise a read depth of at least IX, 5X, 10X, 20X, 30X, 40X, 50X, 60X, 70X, 80X, 90X, 100X, or more.
- Enrichment of target sequences with a controlled stoichiometry probe library increases the efficiency of downstream sequencing, as fewer total reads will be required to obtain an outcome with an acceptable number of reads over a desired % of target sequences.
- 55X theoretical read depth of target sequences results in at least 3 OX coverage of at least 90% of the sequences.
- no more than 55X theoretical read depth of target sequences results in at least 3 OX read depth of at least 80% of the sequences.
- no more than 55X theoretical read depth of target sequences results in at least 3 OX read depth of at least 95% of the sequences.
- no more than 55X theoretical read depth of target sequences results in at least 10x read depth of at least 98% of the sequences. In some instances, 55X theoretical read depth of target sequences results in at least 20X read depth of at least 98% of the sequences. In some instances, no more than 55X theoretical read depth of target sequences results in at least 5X read depth of at least 98% of the sequences.
- Increasing the concentration of probes during hybridization with targets can lead to an increase in read depth. In some instances, the concentration of probes is increased by at least 1.5X, 2. OX, 2.5X, 3X, 3.5X, 4X, 5X, or more than 5X.
- increasing the probe concentration results in at least a 1000% increase, or a 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 750%, 1000%, or more than a 1000% increase in read depth. In some instances, increasing the probe concentration by 3X results in a 1000% increase in read depth.
- Homology refers to the percent identity between two polynucleotides or two polypeptide sequences.
- Two DNA or polypeptide sequences are “homologous” to each other when the sequences exhibit at least about 75% to 85% (including 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, and 85%), at least about 90%, or at least about 95% to 99% (including 95%, 96%, 97%, 98%, 99%) contiguous sequence identity over a defined length of the sequences.
- RAPGEF3 gene amplicons are obtained by using a forward primer having 80% homology or higher to that of the nucleotide sequence 5’ CTTCCTTCATTTCTCCACCTG 3’ (SEQ ID NO: 1) and the reverse primer having 80% homology or higher to that of the nucleotide sequence 5’ TCTGTGTCCTCTTGCCTGC 3’ (SEQ ID NO: 2).
- Identity or homology with respect to a specified amino acid sequence of this invention is defined herein as the percentage of amino acid residues in a candidate sequence that are identical with the specified residues, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent homology, and not considering any conservative substitutions as part of the sequence identity. None of N-terminal, C-terminal or internal extensions, deletions, or insertions into the specified sequence shall be construed as affecting homology. All sequence alignments called for herein are such maximal homology alignments.
- nucleic acid sequence homology between the polynucleotides, oligonucleotides, and fragments disclosed herein and a nucleic acid sequence of interest will be at least 80% or greater, and more typically with preferably increasing homologies of at least 85%, 90%, 91%, 92%, 92%, 94%, 95%, 96%, 97%, 98%, 99%, and/or 100%.
- Two amino acid sequences are homologous if there is a partial or complete identity between their sequences.
- the practice of the present disclosure will employ, together with the methods featured herein, unless otherwise indicated molecular biology, microbiology, recombinant DNA, and immunology techniques. See, e.g., Molecular Cloning A Laboratory Manual, 2nd Ed., ed. by Sambrook, Fritsch and Maniatis (Cold Spring Harbor Laboratory Press, 1989); DNA Cloning, Volumes I and II (D. N. Glover ed., 1985); Culture Of Animal Cells (R. I. Freshney, Alan R. Liss, Inc., 1987); Immobilized Cells And Enzymes (IRL Press, 1986); B.
- Polymorphic sites that are contained in the target nucleic acids include without limitation single nucleotide polymorphisms (SNPs), tandem SNPs, small-scale multi-base deletions or insertions, also referred to as “IN-DELS” or deletion insertion polymorphisms “DIPs”, Multi -Nucleotide Polymorphisms “MNPs”, and Short Tandem Repeats “STRs”.
- SNPs single nucleotide polymorphisms
- tandem SNPs small-scale multi-base deletions or insertions
- DIPs deletion insertion polymorphisms
- MNPs Multi -Nucleotide Polymorphisms
- STRs Short Tandem Repeats
- the nucleic acids in the sample is enriched for target nucleic acids that comprise at least one SNP.
- each target nucleic acid comprises a single i.e. one SNP.
- Target nucleic acid sequences comprising SNPs are available from publically accessible databases including, but not limited to Human SNP Database at world wide web address wi.mit.edu, NCBI dbSNP Home Page at world wide web address ncbi.nlm.nih.gov, world wide web address lifesciences.perkinelmer.com, Celera Human SNP database at world wide web address celera.com, the SNP Database of the Genome Anulysis Group (GAN) at world wide web address gan.iarc.fr.
- GAN Genome Anulysis Group
- SNPs that are encompassed by the method disclosed herein include linked and unlinked SNPs.
- Each target nucleic acid comprises at least one polymorphic site e.g. a single SNP, that differs from that present on another target nucleic acid to generate a panel of polymorphic sites e.g.
- SNPs that contain a sufficient number of polymorphic sites of which at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, a.t least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40 or more are informative.
- a panel of SNPs can be configured, to comprise at. least one informative SNP.
- Amplification of the target sequences can be performed by any method that uses PCR or variations of the method including but not limited to asymmetric PCR, helicasedependent amplification, hot-start PCR, qPCR, solid phase PCR, and touchdown PCR.
- replication of target nucleic acid sequences can be obtained by enzymeindependent methods e.g. chemical solid-phase synthesis using the phosphoramidites.
- Amplification of the target sequences is accomplished using primer pairs each capable of amplifying a target nucleic acid sequence comprising the polymorphic site e.g. SNP, in a multiplex PCR reaction.
- Multiplex PCR reactions include combining at least 2, at least three, at least 3, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30 at least 30, at least 35, at least 40 or more sets of primers in the same reaction to quantify the amplified target nucleic acids comprising at least two, at least three, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 30, at least 35, at least 40 or more polymorphic sites in the same sequencing reaction.
- Any panel of primer sets can be configured to amplify- at least one informative polymorphic sequence.
- the PCR primers used for amplifying target nucleic acids are designed to amplify sequences that are of sufficient length to the bridge amplified and to identify SNPs that are encompassed by the sequence reads.
- the first of two primers in the primer set comprising the forward and the reverse primer for amplifying the target nucleic acid is designed to identify a single SNP present within a sequence read of about 20bp, about 25bp, about 30bp, about 35bp, about 4Ubp, about 45bp, about 50bp, about 55bp, about 60bp, about 65bp, about 70bp, about 75bp, about 80bp, about 85bp, about90bp, about 95bp, about lOObp, about 1 lObp, about 120bp, about 130, about 140bp, about 150bp, about 200bp, about 250bp, about 300bp, about 350bp, about 400bp, about 450bp, or about 500bp.
- the forward and reverse primers are each designed for amplifying target nucleic acids each comprising a set of two tandem SNPs, each being presenl within a sequence read or about 20bp, about 25bp, about 30bp, about 35bp, about 40bp, about 45bp, about 50bp, about 55bp, about 60bp, about 65bp, about 70bp, about 75bp, about 80bp, about 85bp, about 90bp, about 95bp, about lOObp, about HObp, about 120bp, about 130, about 140bp, about 150bp, about 200bp, about 250bp, about 300bp, about 350bp, about 400bp, about 450bp, or about 500bp.
- at least one of the primers is designed to amplify the target nucleic acid comprising a set of two tandem SNPs as an amplicon of sufficient length to allow for bridge amplification.
- the SNPs are contained in amplified target nucleic acid amplicons of at least about lOObp, at least about 150bp, at least about 200bp, at least about 250bp, at least about 300bp, at least about 350bp, or at least about 400bp.
- target nucleic acids comprising a polymorphic site e.g. a SNP
- target nucleic acids comprising two or more polymorphic sites are amplified as amplicons of at least about 110 bp, and that comprise a SNP within 36 bp from the 3’ or 5’ end of the amplicon.
- target nucleic acids comprising two or more polymorphic sites e.g.
- two tandem SNPs are amplified as amplicons of at least about 110 bp, and that comprise the first SNP within 36 bp from the 3’ end of the amplicon, and/or the second SNP within 36 bp from the 5’ end of the amplicon.
- a modulator is a therapeutic agent that alters the expression, level, and/or activity of a gene or gene product (including protein or RNA, e.g., mRNA, miRNA, or piRNA).
- a modulator is administered to a subject having a gene variant, to alter a gene or gene product to similar levels as unaffected subjects who do not have the variant, or to provide similar levels of a gene or gene product as subjects carrying protective gene variants against developing a myeloid malignancy.
- a modulator is administered to a subject to enhance the expression, level, and/or activity of a gene or gene product in subjects.
- a modulator is administered to a subject to restore the expression, level and/or activity of a gene or gene product to baseline or to substantially the same as in a subject carrying protective gene variants, for example to expression, level and/or activity of a non-carrier.
- a modulator is administered to a subject to reduce the expression, level and/or activity of a gene or gene product in subjects.
- the modulator alters gene expression, e.g. by affecting a transcriptional regulator (inhibitor or activator) or inhibitory RNA.
- a modulator alters the gene or gene product levels, e.g. by affecting the RNA product of a gene and affecting its stability or translation.
- a modulator alters the activity of the gene or gene product, e.g. by affecting the protein product directly, activating or inhibiting the protein, or enhancing or preventing its multimerization or binding capacity, which can be through competitive, noncompetitive, or uncompetitive interactions.
- a modulator is an antagonist.
- an antagonist is a therapeutic agent that interferes with or inhibits the physiological action, e.g., expression, level and/or activity of a target, or that inhibits or interferes with the physiological action of a positive-regulator gene or gene product (that increases the expression, level and/or activity of the target gene or gene product or decreases the expression, level and/or activity of an activator of the target).
- the antagonist can inhibit the gene or gene product directly, or activate an inhibitor of the gene or gene product, or inhibit an activator of the gene to reduce the expression, level, and/or activity of the target gene or gene product.
- the antagonist alters expression of the gene, e.g.
- the antagonist alters the gene or gene product levels by destabilizing the RNA of the gene or inhibiting its translation. In some embodiments, the antagonist alters the activity of the gene or gene product, for example by directly inhibiting the protein’s activity, binding, or multimerization, or by competitively, non-competitvely, or uncompetitively binding the protein to decrease its activity.
- the term “antagonist” also refers to enzyme inhibitors.
- a modulator is an agonist.
- an agonist is a therapeutic agent that activates or enhances the the physiological action, e.g., expression, level and/or activity of a target, or that activates a positive regulator of the target (a gene or gene product that increases the expression, level and/or activity of a target or that interferes with or inhibits a negative regulator of the target).
- the agonist can activate the gene or gene product directly, or inhibit an inhibitor of the gene or gene product, or activate or upregulate an activator of the gene to increase the expression, level, and/or activity of the target gene or gene product.
- the agonist alters expression of the gene, e.g.
- the agonist alters the gene or gene product levels by stabilizing the RNA of the gene or enhancing its translation.
- the agonist alters the activity of the gene or gene product, for example by directly activating the protein’s activity, binding, or multimerization, or by competitively, non-competitvely, or uncompetitively binding the protein to increase its activity.
- the term “agonist” also refers to enzyme activators.
- the method of treating a subject having or suspected of having a myeloid malignancy further comprises administering one or more additional active agents or supportive therapies for treating, preventing, or reducing the severity of a myeloid malignancy to the subject.
- the present disclosure provides a method for treating a subject diagnosed with a myeloid malignancy.
- the method involves administering to the subject a therapeutically effective amount of a compound identified using the screening methods described herein.
- the compound may be administered alone or in combination with other therapeutic agents, such as chemotherapy, targeted therapy, or immunotherapy.
- the treatment method involves administering to the subject a compound that targets a specific molecular alteration identified in the subject's myeloid cells using the detection methods described herein. For example, if the subject is found to have a particular gene mutation or fusion protein, a compound that selectively inhibits the activity of that mutant protein may be administered.
- the treatment method involves administering a compound that enhances the immune response against myeloid tumor cells.
- This may include checkpoint inhibitors that release the brakes on T cell activation, or therapeutic vaccines that stimulate the production of tumor-specific T cells.
- the treatment method involves administering a compound that induces differentiation or apoptosis of myeloid tumor cells.
- a compound that induces differentiation or apoptosis of myeloid tumor cells may include agents that target epigenetic regulators, such as DNA methyltransferase or histone deacetylase inhibitors, or that modulate the activity of transcription factors involved in myeloid cell development and survival.
- target epigenetic regulators such as DNA methyltransferase or histone deacetylase inhibitors
- the treatment methods of the present disclosure can be used to achieve various therapeutic goals, such as inducing remission, prolonging survival, reducing tumor burden, alleviating symptoms, or preventing relapse.
- the choice of specific therapeutic agent(s) and dosing regimen will depend on factors such as the type and stage of myeloid malignancy, the molecular profile of the tumor cells, the patient's age and general health status, and the presence of comorbidities or contraindications.
- the treatment method further involves monitoring the patient's response to therapy using the detection methods described herein. Changes in the level or spectrum of leukemic molecular alterations over time can provide an early indication of treatment efficacy or the emergence of drug resistance, allowing for timely adjustment of the therapeutic strategy.
- the treatment methods of the present disclosure can be used in conjunction with conventional supportive care measures, such as transfusion of blood products, administration of prophylactic antibiotics, or management of treatment-related side effects.
- the methods can also be used in the context of hematopoietic stem cell transplantation, either as a means of preparing the patient for transplant or as a post-transplant maintenance therapy.
- the myeloid lineage malignancy treatment can include or exclude one of the following treatments: Arsenic Trioxide, Azacitidine, Cyclophosphamide, Cytarabine, Daunorubicin Hydrochloride, Daunorubicin Hydrochloride and Cytarabine Liposome, Daurismo (Glasdegib Maleate), Dexamethasone, Doxorubicin Hydrochloride, Enasidenib Mesylate, Gemtuzumab Ozogamicin, Gilteritinib Fumarate, Glasdegib Maleate, Idamycin PFS (Idarubicin Hydrochloride), Idarubicin Hydrochloride, Idhifa (Enasidenib Mesylate), Ivosidenib, Midostaurin, Mitoxantrone Hydrochloride, Mylotarg (Gemtuzumab Ozogamicin), Olutasidenib,
- Onureg (Azacitidine), Pemazyre (Pemigatinib), Pemigatinib, Prednisone, Quizartinib Dihydrochloride, Rezlidhia (Olutasidenib), Rituxan (Rituximab), Rituximab, Rydapt (Midostaurin), Tabloid (Thioguanine), Thioguanine, Tibsovo (Ivosidenib), Tisagenlecleucel (Kymriah), Trisenox (Arsenic Trioxide), Vanflyta (Quizartinib Dihydrochloride), Venclexta (Venetoclax), Venetoclax, Vincristine Sulfate, Vyxeos (Daunorubicin Hydrochloride and Cytarabine Liposome), or Xospata (Gilteritinib Fumarate).
- antibody includes, without limitation, a glycoprotein immunoglobulin which binds specifically to an antigen.
- antibody can comprise at least two heavy (H) chains and two light (L) chains interconnected by disulfide bonds, or an antigen-binding portion thereof.
- Each H chain comprises a heavy chain variable region (abbreviated herein as VH) and a heavy chain constant region.
- the heavy chain constant region comprises three constant domains, CHI, CH2 and CH3.
- Each light chain comprises a light chain variable region (abbreviated herein as VL) and a light chain constant region.
- the light chain constant region comprises one constant domain, CL.
- VH and VL regions are further subdivided into regions of hypervariability, termed complementarity determining regions (CDRs), interspersed with regions that are more conserved, termed framework regions (FR).
- CDRs complementarity determining regions
- FR framework regions
- Each VH and VL comprises three CDRs and four FRs, arranged from amino-terminus to carboxy-terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4.
- the variable regions of the heavy and light chains contain a binding domain that interacts with an antigen.
- the constant regions of the Abs may mediate the binding of the immunoglobulin to host tissues or factors, including various cells of the immune system (e.g., effector cells) and the first component (Clq) of the classical complement system.
- An immunoglobulin may derive from any of the commonly known isotypes, including but not limited to IgA, secretory IgA, IgG and IgM.
- IgG subclasses also include but are not limited to human IgGl, IgG2, IgG3 and IgG4.
- “Isotype” refers to the Ab class or subclass (e.g., IgM or IgGl) that is encoded by the heavy chain constant region genes.
- antibody can include or exclude both naturally occurring and non-naturally occurring Abs; monoclonal and polyclonal Abs; chimeric and humanized Abs; human or nonhuman Abs; wholly synthetic Abs; and single chain Abs.
- a nonhuman Ab may be humanized by recombinant methods to reduce its immunogenicity in man.
- the term “antibody” also includes an antigen-binding fragment or an antigen-binding portion of any of the aforementioned immunoglobulins, and includes a monovalent and a divalent fragment or portion, and a single chain Ab.
- dislcosed herein are methods for detecting gene variants in a sample from a subject having or suspected of having a myeloid malignancy.
- dislcosed herein aremethods for preparing the isolated myeloid nucleic acids (IMNA) or in nucleic acids processed therefrom (NAPT).
- IMNA isolated myeloid nucleic acids
- NAPT nucleic acids processed therefrom
- a method of treating a disease or disorder in a subj ect comprising administering to the subj ect a therapeutically effective amount of a complex or a composition as described herein.
- a method of treating a disease or disorder in a subj ect comprising administering to the subj ect a therapeutically effective amount of a complex or a composition as described herein.
- the method comprises collecting whole blood from a patient, isolating peripheral blood mononuclear cells (PBMCs) from the whole blood via Ficoll-Paque extraction, thereby eliminating granulocytes, applying CD3+/CD19+ magnetic dynabeads to eliminate T cells and B cells from the PBMCs, extracting gDNA from the PBMCs, then using the gDNA as input for targeted error-corrected NGS via the Oncomine Myeloid MRD assay to determine the presence of disease-associated variants.
- the Oncomine assay is an Ion Torrent Ampliseq assay.
- Ficoll-Paque solution 15 mL Ficoll-Paque solution is added to a 50 mL centrifuge tube, then the diluted blood is carefully layered onto the Ficoll-Paque, ensuring that the two solutions do not mix.
- the tube is centrifuged at 400 g for 30 to 40 minutes at approximately 18 Celsius.
- the upper layer comprising plasma and platelets is removed, then the lower layer comprising PBMCs is isolated by pipette and transferred to a second tube.
- the total PBMC count may again be estimated by manual or automated method.
- T cells and B cells are removed from the isolated PBMCs via Thermo Fisher Dynabeads.
- Anti-CD3 https://www.thermofisher.com/order/catalog/product/11151D
- anti-CD19 https://www.thermofisher.com/order/catalog/product/11143D
- Dynabeads are mixed in a 1 :1 ratio, then added at a ratio of 25 pL per IxlO 7 input cells. The mixture is incubated for 30 minutes at 8 Celsius with gentle tilting and rotation, then a magnet is applied to the tube for 2 minutes. With the magnet in place, the supernatant is carefully removed by pipette and placed into a new tube.
- the total depleted PBMC cell count may again be estimated by manual or automated method.
- gDNA is isolated from the T and B cell depleted PBMCs via the Qiagen QIAwave DNA Blood and Tissue Kit (Qiagen), using the protocol for cultured cells, eluting the DNA into low TE buffer.
- the concentration of eluted DNA is measured, e.g. via Thermo Fisher NanoDrop or Thermo Fisher Qubit. If the DNA concentration is less than 9 ng/pL, concentrate the sample (e.g. via vacuum centrifugation) prior to proceeding with the next step. Ensure at least approximately 150 ng of gDNA is available before proceeding to the next step.
- process 150 ng of gDNA via the Thermo Fisher Oncomine Myeloid MRD assay following the DNA-only workflow protocol using the S5 550 flow cell e.g. as described in MAN0025670 (Oncomine Myeloid MRD Assay, Thermofisher).
- variant frequencies are adjusted by multiplying each detected variant frequency by the depleted cell count divided by the total blood cell count. Finally, the adjusted variant frequencies are reported for all detected variants.
- a sample may be classified as MRD positive or negative based on the presence and frequency of detected disease-associated variants.
- the gDNA serves as input for a highly sensitive targeted duplex sequencing assay, where both Watson and corresponding Crick strand are sequenced, and variants/mutations are only identified if they are present in both the Watson and corresponding Crick strand.
- such an assay targets for example, regions of from 20 to 35, 40, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or approximately 100 genes.
- multiple probe intervals are used in each region of interest.
- the assay targets for example, approximately regions of approximately 35 to 40, or approximately 36 genes recurrently mutated in AML with over 220 probe intervals.
- such a panel of approximately 36 genes includes, but is not limited to, ASXL1, BCOR, BCORL1, CALR, CBL, CEBPA, CSF3R, DDX41, DNMT3A, ETV6, EZH2, FLT3, GATA2, GNAS, HRAS, IDH1, IDH2, IKZF1, JAK2, KIT, KRAS, KMT2A (MLL), MPL, NF1, NPM1, NRAS, PHF6, PPM1D, PTPN11, RAD21, RUNX1, SETBP1, SF3B1, SRSF2, STAG2, STAT3, TET2, TP53, U2AF1, WT1, and ZRSR2, or a selection thereof relevant for AML MRD detection.
- the probe intervals are designed to cover mutational hotspots, entire coding regions, or specific exons and introns of these genes.
- library preparation is performed using a Duplex Sequencing library preparation kit (e.g., an equivalent to DuplexSeq V2 Library Preparation Kit), followed by sequencing on a suitable NGS platform (e.g., Illumina NovaSeq 6000).
- the error-corrected sequencing data from such an assay detects variants at frequencies below 0.01% VAF, and these ultra-low frequency variant data, after adjustment for cell enrichment as described above, is used to classify a sample as MRD positive or negative with enhanced sensitivity.
- any of the terms “comprising”, “consisting essentially of’, and “consisting of’ may be replaced with either of the other two terms in the specification.
- the terms “comprising”, “including”, containing”, etc. are to be read expansively and without limitation.
- the methods and processes illustratively described herein suitably may be practiced in differing orders of steps, and that they are not necessarily restricted to the orders of steps indicated herein or in the claims. It is also that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural reference unless the context clearly dictates otherwise.
Landscapes
- Chemical & Material Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Organic Chemistry (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Analytical Chemistry (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Engineering & Computer Science (AREA)
- Genetics & Genomics (AREA)
- Immunology (AREA)
- Molecular Biology (AREA)
- General Engineering & Computer Science (AREA)
- Biotechnology (AREA)
- Biophysics (AREA)
- Physics & Mathematics (AREA)
- Biochemistry (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Microbiology (AREA)
- General Health & Medical Sciences (AREA)
- Pathology (AREA)
- Hospice & Palliative Care (AREA)
- Oncology (AREA)
- Chemical Kinetics & Catalysis (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
Described herein are methods for the identification and detection of detecting myeloid malignancies by isolating and analyzing nucleic acids from myeloid lineage cells. Included herewith are methods of obtaining a sample containing myeloid lineage cells, performing negative enrichment to remove lymphoid lineage cells while retaining the myeloid cells, isolating nucleic acids from the enriched myeloid cells, and detecting leukemic molecular alterations in the isolated nucleic acids or nucleic acids processed therefrom. Further, this disclosure provides for enrichment resulting in a several-fold increase in myeloid tumor cells compared to the original sample, enabling sensitive detection of myeloid malignancies, increased on-target detection, and significantly reducing cost of analysis.
Description
METHOD FOR DETECTING MYELOID-LINEAGE MALIGN ANCIES/MYELOID MALIGNANCIES
RELATED APPLICATIONS AND INCORPORATION BY REFERENCE
[0001] This application claims priority to U.S. Provisional Patent Application No. 63/650,807, filed May 22, 2024. The foregoing applications, and all documents cited therein or during their prosecution and all documents cited or referenced in the application cited documents, and all documents cited or referenced herein (herein cited documents), and all documents cited or referenced in herein cited documents, together with any manufacturer’s instructions, descriptions, product specifications, and product sheets for any products mentioned herein or in any document incorporated by reference herein, are hereby incorporated herein by reference, and may be employed in the practice of the disclosed subject matter. More specifically, all referenced documents are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference.
FIELD
[0002] Disclosed herein are subject matters relating to the field of cancer diagnostics, specifically to methods for detecting malignancies of myeloid lineage cells such as acute myeloid leukemia by analyzing nucleic acids isolated from myeloid cell populations enriched through negative selection.
BACKGROUND
[0003] The following includes information that may be useful in understanding the presently disclosed subject matter. It is not an admission that any of the information, publications or documents specifically or implicitly referenced herein is prior art, or essential, to the presently described or claimed subject matter. All publications, patents, related applications, and other written or electronic materials mentioned or identified herein are hereby incorporated herein by reference in their entirety. The information incorporated is as much a part of the application as filed as if all of the text and other content was repeated in the application, and should be treated as part of the text and content of the application as filed.
[0004] Myeloid lineage malignancies, such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), myelodysplastic syndrome (MDS), or myeloproliferative neoplasms (MPNs) are characterized by the accumulation of abnormal myeloid progenitor cells in the bone marrow and peripheral blood. These malignancies are driven by molecular alterations
including mutations, structural variants, and aberrant methylation patterns. Acute myeloid leukemia (AML) is the most common leukemia in adults, accounting for 80% of all cases; despire recent advances in therapy, the prognosis is poor, particularly amongst the elderly. Measurable residual disease (MRD) has been established as an established biomarker for disease prognosis, therapy monitoring, treatment sequencing, and may be used as a surrogate endpoint during drug thereapy clinical trials.
SUMMARY
[0005] Sensitive detection of molecular alterations that are markers of myeloid lineage malignancies such as AML, or other myeloid conditions (CML, MDS, or MPNs) is crucial for diagnosis, prognosis, treatment selection, monitoring treatment, and detecting minimal residual disease (MRD). However, owing to heterogeneous phenotype, genotype, and molecular features of myeloid lineage malignancies, approaches incorporating positive selection for myeloid lineage malignancies are inevitably limited to the subset of patients with myeloid lineage malignancies whose leukemic cells harbor the detected markers. Further, in a clinical setting, essential information regarding the leukemic cell phenotype is often unavailable at the time of the MRD analysis, thus precluding use of positive selection as part of a clinical MRD workflow. Moreover, ccurrent methods for MRD analysis are lacking. For example, flow cytometry -based methods are unable to distinguish circulating leukemic cells from rare, normal cells, resulting in a suboptimal limit-of-detection (LoD) for abnormal cells of 10'3 to 10'4 (i.e., 0.1% to 0.01% abnormal cells). As another example, allele specific oligonucleotide PCR (ASO-PCR, i.e. RT-PCR) or digital droplet PCR (ddPCR) is able to achieve a LoD of 10'6 Variant Allele Frequency (VAF) (www.ncbi.nlm.nih.gov/pmc/articles/PMC8582498) but is only applicable to a subset of patients and may be too operationally complex to deploy in a clinical setting. Additionally, next-generation sequencing (NGS) based methods, by assessing the presence of multiple potential molecular aberrations in parallel, are applicable to the majority of patients and have reduced operational complexity, but are generally cost prohibitive at an LoD below 10'3 owing to the linear relationship between sequencing cost and LoD, and the need for deep sequencing to accommodate error correction methods, e.g. those employing unique or non-unique molecular identifiers (Mis) (e.g. IDs), start-end site coordinates, CODEC, duplex error correction, duplex sequencing, SaferSeq technology, and/or SaferSeqS technology.
[0006] For subjects with myeloid lineage malignancies, the rarity of leukemic cells, especially in early-stage or residual disease, poses challenges for reliable detection. Current methods enrich tumor cells through positive selection using markers expressed on the tumor
cells. However, this can bias the recovered population based on marker expression. Methods are needed to enrich myeloid lineage cells, both normal and transformed, to allow unbiased characterization of molecular alterations associated with myeloid malignancies.
[0007] This disclosure features methods to enable sensitive, on-target, and cost-efficient MRD detection of myeloid lineage malignancies or conditions. In some aspects, this disclosure features an improved method for MRD analysis comprising negative enrichment selection coupled with detecting one or more myeloid condition variants and/or molecular profiling of the enriched myeloid cells. In some aspects, the methods described herein enable cost-effective detection of AML or other myeloid conditions down to 10'6 LoD with an operationally efficient workflow. When coupled with NGS analysis, the negative enrichment workflow can reduce sequencing costs, e.g., by up to 10-fold.
[0008] In one aspect, this disclosure provides methods for preparing nucleic acids from myeloid lineage tumor cells. The methods include obtaining a sample comprising myeloid lineage cells, performing negative selection enrichment to remove lymphoid cells while retaining residual myeloid cells, and isolating nucleic acids from the residual cells of myeloid lineage. In one aspect, the methods featured herein further comprise detecting molecular alterations in the isolated myeloid nucleic acids (IMNA) or nucleic acids processed therefrom (NAPT)
[0009] In one aspect the methods of this disclosure are applied to various sample types, including whole blood, peripheral blood, bone marrow, or fractions thereof such as peripheral blood mononuclear cells (PBMCs). In some aspects, the negative selection enrichment steps may involve density gradient or differential centrifugation, e.g., such as Ficoll density gradient centrifugation, and/or immunomagnetic cell separation using antibodies, e.g., antibody-bead extraction, against lymphoid cell markers, which can include or exclude: B cells, T cells, or NK cells. In an aspect, alternative or complementary negative selection techniques can be employed, including but not limited to, fluorescence-activated cell sorting (FACS) using a panel of antibodies to negatively select for myeloid cells by excluding labeled lymphoid and other non-myeloid populations, or microfluidic chip-based separation technologies that utilize affinity-based depletion or physical property differences to remove lymphoid cells.
[0010] In one aspect, the isolated myeloid nucleic acids are analyzed using next-generation sequencing (NGS) technologies. In one aspect, the NGS assay may be targeted to specific genes or regions frequently mutated in myeloid malignancies. In one aspect, the NGS assay may include one or more myeloid lineage malignancies-informed or myeloid-informed markers selected by comparison of sequencing results from whole genome sequencing or whole exome
sequencing of myeloid lineage cells compared to sequencing results from myeloid lineage cells from patients with AML or other myeloid conditions following diagnosis but prior to any treatment, or before any treatment that significantly reduces leukemic cell levels. The NGS assay may incorporate unique or non-unique molecular identifiers (Mis) and/or strand distinguishing sequences to enable error correction and detection of low-frequency variants. In some aspects, the disclosure provides methods for detecting leukemic alterations using PCR- based assays such as quantitative PCR, digital PCR, or droplet digital PCR. In some aspects, this disclosure provides methods for detecting leukemic alterations using molecular inversion probes, including barcoded molecular inversion probes.
[0011] In another aspect, the disclosure relates to methods for detecting leukemic molecular alterations using an NGS assay on nucleic acids processed from IMNA or NAPT.
[0012] In some aspects, the NGS assay used in the methods of this disclosure includes error correction, such as duplex sequencing or SaferSeqs sequencing.
[0013] In some aspects, the NGS assay is an Ion Torrent AmpliSeq sequencing assay. For example, in one aspect the NGS assay includes highly multiplexed targeted amplification (e.g., up to 24,000 amplicons), sample indexing and sequencing on an Ion Torrent sequencer. In one aspect the NGS assay includes digesting primers following multiplex amplification.
[0014] The methods of this disclosure feature a targeted NGS assay, such as a targeted duplex sequencing assay, where the processed nucleic acids are partially double-stranded and include a double-stranded target sequence, one or more double-stranded Mis located proximal to the ends of the target sequence, and one or more strand-distinguishing sequences located distal to the Mis on each strand. In some aspects the Mis can be unique Mis (UMIs) or nonunique Mis. In some aspects, non-unique Mis are used, but every processed DNA adapter has essentially a unique combination of start and end sites from the original target sequence and MI.
[0015] In one aspect, the disclosure provides a duplex sequencing method for detecting leukemic molecular alterations. The method involves producing IMNA-adapters or NAPT- adapters with specific configurations of Mis, e.g, non-unique Mis or UMIs and stranddistinguishing sequences on the Watson and Crick strands, converting the IMNA-adapters or NAPT-adapters to a library, e.g., by amplifying the IMNA-adapters or NAPT-adapters, optionally converting the IMNA-adapters or NAPT-adapters or amplicons thereof to enriched IMNA-adapters or enriched NAPT-adapters by enriching a plurality of all or a portion of the IMNA-adapters or NAPT-adapters (e.g., by performing one or more targeted amplification steps with at least one target specific primer, or by bait hybridization), and analyzing the
amplicons by NGS to obtain sequence reads. In an aspect, the targeted amplification is achieved using a multiplex PCR approach, using from 10 to 100, 500, 1000, 5000, or 10,000 primer pairs to amplify specific regions of interest. In some aspects the primers may bind to, or flank, specific regions of interest. In another aspect, target enrichment is achieved by hybridization to a panel of oligonucleotide baits (hybrid capture), where said baits are designed to selectively bind to genomic regions relevant to myeloid malignancies, including coding regions, specific exons, introns, or regulatory regions of tens to hundreds of genes. In some aspects, the target specific primer used in a second or successive targeted amplication step is nested (inner) with respect the target specific primer used in a first targeted amplification step.
[0016] In one aspect, the method involves producing a plurality of IMNA adapters or NAPT adapters, each comprising a Watson strand and a Crick strand. In an aspect, the Watson strand of each adapter includes, from 5' to 3', a first Watson single-stranded stranddistinguishing sequence comprising a first universal primer binding site, a first Watson UMI sequence, a Watson strand of a target sequence, a second Watson UMI sequence, and a second Watson single-stranded strand-distinguishing sequence comprising a second, different universal primer binding site. In one aspect, the Crick strand of each adapter includes, from 5' to 3', a first Crick single-stranded strand-distinguishing sequence identical to the first Watson strand-distinguishing sequence, a first Crick UMI sequence complementary to the second Watson UMI sequence, a Crick strand of the target sequence complementary to the Watson strand of the target sequence, a second Crick UMI sequence complementary to the first Watson UMI sequence, and a second Crick single-stranded strand-distinguishing sequence identical to the second Watson strand-distinguishing sequence. In one aspect the methods featured herein further comprise amplifying the Watson and Crick strands of the adapters using universal primers that bind to the strand distinguishing sequences to obtain Watson and Crick amplicons, optionally converting the Watson and Crick amplicons to enriched and/or targeted Watson amplicons and enriched and/or targeted Crick amplicons by, e.g., performing one or more targeted amplification steps, or bait hybridization steps on the Watson and Crick amplicons. In one aspect, the methods featured herein further comprise analyzing the amplicons by nextgeneration sequencing (NGS) to obtain sequence reads, and identifying variants not present in all sequence reads from the Watson and Crick amplicons derived from the same original molecule.
[0017] In some aspects, the duplex sequencing method further includes identifying variants that are not present in all sequence reads from any one target sequence or selected sequence in
Watson amplicons derived from any one Watson IMNA orNAPT and its corresponding Crick amplicons derived from any one Crick IMNA or NAPT.
[0018] Variants not present in all or substantially all (e.g., greater than 80%, 85%, 90%, 95%, 97%, 98%, or 99%) sequence reads from the Watson and Crick amplicons derived from the same original molecule, or, e.g., consensus sequence reads from the Watson and Crick amplicons derived from the same original molecule, are identified as a sequencing process error (from amplification or sequencing) or as a damaged Watson and/or damaged Crick target sequence.'
[0019] The methods of this disclosure can detect various types of leukemic molecular alterations, including single nucleotide variants (SNVs), indels, structural variants, aberrant methylation, and/or copy number variants. These alterations may be present in two or more genomic regions.
[0020] In one aspect, the disclosure relates to methods where the myeloid lineage tumor cells being analyzed are selected from megakaryocytes, platelets, eosinophils, basophils, erythrocytes, monocytes, dendritic cells, macrophages, and/or neutrophils.
[0021] In some aspects, the methods involve analyzing myeloid lineage tumor cells that express AML specific surface markers. In some aspects the AML specific surface markers may include or exclude any of the following markers: CDl lc, CD13, CD14, CD16, CD31, CD33, CD36, CD56, CD64, CD68, CD115, CD116, CD123, CD124, CD135, CD163, CD203c, CD244, CD300a, CD341, CD366, CD371, CD383, CD387 and/or Myeloperoxidase (MPO).
[0022] It is noted that in this disclosure and particularly in the claims and/or paragraphs, terms such as "comprises”, “comprised”, “comprising” and the like can have the meaning attributed to it in U.S. Patent law; e.g., they can mean “includes”, “included”, “including”, and the like; and that terms such as “consisting essentially of’ and “consists essentially of’ have the meaning ascribed to them in U.S. Patent law, e.g., they allow for elements not explicitly recited, but exclude elements that are found in the prior art or that affect a basic or novel characteristic of the present disclosure.
[0023] These and other embodiments are disclosed or are obvious from and encompassed by, the following Detailed Description.
BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The following detailed description, given by way of example, but not intended to limit the invention solely to the specific embodiments described, may best be understood in conjunction with the accompanying drawings.
DETAILED DESCRIPTION
[0025] The inventions described and claimed herein have many attributes and embodiments including, but not limited to, those set forth or described or referenced in this Detailed Description. It is not intended to be all-inclusive and the inventions described and claimed herein are not limited to or by the features or embodiments identified in this Detailed Description, which is included for purposes of illustration only and not restriction.
[0026] A. Certain Exemplary Definitions
[0027] Before the present compounds, compositions, articles, devices, and/or methods are disclosed and described, it is to be understood that they are not limited to specific synthetic methods or specific recombinant biotechnology methods unless otherwise specified, or to particular reagents unless otherwise specified, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
[0028] As used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a pharmaceutical carrier” includes mixtures of two or more such carriers, and the like.
[0029] Ranges can be expressed herein as from “about” one particular value, and/or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and/or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another embodiment. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as “about” that particular value in addition to the value itself. For example, if the value “10” is disclosed, then “about 10” is also disclosed. It is also understood that when a value is disclosed that “less than or equal to” the value, “greater than or equal to the value” and possible ranges between values are also disclosed. For example, if the value “10” is disclosed the “less than or equal to 10” as well as “greater than or equal to 10” is also disclosed. It is also understood that the throughout the application, data is provided in a number of different formats, and that this data, represents endpoints and starting points, and ranges for any combination of the data points. For example, if a particular data point “10” and a particular data point 15 are disclosed, it is understood that greater than, greater than or equal to, less than, less than or equal to, and equal to 10 and 15 are considered disclosed as well as values between 10 and 15. For example, it is also understood
that each unit between two particular units are also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed. In this application, if a data point range is disclosed, it is understood that each unit from the lowest data point to the highest stated datapoint, including the first (lowest) and last (highest) data point is disclosed. For example, if a data point range 1-20 is disclosed, it is understood that data points 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20, and each unit between any two particular units in the range are also disclosed. It is also understood that whenever a series of values are disclosed, that any range falling between any two of the recited values is also understood to be included. [0030] In this specification and in the claims which follow, reference will be made to a number of terms which shall be defined to have the following meanings:
[0031] “Optional” or “optionally” means that the subsequently described event or circumstance may or may not occur, and that the description includes embodiments where said event or circumstance occurs and embodiments where it does not.
[0032] As used herein, the term “separation”, includes any means of substantially purifying one component from another (e.g., by filtration, magnetic attraction, etc.).
[0033] As used herein, the term “isolation” or “isolating”, includes the extraction, purification, and separation of nucleic acids from a biological sample, wherein the nucleic acids are removed from other cellular components, contaminants, or impurities, to obtain a preparation that is suitable for downstream applications such as sequencing, amplification, or analysis.
[0034] The term “subject” refers to any individual who is the target of administration or treatment. The subject can be a vertebrate, for example, a mammal. In one embodiment, the subject can be human, non-human primate, bovine, equine, porcine, canine, or feline. The subject can also be a guinea pig, rat, hamster, rabbit, mouse, or mole. The subject can be a human or veterinary patient. The term subject, in some embodiments, refers to a “patient” under the treatment of a clinician, e.g., physician.
[0035] As used herein, the terms "treat" and "treatment" refer to both therapeutic treatment and prophylactic or preventative measures, wherein the object is to prevent or decrease an undesired physiological change or disorder. For purposes of this disclosure, beneficial or desired clinical results include, but are not limited to, alleviation of symptoms, diminishment of extent of disease, stabilized (i.e., not worsening) state of disease, delay or slowing of disease progression, amelioration or palliation of the disease state, and remission (whether partial or total), whether detectable or undetectable. "Treatment" can also mean prolonging survival as compared to expected survival if not receiving treatment. Those in need of treatment include
those already with the condition or disorder and those prone to have the condition or disorder or those in which the condition or disorder is to be prevented.
[0036] The term “preventing” means preventing in whole or in part, or ameliorating or controlling.
[0037] As used herein, the term "therapeutically effective amount" means an amount of a compound of the present disclosure that (i) treats the particular disease, condition, or disorder, (ii) attenuates, ameliorates, or eliminates one or more symptoms of the particular disease, condition, or disorder, or (iii) prevents or delays the onset of or reduces the intensity of one or more symptoms of the particular disease, condition, or disorder described herein.
[0038] As used herein, the term “substantially the same” means an amount of expression, level, and/or activity of a gene or gene product within 90% of baseline expression, level, and/or activity as in a subject unaffected by a disease. “Substantially the same” can also mean an amount of expression, level, and/or activity of a gene or gene product within 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, to 100% of the baseline expression, level, and/or activity as in a subject unaffected by unaffected by a disease.
[0039] As described herein, any concentration range, percentage range, ratio range or integer range is to be understood to include the value of any integer within the recited range and, when appropriate, fractions thereof (including one-tenth and one-hundredth of an integer), unless otherwise indicated.
[0040] As used herein, the term ‘Limit of Detection’ or ‘LoD,’ as used herein, refers to the lowest quantity or frequency of an analyte or alteration that can be reliably distinguished from its absence. The specific units for LoD depend on the analyte and assay methodology; for example, LoD for molecular variants is commonly expressed as Variant Allele Frequency (VAF), LoD for cell-based assays like flow cytometry may be expressed as the percentage or number of target cells per total cells or unit volume, and LoD for quantitative nucleic acid assays may be expressed as copies per unit of input material (e.g., copies/pg RNA or copies/mL of sample).”
[0041] As used herein, the terms alteration or variant refers to a single nucleotide polymorphism/single nucleotide variants (SNPs/SNVs), and deletions and insertions,. In some embodiments, copy number variations, translocations, inversions, and structural variations are also detected. SNPs/SNVs include synonymous and nonsynonymous mutations. Nonsynonymous mutations include missense mutations, frame-shift mutations, nonsense mutations, and readthrough mutations that can contribute to the development or progression of myeloid malignancies such as acute myeloid leukemia (AML), chronic myeloid leukemia
(CML), myelodysplastic syndromes (MDS), or myeloproliferative neoplasms (MPNs). As used herein, the terms “risk allele”, “risk variant”, and “risk gene variant” include, but are not limited to a gene variant or SNP/SNV that is associated with risk to develop a disease. As used herein, the terms “protective allele”, “protective variant”, and “protective gene variant” include, but are not limited to a gene variant or SNP/SNV that is associated with protection against developing a disease. In some embodiments, a subject having a protective allele or protective variant is a subject that has one protective allele.
[0042] As used herein, the terms “polynucleotide”, “nucleic acid” and “nucleic acid molecule” are used interchangeably herein to refer to a polymeric form of nucleotides of any length, and may comprise ribonucleotides, deoxyribonucleotides, analogs thereof, or mixtures thereof. This term refers only to the primary structure of the molecule. Thus, the term includes triple-, double- and single-stranded deoxyribonucleic acid (“DNA”), as well as triple-, double- and single-stranded ribonucleic acid (“RNA”).
[0043] As used herein, the term "isolated myeloid nucleic acids (IMNA)" refers to nucleic acids that are extracted from enriched myeloid lineage cells and separated from substantially all other components of blood or plasma.
[0044] The term "myeloid lineage" as used herein, in reference to cells, refers to those cells derived from myeloid cell lineage, such as macrophages (and monocytes), granulocytes (including neutrophils, eosinophils, and basophils), erythrocytes, and megakaryocytes/platelets. Myeloid progenitor cells such as myeloblasts and promyelocytes are also included. In an embodiment, these myeloid malignancies encompass acute myeloid leukemia (AML) and its various subtypes (e.g., AML with recurrent genetic abnormalities such as t(8;21)(q22;q22.1); RUNX1-RUNX1T1, inv(16)(pl3.1q22) or t(16;16)(pl3.1;q22); CBFB- MYH11, t(15;17)(q22;ql2); PML-RARA, AML with KMT2A rearrangement, AML with MECOM rearrangement, AML with NUP98 rearrangement, or AML with mutated NPM1, biallelic mutations of CEBPA, mutated RUNX1, mutated ASXL1, mutated TP53, mutated SF3B1, mutated SRSF2, mutated U2AF1, mutated ZRSR2, mutated BCOR, mutated EZH2, or mutated STAG2), myelodysplastic syndromes (MDS) (e.g., MDS with single lineage dysplasia, MDS with multilineage dysplasia, MDS with ring sideroblasts, MDS with excess blasts, MDS with isolated del(5q), MDS with SF3B1 mutation), myeloproliferative neoplasms (MPNs) (e.g., chronic myeloid leukemia (CML) with BCR-ABL1 fusion, polycythemia vera with JAK2 mutations, essential thrombocythemia with JAK2, CALR, or MPL mutations, primary myelofibrosis with JAK2, CALR, or MPL mutations), chronic myelomonocytic leukemia (CMML), and juvenile myelomonocytic leukemia (JMML). In an embodiment, the
methods are also applicable to detecting molecular markers associated with clonal hematopoiesis of indeterminate potential (CHIP) within the enriched myeloid cell fraction, which may indicate a risk for progression to overt myeloid malignancy.
[0045] “AML” refers to an acute myeloid leukemia, with several subtypes, including: myeloid leukemia in cells that produce neutrophils, a white blood cell; acute monocytic leukemia (AML-M5) in cells that produce monocytes, a white blood cell; acute megakaryocytic leukemia (AMLK) in cells that produce red blood cells or platelets, and acute promyelocytic leukemia (APL) in promyelocytes (immature white blood cells).
[0046] The term “nucleic acids processed therefrom (NAPT)” refers to nucleic acids derived from the IMNAs by further processing steps such as fragmentation, end repair (e.g., phosphorylating or dephosphorylatiing, blunting, dA-tailing, or conversion of methyl cytosinse and/or or 5 ’hydroxymethylcytosine by e.g., bisulfite conversion.
[0047] In one embodiment, disclosed herein is a method for detecting one or more myeloid leukemic molecular alterations in nucleic acids from myeloid lineage tumor cells. The method includes the steps of obtaining a sample comprising myeloid lineage cells, performing one or more negative enrichment steps to remove lymphoid lineage cells while retaining the myeloid lineage cells, isolating nucleic acids from the enriched myeloid cells, and detecting leukemic molecular alterations in the isolated nucleic acids or nucleic acids processed therefrom.
[0048] The sample used in the method may be obtained from various sources, including whole blood, peripheral blood or fractions thereof such as PBMCs, or bone marrow aspirate. In some embodiments, the sample is collected from a patient suspected of having a myeloid malignancy, such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), myelodysplastic syndrome (MDS), or myeloproliferative neoplasms (MPNs).
[0049] In some embodiments of any embodiment herein, the methods disclosed herein comprise negative selection enrichment or negative enrichment of myeloid lineage cells, a key step in the method that allows for sensitive detection of myeloid leukemic molecular alterations. By removing the background of normal lymphoid cells, which can dilute out the signal from the myeloid tumor cells, the relative abundance of the myeloid cells is increased, enhancing the ability to detect mutations or other aberrations. In some embodiments, the negative enrichment involves density gradient centrifugation, such as Ficoll-Paque centrifugation, to separate mononuclear cells from granulocytes and red blood cells. In certain embodiments, the mononuclear cell fraction, which includes lymphocytes and monocytes, can then be subjected to further negative selection to remove the lymphocytes.
[0050] In some embodiments, the negative enrichment is performed using a combination of density gradient centrifugation and immunomagnetic cell separation. This two-step procedure can achieve higher purity of the myeloid cell population compared to either method alone. The enriched myeloid cells are then lysed and the nucleic acids are extracted and isolated using standard techniques such as phenol -chloroform extraction, silica-based columns, or magnetic bead-based methods.
[0051] In one embodiment, the lymphocytes are removed using antibodies that specifically bind to cell surface markers expressed on B cells, T cells, and/or NK cells. In some embodiments, these antibodies are conjugated to magnetic beads, allowing the labeled cells to be separated from the unlabeled myeloid cells using a magnetic field. For example, in some embodiments, anti -CD 19 antibodies are used to remove B cells, anti-CD3 antibodies for T cells, and anti-KIR (Killer Immunoglobulin-like Receptor Family) antibodies for NK cells. In an embodiment, the specific markers targeted for depletion may vary depending on the desired purity and recovery of the myeloid cell population. In an embodiment, for robust depletion of T cells, antibodies targeting one or more of CD2, CD3, CD4, CD5, CD7, CD8, or T-cell receptor (TCR) components can be used. In an embodiment, for B cell depletion, antibodies targeting one or more of CD 19, CD20, CD22, CD79a, or CD79b can be employed. In an embodiment, for NK cell depletion, antibodies targeting one or more of CD56, CD 16, or members of the KIR family or NKG2 family receptors can be utilized. In an embodiment, a combination or cocktail of antibodies against multiple markers for each lymphoid lineage is used to maximize depletion efficiency. In another embodiment, the ratio of different antibody- conjugated beads or depletion reagents is optimized based on the typical or expected distribution of lymphoid cell subsets in the specific sample type (e.g., peripheral blood vs. bone marrow).
[0052] In one embodiment, the isolated nucleic acids, which may include genomic DNA, mitochondrial DNA, and/or RNA, are then analyzed for the presence of myeloid leukemic molecular alterations. In one embodiment, the alterations are detected using next-generation sequencing (NGS) technologies. NGS allows for high-throughput, sequencing of multiple genomic regions or even the entire genome or transcriptome. In some embodiments, samples can be multiplexed when using sample indexes or sample barcodes. In an embodiment where RNA is isolated from the enriched myeloid cells, this RNA can include total RNA, messenger RNA (mRNA), or fractions enriched for specific RNA species such as microRNA (miRNA) or long non-coding RNA (IncRNA). Such RNA can be subsequently converted to complementary
DNA (cDNA) for analysis of fusion transcripts, gene expression levels, or RNA-specific variants.
[0053] The isolated nucleic acids can be processed, e.g., fragmented, and/or end repaired (e.g., phosphorylation or dephosphorylation of 5’ or 3’ ends, blunting of 5’ or 3’ overhangs with or without the use of enzyymes such as, e.g., T4 DNA polymerase, removal of 3' overhangs using 3'-5' exonucleases such as, e.g., Exonuclease I, removal of 5' overhangs using 5'-3' exonucleases or endonucleases such as, e.g., lambda exonuclease or mung bean nuclease, dA-tailing, etc). In an embodiment, the isolated nucleic acids (IMNA), particularly DNA, are subjected to fragmentation prior to adapter ligation. In an embodiment, fragmentation is performed using enzymatic methods, for example, utilizing a non-specific endonuclease or a transposase-based system, to generate DNA fragments of a desired size range suitable for downstream sequencing. In other embodiments, mechanical shearing (e.g., sonication, acoustic shearing) or chemical methods are used for fragmentation. In an embodiment, following fragmentation, or if the starting nucleic acids are already suitably fragmented (e.g., NAPT), the nucleic acid ends are repaired and modified to facilitate adapter ligation. In an embodiment, this includes, but is not limited to, 5’ phosphorylation, 3’ adenylation (dA-tailing) to prepare for ligation with adapters having a 3’ thymidine (dT) overhang, blunting of ends, and/or repair of nicks or gaps. In an embodiment, these steps can be performed using a combination of enzymes such as polymerases (e.g., T4 DNA polymerase, Klenow fragment), kinases (e.g., T4 polynucleotide kinase), and ligases. In one embodinent, the fragments may undergo dephosphorylation and blunt ending prior to ligation with a 3’ adapter, and subsequent hybridization with a 5’ adapter, following by extension and ligation of the 5’ adapter. In an embodiment, after initial end repair and before or after adapter ligation, the nucleic acid library undergoes a conditioning step to remove or repair damaged DNA molecules. In an aspect, this can involve enzymatic treatments to excise damaged bases (e.g., uracil DNA glycosylase for uracil removal, FPG for oxidized purines) or to repair other forms of DNA damage such as abasic sites or single-strand breaks, thereby improving the quality and accuracy of the sequencing library. Subsequent to these processing steps, adapters are ligated to the ends of the processed nucleic acid fragments.
[0054] In some embodiments, the processed nucleic acids are converted to IMNA-adapters or NAPT- adapters with specific configurations of Mis, e.g, non-unique Mis or UMIs and strand-distinguishing sequences on the Watson and Crick strands, converting the IMNA- adapters or NAPT-adapters to a library, e.g., by amplifying the IMNA-adapters or NAPT- adapters, optionally converting the IMNA-adapters or NAPT-adapters or amplicons thereof to
enriched IMNA-adapters or enriched NAPT-adapters by enriching a plurality of all or a portion of the IMNA-adapters or NAPT-adapters (e.g., by performing one or more targeted amplification steps with at least one target specific primer, or by bait hybridization), and analyzing the amplicons by NGS to obtain sequence reads. In an embodiment, the targeted amplification is achieved using a multiplex PCR approach, potentially involving thousands of primer pairs to amplify specific regions of interest. In another embodiment, target enrichment is achieved by hybridization to a panel of oligonucleotide baits (hybrid capture), where said baits are designed to selectively bind to genomic regions relevant to myeloid malignancies, including coding regions, specific exons, introns, or regulatory regions of tens to hundreds of genes. In some embodiments, the target specific primer used in a second or successive targeted amplication step is nested (inner) with respect the target specific primer used in a first targeted amplification step.
[0055] In an embodiment, following adapter ligation to the IMNA or NAPT fragments and prior to target enrichment by hybrid capture, a limited number of PCR amplification cycles (pre-capture amplification) are performed using primers that bind to universal sequences within the ligated adapters. This step serves to increase the quantity of adapter-ligated library material available for subsequent hybridization to baits. In an embodiment, the conversion to IMNA- adapters or NAPT-adapters involves the ligation of Y-shaped adapters, which facilitate subsequent amplification and sequencing steps and can incorporate features such as molecular identifiers (Mis) and primer binding sites. In an embodiment, these Mis can be unique molecular identifiers (UMIs) or non-unique Mis where the combination of the MI sequence and the start and end mapping coordinates of the nucleic acid fragment provides a unique signature for each original molecule, a strategy employed in error-correction methodologies. Such adapter structures and MI strategies allow for high-sensitivity sequencing, including, for example, duplex sequencing.
[0056] In one embodiment, the method involves producing a plurality of IMNA adapters or NAPT adapters, each comprising a Watson strand and a Crick strand. In an embodiment, the Watson strand of each adapter includes, from 5' to 3', a first Watson single-stranded stranddistinguishing sequence comprising a first universal primer binding site, a first Watson UMI sequence, a Watson strand of a target sequence, a second Watson UMI sequence, and a second Watson single-stranded strand-distinguishing sequence comprising a second, different universal primer binding site. In one embodiment, the Crick strand of each adapter includes, from 5' to 3', a first Crick single-stranded strand-distinguishing sequence identical to the first Watson strand-distinguishing sequence, a first Crick UMI sequence complementary to the
second Watson UMI sequence, a Crick strand of the target sequence complementary to the Watson strand of the target sequence, a second Crick UMI sequence complementary to the first Watson UMI sequence, and a second Crick single-stranded strand-distinguishing sequence identical to the second Watson strand-distinguishing sequence. In one embodiment the methods featured herein further comprise amplifying the Watson and Crick strands of the adapters using universal primers that bind to the strand distinguishing sequences to obtain Watson and Crick amplicons, optionally converting the Watson and Crick amplicons to enriched and/or targeted Watson amplicons and enriched and/or targeted Crick amplicons by, e.g., performing one or more targeted amplification steps, or bait hybridization steps on the Watson and Crick amplicons. In one embodiment, the methods featured herein further comprise analyzing the amplicons by next-generation sequencing (NGS) to obtain sequence reads, and identifying variants not present in all sequence reads from the Watson and Crick amplicons derived from the same original molecule.
[0057] In some embodiments, the sequencing reads can be aligned to a reference genome and analyzed for variants such as single nucleotide variants (SNVs), insertions/deletions (indels). In some embodiments copy number variations (CNVs), and structural rearrangements are also analyzed.
[0058] In an embodiment, bioinformatic analysis of the sequencing data includes steps such as read alignment to a reference genome, quality control filtering, variant calling for SNVs and indels using algorithms optimized for low-frequency detection, and annotation of identified variants. In an embodiment, for duplex sequencing data, specialized bioinformatic tools are employed to process UMIs, build consensus sequences from single strands and then duplex consensus sequences, and perform error correction to distinguish true low-level somatic mutations from artifacts. In an embodiment, the analysis pipeline also includes algorithms for detecting structural variants, including gene fusions, from DNA-seq or RNA-seq data, and for inferring copy number variations (CNVs) from targeted panel sequencing depth of coverage or from whole exome/genome data. In an embodiment, machine learning algorithms are utilized for refining variant calls, predicting pathogenicity, or for integrating multiple data types to derive prognostic or predictive signatures.
[0059] In one embodiment, the present disclosure provides methods for detecting leukemic molecular alterations using a variety of molecular assays. These assays may include, but are not limited to, Next-Generation Sequencing (NGS), multiplex PCR, Droplet Digital PCR (ddPCR), Quantitative PCR (qPCR), hybrid capture assays, and molecular inversion probe assays. Each of these assays has unique advantages and can be selected based on factors such
as the type of alteration being detected, the desired sensitivity and specificity, the throughput required, and the available resources.
[0060] In some embodiments, an MI, e.g., UMI or non-unique MI, strand-distinguishing sequence, or primer region can be the template for amplification primers utilized in a second or subsequent round of amplification, for example for library preparation. An aliquot of the first amplified sample can be removed and amplified a second time using a second set of amplification primers that are specific to the MI, e.g., UMI or non-unique MI, stranddistinguishing sequence, or primer region, e.g., a MI, e.g., UMI or non-unique MI or an sequencing primer region, of the first amplification primers which may comprise of one or more additional MI, strand- distinguishing sequence, or primer sequences, such as sequence MI, strand-distinguishing sequence, or primer specific for one or more downstream sequencing workflows, and the same or a second PCR master mix.
[0061] Bait hybridization can occur before or after adding adaptors comprising barcodes, UMIs, universal amplification primers, strand-distinguishing sequences, or sequencing primers, or the like. As such, a library of the original DNA sample can be ready for sequencing. [0062] NGS assays are particularly useful for detecting a wide range of alterations across multiple genomic regions simultaneously. NGS technologies allow for high-throughput sequencing of DNA or RNA fragments, generating millions of reads that can be aligned to a reference genome and analyzed for variants. In some embodiments, the NGS assay is performed on nucleic acids converted to enriched and/or targeted Watson amplicons and enriched and/or targeted Crick amplicons prepared from IMNA or NAPT. In some embodiments the processing steps include fragmentation, and end repair. In some embodimnets the preparing steps include converting IMNA or NAPT to IMNA-adapters or NAPT adapters by one or more ligation steps using one or more adapters, or one or more amplicfication steps using one or more adapter primers or one or more adapter primer pairs, one or more library amplifications using universal primers, and/or conversion to enriched and/or targeted Watson amplicons and enriched and/or targeted Crick amplicons using one or more targeted amplification steps or one or more bait enrichment steps
[0063] In one embodiment, the NGS assay includes an error correction method to improve the accuracy of variant calling. One such method is duplex sequencing, which involves converting DNA molecules (e.g., IMNA or NAPT) into DNA-adapters (e.g., IMNA-adapters or NAPT-adapters) using one or more adapters comprising a strand distinguishing sequence and/or a UMI before enrichment and sequencing. The Mis, e.g, non-unique Mis, e.g., UMIs or non-unique Mis and strand distinguishing sequences allow for the identification and consensus
calling of variants on Watson strand reads and Crick strand reads, and the identification of errors introduced during amplification or sequencing that are not present in all or substantially all sequence reads or consensus sequence reads from both strands of the original DNA duplex. By comparing the UMI sequences across multiple reads originating from each strand of the same DNA fragment, true mutations can be distinguished from artifacts, enabling the detection of variants present at very low allele frequencies. In an embodiment, the specific implementation of Mis, whether as statistically unique sequences of a certain length or as shorter, non-unique sequences that derive uniqueness from their combination with other sequence features like fragment start/end points, can be adapted based on the requirements of the sequencing platform and the desired error-correction fidelity. In an embodiment, such error correction utilizing UMIs and analysis of both DNA strands is a feature of Duplex Sequencing methodologies, which are designed to achieve high accuracy in calling low-frequency variants by distinguishing true mutations from PCR-induced or sequencing errors. This allows for a limit of detection (LoD) for variants that can extend to 10'5, 10'6, or even lower, depending on sequencing depth and initial DNA input from the enriched myeloid cells. Duplex sequencing can substantially reduce the error rate and enable the detection of low-frequency variants with high confidence.
[0064] In some embodiments, the NGS assay is a targeted assay that focuses on specific genomic regions of interest. In one embodiment, the NGS assay is designed to target specific genes or genomic regions known to be frequently mutated in myeloid malignancies. In some embodiments, the NGS assay is designed to target one or more myeloid informed genes or genomic regions specific for a subject. Either of these targeted approaches can increase the sensitivity and specificity of mutation detection by focusing the sequencing depth on the regions of interest. Targeted NGS can be achieved by various methods, such as amplicon sequencing, hybrid capture, or molecular inversion probes. In an embodiment, hybrid capture methodologies, sometimes referred to as bait-based enrichment, utilize a library of oligonucleotide probes (baits) designed to hybridize to specific genomic regions of interest within the IMNA or NAPT adapter-ligated library. In some embodiments the baits are biotinylated. Following hybridization, the bait-target hybrids are captured, for example, using streptavidin-coated magnetic beads, washed to remove non-target sequences, and then eluted for sequencing. This approach allows for the enrichment of potentially large and numerous target regions. In some embodiments baits hybridize to regions comprising variants and/or mutations of interest, and can bind under suitable hybridization conditions to both wild-type and variant/mutant sequences. In some emodiments the baits are allele, SNP, and/or mutation
specific. In another embodiment, targeted amplicon sequencing involves PCR amplification using primers specific to the desired genomic regions. Both hybrid capture and amplicon-based approaches are utilized in various commercially available targeted sequencing assays and can be effectively coupled with the negative enrichment methods described herein to enhance the detection of low-frequency variants from the myeloid cell fraction. In an embodiment where hybrid capture is used, the bait panel can be designed to cover a focused set of key myeloid genes (e.g., 20-50 genes) or a more comprehensive panel (e.g., 50-500 genes or more), including genes such as, but not limited to, NPM1, IDH1, IDH2, FLT3, KIT, TP53, RUNX1, ASXL1, DNMT3A, TET2, SF3B1, SRSF2, U2AF1, JAK2, CALR, MPL, CEBPA, BCOR, EZH2, GATA2, PHF6, RAD21, SETBP1, SH2B3, STAG2, and WT1. Such panels can target SNVs, indels, and regions informative for CNV detection. In an embodiment, the design of such bait panels also considers intronic regions flanking exons to detect splice site mutations. The specific panel of genes may be selected based on the type of myeloid malignancy or condition, the clinical context, or the desired level of comprehensive profiling. In one embodiment the NGS assay is a SaferSeq assay.
[0065] In some embodiments, the NGS assay is an Ion Torrent AmpliSeq sequencing assay. For example, in one embodiment the NGS assay includes highly multiplexed targeted amplification (e.g., up to 24,000 amplicons), sample indexing and sequencing on an Ion Torrent sequencer. In one embodiment the NGS assay includes digesting primers following multiplex amplification. The amplicons may be selected based on population myeloid cancer markers or may be informed by a subject’s myeloid markers vs. normal tissue markers.
[0066] The terms "proximal" and "distal" in reference to the location of Mis or other sequences relative to a target sequence. In some embodiments, sequences that are "proximal" or "distal" to each other are a distance of 0-50 nucleotides and greater than 50 nucleotides, respectively. In one embodiment, the targeted NGS assay is a targeted duplex sequencing assay, where the processed nucleic acids are partially double-stranded and include a double-stranded target sequence, one or more double-stranded Mis located proximal to the ends of the target sequence, and one or more strand-distinguishing sequences located distal to the Mis on each strand. This configuration allows for the independent tracking and error correction of the Watson and Crick strands of each original DNA molecule.
[0067] The duplex sequencing method may involve several steps, including: (i) producing IMNA or NAPT adapters with specific configurations of Mis and strand-distinguishing sequences on the Watson and Crick strands, (ii) amplifying the IMNA-adapters or NAPT- adapters to obtain Watson and Crick amplicons, (iii) optionally performing targeted
amplification or bait hybridization on the amplicons, and (iv) analyzing the amplicons by NGS to obtain sequence reads. In some embodiments, an additional step of (v) identifying variants that are not present in all sequence reads from the Watson and Crick amplicons derived from the same original molecule is performed. This step helps to filter out errors and identify true variants based on their presence on both strands of the original DNA duplex.
[0068] Following library production, the library can be optionally purified and quantitated. In some embodiments, purification can be performed by processing the sample through a substrate such as SPRI beads, such as, e.g., AMPURE XP Beads (Beckman Coulter), which serves to purify the DNA fragments away from reaction components. In some methods, the purification can be performed by incorporating a biotin into one of the primers of the second amplification primer set then the library fragments could be capturing using a streptavidin moiety on a bead for example. In some embodiemnts, utilizing the capture strategy the libraries could also be normalized and quantitated using bead-based normalization. In other embodiments, libraries can be purified and quantitated, or pooled and quantitated if multiple reactions are being performed, without the use of bead-based normalization. In an embodiment, libraries are quantitated by gel electrophoretic methods, microfluidics-based automated electrophoresis methods, e.g., using the Agilent Bioanalyzer, qPCR, spectrophotometric methods, quantitation kits (e.g., PicoGreen™, Nanodrop, etc.) and the like as known in the art. In an embodiment, following quantitation, the library can then be sequenced.
[0069] In an embodiment, other molecular assays are used to detect leukemic molecular alterations in addition to NGS. Multiplex PCR allows for the simultaneous amplification of multiple target regions in a single reaction. In some embodiments where primers are designed to bind to regions of interest, e.g., those containing known hotspot mutations, multiplex PCR provides a rapid and cost-effective way to screen for alterations. ddPCR and qPCR are useful for quantifying specific alterations, such as gene fusions or copy number variations. These assays rely on the use of fluorescent probes or intercalating dyes to measure the amount of DNA amplification in real-time. Hybrid capture assays involve the use of biotinylated probes to selectively enrich for target regions before sequencing. This can improve the sensitivity and specificity of the assay by increasing the proportion of sequencing reads that map to the regions of interest.
[0070] In an embodiment, the leukemic molecular alterations detected by the duplex sequencing methods of this disclosure include a wide range of genetic and epigenetic changes relevant to myeloid malignancies. In one embodiment, single nucleotide variants (SNVs) are single base pair changes that occur in coding or non-coding regions of the genome. In an
embodiment, SNVs in genes frequently mutated in AML, MDS, or MPNs, such as, e.g., SNVs in any one or more of AKT1, AKT2, AKT3, ALK, AR, ARAF, AXL, BRAF, BTK, CBL, CCND1, CDK4, CDK6, CHEK2, CSF1R, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, ERBB4, ERCC2, ESRI, EZH2, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, FOXL2, GATA2, GNA11, GNAQ, GNAS, H3F3A, HIST1H3B, HNF1A, HRAS, IDH1, IDH2, JAK1, JAK2, JAK3, KDR, KIT, KNSTRN, KRAS, MAGOH, MAP2K1, MAP2K2, MAP2K4, MAPK1, MAX, MDM4, MED12, MET, MTOR, MYC, MYCN, MYD88, NFE2L2, NRAS, NTRK1, NTRK2, NTRK3, PDGFRA, PDGFRB, PIK3CB, PIK3CA, PPP2R1A, PTPN11, RAC1, RAFI, RET, RHEB, RHOA, ROS1, SF3B1, SMAD4, SMO, SPOP, SRC, STAT3, TERT, TOPI, U2AF1, XPO1, ARID1A, ATM, ATR, ATRX, BAP1, BRCA1, BRCA2, CDK12, CDKN1B, CDKN2A, CDKN2B, CHEK1, CREBBP, FANCA, FANCD2, FANCI, FBXW7, MLH1, MRE11, MSH6, MSH2, NBN, NF1, NF2, NOTCH1, NOTCH2, NOTCH3, PALB2, PIK3R1, PMS2, POLE, PTCHI, PTEN, RAD50, RAD51, RAD51B, RAD51C, RAD51D, RNF43, RBI, SETD2, SLX4, SMARCA4, SMARCB1, STK11, TP53, TSC1, TSC2, CCND2, CCND3, CCNE1, CDK2, FGF19, FGF3, IGF1R, MDM2, MYCL, PPARG, and RICTOR. In one embodiment, the SNVs may include or exclude SNVs in one or a combination of the foregoing genes. In an embodiment, SNVs in genes frequently mutated in AML, MDS, or MPNs, such as, e.g., SNVs in any one or more of FLT3, NPM1, CEBPA, IDH1/2, or TET2, are detected with high accuracy. Indels are small insertions or deletions of one or more base pairs. In an embodiment, structural variants, which include larger-scale changes such as translocations, inversions, or copy number variations, can be detected based on the presence of Watson and Crick reads spanning the breakpoints. Aberrant methylation refers to changes in the DNA methylation patterns that affect gene expression and contribute to leukemogenesis. Copy number variants are changes in the number of copies of a particular gene or genomic region. In an embodiment, these alterations are driver events that contribute to the initiation and progression of myeloid malignancies, or they are passenger events that do not directly confer a selective advantage but may still have diagnostic or prognostic value.
[0071] In an embodiment, the panel of genes interrogated for SNVs and indels can be comprehensive, including, but not limited to, genes involved in signal transduction pathways: (e g., FLT3, KIT, JAK2, MPL, NRAS, KRAS, HRAS, PTPN11, CBL, NF1, SH2B3, CSF3R, MET), transcription factors and regulators: (e.g., NPM1, CEBPA, RUNX1, GATA1, GATA2, ETV6, IKZF1, WT1, CUX1, PHF6, ERG, MYC, TP53, ETV6, STAT3, SOX4, TALI, LYL1), epigenetic modifiers: including those involved in DNA methylation (e.g., DNMT1, DNMT3A, DNMT3B, TET1, TET2, TET3, IDH1, IDH2), histone modification (e.g, ASXL1, ASXL2,
EZH1, EZH2, SUZ12, EED, KMT2A (MLL), KMT2C, KMT2D, KDM5A, KDM5C, KDM6A (UTX), KDM6B, SETD2, SETBP1, NSD1, NSD2, NSD3, BCOR, BCORL1, CREBBP, EP300, HDACs, HATs), and chromatin remodeling (e.g., SMARCA2, SMARCA4, SMARCB1, SMARCD1, SMARCD2, ARID1A, ARID1B, ARID2, ARID5B), RNA splicing machinery: (e.g., SF1, SF3A1, SF3B1, SRSF2, U2AF1, U2AF2, PRPF8, PRPF40B, ZRSR2, LUC7L2, DDX41), cohesin complex: (e.g., STAG1, STAG2, RAD21, SMC1A, SMC3, NIPBL, PDS5A, PDS5B), ribosomal proteins and biogenesis: (e.g., RPL5, RPL10, RPL11, RPL22, RPS7, RPS10, RPS14, RPS15, RPS17, RPS19, RPS20, RPS24, RPS26, RPS27, RPS28, RPS29, SBDS), DNA damage response and repair pathways: (e.g., ATM, ATR, BRCA1, BRCA2, FANCA, FANCC, FANCG, MLH1, MSH2, MSH6, PMS2, POLE, PALB2, CHEK1, CHEK2, RAD50, RAD51), cell cycle regulators: (e.g., CDKN1A, CDKN1B, CDKN2A, CDKN2B, CCND1, CCND2, CCND3, CCNE1, CDK2, CDK4, CDK6, CDK12, RBI), other genes implicated in myeloid development, apoptosis, or leukemogenesis: (e.g., CALR, MYD88, PPM1D, CUX1, N0TCH1, N0TCH2, N0TCH3, FOXL2, MAGOH, MDM2, MDM4, MED12, PPARG, RAC1, RAFI, RHEB, RHOA, ROS1, SMAD4, SMO, SPOP, SRC, TERT (including promoter mutations), TOPI, XPO1, BCL2, MCL1). In some embodiments, a targeted panel comprises a selection of these genes known to be recurrently mutated in myeloid malignancies. In some embodiments, for instance, a panel can include, or be similar to, genes targeted by sensitive assays such as the DuplexSeq AML MRD Assay, which includes but is not limited to: ASXL1, BCOR, BCORL1, CALR, CBL, CEBPA, CSF3R, DDX41, DNMT3A, ETV6, EZH2, FLT3, GATA2, GNAS, HRAS, IDH1, IDH2, IKZF1, JAK2, KIT, KRAS, KMT2A (MLL), MPL, NF1, NPM1, NRAS, PHF6, PPM1D, PTPN11, RAD21, RUNX1, SETBP1, SF3B1, SRSF2, STAG2, STAT3, TET2, TP53, U2AF1, WT1, and ZRSR2. The selection of genes for a particular panel can be tailored based on the specific application, desired sensitivity for particular mutations, and current clinical or research guidelines, such as the European LeukemiaNet (ELN) recommendations.
[0072] In an embodiment, the methods described herein are also capable of detecting alterations in non-coding regions of the genome that have functional consequences in myeloid malignancies. These include, but are not limited to, mutations in promoter regions (e.g., TERT promoter mutations), enhancer regions, silencer regions, insulator elements, 5’ and 3’ untranslated regions (UTRs), or regions encoding non-coding RNAs. In an embodiment, analysis of non-coding RNAs (ncRNAs) isolated from the enriched myeloid cell fraction is performed. This includes, but is not limited to, the quantification of expression levels or sequence analysis of microRNAs (miRNAs, e.g., miR-155 family), long non-coding RNAs
(IncRNAs, e.g., HOTAIR, specific IncRNA signatures associated with AML subtypes or risk), circular RNAs (circRNAs), small nucleolar RNAs (snoRNAs), and PlWI-interacting RNAs (piRNAs), which are known to be dysregulated in various myeloid malignancies and can serve as diagnostic, prognostic, or predictive biomarkers. In another embodiment, epigenetic alterations beyond DNA methylation patterns are assessed. This can encompass the analysis of histone modifications (e.g., specific acetylation or methylation marks on histones such as H3K4me3) by adapting the methods for chromatin immunoprecipitation followed by sequencing (ChlP-seq) on the enriched myeloid cell population, or by analyzing expression levels of histone modifying enzymes. Furthermore, alterations in chromatin accessibility, for example, using Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) on the enriched myeloid cells, can provide insights into the regulatory landscape of the malignant cells.
[0073] The number of genomic regions harboring leukemic alterations can vary depending on the type and stage of the myeloid malignancy. In some embodiments, the leukemic molecular alterations are present in two or more genomic regions. This can reflect the presence of multiple driver events or the accumulation of passenger mutations over time. In one embodiment, the alterations are present in 2 genomic regions. In one embodiment, the alterations are present in 2-60 genomic regions frequently impacted in myeloid malignanices, such as, e.g., any one or more of alterations selected from alterations in FLT3, NPM1, CEBPA, CBFB, IDH1, IDH2, TET2, ASXL1, RUNX1, RUNX1T1, TP53, MYH11, NRAS, KRAS, KIT, CALR, WT1, JAK2, and MPL. In one embodiment, the alteraction can be one or more mutations or alterations selected from alterations in AKT1, AKT2, AKT3, ALK, AR, ARAF, AXL, BRAF, BTK, CBL, CCND1, CDK4, CDK6, CHEK2, CSF1R, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, ERBB4, ERCC2, ESRI, EZH2, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, FOXL2, GATA2, GNA11, GNAQ, GNAS, H3F3A, HIST1H3B, HNF1A, HRAS, IDH1, IDH2, JAK1, JAK2, JAK3, KDR, KIT, KNSTRN, KRAS, MAGOH, MAP2K1, MAP2K2, MAP2K4, MAPK1, MAX, MDM4, MED12, MET, MTOR, MYC, MYCN, MYD88, NFE2L2, NRAS, NTRK1, NTRK2, NTRK3, PDGFRA, PDGFRB, PIK3CB, PIK3CA, PPP2R1A, PTPN11, RAC1, RAFI, RET, RHEB, RHOA, ROS1, SF3B1, SMAD4, SMO, SPOP, SRC, STAT3, TERT, TOPI, U2AF1, XPO1, ARID1A, ATM, ATR, ATRX, BAP1, BRCA1, BRCA2, CDK12, CDKN1B, CDKN2A, CDKN2B, CHEK1, CREBBP, FANCA, FANCD2, FANCI, FBXW7, MLH1, MRE11, MSH6, MSH2, NBN, NF1, NF2, NOTCH1, NOTCH2, NOTCH3, PALB2, PIK3R1, PMS2, POLE, PTCHI, PTEN, RAD50, RAD51, RAD51B, RAD51C, RAD51D, RNF43, RBI, SETD2, SLX4, SMARCA4, SMARCB1, STK11, TP53,
TSC1, TSC2, CCND2, CCND3, CCNE1, CDK2, FGF19, FGF3, IGF1R, MDM2, MYCL, PPARG, and RICTOR. In one embodiment, the alterations include or exclude one or more of the foregoing or combinations thereof. In one embodiment, the alteraction can be one or more mutations, alterations, or gene fusions selected from a gene fusion of one or more of AKT2, ALK, AR, AXL, BRCA1, BRCA2, BRAF, CDKN2A, EGFR, ERBB2, ERBB4, ERG, ESRI, ETV1, ETV4, ETV5, FGFR1, FGFR2, FGFR3, FGR, FLT3, JAK2, KRAS, MDM4, MET, MYB, MYBL1, NF1, N0TCH1, N0TCH4, NRG1, NTRK1, NTRK2, NTRK3, NUTM1, PDGFRA, PDGFRB, PIK3CA, PRKACA, PRKACB, PTEN, PPARG, RAD51B, RAFI, RBI, RELA, RET, ROS1, RSPO2, RSPO3, and TERT. In one embodiment, the alterations include or exclude one or more of the foregoing or combinations thereof.
[0074] The methods of the present disclosure can be applied to various types of myeloid lineage tumor cells, including megakaryocytes, platelets, eosinophils, basophils, erythrocytes, monocytes, dendritic cells, macrophages, and/or neutrophils. In some embodiments, these cell types represent different stages and lineages of myeloid differentiation and give rise to distinct subtypes of myeloid malignancies. In some embodiments, the myeloid lineage tumor cells express specific surface markers that are used for their identification and isolation. These markers may include CDl lc, CD13, CD14, CD31, CD33, CD36, CD64, CD68, CD115, CD116, CD123, CD124, CD135, CD163, CD203c, CD244, CD300a, CD341, CD366, CD371, CD383, CD387 and/or Myeloperoxidase (MPO). Some of these markers are specific to certain myeloid lineages or differentiation stages, while others are more broadly expressed.
[0075] In one embodiment, the methods of the present disclosure are applied to myeloid myeloid lineage cells. In one embodiment, the methods of the present disclosure are applied to acute myeloid leukemia (AML) cells. AML is a heterogeneous malignancy characterized by the clonal expansion of immature myeloid cells in the bone marrow and peripheral blood. In one embodiment, AML cells express various combinations of surface markers, including CD13, CD33, CD34, CD117, and MPO. The genetic alterations in AML are diverse. In some embodiments, the methods of the present disclosure are used to comprehensively profile the genetic and epigenetic landscape of AML cells and to monitor the response to therapy and the emergence of resistant clones.
[0076] In another embodiment, the methods of the present disclosure are applied to myelodysplastic syndromes (MDS). MDS are a group of clonal hematopoietic disorders characterized by ineffective hematopoiesis and an increased risk of progression to AML. In one embodiment, MDS cells show dysplastic morphology and increased apoptosis in one or more myeloid lineages. The genetic alterations in MDS often involve mutations in genes
related to RNA splicing (e.g. SF3B1, SRSF2, U2AF1), epigenetic regulation (e.g. TET2, ASXL1, DNMT3A), and transcriptional regulation (e.g. RUNX1, TP53). In an embodiment, the methods of the present disclosure are used to identify the specific genetic and epigenetic alterations in MDS cells and to guide risk stratification and treatment decisions.
[0077] In yet another embodiment, the methods of the present disclosure are applied to chronic myeloid leukemia (CML) cells. CML is a myeloproliferative neoplasm characterized by the BCR-ABL1 fusion gene, which results from a translocation between chromosomes 9 and 22. The BCR-ABL1 fusion protein has constitutive tyrosine kinase activity and drives the proliferation and survival of myeloid cells. In one embodiment, the methods of the present disclosure are used to monitor the response to tyrosine kinase inhibitor therapy and to detect the emergence of resistance mutations in the BCR-ABL1 kinase domain.
[0078] In some embodiments, the methods of the present disclosure are used to detect minimal residual disease (MRD) in myeloid malignancies. MRD refers to the presence of residual leukemic cells below the threshold of morphologic detection. In one embodiment, MRD is an important prognostic factor and guides decisions about the intensity and duration of therapy. The high sensitivity of the molecular assays used in the present disclosure, such as NGS and ddPCR, enables the detection of MRD at levels as low as 0.01% or lower. In one embodiment, MRD allows for early intervention and personalized management of patients based on their MRD status. In one embodiment, the initial amount of nucleic acid (IMNA or NAPT) input into the molecular assay following enrichment can be varied. In an embodiment, for example, DNA input amounts may have a range that is between any two values from about 0.1 ng to about 2000 ng, or more. In an embodiment, the initial amount of nucleic acid (IMNA or NAPT) input into the molecular assay following enrichment can be varied. In an aspect, DNA input amounts may have a range that is between any two values of about 0.1 ng, about 1 ng, about 5 ng, about 10 ng, about 15 ng, about 20 ng, about 25 ng, about 30 ng, about 40 ng, about 50 ng, about 75 ng, about 100 ng, about 125 ng, about 150 ng, about 175 ng, about 200 ng, about 250 ng, about 300 ng, about 400 ng, about 500 ng, about 750 ng, about 1000 ng (1 pg), about 1500 ng (1.5 pg), about 2000 ng (2 pg), or more. In some embodiments, the DNA input amount is, is about, or is at least 0.1 ng, 1 ng, 10 ng, 25 ng, 50 ng, 100 ng, or 150 ng. In some embodiments, the DNA input amount is, is about, or is not more than 50 ng, 100 ng, 250 ng, 500 ng, 1000 ng, or 2000 ng. In an embodiment, the DNA input amount is within a range of about 10 ng to about 50 ng. In an embodiment, the DNA input amount is within a range of about 50 ng to about 250 ng. In an embodiment, the DNA input amount is within a range of about 200 ng to about 2000 ng. In an embodiment, the selection of a specific input amount can
depend on factors such as the type of nucleic acid, the expected abundance of target sequences, the downstream molecular assay to be performed, and the desired level of sensitivity.
[0079] The size of the nucleic acid molecules may vary greatly using the methods and compositions disclosed herein. It would be appreciated by those skilled in the art that the nucleic acid molecules amplified from a target sequence comprising a tandem repeat (e.g., STR) may have a large size, while the nucleic acid molecules amplified from a target sequence comprising a SNP may have a small size. In some embodiments, the nucleic acid molecules may comprise from less than a hundred nucleotides to hundreds or even thousands of nucleotides. In some embodiments, the size of the nucleic acid molecules may have a range that is between 3- bp to 1 kb. The size of the nucleic acid molecules may have a range that is between any two values of about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1 kb, or more. In some embodiments, the minimal size of the nucleic acid molecules may be a length that is, is about, or is less than, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, or 100 bp. In some embodiments, the maximum size of the nucleic acid molecules may be a length that is, is about, or is more than, 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, or 1 kb.
[0080] In some embodiments, the method involves ligating adapters to the ends of doublestranded IMNA or NAPT fragments to generate adapter-attached, double-stranded DNA fragments. In some embodiments, the adapters comprise a double-stranded portion comprising a UMI and a single-stranded forked portion with a 3' sequence and a non-complementary 5' sequence to create Y-shaped ends. In some embodiments, the 5' sequence can include a binding site for a first sequencing primer (Rl), while the 3' sequence can include a binding site for a second sequencing primer (R2). In some embodiments each Y-shaped adapter comprises: a 3’ single-stranded arm, a 5’ single-stranded arm, an MI sequence (e.g., UMI or non-unique MI) that alone or in combination with a start or end site, or sequence from the target sequence uniquely labels a ligated IMNA-adapter or NAPT-adapter such that each individual ligated IMNA-adapter or NAPT-adapter is distinguishable from other IMNA-adapters or NAPT- adapters; and a strand-distinguishing sequence that, following the attachment of the adapter to the IMNA or NAPT, provides a region of non-complementarity between a first strand of an individual IMNA-adapter or NAPT-adapter and a second strand of the same IMNA-adapter or NAPT-adapter.
[0081] The R1 and R2 primer binding sites are used in sequencing systems that generate paired-end reads from opposite ends of each fragment. The R1 primer is used to produce "Rl" or "Read 1" sequences from one end, while the R2 primer is used to produce "R2" or "Read 2" sequences from the other end. The Rl and R2 reads for each fragment can be aligned as read pairs and grouped by their unique Mis for error correction.
[0082] In an embodiment, the Rl and R2 primer sites allow for independent amplification and sequencing of the Watson and Crick strands of each fragment. In one embodiment, the designation of Rl vs R2, or the particular adapter sequences used, are not critical as long as the strand information is retained, such as, e.g., in duplex sequencing. The Rl and R2 nomenclature is used generically herein to refer to any pair of opposing primers and paired-end reads and does not necessarily imply a specific orientation or adapter sequence.
[0083] In some embodiments, the MI is a UMI unique to each double-stranded DNA fragment in the nucleic acid sample. In some embodiments, the UMI is not unique to each double-stranded DNA fragment. In some embodiments, the MI is essentially unique, in that the number of Mi’s exceeds the expected number of duplicate fragments (fragments having the same start and stop site) by 100 fold, 250 fold, 500 fold, 750 fold, 1000 fold, 2000 fold, 5000 fold, 10000 fold, 25,000 fold, 50,000 fold, 100,000 fold, 250,000 fold, 500,000 fold, 1,000,000 fold, 5,000,000 fold, 10,000,000 fold, 50,000,000 fold, 100,000,000 fold or more,, or any range between any two of the recited numbers. As an example, if the number of expected duplicate fragments (e.g., distinct original molecules that happen to yield the same start and end sites after fragmentation) in a tested sample is estimated to be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or in a range such as about 1 to about 2, about 1 to about 3, about 1 to about 4, about 1 to about 5, about 1 to about 6, about 1 to about 7, about 1 to about 8, about 1 to about 9, about 1 to about 10, about 2 to about 5, about 5 to about 10, about 1 to about 15, about 1 to about 20, or any range between any two of these recited numbers, then the the diversity of the MI pool (total number of different MI sequences available) is selected to be very large to ensure a high probability of assigning a unique MI to each distinct molecule. In an embodiment, in such cases, the total number of available Mis may be, for example, any number between 100, 250, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 100,000, 250,000, 500,000, 1,000,000, 5,000,000, 10,000,000, 25,000,000, 50,000,000, 100,000,000, 150,000,000, 200,000,000, 250,000,000, about 268,000,000, 300,000,000, 500,000,000, or 1,000,000,000 or , or any range between any two of the recited numbers, or more.
[0084] In some embodiments, the MI, e.g., UMI has a length. The length can be about 2- 4000 nt. For example, in some embodiments, the length is from about 2, 3, 4, 5, 6, 7, 8, 9, 10,
11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, to about 4000 nt, or any range between any two of these recited lengths. The length can be about 6-100 nt. In some embodiments, the length is from about 6 nt to about 100 nt. For example, in some embodiments, the length is about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 22, 24, 25, 28, 30, 32, 35, 36, 40, 42, 45, 48, 50, 55, 60, 64, 65, 70, 72, 75, 80, 85, 90, 95, or about 100 nt, or any range between any two of these recited lengths. The length can be about 8-50 nt. For example, in some embodiments, the length is about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or about 50 nt, or any range between any two of these recited lengths. The length can be about 10-20 nt. For example, in some embodiments, the length is about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or about 20 nt, or any range between any two of these recited lengths. The length can be about 12-14 nt. In some embodiments, the MI length is sufficient to uniquely barcode the molecules without interfering with downstream processing steps. In an embodiment, Mis, e.g., UMIs are randomly generated sequences that are confirmed to be distinct from the genomic sequences of the myeloid samples.
[0085] In some embodiments, methods described herein involve unique molecular identifiers (UMI). A UMI comprises a randomly selected stretch of nucleotides that can be used during sequencing to correct for PCR and sequencing errors, thereby adding an additional layer of error correction to sequencing results. In some embodiments, a MI, e.g, UMI or non-unique identier can be associated with a selected target sequence. Mis could be from, for example from 3-50 nucleotides long (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17. 18, 19, 20, 30, 40, or 50, or any range between any of those lengths). In some embodiments, the number of Mis, e.g., UMIs depends on the amount of input DNA. In some embodiments, the length of the UMI is sufficient to uniquely barcode the molecules and the length/sequence of the UMI does not interfere with the downstream amplification steps.
[0086] In some embodiments, all PCR duplicates from the same PCR reaction would have the same UMI sequence, as such the duplicates can be compared and any errors in the sequence such as single base substitutions, deletions, insertions (i.e., stutter in PCR) can be excluded from the sequencing results bioinformatically.
[0087] In some embodiments, the molecular identifiers, e.g., non-unique or unique molecular identifier (UMI) sequences are distinct from any sequences present in the native IMNA or NAPT fragments. This ensures that the Mis can be unambiguously distinguished
from genomic sequences during data analysis. In some embodiments such unique sequences can be randomly generated, e.g., by a computer readable medium, and selected by BLASTing against known nucleotide databases such as, e.g., GenBank. In some embodiments, in other embodiments, the MI sequences may be designed to match a specific sequence in the IMNA or NAPT fragments. In these cases, the position of the UMI within the read (e.g. at a specific distance from the R1 or R2 primer) is used to discriminate it from the genomic sequence during data analysis. In an embodiment, this approach may be used to label individual strands of a double-stranded fragment with the same UMI for applications like error correction.
[0088] In some embodiments, a primer of the present methods can comprise one or more Mis, strand distinguishing sequences, and/or primer sequences. In some embodiemnts, the MI, strand distinguishing sequences, and/or primer sequences can be one or more of primer sequences that are not homologous to the target sequence, but for example can be used as templates for one or more amplification reactions. In some embodiments, the MI, strand distinguishing sequences, and/or primer sequence can be a capture sequence, for example a hapten sequence such as biotin that can be used to purify amplicons away from reaction components. In an embodiment, the UMI, strand distinguishing sequences, and/or primer sequences can be sequences such as adaptor sequences that are advantageous for capturing the library amplicons on a substrate for example for bridge amplification in anticipation of sequence by synthesis technologies as described herein. In a further embodiment, MI, strand distinguishing sequences, and/or primer can be Mis, e.g, UMIs or non-unique Mis of between, for example, 3-20 nucleotides comprised of a randomized stretch of nucleotides that can be used for error correction during library preparation and/or sequencing methods.
[0089] The term "strand-distinguishing sequence" refers to a sequence on one strand of a double-stranded nucleic acid that is not complementary to a sequence at the corresponding position on the other strand. In some embodiments, a strand-distinguishing sequence is incorporated into each strand of the double-stranded molecule, with each strand receiving a different strand-distinguishing sequence. In some embodiments, a strand-distinguishing sequence on the first or Watson strand is not complementary to a strand-distinguishing sequence on the second or Crick strand. In some embodiments, the strand-distinguishing sequences on the first and second strands are independently selected. This means that the choice of strand-distinguishing sequence on one strand does not dictate or limit the choice of stranddistinguishing sequence on the opposite strand.
[0090] In some embodiments, the steps of isolating the nucleic acid, processing the nucleic acids, converting the isolated or processed nucleic acids to isolated DNA-adapters or processed
DNA-adapters, further preparing of the isolated DNA-adapters or processed DNA-adapters, e.g., by enriching isolated DNA-adapters or processed DNA-adapters or amplicons thereof by probe hybridization or by targeted amplification (including multiplex targeted amplification using targeted primers, or using a universal primer and targeted primers, which can be nested), sequencing, and error corrections are performed as any of those embodiments are referred to in one or more of WO2016181128A1, US20220073977A1, US10837063B2, US10557172B2, US20170088887A1, US9487829B2, US 10752951B2, which publications are hereby incorporated by reference.
[0091] B. Certain Exemplary Methods
[0092] The present disclosure also relates to methods for detecting one or more leukemic molecular alterations in nucleic acids from myeloid lineage tumor cells. The methods involve obtaining a sample comprising myeloid lineage cells, performing negative selection enrichment, or negative enrichment, to remove lymphoid cells while retaining myeloid cells, isolating nucleic acids from the residual cells of myeloid lineage, and detecting molecular alterations in the isolated myeloid nucleic acids (IMNA) or nucleic acids processed therefrom (NAPT).
[0093] In some embodiments, this disclosure provides methods for preparing nucleic acids from myeloid lineage tumor cells. The methods include obtaining a sample containing myeloid lineage cells, performing negative enrichment to remove lymphoid cells while retaining myeloid cells, and isolating nucleic acids from the residual cells of myeloid lineage. These negative enrichment steps are designed to remove non-myeloid cells from the sample, thereby increasing the relative proportion of myeloid cells (e.g. AML cells or other myeloid condition cells and other myeloid lineage malignancies cells). In an embodiment, the enrichment of myeloid cells is be quantified by comparing the number of meyloid cells present in the sample before and after the negative enrichment steps.
[0094] In some embodiments, the negative enrichment steps result in a 1.1-fold enrichment of myeloid or myeloid lineage cells relative to the initial number of myeloid or myeloid lineage cells present in the sample. In other embodiments, the negative enrichment steps result in a 1.2- fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, or 2.0-fold enrichment of AML cells.
[0095] In certain embodiments, the negative enrichment steps result in a 2-fold enrichment of myeloid (e.g., AML) cells relative to the initial number of myeloid (e.g., AML) cells present in the sample. In other embodiments, the negative enrichment steps result in a 3 -fold, 4-fold,
5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16- fold, 17-fold, 18-fold, 19-fold, or 20-fold enrichment of AML cells.
[0096] The degree of enrichment achieved by the negative enrichment steps may depend on various factors, such as the initial proportion of myeloid (e.g., AML) cells in the sample, the specific methods used for negative enrichment, and the efficiency of cell separation. In some embodiments, the negative enrichment steps are optimized to achieve a desired level of myeloid (e.g., AML) cell enrichment based on these factors.
[0097] In one embodiment, the negative enrichment steps involve the use of density gradient centrifugation to separate mononuclear cells, including myeloid (e.g., AML) cells, from other cell types based on their density. For example, Ficoll-Paque centrifugation can be used to isolate mononuclear cells from peripheral blood or bone marrow samples. This step can result in a 2-fold to 5-fold enrichment of myeloid (e.g., AML) cells, depending on the initial myeloid (e.g., AML) cell frequency and the efficiency of the separation.
[0098] In another embodiment, the negative enrichment steps involve the use of immunomagnetic beads to deplete non-myeloid (e.g., non-AML) cells from the sample. For example, beads coated with antibodies specific for T cells (e.g., CD3 or any one or more T cell lineage marker featured herein), specific for B cells (e.g., CD19 or any one or more B cell lineage marker featured herein), and specific for NK cells (e.g., KIR or any one or more NK cell lineage markers featuered herein) can be used to remove these cell types from the sample. This step can result in a 5-fold to 20-fold enrichment of myeloid (e.g., AML) cells, depending on the initial myeloid (e.g., AML) cell frequency and the efficiency of the depletion.
[0099] In one embodiment, the one or more B cell lineage markers are selected from: CD19, CD20, CD22, CD23, CD24, CD38, CD40, CD45, CD79a, CD79b, CD138, CD200 IgKappa, IgLambda, IgM, IgD, IgGl, IgG2, IgG3, IgG4, IgE, IgAl, and/or IgA2. In some embodiments, the B cell lineage markers analyzed may include or exclude one or more of the foregoing.
[00100] In one embodiment, the one or more T cell lineage markers are selected from: CD2, CD3, CD4, CD5, CD7, CD8, CD25, CD27, CD28, CD45RA, CD45RO, CD69, TRBC1, TRBC2, and/or CD127. In some embodiments, the T cell lineage markers analyzed may include or exclude one or more of the foregoing.
[00101] In one embodiment, the one or more NK cell lineage markers are selected from: CD57, CD94, CD122, CD158a, CD158b, CD159b, CD161, CD314, NKG2A, NKG2C, NKG2D, NKp30, NKp44, NKp46, NKp80, and KIR Family Receptors. In some embodiments, the NK cell lineage markers analyzed may include or exclude one or more of the foregoing.
[00102] In some embodiments, the negative enrichment steps involve a combination of density gradient centrifugation and immunomagnetic depletion. For example, mononuclear cells can be isolated by Ficoll-Paque centrifugation, followed by depletion of non-myeloid (e.g., non-AML) cells using immunomagnetic beads. This approach can result in a 10-fold to 20-fold enrichment of myeloid (e.g., AML) cells, depending on the initial myeloid (e.g., AML) cell frequency and the efficiency of each step.
[00103] In certain embodiments, the negative enrichment steps are performed using microfluidic devices that allow for the continuous separation of AML cells from non-myeloid (e.g., non-AML) cells based on their physical or biological properties. For example, microfluidic devices can be designed to separate cells based on their size, deformability, or expression of specific surface markers. In certain embodiemnts, the devices can achieve high levels of myeloid cell enrichment, with some studies reporting up to 1000-fold enrichment.
[00104] The effectiveness of the negative enrichment steps can be assessed by measuring the proportion of myeloid cells in the sample before and after enrichment. This can be done using flow cytometry, immunohistochemistry, or other methods that allow for the specific identification and quantification of myeloid cells. In some embodiments, the enriched myeloid (e.g., AML) cell population is further characterized using molecular techniques, such as nextgeneration sequencing or polymerase chain reaction, to detect myeloid -associated genetic alterations.
[00105] Obtaining the sample
[00106] In some embodiments, the sample from a subject is obtained or derived from the subject’ s whole blood, peripheral blood, peripheral blood mononuclear cells (PBMCs), or bone marrow, cells, or tissue. In certain embodiments, the sample contains an amount of nucleic acids. In some embodiments, obtaining or having obtained the sample results in disrupting or lysing cells in the sample to release nucleic acids from the cells. Thus, in some embodiments, the subject’s sample comprises cellular nucleic acids. In some embodiments, cellular nucleic acids make up less than about 1% of the total cellular nucleic acids in the sample. In some embodiments, cellular nucleic acids make up less than about 5% of the total cellular nucleic acids in the sample. In some embodiments, cellular nucleic acids make up less than about 10% of the total cellular nucleic acids in the sample. In some embodiments, cellular nucleic acids make up less than about 20% of the total cellular nucleic acids in the sample. In some embodiments, cellular nucleic acids make up more than about 50% of the total cellular nucleic acids in the sample. In some embodiments, cellular nucleic acids make up less than about 90% of the total cellular nucleic acids in the sample.
[00107] The sample used in the method may be obtained from various sources, including whole blood, peripheral blood, or fractions thereof such as PBMCs, or bone marrow aspirate. In some embodiments, the sample is collected from a patient suspected of having a myeloid malignancy, such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), myelodysplastic syndrome (MDS), or myeloproliferative neoplasms (MPNs). The sample may be collected at diagnosis, during treatment, or at follow-up to monitor for minimal residual disease or relapse.
[00108] In some embodiments, this disclosure provides for methods for obtaining or having obtained a sample for the methods as described herein. A sample may be obtained directly (e.g., a doctor takes a blood sample from a subject). A sample may be obtained indirectly (e.g., through shipping, by a technician from a doctor or a subject). In some embodiments, the sample is a whole blood, peripheral blood, peripheral blood mononuclear cells (PBMCs), or bone marrow.
[00109] In some embodiments, this disclosure provides for methods comprising obtaining or having obtained samples with fragmented nucleic acids. The sample may have been subjected to conditions that are not conducive to preserving the integrity of nucleic acids. The sample may have been exposed to air, heat, light, or chemicals or enzymes that degrade nucleic acids. In some embodiments, methods comprise obtaining or having obtained a tissue sample wherein the tissue sample comprises fragmented nucleic acids. In some embodiments, methods comprise obtaining or having obtained a tissue sample wherein the tissue sample comprises nucleic acids and fragmenting the nucleic acids to produced fragmented nucleic acids. In some embodiments, the tissue sample is a frozen sample. In some embodiments, the sample is a preserved sample. In some embodiments, the tissue sample is a fixed sample (e.g. formaldehyde-fixed). In an embodiment, methods may comprise isolating the (fragmented) nucleic acids from the sample.
[00110] In some embodiments, this disclosure provides for methods comprising obtaining or having obtained a sample from a subject, wherein the sample contains at least one nucleic acid of interest. In some embodiments, the nucleic acid of interest may be DNA from the whole blood, peripheral blood, peripheral blood mononuclear cells (PBMCs), or bone marrow of a subject. In some embodiments, this disclosure provides for methods comprising obtaining or having obtained a sample from the subject, wherein the sample contains about 1 to about 100 nucleic acids. In some embodiments, the at least one nucleic acid is represented by a sequence that is unique to a target gene disclosed herein.
[00111] Isolating and purifying nucleic acids from the sample
[00112] In some embodiments of any embodiment herein, this disclosure provides for methods comprising isolating or having isolated nucleic acids from a biological sample. In some embodiments, this disclosure provides for methods comprising purifying or having purified nucleic acids from a biological sample. As used herein the term “biological sample” means a whole blood, peripheral blood, peripheral blood mononuclear cells (PBMCs), or bone marrow sample obtained from a subject. In some embodiments, this disclosure provides for methods comprising isolating or having isolated and purifying or having purified nucleic acids from a biological sample. In some embodiments, this disclosure provides for methods of isolating nucleic acids comprising removing or having removed non-nucleic acid components from a biological sample described herein. The nucleic acids derived therefrom can include any or all the following manipulations described herein, including any method of isolating, purifying, enriching, and one or more steps involved in preparing a library of nucleic acids for detection from the nucleic acids.
[00113] In some embodiments, isolating or purifying comprises removing unwanted non- nucleic acid components from a biological sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of unwanted non-nucleic acid components from a biological sample, such as, e.g., components of blood or plasma. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 95% of unwanted non-nucleic acid components from a biological sample to obtain a sample derived from the initial sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 97% of unwanted non-nucleic acid components from a biological sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 98% of unwanted non-nucleic acid components from a biological sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 99% of unwanted non-nucleic acid components from a biological sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 95% of unwanted non- nucleic acid components from a biological sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 97% of unwanted non-nucleic acid components from a biological sample. In
some embodiments, isolating or having isolated and purifying or having purified nucleic acids from a biological sample comprises removing at least 98% of unwanted non-nucleic acid components from a biological sample. In some embodiments, isolating or having isolated and purifying or having purified IMNA or NAPT from a biological sample comprises removing at least 99% of unwanted non-nucleic acid components from a biological sample.
[00114] In some embodiments, this disclosure provides for methods comprising isolating or having isolated and purifying or having purified DNA, RNA, or a subset of desired nucleic acids from a mix of nucleic acids derived from a biological sample. In some embodiments, methods may comprise analyzing only DNA. In methods wherein only DNA is analyzed, RNA is unwanted and creates undesirable background noise or contamination to the DNA; therefore, in some embodiments, the removed nucleic acid components are discarded. In some embodiments, this disclosure provides for methods comprising removing RNA from a biological sample. In some embodiments, this disclosure provides for methods comprising removing mRNA from a biological sample. In some embodiments, this disclosure provides for methods comprising removing microRNA from a biological sample. In some embodiments, removing nucleic acid components comprises contacting the nucleic acid components with an oligonucleotide capable of hybridizing to the nucleic acid, wherein the oligonucleotide is conjugated, attached or bound to a capturing device (e.g., bead, column, matrix, nanoparticle, magnetic particle, etc.). In some embodiments, the removed nucleic acid components are discarded.
[00115] In one embodiment, disclosed herein are methods for performing one or more negative enrichment steps to remove lymphoid lineage cells while retaining the myeloid lineage cells. In some embodiments, the negative enrichment involves density gradient centrifugation, such as Ficoll-Paque centrifugation, to separate mononuclear cells from granulocytes and red blood cells. The mononuclear cell fraction, which includes lymphocytes and monocytes, can then be subjected to further negative selection to remove the lymphocytes. In one embodiment, the lymphocytes are removed using antibodies that specifically bind to cell surface markers expressed on B cells, T cells, and/or NK cells. These antibodies can be conjugated to magnetic beads, allowing the labeled cells to be separated from the unlabeled myeloid cells using a magnetic field. For example, anti-CD19 antibodies can be used to remove B cells, anti-CD3 antibodies for T cells, and anti-KIR antibodies for NK cells. In an embodiment, the specific markers targeted for depletion may vary depending on the desired purity and recovery of the myeloid cell population.
[00116] In some embodiments, the negative enrichment is performed using a microfluidic device. In some embodiments, the negative enrichment is performed using a combination of density gradient centrifugation and immunomagnetic cell separation. This two-step procedure can achieve higher purity of the myeloid cell population compared to either method alone. The enriched myeloid cells are then lysed and the nucleic acids are extracted using standard techniques such as phenol-chloroform extraction, silica-based columns, or magnetic beadbased methods.
[00117] In some embodiments, this disclosure provides for methods comprising isolating or having isolated and purifying or having purified nucleic acids from one or more non-nucleic acid components of a biological sample. Non-nucleic acid components may also be considered unwanted substances. Non-limiting examples of non-nucleic acid components include cells (e.g., blood cells), cell fragments, extracellular vesicles, lipids, proteins or a combination thereof. Additional non-nucleic acid components are described herein and throughout. It should be noted that while methods may comprise isolating and/or purifying nucleic acids, they may also comprise analyzing a non-nucleic acid component of a sample that is considered an unwanted substance in a nucleic acid purifying step. Isolating or having isolated and purifying or having purified may comprise removing components of a biological sample that would inhibit, interfere with or otherwise be detrimental to the later process steps such as nucleic acid amplification or detection.
[00118] In some embodiments, purifying or having purified nucleic acids does not comprise washing the nucleic acids with a wash buffer. In some embodiments, purifying or having purified nucleic acids comprises capturing the nucleic acids with a nucleic acid capturing moiety to produce captured nucleic acids. Non-limiting examples of nucleic acid capturing moieties are silica particles and paramagnetic particles. In some embodiments, purifying or having purified nucleic acids comprises passing the sample comprising the captured nucleic acids through a hydrophobic phase (e.g., a liquid or wax). The hydrophobic phase retains impurities in the sample that would otherwise inhibit further manipulation (e.g., amplification, sequencing) of the nucleic acids.
[00119] Isolating and/or purifying may occur with the use of a sample purifier. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids comprises removing non-nucleic acid components from a biological sample described herein. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids comprises discarding non-nucleic acid components from a biological sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids
comprises collecting, processing and analyzing the non-nucleic acid components. In some embodiments, the non-nucleic acid components may be considered biomarkers because they provide additional information about the subject.
[00120] In some embodiments, isolating or having isolated and purifying or having purified nucleic acids comprise lysing a cell. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids avoids lysing a cell. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids does not comprise lysing a cell. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids does not comprise an active step intended to lyse a cell. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids does not comprise intentionally lysing a cell. Intentionally lysing a cell may include mechanically disrupting a cell membrane (e.g., shearing). Intentionally lysing a cell may include contacting the cell with a lysis reagent.
[00121] In some embodiments, isolating or having isolated and purifying or having purified nucleic acids comprises separating components of a biological sample disclosed herein. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids comprises centrifuging the biological sample, filtering the biological sample, contacting the sample with a solid phase support, or using solid phase extraction, in order to separate components of a biological sample, or a combination thereof. In some embodiments, methods comprise subjecting the blood to vertical filtration. In some embodiments, methods comprise subjecting the blood to a sample purifier comprising a filter matrix for receiving whole blood, the filter matrix having a pore size that is prohibitive for cells to pass through, while nucleic acids can pass through the filter matrix uninhibited. Such vertical filtration and filter matrices are described for devices disclosed herein.
[00122] In some embodiments, isolating or having isolated and purifying or having purified nucleic acids comprises filtering the biological sample in order to remove non-nucleic acid components from the biological sample. In some embodiments, isolating or having isolated and purifying or having purified nucleic acids comprises filtering the biological sample in order to capture nucleic acids from the biological sample. In some embodiments, removing non-nucleic acid components may comprise centrifuging the biological sample. In some embodiments, removing non-nucleic acid components may comprise contacting the biological sample with a binding moiety.
[00123] Isolating or having isolated and purifying or having purified may comprise capturing an extracellular vesicle or extracellular microparticle in the biological sample with a
binding moiety. In some embodiments, the extracellular vesicle contains at least one of DNA and RNA.
[00124] Enrichment of nucleic acids
[00125] In some embodiments, this disclosure provides for methods of enriching nucleic acids isolated and purified from a biological sample, e.g. by hybridizing enrichment such as bait capture or by amplification, e.g., PCR amplification, or to provide a sample enriched in selected or target nucleic acids. In some embodiments, methods comprise performing targeted sequencing on the enriched nucleic acids, which may comprise amplicons of nucleic acids in the biological sample, obtained by whole genome or targeted amplification. In some embodiments, targeted amplification is performed using target-specific primers for fully nested or hemi-nested PCR.
[00126] In some embodiments, enriching nucleic acid components comprises separating the nucleic acid components on a gel by size. In some embodiments, this disclosure provides for methods comprising removing DNA from the biological sample. In some embodiments, this disclosure provides for methods comprising capturing DNA from the biological sample. In some embodiments, this disclosure provides for methods comprising selecting DNA from the biological sample. In some embodiments, the genomic DNA has a minimum length. In some embodiments, the minimum length is about 50 base pairs. In some embodiments, the minimum length is about 100 base pairs. In some embodiments, the minimum length is about 110 base pairs. In some embodiments, the minimum length is about 120 base pairs. In some embodiments, the minimum length is about 130, about 140, about 150, about 160, about 170 or about 180 base pairs. In some embodiments, the minimum length is about 130, about 140, about 150, about 160, or about. In some embodiments, the DNA has a maximum length. In some embodiments, the maximum length is about 180 base pairs. In some embodiments, the maximum length is about 200 base pairs. In some embodiments, the maximum length is about 220 base pairs. In some embodiments, the maximum length is about 240 base pairs. In some embodiments, the maximum length is about 300 base pairs. Size based separation would be useful for other categories of nucleic acids having limited size ranges (e.g., microRNAs).
[00127] In some embodiments, enriching methods comprise capturing a nucleosome in a biological sample and analyzing nucleic acids attached to the nucleosome. In some embodiments, methods comprise capturing an exosome in a biological sample and analyzing nucleic acids attached to the exosome. Capturing nucleosomes and/or exosomes may preclude the need for a lysis step or reagent, thereby simplifying the method and reducing time from sample collection to detection. In some embodiments, enriching nucleic acids comprises lysing
and performing sequence specific capture of a target nucleic acid with "bait" in a solution followed by binding of the "bait" to solid supports such as magnetic beads. In some embodiments, methods comprise performing sequence specific capture in the presence of a recombinase or helicase to reduce the need for heat denaturation of a nucleic acid thereby speeding up the detection step.
[00128] In some embodiments, enriching nucleic acids comprises subjecting a biological sample, or a fraction thereof, or a modified version thereof, to a binding moiety. The binding moiety may be capable of binding to a component of a biological sample and removing it to produce a modified sample depleted of cells, cell fragments, nucleic acids or proteins that are unwanted or of no interest. In some embodiments, enriching purified nucleic acids comprises subjecting a biological sample to a binding moiety to reduce unwanted substances or non- nucleic acid components in a biological sample. In some embodiments, enriching purified nucleic acids comprises subjecting a biological sample to a binding moiety to produce a modified sample enriched with target cell, target cell fragments, target nucleic acids or target proteins. The resulting cell-bound binding moieties can be captured and enriched for with antibodies or other methods, e.g., low speed centrifugation.
[00129] In some embodiments, this disclosure provides for methods comprising amplifying at least one nucleic acid in a sample to produce at least one amplification product. The at least one nucleic acid may be a genomic nucleic acid or nucleic acid obtained from a cell in a sample. The sample may be a biological sample disclosed herein or a fraction or portion thereof. In some embodiments, methods comprise producing a copy of the nucleic acid in the sample and amplifying the copy to produce the at least one amplification product. In some embodiments, methods comprise producing a reverse transcript of the nucleic acid in the sample and amplifying the reverse transcript to produce the at least one amplification product.
[00130] In some embodiments, methods comprise performing whole genome amplification. In some embodiments, methods do not comprise performing whole genome amplification. The term, "whole genome amplification" also referred to as “random amplification” may refer to amplifying all of the genomic nucleic acids in a biological sample. The term, "whole genome amplification" may refer to amplifying at least 90% of the genomic nucleic acids in a biological sample. Whole genome may refer to multiple genomes. Whole genome amplification may comprise amplifying genomic nucleic acids from a biological sample of a subject having an infection, wherein the biological sample comprises genomic nucleic acids from the subject and a pathogen.
[00131] In some embodiments, targeted sequencing of the IMNA- or NAPT-adapters is performed using bait-capture methods. In some embodiments, a prepared pool of NGS templates is heat denatured and mixed with a pool of capture probe oligonucleotides (“baits”). In an embodiment, the baits are designed to hybridize to the regions of interest within the target genome and are 60-200 bases in length, for example, about 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or about 200 bases in length, or any range between any two of these recited lengths. In some embodiments, baits are further modified to contain a ligand that permits subsequent capture of these probes. In an embodiment, such oligonucleotide baits comprise deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or analogs thereof, including oligonucleotides with modified backbones or bases to enhance hybridization specificity, stability, or reduce non-specific binding. In an embodiment, the length of the baits can vary widely, for example, depending on the specific application, desired hybridization kinetics, the complexity of the target regions, and the specific enrichment strategy employed. In some embodiments, a commonly used approach involves baits with lengths typically ranging from about 100 nucleotides to about 150 nucleotides, for example, around about 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110,
111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129,
130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148,
149, or about 150 nucleotides, or any range between any two of these recited lengths. In some embodiments, strategies utilizing shorter baits are employed, with bait lengths for example ranging from about 30 nucleotides to about 50 nucleotides, e.g., about 30, 35, 40, 45, or 50 nucleotides; such shorter baits can, in some embodiments, be utilized in single or multiple rounds of hybridization to achieve the desired target enrichment. While these represent common ranges, the baits can also have lengths outside of these specific examples, for instance, generally ranging from about 30 nucleotides up to about 500 nucleotides or more. In some embodiments, this range includes lengths such as about 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, or about 500 nucleotides, or any range between any two recited lengths, depending on the requirements of the assay.
[00132] In an embodiment, the capture method incorporates a biotin group (or groups) on the baits. In an embodiment, other capture ligands are used. In an embodiment, after hybridization is complete to form the DNA template:bait hybrids, capture is performed with a component having affinity for only the bait. In an embodiment, streptavidin-magnetic beads are used to bind the biotin moiety of biotinylated-baits that are hybridized to the desired DNA
targets from the pool of NGS templates. In an embodiment, washing removes unbound nucleic acids, reducing the complexity of the retained material. In an embodiment, the retained material is then eluted from the magnetic beads and introduced into automated sequencing processes, providing for ‘capture enrichment’, where the captured nucleic acids are retained as an enriched pool for subsequent study.
[00133] In an embodiment, the design of bait oligonucleotides accounts for potential sequence variations in the population to ensure efficient capture of target regions across different individuals. In an embodiment, baits can be synthesized with modified nucleotides or backbones (e.g., locked nucleic acids - LNAs) to enhance their binding affinity and specificity for target sequences, or to improve their stability. The density and tiling of baits across a target region can also be optimized to ensure uniform capture and coverage. Furthermore, computational methods can be used to design bait sequences that minimize cross-hybridization to off-target genomic regions, including pseudogenes or repetitive elements.
[00134] In some embodiments, this disclosure provides for methods comprising amplifying a nucleic acid, wherein amplifying comprises performing an isothermal amplification of the nucleic acid. Non-limiting examples of isothermal amplification are as follows: loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicase dependent amplification (HDA), nicking enzyme amplification reaction (NEAR), and recombinase polymerase amplification (RPA). In some embodiments, the isothermal amplification is high throughput involving parallel sample processing. In some embodiments, the high throughput isothermal amplification involves amplifying a nucleic acid in 12, 24, 36, 48, 60, 72, 84, 96, 108, or more samples in parallel. In some embodiments, the high throughput isothermal amplification involves amplifying a nucleic acid in between 12-24, 24-36, 36-48, 48-60, 70-72, 72-84, 84-96, 96-108, 108-120, 120-132, 132-144, 144-156-156-168, 168-180, 180-192, 192-204, 204-216, 216-228, 228-240, 240-252, or 252-264, samples in parallel. In some embodiments, the high throughput isothermal amplification involves amplifying a nucleic acid in at least 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, or 1,500 samples in parallel.
[00135] Amplification methods can include isothermal amplification. In some embodiments, amplification is isothermal with the exception of an initial heating step before isothermal amplification begins. A number of isothermal amplification methods, each having different considerations and providing different advantages, can be used (Zanoli and Spoto, 2013, "Isothermal Amplification Methods for the Detection of Nucleic Acids in Microfluidic Devices," Biosensors 3: 18-43; Fakruddin, et al., 2013, "Alternative Methods of Polymerase
Chain Reaction (PCR)," Journal of Pharmacy and Bioallied Sciences 5(4): 245-252). In some embodiments, any appropriate isothermic amplification method is used. In some embodiments, the isothermic amplification method used is selected from: Loop Mediated Isothermal Amplification (LAMP); Nucleic Acid Sequence Based Amplification (NASBA); Multiple Displacement Amplification (MDA); Rolling Circle Amplification (RCA); Helicase Dependent Amplification (HDA); Strand Displacement Amplification (SDA); Nicking Enzyme Amplification Reaction (NEAR); Ramification Amplification Method (RAM); and Recombinase Polymerase Amplification (RPA).
[00136] In some embodiments, the amplification method is Nucleic Acid Sequence Based Amplification (NASBA). NASBA (also known as 3 SR, and transcription-mediated amplification) is an isothermal transcription-based RNA amplification system. Three enzymes (avian myeloblastosis virus reverse transcriptase, RNase Hand T7 DNA dependent RNA polymerase) are used to generate single-stranded RNA. In certain cases, NASBA can be used to amplify DNA. The amplification reaction is performed at 41 °C, maintaining constant temperature, typically for about 60 to about 90 minutes (see, e.g., Fakruddin, et al., 2012, "Nucleic Acid Sequence Based Amplification (NASBA) Prospects and Applications," Int. J. of Life Science and Pharma Res. 2(l):L106-L121).
[00137] In some embodiments, the NASBA reaction is carried out at about 40 °C to about 42 °C. In some embodiments, the NASBA reaction is carried out at 41 °C. In some embodiments, the NASBA reaction is carried out at at most about 42 °C. In some embodiments, the NASBA reaction is carried out at about 40 °C to about 41 °C, about 40 °C to about 42 °C, or about 41 °C to about 42 °C. In some embodiments, the NASBA reaction is carried out at about 40 °C, about 41 °C, or about 42 °C.
[00138] In some embodiments, the amplification method is Strand Displacement Amplification (SDA). SDA is an isothermal amplification method that uses four different primers. A primer comprising a restriction site (a recognition sequence for Hindi exonuclease) is annealed to the DNA template. An exonuclease-deficient fragment of Eschericia coli DNA polymerase 1 (exoKlenow) elongates the primers. Each SDA cycle consists of (1) primer binding to a displaced target fragment, (2) extension of the primer/target complex by exoKlenow, (3) nicking of the resultant hemiphosphothioate Hindi site, (4) dissociation ofHindl from the nicked site and (5) extension of the nick and displacement of the downstream strand by exo-Klenow.
[00139] In some embodiments, this disclosure provides for methods comprising contacting DNA in a sample with a helicase. In some embodiments, the amplification method is Helicase
Dependent Amplification (HD A). HDA is an isothermal reaction because a helicase, instead of heat, is used to denature DNA.
[00140] In some embodiments, the amplification method is Multiple Displacement Amplification (MDA). The MDA is an isothermal, strand-displacing method based on the use of the highly processive and strand-displacing DNA polymerase from bacteriophage 029, in conjunction with modified random primers to amplify the entire genome with high fidelity. It has been developed to amplify all DNA in a sample from a very small amount of starting material. In MDA 029 DNA polymerase is incubated with dNTPs, random hexamers and denatured template DNA at 30°C for 16 to 18 hours and the enzyme must be inactivated at high temperature (65°C) for 10 min. No repeated recycling is required, but a short initial denaturation step, the amplification step, and a final inactivation of the enzyme are needed.
[00141] In some embodiments, the amplification method is Rolling Circle Amplification (RCA). RCA is an isothermal nucleic acid amplification method which allows amplification of the probe DNA sequences by more than 109 fold at a single temperature, typically about 30°C. Numerous rounds of isothermal enzymatic synthesis are carried out by 029 DNA polymerase, which extends a circle-hybridized primer by continuously progressing around the circular DNA probe. In some embodiments, the amplification reaction is carried out using RCA, at about 28°C to about 32°C. In some embodiments, RCA is used to amplify ligated molecular inversion probes, optionally ligated barcoded molecular inversion probes.
[00142] Additional amplification methods can be used. Ideally, the amplification method is isothermal and fast relative to traditional PCR. In some embodiments, amplifying comprises performing an exponential amplification reaction (EXP AR), which is an isothermal molecular chain reaction in that the products of one reaction catalyze further reactions that create the same products. In some embodiments, amplifying occurs in the presence of an endonuclease. The endonuclease may be a nicking endonuclease. (Wu et al., "Aligner-Mediated Cleavage of Nucleic Acids," Chemical Science (2018)). In some embodiments, amplifying does not require initial heat denaturation of target DNA. (Toley et al., "Isothermal strand displacement amplification (iSDA): a rapid and sensitive method of nucleic acid amplification for point-of- care diagnosis," The Analyst (2015)).
[00143] In some embodiments, this disclosure provides for methods comprising performing multiple cycles of nucleic acid amplification with a pair of primers. The number of amplification cycles is important because amplification may introduce a bias into the representation of regions. Not all regions amplify with the same efficiency and therefore the overall representation may not be uniform which will impact the accuracy of the analysis.
Fewer amplification cycles are ideal if amplification is necessary at all. In some embodiments, methods comprise performing fewer than 50, 45, 40, 35, 30, or 25 cycles of amplification. In some embodiments, methods comprise performing fewer than 25 cycles of amplification. In some embodiments, methods comprise performing fewer than 20 cycles of amplification. In some embodiments, methods comprise performing fewer than 15 cycles of amplification. In some embodiments, methods comprise performing fewer than 12 cycles of amplification. In some embodiments, methods comprise performing fewer than 11 cycles of amplification. In some embodiments, methods comprise performing fewer than 10 cycles of amplification. In some embodiments, methods comprise performing at least 3 cycles of amplification. In some embodiments, methods comprise performing at least 5 cycles of amplification. In some embodiments, methods comprise performing at least 8 cycles of amplification. In some embodiments, methods comprise performing at least 10 cycles of amplification.
[00144] In some embodiments, the amplification reaction is carried for about 30 seconds to about 90 minutes. In some embodiments, the amplification reaction is carried out for at least about 30 minutes. In some embodiments, the amplification reaction is carried out for at most about 90 minutes. In some embodiments, the amplification reaction is carried out for about 30 minutes to about 35 minutes, about 30 minutes to about 40 minutes, about 30 minutes to about 45 minutes, about 30 minutes to about 50 minutes, about 30 minutes to about 55 minutes, about 30 minutes to about 60 minutes, about 30 minutes to about 65 minutes, about 30 minutes to about 70 minutes, about 30 minutes to about 75 minutes, about 30 minutes to about 80 minutes, about 30 minutes to about 90 minutes, about 35 minutes to about 40 minutes, about 35 minutes to about 45 minutes, about 35 minutes to about 50 minutes, about 35 minutes to about 55 minutes, about 35 minutes to about 60 minutes, about 35 minutes to about 65 minutes, about 35 minutes to about 70 minutes, about 35 minutes to about 75 minutes, about 35 minutes to about 80 minutes, about 35 minutes to about 90 minutes, about 40 minutes to about 45 minutes, about 40 minutes to about 50 minutes, about 40 minutes to about 55 minutes, about 40 minutes to about 60 minutes, about 40 minutes to about 65 minutes, about 40 minutes to about 70 minutes, about 40 minutes to about 75 minutes, about 40 minutes to about 80 minutes, about 40 minutes to about 90 minutes, about 45 minutes to about 50 minutes, about 45 minutes to about 55 minutes, about 45 minutes to about 60 minutes, about 45 minutes to about 65 minutes, about 45 minutes to about 70 minutes, about 45 minutes to about 75 minutes, about 45 minutes to about 80 minutes, about 45 minutes to about 90 minutes, about 50 minutes to about 55 minutes, about 50 minutes to about 60 minutes, about 50 minutes to about 65 minutes, about 50 minutes to about 70 minutes, about 50 minutes to about 75 minutes, about 50 minutes to
about 80 minutes, about 50 minutes to about 90 minutes, about 55 minutes to about 60 minutes, about 55 minutes to about 65 minutes, about 55 minutes to about 70 minutes, about 55 minutes to about 75 minutes, about 55 minutes to about 80 minutes, about 55 minutes to about 90 minutes, about 60 minutes to about 65 minutes, about 60 minutes to about 70 minutes, about 60 minutes to about 75 minutes, about 60 minutes to about 80 minutes, about 60 minutes to about 90 minutes, about 65 minutes to about 70 minutes, about 65 minutes to about 75 minutes, about 65 minutes to about 80 minutes, about 65 minutes to about 90 minutes, about 70 minutes to about 75 minutes, about 70 minutes to about 80 minutes, about 70 minutes to about 90 minutes, about 75 minutes to about 80 minutes, about 75 minutes to about 90 minutes, or about 80 minutes to about 90 minutes. In some embodiments, the amplification reaction is carried out for about 30 minutes, about 35 minutes, about 40 minutes, about 45 minutes, about 50 minutes, about 55 minutes, about 60 minutes, about 65 minutes, about 70 minutes, about 75 minutes, about 80 minutes, or about 90 minutes.
[00145] In some embodiments, this disclosure provides for methods comprising amplifying a nucleic acid at at least one temperature. In some embodiments, this disclosure provides for methods comprising amplifying a nucleic acid at a single temperature (e.g., isothermal amplification). In some embodiments, this disclosure provides for methods comprising amplifying a nucleic acid, wherein the amplifying occurs at not more than two temperatures. Amplifying may occur in one step or multiple steps. Non-limiting examples of amplifying steps include double strand denaturing, primer hybridization, and primer extension.
[00146] In some embodiments, at least one step of amplifying occurs at room temperature. In some embodiments, all steps of amplifying occur at room temperature. In some embodiments, at least one step of amplifying occurs in a temperature range. In some embodiments, all steps of amplifying occur in a temperature range. In some embodiments, the temperature range is about 0° C to about 100°C. In some embodiments, the temperature range is about 15°C to about 100°C. In some embodiments, the temperature range is about 25°C to about 100°C. In some embodiments, the temperature range is about 35°C to about 100°C. In some embodiments, the temperature range is about 55°C to about 100°C. In some embodiments, the temperature range is about 65°C to about 100°C. In some embodiments, the temperature range is about 15°C to about 80°C. In some embodiments, the temperature range is about 25°C to about 80°C. In some embodiments, the temperature range is about 35°C to about 80°C. In some embodiments, the temperature range is about 55°C to about 80°C. In some embodiments, the temperature range is about 65°C to about 80°C. In some embodiments, the temperature range is about 15°C to about 60°C. In some embodiments, the temperature range
is about 25°C to about 60°C. In some embodiments, the temperature range is about 35°C to about 60°C. In some embodiments, the temperature range is about 15°C to about 40°C. In some embodiments, the temperature range is about -20°C to about 100°C. In some embodiments, the temperature range is about -20°C to about 90°C. In some embodiments, the temperature range is about -20°C to about 50°C. In some embodiments, the temperature range is about -20°C to about 40°C. In some embodiments, the temperature range is about -20°C to about 10°C. In some embodiments, the temperature range is about 0°C to about 100°C. In some embodiments, the temperature range is about 0°C to about 40°C. In some embodiments, the temperature range is about 0°C to about 30°C. In some embodiments, the temperature range is about 0°C to about 20°C. In some embodiments, the temperature range is about 0°C to about 10°C. In some embodiments, the temperature range is about 15°C to about 100°C. In some embodiments, the temperature range is about 15°C to about 90°C. In some embodiments, the temperature range is about 15°C to about 80°C. In some embodiments, the temperature range is about is about 15°C to about 70°C. In some embodiments, the temperature range is about 15°C to about 60°C. In some embodiments, the temperature range is about 15°C to about 50 °C. In some embodiments, the temperature range is about 15°C to about 30°C. In some embodiments, the temperature range is about 10°C to about 30°C. In some embodiments, methods disclose herein are performed at room temperature, not requiring cooling, freezing or heating. In some embodiments, amplifying comprises contacting the sample with random oligonucleotide primers.
[00147] In some embodiments, amplifying comprises targeted amplification (including that described in US6558928). In some embodiments, amplifying a nucleic acid comprises contacting a nucleic acid with at least one primer having a sequence corresponding to a target gene sequence. In some embodiments, amplifying comprises contacting the nucleic acid with at least one primer having a sequence corresponding to a non-target gene sequence. In some embodiments, amplifying comprises contacting the nucleic acid with not more than one pair of primers, wherein each primer of the pair of primers comprises a sequence corresponding to a sequence on a target gene disclosed herein. In some embodiments, amplifying comprises contacting the nucleic acid with multiple sets of primers, wherein each of a first pair in a first set and each of a pair in a second set are all different.
[00148] In some embodiments, amplifying comprises contacting the sample with at least one primer having a sequence corresponding to a sequence on a target gene disclosed herein. In some embodiments, amplifying comprises contacting the sample with at least one primer having a sequence corresponding to a sequence on a non-target gene disclosed herein. In some
embodiments, amplifying comprises contacting the sample with not more than one pair of primers, wherein each primer of the pair of primers comprises a sequence corresponding to a sequence on a target gene disclosed herein. In some embodiments, amplifying comprises contacting the sample with multiple sets of primers, wherein each of a first pair in a first set and each of a pair in a second set are all different.
[00149] In some embodiments, amplifying comprises multiplexing (nucleic acid amplification of a plurality of nucleic acids in one reaction). In some embodiments, multiplexing comprises contacting nucleic acids of the biological sample with a plurality of oligonucleotide primer pairs. In some embodiments, multiplexing comprising contacting a first nucleic acid and a second nucleic acid, wherein the first nucleic acid corresponds to a first sequence and the second nucleic acid corresponds to a second sequence. In some embodiments, the first sequence and the second sequence are the same. In some embodiments, the first sequence and the second sequence are different. In some embodiments, amplifying does not comprise multiplexing. In some embodiments, amplifying does not require multiplexing. In some instance, amplifying comprises nested primer amplification. Methods may comprise multiplex PCR of multiple regions, wherein each region comprises a SNP. Multiplexing may occur in a single tube. In some embodiments, methods comprise multiplex PCR of more than 100 regions wherein each region comprises a SNP. In some embodiments, methods comprise multiplex PCR of more than 500 regions wherein each region comprises a SNP. In some embodiments, methods comprise multiplex PCR of more than 1000 regions wherein each region comprises a SNP. In some embodiments, methods comprise multiplex PCR of more than 2000 regions wherein each region comprises a SNP. In some embodiments, methods comprise multiplex PCR of more than 300 regions wherein each region comprises a SNP.
[00150] In some embodiments, this disclosure provides for methods comprising amplifying a nucleic acid in the sample, wherein amplifying comprises contacting the sample with at least one oligonucleotide primer, wherein the at least one oligonucleotide primer is not active or extendable until it is in contact with the sample. In some embodiments, amplifying comprises contacting the sample with at least one oligonucleotide primer, wherein the at least one oligonucleotide primer is not active or extendable until it is exposed to a selected temperature. In some embodiments, amplifying comprises contacting the sample with at least one oligonucleotide primer, wherein the at least one oligonucleotide primer is not active or extendable until it is contacted with an activating reagent. In some embodiments, the at least one oligonucleotide primer comprises a blocking group. Using such oligonucleotide primers may minimize primer dimers, allow recognition of unused primer, and/or avoid false results
caused by unused primers. In some embodiments, amplifying comprises contacting the sample with at least one oligonucleotide primer comprising a sequence corresponding to a sequence on a target gene disclosed herein.
[00151] In some embodiments, this disclosure provides for methods comprising detecting an amplification product, wherein the amplification product is produced by amplifying at least a portion of a target gene disclosed herein. In some embodiments, detecting amplification products disclosed herein does not comprise barcoding or labeling the amplification product. In some embodiments, methods detect the amplification product based on its amount. For example, the methods may detect an increase in the amount of double stranded DNA in the sample. In some embodiments, detecting the amplification product is at least partially based on its size. In some embodiments, the amplification product has a length of about 50 base pairs to about 500 base pairs.
[00152] Library preparation
[00153] In some embodiments, this disclosure provides for methods comprising modifying genomic nucleic acids from the biological sample to produce a library of IMNA or NAPT fragments or amplicons for detection. In some embodiments, methods comprise modifying nucleic acids for nucleic acid sequencing. In some embodiments, methods comprise modifying nucleic acids for detection, wherein detection does not comprise nucleic acid sequencing. In some embodiments, methods comprise modifying nucleic acids for detection, to obtain nucleic acids derived from the isolated myeloid nucleic acids (IMNA) or in nucleic acids processed therefrom (NAPT) in the initial sample, wherein detection comprises counting IMNA-adapters or NAPT-adapters nucleic acids based on an occurrence of IMNA-adapters, NAPT-adapters, Mis (e.g., UMIs or non-unique Mis), primers, or primer binding sites. In some embodiments, this disclosure provides for methods comprising converting in one or more stepa, IMNA or NAPT in the biological sample to a library of IMNA-adapters and NAPT-adapters, wherein the method comprises amplifying the nucleic acids. In some embodiments, modifying occurs before amplifying. In some embodiments, modifying occurs after amplifying.
[00154] In some embodiments, modifying the nucleic acids comprises repairing ends of nucleic acids that are fragments of a nucleic acid. In some embodiments, repairing ends may comprise restoring a 5’ phosphate group, a 3’ hydroxy group, or a combination thereof to the nucleic acid. In some embodiments, repairing comprises 5’ phosphorylation, A-tailing, gap filling, closing nick sites or a combination thereof. In some embodiments, repairing may comprise removing overhangs. In some embodiments, repairing may comprise filling in overhangs with complementary nucleotides. In some embodiments, modifying the nucleic
acids for preparing a library comprises use of an adapter. The adapter may also be referred to herein as a sequencing adapter. In some embodiments, the adapter aids in sequencing. The adapter can comprise an oligonucleotide. In some embodiments, the adapter may simplify other steps in the methods, such as amplifying, purification and sequencing because it is a sequence that is universal to multiple, if not all, nucleic acids in a sample after modifying. In some embodiments, modifying the nucleic acids comprises ligating an adapter to the nucleic acids. Ligating may comprise blunt ligation. In some embodiments, modifying the nucleic acids comprises hybridizing an adapter to the nucleic acids. In some embodiments, the sequencing adapter comprises a hairpin or stem-loop adapter. In some embodiments, modifying the nucleic acids comprises hybridizing a hairpin or stem-loop adapter to the nucleic acids, thereby generating a circular library product that is sequenced or analyzed. In some embodiments, the sequencing adapter comprises a blocked 5’ end leaving a nick at the 3’ end. Advantages of this configuration include, but are not limited to, an increase in library efficiency and reduction of unwanted byproducts such as adapter dimers. In some embodiments the adapter has a cleavable replication stop to linearize templates.
[00155] In some embodiments, modifying the nucleic acids for preparing a library comprises use of an IMNA-adapters, NAPT-adapters, Mis, e.g, non-unique Mis or UMIs, primers, or primer binding sites. The IMNA-adapters, NAPT-adapters, Mis, e.g, non-unique Mis or UMIs, primers, or primer binding sites may also be referred to herein as a barcode. In some embodiments, this disclosure provides for methods comprising modifying nucleic acids with a barcode that corresponds to a chromosomal region of interest. In some embodiments, this disclosure provides for methods comprising modifying nucleic acids with a barcode that is specific to a chromosomal region that is not of interest. In some embodiments, this disclosure provides for methods comprising modifying a first portion of nucleic acids with a first barcode that corresponds to at least one chromosomal region that is of interest and a second portion of nucleic acids with a second barcode that corresponds to at least one chromosomal region that is not of interest. In some embodiments, modifying the nucleic acids comprises ligating a barcode to the nucleic acids. Ligating may comprise blunt ligation. In some embodiments, modifying the nucleic acids comprises hybridizing a barcode to the nucleic acids. In some embodiments, the barcodes comprise oligonucleotides. In some embodiments, the barcodes comprise a non-oligonucleotide marker or label that can be detected by means other than nucleic acid analysis. In some embodiments, a non-oligonucleotide marker or label could comprise a fluorescent molecule, a nanoparticle, a dye, a peptide, or other detectable/quantifiable small molecule.
[00156] In some embodiments, modifying the nucleic acids for preparing a library comprises use of a sample index, also simply referred to herein as an index. In some embodiments, the index may comprise an oligonucleotide, a small molecule, a fluorescent molecule, a dye, or other detectable/quantifiable moiety. In some embodiments, a first group of nucleic acids from a first biological sample are labeled with a first index, and a first group of nucleic acids from a first biological sample are labeled with a second index, wherein the first index and the second index are different. Thus, multiple indexes allow for distinguishing nucleic acids from multiple samples when multiple samples are analyzed at once. In some embodiments, methods disclose amplifying nucleic acids wherein an oligonucleotide primer used to amplify the nucleic acids comprises an index.
[00157] In some embodiments, the methods disclosed herein comprise amplification and library preparation in anticipation of downstream sequencing. An assay can include one or two PCR master mixes, one or two thermostable polymerases, one or two primer mixes and library adaptors. In some embodiments, a sample of DNA may be amplified for a number of cycles by using a first set of amplification primers that comprise target specific regions and non-target specific barcode regions and a first PCR master mix. The barcode region can be any sequence, such as a universal barcode region, a capture barcode region, an amplification barcode region, a sequencing barcode region, a UMI (unique molecular identifier) barcode region, and the like. In some embodiments, a barcode region can be the template for amplification primers utilized in a second or subsequent round of amplification, for example for library preparation. In some embodiments, the methods comprise adding single stranded-binding protein to the first amplification products. An aliquot of the first amplified sample can be removed and amplified a second time using a second set of amplification primers that are specific to the barcode region, e.g., a universal barcode region or an amplification barcode region, of the first amplification primers which may comprise of one or more additional barcode sequences, such as sequence barcodes specific for one or more downstream sequencing workflows, and the same or a second PCR master mix. As such, a library of the original DNA sample can be ready for sequencing. [00158] Sequencing
[00159] In certain embodiments, the method of preparing gene amplicons or enriched nucleic acid samples further comprises detecting the clonally amplified gene amplicons by multiplex PCR or sequencing.
[00160] In certain embodiments, clonally amplified gene amplicons are detected using, for example, polymerase chain reaction (PCR), multiplex PCR, or the Luminex assay (Bio-techne, Minneapolis, MN, USA). In some embodiments, clonally amplified gene amplicons are
detected using sequencing. In some embodiments, the sequencing methods can include or exclude parallel sequencing, nanopore sequencing, or pyrosequencing. In some embodiments, the sequencing can be targeted or random sequencing.
[00161] In some embodiments, this disclosure provides for methods comprising sequencing a nucleic acid. The nucleic acid may be a nucleic acid disclosed herein, such as an amplified nucleic acid. In some embodiments the nucleic acid is DNA.In some embodiments, the nucleic acid is a fetal nucleic acid, a newborn nucleic acid (obtained from a subject having an age less than 3, 2, or 1 years) a nucleic acid having a sequence corresponding to a target gene, a nucleic acid having a sequence corresponding to a region of a target gene, a nucleic acid having a sequence corresponding to a non-target gene, or a combination thereof.
[00162] In some embodiments, sequencing comprises targeted sequencing. In some embodiments, sequencing comprises whole genome sequencing. In some embodiments, sequencing comprises targeted sequencing and whole genome sequencing. In some embodiments, whole genome sequencing comprises massive parallel sequencing. In some embodiments, whole genome sequencing comprises random massive parallel sequencing. In some embodiments, sequencing comprises random massive parallel sequencing of target regions captured from a whole genome library.
[00163] In some embodiments, the methods comprise sequencing Watson amplicons and Crick amplicons. In some embodiments, Watson amplicons and Crick amplicons are produced by targeted amplification (e.g., with primers specific to target sequences of interest, or pulldown capture (e.g., bait) sequences first isolating selected regions of a genome followed by the amplification of said pulled-down regions). In some embodiments, Watson amplicons and Crick amplicons are produced by non-targeted amplification (e.g., with random oligonucleotide primers). In some embodiments, methods comprise sequencing amplified nucleic acids.
[00164] The present methods are not limited to any particular sequencing platform. In some embodiments, the sequencing can be performed by sequencing by synthesis (“SBS). In some embodiments, sequencing comprises a flow cell wherein nucleic acids are attached at (or through a hydrogel to) fixed locations in an array such that their relative positions do not change and wherein the array is repeatedly imaged. Examples in which images are obtained in different color channels, for example, coinciding with different labels used to distinguish one nucleotide base type from another are particularly applicable.
[00165] In some embodiments, SBS methods comprise the enzymatic extension of a priming nucleic acid strand through the iterative addition of (optionally labelled-and-3’ blocked)
nucleotides against a template strand. In certain SBS methods, a single nucleotide monomer may be provided to a target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to a target nucleic acid in the presence of a delivery ed polymerase.
[00166] In some embodiments, SBS methods involve nucleotide monomers that comprise a label moiety. In some embodiments, SBS methods involve nucleotide monomers that no not have a label moiety. Nucleotide incorporation events can be detected based on a characteristic of the label, such as fluorescence of the label; a characteristic of the nucleotide monomer such as charge; a byproduct of incorporation of the nucleotide, such as release of pyrophosphate; or the like. In some embodiments, two or more nucleotides are present in a sequencing step, and the different nucleotides can be distinguishable from each other, or alternatively, the two or more different labels can be the indistinguishable under the detection techniques being used. For example, the different nucleotides present in a sequencing reagent can have different labels and they can be distinguished using appropriate optics as exemplified by the sequencing methods presently commercialized by Illumina, Inc. (San Diego, CA), Ultima Genomics, Pacific Biosciences, Complete Genomics (MGI), ThermoFisher, Element Biosciences, or Singular Genomics.
[00167] In some embodiments, SBS involves pyrosequencing. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) or protons as particular nucleotides are incorporated into the nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) “Real-time DNA sequencing using detection of pyrophosphate release.” Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) “Pyrosequencing sheds light on DNA sequencing.” Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) “A sequencing method based on real-time pyrophosphate.” Science 281(5375), 363; U.S. Pat. No. 6,210,891; the disclosures of which are incorporated herein by reference in their entireties).
[00168] The pyrosequencing mechanism involves detecting released PPi by being enzymatically converted to adenosine triphosphate (ATP) by ATP sulfurylase, such that the amount of ATP generated is detected via luciferase-produced photons. The nucleic acids to be sequenced can be attached to features in an array comprising wells to localize the released ATP and reduce or prevent ATP crossover from well to well, and the array can be imaged to capture the chemiluminescent signals that are produced due to incorporation of a nucleotides at the features of the array. An image can be obtained after the array is treated with a particular nucleotide type (e.g. A, T, C or G). Images obtained after addition of each nucleotide type will differ with regard to which features in the array are detected. These differences in the image
reflect the different sequence content of the features on the array. The relative locations of each feature, however, will remain unchanged in the images. The images can be stored, processed and analyzed. In some embodiments, images obtained after treatment of the array with each different nucleotide type can be handled in the same way as for images obtained from different detection channels for reversible terminator-based sequencing methods.
[00169] In some embodiment, SBS methods involve detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are commercially available from Thermo Fisher Scientific as the Ion Torrent instrument or sequencing methods and systems described in US 8262900A1, incorporated herein by reference. Methods set forth herein for amplifying target nucleic acids using kinetic exclusion can be readily applied to substrates used for detecting protons. More specifically, methods set forth herein can be used to produce clonal populations of amplicons that are used to detect protons.
[00170] In some embodiments, SBS involves cycle sequencing by the stepwise addition of reversible terminator nucleotides which comprise a cleavable or photobleachable dye label (as described in U.S. Pat. No. 7,057,026, the disclosure of which is incorporated herein by reference). This approach is commercialized Illumina Inc., and is also described in WO 91/06678 (also referred to as US 427,321, Tsien et al.) and U.S. Patent No. 8241573, each of which is incorporated herein by reference. The availability of fluorescently-labeled terminators in which both the termination can be reversed and the fluorescent label cleaved facilitates efficient cyclic reversible termination sequencing. Polymerases can also be engineered to efficiently incorporate and extend from these modified nucleotides. Additional exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Patent Nos. 7541444, 7566537, 7057026, 8460910, 8623628, 8951781, and 9193996, the disclosures of which are incorporated herein by reference in their entireties.
[00171] For cluster generation, the library fragments are immobilized on a substrate, for example a slide, which comprises homologous oligonucleotide sequences for capturing and immobilizing the DNA library fragments. The immobilized DNA library fragments are amplified using cluster amplification methodologies as exemplified by the disclosures of U.S. Pat. Nos. 7985565 and 7115400, the contents of each of which is incorporated herein by reference in its entirety. The incorporated materials of U.S. Pat. Nos. 7985565 and 7115400 describe methods of solid-phase nucleic acid amplification which allow amplification products to be immobilized on a solid support in order to form arrays comprised of clusters or “colonies”
of immobilized nucleic acid molecules. Each cluster or colony on such an array is formed from a plurality of identical immobilized polynucleotide strands and a plurality of identical immobilized complementary polynucleotide strands. The arrays so-formed are generally referred to as “clustered arrays”. The products of solid-phase amplification reactions such as those described in U.S. Pat. Nos. 7985565 and 7115400 are so-called “bridged” structures formed by annealing of pairs of immobilized polynucleotide strands and immobilized complementary strands, both strands being immobilized on the solid support at the 5' end, preferably via a covalent attachment. Cluster amplification methodologies are examples of methods wherein an immobilized nucleic acid template is used to produce immobilized amplicons. Other suitable methodologies can also be used to produce immobilized amplicons from immobilized DNA fragments produced according to the methods provided herein. For example, one or more clusters or colonies can be formed via solid-phase PCR whether one or both primers of each pair of amplification primers are immobilized. However, the methods described herein are not limited to any particular sequencing preparation methodology or sequencing platform and can be amenable to other parallel sequencing platform preparation methods and associated sequencing platforms.
[00172] In some embodiments, SBS methods involve detection of four different nucleotides using fewer than four different labels. SBS can be performed utilizing methods and systems described in the incorporated materials of U.S. Patent No. 9453258. In some embodiments, a pair of nucleotide types can be detected at the same wavelength, but distinguished based on a difference in intensity for one member of the pair compared to the other, or based on a change to one member of the pair (e.g. via chemical modification, photochemical modification or physical modification) that causes apparent signal to appear or disappear compared to the signal detected for the other member of the pair. In some embodiments, three of four different nucleotide types can be detected under particular conditions while a fourth nucleotide type lacks a label that is detectable under those conditions, or is minimally detected under those conditions (e.g., due to background fluorescence). Incorporation of the first three nucleotide types into a nucleic acid can be determined based on presence of their respective signals and incorporation of the fourth nucleotide type into the nucleic acid can be determined based on absence or minimal detection of any signal. In some embodiments, one nucleotide type can include label(s) that are detected in two different channels, whereas other nucleotide types are detected in no more than one of the channels. The aforementioned three exemplary configurations are not considered mutually exclusive and can be used in various combinations. In some embodiments, SBS methods can involve a first nucleotide type that is detected in a
first channel (e.g. dATP having a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g. dCTP having a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and the second channel (e.g. dTTP having at least one label that is detected in both channels when excited by the first and/or second excitation wavelength) and a fourth nucleotide type that lacks a label that is not, or minimally, detected in either channel (e.g. dGTP having no label).
[00173] Further, as described in the incorporated materials of U.S. Patent No. 9453258, sequencing data can be obtained using a single channel. In such so-called one-dye sequencing approaches, the first nucleotide type is labeled but the label is removed after the first image is generated, and the second nucleotide type is labeled only after a first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.
[00174] In some embodiments, sequencing can be performed by methods capable of single molecule sequencing, with or without amplification or labeling, such as nanopore sequencing (as described in: Deamer, D. W. & Akeson, M. “Nanopores and nucleic acids: prospects for ultrarapid sequencing.” Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, “Characterization of nucleic acids by nanopore analysis”. Ace. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A Golovchenko, “DNA molecules and configurations in a solid-state nanopore microscope” Nat. Mater. 2:611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, the target nucleic acid passes through a nanopore. The nanopore can be a synthetic pore or biological membrane protein, such as alpha-hemolysin. Each base-pair can be identified by measuring fluctuations in the electrical conductance of the pore as the target nucleic acid passes through the nanopore. (U.S. Pat. No. 7001792; Soni, G. V. & Meller, “A Progress toward ultrafast DNA sequencing using solid-state nanopores.” Clin. Chem. 53, 1996- 2001 (2007); Healy, K. “Nanopore-based single-molecule DNA analysis.” Nanomed. 2, 459- 481 (2007); Cockroft, S. L., Chu, J., Amorin, M. & Ghadiri, M. R. “A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution.” J. Am. Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties).
[00175] In some embodiments, sequencing can be performed by methods comprising the real-time monitoring of DNA polymerase activity. Nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-
bearing polymerase and gamma-phosphate-labeled nucleotides as described, for example, in U.S. Pat. No. 7329492 (which is incorporated herein by reference) or nucleotide incorporations can be detected with zero-mode waveguides as described, for example, in U.S. Pat. No. 7315019 (which is incorporated herein by reference) and using fluorescent nucleotide analogs and engineered polymerases as described, for example, in U.S. Pat. No. 7405281 and U.S. Patent No. 8343746 (each of which is incorporated herein by reference). The illumination can be restricted to a zeptoliter-scale volume around a surface-tethered polymerase such that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M. J. et al. “Zero-mode waveguides for single-molecule analysis at high concentrations.” Science 299, 682-686 (2003); Lundquist, P. M. et al. “Parallel confocal detection of single molecules in real time.” Opt. Lett. 33, 1026-1028 (2008); the disclosures of which are incorporated herein by reference in their entireties). Commercial systems employing such technologies are presently commercialized by PacBio (Menlo Park, CA, USA).
[00176] The sequencing methods described herein can be advantageously carried out in multiplex formats such that multiple different target nucleic acids are manipulated simultaneously. In some embodiments, different target nucleic acids can be treated in a common reaction vessel or on a surface of a particular substrate, allowing for the convenient delivery of sequencing reagents, removal of unreacted reagents and detection of incorporation events in a multiplexed manner. In embodiments involving surface-bound (or bound through a hydroglel to a surface) target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids can be typically bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent attachment, attachment to a bead or other particle or binding to a polymerase or other molecule that is attached to the surface. In some embodiments, the target nucleic acids can be bound through a surface-immobilized hydrogel (as described in U.S. Patent No. 9012022, incorporated herein by reference). The array can include a single copy of a target nucleic acid at each site (also referred to as a feature) or multiple copies having the same sequence can be present at each site or feature (also referred to as a colony). Multiple copies can be produced by amplification methods which can include or exclude bridge amplification or emulsion PCR as described in further detail below.
[00177] The methods of the present disclosure may utilize the Illumina, Inc. SBS technology for sequencing the DNA profile libraries created by practicing the methods described herein. The Novaseq 6000 platform (a sequencing instrument) was used for clustering and sequencing
for the examples described herein. However, as understood by a skilled artisan, the present methods are not limited by the type of sequencing platform used.
[00178] In other embodiments, sequencing can be performed using long-read sequencing technologies, such as those developed by Pacific Biosciences (e.g., SMRT sequencing) or Oxford Nanopore Technologies (e.g., nanopore sequencing). Long-read sequencing can be advantageous for resolving complex structural variants, phasing mutations over long distances, distinguishing highly homologous gene regions or pseudogenes, and characterizing full-length RNA transcripts including isoforms and fusion genes, especially from nucleic acids derived from the enriched myeloid cell population.
[00179] In some embodiments, sequencing libraries are prepared prior to sequencing. Sequencing library preparation involves the production of a random collection of adapter- modified DNA fragments, which are ready to be sequenced. Sequencing libraries of polynucleotides can be prepared from DNA or RNA, including equivalents, analogs of either DNA or cDNA, that is complementary or copy DNA produced from an RNA template, for example by the action of reverse transcriptase. The polynucleotides may originate in doublestranded DNA (dsDNA) form (e.g. genomic DNA fragments, PCR and amplification products) or polynucleotides that may have originated in single-stranded form, as DNA or RNA, and been converted to dsDNA form. By way of example, mRNA molecules may be copied into double-stranded cDNAs suitable for use in preparing a sequencing library. The precise sequence of the primary polynucleotide molecules is generally not material to the method of library preparation, and may be known or unknown. Preparation of sequencing libraries for some sequencing platforms may require that the polynucleotides be or a specific range of fragment sizes e.g. 0-1200 bp. Therefore, fragmentation of polynucleotides e.g. genomic DNA may be required.
[00180] Standard protocols e.g. protocols for sequencing using, for example, the Illumina platforms, instruct users to purify the end-repaired products prior to dA-tailing, and to purify the dA-tailing products prior to the adapter-ligating steps of the library preparation. Purification of the end-repaired products and dA-tailed products remove enzymes, buffers, salts and the like to provide favorable reaction conditions for the subsequent enzymatic step. In one embodiment, the steps of end-repairing, dA-tailing and adapter ligating exclude the purification steps. Thus, in one embodiment, the method disclosed herein encompasses preparing a sequencing library that comprises the consecutive steps of end-repairing, dA-tailing and adapter-ligating. In some embodiments, dA-tailing is not performed. In some embodiments, adapters are added via blunt-end ligation to one or both strands of a target sequence.
[00181] As part of the sequencing protocol, an amplification reaction is prepared. The amplification step introduces to the adapter ligated template molecules the oligonucleotide sequences required for hybridization to the flow cell. The contents of an amplification reaction include appropriate substrates (such as dNTPs), enzymes (e.g. a DNA polymerase) and buffer components required for an amplification reaction. Optionally, amplification of adapter-ligated polynucleotides can be omitted. Generally, amplification reactions require at least two amplification primers i.e. primer oligonucleotides, which may be identical, and include an adapter-specific portion, capable of annealing to a primer-binding sequence in the polynucleotide molecule to be amplified (or the complement thereof if the template is viewed as a single strand) during the annealing step. Once formed, the library or templates prepared according to the methods described above can be used for solid-phase nucleic acid amplification. The term ‘solid-phase amplification’ as used herein refers to any nucleic acid amplification reaction carried out on or in association with a solid support such that all or a portion of the amplified products are immobilized on the solid support as they are formed. In particular, the term encompasses solid-phase polymerase chain reaction (solid-phase PCR) and solid phase isothermal amplification which are reactions analogous to standard solution phase amplification, except that one or both of the forward and reverse amplification primers is/are immobilized on the solid support. Solid phase PCR covers systems such as emulsions, wherein one primer is anchored to a bead and the other is in free solution, and colony formation in solid phase gel matrices wherein one primer is anchored to the surface, and one is in free solution. The library of template polynucleotide can be used in sequencing methods. In addition to providing templates for solid-phase sequencing and solid-phase PCR, library templates provide templates for whole genome amplification.
[00182] Sequencing of the amplified libraries can be carried out using any suitable sequencing technique as described herein. In some embodiments, sequencing is performed using sequencing-by-synthesis (SBS), or single molecule sequencing such as nanopore sequencing.
[00183] In some embodiments, this disclosure provides for methods which employ a sequencing technology in which clonally amplified DNA templates or single DNA molecules are sequenced within a flow cell.
[00184] In one embodiment, the method employs sequencing of DNA fragments using Illumina's sequencing-by-synthesis (SBS) and reversible terminator-based sequencing chemistry, as described herein. In some embodiments, template DNA can be genomic DNA e.g. ctDNA. In some embodiments, genomic DNA from isolated cells is used as the template,
and is fragmented into lengths of several hundred base pairs. Illumina's sequencing technology involves the attachment of fragmented genomic DNA to a planar, optically transparent surface on which oligonucleotide anchors are bound to a surface-immobilized hydrogel polymer. Template DNA is end-repaired to generate 5 ’-phosphorylated blunt ends, and the polymerase activity of KI enow fragment is used to add a single A base to the 3’ end of the blunt phosphorylated DNA fragments. The addition prepares the DNA fragments for ligation to oligonucleotide adapters, which have an overhang of a single T base at their 3’ end to increase ligation efficiency. The adapter oligonucleotides are complementary to the flow-cell anchors. Under limiting-dilution conditions, adapter-modified, single-stranded template DNA is added to the flow cell and immobilized by hybridization to the anchors. Attached DNA fragments are extended and bridge amplified to create an ultra-high density sequencing flow cell with hundreds of millions of clusters, each containing -1,000 copies of the same template.
[00185] In one embodiment, the present disclosure provides a duplex sequencing method for detecting leukemic molecular alterations in isolated myeloid nucleic acids (IMNA) or nucleic acids processed therefrom (NAPT). The method involves producing a plurality of IMNA adapters or NAPT adapters, each comprising a Watson strand and a Crick strand. The Watson strand of each adapter includes the following elements in a 5' to 3' orientation:
(a) a first Watson single-stranded strand-distinguishing sequence that includes at least one universal primer binding site;
(b) a first Watson unique molecular identifier (UMI) sequence;
(c) a Watson strand of a target sequence;
(d) a second Watson UMI sequence;
(e) a second Watson single-stranded strand-distinguishing sequence that includes at least one universal primer binding site, which is different from the universal primer binding site in the first Watson single-stranded strand-distinguishing sequence.
[00186] The Crick strand of each adapter can include the following elements in a 5' to 3' orientation:
(a) a first Crick single-stranded strand-distinguishing sequence that is identical to the first Watson single-stranded strand-distinguishing sequence;
(b) a first Crick UMI sequence that is complementary to the second Watson UMI sequence;
(c) a Crick strand of the target sequence that is complementary to the Watson strand of the target sequence;
(d) a second Crick UMI sequence that is complementary to the first Watson UMI sequence;
(e) a second Crick single-stranded strand-distinguishing sequence that is identical to the second Watson single-stranded strand-distinguishing sequence.
[00187] The strand-distinguishing sequences allow for the independent amplification and sequencing of the Watson and Crick strands of each adapter. The UMI sequences enable the identification and consensus calling of sequence reads originating from the same starting molecule, which is essential for error correction.
[00188] After producing the IMNA or NAPT adapters, the method can involve amplifying all or a portion of the Watson and Crick strands using at least one strand-distinguishing sequence that includes at least one universal primer. This generates Watson and Crick amplicons that can be further analyzed.
[00189] In some embodiments, the method includes an optional step of performing either (a) one or more targeted amplification steps on the Watson and Crick amplicons or their amplified products to generate targeted amplicons, or (b) bait hybridization on the Watson and Crick amplicons to generate selected amplicons. Targeted amplification allows for the enrichment of specific regions of interest, while bait hybridization enables the capture of sequences that are complementary to the bait probes.
[00190] The Watson and Crick amplicons, whether targeted or selected, are then analyzed by next-generation sequencing (NGS) to obtain sequence reads. The sequence reads are mapped to the reference sequence of the target region and grouped based on their unique molecular identifiers. Sequence reads with the same UMI are considered to have originated from the same starting molecule and are used to generate a consensus sequence.
[00191] A key step in the duplex sequencing method is the identification of sequence variants that are not present in all sequence reads from the Watson and Crick amplicons derived from the same original molecule. Such variants are likely to be artifacts introduced during amplification or sequencing, while true variants should be present in both the Watson and Crick strands. By comparing the sequence reads from the Watson and Crick amplicons with the same UMI, errors can be identified and removed, increasing the accuracy of variant calling.
[00192] In some embodiments, the method further includes identifying variants that are not present in all sequence reads from any one target sequence or selected sequence in Watson amplicons derived from any one Watson IMNA or NAPT and its corresponding Crick amplicons derived from any one Crick IMNA or NAPT. This allows for the identification of errors that may be specific to certain target sequences or selected regions.
[00193] The duplex sequencing method described herein can be used to detect a wide range of leukemic molecular alterations, including single nucleotide variants, insertions, deletions, and structural rearrangements. The high accuracy of the method enables the detection of low- frequency variants that may be missed by conventional sequencing approaches.
[00194] The duplex sequencing method may involve several steps, including: (i) producing IMNA or NAPT adapters with specific configurations of Mis, e.g, non-unique Mis or UMIs and strand-distinguishing sequences on the Watson and Crick strands, (ii) amplifying the adapters to obtain Watson and Crick amplicons, (iii) optionally performing targeted amplification or bait hybridization on the amplicons, and (iv) analyzing the amplicons by NGS to obtain sequence reads. In some embodiments, an additional step of (v) identifying variants that are not present in all sequence reads from the Watson and Crick amplicons derived from the same original molecule is performed. This step helps to filter out errors and identify true variants based on their presence on both strands of the original DNA duplex.
[00195] In one embodiment, the duplex sequencing method is used to detect minimal residual disease (MRD) in patients with myeloid malignancies. MRD refers to the presence of residual leukemic cells below the limit of detection of conventional morphologic or cytogenetic methods. By analyzing IMNA or NAPT from patient samples using duplex sequencing, MRD can be detected with high sensitivity, allowing for early intervention and personalized treatment decisions.
[00196] In another embodiment, the duplex sequencing method is used to monitor clonal evolution and the emergence of therapy-resistant subclones in myeloid malignancies. By performing serial analyses of IMNA or NAPT from patient samples before, during, and after treatment, changes in the genetic landscape of the leukemic cells can be tracked over time. This can provide insights into the mechanisms of drug resistance and guide the selection of alternative therapies.
[00197] In some embodiments, the duplex sequencing method is used to identify novel driver mutations or cooperating mutations in myeloid malignancies. By analyzing IMNA or NAPT from a large cohort of patients, recurrent mutations that are not detected by conventional sequencing methods may be identified. These mutations may represent novel therapeutic targets or prognostic biomarkers.
[00198] In other embodiments, the duplex sequencing method is used to characterize the clonal architecture and evolution of myeloid malignancies. By analyzing IMNA or NAPT from different cell populations or at different time points, the relative abundance and genetic
composition of different subclones can be determined. This can provide insights into the order of acquisition of mutations and the evolutionary trajectories of the leukemic cells.
[00199] In one embodiment, the strand-distinguishing sequences on the Watson and Crick strands are of different lengths. For example, the strand-distinguishing sequence on the Watson strand may be longer or shorter than the strand-distinguishing sequence on the Crick strand. This length difference can provide an additional level of strand discrimination during amplification and sequencing.
[00200] In another embodiment, the strand-distinguishing sequences on the Watson and Crick strands have different base compositions. For example, the strand-distinguishing sequence on the Watson strand may be rich in G/C bases, while the strand-distinguishing sequence on the Crick strand may be rich in A/T bases. This difference in base composition can affect the melting temperature and hybridization specificity of the sequences, allowing for selective amplification or sequencing of one strand over the other.
[00201] In yet another embodiment, the strand-distinguishing sequences on the Watson and Crick strands are artificial sequences that are not derived from or homologous to any naturally occurring genomic sequences. This can minimize the risk of non-specific hybridization or amplification of unintended targets.
[00202] In some embodiments, the strand-distinguishing sequences on the Watson and Crick strands are selected to have minimal self-complementarity or secondary structure. This can improve the efficiency and specificity of primer binding and extension during amplification and sequencing.
[00203] In certain embodiments, the strand-distinguishing sequences on the Watson and Crick strands are designed to be compatible with specific sequencing platforms or chemistries. For example, the sequences may be selected to optimize cluster generation, primer hybridization, or nucleotide incorporation on a particular sequencing instrument.
[00204] In one embodiment, the strand-distinguishing sequences on the Watson and Crick strands are used as binding sites for strand-specific sequencing primers. By using different primers for the Watson and Crick strands, the sequences of the two strands can be determined independently and used for error correction or haplotype phasing.
[00205] In another embodiment, the strand-distinguishing sequences on the Watson and Crick strands are used as binding sites for strand-specific probes or baits. This can allow for the selective capture or enrichment of one strand over the other, which can be useful for applications such as targeted sequencing or strand-specific gene expression analysis.
[00206] In some embodiments, the strand-distinguishing sequences on the Watson and Crick strands are used as primer binding sites for strand-specific PCR. By using different primers for the Watson and Crick strands, the two strands can be amplified separately and used for downstream applications such as sequencing or cloning.
[00207] In certain embodiments, the strand-distinguishing sequences on the Watson and Crick strands are used as barcodes for strand-specific labeling or detection. For example, the sequences may be conjugated to different fluorescent dyes or affinity tags that allow for the visual differentiation or physical separation of the two strands.
[00208] In some embodiments, the methods comprise generating a library from the IMNA or NAPT in the biological sample, e.g., a library of IMNA or NAPT, adapter-IMNA or NAPT and/or enriched IMNA or NAPT or enriched adapter-IMNA or NAPT, e.g., bait enriched or amplification products with or without adapters. In some embodiments, the methods comprise determining the sequences of the barcode library. In some embodiments, the barcode sample is obtained or derived from a sample obtained from a human.
[00209] In some embodiments, each of the plurality of primers comprises one or more barcode sequences. In some embodiments, the one or more barcode sequences comprise a primer barcode, a capture barcode, a sequencing barcode, a unique molecular identifier barcode, or a combination thereof. In some embodiments, the one or more barcode sequences comprise a primer barcode. In some embodiments, the one or more barcode sequences comprise a unique molecular identifier (UMI) barcode.
[00210] In some embodiments, nucleic acid sequencing may comprise sequencing at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300 or more nucleotides or base pairs of the nucleic acid molecule sequences. In some embodiments, sequencing may comprise sequencing at least about 200, 300, 400, 500, 600, 700, 800, 900, 1,000 or more nucleotides or base pairs of the nucleic acid molecule sequences. In other embodiments, sequencing may comprise sequencing at least about 1,500; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; 10,000; 50,000 or more nucleotides or base pairs of the nucleic acid molecule sequences.
[00211] In some embodiments, nucleic acid sequencing may comprise at least about 200, 300, 400, 500, 600, 700, 800, 900, 1,000 or more sequencing reads per run. In some embodiments, sequencing may comprise sequencing at least about 1,500; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; or 10,000 or more sequencing reads per run. In some embodiments, nucleic acid sequencing may comprise at least about 10,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; or 100,000 or more sequencing reads per run.
In some embodiments, nucleic acid sequencing may comprise at least about 250,000; 500,000; 1,000,000; 10,000,000; 100,000,000; or 1,000,000,000 or more sequencing reads per run. In some embodiments, nucleic acid sequencing may comprise less than or equal to about 1,600,000,000 sequencing reads per run. In some embodiments, nucleic acid sequencing may comprise less than or equal to about 200,000,000 reads per run.
[00212] The number of times a single nucleotide or polynucleotide is identified or “read” is defined as the sequencing depth or read depth, or fold coverage, optionally describing a percentage of bases. Read depth (sequencing depth, or sampling) represents the total number of times a sequenced nucleic acid fragment (a “read”) is obtained for a sequence. Theoretical read depth is defined as the expected number of times the same nucleotide is read, assuming reads are perfectly distributed throughout an idealized genome. Read depth is expressed as function of percentage coverage (or coverage breadth). For example, 10 million reads of a 1 million base genome, perfectly distributed, theoretically results in 10X read depth of 100% of the sequences. In practice, a greater number of reads (higher theoretical read depth, or oversampling) may be needed to obtain the desired read depth for a percentage of the target sequences.
[00213] In some embodiments, sequencing is peformed at a read depth of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550. 600, 650, 700, 750, 800, 850, 900, 950, or at least 1000, at least 10,000, at least 20,000, at least 30,000, at least 50,0000, at least 75,000, or least 100,000 unique reads per base. In some embodiments, sequencing may comprise a read depth of at least IX, 5X, 10X, 20X, 30X, 40X, 50X, 60X, 70X, 80X, 90X, 100X, or more.
[00214] Enrichment of target sequences with a controlled stoichiometry probe library increases the efficiency of downstream sequencing, as fewer total reads will be required to obtain an outcome with an acceptable number of reads over a desired % of target sequences. For example, in some instances, 55X theoretical read depth of target sequences results in at least 3 OX coverage of at least 90% of the sequences. In some instances, no more than 55X theoretical read depth of target sequences results in at least 3 OX read depth of at least 80% of the sequences. In some instances, no more than 55X theoretical read depth of target sequences results in at least 3 OX read depth of at least 95% of the sequences. In some instances, no more than 55X theoretical read depth of target sequences results in at least 10x read depth of at least 98% of the sequences. In some instances, 55X theoretical read depth of target sequences results in at least 20X read depth of at least 98% of the sequences. In some instances, no more than 55X theoretical read depth of target sequences results in at least 5X read depth of at least 98%
of the sequences. Increasing the concentration of probes during hybridization with targets can lead to an increase in read depth. In some instances, the concentration of probes is increased by at least 1.5X, 2. OX, 2.5X, 3X, 3.5X, 4X, 5X, or more than 5X. In some instances, increasing the probe concentration results in at least a 1000% increase, or a 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 750%, 1000%, or more than a 1000% increase in read depth. In some instances, increasing the probe concentration by 3X results in a 1000% increase in read depth.
[00215] “Homology” refers to the percent identity between two polynucleotides or two polypeptide sequences. Two DNA or polypeptide sequences are “homologous” to each other when the sequences exhibit at least about 75% to 85% (including 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, and 85%), at least about 90%, or at least about 95% to 99% (including 95%, 96%, 97%, 98%, 99%) contiguous sequence identity over a defined length of the sequences.
[00216] In some embodiments, RAPGEF3 gene amplicons are obtained by using a forward primer having 80% homology or higher to that of the nucleotide sequence 5’ CTTCCTTCATTTCTCCACCTG 3’ (SEQ ID NO: 1) and the reverse primer having 80% homology or higher to that of the nucleotide sequence 5’ TCTGTGTCCTCTTGCCTGC 3’ (SEQ ID NO: 2).
[00217] Identity or homology with respect to a specified amino acid sequence of this invention is defined herein as the percentage of amino acid residues in a candidate sequence that are identical with the specified residues, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent homology, and not considering any conservative substitutions as part of the sequence identity. None of N-terminal, C-terminal or internal extensions, deletions, or insertions into the specified sequence shall be construed as affecting homology. All sequence alignments called for herein are such maximal homology alignments. Generally, the nucleic acid sequence homology between the polynucleotides, oligonucleotides, and fragments disclosed herein and a nucleic acid sequence of interest will be at least 80% or greater, and more typically with preferably increasing homologies of at least 85%, 90%, 91%, 92%, 92%, 94%, 95%, 96%, 97%, 98%, 99%, and/or 100%. Two amino acid sequences are homologous if there is a partial or complete identity between their sequences.
[00218] In some embodiments, the practice of the present disclosure will employ, together with the methods featured herein, unless otherwise indicated molecular biology, microbiology, recombinant DNA, and immunology techniques. See, e.g., Molecular Cloning A Laboratory Manual, 2nd Ed., ed. by Sambrook, Fritsch and Maniatis (Cold Spring Harbor Laboratory
Press, 1989); DNA Cloning, Volumes I and II (D. N. Glover ed., 1985); Culture Of Animal Cells (R. I. Freshney, Alan R. Liss, Inc., 1987); Immobilized Cells And Enzymes (IRL Press, 1986); B. Perbal, A Practical Guide To Molecular Cloning (1984); the treatise, Methods In Enzymology (Academic Press, Inc., N.Y.); Gene Transfer Vectors For Mammalian Cells (J. H. Miller and M. P. Calos eds., 1987, Cold Spring Harbor Laboratory); Methods In Enzymology, Vols. 154 and 155 (Wu et al. eds.), Immunochemical Methods In Cell And Molecular Biology (Mayer and Walker, eds., Academic Press, London, 1987); Antibodies: A Laboratory Manual, by Harlow and Lane s (Cold Spring Harbor Laboratory Press, 1988); and Handbook Of Experimental Immunology, Volumes I-FV (D. M. Weir and C. C. Blackwell, eds., 1986). [00219] Polymorphic Sequences
[00220] Polymorphic sites that are contained in the target nucleic acids include without limitation single nucleotide polymorphisms (SNPs), tandem SNPs, small-scale multi-base deletions or insertions, also referred to as “IN-DELS” or deletion insertion polymorphisms “DIPs”, Multi -Nucleotide Polymorphisms “MNPs”, and Short Tandem Repeats “STRs”.
[00221] In one embodiment, the nucleic acids in the sample is enriched for target nucleic acids that comprise at least one SNP. In some embodiments, each target nucleic acid comprises a single i.e. one SNP. Target nucleic acid sequences comprising SNPs are available from publically accessible databases including, but not limited to Human SNP Database at world wide web address wi.mit.edu, NCBI dbSNP Home Page at world wide web address ncbi.nlm.nih.gov, world wide web address lifesciences.perkinelmer.com, Celera Human SNP database at world wide web address celera.com, the SNP Database of the Genome Anulysis Group (GAN) at world wide web address gan.iarc.fr. In one embodiment, the SNPs chosen selected from the group of 92 individual identification SNPs (IISNPs) described by Pakstis el al. (Pakstis et el. Hum Genet 127:315-324 [2010]), which have been shown to have a very small variation in frequency across populations (Fst <0.06), and to be highly in formative around the world having an average heterozygosity >0.4. SNPs that are encompassed by the method disclosed herein include linked and unlinked SNPs. Each target nucleic acid comprises at least one polymorphic site e.g. a single SNP, that differs from that present on another target nucleic acid to generate a panel of polymorphic sites e.g. SNPs, that contain a sufficient number of polymorphic sites of which at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, a.t least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40 or more are informative. For example, a panel of SNPs can be configured, to comprise at. least one informative SNP.
[00222] Enrichment of the sample for the target nucleic acids is accomplished by methods that comprise specifically amplifying target nucleic acid sequences that comprise the polymorphic site. Amplification of the target sequences can be performed by any method that uses PCR or variations of the method including but not limited to asymmetric PCR, helicasedependent amplification, hot-start PCR, qPCR, solid phase PCR, and touchdown PCR. Alternatively, replication of target nucleic acid sequences can be obtained by enzymeindependent methods e.g. chemical solid-phase synthesis using the phosphoramidites. Amplification of the target sequences is accomplished using primer pairs each capable of amplifying a target nucleic acid sequence comprising the polymorphic site e.g. SNP, in a multiplex PCR reaction. Multiplex PCR reactions include combining at least 2, at least three, at least 3, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30 at least 30, at least 35, at least 40 or more sets of primers in the same reaction to quantify the amplified target nucleic acids comprising at least two, at least three, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 30, at least 35, at least 40 or more polymorphic sites in the same sequencing reaction. Any panel of primer sets can be configured to amplify- at least one informative polymorphic sequence.
[00223] Amplification of polymorphic sequences
[00224] Amplification of the target nucleic acids is performed using sequence-specific primers that allow for sequence specific amplification. For example, the PCR primers are designed to discriminate against the amplification of similar genes or paralogs that are on other chromosomes by taking advantage of sequence differences between the target nucleic acid and any paralogs from other chromosomes. The forward or reverse PCR primers are designed to anneal close to the SNP site and to amplify a nucleic acid sequence of sufficient length to be encompassed in the reads generated by sequencing methods. Some sequencing methods require that nucleic: acid sequence have a minimum length (bp) to enable bridging amplification that may optionally be used prior to sequencing. Thus the PCR primers used for amplifying target nucleic acids are designed to amplify sequences that are of sufficient length to the bridge amplified and to identify SNPs that are encompassed by the sequence reads. In some embodiments, the first of two primers in the primer set comprising the forward and the reverse primer for amplifying the target nucleic acid is designed to identify a single SNP present within a sequence read of about 20bp, about 25bp, about 30bp, about 35bp, about 4Ubp, about 45bp, about 50bp, about 55bp, about 60bp, about 65bp, about 70bp, about 75bp, about 80bp, about 85bp, about90bp, about 95bp, about lOObp, about 1 lObp, about 120bp, about 130, about 140bp, about 150bp, about 200bp, about 250bp, about 300bp, about 350bp, about 400bp, about 450bp,
or about 500bp. It is expected that technological advances in sequencing technologies will enable single-end reads of greater than 500bp. In one embodiment, one of the PCR primers is designed to amplify SNPs that are encompassed in sequence reads of 36 bp. The second primer is designed to amplify the target nucleic acid as an amplicon of sufficient length to allow for bridge amplification. In one embodiment, the exemplary PCR primers are designed to amplify target nucleic acids that contain a single SNP selected from SNPs rs8082254, rs72847785, rs3181254, rsl7256902, rs2240079, rsl 1168215, rs61917617, rsl 1168214, rs55683248, and rs2072341. In other embodiments, the forward and reverse primers are each designed for amplifying target nucleic acids each comprising a set of two tandem SNPs, each being presenl within a sequence read or about 20bp, about 25bp, about 30bp, about 35bp, about 40bp, about 45bp, about 50bp, about 55bp, about 60bp, about 65bp, about 70bp, about 75bp, about 80bp, about 85bp, about 90bp, about 95bp, about lOObp, about HObp, about 120bp, about 130, about 140bp, about 150bp, about 200bp, about 250bp, about 300bp, about 350bp, about 400bp, about 450bp, or about 500bp. In one embodiment, at least one of the primers is designed to amplify the target nucleic acid comprising a set of two tandem SNPs as an amplicon of sufficient length to allow for bridge amplification.
[00225] The SNPs, single or tandem SNPs, are contained in amplified target nucleic acid amplicons of at least about lOObp, at least about 150bp, at least about 200bp, at least about 250bp, at least about 300bp, at least about 350bp, or at least about 400bp. In one embodiment, target nucleic acids comprising a polymorphic site e.g. a SNP, are amplified as amplicons of at least about 110 bp, and that comprise a SNP within 36 bp from the 3’ or 5’ end of the amplicon. In another embodiment, target nucleic acids comprising two or more polymorphic sites e.g. two tandem SNPs, are amplified as amplicons of at least about 110 bp, and that comprise the first SNP within 36 bp from the 3’ end of the amplicon, and/or the second SNP within 36 bp from the 5’ end of the amplicon.
[00226] Methods of Treatment
[00227] The present disclosure also relates to methods of treating a subject having or suspected of having a myeloid malignancy. In some embodiments, this disclosure relates to methods of treating a subject, e.g., patient having myeloid malignancy.
[00228] As used herein, a modulator is a therapeutic agent that alters the expression, level, and/or activity of a gene or gene product (including protein or RNA, e.g., mRNA, miRNA, or piRNA). In embodiments, a modulator is administered to a subject having a gene variant, to alter a gene or gene product to similar levels as unaffected subjects who do not have the variant, or to provide similar levels of a gene or gene product as subjects carrying protective gene
variants against developing a myeloid malignancy. In embodiments, a modulator is administered to a subject to enhance the expression, level, and/or activity of a gene or gene product in subjects. In embodiments, a modulator is administered to a subject to restore the expression, level and/or activity of a gene or gene product to baseline or to substantially the same as in a subject carrying protective gene variants, for example to expression, level and/or activity of a non-carrier. In some embodiments, a modulator is administered to a subject to reduce the expression, level and/or activity of a gene or gene product in subjects. In some embodiments, the modulator alters gene expression, e.g. by affecting a transcriptional regulator (inhibitor or activator) or inhibitory RNA. In some embodiments, a modulator alters the gene or gene product levels, e.g. by affecting the RNA product of a gene and affecting its stability or translation. In some embodiments, a modulator alters the activity of the gene or gene product, e.g. by affecting the protein product directly, activating or inhibiting the protein, or enhancing or preventing its multimerization or binding capacity, which can be through competitive, noncompetitive, or uncompetitive interactions.
[00229] In some embodiments, a modulator is an antagonist. As used herein, an antagonist is a therapeutic agent that interferes with or inhibits the physiological action, e.g., expression, level and/or activity of a target, or that inhibits or interferes with the physiological action of a positive-regulator gene or gene product (that increases the expression, level and/or activity of the target gene or gene product or decreases the expression, level and/or activity of an activator of the target). In some embodiments, the antagonist can inhibit the gene or gene product directly, or activate an inhibitor of the gene or gene product, or inhibit an activator of the gene to reduce the expression, level, and/or activity of the target gene or gene product. In some embodiments, the antagonist alters expression of the gene, e.g. by inhibiting an activator of the target, or enhancing or stabilizing an inhibitor of the target. In some embodiments, the antagonist alters the gene or gene product levels by destabilizing the RNA of the gene or inhibiting its translation. In some embodiments, the antagonist alters the activity of the gene or gene product, for example by directly inhibiting the protein’s activity, binding, or multimerization, or by competitively, non-competitvely, or uncompetitively binding the protein to decrease its activity. As used herein, the term “antagonist” also refers to enzyme inhibitors.
[00230] In some embodiments, a modulator is an agonist. As used herein, an agonist is a therapeutic agent that activates or enhances the the physiological action, e.g., expression, level and/or activity of a target, or that activates a positive regulator of the target (a gene or gene product that increases the expression, level and/or activity of a target or that interferes with or
inhibits a negative regulator of the target). In some embodiments, the agonist can activate the gene or gene product directly, or inhibit an inhibitor of the gene or gene product, or activate or upregulate an activator of the gene to increase the expression, level, and/or activity of the target gene or gene product. In some embodiments, the agonist alters expression of the gene, e.g. by enhancing an activator of the target, or inhibiting or destabilizing an inhibitor of the target. In some embodiments, the agonist alters the gene or gene product levels by stabilizing the RNA of the gene or enhancing its translation. In some embodiments, the agonist alters the activity of the gene or gene product, for example by directly activating the protein’s activity, binding, or multimerization, or by competitively, non-competitvely, or uncompetitively binding the protein to increase its activity. As used herein, the term “agonist” also refers to enzyme activators.
[00231] In some embodiments, the method of treating a subject having or suspected of having a myeloid malignancy further comprises administering one or more additional active agents or supportive therapies for treating, preventing, or reducing the severity of a myeloid malignancy to the subject.
[00232] In one embodiment, the present disclosure provides a method for treating a subject diagnosed with a myeloid malignancy. The method involves administering to the subject a therapeutically effective amount of a compound identified using the screening methods described herein. The compound may be administered alone or in combination with other therapeutic agents, such as chemotherapy, targeted therapy, or immunotherapy.
[00233] In another embodiment, the treatment method involves administering to the subject a compound that targets a specific molecular alteration identified in the subject's myeloid cells using the detection methods described herein. For example, if the subject is found to have a particular gene mutation or fusion protein, a compound that selectively inhibits the activity of that mutant protein may be administered.
[00234] In some embodiments, the treatment method involves administering a compound that enhances the immune response against myeloid tumor cells. This may include checkpoint inhibitors that release the brakes on T cell activation, or therapeutic vaccines that stimulate the production of tumor-specific T cells.
[00235] In certain embodiments, the treatment method involves administering a compound that induces differentiation or apoptosis of myeloid tumor cells. This may include agents that target epigenetic regulators, such as DNA methyltransferase or histone deacetylase inhibitors, or that modulate the activity of transcription factors involved in myeloid cell development and survival.
[00236] The treatment methods of the present disclosure can be used to achieve various therapeutic goals, such as inducing remission, prolonging survival, reducing tumor burden, alleviating symptoms, or preventing relapse. The choice of specific therapeutic agent(s) and dosing regimen will depend on factors such as the type and stage of myeloid malignancy, the molecular profile of the tumor cells, the patient's age and general health status, and the presence of comorbidities or contraindications.
[00237] In some embodiments, the treatment method further involves monitoring the patient's response to therapy using the detection methods described herein. Changes in the level or spectrum of leukemic molecular alterations over time can provide an early indication of treatment efficacy or the emergence of drug resistance, allowing for timely adjustment of the therapeutic strategy.
[00238] The treatment methods of the present disclosure can be used in conjunction with conventional supportive care measures, such as transfusion of blood products, administration of prophylactic antibiotics, or management of treatment-related side effects. The methods can also be used in the context of hematopoietic stem cell transplantation, either as a means of preparing the patient for transplant or as a post-transplant maintenance therapy.
[00239] In some embodiments, the myeloid lineage malignancy treatment can include or exclude one of the following treatments: Arsenic Trioxide, Azacitidine, Cyclophosphamide, Cytarabine, Daunorubicin Hydrochloride, Daunorubicin Hydrochloride and Cytarabine Liposome, Daurismo (Glasdegib Maleate), Dexamethasone, Doxorubicin Hydrochloride, Enasidenib Mesylate, Gemtuzumab Ozogamicin, Gilteritinib Fumarate, Glasdegib Maleate, Idamycin PFS (Idarubicin Hydrochloride), Idarubicin Hydrochloride, Idhifa (Enasidenib Mesylate), Ivosidenib, Midostaurin, Mitoxantrone Hydrochloride, Mylotarg (Gemtuzumab Ozogamicin), Olutasidenib,
[00240] Onureg (Azacitidine), Pemazyre (Pemigatinib), Pemigatinib, Prednisone, Quizartinib Dihydrochloride, Rezlidhia (Olutasidenib), Rituxan (Rituximab), Rituximab, Rydapt (Midostaurin), Tabloid (Thioguanine), Thioguanine, Tibsovo (Ivosidenib), Tisagenlecleucel (Kymriah), Trisenox (Arsenic Trioxide), Vanflyta (Quizartinib Dihydrochloride), Venclexta (Venetoclax), Venetoclax, Vincristine Sulfate, Vyxeos (Daunorubicin Hydrochloride and Cytarabine Liposome), or Xospata (Gilteritinib Fumarate). [00241] As used herein, the term “antibody” (Ab) includes, without limitation, a glycoprotein immunoglobulin which binds specifically to an antigen. In general, and antibody can comprise at least two heavy (H) chains and two light (L) chains interconnected by disulfide bonds, or an antigen-binding portion thereof. Each H chain comprises a heavy chain variable
region (abbreviated herein as VH) and a heavy chain constant region. The heavy chain constant region comprises three constant domains, CHI, CH2 and CH3. Each light chain comprises a light chain variable region (abbreviated herein as VL) and a light chain constant region. The light chain constant region comprises one constant domain, CL. The VH and VL regions are further subdivided into regions of hypervariability, termed complementarity determining regions (CDRs), interspersed with regions that are more conserved, termed framework regions (FR). Each VH and VL comprises three CDRs and four FRs, arranged from amino-terminus to carboxy-terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4. The variable regions of the heavy and light chains contain a binding domain that interacts with an antigen. The constant regions of the Abs may mediate the binding of the immunoglobulin to host tissues or factors, including various cells of the immune system (e.g., effector cells) and the first component (Clq) of the classical complement system. An immunoglobulin may derive from any of the commonly known isotypes, including but not limited to IgA, secretory IgA, IgG and IgM. IgG subclasses also include but are not limited to human IgGl, IgG2, IgG3 and IgG4. “Isotype” refers to the Ab class or subclass (e.g., IgM or IgGl) that is encoded by the heavy chain constant region genes. The term “antibody” can include or exclude both naturally occurring and non-naturally occurring Abs; monoclonal and polyclonal Abs; chimeric and humanized Abs; human or nonhuman Abs; wholly synthetic Abs; and single chain Abs. A nonhuman Ab may be humanized by recombinant methods to reduce its immunogenicity in man. The term “antibody” also includes an antigen-binding fragment or an antigen-binding portion of any of the aforementioned immunoglobulins, and includes a monovalent and a divalent fragment or portion, and a single chain Ab.
[00242] Methods of Screening
[00243] Embodiments of the Invention
[00244] In certain embodiments, dislcosed herein are methods for detecting gene variants in a sample from a subject having or suspected of having a myeloid malignancy.
[00245] In certain embodiments, dislcosed herein aremethods for preparing the isolated myeloid nucleic acids (IMNA) or in nucleic acids processed therefrom (NAPT).
[00246] In certain embodiments, disclosed herein are a method of treating a disease or disorder in a subj ect, comprising administering to the subj ect a therapeutically effective amount of a complex or a composition as described herein.
[00247] In certain embodiments, disclosed herein are a method of treating a disease or disorder in a subj ect, comprising administering to the subj ect a therapeutically effective amount of a complex or a composition as described herein.
[00248] Although the present disclosure and its advantages have been described in detail, it should be understood that various changes, substitutions and alterations can be made herein without departing from the spirit and scope of the invention as defined in the appended claims. [00249] The present disclosure will be further illustrated in the following Examples which are given for illustration purposes only and are not intended to limit the invention in any way.
EXAMPLES
[00250] The specific methods and compositions described herein are representative of preferred embodiments and are exemplary and not intended as limitations on the scope of the invention. Other objects, aspects, and embodiments will occur to those skilled in the art upon consideration of this specification, and are encompassed within the spirit of the invention as defined by the scope of the claims. Thus, for example, in each instance herein, and in embodiments or examples of the present invention, any of the terms “comprising”, “consisting essentially of’, and “consisting of’ may be replaced with either of the other two terms in the specification. The methods and processes illustratively described herein suitably may be practiced in differing orders of steps, and that they are not necessarily restricted to the orders of steps indicated herein or in the claims. It is also that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural reference unless the context clearly dictates otherwise. Under no circumstances may the patent be interpreted to be limited to the specific examples or embodiments or methods specifically disclosed herein. Under no circumstances may the patent be interpreted to be limited by any statement made by any Examiner or any other official or employee of the Patent and Trademark Office unless such statement is specifically and without qualification or reservation agreed to and expressly adopted in a responsive writing by Applicants.
Example 1:
[00251] In some embodiments, the method comprises collecting whole blood from a patient, isolating peripheral blood mononuclear cells (PBMCs) from the whole blood via Ficoll-Paque extraction, thereby eliminating granulocytes, applying CD3+/CD19+ magnetic dynabeads to eliminate T cells and B cells from the PBMCs, extracting gDNA from the PBMCs, then using the gDNA as input for targeted error-corrected NGS via the Oncomine Myeloid MRD assay to determine the presence of disease-associated variants. The Oncomine assay is an Ion Torrent Ampliseq assay.
[00252] Briefly, 10 mL of peripheral whole blood (18-20 Celsius) preserved in EDTA is added to a 25ml centrifuge tube with an equal volume Ca2+/Mg2+ phosphate buffered saline (PBS), then the tube is mixed by inversion to create a diluted blood sample. The total cell number is estimated via manual (e.g. hemocytometer) or automated (e.g. electrical impedancebased) cytometer.
[00253] Next, 15 mL Ficoll-Paque solution is added to a 50 mL centrifuge tube, then the diluted blood is carefully layered onto the Ficoll-Paque, ensuring that the two solutions do not mix. The tube is centrifuged at 400 g for 30 to 40 minutes at approximately 18 Celsius. Following centrifugation, the upper layer comprising plasma and platelets is removed, then the lower layer comprising PBMCs is isolated by pipette and transferred to a second tube. For quality assessment purposes and to facilitate downstream processing, the total PBMC count may again be estimated by manual or automated method.
[00254] Next, T cells and B cells are removed from the isolated PBMCs via Thermo Fisher Dynabeads. Anti-CD3 (https://www.thermofisher.com/order/catalog/product/11151D) and anti-CD19 (https://www.thermofisher.com/order/catalog/product/11143D) Dynabeads are mixed in a 1 :1 ratio, then added at a ratio of 25 pL per IxlO7 input cells. The mixture is incubated for 30 minutes at 8 Celsius with gentle tilting and rotation, then a magnet is applied to the tube for 2 minutes. With the magnet in place, the supernatant is carefully removed by pipette and placed into a new tube. For quality assessment purposes and to facilitate downstream processing, the total depleted PBMC cell count may again be estimated by manual or automated method.
[00255] Next, gDNA is isolated from the T and B cell depleted PBMCs via the Qiagen QIAwave DNA Blood and Tissue Kit (Qiagen), using the protocol for cultured cells, eluting the DNA into low TE buffer. The concentration of eluted DNA is measured, e.g. via Thermo Fisher NanoDrop or Thermo Fisher Qubit. If the DNA concentration is less than 9 ng/pL, concentrate the sample (e.g. via vacuum centrifugation) prior to proceeding with the next step. Ensure at least approximately 150 ng of gDNA is available before proceeding to the next step. [00256] Next, process 150 ng of gDNA via the Thermo Fisher Oncomine Myeloid MRD assay, following the DNA-only workflow protocol using the S5 550 flow cell e.g. as described in MAN0025670 (Oncomine Myeloid MRD Assay, Thermofisher). Following variant reporting via the Oncomine Myeloid MRD workflow, variant frequencies are adjusted by multiplying each detected variant frequency by the depleted cell count divided by the total blood cell count. Finally, the adjusted variant frequencies are reported for all detected variants.
A sample may be classified as MRD positive or negative based on the presence and frequency of detected disease-associated variants.
[00257] In one embodiment, in an exemplary workflow, following gDNA isolation, the gDNA serves as input for a highly sensitive targeted duplex sequencing assay, where both Watson and corresponding Crick strand are sequenced, and variants/mutations are only identified if they are present in both the Watson and corresponding Crick strand. In an embodiment, such an assay targets, for example, regions of from 20 to 35, 40, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or approximately 100 genes. In some embodiments, multiple probe intervals are used in each region of interest. In one embodiment, the assay targets, for example, approximately regions of approximately 35 to 40, or approximately 36 genes recurrently mutated in AML with over 220 probe intervals. In an embodiment, such a panel of approximately 36 genes includes, but is not limited to, ASXL1, BCOR, BCORL1, CALR, CBL, CEBPA, CSF3R, DDX41, DNMT3A, ETV6, EZH2, FLT3, GATA2, GNAS, HRAS, IDH1, IDH2, IKZF1, JAK2, KIT, KRAS, KMT2A (MLL), MPL, NF1, NPM1, NRAS, PHF6, PPM1D, PTPN11, RAD21, RUNX1, SETBP1, SF3B1, SRSF2, STAG2, STAT3, TET2, TP53, U2AF1, WT1, and ZRSR2, or a selection thereof relevant for AML MRD detection. In some embodiments, the probe intervals are designed to cover mutational hotspots, entire coding regions, or specific exons and introns of these genes. In an embodiment, library preparation is performed using a Duplex Sequencing library preparation kit (e.g., an equivalent to DuplexSeq V2 Library Preparation Kit), followed by sequencing on a suitable NGS platform (e.g., Illumina NovaSeq 6000). In one embodiment, the error-corrected sequencing data from such an assay detects variants at frequencies below 0.01% VAF, and these ultra-low frequency variant data, after adjustment for cell enrichment as described above, is used to classify a sample as MRD positive or negative with enhanced sensitivity.
[00258] The terms and expressions that have been employed are used as terms of description and not of limitation, and there is no intent in the use of such terms and expressions to exclude any equivalent of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Having thus described in detail preferred embodiments of the present invention, it is to be understood that the invention defined by the above paragraphs is not to be limited to particular details set forth in the above description as many apparent variations thereof are possible without departing from the spirit or scope of the present invention.
[00259] The inventions described and claimed herein have many attributes and embodiments including, but not limited to, those set forth or described or referenced in this
Detailed Disclosure. It is not intended to be all-inclusive and the inventions described and claimed herein are not limited to or by the features or embodiments identified in this Detailed Disclosure, which is included for purposes of illustration only and not restriction. A person having ordinary skill in the art will readily recognise that many of the components and parameters may be varied or modified to a certain extent or substituted for known equivalents without departing from the scope of the invention. It should be appreciated that such modifications and equivalents are herein incorporated as if individually set forth. The invention also includes all of the steps, features, compositions and compounds referred to or indicated in this specification, individually or collectively, and any and all combinations of any two or more of said steps or features.
[00260] All patents, publications, scientific articles, web sites, and other documents and materials referenced or mentioned herein are indicative of the levels of skill of those skilled in the art to which the invention pertains, and each such referenced document and material is hereby incorporated by reference to the same extent as if it had been incorporated by reference in its entirety individually or set forth herein in its entirety. Applicants reserve the right to physically incorporate into this specification any and all materials and information from any such patents, publications, scientific articles, web sites, electronically available information, and other referenced materials or documents. Reference to any applications, patents and publications in this specification is not, and should not be taken as, an acknowledgment or any form of suggestion that they constitute valid prior art or form part of the common general knowledge in any country in the world.
[00261] The specific methods and compositions described herein are representative of preferred embodiments and are exemplary and not intended as limitations on the scope of the invention. Other objects, aspects, and embodiments will occur to those skilled in the art upon consideration of this specification, and are encompassed within the spirit of the invention as defined by the scope of the claims. It will be readily apparent to one skilled in the art that varying substitutions and modifications may be made to the invention disclosed herein without departing from the scope and spirit of the invention. The invention illustratively described herein suitably may be practiced in the absence of any element or elements, or limitation or limitations, which is not specifically disclosed herein as essential. Thus, for example, in each instance herein, in embodiments or examples of the present invention, any of the terms “comprising”, “consisting essentially of’, and “consisting of’ may be replaced with either of the other two terms in the specification. Also, the terms “comprising”, “including”, containing”, etc. are to be read expansively and without limitation. The methods and processes
illustratively described herein suitably may be practiced in differing orders of steps, and that they are not necessarily restricted to the orders of steps indicated herein or in the claims. It is also that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural reference unless the context clearly dictates otherwise. Under no circumstances may the patent be interpreted to be limited to the specific examples or embodiments or methods specifically disclosed herein. Under no circumstances may the patent be interpreted to be limited by any statement made by any Examiner or any other official or employee of the Patent and Trademark Office unless such statement is specifically and without qualification or reservation expressly adopted in a responsive writing by Applicants. Furthermore, titles, headings, or the like are provided to enhance the reader’ s comprehension of this document, and should not be read as limiting the scope of the present invention. Any examples of aspects, embodiments or components of the invention referred to herein are to be considered nonlimiting.
[00262] The terms and expressions that have been employed are used as terms of description and not of limitation, and there is no intent in the use of such terms and expressions to exclude any equivalent of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, it will be understood that although the present invention has been specifically disclosed by preferred embodiments and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.
[00263] The invention has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also form part of the invention. This includes the generic description of the invention with a proviso or negative limitation removing any subject matter from the genus, regardless of whether or not the excised material is specifically recited herein.
[00264] Other embodiments are within the following claims. In addition, where features or aspects of the invention are described in terms of Markush groups, those skilled in the art will recognize that the invention is also thereby described in terms of any individual member or subgroup of members of the Markush group.
Claims
1. A method for detecting one or more leukemic molecular alterations in nucleic acids from myeloid lineage tumor cells, the method comprising: a. obtaining a sample comprising myeloid lineage cells; b. performing one or a plurality of negative enrichment steps on the sample to remove cells from one or more lymphoid lineages, while retaining residual cells of myeloid lineage; c. isolating nucleic acids from the residual cells of myeloid lineage to obtain a sample comprising isolated myeloid nucleic acids; and d. detecting the presence of one or more leukemic molecular alterations in the isolated myeloid nucleic acids (IMNA) or in nucleic acids processed therefrom (NAPT).
2. A method for preparing nucleic acids from myeloid lineage tumor cells comprising: a. obtaining a sample comprising myeloid lineage cells; b. performing one or a plurality of negative enrichment steps on the sample to remove cells of one or more lymphoid lineages, while retaining residual cells of myeloid lineage; and c. isolating nucleic acids from the residual cells of myeloid lineage to obtain a sample comprising isolated myeloid nucleic acids.
3. The method of any of claims 1 or 2, wherein the sample is selected from: whole blood, peripheral blood, peripheral blood mononuclear cells (PBMCs), or bone marrow.
4. The method of claim 3, wherein the one or plurality of negative enrichment steps is selected from: differential centrifugation, and antibody -bead extraction, and wherein said one or plurality of negative enrichment steps result in a reduction of one or more cells selected from B cells, T cells, or NK cells.
5. The method of claim 4, wherein the differential centrifugation is a Ficoll density gradient centrifugation.
6. The method of claim 5, wherein the Ficoll density gradient centrifugation selects for PBMCs.
7. The method of claim 4, wherein the one or plurality of negative enrichment steps comprises a first differential centrifugation step, and a second antibody-bead extraction step.
8. The method of claim 4, wherein the differential centrifugation step negatively selects granulocytes.
9. The method of claim 4, wherein the antibody-bead extraction step negatively selects one or plurality of: B cells, T cells, or NK cells.
10. The method of claim 4, wherein the antibody -bead extraction step is performed using one or more types of antibody-conjugated beads.
11. The method of claim 10, wherein the one or more types of antibody-conjugated beads comprise antibodies which independently target markers selected from one or more of B cell, T cell or NK cell lineage markers.
12. The method of claim 11, wherein the one or more B cell lineage markers is selected from: CD19, CD20, CD22, CD23, CD24, CD38, CD40, CD45, CD79a, CD79b, CD138, CD200 IgKappa, IgLambda, IgM, IgD, IgGl, IgG2, IgG3, IgG4, IgE, IgAl, and/or IgA2.
13. The method of claim 11, wherein the one or more T cell lineage markers is selected from: CD2, CD3, CD4, CD5, CD7, CD8, CD25, CD27, CD28, CD45RA, CD45RO, CD69, TRBC1, TRBC2, and/or CD127.
14. The method of claim 11, wherein the one or more NK cell lineage markers is selected from: CD57, CD94, CD122, CD158a, CD158b, CD159b, CD161, CD314, NKG2A, NKG2C, NKG2D, NKp30, NKp44, NKp46, NKp80, and/or KIR Family Receptors.
15. The method of claim 11, wherein the one or more types of antibody-conjugated beads targeting the B cell, T cell and NK cell lineage markers are present at a ratio that is substantially the same as the relative abundance of the B cells, T cells and NK cells.
16. The method of any of claims 1 or 2, wherein the one or a plurality of negative enrichment steps result in a 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16-fold enrichment of acute myeloid leukemia (AML) cells relative to the initial number of AML cells present in the sample.
17. The method of claim 16, wherein the one or a plurality of negative enrichment steps is performed using a microfluidic device.
18. The method of claim 1, wherein detecting the presence of one or more leukemic molecular alterations is performed using one or a plurality of methods selected from: a Next- Generation Sequencing (NGS) assay, a multiplex PCR assay, a Droplet Digital PCR (ddPCR) assay, a Quantitative PCR (qPCR) assay, and/or a hybrid capture assay.
19. The method of claim 18, wherein the method is an NGS assay on the nucleic acids processed therefrom.
20. The method of claim 19, wherein the NGS assay comprises error correction.
21. The method of claim 20, wherein the error correction is a duplex sequencing method.
22. The method of claim 19, wherein the NGS assay is a targeted NGS assay.
23. The method of claim 22, wherein the targeted NGS assay is a targeted duplex sequencing method, and further wherein the nucleic acids processed therefrom are partially double-stranded and comprise:
(i) a double-stranded target sequence,
(ii) one or more double stranded Mis selected from unique Mis or non-unique Mis located proximal to one or both ends, respectively, of the double-stranded target sequence, and
(iii) on each strand of said double-stranded nucleic acid processed therefrom, one or more strand-distinguishing sequences located distal to the Mis.
24. The method of any one of claims 19, 21 or 22, wherein the NGS assay or the targeted NGS assay, respectively, comprises error correction.
25. The method of claim 21, wherein the duplex sequencing method comprises:
(i) producing a plurality of IMNA-adapters or NAPT-adapters -, wherein a Watson strand of each of the plurality of IMNA-adapters or NAPT-adapters comprises from 5’ to 3’ :
(a) a first Watson single-stranded strand-distinguishing sequence comprising at least one universal primer binding site,
(b) a first Watson MI sequence,
(c) a Watson strand of a target sequence,
(d) a second Watson MI sequence,
(e) a second Watson single-stranded strand-distinguishing sequence comprising at least one universal primer binding site which is not the same as the universal primer binding site in the first Watson single-stranded stranddistinguishing sequence; and wherein a Crick strand of each of the plurality of IMNA-adapters or NAPT-adapters comprises from 5’ to 3’:
(a) a first Crick single-stranded strand-distinguishing sequence which is the same as the first Watson single-stranded strand-distinguishing sequence,
(b) a first Crick MI sequence which is complementary to the second Watson MI sequence,
(c) a Crick strand of the target sequence which is complementary to the Watson strand of the target sequence,
(d) a second Crick MI sequence which is complementary to the first Watson MI sequence,
(e) a second Crick single-stranded strand-distinguishing sequence which is the same as the second Watson single-stranded strand-distinguishing sequence;
(ii) amplifying all or a portion of the plurality of the Watson strands and Crick strands of the IMNA-adapters or NAPT-adapters to obtain Watson amplicons and Crick amplicons using at least one strand-distinguishing sequence comprising at least one universal primer to obtain Watson and Crick IMNA-amplicons or Watson and Crick NAPT-amplicons;
(iii) optionally, performing one of:
(a) one or more targeted amplification steps on a plurality of the Watson and Crick IMNA-amplicons or Watson and Crick NAPT-amplicons or amplified products thereof to yield targeted Watson and Crick IMNA-amplicons or targeted Watson and Crick NAPT-amplicons; or
(b) bait hybridization on the Watson and Crick IMNA-amplicons or Watson and Crick NAPT-amplicons to yield selected Watson and Crick IMNA-amplicons or selected Watson and Crick NAPT-amplicons; and
(iv) analyzing the targeted Watson and Crick IMNA-amplicons or targeted Watson and Crick NAPT-amplicons or the selected Watson and Crick IMNA-amplicons or selected Watson and Crick NAPT-amplicons by performing next-generation sequencing to obtain sequence reads from the targeted Watson and Crick IMNA- amplicons or targeted Watson and Crick NAPT-amplicons or the selected Watson and Crick IMNA-amplicons or selected Watson and Crick NAPT-amplicons.
26. The method of claim 25, further comprising the step of:
(v) identifying one or more variants that are not present in all sequence reads from:
(a) any one target sequence in Watson IMNA amplicons derived from any one Watson IMNA and its corresponding Crick IMNA amplicons derived from any one Crick IMNA;
(b) any one selected sequence in Watson IMNA amplicons derived from any one Watson IMNA and its corresponding Crick IMNA amplicons derived from any one Crick IMNA;
(c) any one target sequence in Watson NAPT amplicons derived from any one Watson NAPT and its corresponding Crick NAPT amplicons derived from any one Crick NAPT; or
(d) any one selected sequence in Watson NAPT amplicons derived from any one Watson NAPT and its corresponding Crick NAPT amplicons derived from any one Crick NAPT.
27. The method of claim 1, wherein the one or more leukemic molecular alterations comprise one or more of genetic signatures selected from: single nucleotide variants (SNVs), indels, structural variants, aberrant methylation, and/or copy number variants.
28. The method of claim 27, wherein the one or more leukemic molecular alterations is present in two or more genomic regions.
29. The method of claim 28, wherein the one or more leukemic molecular alterations is present in 2-60 genomic regions.
30. The method of any of claims 1 or 2, wherein the myeloid lineage tumor cells are selected from one or more cell types selected from: megakaryocytes, platelets, eosinophils, basophils, erythrocytes, monocytes, dendritic cells, macrophages, and/or neutrophils.
31. The method of claim 30, wherein the myeloid lineage tumor cells express one or more surface cell markers selected from: CDl lc, CD13, CD14, CD16, CD31, CD33, CD36, CD56, CD64, CD68, CD115, CD116, CD123, CD124, CD135, CD163, CD203c, CD244, CD300a, CD341, CD366, CD371, CD383, CD387 and/or Myeloperoxidase (MPO).
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463650807P | 2024-05-22 | 2024-05-22 | |
| US63/650,807 | 2024-05-22 |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| WO2025245286A2 true WO2025245286A2 (en) | 2025-11-27 |
| WO2025245286A3 WO2025245286A3 (en) | 2026-01-22 |
Family
ID=97796020
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2025/030454 Pending WO2025245286A2 (en) | 2024-05-22 | 2025-05-21 | Method for detecting myeloid-lineage malignancies/myeloid malignancies |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025245286A2 (en) |
-
2025
- 2025-05-21 WO PCT/US2025/030454 patent/WO2025245286A2/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2025245286A3 (en) | 2026-01-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11390905B2 (en) | Methods of nucleic acid sample preparation for analysis of DNA | |
| EP3535405B1 (en) | Methods of nucleic acid sample preparation for immune repertoire sequencing | |
| US20220056534A1 (en) | Methods for analysis of circulating cells | |
| JP6054303B2 (en) | Optimization of multigene analysis of tumor samples | |
| JP2021104031A (en) | Tm-ENHANCED BLOCKING OLIGONUCLEOTIDES AND BAITS FOR IMPROVED TARGET ENRICHMENT AND REDUCED OFF-TARGET SELECTION | |
| EP3512947B1 (en) | Methods of nucleic acid sample preparation | |
| JP2021166532A (en) | Multiple gene analysis of tumor samples | |
| CN110741096B (en) | Compositions and methods for detecting circulating tumor DNA | |
| JP2023511200A (en) | Immune repertoire biomarkers in autoimmune and immunodeficiency diseases | |
| EP3768864A1 (en) | Immune repertoire monitoring | |
| US20220282305A1 (en) | Methods of nucleic acid sample preparation | |
| WO2025245286A2 (en) | Method for detecting myeloid-lineage malignancies/myeloid malignancies | |
| HK40066065A (en) | Methods of nucleic acid sample preparation | |
| WO2026015794A1 (en) | Methods for modifying dna using cpg-specific deamination and uracil base excision | |
| EP4051808A1 (en) | Method for identifying transplant donors for a transplant recipient | |
| HK40063633B (en) | Methods for analysis of circulating cells | |
| HK40063633A (en) | Methods for analysis of circulating cells |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25808618 Country of ref document: EP Kind code of ref document: A2 |