EP4616005A2 - Systeme für mutationsanrufer und verfahren zur verwendung davon - Google Patents
Systeme für mutationsanrufer und verfahren zur verwendung davonInfo
- Publication number
- EP4616005A2 EP4616005A2 EP23889780.5A EP23889780A EP4616005A2 EP 4616005 A2 EP4616005 A2 EP 4616005A2 EP 23889780 A EP23889780 A EP 23889780A EP 4616005 A2 EP4616005 A2 EP 4616005A2
- Authority
- EP
- European Patent Office
- Prior art keywords
- sample
- cancer
- nullomers
- subject
- cell
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/20—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6876—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
- C12Q1/6883—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
- C12Q1/6886—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material for cancer
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2600/00—Oligonucleotides characterized by their use
- C12Q2600/156—Polymorphic or mutational markers
Definitions
- the present disclosure relates to the identification of prognostic and diagnostic cancer biomarkers in biological material and the characterization of tumor subtype, vulnerabilities and therapeutic strategies, from the resurfacing of nullomers.
- Cancer is the second leading cause of death worldwide (“Cancer” n.d.), and for most cancer types, survivability is significantly higher if the tumor is detected at an early stage (Hawkes 2019; Etzioni et al. 2003).
- mass population screening is applicable only for breast and cervical cancers and utilizes physical tests like mammography and cytology screens. Detection for other cancer types, done both en masse and in a low and affordable resource setting, still poses a major challenge for the scientific and clinical communities (“Cancer” n.d.).
- a major hurdle is to single-out cancer biomarkers for the detection of cancer development at its earliest stage for patient stratification and improvement of patients’ outcome by providing personalized treatments.
- Some of the major hurdles include: 1) cfDNA is fragmented (180-360 base pairs) making its collection and extraction more challenging and the tumor-derived DNA makes up only a small portion (estimated to be around 0.4%) warranting the need for extremely sensitive biomarkers that can easily detect the presence of cancerous cells; 2) prior knowledge of specific mutations or methylation marks is required for targeted screening, and consequently the main focus has been on coding mutations which only constitute a small fraction of mutations; 3) cfDNA mutation and epigenetic diagnosis could be confounded by somatic alterations in white blood cells (Razavi et al. 2019); 4) the diagnostic techniques used to detect methylation or histone marks are technologically complex and can have low sensitivity and specificity (Ji et al.
- nullomers do not exist in a human genome, their appearance due to mutagenesis followed by clonal expansion could be exploited as a diagnostic method for diseases associated with a mutational burden, such as cancer.
- neomers alter regulatory activity of tumors and can be used to detect cancer-associated mutations in gene regulatory elements.
- MPRA massively parallel reporter assay
- the disclosure relates to methods and compositions for the detection, identification, classification and characterization of cancer in general and cancer types in biological material such as solid tumors.
- the disclosure also relates to a method of identifying a neomer from a sample from subject comprising:
- the disclosure relates to a method of creating a library of neomers that correspond to a cancer type from a sample of a subject comprising:
- nullomers from a sample of a subject (b) identifying nullomers from a sample of a subject; (c) annotating the nullomers from the subject as being specific to the subject by comparing the nullomers from the sample with nucleic acid sequences from a control subject;
- the one or plurality of probes comprise a complementary nucleic acid sequence bound to or associated with a fluorescent molecule, radioactive isotope or chemiluminescent molecule.
- the step of detecting is performed by mass spectrometry.
- the step of detecting and/or correlating the presence or quantity of probes with the likelihood of the presence or quantity of neomer in the sample is performed by analyzing the sample for the presence of a neomer.
- the disclosure provides a method of identifying one or a plurality of nullomers in a sample comprising: (a) isolating a plurality of nucleic acids from the sample; (b) contacting the nucleic acids to one or a plurality of probes specific for one or a plurality of nullomers; (c) detecting the presence of the probes associated with the one or plurality of nullomers; and (d) correlating the presence or quantity of probes with the likelihood of the presence or quantity of nullomers in the sample.
- the one or plurality of probes comprise a complementary nucleic acid sequence bound to or associated with a fluorescent molecule, radioactive isotope or chemiluminescent molecule.
- the step of detecting is performed by mass spectrometry.
- the method further comprises, prior to step (b), disassociating a plurality of double stranded nucleic acid sequences comprising at least one nullomer by exposing the double-stranded nucleic acid sequences to a predetermined melting temperature for a period of time sufficient to create single stranded nullomer, annealing at least one primer to the nullomer, and allowing a sufficient period of time to extend the primer in the presence of nucleotide triphosphates (NTPs) and a polymerase or functional fragment thereof.
- NTPs nucleotide triphosphates
- the steps of disassociating a plurality of double stranded nucleic acid sequences comprising at least one nullomer by exposing the double-stranded nucleic acid sequences to a predetermined melting temperature for a period of time sufficient to create single stranded nullomer, annealing at least one primer to the nullomer, and allowing a sufficient period of time to extend the primer in the presence of NTPs and the polymerase are repeated multiple times such that copies of the at least one nullomer are produced.
- the disclosure further provides a method of identifying one or plurality of nullomers in a sample comprising: (a) isolating a plurality of nucleic acids from the sample; (b) contacting the nucleic acids to one or a plurality of probes specific for one or a plurality of nullomers; (c) detecting the presence of the probes associated with the one or plurality of neomers; (d) correlating the presence or quantity of probes with the likelihood or the presence or quantity of neomers in the sample; and (e) comparing the sequence of the neomer with the sequence of a library of known neomer sequences.
- the probe or plurality of probes comprise a complementary nucleic acid sequence bound to or associated with a fluorescent molecule, radioactive isotope or chemiluminescent molecule.
- the method further comprises a step of performing polymerase chain reaction (PCR) with one or a plurality of primers specific for the one or plurality of neomers.
- PCR polymerase chain reaction
- the disclosure relates to a method of diagnosing a subject with a cancer comprising:
- the sample in any of the disclosed methods is a cell free nucleic acid sample.
- the disclosure also provides a computer-implemented method of identifying a mutation associated with a comprising: (a) isolating one or a plurality of nucleic acid molecules from a sample associated with the hyperproliferative disorder; (b) contacting the nucleic acids to one or a plurality of probes specific for one or a plurality of neomers; (c) in a system configured to compile data and detect the presence or quantify the presence of a nucleic acid sequence, detecting the presence of the probes associated with the one or plurality of neomers; (d) correlating the presence or quantity of the neomer to the likelihood of a specific mutation serving as a biomarker for a hyperproliferative disorder.
- the method further comprises, prior to step (a), in a system configured to compile data and detect the presence or quantity of nucleic acids in a sample: compiling genetic data about a population of subjects including the subject that has a mutation candidate that is a biomarker for a hyperproliferative disorder.
- the method further comprises, after step (d), a step of: (e) selecting a cancer treatment for the subject based upon identification of the hyperproliferative disorder.
- the hyperproliferative disorder is breast cancer, ovarian cancer, lung cancer, pancreatic cancer, or liver cancer.
- the hyperproliferative disorder is breast cancer, pancreatic cancer, esophagus cancer, lymphoid cancer, kidney cancer, ovarian cancer, head and neck cancer, lung cancer, stomach cancer, CNS cancer, uterus cancer, skin cancer, colorectal cancer, prostate cancer, bladder cancer, bone and soft tissue cancer, biliary cancer, cervix cancer, thyroid cancer, myeloid cancer, or liver cancer.
- the hyperproliferative disorder is a malignant tumor.
- the sample is a brush biopsy, puncture biopsy, fluid from a needle biopsy, blood, blood cells, cells from a hair sample, nucleic acids from a hair sample, saliva, or spit.
- the probe or plurality of probes comprise a complementary nucleic acid sequence bound to or associated with a fluorescent molecule, radioactive isotope or chemiluminescent molecule.
- the method further comprises a step of performing PCR with one or a plurality of primers specific for the one or plurality of nullomers.
- the disclosure additionally provides a method of treating a hyperproliferative disorder in a subject in need thereof comprising: (a) exposing a sample from the subject to a probe specific for at least one neomer chosen from Table 1; (b) detecting the presence, absence or quantity of the at least one neomer in the sample; (c) normalizing the presence, absence, or quantity of the at least one neomer in the sample against the presence, absence or quantity of the at least one neomer in a sample of a healthy subject or a sample of a subject known to have the hyperproliferative disorder; (d) correlating the presence, absence, or quantity of the at least one neomer in the sample to the subject having the hyperproliferative disorder; and (e) administering a therapeutically effective amount of one or a plurality of active agents to the subject.
- the method further comprises obtaining the sample from the subject prior to the step of exposing.
- the one or plurality of active agents is chosen from one or a combination of the agents identified in Table 3.
- the sample is plasma, serum, whole blood, respiratory tissue, respiratory mucosal sample, saliva, urine, blood cells, cells from a hair sample, nucleic acids from a hair sample, or spit.
- the sample is a blood sample comprising cell -free genomic DNA or RNA from a solid tumor or cell-free genomic DNA or RNA from a circulating tumor cell from the subject.
- the sample is a blood sample comprising a circulating tumor cell from the subject.
- step (b) further comprises calculating one or more scores based upon the presence, absence, or quantity of the at least one neomer
- step (d) further comprises correlating the one or more scores to the presence, absence, or quantity of the at least one nullomer such that, if the amount of the at least one neomer is greater than the quantity of the at least one neomer in a control sample; or, if the amount of the at least one neomer is substantially equal to the quantity of the at least one neomer in a sample taken from a subject known to have a hyperproliferative disorder, then the subject is diagnosed as having a hyperproliferative disorder.
- the probe is a radioactive probe, a chemiluminescent probe, or a fluorescent probe.
- the sample is free of cells.
- the at least one nullomer is detected by next generation sequencing, quantitative real-time reverse transcript! on-PCR (qRT-PCR), isothermal amplification, microarray, multiplex nullomer profiling assay, RNA in situ hybridization (RNA-ish), or northern blotting.
- the at least one nullomer is detected by qRT-PCR.
- the step of quantifying at least one quantity of the at least one nullomer in the sample comprises using a fluorescence and/or digital imaging.
- the step of analyzing comprises detecting a presence, absence, or quantity of at least 2 different neomers. In some embodiments, the step of analyzing comprises detecting the presence, absence, or quantity of the at least one neomer by PCR amplification using one or a plurality of primers specific for the at least one neomer chosen from Table 1. In some embodiments, the step of analyzing comprises detecting presence, absence, or quantity of the at least one neomer by a probe comprising a nucleic acid sequence complementary to the nucleic acid sequence of the at least one nullomer.
- the disclosure further provide a method of diagnosing a subject with cancer comprising: (a) contacting a plurality of nucleic acids from a sample to a system comprising a probe specific for one or a plurality of neomers; and (b) detecting the presence of or quantifying the amount of one or more nucleic acids from the sample.
- the method comprises detecting the presence, absence or quantity of one or a plurality of the neomers provided in Table 1.
- the method comprises detecting the presence, absence or quantity of nullomers that comprise at least 93% sequence identify to one or a plurality of the nullomers provided in Table 1.
- the at least one nullomer is detected by qRT-PCR.
- the at least one nullomer is detected by CRISPR diagnosis.
- the at least one nullomer is detected by CRISPR diagnosis and Cas9, Casl2 or Casl3 protein is used.
- the method further comprises, after the step of detecting, normalizing the quantity of the probe as compared to a quantity of signal from a negative control. In some embodiments, the method further comprises, after the step of detecting, correlating the one or more scores to the presence, absence, or quantity of the at least one neomer such that, if the amount of the at least one neomer is greater than the quantity of the at least one neomer in a control sample; or, if the amount of the at least one neomer is substantially equal to the quantity of the at least one neomer in a sample taken from a subject known to have a hyperproliferative disorder, then the subject is diagnosed as having a hyperproliferative disorder.
- the hyperproliferative disorder is solid tumor of the breast, pancreas, ovary, lung or liver.
- the hyperproliferative disorder is breast cancer, pancreatic cancer, esophagus cancer, lymphoid cancer, kidney cancer, ovary cancer, head and neck cancer, lung cancer, stomach cancer, CNS cancer, uterus cancer, skin cancer, colorectal cancer, prostate cancer, bladder cancer, bone and soft tissue cancer, biliary cancer, cervix cancer, thyroid cancer, myeloid cancer, or liver cancer.
- the hyperproliferative disorder is a metastatic tumor if the presence or quantity of the neomer corresponds to the presence of a circulating tumor cell in a blood sample.
- kits comprising one or more probes or primers for detecting the presence, absence or quantity of one or a plurality of the neomers provided in Table 1 or neomers that comprise at least 93% sequence identify to one or a plurality of the neomers provided in Table 1.
- the one or more probes comprised in the disclosed kit comprise one or a combination of the neomer sequences of Table 1 or complementary thereof.
- a computer program product encoded on a computer-readable storage medium, wherein the computer program product comprises instructions for: a) detecting the presence, absence or quantity of at least one neomer in a sample of a subject; b) normalizing the presence, absence, or quantity of the at least one neomer in the sample against the presence, absence or quantity of the at least one neomer in a control sample; and c) correlating the presence, absence, or quantity of the at least one neomer in the sample to a likelihood that the subject having a hyperproliferative disorder.
- the computer program product further comprises instructions for calculating a score associated with the presence, absence or quantity of the at least one neomer in the sample and correlating the score to a likelihood that the subject has a hyperproliferative disorder.
- the computer program product further comprises instructions for: a) detecting and normalizing the presence, absence or quantity of a second neomer in the sample; b) calculating a combined score associated with the presence, absence or quantity of the at least one neomer and the second neomer in the sample; and c) correlating the combined score to a likelihood that the subject having a hyperproliferative disorder.
- At least 2 different neomers in the sample are detected, normalized and correlated by the computer program product.
- the computer program product detects the presence, absence, or quantity of the at least one neomer by qRT-PCR amplification.
- the control sample used in the computer program product is obtained from a subject free of a hyperproliferative disorder.
- the disclosure further provides a system for detecting the presence or quantity of neomer in a sample of a subject comprising: a processor operable to execute programs, a memory associated with the processor, a database associated with said processor and said memory, and a program stored in the memory and executable by the processor, the program being operable for: a) detecting the presence, absence or quantity of at least one neomer in a sample of a subject; b) normalizing the presence, absence, or quantity of the at least one neomer in the sample against the presence, absence or quantity of the at least one neomer in a control sample; and c) correlating the presence, absence, or quantity of the at least one neomer in the sample to a likelihood that the subject having a hyperproliferative disorder.
- the program is further operable for calculating a score associated with the presence, absence or quantity of the at least one neomer in the sample and correlating the score to a likelihood that the subject has a hyperproliferative disorder. In some embodiments, the program is further operable for detecting and normalizing the presence, absence or quantity of a second neomer in the sample.
- the one or plurality of probes used in any of the disclosed methods, systems, or computer program product, or comprised in any of the disclosed kits comprise a nucleic acid sequence chosen from Table 1. In some embodiments, the one or plurality of probes used in any of the disclosed methods, systems, or computer program product, or comprised in any of the disclosed kits comprise a nucleic acid sequence comprising at least about 93% sequence identity to any of the sequences in Table 1.
- the one or plurality of probes used in any of the disclosed methods, systems, or computer program product, or comprised in any of the disclosed kits comprise a nucleic acid sequence that is complementary to any of the neomer sequences provided in Table 1, or a fragment thereof. In some embodiments, the one or plurality of probes used in any of the disclosed methods, systems, or computer program product, or comprised in any of the disclosed kits comprise a nucleic acid sequence that is complementary to a neomer comprising at least about 93% sequence identity to any of the neomer sequences provided in Table 1, or a fragment thereof.
- the one or plurality of primers specific for the one or plurality of neomers used in any of the disclosed methods, systems, or computer program product, or comprised in any of the disclosed kits comprise a nucleic acid sequence that is complementary to any of the neomer sequences provided in Table 1, or a fragment thereof.
- the one or plurality of primers specific for the one or plurality of neomers used in any of the disclosed methods, systems, or computer program product, or comprised in any of the disclosed kits comprise a nucleic acid sequence that is complementary to a neomer comprising at least about 93% sequence identity to any of the neomer sequences provided in Table 1, or a fragment thereof.
- FIGS. 1A-1E Neomers can detect cancer tissue of origin.
- Fig. 1A Schematic overview of neomer cancer diagnostic pipeline.
- Fig. IB Number of neomers per patient sample across tissues. Each dot represents a patient sample.
- Fig. ID Heatmap showing the Jaccard index for the overlap of neomer sets associated with different cancer types.
- Fig. IE Heatmap showing the occurrence of neomers across patients for each cancer type. Each row represents a cancer type and each column a patient. The intensity of the heatmap (log2-scale) shows the number of neomers for each tissue set.
- Figs. 2A-2D Neomers can distinguish cancer features.
- Fig. 2A-B Classifier accuracy (A) using an unsupervised classifier and F l (B) score for the same classifier for each of the twenty- one cancer types.
- Fig. 2C Separation of MSI and MSS samples using a supervised selection of neomers.
- Fig. 2D Separation of POLE proficient and deficient samples using nullomers.
- the vertical line displays the harmonic mean.
- Figs. 3A-3G Identification of cancer in liquid biopsy samples using neomers.
- Fig. 3A Cancer status detection in lung patients and healthy controls from whole-genome sequencing of liquid biopsy samples (***p-value ⁇ 0.0005, Mann-Whitney U). Number of neomers detected in cfDNA from healthy controls, lung cancer patients, and matching tumors.
- Fig. 3B Number of neomers for lung cancer stratified by tumor stage (p-vahie ⁇ 0.006, Kruskal-Wallis test).
- Fig. 3C Number of neomers observed in ovarian samples and healthy controls
- Fig. 3D Cancer status detection in ovarian cancer and controls using rare 13mers. (*p-value ⁇ 0.03, Mann-Whitney U).
- Fig. 3A Cancer status detection in lung patients and healthy controls from whole-genome sequencing of liquid biopsy samples (***p-value ⁇ 0.0005, Mann-Whitney U). Number of neomers detected in cfDNA from healthy controls, lung cancer patients, and matching
- Fig. 3E Number of first order nullomers detected in ovarian samples and healthy controls (*p- value ⁇ 0.01, Mann-Whitney U), Fig. 3F. Number of neomers for prostate and control. Fig. 3G. Jaccard similarity between prostate neomers found in cfDNA and tumor samples.
- Figs. 4A-C Neomers alter the activity of gene regulatory elements.
- Fig. 4C Relative luciferase units from a luciferase reporter assay for promoters containing either the reference or nullomer variant. Transfection efficiency was normalized using renilla luciferase and significance is calculated using a two-way ANOVA with multiple testing and Sidak correction.
- Figs. 5A-E Fig. 5A. Number of patients per cancer tissue.
- Figs. 5B-C Number of nullomers resurfaced due to indels or substitutions observed for each tumor sample (per patient).
- Fig. 5D Association between number of mutations and number of nullomers observed.
- Fig. 5E Number of nullomers of different lengths observed per patient. DETAILED DESCRIPTION OF EMBODIMENTS
- a reference to “A and/or B,” when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A without B (optionally including elements other than B); in another embodiment, to B without A (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
- the term “animal” includes, but is not limited to, humans and non-human vertebrates such as wild animals, rodents, such as rats, ferrets, and domesticated animals, and farm animals, such as dogs, cats, horses, pigs, cows, sheep, and goats.
- the animal is a mammal.
- the animal is a human.
- the animal is a non-human mammal.
- an “algorithm,” “formula,” or “model” is any mathematical equation, algorithmic, analytical or programmed process, or statistical technique that takes one or more continuous or categorical inputs (herein called “parameters”) and calculates an output value, sometimes referred to as an “index” or “index value.”
- “formulas” include sums, ratios, and regression operators, such as coefficients or exponents, biomarker (e.g., nullomers disclosed herein) value transformations and normalizations (including, without limitation, those normalization schemes based on clinical parameters, such as gender, age, or ethnicity), rules and guidelines, statistical classification models, and neural networks trained on historical populations.
- markers Of particular use in combining markers are linear and non-linear equations and statistical classification analyses to determine the relationship between levels of the biomarkers detected in a subject sample and the subject’s risk of disease (for example).
- panel and combination construction of particular interest are structural and syntactic statistical classification algorithms, and methods of risk index construction, utilizing pattern recognition features, including established techniques such as cross correlation, Principal Components Analysis (PCA), factor rotation, Logistic Regression (LogReg), Linear Discriminant Analysis (LDA), Eigengene Linear Discriminant Analysis (ELDA), Support Vector Machines (SVM), Random Forest (RF), Recursive Partitioning Tree (RPART), as well as other related decision tree classification techniques, Shruken Centroids (SC), StepAIC, Kth-Nearest Neighbor, Boosting, Decision Trees, Neural Networks, Bayesion Networks, Support Vector Machines, and Hidden Markov Models, among others.
- PCA Principal Components Analysis
- LogReg Logistic Regression
- LDA Linear Discriminant Analysis
- biomarker selection techniques are useful either combined with a biomarker selection technique, such as forward selection, backwards selection, or stepwise selection, complete enumeration of all potential panels of a given size, genetic algorithms, or they may themselves include biomarker selection methodologies in their own technique.
- biomarker selection methodologies such as Akaike’s Information Criterion (AIC) or Bayes Information Criterion (BIC), in order to quantify the trade-off between additional biomarkers and model improvement, and to aid in minimizing overfit.
- AIC Information Criterion
- BIC Bayes Information Criterion
- the resulting predictive models may be validated in other studies, or cross-validated in the study they were originally trained in, using such techniques as Leave- One-Out (LOO) and 10-Fold cross-validation (10-Fold-CV).
- LEO Leave- One-Out
- 10-Fold cross-validation 10-Fold-CV
- At least prior to a number or series of numbers (e.g. “at least two”) is understood to include the number adjacent to the term “at least,” and all subsequent numbers or integers that could logically be included, as clear from context.
- at least is present before a series of numbers or a range, it is understood that “at least” can modify each of the numbers in the series or range.
- biomarker refers to a biological molecule present in an individual at varying concentrations useful in predicting the cancer status of an individual.
- a biomarker may include but is not limited to, nucleic acids, proteins and variants and fragments thereof.
- a biomarker may be DNA comprising the entire or partial nucleic acid sequence encoding the biomarker, or the complement of such a sequence.
- Biomarker nucleic acids useful in the disclosure are considered to include both DNA and RNA comprising the entire or partial sequence of any of the nucleic acid sequences of interest.
- the biomarker of the disclosure is any of the nullomers disclosed herein.
- the term “bodily fluid” as used herein refers to a bodily fluid including blood (or a fraction of blood such as plasma or serum), lymph, mucus, tears, saliva, sweat, sputum, urine, semen, stool, cerebrospinal fluid (CSF), breast milk, and, ascites fluid.
- the bodily fluid is blood.
- the bodily fluid is a fraction of blood.
- the bodily fluid is plasma.
- the bodily fluid is serum.
- the bodily fluid is urine.
- the bodily fluid is free of cells.
- the bodily fluid comprises a circulating tumor cell.
- the sample comprises cell-free nucleic acids.
- cancer and “cancerous” as used herein refer to or describe a physiological condition in mammals in which a population of cells are characterized by unregulated cell growth.
- cancer refers to a group of diseases involving abnormal cell growth with the potential to invade or spread to other parts of the body.
- cancer examples include, but not limited to, lung cancer, bone cancer, blood cancer, chronic myelomonocytic leukemia (CMML), bile duct cancer, cervical cancer, liver cancer, pancreatic cancer, skin cancer, cancer of the head and neck, cancer of the eye, cutaneous or intraocular melanoma, uterine cancer, ovarian cancer, rectal cancer, cancer of the anal region, stomach cancer, colon cancer, breast cancer, testicular cancer, gynecologic tumors (e.g., uterine sarcomas, carcinoma of the fallopian tubes, carcinoma of the endometrium, carcinoma of the cervix, carcinoma of the vagina or carcinoma of the vulva), Hodgkin’s disease, cancer of the esophagus, cancer of the small intestine, cancer of the endocrine system (e.g., cancer of the thyroid, parathyroid or adrenal glands), sarcomas of soft tissues, cancer of the urethra, cancer of the penis
- the term “characterizing cancer in a subject” refers to the identification of one or more properties of a cancer sample in a subject, including but not limited to, the presence of benign, pre-cancerous or cancerous tissue, the stage of the cancer, the type of the cancer, the tissue of origin of the cancer, and the subject’s prognosis. Cancers may be characterized by the identification of the expression of one or more cancer marker genes, including but not limited to, the nullomers disclosed herein. As used herein, the term “stage of cancer” refers to a qualitative or quantitative assessment of the level of advancement of a cancer.
- Criteria used to determine the stage of a cancer include, but are not limited to, the size of the tumor and the extent of metastases (e.g., localized or distant).
- the subject has been previously diagnosed with having a cancer and received, or is currently receiving, cancer treatment, including but not limited to surgical intervention and cancer therapy, and in such embodiments, the term “characterizing cancer in a subject” refers to monitoring the progress of the cancer treatment.
- complementarity refers to polynucleotides (i.e., a sequence of nucleotides) related by base-pairing rules, for example, the sequence “5’-AGT-3’,” is complementary to the sequence “5’-ACT-3’ ”
- Complementarity may be “partial,” in which only some of the nucleic acids’ bases are matched according to the base pairing rules, or there may be “complete” or “total” complementarity between the nucleic acids.
- the degree of complementarity between nucleic acid strands can have significant effects on the efficiency and strength of hybridization between nucleic acid strands under defined conditions. This is of particular importance for methods that depend upon binding between nucleic acid bases.
- the terms “comprising” (and any form of comprising, such as “comprise,” “comprises,” and “comprised”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”), or “containing” (and any form of containing, such as “contains” and “contain”), are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.
- correlate refers to a statistical association between instances of two events, where events may include numbers, data sets, and the like.
- a positive correlation also referred to herein as a “direct correlation” means that as one increases, the other increases as well.
- a negative correlation also referred to herein as an “inverse correlation” means that as one increases, the other decreases.
- nullomers the levels of which are correlated with a particular outcome measure, such as between the presence of a particular nullomer and the likelihood of developing a particular type of cancer. For example, the increased level of a nullomer may be negatively correlated with a likelihood of good clinical outcome for the patient.
- the patient may have a decreased likelihood of long-term survival without recurrence of the cancer and/or a positive response to a chemotherapy, and the like.
- a negative correlation indicates that the patient likely has a poor prognosis or will respond poorly to a chemotherapy, and this may be demonstrated statistically in various ways, e.g., by a high hazard ratio.
- Detecting a composition may comprise determining the presence or absence of a composition. Detecting may comprise quantifying a composition. For example, detecting comprises determining the expression level of a composition.
- the composition may comprise a nucleic acid molecule.
- the composition may comprise one or a plurality of the nullomers disclosed herein. Alternatively, or additionally, the composition may be a detectably labeled composition.
- diagnosis or “prognosis” as used herein refers to the use of information (e g., genetic information or data from other molecular tests on biological samples, signs and symptoms, physical exam findings, cognitive performance results, etc.) to anticipate the most likely outcomes, timeframes, and/or response to a particular treatment for a given disease, disorder, or condition, based on comparisons with a plurality of individuals sharing common nucleotide sequences, symptoms, signs, family histories, or other data relevant to consideration of a patient’s health status.
- information e e g., genetic information or data from other molecular tests on biological samples, signs and symptoms, physical exam findings, cognitive performance results, etc.
- a functional fragment means any portion of a polypeptide or nucleic acid sequence from which the respective full-length polypeptide or nucleic acid relates that is of a sufficient length and has a sufficient structure to confer a biological affect that is similar or substantially similar to the full-length polypeptide or nucleic acid upon which the fragment is based.
- a functional fragment is a portion of a full-length or wild-type nucleic acid sequence that encodes any one of the nucleic acid sequences disclosed herein, and said portion encodes a polypeptide of a certain length and/or structure that is less than full-length but encodes a domain that still biologically functional as compared to the full-length or wild-type protein.
- the functional fragment may have a reduced biological activity, about equivalent biological activity, or an enhanced biological activity as compared to the wildtype or full-length polypeptide sequence upon which the fragment is based.
- the functional fragment is derived from the sequence of an organism, such as a human.
- the functional fragment may retain about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% sequence identity to the wild-type or given sequence upon which the sequence is derived.
- the functional fragment may retain about 85%, 80%, 75%, 70%, 65%, or 60% sequence identity to the wild-type sequence upon which the sequence is derived.
- the given sequence is a nullomer sequence of Table 1 or Table B. In other embodiments, the given sequence is a complementary sequence of any of the nullomer sequences of Table 1 or Table B.
- hypoproliferation as used herein is defined as clonal expansion, in which daughter cells share a set of somatic mutations that were not originally present in the germline and which could include but are not limited to driver mutations. Clonal expansion could include but is not limited to resistance to cell death, evasion of growth suppressors, sustaining proliferate signaling, enabling replicative immortality, activating invasion and metastasis or inducing angiogenesis.
- hyperproliferative cell refers to a cell located in a tissue or organ having a “hyperproliferative disorder,” a disease or disorder characterized by abnormal proliferation, abnormal growth, abnormal senescence, abnormal quiescence, or abnormal removal of cells in an organism, and includes all forms of hyperplasias, neoplasias, and cancer.
- the “hyperproliferative cell” is a precancerous cell in form of hyperplasias.
- the “hyperproliferative cell” is precancerous cell in form of neoplasias.
- the “hyperproliferative cell” is a cancerous cell.
- the hyperproliferative disorder or disease is a cancer derived from the gastrointestinal tract or urinary system.
- a hyperproliferative disorder or disease is a cancer of the adrenal gland, bile ducts, bladder, blood, bone, bone marrow, brain, breast, cervix, colon, esophagus, eye, gall bladder, ganglia, gastrointestinal tract, heart, lymphatic system, liver, lung, kidney, muscle, ovary, pancreas, parathyroid, penis, prostate, prostate glands, rectum, salivary glands, skin, spine, stomach, spleen, testis, thymus, thyroid, or uterus.
- the term hyperproliferative disorder or disease is a cancer chosen from: lung cancer, bone cancer, blood cancer, chronic myelomonocytic leukemia (CMML), bile duct cancer, cervical cancer, liver cancer, pancreatic cancer, skin cancer, cancer of the head and neck, cancer of the eye, cutaneous or intraocular melanoma, uterine cancer, ovarian cancer, rectal cancer, cancer of the anal region, stomach cancer, colon cancer, breast cancer, testicular cancer, gynecologic tumors (e.g., uterine sarcomas, carcinoma of the fallopian tubes, carcinoma of the endometrium, carcinoma of the cervix, carcinoma of the vagina or carcinoma of the vulva), Hodgkin’s disease, cancer of the esophagus, cancer of the small intestine, cancer of the endocrine system (e.g., cancer of the thyroid, parathyroid or adrenal glands), sarcomas of soft tissues, cancer of
- the hyperproliferative disorder or disease is a breast cancer, pancreatic cancer, esophagus cancer, lymphoid cancer, kidney cancer, ovary cancer, head and neck cancer, lung cancer, stomach cancer, CNS cancer, uterus cancer, skin cancer, colorectal cancer, prostate cancer, bladder cancer, bone and soft tissue cancer, biliary cancer, cervix cancer, thyroid cancer, myeloid cancer, or liver cancer.
- the hyperproliferative disorder or disease comprises one or a plurality of mutations in one or a plurality of genes selected from Table A.
- the phrase “in need thereof’ means that the animal or mammal has been identified or suspected as having a need for the particular method or treatment. In some embodiments, the identification can be by any means of diagnosis or observation. In any of the methods and treatments described herein, the animal or mammal can be in need thereof.
- a label may be a charged moiety (positive or negative charge) or alternatively, may be charge neutral.
- Labels can include or consist of nucleic acid or protein sequence, so long as the sequence comprising the label is detectable. In some embodiments, nucleic acids are detected directly without a label (e.g., directly reading a sequence).
- level refers to qualitative or quantitative amount of the number of copies of a nullomer.
- a nullomer exhibits an “increased level” when the level of the nullomer is higher in a first sample, such as in a clinically relevant subpopulation of patients (e.g., patients who have cancer), than in a second control sample, such as in a related subpopulation (e.g., patients who do not have cancer).
- a nullomer exhibits “increased level” when the level of the nullomer in the subject trends toward, or more closely approximates, the level characteristic of a clinically relevant subpopulation of patients.
- measuring means assessing the presence, absence, quantity or amount (which can be an effective amount) of either a given substance within a clinical or subject-derived sample, including the derivation of qualitative or quantitative concentration levels of such substances, or otherwise evaluating the values or categorization of a subject’s clinical parameters.
- detecting or “detection” may be used and is understood to cover all measuring or measurement as described herein.
- metalastasis refers to the process by which a cancer spreads or transfers from the site of origin to other regions of the body.
- a “metastatic” or “metastasizing” cell is one that loses adhesive contacts with neighboring cells and migrates (e.g., via the bloodstream or lymph) from the primary site of disease to secondary sites.
- nucleic acid refers to any nucleic acid
- oligonucleotide refers to any nucleic acid molecules
- polynucleotide refers to any combination of nucleic acid molecules.
- Both terms are used to denote a DNA, RNA, modified or synthetic DNA or RNA sequence (including, but not limited to nucleic acids comprising synthetic and naturally-occurring base analogs, dideoxy or other sugars, thiols or other non-natural or natural polymer backbones), or other nucleobase containing polymers capable of hybridizing to DNA and/or RNA. Accordingly, the terms should not be construed to define or limit the length of the nucleic acids referred to and used herein, nor should the terms be used to limit the nature of the polymer backbone to which the nucleobases are attached.
- nucleic acid sequence or “polynucleotide sequence” refers to a contiguous string of nucleotide bases and in particular contexts also refers to the particular placement of nucleotide bases in relation to each other as they appear in a polynucleotide.
- Nucleobase means a heterocyclic moiety capable of non-covalently pairing with another nucleobase.
- Nucleoside means a nucleobase linked to a sugar moiety.
- Nucleotide means a nucleoside having a phosphate group covalently linked to the sugar portion of a nucleoside. In some embodiments, the nucleotide is characterized as being modified if the 3' phosphate group is covalently linked to a contiguous nucleotide by any linkage other than a phosphodiester bond.
- “Compound comprising a modified oligonucleotide consisting of a number of linked nucleosides” means a compound that includes a modified oligonucleotide having the specified number of linked nucleosides. Thus, the compound may include additional substituents or conjugates. Unless otherwise indicated, the compound does not include any additional nucleosides beyond those of the modified oligonucleotide.
- Modified oligonucleotide means an oligonucleotide having one or more modifications relative to a naturally occurring terminus, sugar, nucleobase, and/or internucleoside linkage.
- a modified oligonucleotide may comprise unmodified nucleosides.
- Single-stranded modified oligonucleotide means a modified oligonucleotide which is not hybridized to a complementary nucleic acid strand.
- Modified nucleoside means a nucleoside having any change from a naturally occurring nucleoside.
- a modified nucleoside may have a modified sugar, and an unmodified nucleobase.
- a modified nucleoside may have a modified sugar and a modified nucleobase.
- a modified nucleoside may have a natural sugar and a modified nucleobase.
- a modified nucleoside is a bicyclic nucleoside.
- a modified nucleoside is a non-bicyclic nucleoside.
- nullomers refers to expressed oligonucleotide sequences in a species, the genetic templates of which are congenitally absent in the species.
- nullomers of the disclosure are nullomers not present in the published human genome sequences.
- nullomers of the disclosure are nullomers not present in the published human genome sequences and associated with one or a plurality of cancers.
- one or more of includes at least one of the recited components, or 2, 3, 4, 5, or 5 etc. of the recited components.
- the phase includes all of the recited components.
- Ranges provided herein are understood to include all individual integer values and all subranges within the ranges.
- sample refers to a biological sample obtained or derived from a source of interest, as described herein.
- a source of interest comprises an organism, such as an animal or human.
- a biological sample comprises biological tissue or fluid.
- a biological sample may be or comprise bone marrow, blood, blood cells, cells from a hair sample, ascites, tissue or fine needle biopsy samples, cell-containing body fluids, free floating nucleic acids, sputum, saliva or spit, urine, cerebrospinal fluid, peritoneal fluid, pleural fluid, feces, lymph, gynecological fluids, skin swabs, vaginal swabs, oral swabs, nasal swabs, washings or lavages such as a ductal lavages or broncheoalveolar lavages, aspirates, scrapings, bone marrow specimens, tissue biopsy specimens, surgical specimens, feces, other body fluids, secretions and/or excretions, and/or cells therefrom, etc.
- the sample is a brush biopsy, puncture biopsy, or fluid from a needle biopsy.
- the sample is blood or blood cells.
- the sample is cells from a hair sample or nucleic acids from a hair sample.
- the sample is sputum, saliva or spit.
- a biological sample is or comprises cells obtained from an individual.
- a sample is a “primary sample” obtained directly from a source of interest by any appropriate means.
- a primary biological sample is obtained by methods selected from the group consisting of biopsy (e.g., fine needle aspiration or tissue biopsy), surgery, collection of body fluid (e.g., blood, lymph, feces etc.), etc.
- sample refers to a preparation that is obtained by processing (e.g., by removing one or more components of and/or by adding one or more agents to) a primary sample. For example, filtering using a semi-permeable membrane.
- processing e.g., by removing one or more components of and/or by adding one or more agents to
- a primary sample may comprise, for example nucleic acids or proteins extracted from a sample or obtained by subjecting a primary sample to techniques such as amplification or reverse transcription of mRNA, isolation and/or purification of certain components, etc.
- minimal residual disease refers to a small number of cancer cells remaining in the body after treatment or surgical intervention. These cells cannot usually be detected by standard scans or tests, due to lower abundance than detection sensitivity thresholds.
- a “score” is a value or set of values selected so as to provide a normalized quantitative measure of a variable or characteristic of a subject’ s condition, and/or to discriminate, differentiate or otherwise characterize a subject’s condition.
- the value(s) comprising the score can be based on, for example, quantitative data resulting in a measured amount of one or more sample constituents obtained from the subject, or from clinical parameters, or from clinical assessments, or any combination thereof.
- the score can be derived from a single constituent, parameter or assessment, while in other embodiments the score is derived from multiple constituents, parameters and/or assessments.
- the score can be based upon or derived from an interpretation function; e.g., an interpretation function derived from a particular predictive model using any of various statistical algorithms known in the art.
- a “change in score” can refer to the absolute change in score, e.g. from one time point to the next, or the percent change in score, or the change in the score per unit time (i.e., the rate of score change).
- the score is calculated through an interpretation function or algorithm.
- the subject is suspected of having expression of a gene that promotes or contributes to the likelihood of acquiring a disease state or whose expression is correlative to the presence of a pathogen. Calculation of score can be accomplished using known algorithms executable in computer program products within equipment used in sequencing or analyzing samples.
- the methods disclosed herein comprise substeps of detecting the presence, absence or quantity of a given biomarker by calculating the quantity of a probe in a control sample, calculating the quantity of a probe in the subject sample, and normalizing the signal obtained from the subject sample by subtracting the signal obtained from the control sample.
- sequence identity is determined by using the stand-alone executable BLAST engine program for blasting two sequences (bl2seq), which can be retrieved from the National Center for Biotechnology Information (NCBI) ftp site, using the default parameters (Tatusova and Madden, FEMS Microbiol Lett., 1999, 174, 247-250; which is incorporated herein by reference in its entirety).
- NCBI National Center for Biotechnology Information
- % sequence identity can be determined using the EMBOSS Pairwise Alignment Algorithms tool available from The European Bioinformatics Institute (EMBL-EBI), which is part of the European Molecular Biology Laboratory (EMBL).
- This tool is accessible at the website ebi.ac.uk/Tools/emboss/align/.
- This tool utilizes the Needleman-Wunsch global alignment algorithm (Needleman, S. B. and Wunsch, C. D. (1970) J. Mol. Biol. 48, 443-453; Kruskal, J. B. (1983) An overview of sequence comparison, In D. Sankoff and B. Kruskal, (ed.), Time warps, string edits and macromolecules: the theory and practice of sequence comparison, pp. 1-44, Addison Wesley). Default settings are utilized which include Gap Open: 10.0 and Gap Extend 0.5. The default matrix “Blosum62” is utilized for amino acid sequences and the default matrix “DNAfull” is utilized for nucleic acid sequences.
- the term “statistically significant” means an observed alteration is greater than what would be expected to occur by chance alone (e.g., a “false positive”).
- Statistical significance can be determined by any of various methods well-known in the art. An example of a commonly used measure of statistical significance is the p-value. The p-value represents the probability of obtaining a given result equivalent to a particular datapoint, where the datapoint is the result of random chance alone. A result is often considered highly significant (not random chance) at a p-value less than or equal to about 0.05.
- subject refers to a vertebrate, preferably a mammal, more preferably a human.
- Mammals include, but are not limited to, murine, simians, humans, farm animals, cows, pigs, goats, sheep, horses, dogs, sport animals, and pets.
- Tissues, cells and their progeny obtained in vivo or cultured in vitro are also encompassed by the definition of the term “subject.”
- the subject is a human.
- the term “patient” may be interchangeably used for treatment of those conditions which are specific for a specific subject, such as a human being.
- the term “patient” will refer to human patients suffering from a particular disease or disorder.
- the subject may be a non-human animal.
- the term “mammal” encompasses both humans and non-humans and includes but is not limited to humans, non-human primates, canines, felines, murine, bovines, equines, caprine, and porcines.
- nucleic acid molecule comprises at least about 50% sequence identity to a reference nucleic acid sequence (for example, any one of the nucleic acid sequences described herein) or amino acid sequence. In some embodiments, such a sequence is at least about 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, or even 99% identical at the nucleic acid level or amino acid level to the reference sequence used for comparison.
- terapéutica means an agent utilized to treat, combat, ameliorate, prevent or improve an unwanted condition or disease of a patient.
- terapéuticaally effective amount means a quantity sufficient to achieve a desired therapeutic effect, for example, an amount which results in the prevention or amelioration of or a decrease in the symptoms associated with a disease that is being treated, e.g., disorders associated with cancer growth or a hyperproliferative disorder.
- the amount of compound administered to the subject will depend on the type and severity of the disease and on the characteristics of the individual, such as general health, age, sex, body weight and tolerance to drugs. It will also depend on the degree, severity and type of disease. The skilled artisan will be able to determine appropriate dosages depending on these and other factors.
- the regimen of administration can affect what constitutes an effective amount.
- an effective amount of the compounds of the present disclosure sufficient for achieving a therapeutic effect, range from about 0.000001 mg per kilogram body weight per day to about 10,000 mg per kilogram body weight per day.
- the dosage ranges are from about 0.0001 mg per kilogram body weight per day to about 100 mg per kilogram body weight per day.
- the compounds disclosed herein can also be administered in combination with each other, or with one or more additional therapeutic compounds.
- beneficial or desired clinical results include, but are not limited to, one or more of the following: (1) preventing or delaying the appearance of clinical symptoms of the state, disorder, or condition developing in a person who may be afflicted with or predisposed to the state, disorder or condition but does not yet experience or display clinical symptoms of the state, disorder or condition; (2) inhibiting the state, disorder or condition, i.e., arresting, reducing or delaying the development of the disease or a relapse thereof (in case of maintenance treatment) or at least one clinical symptom, sign, or test, thereof; or (3) relieving the disease, i.e., causing regression of the state, disorder or condition or at least one of its clinical or sub-clinical symptoms or signs.
- a subject is successfully “treated” according to the methods of the present disclosure if the patient shows one or more of the following: a reduction in the number of and/or complete absence of cancer cells; a reduction in the tumor size; an inhibition of tumor growth; inhibition of and/or an absence of cancer cell infiltration into peripheral organs including the spread of cancer cells into soft tissue and bone; inhibition of and/or an absence of tumor or cancer cell metastasis; inhibition and/or an absence of cancer growth; relief of one or more symptoms associated with the specific cancer; reduced morbidity and mortality; improvement in quality of life; reduction in tumorigenicity; reduction in the number or frequency of cancer stem cells; or some combination of such effects.
- tumor refers to all neoplastic cell growth and proliferation, whether malignant or benign, and all pre-cancerous and cancerous cells and tissues.
- a “benign” tumor is not cancerous and it does not invade nearby tissue or spread to other parts of the body.
- a “premalignant” tumor is a tumor which is not yet cancerous but has the potential to become malignant.
- a “malignant” tumor is cancerous and can grow and spread to other parts of the body.
- tumor sample refers to a sample comprising tumor material obtained from a cancer patient.
- the term encompasses tumor tissue samples, for example, tissue obtained by surgical resection and tissue obtained by biopsy, such as for example, a core biopsy or a fine needle biopsy.
- the tumor sample is a fixed, wax-embedded tissue sample, such as a formalin-fixed, paraffin-embedded tissue sample.
- tumor sample encompasses a sample comprising tumor cells obtained from sites other than the primary tumor, e.g., circulating tumor cells.
- the term also encompasses cells that are the progeny of the patient’s tumor cells, e.g. cell culture samples derived from primary tumor cells or circulating tumor cells.
- the term further encompasses samples that may comprise protein or nucleic acid material shed from tumor cells in vivo, e.g., bone marrow, blood, plasma, serum, and the like.
- the identification of nullomers can be performed using any methods known in the art.
- the identification of nullomers of the disclosure is performed as previously described in Georgakopoulos-Soares et al., published in bioRxiv, available at biorxiv.org/content/10.1101/2020.03.02.972422vl, incorporated by reference herein.
- a dataset is obtained.
- the dataset is obtained from WGS cancers from ICGC under the project PanCancer Analysis of Whole Genomes (ICGC/TCGA Pan-Cancer Analysis of Whole Genomes Consortium. Pan-cancer analysis of whole genomes, Nature, 2020, 578:82-93), which includes 46 cancer projects from 21 organs.
- WGS patients were analyzed using the GRCh37 (hg 19) reference assembly of the human genome.
- somatic indel calls are performed using three pipelines from four somatic variant callers. These are the Wellcome Sanger Institute pipeline, the DKFZ/ EMBL pipeline and the Broad Institute pipeline, with somatic variant false discovery rate of about 2.5%.
- indel calling is performed by those algorithms and only indels called by at least two of the callers were analyzed, therefore generating a conservative dataset. As a result, the false negative rate of indel detection can be higher than that of other methods, and of each pipeline separately, which implies that many indels present in the samples were not identified successfully.
- the indel calls are visually examined using JBrowse Genome Browser32, to inspect the number of reads reporting the indel, if the indel calls are biased towards the end of the sequencing reads or if there were other systematic biases between the normal and tumor sequencing reads; such biases could not be identified.
- Bedtools intersect utility is used to measure overlap between indels and polyN tracts.
- overlap in this context refers to deleted bases occurring at any position across the entire length of the repeat or inserted bases occurring at any position across the length of the repeat and immediately before or after the repeat.
- Indel density is defined as the number of indel mutations for a given number of bases.
- the distance between each pair of consecutive indels is calculated per patient. In some embodiments, indels in different chromosomes are excluded because their pairwise distance cannot be defined. In some embodiments, the same analysis is performed separately for insertions and deletions.
- substitution calling is performed using four somatic mutationcalling algorithms, with mutation calls being shared by at least two algorithms.
- C > A substitutions can be examined with respect to transcriptional strand asymmetries at polyG tracts and replication timing.
- the numbers of indels overlapping motifs found in the template or non-template strands are obtained using the bedtools intersect command.
- strand bias is calculated for the vector of genes, reporting the number of polyN motif occurrences and the number of overlapping motifs as:
- A (indels overlapping motif at non-template)/(motif occurrences at non-template)
- B (indels overlapping motif at template)/(motif occurrences at template)
- Strand bias A/(A + B) with motifs representing polyN repeat tracts of size 2-10 bp and dinucleotide repeat tracts of 1-5 repeated units, at genic regions.
- bootstrapping with replacement randomly selecting the indels overlapping motifs at template and non-template strands from each randomly selected gene are performed for equal number of genes in multiple iterations, from which the standard deviation for the strand bias can be calculated.
- the nullomers can be of any length. In some embodiments, the nullomers are in a length of from about 8 to about 50 nucleotides. In some embodiments, the nullomers are in a length of from about 10 to about 45 nucleotides. In some embodiments, the nullomers are in a length of from about 12 to about 40 nucleotides. In some embodiments, the nullomers are in a length of from about 14 to about 30 nucleotides. In some embodiments, the nullomers are in a length of from about 16 to about 20 nucleotides. In some embodiments, the nullomers are in a length of from about 8 nucleotides.
- the nullomers are in a length of about 18 nucleotides. In some embodiments, the nullomers are in a length of about 19 nucleotides. In some embodiments, the nullomers are in a length of about 20 nucleotides. In some embodiments, the nullomers are in a length of about 25 nucleotides. In some embodiments, the nullomers are in a length of about 30 nucleotides. In some embodiments, the nullomers are in a length of about 35 nucleotides. In some embodiments, the nullomers are in a length of about 40 nucleotides. In some embodiments, the nullomers are in a length of about 45 nucleotides. In some embodiments, the nullomers are in a length of about 50 nucleotides. In some embodiments, the nullomers are in a length of more than about 50 nucleotides.
- the disclosure provides nullomers identified in cancers of numerous organs or tissues, including pancreas, esophagus, lymphoid, kidney, ovary, head and neck, lung, stomach, liver, CNS, uterus, skin, colorectal, prostate, bladder, bone and soft tissue, breast, biliary, cervix, thyroid and myeloid.
- the neomers of the disclosure are provided in Table 1 or Table B.
- the disclosure relates to a nullomer comprising at least about 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89% 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 97%, 98, 99% or 100% sequence identity to any of the sequences provided in Table 1 or Table B.
- the disclosure relates to a neomer comprising any of the sequences provided in Table 1 or Table B.
- the disclosure relates to a nucleic acid sequence that is complementary to any of the sequences provided in Table 1 or Table B.
- the disclosure relates to relates to a system comprising a solid support and one or a plurality of nucleic acid sequences that is complementary to any of the sequences provided in Table 1. In some embodiments, the disclosure relates to relates to a system comprising a solid support and one or a plurality of nucleic acid sequences that is complementary to any of the sequences provided in Table 1 immobilized to a surface of the solid support.
- PCT/US2022/027536 describes nullomer detection steps and is incorporated by reference in its entirety herein.
- the expression level of one or more disclosed nullomers can be determined in a biological sample obtained from a subject.
- a sample of a subject is one that originates from a subject. Such a sample may be further processed after it is obtained from the subject.
- DNA or RNA may be isolated from a sample.
- the DNA or RNA isolated from the sample is also a sample obtained from the subject.
- a biological sample useful for determining the level of one or more disclosed nullomers may be obtained from essentially any source, including cells, blood, hair, tissues, and fluids throughout the body.
- the biological sample used for determining the level of one or more disclosed nullomers is a sample.
- the sample comprises circulating nullomers, e.g., extracellular nullomers.
- Extracellular nullomers freely circulate in a wide range of biological material, including bodily fluids, such as fluids from the circulatory system, e.g., a blood sample or a lymph sample, or from another bodily fluid such as urine or saliva or serum.
- the biological sample used for determining the level of one or more disclosed nullomers is a bodily fluid, for example, blood, fractions thereof, serum, plasma, urine, saliva, tears, sweat, semen, vaginal secretions, lymph, bronchial secretions, CSF, whole blood, etc.
- the sample is a sample that is obtained non-invasively.
- the sample is whole blood or blood cells.
- the sample is cells from a hair sample or nucleic acids from a hair sample.
- the sample is sputum, saliva or spit.
- the sample is a serum sample from a human.
- the sample is a bodily fluid from a human.
- the sample is a liquid biopsy from a human.
- the sample is free of cells but comprises cell free DNA or RNA.
- any of the methods disclosed herein comprise using a small volume of sample for detection and/or diagnosis.
- the sample used in any of the disclosed methods has a volume of no more than about 100 microliters of fluid. In some embodiments, the sample has a volume of no more than about 90 microliters of fluid. In some embodiments, the sample has a volume of no more than about 80 microliters of fluid. In some embodiments, the sample has a volume of no more than about 70 microliters of fluid. In some embodiments, the sample has a volume of no more than about 60 microliters of fluid. In some embodiments, the sample has a volume of no more than about 50 microliters of fluid.
- the sample has a volume of no more than about 40 microliters of fluid. In some embodiments, the sample has a volume of no more than about 30 microliters of fluid. In some embodiments, the sample has a volume of no more than about 20 microliters of fluid. In some embodiments, the sample has a volume of no more than about 10 microliters of fluid. In some embodiments, the sample has a volume of no more than about 5 microliters of fluid. In some embodiments, the sample has a volume of no more than about 1 microliters of fluid.
- the disclosed methods comprise isolating total DNA or RNA and/or amplifying nullomers in a sample of no more than about 5 microliters, no more than about 10 microliters, no more than about 20 microliters, no more than about 40 microliters, no more than about 80 microliters, no more than about 100 microliters, no more than about 200 microliters, no more than about 300 microliters, no more than about 400 microliters, no more than about 500 microliters, no more than about 600 microliters, no more than about 700 microliters, no more than about 800 microliters, no more than about 900 microliters, no more than about 1 milliliter, no more than about 1.1 milliliters, no more than about 1.2 milliliters, no more than about 1.3 milliliters, no more than about 1.4 milliliters, no more than about 1.5 milliliters, no more than about 1.6 milliliters, no more than about 1.7 milliliters, no more than about 1.8 milliliters
- the sample size is from about 1 microliters to about 2 milliliters, from about 20 microliters to about 2 milliliters, from about 5 microliters to about 1.5 milliliters, from about 10 microliters to about 500 microliters, from about 15 microliters to about 300 microliters, from about 20 microliters to about 200 microliters, from about 30 microliters to about 100 microliters, from about 1 microliters to about
- microliters from about 5 microliters to about 75 microliters, or from about 10 microliters to about 50 microliters of liquid sample in the form of subject plasma, whole blood, blood cells, cells from a hair sample, saliva or spit, or serum.
- the methods disclosed herein comprise isolating total DNA or RNA and/or amplifying nullomers in a sample of no more than about 5 microliters of serum, no more than about 10 microliters of serum, no more than about 20 microliters of serum, no more than about 40 microliters of serum, no more than about 80 microliters of serum, no more than about 100 microliters of serum, no more than about 200 microliters of serum, no more than about 300 microliters of serum, no more than about 400 microliters of serum, no more than about 500 microliters of serum, no more than about 600 microliters of serum, no more than about 700 microliters of serum, no more than about 800 microliters of serum, no more than about 900 microliters of serum, no more than about 1 milliliter of serum, no more than about 1.1 milliliters of serum, no more than about 1.2 milliliters of serum, no more than about 1.3 milliliters of serum, no more than about 1.4 milliliters of serum, no more
- Circulating nullomers include nullomers in cells, extracellular nullomers in microvesicles, in exosomes and extracellular nullomers that are not associated with cells or microvesicles (extracellular, non-vesicular nullomers).
- the biological sample used for determining the level of one or more nullomers may contain cells.
- the biological sample may be free or substantially free of cells (e.g., a serum sample).
- a sample containing circulating nullomers, e.g., extracellular nullomers is a blood-derived sample.
- Exemplary blood-derived sample types include, e.g., a plasma sample, a serum sample, a blood sample, etc.
- a sample containing circulating nullomers is a lymph sample. Circulating nullomers are also found in urine and saliva, and biological samples derived from these sources are likewise suitable for determining the level of one or more disclosed nullomers.
- any of the methods of the disclosure comprises a step of isolating total DNA or RNA from a sample or cell or exosome or microvesicle.
- Methods of isolating DNA or RNA for expression analysis from blood, plasma and/or serum see for example, Tsui NB et al. (2002) Clin. Chem. 48,1647-53, incorporated by reference in its entirety herein
- urine see for example, Boom R et al. (1990) J Clin Microbiol. 28, 495-503, incorporated by reference in its entirety herein
- the level of one or more disclosed nullomers in a biological sample can be determined by any suitable method. Any reliable method for measuring the level or amount of a nullomer in a sample can be used.
- nullomers can be detected and quantified from a sample (including fractions thereof), such as samples of isolated DNA or RNA by various methods known for DNA or mRNA, including, for example, amplification-based methods (e.g., Polymerase Chain Reaction (PCR), Real-Time Polymerase Chain Reaction (RT-PCR), Quantitative Polymerase Chain Reaction (qPCR), rolling circle amplification, etc.), hybridization-based methods (e.g., hybridization arrays (e.g., microarrays), NanoString analysis, Northern Blot analysis, branched DNA (bDNA) signal amplification, in situ hybridization, etc.), and sequencing-based methods (e.g., next-generation sequencing methods, for example, using the Illumina or lonTorrent platforms).
- Other exemplary techniques include ribonucleas
- RNA is converted to DNA (cDNA) prior to analysis.
- cDNA can be generated by reverse transcription of isolated RNA using conventional techniques.
- nullomer is amplified prior to measurement.
- the level of nullomer is measured during the amplification process.
- the level of nullomer is not amplified prior to measurement.
- amplification-based methods exist for detecting the level of nullomers, including, but not limited to, PCR, RT-PCR, qPCR, and rolling circle amplification.
- Other amplificationbased techniques include, for example, ligase chain reaction, multiplex ligatable probe amplification, in vitro transcription (IVT), strand displacement amplification, transcription- mediated amplification, RNA (Eberwine) amplification, and other methods that are known to persons skilled in the art.
- a typical PCR reaction includes multiple steps, or cycles, that selectively amplify target nucleic acid species: a denaturing step, in which a target nucleic acid is denatured; an annealing step, in which a set of PCR primers (i.e., forward and reverse primers) anneal to complementary DNA strands, and an elongation step, in which a thermostable DNA polymerase elongates the primers. By repeating these steps multiple times, a DNA fragment is amplified to produce an amplicon, corresponding to the target sequence.
- Typical PCR reactions include 20 or more cycles of denaturation, annealing, and elongation.
- a reverse transcription reaction (which produces a cDNA sequence having complementarity to a RNA) may be performed prior to PCR amplification.
- Reverse transcription reactions include the use of, e.g., a RNA-based DNA polymerase (reverse transcriptase) and a primer.
- Kits for quantitative real time PCR of nullomers are known, and are commercially available. Examples of suitable kits include, but are not limited to, the TaqMan mRNA Assay (Applied Biosystems) and the mirVana qRT-PCR nullomer detection kit (Ambion).
- the RNA can be ligated to a single stranded oligonucleotide containing universal primer sequences, a polyadenylated sequence, or adaptor sequence prior to reverse transcriptase and amplified using a primer complementary to the universal primer sequence, poly(T) primer, or primer comprising a sequence that is complementary to the adaptor sequence.
- custom qRT-PCR assays can be developed for determination of nullomer levels.
- Custom qRT-PCR assays to measure nullomers in a biological sample e.g., a body fluid
- Custom nullomer assays can be tested by running the assay on a dilution series of chemically synthesized nullomer corresponding to the target sequence. This permits determination of the limit of detection and linear range of quantitation of each assay.
- these data permit an estimate of the absolute abundance of nullomers measured in biological samples.
- Amplification curves may optionally be checked to verify that Ct values are assessed in the linear range of each amplification plot.
- the linear range spans several orders of magnitude.
- a chemically synthesized version of the nullomer can be obtained and analyzed in a dilution series to determine the limit of sensitivity of the assay, and the linear range of quantitation.
- Relative expression levels may be determined, for example, as described by Livak et al., Methods (2001) December; 25(4):402-8.
- two or more nullomers are amplified in a single reaction volume.
- multiplex q-PCR such as qRT-PCR, enables simultaneous amplification and quantification of at least two nullomers of interest in one reaction volume by using more than one pair of primers and/or more than one probe.
- the primer pairs comprise at least one amplification primer that specifically binds each nullomer, and the probes are labeled such that they are distinguishable from one another, thus allowing simultaneous quantification of multiple nullomers.
- Rolling circle amplification is a DNA-polymerase driven reaction that can replicate circularized oligonucleotide probes with either linear or geometric kinetics under isothermal conditions (see, for example, Lizardi et al., Nat. Gen. (1998) 19(3):225-232; Gusev et al., Am. J. Pathol. (2001) 159(l):63-69; Nallur et al., Nucleic Acids Res. (2001) 29(23):E118).
- a complex pattern of strand displacement results in the generation of over 10 9 copies of each DNA molecule in 90 minutes or less.
- Tandemly linked copies of a closed circle DNA molecule may be formed by using a single primer. The process can also be performed using a matrix-associated DNA. The template used for rolling circle amplification may be reverse transcribed. This method can be used as a highly sensitive indicator of nullomer sequence and expression level at very low nullomer concentrations (see, for example, Cheng et al., Angew Chem. Int. Ed. Engl. (2009) 48(18)3268-72; Neubacher et al., Chembiochem. (2009) 10(8): 1289-
- the disclosure provide a method for identifying the presence, absence, or quantity of one or a plurality of the disclosed nullomers comprising: a) isolating nucleic acids from a sample; and b) mixing the nucleic acids with one or a plurality of primers under conditions and for a period of time sufficient to allow amplification of the one or plurality nullomers, wherein the one or plurality of primers comprises sequences that are complementary to any of the nullomers provided in Table 1.
- the nucleic acid from a sample is cell-free (cfDNA).
- the nucleic acid from a sample is circulating tumor
- the primer used in the disclosed method comprises from about 6 to about 16 nucleotides. In some embodiments, the primer used in the disclosed method comprises from about 7 to about 15 nucleotides. In some embodiments, the primer used in the disclosed method comprises from about 8 to about 14 nucleotides. In some embodiments, the primer used in the disclosed method comprises about 6 nucleotides.
- the primer used in the disclosed method comprises about 7 nucleotides, In some embodiments, the primer used in the disclosed method comprises about 8 nucleotides, In some embodiments, the primer used in the disclosed method comprises about 9 nucleotides, In some embodiments, the primer used in the disclosed method comprises about 10 nucleotides, In some embodiments, the primer used in the disclosed method comprises about 11 nucleotides, In some embodiments, the primer used in the disclosed method comprises about 12 nucleotides, In some embodiments, the primer used in the disclosed method comprises about 13 nucleotides, In some embodiments, the primer used in the disclosed method comprises about 14 nucleotides, In some embodiments, the primer used in the disclosed method comprises about 15 nucleotides, In some embodiments, the primer used in the disclosed method comprises about 16 nucleotides.
- the identification of the presence or quantity of one or a plurality of the disclosed nullomers is indicative that the subject from which the sample is obtained has the cancer type corresponding to the particular nullomer identified in Table 1.
- Hybridization -Based Methods Nullomers may be detected using hybridization-based methods, including but not limited to hybridization arrays (e.g., microarrays), NanoString analysis, Southern Blot analysis, Northern Blot analysis, branched DNA (bDNA) signal amplification, and in situ hybridization.
- hybridization arrays e.g., microarrays
- NanoString analysis e.g., Southern Blot analysis
- Northern Blot analysis e.g., Northern Blot analysis
- bDNA branched DNA
- Microarrays can be used to measure the levels of large numbers of nullomers simultaneously.
- Microarrays can be fabricated using a variety of technologies, including printing with fine-pointed pins onto glass slides, photolithography using pre-made masks, photolithography using dynamic micromirror devices, ink-jet printing, or electrochemistry on microelectrode arrays.
- microfluidic TaqMan Low-Density Arrays which are based on an array of microfluidic qRT-PCR reactions, as well as related microfluidic qRT-PCR based methods.
- Axon B-4000 scanner and Gene-Pix Pro 4.0 software or other suitable software can be used to scan images. Non-positive spots after background subtraction, and outliers detected by the ESD procedure, are removed. The resulting signal intensity values are normalized to per-chip median values and then used to obtain geometric means and standard errors for each nullomer. Each signal can be transformed to log base 2, and a one-sample t test can be conducted. Independent hybridizations for each sample can be performed on chips with each nullomer spotted multiple times to increase the robustness of the data.
- Microarrays can be used for the expression profiling of nullomers in diseases.
- DNA or RNA can be extracted from a sample and, optionally, the nullomers are size- selected from total DNA or RNA.
- Oligonucleotide linkers can be attached to the 5’ and 3’ ends of the nullomers and the resulting ligation products are used as templates for an RT-PCR reaction.
- the sense strand PCR primer can have a fluorophore attached to its 5’ end, thereby labeling the sense strand of the PCR product.
- the PCR product is denatured and then hybridized to the microarray.
- probes of the disclosure are nucleic acid sequences comprising from about 10 to about 20 nucleotides in length and are DNA or RNA or NDA/RNA hybrid seqeunces complementary to a nullomer of Table 1, Table 5, Table 6 or Table B.
- the disclosure relate to composition comprising one or a plurality f such probes.
- those probes comprise a fluorescent probe detectable when exposed to light emitted onto the probe.
- the fluorescence intensity of each spot is then evaluated in terms of the number of copies of a particular nullomer, using a number of positive and negative controls and array data normalization methods, which will result in assessment of the level of expression of a particular nullomer.
- Total RNA containing the nullomers extracted from a body fluid sample can also be used directly without size-selection of the nullomers.
- the RNA can be 3’ end labeled using T4 RNA ligase and a fluorophore-labeled short RNA linker.
- Fluorophore-labeled nullomers complementary to the corresponding nullomer capture probe sequences on the array hybridize, via base pairing, to the spot at which the capture probes are affixed. The fluorescence intensity of each spot is then evaluated in terms of the number of copies of a particular nullomer, using a number of positive and negative controls and array data normalization methods, which will result in assessment of the level of expression of a particular nullomer.
- microarrays can be employed including, but not limited to, spotted oligonucleotide microarrays, pre-fabricated oligonucleotide microarrays or spotted long oligonucleotide arrays.
- Nullomers can also be detected without amplification using the nCounter Analysis System (NanoString Technologies, Seattle, Wash.). This technology employs two nucleic acid-based probes that hybridize in solution (e.g., a reporter probe and a capture probe). After hybridization to a nullomers disclosed herein, excess probes are removed, and probe/target complexes are analyzed in accordance with the manufacturer’s protocol. nCounter nullomer assay kits are available from NanoString Technologies, which are capable of distinguishing between highly similar nullomers with great specificity.
- Nullomers can also be detected using branched DNA (bDNA) signal amplification (see, for example, Urdea, Nature Biotechnology (1994), 12:926-928).
- RNA assays based on bDNA signal amplification are commercially available.
- One such assay is the QuantiGene.RTM. 2.0 nullomer Assay (Affymetrix, Santa Clara, Calif.).
- Southern Blot, Northern Blot and in situ hybridization may also be used to detect nullomers. Suitable methods for performing Southern Blot, Northern Blot and in situ hybridization are known in the art.
- biomarker expression is determined by an assay known to those of skill in the art, including but not limited to, multi-analyte profile test, enzyme-linked immunosorbent assay (ELISA), radioimmunoassay, Western blot assay, immunofluorescent assay, enzyme immunoassay, immunoprecipitation assay, chemiluminescent assay, immunohistochemical assay, dot blot assay, or slot blot assay.
- an antibody is used in the assay the antibody is detectably labeled.
- the antibody labels may include, but are not limited to, immunofluorescent label, chemiluminescent label, phosphorescent label, enzyme label, radiolabel, avidin/biotin, colloidal gold particles, colored particles, and magnetic particles.
- biomarker expression is determined by an IHC assay.
- biomarker expression is determined using an agent that specifically binds the biomarker.
- Any molecular entity that displays specific binding to a biomarker can be employed to determine the level of that biomarker protein in a sample.
- Specific binding agents include, but are not limited to, antibodies, antibody fragments, antibody mimetics, and polynucleotides (e.g., aptamers).
- polynucleotides e.g., aptamers
- the disclosure relates to a system comprising a solid support (such as an ELISA plate, gel, bead or column comprising an antibody, antibody fragment, antibody mimetic, and/or polynucleotides capable of binding to T3p or a salt thereof.
- a solid support such as an ELISA plate, gel, bead or column comprising an antibody, antibody fragment, antibody mimetic, and/or polynucleotides capable of binding to T3p or a salt thereof.
- nullomers can be detected using Illumina. Next Generation Sequencing (e.g., Sequencing-By-Synthesis or TruSeq methods, using, for example, the HiSeq, HiScan, GenomeAnalyzer, or MiSeq systems (Illumina, Inc., San Diego, Calif.)). Nullomers can also be detected using Ion Torrent Sequencing (Ion Torrent Systems, Inc., Gulliford, Conn.), or other suitable methods of semiconductor sequencing.
- Next Generation Sequencing e.g., Sequencing-By-Synthesis or TruSeq methods, using, for example, the HiSeq, HiScan, GenomeAnalyzer, or MiSeq systems (Illumina, Inc., San Diego, Calif.)
- Nullomers can also be detected using Ion Torrent Sequencing (Ion Torrent Systems, Inc., Gulliford, Conn.), or other suitable methods of semiconductor sequencing.
- RNA endonucleases RNases
- MS/MS tandem MS
- the first approach developed utilized the on-line chromatographic separation of endonuclease digests by reversed phase HPLC coupled directly to ESLMS.
- the presence of posttranscriptional modifications can be revealed by mass shifts from those expected based upon the RNA sequence. Ions of anomalous mass/charge values can then be isolated for tandem MS sequencing to locate the sequence placement of the posttranscriptionally modified nucleoside.
- MALDI-MS Matrix -assisted laser desorption/ionization mass spectrometry
- MALDI-MS has also been used as an analytical approach for obtaining information about posttranscriptionally modified nucleosides.
- MALDI-based approaches can be differentiated from ESI-based approaches by the separation step.
- the mass spectrometer is used to separate the nullomers.
- a system of capillary LC coupled with nanoESI-MS can be employed, by using a linear ion trap-orbitrap hybrid mass spectrometer (LTQ Orbitrap XL, Thermo Fisher Scientific) or a tandem-quadrupole time-of-flight mass spectrometer (QSTAR XL, Applied Biosystems) equipped with a custom-made nanospray ion source, a Nanovolume Valve (Valeo Instruments), and a splitless nano HPLC system (DiNa, KYA Technologies). Analyte/TEAA is loaded onto a nano-LC trap column, desalted, and then concentrated.
- LTQ Orbitrap XL linear ion trap-orbitrap hybrid mass spectrometer
- QSTAR XL tandem-quadrupole time-of-flight mass spectrometer
- Analyte/TEAA is loaded onto a nano-LC trap column, desalted, and then concentrated.
- Intact nullomers are eluted from the trap column and directly injected into a Cl 8 capillary column, and chromatographed by RP-HPLC using a gradient of solvents of increasing polarity.
- the chromatographic eluent is sprayed from a sprayer tip attached to the capillary column, using an ionization voltage that allows ions to be scanned in the negative polarity mode.
- nullomer detection and measurement include, for example, strand invasion assay (Third Wave Technologies, Inc.), surface plasmon resonance (SPR), cDNA, MTDNA (metallic DNA; Advance Technologies, Saskatoon, SK), and single-molecule methods such as the one developed by US Genomics.
- Multiple nullomers can be detected in a microarray format using a novel approach that combines a surface enzyme reaction with nanoparticle- amplified SPR imaging (SPRI).
- SPRI nanoparticle- amplified SPR imaging
- the surface reaction of poly(A) polymerase creates poly(A) tails on nullomers hybridized onto locked nucleic acid (LNA) microarrays. DNA-modified nanoparticles are then adsorbed onto the poly(A) tails and detected with SPRI.
- CRISPR-Cas9 complexes can be used to detect the presence of nullomers in vitro based upon exposure of a sample from a patient to sgRNA-Cas protein complex, wherein the sgRNA is complementary to at least a portion of the nullomer sequence.
- the exposure is to genomic DNA within a cancer cell.
- the disclosure relates to a composition or system comprising one or a plurality of sgRNAs that comprise about 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to the sequences of Table 6 or Table B.
- the disclosure relates to a composition or system comprising one or a plurality of sgRNAs that comprise from about 98 to about 110 nucleotides in length with at least one portion of the sgRNA complementary to a nucleic sequence from about 8 to about 18 nucleotides of any nullomer disclosed in Table 1, Table 5, Table 6 or Table B.
- the term “mutagen” means any molecule, a nucleic acid sequence, amino acid sequence, or hybrid amino acid or nucleic acid sequence that causes a mutation or modification in one or more regions of endogenous nucleic acid when exposed for a time period sufficient to cause the mutation.
- the mutation is a point mutation, frameshift mutation, deletion, truncation, or addition.
- the mutagen is a vector or a gene-modifying enzyme.
- vector refers to any genetic element, such as a plasmid, phage, transposon, cosmid, chromosome, artificial chromosome, virus, virion, etc., which is capable of replication when associated with the proper control elements and which can transfer gene sequences between cells.
- vector includes cloning and expression vehicles, as well as viral vectors.
- gene-modifying enzyme refers to an enzyme that is capable of modifying a gene by introducing a mutation (e.g., point mutation, frameshift mutation, deletion, or truncation) causing gene inactivation or introducing heterologous nucleotides (e.g., genes) through non-homologous end joining or homologous recombination.
- exemplary gene-modifying enzymes include but not limited to, a Cas protein, a meganuclease, a transcription activator-like effector nucleases (TALEN), a transposon, a zinc-finger nuclease (ZFN), or a recombinase.
- the gene-modifying enzyme suitable for the methods disclosed herein is a Cas protein, a meganuclease, a TALEN, a ZFN, or a recombinase. In some embodiments, the genemodifying enzyme suitable for the methods disclosed herein is a Cas protein. In some preferred embodiments, the gene-modifying enzyme suitable for the methods disclosed herein is a Cas9 protein.
- Cas9 protein refers to the “clustered, regularly interspaced, short palindromic repeats (CRlSPR)-associated protein 9.” This term is well known in the art and has been described, e.g. in Makarova et al. (2011) Nat. Rev. Microbiol., 9:467-477, and in Makarova et al. (2011) Biol. Direct., 6:38. Cas proteins are endonuclease that form part of an adaptive defense mechanism evolved by bacteria and archaea to protect them from invading viruses and plasmids. Cas9 protein or gene information can be obtained from a known database such as the GenBank of NCBI (National Center for Biotechnology Information), but is not limited thereto.
- the Cas9 protein may comprise not only wild-type Cas9, but also deactivated Cas9 (dCas9), or Cas9 variants such as Cas9 nickase.
- the deactivated Cas9 may be RFN (RNA-guided FokI nuclease) comprising a FokI nuclease domain bound to dCas9, or may be dCas9 to which a transcription activator or repressor domain is bound.
- the Cas9 protein is not limited in its origin.
- the Cas9 protein may be derived from Streptococcus pyogenes, Francisella novicida, Streptococcus thermophilus, Legionella pneumophila, Listeria innocua, or Streptococcus mutans.
- Cas9 protein is the major protein element of the CRISPR/Cas9 system, which forms a complex with crRNA (CRISPR RNA) and tracrRNA (trans-activating crRNA) to form activated endonuclease or nickase.
- CRISPR system refers collectively to transcripts or synthetically produced transcripts and other elements involved in the expression of or directing the activity of CRISPR-associated (“Cas”) genes, including sequences encoding a Cas gene, a tracr (trans- activating CRISPR) sequence (e.g.
- tracrRNA or an active partial tracrRNA encompassing a “direct repeat” and a tracrRNA-processed partial direct repeat in the context of an endogenous CRISPR system
- a guide sequence also referred to as a “spacer” in the context of an endogenous CRISPR system
- other sequences and transcripts from a CRISPR locus e.g., one or more elements of a CRISPR system is derived from a type I, type II, or type III CRISPR system.
- one or more elements of a CRISPR system is derived from a particular organism comprising an endogenous CRISPR system, such as Streptococcus pyogenes.
- a CRISPR system is characterized by elements that promote the formation of a CRISPR complex at the site of a target sequence (also referred to as a protospacer in the context of an endogenous CRISPR system).
- target sequence refers to a nucleic acid sequence to which a guide sequence is designed to have complementarity, where hybridization between a target sequence and a guide sequence promotes the formation of a CRISPR complex.
- Full complementarity is not necessarily required, provided there is sufficient complementarity to cause hybridization and promote formation of a CRISPR complex.
- a target sequence may comprise any polynucleotide, such as DNA or RNA polynucleotides, but in some embodiments, the target sequence is a nullomer or a region of a nullomer that is from about 10 to about 35 nucleotides of the nullomer sequence of any nullomer from Table 1 . In some embodiments, the target sequence is a DNA polynucleotide and is referred to a DNA target sequence.
- a target sequence comprises at least three nucleic acid sequences that are recognized by a Cas-protein when the Cas protein is associated with a CRISPR complex or system which comprises at least one sgRNA or one tracrRNA/crRNA duplex at a concentration and within an microenvironment suitable for association of such a system.
- the target DNA comprises at least one or more proto-spacer adjacent motifs which sequences are known in the art and are dependent upon the Cas protein system being used in conjunction with the sgRNA or crRNA/tracrRNAs employed by this work.
- the target DNA comprises NNG, where G is a guanine and N is any naturally occurring nucleic acid.
- the target DNA comprises any one or combination of NNG, NNA, GAA, NNAGAAW and NGGNG, where G is an guanine, A is adenine, and N is any naturally occurring nucleic acid from one nullomer in Table 1.
- a CRISPR complex comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins
- formation of a CRISPR complex results in cleavage of one or both strands in or near (e.g. within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs from) the target sequence, without wishing to be bound by theory, the tracr sequence, which may comprise or consist of all or a portion of a wild-type tracr sequence (e.g.
- a wild-type tracr sequence may also form part of a CRISPR complex, such as by hybridization along at least a portion of the tracr sequence to all or a portion of a tracr mate sequence that is operably linked to the guide sequence.
- the tracr sequence has sufficient complementarity to a tracr mate sequence to hybridize and participate in formation of a CRISPR complex. As with the target sequence, it is believed that complete complementarity is not needed, provided there is sufficient to be functional (bind the Cas protein or functional fragment thereof).
- the tracr sequence has at least 50%, 60%, 70%, 80%, 90%, 95% or 99% of sequence complementarity along the length of the tracr mate sequence when optimally aligned.
- one or more vectors driving expression of one or more elements of a CRISPR system are introduced into a host cell such that the presence and/or expression of the elements of the CRISPR system direct formation of a CRISPR complex at one or more target sites.
- a Cas enzyme, a guide sequence linked to a tracr-mate sequence, and a tracr sequence could each be operably linked to separate regulatory elements on separate vectors.
- the target site is a genomic DNA of a cancer cell within the host or a cancer cell isolated from the subject in a sample or within a system independent of a tumor.
- the guide sequence or RNA or DNA sequences that form a CRISPR complex are at least partially synthetic.
- the CRISPR system elements that are combined in a single vector may be arranged in any suitable orientation, such as one element located 5’ with respect to (“upstream” of) or 3 ’ with respect to (“downstream” of) a second element.
- the disclosure relates to a composition comprising a chemically synthesized guide sequence.
- the chemically synthesized guide sequence is used in conjunction with a vector comprising a coding sequence that encodes a CRISPR enzyme, such as a type II Cas9 protein.
- the chemically synthesized guide sequence is used in conjunction with one or more vectors, wherein each vector comprises a coding sequence that encodes a CRISPR enzyme, such as a type II Cas9 protein.
- the coding sequence of one element may be located on the same or opposite strand of the coding sequence of a second element, and oriented in the same or opposite direction.
- a single promoter drives expression of a transcript encoding a CRISPR enzyme and one or more additional (second, third, fourth, etc.) guide sequences, tracr mate sequence (optionally operably linked to the guide sequence), and a tracr sequence embedded within one or more intron sequences (e.g.
- the CRISPR enzyme, one or more additional guide sequence, tracr mate sequence, and tracr sequence are each a component of different nucleic acid sequences.
- the disclosure relates to a composition
- a composition comprising at least a first and second nucleic acid sequence, wherein the first nucleic acid sequence comprises a tracr sequence and the second nucleic acid sequence comprises a tracr mate sequence, wherein the first nucleic acid sequence is at least partially complementary to the second nucleic acid sequence such that the first and second nucleic acid for a duplex and wherein the first nucleic acid and the second nucleic acid either individually or collectively comprise a DNA-targeting domain, a Cas protein binding domain, and a transcription terminator domain.
- the CRISPR enzyme, one or more additional guide sequence, tracr mate sequence, and tracr sequence are operably linked to and expressed from the same promoter.
- the disclosure relates to compositions comprising any one or combination of the disclosed domains on one guide sequence or two separate tracrRNA/crRNA sequences with or without any of the disclosed modifications. Any methods disclosed herein also relate to the use of tracrRNA/crRNA sequence interchangeably with the use of a guide sequence, such that a composition may comprise a single synthetic guide sequence and/or a synthetic tracrRNA/crRNA with any one or combination of modified domains disclosed herein.
- the CRISPR system suitable for the present disclosure can also comprise a modified CRISPR enzyme (or “Cas protein”) or a nucleotide sequence encoding one or more Cas proteins.
- a Cas protein Any protein capable of enzymatic activity in cooperation with a guide sequence is a Cas protein.
- the disclosure relates to a system comprises a vector comprising a regulatory element operably linked to an enzyme-coding sequence encoding a CRISPR enzyme, such as a Cas protein from the Cas family of enzymes.
- the disclosure relates to a system, composition, or pharmaceutical composition comprising any one or plurality of Cas proteins either individually or in combination with one or a plurality of guide sequences.
- the amino acid sequence of S. pyogenes Cas9 protein may be found in the SwissProt database under accession number Q99ZW2.
- the unmodified CRISPR enzyme has DNA cleavage activity, such as Cas9.
- the CRISPR enzyme is Cas9, and may be Cas9 from 5. pyogenes or 5. pneumoniae .
- the CRISPR enzyme directs cleavage of one or both strands at the location of a target sequence, such as within the target sequence and/or within the complement of the target sequence.
- the CRISPR enzyme directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of a target sequence.
- a vector encodes a CRISPR enzyme or Cas protein that is mutated to with respect to a corresponding wild-type enzyme such that the mutated CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing a target sequence.
- D10A aspartate-to-alanine substitution
- pyogenes converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaves a single strand).
- Other examples of mutations that render Cas9 a nickase include, without limitation, H840A, N854A, and N863A.
- a Cas9 nickase may be used in combination with guide sequence(s), e.g., two guide sequences, which target respectively sense and antisense strands of the DNA target. This combination allows both strands to be nicked and used to induce NHEJ.
- two or more catalytic domains of Cas9 may be mutated to produce a mutated Cas9 substantially lacking all DNA cleavage activity.
- a D10A mutation is combined with one or more of H840A, N854A, or N863A mutations to produce a Cas9 enzyme substantially lacking all DNA cleavage activity.
- a CRISPR enzyme is considered to substantially lack all DNA cleavage activity when the DNA cleavage activity of the mutated enzyme is less than about 25%, 10%, 5%, 1%, 0.1%, 0.01%, or lower with respect to its non-mutated form.
- Other mutations may be useful; where the Cas9 or other CRISPR enzyme is from a species other than S. pyogenes, mutations in corresponding amino acids may be made to achieve similar effects.
- the disclosure relates to a method of detecting the presence of a nullomer by exposing a Cas protein and sgRNA specific to a target nullomer sequence to a nullomer target sequence.
- the nullomer target sequence is any nullomer from Table 1 and the sgRNA sequence specific for the nullomer is any RNA molecule that comprises from about 10 to about 35 nucleotides complementary to a nullomer in Table 1.
- the method further comprises allowing a time period sufficient for the sgRNA to associate with the nullomer and the Cas protein to excise the nullomer from the genomic DNA of a host cell or cell within a sample. Detection of the nullomer can further comprise identifying the nullomer sequence excised from the cell by amplification through PCR or a non-amplification event such as those disclosed herein.
- labels, dyes, or labeled probes and/or primers are used to detect amplified or unamplified nullomers.
- detection methods are appropriate based on the sensitivity of the detection method and the abundance of the target.
- amplification may or may not be required prior to detection.
- nullomer amplification is preferred.
- a probe or primer may include standard (A, T or U, G and C) bases, or modified bases.
- Modified bases include, but are not limited to, the AEGIS bases (from Eragen Biosciences), which have been described, e.g., in U.S. Pat. Nos. 5,432,272, 5,965,364, and 6,001,983.
- bases are joined by a natural phosphodiester bond or a different chemical linkage.
- Different chemical linkages include, but are not limited to, a peptide bond or a Locked Nucleic Acid (LNA) linkage, which is described, e.g., in U.S. Pat. No. 7,060,809.
- LNA Locked Nucleic Acid
- oligonucleotide probes or primers present in an amplification reaction are suitable for monitoring the amount of amplification product produced as a function of time.
- probes having different single stranded versus double stranded character are used to detect the nucleic acid.
- Probes include, but are not limited to, the 5 ’-exonuclease assay (e.g., TAQMAN) probes (see U.S. Pat. No. 5,538,848), stem-loop molecular beacons (see, e.g., U.S. Pat. Nos. 6,103,476 and 5,925,517), stemless or linear beacons (see, e.g., WO 9921881, U.S.
- one or more of the primers in an amplification reaction can include a label.
- different probes or primers comprise detectable labels that are distinguishable from one another.
- a nucleic acid, such as the probe or primer may be labeled with two or more distinguishable labels.
- a label is attached to one or more probes and has one or more of the following properties: (i) provides a detectable signal; (ii) interacts with a second label to modify the detectable signal provided by the second label, e g., FRET (Fluorescent Resonance Energy Transfer); (iii) stabilizes hybridization, e.g., duplex formation; and (iv) provides a member of a binding complex or affinity set, e.g., affinity, antibody-antigen, ionic complexes, hapten-ligand (e.g., biotin-avidin).
- use of labels can be accomplished using any one of a large number of known techniques employing known labels, linkages, linking groups, reagents, reaction conditions, and analysis and purification methods.
- Nullomers can be detected by direct or indirect methods.
- a direct detection method one or more nullomers are detected by a detectable label that is linked to a nucleic acid molecule.
- the nullomers may be labeled prior to binding to the probe. Therefore, binding is detected by screening for the labeled nullomer that is bound to the probe.
- the probe is optionally linked to a bead in the reaction volume.
- nucleic acids are detected by direct binding with a labeled probe, and the probe is subsequently detected.
- the nucleic acids such as amplified nullomers, are detected using FlexMAP Microspheres (Luminex) conjugated with probes to capture the desired nucleic acids.
- FlexMAP Microspheres Luminex
- Some methods may involve detection with polynucleotide probes modified with fluorescent labels or branched DNA (bDNA) detection, for example.
- biomarker expression is determined using a PCR-based assay comprising specific primers and/or probes for each biomarker.
- probe refers to any molecule that is capable of selectively binding a specifically intended target biomolecule.
- probe refers to any molecule that may bind or associate, indirectly or directly, covalently or non-covalently, to any of the substrates and/or reaction products and/or proteases disclosed herein and whose association or binding is detectable using the methods disclosed herein.
- the term “probe” refers to any molecule comprising a nucleic acid sequence that is complementary to any of the nucleic acid sequences disclosed in TABLE 1 or one comprising at least about 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to any of the nucleic acid sequences disclosed in TABLE 1.
- the term “probe” refers to any molecule comprising a nucleic acid sequence that is complementary to a fragment of any of the nucleic acid sequences disclosed in TABLE 1 or one comprising at least about 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to a fragment of any of the nucleic acid sequences disclosed in TABLE 1 .
- the term “probe” refers to a sgRNA molecule comprising a nucleic acid sequence that is complementary to a fragment of any of the nucleic acid sequences disclosed in TABLE 1 or one comprising at least about 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to a fragment of any of the nucleic acid sequences disclosed in TABLE 1.
- the probe is a fluorogenic probe, antibody or absorbance-based probes.
- the chromophore pNA may be used as a probe for detection and/or quantification of a target nucleic acid sequence disclosed herein.
- the probe may comprise a nucleic acid sequence labeled with a fluorogenic molecule or a substrate that when exposed to an enzyme becomes fluorogenic and the nucleic acid sequence is complementary to any of the nucleic acid sequences disclosed in TABLE 1 or one comprising at least about 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to any of the nucleic acid sequences disclosed in TABLE 1 or Table B.
- Probes can be synthesized by one of skill in the art using known techniques, or derived from biological preparations. Probes may include but are not limited to, RNA, DNA, proteins, peptides, aptamers, antibodies, and organic molecules.
- the term “primer” or “probe” encompasses oligonucleotides that have a specific sequence or oligoribonucleotides that have a specific sequence.
- the probe are from about 5 to about 20 nucleotides in length and are complementary to the nucleic acid sequences in TABLE 1 or in Table B and comprise at least about 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or about 100% sequence identity to any one or combination of nucleic acid sequences complementary to those provided in TABLE 1 or Table B.
- the probe are from about 5 to about 20 nucleotides in length and are complementary to the nucleic acid sequences in TABLE 1 and comprise at least about 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or about 100% sequence identity to any one or combination of nucleic acid sequences complementary to those provided in TABLE 7.
- the probe are from about 5 to about 20 nucleotides in length and are complementary to the nucleic acid sequences in TABLE 1 and comprise at least about 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or about 100% sequence identity to any one or combination of nucleic acid sequences complementary to those provided in TABLE 8.
- the target molecule could be any one or a combination of nucleic acid sequences identified in TABLE 1.
- the target molecule is a nucleic acid sequence comprising at least about 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or about 99% sequence identity to any one or combination of nucleic acid sequences provided in TABLE 1.
- the target molecule is any amplified fragment of any one or combination of nucleic acid sequences identified in TABLE 1, and/or any one or combination of nucleic acid sequence comprising at least about 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or about 99% sequence identity to any one or combination of nucleic acid sequences in TABLE 1.
- nucleic acids are detected by indirect detection methods.
- a biotinylated probe may be combined with a streptavidin-conjugated dye to detect the bound nucleic acid.
- the streptavidin molecule binds a biotin label on amplified nullomer, and the bound nullomer is detected by detecting the dye molecule attached to the streptavidin molecule.
- the streptavidin-conjugated dye molecule comprises PHYCOLINK. Streptavidin R-Phycoerythrin (PROzyme). Other conjugated dye molecules are known to persons skilled in the art.
- Labels include, but are not limited to, light-emitting, light-scattering, and light-absorbing compounds which generate or quench a detectable fluorescent, chemiluminescent, or bioluminescent signal (see, e.g., Kricka, L., Nonisotopic DNA Probe Techniques, Academic Press, San Diego (1992) and Garman A., Non-Radioactive Labeling, Academic Press (1997).).
- a dual labeled fluorescent probe that includes a reporter fluorophore and a quencher fluorophore is used in some embodiments. It will be appreciated that pairs of fluorophores are chosen that have distinct emission spectra so that they can be easily distinguished.
- labels are hybridization-stabilizing moieties which serve to enhance, stabilize, or influence hybridization of duplexes, e.g., intercalators and intercalating dyes (including, but not limited to, ethidium bromide and SYBR-Green), minor-groove binders, and cross-linking functional groups (see, e.g., Blackbum et al., eds. “DNA and RNA Structure” in Nucleic Acids in Chemistry and Biology (1996)).
- intercalators and intercalating dyes including, but not limited to, ethidium bromide and SYBR-Green
- minor-groove binders include, but not limited to, ethidium bromide and SYBR-Green
- cross-linking functional groups see, e.g., Blackbum et al., eds. “DNA and RNA Structure” in Nucleic Acids in Chemistry and Biology (1996)).
- methods relying on hybridization and/or ligation to quantify nullomers may be used, including oligonucleotide ligation (OLA) methods and methods that allow a distinguishable probe that hybridizes to the target nucleic acid sequence to be separated from an unbound probe.
- OLA oligonucleotide ligation
- HARP-like probes as disclosed in U.S. Publication No. 2006/0078894 may be used to measure the quantity of nullomers.
- the probe after hybridization between a probe and the targeted nucleic acid, the probe is modified to distinguish the hybridized probe from the unhybridized probe. Thereafter, the probe may be amplified and/or detected.
- a probe inactivation region comprises a subset of nucleotides within the target hybridization region of the probe.
- a post-hybridization probe inactivation step is carried out using an agent which is able to distinguish between a HARP probe that is hybridized to its targeted nucleic acid sequence and the corresponding unhybridized HARP probe.
- the agent is able to inactivate or modify the unhybridized HARP probe such that it cannot be amplified.
- a probe ligation reaction may also be used to quantify nullomers.
- MLP A Multiplex Ligation-dependent Probe Amplification
- the nullomers described herein can be used individually or in combination in diagnostic tests to assess the type of cancer, tissue of origin, and status or stage of the cancer in a subject.
- Cancer status or stage includes the presence or absence of the cancer. Cancer status or stage may also include monitoring the course of the cancer, for example, monitoring disease progression. Based on the cancer status or stage of a subject, additional procedures may be indicated, including, for example, additional diagnostic tests or therapeutic procedures.
- the power of a diagnostic test to correctly predict disease status is commonly measured in terms of the accuracy of the assay, the sensitivity of the assay, the specificity of the assay, or the “Area Under a Curve” (AUC), for example, the area under a Receiver Operating Characteristic (ROC) curve.
- accuracy is a measure of the fraction of misclassified samples. Accuracy may be calculated as the total number of correctly classified samples divided by the total number of samples, e.g., in a test population.
- Sensitivity is a measure of the “true positives” that are predicted by a test to be positive, and may be calculated as the number of correctly identified cancer samples divided by the total number of cancer samples.
- Specificity is a measure of the “true negatives” that are predicted by a test to be negative, and may be calculated as the number of correctly identified normal samples divided by the total number of normal samples.
- AUC is a measure of the area under a Receiver Operating Characteristic curve, which is a plot of sensitivity vs. the false positive rate (1-specificity). The greater the AUC, the more powerful the predictive value of the test.
- Other useful measures of the utility of a test include the “positive predictive value,” which is the percentage of actual positives who test as positives, and the “negative predictive value,” which is the percentage of actual negatives who test as negatives.
- diagnostic tests that use nullomers described herein individually or in combination show an accuracy of at least about 75%, e.g., an accuracy of at least about 75%, about 80%, about 85%, about 90%, about 95%, about 97%, about 99% or about 100%.
- diagnostic tests that use nullomers described herein individually or in combination show a specificity of at least about 75%, e.g., a specificity of at least about 75%, about 80%, about 85%, about 90%, about 95%, about 97%, about 99% or about 100%.
- diagnostic tests that use nullomers described herein individually or in combination show a sensitivity of at least about 75%, e.g., a sensitivity of at least about 75%, about 80%, about 85%, about 90%, about 95%, about 97%, about 99% or about 100%.
- diagnostic tests that use nullomers described herein individually or in combination show a specificity and sensitivity of at least about 75% each, e.g., a specificity and sensitivity of at least about 75%, about 80%, about 85%, about 90%, about 95%, about 97%, about 99% or about 100% (for example, a specificity of at least about 80% and sensitivity of at least about 80%, or for example, a specificity of at least about 80% and sensitivity of at least about 95%).
- Each nullomer listed in TABLE 1 is identified as being associated with certain type(s) of cancer as provided. In some instances, one particular nullomer may be associated with more than one types of cancers. In other instances, one particular nullomer may be associated with only one type of cancer.
- Each nullomer listed in TABLE 1 is differentially present in biological samples derived from subjects having certain types of cancers as compared with normal subjects, and thus each is individually useful in facilitating the determination of those types of cancer in a test subject.
- Such a method involves determining the level of the nullomer in a sample obtained from the subject. Determining the level of the nullomer in a sample may include measuring, detecting, or assaying the level of the nullomer in the sample using any suitable method, for example, the methods set forth herein. Determining the level of the nullomer in a sample may also include examining the results of an assay that measured, detected, or assayed the level of the nullomer in the sample.
- the method may also involve comparing the level of the nullomer in a sample with a suitable control.
- a change in the level of the nullomer relative to that in a normal subject as assessed using a suitable control is indicative of the cancer status or stage of the subject.
- a diagnostic amount of a nullomer that represents an amount of the nullomer above or below which a subject is classified as having a particular cancer status or stage can be used. For example, if the nullomer is upregulated in samples from an individual having cancer as compared to a normal individual, a measured amount above the diagnostic cutoff provides a diagnosis of the type of cancer that individual has.
- the nullomers in TABLE 1 and Table 7 are upregulated in cancer samples relative to samples obtained from normal individuals.
- adjusting the particular diagnostic cut-off used in an assay allows one to adjust the sensitivity and/or specificity of the diagnostic assay as desired.
- the particular diagnostic cut-off can be determined, for example, by measuring the amount of the nullomer in a statistically significant number of samples from subjects with different cancer statuses, and drawing the cut-off at the desired level of accuracy, sensitivity, and/or specificity.
- the diagnostic cut-off can be determined with the assistance of a classification algorithm, as described elsewhere herein.
- methods for diagnosing cancer in a subject, by determining the level of at least one nullomer in a sample from the subject, wherein a difference in the level of the at least one nullomer versus that in a normal subject (as determined relative to a suitable control) is indicative of cancer in the subject.
- the at least one nullomer includes one or more nullomers from TABLE 1.
- a difference in the level of the at least one nullomer versus that in a normal subject is indicative of the type(s) of cancer identified as being associated with the detected at least one nullomer in the subject.
- the disclosed method of determining the level of at least one nullomer in a sample from a subject, wherein an increase in the level of the at least one nullomer relative to a control is indicative of cancer in the subject, particularly of the type(s) of cancer identified as being associated with the at least one nullomer detected.
- the subject is diagnosed with having breast cancer, pancreatic cancer, esophagus cancer, lymphoid cancer, kidney cancer, ovary cancer, head and neck cancer, lung cancer, stomach cancer, CNS cancer, uterus cancer, skin cancer, colorectal cancer, prostate cancer, bladder cancer, bone and soft tissue cancer, biliary cancer, cervix cancer, thyroid cancer, myeloid cancer, or liver cancer by the disclosed method.
- the method may further comprise providing a diagnosis that the subject has or does not have cancer based on the level of at least one nullomer in the sample.
- the method may further comprise correlating a difference in the level or levels of at least one nullomer relative to a suitable control with a diagnosis of cancer in the subject.
- a diagnosis may be provided directly to the subject, or it may be provided to another party involved in the subject’s care.
- nullomers While individual nullomers are useful in diagnostic applications for various types of cancer, as shown herein, a combination of nullomers may provide greater predictive value of cancer status or stage than the nullomers when used alone. Specifically, the detection of a plurality of nullomers can increase the accuracy, sensitivity, and/or specificity of a diagnostic test. The detection of a plurality of nullomers can also assist in narrowing down the type of cancer and/or status or stage thereof in a subject. This is particular useful when a given nullomer is identified as being associated with more than one type of cancer.
- nullomer A is identified as being associated with cancers X, Y and Z
- nullomer B is identified as being associated with cancers X and Y
- nullomer C is identified as being associated with cancers X and Z
- a detection of the presence of nullomers A, B and C in a subject is indicative that the subject has cancer X.
- the disclosure thus includes the individual nullomer provided in TABLE 1 and nullomer combinations as set forth herein, and their use in methods and kits described herein.
- nullomers include one or more of nullomers provided in TABLEI .
- the type(s) of cancer thus diagnosed is/are the one(s) provided in TABLE 1 as being associated with each individual nullomer provided in TABLE 1.
- the set of data serves as a suitable control or reference standard for comparison with the sample from the subject.
- Comparison of the sample from the subject with the set of data may be assisted by a classification algorithm, which computes whether or not a statistically significant difference exists between the collective levels of the two or more nullomers in the sample, and the levels of the same nullomers present in normal subjects or subjects having cancer.
- data that are generated using samples such as “known samples” can then be used to “train” a classification model.
- a “known sample” is a sample that has been preclassified, e.g., classified as being derived from a normal subject or from a subject having a particular type of cancer.
- the data that are derived from the spectra and are used to form the classification model can be referred to as a “training data set.”
- the classification model can recognize patterns in data derived from spectra generated using unknown samples.
- the classification model can then be used to classify the unknown samples into classes. This can be useful, for example, in predicting whether or not a particular biological sample is associated with a certain biological condition (e.g., diseased versus non-diseased).
- data for the training data set that is used to form the classification model can be obtained directly from quantitative PCR (for example, Ct values obtained using the double delta Ct method), or from high-throughput expression profiling, such as microarray analysis (for example, total counts or normalized counts from a nullomer or neomer expression assay).
- quantitative PCR for example, Ct values obtained using the double delta Ct method
- high-throughput expression profiling such as microarray analysis (for example, total counts or normalized counts from a nullomer or neomer expression assay).
- Classification models can be formed using any suitable statistical classification (or “learning”) method that attempts to segregate bodies of data into classes based on objective parameters present in the data.
- Classification methods may be either supervised or unsupervised. Examples of supervised and unsupervised classification processes are described in Jain, “Statistical Pattern Recognition: A Review,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 22, No. 1, January 2000, the teachings of which are incorporated by reference in its entirety.
- supervised classification training data containing examples of known categories are presented to a learning mechanism, which learns one or more sets of relationships that define each of the known classes. New data may then be applied to the learning mechanism, which then classifies the new data using the learned relationships.
- supervised classification processes include linear regression processes (e.g., multiple linear regression (MLR), partial least squares (PLS) regression and principal components regression (PCR)), binary decision trees (e.g., recursive partitioning processes such as CART - classification and regression trees), artificial neural networks such as back propagation networks, discriminant analyses (e.g., Bayesian classifier or Fischer analysis), logistic classifiers, and support vector classifiers (support vector machines).
- linear regression processes e.g., multiple linear regression (MLR), partial least squares (PLS) regression and principal components regression (PCR)
- binary decision trees e.g., recursive partitioning processes such as CART - classification and regression trees
- artificial neural networks such as back propagation networks
- discriminant analyses e.g.,
- the classification models that are created can be formed using unsupervised learning methods.
- Unsupervised classification attempts to learn classifications based on similarities in the training data set, without pre-classifying the spectra from which the training data set was derived.
- Unsupervised learning methods include cluster analyses. A cluster analysis attempts to divide the data into “clusters” or groups that ideally should have members that are very similar to each other, and very dissimilar to members of other clusters. Similarity is then measured using some distance metric, which measures the distance between data items, and clusters together data items that are closer to each other.
- Clustering techniques include the MacQueen’s K-means algorithm and the Kohonen’s Self-Organizing Map algorithm.
- the classification models can be formed on and used on any suitable digital computer.
- Suitable digital computers include micro, mini, or large computers using any standard or specialized operating system, such as a Unix, WINDOWS or LINUX based operating system.
- the training data set(s) and the classification models can be embodied by computer code that is executed or used by a digital computer.
- the computer code can be stored on any suitable computer readable media including optical or magnetic disks, sticks, tapes, etc., and can be written in any suitable computer programming language including C, C++, visual basic, etc.
- the learning algorithms described herein can be used for developing classification algorithms for nullomers or meomers for various types of tumors.
- the classification algorithms can, in turn, be used in diagnostic tests by providing diagnostic values (e.g., cut-off points) for neomers used singly or in combination.
- the algorithms can also be used to correlate the presence or absence of a neomer in a sample to a presence of a mutation or presence of a functional error in a particular genetic element of the cell in a subject.
- the presence of the neomer can indicate the liklehood of the presence of a mutation at a particular locus within the genome of the cancer cell or the likelihood of the presence of a dysfunction of a particular regulatory element within the genome of the cancer cell.
- a calculaoin or determination of the likelihood can establish a recommendation of therapy for the subject, such that there is a greater likelihood the subject is responsive to the therapy.
- Table C lists the type of cancer associated with the Genes and Neomers identified in Table B.
- Table C also lists the types of treatments recommended for the particular cancer types lists.
- the methods comprise a step of correlating the presence of the neomer to a mutation or dysfunctional phenotype of the cancer cell. In some embodiments, the methods further comprise pairing the mutation or dysfunctional phenotype to a therapy recommendation for the subject or a likelihood that the subject would be responsive to a certain therapy.
- the presence of a neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 1 through SEQ ID NO: 2, is correlated to a mutation at sequence CCGTGCAGCTCATCATGCAGCTCATGCCCTT of the EGFR1 locus or regulatory element within the cancer cell of the subject.
- methods comprise a step of recommending or determining a therapy for the subject based upon the presence of the neomer.
- the subject has non-small cell lung cancer.
- methods further comprise a step of recommending therapy of one or a combination of erlotinib, osimertinib, gefitinib, afatinib, amivantamb, dacomitinib or mobocertinib if the sample of the subject comprises the presence of the neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 1 through SEQ ID NO: 2.
- the presence of a neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 1 through SEQ ID NO: 2, is correlated to a mutation at sequence CCGTGCAGCTCATCATGCAGCTCATGCCCTT of the EGFR1 locus or regulatory element within the cancer cell of the subject.
- methods comprise a step of recommending or determining a therapy for the subject based upon the presence of the neomer.
- the subject has colorectal cancer.
- methods further comprise a step of recommending therapy of one or a combination of cetuximab and/or panitumumab if the sample of the subject comprises the presence of the neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 1 through SEQ ID NO: 2.
- the presence of a neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 3 through SEQ ID NO: 28, is correlated to a mutation at sequence TCACAGATTTTGGGCGGGCCAAACTGCTGGG of the EGFR1 locus or regulatory element within the cancer cell of the subject.
- methods comprise a step of recommending or determining a therapy for the subject based upon the presence of the neomer.
- the subject has non-small cell lung cancer.
- methods further comprise a step of recommending therapy of one or a combination of erlotinib, osimertinib, gefitinib, afatinib, amivantamb, dacomitinib or mobocertinib if the sample of the subject comprises the presence of the neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 3 through SEQ ID NO: 28.
- the presence of a neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 3 through SEQ ID NO: 28, is correlated to a mutation at sequence TCACAGATTTTGGGCGGGCCAAACTGCTGGG of the EGFR1 locus or regulatory element within the cancer cell of the subject.
- methods comprise a step of recommending or determining a therapy for the subject based upon the presence of the neomer.
- the subject has colorectal cancer.
- methods further comprise a step of recommending therapy of one or a combination of cetuximab and/or panitumumab if the sample of the subject comprises the presence of the neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 3 through SEQ ID NO: 28.
- the presence of a neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 29 through SEQ ID NO: 66 is correlated to a mutation at sequence identified in Table B of the BRCA2 locus or regulatory element within the cancer cell of the subject.
- methods comprise a step of recommending or determining a therapy for the subject based upon the presence of the neomer. Therefore in some embodiments, methods further comprise a step of recommending therapy of one or a combination of the therapies listed in Table C if the sample of the subject comprises the presence of the neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 29 through SEQ ID NO: 66.
- the presence of a neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 67 through SEQ ID NO: 138 is correlated to a mutation at sequence identified in Table B of the EGFR locus or regulatory element within the cancer cell of the subject.
- methods comprise a step of recommending or determining a therapy for the subject based upon the presence of the neomer. Therefore in some embodiments, methods further comprise a step of recommending therapy of one or a combination of the therapies listed in Table C if the sample of the subject comprises the presence of the neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 67 through SEQ ID NO: 138.
- the presence of a neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 139 through SEQ ID NO: 156 is correlated to a mutation at sequence identified in Table B of the TP53 locus or regulatory element within the cancer cell of the subject.
- methods comprise a step of recommending or determining a therapy for the subject based upon the presence of the neomer. Therefore in some embodiments, methods further comprise a step of recommending therapy of one or a combination of the therapies listed in Table C if the sample of the subject comprises the presence of the neomer that comprises at least about 80% sequence identity to any one of SEQ ID NO: 139 through SEQ ID NO: 156.
- the quantity of neomers is indicative of various types of cancer may be used as a stand-alone diagnostic indicator of cancer in a subject.
- the methods may include the performance of at least one additional test to facilitate the diagnosis of cancer.
- other tests in addition to determining the level of one or more nullomers and/or neomers in order to facilitate a diagnosis of cancer may be performed. Any other test or combination of tests used in clinical practice to facilitate a diagnosis of cancer may be used in conjunction with the neomers.
- the method of diagnosis comprise identifying the presence or quantity of the amount of neomer in a sample from a subject.
- the disclosure further provides methods of treating the subject identified as having a cancer or a method of propsing a therapy for better responsiveness of the subject to cancer therapy. Accordingly, in some embodiments, the disclosure relates to a method of treating cancer in a subject, comprising determining the level of at least one neomer in a sample from the subject, wherein a difference in the level of at least one neomer versus that in a normal subject as determined relative to a suitable control is indicative of cancer in the subject, and administering a therapeutically effective amount of a cancer therapeutic to the subject.
- the disclosure relates to a method of treating a subject having cancer, comprising identifying a subject having cancer in which the level of at least one neomer in a sample from the subject is different (e.g., increased) versus that in a normal subject as determined relative to a suitable control, and administering a therapeutically effective amount of a cancer therapeutic to the subj ect.
- cancer therapeutic includes, for example, substances approved by the U.S. Food and Drug Administration for the treatment of cancer.
- drugs approved to treat breast cancer include, but are not limited to, Abemaciclib, Abitrexate (Methotrexate), Abraxane (Paclitaxel Albumin-stabilized Nanoparticle Formulation), Ado-Trastuzumab Emtansine, Afinitor (Everolimus), Anastrozole, Aredia (Pamidronate Disodium), Arimidex (Anastrozole), Aromasin (Exemestane), Capecitabine, Clafen (Cyclophosphamide), Cyclophosphamide, Cytoxan (Cyclophosphamide), Docetaxel, Doxorubicin Hydrochloride, Ellence (Epirubicin Hydrochloride), Epirubicin Hydrochloride, Eribulin Mesylate, Everolimus, Exemestane, 5-FU (Fluorouracil Injection),
- the cancer therapeutics may be administered to a subject using a pharmaceutical composition.
- suitable pharmaceutical compositions comprise a pharmaceutically effective amount of a cancer therapeutic (or a pharmaceutically acceptable salt or ester thereof), and optionally comprise a pharmaceutically acceptable carrier. In certain embodiments, these compositions optionally further comprise one or more additional therapeutic agents.
- the term “pharmaceutically acceptable salt” refers to those salts which are, within the scope of sound medical judgment, suitable for use in contact with the tissues of humans and lower animals without undue toxicity, irritation, allergic response and the like, and are commensurate with a reasonable benefit/risk ratio.
- Pharmaceutically acceptable salts of amines, carboxylic acids, and other types of compounds are well known in the art. For example, S. M. Berge, et al. describe pharmaceutically acceptable salts in detail in J. Pharmaceutical Sciences, 66: 1-19 (1977), incorporated herein by reference.
- the salts can be prepared in situ during the final isolation and purification of the compounds, or separately by reacting a free base or free acid function with a suitable reagent.
- a free base function can be reacted with a suitable acid.
- suitable pharmaceutically acceptable salts thereof may, include metal salts such as alkali metal salts, e.g., sodium or potassium salts, and alkaline earth metal salts, e.g., calcium or magnesium salts.
- the cancer therapeutic is a pharmaceutically acceptable salt.
- ester refers to esters that hydrolyze in vivo and include those that break down readily in the human body to leave the parent compound or a salt thereof.
- Suitable ester groups include, for example, those derived from pharmaceutically acceptable aliphatic carboxylic acids, particularly alkanoic, alkenoic, cycloalkanoic and alkanedioic acids, in which each alkyl or alkenyl moiety advantageously has not more than 6 carbon atoms.
- the cancer therapeutic is a pharmaceutically acceptable ester.
- the pharmaceutical compositions may additionally comprise a pharmaceutically acceptable carrier.
- pharmaceutically acceptable carrier includes any and all solvents, diluents, or other liquid vehicle, dispersion or suspension aids, surface active agents, isotonic agents, thickening or emulsifying agents, preservatives, solid binders, lubricants and the like, suitable for preparing the particular dosage form desired.
- Remington s Pharmaceutical Sciences, Sixteenth Edition, E. W. Martin (Mack Publishing Co., Easton, Pa., 1980) discloses various carriers used in formulating pharmaceutical compositions and known techniques for the preparation thereof.
- materials which can serve as pharmaceutically acceptable carriers include, but are not limited to, sugars such as lactose, glucose and sucrose; starches such as corn starch and potato starch; cellulose and its derivatives such as sodium carboxymethyl cellulose, ethyl cellulose and cellulose acetate; powdered tragacanth; malt; gelatine; talc; excipients such as cocoa butter and suppository waxes; oils such as peanut oil, cottonseed oil; safflower oil, sesame oil; olive oil; corn oil and soybean oil; glycols; such as propylene glycol; esters such as ethyl oleate and ethyl laurate; agar; buffering agents such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogenfree water; isotonic saline; Ringer’s solution; ethyl alcohol, and phosphate buffer solutions, as well as other non-toxic compatible lubricants such as sodium
- compositions for use in the present disclosure may be formulated to have any concentration of the cancer therapeutic desired.
- the composition is formulated such that it comprises a therapeutically effective amount of the cancer therapeutic.
- the disclosure generally relates to a method of diagnosing a subject with a benign, pre- malignant, or malignant hyperproliferative cell comprising: detecting the presence, absence, and/or quantity of at least one neomer in a sample.
- the step of detecting comprise exposing a sample from a subject (e.g., a human subject) to one or a plurality of probes, each probe capable of binding one or a plurality of neomers in the sample.
- the probe is a labeled nucleic acid molecule (DNA, RNA or hybrid thereof) that comprises at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the complement of any nucleic acid sequences of Table 1, Table 5, Table 6 or Table B.
- the probe is a labeled nucleic acid molecule (DNA, RNA or hybrid thereof) that is an RNA sequence comprising at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the complement of any nucleic acid sequences of TABLE 1, where each thymine is replaced with a uracil.
- the plurality of probes are one or a combination of labeled nucleic acid sequences that are an RNA complementary to a nucleic acid sequence comprising at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any nucleic acid sequences of TABLE 1 or Table B.
- the plurality of probes are one or a combination of labeled nucleic acid sequences chosen from any nucleic acid sequences of TABLE 1.
- the plurality of probes comprise one or a combination of nucleic acid sequences complementary to the nucleic acid sequences chosen from any nucleic acid sequences of TABLE 1.
- the probe is a labeled nucleic acid molecule (DNA, RNA or hybrid thereof) that comprises at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the complement of any nucleic acid sequences of TABLE 1.
- the probe is a labeled nucleic acid molecule (DNA, RNA or hybrid thereof) that is an RNA sequence comprising at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the complement of any nucleic acid sequences of TABLE 1, where each thymine is replaced with a uracil.
- the plurality of probes are one or a combination of labeled nucleic acid sequences that are an RNA complementary to a nucleic acid sequence comprising at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any nucleic acid sequences of TABLE 1.
- the plurality of probes are one or a combination of labeled nucleic acid sequences chosen from any nucleic acid sequences of TABLE 1.
- the plurality of probes comprise one or a combination of nucleic acid sequences complementary to the nucleic acid sequences chosen from any nucleic acid sequences of TABLE 1.
- the probe is a labeled nucleic acid molecule (DNA, RNA or hybrid thereof) that comprises at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the complement of any nucleic acid sequences of TABLE 1.
- the probe is a labeled nucleic acid molecule (DNA, RNA or hybrid thereof) that is an RNA sequence comprising at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the complement of any nucleic acid sequences of TABLE 5, where each thymine is replaced with a uracil.
- the plurality of probes are one or a combination of labeled nucleic acid sequences that are an RNA complementary to a nucleic acid sequence comprising at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any nucleic acid sequences of TABLE 5.
- the plurality of probes are one or a combination of labeled nucleic acid sequences chosen from any nucleic acid sequences of TABLE
- the plurality of probes comprise one or a combination of nucleic acid sequences complementary to the nucleic acid sequences chosen from any nucleic acid sequences of TABLE 5.
- the probe is a labeled nucleic acid molecule (DNA, RNA or hybrid thereof) that comprises at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the complement of any nucleic acid sequences of TABLE 6.
- the probe is a labeled nucleic acid molecule (DNA, RNA or hybrid thereof) that is an RNA sequence comprising at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the complement of any nucleic acid sequences of TABLE 6, where each thymine is replaced with a uracil.
- the plurality of probes are one or a combination of labeled nucleic acid sequences that are an RNA complementary to a nucleic acid sequence comprising at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any nucleic acid sequences of TABLE 6.
- the plurality of probes are one or a combination of labeled nucleic acid sequences chosen from any nucleic acid sequences of TABLE
- the plurality of probes comprise one or a combination of nucleic acid sequences complementary to the nucleic acid sequences chosen from any nucleic acid sequences of TABLE 6.
- the probe is a labeled nucleic acid molecule (DNA, RNA or hybrid thereof) that comprises at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the complement of any nucleic acid sequences of TABLE B.
- the probe is a labeled nucleic acid molecule (DNA, RNA or hybrid thereof) that is an RNA sequence comprising at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the complement of any nucleic acid sequences of TABLE B, where each thymine is replaced with a uracil.
- the plurality of probes are one or a combination of labeled nucleic acid sequences that are an RNA complementary to a nucleic acid sequence comprising at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any nucleic acid sequences of TABLE B.
- the plurality of probes are one or a combination of labeled nucleic acid sequences chosen from any nucleic acid sequences of TABLE B.
- the plurality of probes comprise one or a combination of nucleic acid sequences complementary to the nucleic acid sequences chosen from any nucleic acid sequences of TABLE B
- the subject may be a human diagnosed with or suspected as having cancer.
- the step of detecting is preceded by a step of acquiring a sample from the subject.
- the probe or plurality of probes are one or a plurality of antibodies or antibody fragments comprising a CDR that binds to a nucleic acid molecule (DNA, RNA or hybrid thereof) that comprises at least 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any nucleic acid sequences of TABLE 1.
- a nucleic acid molecule DNA, RNA or hybrid thereof
- the probe or plurality of probes are one or a plurality of antibodies or antibody fragments comprising a CDR that binds to a nucleic acid molecule (DNA, RNA or hybrid thereof) that comprises at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any nucleic acid sequences of TABLE 1, wherein each of sequences are modified such that the thymines in each sequence are replaced with a uracil.
- the methods further comprise isolating RNA from the sample before exposing the sample to one or a plurality of probes.
- the method comprises detecting or quantifying an amount of neomers in a sample by performing semiquantitative or quantitative PCR or sequencing analysis of the neomers in a sample.
- Probes may be immobilized to a solid support such as an ELISA plate, plastic, slide, microarray, silica chip or other surface such that the single-strand nucleotide sequences are exposed to a sample comprising neomers from a subject.
- the probes may comprise, in some embodiments, from about 5 to about 100 nucleotides in length and comprise any of the sequences provided in TABLE 1 or any complementary sequence in RNA or DNA form of the sequences set forth in TABLE 1.
- the step of detecting the presence, absence, and/or quantity of at least one neomer having at least about 70% sequence identity to one of the neomers in a sample comprises using a chemoluminescent probe, fluorescent probe, and/or fluorescence microscopy, calculating the presence or quantity by correlating the signal of the detectable probe to the presence of the neomer.
- any of the methods disclosed herein further comprise a step of correlating the presence or quantity of one or more neomers, such as those disclosed in TABLE 1 or any combination thereof, to the likelihood that the subject has cancer.
- the disclosure relates to a method of preparing, isolating or assessing a nucleic acid or ribonucleic acid fraction from a subject useful for analyzing a neomer involved in cancer comprising: extracting DNA or RNA from a substantially cell-free sample of blood plasma or blood serum of a subject to obtain DNA or RNA pools; (b) producing a fraction of the DNA or RNA extracted in (a) by: (i) sequence discrimination of the DNA or RNA; and (ii) selectively removing neomers by exposing one or a plurality of probes to the neomers, wherein the neomers after (b) comprises one or a plurality of neomers disclosed in TABLE 1 ; and (c) analyzing the
- the step of analyzing comprises normalizing the amount of neomers in the sample as compared to a control amount of neomers from a control sample and determining whether the subject has cancer by comparing the normalized presence, absence or quantity of neomers in the sample to the presence, absence or quantity of neomers in a control sample.
- kits for diagnosing type of cancer, tissue of origin, and status or stage of the cancer in a subject which kits are useful for determining the level of one or more neomers from TABLE 1 or Table B, wherein the sequences optionally comprise uracils in place of one, more than one, or all of the disclosed thymines), and combinations thereof.
- the one or more neomers are selected from the neomers listed in TABLE 1 or Table B.
- Kits may include materials and reagents adapted to selectively detect the presence of a neomer or group of neomers diagnostic for cancer in a sample of a subject.
- the kit may include a reagent that specifically hybridizes to a neomer.
- a reagent may be a nucleic acid molecule in a form suitable for detecting the neomer, for example, a probe or a primer.
- the kit may include reagents useful for performing an assay to detect one or more neomers, for example, reagents which may be used to detect one or more neomers in a qPCR reaction.
- the kit may likewise include a microarray useful for detecting one or more neomers.
- the kit may contain instructions for suitable operational parameters in the form of a label or product insert.
- the instructions may include information or directions regarding how to collect a sample, how to determine the level of one or more neomers in a sample, and/or how to correlate the level of one or more neomers in a sample with the type of cancer, tissue of origin, and status or stage of the cancer of a subject.
- the kit can contain one or more containers with neomer samples, to be used as reference standards, suitable controls, or for calibration of an assay to detect the neomers in a test sample.
- Radioisotopes that may be incorporated into pharmaceutical compositions or used as probes or labels with neomers.
- the agent in selected from one or a plurality of agents chosen from Table 3.
- the embodiments may be implemented using a computer program product (i.e. software), hardware, software or a combination thereof.
- the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers.
- a computer may be embodied in any of a number of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, a computer may be embedded in a device not generally regarded as a computer but with suitable processing capabilities, including a Personal Digital Assistant (PDA), a smart phone or any other suitable portable or fixed electronic device.
- PDA Personal Digital Assistant
- a computer may have one or more input and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that can be used to provide a user interface include printers or display screens for visual presentation of output and speakers or other sound generating devices for audible presentation of output. Examples of input devices that can be used for a user interface include keyboards, and pointing devices, such as mice, touch pads, and digitizing tablets. As another example, a computer may receive input information through speech recognition or in other audible format.
- Such computers may be interconnected by one or more networks in any suitable form, including a local area network or a wide area network, such as an enterprise network, and intelligent network (IN) or the Internet.
- networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks, wired networks or fiber optic networks.
- a computer employed to implement at least a portion of the functionality described herein may include a memory, coupled to one or more processing units (also referred to herein simply as “processors”), one or more communication interfaces, one or more display units, and one or more user input devices.
- the memory may include any computer-readable media, and may store computer instructions (also referred to herein as “processor-executable instructions”) for implementing the various functionalities described herein.
- the processing unit(s) may be used to execute the instructions.
- the communication interface(s) may be coupled to a wired or wireless network, bus, or other communication means and may therefore allow the computer to transmit communications to and/or receive communications from other devices.
- the display unit(s) may be provided, for example, to allow a user to view various information in connection with execution of the instructions.
- the user input device(s) may be provided, for example, to allow the user to make manual adjustments, make selections, enter data or various other information, and/or interact in any of a variety of manners with the processor during execution of the instructions.
- the various methods or processes outlined herein may be coded as software that is executable on one or more processors that employ any one of a variety of operating systems or platforms. Additionally, such software may be written using any of a number of suitable programming languages and/or programming or scripting tools, and also may be compiled as executable machine language code or intermediate code that is executed on a framework or virtual machine.
- inventive concepts may be embodied as a computer readable storage medium (or multiple computer readable storage media) (e.g., a computer memory, one or more floppy discs, compact discs, optical discs, magnetic tapes, flash memories, circuit configurations in Field Programmable Gate Arrays or other semiconductor devices, or other non-transitory medium or tangible computer storage medium) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods that implement the various embodiments of the invention disclosed herein.
- the computer readable medium or media can be transportable, such that the program or programs stored thereon can be loaded onto one or more different computers or other processors to implement various aspects of the present invention as discussed above.
- the system comprises cloud-based software that executes one or all of the steps of each disclosed method instruction.
- program or “software” are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects of embodiments as discussed above. Additionally, it should be appreciated that according to one aspect, one or more computer programs that when executed perform methods of the present disclosure need not reside on a single computer or processor, but may be distributed in a modular fashion amongst a number of different computers or processors to implement various aspects of the present invention.
- Computer-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices.
- program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types.
- functionality of the program modules may be combined or distributed as desired in various embodiments.
- data structures may be stored in computer-readable media in any suitable form.
- data structures may be shown to have fields that are related through location in the data structure. Such relationships may likewise be achieved by assigning storage for the fields with locations in a computer-readable medium that convey relationship between the fields.
- any suitable mechanism may be used to establish a relationship between information in fields of a data structure, including through the use of pointers, tags or other mechanisms that establish relationship between data elements.
- the disclosure relates to various embodiments in which one or more methods.
- the acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.
- the disclosure relates to a system that comprises at least one processor, a program storage, such as memory, for storing program code executable on the processor, and one or more input/output devices and/or interfaces, such as data communication and/or peripheral devices and/or interfaces.
- the user device and computer system or systems are communicably connected by a data communication network, such as a Local Area Network (LAN), the Internet, or the like, which may also be connected to a number of other client and/or server computer systems.
- the user device and client and/or server computer systems may further include appropriate operating system software.
- components and/or units of the devices described herein may be able to interact through one or more communication channels or mediums or links, for example, a shared access medium, a global communication network, the Internet, the World Wide Web, a wired network, a wireless network, a combination of one or more wired networks and/or one or more wireless networks, one or more communication networks, an a-synchronic or asynchronous wireless network, a synchronic wireless network, a managed wireless network, a non-managed wireless network, a burstable wireless network, a non-burstable wireless network, a scheduled wireless network, a non-scheduled wireless network, or the like.
- a shared access medium for example, a shared access medium, a global communication network, the Internet, the World Wide Web, a wired network, a wireless network, a combination of one or more wired networks and/or one or more wireless networks, one or more communication networks, an a-synchronic or asynchronous wireless network, a synchronic wireless network, a managed wireless network
- Discussions herein utilizing terms such as, for example, “processing,” “computing,” “calculating,” “determining,” or the like, may refer to operation(s) and/or process(es) of a computer, a computing platform, a computing system, or other electronic computing device, that manipulate and/or transform data represented as physical (e g., electronic) quantities within the computer's registers and/or memories into other data similarly represented as physical quantities within the computer’s registers and/or memories or other information storage medium that may store instructions to perform operations and/or processes.
- processing may refer to operation(s) and/or process(es) of a computer, a computing platform, a computing system, or other electronic computing device, that manipulate and/or transform data represented as physical (e g., electronic) quantities within the computer's registers and/or memories into other data similarly represented as physical quantities within the computer’s registers and/or memories or other information storage medium that may store instructions to perform operations and/or processes.
- Some embodiments may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment including both hardware and software elements. Some embodiments may be implemented in software, which includes but is not limited to firmware, resident software, microcode, or the like.
- some embodiments may take the form of a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system.
- a computer-usable or computer-readable medium may be or may include any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
- the system performs a computer-implemented method of selecting a neomer sequence in some embodiments.
- the methods are computer-implemented methods of selecting a therapy, analyzing data from a sample or diagnosing a subject comprising:
- the therapy is chosen or the diagnosis is made based opon the presence of the neomer.
- the sample is cell free.
- the neomer are chosen from one or a plurality of neomers identified in Table disclosed herein or one or a plurality of sequences that comprise at least about 85%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to any of the sequences identified in Table 1.
- the medium may be or may include an electronic, magnetic, optical, electromagnetic, InfraRed (IR), or semiconductor system (or apparatus or device) or a propagation medium.
- a computer-readable medium may include a semiconductor or solid state memory, magnetic tape, a removable computer diskette, a Random Access Memory (RAM), a Read-Only Memory (ROM), a rigid magnetic disk, an optical disk, or the like.
- RAM Random Access Memory
- ROM Read-Only Memory
- optical disks include Compact Disk-Read-Only Memory (CD-ROM), Compact Di sk-Read/Write (CD-R/W), DVD, or the like.
- a data processing system suitable for storing and/or executing program code may include at least one processor coupled directly or indirectly to memory elements, for example, through a system bus.
- the memory elements may include, for example, local memory employed during actual execution of the program code, bulk storage, and cache memories which may provide temporary storage of at least some program code in order to reduce the number of times code must be retrieved from bulk storage during execution.
- I/O devices including but not limited to keyboards, displays, pointing devices, etc.
- I/O controllers may be coupled to the system either directly or through intervening I/O controllers.
- network adapters may be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices, for example, through intervening private or public networks.
- modems, cable modems and Ethernet cards are demonstrative examples of types of network adapters. Other suitable components may be used.
- Some embodiments may be implemented by software, by hardware, or by any combination of software and/or hardware as may be suitable for specific applications or in accordance with specific design requirements. Some embodiments may include units and/or sub-units, which may be separate of each other or combined together, in whole or in part, and may be implemented using specific, multi-purpose or general processors or controllers. Some embodiments may include buffers, registers, stacks, storage units and/or memory units, for temporary or long-term storage of data or in order to facilitate the operation of particular implementations. Some embodiments may be implemented, for example, using a machine-readable medium or article which may store an instruction or a set of instructions that, if executed by a machine, cause the machine to perform a method steps and/or operations described herein.
- Such machine may include, for example, any suitable processing platform, computing platform, computing device, processing device, electronic device, electronic system, computing system, processing system, computer, processor, or the like, and may be implemented using any suitable combination of hardware and/or software.
- the machine-readable medium or article may include, for example, any suitable type of memory unit, memory device, memory article, memory medium, storage device, storage article, storage medium and/or storage unit; for example, memory, removable or non-removable media, erasable or non-erasable media, writeable or re-writeable media, digital or analog media, hard disk drive, floppy disk, Compact Disk Read Only Memory (CD-ROM), Compact Disk Recordable (CD-R), Compact Disk Re-Writeable (CD-RW), optical disk, magnetic media, various types of Digital Versatile Disks (DVDs), a tape, a cassette, or the like.
- CD-ROM Compact Disk Read Only Memory
- CD-R Compact Disk Recordable
- CD-RW Compact Disk Re-Write
- the instructions may include any suitable type of code, for example, source code, compiled code, interpreted code, executable code, static code, dynamic code, or the like, and may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled and/or interpreted programming language, e.g., C, C++, JavaTM, BASIC, Pascal, Fortran, Cobol, assembly language, machine code, or the like.
- code for example, source code, compiled code, interpreted code, executable code, static code, dynamic code, or the like
- suitable high-level, low-level, object-oriented, visual, compiled and/or interpreted programming language e.g., C, C++, JavaTM, BASIC, Pascal, Fortran, Cobol, assembly language, machine code, or the like.
- circuits may be implemented as a hardware circuit comprising custom very-large-scale integration (VLSI) circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components.
- VLSI very-large-scale integration
- a circuit may also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices or the like.
- the circuits may also be implemented in machine-readable medium for execution by various types of processors.
- An identified circuit of executable code may, for instance, comprise one or more physical or logical blocks of computer instructions, which may, for instance, be organized as an object, procedure, or function. Nevertheless, the executables of an identified circuit need not be physically located together, but may comprise disparate instructions stored in different locations which, when joined logically together, comprise the circuit and achieve the stated purpose for the circuit.
- a circuit of computer readable program code may be a single instruction, or many instructions, and may even be distributed over several different code segments, among different programs, and across several memory devices.
- operational data may be identified and illustrated herein within circuits, and may be embodied in any suitable form and organized within any suitable type of data structure.
- the operational data may be collected as a single data set, or may be distributed over different locations including over different storage devices, and may exist, at least partially, merely as electronic signals on a system or network.
- the computer readable medium (also referred to herein as machine-readable media or machine-readable content) may be a tangible computer readable storage medium storing the computer readable program code.
- the computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, holographic, micromechanical, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
- examples of the computer readable storage medium may include but are not limited to a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), an optical storage device, a magnetic storage device, a holographic storage medium, a micromechanical storage device, or any suitable combination of the foregoing.
- a computer readable storage medium may be any tangible medium that can contain, and/or store computer readable program code for use by and/or in connection with an instruction execution system, apparatus, or device.
- the computer readable medium may also be a computer readable signal medium.
- a computer readable signal medium may include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electrical, electro-magnetic, magnetic, optical, or any suitable combination thereof.
- a computer readable signal medium may be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport computer readable program code for use by or in connection with an instruction execution system, apparatus, or device.
- computer readable program code embodied on a computer readable signal medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, Radio Frequency (RF), or the like, or any suitable combination of the foregoing.
- the computer readable medium may comprise a combination of one or more computer readable storage mediums and one or more computer readable signal mediums.
- computer readable program code may be both propagated as an electro-magnetic signal through a fiber optic cable for execution by a processor and stored on RAM storage device for execution by the processor.
- Computer readable program code for carrying out operations for aspects of the present invention may be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages.
- the computer readable program code may execute entirely on a user's computer, partly on the user’s computer, as a stand-alone computer-readable package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server.
- the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).
- LAN local area network
- WAN wide area network
- Internet Service Provider for example, AT&T, MCI, Sprint, EarthLink, MSN, GTE, etc.
- the program code may also be stored in a computer readable medium that can direct a computer, other programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions which implement the function/act specified in the schematic flowchart diagrams and/or schematic block diagrams block or blocks.
- Recurrent nullomers (neomers) (ri) were annotated as those that resulted from substitutions or indels across two or more patients within a cancer type. When possible, ri was chosen to get -10,000 neomers from each tissue, otherwise it was set to 2 (Table 4). Driver mutation-derived neomers were defined as nullomers detected from driver mutations and were identified using the database (REF). ii. Classification of tumor tissue of origin using neomers
- the MMR status of each biopsy sample was derived from 85 .
- the model was trained on neomers identified in MSI samples and the performance of the algorithm evaluated. For the MSI versus the MS S samples, we counted the number of neomers that contained either AAAAAAAA or TTTTTTTT re p e ts, since MSI cancers have been associated with mutations of polyA/T repeats 35 .
- the threshold for determining MSI or MSS was set as the harmonic mean of the maximum number of counts in the MSS set and the minimum number of counts in the MSI set.
- the POLE deficiency status of each biopsy sample was derived from 85 and we used a similar strategy to that of MMR status, but instead counted neomers created through either a TCT>TAT or TCG>TTG mutation. Since the number of patients in each category was limited, we only used a 5-fold cross validation.
- Prostate cancer samples described in 86 , were obtained from the Witte lab at UCSF as extracted cfDNA kept at -80C. Extracted dsDNA concentration was measured using the Qubit High-Sensitivity dsDNA kit. For ovarian and lung cancer samples, cfDNA was extracted from ImL of plasma, following centrifugation at 4C at 600rpm for 3 minutes, to remove larger debris, using the QIAamp Circulating Nucleic Acid Kit (Qiagen).
- cfDNA was eluted in 50uL of elution buffer and measured using the Qubit High-Sensitivity dsDNA kit, and validated for size distribution (160-180bp) using an Agilent BioAnalyzer 2100 Sensitivity DNA chip.
- Up to lOng of cfDNA was used for sequencing library construction using the library preparation enzymatic fragmentation kit 2.0 (Twist Bioscience) adjusted with IDT’s xGen UDI-UMI 96 barcodes system (IDT) to replace the Twist universal adapter, and using KAPA HiFi polymerase instead of the polymerase provided by the kit.
- Frozen solid tumor samples from both ovarian and prostate cohorts were received from the Chapman and Witte labs respectively. Tumor masses were kept in dry ice and excised on a frozen tray, to yield out pieces of 1-8 mm A 3, and gDNA was extracted using the DNeasy Blood and Tissue kit (Qiagen), according to manufacturer instructions. gDNA concentrations were measured using NanoDrop. vi. Neomer identification in cfDNA samples
- Promoter sequences with and without the neomer were synthetically generated and cloned into the modified Promega promoter assay luciferase vector pGL4.1 lb (a gift from Dr. Rick Myers, HudsonAlpha) by BioMatik Inc and Sanger sequence verified.
- LNCaP cells were plated at an initial density of 2*10 A 5 cells/well in 24-well tissue culture plates and maintained in RPMI medium, 10% FBS supplemented with L-Glutamine and Penicillin/Streptomycin.
- Plasmids together with a renilla expressing plasmid, pGL4.74 (Promega), at a ratio of 10: 1 luciferase:renilla were transfected using the X-tremeGENETM HP DNA Transfection Reagent (Roche) using 1:4 ratio of DNA (ug) to reagent (ul). 72 hours post transfection luciferase and renilla levels were measured using the Dual-Luciferase Reporter Assay System (Promega) following the manufacturer’s protocol using a GloMax Explorer Multimode Microplate Reader (Promega). Luciferase activity was normalized to renilla levels and presented as Relative Luciferase Units (RLU). Statistical analysis was performed using Prism version 9.0.2 (GraphPad). All values were reported as means (AVG) and standard errors (SE). p values ⁇ 0.05 were considered statistically significant. viii. Software availability
- the package is composed of six functions: 1) EnumerateNullomers, which extracts all nullomers of specified kmer lengths in a FASTA sample; 2) ExtractMutationNullomers, which finds all mutations that cause the resurfacing of a list of nullomers; 3) IdentifyRecurrentNull omers, which identifies nullomers that recur in a dataset through mutagenesis; 4) FindAlmostNullomers, which identifies the positions that can create a list of nullomers genome-wide for every possible substitution and single base-pair insertion and deletion; 5) FindNullomerVariants, which removes nullomers that are likely to result from common variants in a user specified variant VCF file; 6) FindDNANullomersFromReads, which performs the identification of nullomers in raw read samples.
- the package can be found at: https://github.com/Ahitu
- Lentiviral bound MPRA was done as described previously 60 .
- An oligonucleotide library of 230bp long fragments bearing 1) 4,609 loci of recurrent mutations across prostate cancer patients, which cause neomer resurfacing and their reference genome pair; 2) 64 fragments to tile the 350bp long TMEM127-CIAO1 and RPS2-SNHG9 promoters used in the luciferase assay; as well as 3) 100 scrambled control sequences.
- Lentivirus was produced and titered, later to be used to infect prostate cancer cell line (LNCaP, DU-145 and PC-3) at an MOI of 50 virus particles per cell.
- DNA and RNA were extracted from the 3 replicates of the cell culture experiment used for library construction and multiplexed for NGS. All MPRA-related sequencing was performed using an Illumina NextSeq 500 (Novogene) with either PEI 50 for the CRS-BC association library or PEI 5 for the DNA/RNA BC count portion of the protocol.
- nullomers As cancer is associated with a large number of somatic DNA mutations, we investigated if they can result in the resurfacing of nullomers (Fig. 1A). Using our previously characterized human nullomers 19 , we analyzed WGS results from 2,577 patients across 21 different cancer types from TCGA22 for resurfacing nullomers (Fig. 5A). We focused on 16bp nullomers, as it is the shortest length where we detect a sufficient number of nullomers per patient, with the human reference genome having only 37.24% of all possible 16mers. The majority of the 44,599,472 single nucleotide substitutions give rise to multiple nullomers, allowing us to identify 213, 164,038 resurfacing nullomers across all cancer types.
- nullomers that could be used as cancer biomarkers
- the number of neomers was proportional to the total number of mutations (Fig. 1C, Table 4). As both the number of patients per cancer type and the mutational load varied, the median number of neomers for each tissue type ranged from 0-98. Analysis of the most frequent neomers revealed several previously known cancer-associated mutations (Table 1).
- KRAS KRAS proto-oncogene GTPase
- telomerase reverse transcriptase TERT
- TERT tumor protein p53
- BRAF B-Raf proto-oncogene serine/threonine kinase
- PJK3CA phosphatidylinositol-4,5-bisphosphate 3-kinase catalytic subunit alpha
- This mutation is extremely common in numerous cancer types 26 and is thought to disrupt a G-quadruplex 27 leading to the binding of GAPB 28 , an ETS transcription factor, resulting in increased TERT expression.
- GAPB 28 an ETS transcription factor
- Table 1 Common cancer-associated neomers. Six of the most common neomers created by a single mutation. The nucleotide in red is the neomer causing mutation. Table 4. Minimal recurrency thresholds and associated number of neomers per tissue type. We also identified several neomers that are frequently created by different mutations (Table 5). Interestingly, some of these frequently recurrent neomers are created by different mutations, yet are predominantly found in one cancer. For example, GTTTTTCTCCTAGACC is found 40 times in skin cancer at 31 different loci while CTGGCAGTGAGCCACG is found 21 times in liver cancer across 18 loci.
- CGACGTTCTGCCCACT is found in 32 loci, primarily in pancreatic and stomach cancer. Of those loci, 21/32 (65.6%) were found in noncoding regions nearby pancreatic cancer associated genes.
- CCL4 C-C motif chemokine ligand
- POM121L12 POM121 transmembrane nucleoporin like 12
- KCNV1 potassium voltage-gated channel modifier subfamily V member 1
- driver mutation-derived neomers the neomers detected at driver mutation loci.
- driver mutation-derived neomers the neomers detected at driver mutation loci.
- driver mutation-derived neomers the neomers detected at driver mutation loci.
- Fig. 1 the pan-cancer analysis, we identified 19,594,212 neomers resulting from 50,167 putative driver mutations (Fig. 1).
- 81% of driver mutations resulted in one or more neomers, ranging between 63.88% and 86.50% in pancreatic and lung squamous adenocarcinoma respectively.
- neomers also varied by cancer type, ranging between 92 in pancreatic cancer and 9,434 in cutaneous melanoma (Fig. ID).
- driver mutations were 1.4-fold more likely to result in the creation of a neomer.
- MSI microsatellite unstable
- MSS microsatellite stable
- Neomers detect cancer in cfDNA
- neomers could be used to diagnose cancer in cfDNA.
- For each lung cancer associated neomer we characterized all possible single nucleotide substitutions in the reference genome that could give rise to this neomer.
- an important computational advantage compared to conventional mutation calling pipelines is that only a single pass is made across the reads to identify those containing neomers, and only those reads are aligned to the reference genome.
- a median neomer number of 396 for the controls and 639 for the cancer patients p-value ⁇ 0.005, Mann-Whitney U
- a classifier that compares the number of detected neomers to a threshold, achieving an Fl -score of 0.82 using 2-fold cross validation (Fig. 3B).
- our classifier also performs well for early stages with only a slight drop in performance for stage I.
- localized prostate cancer has a low abundance of ctDNA making it difficult to detect by ultra low pass WGS or targeted cfDNA sequencing 48 or via methylation 49 compared to metastatic 50 , providing a challenging test for our neomer approach.
- We generated cfDNA WGS datasets from twelve controls and eight localized prostate cancer patients. Searching for 4,621 neomers, we identified a median of 4 in the controls and 8 in the patients (p-value 0.069, Mann-Whitney U) (Fig. 3F).
- the neomers can serve as a sensitive and specific indicator for this type of cancer, as the classifier achieved an Fl score of 0.67.
- neomers in: 1) a promoter between two divergent genes, RPS2 and the IncRNA gene SNHG9 (Fig. 4A), both of which are overexpressed in prostate cancer 54 ; 2) a promoter between two divergent genes, TMEM127 and CIAO1 (Fig.
- neomers could be used to identify driver mutations in enhancers.
- lentiMPRA lentivirus-based MPRA
- Fig. 5A Two hundred base pair sequences, where the position of the neomer is used as a center, were synthesized and cloned upstream of a minimal promoter followed by a GFP reporter gene.
- RNA/DNA ratio > 1.5).
- Graph. 5B GO analysis of sequences leading to increased activity found enrichment for terms related to cell-to-cell junctions and gamma-catenin binding (G0:0045295) (Fig. 5C), which is important in cell-cell adhesion and has been associated with prostate cancer progression through interaction with the beta-catenin Wnt signaling axis 62.
- the neomer that showed the highest level of increased activity compared to reference (4.09 fold) is located in an intron of the catenin alpha 1 (CTNNA1) gene.
- This gene is a core member of the cadherin/catenin complex and is involved in the regulation of the Wnt/beta-catenin pathway, which has been widely studied in cancer 63 , including prostate cancer64, and is the target of several therapies 63,66 .
- TFBS analysis of the neomer found that it leads to a gain of a TCF7L1 motif and loss of a STAT2 TFBS (Fig. 5D), both of which were shown to play a role in prostate cancer malignancy 67,68 .
- a neomer residing in the 3rd intron of HERC3 gene resulted in 2.7 fold downregulation of the reporter activity in our MPRA.
- HERC3 gene had been shown to inhibit metastasis of colorectal cancer and was indicated to be downregulated in colorectal cancer and its downregulation showed poor overall survival (OS) and disease-free survival (DFS) in colorectal patients (Zhang et al. 2022). Taken together, these results demonstrate that neomers could be utilized to identify gene regulatory driver mutations in cancer.
- Cancer is a DNA mutation associated disease.
- WGS cancer-associated DNA mutations
- neomers short sequences that are predominantly absent from genomes of healthy individuals.
- Further analyses of these sets of neomers show that they can be used not only to classify cancer tissue of origin, but also additional cancer features, such as MSI or POLE deficiency with high accuracy.
- MSI MSI
- POLE additional cancer features
- Analysis of cfDNA WGS datasets finds that neomers could be used to tease out patients from controls in several cancers, including those with a low mutational burden.
- reporter assays we show that neomers have a functional effect on regulatory sequences.
- cfDNA detection approach has several advantages over current methods: 1) Detecting short DNA sequences that are enriched in cancer samples provides an easy to use diagnostic that could allow detection from low amounts of ctDNA. In addition to sequencing-based assays, alternate techniques could potentially be used, such as CRISPR-based detection tools that utilize Casl2 or Casl3 74 , that can also allow the testing of thousands of sequences in parallel 75 . In addition, with neomer-based diagnostics potentially not needing large amounts of starting material, cfDNA could be collected from urine, sputum, saliva or other bodily fluids, which were shown to be a viable but reduced source of cfDNA 76,77 .
- cfDNA fragmentation patterns combined with CT imaging, clinical risk factor and serum levels of carcinoembryonic antigen significantly increased the ability to diagnose lung cancer 79 .
- Adding neomers to known cancer-associated coding mutations in the screening of cfDNA could increase sensitivity and specificity.
- coupling neomer-based diagnostics to existing cancer biomarkers and risk factors could improve the power to detect various cancer subtypes.
- nullomers/neomers do not exist in the human genome they could also be exceptional candidates for neoantigens, to be targeted via immunotherapy.
- Previous work has shown that minimal absent words, short sequences that are absent from a genome or proteome, could be used to identify phosphorylation sites of high confidence, some of which could be associated with cancer 80 .
- Analysis of the Immune Epitope Database of validated antigens 81 found that 13 of the recurrent coding neomers can create neoantigens with predicted strong binding levels that were subsequently validated (Table 7).
- Neomers can be used as a novel tool to identify cancer-associated gene regulatory mutations.
- Our MPRA library of 4,609 neomer causing mutations in enhancers revealed that 2.6% can change gene expression by >1.5-fold, suggesting that a subset of these mutations could have important functional consequences in cancer.
- neomer-based screening with clinical characteristics and additional diagnostic tools/features could increase the positive predictive value.
- cfDNA could also be isolated from urine and saliva, and detection of these sequences only requires a relatively small amount of DNA, neomer-based diagnosis could be carried out in a non-invasive manner.
- neomers could be used to highlight cancer-associated gene regulatory mutations which have been difficult to identify. Further high-throughput characterization of these mutations could allow the detection of bona fide cancer-associated functional regulatory mutations that could be used for diagnosis and treatment.
- TMEM127 The tumor susceptibility gene TMEM127 is mutated in renal cell carcinomas and modulates endolysosomal function. Hum. Mol. Genet. 23, 2428-2439 (2014).
- Li, D. et al. FOXD3 is a novel tumor suppressor that affects growth, invasion, metastasis and angiogenesis of neuroblastoma.
Landscapes
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Physics & Mathematics (AREA)
- Medical Informatics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Health & Medical Sciences (AREA)
- Pathology (AREA)
- Public Health (AREA)
- Analytical Chemistry (AREA)
- Biotechnology (AREA)
- Organic Chemistry (AREA)
- Genetics & Genomics (AREA)
- Biophysics (AREA)
- Biomedical Technology (AREA)
- Theoretical Computer Science (AREA)
- Immunology (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Data Mining & Analysis (AREA)
- Databases & Information Systems (AREA)
- Epidemiology (AREA)
- Molecular Biology (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Computation (AREA)
- Oncology (AREA)
- Hospice & Palliative Care (AREA)
- Microbiology (AREA)
- Software Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Bioethics (AREA)
- Biochemistry (AREA)
- General Engineering & Computer Science (AREA)
- Artificial Intelligence (AREA)
- Primary Health Care (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202263424478P | 2022-11-10 | 2022-11-10 | |
| PCT/US2023/079380 WO2024103003A2 (en) | 2022-11-10 | 2023-11-10 | Systems for mutation caller and methods of using the same |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4616005A2 true EP4616005A2 (de) | 2025-09-17 |
Family
ID=91033480
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23889780.5A Pending EP4616005A2 (de) | 2022-11-10 | 2023-11-10 | Systeme für mutationsanrufer und verfahren zur verwendung davon |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4616005A2 (de) |
| WO (1) | WO2024103003A2 (de) |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| AU2002363628A1 (en) * | 2001-07-17 | 2003-05-26 | Stratagene | Methods for detection of a target nucleic acid by capture using multi-subunit probes |
| US8927213B2 (en) * | 2004-12-23 | 2015-01-06 | Greg Hampikian | Reference markers for biological samples |
-
2023
- 2023-11-10 EP EP23889780.5A patent/EP4616005A2/de active Pending
- 2023-11-10 WO PCT/US2023/079380 patent/WO2024103003A2/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024103003A3 (en) | 2025-02-27 |
| WO2024103003A2 (en) | 2024-05-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP3421613B1 (de) | Identifikation und verwendung von zirkulierenden nukleinsäure-tumormarkern | |
| US12398429B2 (en) | Methods and systems for sequencing polynucleotides | |
| US20240229157A1 (en) | Compositions comprising nullomers and methods of using the same for cancer detection and diagnosis | |
| Xiao et al. | Non‐invasive diagnosis and surveillance of bladder cancer with driver and passenger DNA methylation in a prospective cohort study | |
| US20220411878A1 (en) | Methods for disease detection | |
| WO2019232483A1 (en) | Detection method | |
| CA3152887A1 (en) | Novel biomarkers and diagnostic profiles for prostate cancer integrating clinical variables and gene expression data | |
| CA3208638A1 (en) | Cell-free dna methylation test | |
| JP2019514344A (ja) | 癌のエピジェネティックプロファイリング | |
| CN116261600A (zh) | 用于癌症诊断的oncrna的检测的系统和方法 | |
| Michel et al. | Noninvasive multicancer detection using DNA hypomethylation of LINE-1 retrotransposons | |
| ES3061610T3 (en) | Methods for detecting nucleic acid variants | |
| US20250297320A1 (en) | Methylation signatures in cell-free dna for tumor classification and early detection | |
| Strauss et al. | Analysis of tumor template from multiple compartments in a blood sample provides complementary access to peripheral tumor biomarkers | |
| Shi et al. | Field-effect-informed urine liquid biopsy for bladder cancer | |
| WO2024103003A2 (en) | Systems for mutation caller and methods of using the same | |
| US9476096B2 (en) | Recurrent gene fusions in hemangiopericytoma | |
| WO2019178214A1 (en) | Methods and compositions related to methylation and recurrence in gastric cancer patients | |
| US20250101510A1 (en) | Methods and systems for sequencing polynucleotides | |
| Vasseur et al. | Transcription Factor Subtype Governs Response and Resistance to DLL3-Directed T-Cell Engagement in Small Cell Lung Cancer | |
| De Michino | Exploration of Epigenetic Profiles in Circulating Cell-Free Chromatin to Identify Predictive Cancer Biomarkers | |
| WO2025160307A1 (en) | Treating cancer using biomarkers and gene signatures | |
| Xu et al. | Identification of an Excellent PCR-Based Classifier to Predict Tumor Relapse in Stage II/III Colorectal Cancer and Its Clinical Application Irrespective of Consensus Molecular Subtypes | |
| EP4298250A1 (de) | Marker zur vorhersage der reaktion auf eine car-t-zelltherapie | |
| HK40032567A (en) | Non-coding rna for detection of cancer |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250520 |
|
| AK | Designated contracting states |
Kind code of ref document: A2 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |