WO2023178152A1 - Improved methods of predicting response to immune checkpoint blockade therapies and uses thereof - Google Patents

Improved methods of predicting response to immune checkpoint blockade therapies and uses thereof Download PDF

Info

Publication number
WO2023178152A1
WO2023178152A1 PCT/US2023/064398 US2023064398W WO2023178152A1 WO 2023178152 A1 WO2023178152 A1 WO 2023178152A1 US 2023064398 W US2023064398 W US 2023064398W WO 2023178152 A1 WO2023178152 A1 WO 2023178152A1
Authority
WO
WIPO (PCT)
Prior art keywords
tmb
genes
ddr
subject
bets
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2023/064398
Other languages
French (fr)
Inventor
William Y. KIM
Peter J. MUCHA
William H. WEIR
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
University of North Carolina at Chapel Hill
Original Assignee
University of North Carolina at Chapel Hill
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by University of North Carolina at Chapel Hill filed Critical University of North Carolina at Chapel Hill
Publication of WO2023178152A1 publication Critical patent/WO2023178152A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/30ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for calculating health indices; for individual health risk assessment
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6883Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
    • C12Q1/6886Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material for cancer
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B25/00ICT specially adapted for hybridisation; ICT specially adapted for gene or protein expression
    • G16B25/10Gene or protein expression profiling; Expression-ratio estimation or normalisation
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/106Pharmacogenomics, i.e. genetic variability in individual responses to drugs and drug metabolism
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/118Prognosis of disease development
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/156Polymorphic or mutational markers

Definitions

  • Program 2 is entitled “150-33-2_bipartite_helper_functions.txt” and 8,485 bytes in size.
  • Program 3 is entitled “150-33- 3_bipartite_matching.txt” and 13,493 bytes in size.
  • Program 4 is entitled “150-33- 4_ddr_data_object.txt” and 6,623 bytes in size.
  • Program 5 is entitled “150-33-5_file_locations.txt” and 1,649 bytes in size.
  • Program 6 is entitled “150-33-6_load_clinical_datasets.txt” and 39,528 bytes in size.
  • Program 7 is entitled “150-33-7_load_pmec.txt” and 11,478 bytes in size.
  • Program 8 is entitled “150-33-8_load_tcga_dataset.txt” and 19,905 bytes in size.
  • Program 9 is entitled “150- 33-9_name_matching_scripts.txt” and 1,375 bytes in size.
  • Supplemental Table 1 is entitled 150-33- PCT_SUPP_TABLE_1 and is 1,679,792 bytes in size.
  • Supplemental Table 2 is entitled 150-33- PCT_SUPP_TABLE_2 and is 7,058,273 bytes in size. They are hereby incorporated by reference in their entireties. 1.
  • the present disclosure provides an improved method of selecting patients for immune checkpoint blockade (ICB) treatment that complements existing methods of measuring total mutational burden (TMB) of a particular cancer in a subject.
  • TMB total mutational burden
  • the “background” description provided herein is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in this background section, as well as aspects of the description which may not otherwise qualify as prior art at the time of filing, are neither expressly nor impliedly admitted as prior art against the present disclosure.
  • Immune checkpoint blockade (ICB) has achieved remarkable success in many solid tumors.
  • TMB tumor mutational burden
  • RNA expression signatures i.e. a T cell inflamed gene expression profile
  • TMB represents the balance between a tumor's exposure to a mutagenic process (i.e. UV radiation, carcinogen, etc.) and the integrity of the cellular DNA Damage Repair (DDR) pathways.
  • a mutagenic process i.e. UV radiation, carcinogen, etc.
  • DDR DNA Damage Repair
  • Bipartite Graph-Based Expected TMB Score (BiG-BETS), that resolves the TMB Paradox, accurately defines genes associated with elevated TMB, and remarkably delineates a cohort of subjects (TMB-High, Low BiG-BETS DDR mutant) with high predictive power for ICB response and prolonged overall survival.
  • the present disclosure provides a method for predicting response to immune checkpoint blockade (ICB) therapy or overall survival for a subject with cancer, comprising independently measuring or obtaining (i) a tumor mutational burden (TMB) level and (ii) defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5 in a sample from the subject, normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample, where
  • a method to select a subject with cancer for immune blockade therapy comprises independently measuring or obtaining both (i) a tumor mutational burden (TMB) level and (ii) defects in DNA Damage Repair (DDR) genes or their expression products in a sample from the subject; calculating a Bipartite Graph-Based Expected TMB Score (BiG-BETS) to determine which DDR genes or expression products are associated with elevated TMB; if the subject has a high BiG-BETS determining that the defect in the DDR genes or their expression products do not add to the TMB level over measurement of TMB levels alone; and using the TMB levels alone to select a subject for ICB.
  • ICB immune blockade therapy
  • the DDR genes or their expression products may be at least ten biomarkers selected from the group consisting of BARD1, CHEK1, CUL5, EME1, ERCC4, ERCC5, ERCC6, FANCA, FANCC, FANCI, FEN1, GEN1, LIG4, MDC1, MLH1, MLH3, MRE11, MSH2, MSH3, MSH6, NBN, PALB2, PARP1, PMS1, PMS2, POLE, POLE3, POLL, POLM, RAD50, RNMT, SEM1, SLX1A, TDG, TOP3A, TOPBP1, UNG, XPC, XRCC2, XRCC4, and XRCC6.
  • biomarkers selected from the group consisting of BARD1, CHEK1, CUL5, EME1, ERCC4, ERCC5, ERCC6, FANCA, FANCC, FANCI, FEN1, GEN1, LIG4, MDC1, MLH1, MLH3, MRE11, MSH2, MSH3, MSH6,
  • a method to select a subject with cancer for immune blockade therapy comprises independently measuring or obtaining both (i) a tumor mutational burden (TMB) level and (ii) defects in DNA Damage Repair (DDR) genes or their expression products in a sample from the subject; calculating a Bipartite Graph-Based Expected TMB Score (BiG-BETS) to determine which DDR genes or expression products are associated with elevated TMB; if the subject has a low BiG-BETS determining that the defects in the DDR genes or their expression products contribute to the likelihood of benefit from ICB over measurement of TMB levels alone; and using both the TMB levels and the defects in the DDR genes or their expression products to select a subject for ICB.
  • ICB immune blockade therapy
  • the DDR genes or their expression products may be at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5.
  • biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1,
  • the cancer may be bladder urothelial carcinoma, colon adenocarcinoma, esophageal carcinoma, invasive breast carcinoma, head and neck squamous cell carcinoma, kidney renal clear cell carcinoma, lung adenocarcinoma, melanoma, or squamous cell lung carcinoma and the defects may be mutations or copy number alterations.
  • the mutations may be deletions, frameshift mutations, insertions, missense mutations, nonsense mutations, start codon loss, stop codon loss or gain, or a combination thereof.
  • the detecting defects in nucleic acids encoding genes, or their expression products may comprise performing next generation sequencing (NGS), nucleic acid hybridization, quantitative RT-PCR, immunohistochemistry (IHC), immunocytochemistry (ICC), or immunofluorescence (IF).
  • NGS next generation sequencing
  • IHC immunohistochemistry
  • ICC immunocytochemistry
  • IF immunofluorescence
  • the method may further comprises assessment of a medical history, a family history, a physical examination, an endoscopic examination, imaging, a biopsy result, or a combination thereof.
  • the method may be used to develop a treatment strategy for the subject with cancer.
  • the nucleic acids encoding genes may be isolated from a fixed, paraffin-embedded sample from the subject.
  • the nucleic acids encoding genes are isolated from core biopsy tissue or fine needle aspirate cells from the subject.
  • the method may further comprises treating the subject with a combination of ICB therapy and kinase inhibitor therapy.
  • the disclosure provides a method for treating a subject with cancer which comprises independently measuring or obtaining a tumor mutational burden (TMB) level; and defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of low BIG- BETS DDR genes normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample; and if the subject has a high TMB level and wild type low BIG- BETs genes, treating the subject with a combination of immune checkpoint blockcade (ICB) therapy and an inhibitor of a low BIG-BETs kinase so as to reduce the activity of the low BIG-BET kinase and thereby treat the subject
  • TMB tumor mutational burden
  • the low BIG-BETs kinase may be ATR, CHEK1, or WEE1.
  • a kit comprising at least ten nucleic acid probes, wherein each of said probes specifically binds to one of ten distinct biomarker nucleic acids or fragments thereof selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5.
  • Fig. 1A-Fig. 1F Univariate testing inappropriately associates most genes with elevated TMB and recasting samples and mutations as a bipartite network overcomes this limitation.
  • Fig.1A Distribution of Mann-Whitney U test p-values (with multiple test correction) on reversed log scale across all genes in Pan-TCGA dataset (light gray) vs DDR genes only (dark gray). For each gene, the MWU test compares distribution of TMB values for samples with a mutation in the gene vs all samples in the cohort. Right of dashed black line represents p ⁇ 0.05.
  • FIG.1B Percentage of genes in which mutations are significantly associated with elevated TMB by the MWU test (with FDR correction) broken down by DDR genes (dark gray) and non-DDR genes (light gray)
  • FIG.1C- Fig.1D Distribution of mean TMB values for mutated sample set for all genes (light gray) vs DDR genes (dark gray) in TCGA. Vertical dotted black line in Fig. 1D denotes the overall mean TMB for the cohort of samples.
  • FIG. 1E Schematic representation of converting the mutational data in a matrix to a bipartite network.
  • FIG. 1F Schematic representation of BiG-BETS network rewiring process to sample from the bipartite configuration model.
  • Mann-Whitney U p-values are calculated by comparing the distribution of TMB for samples with any mutation in the genes that define a pathway (counting only once if they have multiple mutations) to the distribution of all samples, including those with mutations in the same DDR pathway.
  • Fig.2B Schematic and example of the TMB paradox. Note how the addition of a single, highly mutated sample (right) only raises the average degree from 1.3 to 2 ( ⁇ 50% increase), while raising the average gene’s neighbor’s degree from 1.25 to 3 (140% increase). This effect is even larger network with a heavy tailed degree distribution.
  • Fig.3A-Fig.3C Distributions of p-values for the BiG-BETS permutation test applied to all 18,000 genes in the TCGA dataset, split by non-DDR genes (light gray) vs DDR genes (dark gray).
  • Fig. 3B Expected distribution of TMB values (shown in gray) for samples with mutation in ATR (top) and CHEK1 (bottom). Observed distribution of TMB for mutated tumors is shown by lighter gray curve. Distribution of means across the 400 samples from the network model shown in figure inset with dashed vertical line denoting observed mean TMB.
  • Fig. 3C Comparison of z-scores derived from Samstein et al.
  • Fig. 4A-Fig. 4D A networks-based model and permutation test (BiG-BETS) are superior to univariate model.
  • Fig. 4A Percentage of genes that are significant (with multiple test correction) using MWU test (left two bars) and significant by the networks-based test (right two bars). Differences between DDR vs non-DDR genes computed assessed with Chi-squared test with p-value shown above the corresponding bars.
  • Fig. 4B (Distribution of BiG-BETS z-scores for the DDR genes.
  • Fig. 5A-Fig. 5C Fig. 5A Schematic depicting criteria for filtering variants in the TCGA cohort used to calculate the BiG-BET scores. Variants were filtered on basis of impact as defined by the sequence ontology and polphen status (see Methods for full details).
  • Fig. 5B Comparison of BiG-BET scores obtained from TCGA cohort based on High Consequence mutations only (x-axis) and Moderate+High Consequences (y-axis). From left to right all genes (with hotspot genes shown in gray), DDR genes only, and the DDR pathways.
  • FIG. 6B Percentage of patients that have complete or partial response in TMB-H vs TMB-L for IMvigor210. Significance assessed using Chi-square test.
  • Fig. 6C Kaplan-Meier curves depicting overall survival (OS) in the IMvigor210 (left) and Samstein (right) cohorts broken down by TMB-H (light gray-lines) vs TMB-L (dark gray-lines) and into samples with a mutation in High-BiG-BETS DDR gene (bold lines) and High-BiG-BETS DDR WT tumors (dotted lines).
  • FIG. 6D Percentage of patients that have complete or partial response in High BiG-BETS DDR mutant vs wild type in TMB-H and TMB-L tumors from IMvigor210. Significance assessed using Chi-square test.
  • Fig. 7A-Fig. 7I TMB high tumors with mutation in a Low BiG-BETS DDR gene have improved survival and response.
  • FIG.7A Kaplan-Meier curves depicting overall survival (OS) in the IMvigor210 cohort broken down by TMB-High (light gray-lines) vs TMB-Low (dark gray-lines) and into samples with a mutation in Low BiG-BETS DDR genes (bold lines) and Low BiG-BETS DDR WT tumors (dotted lines).
  • Table underneath shows forest plot of coefficients for CPH model jointly testing TMB (as continuous variable), mutation in Low BiG-BETS DDR genes, as well as an interaction term between the two variables (denoted by Low BiG-BETS; TMB-H).
  • TMB TMB-High vs TMB-Low
  • WT Low BiG-BETS DDR gene mutation or not
  • the number of patients who were responders (CR/PR) in each category from left to right is 12, 1, 28, and 10 respectively.
  • CR/PR responders
  • Fig. 7D KM curves depicting OS in the Weir metadataset (see Methods for full description) broken down along the same lines as (A) and (C) with corresponding coefficients in CPH model below.
  • Patient counts for each category in TMB-H_MUT, TMB-H_WT, TMB-L_MUT, and TMB-L_WT were 60, 117, 33, and 201 respectively.
  • FIG.7E Response rates by Low BiG-BETS DDR mutations in the Weir Metadataset.
  • Fig. 7G- Fig. 7I KM curves depicting OS in a combined dataset that includes IMVigor210, Samstein et al., and Weir metadataset split out by tumor type including (Fig. 7G) bladder cancer, (Fig.7H) non-small cell lung cancer, and (Fig.7I) melanoma.
  • FIG. 7A bladder cancer
  • FIG.7H non-small cell lung cancer
  • Fig.7I melanoma.
  • KM Kaplan-Meier (KM) curves depicting overall survival (OS) in the Samstein cohort broken down by TMB-High (light gray-lines) vs TMB-Low (dark gray-lines) and into samples with a mutation in Low BiG-BETS DDR genes (bold lines) using only 17 DDR genes that overlap with IMvigor210 mutations (see Fig.5C) (bold lines) vs those that are Low BiG- BETS DDR WT (dashed line).
  • Fig. 9A-Fig. 9D Mutation of Low BiG-BETS DDR genes in TMB High tumors is associated with elevated STING and IRF3 gene signatures.
  • FIG. 9A- Fig. 9D Box plots of indicated gene signatures in IMVigor210 patients stratified by TMB (TMB-High vs TMB-Low) and Low BIG-BETS DDR gene mutation or not (WT). For each signature, each sample is assigned a z-score based on the average expression level of all genes in the signature compared to the average across all samples (See Methods). Significance calculated using the Mann-Whitney U test.
  • Fig.10A-Fig.10D Box plots of indicated gene signatures in TCGA tumors subsetted to match Samstein (see Methods) stratified by TMB (TMB-High vs TMB-Low) and Low BIG-BETS DDR gene mutation or not (WT). For each signature, each sample is assigned a z-score based on the average expression level of all genes in the signature compared to the average across all samples (See Methods). Significance calculated using the Mann-Whitney U test. (Fig.10E) Proportion of genes in each DDR pathway that are categorized as High and Low BIG-BETS genes. [0035] Fig.11A and Fig. 11B shows a graphical abstract of the disclosure.
  • Fig.11A Genes and samples are represented as a bipartite network.
  • a Bipartite Graph-Based Expected TMB (Big- BET) score is determined by a random rewiring of the bipartite network.
  • Fig.11B for subjects with mutations in BiG-BETs high genes, their ICB response is driven by their TMB levels. Subjects with mutations in the BiG-BETs low genes and high TMB, are more likely to respond to ICB and show increased overall survival (OS). 5.
  • DETAILED DESCRIPTION OF THE DISCLOSURE [0036] Immune checkpoint blockade (ICB) has had remarkable success for treatment of solid tumors.
  • TMB tumor mutational burden
  • DDR DNA Damage Repair
  • Examples of monoclonal antibody kinase inhibitors are trastuzumab (Herceptin®), an inhibitor of ERB-B2 and approved for breast cancer or bevacizumab (Avastin®), an inhibitor of vascular endothelial growth factor (VEGF) approved for colorectal cancer.
  • Other examples of drugs approved with a companion diagnostic include drugs approved for BRCA1/2 mutations, KRAS mutations and cKIT expression. Table 1 lists a number of approved drugs including a number of kinase inhibitors. See, Janne et al., 2009 Nat. Rev. Drug Disc.8709-723; Levitzki and Klein, 2010 Mol. Aspects Med.31, 287-329; and Mellor et al.2011 Tox. Sci.120(1) 14-32; and the package inserts for the specific drugs. [0039] TABLE 1
  • ALL acute lymphoblastic leukemia
  • AML acute myeloid leukemia
  • BrCA breast cancer
  • BCC basal cell carcinoma
  • CML chronic myeloid leukemia
  • CMML chronic myelomonocytic leukemia
  • CRC colorectal cancer
  • CSCC cutaneous squamous cell carcinoma
  • GIST gastrointestinal stromal tumor
  • HCC hepatocellular carcinoma
  • HNSCC head and neck squamous cell carcinoma
  • MCC Merkle cell carcinoma
  • MDS/MPD myelodysplastic syndrome/myeloproliferative disease
  • NSCLC non-small cell lung cancer
  • OvCA ovarian cancer
  • RCC renal cell carcinoma
  • STS soft tissue sarcoma
  • TNBC triple negative breast cancer
  • UC urothelial carcinoma.
  • trastuzumab (Herceptin®) is approved for breast cancer over expressing ERB-B2 and cetuximab (Erbitux®) for patients with wild-type KRAS.
  • trastuzumab (Herceptin®) is approved for breast cancer over expressing ERB-B2 and cetuximab (Erbitux®) for patients with wild-type KRAS.
  • trastuzumab (Herceptin®) is approved for breast cancer over expressing ERB-B2 and cetuximab (Erbitux®) for patients with wild-type KRAS.
  • kinase inhibitor approved for use with a diagnostic is crizotinib (Xalkori®) approved with a fluorescent in situ hybridization (FISH) test for ALK rearrangements (Vysis LSI ALK Dual Color, Break Apart Rearrangement Probe; Abbott Molecular, Abbott Park, IL).
  • FISH fluorescent in situ hybridization
  • Vemurafenib (Zelboraf®) is approved for use in patients with BRAF V600E mutation (Cobas 4800 BRAF V600 Mutation Test, Roche Molecular Diagnostics, Pleasanton, CA). Chapman et al., 2011 NEJM 3642507- 2516.
  • Non-limiting examples for bladder cancer include erdafitinib (BALVERSATM) or pembrolizumab (KEYTRUDA®) in Table 1, additional therapies include avelumab (BAVENCIO®), durvalumab (IMFINZITM), or nivolumab (OPDIVO®).
  • Non-limiting examples for BrCA include abemaciclib (VERZENIO®), ado- trastuzumab emtansine (KADCYLA®), alpelisib (PIQRAY®), atezolizumab (TECENTRIQ®), Everolimus (AFINITOR®), lapatinib (TYKERB®), olaparib (LYNPARZA®), palbociclib (IBRANCE®), pertuzumab (PERJETA®), ribociclib (KISQALI®) or trastuzumab (HERCEPTIN®), or trastuzumab (HERCEPTIN HYLECTA TM ) in Table 1, additional therapies include anastrozole (ARIMIDEX®), exemestane (AROMASIN®), fulvestrant (FASLODEX®), letrozole (FEMARA®), neratinib (NERLYNXTM), tamoxifen (SOLTAMOX®), or toremifene
  • Non-limiting examples for CRC include bevacizumab (AVASTIN®), Cetuximab (ERBITUX®), panitumumab (VECTIBIX®), ramucirumab (CYRAMZA®), or regorafenib (STIVARGA®) in Table 1, additional therapies include ipilimumab (YERVOY®), nivolumab (OPDIVO®), or ziv-aflibercept (ZALTRAP®).
  • HCC include pembrolizumab (KEYTRUDA®), ramucirumab (CYRAMZA®), regorafenib (STIVARGA®), or sorafenib (NEXAVAR®) in Table 1, additional therapies include cabozantinib (CABOMETYXTM), lenvatinib (LENVIMA®), or nivolumab (OPDIVO®).
  • kidney cancer examples include axitinib (INLYTA®), bevacizumab (AVASTIN®), cabozantinib (CABOMETYX®), Everolimus (AFINITOR®), pazopanib (VOTRIENT®), pembrolizumab (KEYTRUDA®), sorafenib (NEXAVAR®), sunitinib (SUTENT®), temsirolimus (TORISEL®) in Table 1, additional therapies include avelumab (BAVENCIO®), ipilimumab (YERVOY®), lenvatinib mesylate (LENVIMA®), or nivolumab (OPDIVO®).
  • Non- limiting examples for leukemia include dasatinib (SPRYCEL®), enasidenib (IDHIFA®), gilteritinib (XOSPATA®), imatinib (GLEEVEC®), ivosidenib (TIBSOVO®), midostaurin (RYDAPT®), nilotinib (TASIGNA®), or venetoclax (VENCLEXTA®) in Table 1, additional therapies include alemtuzumab (CAMPATH®), blinatumomab (BLINCYTO®), bosutinib (BOSULIF®), duvelisib (COPIKTRATM), gemtuzumab ozogamicin (MYLOTARGTM), glasdegib (DAURISMOTM), ibrutinib (IMBRUVICA®), idelalisib (ZYDELIG®), inotuzumab ozogamicin (BESPONSA®), moxe
  • Non-limiting examples for lung cancers include in Table 1 afatinib (GILORAF®), alectinib (ALECENSA®), atezolizumab (TECENTRIQ®), bevacizumab (AVASTIN®), ceritinib (LDK378/ZYKADIA®), crizotinib (XALKORI®), dabrafenib (TAFINAR®), dacomitinib (VIZIMPRO®), erlotinib (TARCEVA®), gefitinib (IRESSA®), osimertinib (TAGRISSO®), pembrolizumab (KEYTRUDA®), pemetrexed (ALIMTA®), ramucirumab (CYRAMZA®), trametinib (MEKANIST®), additional therapies include brigatinib (ALUNBRIGTM), durvalumab (IMFINZITM), lorlatinib (LORBRENA®),
  • Non-limiting examples for lymphoma include acalabrutinib (CALQUENCE®), pembrolizumab (KEYTRUDA®), venetoclax (VENCLEXTA®) in Table 1, additional therapies include axicabtagene ciloleucel (YESCARTATM), belinostat (BELEODAQ®), bexarotene (TARGRETIN®), bortezomib (VELCADE®), brentuximab vedotin (ADCETRIS®), copanlisib (ALIQOPATM), denileukin diftitox (ONTAK®), duvelisib (COPIKTRATM), Ibritumomab tiuxetan (ZEVALIN®), ibrutinib (IMBRUVICA®), idelalisib (ZYDELIG®), mogamulizumab-kpkc (POTELIGEO®), nivolumab (
  • Non-limiting examples for melanoma include alitretinoin (PANRETIN®), binimetinib (MEKTOVI®), cobimetinib (COTELLIC®), dabrafenib (TAFINAR®), encorafenib (BRAFTOVITM), pembrolizumab (KEYTRUDA®), trametinib (MEKANIST®), or vemurafenib (ZELBORAF®) in Table 1, additional therapies include avelumab (BAVENCIO®), cemiplimab- rwlc (LIBTAYO®), ipilimumab (YERVOY®), nivolumab (OPDIVO®), sonidegib (ODOMZO®), or vismodegib (ERIVEDGE®).
  • PANRETIN® alitretinoin
  • MEKTOVI® binimetinib
  • COTELLIC® dabrafenib
  • MM multiple myeloma
  • MM multiple myeloma
  • VELCADE® Bortezomib
  • KYPROLIS® carfilzomib
  • DARZALEXTM daratumumab
  • EMPLICITITM elotuzumab
  • ixazomib NINLARO®
  • panobinostat FARYDAK®
  • selinexor XPOVIOTM
  • Non-limiting examples for prostate cancer include abiraterone acetate (ZYTIGA®) in Table 1, additional therapies include apalutamide (ERLEADATM), Cabazitaxel (JEVTANA®), darolutamide (NUBEQA®), enzalutamide (XTANDI®), radium 223 dichloride (XOFIGO®).
  • ERLEADATM apalutamide
  • JEVTANA® Cabazitaxel
  • NUBEQA® darolutamide
  • XTANDI® enzalutamide
  • XOFIGO® radium 223 dichloride
  • Additional drugs that may be used for cancer treatment include Denosumab (XGEVA®), Dinutuximab (UNITUXINTM), iobenguane I 131 (AZEDRA®), Lanreotide acetate (SOMATULINE® Depot), lutetium Lu 177-dotatate (LUTATHERA®), niraparib (ZEJULATM), rucaparib camsylate (RUBRACATM), ruxolitinib phosphate (JAKAFI®), Sirolimus (RAPAMUNE®), or Talazoparib (TALZENNA®).
  • XGEVA® Denosumab
  • Dinutuximab UNITUXINTM
  • iobenguane I 131 AZEDRA®
  • AZEDRA® Lanreotide acetate
  • LTATHERA® Lanreotide acetate
  • LUTATHERA® lutetium Lu 177-dotatate
  • “about 40 [units]” may mean within ⁇ 25% of 40 (e.g., from 30 to 50), within ⁇ 20%, ⁇ 15%, ⁇ 10%, ⁇ 9%, ⁇ 8%, ⁇ 7%, ⁇ 6%, ⁇ 5%, ⁇ 4%, ⁇ 3%, ⁇ 2%, ⁇ 1%, less than ⁇ 1%, or any other value or range of values therein or there below.
  • the term “about” may mean ⁇ one half a standard deviation, ⁇ one standard deviation, or ⁇ two standard deviations.
  • the phrases “less than about [a value]” or “greater than about [a value]” should be understood in view of the definition of the term “about” provided herein.
  • clinical signs of cancer means and includes any sign or indication of the existence of cancer in a subject, which sign or indication would be well known to the skilled artisan (e.g., oncologist, nurse practitioner).
  • the clinical signs of cancer may be any symptom known to be associated with the cancer.
  • Clinical signs of some cancers include, for example, chronic pain, nausea, vomiting, abnormal taste sensation, constipation, urinary symptoms (e.g., bladder spasm), respiratory symptoms, skin problems (e.g., pruritus, hair loss), or fever, among others.
  • the term “reference set” may be an internal, external, or a universal reference set of nucleic acids or expression products used to calibrate a particular sample.
  • an internal reference set of nucleic acids may be obtained using normal tissue or a blood sample from the subject.
  • an internal reference set may based on the total RNA in the sample.
  • the reference set may be a set of one or more housekeeping genes, e.g., human acidic ribosomal protein (HuPO), ⁇ -actin (BA), cyclophylin (CYC), glyceraldehyde-3- phosphate dehydrogenase (GAPDH), phosphoglycerokinase (PGK), ⁇ 2-microglobulin (B2M), ⁇ - glucuronidase (GUS), hypoxanthine phosphoribosyltransferase (HPRT), transcription factor IID TATA binding protein (TBP), transferrin receptor (TfR), human acidic ribosomal protein (HuPO), elongation factor-1- ⁇ (EF-1- ⁇ ), metastatic lymph node 51(MLN51), or ubiquitin conjugating enzyme (UbcH5B).
  • HuPO human acidic ribosomal protein
  • BA ⁇ -actin
  • CYC cyclophylin
  • GPDH glycer
  • remission means and includes a period during which the symptoms of a cancer have been reduced or eliminated, as remission is ordinarily defined in the oncology art.
  • “serially monitoring" levels of a biomarker in a sample refers to measuring levels of a biomarker in a sample more than once, e.g., quarterly, bimonthly, monthly, biweekly, weekly, every three days, daily, or several times per day. Serial monitoring of a level includes periodically measuring levels of biomarkers at regular intervals as deemed necessary by the skilled artisan.
  • standard level refers to a baseline level of a biomarker as determined in one or more normal subjects.
  • the measurement of biomarker levels may be carried out using the multiplexed copy number as described.
  • "elevation" of a measured level of a biomarker relative to a standard level means that the amount or concentration of a biomarker in a sample is sufficiently greater in a subject relative to the standard to be detected by the methods described herein.
  • elevation of the measured level relative to a standard level may be any statistically significant elevation which is detectable.
  • Such an elevation may include, but is not limited to, about a 1%, about a 10%, about a 20%, about a 40%, about an 80%, about a 2-fold, about a 4-fold, about an 8- fold, about a 20-fold, or about a 100-fold elevation, or more, relative to the standard.
  • Non-limiting examples of signaling pathway modulators or chemotherapeutic agents known in the art are 5-fluorouracil; asparaginase; bevacizumab (AVASTIN®); bleomycin; campathecins; cetuximab (ERBITUX®); crizotinib (XALKORI®); cyclophosphamide; cytarabine; dacarbazine; dactinomycin; dasatinib (SPRYCEL®); daunorubicin; DNA methyltransferase inhibitors (DNMTs) such as azacitidine (VIDAZA®) and decitabine; doxorubicin; doxorubicin; epirubicin; erbstatin; erlotinib (TARCEVA®); estramustine; etoposide; etoposide; gefitinib (IRESSA®), gemcitabine, genistein, histone acetyl transferase inhibitor
  • the chemotherapeutic agent is bevacizumab (AVASTIN®), cetuximab (ERBITUX®), crizotinib (XALKORI®), dasatinib (SPRYCEL®), erlotinib (TARCEVA®), everolimus (AFINITOR®), gefitinib (IRESSA®), imatinib (GLEEVEC®), lapatinib (TYKERB®), nilotinib (TASIGNA®), panitumumab (VECTIBIX®), pazopanib (VOTRIENT®), sirolimus (RAPAMUNE®), sorafenib (NEXAVAR®), sunitinib (SUTENT®), temsirolimus (TORISEL®), trastuzumab (HERCEPTIN®), vandetanib (CAPRELSA®), or vemurafenib (ZELBORAF®).
  • AVASTIN® cetuximab
  • a computing device may be implemented in programmable hardware devices such as processors, digital signal processors, central processing units, field programmable gate arrays, programmable array logic, programmable logic devices, cloud processing systems, or the like.
  • the computing devices may also be implemented in software for execution by various types of processors.
  • An identified device may include executable code and may, for instance, comprise one or more physical or logical blocks of computer instructions, which may, for instance, be organized as an object, procedure, function, or other construct. Nevertheless, the executable of an identified device need not be physically located together but may comprise disparate instructions stored in different locations which, when joined logically together, comprise the computing device and achieve the stated purpose of the computing device.
  • a computing device may be a server or other computer located within a hospital or out-patient environment and communicatively connected to other computing devices (e.g., POS equipment or computers) for managing accounting, purchase transactions, and other processes within the hospital or out-patient environment.
  • a computing device may be a mobile computing device such as, for example, but not limited to, a smart phone, a cell phone, a pager, a personal digital assistant (PDA), a mobile computer with a smart phone client, or the like.
  • a computing device may be any type of wearable computer, such as a computer with a head-mounted display (HMD), or a smart watch or some other wearable smart device. Some of the computer sensing may be part of the fabric of the clothes the user is wearing.
  • a computing device can also include any type of conventional computer, for example, a laptop computer or a tablet computer.
  • a typical mobile computing device is a wireless data access-enabled device (e.g., an iPHONE ® smart phone, a BLACKBERRY ® smart phone, a NEXUS ONETM smart phone, an iPAD ® device, smart watch, or the like) that is capable of sending and receiving data in a wireless manner using protocols like the Internet Protocol, or IP, and the wireless application protocol, or WAP.
  • a wireless data access-enabled device e.g., an iPHONE ® smart phone, a BLACKBERRY ® smart phone, a NEXUS ONETM smart phone, an iPAD ® device, smart watch, or the like
  • IP Internet Protocol
  • WAP wireless application protocol
  • Wireless data access is supported by many wireless networks, including, but not limited to, Bluetooth, Near Field Communication, CDPD, CDMA, GSM, PDC, PHS, TDMA, FLEX, ReFLEX, iDEN, TETRA, DECT, DataTAC, Mobitex, EDGE and other 2G, 3G, 4G, 5G, and LTE technologies, and it operates with many handheld device operating systems, such as PalmOS, EPOC, Windows CE, FLEXOS, OS/9, JavaOS, iOS and Android.
  • these devices use graphical displays and can access the Internet (or other communications network) on so-called mini- or micro-browsers, which are web browsers with small file sizes that can accommodate the reduced memory constraints of wireless networks.
  • the mobile device is a cellular telephone or smart phone or smart watch that operates over GPRS (General Packet Radio Services), which is a data technology for GSM networks or operates over Near Field Communication e.g. Bluetooth.
  • GPRS General Packet Radio Services
  • a given mobile device can communicate with another such device via many different types of message transfer techniques, including Bluetooth, Near Field Communication, SMS (short message service), enhanced SMS (EMS), multi-media message (MMS), email WAP, paging, or other known or later-developed wireless data formats.
  • SMS short message service
  • EMS enhanced SMS
  • MMS multi-media message
  • email WAP paging
  • paging or other known or later-developed wireless data formats.
  • An executable code of a computing device may be a single instruction, or many instructions, and may even be distributed over several different code segments, among different applications, and across several memory devices.
  • operational data may be identified and illustrated herein within the computing device, and may be embodied in any suitable form and organized within any suitable type of data structure. The operational data may be collected as a single data set, or may be distributed over different locations including over different storage devices, and may exist, at least partially, as electronic signals on a system or network.
  • the term “memory” is generally a storage device of a computing device. Examples include, but are not limited to, read-only memory (ROM) and random access memory (RAM). [0060]
  • the device or system for performing one or more operations on a memory of a computing device may be a software, hardware, firmware, or combination of these.
  • the device or the system is further intended to include or otherwise cover all software or computer programs capable of performing the various heretofore-disclosed determinations, calculations, or the like for the disclosed purposes.
  • exemplary embodiments are intended to cover all software or computer programs capable of enabling processors to implement the disclosed processes.
  • Exemplary embodiments are also intended to cover any and all currently known, related art or later developed non-transitory recording or storage mediums (such as a CD-ROM, DVD-ROM, hard drive, RAM, ROM, floppy disc, magnetic tape cassette, etc.) that record or store such software or computer programs.
  • Exemplary embodiments are further intended to cover such software, computer programs, systems and/or processes provided through any other currently known, related art, or later developed medium (such as transitory mediums, carrier waves, etc.), usable for implementing the exemplary operations disclosed below.
  • the disclosed computer programs can be executed in many exemplary ways, such as an application that is resident in the memory of a device or as a hosted application that is being executed on a server and communicating with the device application or browser via a number of standard protocols, such as TCP/IP, HTTP, XML, SOAP, REST, JSON and other sufficient protocols.
  • the disclosed computer programs can be written in exemplary programming languages that execute from memory on the device or from a hosted server, such as BASIC, COBOL, C, C++, Java, Pascal, or scripting languages such as JavaScript, Python, Ruby, PHP, Perl, or other suitable programming languages.
  • the terms “computing device” and “entities” should be broadly construed and should be understood to be interchangeable. They may include any type of computing device, for example, a server, a desktop computer, a laptop computer, a smart phone, a cell phone, a pager, a personal digital assistant (PDA, e.g., with GPRS NIC), a mobile computer with a smartphone client, or the like.
  • PDA personal digital assistant
  • a user interface is generally a system by which users interact with a computing device.
  • a user interface can include an input for allowing users to manipulate a computing device, and can include an output for allowing the system to present information and/or data, indicate the effects of the user’s manipulation, etc.
  • An example of a user interface on a computing device includes a graphical user interface (GUI) that allows users to interact with programs in more ways than typing.
  • GUI graphical user interface
  • a GUI typically can offer display objects, and visual indicators, as opposed to text-based interfaces, typed command labels or text navigation to represent information and actions available to a user.
  • an interface can be a display window or display object, which is selectable by a user of a mobile device for interaction.
  • a user interface can include an input for allowing users to manipulate a computing device, and can include an output for allowing the computing device to present information and/or data, indicate the effects of the user’s manipulation, etc.
  • An example of a user interface on a computing device includes a graphical user interface (GUI) that allows users to interact with programs or applications in more ways than typing.
  • GUI graphical user interface
  • a GUI typically can offer display objects, and visual indicators, as opposed to text-based interfaces, typed command labels or text navigation to represent information and actions available to a user.
  • a user interface can be a display window or display object, which is selectable by a user of a computing device for interaction.
  • the display object can be displayed on a display screen of a computing device and can be selected by and interacted with by a user using the user interface.
  • the display of the computing device can be a touch screen, which can display the display icon. The user can depress the area of the display screen where the display icon is displayed for selecting the display icon.
  • the user can use any other suitable user interface of a computing device, such as a keypad, to select the display icon or display object.
  • the user can use a track ball or arrow keys for moving a cursor to highlight and select the display object.
  • the display object can be displayed on a display screen of a mobile device and can be selected by and interacted with by a user using the interface.
  • the display of the mobile device can be a touch screen, which can display the display icon.
  • the user can depress the area of the display screen at which the display icon is displayed for selecting the display icon.
  • the user can use any other suitable interface of a mobile device, such as a keypad, to select the display icon or display object.
  • the user can use a track ball or times program instructions thereon for causing a processor to carry out aspects of the present disclosure.
  • a computer network may be any group of computing systems, devices, or equipment that are linked together. Examples include, but are not limited to, local area networks (LANs) and wide area networks (WANs).
  • a network may be categorized based on its design model, topology, or architecture.
  • a network may be characterized as having a hierarchical internetworking model, which divides the network into three layers: access layer, distribution layer, and core layer.
  • the access layer focuses on connecting client nodes, such as workstations to the network.
  • the distribution layer manages routing, filtering, and quality-of-server (QoS) policies.
  • QoS quality-of-server
  • the core layer can provide high-speed, highly-redundant forwarding services to move packets between distribution layer devices in different regions of the network.
  • the core layer typically includes multiple routers and switches.
  • the computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present subject matter.
  • the computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device.
  • the computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.
  • a non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing.
  • a computer readable storage medium is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
  • Computer readable program instructions described herein can be downloaded to respective computing/processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and/or a wireless network, or Near Field Communication.
  • the network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and/or edge servers.
  • Computer readable program instructions for carrying out operations of the present subject matter may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state- setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++, Javascript or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages.
  • ISA instruction-set-architecture
  • machine instructions machine dependent instructions
  • microcode firmware instructions
  • state- setting data or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++, Javascript or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages.
  • the computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.
  • the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).
  • LAN local area network
  • WAN wide area network
  • Internet Service Provider for example, AT&T, MCI, Sprint, EarthLink, MSN, GTE, etc.
  • electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present subject matter.
  • FPGA field-programmable gate arrays
  • PLA programmable logic arrays
  • These computer readable program instructions may be provided to a processor of a computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions/acts specified in the flowchart and/or block diagram block or blocks.
  • These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and/or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function/act specified in the flowchart and/or block diagram block or blocks.
  • the computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions/acts specified in the flowchart and/or block diagram block or blocks.
  • the description illustrates the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present subject matter.
  • each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s).
  • each block of the block diagrams and/or flowchart illustration, and combinations of blocks in the block diagrams and/or flowchart illustration can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
  • the sample may be from a subject suspected of having a particular cancer or from a patient diagnosed with cancer, e.g., for confirmation of diagnosis or establishing a clear margin or for the detection of cancer cells in other tissues such as lymph nodes, or circulating tumor cells.
  • the biological sample may also be from a subject with an ambiguous diagnosis in order to clarify the diagnosis.
  • the sample may be obtained for the purpose of differential diagnosis, e.g., a subject with a histopathologically benign lesion to confirm the diagnosis.
  • the sample may also be obtained for the purpose of prognosis, i.e., determining the course of the disease and selecting primary treatment options. Tumor staging and grading are examples of prognosis.
  • Samples may be obtained using any of a number of methods in the art. Examples of biological samples comprising potential cancer cells include those obtained from excised skin biopsies, such as punch biopsies, shave biopsies, core needle biopsies, fine needle aspirates (FNA), or surgical excisions; or biopsy from non- cutaneous tissues such as lymph node tissue, mucosa, other embodiments.
  • excised skin biopsies such as punch biopsies, shave biopsies, core needle biopsies, fine needle aspirates (FNA), or surgical excisions
  • FNA fine needle aspirates
  • the sample may be from a distant metastatic site, a soft tissue, e.g., lung, liver, bone, skin, or brain.
  • Representative biopsy techniques include, but are not limited to, excisional biopsy, incisional biopsy, pinch biopsy, forceps biopsy, needle biopsy, or surgical biopsy.
  • An "excisional biopsy” refers to the removal of an entire tumor mass with a small margin of normal tissue surrounding it.
  • An “incisional biopsy” refers to the removal of a wedge of tissue that includes a cross-sectional diameter of the tumor.
  • a diagnosis or prognosis made by endoscopy or fluoroscopy may require a "core-needle biopsy" of the tumor mass, or a "fine-needle aspiration biopsy” which generally contains a suspension of cells from within the tumor mass.
  • the biological sample may be a microdissected sample, such as a PALM-laser (Carl Zeiss MicroImaging GmbH, Germany) capture microdissected sample.
  • a sample may also be a sample of muscosal surfaces, blood and blood fractions or products (e.g., serum, plasma, platelets, red blood cells, white blood cells, circulating tumor cells isolated from blood, free DNA isolated from blood, and the like), sputum, saliva, lymph and tongue tissue, cultured cells, e.g., primary cultures, explants, and transformed cells, stool, urine, etc.
  • the sample may also be vascular tissue or cells from blood vessels such as microdissected blood vessel cells of endothelial origin.
  • a sample is typically obtained from a eukaryotic organism, most preferably a mammal such as a primate e.g., chimpanzee or human, cow, dog, cat; or a rodent, e.g., guinea pig, rat, mouse, rabbit.
  • a sample can be treated with a fixative such as formaldehyde and embedded in paraffin (FFPE) and sectioned for use in the methods of the invention.
  • FFPE formaldehyde and embedded in paraffin
  • fresh or frozen tissue may be used.
  • These cells may be fixed, e.g., in alcoholic solutions such as 100% ethanol or 3:1 methanol:acetic acid.
  • Nuclei can also be extracted from thick sections of paraffin-embedded specimens to reduce truncation artifacts and eliminate extraneous embedded material.
  • biological samples once obtained, are harvested and processed prior to nucleic acid analysis using standard methods known in the art. Such processing typically includes protease treatment and additional fixation in an aldehyde solution such as formaldehyde. 5.3.1. Polynucleotide Sequence Amplification and Determination [0078] In many instances, it is desirable to amplify a nucleic acid sequence using any of several nucleic acid amplification procedures which are well known in the art.
  • nucleic acid amplification is the chemical or enzymatic synthesis of nucleic acid copies which contain a sequence that is complementary to a nucleic acid sequence being amplified (template).
  • the methods and kits of the invention may use any nucleic acid amplification or detection methods known to one skilled in the art, such as those described in U.S. Pat. Nos. 5,525,462 (Takarada et al.); 6,114,117 (Hepp et al.); 6,127,120 (Graham et al.); 6,344,317 (Urnovitz); 6,448,001 (Oku); 6,528,632 (Catanzariti et al.); and PCT Pub. No.
  • the nucleic acids may be amplified by PCR amplification using methodologies known to one skilled in the art.
  • amplification can be accomplished by other known methods, such as ligase chain reaction (LCR), Q ⁇ -replicase amplification, rolling circle amplification, transcription amplification, self-sustained sequence replication, nucleic acid sequence-based amplification (NASBA), each of which provides sufficient amplification.
  • Branched-DNA technology may also be used to qualitatively demonstrate the presence of a sequence of the technology which may quantitatively determine the amount of this particular genomic sequence in a sample.
  • Nolte reviews branched-DNA signal amplification for direct quantitation of nucleic acid sequences in clinical samples (Nolte, 1998, Adv. Clin. Chem. 33:201-235).
  • the PCR process is well known in the art and is thus not described in detail herein. For a review of PCR methods and protocols, see, e.g., Innis et al., eds., PCR Protocols, A Guide to Methods and Application, Academic Press, Inc., San Diego, Calif. 1990; U.S. Pat. No.
  • PCR reagents and protocols are also available from commercial vendors, such as Roche Molecular Systems. PCR may be carried out as an automated process with a thermostable enzyme. In this process, the temperature of the reaction mixture is cycled through a denaturing region, a primer annealing region, and an extension reaction region automatically. Machines specifically adapted for this purpose are commercially available. 5.3.2. High Throughput and Single Molecule Sequencing Technology [0081] Suitable next generation sequencing technologies are widely available.
  • Examples include the 454 Life Sciences platform (Roche, Branford, CT) (Margulies et al.2005 Nature, 437, 376-380); lllumina’s Genome Analyzer, Illumina’s MiSeq System, Illumina’s NextSeq System, Illumina’s MiniSeq System, (Illumina, San Diego, CA; Bibkova et al., 2006, Genome Res.16, 383- 393; U.S. Pat. Nos. 6,306,597 and 7,598,035 (Macevicz); 7,232,656 (Balasubramanian et al.)); or DNA Sequencing by Ligation, SOLiD System (Applied Biosystems/Life Technologies; U.S. Pat.
  • Chem.53, 1996-2001 which are incorporated herein by reference in their entirety.
  • These systems allow the sequencing of many nucleic acid molecules isolated from a specimen at high orders of multiplexing in a parallel fashion (Dear, 2003, Brief Funct. Genomic Proteomic, 1(4), 397-416 and McCaughan and Dear, 2010, J. Pathol., 220, 297- 306).
  • Each of these platforms allow sequencing of clonally expanded or non-amplified single molecules of nucleic acid fragments.
  • Certain platforms involve, for example, (i) sequencing by ligation of dye-modified probes (including cyclic ligation and cleavage), (ii) pyrosequencing, (iii) targeted next-generation sequencing from bisulfite treated DNA and (iv) single-molecule sequencing.
  • Pyrosequencing is a nucleic acid sequencing method based on sequencing by synthesis, which relies on detection of a pyrophosphate released on nucleotide incorporation.
  • sequencing by synthesis involves synthesizing, one nucleotide at a time, a DNA strand complimentary to the strand whose sequence is being sought.
  • Study nucleic acids may be immobilized to a solid support, hybridized with a sequencing primer, incubated with DNA polymerase, ATP sulfurylase, luciferase, apyrase, adenosine 5' phosphsulfate and luciferin. Nucleotide solutions are sequentially added and removed. Correct incorporation of a nucleotide releases a pyrophosphate, which interacts with ATP sulfurylase and produces ATP in the presence of adenosine 5' phosphosulfate, fueling the luciferin reaction, which produces a chemiluminescent signal allowing sequence determination.
  • Machines for pyrosequencing are available from Qiagen, Inc. (Valencia, CA).
  • An example of a system that can be used by a person of ordinary skill based on pyrosequencing generally involves the following steps: ligating an adaptor nucleic acid to a study nucleic acid and hybridizing the study nucleic acid to a bead; amplifying a nucleotide sequence in the study nucleic acid in an emulsion; sorting beads using a picoliter multiwell solid support; and sequencing amplified nucleotide sequences by pyrosequencing methodology (e.g., Nakano et al., 2003, J. Biotech.102, 117-124).
  • NGS Next-generation sequencing
  • dNTPs deoxyribonucleotide triphosphates
  • sequencing by synthesis involves synthesizing, one nucleotide at a time, a DNA strand complimentary to the strand whose sequence is being sought.
  • Study nucleic acids may be immobilized to a solid support, hybridized with a sequencing primer, and incubated with DNA polymerase in the presence of fluorescently labeled dNTPS. After each cycle, the image is scanned and the emission wavelength and intensity are recorded and used to identify the base incorporated. This process is repeated multiple times to create a specific read length of bases.
  • Certain single-molecule sequencing embodiments are based on the principal of sequencing by synthesis, and utilize single-pair Fluorescence Resonance Energy Transfer (single pair FRET) as a mechanism by which photons are emitted as a result of successful nucleotide incorporation.
  • the emitted photons often are detected using intensified or high sensitivity cooled charge-couple-devices in conjunction with total internal reflection microscopy (TIRM). Photons are only emitted when the introduced reaction solution contains the correct nucleotide for incorporation into the growing nucleic acid chain that is synthesized as a result of the sequencing process.
  • TIRM total internal reflection microscopy
  • FRET FRET based single-molecule sequencing or detection
  • energy is transferred between two fluorescent dyes, sometimes polymethine cyanine dyes Cy3 and Cy5, through long-range dipole interactions.
  • the donor is excited at its specific excitation wavelength and the excited state energy is transferred, non-radiatively to the acceptor dye, which in turn becomes excited.
  • the acceptor dye eventually returns to the ground state by radiative emission of a photon.
  • the two dyes used in the energy transfer process represent the "single pair", in single pair FRET. Cy3 often is used as the donor fluorophore and often is incorporated as the first labeled nucleotide.
  • Cy5 often is used as the acceptor fluorophore and is used as the nucleotide label for successive nucleotide additions after incorporation of a first Cy3 labeled nucleotide.
  • the fluorophores generally are within 10 nanometers of each other for energy transfer to occur successfully.
  • An example of a system that can be used based on single-molecule sequencing generally involves hybridizing a primer to a study nucleic acid to generate a complex; associating the complex with a solid phase; iteratively extending the primer by a nucleotide tagged with a fluorescent molecule; and capturing an image of fluorescence resonance energy transfer signals after each iteration (e.g., Braslavsky et al., PNAS 100(7): 3960-3964 (2003); U.S. Pat. No.7,297,518 (Quake et al.) which are incorporated herein by reference in their entirety).
  • Such a system can be used to directly sequence amplification products generated by processes described herein.
  • the released linear amplification product can be hybridized to a primer that contains sequences complementary to immobilized capture sequences present on a solid support, a bead or glass slide for example. Hybridization of the primer-released linear amplification product complexes with the immobilized capture sequences, immobilizes released linear amplification products to solid supports for single pair FRET based sequencing by synthesis.
  • the primer often is fluorescent, so that an initial reference image of the surface of the slide with immobilized nucleic acids can be generated. The initial reference image is useful for determining locations at which true nucleotide incorporation is occurring. Fluorescence signals detected in array locations not initially identified in the "primer only" reference image are discarded as non-specific fluorescence.
  • the bound nucleic acids often are sequenced in parallel by the iterative steps of, a) polymerase extension in the presence of one fluorescently labeled nucleotide, b) detection of fluorescence using appropriate microscopy, TIRM for example, c) removal of fluorescent nucleotide, and d) return to step a with a different fluorescently labeled nucleotide.
  • TIRM microscopy
  • c) removal of fluorescent nucleotide and d) return to step a with a different fluorescently labeled nucleotide.
  • nucleotide sequencing may be by solid phase single nucleotide sequencing methods and processes.
  • Solid phase single nucleotide sequencing methods involve contacting sample nucleic acid and solid support under conditions in which a single molecule of sample nucleic acid hybridizes to a single molecule of a solid support. Such conditions can include providing the solid support molecules and a single molecule of sample nucleic acid in a "microreactor.” Such conditions also can include providing a mixture in which the sample nucleic acid molecule can hybridize to solid phase nucleic acid on the solid support.
  • Single nucleotide sequencing methods useful in the embodiments described herein are described in PCT Pub. No. WO 2009/091934 (Cantor).
  • nanopore sequencing detection methods include (a) contacting a nucleic acid for sequencing ("base nucleic acid,” e.g., linked probe molecule) with sequence- specific detectors, under conditions in which the detectors specifically hybridize to substantially complementary subsequences of the base nucleic acid; (b) detecting signals from the detectors and (c) determining the sequence of the base nucleic acid according to the signals detected.
  • the detectors hybridized to the base nucleic acid are disassociated from the base nucleic acid (e.g., sequentially dissociated) when the detectors interfere with a nanopore structure as the base nucleic acid passes through a pore, and the detectors disassociated from the base sequence are detected.
  • a detector also may include one or more regions of nucleotides that do not hybridize to the base nucleic acid.
  • a detector is a molecular beacon.
  • a detector often comprises one or more detectable labels independently selected from those described herein. Each detectable label can be detected by any convenient detection process capable of detecting a signal generated by each label (e.g., magnetic, electric, chemical, optical and the like). For example, a CD camera can be used to detect signals from one or more distinguishable quantum dots linked to a detector.
  • the invention encompasses methods known in the art for enhancing the sensitivity of the detectable signal in such assays, including, but not limited to, the use of cyclic probe technology (Bakkaoui et al., 1996, BioTechniques 20: 240-8, which is incorporated herein by reference in its entirety); and the use of branched probes (Urdea et al., 1993, Clin. Chem. 39, 725-6; which is incorporated herein by reference in its entirety).
  • the hybridization complexes are detected according to well-known techniques in the art.
  • Reverse transcribed or amplified nucleic acids may be modified nucleic acids.
  • Modified nucleic acids can include nucleotide analogs, and in certain embodiments include a detectable label and/or a capture agent.
  • detectable labels include, without limitation, fluorophores, radioisotopes, colorimetric agents, light emitting agents, chemiluminescent agents, light scattering agents, enzymes and the like.
  • capture agents include, without limitation, an agent from a binding pair selected from antibody/antigen, antibody/antibody, antibody/antibody fragment, antibody/antibody receptor, antibody/protein A or protein G, hapten/anti-hapten, biotin/avidin, biotin/streptavidin, folic acid/folate binding protein, vitamin B12/intrinsic factor, chemical reactive group/complementary chemical reactive group (e.g., sulfhydryl/maleimide, sulfhydryl/haloacetyl derivative, amine/isotriocyanate, amine/succinimidyl ester, and amine/sulfonyl halides) pairs, and the like.
  • an agent from a binding pair selected from antibody/antigen, antibody/antibody, antibody/antibody fragment, antibody/antibody receptor, antibody/protein A or protein G, hapten/anti-hapten, biotin/avidin, biotin/streptavidin, folic acid/
  • Modified nucleic acids having a capture agent can be immobilized to a solid support in certain embodiments.
  • Next generation sequencing techniques may be applied to measure expression levels or count numbers of transcripts using RNA-seq or whole transcriptome shotgun sequencing. See, e.g., Mortazavi et al. 2008 Nat Meth 5(7) 621-627 or Wang et al. 2009 Nat Rev Genet 10(1) 57-63. Nucleic acids in the invention may be counted using methods known in the art. In one embodiment, NanoString’s nCounter® system may be used (Seattle, WA). Geiss et al. 2008 Nat Biotech 26(3) 317-325; U.S. Pat. No. 7,473,767 (Dimitrov).
  • NanoString Digital Spatial Profiling (DSP) platform may be used for nucleic acid or protein detection. Blank et al., 2018 Nature Medicine 24 1655–1661; Amaria et al., 2018 Nature Medicine 24 1649–1654.
  • Fluidigm Dynamic Array system may be used (South San Francisco, CA). Byrne et al. 2009 PLoS ONE 4 e7118; Helzer et al. 2009 Can Res 697860-7866.
  • compositions and Kits [0093] The invention provides compositions and kits detecting the biomarkers described herein using antibodies or other reagents specific for the nucleic acids specific for the polynucleotides.
  • Kits for carrying out the diagnostic assays of the invention typically include, in suitable container means, (i) a probe that comprises an antibody or nucleic acid sequence that specifically binds to the marker polynucleotides of the invention, (ii) a label for detecting the presence of the probe and (iii) instructions for how to measure the level the polynucleotide.
  • kits may include several antibodies or polynucleotide sequences encoding biomarkers disclosed herein, e.g., a first antibody and/or second and/or third and/or additional antibodies that recognize the biomarkers or specific nucleic acids.
  • the nucleic acids in the kit are the forward and reverse PCR primers for the biomarkers disclosed herein.
  • the container means of the kits will generally include at least one vial, test tube, flask, bottle, syringe and/or other container into which a first antibody specific for one of the polypeptides or a first nucleic acid specific for one of the polynucleotides of the present invention may be placed and/or suitably aliquoted.
  • kits will also generally contain a second, third and/or other additional container into which this component may be placed.
  • a container may contain a mixture of more than one antibody or nucleic acid reagent, each reagent specifically binding a different marker in accordance with the present invention.
  • the kits of the present invention will also typically include means for containing the antibody or nucleic acid probes in close confinement for commercial sale. Such containers may include injection and/or blow-molded plastic containers into which the desired vials are retained.
  • the kits may further comprise positive and negative controls, as well as instructions for the use of kit components contained therein, in accordance with the methods of the present invention.
  • the friendship paradox holds that in a social network, most people have fewer friends than their friends do (14). In other words, for most nodes in a network, their neighbors or “friends” will on average have a higher degree (number of connections) than the node itself. This arises because higher degree nodes count towards the degree in multiple neighboring nodes and thus are oversampled (see Methods: proof of friendship paradox).
  • highly mutated samples contribute toward the average TMB for many of the genes in the dataset, resulting in an outsized effect (Supplemental Figure 1B). Since the univariate t-test and Mann-Whitney U tests are testing for differences in central tendencies (i.e.
  • the TMB Paradox explains why the majority of genes have a highly significant association with elevated TMBs.
  • the TMB paradox is therefore a manifestation of the oversampling bias introduced by highly connected nodes (i.e. High TMB tumors) and this oversampling bias is what makes univariate tests (T-test and MWU) inappropriate to identify which genes are associated with higher levels of TMB. Therefore, the majority of the 98% of genes associated with elevated TMB by the univariate approach (Figure 1A) are likely a result of the TMB paradox rather than the underlying biology.
  • BiG-BETS Bipartite Graph-Based Expected TMB Score
  • TCGA Pan-Cancer dataset found that BiG-BETS returned a more uniform distribution of p-values (Supplemental Figure 2A), and only a subset of DDR genes when mutated have a significant association with elevated TMB, (Figure 2A) hereafter referred to as “High BiG-BETS DDR gene” ( Figure 2B). It was reassuring to see that MMR genes such as MSH3, MLH1, MSH2 had some of the highest BiG-BET scores (Figure 2B).
  • the TCGA MC3 dataset was then filtered to include only “High Consequence” non-synonymous mutations, which were defined as being categorized as a high consequence mutation by the Sequence Ontology and summarized by Ensembl here (https://m.ensembl.org/info/genome/variation/prediction/predicted_data.html) ['stop lost', 'stop gained', 'transcript ablation', 'start lost', 'frameshift variant’, ’splice_site’, ‘translation_start_site’] and were also categorized as having a Polyphen score of ‘probably damaging’, ‘possibly damaging’, or ‘unknown’.
  • the bipartite network representation was used to derive a null model of TMB distribution for each gene. Specifically, random sampling (permutations) of the bipartite network was performed by stochastically rewiring the network while maintaining the degree distribution (number of edges of each node) of the original dataset (this null model is known as the configuration model, which for a bipartite network is further constrained to maintain the bipartite nature of the network) (15,16).
  • This null model is known as the configuration model, which for a bipartite network is further constrained to maintain the bipartite nature of the network
  • the BiG-BET score consists of comparing the observed mean TMB for each gene or pathway's mutated sample set against the expected distribution under random sampling of bipartite networks that match the degree distribution of the original dataset.
  • the null model for networks in which all networks with a given degree sequence are uniformly likely is known as the configuration model (24), which has also been extended to bipartite graphs (15).
  • the bipartite configuration model can be envisioned by cutting across the edges in the original network and reconnecting the “stubs" at random with each possible set of pairings respecting the bipartite structure of the original network and being equally likely under the model (visualized in Figure 1F).
  • TMB is treated as a binary variable with a threshold of TMB>10 defining the high TMB group.
  • TMB is treated as a continuous variable. Comparison of response across groups is conducted using a Chi-squared test. Samples were divided into groups based on the presence of a mutation within a High DDR z-score gene or Low z-score DDR gene, shown in Figure 2B.
  • the High BIG-BET z-score DDR genes were: [00148] BARD1, CHEK1, CUL5, EME1, ERCC4, ERCC5, ERCC6, FANCA, FANCC, FANCI, FEN1, GEN1, LIG4, MDC1, MLH1, MLH3, MRE11, MSH2, MSH3, MSH6, NBN, PALB2, PARP1, PMS1, PMS2, POLE, POLE3, POLL, POLM, RAD50, RNMT, SEM1, SLX1A, TDG, TOP3A, TOPBP1, UNG, XPC, XRCC2, XRCC4, XRCC6.
  • RNAseq expression data was obtained and filtered to the corresponding samples with mutational data. We log(1+x) transformed the data and used a robust scaling (median centered and scaled by inter-quartile range) to normalize across samples. For each signature, we calculate the average expression of all genes within the signature and then assign each sample a z-score of the basis of its expression relative to the entire cohort. Signatures used in analysis are given in below.
  • GRANDVAUX_IRF3_TARGETS_UP ARG2, B4GALT5, F13B, GBP1, IFI44, IFIT1, IFIT3, ISG15, LILRB1, NR3C1, OAS2, PLCG2 PMAIP1, RSAD2 13_T_Cell_EntrezID ACTN1, ACVR2B, ADA, AKTIP, ANXA1, AOC1, APBA2, APBB1, AQP3, ARL4C, ATP13A4, ATP1A1, BCL11B, BIN2, BUB1B, C15orf62, C9orf164, CAMK4, CAPZB, CCL5, CCND2, CD2, CD247, CD28, CD3D, CD3E, CD3G, CD5, CD6, CDC14A, CEP41, CEP85L, CISH, CTSW, DISC1, DNAJB1, DNASE1L3, DOCK9, DPP4, DUSP16, DUSP2, FAM102A, FAM134B,
  • Supplemental Figure 3B demonstrates excellent correlation in the BiG-BET score between the high impact and the high+moderate impact datasets, especially with regards to the DDR genes and pathways.
  • TMB values for TCGA were obtained from (https://gdc.cancer.gov/about- data/publications/PanCan-CellOfOrigin) combining both Silent and Non-Silent scores for each sample.
  • Clinical data used for survival analysis of the TCGA were obtained from the TCGA clinical data resource as detailed in Liu et al 2018. We looked at the effect of low and high BiG- BET DDR mutations on overall 5-year survival as detailed above.
  • IMvigor210 [00158] The IMvigor210 trial is a Phase II single arm study examining the response of patients with locally advanced or metastatic urothelial bladder cancer to atezoliziumab (anti PD-L1). A full description of the characteristics of the patient cohort can be found in (7).
  • the cohort consists of 260 patients with 1249 short variants across 160 different genes. Because less detailed annotations were available, we did not filter any of the mutations from this cohort.
  • the class specific degree distributions as and to represent the fraction of nodes within each class with a given degree.
  • the overall degree distribution, , and the class specific degree distributions are related by [00172] .
  • [00173] In our gene-sample network, we are interested in the average degree across all samples with a mutation in a given gene. We show that this value, the average neighbor-of-a-gene degree, is typically greater than or equal to the average degree of the sample nodes in the network, following a proof similar to that for unipartite networks in 35 .
  • class 1 is our class of interest (the sample nodes). There are m edges connected to nodes of class 1, so the probability of ending at a particular node with degree is . Since there are such nodes with degree , the probability of following an edge to a class 1 node of degree is [00175] [00176] where gives the average degree for nodes of class 1. That is, the average neighbor degree distribution is weighted by a factor of . We are more likely to choose a higher degree vertex by virtue of the simple fact that it has more edges connected to it.
  • TGF ⁇ attenuates tumour response to PD-L1 blockade by contributing to exclusion of T cells. Nature. 2018;554:544–8.
  • Knijnenburg TA Wang L, Zimmermann MT, Chambwe N, Gao GF, Cherniack AD, et al. Genomic and Molecular Landscape of DNA Damage Repair Deficiency across The Cancer Genome Atlas. Cell reports.2018;23:239-254.e6. 11. Chae YK, Anker JF, Carneiro BA, Chandra S, Kaplan J, Kalyan A, et al. Genomic landscape of DNA repair genes in cancer. Oncotarget.2016;7:23312–21. 12. Chalmers ZR, Connelly CF, Fabrizio D, Gay L, Ali SM, Ennis R, et al. Analysis of 100,000 human cancer genomes reveals the landscape of tumor mutational burden. Genome medicine [Internet].
  • Fibroblast growth factor receptor 3 alterations and response to immune checkpoint inhibition in metastatic urothelial cancer a real world experience.
  • 33. Braun DA, Hou Y, Bakouny Z, Ficial M, Angelo MS, Forman J, et al. Interplay of somatic alterations and immune infiltration modulates response to PD-1 blockade in advanced clear cell renal cell carcinoma. Nat Med.2020;26:909–18.
  • a method for predicting response to immune checkpoint blockade (ICB) therapy for a subject with cancer comprising independently measuring or obtaining (i) a tumor mutational burden (TMB) level and (ii) defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5 in a sample from the subject, normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample, wherein subjects with both
  • a method for predicting overall survival for a subject with cancer comprising independently measuring or obtaining (i) a tumor mutational burden (TMB) level and (ii) defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5 in a sample from the subject, normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample, wherein subjects with both a high TMB level and defects in at least
  • Statement 3 A method to select a subject with cancer for immune blockade therapy (ICB) which comprises independently measuring or obtaining both (i) a tumor mutational burden (TMB) level and (ii) defects in DNA Damage Repair (DDR) genes or their expression products in a sample from the subject; calculating a Bipartite Graph-Based Expected TMB Score (BiG-BETS) to determine which DDR genes or expression products are associated with elevated TMB; if the subject has a high BiG-BETS determining that the defect in the DDR genes or their expression products do not add to the TMB level over measurement of TMB levels alone; and using the TMB levels alone to select a subject for ICB.
  • ICB immune blockade therapy
  • Statement 4 The method of Statement 3, wherein the DDR genes or their expression products are at least ten biomarkers selected from the group consisting of BARD1, CHEK1, CUL5, EME1, ERCC4, ERCC5, ERCC6, FANCA, FANCC, FANCI, FEN1, GEN1, LIG4, MDC1, MLH1, MLH3, MRE11, MSH2, MSH3, MSH6, NBN, PALB2, PARP1, PMS1, PMS2, POLE, POLE3, POLL, POLM, RAD50, RNMT, SEM1, SLX1A, TDG, TOP3A, TOPBP1, UNG, XPC, XRCC2, XRCC4, and XRCC6.
  • biomarkers selected from the group consisting of BARD1, CHEK1, CUL5, EME1, ERCC4, ERCC5, ERCC6, FANCA, FANCC, FANCI, FEN1, GEN1, LIG4, MDC1, MLH1, MLH3, MRE11, M
  • a method to select a subject with cancer for immune blockade therapy which comprises independently measuring or obtaining both (i) a tumor mutational burden (TMB) level and (ii) defects in DNA Damage Repair (DDR) genes or their expression products in a sample from the subject; calculating a Bipartite Graph-Based Expected TMB Score (BiG-BETS) to determine which DDR genes or expression products are associated with elevated TMB; if the subject has a low BiG-BETS determining that the defects in the DDR genes or their expression products contribute to the likelihood of benefit from ICB over measurement of TMB levels alone; and using both the TMB levels and the defects in the DDR genes or their expression products to select a subject for ICB.
  • ICB immune blockade therapy
  • Statement 6 The method of Statement 5, wherein the DDR genes or their expression products are at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5.
  • biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1,
  • Statement 7 The method of any of Statements 1-6, wherein the cancer is bladder urothelial carcinoma, colon adenocarcinoma, esophageal carcinoma, invasive breast carcinoma, head and neck squamous cell carcinoma, kidney renal clear cell carcinoma, lung adenocarcinoma, melanoma, or squamous cell lung carcinoma.
  • Statement 8 The method of any of Statements 1-6, wherein the defects are mutations or copy number alterations.
  • Statement 9 The method of Statement 8, wherein the mutations are deletions, frameshift mutations, insertions, missense mutations, nonsense mutations, start codon loss, stop codon loss or gain, or a combination thereof.
  • Statement 10 The method of any of Statements 1-6, wherein the detecting defects in nucleic acids encoding genes, or their expression products, for the biomarkers comprises performing next generation sequencing (NGS), nucleic acid hybridization, quantitative RT-PCR, immunohistochemistry (IHC), immunocytochemistry (ICC), or immunofluorescence (IF).
  • NGS next generation sequencing
  • IHC immunohistochemistry
  • ICC immunocytochemistry
  • IF immunofluorescence
  • Statement 11 The method of any of Statements 1-6, wherein the method further comprises assessment of a medical history, a family history, a physical examination, an endoscopic examination, imaging, a biopsy result, or a combination thereof.
  • Statement 12 The method of Statement 11, wherein the method is used to develop a treatment strategy for the subject with cancer.
  • Statement 13 The method of any of Statements 1-6, wherein the nucleic acids encoding genes are isolated from a fixed, paraffin-embedded sample from the subject.
  • Statement 14 The method of any of Statements 1-6, wherein the nucleic acids encoding genes are isolated from core biopsy tissue or fine needle aspirate cells from the subject.
  • Statement 15 The method of Statements 5 or 6, which further comprises treating the subject with a combination of ICB therapy and kinase inhibitor therapy.
  • a method for treating a subject with cancer which comprises independently measuring or obtaining a tumor mutational burden (TMB) level; and defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of low BIG-BETS DDR genes normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample; and if the subject has a high TMB level and wild type low BIG-BETs genes, treating the subject with a combination of immune checkpoint blockcade (ICB) therapy and an inhibitor of a low BIG- BETs kinase so as to reduce the activity of the low BIG-BET kinase and thereby treat the subject with cancer.
  • TMB tumor mutational burden
  • DDR DNA damage repair
  • Statement 17 The method of Statement 16, wherein the low BIG-BETs kinase is ATR, CHEK1, or WEE1.
  • Statement 18 A kit comprising at least ten nucleic acid probes, wherein each of said probes specifically binds to one of ten distinct biomarker nucleic acids or fragments thereof selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5.

Landscapes

  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Chemical & Material Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Physics & Mathematics (AREA)
  • Pathology (AREA)
  • Medical Informatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Organic Chemistry (AREA)
  • Immunology (AREA)
  • Public Health (AREA)
  • Wood Science & Technology (AREA)
  • Biophysics (AREA)
  • Molecular Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Analytical Chemistry (AREA)
  • Biotechnology (AREA)
  • Zoology (AREA)
  • Epidemiology (AREA)
  • Theoretical Computer Science (AREA)
  • Hospice & Palliative Care (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Oncology (AREA)
  • Evolutionary Biology (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Primary Health Care (AREA)
  • Microbiology (AREA)
  • Databases & Information Systems (AREA)
  • Data Mining & Analysis (AREA)
  • Biomedical Technology (AREA)
  • Biochemistry (AREA)
  • General Engineering & Computer Science (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

This disclosure is directed to a novel test, the Bipartite Graph-Based Expected TMB Score (BiG-BETS), that resolves the TMB Paradox, accurately defines genes associated with elevated TMB, and remarkably delineates a cohort of subjects (TMB-High, Low BiG-BETS DDR mutant) with high predictive power for ICB response and prolonged overall survival for cancer treatment.

Description

IMPROVED METHODS OF PREDICTING RESPONSE TO IMMUNE CHECKPOINT BLOCKADE THERAPIES AND USES THEREOF CROSS REFERENCE TO RELATED APPLICATIONS [0001] This application claims the benefit of U.S. Provisional Appn. No. 63/320,169 filed March 15, 2022, Kim et al., entitled “IMPROVED METHODS OF PREDICTING RESPONSE TO IMMUNE CHECKPOINT BLOCKADE THERAPIES AND USES THEREOF”, Atty. Dkt. 150- 33-PROV, which is hereby incorporated by reference in its entirety. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT [0002] This invention was made with government support under Grant Number DK111930 awarded by the National Institutes of Health. The government has certain rights in the invention. REFERENCE TO A COMPUTER PROGRAM LISTING AND TABLES APPENDIX SUBMITTED AS AN ASCII TEXT FILE [0003] This application contains nice computer program appendices. They have been submitted electronically via EFS-Web as an ASCII text files. The program text files were all created on March 14, 2022. Program 1 is entitled “150-33-1_INIT.txt” and 1,088 bytes in size. Program 2 is entitled “150-33-2_bipartite_helper_functions.txt” and 8,485 bytes in size. Program 3 is entitled “150-33- 3_bipartite_matching.txt” and 13,493 bytes in size. Program 4 is entitled “150-33- 4_ddr_data_object.txt” and 6,623 bytes in size. Program 5 is entitled “150-33-5_file_locations.txt” and 1,649 bytes in size. Program 6 is entitled “150-33-6_load_clinical_datasets.txt” and 39,528 bytes in size. Program 7 is entitled “150-33-7_load_pmec.txt” and 11,478 bytes in size. Program 8 is entitled “150-33-8_load_tcga_dataset.txt” and 19,905 bytes in size. Program 9 is entitled “150- 33-9_name_matching_scripts.txt” and 1,375 bytes in size. Supplemental Table 1 is entitled 150-33- PCT_SUPP_TABLE_1 and is 1,679,792 bytes in size. Supplemental Table 2 is entitled 150-33- PCT_SUPP_TABLE_2 and is 7,058,273 bytes in size. They are hereby incorporated by reference in their entireties. 1. FIELD [0004] The present disclosure provides an improved method of selecting patients for immune checkpoint blockade (ICB) treatment that complements existing methods of measuring total mutational burden (TMB) of a particular cancer in a subject. 2. BACKGROUND 2.1. Introduction [0005] The “background” description provided herein is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in this background section, as well as aspects of the description which may not otherwise qualify as prior art at the time of filing, are neither expressly nor impliedly admitted as prior art against the present disclosure. [0006] Immune checkpoint blockade (ICB) has achieved remarkable success in many solid tumors. Nonetheless, only a minority of patients respond and ICB is associated with significant financial toxicity (1); therefore, the ability to better predict ICB response has the potential to impact both patient survival and quality of life. Several genomic markers have demonstrated consistent predictive power for ICB response including tumor intrinsic properties such as tumor mutational burden (TMB) as well as RNA expression signatures (i.e. a T cell inflamed gene expression profile) (2). While high levels of TMB in particular have been consistently associated with ICB response (3), even patients with high TMB levels have response rates below 40%. Therefore, despite these important observations, there remains significant patient heterogeneity in ICB response that is not explained by existing biomarkers. [0007] TMB represents the balance between a tumor's exposure to a mutagenic process (i.e. UV radiation, carcinogen, etc.) and the integrity of the cellular DNA Damage Repair (DDR) pathways. Consistent with this notion, an elevated TMB is frequently seen in tumors associated with carcinogens (i.e. melanoma & UV radiation, lung cancer & cigarette smoke) (3,4) and has also been associated with mutations in some DDR genes (5–9). In contrast, in a comprehensive study the TCGA DDR working group assessed whether mutations in DDR genes were associated with elevated TMB. They found only two DDR genes that when mutated were significantly associated with a higher TMB than other genes in the cohort (10). We sought to resolve this discrepancy as well as to assess whether mutations in DDR genes have predictive power for ICB response independent of TMB. 3. SUMMARY OF THE DISCLOSURE [0008] This disclosure is directed to a novel test, the Bipartite Graph-Based Expected TMB Score (BiG-BETS), that resolves the TMB Paradox, accurately defines genes associated with elevated TMB, and remarkably delineates a cohort of subjects (TMB-High, Low BiG-BETS DDR mutant) with high predictive power for ICB response and prolonged overall survival. [0009] The present disclosure provides a method for predicting response to immune checkpoint blockade (ICB) therapy or overall survival for a subject with cancer, comprising independently measuring or obtaining (i) a tumor mutational burden (TMB) level and (ii) defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5 in a sample from the subject, normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample, wherein subjects with both a high TMB level and defects in at least ten of the nucleic acids encoding DDR genes, or their expression products, is indicative of an increased response to ICB therapy for the subject with cancer. [0010] Also provided is a method to select a subject with cancer for immune blockade therapy (ICB) which comprises independently measuring or obtaining both (i) a tumor mutational burden (TMB) level and (ii) defects in DNA Damage Repair (DDR) genes or their expression products in a sample from the subject; calculating a Bipartite Graph-Based Expected TMB Score (BiG-BETS) to determine which DDR genes or expression products are associated with elevated TMB; if the subject has a high BiG-BETS determining that the defect in the DDR genes or their expression products do not add to the TMB level over measurement of TMB levels alone; and using the TMB levels alone to select a subject for ICB. [0011] In the methods above, the DDR genes or their expression products may be at least ten biomarkers selected from the group consisting of BARD1, CHEK1, CUL5, EME1, ERCC4, ERCC5, ERCC6, FANCA, FANCC, FANCI, FEN1, GEN1, LIG4, MDC1, MLH1, MLH3, MRE11, MSH2, MSH3, MSH6, NBN, PALB2, PARP1, PMS1, PMS2, POLE, POLE3, POLL, POLM, RAD50, RNMT, SEM1, SLX1A, TDG, TOP3A, TOPBP1, UNG, XPC, XRCC2, XRCC4, and XRCC6. [0012] Also provided is a method to select a subject with cancer for immune blockade therapy (ICB) which comprises independently measuring or obtaining both (i) a tumor mutational burden (TMB) level and (ii) defects in DNA Damage Repair (DDR) genes or their expression products in a sample from the subject; calculating a Bipartite Graph-Based Expected TMB Score (BiG-BETS) to determine which DDR genes or expression products are associated with elevated TMB; if the subject has a low BiG-BETS determining that the defects in the DDR genes or their expression products contribute to the likelihood of benefit from ICB over measurement of TMB levels alone; and using both the TMB levels and the defects in the DDR genes or their expression products to select a subject for ICB. [0013] In this method, the DDR genes or their expression products may be at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5. [0014] In the methods above, the cancer may be bladder urothelial carcinoma, colon adenocarcinoma, esophageal carcinoma, invasive breast carcinoma, head and neck squamous cell carcinoma, kidney renal clear cell carcinoma, lung adenocarcinoma, melanoma, or squamous cell lung carcinoma and the defects may be mutations or copy number alterations. The mutations may be deletions, frameshift mutations, insertions, missense mutations, nonsense mutations, start codon loss, stop codon loss or gain, or a combination thereof. [0015] For the methods above the detecting defects in nucleic acids encoding genes, or their expression products, for the biomarkers may comprise performing next generation sequencing (NGS), nucleic acid hybridization, quantitative RT-PCR, immunohistochemistry (IHC), immunocytochemistry (ICC), or immunofluorescence (IF). The method may further comprises assessment of a medical history, a family history, a physical examination, an endoscopic examination, imaging, a biopsy result, or a combination thereof. The method may be used to develop a treatment strategy for the subject with cancer. [0016] In the methods above, the nucleic acids encoding genes may be isolated from a fixed, paraffin-embedded sample from the subject. Alternatively, the nucleic acids encoding genes are isolated from core biopsy tissue or fine needle aspirate cells from the subject. [0017] The method may further comprises treating the subject with a combination of ICB therapy and kinase inhibitor therapy. [0018] In an alternative embodiment, the disclosure provides a method for treating a subject with cancer which comprises independently measuring or obtaining a tumor mutational burden (TMB) level; and defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of low BIG- BETS DDR genes normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample; and if the subject has a high TMB level and wild type low BIG- BETs genes, treating the subject with a combination of immune checkpoint blockcade (ICB) therapy and an inhibitor of a low BIG-BETs kinase so as to reduce the activity of the low BIG-BET kinase and thereby treat the subject with cancer. The low BIG-BETs kinase may be ATR, CHEK1, or WEE1. [0019] In addition, a kit is provided. Specifically, a kit comprising at least ten nucleic acid probes, wherein each of said probes specifically binds to one of ten distinct biomarker nucleic acids or fragments thereof selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5. 4. BRIEF DESCRIPTION OF THE FIGURES [0020] Fig. 1A-Fig. 1F. Univariate testing inappropriately associates most genes with elevated TMB and recasting samples and mutations as a bipartite network overcomes this limitation. [0021] (Fig.1A) Distribution of Mann-Whitney U test p-values (with multiple test correction) on reversed log scale across all genes in Pan-TCGA dataset (light gray) vs DDR genes only (dark gray). For each gene, the MWU test compares distribution of TMB values for samples with a mutation in the gene vs all samples in the cohort. Right of dashed black line represents p<0.05. (Fig.1B) Percentage of genes in which mutations are significantly associated with elevated TMB by the MWU test (with FDR correction) broken down by DDR genes (dark gray) and non-DDR genes (light gray) (Fig.1C- Fig.1D) Distribution of mean TMB values for mutated sample set for all genes (light gray) vs DDR genes (dark gray) in TCGA. Vertical dotted black line in Fig. 1D denotes the overall mean TMB for the cohort of samples. (Fig. 1E) Schematic representation of converting the mutational data in a matrix to a bipartite network. (Fig. 1F) Schematic representation of BiG-BETS network rewiring process to sample from the bipartite configuration model. Random pairs of edges are selected to be exchanged to generate new samples. [0022] Fig. 2A-Fig. 2B (Fig. 2A) Violin plots of tumor mutational burden (TMB) by DDR pathway using Pan-TCGA dataset, DDR (DNA Damage Repair) genes were categorized into DDR pathways according to TCGA DDR working group. Violin plots show the distribution of tumor mutational burdens (TMB) for samples with a mutation in each of the DDR pathways. Dashed line = median TMB for all TCGA samples with the horizonal light gray strip showing the interquartile range. Mann-Whitney U p-values are calculated by comparing the distribution of TMB for samples with any mutation in the genes that define a pathway (counting only once if they have multiple mutations) to the distribution of all samples, including those with mutations in the same DDR pathway. (Fig.2B) Schematic and example of the TMB paradox. Note how the addition of a single, highly mutated sample (right) only raises the average degree from 1.3 to 2 (~50% increase), while raising the average gene’s neighbor’s degree from 1.25 to 3 (140% increase). This effect is even larger network with a heavy tailed degree distribution. [0023] Fig.3A-Fig.3C (Fig.3A) Distributions of p-values for the BiG-BETS permutation test applied to all 18,000 genes in the TCGA dataset, split by non-DDR genes (light gray) vs DDR genes (dark gray). (Fig. 3B) Expected distribution of TMB values (shown in gray) for samples with mutation in ATR (top) and CHEK1 (bottom). Observed distribution of TMB for mutated tumors is shown by lighter gray curve. Distribution of means across the 400 samples from the network model shown in figure inset with dashed vertical line denoting observed mean TMB. (Fig. 3C) Comparison of z-scores derived from Samstein et al. (x-axis) with those from TCGA using only genes present in both sets (468 genes). DDR genes shown as gray stars. [0024] Fig. 4A-Fig. 4D. A networks-based model and permutation test (BiG-BETS) are superior to univariate model. [0025] (Fig. 4A) Percentage of genes that are significant (with multiple test correction) using MWU test (left two bars) and significant by the networks-based test (right two bars). Differences between DDR vs non-DDR genes computed assessed with Chi-squared test with p-value shown above the corresponding bars. (Fig. 4B) (Distribution of BiG-BETS z-scores for the DDR genes. Low vs High z-score DDR genes defined using z-score < 0 and z-score > 0 (dashed vertical line) respectively, with individual genes in each bin listed about the plot. (Fig. 4C) Application of bipartite configuration test to the DDR pathways in the TCGA data. Each subplot shows the observed cumulative distribution of TMB for samples with a mutation in the genes of the specified pathway by the solid line. The dark line shows the average cumulative distribution across 400 sampled networks, with the light gray band showing the 99% confidence interval. Horizontal line at y=0.5 denotes the median TMB for the distributions. Inset figures show a histogram of the means of the sampled distributions of TMB for samples with a mutation in the corresponding DDR pathway. The vertical dashed line within the inset depicts the observed mean TMB in the actual data set. Z-scores were constructed by comparing the observed mean TMB to the sampled means. (Fig. 4D) Significant gene ontology (GO) terms identified in the 50 lowest BiG-BET genes from the Pan-TCGA dataset. [0026] Fig. 5A-Fig. 5C (Fig. 5A) Schematic depicting criteria for filtering variants in the TCGA cohort used to calculate the BiG-BET scores. Variants were filtered on basis of impact as defined by the sequence ontology and polphen status (see Methods for full details). (Fig. 5B) Comparison of BiG-BET scores obtained from TCGA cohort based on High Consequence mutations only (x-axis) and Moderate+High Consequences (y-axis). From left to right all genes (with hotspot genes shown in gray), DDR genes only, and the DDR pathways. (Fig. 5C) Venn- diagram of overlap of DDR genes between the Samstein and IMvigor210 cohorts. Outer dark gray circle depicts all 72 of the core-DDR genes. [0027] Fig.6A-Fig.6D (Fig.6A) Overall survival in IMvigor210 (left) and Samstein (right) by TMB-H (defined as TMB >=10) vs TMB-L (TMB<10). Tables to below shows CPH coefficients for model with TMB (as continuous variable) as well as mutation in DDR gene. (Fig. 6B) Percentage of patients that have complete or partial response in TMB-H vs TMB-L for IMvigor210. Significance assessed using Chi-square test. (Fig. 6C) Kaplan-Meier curves depicting overall survival (OS) in the IMvigor210 (left) and Samstein (right) cohorts broken down by TMB-H (light gray-lines) vs TMB-L (dark gray-lines) and into samples with a mutation in High-BiG-BETS DDR gene (bold lines) and High-BiG-BETS DDR WT tumors (dotted lines). Tables below shows forest plot of coefficients for CPH model jointly testing TMB (as continuous variable), mutation in High- BiG-BETS DDR gene, as well as an interaction term between the two variables (denoted by High BiG-BETS: TMB-H). (Fig. 6D) Percentage of patients that have complete or partial response in High BiG-BETS DDR mutant vs wild type in TMB-H and TMB-L tumors from IMvigor210. Significance assessed using Chi-square test. [0028] Fig. 7A-Fig. 7I. TMB high tumors with mutation in a Low BiG-BETS DDR gene have improved survival and response. [0029] (Fig.7A) Kaplan-Meier curves depicting overall survival (OS) in the IMvigor210 cohort broken down by TMB-High (light gray-lines) vs TMB-Low (dark gray-lines) and into samples with a mutation in Low BiG-BETS DDR genes (bold lines) and Low BiG-BETS DDR WT tumors (dotted lines). Table underneath shows forest plot of coefficients for CPH model jointly testing TMB (as continuous variable), mutation in Low BiG-BETS DDR genes, as well as an interaction term between the two variables (denoted by Low BiG-BETS; TMB-H). Patient counts for each category in TMB-H_MUT, TMB-H_WT, TMB-L_MUT, and TMB-L_WT were 19, 84, 12, and 159 respectively. (Fig. 7B) KM curves depicting OS in the Samstein et al cohort broken down along the same lines as (A) with corresponding coefficients in CPH model below. Patient counts for each category in TMB-H_MUT, TMB-H_WT, TMB-L_MUT, and TMB-L_WT were 67, 307, 72, and 861 respectively. (Fig. 7C) Percentage of patients with response (complete or partial response) to ICB in the IMvigor210 dataset. Patients are stratified into by TMB (TMB-High vs TMB-Low) and into samples with a Low BiG-BETS DDR gene mutation or not (WT). The number of patients who were responders (CR/PR) in each category from left to right is 12, 1, 28, and 10 respectively. Significant differences between groups tested using Chi-squared. (Fig. 7D) KM curves depicting OS in the Weir metadataset (see Methods for full description) broken down along the same lines as (A) and (C) with corresponding coefficients in CPH model below. Patient counts for each category in TMB-H_MUT, TMB-H_WT, TMB-L_MUT, and TMB-L_WT were 60, 117, 33, and 201 respectively. (Fig.7E) Response rates by Low BiG-BETS DDR mutations in the Weir Metadataset. (Fig. 7G- Fig. 7I) KM curves depicting OS in a combined dataset that includes IMVigor210, Samstein et al., and Weir metadataset split out by tumor type including (Fig. 7G) bladder cancer, (Fig.7H) non-small cell lung cancer, and (Fig.7I) melanoma. [0030] Each plot is broken down along the same lines as (Fig. 7A) and (Fig. 7B) with corresponding coefficients in CPH model below. [0031] Fig. 8A-Fig. 8C (Fig. 8A) Kaplan-Meier (KM) curves depicting overall survival (OS) in the Samstein cohort broken down by TMB-High (light gray-lines) vs TMB-Low (dark gray-lines) and into samples with a mutation in Low BiG-BETS DDR genes (bold lines) using only 17 DDR genes that overlap with IMvigor210 mutations (see Fig.5C) (bold lines) vs those that are Low BiG- BETS DDR WT (dashed line). (Fig. 8B) KM curves depicting OS in the Samstein et al and IMvigor210 cohort broken down along the same lines (TMB-H vs TMB-L and Low-BIG-BETS chromatin remodeling pathway genes MUT vs WT), with corresponding coefficients in CPH model below. Response rates in IMvigor210 by TMB-H vs TMB-L and mutation in low-BiG-BET chromatin remodeling genes (Fig. 8C) KM curves depicting overall survival (OS) in the Braun et al. 2020 ccRCC cohort broken down along the same lines as (Fig.8A) and percentage of patients that have complete or partial response in TMB-H vs TMB-L for Braun et al 2020. Significance assessed using Chi-square test. [0032] Fig. 9A-Fig. 9D. Mutation of Low BiG-BETS DDR genes in TMB High tumors is associated with elevated STING and IRF3 gene signatures. [0033] (Fig. 9A- Fig. 9D) Box plots of indicated gene signatures in IMVigor210 patients stratified by TMB (TMB-High vs TMB-Low) and Low BIG-BETS DDR gene mutation or not (WT). For each signature, each sample is assigned a z-score based on the average expression level of all genes in the signature compared to the average across all samples (See Methods). Significance calculated using the Mann-Whitney U test. [0034] Fig.10A-Fig.10D Box plots of indicated gene signatures in TCGA tumors subsetted to match Samstein (see Methods) stratified by TMB (TMB-High vs TMB-Low) and Low BIG-BETS DDR gene mutation or not (WT). For each signature, each sample is assigned a z-score based on the average expression level of all genes in the signature compared to the average across all samples (See Methods). Significance calculated using the Mann-Whitney U test. (Fig.10E) Proportion of genes in each DDR pathway that are categorized as High and Low BIG-BETS genes. [0035] Fig.11A and Fig. 11B shows a graphical abstract of the disclosure. In Fig.11A Genes and samples are represented as a bipartite network. A Bipartite Graph-Based Expected TMB (Big- BET) score is determined by a random rewiring of the bipartite network. In Fig.11B for subjects with mutations in BiG-BETs high genes, their ICB response is driven by their TMB levels. Subjects with mutations in the BiG-BETs low genes and high TMB, are more likely to respond to ICB and show increased overall survival (OS). 5. DETAILED DESCRIPTION OF THE DISCLOSURE [0036] Immune checkpoint blockade (ICB) has had remarkable success for treatment of solid tumors. However, as only a subset of patients exhibit responses, there is a continued need for biomarker development. Numerous reports have shown a link between tumor mutational burden (TMB) and ICB response, while others have identified a link between ICB response and mutation in DNA Damage Repair (DDR) genes. However, it remains unclear to what extent mutations in DDR genes hold predictive value above and beyond their association with TMB. [0037] This disclosure provides present a novel, networks-based test and Bipartite Graph-Based Expected TMB Score (BiG-BETS) with higher specificity for discriminating DDR genes and pathways that are associated with elevated TMB. Moreover, mutations in certain DDR genes that are not associated with elevated TMB (Low BiG-BETS) are found nevertheless predictive of ICB benefit in high TMB patients, demonstrating that their inactivation contributes to ICB response in a TMB independent manner. [0038] There are many examples of approved drugs that may be used in treatment to accompany the methods disclosed herein including small molecule kinase inhibitors such as imatinib (Gleevac®) an inhibitor of breakpoint cluster region-abelson (BCR-ABL) approved initially for chronic myelogenous leukemia (CML). Examples of monoclonal antibody kinase inhibitors are trastuzumab (Herceptin®), an inhibitor of ERB-B2 and approved for breast cancer or bevacizumab (Avastin®), an inhibitor of vascular endothelial growth factor (VEGF) approved for colorectal cancer. Other examples of drugs approved with a companion diagnostic include drugs approved for BRCA1/2 mutations, KRAS mutations and cKIT expression. Table 1 lists a number of approved drugs including a number of kinase inhibitors. See, Janne et al., 2009 Nat. Rev. Drug Disc.8709-723; Levitzki and Klein, 2010 Mol. Aspects Med.31, 287-329; and Mellor et al.2011 Tox. Sci.120(1) 14-32; and the package inserts for the specific drugs. [0039] TABLE 1
Figure imgf000012_0001
Figure imgf000013_0001
Figure imgf000014_0001
Figure imgf000015_0001
[0040] Abbreviations: For gene targets see Gene Cards (http://www.genecards.org/). For indications: ALL, acute lymphoblastic leukemia; AML, acute myeloid leukemia; BrCA, breast cancer; BCC, basal cell carcinoma; CML, chronic myeloid leukemia; CMML, chronic myelomonocytic leukemia; CRC, colorectal cancer; CSCC, cutaneous squamous cell carcinoma; GIST, gastrointestinal stromal tumor; HCC, hepatocellular carcinoma; HNSCC, head and neck squamous cell carcinoma; MCC, Merkle cell carcinoma; MDS/MPD, myelodysplastic syndrome/myeloproliferative disease; NSCLC, non-small cell lung cancer; OvCA, ovarian cancer; RCC, renal cell carcinoma; STS, soft tissue sarcoma; TNBC, triple negative breast cancer; and UC, urothelial carcinoma. [0041] Many of the drugs in the above Table are approved for use with a companion diagnostic. For example, trastuzumab (Herceptin®) is approved for breast cancer over expressing ERB-B2 and cetuximab (Erbitux®) for patients with wild-type KRAS. Amado et al., 2008, J Clin Oncol 26 (10): 1626–1634; Allegra et al., 2009 J Clin Oncol 272091-2096. Another kinase inhibitor approved for use with a diagnostic is crizotinib (Xalkori®) approved with a fluorescent in situ hybridization (FISH) test for ALK rearrangements (Vysis LSI ALK Dual Color, Break Apart Rearrangement Probe; Abbott Molecular, Abbott Park, IL). Shah et al., 2011 Lancet Oncol 121004-1012; Shaw et al., 2009 J Clin Oncol 274247-4253. Vemurafenib (Zelboraf®) is approved for use in patients with BRAF V600E mutation (Cobas 4800 BRAF V600 Mutation Test, Roche Molecular Diagnostics, Pleasanton, CA). Chapman et al., 2011 NEJM 3642507- 2516. Additional details may be found at the US FDA website for companion diagnostics (https://www.fda.gov/MedicalDevices/ProductsandMedicalProcedures/InVitroDiagnostics/ucm30 1431.htm). See the Biomarker column in Table 1 for additional companion diagnostics. [0042] The methods disclosed herein may be used as an aid in the diagnostics and treatment of a number of cancers. Once a particular cancer is diagnosed there are a variety of targeted therapies that a clinician may use to treat the patient. Non-limiting examples for bladder cancer include erdafitinib (BALVERSA™) or pembrolizumab (KEYTRUDA®) in Table 1, additional therapies include avelumab (BAVENCIO®), durvalumab (IMFINZI™), or nivolumab (OPDIVO®). Non-limiting examples for BrCA include abemaciclib (VERZENIO®), ado- trastuzumab emtansine (KADCYLA®), alpelisib (PIQRAY®), atezolizumab (TECENTRIQ®), Everolimus (AFINITOR®), lapatinib (TYKERB®), olaparib (LYNPARZA®), palbociclib (IBRANCE®), pertuzumab (PERJETA®), ribociclib (KISQALI®) or trastuzumab (HERCEPTIN®), or trastuzumab (HERCEPTIN HYLECTATM) in Table 1, additional therapies include anastrozole (ARIMIDEX®), exemestane (AROMASIN®), fulvestrant (FASLODEX®), letrozole (FEMARA®), neratinib (NERLYNX™), tamoxifen (SOLTAMOX®), or toremifene (FARESTON®). Non-limiting examples for CRC include bevacizumab (AVASTIN®), Cetuximab (ERBITUX®), panitumumab (VECTIBIX®), ramucirumab (CYRAMZA®), or regorafenib (STIVARGA®) in Table 1, additional therapies include ipilimumab (YERVOY®), nivolumab (OPDIVO®), or ziv-aflibercept (ZALTRAP®). Non-limiting examples for HCC include pembrolizumab (KEYTRUDA®), ramucirumab (CYRAMZA®), regorafenib (STIVARGA®), or sorafenib (NEXAVAR®) in Table 1, additional therapies include cabozantinib (CABOMETYX™), lenvatinib (LENVIMA®), or nivolumab (OPDIVO®). Non- limiting examples for kidney cancer include axitinib (INLYTA®), bevacizumab (AVASTIN®), cabozantinib (CABOMETYX®), Everolimus (AFINITOR®), pazopanib (VOTRIENT®), pembrolizumab (KEYTRUDA®), sorafenib (NEXAVAR®), sunitinib (SUTENT®), temsirolimus (TORISEL®) in Table 1, additional therapies include avelumab (BAVENCIO®), ipilimumab (YERVOY®), lenvatinib mesylate (LENVIMA®), or nivolumab (OPDIVO®). Non- limiting examples for leukemia include dasatinib (SPRYCEL®), enasidenib (IDHIFA®), gilteritinib (XOSPATA®), imatinib (GLEEVEC®), ivosidenib (TIBSOVO®), midostaurin (RYDAPT®), nilotinib (TASIGNA®), or venetoclax (VENCLEXTA®) in Table 1, additional therapies include alemtuzumab (CAMPATH®), blinatumomab (BLINCYTO®), bosutinib (BOSULIF®), duvelisib (COPIKTRA™), gemtuzumab ozogamicin (MYLOTARG™), glasdegib (DAURISMO™), ibrutinib (IMBRUVICA®), idelalisib (ZYDELIG®), inotuzumab ozogamicin (BESPONSA®), moxetumomab pasudotox-tdfk (LUMOXITI™), obinutuzumab (GAZYVA®), ofatumumab (ARZERRA®), ponatinib (ICLUSIG®), rituximab (RITUXAN®), rituximab and hyaluronidase human (RITUXAN HYCELA™), tagraxofusp-erzs (ELZONRIS™), tisagenlecleucel (KYMRIAH®), or tretinoin (VESANOID®). Non-limiting examples for lung cancers, e.g., NSCLC, include in Table 1 afatinib (GILORAF®), alectinib (ALECENSA®), atezolizumab (TECENTRIQ®), bevacizumab (AVASTIN®), ceritinib (LDK378/ZYKADIA®), crizotinib (XALKORI®), dabrafenib (TAFINAR®), dacomitinib (VIZIMPRO®), erlotinib (TARCEVA®), gefitinib (IRESSA®), osimertinib (TAGRISSO®), pembrolizumab (KEYTRUDA®), pemetrexed (ALIMTA®), ramucirumab (CYRAMZA®), trametinib (MEKANIST®), additional therapies include brigatinib (ALUNBRIG™), durvalumab (IMFINZI™), lorlatinib (LORBRENA®), necitumumab (PORTRAZZA™), nivolumab (OPDIVO®). Non-limiting examples for lymphoma include acalabrutinib (CALQUENCE®), pembrolizumab (KEYTRUDA®), venetoclax (VENCLEXTA®) in Table 1, additional therapies include axicabtagene ciloleucel (YESCARTA™), belinostat (BELEODAQ®), bexarotene (TARGRETIN®), bortezomib (VELCADE®), brentuximab vedotin (ADCETRIS®), copanlisib (ALIQOPA™), denileukin diftitox (ONTAK®), duvelisib (COPIKTRA™), Ibritumomab tiuxetan (ZEVALIN®), ibrutinib (IMBRUVICA®), idelalisib (ZYDELIG®), mogamulizumab-kpkc (POTELIGEO®), nivolumab (OPDIVO®), obinutuzumab (GAZYVA®), polatuzumab vedotin-piiq (POLIVY™), pralatrexate (FOLOTYN®), rituximab (Rituxan®), rituximab and hyaluronidase human (RITUXAN HYCELA™), romidepsin (ISTODAX®), siltuximab (SYLVANT®), tisagenlecleucel (KYMRIAH®), vorinostat (ZOLINZA®). Non-limiting examples for melanoma include alitretinoin (PANRETIN®), binimetinib (MEKTOVI®), cobimetinib (COTELLIC®), dabrafenib (TAFINAR®), encorafenib (BRAFTOVI™), pembrolizumab (KEYTRUDA®), trametinib (MEKANIST®), or vemurafenib (ZELBORAF®) in Table 1, additional therapies include avelumab (BAVENCIO®), cemiplimab- rwlc (LIBTAYO®), ipilimumab (YERVOY®), nivolumab (OPDIVO®), sonidegib (ODOMZO®), or vismodegib (ERIVEDGE®). Non-limiting examples for multiple myeloma (MM) include Bortezomib (VELCADE®), carfilzomib (KYPROLIS®), daratumumab (DARZALEX™), elotuzumab (EMPLICITI™), ixazomib (NINLARO®), panobinostat (FARYDAK®), selinexor (XPOVIO™). Non-limiting examples for prostate cancer include abiraterone acetate (ZYTIGA®) in Table 1, additional therapies include apalutamide (ERLEADA™), Cabazitaxel (JEVTANA®), darolutamide (NUBEQA®), enzalutamide (XTANDI®), radium 223 dichloride (XOFIGO®). Additional drugs that may be used for cancer treatment include Denosumab (XGEVA®), Dinutuximab (UNITUXIN™), iobenguane I 131 (AZEDRA®), Lanreotide acetate (SOMATULINE® Depot), lutetium Lu 177-dotatate (LUTATHERA®), niraparib (ZEJULA™), rucaparib camsylate (RUBRACA™), ruxolitinib phosphate (JAKAFI®), Sirolimus (RAPAMUNE®), or Talazoparib (TALZENNA®). 5.1. Definitions [0043] While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject matter. [0044] Throughout the present specification, the terms “about” and/or “approximately” may be used in conjunction with numerical values and/or ranges. The term “about” is understood to mean those values near to a recited value. For example, “about 40 [units]” may mean within ± 25% of 40 (e.g., from 30 to 50), within ± 20%, ± 15%, ± 10%, ± 9%, ± 8%, ± 7%, ± 6%, ± 5%, ± 4%, ± 3%, ± 2%, ± 1%, less than ± 1%, or any other value or range of values therein or there below. Alternatively, depending on the context, the term “about” may mean ± one half a standard deviation, ± one standard deviation, or ± two standard deviations. Furthermore, the phrases “less than about [a value]” or “greater than about [a value]” should be understood in view of the definition of the term “about” provided herein. The terms “about” and “approximately” may be used interchangeably. [0045] Throughout the present specification, numerical ranges are provided for certain quantities. It is to be understood that these ranges comprise all subranges therein. Thus, the range “from 50 to 80” includes all possible ranges therein (e.g., 51-79, 52-78, 53-77, 54-76, 55-75, 60- 70, etc.). Furthermore, all values within a given range may be an endpoint for the range encompassed thereby (e.g., the range 50-80 includes the ranges with endpoints such as 55-80, 50- 75, etc.). [0046] As used herein, the verb “comprise” as used in this description and in the claims and its conjugations are used in its non-limiting sense to mean that items following the word are included, but items not specifically mentioned are not excluded. [0047] Throughout the specification the word “comprising,” or variations such as “comprises” or “comprising,” will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers or steps, but not the exclusion of any other element, integer or step, or group of elements, integers or steps. The present disclosure may suitably “comprise”, “consist of”, or “consist essentially of”, the steps, elements, and/or reagents described in the claims. [0048] It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely", "only" and the like in connection with the recitation of claim elements, or the use of a "negative" limitation. [0049] As used herein, "clinical signs of cancer" means and includes any sign or indication of the existence of cancer in a subject, which sign or indication would be well known to the skilled artisan (e.g., oncologist, nurse practitioner). The clinical signs of cancer may be any symptom known to be associated with the cancer. Clinical signs of some cancers include, for example, chronic pain, nausea, vomiting, abnormal taste sensation, constipation, urinary symptoms (e.g., bladder spasm), respiratory symptoms, skin problems (e.g., pruritus, hair loss), or fever, among others. [0050] As used herein, the term “reference set” may be an internal, external, or a universal reference set of nucleic acids or expression products used to calibrate a particular sample. For example, an internal reference set of nucleic acids may be obtained using normal tissue or a blood sample from the subject. Alternatively, an internal reference set may based on the total RNA in the sample. In another embodiment, the reference set may be a set of one or more housekeeping genes, e.g., human acidic ribosomal protein (HuPO), β-actin (BA), cyclophylin (CYC), glyceraldehyde-3- phosphate dehydrogenase (GAPDH), phosphoglycerokinase (PGK), β2-microglobulin (B2M), β- glucuronidase (GUS), hypoxanthine phosphoribosyltransferase (HPRT), transcription factor IID TATA binding protein (TBP), transferrin receptor (TfR), human acidic ribosomal protein (HuPO), elongation factor-1-α (EF-1-α), metastatic lymph node 51(MLN51), or ubiquitin conjugating enzyme (UbcH5B). See Dheda et al. 2004 BioTechniques 37:112-119. An external reference set may be obtained from clinical studies to determine normal ranges and ranges for a particular cancer. Alternatively, the reference set may be based on a particular patient population such as smokers, gender or race. In yet another embodiment, the reference set may be a universal reference set. Many commercial vendors sell cDNA and RNA reference sets of genes or reference libraries. [0051] As used herein, "remission" means and includes a period during which the symptoms of a cancer have been reduced or eliminated, as remission is ordinarily defined in the oncology art. [0052] As used herein "serially monitoring" levels of a biomarker in a sample, refers to measuring levels of a biomarker in a sample more than once, e.g., quarterly, bimonthly, monthly, biweekly, weekly, every three days, daily, or several times per day. Serial monitoring of a level includes periodically measuring levels of biomarkers at regular intervals as deemed necessary by the skilled artisan. [0053] The term "standard level" as used herein refers to a baseline level of a biomarker as determined in one or more normal subjects. For example, a baseline may be obtained from at least one subject and preferably is obtained from an average of subjects (e.g., n=2 to 100 or more), wherein the subject or subjects have no prior history of cancer. In the present invention, the measurement of biomarker levels may be carried out using the multiplexed copy number as described. [0054] As used herein, "elevation" of a measured level of a biomarker relative to a standard level means that the amount or concentration of a biomarker in a sample is sufficiently greater in a subject relative to the standard to be detected by the methods described herein. For example, elevation of the measured level relative to a standard level may be any statistically significant elevation which is detectable. Such an elevation may include, but is not limited to, about a 1%, about a 10%, about a 20%, about a 40%, about an 80%, about a 2-fold, about a 4-fold, about an 8- fold, about a 20-fold, or about a 100-fold elevation, or more, relative to the standard. The term "about" as used herein, refers to a numerical value plus or minus 10% of the numerical value. [0055] Non-limiting examples of signaling pathway modulators or chemotherapeutic agents known in the art are 5-fluorouracil; asparaginase; bevacizumab (AVASTIN®); bleomycin; campathecins; cetuximab (ERBITUX®); crizotinib (XALKORI®); cyclophosphamide; cytarabine; dacarbazine; dactinomycin; dasatinib (SPRYCEL®); daunorubicin; DNA methyltransferase inhibitors (DNMTs) such as azacitidine (VIDAZA®) and decitabine; doxorubicin; doxorubicin; epirubicin; erbstatin; erlotinib (TARCEVA®); estramustine; etoposide; etoposide; gefitinib (IRESSA®), gemcitabine, genistein, histone acetyl transferase inhibitors (HATs); histone deacetyl transferase inhibitors (HDACs) such as belinostat, entinostat (MS-275), panobinostat, PCI-24781, romidepsin (depsipeptide, FK-228), valproic acid, vorinostat (ZOLINZA®, SAHA) or heat shock protein inhibitors, including HSP90 inhibitors such as alvespimycin (IPI-493), AT13387, AUY922 (resorcinolic isoxazole amide), CNF2024 (BIIB021), HSP990, MPC-3100, retaspimycin (IPI-504), SNX-2112, SNX-5422, STA-9090, tanespimycin (17-AAG; KOS-953), or XL888; herbimycin A; hexamethylmelamine; hedgehog pathway inhibitors such as saridegib (IPL-926), vismodegib (ERIVEDGE™); hydroxyurea, idarubicin, ifosfamide, imatinib (GLEEVEC®), irinotecan, lapatinib (TYKERB®), lavendustin A, leucovorin, levamisole, mercaptopurine, methotrexate, mitomycin, mitoxantrone, mTOR inhibitors such as everolimus (AFINITOR®), sirolimus (RAPAMUNE®), temsirolimus (TORISEL®); nilotinib (TASIGNA®); nitrosoureas such as carmustine and lomustine; paclitaxel; panitumumab (VECTIBIX®); pazopanib (VOTRIENT®); pegaptanib (MACUGEN®); platinum compounds such as carboplatin, cisplatin, oxaplatin; plicamycin; procarbizine; proteasome inhibitors such as bortezomib (VELCADE®); ranibizumab (LUCENTIS®); sorafenib (NEXAVAR®); sunitinib (SUTENT®); taxanes such as docetaxel,paclitaxel, taxol; thioguanine; topotecan; trastuzumab (HERCEPTIN®); tyrosine kinase inhibitors; tyrphostins; vandetanib (CAPRELSA®); vemurafenib (ZELBORAF®); vinblastine; vinca alkaloids; vincristine; or vinorelbine. In a preferred embodiment, the chemotherapeutic agent is bevacizumab (AVASTIN®), cetuximab (ERBITUX®), crizotinib (XALKORI®), dasatinib (SPRYCEL®), erlotinib (TARCEVA®), everolimus (AFINITOR®), gefitinib (IRESSA®), imatinib (GLEEVEC®), lapatinib (TYKERB®), nilotinib (TASIGNA®), panitumumab (VECTIBIX®), pazopanib (VOTRIENT®), sirolimus (RAPAMUNE®), sorafenib (NEXAVAR®), sunitinib (SUTENT®), temsirolimus (TORISEL®), trastuzumab (HERCEPTIN®), vandetanib (CAPRELSA®), or vemurafenib (ZELBORAF®). Further examples of chemotherapeutic agents may be found Table 1 above in standard publications and texts. See e.g., National Comprehensive Cancer Network (NCCN GuidelineTM) or Manual of Clinical Oncology, Dennis A. Casciato and Barry B. Lowitz, ed., 4th edition, Jul.15, 2000, Little, Brown and Company, U.S. 5.2. Computing Devices [0056] A computing device may be implemented in programmable hardware devices such as processors, digital signal processors, central processing units, field programmable gate arrays, programmable array logic, programmable logic devices, cloud processing systems, or the like. The computing devices may also be implemented in software for execution by various types of processors. An identified device may include executable code and may, for instance, comprise one or more physical or logical blocks of computer instructions, which may, for instance, be organized as an object, procedure, function, or other construct. Nevertheless, the executable of an identified device need not be physically located together but may comprise disparate instructions stored in different locations which, when joined logically together, comprise the computing device and achieve the stated purpose of the computing device. In another example, a computing device may be a server or other computer located within a hospital or out-patient environment and communicatively connected to other computing devices (e.g., POS equipment or computers) for managing accounting, purchase transactions, and other processes within the hospital or out-patient environment. In another example, a computing device may be a mobile computing device such as, for example, but not limited to, a smart phone, a cell phone, a pager, a personal digital assistant (PDA), a mobile computer with a smart phone client, or the like. In another example, a computing device may be any type of wearable computer, such as a computer with a head-mounted display (HMD), or a smart watch or some other wearable smart device. Some of the computer sensing may be part of the fabric of the clothes the user is wearing. A computing device can also include any type of conventional computer, for example, a laptop computer or a tablet computer. A typical mobile computing device is a wireless data access-enabled device (e.g., an iPHONE® smart phone, a BLACKBERRY® smart phone, a NEXUS ONE™ smart phone, an iPAD® device, smart watch, or the like) that is capable of sending and receiving data in a wireless manner using protocols like the Internet Protocol, or IP, and the wireless application protocol, or WAP. This allows users to access information via wireless devices, such as smart watches, smart phones, mobile phones, pagers, two-way radios, communicators, and the like. Wireless data access is supported by many wireless networks, including, but not limited to, Bluetooth, Near Field Communication, CDPD, CDMA, GSM, PDC, PHS, TDMA, FLEX, ReFLEX, iDEN, TETRA, DECT, DataTAC, Mobitex, EDGE and other 2G, 3G, 4G, 5G, and LTE technologies, and it operates with many handheld device operating systems, such as PalmOS, EPOC, Windows CE, FLEXOS, OS/9, JavaOS, iOS and Android. Typically, these devices use graphical displays and can access the Internet (or other communications network) on so-called mini- or micro-browsers, which are web browsers with small file sizes that can accommodate the reduced memory constraints of wireless networks. In a representative embodiment, the mobile device is a cellular telephone or smart phone or smart watch that operates over GPRS (General Packet Radio Services), which is a data technology for GSM networks or operates over Near Field Communication e.g. Bluetooth. In addition to a conventional voice communication, a given mobile device can communicate with another such device via many different types of message transfer techniques, including Bluetooth, Near Field Communication, SMS (short message service), enhanced SMS (EMS), multi-media message (MMS), email WAP, paging, or other known or later-developed wireless data formats. Although many of the examples provided herein are implemented on smart phones, the examples may similarly be implemented on any suitable computing device, such as a computer. [0057] An executable code of a computing device may be a single instruction, or many instructions, and may even be distributed over several different code segments, among different applications, and across several memory devices. Similarly, operational data may be identified and illustrated herein within the computing device, and may be embodied in any suitable form and organized within any suitable type of data structure. The operational data may be collected as a single data set, or may be distributed over different locations including over different storage devices, and may exist, at least partially, as electronic signals on a system or network. [0058] The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided, to provide a thorough understanding of embodiments of the disclosed subject matter. One skilled in the relevant art will recognize, however, that the disclosed subject matter can be practiced without one or more of the specific details, or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the disclosed subject matter. [0059] As used herein, the term “memory” is generally a storage device of a computing device. Examples include, but are not limited to, read-only memory (ROM) and random access memory (RAM). [0060] The device or system for performing one or more operations on a memory of a computing device may be a software, hardware, firmware, or combination of these. The device or the system is further intended to include or otherwise cover all software or computer programs capable of performing the various heretofore-disclosed determinations, calculations, or the like for the disclosed purposes. For example, exemplary embodiments are intended to cover all software or computer programs capable of enabling processors to implement the disclosed processes. Exemplary embodiments are also intended to cover any and all currently known, related art or later developed non-transitory recording or storage mediums (such as a CD-ROM, DVD-ROM, hard drive, RAM, ROM, floppy disc, magnetic tape cassette, etc.) that record or store such software or computer programs. Exemplary embodiments are further intended to cover such software, computer programs, systems and/or processes provided through any other currently known, related art, or later developed medium (such as transitory mediums, carrier waves, etc.), usable for implementing the exemplary operations disclosed below. [0061] In accordance with the exemplary embodiments, the disclosed computer programs can be executed in many exemplary ways, such as an application that is resident in the memory of a device or as a hosted application that is being executed on a server and communicating with the device application or browser via a number of standard protocols, such as TCP/IP, HTTP, XML, SOAP, REST, JSON and other sufficient protocols. The disclosed computer programs can be written in exemplary programming languages that execute from memory on the device or from a hosted server, such as BASIC, COBOL, C, C++, Java, Pascal, or scripting languages such as JavaScript, Python, Ruby, PHP, Perl, or other suitable programming languages. [0062] As referred to herein, the terms “computing device” and “entities” should be broadly construed and should be understood to be interchangeable. They may include any type of computing device, for example, a server, a desktop computer, a laptop computer, a smart phone, a cell phone, a pager, a personal digital assistant (PDA, e.g., with GPRS NIC), a mobile computer with a smartphone client, or the like. [0063] As referred to herein, a user interface is generally a system by which users interact with a computing device. A user interface can include an input for allowing users to manipulate a computing device, and can include an output for allowing the system to present information and/or data, indicate the effects of the user’s manipulation, etc. An example of a user interface on a computing device (e.g., a mobile device) includes a graphical user interface (GUI) that allows users to interact with programs in more ways than typing. A GUI typically can offer display objects, and visual indicators, as opposed to text-based interfaces, typed command labels or text navigation to represent information and actions available to a user. For example, an interface can be a display window or display object, which is selectable by a user of a mobile device for interaction. A user interface can include an input for allowing users to manipulate a computing device, and can include an output for allowing the computing device to present information and/or data, indicate the effects of the user’s manipulation, etc. An example of a user interface on a computing device includes a graphical user interface (GUI) that allows users to interact with programs or applications in more ways than typing. A GUI typically can offer display objects, and visual indicators, as opposed to text-based interfaces, typed command labels or text navigation to represent information and actions available to a user. For example, a user interface can be a display window or display object, which is selectable by a user of a computing device for interaction. The display object can be displayed on a display screen of a computing device and can be selected by and interacted with by a user using the user interface. In an example, the display of the computing device can be a touch screen, which can display the display icon. The user can depress the area of the display screen where the display icon is displayed for selecting the display icon. In another example, the user can use any other suitable user interface of a computing device, such as a keypad, to select the display icon or display object. For example, the user can use a track ball or arrow keys for moving a cursor to highlight and select the display object. [0064] The display object can be displayed on a display screen of a mobile device and can be selected by and interacted with by a user using the interface. In an example, the display of the mobile device can be a touch screen, which can display the display icon. The user can depress the area of the display screen at which the display icon is displayed for selecting the display icon. In another example, the user can use any other suitable interface of a mobile device, such as a keypad, to select the display icon or display object. For example, the user can use a track ball or times program instructions thereon for causing a processor to carry out aspects of the present disclosure. [0065] As referred to herein, a computer network may be any group of computing systems, devices, or equipment that are linked together. Examples include, but are not limited to, local area networks (LANs) and wide area networks (WANs). A network may be categorized based on its design model, topology, or architecture. In an example, a network may be characterized as having a hierarchical internetworking model, which divides the network into three layers: access layer, distribution layer, and core layer. The access layer focuses on connecting client nodes, such as workstations to the network. The distribution layer manages routing, filtering, and quality-of-server (QoS) policies. The core layer can provide high-speed, highly-redundant forwarding services to move packets between distribution layer devices in different regions of the network. The core layer typically includes multiple routers and switches. [0066] The present subject matter may be a system, a method, and/or a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present subject matter. [0067] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire. [0068] Computer readable program instructions described herein can be downloaded to respective computing/processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and/or a wireless network, or Near Field Communication. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and/or edge servers. A network adapter card or network interface in each computing/processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing/processing device. [0069] Computer readable program instructions for carrying out operations of the present subject matter may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state- setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++, Javascript or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present subject matter. [0070] Aspects of the present subject matter are described herein with reference to flowchart illustrations and/or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the subject matter. It will be understood that each block of the flowchart illustrations and/or block diagrams, and combinations of blocks in the flowchart illustrations and/or block diagrams, can be implemented by computer readable program instructions. [0071] These computer readable program instructions may be provided to a processor of a computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions/acts specified in the flowchart and/or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and/or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function/act specified in the flowchart and/or block diagram block or blocks. [0072] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions/acts specified in the flowchart and/or block diagram block or blocks. [0073] The description illustrates the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present subject matter. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the herein. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and/or flowchart illustration, and combinations of blocks in the block diagrams and/or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions. 5.3. Samples [0074] The sample may be from a subject suspected of having a particular cancer or from a patient diagnosed with cancer, e.g., for confirmation of diagnosis or establishing a clear margin or for the detection of cancer cells in other tissues such as lymph nodes, or circulating tumor cells. The biological sample may also be from a subject with an ambiguous diagnosis in order to clarify the diagnosis. The sample may be obtained for the purpose of differential diagnosis, e.g., a subject with a histopathologically benign lesion to confirm the diagnosis. The sample may also be obtained for the purpose of prognosis, i.e., determining the course of the disease and selecting primary treatment options. Tumor staging and grading are examples of prognosis. The sample may also be evaluated to select or monitor therapy, selecting likely responders in advance from non-responders or monitoring response in the course of therapy. In addition, the sample may be evaluated as part of post-treatment ongoing surveillance of patients who have had head and neck cancer. [0075] Samples may be obtained using any of a number of methods in the art. Examples of biological samples comprising potential cancer cells include those obtained from excised skin biopsies, such as punch biopsies, shave biopsies, core needle biopsies, fine needle aspirates (FNA), or surgical excisions; or biopsy from non- cutaneous tissues such as lymph node tissue, mucosa, other embodiments. In addition, the sample may be from a distant metastatic site, a soft tissue, e.g., lung, liver, bone, skin, or brain. Representative biopsy techniques include, but are not limited to, excisional biopsy, incisional biopsy, pinch biopsy, forceps biopsy, needle biopsy, or surgical biopsy. An "excisional biopsy" refers to the removal of an entire tumor mass with a small margin of normal tissue surrounding it. An "incisional biopsy" refers to the removal of a wedge of tissue that includes a cross-sectional diameter of the tumor. A diagnosis or prognosis made by endoscopy or fluoroscopy may require a "core-needle biopsy" of the tumor mass, or a "fine-needle aspiration biopsy" which generally contains a suspension of cells from within the tumor mass. The biological sample may be a microdissected sample, such as a PALM-laser (Carl Zeiss MicroImaging GmbH, Germany) capture microdissected sample. [0076] A sample may also be a sample of muscosal surfaces, blood and blood fractions or products (e.g., serum, plasma, platelets, red blood cells, white blood cells, circulating tumor cells isolated from blood, free DNA isolated from blood, and the like), sputum, saliva, lymph and tongue tissue, cultured cells, e.g., primary cultures, explants, and transformed cells, stool, urine, etc. The sample may also be vascular tissue or cells from blood vessels such as microdissected blood vessel cells of endothelial origin. A sample is typically obtained from a eukaryotic organism, most preferably a mammal such as a primate e.g., chimpanzee or human, cow, dog, cat; or a rodent, e.g., guinea pig, rat, mouse, rabbit. [0077] A sample can be treated with a fixative such as formaldehyde and embedded in paraffin (FFPE) and sectioned for use in the methods of the invention. Alternatively, fresh or frozen tissue may be used. These cells may be fixed, e.g., in alcoholic solutions such as 100% ethanol or 3:1 methanol:acetic acid. Nuclei can also be extracted from thick sections of paraffin-embedded specimens to reduce truncation artifacts and eliminate extraneous embedded material. Typically, biological samples, once obtained, are harvested and processed prior to nucleic acid analysis using standard methods known in the art. Such processing typically includes protease treatment and additional fixation in an aldehyde solution such as formaldehyde. 5.3.1. Polynucleotide Sequence Amplification and Determination [0078] In many instances, it is desirable to amplify a nucleic acid sequence using any of several nucleic acid amplification procedures which are well known in the art. Specifically, nucleic acid amplification is the chemical or enzymatic synthesis of nucleic acid copies which contain a sequence that is complementary to a nucleic acid sequence being amplified (template). The methods and kits of the invention may use any nucleic acid amplification or detection methods known to one skilled in the art, such as those described in U.S. Pat. Nos. 5,525,462 (Takarada et al.); 6,114,117 (Hepp et al.); 6,127,120 (Graham et al.); 6,344,317 (Urnovitz); 6,448,001 (Oku); 6,528,632 (Catanzariti et al.); and PCT Pub. No. WO 2005/111209 (Nakajima et al.); all of which are incorporated herein by reference in their entirety. [0079] In some embodiments, the nucleic acids may be amplified by PCR amplification using methodologies known to one skilled in the art. One skilled in the art will recognize, however, that amplification can be accomplished by other known methods, such as ligase chain reaction (LCR), Qβ-replicase amplification, rolling circle amplification, transcription amplification, self-sustained sequence replication, nucleic acid sequence-based amplification (NASBA), each of which provides sufficient amplification. Branched-DNA technology may also be used to qualitatively demonstrate the presence of a sequence of the technology which may quantitatively determine the amount of this particular genomic sequence in a sample. Nolte reviews branched-DNA signal amplification for direct quantitation of nucleic acid sequences in clinical samples (Nolte, 1998, Adv. Clin. Chem. 33:201-235). [0080] The PCR process is well known in the art and is thus not described in detail herein. For a review of PCR methods and protocols, see, e.g., Innis et al., eds., PCR Protocols, A Guide to Methods and Application, Academic Press, Inc., San Diego, Calif. 1990; U.S. Pat. No. 4,683,202 (Mullis); which are incorporated herein by reference in their entirety. PCR reagents and protocols are also available from commercial vendors, such as Roche Molecular Systems. PCR may be carried out as an automated process with a thermostable enzyme. In this process, the temperature of the reaction mixture is cycled through a denaturing region, a primer annealing region, and an extension reaction region automatically. Machines specifically adapted for this purpose are commercially available. 5.3.2. High Throughput and Single Molecule Sequencing Technology [0081] Suitable next generation sequencing technologies are widely available. Examples include the 454 Life Sciences platform (Roche, Branford, CT) (Margulies et al.2005 Nature, 437, 376-380); lllumina’s Genome Analyzer, Illumina’s MiSeq System, Illumina’s NextSeq System, Illumina’s MiniSeq System, (Illumina, San Diego, CA; Bibkova et al., 2006, Genome Res.16, 383- 393; U.S. Pat. Nos. 6,306,597 and 7,598,035 (Macevicz); 7,232,656 (Balasubramanian et al.)); or DNA Sequencing by Ligation, SOLiD System (Applied Biosystems/Life Technologies; U.S. Pat. Nos.6,797,470, 7,083,917, 7,166,434, 7,320,865, 7,332,285, 7,364,858, and 7,429,453 (Barany et al.); or the Helicos True Single Molecule DNA sequencing technology (Harris et al., 2008 Science, 320, 106-109; U.S. Pat. Nos.7,037,687 and 7,645,596 (Williams et al.); 7,169,560 (Lapidus et al.); 7,769,400 (Harris)), the single molecule, real-time (SMRTTM) technology of Pacific Biosciences, and sequencing (Soni and Meller, 2007, Clin. Chem.53, 1996-2001) which are incorporated herein by reference in their entirety. These systems allow the sequencing of many nucleic acid molecules isolated from a specimen at high orders of multiplexing in a parallel fashion (Dear, 2003, Brief Funct. Genomic Proteomic, 1(4), 397-416 and McCaughan and Dear, 2010, J. Pathol., 220, 297- 306). Each of these platforms allow sequencing of clonally expanded or non-amplified single molecules of nucleic acid fragments. Certain platforms involve, for example, (i) sequencing by ligation of dye-modified probes (including cyclic ligation and cleavage), (ii) pyrosequencing, (iii) targeted next-generation sequencing from bisulfite treated DNA and (iv) single-molecule sequencing. [0082] Pyrosequencing is a nucleic acid sequencing method based on sequencing by synthesis, which relies on detection of a pyrophosphate released on nucleotide incorporation. Generally, sequencing by synthesis involves synthesizing, one nucleotide at a time, a DNA strand complimentary to the strand whose sequence is being sought. Study nucleic acids may be immobilized to a solid support, hybridized with a sequencing primer, incubated with DNA polymerase, ATP sulfurylase, luciferase, apyrase, adenosine 5' phosphsulfate and luciferin. Nucleotide solutions are sequentially added and removed. Correct incorporation of a nucleotide releases a pyrophosphate, which interacts with ATP sulfurylase and produces ATP in the presence of adenosine 5' phosphosulfate, fueling the luciferin reaction, which produces a chemiluminescent signal allowing sequence determination. Machines for pyrosequencing are available from Qiagen, Inc. (Valencia, CA). An example of a system that can be used by a person of ordinary skill based on pyrosequencing generally involves the following steps: ligating an adaptor nucleic acid to a study nucleic acid and hybridizing the study nucleic acid to a bead; amplifying a nucleotide sequence in the study nucleic acid in an emulsion; sorting beads using a picoliter multiwell solid support; and sequencing amplified nucleotide sequences by pyrosequencing methodology (e.g., Nakano et al., 2003, J. Biotech.102, 117-124). Such a system can be used to exponentially amplify amplification products generated by a process described herein, e.g., by ligating a heterologous nucleic acid to the first amplification product generated by a process described herein. [0083] Next-generation sequencing (NGS) is a nucleic acid sequencing method based on sequencing by synthesis, where fluorescently labeled deoxyribonucleotide triphosphates (dNTPs) catalyzed by DNA polymerase are incorporated into a DNA temple through cycles of DNA synthesis and nucleotides are identified by fluorophore excitation at each incorporation step. NGS allows this process to take place in a multiplex reaction across millions of DNA fragments in parallel. Generally, sequencing by synthesis involves synthesizing, one nucleotide at a time, a DNA strand complimentary to the strand whose sequence is being sought. Study nucleic acids may be immobilized to a solid support, hybridized with a sequencing primer, and incubated with DNA polymerase in the presence of fluorescently labeled dNTPS. After each cycle, the image is scanned and the emission wavelength and intensity are recorded and used to identify the base incorporated. This process is repeated multiple times to create a specific read length of bases. [0084] Certain single-molecule sequencing embodiments are based on the principal of sequencing by synthesis, and utilize single-pair Fluorescence Resonance Energy Transfer (single pair FRET) as a mechanism by which photons are emitted as a result of successful nucleotide incorporation. The emitted photons often are detected using intensified or high sensitivity cooled charge-couple-devices in conjunction with total internal reflection microscopy (TIRM). Photons are only emitted when the introduced reaction solution contains the correct nucleotide for incorporation into the growing nucleic acid chain that is synthesized as a result of the sequencing process. In FRET based single-molecule sequencing or detection, energy is transferred between two fluorescent dyes, sometimes polymethine cyanine dyes Cy3 and Cy5, through long-range dipole interactions. The donor is excited at its specific excitation wavelength and the excited state energy is transferred, non-radiatively to the acceptor dye, which in turn becomes excited. The acceptor dye eventually returns to the ground state by radiative emission of a photon. The two dyes used in the energy transfer process represent the "single pair", in single pair FRET. Cy3 often is used as the donor fluorophore and often is incorporated as the first labeled nucleotide. Cy5 often is used as the acceptor fluorophore and is used as the nucleotide label for successive nucleotide additions after incorporation of a first Cy3 labeled nucleotide. The fluorophores generally are within 10 nanometers of each other for energy transfer to occur successfully. [0085] An example of a system that can be used based on single-molecule sequencing generally involves hybridizing a primer to a study nucleic acid to generate a complex; associating the complex with a solid phase; iteratively extending the primer by a nucleotide tagged with a fluorescent molecule; and capturing an image of fluorescence resonance energy transfer signals after each iteration (e.g., Braslavsky et al., PNAS 100(7): 3960-3964 (2003); U.S. Pat. No.7,297,518 (Quake et al.) which are incorporated herein by reference in their entirety). Such a system can be used to directly sequence amplification products generated by processes described herein. In some embodiments, the released linear amplification product can be hybridized to a primer that contains sequences complementary to immobilized capture sequences present on a solid support, a bead or glass slide for example. Hybridization of the primer-released linear amplification product complexes with the immobilized capture sequences, immobilizes released linear amplification products to solid supports for single pair FRET based sequencing by synthesis. The primer often is fluorescent, so that an initial reference image of the surface of the slide with immobilized nucleic acids can be generated. The initial reference image is useful for determining locations at which true nucleotide incorporation is occurring. Fluorescence signals detected in array locations not initially identified in the "primer only" reference image are discarded as non-specific fluorescence. Following immobilization of the primer-released linear amplification product complexes, the bound nucleic acids often are sequenced in parallel by the iterative steps of, a) polymerase extension in the presence of one fluorescently labeled nucleotide, b) detection of fluorescence using appropriate microscopy, TIRM for example, c) removal of fluorescent nucleotide, and d) return to step a with a different fluorescently labeled nucleotide. [0086] The technology described herein may be practiced with digital PCR. Digital PCR was developed by Kalinina and colleagues (Kalinina et al., 1997, Nucleic Acids Res. 25; 1999-2004) and further developed by Vogelstein and Kinzler (1999, Proc. Natl. Acad. Sci. U.S.A. 96; 9236- 9241). The application of digital PCR is described by Cantor et al. (PCT Pub. Nos. WO 2005/023091A2 (Cantor et al.); WO 2007/092473 A2, (Quake et al.)), which are hereby incorporated by reference in their entirety. Digital PCR takes advantage of nucleic acid (DNA, cDNA or RNA) amplification on a single molecule level, and offers a highly sensitive method for quantifying low copy number nucleic acid. Fluidigm® Corporation offers systems for the digital analysis of nucleic acids. [0087] In some embodiments, nucleotide sequencing may be by solid phase single nucleotide sequencing methods and processes. Solid phase single nucleotide sequencing methods involve contacting sample nucleic acid and solid support under conditions in which a single molecule of sample nucleic acid hybridizes to a single molecule of a solid support. Such conditions can include providing the solid support molecules and a single molecule of sample nucleic acid in a "microreactor." Such conditions also can include providing a mixture in which the sample nucleic acid molecule can hybridize to solid phase nucleic acid on the solid support. Single nucleotide sequencing methods useful in the embodiments described herein are described in PCT Pub. No. WO 2009/091934 (Cantor). [0088] In certain embodiments, nanopore sequencing detection methods include (a) contacting a nucleic acid for sequencing ("base nucleic acid," e.g., linked probe molecule) with sequence- specific detectors, under conditions in which the detectors specifically hybridize to substantially complementary subsequences of the base nucleic acid; (b) detecting signals from the detectors and (c) determining the sequence of the base nucleic acid according to the signals detected. In certain embodiments, the detectors hybridized to the base nucleic acid are disassociated from the base nucleic acid (e.g., sequentially dissociated) when the detectors interfere with a nanopore structure as the base nucleic acid passes through a pore, and the detectors disassociated from the base sequence are detected. [0089] A detector also may include one or more regions of nucleotides that do not hybridize to the base nucleic acid. In some embodiments, a detector is a molecular beacon. A detector often comprises one or more detectable labels independently selected from those described herein. Each detectable label can be detected by any convenient detection process capable of detecting a signal generated by each label (e.g., magnetic, electric, chemical, optical and the like). For example, a CD camera can be used to detect signals from one or more distinguishable quantum dots linked to a detector. [0090] The invention encompasses methods known in the art for enhancing the sensitivity of the detectable signal in such assays, including, but not limited to, the use of cyclic probe technology (Bakkaoui et al., 1996, BioTechniques 20: 240-8, which is incorporated herein by reference in its entirety); and the use of branched probes (Urdea et al., 1993, Clin. Chem. 39, 725-6; which is incorporated herein by reference in its entirety). The hybridization complexes are detected according to well-known techniques in the art. [0091] Reverse transcribed or amplified nucleic acids may be modified nucleic acids. Modified nucleic acids can include nucleotide analogs, and in certain embodiments include a detectable label and/or a capture agent. Examples of detectable labels include, without limitation, fluorophores, radioisotopes, colorimetric agents, light emitting agents, chemiluminescent agents, light scattering agents, enzymes and the like. Examples of capture agents include, without limitation, an agent from a binding pair selected from antibody/antigen, antibody/antibody, antibody/antibody fragment, antibody/antibody receptor, antibody/protein A or protein G, hapten/anti-hapten, biotin/avidin, biotin/streptavidin, folic acid/folate binding protein, vitamin B12/intrinsic factor, chemical reactive group/complementary chemical reactive group (e.g., sulfhydryl/maleimide, sulfhydryl/haloacetyl derivative, amine/isotriocyanate, amine/succinimidyl ester, and amine/sulfonyl halides) pairs, and the like. Modified nucleic acids having a capture agent can be immobilized to a solid support in certain embodiments. [0092] Next generation sequencing techniques may be applied to measure expression levels or count numbers of transcripts using RNA-seq or whole transcriptome shotgun sequencing. See, e.g., Mortazavi et al. 2008 Nat Meth 5(7) 621-627 or Wang et al. 2009 Nat Rev Genet 10(1) 57-63. Nucleic acids in the invention may be counted using methods known in the art. In one embodiment, NanoString’s nCounter® system may be used (Seattle, WA). Geiss et al. 2008 Nat Biotech 26(3) 317-325; U.S. Pat. No. 7,473,767 (Dimitrov). In addition, NanoString’s Digital Spatial Profiling (DSP) platform may be used for nucleic acid or protein detection. Blank et al., 2018 Nature Medicine 24 1655–1661; Amaria et al., 2018 Nature Medicine 24 1649–1654. Alternatively, Fluidigm’s Dynamic Array system may be used (South San Francisco, CA). Byrne et al. 2009 PLoS ONE 4 e7118; Helzer et al. 2009 Can Res 697860-7866. For reviews, see also Zhao et al. 2011 Sci China Chem 54(8) 1185-1201 and Ozsolak and Milos 2011 Nat Rev Genet 1287-98. 5.4. Compositions and Kits [0093] The invention provides compositions and kits detecting the biomarkers described herein using antibodies or other reagents specific for the nucleic acids specific for the polynucleotides. Kits for carrying out the diagnostic assays of the invention typically include, in suitable container means, (i) a probe that comprises an antibody or nucleic acid sequence that specifically binds to the marker polynucleotides of the invention, (ii) a label for detecting the presence of the probe and (iii) instructions for how to measure the level the polynucleotide. The kits may include several antibodies or polynucleotide sequences encoding biomarkers disclosed herein, e.g., a first antibody and/or second and/or third and/or additional antibodies that recognize the biomarkers or specific nucleic acids. In one embodiment the nucleic acids in the kit are the forward and reverse PCR primers for the biomarkers disclosed herein. The container means of the kits will generally include at least one vial, test tube, flask, bottle, syringe and/or other container into which a first antibody specific for one of the polypeptides or a first nucleic acid specific for one of the polynucleotides of the present invention may be placed and/or suitably aliquoted. Where a second and/or third and/or additional component is provided, the kit will also generally contain a second, third and/or other additional container into which this component may be placed. Alternatively, a container may contain a mixture of more than one antibody or nucleic acid reagent, each reagent specifically binding a different marker in accordance with the present invention. The kits of the present invention will also typically include means for containing the antibody or nucleic acid probes in close confinement for commercial sale. Such containers may include injection and/or blow-molded plastic containers into which the desired vials are retained. [0094] The kits may further comprise positive and negative controls, as well as instructions for the use of kit components contained therein, in accordance with the methods of the present invention. [0095] Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Preferred methods, devices, and materials are described, although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure. All references cited herein are incorporated by reference in their entirety. [0096] The following Examples further illustrate the disclosure and are not intended to limit the scope. In particular, it is to be understood that this disclosure is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present disclosure will be limited only by the appended claims. 6. EXAMPLES [0097] In this study we address the dilemma as to whether DDR gene mutations are a cause or consequence of elevated TMB. We show that, using traditional univariate tests, the vast majority of genes when mutated are associated with elevated TMB, illustrating that these approaches are ill suited to define relationships between gene mutation and TMB. This is explained by the fact that the readout of interest (TMB) is confounded by the variable it is being associated to (mutations in a gene). We illustrate how the “TMB paradox”, refashioned from the well-established friendship paradox in network science, accounts for why the vast majority of genes are associated with an elevated TMB when using univariate testing. Furthermore, we show that representing tumors and their mutated genes as a bipartite network allows for development of a Bipartite Graph-Based Expected TMB Score (BiG-BETS) that more accurately defines DDR genes associated with high or low TMB, termed High and Low BiG-BETS DDR genes respectively. Finally, while we note that having a mutation in a High BiG-BETS DDR gene did not add predictive power to ICB benefit for patients with TMB High tumors, Low BiG-BETS DDR gene mutation in TMB High tumors enriched for patients with elevated STING pathway activity, significantly increased ICB response, and overall survival benefit to ICB. [0098] The majority of genes, when mutated, are associated with elevated TMB by univariate test. [0099] While it is logical that ineffective DNA damage or repair mechanisms would result in increased mutational load, all human genomic studies to date describing an association between the presence of DDR mutations and elevated TMB are correlative. Indeed, an equally plausible explanation is that mutation in a DDR gene (or any gene) is simply more likely in TMB high tumors because more genes are mutated. We therefore sought to understand whether tumors with DDR gene inactivation associate with elevated TMB as a cause or consequence of DDR gene mutation. [00100] We noted that the majority of studies linking DDR gene inactivation to elevated TMB assessed statistical significance using a univariate test (t-test or Mann-Whitney U) comparing mutant tumors to non-mutant tumors (7–9,11,12). This approach is intrinsically biased because the readout of interest (TMB) is confounded with the variable it is being associated to (mutations in a gene). By using a univariate test, tumors with a higher TMB will have a higher likelihood of having mutations in the genes being tested (i.e. DDR pathways). Moreover, aggregating mutations across genes into a pathway further increases the effect of this bias. [00101] To demonstrate this bias, we applied the classic univariate approach (Mann-Whitney U [MWU] test) to the Pan-Cancer TCGA dataset, assessing whether mutation of a gene correlates with elevated TMB. Strikingly, we found that 98% (17,860/18,151) of genes when mutated were associated with elevated TMB relative to their respective non-mutated out-groups, even when corrected for multiple comparisons (Figure 1A), and that DDR genes as a group did not have a significantly higher proportion of genes that when mutated were associated with elevated TMB (Figure 1B). Moreover, when we applied the classic univariate approach (MWU test) to query whether inactivation of any of the DDR pathways (NER, BER, NHEJ, MMR, FA, HR, DS) as defined by the TCGA DDR group (10) were associated with increased TMB, we found that inactivation of each DDR pathway was highly significantly associated with elevated TMB (Supplemental Figure 1A) in keeping with previously published work (7–9,12,13). Furthermore, examination of the distribution of the mean TMB values of tumors with a mutation in each gene, found that the average mean TMB linked to DDR mutations was only slightly higher (24.1 vs 23.26) than the average mean TMB for non-DDR genes (Figure 1C). The fact that this univariate approach rejects the null hypothesis for the vast majority of genes and that the mean TMB for DDR gene mutations is not different than that of mutations in non-DDR genes suggests that the application of univariate statistical testing is invalid. [00102] The friendship paradox accounts for why the majority of genes associate with elevated TMB by univariate testing. [00103] It is counterintuitive that the majority of genes (when mutated) have a higher mean TMB (~24) than the mean TMB of all samples (~5, Figure 1D, vertical dotted line). We have labeled this phenomenon the “TMB paradox” because it reflects the well-established “friendship paradox" found in network analysis. The friendship paradox holds that in a social network, most people have fewer friends than their friends do (14). In other words, for most nodes in a network, their neighbors or “friends” will on average have a higher degree (number of connections) than the node itself. This arises because higher degree nodes count towards the degree in multiple neighboring nodes and thus are oversampled (see Methods: proof of friendship paradox). Analogously, highly mutated samples contribute toward the average TMB for many of the genes in the dataset, resulting in an outsized effect (Supplemental Figure 1B). Since the univariate t-test and Mann-Whitney U tests are testing for differences in central tendencies (i.e. the mean or the median respectively), the TMB Paradox explains why the majority of genes have a highly significant association with elevated TMBs. The TMB paradox is therefore a manifestation of the oversampling bias introduced by highly connected nodes (i.e. High TMB tumors) and this oversampling bias is what makes univariate tests (T-test and MWU) inappropriate to identify which genes are associated with higher levels of TMB. Therefore, the majority of the 98% of genes associated with elevated TMB by the univariate approach (Figure 1A) are likely a result of the TMB paradox rather than the underlying biology. [00104] Representing tumors and their mutated genes as a bipartite network facilitates development of a Bipartite Graph-Based Expected TMB (BIG-BET) Score to define DDR genes associated with high TMB. [00105] We hypothesized that a networks-based approach and a more appropriate, joint statistical test for whether mutation in a given gene is specifically associated with higher levels of TMB would better discriminate which DDR genes are truly associated with an elevated TMB. To this end, we accounted for the TMB paradox by converting the Pan-Cancer TCGA tumors and their respective mutated genes (Moderate+High consequence, see Methods) into a bipartite network, where genes and tumors represent the two classes of nodes and the edges (connecting two nodes) indicate the mutated genes within a tumor (Figure 1E). By re-casting our data into a bipartite network, we see in Figure 1E that the TMB for a sample is equivalent to the number of edges it has (i.e. it’s degree) and that the average TMB associated with mutation in a given gene is the average degree of its neighbors (Figure 1E and Supplemental Figure 1B). [00106] We then leveraged this bipartite network representation to derive a null model of TMB distribution for each DDR gene or pathway. Specifically, random sampling (permutations) of the bipartite network was performed by stochastically rewiring the network while maintaining the degree distribution (number of edges of each node) of the original dataset (this null model is known as the configuration model. See methods) (15,16). Generation of a null model through permutation allowed us to compare the actual mean TMB for tumors with mutation in a given gene or pathway against the expected distribution under random sampling (Figure 1F). [00107] Application of this Bipartite Graph-Based Expected TMB Score (BiG-BETS) to the TCGA Pan-Cancer dataset found that BiG-BETS returned a more uniform distribution of p-values (Supplemental Figure 2A), and only a subset of DDR genes when mutated have a significant association with elevated TMB, (Figure 2A) hereafter referred to as “High BiG-BETS DDR gene” (Figure 2B). It was reassuring to see that MMR genes such as MSH3, MLH1, MSH2 had some of the highest BiG-BET scores (Figure 2B). We noted that several DDR genes such as ATR or CHEK1 which had previously been associated with elevated TMB by others (9) were no longer found to have significant associations with elevated TMB by BiG-BETS (z-scores of -1.90 and -0. 058 respectively) [Supplemental Figure 2B]. A comparison of the number of genes significantly associated with elevated TMB by univariate (MWU) and our novel BiG-BETS permutation test, demonstrates that BiG-BETS has drastically fewer genes that are significantly associated with elevated TMB (Figure 2A). [00108] We next used BiG-BETS to compute DDR pathway level z-scores (Figure 2C). Unlike the univariate MWU approach, which found that all DDR pathways were significantly associated with elevated TMB (Supplemental Figure 1A), BiG-BETS found that only mutations in the mismatch repair (MMR) and nucleotide excision repair (NER) pathways were significantly associated with higher levels of TMB (z score = 2.23 and z score = 3.04 respectively) (Figure 2C). [00109] Interestingly, we found that across all genes, the genes with the lowest BiG-BET z- scores were highly enriched for a number of biological processes (Figure 2D), including the negative regulation of cell proliferation and chromatin remodeling. This finding is congruent with the fact that cancer types with low overall levels of mutational burden are often driven by epigenetic changes and disruption in chromatin architecture (17). Reassuringly, we see independent validation using a real-world clinically sequenced dataset (Samstein et al., see Methods) (18). The BiG-BET permutation test derived z-scores obtained using our approach correlated well between the Pan- TCGA and Samstein datasets (using the 468 genes in common between the two datasets) [Supplemental Figure 2C]. Finally, while the above analysis was performed using non-synonymous mutations defined to be of “Moderate+High_Consequence” (see Methods and Supplemental Figure 3A), we found a high level of correlation between these BiG-BET scores and those derived from a more restrictive “High Consequence” set of mutations (Supplemental Figure3A and 3B). There was also excellent agreement across the DDR genes and pathways (Supplemental Figure 3B). Interestingly, a set of genes had BiG-BET scores that were noticeably lower when using the Moderate+High Consequence categorization (Supplemental Figure 3B). We observed that the vast majority of them were genes with hotspot mutations [defined using a curated list from (19)]. These hotspot mutations (i.e. BRAF V600E, IDH1 R132H, etc.) are included in the Moderate+High_Consequence set but excluded from the High_Consequence set. Further validating our method, past work has shown that some of these hotspot mutations such as EGFR, BRAF V600E, and IDH1 R132H are associated with a lower TMB (20–22). [00110] Mutations in Low BiG-BETS DDR genes add predictive power for ICB response in TMB High tumors. [00111] Given TMB’s correlation with ICB response, it is not surprising that inactivation of some DDR genes or pathways has been associated with clinical benefit from ICB. For example, cancers with loss of the mismatch-repair (MMR) machinery have increased response to pembrolizumab (23–25), and mutations in specific genes such as POLE (26) or BRCA2 (27) have also been shown to correlate with ICB response. However, whether DDR gene inactivation increases predictive power to ICB response over TMB alone is unclear with one retrospective study suggesting increased predictive power of DDR mutation on top of TMB (6), while another larger study from a large prospective phase II trial (IMvigor210) found that DDR mutations did not have predictive power for ICB response beyond its relationship to TMB (7). [00112] To assess whether inactivation of a specific DDR gene adds predictive power to ICB response over TMB alone we first analyzed two large, annotated genomic datasets of ICB treated patients, IMvigor210 (Phase II trial of atezolizumab in urothelial cancer) and Samstein (real-world, pan-tumor, MSKCC) (7,18). There was reasonable overlap in DDR genes between the IMvigor210 and Samstein datasets. IMvigor210 and Samstein contained 19 and 31 DDR genes respectively with 17 DDR genes contained in both datasets (Supplemental Figure 3C). We first validated that TMB correlated with OS (IMvigor210 and Samstein datasets) [Supplemental Figure 4A, 4B] and response (IMvigor210) [Supplemental Figure 4B]. Given current dogma that DDR mutations enhance ICB response through their potential to increase TMB (3,5) we tested the hypothesis that mutations in High BiG-BETS DDR genes (i.e. those associated with high TMB) would correlate with increased ICB benefit merely because of their association with high TMB. In keeping with this notion, mutation in a High BiG-BETS DDR gene did not add any predictive power for response or OS in TMB High tumors (Supplemental Figure 4C, 4D, IMvigor210 CPH p=0.34 and Samstein CPH p=0.11). [00113] In contrast, there was a strong interaction between mutation in a Low BiG-BETS DDR gene and both response and OS with TMB as a covariate. Specifically, we found that in patients with TMB-High tumors, a mutation in a Low BIG-BETS DDR gene was associated with the longer OS (Figure 3A and 3B) and increased response rate to ICB (Figure 3C). This effect on OS is seen within both the IMvigor210 and Samstein datasets (median survival 20.0 vs 10.7 months in IMvigor210 [Figure 3A] and 14.5 vs 10.0 months in Samstein [Figure 3B] for TMB-High tumors with a Low BiG-BETS DDR mutation vs WT), and is reinforced in the clinical response data for IMvigor210 as 67% of TMB-High tumors with a Low BiG-BETS DDR mutation had ICB response vs a 37% response rate for TMB-High tumors that were Low BiG-BETS DDR WT (Figure 3C, p = 0.035). Similar results were seen when the analysis was done with only the subset of DDR genes that overlapped between the Samstein and IMvigor210 datasets (n=17) [Supplemental Figure 3C, 5A]. [00114] Given the importance of this finding, we wish to further validate our BiG-BETS method. We therefore generated a metadataset (hereafter called “Weir_metadataset”) of 424 patients from available cohorts on cBioportal, including melanoma (27–29), non-small-cell lung cancer [NSCLC] (30,31), as well as a recently published real world cohort of bladder cancer patient from our own institution (32). Using the Weir_metadataset, we observed similar findings as we saw in the Samstein and IMvigor210 datasets. Specifically, patients with TMB-high tumors and mutation in a Low BiG-BETS DDR gene had prolonged OS (median OS 19.3 vs 14.5 months, Figure 3D) and increased response (RR 48% vs 29%, Figure 3E) relative to those with TMB-high tumors wild type for a Low BiG-BETS DDR gene. Therefore, mutation in a Low BiG-BETS DDR gene has additional predictive power for ICB benefit over TMB alone. This observation seems to be specific to DDR genes as when we performed a parallel analysis using Chromatin Remodeling genes from the Gene Ontology Resource (GO Term, Chromatin Remodeling, GO:0006338) we found no association between having a mutation in a Low BiG-BETS, chromatin remodeling gene and improved clinical outcome from ICB in High TMB tumors (Supplemental Figure 5B). [00115] Mutation in Low BiG-BETS DDR genes in TMB High tumors is not merely prognostic and is predictive across individual tumor types. [00116] To ensure that mutation in a Low BiG-BETS DDR gene in a TMB High tumor is not merely prognostic, we analyzed the subset of TCGA tumors that overlap with the ICB treated tumor types from the Samstein dataset (see Methods). Since the vast majority of the TCGA tumors were collected and sequenced prior to the widespread use of ICB they should serve as an ICB naïve cohort. To this end, we looked at OS by TMB and BiG-BETS DDR mutation status in the TCGA Samstein overlap cohort (Figure 3F). In this non-ICB treated cohort, there was no differences in survival in TMB High tumors based on Low BiG-BETS DDR gene mutation status (HR = 0.73 [0.46-1.15], CPH p=0.177). Therefore, the BiG-BETS score is not merely prognostic. [00117] We wanted to understand whether the clinical benefit from ICB seen in TMB High, Low BiG-BETS DDR mutant tumors is present within individual tumor types. Of the individual tumor types within the Samstein dataset, most tumor types did not have enough TMB High, Low BiG- BETS DDR mutant samples to see the conditional effect. Only NSCLC and bladder cancer had more than n=5 samples with High-TMB, Low BiG-BETS DDR mutations. We therefore combined patients from the Weir_metadataset with the Samstein and the IMvigor210 cohorts (hereafter called Weir_combined). Analysis of the Weir_combined dataset showed that NSCLC patients with TMB High, Low BiG-BETS DDR mutations had a significantly prolonged OS (CPH p=0.001), while melanoma and bladder cancer while not significant showed a similar trend (Figure 3G, 3I). We also assessed our method on a kidney cancer dataset (33) as well given the frequent treatment of RCC patients with ICB. Interestingly, RCC patients did not appear to show any difference in OS or response on the basis of either TMB or low BiG-BETS DDR mutation status (Supplemental Figure 5C). We hypothesize that this might reflect that ICB response in RCC is because neoantigens are driven by frameshift mutations rather than SNVs (34) or that other tumor associated antigens (like cancer testis antigens or endogenous retroviruses) mediate the responses seen in this unique tumor type. Therefore, we believe that mutations in Low BiG-BETS DDR genes in TMB high tumors correlates with enhanced clinical benefit to ICB across individual tumor types. [00118] Tumors with mutation in Low BiG-BETS DDR genes have enhanced STING activity. [00119] We hypothesized that alterations in these Low BiG-BETS DDR genes may mediate an anti-tumor immune response through enhanced innate immunity or antigen presentation or possibly through production of neo-antigens in a way that is not captured by TMB alone (3,5). Indeed, we saw that in TMB High patients from the IMvigor210 dataset, Low BiG-BETS DDR mutant tumors had elevated gene signatures scores of both STING as well as its key downstream transcriptional mediator, interferon regulatory factor 3 (IRF3) [Figure 4A] but interestingly, did not show significant changes in other gene signatures associated with ICB response (i.e. CD8 T cell, EMT- Stroma, TGFB, or IFNG signatures) [Figure 4B, 4C, 4D]. We noted that TMB High, Low BiG- BETS DDR mutant tumors also had significantly elevated STING gene signature in the TCGA (with cancer types matched to Samstein) [Supplemental Figure 6A, 6B, 6C, 6D], though EMT_stroma and FTRBS signatures were also significantly different in this case. In aggregate these data demonstrate that TMB-High patients with a mutation in a Low BiG-BETS DDR gene have increased response (67%) and prolonged OS when treated with ICB, potentially due to enhanced baseline STING activity. [00120] Discussion [00121] In summary, we identified a sampling bias in currently used methods to test for an association between mutation in a specific gene or pathway and elevated TMB. Application of our novel method, BiG-BETS, accurately defines DDR genes associated with high and low TMB (High BiG-BETS DDR and Low BiG-BETS DDR genes respectively) and identifies that only inactivation of the MMR and NER DDR pathways are truly associated with elevated TMB. Moreover, we demonstrate that mutation in High BiG-BETS DDR genes do not hold predictive value for ICB response in TMB High tumors, because their predictive value is driven by their association with elevated TMB. In contrast, and of clinical importance, is that in TMB High tumors, mutation in a Low BiG-BETS DDR gene is significantly associated with elevated STING pathway activity, increased ICB response (up to 67%), and prolonged overall survival benefit from ICB treatment. [00122] Recent work by Hsiehchen and colleagues also examined correlation between DDR pathways and ICB benefit across a large dataset of tumors with genomic annotation and found that patients whose tumors harbored mutations in the NER and HR pathways were associated with higher RR and OS (35). Our work is distinct, as our primary objective was to faithfully define DDR genes that associate with Low and High TMB status. We have found that patients with TMB High tumors and mutation in Low BiG-BETS DDR gene have the best OS and RR. There are a number of important differences between our study and Hsiehchen and colleagues including the number of DDR genes used (Hsiehchen N=40, our study N=72) as well as the classification of DDR genes into specific DDR pathways. Finally, while the authors suggest that the NER and HR mutations predict OS independent of TMB, we note that all of the NER genes from their paper have High BiG-BETS scores in our analysis. Therefore, by our analysis, mutations of genes in the NER pathway are associated with elevation in TMB. [00123] Our novel BiG-BETS method and clinical observations have potentially important implications for patient care. First, as mentioned above, patients with TMB High, Low BiG-BETS DDR mutant tumors have significantly increased ICB response (67%) and prolonged overall survival and should therefore be assessed in future clinical trials. Moreover, this tandem predictive biomarker can likely be improved upon with integration of other immunogenomic features such as the integration of a T cell inflamed gene expression profile, which has shown predictive value for ICB response in pan tumor analysis (2), and is not correlated with TMB (36). Second, a number of the Low BiG-BETS DDR genes are kinases, multiple of which have small molecule inhibitors in late stage clinical trials (i.e. ATM, CHEK1, WEE1). One would predict that treatment of a TMB High tumor that is Low BiG-BETS DDR wild-type, with one of these kinase inhibitors (in an attempt to replicate a Low BiG-BETS DDR mutation) and ICB might mimic our genetic findings. In contrast, our method predicts that inhibition of High BiG-BETS DDR genes that are kinases (i.e. ATR) in combination with ICB would not necessarily benefit patients with TMB High tumors. [00124] There is much clinical interest in examining the efficacy of ICB in patients with DDR mutations. Moreover, there are numerous clinical trials underway combining DDR inhibitors with ICB. These trials often use broad panels of DDR genes as inclusion criteria or focus on a specific DDR pathway (i.e. homologous recombination). Our data suggests that restricting biomarker selection to a specific DDR pathway is not wise as our Low BiG-BETS DDR genes are relatively evenly spread across all of the DDR pathways (Supplemental Figure 6E). While a unifying explanation for how our Low BiG-BETS DDR genes, which are spread across all DDR pathways, are associated with enhanced STING activity is challenging to envision, we see and validate this finding in both the IMvigor210 and TCGA datasets. Finally, our work is cautionary as it suggests that without selection of both 1) the subset of DDR genes with predictive power for ICB response (Low BiG-BETS DDR genes) as well as 2) patients with TMB High tumors, trials integrating DDR mutations / inhibitors with ICB may show a lack of clinical benefit.
Methods [00125] Definition of High and Moderate+High Consequence mutations [00126] We started with the MC3 mutational dataset provided by the TCGA (https://gdc.cancer.gov/about-data/publications/pancanatlas). The TCGA MC3 dataset was then filtered to include only “High Consequence” non-synonymous mutations, which were defined as being categorized as a high consequence mutation by the Sequence Ontology and summarized by Ensembl here (https://m.ensembl.org/info/genome/variation/prediction/predicted_data.html) ['stop lost', 'stop gained', 'transcript ablation', 'start lost', 'frameshift variant’, ’splice_site’, ‘translation_start_site’] and were also categorized as having a Polyphen score of ‘probably damaging’, ‘possibly damaging’, or ‘unknown’. [00127] We defined non-synonymous mutations of “Moderate+High_Consequence” as moderate [‘inframe insertion’, inframe deletion’, missense variant’, and ‘protein altering variant’] or high ['stop lost', 'stop gained', 'transcript ablation', 'start lost', 'frameshift variant’, ’splice_site, ’translation_start_site’] consequence mutations by the Sequence Ontology and were also categorized as having a Polyphen score of ‘probably damaging’, ‘possibly damaging’, or ‘unknown’ (Supplemental Figure 3A). [00128] Bipartite-Graph Based-Expected TMB Score (BiG-BETS) [00129] We converted the Pan-Cancer TCGA tumors and their respective mutated genes into a bipartite network, where genes and tumors represent the two classes of nodes and the edges (connecting two nodes) indicate the mutated genes within a given tumor (Figure 1E). By re-casting our data into a bipartite network, we see in Figure 1E that the TMB for a sample is proportional to the number of edges it has (i.e. its degree) and that the average TMB associated with mutation in a given gene is essentially average degree of its neighbors (Supplemental Figure 1B). [00130] The bipartite network representation was used to derive a null model of TMB distribution for each gene. Specifically, random sampling (permutations) of the bipartite network was performed by stochastically rewiring the network while maintaining the degree distribution (number of edges of each node) of the original dataset (this null model is known as the configuration model, which for a bipartite network is further constrained to maintain the bipartite nature of the network) (15,16). Generation of a null model through permutation allowed us to compare the actual mean TMB for tumors with mutation in a given gene or pathway against the expected distribution under random sampling (Figure 1F). [00131] The BiG-BET score consists of comparing the observed mean TMB for each gene or pathway's mutated sample set against the expected distribution under random sampling of bipartite networks that match the degree distribution of the original dataset. The null model for networks in which all networks with a given degree sequence are uniformly likely is known as the configuration model (24), which has also been extended to bipartite graphs (15). The bipartite configuration model can be envisioned by cutting across the edges in the original network and reconnecting the “stubs" at random with each possible set of pairings respecting the bipartite structure of the original network and being equally likely under the model (visualized in Figure 1F). We note that we have applied a rewiring procedure as described in (15) to sample this bipartite configuration model rather than the more direct “stub matching" approach to ensure that we sample the appropriate, more restricted model without self-loops and multi-edges (3-5). We iterate the following steps: [00132] Select two edges at random in the network:
Figure imgf000045_0001
and
Figure imgf000045_0002
[00133] Confirm that each edge is connected to a distinct pair of nodes:
Figure imgf000045_0003
and
Figure imgf000045_0004
. If the two edges involve either the same sample node, or the same gene node, repeat 1. [00134] Swap which gene is connected to which sample from these two edges. Add
Figure imgf000045_0005
and to the set of edges while removing original edges.
Figure imgf000045_0006
[00135] Repeat 1-3 [00136] Steps 1-3 constitute a single rewire of the network. To obtain a BiG-BET score for each individual gene, we rewire the network many times in sequence, keeping track of the edge changes so that the network becomes unrecognizable from the original network and the previous sample. This process is a Markov chain, generating a random network at each step conditioned only on its immediate predecessor that is independent of earlier networks. If run long enough, the process will generate all possible networks from the model with uniform probability. Prior to drawing samples from the Markov chain, we conduct
Figure imgf000045_0012
``burn-in" rewires, where is the number of edges in the network, to give the process freedom to sample high likelihood regions under the model independent of the initialization at our observed data. We perform at least
Figure imgf000045_0013
rewires between samples to ensure that most edges in the network will have the opportunity to rewire. [00137] We define the BiG-BET score as follows; Let the gene
Figure imgf000045_0007
have degree
Figure imgf000045_0008
. Let
Figure imgf000045_0009
denote the set of samples connected to
Figure imgf000045_0010
in the bipartite representation of the original data, (that is the neighbors of ). Let
Figure imgf000045_0011
Figure imgf000046_0001
be the average TMB for all samples connected to
Figure imgf000046_0002
in the observed dataset. We derive a z-score for a significant association between TMB and
Figure imgf000046_0008
as follows: [00138] We sample
Figure imgf000046_0003
independent realizations of the bipartite network with fixed degrees sequences, according to the process described above. [00139] For each sampled network,
Figure imgf000046_0004
, we compute
Figure imgf000046_0005
, the average TMB for all neighbors of
Figure imgf000046_0007
in each sampled bipartite network
Figure imgf000046_0006
. [00140] We compute the z-score for the observed graph using: [00141]
Figure imgf000046_0010
[00142] where
Figure imgf000046_0011
is the empirical standard deviation for the distribution of sampled
Figure imgf000046_0009
. All results in the paper have been derived using R=400 samples from the bipartite configuration model. All BiG-BET scores for genes in the TCGA cohort are listed in Supplemental Table 1. [00143] Survival and Response Analyses [00144] All survival analyses were conducted using the Cox-proportional hazard model with a log-likelihood ratio test (LLR) to test for the overall significance of the model and a t-test to assess the significance of individual variables in the model. For the depiction of Kaplan-Meier curves, TMB is treated as a binary variable with a threshold of TMB>10 defining the high TMB group. However, in the joint models depicted by the forest plots, TMB is treated as a continuous variable. Comparison of response across groups is conducted using a Chi-squared test. Samples were divided into groups based on the presence of a mutation within a High DDR z-score gene or Low z-score DDR gene, shown in Figure 2B. [00145] The Low BIG-BET z-score DDR genes included: [00146] APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, XRCC5. [00147] The High BIG-BET z-score DDR genes were: [00148] BARD1, CHEK1, CUL5, EME1, ERCC4, ERCC5, ERCC6, FANCA, FANCC, FANCI, FEN1, GEN1, LIG4, MDC1, MLH1, MLH3, MRE11, MSH2, MSH3, MSH6, NBN, PALB2, PARP1, PMS1, PMS2, POLE, POLE3, POLL, POLM, RAD50, RNMT, SEM1, SLX1A, TDG, TOP3A, TOPBP1, UNG, XPC, XRCC2, XRCC4, XRCC6. [00149] Tumors with mutations in both a High and a Low DDR gene were considered in the High z-score category and excluded from the Low z-score category. [00150] Gene Expression Signatures Analysis [00151] To calculate gene signatures profiles for each dataset, RNAseq expression data was obtained and filtered to the corresponding samples with mutational data. We log(1+x) transformed the data and used a robust scaling (median centered and scaled by inter-quartile range) to normalize across samples. For each signature, we calculate the average expression of all genes within the signature and then assign each sample a z-score of the basis of its expression relative to the entire cohort. Signatures used in analysis are given in below. GRANDVAUX_IRF3_TARGETS_UP ARG2, B4GALT5, F13B, GBP1, IFI44, IFIT1, IFIT3, ISG15, LILRB1, NR3C1, OAS2, PLCG2 PMAIP1, RSAD2 13_T_Cell_EntrezID ACTN1, ACVR2B, ADA, AKTIP, ANXA1, AOC1, APBA2, APBB1, AQP3, ARL4C, ATP13A4, ATP1A1, BCL11B, BIN2, BUB1B, C15orf62, C9orf164, CAMK4, CAPZB, CCL5, CCND2, CD2, CD247, CD28, CD3D, CD3E, CD3G, CD5, CD6, CDC14A, CEP41, CEP85L, CISH, CTSW, DISC1, DNAJB1, DNASE1L3, DOCK9, DPP4, DUSP16, DUSP2, FAM102A, FAM134B, FBLN5, FHIT, FLT3LG, FYB, FYN, GABARAPL1, GALT, GATA3, GBP1, GBP2, GIPC1, GPSM3, GZMK, HOXB2, HSPA1L, ID2, IFITM1, IL18R1, IL32, IL6ST, IL7R, INPP4B, ITGA6, ITK, ITM2A, ITPKB, JAKMIP1, KLRB1, KLRC4, KLRG1, LAT, LCK, LCP2, LDHB, LEF1, LEPROTL1, LIMA1, LINS, LOC283666, LOC642808, LOC647217, LPAR2, LPAR6, LPIN2, LRIG1, LYAR, MAL, MAN1C1, MAP7D1, MAPKAPK5, MAST4, MATN2, MEN1, MGAT4A, MLLT3, MORC2-AS1, MPP7, MRPL27, MYBL1, NAP1L5, NELL2, NGFRAP1, NOL4L, NPDC1, NPTXR, NR4A2, NSG1, OLAH, OSBP2, PCSK5, PCYT2, PDE4D, PDE9A, PIK3IP1, PIK3R1, PKM, PRKCA, PRKCI, PRKCQ, PRKCQ-AS1, PRSS1, PTGER2, PXN, RAB43, RARRES3, RASGRP1, RBMS1, RGS10, RMDN1, RNF213, RORA, RSU1, RTKN2, RUNX2, S100A10, S100A8, SATB1, SBK1, SELPLG, SH2D1A, SHFM1, SLC35D2, SLC39A8, SLCO3A1, SLFN5, SNPH, SOCS3, SORL1, SPEG, SPOCK2, STAT4, SYNE2, SYT1, TACC3, TARP, TCF7, TESPA1, TIAM1, TMEM173, TNFAIP3, TNFRSF25, TNFSF8, TNIK, TOB1, TOMM40, TRA, TRABD2A, TRAT1, TSEN54, TXK, UPP1, VIPR1, WNT10B, WWP1, ZAP70 18_Iglesia_CD8_cluster_EntrezID, ACAP1, BTLA, C16orf54, CCL5, CD2, CD247, CD27, CD3D, CD3E, CD3G, CD48, CD5, CD6, CD8A, CD96, CXCR3, CXCR6, GPR171, GZMA, GZMK, IL2RG, ITK, KLRK1, LCK, LY9, NKG7, PRF1, PRKCB, PTPN7, PTPRCAP, PYHIN1, S1PR4, SAMD3, SCML4, SH2D1A, SIRPG, SIT1, SLA2, SLAMF1, SLAMF6, TBC1D10C, TBX21, TIGIT, TRAT1, UBASH3A, ZAP70, ZNF831, 25_Bindea_CD8_T_cells_EntrezID, ABT1, AES, APBA2, ARHGAP8, CAMLG, CD8A, CD8B, CDKN2AIP, DNAJB1, FLT3LG, GADD45A, GZMM, HAUS3, KAT6A, KLF9, LEPROTL1, LIME1, MAPKAPK5-AS1, PF4, PPP1R2, PRF1, PRR5, RBM3, SF1, SLC16A7, SRSF7, TBCC, THUMPD1, TMC6, TMEM259, TSC22D3, VAMP2, ZEB1, ZFP36L2, ZNF22, ZNF609, ZNF91 EMT_stroma_core_18, FLNA, CTHRC1, COL6A3, COL6A2, MFAP5, EMP3, CALD1, FN1, FOXC2, LOX, FBN1, TNC, LRP1, SERPINE2, ECM1, LAMA2, VIM, CALU, FTBRS, ACTA2, ACTG2, ADAM12, ADAM19, CNN1, COL4A1, CTGF, CTPS1, FAM101B, FSTL3, HSPB1, IGFBP3, PXDC1, SEMA7A, SH3PXD2A , TAGLN, TGFBI, TNS1 , TPM1 50_HALLMARK_INFN_GAMMA_RESPONSE, ADAR, APOL6, ARID5B, ARL4A, AUTS2, B2M, BANK1, BATF2, BPGM, BST2, BTG1, C1R, C1S, CASP1, CASP3, CASP4, CASP7, CASP8, CCL2, CCL5, CCL7, CD274, CD38, CD40, CD69, CD74, CD86, CDKN1A, CFB, CFH, CIITA, CMKLR1, CMPK2, CSF2RB, CXCL10, CXCL11, CXCL9, DDX58, DDX60, DHX58, EIF2AK2, EIF4E3, EPSTI1, FAS, FCGR1A, FGL2, FPR1, FTSJD2, GBP4, GBP6, GCH1, GPR18, GZMA, HERC6, HIF1A, HLA-A, HLA-B, HLA-DMA, HLA-DQA1, HLA-DRB1, HLA-G, ICAM1, IDO1, IFI27, IFI30, IFI35, IFI44, IFI44L, IFIH1, IFIT1, IFIT2, IFIT3, IFITM2, IFITM3, IFNAR2, IL10RA, IL15, IL15RA, IL18BP, IL2RB, IL4R, IL6, IL7, IRF1, IRF2, IRF4, IRF5, IRF7, IRF8, IRF9, ISG15, ISG20, ISOC1, ITGB7, JAK2, KLRK1, LAP3, LATS2, LCP2, LGALS3BP, LY6E, LYSMD2, 1-Mar, METTL7B, MT2A, MTHFD2, MVP, MX1, MX2, MYD88, NAMPT, NCOA3, NFKB1, NFKBIA, NLRC5, NMI, NOD1, NUP93, OAS2, OAS3, OASL, OGFR, P2RY14, PARP12, PARP14, PDE4B, PELI1, PFKP, PIM1, PLA2G4A, PLSCR1, PML, PNP, PNPT1, PRIC285, PSMA2, PSMA3, PSMB10, PSMB2, PSMB8, PSMB9, PSME1, PSME2, PTGS2, PTPN1, PTPN2, PTPN6, RAPGEF6, RBCK1, RIPK1, RIPK2, RNF213, RNF31, RSAD2, RTP4, SAMD9L, SAMHD1, SECTM1, SELP, SERPING1, SLAMF7, SLC25A28, SOCS1, SOCS3, SOD2, SP110, SPPL2A, SRI, SSPN, ST3GAL5, ST8SIA4, STAT1, STAT2, STAT3, STAT4, TAP1, TAPBP, TDRD7, TNFAIP2, TNFAIP3, TNFAIP6, TNFSF10, TOR1B, TRAFD1, TRIM14, TRIM21, TRIM25, TRIM26, TXNIP, UBE2L6, UPP1, USP18, VAMP5, VAMP8, VCAM1, WARS, XAF1, XCL1, ZBP1, ZNFX1 Ayers_IFNG, HLA-DRA, IFNG, IDO1, CXCL10, CXCL9, STAT1 Higgs_IFNG, IFNG, LAG3, CXCL9, PD-L1 [00152] Description of the Datasets [00153] The Cancer Genome Atlas (TCGA) Pan-Cancer MC3 Dataset [00154] The primary dataset we used for developing our method was the TCGA-pancan unified ensemble MC3 variant call set (downloadable at https://gdc.cancer.gov/about- data/publications/pancanatlas). See https://www.synapse.org/#!Synapse:syn7214402/wiki/405297 for further description.). This dataset used Whole Exome Sequence (WES) tumor samples from all TCGA centers and variants were re-called in a uniform pipeline. This dataset includes 3.6 million, small variants from 10,295 tumor samples. This dataset is described further in (37). Filtering this list of variants on Moderate+High consequence (defined above) resulted in 1,000,011 variants in 19,255 genes in 10,164 different samples, while keeping High consequence variants resulted in 208,682 remaining variants in 18,284 different genes from 9,530 different samples. Supplemental Figure 3B demonstrates excellent correlation in the BiG-BET score between the high impact and the high+moderate impact datasets, especially with regards to the DDR genes and pathways. [00155] TMB values for TCGA were obtained from (https://gdc.cancer.gov/about- data/publications/PanCan-CellOfOrigin) combining both Silent and Non-Silent scores for each sample. [00156] Clinical data used for survival analysis of the TCGA were obtained from the TCGA clinical data resource as detailed in Liu et al 2018. We looked at the effect of low and high BiG- BET DDR mutations on overall 5-year survival as detailed above. We filtered the cohort to the cancer types that best reflected the composition of Samstein et al, keeping the following TCGA types: LUAD, LUSC, BRCA, SKCM, COAD, ESCA, KIRC, BLCA, and HNSC. This resulted in 4,287 samples. To calculate gene signatures profiles, the TCGA-PanCan expression data was obtained from https://gdc.cancer.gov/about-data/publications/pancanatlas and processed as described above in Gene Expression Signatures Analysis. [00157] IMvigor210 [00158] The IMvigor210 trial is a Phase II single arm study examining the response of patients with locally advanced or metastatic urothelial bladder cancer to atezoliziumab (anti PD-L1). A full description of the characteristics of the patient cohort can be found in (7). We have used the publicly available dataset released by Mariathansan et al. which can be accessed at http://research- pub.gene.com/IMvigor210CoreBiologies/. The cohort consists of 260 patients with 1249 short variants across 160 different genes. Because less detailed annotations were available, we did not filter any of the mutations from this cohort. [00159] Samstein et al. Cohort [00160] To validate our clinical findings, we used a large, multi-trial cohort consisting of 1661 patients treated with different Immune Checkpoint Blockade (ICB) therapies and with targeted clinical sequencing, first compiled and analyzed in (18). Sequencing was performed using the Memorial Sloan Kettering Integrated Mutation Profiling of Actionable Cancer Targets (MSK- IMPACT) panel. We downloaded the data from cbioportal using the following link: https://www.cbioportal.org/study/summary?id=tmb_mskcc_2018. We filtered down to the 1307 patients that received anti PD-1/PD-L1 therapies and kept the variants with one of the following high impact consequences: missense mutation, nonsense mutation, frameshift deletion, frameshift insertion, translation start site, or nonstop mutation. This resulted in a total of 19,057 variants in 468 different genes. [00161] Weir meta-dataset [00162] As an additional validation set, we also compiled a meta-dataset from several different studies available on cBioportal as well as a recently published study from our group. This meta- dataset included melanoma (27–29), non-small cell lung cancer (30,31) and clear cell renal carcinoma (38). The datasets from cbioportal can be accessed using the following link: https://www.cbioportal.org/study/summary?id=ccrcc_dfci_2019,skcm_mskcc_2014,skcm_dfci_2 015,mel_ucla_2016,nsclc_mskcc_2018,nsclc_mskcc_2015. Additionally, we included a real- world metastatic urothelial carcinoma cohort from UNC (32). The Weir metadataset included 407 total samples with 16,250 variants in 599 genes. As most of the datasets on cBioportal did not include TMB values, we used the total number of mutations for each sample as a surrogate. For our analyses, we considered a tumor to be TMB-H if it was in the top 50% of tumors by total mutation count. Data is included in Supplemental Table 2. [00163] DNA Damage Repair Pathways [00164] We have relied on the core DNA Damage Repair pathways defined by Knijnenburg et al. (10) to conduct all of our pathway level analysis. The pathways are defined as follows: [00165] Table 2 DNA Damage Repair Pathways
Figure imgf000051_0001
[00166] Chromatin Remodeling Pathway [00167] We also looked at genes associated with chromatin remodeling as annotated by the Gene Ontology (GO) project (http://geneontology.org/). We selected all genes associated with the GO:0006338 – ‘chromatin_remodeling’ or any of its associated sub-term, resulting in a set of 267 genes given in Table 3. Table 3. Chromatin Remodeling genes GO:0006338
1
Figure imgf000052_0001
Figure imgf000053_0001
[00168] Proof of the friendship paradox for bipartite network [00169] We can represent a network of
Figure imgf000054_0004
nodes and
Figure imgf000054_0005
edges with an
Figure imgf000054_0006
adjacency matrix, , where the entries of
Figure imgf000054_0019
are defined as follows [00170]
Figure imgf000054_0001
[00171] where we use
Figure imgf000054_0007
to denote the set of edges present in the graph, indexed by the pair of nodes connected by each edge. For bipartite networks, we denote
Figure imgf000054_0009
to be the number of nodes of class 1 and likewise
Figure imgf000054_0008
the number in class 2, with
Figure imgf000054_0010
. In a bipartite network, each edge connects a node from class 1 with one from class 2. The degree
Figure imgf000054_0011
of node
Figure imgf000054_0012
is given by the number of edges connected to that node:
Figure imgf000054_0013
. We let the degree distribution
Figure imgf000054_0015
give the fraction of nodes with degree
Figure imgf000054_0014
, representing the probability that a randomly chosen node will have that degree. We denote the class specific degree distributions as
Figure imgf000054_0016
and
Figure imgf000054_0017
to represent the fraction of nodes within each class with a given degree. The overall degree distribution,
Figure imgf000054_0018
, and the class specific degree distributions are related by [00172]
Figure imgf000054_0002
. [00173] In our gene-sample network, we are interested in the average degree across all samples with a mutation in a given gene. We show that this value, the average neighbor-of-a-gene degree, is typically greater than or equal to the average degree of the sample nodes in the network, following a proof similar to that for unipartite networks in 35. [00174] We begin by computing the probability that, after following a randomly chosen edge in our bipartite network, we arrive at a node of a given class with degree
Figure imgf000054_0020
. Without loss of generality we assume class 1 is our class of interest (the sample nodes). There are m edges connected to nodes of class 1, so the probability of ending at a particular node with degree
Figure imgf000054_0021
is
Figure imgf000054_0022
. Since there are such nodes with degree
Figure imgf000054_0023
, the probability of following an edge to a class 1 node of degree
Figure imgf000054_0024
is [00175]
Figure imgf000054_0003
[00176] where
Figure imgf000055_0001
gives the average degree for nodes of class 1. That is, the average neighbor degree distribution is weighted by a factor of
Figure imgf000055_0005
. We are more likely to choose a higher degree vertex by virtue of the simple fact that it has more edges connected to it. We can then compute the average neighbor degree by [00177]
Figure imgf000055_0002
[00178] We can compute the difference between the average neighbor degree and the average degree, restricted to nodes in class 1: [00179]
Figure imgf000055_0003
[00180] where
Figure imgf000055_0004
is the variance of the degree distribution restricted to nodes of class 1. This is strictly non-negative and is zero only in the case where all nodes of class 1 have the same degree. That is, except for the case where all nodes in the class have the same degree, the average neighbor degree is greater than the average degree. Furthermore, we see that this difference is proportional to the variance of the degree distribution of class 1, meaning that heavier-tailed class- restricted degree distributions give even bigger differences, and thus might be even more likely to be misanalysed by a univariate statistic that inadvertently mixes the roles of these two averages. [00181] In striving for transparency and replicability, we have made a repository of our code and analyses available at https://github.com/wweir827/BIGBETS. 7. REFERENCES 1. Tran G, Zafar SY. Financial toxicity and implications for cancer care in the era of molecular and immune therapies. Ann Transl Medicine.2018;6:166–166. 2. Cristescu R, Mogg R, Ayers M, Albright A, Murphy E, Yearley J, et al. Pan-tumor genomic biomarkers for PD-1 checkpoint blockade-based immunotherapy. Science.2018;362:eaar3593. 3. Keenan TE, Burke KP, Allen EMV. Genomic correlates of response to immune checkpoint blockade. Nature Medicine. 2019;25:389–402. 4. Alexandrov LB, Nik-Zainal S, Wedge DC, Aparicio SAJR, Behjati S, Biankin AV, et al. Signatures of mutational processes in human cancer. Nature.2013;500:415–21. 5. Mouw KW, Goldberg MS, Konstantinopoulos PA, D’Andrea AD. DNA Damage and Repair Biomarkers of Immunotherapy Response. Cancer Discovery [Internet].2017;7:675–93. Available from: http://cancerdiscovery.aacrjournals.org/content/7/7/675.article-info 6. Teo MY, Seier K, Ostrovnaya I, Regazzi AM, Kania BE, Moran MM, et al. Alterations in DNA Damage Response and Repair Genes as Potential Marker of Clinical Benefit From PD-1/PD-L1 Blockade in Advanced Urothelial Cancers. Journal of Clinical Oncology [Internet]. 2018;JCO.2017.75.774. Available from: http://ascopubs.org/doi/abs/10.1200/JCO.2017.75.7740#affiliationsContainer 7. Mariathasan S, Turley SJ, Nickles D, Castiglioni A, Yuen K, Wang Y, et al. TGFβ attenuates tumour response to PD-L1 blockade by contributing to exclusion of T cells. Nature. 2018;554:544–8. 8. Wang J, wang zhijie, Zhao J, Wang G, Zhang F, Zhang Z, et al. Co-mutations in DNA damage response pathways serve as potential biomarkers for immune checkpoint blockade. Cancer Res. 2018;canres.1814.2018. 9. Parikh AR, He Y, Hong TS, Corcoran RB, Clark JW, Ryan DP, et al. Analysis of DNA Damage Response Gene Alterations and Tumor Mutational Burden Across 17,486 Tubular Gastrointestinal Carcinomas: Implications for Therapy. Oncol.2019;24:1340–7. 10. Knijnenburg TA, Wang L, Zimmermann MT, Chambwe N, Gao GF, Cherniack AD, et al. Genomic and Molecular Landscape of DNA Damage Repair Deficiency across The Cancer Genome Atlas. Cell reports.2018;23:239-254.e6. 11. Chae YK, Anker JF, Carneiro BA, Chandra S, Kaplan J, Kalyan A, et al. Genomic landscape of DNA repair genes in cancer. Oncotarget.2016;7:23312–21. 12. Chalmers ZR, Connelly CF, Fabrizio D, Gay L, Ali SM, Ennis R, et al. Analysis of 100,000 human cancer genomes reveals the landscape of tumor mutational burden. Genome medicine [Internet]. 2017;9:480. Available from: https://genomemedicine.biomedcentral.com/articles/10.1186/s13073-017-0424-2 13. Chae YK, Anker JF, Bais P, Namburi S, Giles FJ, Chuang JH. Mutations in DNA repair genes are associated with increased neo-antigen load and activated T cell infiltration in lung adenocarcinoma. Oncotarget.2017;9:7949–60. 14. Feld SL. Why Your Friends Have More Friends Than You Do. JSTOR [Internet]. 1991;96:1464–77. Available from: http://www.jstor.org/stable/2781907 15. Fosdick BK, Larremore DB, Nishimura J, Ugander J. Configuring Random Graph Models with Fixed Degree Sequences. Siam Rev. 2018;60:315–55. 16. Saracco F, Clemente RD, Gabrielli A, Squartini T. Randomizing bipartite networks: the case of the World Trade Web. Sci Rep-uk.2015;5:10595. 17. Gonzalez-Perez A, Jene-Sanz A, Lopez-Bigas N. The mutational landscape of chromatin regulatory factors across 4,623 tumor samples. Genome Biol.2013;14:r106. 18. Samstein RM, Lee C-H, Shoushtari AN, Hellmann MD, Shen R, Janjigian YY, et al. Tumor mutational load predicts survival after immunotherapy across multiple cancer types. Nature Genetics.2019;51:202–6. 19. Miao D, Margolis CA, Vokes NI, Liu D, Taylor-Weiner A, Wankowicz SM, et al. Genomic correlates of response to immune checkpoint blockade in microsatellite-stable solid tumors. Nat Genet. 2018;50:1271–81. 20. Battaglin F, Xiu J, Baca Y, Shields AF, Goldberg RM, Seeber A, et al. Comprehensive molecular profiling of IDH1/2 mutant biliary cancers (BC). J Clin Oncol.2020;38:479–479. 21. Mar VJ, Wong SQ, Li J, Scolyer RA, McLean C, Papenfuss AT, et al. BRAF/NRAS Wild- Type Melanomas Have a High Mutation Load Correlating with Histologic and Molecular Signatures of UV Damage. Clin Cancer Res.2013;19:4589–98. 22. Negrao MV, Skoulidis F, Montesion M, Schulze K, Bara I, Shen V, et al. Oncogene-specific differences in tumor mutational burden, PD-L1 expression, and outcomes from immunotherapy in non-small cell lung cancer. J Immunother Cancer.2021;9:e002891. 23. Lemery S, Keegan P, Pazdur R. First FDA Approval Agnostic of Cancer Site — When a Biomarker Defines the Indication. New Engl J Medicine.2017;377:1409–12. 24. Le DT, Uram JN, Wang H, Bartlett BR, Kemberling H, Eyring AD, et al. PD-1 Blockade in Tumors with Mismatch-Repair Deficiency. New England Journal of Medicine.2015;372:2509– 20. 25. Le DT, Durham JN, Smith KN, Wang H, Bartlett BR, Aulakh LK, et al. Mismatch-repair deficiency predicts response of solid tumors to PD-1 blockade. Science [Internet].2017;eaan6733. Available from: http://science.sciencemag.org/content/early/2017/06/07/science.aan6733?utm_campaign=fr_sci_2 017-06-08&et_rid=35386688&et_cid=1373712 26. Mehnert JM, Panda A, Zhong H, Hirshfield K, Damare S, Lane K, et al. Immune activation and response to pembrolizumab in POLE-mutant endometrial cancer. J Clin Invest. 2016;126:2334–40. 27. Hugo W, Zaretsky JM, Sun L, Song C, Moreno BH, Hu-Lieskovan S, et al. Genomic and Transcriptomic Features of Response to Anti-PD-1 Therapy in Metastatic Melanoma. Cell. 2016;165:35–44. 28. Snyder A, Makarov V, Merghoub T, Yuan J, Zaretsky JM, Desrichard A, et al. Genetic basis for clinical response to CTLA-4 blockade in melanoma. New England Journal of Medicine. 2014;371:2189–99. 29. Allen EMV, Miao D, Schilling B, Shukla SA, Blank C, Zimmer L, et al. Genomic correlates of response to CTLA-4 blockade in metastatic melanoma. Science.2015;350:207–11. 30. Rizvi NA, Hellmann MD, Snyder A, Kvistborg P, Makarov V, Havel JJ, et al. Cancer immunology. Mutational landscape determines sensitivity to PD-1 blockade in non-small cell lung cancer. Science.2015;348:124–8. 31. Hellmann MD, Nathanson T, Rizvi H, Creelan BC, Sanchez-Vega F, Ahuja A, et al. Genomic Features of Response to Combination Immunotherapy in Patients with Advanced Non-Small-Cell Lung Cancer. Cancer Cell.2018;33:843-852.e4. 32. Rose TL, Weir WH, Mayhew GM, Shibata Y, Eulitt P, Uronis JM, et al. Fibroblast growth factor receptor 3 alterations and response to immune checkpoint inhibition in metastatic urothelial cancer: a real world experience. Brit J Cancer.2021;1–10. 33. Braun DA, Hou Y, Bakouny Z, Ficial M, Angelo MS, Forman J, et al. Interplay of somatic alterations and immune infiltration modulates response to PD-1 blockade in advanced clear cell renal cell carcinoma. Nat Med.2020;26:909–18. 34. Turajlic S, Litchfield K, Xu H, Rosenthal R, McGranahan N, Reading JL, et al. Insertion-and- deletion-derived tumour-specific neoantigens and the immunogenic phenotype: a pan-cancer analysis. The lancet oncology.2017;18:1009–21. 35. Hsiehchen D, Hsieh A, Samstein RM, Lu T, Beg MS, Gerber DE, et al. DNA Repair Gene Mutations as Predictors of Immune Checkpoint Inhibitor Response beyond Tumor Mutation Burden. Cell Reports Medicine. 2020;1:100034. 36. Spranger S, Luke JJ, Bao R, Zha Y, Hernandez KM, Li Y, et al. Density of immunogenic antigens does not explain the presence or absence of the T-cell-inflamed tumor microenvironment in melanoma. Proceedings of the National Academy of Sciences.2016;201609376. 37. Ellrott K, Bailey MH, Saksena G, Covington KR, Kandoth C, Stewart C, et al. Scalable Open Science Approach for Mutation Calling of Tumor Exomes Using Multiple Genomic Pipelines. Cell Syst.2018;6:271-281.e7. 38. Miao D, Margolis CA, Gao W, Voss MH, Li W, Martini DJ, et al. Genomic correlates of response to immune checkpoint therapies in clear cell renal cell carcinoma. Science [Internet]. 2018;eaan5951. Available from: http://science.sciencemag.org/content/early/2018/01/03/science.aan5951?utm_campaign=fr_sci_2 018-01-04&et_rid=35386688&et_cid=1771523 8. GENERALIZED STATEMENTS OF THE DISCLOSURE [00182] The following numbered statements provide a general description of the disclosure and are not intended to limit the appended claims. [00183] Statement 1: A method for predicting response to immune checkpoint blockade (ICB) therapy for a subject with cancer, comprising independently measuring or obtaining (i) a tumor mutational burden (TMB) level and (ii) defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5 in a sample from the subject, normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample, wherein subjects with both a high TMB level and defects in at least ten of the nucleic acids encoding DDR genes, or their expression products, is indicative of an increased response to ICB therapy for the subject with cancer. [00184] Statement 2: A method for predicting overall survival for a subject with cancer, comprising independently measuring or obtaining (i) a tumor mutational burden (TMB) level and (ii) defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5 in a sample from the subject, normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample, wherein subjects with both a high TMB level and defects in at least ten of the nucleic acids encoding DDR genes, or their expression products, is indicative of an increased overall survival for the subject with cancer. [00185] Statement 3: A method to select a subject with cancer for immune blockade therapy (ICB) which comprises independently measuring or obtaining both (i) a tumor mutational burden (TMB) level and (ii) defects in DNA Damage Repair (DDR) genes or their expression products in a sample from the subject; calculating a Bipartite Graph-Based Expected TMB Score (BiG-BETS) to determine which DDR genes or expression products are associated with elevated TMB; if the subject has a high BiG-BETS determining that the defect in the DDR genes or their expression products do not add to the TMB level over measurement of TMB levels alone; and using the TMB levels alone to select a subject for ICB. [00186] Statement 4: The method of Statement 3, wherein the DDR genes or their expression products are at least ten biomarkers selected from the group consisting of BARD1, CHEK1, CUL5, EME1, ERCC4, ERCC5, ERCC6, FANCA, FANCC, FANCI, FEN1, GEN1, LIG4, MDC1, MLH1, MLH3, MRE11, MSH2, MSH3, MSH6, NBN, PALB2, PARP1, PMS1, PMS2, POLE, POLE3, POLL, POLM, RAD50, RNMT, SEM1, SLX1A, TDG, TOP3A, TOPBP1, UNG, XPC, XRCC2, XRCC4, and XRCC6. [00187] Statement 5: A method to select a subject with cancer for immune blockade therapy (ICB) which comprises independently measuring or obtaining both (i) a tumor mutational burden (TMB) level and (ii) defects in DNA Damage Repair (DDR) genes or their expression products in a sample from the subject; calculating a Bipartite Graph-Based Expected TMB Score (BiG-BETS) to determine which DDR genes or expression products are associated with elevated TMB; if the subject has a low BiG-BETS determining that the defects in the DDR genes or their expression products contribute to the likelihood of benefit from ICB over measurement of TMB levels alone; and using both the TMB levels and the defects in the DDR genes or their expression products to select a subject for ICB. [00188] Statement 6: The method of Statement 5, wherein the DDR genes or their expression products are at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5. [00189] Statement 7: The method of any of Statements 1-6, wherein the cancer is bladder urothelial carcinoma, colon adenocarcinoma, esophageal carcinoma, invasive breast carcinoma, head and neck squamous cell carcinoma, kidney renal clear cell carcinoma, lung adenocarcinoma, melanoma, or squamous cell lung carcinoma. [00190] Statement 8: The method of any of Statements 1-6,, wherein the defects are mutations or copy number alterations. [00191] Statement 9: The method of Statement 8, wherein the mutations are deletions, frameshift mutations, insertions, missense mutations, nonsense mutations, start codon loss, stop codon loss or gain, or a combination thereof. [00192] Statement 10: The method of any of Statements 1-6, wherein the detecting defects in nucleic acids encoding genes, or their expression products, for the biomarkers comprises performing next generation sequencing (NGS), nucleic acid hybridization, quantitative RT-PCR, immunohistochemistry (IHC), immunocytochemistry (ICC), or immunofluorescence (IF). [00193] Statement 11: The method of any of Statements 1-6, wherein the method further comprises assessment of a medical history, a family history, a physical examination, an endoscopic examination, imaging, a biopsy result, or a combination thereof. [00194] Statement 12: The method of Statement 11, wherein the method is used to develop a treatment strategy for the subject with cancer. [00195] Statement 13: The method of any of Statements 1-6, wherein the nucleic acids encoding genes are isolated from a fixed, paraffin-embedded sample from the subject. [00196] Statement 14: The method of any of Statements 1-6, wherein the nucleic acids encoding genes are isolated from core biopsy tissue or fine needle aspirate cells from the subject. [00197] Statement 15: The method of Statements 5 or 6, which further comprises treating the subject with a combination of ICB therapy and kinase inhibitor therapy. [00198] Statement 16: A method for treating a subject with cancer which comprises independently measuring or obtaining a tumor mutational burden (TMB) level; and defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of low BIG-BETS DDR genes normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample; and if the subject has a high TMB level and wild type low BIG-BETs genes, treating the subject with a combination of immune checkpoint blockcade (ICB) therapy and an inhibitor of a low BIG- BETs kinase so as to reduce the activity of the low BIG-BET kinase and thereby treat the subject with cancer. [00199] Statement 17: The method of Statement 16, wherein the low BIG-BETs kinase is ATR, CHEK1, or WEE1. [00200] Statement 18: A kit comprising at least ten nucleic acid probes, wherein each of said probes specifically binds to one of ten distinct biomarker nucleic acids or fragments thereof selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5. [00201] It should be understood that the above description is only representative of illustrative embodiments and examples. For the convenience of the reader, the above description has focused on a limited number of representative examples of all possible embodiments, examples that teach the principles of the disclosure. The description has not attempted to exhaustively enumerate all possible variations or even combinations of those variations described. That alternate embodiments may not have been presented for a specific portion of the disclosure, or that further undescribed alternate embodiments may be available for a portion, is not to be considered a disclaimer of those alternate embodiments. One of ordinary skill will appreciate that many of those undescribed embodiments, involve differences in technology and materials rather than differences in the application of the principles of the disclosure. Accordingly, the disclosure is not intended to be limited to less than the scope set forth in the following claims and equivalents. [00202] INCORPORATION BY REFERENCE [00203] All references, articles, publications, patents, patent publications, and patent applications cited herein are incorporated by reference in their entireties for all purposes. However, mention of any reference, article, publication, patent, patent publication, and patent application cited herein is not, and should not be taken as an acknowledgment or any form of suggestion that they constitute valid prior art or form part of the common general knowledge in any country in the world. It is to be understood that, while the disclosure has been described in conjunction with the detailed description, thereof, the foregoing description is intended to illustrate and not limit the scope. Other aspects, advantages, and modifications are within the scope of the claims set forth below. All publications, patents, and patent applications cited in this specification are herein incorporated by reference as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference.

Claims

CLAIMS What is claimed is: 1. A method for predicting response to immune checkpoint blockade (ICB) therapy for a subject with cancer, comprising independently measuring or obtaining (i) a tumor mutational burden (TMB) level and (ii) defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5 in a sample from the subject, normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample, wherein subjects with both a high TMB level and defects in at least ten of the nucleic acids encoding DDR genes, or their expression products, is indicative of an increased response to ICB therapy for the subject with cancer.
2. A method for predicting overall survival for a subject with cancer, comprising independently measuring or obtaining (i) a tumor mutational burden (TMB) level and (ii) defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5 in a sample from the subject, normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample, wherein subjects with both a high TMB level and defects in at least ten of the nucleic acids encoding DDR genes, or their expression products, is indicative of an increased overall survival for the subject with cancer.
3. A method to select a subject with cancer for immune blockade therapy (ICB) which comprises (a) independently measuring or obtaining both (i) a tumor mutational burden (TMB) level and (ii) defects in DNA Damage Repair (DDR) genes or their expression products in a sample from the subject; (b) calculating a Bipartite Graph-Based Expected TMB Score (BiG-BETS) to determine which DDR genes or expression products are associated with elevated TMB; (c) if the subject has a high BiG-BETS determining that the defect in the DDR genes or their expression products do not add to the TMB level over measurement of TMB levels alone; and (d) using the TMB levels alone to select a subject for ICB.
4. The method of claim 3, wherein the DDR genes or their expression products are at least ten biomarkers selected from the group consisting of BARD1, CHEK1, CUL5, EME1, ERCC4, ERCC5, ERCC6, FANCA, FANCC, FANCI, FEN1, GEN1, LIG4, MDC1, MLH1, MLH3, MRE11, MSH2, MSH3, MSH6, NBN, PALB2, PARP1, PMS1, PMS2, POLE, POLE3, POLL, POLM, RAD50, RNMT, SEM1, SLX1A, TDG, TOP3A, TOPBP1, UNG, XPC, XRCC2, XRCC4, and XRCC6.
5. A method to select a subject with cancer for immune blockade therapy (ICB) which comprises (a) independently measuring or obtaining both (i) a tumor mutational burden (TMB) level and (ii) defects in DNA Damage Repair (DDR) genes or their expression products in a sample from the subject; (b) calculating a Bipartite Graph-Based Expected TMB Score (BiG-BETS) to determine which DDR genes or expression products are associated with elevated TMB; (c) if the subject has a low BiG-BETS determining that the defects in the DDR genes or their expression products contribute to the likelihood of benefit from ICB over measurement of TMB levels alone; and (d) using both the TMB levels and the defects in the DDR genes or their expression products to select a subject for ICB.
6. The method of claim 5, wherein the DDR genes or their expression products are at least ten biomarkers selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5.
7. The method of any of claims 1-6, wherein the cancer is bladder urothelial carcinoma, colon adenocarcinoma, esophageal carcinoma, invasive breast carcinoma, head and neck squamous cell carcinoma, kidney renal clear cell carcinoma, lung adenocarcinoma, melanoma, or squamous cell lung carcinoma.
8. The method of any of claims 1-6, wherein the defects are mutations or copy number alterations.
9. The method of claim 8, wherein the mutations are deletions, frameshift mutations, insertions, missense mutations, nonsense mutations, start codon loss, stop codon loss or gain, or a combination thereof.
10. The method of any of claims 1-6, wherein the detecting defects in nucleic acids encoding genes, or their expression products, for the biomarkers comprises performing next generation sequencing (NGS), nucleic acid hybridization, quantitative RT-PCR, immunohistochemistry (IHC), immunocytochemistry (ICC), or immunofluorescence (IF).
11. The method of any of claims 1-6, wherein the method further comprises assessment of a medical history, a family history, a physical examination, an endoscopic examination, imaging, a biopsy result, or a combination thereof.
12. The method of claim 11, wherein the method is used to develop a treatment strategy for the subject with cancer.
13. The method of any of claims 1-6, wherein the nucleic acids encoding genes are isolated from a fixed, paraffin-embedded sample from the subject.
14. The method of any of claims 1-6, wherein the nucleic acids encoding genes are isolated from core biopsy tissue or fine needle aspirate cells from the subject.
15. The method of claims 5 or 6 which further comprises treating the subject with a combination of ICB therapy and kinase inhibitor therapy.
16. A method for treating a subject with cancer which comprises independently measuring or obtaining (a) a tumor mutational burden (TMB) level; and (b) defects in nucleic acids encoding DNA damage repair (DDR) genes, or their expression products, for at least ten biomarkers selected from the group consisting of low BIG-BETS DDR genes normalized against a reference set of nucleic acids encoding genes, or their expression products, in the sample; and (c) if the subject has a high TMB level and wild type low BIG-BETs genes, treating the subject with a combination of immune checkpoint blockcade (ICB) therapy and an inhibitor of a low BIG-BETs kinase so as to reduce the activity of the low BIG- BET kinase and thereby treat the subject with cancer.
17. The method of claim 16, wherein the low BIG-BETs kinase is ATR, CHEK1, or WEE1.
18. A kit comprising at least ten nucleic acid probes, wherein each of said probes specifically binds to one of ten distinct biomarker nucleic acids or fragments thereof selected from the group consisting of APEX1, APEX2, ATM, ATR, ATRIP, BLM, BRCA1, BRCA2, BRIP1, CHEK2, ERCC1, ERCC2, EXO1, FANCB, FANCD2, FANCL, FANCM, MUS81, NHEJ1, POLB, PRKDC, RAD51, RAD52, RBBP8, TDP1, TP53BP1, TREX1, UBE2T, XPA, XRCC3, and XRCC5.
PCT/US2023/064398 2022-03-15 2023-03-15 Improved methods of predicting response to immune checkpoint blockade therapies and uses thereof Ceased WO2023178152A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202263320169P 2022-03-15 2022-03-15
US63/320,169 2022-03-15

Publications (1)

Publication Number Publication Date
WO2023178152A1 true WO2023178152A1 (en) 2023-09-21

Family

ID=88024413

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2023/064398 Ceased WO2023178152A1 (en) 2022-03-15 2023-03-15 Improved methods of predicting response to immune checkpoint blockade therapies and uses thereof

Country Status (1)

Country Link
WO (1) WO2023178152A1 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN121249890A (en) * 2025-12-04 2026-01-02 嘉兴市中医医院 Pulmonary adenocarcinoma postoperative detection gene combination for recurrence risk assessment and application thereof

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20210397995A1 (en) * 2012-06-21 2021-12-23 Philip Morris Products S.A. Systems and methods relating to network-based biomarker signatures

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20210397995A1 (en) * 2012-06-21 2021-12-23 Philip Morris Products S.A. Systems and methods relating to network-based biomarker signatures

Non-Patent Citations (5)

* Cited by examiner, † Cited by third party
Title
IRANZO JAIME, MARTINCORENA IÑIGO, KOONIN EUGENE V.: "Cancer-mutation network and the number and specificity of driver mutations", PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES, NATIONAL ACADEMY OF SCIENCES, vol. 115, no. 26, 26 June 2018 (2018-06-26), XP093094102, ISSN: 0027-8424, DOI: 10.1073/pnas.1803155115 *
LABRIOLA MATTHEW KYLE, ZHU JASON, GUPTA RAJAN, MCCALL SHANNON, JACKSON JENNIFER, KONG ERIC F, WHITE JAMES R, CERQUEIRA GUSTAVO, GE: "Characterization of tumor mutation burden, PD-L1 and DNA repair genes to assess relationship to immune checkpoint inhibitors response in metastatic renal cell carcinoma", JOURNAL FOR IMMUNOTHERAPY OF CANCER, vol. 8, no. 1, 1 March 2020 (2020-03-01), pages e000319, XP055869090, DOI: 10.1136/jitc-2019-000319 *
TEO MIN YUEN, SEIER KENNETH, OSTROVNAYA IRINA, REGAZZI ASHLEY M., KANIA BROOKE E., MORAN MEREDITH M., CIPOLLA CATHARINE K., BLUTH : "Alterations in DNA Damage Response and Repair Genes as Potential Marker of Clinical Benefit From PD-1/PD-L1 Blockade in Advanced Urothelial Cancers", JOURNAL OF CLINICAL ONCOLOGY, AMERICAN SOCIETY OF CLINICAL ONCOLOGY, US, vol. 36, no. 17, 10 June 2018 (2018-06-10), US , pages 1685 - 1694, XP093094099, ISSN: 0732-183X, DOI: 10.1200/JCO.2017.75.7740 *
VENKATRAMAN DIVYA LAKSHMI, PULIMAMIDI DEEPSHIKA, SHUKLA HARSH G., HEGDE SHUBHADA R.: "Tumor relevant protein functional interactions identified using bipartite graph analyses", SCIENTIFIC REPORTS, vol. 11, no. 1, XP093094101, DOI: 10.1038/s41598-021-00879-2 *
WEIR WILLIAM H., MUCHA PETER J., KIM WILLIAM Y.: "A bipartite graph-based expected networks approach identifies DDR genes not associated with TMB yet predictive of immune checkpoint blockade response", CELL REPORTS MEDICINE, vol. 3, no. 5, 1 May 2022 (2022-05-01), pages 100602, XP093094103, ISSN: 2666-3791, DOI: 10.1016/j.xcrm.2022.100602 *

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN121249890A (en) * 2025-12-04 2026-01-02 嘉兴市中医医院 Pulmonary adenocarcinoma postoperative detection gene combination for recurrence risk assessment and application thereof

Similar Documents

Publication Publication Date Title
JP7595619B2 (en) Methods and materials for assessing loss of heterozygosity
Parry et al. Evolutionary history of transformation from chronic lymphocytic leukemia to Richter syndrome
Zhang et al. Genetic subtype-guided immunochemotherapy in diffuse large B cell lymphoma: The randomized GUIDANCE-01 trial
Desch et al. Genotyping circulating tumor DNA of pediatric Hodgkin lymphoma
JP7232476B2 (en) Methods and agents for evaluating and treating cancer
US20220010385A1 (en) Methods for detecting inactivation of the homologous recombination pathway (brca1/2) in human tumors
Shukla et al. Plasma DNA-based molecular diagnosis, prognostication, and monitoring of patients with EWSR1 fusion-positive sarcomas
Martins et al. A cluster of evolutionarily recent KRAB zinc finger proteins protects cancer cells from replicative stress–induced inflammation
Lak et al. Cell-free DNA as a diagnostic and prognostic biomarker in pediatric rhabdomyosarcoma
JP2017516501A (en) Lung cancer typing method
WO2023150627A1 (en) Systems and methods for monitoring of cancer using minimal residual disease analysis
Kennedy et al. RAS pathway activation drives clonal selection and monocytic differentiation in FLT3 and BCL2 inhibitor resistance
Chen et al. Circulating tumor DNA in colorectal cancer: biology, methods and applications
WO2025235468A1 (en) Cancer-associated fibroblast subtypes for diagnosis, prognosis, and treatment
WO2019178214A1 (en) Methods and compositions related to methylation and recurrence in gastric cancer patients
Filser et al. Nanopore sequencing as a cutting-edge technology for medulloblastoma classification
Roostee et al. Stand-alone transcriptional immune response prediction in primary triple-negative breast cancer
Fujikura et al. Genome-wide analysis of somatic non-coding mutation patterns and mitochondrial heteroplasmy in type B1 and B2 thymomas
Golubickaitė Breast, Cervical, Head and Neck Cancer: TFAM and POLG Variants, Mitochondrial DNA Alterations and Mirna-210-3P Expression Effect on Tumor Phenotype and Disease Outcome
Konigsberg Altered Epigenetic Regulation of Immune and Repair Processes in Interstitial Lung Disease
CN113278704A (en) Marker and product for diagnosing oral squamous cell carcinoma
HK40092784A (en) Multimodal analysis of circulating tumor nucleic acid molecules
HK40038488A (en) Methods and materials for assessing and treating cancer

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23771613

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 23771613

Country of ref document: EP

Kind code of ref document: A1

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 20/02/2025)

122 Ep: pct application non-entry in european phase

Ref document number: 23771613

Country of ref document: EP

Kind code of ref document: A1