WO2012135635A2 - Ovarian cancer biomarkers - Google Patents

Ovarian cancer biomarkers Download PDF

Info

Publication number
WO2012135635A2
WO2012135635A2 PCT/US2012/031484 US2012031484W WO2012135635A2 WO 2012135635 A2 WO2012135635 A2 WO 2012135635A2 US 2012031484 W US2012031484 W US 2012031484W WO 2012135635 A2 WO2012135635 A2 WO 2012135635A2
Authority
WO
WIPO (PCT)
Prior art keywords
sample
ovarian cancer
snv
nucleic acid
mutations
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2012/031484
Other languages
French (fr)
Other versions
WO2012135635A3 (en
Inventor
Jian-Bing Fan
Russell GROCOCK
Keira CHEETHAM
Richard Shaw
Jeremy R. CHIEN
Vijayalakshmi Shridhar
Lynn Hartmann
Dirk Evers
John Peden
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Mayo Foundation for Medical Education and Research
Illumina Inc
Mayo Clinic in Florida
Original Assignee
Mayo Foundation for Medical Education and Research
Illumina Inc
Mayo Clinic in Florida
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Mayo Foundation for Medical Education and Research, Illumina Inc, Mayo Clinic in Florida filed Critical Mayo Foundation for Medical Education and Research
Publication of WO2012135635A2 publication Critical patent/WO2012135635A2/en
Publication of WO2012135635A3 publication Critical patent/WO2012135635A3/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6883Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
    • C12Q1/6886Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material for cancer
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/156Polymorphic or mutational markers

Definitions

  • Ovarian cancer is the among the top ten most common cancers among women, and the fifth leading cause of death for women with cancer in the United States.
  • ovarian cancer the most lethal gynecological disease among women in developed countries is ovarian cancer.
  • the American Cancer Society estimates that about 22,000 new cases will be diagnosed this year and approximately 15,000 women will die from ovarian cancer in the United States alone.
  • the incidence rate of ovarian cancer is roughly 13 per 100,000 women per year and even though the median age at diagnosis is around 63, no age group is immune to the disease. Further, even though incidence of ovarian cancer is slightly higher among white women, no race or ethnic background is immune.
  • Survival rate once a diagnosis is made is typically dismal, for example if ovarian cancer is detected and effectively treated prior to metastasis the 5 year survival rate can be as high as 73%, however if the cancer is not detected until it has metastasized then the long term survival rate drops to ⁇ 30%.
  • ovarian cancer often goes undetected until it has metastasized within the pelvis and abdomen, therefore the outcome for most women is grim as the cancer is difficult to treat and is often fatal.
  • early detection ovarian cancer is crucial to benefit those patients, for example, that present with no or vague symptoms or with tumors that are below the level of detection during a physical examination.
  • a considerable amount of research effort has been focused in discovering and developing early detection systems, however to date no effective screening method has been developed.
  • what are needed are ways to detect ovarian cancer, preferably at an early stage, for example before metastasis, thereby improving the long term survival of those women afflicted with this disease.
  • the present disclosure identifies biological markers, or biomarkers, indicative of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer.
  • the disclosed biomarkers allow for identification of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject.
  • the methods described herein utilize the biomarkers and provide alternatives to currently available ovarian cancer determinative, diagnostic and prognostic methodologies.
  • the biomarkers and methods of their use as disclosed herein can be applied to the characterization, classification, differentiation, grading, staging, diagnosis, or prognosis of ovarian cancer, ovarian cancer type and/or ovarian cancer stage.
  • Embodiments as described herein are based in part on the identification of reliable biomarkers for the improved determination, screening, diagnosis and/or prognosis of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject.
  • the disclosure provides a population of gene targets or gene related targets (i.e., biomarkers) and methods of use as described herein.
  • Biomarkers and methods of their use for determining ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject comprise TP53, MUC2, MY016, ARID1A, CT B 1, CSMD3, TRRAP, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1 S, CACNA2D1, PIKFYVE, PIK3CA, NIPBL, PKHD1, MLL3, ATBF 1, TPO, DNAH7, LRRIQ1, SDHA, TIAM1, TTN, SLC16A7, COL3A1 and HNRNPUL1.
  • methods for determining the presence of ovarian cancer in a subject comprises providing a nucleic acid sample from a subject, detecting the presence of one or more variant mutations in two or more genes selected from the group consisting essentially of ARID 1 A, CACNA2D1, CTTNB1, PKHD1, DNAH7, PIKC3A, TRRAP and CSMD3 in the sample, evaluating the probability that that the two or more genes are correlated with ovarian cancer, and determining the presence of ovarian cancer in the sample based on the probability of correlation.
  • evaluating the probability comprises comparing the nucleic acid sample suspected of having ovarian cancer to a matched normal nucleic acid sample, wherein both the nucleic acid samples (i.e., test and normal) are genomic DNA samples.
  • a gene is correlated with ovarian cancer at p ⁇ 0.05.
  • the gene DNAH7 is one of the two or more genes identified as being correlated with ovarian cancer at p ⁇ 0.05.
  • genomic DNA for testing for presence of ovarian cancer from a subject is isolated from a sample selected from a group consisting of a tissue sample, a biopsy sample, a cell sample, a circulating tumor cell sample, a fixed tissue sample, a frozen tissue sample or a lavage sample
  • methods for determining the presence of ovarian cancer in a subject comprises providing a nucleic acid sample from a subject, detecting the presence of one or more variant mutations in two or more genes selected from the group consisting of ARID1A, CACNA2D1, CTTNB1, PKHD 1 , DNAH7, PIKC3 A, TRRAP and CSMD3 , evaluating the probability that that the two or more genes are correlated with ovarian cancer, and determining the presence of ovarian cancer in the sample based on the probability of correlation.
  • detecting mutations in genes correlated with ovarian cancer comprises sequencing, such as sequence by synthesis methodologies, microarray analysis and/or polymerase chain reaction
  • methods are described herein for determining the presence of ovarian cancer in a sample comprising creating a DNA library from nucleic acids derived from a sample suspected of having ovarian cancer, sequencing the DNA library to identify one or more mutations in two or more genes from the list consisting essentially of ARID 1 A, CACNA2D1, CTTNB1, PKHD l, DNAH7, PIKC3A, TRRAP and CSMD3, computationally determining the probability that the two or more genes are correlated with ovarian cancer, and determining the presence of ovarian cancer in said sample based on the probability of correlation.
  • computational determination of ovarian cancer correlated comprising comparing the sequence of the nucleic acid sample suspected of having ovarian cancer to a matched normal sample.
  • a gene is correlated with ovarian cancer at p ⁇ 0.05.
  • DNAH7 is one of the two or more genes identified as being correlated with ovarian cancer.
  • a test sample suspected of having ovarian cancer is isolated from a sample selected from the group consisting of a tissue sample, a biopsy sample, a cell sample, a circulating tumor cell sample, a fixed tissue sample, a frozen tissue sample or a lavage sample, wherein a matched normal sample may be isolated from either a blood sample or a tissue sample that does not have ovarian cancer.
  • a DNA library is sequenced using sequence by synthesis
  • a method for confirming the presence of ovarian cancer in a sample comprises comparing the sequence of a nucleic acid sample derived from a tissue suspected of having ovarian cancer with the sequence of a nucleic acid from a tissue not suspected of having ovarian cancer, wherein the presence of one or more variant sequences identified in two or more of ARIDIA, CACNA2D1, CTTNB l, PKHDl, DNAH7, PIKC3A, TRRAP and CSMD3 in the test sample as compared to the normal sample indicates the presence of ovarian cancer in a sample.
  • methods are described herein for determining the presence of ovarian cancer in a sample comprising evaluating a sample for the presence of one or more variant mutations in two or more genes selected from the group comprising ARIDIA, CTNNBl, CSMD3, TRRAP, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
  • a sample under evaluation is compared to a matched normal nucleic acid sample.
  • the nucleic acid sample is genomic DNA from a human subject as is the matched normal nucleic acid sample.
  • a genomic DNA sample is isolated from a tissue sample, a biopsy sample a cell sample, a circulating tumor cell sample, a fixed tissue sample, a frozen tissue sample or a lavage sample.
  • the presence of one or more variant mutations is found in two or more genes selected from the group consisting essentially of MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl, TTN, SLC16A7, COL3A1 and HNRNPULl.
  • the presence of one or more variant mutations is found in two or more genes selected from the group consisting of MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
  • determining ovarian cancer comprises determining the presence of high grade serous ovarian cancer, whereas in other embodiments determining ovarian cancer comprises determining the presence of non-high grade serous ovarian cancer. In some embodiments, determining ovarian cancer comprises determining the presence of serous ovarian cancer.
  • the presence of one or more variant mutations is found in two or more genes selected from the group consisting of MUC2, MY016,
  • OR4A47 CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl, and TTN, further comprising determining the presence of high grade serous ovarian cancer based on the presence of the variant mutations.
  • the presence of one or more variant mutations is found in a group comprising SLC16A7, wherein the presence of one or more variant mutation determines the presence of non-high grade serous ovarian cancer in a sample.
  • the presence of one or more variant mutations is found in two or more genes selected from the group consisting essentially of ARIDIA, CTNNB1, CSMD3, TRRAP, PIKC3A, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl, TTN, SLC16A7, COL3A1 and HNRNPULl, the presence of which determines the presence of serous ovarian cancer.
  • a sample is evaluated for the presence of one or more variant mutations in the group comprising MUC2, MYO 16, HNRNPULl and COL3A1, the presence of which determines the presence of serous ovarian cancer. In some embodiments, a sample is evaluated for the presence of one or more variant mutations in the group consisting of MUC2, MYO 16, HNRNPULl and COL3A1, the presence of which determines the presence of serous ovarian cancer.
  • methods are described herein for determining the presence of high grade serous ovarian cancer in a sample comprising evaluating the sample for the presence of one or more variant mutations in two or more genes selected from the group comprising MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl and TTN.
  • evaluating the sample for the presence of one or more variant mutations comprises evaluating the nucleic acid sample for one or more variant mutations in one or more of MUC2 and MY016 and one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
  • the evaluation further comprises evaluating one or more of DNAH7, LRRIQl, SDHA, TIAMl and TTN.
  • methods are described herein for determining the presence of serous ovarian cancer in a sample comprising evaluating the sample for the presence of one or more variant mutations in two or more genes selected from the group comprising MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, COL3A1, CACNA1S, CACNA2D1, PIKFYVE, NIPBL,
  • evaluating the sample for the presence of one or more variant mutations comprises evaluating the nucleic acid sample for one or more variant mutations in one or more of HNRNPULl and COL3A1 and one or more of MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl and TTN.
  • evaluating the sample for the presence of one or more variant mutations comprises evaluating the nucleic acid sample for one or more variant mutations in one or more of HNRNPULl and COL3 A 1 , one or more of MUC2 and MYO 16 and one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
  • evaluating the sample for the presence of one or more variant mutations comprises evaluating the nucleic acid sample for one or more variant mutations in one or more of HNRNPUL1 and COL3A1, one or more of MUC2 and MY016, one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1 S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF 1 , TPO and one or more of DNAH7, LRRIQ 1 , SDHA, TIAM 1 and TTN.
  • evaluating a nucleic acid sample for the presence of ovarian cancer, high grade serous ovarian cancer or serous ovarian cancer comprises sequencing the nucleic acid sample. In some embodiments, sequencing a sample comprises sequence by synthesis methodologies. In some embodiments, evaluation of a nucleic acid sample for the presence of ovarian cancer, high grade serous ovarian cancer or serous ovarian cancer is carried out on a microarray. In some embodiments, evaluating a nucleic acid sample for the presence of ovarian cancer, high grade serous ovarian cancer or serous ovarian cancer comprises performing polymerase chain reaction on the sample, for example quantitative or real-time PCR.
  • Figure 2 is exemplary of the number of serous ovarian cancer patient samples
  • test sample is intended to mean any biological fluid, cell, tissue, organ or portion thereof that contains genomic nucleic acids, for example genomic DNA or RNA, suitable for mutational detection via the disclosed methods.
  • a test sample can include or be suspected to include a cell, such as a cell from an ovary, uterus, fallopian tube, vagina, or other organ or tissue that contains or is suspected to contain a cancerous cell.
  • the term includes samples present from an individual as well as samples obtained or derived from an individual.
  • a sample can be a histologic section of a specimen obtained by biopsy, cell scraping, etc. or cells that are placed in or adapted to tissue culture.
  • a sample further can be a sub-cellular fraction or extract, or a crude or isolated nucleic acid molecule.
  • a patient matched normal sample can be used to establish a mutational background for comparison to a patient test sample.
  • a sample may be obtained in a variety of ways known in the art. Samples may be obtained according to standard techniques from all types of biological sources that are usual sources of genomic DNA including, but not limited to cells or cellular components which contain DNA, cell lines, circulating tumor cells, biopsies, bodily fluids such as blood, lavage specimens, tissue samples such as tissue that are formalin fixed and embedded in paraffin such as tissue from ovaries, endometrium, cervix, fallopian tubes, omentum, histological object slides, and all possible combinations thereof. Further, tissues can be fresh, fresh frozen, etc. Accordingly, a sample can be from an archived, stored or fresh source as suits a particular application of the methods set forth herein.
  • the methods described herein can be performed on one or more samples from ovarian cancer patients such as samples obtained by vaginal lavage, endometrial biopsy, ovarian biopsy, and/or blood draw.
  • Sample analysis can be applied, for example, to the presence or absence of ovarian cancer, differentiation between early and/or late stage ovarian cancer types, ovarian cancer epithelial type differentiation, or to monitor cancer progression or response to treatment.
  • a suitable sample can be collected and acquired that is either known to comprise ovarian cancer cells or is subsequent to the formulation of the diagnostic aim of a biomarker as disclosed herein.
  • a sample can be derived from a population of cells or from a tissue that is predicted to be afflicted with or phenotypic of ovarian cancer.
  • the genomic DNA can be derived from a high-quality source such that the sample contains only the tissue type of interest, minimum contamination and minimum DNA fragmentation.
  • samples are contemplated to be representative of the tissue or cell type of interest that is to be handled by an assay.
  • a population or set of samples from an individual source can be analyzed to maximize confidence in the results for an individual.
  • a sample from an individual is matched and compared to a normal sample from that same individual to identify the mutational status of biomarkers for that individual.
  • the normal sample, or patient matched normal sample can be from the same or similar organ, tissue or fluid as the sample to which it is compared.
  • the normal sample will typically display a phenotype that is different from a phenotype of the sample to which it is compared.
  • isolated or purified when used in relation to a nucleic acid refers to a nucleic acid sequence that is extracted and separated from at least one component or contaminant with which it is ordinarily associated in its natural source. As such, an isolated or purified nucleic acid is present in a form or setting that is different from that in which it is found in nature.
  • a biomarker can be DNA or RNA, proteins, polypeptides, variants, fragments or functional equivalents thereof.
  • a biomarker is generally associated with a genomic nucleic acid such as a gene or gene associated region or location unless specified otherwise.
  • Biomarkers disclosed herein that are associated with ovarian cancer, a particular type of ovarian cancer and/or a particular stage of ovarian cancer comprise one or more single nucleotide variants and/or insertions/deletions (indels) located in a gene or gene associated region as compared to its equivalent in a normal sample.
  • a gene that contains one or more somatic mutations, such as variant mutations, identified in two or more patient samples is contemplated to be a biomarker that is useful in detecting, diagnosing or prognosing ovarian cancer, a particular type of ovarian cancer and/or a stage of ovarian cancer.
  • Somatic mutation is an alteration in the genome that occurs after conception resulting in a genetic difference of the genome at that particular location. Somatic mutations can occur in any cell in the body except the germ cells and are passed to the cell progeny during cell division. Somatic mutations include, but are not limited to, point mutations such as single nucleotide variants (SNVs), gene amplification or duplication, genetic insertions and/or deletions (indels), chromosomal translocations, chromosomal inversions and single nucleotide polymorphisms (SNPs). Somatic mutations can result in phenotypic changes, disease formation, cancer, etc. "Somatic mutations" as used herein, unless otherwise stated, include SNVs and/or indels present in genomic DNA, and are considered variant mutations, or mutations that may result in phenotypic changes, disease formation, cancer, etc.
  • Identification of somatic variants described herein was performed by looking for a SNV or indel in the patient cancer sample which was not present in the patient matched normal sample. If a somatic mutation was found in the cancer sample which was not present in the normal sample, then that mutation was identified as a SNV or indel, as the case may be. If a somatic mutation was present in both the cancer sample and the normal sample, then that mutation was considered part of that patient's genetic background and was not considered a variant.
  • a gene from two or more patient samples that has one or more variant mutations is considered a biomarker and useful in detecting, diagnosing and prognosing ovarian cancer, a particular type of ovarian cancer and/or a stage of ovarian cancer.
  • gene refers to a nucleic acid sequence, such as DNA, that comprises coding sequences associated with the production of a polypeptide, precursor, or RNA (e.g., rRNA, tRNA). Typically, a gene also includes non-coding and intergenic sequences. The term can encompass the coding region of a gene and the sequences located adjacent to the coding region on both the 5' and 3' such that the gene corresponds to the length of the full-length mRNA. Sequences located 5' of the coding region and present on the mRNA are referred to as 5' non-translated sequences. Sequences located 3' or downstream of the coding region and present on the mRNA are referred to as 3' non-translated sequences.
  • genomic form or clone of a gene contains the coding region interrupted with non-coding sequences such as introns, intervening regions, intervening sequences or intergenic regions.
  • Ovarian cancer is often referred to as a silent killer because of its subtle symptoms that lead to delayed discovery, diagnosis and treatment.
  • the majority of ovarian cancers are diagnosed when the cancer has already reached an advanced stage, for example >80% of serous ovarian cancers are diagnosed at Stage III or Stage IV leading to a very low chance of long-term survival in these patients.
  • Screening and/or detecting ovarian cancer in women who might be at higher risk of developing ovarian cancer, such as those with a strong family history of such cancer is problematic.
  • the two most common screening tests for ovarian cancer include transvaginal sonography and identification of a protein marker, CA-125. However, both tests have limitations.
  • transvaginal sonography can identify a mass in the ovary however the sonogram is unable to distinguish whether the mass is cancerous or not.
  • the protein marker CA-125 is not specific to the presence of ovarian cancer as other cancers also exhibit high levels of CA-125.
  • the majority of ovarian tumor cancers are of the epithelial histologic type, which can be further divided into different tumor subtypes, for example serous, endometrioid, clear cell, mucinous, Brenner or transitional cell, squamous cell, undifferentiated and mixed epithelial cell types (AJCC Cancer Staging Manual 7 th Ed., p.422).
  • TAM tumor node metastasis
  • FIGO Federation Internationale de Gynecologie et d'Obstetrique
  • Ovarian cancer can also be classified into two groups based on molecular progression. For example, Type I ovarian tumors of mucinous, clear cell, endometrioid, and low-grade serous type develop in stepwise fashion from adenomas to carcinomas, whereas Type II tumors of high-grade serous develop de novo from undefined precursor lesions and progress rapidly with no apparent stepwise progression (Ie and Kurman, 2004, Am J Pathol 164: 151 1-1518). Further explanation of cancer staging and grading can be found at, for example, AJCC Cancer Staging Manual, Edge, SB et al, Eds., Springer-Verlag, New York. The vast majority of
  • Type II, high-grade serous ovarian cancers are diagnosed at advanced stages and represent a major challenge in early detection (Chan et al, 2006, Obs and Gyn 108: 521-528).
  • ovarian cancer arises from epithelial cells that line the surface of the ovary. Approximately 50% of epithelial ovarian tumors are classified as serous, or tumors with glandular features, and make up approximately 80% of all ovarian tumors. Other types of ovarian cancers can arise from germ cells (e.g., cancer of the ovarian egg-making cells) and sarcomas. High-grade serous tumors denote highly aggressive, invasive tumors as compared to low malignant potential (LMP) tumors. Whether an invasive serous tumor is classified as either high or low grade is based on the clinical course of the disease.
  • LMP low malignant potential
  • high grade serous tumors were found to over express genes that control various cellular functions associated with cancer cells, for example genes that control cell growth, DNA stability (or lack thereof) and genes that silence other genes.
  • LMP tumors were not found to overexpress these types of genes and LMP tumors were alternatively characterized by expression of growth control pathways, such as tumor protein 53 (TP53 or p53) pathways.
  • a high grade serous tumor has a high degree of chromosomal instability, has mutated TP53 gene, demonstrates very fast tumor development, typically has nuclei that are non-uniform, enlarged and irregularly shaped and has high mitotic index.
  • staging and grading cancers are subjective and rely on a diagnostician to interpret morphology, histology, anatomy and other related indices. Further, as ovarian cancer is typically left undiagnosed until late stage cancer due to, for example, its asymptomatic phenotype, the staging and grading do nothing to identify early stage cancer, or identify ovarian cancer earlier in the disease progression in the absence of disease related symptoms.
  • a patient cohort was sequenced as described herein and a subset of the candidate list was created (Table 2).
  • Gene locations as identified in the tables herein, for example in Table 2 are as found in the Archive Ensembl Human database, release 59-Aug 2010 (http://aug2010.archive.ensembl.org/Homo_sapiens/Info/Index) which provides human genomic data as assembled from the Genome Reference Consortium (GRC).
  • GRC consists of the Wellcome Trust Sanger Institute, the Genome Center at Washington University, the European Bioinformatics Institute and the National Center for Biotechnology Information.
  • the subset comprises biomarkers identified during sequencing that were correlated with ovarian cancer, such that one or more of the patient samples were identified to have one or more variant mutations in a gene, thereby classifying that gene as a biomarker for detecting, diagnosing, and prognosing ovarian cancer, a particular type of ovarian cancer and/or a stage of ovarian cancer.
  • Table 2-List of an exemplary subset from candidate list
  • CACNA1S IS subunit chrl:201008642-201081694 calcium channel, voltage-dependent, alpha 2/delta
  • CNBD1 cyclic nucleotide binding domain containing 1 chr8:88218216-88394955
  • FRG2C FSHD region gene 2 family member C chr3:75713481-75716371
  • HNRNPUL1 heterogeneous nuclear ribonucleoprotein U-like 1 chrl9:41768391-41813811
  • IL13RA1 interleukin 13 receptor, alpha 1 chrX:117861535-117928502
  • Phosphoinositide-3-kinase catalytic, alpha
  • solute carrier family 16 member 7 (monocarboxylic
  • the Table 2 list of biomarkers was collated based on data from a retrospective study comprising genomic samples from high-grade serous, low-grade serous, endometrioid, mucinous and clear cell epithelial type ovarian cancer tumors from the first sample cohort of 25 patients (Table 3, and Example 1).
  • the tier grade was determined as described in Vang 2009.
  • samples were selected with high tumor content prior to sequencing by selecting the tumor samples with at least 70% tumor nuclei during pathological review of fresh frozen tumor samples. Average coverage of a haploid genome was 40 fold with approximately 93% of exomic sequence and approximately 86% of the whole genomic sequence covered by at least 10 reads (Table 4).
  • the tumor with 385,369 mutations was from a patient with Stage IC high-grade serous ovarian cancer and was determined to represent a hypermutator phenotype as the patient had neither prior chemotherapy nor radiation therapy treatment and no pre-existing cancer prior to ovarian cancer diagnosis. Excluding the hypermutated sample, approximately 2130 mutations within gene coding regions and 36,936 mutations outside gene coding regions (i.e., non-coding regions) were identified in approximately 256 genes in one or more tumor samples in the study.
  • the tumor protein 53 (TP53 or p53) was identified as the most frequently mutated gene in high-grade ovarian cancer, with 15 out of 16 tumors showing mutations in TP53. Mutations of TP53 were identified throughout the coding regions, consistent with the mutation profiles of tumor suppressor genes (Jones et al, 2010, Science 330:228-31; Vogelstein and Kinzler, 2004, Nat Med 10:789-99).
  • experimentation identified 338 gene associated locations having statistically significant mutations in non-coding regions when compared to random background mutations. Since previous studies have demonstrated that distal enhancer elements could be as far away as 500 kb in a genome, it is contemplated that mutations in non-coding regions several hundred kilobases away from a promoter may possess functional significance in regulating gene expression and therefore also serve as viable biomarkers indicative of cancer, such as ovarian cancer.
  • Validation studies were performed on a different second patient cohort (i.e., different from the fist patient cohort) of 40 tumor tissue samples and matched normal tissue samples (i.e., validation sample cohort).
  • a validation cohort of 40 tumor normal and normal matched samples was obtained.
  • the validation tumor samples were formalin fixed, paraffin embedded (FFPE) samples and normal matched FFPE samples were obtained from fallopian tubes, ovary, or normal lymph nodes.
  • FFPE formalin fixed, paraffin embedded
  • the validation data identified nine biomarkers in particular that can be considered hot spots for mutations in ovarian cancer.
  • the nine biomarkers as listed in Table 7 comprise ARID 1 A, CACNA2D1, CSMD3, CTN B1, DNAH7, PIKC3A, PKHDl, TP53 and TRRAP and were identified upon data analysis from the validation sample cohort, wherein mutations in these genes were found to be correlated (p ⁇ 0.05) with the presence of ovarian cancer. Gene ID and gene location are reported as found in the Ensembl database as previously described.
  • biomarkers and their methods of use are described below.
  • the biomarkers and their methods of use are not limited to these embodiments.
  • Biomarkers as described herein find utility, either alone or in combination, in methods for diagnosing ovarian cancer, a type of ovarian cancer and/or a stage of ovarian cancer. Biomarkers as described herein find utility, either alone or in combination, in methods for prognosing patient outcome diagnosed with ovarian cancer, a type of ovarian cancer and/or a stage of ovarian cancer.
  • Biomarkers as described herein find utility, either alone or in combination, in methods for screening patients for the presence or absence of ovarian cancer, a type of ovarian cancer and/or a stage of ovarian cancer, for example for patients that might be part of a high risk population predisposed to developing ovarian cancer (e.g., family history, genetic predisposition, etc.).
  • the biomarkers as described herein, either alone or in combination, find utility as diagnostic, prognostic, or screening tools in conjunction with additional tests and methods for identifying ovarian cancer.
  • Additional tests and methods for identifying ovarian cancer include, but are not limited to, transvaginal sonography, protein staining methods such as IHC or histopathological staining such as H&E, genetic probe assays such as in situ hybridization (ISH), TNM and/or FIGO staging, clinical staging, pathological staging, etc., for example as recognized by the American Joint Committee on Cancer (AJCC) and/or the World Health Organization (AJCC Cancer Staging Manual). Additional tests and method for identifying ovarian cancer may include experimental or discovery related tests and methods that are not yet recognized as mainstream, however find utility in providing support for a diagnosis of ovarian cancer nonetheless.
  • a biomarker is a gene or genetic location that was identified to comprise one or more variant mutations in patient test samples as compared to the gene or gene location in a patient matched normal sample.
  • a biomarker that was identified to have variants mutations in at least two patient samples is contemplated to represent a "hot spot", or gene that comprises variants mutations as compared to other genes in an ovarian cancer test sample (e.g., tissue, cell, circulating tumor cells, etc.).
  • Table 2 and Table 7 are exemplary of genes that were identified in two or more patient samples to have variant mutations compared to the same gene is a patient matched normal sample, thereby identifying them as potential biomarkers for the presence or absence of ovarian cancer, type of ovarian cancer, and/or stage of ovarian cancer.
  • variant mutations in biomarkers as described herein are located in a coding region of a gene.
  • biomarkers as described herein are located in non-coding regions of a gene.
  • biomarkers as described herein are located in intergenic regions.
  • biomarkers as described herein comprise single nucleotide variants (SNVs).
  • biomarkers as described herein comprise insertions and/or deletions (indels) of one or more genomic sequences.
  • a biomarker may comprise both SNVs and indels.
  • the biomarkers as disclosed herein are useful in detecting the presence or absence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject.
  • the biomarkers as disclosed herein are useful in diagnosing the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject.
  • the biomarkers as disclosed herein are useful in prognosing disease progression, treatment outcome and/or treatment regimen progress of a subject diagnosed with ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer. In some
  • the biomarkers as disclosed herein are useful in screening a subject for the possibility of developing ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject.
  • the biomarkers as described herein are useful in screening potential therapeutic options for treating a patient having ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject.
  • the present disclosure provides biomarkers comprising non-coding region variant mutations for detecting the presence or absence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer.
  • ROB02 one or more of ROB02, CNTNAP2, PTPRN2, GPC5, LRP 1B, CSMD3, DLG2, CNTNAP5, DPP6, PCDH15, and CDH12 are biomarkers indicative of ovarian cancer as demonstrated in Table 9.
  • ROB02 codes for a member of the roundabout family of proteins involved in axonal guidance and neuronal migration.
  • CNTNAP2 and CNTNAP5 contactin associated protein genes are members of the neurexin family which function as cell adhesion molecules.
  • PTPR 2 is a member of the receptor protein tyrosine
  • phosphatase gene family the proteins of which are involved in regulation of a variety of cellular processes including, but not limited to, cell growth and oncogenic
  • GPC5 a glypican gene, codes for a cell surface proteoglycan
  • biomarkers associated with the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer comprise one or more of ROB02, PTPRN2, GPC5, CSMD3, HNT and CNTNAP5.
  • biomarkers associated with the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer further comprise one or more of DPP6, KCNIP4 and PCDH15.
  • biomarkers associated with the presence of ovarian cancer in a sample further comprise CNTNAP2, LRP1B, and DLG2.
  • biomarkers associated with the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer TP53, MUC2, MY016, OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, PIK3CA, ARID 1 A, CTNNB1, CSMD3, TRRAP, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl, TTN, SLC16A7, COL3A1 and HNRNPULl.
  • biomarkers comprising variant mutations in ovarian cancer patient samples comprise one or more of, two or more of, TP53, MUC2, MY016, OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, PIK3CA, ARID 1 A, CTNNB1, CSMD3, TRRAP, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl, TTN, SLC16A7, COL3A1 and
  • biomarkers comprising variant mutations in ovarian cancer patient samples and associated with the presence of high grade serous ovarian cancer in a sample comprise one or more of TP53, MUC2 and MY016 and one or more of OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl and TTN ( Figure 1).
  • biomarkers comprising variant mutations in three or more patient samples and associated with the presence of high grade serous ovarian cancer in a sample comprise TP53, MUC2 and MY016.
  • biomarkers comprising variant mutations in two or more patient samples and associated with the presence of high grade serous ovarian cancer in a sample comprise one or more OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl and TTN.
  • a method for diagnosing the presence of high grade serous ovarian cancer in a subject comprises identifying variant mutations in one or more of TP53, MUC2, MY016,OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl and TTN in a sample from a subject.
  • a method for diagnosing the presence of high grade serous ovarian cancer in a subject comprises identifying variant mutations in one or more of TP53, MUC2 and MY016 and one or more of OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl and TTN in a sample from a subject.
  • a method for diagnosing the presence of high grade serous ovarian cancer in a subject comprises identifying variant mutations in one or more of TP53, MUC2 and MY016 and one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1 S, CACNA2D 1 , PIKFYVE, NIPBL, PKHD 1 , MLL3 , ATBF 1 , TPO and one or more of DNAH7, LRRIQ1, SDHA, TIAMl and TTN in a sample from a subject.
  • a sample from a subject used in methods for diagnosing ovarian cancer as described herein is a tissue sample, for example a biopsy tissue sample, for example an ovarian tissue biopsy sample.
  • a biopsy tissue sample used in diagnostic methods described herein is a fresh sample, or a sample that has been frozen or modified.
  • a modified sample is, for example, a sample that has been preserved or modified for storage and/or for use in
  • a sample from a subject is a liquid sample, such as a blood sample containing white blood cells.
  • a liquid sample contains circulating tumor cells.
  • the liquid sample is a lavage, for example a vaginal lavage, a cervical lavage, wherein cells are harvested from the lavage sample.
  • a sample is a cell sample, such as a cervical cell sample, for example as derived from a Papanicolaou test (PAP smear) slide.
  • nucleic acids are extracted and isolated from a sample, or portion thereof for subsequent use in methods as described herein for determining the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample.
  • white blood cells from a blood sample were utilized as a patient matched normal sample for comparison to a tissue sample.
  • blood samples were obtained from a patient and matched to that patient's tissue sample for evaluation of biomarker mutations as described herein (Example 1).
  • the genomic DNA was isolated from the white blood cells and served as a patient baseline (normal) for comparing mutations present in the test sample.
  • tissue sample that is free from the cancerous phenotype can also be utilized as a source of comparative normal genomic DNA for a patient where appropriate and available.
  • nucleic acids isolated from a sample, or a portion thereof are used in diagnostic and/or prognostic methods.
  • the one or more biomarkers as described herein can be used in methods for determining ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject. Further, the biomarkers find utility in combination with other biomarkers and/or other diagnostic tests in providing a diagnostician additional tools to determine ovarian cancer status of a subject. The methods as described herein find particular utility as diagnostic and prognostic tools. In embodiments of the present disclosure, methods described herein can be used to diagnose ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject.
  • methods comprising biomarkers as described herein are useful in differentiating early stage, high grade serous ovarian cancer from late stage, high grade serous ovarian cancer. In some embodiments, methods comprising biomarkers as described herein are prognostic for patient survival due to said differentiating early stage, high grade serous ovarian cancer from late stage, high grade serous ovarian cancer.
  • the biomarkers comprise variant genetic sequences in a genomic DNA sample compared to a genomic DNA normal sample.
  • Variant genetic sequences comprise single nucleotide variants, sequence insertions, sequence deletions, within genes that differ from a normal sample and indicate mutations that may indicate phenotypic changes, disease formation, cancer, etc.
  • the methods detect one or more altered genetic sequences as compared to a normal sample.
  • a comparison between gene sequences in a test sample (i.e., collected from a patient, subject, individual, etc.) to a normal or control sample (i.e., from the sample patient, subject, individual from which the test sample is collected) identifies the number of mutations associated with a particular gene, wherein the presence of a variant gene over a normal may associate that gene with ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer as described herein.
  • methods disclosed herein detect the insertion of one or more genetic sequences into a gene, deletion of one or more genetic sequences from a gene, or both as compared to a normal, or control, sample.
  • the methods detect one or more of single nucleotide variant(s) and/or insertion(s) and/or deletion(s) altered genetic sequences in a sample compared to a normal, or control sample.
  • Biomarkers and methods of their use as described herein can be used to differentiate between high grade serous and non-high grade serous (i.e., low grade serous, mucinous, endometrioid, clear cell type ovarian cancer) ovarian cancer
  • Biomarkers useful for differentiating between high grade serous and non- high grade serous ovarian cancer comprise TP53, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
  • a sample of known ovarian cancer is further identified as high grade serous ovarian cancer by detecting variant mutations in one or more of TP53, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQI, SDHA, TIAMl and TTN as compared to a normal, or control sample.
  • a sample of known ovarian cancer is further identified as non-high grade serous ovarian cancer by detection variant mutations in SLC16A7.
  • biomarkers and methods of their use as described herein differentiate between serous and non-serous (i.e., mucinous, endometrioid, clear cell type ovarian cancer) ovarian cancer in a sample ( Figure 2).
  • Biomarkers useful for differentiating between serous and non-serous ovarian cancer comprise one or more of TP53, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, COL3A1, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, HNRNPULl, TPO, DNAH7, LRRIQI, SDHA, TIAMl and TTN.
  • a method for diagnosing the presence of serous ovarian cancer in a subject comprises identifying variant mutations in one or more of TP53, MUC2 and MY016 and one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, COL3A1, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, HNRNPULl, TPO, DNAH7, LRRIQI, SDHA, TIAMl and TTN in a sample from a subject.
  • a method for diagnosing the presence of serous ovarian cancer in a subject comprises identifying variant mutations in one or more of TP53, MUC2 and MY016, one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, COL3A1, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF 1,
  • a test sample i.e., a sample to be assayed for presence of ovarian cancer
  • a second sample e.g., a normal or control sample (e.g., blood sample, tissue sample known not to have a cancerous phenotype) is collected from the same individual.
  • Genomic DNA is isolated from the sample(s) by techniques known in the art (for example, as found in Molecular Cloning, a Laboratory Manual, Eds. Sambrook, et al, Cold Spring Harbor Press.).
  • the isolated DNA from a sample is used in methods as described herein for detecting biomarkers indicative of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample.
  • the isolated DNA from the test and control samples are subjected to sequencing, for example next generation sequencing methodologies.
  • Sequence data from the test and the control DNA samples are compared, for example by aligning the two sequences, variant sequences are identified in the test sequence over the control sequence and the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample is identified based on said comparison.
  • isolated genomic DNA from a sample is used to identify variant mutations in a genetic sequence, wherein genes comprising variant mutations relative to a normal sample are associated with the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample.
  • a subset of biomarkers as found in Table 1 is used to determine the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample in a sample.
  • the subset can represent one or more, two or more, three or more, four or more, five or more, or six or more biomarkers with variant mutations from the subset of which is indicative of the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample in a sample.
  • the subset of biomarkers comprising AC008021.1, ADAMTS20, ANK3, ANKRD30A, ARID 1 A, ASPN, ATBFl, BRCA2, CACNA1S, CACNA2D1, CCDC141, CDH11, CHRM3, CNBD1, CAL3A1, CSMD3, CTNNB1, DNAH11, DNAH7, FAT3, GPR112, GRP115, GPR179, GRM3, HNRNPUL1, HSDL2, IL13RA1, KIAA1109, KIF24, KNDC1, LRRC53, LRRIQl, MFN1, MLL3, MUC16, MUC2, MUC4, MYH6, MY016, NEB, NHS, NIPBL, OR4A47, PIK3CA, PIKFYVE, PKHDl, RPGR, SACS, SDHA, SLC16A7, SLC5A7, THOC2, TIAM1, TP53, TPO, TRRAP, TTN, USP17
  • An additional subset comprising biomarkers TP53, PKHDl, PIK3CA, CACNA2D1, DNAH7, ARID 1 A, CTNNB1, CSMD3 and TRRAP may be useful for indicating the presence of ovarian cancer.
  • An additional subset encompassing biomarkers comprising TP53, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
  • CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl and TTN may be useful for indicating the presence of high grade serous type ovarian cancer.
  • An additional subset encompassing biomarkers comprising SLC16A7 is useful for indicating the presence of non-high grade serous type ovarian cancer.
  • biomarkers with variant mutations as identified in two or more patient samples from the sample cohort or a subset thereof as described herein may be correlated with the presence of ovarian cancer and/or a type of ovarian cancer (e.g., high grade serous, serous, non- high grade serous, non-serous, endometrioid) and/or stage of ovarian cancer (e.g., early stage serous versus late stage serous, etc.).
  • a type of ovarian cancer e.g., high grade serous, serous, non- high grade serous, non-serous, endometrioid
  • stage of ovarian cancer e.g., early stage serous versus late stage serous, etc.
  • the gene Mucin 2 (MUC2) is characterized herein as a biomarker that was mutated in at least four patient samples and is herein associated with ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer. This gene has not been previously associated with the presence of ovarian cancer.
  • the gene Cyclic Nucleotide Binding Domain Containing Protein 1 (CNBD1) has been characterized herein as a biomarker that was mutated in at least three patient samples and is herein associated with ovarian cancer, a type of ovarian cancer, and/or stage of ovarian. This gene has not been previously associated with the presence of ovarian cancer.
  • DNAH7 The gene dynein, axonemal, heavy chain 7 (DNAH7) is characterized herein as a biomarker that was mutated in at least five validation tumor samples and had not been previously statistically associated with ovarian cancer.
  • DNAH7 belongs to the dynein heavy chain family and is a component of the inner dynein arm of ciliary axonemes. These axomenes are expressed in the cytoplasm of bronchial epithelial cells and serve as the force generating protein of respiratory cilia by the motion of their microtubules.
  • Dynein has ATPase activity and the force producing power strike is contemplated to occur on release of ADP.
  • the fallopian tube consists
  • ovarian fimbriae predominately of two types of cells of which ciliated cells predominate throughout the tube, particularly in the ampulla and infundibulum which is surrounded by fimbriae and this includes the ovarian fimbriae that is attached to the ovary.
  • the majority of ovarian cancers can be classified as "epithelial” in type, however it has been suggested that the fallopian tube could also be the source of some types of ovarian cancer (2008, Piek et al, Adv Exp Med Biol 622:79-87; incorporated herein by reference in its entirety).
  • tubal fimbria is viewed as being a preferred site for early adenocarcinoma in women with familial ovarian cancer (2006, Medeiros et al, Am J Surg Pathol 30:230-236); incorporated herein by reference in its entirety). It is herein contemplated that, as DNAH7 is expressed in the ciliated cells of the fallopian tube and fimbriae, somatic variations within the gene may give rise to ovarian cancer and/or increase the aggressiveness of ovarian cancer. Diagnostic methods utilizing biomarkers as described herein are contemplated to be useful for identifying the presence or absence of ovarian cancer in a patient, the type of ovarian cancer present and/or the stage of ovarian cancer present in a patient.
  • epithelial ovarian carcinomas e.g., serous, mucinous, endometrioid, clear cell, transitional cell, squamous cell, undifferentiated and mixed epithelial tumors
  • the difficulty in diagnosing early stage epithelial ovarian carcinomas is high and the disease is typically left undiagnosed thereby leading to poor overall prognosis once the late stage carcinoma is diagnosed.
  • Diagnostic methods utilizing biomarkers as described herein are contemplated to provide for early stage disease diagnosis thereby providing a patient with a more favorable prognostic outcome.
  • Prognostic methods utilizing biomarkers as described herein are contemplated to be useful for determining a proper course of treatment for a patient having ovarian cancer.
  • a course of treatment refers to the therapeutic measures taken for a patient after diagnosis or after treatment for ovarian cancer. For example, a determination of the likelihood for cancer recurrence, spread, or patient survival, can assist in determining whether a more conservative or more radical approach to therapy should be taken, or whether treatment modalities should be combined. For example, when ovarian cancer recurrence is likely, it can be advantageous to precede or follow surgical treatment with chemotherapy, radiation, immunotherapy, biological modifier therapy, gene therapy, vaccines, and the like, or adjust the span of time during which the patient is treated.
  • a diagnosis or prognosis of an ovarian cancer state is contemplated to be correlated with one or more, for example a particular combination, of biomarkers described herein comprising two or more gene mutations.
  • Methods utilizing biomarkers as described herein are contemplated to be useful for monitoring recurrence of ovarian cancer following or during a course of treatment in a patient diagnosed with ovarian cancer.
  • Treatment of a patient diagnosed with ovarian cancer includes, but is not limited to, surgery, chemotherapy, radiation, immunotherapy, biological modifier therapy and gene therapy.
  • the methods utilizing biomarkers as described here are contemplated to be useful in determining the success, failure and/or progress of the treatment regimen.
  • Such information can be used by a clinician to determine a change in treatment course and/or future action or treatment for a patient.
  • a sample is obtained, nucleic acids for example DNA are isolated from the sample by established means known in the art, and the isolated nucleic acids are assayed by methods described herein.
  • a normal or control sample is typically obtained for comparison with the test sample.
  • Methods described herein are contemplated for use in, for example, characterizing the variant mutational status of one or more biomarkers as found in Table 1, wherein the variant mutational status of one or more biomarkers is useful in determining ovarian cancer status.
  • methods for characterization comprise sequencing technologies, for example next generation sequencing technologies.
  • microarray based technologies are utilized to characterize the mutational status of a biomarker as described herein for determining the status of ovarian cancer in a sample.
  • a sample is assayed for methylation status, the data of which is used to characterize a sample for ovarian cancer status.
  • Isolated genomic DNA from samples is typically modified prior to characterization.
  • genomic DNA libraries are created which are applied to downstream detection applications.
  • a library is produced, for example, by performing the methods as described in the NexteraTM DNA Sample Prep Kit (Epicentre® Biotechnologies, Madison WI), GL FLX Titanium Library Preparation Kit (454 Life Sciences, Branford CT), SOLiDTM Library Preparation Kits (Applied BiosystemsTM Life Technologies, Carlsbad CA), and the like.
  • the sample as described herein may be further amplified for sequencing by, for example, multiple stand displacement amplification (MDA) techniques.
  • MDA multiple stand displacement amplification
  • an amplified sample library is, for example, prepared by creating a DNA library as described in Mate Pair Library Prep kit, Genomic DNA Sample Prep kits or TruSeqTM Sample Preparation and Exome Enrichment kits (Illumina®, Inc., San Diego CA).
  • Useful cluster amplification methods are described, for example, in U.S. Patent No. 5,641,658; U.S. Patent Publ. No. 2002/0055100; U.S. Patent No. 7, 115,400; U.S. Patent Publ. No. 2004/0096853; U.S. Patent Publ. No. 2004/0002090; U.S. Patent Publ. No. 2007/0128624; and U.S. Patent Publ. No.
  • Genomic DNA libraries derived from a sample as described herein can be characterized for ovarian cancer status by sequencing for the presence of gene mutations. In one embodiment, sequencing can be performed following
  • BiosystemsTM Life Technologies (ABI PRISM® Sequence detection systems, SOLiDTM System), Ion Torrent® Life Technologies (Personal Genome Machine sequencer) further as those described in, for example, in United States patents and patent applications 5,888,737, 6, 175,002, 5,695,934, 6,140,489, 5,863,722,
  • current technology typically utilizes a light generating readable output, such as fluorescence or luminescence, however the present methods for detecting mutations in a biomarker for determining ovarian cancer status in a sample is not necessarily limited to the type of readable output as long as differences in output signal for a particular sequence of interest can be determined.
  • analysis software examples include, but are not limited to, Pipeline, CASAVA and GenomeStudio data analysis software (Illumina®, Inc.), SignalMap and NimbleScan data analysis software (Roche NimbleGen), GS Analyzer analysis software (454 Life Sciences), SOLiDTM, DNASTAR® SeqMan® NGen® and Partek® Genomics SuiteTM data analysis software (Life Technologies), Feature Extraction and Agilent Genomics Workbench data analysis software (Agilent Technologies), Genotyping ConsoleTM, Chromosome Analysis Suite data analysis software (Affymetrix®).
  • a skilled artisan will know of additional numerous commercially and academically available software alternatives for data analysis for sequencing generated output.
  • Embodiments described herein are not limited to any data analysis method.
  • the number of mutations in a biomarker can be detected using microarray methodologies.
  • a plurality of different probe molecules can be attached to a substrate or otherwise spatially distinguished in an array.
  • Exemplary arrays that can be used to detect the number of mutations in a biomarker include, but are not limited to, slide arrays, silicon wafer arrays, liquid arrays, bead-based arrays and others known in the art or set forth in further detail below.
  • the methods can be practiced with array technology that combines a miniaturized array platform, a high level of assay multiplexing, and scalable automation for sample handling and data processing.
  • Exemplary methods and systems for microarray analysis includes, but is not limited to, those methods and systems commercialized by Roche NimbleGen,
  • An array of beads can also be in a fluid format such as a fluid stream of a flow cytometer or similar device.
  • fluid formats for distinguishing beads include, for example, those used in XMAPTM technologies from Luminex or MPSSTM methods from Lynx Therapeutics.
  • microarray methods and systems can be found in, for example, US patents 5,856,101, 5,981,733; 6,001,309; 6,023,540, 6, 110,426, 6,200,737, 6,221,653; 6,232,072, 6,266,459, 6,327,410, 6,355,431, 6,379,895, 6,429,027, 6,458,583, 6,667,394 6,770,441, 6,489,606 and 6,859,570, 7, 106,513, 7,126,755, and 7, 164,533, US patent applications 2005/0227252, 2006/0023310, 2006/006327, 2006/0071075, 2006/0119913 and PCT publications WO98/40726, W099/18434, WO98/50782, WO00/63437, WO04/024328 and WO05/033681 (each of which is incorporated herein by reference in their entireties). Microarray based technologies for characterizing ovarian cancer are contemplated to be useful either
  • PCR polymerase chain reaction
  • a plurality of probes may be utilized to amplify genomic regions contemplated to comprise one or more somatic variants wherein amplification data are used to determine the presence or absence of variants thereby associating a sample with ovarian cancer, or not, as the case may be.
  • Exemplary methods and systems for PCR analysis include, but are not limited to those methods and systems commercialized by Illumina®, Inc., Roche Applied Science and Applied BiosystemsTM. Polymerase chain reaction based technologies for characterizing ovarian cancer are contemplated to be useful either alone or in combination with other diagnostic and/or prognostic assays.
  • the genes described herein were identified from a group of 25 early-stage ovarian cancers, of predominantly serous histology. Twenty-five patients with known ovarian cancer were enrolled in a study to identify ovarian cancer related biomarkers under approved protocol from the Internal Review Board at the Mayo Foundation for Medical Education and Research, Rochester, MN. One of the 25 patient tumor samples was determined to represent a hypermutated tumor phenotype. Blood samples were collected from the enrolled patients and processed under approved IRB protocols. The blood sample from a patient was matched with its corresponding tissue sample and served as the patient matched normal sample for that patient in determining gene associated variant mutations.
  • Genomic DNA was extracted and isolated from a portion of a fresh frozen biopsy tissue sample using the Gentra® Puregene® Tissue Kit (QIAGEN, Inc., Valencia CA) following manufacturer's protocol.
  • a blood sample was drawn from each patient by venous puncture into a Vacutainer® tube containing EDTA. The tube was centrifuged to separate the white blood cells, or buffy coat, from the red blood cells.
  • DNA was extracted from 100- 200 ⁇ 1 of buffy coat cells using the AutoGen prep 965 (AutoGen, Holliston MA) or Gentra® Puregene® Tissue Kit following manufacturer's protocol.
  • the genomic blood DNA was matched with its respective tissue isolated genomic DNA for each patient.
  • Genomic DNA libraries were generated by adding 4 ⁇ g of sample DNA to the Paired End Sample prep kit PE-102-1001 (Illumina®, Inc.) following manufacturer's protocol. Briefly, DNA fragments are generated by random shearing and conjugated to a pair of oligonucleotides in a forked adaptor configuration. The ligated products are amplified using two oligonucleotide primers, resulting in double-stranded blunt- ended products having a different adaptor sequence on either end.
  • Clusters were formed prior to sequencing using the V3 cluster kit (Illumina®, Inc.). Briefly, products from a DNA library preparation are denatured and single strands annealed to complementary oligonucleotides on the flow-cell surface. A new strand is copied from the original strand in an extension reaction and the original strand is removed by denaturation. The adaptor sequence of the copied strand is annealed to a surface-bound complementary oligonucleotide, forming a bridge and generating a new site for synthesis of a second strand. Multiple cycles of annealing, extension and denaturation in isothermal conditions resulted in growth of clusters, each approximately 1 ⁇ in physical diameter.
  • the DNA in each cluster is linearized by cleavage within one adaptor sequence and denatured, generating single- stranded template for sequencing by synthesis (SBS) to obtain a sequence read.
  • SBS sequencing by synthesis
  • the products of read 1 are removed by denaturation, the template is used to generate a bridge, the second strand is re-synthesized and the opposite strand is then cleaved to provide the template for the second read.
  • the Genome Analyzer IIx is designed to perform multiple cycles of sequencing chemistry and imaging to collect sequence data automatically from each cluster on the surface of each lane of an eight-lane flow cell.
  • the aligned reads were aggregated and sorted into chromosomes based on alignment positions.
  • the sorted reads were used to call variants using Hyrax, a Bayesian S V caller and GROUPER.
  • the callers are part of the standard CASAVA 1.8 distribution and were run with default parameters. This process was carried out for the tumor and normal genomes.
  • Somatic single nucleotide variant subtraction for calling somatic SNVs was performed by taking the list of positions in the tumor genome with snp quality values greater than 15 (Q(snp)tumor >15) and high confidence of the assigned genotype given the polymorphic prior (Q(max_gt)tumor >20). For each putative SNV the normal sample was investigated. If a call was present in the normal sample at the same position as a putative SNV, and if the call had a quality value greater than 0 (Q(snp) normal >0), the position was filtered out as background.
  • the putative SNVs were recalled (using Hyrax) in the tumor sample, however for recalling additional information from candidate indel contigs constructed from the normal sample was used. This process was utilized to avoid any indels that were initially missed in the tumor due to low supporting evidence. A candidate SNV was called when there was complete agreement between the initial SNV call and the recall.
  • variants were also recalled in the normal sample, using Hyrax with all read filtering turned off. If the posterior probability of the tumor genotype was higher than a non-reference genotype, then that SNV was considered to have low confidence evidence in the normal sample and was discarded.
  • insertion/deletions only those indels that were confidently called in the tumor sample and not present in the normal sample were considered. Indels in the tumor with a Q score less than 30 (Q(indel) ⁇ 30) and those positions that had less than 10 reads coverage in the normal sample (for a 3 OX build) were filtered out. To be considered as evidence, a read had a single read alignment score >10 and a paired read alignment score >90. Positions were excluded if they mapped within 1000 bases of a known centromere or telomere (as obtained from the reference genome
  • the indel position was matched with the region and calls present in the normal sample. If a putative somatic indel position overlapped with an indel call originating from the normal sample, the indel was considered to be present in normal germline and hence the position was filtered out.
  • the putative somatic indel region was characterized by finding the shortest sequence around the indel that extended outside any repeats and that region was matched with each intersecting read in the normal sample. If there was evidence in the normal sample having the same pattern in the intersecting normal reads the candidate somatic indel was discarded. To account for homopolymer slippage (i.e., the sequencing polymerase misincorporates, or misses, one or more bases in a sequence of identical bases) during a run, one read was allowed in the normal sample to have the same indel as found in the tumor. After sequence calling, each class of variant was annotated against the Ensembl database release e59. Each somatic variant was queried for overlapping annotated features.
  • Regulatory regions e.g., 3' and 5' untranslated regions
  • Table 10 is exemplary of the variant mutations identified in gene coding regions from the different patient samples.
  • biomarkers identified in Table 9 wherein non-coding regions variants identify one or more biomarkers, to visualize the entire spectrum of genomic somatic mutations (e.g., coding regions, conserved regions, non-coding regions, repeat
  • EXAMPLE 4- Validation cohort sample sequencing and data analysis A validation cohort of 40 tumor and matched normal sample pairs (FFPE) were obtained and processed as described in Example 1 to provide genomic DNA for sequencing.
  • the study population comprised female adults from 18-80 years of age.
  • Some cytosine bases in DNA extracted from FFPE samples can be subject to random deamination resulting in a cytosine base being converted to uracil and in turn misread as a thymidine base. This conversion typically occurs randomly in approximately 0.5- 1.5% of all cytosine bases found in FFPE derived DNA.
  • the random deamination can lead to false positive predictions of somatic OT or G>A mutations during sequencing.
  • the validation design was a 1000X deep sequencing of a targeted pull down (i.e., enrichment).
  • a targeted pull down is an experimental design which targets a subset of the genome by designing complimentary DNA probes which hybridize to, or pull down, randomly fractionated genomic DNA (i.e., a DNA library as described in Example 1).
  • the hybridized DNA can be enriched for fragments that overlap the probe region and the enriched DNA eluted and sequenced to high depth.
  • the targeted pull down was achieved by using an Illumina, Inc. TruSeq Custom Enrichment Array Kit for in solution capture for targeting and enriching user selected human genomic regions.
  • AVP3, AVP4 and AVP5 were designed, each containing the content of previous generations plus additional probes to additional targeted genes.
  • the probes were designed to target previously defined DNA gene regions (i.e., such as those genes identified in Tables 1 and 2), or regions in genes subsequently identified as potential ovarian cancer gene biomarkers.
  • Table 1 1 describes the enrichment design for targeted genes in the validation cohort.
  • the enrichment design AVP3 utilized 4071 probes which targeted 1200213 bases over 215 genes (Table 12).
  • the enrichment design AVP4 utilized 5097 probes which targeted 1473936 bases over 245 genes (Table 12) and enrichment design AVP5 utilized 5490 probes which targeted a total of 1578057 bases over 271 genes (Table 12).
  • the probe designs for the TruSeqTM Custom Enrichment Array Kit consisted of a subset of probes taken from the TruSeqTM Exome Enrichment Kit, the probe design data of which is available within the kit. Further, a skilled artisan following known and established protocols will understand how to define probes for targeted enrichment assays. Table 12 describes the genes targeted, the enrichment assay and whether they were enriched in assay design AVP3, 4 and/or 5. Table 1 1- Enrichment assay design for validation cohort samples
  • sequence data was aligned to human reference genome HG19/GRCh37 using CASAVA-1.8 data analysis software and all optical and PCR duplicates were tagged as duplicates by the software using default parameters.
  • Tumour specific mutations were detected in each sample by selecting those reads with properly matched Rl and R2 read pairs as determined by CASAVA (i.e., those reads wherein the orientation of Rl and R2 read pairs and the size of the DNA fragment were within normal limits) and which were not tagged as duplicate reads, and by subtracting matching normal tissue variants from the tumor variants using the Strelka method for somatic SNV and small indel detection from sequencing data of matched
  • VEP Variant Effect Predictor
  • splice region variant were considered as having a functional impact on the downstream protein of any gene overlapping the mutations (see VEP documentation http://www.ensembl.org/info/docs/variation/predicted_data.html. All other exonic variants were classified as non- functional P-values from Fisher's exact test using number of functional versus all (functional and non- functional) mutations annotated to the gene (excluding intronic mutations).
  • a Condel score was also utilized to integrate the output of computational tools aimed at assessing the impact of non-synonymous SNVs on protein function by computing a weighted average of the scores (WAS) of computational tools, such as SIFT, Polyphen2, MAPP, LogR Pfam e-value (2004, Clifford et al, Bioinformatics 20: 1006-1014; incorporated herein by reference in its entirety) and MutationAssessor.
  • WAS weighted average of the scores
  • the scores of different methods are weighted using the complementary cumulative distributions produced by the five methods on a dataset of approximately 20000 missense mutations, both deleterious and neutral.
  • the probability that a predicted deleterious mutation is not a false positive of the method and the probability that a predicted neutral mutation is not a false negative are employed as weights.
  • a high Condel score, for example above 0.5, represents mutations more likely than not to be deleterious whereas a low Condel score represents the opposite (2011,

Landscapes

  • Chemical & Material Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Organic Chemistry (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Engineering & Computer Science (AREA)
  • Immunology (AREA)
  • Pathology (AREA)
  • Analytical Chemistry (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • Genetics & Genomics (AREA)
  • Hospice & Palliative Care (AREA)
  • Biochemistry (AREA)
  • Microbiology (AREA)
  • Molecular Biology (AREA)
  • Biophysics (AREA)
  • Physics & Mathematics (AREA)
  • Oncology (AREA)
  • Biotechnology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
  • Other Investigation Or Analysis Of Materials By Electrical Means (AREA)
  • Investigating Or Analysing Biological Materials (AREA)

Abstract

The present disclosure provides biomarkers and methods of their use in determining the presence or absence of cancer. In preferred embodiments, biomarkers and methods of their use determine the presence or absence of ovarian cancer, a type of ovarian cancer and/or a stage of ovarian cancer in a sample from an individual.

Description

OVARIAN CANCER BIOMARKERS
RELATED APPLICATIONS This application claims priority to U.S. Provisional Application Serial No.
61/469,510 filed on March 30, 2011, which is incorporated herein by reference in its entirety.
BACKGROUND
Ovarian cancer is the among the top ten most common cancers among women, and the fifth leading cause of death for women with cancer in the United States.
Worldwide, the most lethal gynecological disease among women in developed countries is ovarian cancer. The American Cancer Society estimates that about 22,000 new cases will be diagnosed this year and approximately 15,000 women will die from ovarian cancer in the United States alone. The incidence rate of ovarian cancer is roughly 13 per 100,000 women per year and even though the median age at diagnosis is around 63, no age group is immune to the disease. Further, even though incidence of ovarian cancer is slightly higher among white women, no race or ethnic background is immune.
Survival rate once a diagnosis is made is typically dismal, for example if ovarian cancer is detected and effectively treated prior to metastasis the 5 year survival rate can be as high as 73%, however if the cancer is not detected until it has metastasized then the long term survival rate drops to <30%. Unfortunately, ovarian cancer often goes undetected until it has metastasized within the pelvis and abdomen, therefore the outcome for most women is grim as the cancer is difficult to treat and is often fatal.
As such, early detection ovarian cancer is crucial to benefit those patients, for example, that present with no or vague symptoms or with tumors that are below the level of detection during a physical examination. A considerable amount of research effort has been focused in discovering and developing early detection systems, however to date no effective screening method has been developed. As such, what are needed are ways to detect ovarian cancer, preferably at an early stage, for example before metastasis, thereby improving the long term survival of those women afflicted with this disease.
SUMMARY
The present disclosure identifies biological markers, or biomarkers, indicative of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer. The disclosed biomarkers allow for identification of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject. The methods described herein utilize the biomarkers and provide alternatives to currently available ovarian cancer determinative, diagnostic and prognostic methodologies. The biomarkers and methods of their use as disclosed herein can be applied to the characterization, classification, differentiation, grading, staging, diagnosis, or prognosis of ovarian cancer, ovarian cancer type and/or ovarian cancer stage. Embodiments as described herein are based in part on the identification of reliable biomarkers for the improved determination, screening, diagnosis and/or prognosis of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject. The disclosure provides a population of gene targets or gene related targets (i.e., biomarkers) and methods of use as described herein. Biomarkers and methods of their use for determining ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject comprise TP53, MUC2, MY016, ARID1A, CT B 1, CSMD3, TRRAP, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1 S, CACNA2D1, PIKFYVE, PIK3CA, NIPBL, PKHD1, MLL3, ATBF 1, TPO, DNAH7, LRRIQ1, SDHA, TIAM1, TTN, SLC16A7, COL3A1 and HNRNPUL1.
In some embodiments, methods for determining the presence of ovarian cancer in a subject comprises providing a nucleic acid sample from a subject, detecting the presence of one or more variant mutations in two or more genes selected from the group consisting essentially of ARID 1 A, CACNA2D1, CTTNB1, PKHD1, DNAH7, PIKC3A, TRRAP and CSMD3 in the sample, evaluating the probability that that the two or more genes are correlated with ovarian cancer, and determining the presence of ovarian cancer in the sample based on the probability of correlation. In some embodiments, evaluating the probability comprises comparing the nucleic acid sample suspected of having ovarian cancer to a matched normal nucleic acid sample, wherein both the nucleic acid samples (i.e., test and normal) are genomic DNA samples. In preferred embodiments, a gene is correlated with ovarian cancer at p<0.05. In preferred embodiments, the gene DNAH7 is one of the two or more genes identified as being correlated with ovarian cancer at p<0.05. In some embodiments, genomic DNA for testing for presence of ovarian cancer from a subject is isolated from a sample selected from a group consisting of a tissue sample, a biopsy sample, a cell sample, a circulating tumor cell sample, a fixed tissue sample, a frozen tissue sample or a lavage sample, In some embodiments, methods for determining the presence of ovarian cancer in a subject comprises providing a nucleic acid sample from a subject, detecting the presence of one or more variant mutations in two or more genes selected from the group consisting of ARID1A, CACNA2D1, CTTNB1, PKHD 1 , DNAH7, PIKC3 A, TRRAP and CSMD3 , evaluating the probability that that the two or more genes are correlated with ovarian cancer, and determining the presence of ovarian cancer in the sample based on the probability of correlation. In preferred embodiments, detecting mutations in genes correlated with ovarian cancer comprises sequencing, such as sequence by synthesis methodologies, microarray analysis and/or polymerase chain reaction methodologies including, but not limited to real time PCR and/or quantitative PCR.
In some embodiments, methods are described herein for determining the presence of ovarian cancer in a sample comprising creating a DNA library from nucleic acids derived from a sample suspected of having ovarian cancer, sequencing the DNA library to identify one or more mutations in two or more genes from the list consisting essentially of ARID 1 A, CACNA2D1, CTTNB1, PKHD l, DNAH7, PIKC3A, TRRAP and CSMD3, computationally determining the probability that the two or more genes are correlated with ovarian cancer, and determining the presence of ovarian cancer in said sample based on the probability of correlation. In some embodiments, computational determination of ovarian cancer correlated comprising comparing the sequence of the nucleic acid sample suspected of having ovarian cancer to a matched normal sample. In preferred embodiments, a gene is correlated with ovarian cancer at p<0.05. In other preferred embodiments, DNAH7 is one of the two or more genes identified as being correlated with ovarian cancer. In some embodiments, a test sample suspected of having ovarian cancer is isolated from a sample selected from the group consisting of a tissue sample, a biopsy sample, a cell sample, a circulating tumor cell sample, a fixed tissue sample, a frozen tissue sample or a lavage sample, wherein a matched normal sample may be isolated from either a blood sample or a tissue sample that does not have ovarian cancer. In preferred embodiments, a DNA library is sequenced using sequence by synthesis
methodologies.
In embodiments as described herein, a method for confirming the presence of ovarian cancer in a sample comprises comparing the sequence of a nucleic acid sample derived from a tissue suspected of having ovarian cancer with the sequence of a nucleic acid from a tissue not suspected of having ovarian cancer, wherein the presence of one or more variant sequences identified in two or more of ARIDIA, CACNA2D1, CTTNB l, PKHDl, DNAH7, PIKC3A, TRRAP and CSMD3 in the test sample as compared to the normal sample indicates the presence of ovarian cancer in a sample.
In some embodiments, methods are described herein for determining the presence of ovarian cancer in a sample comprising evaluating a sample for the presence of one or more variant mutations in two or more genes selected from the group comprising ARIDIA, CTNNBl, CSMD3, TRRAP, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl, TTN, SLC16A7, COL3A1 and HNRNPUL1 and determining the presence of ovarian cancer in the sample based on said evaluation. In some embodiments, a sample under evaluation is compared to a matched normal nucleic acid sample. In some embodiments, the nucleic acid sample is genomic DNA from a human subject as is the matched normal nucleic acid sample. In some embodiments, a genomic DNA sample is isolated from a tissue sample, a biopsy sample a cell sample, a circulating tumor cell sample, a fixed tissue sample, a frozen tissue sample or a lavage sample.
In some embodiments, the presence of one or more variant mutations is found in two or more genes selected from the group consisting essentially of MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl, TTN, SLC16A7, COL3A1 and HNRNPULl. In some embodiments, the presence of one or more variant mutations is found in two or more genes selected from the group consisting of MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl, TTN, SLC16A7, COL3A1 and HNRNPULl. In some embodiments, determining ovarian cancer comprises determining the presence of high grade serous ovarian cancer, whereas in other embodiments determining ovarian cancer comprises determining the presence of non-high grade serous ovarian cancer. In some embodiments, determining ovarian cancer comprises determining the presence of serous ovarian cancer.
In some embodiments, the presence of one or more variant mutations is found in two or more genes selected from the group consisting of MUC2, MY016,
OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl, and TTN, further comprising determining the presence of high grade serous ovarian cancer based on the presence of the variant mutations. In some embodiments, the presence of one or more variant mutations is found in a group comprising SLC16A7, wherein the presence of one or more variant mutation determines the presence of non-high grade serous ovarian cancer in a sample.
In some embodiments, the presence of one or more variant mutations is found in two or more genes selected from the group consisting essentially of ARIDIA, CTNNB1, CSMD3, TRRAP, PIKC3A, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl, TTN, SLC16A7, COL3A1 and HNRNPULl, the presence of which determines the presence of serous ovarian cancer. In some embodiments, a sample is evaluated for the presence of one or more variant mutations in the group comprising MUC2, MYO 16, HNRNPULl and COL3A1, the presence of which determines the presence of serous ovarian cancer. In some embodiments, a sample is evaluated for the presence of one or more variant mutations in the group consisting of MUC2, MYO 16, HNRNPULl and COL3A1, the presence of which determines the presence of serous ovarian cancer.
In some embodiments, methods are described herein for determining the presence of high grade serous ovarian cancer in a sample comprising evaluating the sample for the presence of one or more variant mutations in two or more genes selected from the group comprising MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl and TTN. In some embodiments, evaluating the sample for the presence of one or more variant mutations comprises evaluating the nucleic acid sample for one or more variant mutations in one or more of MUC2 and MY016 and one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl and TPO. In some embodiments, the evaluation further comprises evaluating one or more of DNAH7, LRRIQl, SDHA, TIAMl and TTN.
In some embodiments, methods are described herein for determining the presence of serous ovarian cancer in a sample comprising evaluating the sample for the presence of one or more variant mutations in two or more genes selected from the group comprising MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, COL3A1, CACNA1S, CACNA2D1, PIKFYVE, NIPBL,
PKHDl, MLL3, ATBFl, HNRNPULl, TPO, DNAH7, LRRIQl, SDHA, TIAMl and TTN. In some embodiments, evaluating the sample for the presence of one or more variant mutations comprises evaluating the nucleic acid sample for one or more variant mutations in one or more of HNRNPULl and COL3A1 and one or more of MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl and TTN. In some embodiments, evaluating the sample for the presence of one or more variant mutations comprises evaluating the nucleic acid sample for one or more variant mutations in one or more of HNRNPULl and COL3 A 1 , one or more of MUC2 and MYO 16 and one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
CACNA2D 1 , PIKFYVE, NIPBL, PKHD 1 , MLL3 , ATBF 1 and TPO. In some embodiments, evaluating the sample for the presence of one or more variant mutations comprises evaluating the nucleic acid sample for one or more variant mutations in one or more of HNRNPUL1 and COL3A1, one or more of MUC2 and MY016, one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1 S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF 1 , TPO and one or more of DNAH7, LRRIQ 1 , SDHA, TIAM 1 and TTN.
In some embodiments, evaluating a nucleic acid sample for the presence of ovarian cancer, high grade serous ovarian cancer or serous ovarian cancer comprises sequencing the nucleic acid sample. In some embodiments, sequencing a sample comprises sequence by synthesis methodologies. In some embodiments, evaluation of a nucleic acid sample for the presence of ovarian cancer, high grade serous ovarian cancer or serous ovarian cancer is carried out on a microarray. In some embodiments, evaluating a nucleic acid sample for the presence of ovarian cancer, high grade serous ovarian cancer or serous ovarian cancer comprises performing polymerase chain reaction on the sample, for example quantitative or real-time PCR.
FIGURES
Figure 1 is exemplary of the number of high grade serous ovarian cancer patient samples (n=17) from a patient cohort exhibiting variant mutations in a gene compared to the patient samples that were not high grade serous (n=8). Figure 2 is exemplary of the number of serous ovarian cancer patient samples
(n=19) from the patient cohort exhibiting variant mutations in a gene compared to the patient samples that were not serous type (n=6).
DEFINITIONS As used herein, the term "sample" is intended to mean any biological fluid, cell, tissue, organ or portion thereof that contains genomic nucleic acids, for example genomic DNA or RNA, suitable for mutational detection via the disclosed methods. A test sample can include or be suspected to include a cell, such as a cell from an ovary, uterus, fallopian tube, vagina, or other organ or tissue that contains or is suspected to contain a cancerous cell. The term includes samples present from an individual as well as samples obtained or derived from an individual. For example, a sample can be a histologic section of a specimen obtained by biopsy, cell scraping, etc. or cells that are placed in or adapted to tissue culture. A sample further can be a sub-cellular fraction or extract, or a crude or isolated nucleic acid molecule. A patient matched normal sample can be used to establish a mutational background for comparison to a patient test sample.
A sample may be obtained in a variety of ways known in the art. Samples may be obtained according to standard techniques from all types of biological sources that are usual sources of genomic DNA including, but not limited to cells or cellular components which contain DNA, cell lines, circulating tumor cells, biopsies, bodily fluids such as blood, lavage specimens, tissue samples such as tissue that are formalin fixed and embedded in paraffin such as tissue from ovaries, endometrium, cervix, fallopian tubes, omentum, histological object slides, and all possible combinations thereof. Further, tissues can be fresh, fresh frozen, etc. Accordingly, a sample can be from an archived, stored or fresh source as suits a particular application of the methods set forth herein. In particular embodiments, the methods described herein can be performed on one or more samples from ovarian cancer patients such as samples obtained by vaginal lavage, endometrial biopsy, ovarian biopsy, and/or blood draw. Sample analysis can be applied, for example, to the presence or absence of ovarian cancer, differentiation between early and/or late stage ovarian cancer types, ovarian cancer epithelial type differentiation, or to monitor cancer progression or response to treatment.
A suitable sample can be collected and acquired that is either known to comprise ovarian cancer cells or is subsequent to the formulation of the diagnostic aim of a biomarker as disclosed herein. A sample can be derived from a population of cells or from a tissue that is predicted to be afflicted with or phenotypic of ovarian cancer. The genomic DNA can be derived from a high-quality source such that the sample contains only the tissue type of interest, minimum contamination and minimum DNA fragmentation. In particular, samples are contemplated to be representative of the tissue or cell type of interest that is to be handled by an assay. In addition, a population or set of samples from an individual source can be analyzed to maximize confidence in the results for an individual. In some embodiments, a sample from an individual is matched and compared to a normal sample from that same individual to identify the mutational status of biomarkers for that individual. The normal sample, or patient matched normal sample, can be from the same or similar organ, tissue or fluid as the sample to which it is compared. The normal sample will typically display a phenotype that is different from a phenotype of the sample to which it is compared.
As used herein, the term "isolated" or "purified" when used in relation to a nucleic acid refers to a nucleic acid sequence that is extracted and separated from at least one component or contaminant with which it is ordinarily associated in its natural source. As such, an isolated or purified nucleic acid is present in a form or setting that is different from that in which it is found in nature.
As used herein, the terms "marker" or "biomarker" can be DNA or RNA, proteins, polypeptides, variants, fragments or functional equivalents thereof. In the present disclosure, a biomarker is generally associated with a genomic nucleic acid such as a gene or gene associated region or location unless specified otherwise.
Biomarkers disclosed herein that are associated with ovarian cancer, a particular type of ovarian cancer and/or a particular stage of ovarian cancer comprise one or more single nucleotide variants and/or insertions/deletions (indels) located in a gene or gene associated region as compared to its equivalent in a normal sample. A gene that contains one or more somatic mutations, such as variant mutations, identified in two or more patient samples is contemplated to be a biomarker that is useful in detecting, diagnosing or prognosing ovarian cancer, a particular type of ovarian cancer and/or a stage of ovarian cancer.
As used herein, the term "somatic mutation" is an alteration in the genome that occurs after conception resulting in a genetic difference of the genome at that particular location. Somatic mutations can occur in any cell in the body except the germ cells and are passed to the cell progeny during cell division. Somatic mutations include, but are not limited to, point mutations such as single nucleotide variants (SNVs), gene amplification or duplication, genetic insertions and/or deletions (indels), chromosomal translocations, chromosomal inversions and single nucleotide polymorphisms (SNPs). Somatic mutations can result in phenotypic changes, disease formation, cancer, etc. "Somatic mutations" as used herein, unless otherwise stated, include SNVs and/or indels present in genomic DNA, and are considered variant mutations, or mutations that may result in phenotypic changes, disease formation, cancer, etc.
Identification of somatic variants described herein was performed by looking for a SNV or indel in the patient cancer sample which was not present in the patient matched normal sample. If a somatic mutation was found in the cancer sample which was not present in the normal sample, then that mutation was identified as a SNV or indel, as the case may be. If a somatic mutation was present in both the cancer sample and the normal sample, then that mutation was considered part of that patient's genetic background and was not considered a variant. A gene from two or more patient samples that has one or more variant mutations is considered a biomarker and useful in detecting, diagnosing and prognosing ovarian cancer, a particular type of ovarian cancer and/or a stage of ovarian cancer.
The term "gene" refers to a nucleic acid sequence, such as DNA, that comprises coding sequences associated with the production of a polypeptide, precursor, or RNA (e.g., rRNA, tRNA). Typically, a gene also includes non-coding and intergenic sequences. The term can encompass the coding region of a gene and the sequences located adjacent to the coding region on both the 5' and 3' such that the gene corresponds to the length of the full-length mRNA. Sequences located 5' of the coding region and present on the mRNA are referred to as 5' non-translated sequences. Sequences located 3' or downstream of the coding region and present on the mRNA are referred to as 3' non-translated sequences. The term "gene" encompasses both cDNA and genomic forms of a gene. A genomic form or clone of a gene contains the coding region interrupted with non-coding sequences such as introns, intervening regions, intervening sequences or intergenic regions.
DETAILED DESCRIPTION
Ovarian cancer is often referred to as a silent killer because of its subtle symptoms that lead to delayed discovery, diagnosis and treatment. The majority of ovarian cancers are diagnosed when the cancer has already reached an advanced stage, for example >80% of serous ovarian cancers are diagnosed at Stage III or Stage IV leading to a very low chance of long-term survival in these patients. Screening and/or detecting ovarian cancer in women who might be at higher risk of developing ovarian cancer, such as those with a strong family history of such cancer, is problematic. The two most common screening tests for ovarian cancer include transvaginal sonography and identification of a protein marker, CA-125. However, both tests have limitations. For example, transvaginal sonography can identify a mass in the ovary however the sonogram is unable to distinguish whether the mass is cancerous or not. The protein marker CA-125 is not specific to the presence of ovarian cancer as other cancers also exhibit high levels of CA-125. The majority of ovarian tumor cancers are of the epithelial histologic type, which can be further divided into different tumor subtypes, for example serous, endometrioid, clear cell, mucinous, Brenner or transitional cell, squamous cell, undifferentiated and mixed epithelial cell types (AJCC Cancer Staging Manual 7th Ed., p.422). There are several different methods for grading or staging cancers. Perhaps the most clinically applied is the tumor node metastasis (TNM) staging system and/or the staging system as described by the Federation Internationale de Gynecologie et d'Obstetrique (FIGO).
Ovarian cancer can also be classified into two groups based on molecular progression. For example, Type I ovarian tumors of mucinous, clear cell, endometrioid, and low-grade serous type develop in stepwise fashion from adenomas to carcinomas, whereas Type II tumors of high-grade serous develop de novo from undefined precursor lesions and progress rapidly with no apparent stepwise progression (Ie and Kurman, 2004, Am J Pathol 164: 151 1-1518). Further explanation of cancer staging and grading can be found at, for example, AJCC Cancer Staging Manual, Edge, SB et al, Eds., Springer-Verlag, New York. The vast majority of
Type II, high-grade serous ovarian cancers (OvCa) are diagnosed at advanced stages and represent a major challenge in early detection (Chan et al, 2006, Obs and Gyn 108: 521-528).
The most common type of ovarian cancer arises from epithelial cells that line the surface of the ovary. Approximately 50% of epithelial ovarian tumors are classified as serous, or tumors with glandular features, and make up approximately 80% of all ovarian tumors. Other types of ovarian cancers can arise from germ cells (e.g., cancer of the ovarian egg-making cells) and sarcomas. High-grade serous tumors denote highly aggressive, invasive tumors as compared to low malignant potential (LMP) tumors. Whether an invasive serous tumor is classified as either high or low grade is based on the clinical course of the disease. For example, high grade serous tumors were found to over express genes that control various cellular functions associated with cancer cells, for example genes that control cell growth, DNA stability (or lack thereof) and genes that silence other genes. Conversely, LMP tumors were not found to overexpress these types of genes and LMP tumors were alternatively characterized by expression of growth control pathways, such as tumor protein 53 (TP53 or p53) pathways.
More recently, a two-tiered system of characterizing serous ovarian tumors has been described (Vang, et al, 2009, Adv Anat Pathol 16:267-282) based on studies performed Johns Hopkins Hospital and M.D. Anderson Cancer Center. Briefly, low grade serous ovarian tumors are characterized based in a number of criteria, for example low grade serous tumors have low to no chromosomal instability, typically have mutated KRAS, BRAF and ERBB2 genes, demonstrate slow tumor
development, typically have cell nuclei that are uniform, small and round and generally have low mitotic index. Conversely, a high grade serous tumor has a high degree of chromosomal instability, has mutated TP53 gene, demonstrates very fast tumor development, typically has nuclei that are non-uniform, enlarged and irregularly shaped and has high mitotic index.
However, staging and grading cancers are subjective and rely on a diagnostician to interpret morphology, histology, anatomy and other related indices. Further, as ovarian cancer is typically left undiagnosed until late stage cancer due to, for example, its asymptomatic phenotype, the staging and grading do nothing to identify early stage cancer, or identify ovarian cancer earlier in the disease progression in the absence of disease related symptoms. As such, there is a critical need for tools, methods and strategies that can be used for detecting, diagnosing, and prognosing ovarian cancer, a particular type of ovarian cancer (e.g., high grade serous, non-high grade serous, serous, non-serous, etc.) and/or a stage (e.g., early stage vs late stage) of ovarian cancer in a patient. Experiments were conducted as disclosed herein to identify biomarkers useful for detecting, diagnosing, and prognosing ovarian cancer, a particular type of ovarian cancer and/or a stage of ovarian cancer in a patient. Exemplary biomarkers are described in the Figures and Tables herein. An initial candidate list of approximately 268 gene markers (Table 1) was compiled from an original list of approximately 5900 genes that were investigated for their potential use as ovarian cancer related biomarkers.
Table 1-List of candidate 268 genes
Figure imgf000015_0001
CACNA1S chrl:201008642-201081694 MY016 chrl3:109248500-109860355
CACNA2D1 chr7:81575760-82073114 MY03A chrl0:26223196-26501456
CCDC141 chr2:179694996-179914813 NBN chrl:44401329-44402437
CCNE1 chrl9:30302901-30315216 NEB chr2:152341850-152591001
CCNYL2 chrl0:42903616-42970797 NF1 chrl7:29422328-29549008
CD47 chr3:107761941-107809935 NHS chrX:17393543-17754114
CDH1 chrl9:3506295-3538325 NIPBL chr5:36876861-37066515
CDH11 chrl6:64977656-65156101 NOTCH1 chr9:139388896-139440314
CDH12 chr5:21750777-22853731 NOTCH3 chr6:32162620-32191844
CDK12 chrl7:37617739-37690800 NRAS chrl:115247090-115259515
CDK4 chrl2:58142005-58146164 NTS chrl2:86268073-86276763
CDKN2A chr9:21967752-21995300 OLFML2B chrl:161952982-161993644
CHEK1 chrll:125495036-125546148 OR4A47 chrll:48510269-48511332
CHEK2 chr22:29083731-29138410 PALB2 chrl6:23614483-23652678
CHRM3 chrl:239549865-240078750 PARP4 chrl3:24995064-25086948
CHST2 chr3:142838173-142841800 PCDH15 chrl0:55562531-57387702
CNBD1 chr8:88218216-88394955 PCP4 chr21:41239243-41301322
CNGA1 chr4:47937994-48018689 PDGFRA chr4:55095264-55164414
CNTNAP2 chr7:145813453-148118090 PDGFRB chr5:149493400-149535423
CNTNAP5 chr2:124782864-125672864 PDK1 chr2:173420101-173463862
C0L11A1 chrl:103342023-103574052 PGR chrll:100900355-101000544
COL1A1 chrl7:48260650-48278993 PI3 chr20:43803517-43805185
COL3A1 chr2:189839046-189877472 PIK3C2B chrl:204391756-204463852
COL5A1 chr9:137533620-137736686 PIK3CA chr3:178865902-178957881
COMT chr22:19929130-19957498 PIK3CB chr3:138372860-138553780
CRABP1 chrl5:78632666-78640573 PIK3CD chrl:9711803-9788977
CRABP2 chrl:156669398-156675608 PIK3CG chr7:106505723-106547590
CSMD3 chr8:113235157-114449328 PIK3R1 chr5:67522462-67597649
CTNNB1 chr3:41236328-41301587 PIK3R2 chrl9:18263928-18281343
CUTC chrl0:101462315-101515891 PIK3R3 chrl:46505812-46642143
CXCL10 chr4:76942273-76944650 PIKFYVE chr2:209130991-209223475
DI02 chrl4:80664089-80677840 PIM2 chrX:48770459-48776301
DLG2 chrll:83166055-85338966 PKHD1 chr6:51480098-51952423
DNAH11 chr7:21582833-21941457 PKHD1L1 chr8:110374706-110543500
DNAH7 chr2:196602427-196935730 PLAU chrl0:75668935-75677255
DNMT1 chrl9:10244023-10305755 PMS1 chr2:190648811-190742355
DPP6 chr7:153584182-154685995 PMS2 chr7:6012870-6048756
DSP chr6:7541808-7586950 POSTN chrl3:38136722-38172981
EGFR chr7:55086714-55324313 PPP2R1A chrl9:52693274-52730687
EIF5A2 chr3:170606204-170626482 PRKCI chr3:169940153-170023769
EMR1 chrl9:6887582-6940463 PRRX1 chrl:170632323-170708560
ERBB2 chrl7:37844393-37884915 PTEN chrl0:89622870-89731687
ERBB3 chrl2:56473892-56497128 PTK2B chr8:27168999-27316903
ERCC1 chrl9:45912872-45927177 PTPRN2 chr7:157331750-158380480
ESR1 chr6:151977826-152450754 RAB25 chrl:156030951-156040295 EVI1 chr3:168801287-169381406 RAD50 chr5:131891711-131979752
FAP chrl:92711959-92764544 RAD51C chrl7:56769934-56811703
FAT3 chrll:92085262-92629636 RASAL1 chrl2:113537318-113574021
FBXW7 chr4:153242410-153456172 RBI chrl3:48877911-49056122
FGF1 chr5:141971743-142077635 RB1CC1 chr8:53535016-53627026
FGFR1 chr8:38281787-38314964 ROB02 chr3:75955826-77699115
FGFR2 chrl0:123237848-123357972 RP11-42302.6 chrl:142803532-142827235
FGFR3 chr4:1795034-1810598 RPGR chrX:38128424-38186817
FGFR4 chr5:176513887-176525145 RPS6KA2 chr6:166822852-167319939
FLNA chrX:153576892-153603006 RRM1 chrll:4115924-4160106
FN1 chr2:216225163-216300895 S100A1 chrl:153600402-153604513
F0LR1 chrll:71900602-71907342 SACS chrl3:23902965-24007841
FOX03 chr6:108881038-109005977 SCAP chr3:47455184-47518616
FRG1 chr4:190861943-190884359 SDHA chr5:218356-256815
FRG2B chrl0:135437399-135440299 SGCZ chr8:13947373-15095848
FRG2C chr3:75713481-75716371 SLC16A7 chrl2:60083126-60175407
GABRA6 chr5:161112721-161129112 SLC5A7 chr2:108602979-108630450
GJB2 chrl3:20761609-20767037 SLPI chr20:43880880-43883205
GLI2 chr2:121493199-121750229 SMAD4 chrl8:48556583-48611415
GPC5 chrl3:92050929-93519490 SPARC chr5:151040657-151066726
GPR112 chrX:135383122-135499047 SRC chr20:35973088-36034453
GPR115 chr6:47653600-47689757 STK11 chrl9:1205798-1228434
GPR179 chrl7:36481493-36499730 STK36 chr2:219536749-219567439
GRM3 chr7:86273224-86494200 TACC3 chr4:1723227-1746898
GSH2 chr4:54965690-54968672 TET1 chrl0:70320413-70454239
HERC1 chrl5:63900818-64126147 THBS2 chr6:169615875-169654139
HIF1A chrl4:62162119-62214976 TH0C2 chrX:122734412-122866906
HJURP chr2:234742062-234763212 TIAM1 chr21:32361860-32932290
HMGA2 chrl2:66218240-66360068 TMEM158 chr3:45265958-45267770
HNRNPUL1 chrl9:41768391-41813811 TMPRSS4 chrll:117947727-117990554
HNT chrll:131240373-132206716 TNK2 chr3:195590235-195638816
HOXB3 chrl7:46626232-46667634 TOPI chr20:39657458-39753127
HP chrl6:72088508-72094953 TOP2A chrl7:38544768-38574408
HPR chrl6:72097111-72111144 TOP2B chr3:25639475-25706398
HSDL2 chr9:115142217-115234690 TP53 chrl7:7565257-7590856
IGF1R chrl5:99192200-99507759 TPD52 chr8:80947107-81083836
IGFBP1 chr7:45927956-45933267 TPMT chr6:18128542-18155305
IGFBP2 chr2:217497551-217529159 TPO chr3:184089723-184095932
IGFBP3 chr7:45951951-45961473 TRRAP chr7:98475556-98610866
IGFBP4 chrl7:38599713-38613983 TSC1 chr9:135766735-135820020
IGFBP5 chr2:217536828-217560248 TSC2 chrl6:2097466-2138721
IL13RA1 chrX:117861535-117928502 TTN chr2:179390716-179695529
ILKAP chr2:239079042-239112370 TYMS chrl8:657604-673578
INPP4B chr4:142944313-143768585 UBAP2 chr9:33921691-34048947
ITGA5 chrl2:54789047-54813050 UBE2C chr20:44441215-44445596 JAK2 chr9:4985033-5128183 USP17L1P chr8:7189909-7191501
KCNIP4 chr4:20730239-21950422 VCAN chr5:82767284-82878122
KCNJ12 chrl7:21279699-21320404 VEGFA chr6:43737921-43754224
KIAA1109 chr4:123073488-123283907 WDFY4 chrl0:49892921-50191001
KIF24 chr9:34252379-34329198 WFDC2 chr20:44098346-44110172
KIT chr4:55524085-55606881 WT1 chrll:32409321-32457176
KLK5 chrl9:51446561-51456344 XIRP2 chr2:167744997-168116263
KLK7 chrl9:51461888-51472929 XRCC3 chrl4:104163955-104181823
KNDC1 chrl0:134973951-135039916 YWHAZ chr8:101930804-101965616
KRAS chrl2:25358182-25403854 ZBTB41 chrl:197127572-197169672
KRT17 chrl7:39775689-39781094 ZNF385D chr3:21459915-22414812
KSR1 chrl7:25799036-25953461 ZNF568 chrl9:37407231-37488501
A patient cohort was sequenced as described herein and a subset of the candidate list was created (Table 2). Gene locations as identified in the tables herein, for example in Table 2, are as found in the Archive Ensembl Human database, release 59-Aug 2010 (http://aug2010.archive.ensembl.org/Homo_sapiens/Info/Index) which provides human genomic data as assembled from the Genome Reference Consortium (GRC). The GRC consists of the Wellcome Trust Sanger Institute, the Genome Center at Washington University, the European Bioinformatics Institute and the National Center for Biotechnology Information. The subset comprises biomarkers identified during sequencing that were correlated with ovarian cancer, such that one or more of the patient samples were identified to have one or more variant mutations in a gene, thereby classifying that gene as a biomarker for detecting, diagnosing, and prognosing ovarian cancer, a particular type of ovarian cancer and/or a stage of ovarian cancer. Table 2-List of an exemplary subset from candidate list
Figure imgf000018_0001
calcium channel, voltage-dependent, L type, alpha
CACNA1S IS subunit chrl:201008642-201081694 calcium channel, voltage-dependent, alpha 2/delta
CACNA2D1 subunit 1 chr7:81575760-82073114
CCDC141 coiled-coil domain containing 141 chr2:179694996-179914813
CCNYL2 cyclin Y-like 2 chrl0:42903616-42970797
CDH11 Cadherin 11, type 2, osteoblast chrl6:64977656-65156101
CHRM3 cholinergic receptor, muscarinic 3 chrl:239549865-240078750
CNBD1 cyclic nucleotide binding domain containing 1 chr8:88218216-88394955
C0L3A1 Collagen, type III, alpha 1 chr2:189839046-189877472
DNAH11 dynein, axonemal, heavy chain 11 chr7:21582833-21941457
DNAH7 Dynein, axonemal, heavy chain 7 chr2:196602427-196935730
FAT3 FAT tumor suppressor homolog 3 chrll:92085262-92629636
FGFR2 Fibroblast growth factor receptor 2 chrl0:123237848-123357972
FRG1 FSHD region gene 1 chr4:190861943-190884359
FRG2C FSHD region gene 2 family, member C chr3:75713481-75716371
GPR112 G protein-coupled receptor 112 chrX:135383122-135499047
GPR115 G protein-coupled receptor 115 chr6:47653600-47689757
GPR179 G protein-coupled receptor 179 chrl7:36481493-36499730
GRM3 glutamate receptor, metabotropic 3 chr7:86273224-86494200
HNRNPUL1 heterogeneous nuclear ribonucleoprotein U-like 1 chrl9:41768391-41813811
HSDL2 hydroxysteroid dehydrogenase like 2 chr9:115142217-115234690
IL13RA1 interleukin 13 receptor, alpha 1 chrX:117861535-117928502
KIAA1109 KIAA1109 chr4:123073488-123283907
KIF24 kinesin family member 24 chr9:34252379-34329198
KNDC1 kinase non-catalytic C-lobe domain containing 1 chrl0:134973951-135039916
LRRC53 leucine rich repeat containing 53 chrl:74935562-74978298
LRRIQ1 Leucine-rich repeats and IQ motif containing 1 chrl2:85430099-85638881
MFN1 mitofusin 1 chr3:179065480-179112719
MLL3 B melanoma antigen family, member 3 chr7:151832007-152133628
MUC16 Mucin 16 chrl9:9038078-9091814
MUC2 Mucin 2 chrll:1074875-1104416
MUC4 Mucin 4 chr3:195473636-195539148
MYH6 myosin, heavy chain 6, cardiac muscle, alpha chrl4:23851049-23877482
MY016 myosin XVI chrl3:109248500-109860355
NEB nebulin chr2:152341850-152591001
Nance-Horan syndrome (congenital cataracts and
NHS dental anomalies) chrX:17393543-17754114
NI PBL nipped-B homolog chr5:36876861-37066515 olfactory receptor, family 4, subfamily A, member
OR4A47 47 chrll:48510269-48511332
Phosphoinositide-3-kinase, catalytic, alpha
PIK3CA polypeptide chr3:178865902-178957881
PIKFYVE Phosphionositide kinase, FYVE finger containing chr2:209130991-209223475
PKHD1 Polycystic kidney and hepatic disease 1 chr6:51480098-51952423
RP11-42302.6 undefined chrl:142803532-142827235
RPGR retinitis pigmentosa GTPase regulator chrX:38128424-38186817
SACS Spastic ataxia of Charlevoix-Saguenay chrl3:23902965-24007841 succinate dehydrogenase complex, subunit A,
SDHA flavoprotein (Fp) chr5:218356-256815
solute carrier family 16, member 7 (monocarboxylic
SLC16A7 acid transporter 2) chrl2:60083126-60175407 SLC5A7 Solute carrier family 5, member 7 chr2:108602979-108630450
THOC2 THO complex 2 chrX:122734412-122866906
TIAM1 T-cell lymphoma invasion and metastatis 1 chr21:32361860-32932290
TP53 Tumor protein p53 chrl7:7565257-7590856
TPO thyroid peroxidase chr2:1377995-1547483
TTN titin chr2:179390716-179695529
USP17L1P ubiquitin specific peptidase 17-like 1 (pseudogene) chr8:7189909-7191501
WDFY4 WDFY family member 4 chrl0:49892921-50191001
XIRP2 xin actin-binding repeat containing 2 chr2:167744997-168116263
ZBTB41 zinc finger and BTB domain containing 41 chrl:197127572-197169672
ZNF385D zinc finger protein 385D chr3:21459915-22414812
ZNF568 zinc finger protein 568 chrl9:37407231-37488501
Experiments were performed on a first 25 sample patient cohort to identify what genes, if any, had variant mutations. Of variant mutations present in a gene, it was further determined how many of the patient samples in the cohort had a variant mutation in that particular gene. For example, whether one or more, two or more, three or more, four or more, five or more, six or more or seven or more of the patient samples had a variant mutation in a particular gene. Table 2 comprises a list of genes wherein two or more patient samples had variant mutations in that gene. Those genes are recognized herein as biomarkers for detecting, diagnosing or prognosing ovarian cancer, a type of ovarian cancer and/or a stage of ovarian cancer in a patient.
The Table 2 list of biomarkers was collated based on data from a retrospective study comprising genomic samples from high-grade serous, low-grade serous, endometrioid, mucinous and clear cell epithelial type ovarian cancer tumors from the first sample cohort of 25 patients (Table 3, and Example 1). With regards to Table 3, tissue types are S=serous, M=mucinous, E=endometrioid and C=clear cell; ND=not done and NA=not applicable. The tier grade was determined as described in Vang 2009.
Table 3 -Tissue characterization for first patient cohort samples
Figure imgf000020_0001
S106 E 0 1A 4 NA
S108 E 0 1A 3 NA
S110 S ND 1C 3 HGSI
Sill C 0 1C 3 NA
S112 S 3 2C 3 HGSII
S113 S 3 2B 4 HGSII
S114 S 3 2C 3 HGSII
S115 S 0 2B 3 HGSII
S116 S 3 2B 4 HGSII
S117 M 0 1C 3 NA
S118 M 0 1A 3 NA
S119 S 3 2C 4 HGSII
S120 S 3 2C 4 HGSII
S121 S 3 2C 3 HGSII
S122 S 3 1A 4 HGSI
S123 S 3 1A 3 HGSI
S124 S 3 1C 3 HGSI
S125 S 3 3B 3 HGSIII
S126 S 2 2C 1 LGSII
S127 S 3 2C 2 LGSII
To maximize the potential of uncovering a full spectrum of somatic mutations in heterogeneous tumor samples, samples were selected with high tumor content prior to sequencing by selecting the tumor samples with at least 70% tumor nuclei during pathological review of fresh frozen tumor samples. Average coverage of a haploid genome was 40 fold with approximately 93% of exomic sequence and approximately 86% of the whole genomic sequence covered by at least 10 reads (Table 4).
Table 4-First patient cohort tumor sample genome coverage
Figure imgf000021_0001
S113 2607556487 86 60890035 93
S114 2569200848 85 58075374 88
S115 2606916877 86 59597579 91
S116 2593767667 86 59663735 91
S117 2623580591 87 61719219 94
S118 2625519108 87 61189644 93
S119 2627344990 87 63357272 96
S120 2615628702 86 63191504 96
S121 2548632489 84 61915340 94
S122 2635263807 87 62701105 95
S123 2624096726 87 62529748 95
S124 2624372709 87 62513832 95
S125 6230554108 87 62681703 95
S126 2632931952 87 62447325 95
S127 6215150865 86 62137949 94
To supply a patient matched normal DNA sample for each patient genomic DNA tumor tissue sample from the first cohort, blood was obtained from each patient, buffy coat DNA was isolated and the genomic blood DNA was matched with the respective patient sample (Table 5).
Table 5-First patient cohort patient matched normal sample genome coverage
Figure imgf000022_0001
S120 2625822432 87 62663814 95
S121 2603816991 86 62880739 96
S122 2640396083 87 62778534 95
S123 2629871745 87 62206245 95
S124 2631598384 87 62903055 96
S125 2659291712 88 63034991 96
S126 2633657450 87 63159404 96
S127 2628740830 87 61226777 93
Differences in sequence between the patient matched normal sample and tumor sample from a patient was determined under stringent criteria as set forth in Pleasance et al, 2010, Nature 463: 191-196, wherein a baseline set of somatic sequence alterations obtained from each blood sample genome was subtracted from its corresponding tumor genome sequence, thereby allowing identification of the sequence variants in each tissue sample above background mutations. Preliminary analyses identified approximately 424,435 somatic mutations in approximately 5,916 genes among the patient tumor samples. The range of somatic mutations per tumor was approximately 124 to about 385,369 genetic alterations. The tumor with 385,369 mutations was from a patient with Stage IC high-grade serous ovarian cancer and was determined to represent a hypermutator phenotype as the patient had neither prior chemotherapy nor radiation therapy treatment and no pre-existing cancer prior to ovarian cancer diagnosis. Excluding the hypermutated sample, approximately 2130 mutations within gene coding regions and 36,936 mutations outside gene coding regions (i.e., non-coding regions) were identified in approximately 256 genes in one or more tumor samples in the study.
The tumor protein 53 (TP53 or p53) was identified as the most frequently mutated gene in high-grade ovarian cancer, with 15 out of 16 tumors showing mutations in TP53. Mutations of TP53 were identified throughout the coding regions, consistent with the mutation profiles of tumor suppressor genes (Jones et al, 2010, Science 330:228-31; Vogelstein and Kinzler, 2004, Nat Med 10:789-99). Consistent with the presence of TP53 mutations, these tumors displayed high levels of TP53 staining by immunohistochemistry (IHC) consistent with TP53 signatures found in putative precursor lesions (Crum et al, 2007, Curr Op in Obs & Gyn 19:3-9; Lee et al, 2007, J Path 211 :26-35). Analysis of whole-exome sequencing data sets from 26 tumors available from The Cancer Genome Atlas (TCGA) identified 10 out of 26 early stage, high-grade serous ovarian cancer have mutations in TP53. In addition, 277 out of 318 late-stage high-grade serous ovarian cancer had mutations in TP53. These data are consistent with previous studies (Ahmed et al, 2010, J Pathol 221 :49-56), and support the possibility that mutations in TP53 may represent the earliest documented genetic lesions in high-grade serous ovarian cancer (Kindelberger et al, 2007, Am J Surg Pathol 31 : 161-69). The Cancer Genome Atlas data identified the genes FAT3, DNAH7, EMR1, MUC4, and SLC5A7 as frequently mutated, although no statistical correlation with ovarian cancer was reported. Mutations in those genes were identified in several of the subject test samples and the TCGA dataset from early- stage, high-grade serous ovarian cancer. Although the cohort 24 tissue samples were obtained from subjects without germline BRCAl and BRCA2 mutations, three of the tissue samples exhibited somatic mutations in BRCAl and BRCA2, mutation frequencies which were within expected limits.
In the non-coding regions, experimentation identified 338 gene associated locations having statistically significant mutations in non-coding regions when compared to random background mutations. Since previous studies have demonstrated that distal enhancer elements could be as far away as 500 kb in a genome, it is contemplated that mutations in non-coding regions several hundred kilobases away from a promoter may possess functional significance in regulating gene expression and therefore also serve as viable biomarkers indicative of cancer, such as ovarian cancer.
Validation studies were performed on a different second patient cohort (i.e., different from the fist patient cohort) of 40 tumor tissue samples and matched normal tissue samples (i.e., validation sample cohort). A validation cohort of 40 tumor normal and normal matched samples was obtained. The validation tumor samples were formalin fixed, paraffin embedded (FFPE) samples and normal matched FFPE samples were obtained from fallopian tubes, ovary, or normal lymph nodes.
Validation tumor samples were characterized as found in Table 6. However, validation sample AVP-76 was excluded from the final analysis as the sample was determined to be hypermutated (having roughly 15X more mutations than the second highest sample). With regards to Table 6, tissue types are S=serous, HG=high grade serous, LG=low grade serous, M=mucinous, E=endometrioid, C=clear cell;
A=anaplastic, PD=poorly differentiated, B=borderline, focally invasive, P=papillary, LN=lymph node, ND=not done and NA=not applicable.
Table 6-Validation tumor sample characterization
Figure imgf000025_0001
V4-79 ND NA NA
V4-82 ND NA NA
V4-83 ND NA NA
V4-84 ND NA NA
V4-85 ND NA NA
V4-88 ND NA NA
V4-89 ND NA NA
The validation data identified nine biomarkers in particular that can be considered hot spots for mutations in ovarian cancer. The nine biomarkers as listed in Table 7 comprise ARID 1 A, CACNA2D1, CSMD3, CTN B1, DNAH7, PIKC3A, PKHDl, TP53 and TRRAP and were identified upon data analysis from the validation sample cohort, wherein mutations in these genes were found to be correlated (p<0.05) with the presence of ovarian cancer. Gene ID and gene location are reported as found in the Ensembl database as previously described.
Table 7- Biomarkers identified as correlated with ovarian cancer in validation cohort
Figure imgf000026_0001
Validation data analysis results are reported in Table 8 for the genes TP53, ARID 1 A, CT B 1, PIKC3A, TRRAP, DNAH7, CACN12D, CSMD3 and PKHDl . Table 8-Validation data for ovarian cancer associated biomarkers
Figure imgf000027_0001
As shown in Table 8, mutations in TP53 were found in 24 of the validation cohort samples, which were expected and serve as a positive control for the testing processes. However, mutations in other genes were also identified which correlate to the presence of ovarian cancer. For example, mutations in DNAH7 were identified in five of the patient samples, mutations of which correlated with ovarian cancer presence (p=0.009) and which were determined to be more likely than not to be deleterious to the DNAH7 encoded protein (Condel=0.915).
Certain illustrative embodiments of biomarkers and their methods of use are described below. The biomarkers and their methods of use are not limited to these embodiments.
As disclosed herein, the identified biomarkers and methods of their use have numerous diagnostic and prognostic applications. Biomarkers as described herein find utility, either alone or in combination, in methods for diagnosing ovarian cancer, a type of ovarian cancer and/or a stage of ovarian cancer. Biomarkers as described herein find utility, either alone or in combination, in methods for prognosing patient outcome diagnosed with ovarian cancer, a type of ovarian cancer and/or a stage of ovarian cancer. Biomarkers as described herein find utility, either alone or in combination, in methods for screening patients for the presence or absence of ovarian cancer, a type of ovarian cancer and/or a stage of ovarian cancer, for example for patients that might be part of a high risk population predisposed to developing ovarian cancer (e.g., family history, genetic predisposition, etc.). The biomarkers as described herein, either alone or in combination, find utility as diagnostic, prognostic, or screening tools in conjunction with additional tests and methods for identifying ovarian cancer. Additional tests and methods for identifying ovarian cancer include, but are not limited to, transvaginal sonography, protein staining methods such as IHC or histopathological staining such as H&E, genetic probe assays such as in situ hybridization (ISH), TNM and/or FIGO staging, clinical staging, pathological staging, etc., for example as recognized by the American Joint Committee on Cancer (AJCC) and/or the World Health Organization (AJCC Cancer Staging Manual). Additional tests and method for identifying ovarian cancer may include experimental or discovery related tests and methods that are not yet recognized as mainstream, however find utility in providing support for a diagnosis of ovarian cancer nonetheless.
In one embodiment, the present disclosure provides genes or gene associated locations useful as biomarkers for ovarian cancer. In some embodiments, a biomarker is a gene or genetic location that was identified to comprise one or more variant mutations in patient test samples as compared to the gene or gene location in a patient matched normal sample. In some embodiments, a biomarker that was identified to have variants mutations in at least two patient samples is contemplated to represent a "hot spot", or gene that comprises variants mutations as compared to other genes in an ovarian cancer test sample (e.g., tissue, cell, circulating tumor cells, etc.). For example, Table 2 and Table 7 are exemplary of genes that were identified in two or more patient samples to have variant mutations compared to the same gene is a patient matched normal sample, thereby identifying them as potential biomarkers for the presence or absence of ovarian cancer, type of ovarian cancer, and/or stage of ovarian cancer. In embodiments of the present disclosure, variant mutations in biomarkers as described herein are located in a coding region of a gene. In some embodiments, biomarkers as described herein are located in non-coding regions of a gene. In other embodiments, biomarkers as described herein are located in intergenic regions. In some embodiments, biomarkers as described herein comprise single nucleotide variants (SNVs). In other embodiments, biomarkers as described herein comprise insertions and/or deletions (indels) of one or more genomic sequences. In further embodiments, a biomarker may comprise both SNVs and indels. In some embodiments, the biomarkers as disclosed herein are useful in detecting the presence or absence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject. In some embodiments, the biomarkers as disclosed herein are useful in diagnosing the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject. In some embodiments, the biomarkers as disclosed herein are useful in prognosing disease progression, treatment outcome and/or treatment regimen progress of a subject diagnosed with ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer. In some
embodiments, the biomarkers as disclosed herein are useful in screening a subject for the possibility of developing ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject. In some embodiments, the biomarkers as described herein are useful in screening potential therapeutic options for treating a patient having ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject. In one embodiment, the present disclosure provides biomarkers comprising non-coding region variant mutations for detecting the presence or absence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer. In some
embodiments, one or more of ROB02, CNTNAP2, PTPRN2, GPC5, LRP 1B, CSMD3, DLG2, CNTNAP5, DPP6, PCDH15, and CDH12 are biomarkers indicative of ovarian cancer as demonstrated in Table 9.
Table 9-Genes with non-coding region variant mutations
Figure imgf000029_0001
SGCZ sarcoglycan zeta NM_139167 21 10.01 1.00E-05
CDH12 N-cadherin 2 NM_004061 19 8.92 1.00E-05
Kv channel interacting
KCNIP4 protein 4 NM_147182 19 9.32 3.00E-05
DPP6 dipeptidyl-peptidase 6 NM_130797 20 11.51 5.00E-05
PCDH15 protocadherin 15 NM_033056 22 14.49 8.00E-05 low density lipoprotein-
LRP1B related protein IB NM_018557 21 13.32 1.80E-04 contactin associated
CNTNAP2 protein-like 2 NM_014141 23 16.41 3.40E-04 discs, large homolog 2,
DLG2 chapsyn-110 NM_001142702 21 14.76 1.70E-03
ROB02 codes for a member of the roundabout family of proteins involved in axonal guidance and neuronal migration. CNTNAP2 and CNTNAP5, contactin associated protein genes are members of the neurexin family which function as cell adhesion molecules. PTPR 2 is a member of the receptor protein tyrosine
phosphatase gene family, the proteins of which are involved in regulation of a variety of cellular processes including, but not limited to, cell growth and oncogenic
transformation. GPC5, a glypican gene, codes for a cell surface proteoglycan
involved in cell adhesion, growth and proliferation. Reduced expression of GPC5 has been implicated in lung cancer pathogenesis. LRP1B, a low density lipoprotein receptor protein gene is frequently hypermethylated and inactivated in gastric, oral, and lung cancers. CSMD3 gene codes for a putative membrane protein that is found mutated in colorectal cancer. PCDH15 and CDH12 are cadherin related genes, the proteins of which are involved in cell adhesion. In some embodiments, biomarkers associated with the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer comprise one or more of ROB02, PTPRN2, GPC5, CSMD3, HNT and CNTNAP5. In some embodiments, biomarkers associated with the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer further comprise one or more of DPP6, KCNIP4 and PCDH15. In some embodiments, biomarkers associated with the presence of ovarian cancer in a sample further comprise CNTNAP2, LRP1B, and DLG2.
The following genes are disclosed herein as biomarkers associated with the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer: TP53, MUC2, MY016, OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, PIK3CA, ARID 1 A, CTNNB1, CSMD3, TRRAP, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl, TTN, SLC16A7, COL3A1 and HNRNPULl. In some embodiments, biomarkers comprising variant mutations in ovarian cancer patient samples comprise one or more of, two or more of, TP53, MUC2, MY016, OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, PIK3CA, ARID 1 A, CTNNB1, CSMD3, TRRAP, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl, TTN, SLC16A7, COL3A1 and
HNRNPULl. In preferred embodiments, biomarkers comprising variant mutations in ovarian cancer patient samples and associated with the presence of high grade serous ovarian cancer in a sample comprise one or more of TP53, MUC2 and MY016 and one or more of OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl and TTN (Figure 1). In some embodiments, biomarkers comprising variant mutations in three or more patient samples and associated with the presence of high grade serous ovarian cancer in a sample comprise TP53, MUC2 and MY016. In some embodiments, biomarkers comprising variant mutations in two or more patient samples and associated with the presence of high grade serous ovarian cancer in a sample comprise one or more OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl and TTN.
In some embodiments, a method for diagnosing the presence of high grade serous ovarian cancer in a subject comprises identifying variant mutations in one or more of TP53, MUC2, MY016,OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl and TTN in a sample from a subject. In some embodiments, a method for diagnosing the presence of high grade serous ovarian cancer in a subject comprises identifying variant mutations in one or more of TP53, MUC2 and MY016 and one or more of OR4A47, CNBDl, GRMl, BRCA2, SLC5A7, MFNl, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF1, TPO, DNAH7, LRRIQ1, SDHA, TIAMl and TTN in a sample from a subject. In some embodiments, a method for diagnosing the presence of high grade serous ovarian cancer in a subject comprises identifying variant mutations in one or more of TP53, MUC2 and MY016 and one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1 S, CACNA2D 1 , PIKFYVE, NIPBL, PKHD 1 , MLL3 , ATBF 1 , TPO and one or more of DNAH7, LRRIQ1, SDHA, TIAMl and TTN in a sample from a subject.
In one embodiment, a sample from a subject used in methods for diagnosing ovarian cancer as described herein is a tissue sample, for example a biopsy tissue sample, for example an ovarian tissue biopsy sample. In some embodiments, a biopsy tissue sample used in diagnostic methods described herein is a fresh sample, or a sample that has been frozen or modified. A modified sample is, for example, a sample that has been preserved or modified for storage and/or for use in
histopathology, cytopathology, immunocytochemistry, immunohistochemistry, in situ hybridization, for example formalin fixed, paraffin embedded tissues or other methods used for tissue preservation. In one embodiment, a sample from a subject is a liquid sample, such as a blood sample containing white blood cells. In one embodiment, a liquid sample contains circulating tumor cells. In one embodiment, the liquid sample is a lavage, for example a vaginal lavage, a cervical lavage, wherein cells are harvested from the lavage sample. In some embodiments, a sample is a cell sample, such as a cervical cell sample, for example as derived from a Papanicolaou test (PAP smear) slide.
In embodiments of the present disclosure, nucleic acids are extracted and isolated from a sample, or portion thereof for subsequent use in methods as described herein for determining the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample. In embodiments of the present disclosure, white blood cells from a blood sample were utilized as a patient matched normal sample for comparison to a tissue sample. For example, blood samples were obtained from a patient and matched to that patient's tissue sample for evaluation of biomarker mutations as described herein (Example 1). The genomic DNA was isolated from the white blood cells and served as a patient baseline (normal) for comparing mutations present in the test sample. However, it is contemplated that a tissue sample that is free from the cancerous phenotype can also be utilized as a source of comparative normal genomic DNA for a patient where appropriate and available. In some embodiments, nucleic acids isolated from a sample, or a portion thereof, are used in diagnostic and/or prognostic methods.
It is contemplated that the one or more biomarkers as described herein, either alone or in combination, can be used in methods for determining ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject. Further, the biomarkers find utility in combination with other biomarkers and/or other diagnostic tests in providing a diagnostician additional tools to determine ovarian cancer status of a subject. The methods as described herein find particular utility as diagnostic and prognostic tools. In embodiments of the present disclosure, methods described herein can be used to diagnose ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a subject. In particular embodiments, methods comprising biomarkers as described herein are useful in differentiating early stage, high grade serous ovarian cancer from late stage, high grade serous ovarian cancer. In some embodiments, methods comprising biomarkers as described herein are prognostic for patient survival due to said differentiating early stage, high grade serous ovarian cancer from late stage, high grade serous ovarian cancer.
In the methods provided herein, the biomarkers comprise variant genetic sequences in a genomic DNA sample compared to a genomic DNA normal sample. Variant genetic sequences comprise single nucleotide variants, sequence insertions, sequence deletions, within genes that differ from a normal sample and indicate mutations that may indicate phenotypic changes, disease formation, cancer, etc. In some embodiments, the methods detect one or more altered genetic sequences as compared to a normal sample. In some embodiments, a comparison between gene sequences in a test sample (i.e., collected from a patient, subject, individual, etc.) to a normal or control sample (i.e., from the sample patient, subject, individual from which the test sample is collected) identifies the number of mutations associated with a particular gene, wherein the presence of a variant gene over a normal may associate that gene with ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer as described herein. In some embodiments, methods disclosed herein detect the insertion of one or more genetic sequences into a gene, deletion of one or more genetic sequences from a gene, or both as compared to a normal, or control, sample. In some embodiments, the methods detect one or more of single nucleotide variant(s) and/or insertion(s) and/or deletion(s) altered genetic sequences in a sample compared to a normal, or control sample.
Biomarkers and methods of their use as described herein can be used to differentiate between high grade serous and non-high grade serous (i.e., low grade serous, mucinous, endometrioid, clear cell type ovarian cancer) ovarian cancer
(Figure 1). Biomarkers useful for differentiating between high grade serous and non- high grade serous ovarian cancer comprise TP53, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQI, SDHA, TIAMl TTN and SLC16A7. In some embodiments, a sample of known ovarian cancer is further identified as high grade serous ovarian cancer by detecting variant mutations in one or more of TP53, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQI, SDHA, TIAMl and TTN as compared to a normal, or control sample. Conversely, in some embodiments a sample of known ovarian cancer is further identified as non-high grade serous ovarian cancer by detection variant mutations in SLC16A7.
In some embodiments, biomarkers and methods of their use as described herein differentiate between serous and non-serous (i.e., mucinous, endometrioid, clear cell type ovarian cancer) ovarian cancer in a sample (Figure 2). Biomarkers useful for differentiating between serous and non-serous ovarian cancer comprise one or more of TP53, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, COL3A1, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, HNRNPULl, TPO, DNAH7, LRRIQI, SDHA, TIAMl and TTN. In some embodiments, a method for diagnosing the presence of serous ovarian cancer in a subject comprises identifying variant mutations in one or more of TP53, MUC2 and MY016 and one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, COL3A1, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, HNRNPULl, TPO, DNAH7, LRRIQI, SDHA, TIAMl and TTN in a sample from a subject. In some embodiments, a method for diagnosing the presence of serous ovarian cancer in a subject comprises identifying variant mutations in one or more of TP53, MUC2 and MY016, one or more of OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, COL3A1, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHD1, MLL3, ATBF 1,
HNRNPUL1, TPO and one or more of DNAH7, LRRIQ1, SDHA, TIAMl and TTN in a sample from a subject. In some embodiments of methods for determining ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer, a test sample (i.e., a sample to be assayed for presence of ovarian cancer) is collected from an individual. In some embodiments, a second sample, a normal or control sample (e.g., blood sample, tissue sample known not to have a cancerous phenotype) is collected from the same individual.
Genomic DNA is isolated from the sample(s) by techniques known in the art (for example, as found in Molecular Cloning, a Laboratory Manual, Eds. Sambrook, et al, Cold Spring Harbor Press.). The isolated DNA from a sample is used in methods as described herein for detecting biomarkers indicative of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample. In some embodiments, the isolated DNA from the test and control samples are subjected to sequencing, for example next generation sequencing methodologies. Sequence data from the test and the control DNA samples are compared, for example by aligning the two sequences, variant sequences are identified in the test sequence over the control sequence and the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample is identified based on said comparison. As such, isolated genomic DNA from a sample is used to identify variant mutations in a genetic sequence, wherein genes comprising variant mutations relative to a normal sample are associated with the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample.
In some embodiments, a subset of biomarkers as found in Table 1 is used to determine the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample in a sample. The subset can represent one or more, two or more, three or more, four or more, five or more, or six or more biomarkers with variant mutations from the subset of which is indicative of the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer in a sample in a sample. For example, the subset of biomarkers comprising AC008021.1, ADAMTS20, ANK3, ANKRD30A, ARID 1 A, ASPN, ATBFl, BRCA2, CACNA1S, CACNA2D1, CCDC141, CDH11, CHRM3, CNBD1, CAL3A1, CSMD3, CTNNB1, DNAH11, DNAH7, FAT3, GPR112, GRP115, GPR179, GRM3, HNRNPUL1, HSDL2, IL13RA1, KIAA1109, KIF24, KNDC1, LRRC53, LRRIQl, MFN1, MLL3, MUC16, MUC2, MUC4, MYH6, MY016, NEB, NHS, NIPBL, OR4A47, PIK3CA, PIKFYVE, PKHDl, RPGR, SACS, SDHA, SLC16A7, SLC5A7, THOC2, TIAM1, TP53, TPO, TRRAP, TTN, USP17L1P, WDFY4, XIRP2, ZBTB41, ZNF385D and ZNF568 may be useful for indicating the presence of ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer. An additional subset comprising biomarkers TP53, PKHDl, PIK3CA, CACNA2D1, DNAH7, ARID 1 A, CTNNB1, CSMD3 and TRRAP may be useful for indicating the presence of ovarian cancer. An additional subset encompassing biomarkers comprising TP53, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, CACNA1S,
CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, TPO, DNAH7, LRRIQl, SDHA, TIAMl and TTN may be useful for indicating the presence of high grade serous type ovarian cancer. An additional subset encompassing biomarkers comprising TP53, MUC2, MY016, OR4A47, CNBD1, GRM1, BRCA2, SLC5A7, MFN1, CHRM3, NEB, COL3A1, CACNA1S, CACNA2D1, PIKFYVE, NIPBL, PKHDl, MLL3, ATBFl, HNRNPUL1, TPO, DNAH7, LRRIQl, SDHA, TIAMl, TTN is useful for indicating the presence of serous type ovarian cancer. An additional subset encompassing biomarkers comprising SLC16A7 is useful for indicating the presence of non-high grade serous type ovarian cancer.
Therefore, as described herein with reference to ovarian cancer, biomarkers with variant mutations as identified in two or more patient samples from the sample cohort or a subset thereof as described herein may be correlated with the presence of ovarian cancer and/or a type of ovarian cancer (e.g., high grade serous, serous, non- high grade serous, non-serous, endometrioid) and/or stage of ovarian cancer (e.g., early stage serous versus late stage serous, etc.).
The gene Mucin 2 (MUC2) is characterized herein as a biomarker that was mutated in at least four patient samples and is herein associated with ovarian cancer, a type of ovarian cancer, and/or stage of ovarian cancer. This gene has not been previously associated with the presence of ovarian cancer. The gene Cyclic Nucleotide Binding Domain Containing Protein 1 (CNBD1) has been characterized herein as a biomarker that was mutated in at least three patient samples and is herein associated with ovarian cancer, a type of ovarian cancer, and/or stage of ovarian. This gene has not been previously associated with the presence of ovarian cancer. The gene dynein, axonemal, heavy chain 7 (DNAH7) is characterized herein as a biomarker that was mutated in at least five validation tumor samples and had not been previously statistically associated with ovarian cancer. DNAH7 belongs to the dynein heavy chain family and is a component of the inner dynein arm of ciliary axonemes. These axomenes are expressed in the cytoplasm of bronchial epithelial cells and serve as the force generating protein of respiratory cilia by the motion of their microtubules. Dynein has ATPase activity and the force producing power strike is contemplated to occur on release of ADP. The fallopian tube consists
predominately of two types of cells of which ciliated cells predominate throughout the tube, particularly in the ampulla and infundibulum which is surrounded by fimbriae and this includes the ovarian fimbriae that is attached to the ovary. The majority of ovarian cancers can be classified as "epithelial" in type, however it has been suggested that the fallopian tube could also be the source of some types of ovarian cancer (2008, Piek et al, Adv Exp Med Biol 622:79-87; incorporated herein by reference in its entirety). Indeed, the tubal fimbria is viewed as being a preferred site for early adenocarcinoma in women with familial ovarian cancer (2006, Medeiros et al, Am J Surg Pathol 30:230-236); incorporated herein by reference in its entirety). It is herein contemplated that, as DNAH7 is expressed in the ciliated cells of the fallopian tube and fimbriae, somatic variations within the gene may give rise to ovarian cancer and/or increase the aggressiveness of ovarian cancer. Diagnostic methods utilizing biomarkers as described herein are contemplated to be useful for identifying the presence or absence of ovarian cancer in a patient, the type of ovarian cancer present and/or the stage of ovarian cancer present in a patient. For example, epithelial ovarian carcinomas (e.g., serous, mucinous, endometrioid, clear cell, transitional cell, squamous cell, undifferentiated and mixed epithelial tumors) account for approximately 80% of all patients with ovarian cancer. The difficulty in diagnosing early stage epithelial ovarian carcinomas is high and the disease is typically left undiagnosed thereby leading to poor overall prognosis once the late stage carcinoma is diagnosed. Diagnostic methods utilizing biomarkers as described herein are contemplated to provide for early stage disease diagnosis thereby providing a patient with a more favorable prognostic outcome.
Prognostic methods utilizing biomarkers as described herein are contemplated to be useful for determining a proper course of treatment for a patient having ovarian cancer. A course of treatment refers to the therapeutic measures taken for a patient after diagnosis or after treatment for ovarian cancer. For example, a determination of the likelihood for cancer recurrence, spread, or patient survival, can assist in determining whether a more conservative or more radical approach to therapy should be taken, or whether treatment modalities should be combined. For example, when ovarian cancer recurrence is likely, it can be advantageous to precede or follow surgical treatment with chemotherapy, radiation, immunotherapy, biological modifier therapy, gene therapy, vaccines, and the like, or adjust the span of time during which the patient is treated. A diagnosis or prognosis of an ovarian cancer state is contemplated to be correlated with one or more, for example a particular combination, of biomarkers described herein comprising two or more gene mutations.
Methods utilizing biomarkers as described herein are contemplated to be useful for monitoring recurrence of ovarian cancer following or during a course of treatment in a patient diagnosed with ovarian cancer. Treatment of a patient diagnosed with ovarian cancer includes, but is not limited to, surgery, chemotherapy, radiation, immunotherapy, biological modifier therapy and gene therapy. For example, following or during a prescribed treatment regimen for ovarian cancer the methods utilizing biomarkers as described here are contemplated to be useful in determining the success, failure and/or progress of the treatment regimen. Such information can be used by a clinician to determine a change in treatment course and/or future action or treatment for a patient.
For determination of ovarian cancer, a sample is obtained, nucleic acids for example DNA are isolated from the sample by established means known in the art, and the isolated nucleic acids are assayed by methods described herein. A normal or control sample is typically obtained for comparison with the test sample. Methods described herein are contemplated for use in, for example, characterizing the variant mutational status of one or more biomarkers as found in Table 1, wherein the variant mutational status of one or more biomarkers is useful in determining ovarian cancer status. In some embodiments, methods for characterization comprise sequencing technologies, for example next generation sequencing technologies. In some embodiments, microarray based technologies are utilized to characterize the mutational status of a biomarker as described herein for determining the status of ovarian cancer in a sample. In some embodiments, a sample is assayed for methylation status, the data of which is used to characterize a sample for ovarian cancer status.
Isolated genomic DNA from samples is typically modified prior to characterization. For example, genomic DNA libraries are created which are applied to downstream detection applications. A library is produced, for example, by performing the methods as described in the Nextera™ DNA Sample Prep Kit (Epicentre® Biotechnologies, Madison WI), GL FLX Titanium Library Preparation Kit (454 Life Sciences, Branford CT), SOLiD™ Library Preparation Kits (Applied Biosystems™ Life Technologies, Carlsbad CA), and the like. The sample as described herein may be further amplified for sequencing by, for example, multiple stand displacement amplification (MDA) techniques. For sequencing after MDA, an amplified sample library is, for example, prepared by creating a DNA library as described in Mate Pair Library Prep kit, Genomic DNA Sample Prep kits or TruSeq™ Sample Preparation and Exome Enrichment kits (Illumina®, Inc., San Diego CA). Useful cluster amplification methods are described, for example, in U.S. Patent No. 5,641,658; U.S. Patent Publ. No. 2002/0055100; U.S. Patent No. 7, 115,400; U.S. Patent Publ. No. 2004/0096853; U.S. Patent Publ. No. 2004/0002090; U.S. Patent Publ. No. 2007/0128624; and U.S. Patent Publ. No. 2008/0009420, each of which is incorporated herein by reference in its entirety. Another useful method for amplifying nucleic acids on a surface is rolling circle amplification (RCA), for example, as described in Lizardi et al, Nat. Genet. 19:225-232 (1998) and US 2007/0099208 Al, each of which is incorporated herein by reference in its entirety. Emulsion PCR methods are also useful, exemplary methods for which are described in Dressman et al, Proc. Natl. Acad. Sci. USA 100:8817-8822 (2003), WO 05/010145, or U.S. Patent Publ. Nos. 2005/0130173 or 2005/0064460, each of which is incorporated herein by reference in its entirety. Sample preparation methods of the present disclosure are not necessarily limited by any particular library preparation or amplification method as a sample as described herein is contemplated to be amenable to any of a variety of methods known in the art and/or commercially available for such purposes.
Genomic DNA libraries derived from a sample as described herein can be characterized for ovarian cancer status by sequencing for the presence of gene mutations. In one embodiment, sequencing can be performed following
manufacturer's protocols on a system such as those provided by Illumina®, Inc. (HiSeq 1000, HiSeq 2000, Genome Analyzers, MiSeq, HiScan, iScan, BeadExpress systems), 454 Life Sciences (FLX Genome Sequencer, GS Junior), Applied
Biosystems™ Life Technologies (ABI PRISM® Sequence detection systems, SOLiD™ System), Ion Torrent® Life Technologies (Personal Genome Machine sequencer) further as those described in, for example, in United States patents and patent applications 5,888,737, 6, 175,002, 5,695,934, 6,140,489, 5,863,722,
2007/007991, 2009/0247414, 2010/01 11768 and PCT application WO2007/123744, each of which is incorporated herein by reference in its entirety. Methods disclosed herein are not necessarily limited by any particular sequencing system as the particular sample preparation required for a particular instrument is contemplated to be amenable for use with any sample as described herein. Sequencing methodologies for characterizing ovarian cancer are contemplated to be useful either alone or in combination with other assays for ovarian cancer determination. Output from a sequencing instrument can be of any sort. For example, current technology typically utilizes a light generating readable output, such as fluorescence or luminescence, however the present methods for detecting mutations in a biomarker for determining ovarian cancer status in a sample is not necessarily limited to the type of readable output as long as differences in output signal for a particular sequence of interest can be determined. Examples of analysis software that may be used to characterize output derived from practicing methods as described herein include, but are not limited to, Pipeline, CASAVA and GenomeStudio data analysis software (Illumina®, Inc.), SignalMap and NimbleScan data analysis software (Roche NimbleGen), GS Analyzer analysis software (454 Life Sciences), SOLiD™, DNASTAR® SeqMan® NGen® and Partek® Genomics Suite™ data analysis software (Life Technologies), Feature Extraction and Agilent Genomics Workbench data analysis software (Agilent Technologies), Genotyping Console™, Chromosome Analysis Suite data analysis software (Affymetrix®). A skilled artisan will know of additional numerous commercially and academically available software alternatives for data analysis for sequencing generated output. Embodiments described herein are not limited to any data analysis method. In one embodiment, once the sample is appropriately processed the number of mutations in a biomarker can be detected using microarray methodologies. For example, a plurality of different probe molecules can be attached to a substrate or otherwise spatially distinguished in an array. Exemplary arrays that can be used to detect the number of mutations in a biomarker include, but are not limited to, slide arrays, silicon wafer arrays, liquid arrays, bead-based arrays and others known in the art or set forth in further detail below. In some embodiments, the methods can be practiced with array technology that combines a miniaturized array platform, a high level of assay multiplexing, and scalable automation for sample handling and data processing. Exemplary methods and systems for microarray analysis includes, but is not limited to, those methods and systems commercialized by Roche NimbleGen,
Inc., Illumina®, Inc., Affymetrix® and Agilent Technologies. An array of beads can also be in a fluid format such as a fluid stream of a flow cytometer or similar device. Commercially available fluid formats for distinguishing beads include, for example, those used in XMAP™ technologies from Luminex or MPSS™ methods from Lynx Therapeutics.
Exemplary microarray methods and systems can be found in, for example, US patents 5,856,101, 5,981,733; 6,001,309; 6,023,540, 6, 110,426, 6,200,737, 6,221,653; 6,232,072, 6,266,459, 6,327,410, 6,355,431, 6,379,895, 6,429,027, 6,458,583, 6,667,394 6,770,441, 6,489,606 and 6,859,570, 7, 106,513, 7,126,755, and 7, 164,533, US patent applications 2005/0227252, 2006/0023310, 2006/006327, 2006/0071075, 2006/0119913 and PCT publications WO98/40726, W099/18434, WO98/50782, WO00/63437, WO04/024328 and WO05/033681 (each of which is incorporated herein by reference in their entireties). Microarray based technologies for characterizing ovarian cancer are contemplated to be useful either alone or in combination with other diagnostic and/or prognostic assays.
In one embodiment, once the sample is appropriately processed the number of mutations in a biomarker can be detected using polymerase chain reaction (PCR) methodologies (including, but not limited to, US Patents 4,683, 195, 4,683,202, 4,965, 188, 5,075,216, 5,210,015, 5,333,675, 5,475,610, 5,487,972, 5,538,848, 5,602,756, 5,656,493, 5,804,375, 5,994,056, 6,030,787, 6,703,236, 6,814,934;
incorporated herein by reference in their entireties), for example qualitative PCR, real-time PCR, quantitative PCR, and the like. For example, a plurality of probes may be utilized to amplify genomic regions contemplated to comprise one or more somatic variants wherein amplification data are used to determine the presence or absence of variants thereby associating a sample with ovarian cancer, or not, as the case may be. Exemplary methods and systems for PCR analysis include, but are not limited to those methods and systems commercialized by Illumina®, Inc., Roche Applied Science and Applied Biosystems™. Polymerase chain reaction based technologies for characterizing ovarian cancer are contemplated to be useful either alone or in combination with other diagnostic and/or prognostic assays.
The following examples are provided in order to demonstrate and further illustrate certain preferred embodiments and aspects of the present disclosure and are not to be construed as limiting the scope thereof.
EXAMPLES
EXAMPLE 1-Study population, sample collection and processing
The genes described herein were identified from a group of 25 early-stage ovarian cancers, of predominantly serous histology. Twenty-five patients with known ovarian cancer were enrolled in a study to identify ovarian cancer related biomarkers under approved protocol from the Internal Review Board at the Mayo Foundation for Medical Education and Research, Rochester, MN. One of the 25 patient tumor samples was determined to represent a hypermutated tumor phenotype. Blood samples were collected from the enrolled patients and processed under approved IRB protocols. The blood sample from a patient was matched with its corresponding tissue sample and served as the patient matched normal sample for that patient in determining gene associated variant mutations.
Genomic DNA was extracted and isolated from a portion of a fresh frozen biopsy tissue sample using the Gentra® Puregene® Tissue Kit (QIAGEN, Inc., Valencia CA) following manufacturer's protocol. A blood sample was drawn from each patient by venous puncture into a Vacutainer® tube containing EDTA. The tube was centrifuged to separate the white blood cells, or buffy coat, from the red blood cells. DNA was extracted from 100- 200μ1 of buffy coat cells using the AutoGen prep 965 (AutoGen, Holliston MA) or Gentra® Puregene® Tissue Kit following manufacturer's protocol. The genomic blood DNA was matched with its respective tissue isolated genomic DNA for each patient.
Genomic DNA libraries were generated by adding 4μg of sample DNA to the Paired End Sample prep kit PE-102-1001 (Illumina®, Inc.) following manufacturer's protocol. Briefly, DNA fragments are generated by random shearing and conjugated to a pair of oligonucleotides in a forked adaptor configuration. The ligated products are amplified using two oligonucleotide primers, resulting in double-stranded blunt- ended products having a different adaptor sequence on either end.
Clusters were formed prior to sequencing using the V3 cluster kit (Illumina®, Inc.). Briefly, products from a DNA library preparation are denatured and single strands annealed to complementary oligonucleotides on the flow-cell surface. A new strand is copied from the original strand in an extension reaction and the original strand is removed by denaturation. The adaptor sequence of the copied strand is annealed to a surface-bound complementary oligonucleotide, forming a bridge and generating a new site for synthesis of a second strand. Multiple cycles of annealing, extension and denaturation in isothermal conditions resulted in growth of clusters, each approximately 1 μιη in physical diameter. The DNA in each cluster is linearized by cleavage within one adaptor sequence and denatured, generating single- stranded template for sequencing by synthesis (SBS) to obtain a sequence read. To perform paired-read sequencing, the products of read 1 are removed by denaturation, the template is used to generate a bridge, the second strand is re-synthesized and the opposite strand is then cleaved to provide the template for the second read.
EXAMPLE 2- Sequencing and data analysis Sequencing was performed using the Illumina®, Inc. V4 SBS kit with lOObp paired end reads on the Genome Analyzer IIx. Briefly, DNA templates are sequenced by repeated cycles of polymerase-directed single base extension. To ensure base-by- base nucleotide incorporation in a stepwise manner, a set of four reversible terminators, A, C, G and T each labelled with a different removable fluorophore are used. The use of modified nucleotides allows incorporation to be driven essentially to completion without risk of over-incorporation. It also enables addition of all four nucleotides simultaneously minimizing risk of misincorporation. After each cycle of incorporation, the identity of the inserted base is determined by laser- induced excitation of the fluorophores and fluorescence imaging is recorded. The fluorescent dye and linker is removed to regenerate an available group ready for the next cycle of nucleotide addition. The Genome Analyzer IIx is designed to perform multiple cycles of sequencing chemistry and imaging to collect sequence data automatically from each cluster on the surface of each lane of an eight-lane flow cell.
For single genome read aggregation and variant calling image analysis, base calling and Phred quality scoring were carried out using the Illumina analysis pipeline (RTA vl .4-1.8). Sequence reads were ignored from those clusters whose proximity to others resulted in mixed sequence data (purity-filtering).
Sequences were aligned using Elandv2e from CASAVA version 1.8
(Illumina®, Inc.) with full repeat resolution and orphan rescue (sensitive mode), to the human GRCh37.1 reference sequence. The aligned reads were aggregated and sorted into chromosomes based on alignment positions. The sorted reads were used to call variants using Hyrax, a Bayesian S V caller and GROUPER. The callers are part of the standard CASAVA 1.8 distribution and were run with default parameters. This process was carried out for the tumor and normal genomes.
Somatic single nucleotide variant subtraction for calling somatic SNVs was performed by taking the list of positions in the tumor genome with snp quality values greater than 15 (Q(snp)tumor >15) and high confidence of the assigned genotype given the polymorphic prior (Q(max_gt)tumor >20). For each putative SNV the normal sample was investigated. If a call was present in the normal sample at the same position as a putative SNV, and if the call had a quality value greater than 0 (Q(snp) normal >0), the position was filtered out as background.
To minimize false positive somatic SNV calls the putative SNVs were recalled (using Hyrax) in the tumor sample, however for recalling additional information from candidate indel contigs constructed from the normal sample was used. This process was utilized to avoid any indels that were initially missed in the tumor due to low supporting evidence. A candidate SNV was called when there was complete agreement between the initial SNV call and the recall. To further minimize false positive somatic SNV calls, variants were also recalled in the normal sample, using Hyrax with all read filtering turned off. If the posterior probability of the tumor genotype was higher than a non-reference genotype, then that SNV was considered to have low confidence evidence in the normal sample and was discarded.
When considering somatic indel subtraction for calling somatic
insertion/deletions (indels) only those indels that were confidently called in the tumor sample and not present in the normal sample were considered. Indels in the tumor with a Q score less than 30 (Q(indel) <30) and those positions that had less than 10 reads coverage in the normal sample (for a 3 OX build) were filtered out. To be considered as evidence, a read had a single read alignment score >10 and a paired read alignment score >90. Positions were excluded if they mapped within 1000 bases of a known centromere or telomere (as obtained from the reference genome
GRCh37.1) as these locations typically contain highly repetitive regions and read alignments are problematic.
For each putative somatic indel position in a tumor sample the indel position was matched with the region and calls present in the normal sample. If a putative somatic indel position overlapped with an indel call originating from the normal sample, the indel was considered to be present in normal germline and hence the position was filtered out.
Given the repetitive nature of the human genome, the putative somatic indel region was characterized by finding the shortest sequence around the indel that extended outside any repeats and that region was matched with each intersecting read in the normal sample. If there was evidence in the normal sample having the same pattern in the intersecting normal reads the candidate somatic indel was discarded. To account for homopolymer slippage (i.e., the sequencing polymerase misincorporates, or misses, one or more bases in a sequence of identical bases) during a run, one read was allowed in the normal sample to have the same indel as found in the tumor. After sequence calling, each class of variant was annotated against the Ensembl database release e59. Each somatic variant was queried for overlapping annotated features. For all gene features, it was considered whether a consequence of the somatic variant was synonymous, non-synonymous, or nonsense or if the variant could disrupt a canonical splice site at an intron/exon boundary. For variants that fell in a coding exon, the consequence of the change was analyzed and reported.
Regulatory regions (e.g., 3' and 5' untranslated regions) of the gene feature were also reported.
Greater than 88% coding regions and non-coding regions were sequenced at 10X coverage (Tables 4 and 5), with the majority depth of read of 40X depth. Table 10 is exemplary of the variant mutations identified in gene coding regions from the different patient samples.
Table 10-Exemplary variant mutation gene data identified in patient samples
Figure imgf000046_0001
indel:95221949
ASPN snv:95227251
CNBD1 snv:88249287
GRM3 snv:86493611
BRCA2 snv:32912067
SLC5A7 snv:108614422
MFN1 snv:179104323 indel:179069729
USP17L1P snv:7190757 snv:7190674
NHS snv:17744822 snv:17745899
CHRM3 snv:240071188
indel:72021506
AC008021.1 indel:72020572 snv:72021103
ZNF385D snv:21465532
SLC16A7 snv:60098793
indel:179730646
CCDC141 snv:179718336
MUC16 snv:9064320 snv:9072670 indel:135441605
GPR112 snv:135498627
ADAMTS20 snv:43822554
NEB snv:152424658
C0L3A1 snv:189868814 snv:189868986
indel:36483604
GPR179 snv:36485977
PIK3CA snv:178951949
CACNA1S snv:201021749 snv:201010683
indel:34297103
KIF24 snv:34311050
indel:38147055
RPGR snv:38158323
ANK3 snv:61865750
CACNA2D1 snv:81634750
PIKFYVE snv:209169030
NIPBL snv:36961602 snv:36985977 indel:37064682
PKHD1 snv:51824785
MLL3 snv:151949069
IL13RA1 snv:117880939
ATBF1 snv:72992304 snv:72832418
KIAA1109 snv:123176076
ANKRD30A snv:37508677
indel:115181232
HSDL2 snv:115167940
TH0C2 snv:122754804
WDFY4 snv:50151419
HNRNPUL1 snv:41787174
TPO snv:1499939
FAT3 snv:92600076 XIRP2 snv:168107188
S106 S108 S110 Sill S112
TP53 indel:7578275 snv:7578263 indel:7574003
MUC4 snv:195507778 snv:195513574 snv:195507660
MUC2
DNAH7 snv:196682555
LRRIQ1
SDHA snv:251111 snv:224623
MY016
TIAM1 snv:32639209
SACS snv:23908255 snv:23910318
GPR115 snv:47678579 snv:47661435
TTN snv:179476806
KNDC1
OR4A47
MYH6
DNAH11
LRRC53
ZNF568
CDH11
ZBTB41
ASPN
CNBD1 snv:88218358
GRM3
BRCA2 snv:32936688
SLC5A7
MFN1
USP17L1P snv:7190703
NHS
CHRM3
AC008021.1
ZNF385D snv:21478473 snv:21467136
SLC16A7 snv:60169238 snv:60168911
CCDC141
MUC16
GPR112
ADAMTS20
NEB
C0L3A1
GPR179
PIK3CA snv:178938934 snv:178917478
CACNA1S
KIF24
RPGR
ANK3 snv:61844398 CACNA2D1 snv:81765955
PIKFYVE snv:209204208
NIPBL
PKHD1
MLL3
IL13RA1 snv:117875085
ATBF1 snv:72829404
KIAA1109 snv:123156036
ANKRD30A snv:37455581
HSDL2 snv:115171187
TH0C2 snv:122778487
WDFY4
HNRNPUL1
TPO
FAT3 snv:92600139
XIRP2 snv:168099250
S113 S114 S115 S116 S117
TP53 snv:7577539 indel:7573995
MUC4 snv:195505835 snv:195510156
MUC2
DNAH7 snv:196825139
LRRIQ1
SDHA snv:251111
MY016
TIAM1
SACS
GPR115
TTN snv:179496907
KNDC1 snv:135009218
OR4A47
MYH6
DNAH11
LRRC53 snv:74941062
ZNF568
CDH11 snv:64981618
ZBTB41
ASPN
CNBD1
GRM3 snv:86394724
BRCA2
SLC5A7
MFN1
USP17L1P
NHS
CHRM3 AC008021.1
ZNF385D
SLC16A7
CCDC141
MUC16 snv:9089414
GPR112
ADAMTS20
NEB
COL3A1
GPR179
PIK3CA
CACNA1S snv:201044622
KIF24
RPGR
ANK3
CACNA2D1
PIKFYVE
NIPBL
PKHD1
MLL3
IL13RA1
ATBF1
KIAA1109
ANKRD30A
HSDL2
TH0C2
WDFY4 snv:49951404
HNRNPUL1
TPO snv:1499927 snv:1500496
FAT3 snv:92577883
XIRP2 indel:168101130
S118 S119 S120 S121 S122
TP53 snv:7578555 snv:7577545 indel:7577062 indel:7574030
MUC4 snv:195507827
MUC2 snv:1093904
DNAH7
LRRIQ1 snv:85546097
SDHA snv:226065
MY016
TIAM1 snv:32595841 indel:32525031
SACS
GPR115
TTN
KNDC1
OR4A47 snv:48510546 snv:48510718 MYH6 snv:23876348
DNAH11
LRRC53
ZNF568
CDH11
ZBTB41 snv:197128578
ASPN
CNBD1 snv:88364008
GRM3 indel:86469058
BRCA2 snv:32914542
SLC5A7 snv:108627063
MFN1
USP17L1P
NHS snv:17745485
CHRM3 snv:240071787 snv:240071925
AC008021.1
ZNF385D
SLC16A7
CCDC141 indel:179721095
MUC16
GPR112 snv:135432005
ADAMTS20 snv:43847747
NEB snv:152381056 snv:152484212
C0L3A1
GPR179
PIK3CA
CACNA1S
KIF24
RPGR indel:38128962
ANK3 indel:61831806
CACNA2D1
PIKFYVE indel:209165748
NIPBL
PKHD1 indel:51930808 snv:51774274
MLL3 snv:151904481
IL13RA1
ATBF1
KIAA1109 snv:123187946
ANKRD30A
HSDL2
TH0C2
WDFY4 snv:49983792
HNRNPUL1 indel:41778038
TPO
FAT3 XIRP2
S123 S124 S125 S126 S127
TP53 snv:7578208 snv:7577568 snv:7578555 snv:7577121
MUC4
MUC2 snv:1098778
DNAH7
LRRIQ1
SDHA
MY016 snv:109518585
TIAM1
SACS
GPR115
TTN
KNDC1 indel:134999861
OR4A47
MYH6 snv:23859484
DNAH11
LRRC53
ZNF568 snv:37487980
CDH11
ZBTB41
ASPN snv:95237027
CNBD1
GRM3
BRCA2
SLC5A7 snv:108608573
MFN1 snv:179085231
USP17L1P
NHS
CHRM3
AC008021.1
ZNF385D
SLC16A7
CCDC141
MUC16
GPR112
ADAMTS20 snv:43822246
NEB
C0L3A1 snv:189870112
GPR179 snv:36499099
PIK3CA
CACNA1S
KIF24 snv:34256021
RPGR
ANK3 CACNA2D1 snv:81626509
PIKFYVE
NIPBL
PKHD1
MLL3 snv:152008946
IL13RA1 snv:117895147
ATBF1
KIAA1109
ANKRD30A snv:37418912
HSDL2
TH0C2 snv:122757498
WDFY4
HNRNPUL1 snv:41798379
TPO
FAT3
XIRP2
For biomarkers identified in Table 9, wherein non-coding regions variants identify one or more biomarkers, to visualize the entire spectrum of genomic somatic mutations (e.g., coding regions, conserved regions, non-coding regions, repeat
regions, etc.) in the samples, genetic analysis was performed on each base in a 5kb window and somatic mutations detected within the 10 kb windows. It is contemplated that utilizing this method for mutation identification allowed for the objective
identification of regions of interest with significantly higher mutation rates without any prejudice to gene-centric regions. Poisson test statistics were calculated by
permuting the number of mutations within a given 10 kb window and the results were utilized to identify regions of interest with statistically significant mutations, those regions exhibiting higher mutation rates than expected by chance alone.
EXAMPLE 3-Immunohistochemical staining of tumor protein 53 Immunohistochemical staining of tissues for evaluation of p53 expression was performed using monoclonal anti-p53 (DO-1, Dako, CA) followed by visualization using the Universal LSAB+ Kit (Dako), following manufacturer's protocols.
EXAMPLE 4- Validation cohort sample sequencing and data analysis A validation cohort of 40 tumor and matched normal sample pairs (FFPE) were obtained and processed as described in Example 1 to provide genomic DNA for sequencing. The study population comprised female adults from 18-80 years of age. Some cytosine bases in DNA extracted from FFPE samples can be subject to random deamination resulting in a cytosine base being converted to uracil and in turn misread as a thymidine base. This conversion typically occurs randomly in approximately 0.5- 1.5% of all cytosine bases found in FFPE derived DNA. The random deamination can lead to false positive predictions of somatic OT or G>A mutations during sequencing. Any random deamination which might occur in the FFPE samples of the present investigation was mitigated by filtering somatic OT or G>A variants to remove those predicted variants where the somatic allele was detected at a frequency of 15% or less; the estimated false positive rate at an average read depth of 66X determined to be roughly one in 250 samples (based on a read depth of 66X, the average read depth was >66X, 50% GC content, 1.6 MBases targeted, with a 1.5% deamination rate; the binomial theory predicts p=0.0037 of encountering 10
(10/66=15.2%) or more deamination events).
The validation design was a 1000X deep sequencing of a targeted pull down (i.e., enrichment). A targeted pull down is an experimental design which targets a subset of the genome by designing complimentary DNA probes which hybridize to, or pull down, randomly fractionated genomic DNA (i.e., a DNA library as described in Example 1). The hybridized DNA can be enriched for fragments that overlap the probe region and the enriched DNA eluted and sequenced to high depth. The targeted pull down was achieved by using an Illumina, Inc. TruSeq Custom Enrichment Array Kit for in solution capture for targeting and enriching user selected human genomic regions.
Over the duration of the validation experiments three generations of targeted enrichment arrays were designed, AVP3, AVP4 and AVP5, each containing the content of previous generations plus additional probes to additional targeted genes. The probes were designed to target previously defined DNA gene regions (i.e., such as those genes identified in Tables 1 and 2), or regions in genes subsequently identified as potential ovarian cancer gene biomarkers. Table 1 1 describes the enrichment design for targeted genes in the validation cohort. The enrichment design AVP3 utilized 4071 probes which targeted 1200213 bases over 215 genes (Table 12). The enrichment design AVP4 utilized 5097 probes which targeted 1473936 bases over 245 genes (Table 12) and enrichment design AVP5 utilized 5490 probes which targeted a total of 1578057 bases over 271 genes (Table 12). The probe designs for the TruSeq™ Custom Enrichment Array Kit consisted of a subset of probes taken from the TruSeq™ Exome Enrichment Kit, the probe design data of which is available within the kit. Further, a skilled artisan following known and established protocols will understand how to define probes for targeted enrichment assays. Table 12 describes the genes targeted, the enrichment assay and whether they were enriched in assay design AVP3, 4 and/or 5. Table 1 1- Enrichment assay design for validation cohort samples
Figure imgf000055_0001
Table 12-Enrichment assay correlating targeted genes and assay design
Figure imgf000055_0002
ATP7B Y Y Y
AT Y Y Y
AURKA Y Y Y
BARD1 Y Y Y
BRAF Y Y Y
BRCA1 Y Y Y
BRCA2 Y Y Y
BRIP1 Y Y Y
Cllorf41 Y Y Y
CACNA1S Y Y
CACNA2D1 Y Y
CCDC141 Y Y
CCNE1 Y Y Y
CD47 Y Y Y
CDH 1 Y Y Y
CDH 11 Y Y Y
CDH 12 Y
CDK4 Y Y Y
CDKN2A Y Y Y
CHEK1 Y Y Y
CHEK2 Y Y Y
CHRM3 Y
CHST2 Y Y Y
CNBD1 Y Y
CNGA1 Y Y Y
CNTNAP2 Y
CNTNAP5 Y
C0L11A1 Y Y Y
COL1A1 Y Y Y
COL3A1 Y Y Y
COL5A1 Y Y Y
COMT Y Y Y
CRABP1 Y Y Y
CRABP2 Y Y Y
CRKRS Y Y Y
CSMD3 Y Y Y
CTNNB1 Y Y Y
CUTC Y Y Y
CXCL10 Y Y Y
DI02 Y Y Y
DLG2 Y
DNAH 11 Y Y DNAH7 Y Y Y
DNMT1 Y Y Y
DPP6 Y
DSP Y Y Y
EGFR Y Y Y
EIF5A2 Y Y Y
EMR1 Y Y Y
ERBB2 Y Y Y
ERBB3 Y Y Y
ERCC1 Y Y Y
ESR1 Y Y Y
FAP Y Y Y
FAT3 Y Y Y
FBXW7 Y Y Y
FGF1 Y Y Y
FGFR1 Y Y Y
FGFR2 Y Y Y
FGFR3 Y Y Y
FGFR4 Y Y Y
FLNA Y Y Y
FN 1 Y Y Y
F0LR1 Y Y Y
FOX03 Y Y Y
FRG 1 Y
FRG2B Y Y Y
FRG2C Y
GABRA6 Y Y Y
GJB2 Y Y Y
GLI2 Y Y Y
GPC5 Y
GPR112 Y Y
GPR115 Y Y Y
GPR179 Y Y
GRM3 Y Y
GSX2 Y Y Y
HERC1 Y Y Y
HIF1A Y Y Y
HJURP Y Y Y
HLA Y Y Y
HMGA2 Y Y Y
HNRNPULl Y Y
HOXB3 Y Y Y HP Y Y Y
HP Y Y Y
HSDL2 Y Y
IGF1R Y Y Y
IGFBP1 Y Y Y
IGFBP2 Y Y Y
IGFBP3 Y Y Y
IGFBP4 Y Y Y
IGFBP5 Y Y Y
IGFBP7 Y Y Y
IL13RA1 Y Y
ILKAP Y Y Y
INPP4B Y Y Y
ITGA5 Y Y Y
JAK2 Y Y Y
KCNIP4 Y
KCNJ12 Y Y Y
KIAA1109 Y Y Y
KIF24 Y Y
KIT Y Y Y
KLK5 Y Y Y
KLK7 Y Y Y
KNDC1 Y Y
KRAS Y Y Y
KRT17 Y Y Y
KSR1 Y Y Y
LAPTM4B Y Y Y
LATS1 Y Y Y
LCN2 Y Y Y
LOX Y Y Y
LRP1B Y
LRRC53 Y Y
LRRIQ1 Y Y
LUM Y Y Y
MAP2K4 Y Y Y
MAP3K1 Y Y Y
MDM2 Y Y Y
MECOM Y Y Y
MET Y Y Y
MFN 1 Y Y Y
MGMT Y Y Y
MKI67 Y Y Y MLH 1 Y Y Y
MLL3 Y Y
MLL4 Y Y Y
MMP11 Y Y Y
MMP7 Y Y Y
MREllA Y Y Y
MSH2 Y Y Y
MSH6 Y Y Y
MST1 Y Y Y
MTOR Y Y Y
MUC1 Y Y Y
MUC16 Y Y Y
MUC2 Y Y
MUC4 Y Y Y
MUTYH Y Y Y
MYC Y Y Y
MYH6 Y Y
MY016 Y Y
MY03A Y Y Y
NALCN Y Y
NBN Y Y Y
NEB Y Y
NF1 Y Y Y
NHS Y Y Y
NIPBL Y Y Y
NOTCH 1 Y Y Y
NOTCH 3 Y Y Y
NRAS Y Y Y
NTS Y Y Y
OLFML2B Y Y Y
OR4A47 Y Y
OR5V1 Y Y
PALB2 Y Y Y
PARP4 Y Y Y
PCDH 15 Y
PCP4 Y Y Y
PDGFRA Y Y Y
PDGFRB Y Y Y
PDK1 Y Y Y
PGR Y Y Y
PI3 Y Y Y
PIK3C2B Y Y Y PIK3CA Y Y Y
PIK3CB Y Y Y
PIK3CD Y Y Y
PIK3CG Y Y Y
PIK3R1 Y Y Y
PIK3R2 Y Y Y
PIK3R3 Y Y Y
PIKFYVE Y Y Y
PIM2 Y Y Y
PKHD1 Y Y Y
PKHDILI Y Y Y
PLAU Y Y Y
PMS1 Y Y Y
PMS2 Y Y Y
POSTN Y Y Y
PPP2R1A Y Y Y
PRKCI Y Y Y
PRRX1 Y Y Y
PTEN Y Y Y
PTK2B Y Y Y
PTMA Y Y Y
PTPRN2 Y
RAB25 Y Y Y
RAD50 Y Y Y
RAD51C Y Y Y
RASAL1 Y Y Y
RBI Y Y Y
RB1CC1 Y Y Y
ROB02 Y
RPGR Y Y
RPS6KA2 Y Y Y
RREB1 Y
RRM 1 Y Y Y
S100A1 Y Y Y
SACS Y Y Y
SCAP Y Y Y
SDHA Y Y
SGCZ Y
SLC16A7 Y Y
SLC5A7 Y Y Y
SLPI Y Y Y
SMAD4 Y Y Y SPARC Y Y Y
SRC Y Y Y
STK11 Y Y Y
STK36 Y Y Y
SULF1 Y Y
SULF2 Y Y
TACC3 Y Y Y
TET1 Y Y Y
THBS2 Y Y Y
TH0C2 Y Y
TIAM 1 Y Y Y
TMEM 158 Y Y Y
TMPRSS4 Y Y Y
TNK2 Y Y Y
TOPI Y Y Y
TOP2A Y Y Y
TOP2B Y Y Y
TP53 Y Y Y
TPD52 Y Y Y
TPMT Y Y Y
TPO Y Y
TRRAP Y Y Y
TSC1 Y Y Y
TSC2 Y Y Y
TTN Y Y
TYMS Y Y Y
UBAP2 Y Y Y
UBE2C Y Y Y
USP17L1P Y Y
USP17L2 Y Y Y
VCAN Y Y Y
VEGFA Y Y Y
WDFY4 Y Y
WFDC2 Y Y Y
WT1 Y Y Y
XIRP2 Y Y Y
XRCC3 Y Y Y
YWHAZ Y Y Y
ZBTB41 Y Y
ZFHX3 Y Y Y
ZNF385D Y Y
ZNF568 Y Y The validation enriched samples were sequenced using a paired end protocol on the HiSeq 2000 sequencer. Data analysis was performed on the sequenced sample data. Briefly, samples were sequenced on the HiSeq 2000 (to a target depth of 1000X sequence depth per sample) and the tumor sequence data was compared to the normal sequence data thereby identifying those SNVs present in the tumor samples only.
The sequence data was aligned to human reference genome HG19/GRCh37 using CASAVA-1.8 data analysis software and all optical and PCR duplicates were tagged as duplicates by the software using default parameters. Tumour specific mutations were detected in each sample by selecting those reads with properly matched Rl and R2 read pairs as determined by CASAVA (i.e., those reads wherein the orientation of Rl and R2 read pairs and the size of the DNA fragment were within normal limits) and which were not tagged as duplicate reads, and by subtracting matching normal tissue variants from the tumor variants using the Strelka method for somatic SNV and small indel detection from sequencing data of matched
tumor/normal samples (for example, the Strelka method as found in Saunders et al, Bioinformatics, in press; the Strelka workflow source code is publically available at ftp://strelka@ftp.illumina.com) Both single nucleotide variants (SNV) and somatic insertions or deletions (sINDEL) were predicted. The somatic variants were annotated using Variant Effect Predictor (VEP) from Ensembl API (2012, Flicek et al, Nucl Acids Res 40:D84-D90). Mutations annotated by VEP as
splice_acceptor_variant, splice_donor_variant, stop_gained, stop_lost,
complex change in transcript, initiator codon change, inframe codon gain, inframe_codon_loss, non_synonymous_codon, frameshif t variant or
splice region variant were considered as having a functional impact on the downstream protein of any gene overlapping the mutations (see VEP documentation http://www.ensembl.org/info/docs/variation/predicted_data.html. All other exonic variants were classified as non- functional P-values from Fisher's exact test using number of functional versus all (functional and non- functional) mutations annotated to the gene (excluding intronic mutations). Only SNVs predicted by Strelka with QSSscores >= 15 (i.e., quality score for any somatic SNV, for example for the alternate allele to be present at a significantly different frequency in the tumor and normal) or indels with QSI scores >= 25 (i.e, quality score for any somatic indel, for example for the alternate haplotype to be present at a significantly different frequency in the tumor and normal) were considered. A filter was placed on predicted C->T, G->A, T->C and A->G SNVs such that only SNVs with an allele frequency (AF) > 15% were considered, to remove some observed noise in the indel prediction. Additionally, a further filter was placed on insertions to restrict them to less than lObp in length. Approximately 7792 putative variants passed the QC criteria and were further tested for enrichment wherein each gene was tested for enrichment of functional markers (described above) thereby reporting the significance of whether mutations in a biomarker were enriched for deleterious functional mutations. A Condel score was also utilized to integrate the output of computational tools aimed at assessing the impact of non-synonymous SNVs on protein function by computing a weighted average of the scores (WAS) of computational tools, such as SIFT, Polyphen2, MAPP, LogR Pfam e-value (2004, Clifford et al, Bioinformatics 20: 1006-1014; incorporated herein by reference in its entirety) and MutationAssessor. Briefly, the scores of different methods are weighted using the complementary cumulative distributions produced by the five methods on a dataset of approximately 20000 missense mutations, both deleterious and neutral. The probability that a predicted deleterious mutation is not a false positive of the method and the probability that a predicted neutral mutation is not a false negative are employed as weights. A high Condel score, for example above 0.5, represents mutations more likely than not to be deleterious whereas a low Condel score represents the opposite (2011,
Gonzalez-Perez and Lopez-Bigas, Am J Hum Gen 88:440-449; incorporated herein by reference in its entirety).
Data analysis of the validation cohort sequencing data identified nine biomarkers which are considered mutational hotspots indicative of ovarian cancer wherein the biomarkers correlate with the presence of ovarian cancer with high probability (p<0.05) (Table 8). The location of the mutations identified in the biomarkers as found in the validation data cohort is described in Table 13. Further, the proposed consequence of the identified mutations (i.e., whether the mutations listed may be deleterious to the gene protein) may be numerically represented by the Condel score as found in Table 8. For example, the impact of mutations identified in TP53, DNAH7, CTNNB1 and ARID 1 A on the encoded proteins were computed to be more likely than not deleterious to the protein (Condel scores (>0.800)). For example, as shown in Table 13 mutations in these genes resulted in a variety of potentially deleterious consequences to the protein including premature stops, frame shifts, splice variants and non-synonymous codons.
Table 13 -Location of mutations in validation samples
Figure imgf000064_0001
PKHD1 chr6:51503649_G>A V418 non_synonymous_codon,splice_region_variant
PKHD1 chr6:51618148_C>T V423 non_synonymous_codon
PKHD1 chr6:51656150_G>A V401 non_synonymous_codon
PKHD1 chr6:51712645_G>T V416 non_synonymous_codon
PKHD1 chr6:51747931_G>A AVP3-37 non_synonymous_codon
PKHD1 chr6:51750766_C>A V419 non_synonymous_codon
PKHD1 chr6:51908440_G>A V407 non_synonymous_codon
PKHD1 chr6:51908453_G>A V407 non_synonymous_codon
PKHD1 chr6:51930782_G>T V405 non_synonymous_codon
PKHD1 chr6:51949821_T>A V405 splice_region_variant,intron_variant
TP53 chrl7:7573995_C>- AVP4-82 frameshift_variant
TP53 chrl7:7574000_C>A V412 stop_gained
TP53 chrl7:7574003_G>A V406 stop_gained
TP53 chrl7:7574030_G>- V422 frameshift_variant
TP53 chrl7:7576870_C>A V413 stop_gained
TP53 chrl7:7577062_T>- AVP4-89 frameshift_variant
TP53 chrl7:7577120_C>A AVP3-64 non_synonymous_codon
TP53 chrl7:7577120_C>T V414 non_synonymous_codon
TP53 chrl7:7577121_G>A V417 non_synonymous_codon
TP53 chrl7:7577538_C>T V421 non_synonymous_codon
TP53 chrl7:7577568_C>T AVP3-39 non_synonymous_codon
TP53 chrl7:7578195_CAC>- V416 inframe_codon_loss
TP53 chrl7:7578208_T>C V408 non_synonymous_codon
TP53 chrl7:7578265_A>G V402 non_synonymous_codon
TP53 chrl7:7578274_->G AVP4-79 frameshift_variant
TP53 chrl7:7578284_C>- V419 frameshift_variant
TP53 chrl7:7578394_T>C V420 non_synonymous_codon
TP53 chrl7:7578395_G>C V415 non_synonymous_codon
TP53 chrl7:7578404_A>T AVP3-23 non_synonymous_codon
TP53 chrl7:7578406_C>T V411 non_synonymous_codon
TP53 chrl7:7578413_C>T V403 non_synonymous_codon
TP53 chrl7:7579523_GT>- AVP4-83 frameshift_variant
TP53 chrl7:7578406_C>T V418 non_synonymous_codon
TP53 chrl7:7578406_C>T V401 non_synonymous_codon
TRRAP chr7:98506386_A>- V403 frameshift_variant
TRRAP chr7:98524941_G>A AVP4-88 non_synonymous_codon
TRRAP chr7:98575861_G>C AVP3-37 non_synonymous_codon
TRRAP chr7:98580920_G>A AVP3-37 non_synonymous_codon All publications and patents mentioned in the present application are herein incorporated by reference. Various modification and variation of the described methods and compositions of the invention will be apparent to those skilled in the art without departing from the scope and spirit of the invention. Although the invention has been described in connection with specific preferred embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention that are obvious to those skilled in the relevant fields are intended to be within the scope of the following claims.

Claims

A method for determining the presence of ovarian cancer in a subject comprising:
a) providing a nucleic acid sample from a subject,
b) detecting the presence of one or more variant mutations in two or more genes selected from the group consisting essentially of ARID 1 A, CACNA2D1, CT B1, PKHD1, DNAH7, PIKC3A, TRRAP and CSMD3,
c) evaluating the probability that the two or more genes are correlated with ovarian cancer, and
d) determining the presence of ovarian cancer in the sample based on the probability of correlation.
The method of claim 1, wherein said evaluating comprises comparing the nucleic acid sample to a matched normal nucleic acid sample.
The method of claim 1, wherein said nucleic acid sample is a genomic DNA sample.
The method of claim 1, wherein a gene is correlated with ovarian cancer at probability <0.05.
The method of claim 1, wherein said two or more genes comprises DNAH7.
The method of claim 3, wherein said genomic DNA sample is isolated from a sample selected from the group consisting of a tissue sample, a biopsy sample a cell sample, a circulating tumor cell sample, a fixed tissue sample, a frozen tissue sample or a lavage sample.
The method of claim 2, wherein the matched normal nucleic acid sample is a genomic DNA sample.
8. The method of claim 1, wherein the presence of one or more variant mutations is found in two or more genes selected from the group consisting of ARID 1 A, CACNA2D1, CTN B 1, PKHD1, DNAH7, PIKC3A, TRRAP, and CSMD3.
9. The method of claim 1, wherein said detecting comprises sequencing the nucleic acid sample.
10. The method of claim 9, wherein said sequencing comprises sequence by
synthesis.
11. The method of claim 1, wherein said detecting is carried out on a microarray.
12. The method of claim 1, wherein said detecting comprises performing
polymerase chain reaction on the nucleic acid sample.
13. The method of claim 12, wherein polymerase chain reaction is quantitative polymerase chain reaction.
14. The method of claim 12, wherein polymerase chain reaction is real time
polymerase chain reaction.
15. The method of claim 8, wherein the two or more genes comprises DNAH7.
16. A method for determining the presence of ovarian cancer in a sample
comprising:
a) creating a DNA library from nucleic acids derived from a sample
suspected of having ovarian cancer,
b) sequencing said DNA library to generate sequencing data comprising one or more variant mutations as identified in two or more genes from the list consisting essentially of ARID 1 A, CACNA2D1, CTNNB1, PKHD1, DNAH7, PIKC3A, TRRAP and CSMD3,
c) computationally determining the probability that the two or more genes are correlated with ovarian cancer, and d) determining the presence of ovarian cancer in said sample based on the probability of correlation.
17. The method of claim 16, wherein said computation determination comprises comparing the sequence of the nucleic acid sample suspected of having ovarian cancer to the sequence of a matched normal nucleic acid sample.
18. The method of claim 16, wherein a gene is correlated with ovarian cancer at probability <0.05.
19. The method of claim 16, wherein said two or more genes comprises DNAH7.
20. The method of claim 16, wherein said sample suspected of having ovarian cancer is a sample selected from the group consisting of a tissue sample, a biopsy sample, a cell sample, a circulating tumor cell sample, a fixed tissue sample, a frozen tissue sample or a lavage sample.
21. The method of claim 17, wherein said matched normal nucleic acid sample is from a blood sample or a tissue sample not suspected of having ovarian cancer.
22. The method of claim 16, wherein said sequencing comprises sequence by synthesis.
23. A method for confirming the presence of ovarian cancer in a sample
comprising comparing the sequence of a nucleic acid sample derived from a tissue suspected of having ovarian cancer with the sequence of a nucleic acid from a tissue not suspected of having ovarian cancer, wherein the presence of one or more variant sequences identified in two or more of ARID 1 A, CACNA2D1, CTN B 1, PKHD1, DNAH7, PIKC3A, TRRAP and CSMD3 in the nucleic sample suspected of having ovarian cancer compared to the nucleic acid sample not suspected of having ovarian cancer indicates the presence of ovarian cancer in a sample.
24. The method of claim 23, wherein said one or more variant sequences is identified in DNAH7.
PCT/US2012/031484 2011-03-30 2012-03-30 Ovarian cancer biomarkers Ceased WO2012135635A2 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US201161469510P 2011-03-30 2011-03-30
US61/469,510 2011-03-30

Publications (2)

Publication Number Publication Date
WO2012135635A2 true WO2012135635A2 (en) 2012-10-04
WO2012135635A3 WO2012135635A3 (en) 2013-03-07

Family

ID=46932377

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2012/031484 Ceased WO2012135635A2 (en) 2011-03-30 2012-03-30 Ovarian cancer biomarkers

Country Status (1)

Country Link
WO (1) WO2012135635A2 (en)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR101432173B1 (en) * 2013-01-22 2014-08-22 한국원자력의학원 Composition for diagnosis of radio-resistance or radio-sensitive HNRNPUL1 marker and use thereof
WO2015156740A1 (en) * 2014-04-08 2015-10-15 Agency For Science, Technology And Research Markers for ovarian cancer and the uses thereof
CN111635942A (en) * 2020-06-22 2020-09-08 中国人民解放军空军军医大学 Mutant gene groups, libraries and kits for assessing the risk of malignancy in women
CN114966027A (en) * 2022-04-30 2022-08-30 重庆大学附属肿瘤医院 Application of detection reagent for expression levels of three genes in preparation of ovarian cancer sample dryness identification reagent

Non-Patent Citations (5)

* Cited by examiner, † Cited by third party
Title
DANSONKA-MIESZKOWSKA, A. ET AL.: 'A novel germline PALB2 deletion in Polish breast and ovarian cancer patients' BMC MED. GENET. vol. 11, no. 20, 02 February 2010, *
JAKUBOWSKA, A. ET AL.: 'The Leu33Pro polymorphism in the ITGB3 gene does not modify BRCA1/2-associated breast or ovarian cancer risks: results from a multicenter study among 15,542 BRCA1 and BRCA2 mutation carriers' BREAST CANCER RES. TREAT. vol. 121, no. 3, June 2010, pages 639 - 649 *
MADORE, J. ET AL.: 'Characterization of the molecular differences between ovarian endometrioid carcinoma and ovarian serous carcinoma' J. PATHOL. vol. 220, no. 3, February 2010, pages 392 - 400 *
OLIVA, E. ET AL.: 'High frequency of beta-catenin mutations in borderline endometrioid tumors of the ovary' J. PATHOL. vol. 208, no. 5, April 2006, pages 708 - 713 *
WIEGAND, K. C. ET AL.: 'ARID1A mutations in endometriosis-associated ovarian carcinomas' N. ENGL. J. MED. vol. 363, no. 16, 14 October 2010, pages 1532 - 1543 *

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR101432173B1 (en) * 2013-01-22 2014-08-22 한국원자력의학원 Composition for diagnosis of radio-resistance or radio-sensitive HNRNPUL1 marker and use thereof
WO2015156740A1 (en) * 2014-04-08 2015-10-15 Agency For Science, Technology And Research Markers for ovarian cancer and the uses thereof
CN106460064A (en) * 2014-04-08 2017-02-22 新加坡科技研究局 Ovarian cancer markers and their uses
CN111635942A (en) * 2020-06-22 2020-09-08 中国人民解放军空军军医大学 Mutant gene groups, libraries and kits for assessing the risk of malignancy in women
CN114966027A (en) * 2022-04-30 2022-08-30 重庆大学附属肿瘤医院 Application of detection reagent for expression levels of three genes in preparation of ovarian cancer sample dryness identification reagent

Also Published As

Publication number Publication date
WO2012135635A3 (en) 2013-03-07

Similar Documents

Publication Publication Date Title
JP6317354B2 (en) Non-invasive determination of fetal or tumor methylomes by plasma
US20120115735A1 (en) Pathways Underlying Pancreatic Tumorigenesis and an Hereditary Pancreatic Cancer Gene
CN106715723A (en) Method of determining PIK3CA mutational status in a sample
EP2524055A2 (en) Method to use gene expression to determine likelihood of clinical outcome of renal cancer
TWI730429B (en) HOXA7 methylation detection reagent
EP4630585A2 (en) Systems and methods for cell-free nucleic acids methylation assessment
TWI789551B (en) Application of HOXA9 methylation detection reagent in the preparation of lung cancer diagnostic reagent
WO2012135635A2 (en) Ovarian cancer biomarkers
US20180023147A1 (en) Methods and biomarkers for detection of bladder cancer
CN110283910B (en) Application of target gene DNA methylation as a molecular marker in the preparation of a kit for discriminating the progression of colorectal tissue carcinogenesis
CN111788317B (en) Compositions and methods for characterizing cancer
US11130998B2 (en) Unbiased DNA methylation markers define an extensive field defect in histologically normal prostate tissues associated with prostate cancer: new biomarkers for men with prostate cancer
JP5865241B2 (en) Prognostic molecular signature of sarcoma and its use
WO2013144202A1 (en) Biomarkers for discriminating healthy and/or non-malignant neoplastic colorectal cells from colorectal cancer cells
TW202028464A (en) Application of HOXA7 methylation detection reagent to preparation of lung cancer diagnosis reagent
US9920376B2 (en) Method for determining lymph node metastasis in cancer or risk thereof and rapid determination kit for the same
TWI789550B (en) HOXA9 Methylation Detection Reagent
US10308980B2 (en) Methods and biomarkers for analysis of colorectal cancer
WO2017106365A1 (en) Methods for measuring mutation load
CN106119398B (en) Biomarkers to predict responsiveness to pyrotinib therapy in breast cancer patients
KR20240059529A (en) Methylation markers for diagnosing lung cancer and combinations thereof
WO2022188776A1 (en) Gene methylation marker or combination thereof that can be used for gastric carcinoma her2 companion diagnostics, and use thereof
Zhang The integrated genomic analyses of human cancers
Willis Apobec3B and breast cancer Article 196

Legal Events

Date Code Title Description
NENP Non-entry into the national phase in:

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 12762938

Country of ref document: EP

Kind code of ref document: A2