WO2016178236A1 - Methods and kits for breast cancer prognosis - Google Patents

Methods and kits for breast cancer prognosis Download PDF

Info

Publication number
WO2016178236A1
WO2016178236A1 PCT/IL2016/050480 IL2016050480W WO2016178236A1 WO 2016178236 A1 WO2016178236 A1 WO 2016178236A1 IL 2016050480 W IL2016050480 W IL 2016050480W WO 2016178236 A1 WO2016178236 A1 WO 2016178236A1
Authority
WO
WIPO (PCT)
Prior art keywords
protein
biomarker
proteins
expression
sample
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/IL2016/050480
Other languages
French (fr)
Inventor
Tamar Geiger
Yair POZNIAK
Iris Barshack
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ramot at Tel Aviv University Ltd
Tel HaShomer Medical Research Infrastructure and Services Ltd
Original Assignee
Ramot at Tel Aviv University Ltd
Tel HaShomer Medical Research Infrastructure and Services Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ramot at Tel Aviv University Ltd, Tel HaShomer Medical Research Infrastructure and Services Ltd filed Critical Ramot at Tel Aviv University Ltd
Publication of WO2016178236A1 publication Critical patent/WO2016178236A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N33/00Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
    • G01N33/48Biological material, e.g. blood, urine; Haemocytometers
    • G01N33/50Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
    • G01N33/53Immunoassay; Biospecific binding assay; Materials therefor
    • G01N33/575Immunoassay; Biospecific binding assay; Materials therefor for cancer
    • G01N33/57515Immunoassay; Biospecific binding assay; Materials therefor for cancer of the breast
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6883Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
    • C12Q1/6886Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material for cancer
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/106Pharmacogenomics, i.e. genetic variability in individual responses to drugs and drug metabolism
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/158Expression markers
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N2800/00Detection or diagnosis of diseases
    • G01N2800/52Predicting or monitoring the response to treatment, e.g. for selection of therapy based on assay results in personalised medicine; Prognosis

Definitions

  • the invention relates to personalized medicine. More particularly, the invention relates to diagnostic and prognostic methods and kits for detecting metastatic breast cancer and for monitoring breast cancer progression.
  • MCT4 is a marker of oxidative stress in cancer-associated fibroblasts. Cell cycle, 10(11), 1772- 1783. Wisniewski, J.R., Zougman, A., and Mann, M. (2009a). Combination of FASP and StageTip-Based Fractionation Allows In-Depth Analysis of the Hippocampal Membrane Proteome. Journal of Proteome Research 8, 5674-5678.
  • genomic and transcriptomic efforts expanded and refined the original signatures to slightly alter the classification. These efforts culminated in studies which make up the largest breast cancer genomic-profiling done to date, combining data from multiple platforms and utilizing next-generation techniques to study up to 2,000 breast tumors. These genomic and transcriptomic data serve as an invaluable resource of breast cancer associated mutations, chromosomal aberrations and further expanded the classification to additional subtypes. However, the actual manifestation of such genomic changes in the cancer phenotype is far from obvious. Proteomics makes a natural complement to the genomic and transcriptomic studies. As proteins convey the actual functional properties of cells, they represent the final combined effect of all genetic abnormalities, including mutations, copy-number variations, epigenetic and transcription- level regulation.
  • MS-based proteomic analyses have undergone a revolution in the past decade, owing to improvements in instrumentation, sample preparation and quantification methods.
  • the advanced MS instruments combine high resolution, high mass accuracy and high speed, and are capable of comprehensively cataloguing proteomes of yeast, mouse, and most recently, human.
  • the implementation of proteomics to cancer studies is increasing; however, many of these studies are still limited in scope and quantification accuracy.
  • a major improvement in quantification technique has been the introduction of Stable Isotope Labeling with Amino acids in Cell culture (SILAC)(Ong et al., 2002).
  • SILAC-labeled cell line can then be used as a common, 'spike-in', internal standard for comparing a theoretically unlimited number of samples, including samples that cannot be metabolically labeled such as clinical tumor samples (Geiger et al., 2011). Owing to the great complexity of clinical tumor samples and the wide repertoire of expressed proteins, a single SILAC cell line was found to be insufficient for accurate quantification.
  • a mixture of cell lines termed a super-SILAC mix
  • the super-SILAC mix is further disclosed in WO 2011/042467 that is a previous application by one of the inventors.
  • Luminal tumors make up the vast majority of breast tumors and are characterized by expression of the estrogen receptor (ER). As such, they can be treated by endocrine therapy such as Tamoxifen.
  • Luminal A tumors show overall favorable prognosis, while luminal B, which also express higher levels of the proliferation marker ki67 or Her2, have somewhat poorer prognosis.
  • risk of recurrence rises substantially if the cancer had already metastasized to nearby lymph nodes by the time of diagnosis (lymph node positive, LNP).
  • lymph node negative patients benefit from endocrine therapy alone
  • LNP patients are not likely to benefit from such therapy despite high ER expression levels.
  • increasing evidence shows that the added value from adjuvant chemotherapy is also questionable (Ellis and Perou, 2013), making the lymph node status a crucial component in treatment decision-making.
  • the invention relates to a diagnostic and/or prognostic method for determining the progression of breast cancer in a subject.
  • the method of the invention comprises the steps of: (a) determining the expression level of at least one biomarker protein in at least one biological sample of said subject, to obtain an expression value for each of the at least one biomarker protein/s.
  • the biomarker proteins of the invention may be at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof; (b) calculating and determining if the expression value obtained in step (a) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample.
  • the invention relates to a diagnostic and/or prognostic composition
  • a diagnostic and/or prognostic composition comprising at least one detecting molecule or any combination or mixture of plurality of detecting molecules specific for determining the level of expression of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof.
  • a third aspect of the invention relates to a kit comprising (a) detecting molecules specific for determining the level of expression of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof in a biological sample.
  • a kit comprising (a) detecting molecules specific for determining the level of expression of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof in a biological sample.
  • Fig. 1A is a workflow depicting cohort assembly and sample preparation and analysis.
  • Fig. IB is a bar graph showing the number of proteins quantified in this study and in each of the clinical groups.
  • Fig.lC. is a graph showing the distribution of expression intensities of the quantified proteins (filtered for proteins that were quantified in >10 samples) shows a large dynamic range of abundance, still the vast majority of the proteins are expressed within 4 orders of magnitude.
  • Fig 2A shows the number of proteins quantified in each sample.
  • Fig 2B shows the protein occurrence by sample.
  • Fig 2C. is an intensity distribution of 'light' versus 'heavy' proteins shows that >90% of the protein intensities are within 5-fold from the heavy standard in all clinical groups (healthy tissue 90.9%, tumor tissue 91.7% and metastasic tissue 92.7% of the proteins with 5-fold (log 2.3) ratio).
  • Figure shows Pearson correlation for 2974 proteins that were present in >70% of the clinical samples.
  • Cluster 1 DNA replication, rRNA processing, Pyrimidine and Purin metabolism, DNA repair
  • Cluster 2 proteosome, protein export, spliceosome, regulation of cell cycle
  • Cluster 3 TCA cyclate, Glutathion metabolism, protein glycosylation, RNA Splicing
  • Cluster 4 Food acid ⁇ -oxidation, Peroxisome, Oxidative phosphorylation),
  • Cluster 5 RNA polyadenylation
  • Cluster 6 (Ribosome, Oxidative phosphorylation)
  • Cluster 7 (Ribosome, Oxidative phosphorylation, Type-I interferon signaling, Antioxidant activity, Lysosome
  • Cluster 8 Proteasome, tRNA aminoacylayion, Glycolysis/gluconeogenesis, Focal adhesion
  • Fig 4A shows a principal component analysis of healthy and cancer samples.
  • Fig 4B is a hierarchical clustering of sample Pearson correlations.
  • Fig 4D. is a pulsed-SILAC experiment in normal mammary epithelial cells (HMEC) and ER- positive breast cancer cell line (MCF7) shows higher rates of synthesis, degradation and turnover in the cancer cell line.
  • Gradient describes the change in H/L (left), M/L (middle) and H/M (right) ratios at Oh, 2h, 4h, 9, 12h and 24h after the pulse.
  • Graphs represent mean of two biological replicates.
  • Cluster 1 Ribosome, Translation
  • Cluster 5 Ribosome, Nonsense-mediated decay, Protein targeting to ER, Translation, Lysosome, Oxidative phosporylation, Spliceosome
  • Cluster 6 MCM complex, DNA replication, DNA repair, Protein folding, Locomotion, Nucleoulus
  • Fig 5B shows that the correlation among the cancer samples (primary tumors and metastases is higher than among the healthy tissue samples. (Mann- Whitney p ⁇ 0.0001).
  • Fig 5D shows the fraction of the total intensity of ECM part proteins is significantly higher in healthy tissue compared to tumor and metastatic tissue.
  • Fig 5E shows the fraction of the total intensity of ECM part proteins is significantly higher in primary LNN compared to primary LNP tumors.
  • FIG. 6A to 6F Proteomic differences between healthy breast duct epithelia and cancer cells
  • the box in the figure represents area of accurate quantification. Percentages denote the number of ratios within that range for every group.
  • Fig 6E shows activity assay of mitochondrial complex I and complex IV, evaluated by histochemistry on frozen breast cancer tumor arrays (duct carcinoma in situ). Representative tissue cores are presented.
  • Fig 6F showing staining intensity of complex I and IV, one and three normal cores, respectively, and fifteen ER-positive, Her2-negative breast cancer cores were evaluated. Bars represent average scores +SEM.
  • Fig. 7C The overexpression of three metabolic enzymes - glutamine synthethase (GLUL), acyl- CoA thioesterase (ACOT)-l/2 and oxoglutarate/malate carrier (SLC25A11) was validated using immunohistochemistry on tumor arrays.
  • GLUL glutamine synthethase
  • ACOT acyl- CoA thioesterase
  • SLC25A11 oxoglutarate/malate carrier
  • Figure shows upregulated and downregulated proteins of Recon 1 pathways. Each number represents a protein: 1 - Pentose and Glucuronate Interconversions, 2 - Heparan sulfate degradation, 3 - Tetrahydrobiopterin, 4 - Hyaluronan Metabolism, 5 - N-Glycan Degradation, 6 - Chondroitin sulfate degradation, 7 - Keratan sulfate degradation, 8 - Oxidative Phosphorylation, 9 - Pyrimidine Catabolism, 10 - Heme Degradation, 11 - ROS Detoxification, 12 - Sphingolipid Metabolism, 13 - Glutathione Metabolism, 14 - Transport, Lysosomal, 15 - Heme Biosynthesis, 16 - Galactose metabolism, 17 - Nucleotides, 18 - Fatty Acid Metabolism, 19 - Glycolysis/Gluconeogenesis, 20 - Arginine and Proline Metabolism, 21
  • Fig 9A shows staining intensity of Acyl-CoA thioesterase-1 (ACOT1), glutamine-synthetase (GLUL) and oxoglutarate carrier (SLC25A11) antibodies.
  • ACOT1 Acyl-CoA thioesterase-1
  • GLUL glutamine-synthetase
  • SLC25A11 oxoglutarate carrier
  • Fig 9B and Fig 9C show representative figures from the Human Protein Atlas database showing differential staining between healthy and tumor cores.
  • Fig 9D shows validation of 12 downregulated proteins using the Human Protein Atlas database.
  • Fig 9E shows validation of 24 upregulated proteins using the Human Protein Atlas database.
  • Figure shows a protein ranking according to their log fold change (healthy/tumor).
  • the barcode plots show the ranking distribution of proteins in the category.
  • FIG. 12A to 12B Changes in protein expression during cancer progression
  • Fig. 12A discloses that Ki67 does not discriminate between LNN and LNP tumors (Welch's t-test,
  • Fig. 13B Unsupervised clustering of 20 matched pairs of primary tumors (T) and lymph node metastases (N) shows a high incident of co-clustering.
  • Fig. 13D significantly upregulated proteins in primary tumors compared to healthy samples (upper panel) and downregulated proteins (lower panel) show no change in the pattern of expression between primary tumors and metastases.
  • White lines indicate z-scored median expression.
  • FIG. 14A-14C A 15-protein signature predicts LNP primary tumors
  • Fig. 14A A support vector machines (SVM)-based classifier was trained and tested on the 85 significantly changing proteins between LNN and LNP primary tumors, with different numbers of features (proteins) for classification; top- 15 ranked proteins were chosen as an optimal number for prediction.
  • SVM support vector machines
  • the work described in the present invention presents the first genome- scale analysis of breast cancer progression, which is able to capture novel aspects of cancer development.
  • the inventors capture the functional difference between the healthy control tissue and the tumors, between the primary tumors and the metastases, and between pre-metastatic and metastatic breast cancer. While the inventor's data captured hundreds of regulated proteins in the comparison to the healthy tissues, a more challenging comparison was between the groups of tumor tissues. Surprisingly, despite the distinct microenvironment, the inventors found greater similarity between the primary tumors and the lymph node metastases than between two groups of primary tumors, associated with the tumor stage.
  • the most prominent network of upregulated proteins in tumors consisted of structural ribosomal proteins, with a concurrent down regulation of several of the most important co-players of the translational machinery - the tRNA aminoacyl synthetases (ARSs), and also of the auxiliary protein AEVIP2.
  • ARSs tRNA aminoacyl synthetases
  • the canonical role of the ARSs is to ligate an amino acid to its cognate tRNA, later to be added to the nascent polypeptide chain. Improper activity of the ARSs may impair the accuracy of protein synthesis, and not allow for appropriate folding of proteins. Furthermore, even if the proteins are correctly translated, the marked decrease in the expression of important chaperons may also have an adverse effect on their function.
  • misfolded proteins are tumor suppressors, as have been demonstrated for p53 and VHL. Alternatively, it can induce gain-of-function activities or interfere with protein localization.
  • AEVIP2 mediates anti-proliferative and pro-apoptotic functions through regulation of ubiquitination, and as an outcome, mice lacking this protein died neonatally due to severe over-proliferation of lung epithelia, making AIMP2 a bona-fide tumor suppressor.
  • the synthetase YARS has been shown to be secreted and cleaved into two fragments, one of which acts as a cytokine to induce angiogenesis.
  • Secretion of YARS and possibly other ARSs by cancer cells may explain their reduced intracellular levels and may directly affect tumor progression through interaction with its microenvironment. Further efforts will be needed in order to elucidate the roles of tRNA aminoacyl synthetases in tumorigenesis.
  • the inventors propose that the elevated rate of protein production and degradation, together with down regulation of several quality control systems, impairs protein homeostasis in the cancer cells. Accumulation of DNA damage in the tumors, potentially due to high levels of ROS and impairment of repair mechanisms may lead to the accumulation of mutated transcripts; higher ribosome levels fail to produce functional proteins due to translation of damaged transcripts and reduced activity of ARSs; and the last line of defense against the accumulation of such proteins - the chaperones of the unfolded protein response - fail to launch a protective campaign. Damaged proteins may eventually be degraded by the proteasome or by lysosomal proteases, overall increasing protein turnover rates. Despite the tremendous energetic demand of such a mechanism, we speculate that it provides the system the necessary adaptability to changing conditions, and confers an evolutionary advantage to the cancer cells.
  • SLC25A11 oxoglutarate/malate carrier
  • GAT2 mitochondrial aspartate aminotransferase
  • the inventors propose that the decrease in glycolysis rates together with the increase in oxidative phosphorylation stem from the proximity to adipose tissue within the breast and potentially, reduced glucose levels. Increased fatty acid catabolism may suggest an adaptation of the invading cancer cells to their surroundings and selection for cells that adequately change their metabolism.
  • a 15 -protein signature predicts lymph node involvement based on the primary tumor
  • the present invention presents a 15-protein signature (shown in Table 4) that predicts the involvement of lymph node based on expression levels in the primary tumors, using supervised classification algorithms on proteins that significantly change in expression between LNN and LNP luminal tumors. To the inventor's knowledge, this is the first proteomic signature that allows such a classification.
  • the proteomic signature of the invention combines biological importance with high AUC (0.93) and low error rates (14% for LNN primary tumors and 20% for LNP tumors).
  • Several of the proteins found in the signature of the present invention have been reported to have a role in cancer progression, while others are novel indicators of aggressiveness.
  • the highest-changing protein between LNN and LNP tumors has recently been identified by a proteomic study as a marker of poor prognosis in endometrial cancer (Li et al., 2008).
  • the splicing factor SRSF1 is overexpressed in lung cancer and is a transcriptional target of MYC, and the cathepsin inhibitor cystatin B (CSTB) is overexpressed in ovarian cancer.
  • CSTB cathepsin inhibitor cystatin B
  • the present invention therefore presents a deep proteomic analysis of breast cancer progression and provide a novel, high-quality proteomic database.
  • the invention shows that the regulation on protein production is severely impaired in tumors, and that luminal tumors display an anti-Warburg effect, characterized by high levels of cellular respiration and low glycolytic activity.
  • the invention shows that the proteomic landscape of LNN and LNP primary tumors is similar, and that lymph node metastases retain the expression patterns of their original tumor.
  • the present invention provides a 15 -protein signature that accurately predicts lymph node involvement based on the primary tumor.
  • Predicting the onset of a disease is highly valuable and clinically desired particularly for diseases that are often detected at advanced stages.
  • breast cancer diagnosis and prognosis are highly important and crucial in management patient's life quality.
  • Providing tools and specifically non-invasive tools to accurately diagnose at early stage breast cancer and predict its progression may assist in reducing mortality and increase life quality of patients.
  • determining treatment regimen and monitoring patient response to treatment is highly valuable and clinically desired as it can provide information regarding suitable and successful treatment protocols enabling personalized medicine. This is appreciated in view of the fact that treatment protocols are often associated with some extent of undesired side effects.
  • the inventors used computational analysis to provide novel, unique, comprehensive and deep proteomics-scale analysis of breast cancer progression providing a high quality proteomic database.
  • the inventors have identified an arsenal of proteins that was differently expressed in healthy tissues and at different stages of breast cancer development such as metastatic and non-metastatic breast cancer and successfully characterized different protein signatures that are unique for healthy tissue, tumor tissues at different stages and metastatic tissues.
  • the inventors differentiated healthy breast tissue from breast tumors, pre-metastatic breast tumors from metastatic breast cancer and primary breast tumors from metastases.
  • the inventors found significant differences in the expression of proteins in different stages of breast tumors. Specifically, as shown in Example 6 herein, the inventors identified a specific set of signature proteins, specifically, RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB that are differently expressed in primary tumor samples obtained from patients diagnosed as having lymph node negative breast tumors compared to patients diagnosed as having lymph node positive breast tumors.
  • a specific set of signature proteins specifically, specifically, RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB that are differently expressed in primary tumor samples obtained from patients diagnosed as having lymph node negative breast tumors compared to patients diagnosed as having lymph node positive breast tumors.
  • the inventors have therefore suggested that the identified signatory proteins described herein are suitable as a powerful tool for early diagnosis and prognosis of breast cancer metastasis.
  • the invention relates to a diagnostic and prognostic method for determining the progression of breast cancer in a subject.
  • the method of the invention comprises the steps of:
  • step (a) involves determining the expression level of at least one biomarker protein in at least one biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein/s, wherein said biomarker proteins are selected from RPS24 (40S ribosomal protein S24), LSM4 (U6 snRNA-associated Sm-like protein LSm4), RBM12B (RNA-binding protein 12B), RPS29 (40S ribosomal protein S29); RBM3 (Putative RNA-binding protein 3), PNP (Purine nucleoside phosphorylase), METAP2 (Methionine aminopeptidase 2;Methionine aminopeptidase), CAPS (Calcyphosin), EIF4A3 (Eukaryotic initiation factor 4A-III;Eukaryotic initiation factor 4A-III, N-terminally processed), SNX12 (Sorting nexin-12), SRSF1 (Serine/arginine
  • the present invention provides a diagnostic and prognostic method for determining the progression of breast cancer in a subject, the method comprising the steps of: (a) providing at least one detecting molecule/s each specific for at least one biomarker protein, specifically, at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB.
  • the detecting molecules may be provided in a diagnostic composition or in a kit either attached to a solid support or alternatively, in a mixture.
  • the method of the invention encompasses in certain embodiments also the provision of a composition, kit, solid support or mixture comprising at least one detecting molecule specific for at least one of said biomarker proteins of the invention.
  • the next step (b) requires determining the expression level of at least one of said biomarker protein/s in at least one biological sample of the diagnosed subject, to obtain an expression value for each of said at least one biomarker protein/s.
  • the final step (c) requires determining if the expression value obtained in step (b) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample.
  • the method of the invention may use as diagnostic and prognostic tool, the expression values of any one of the marker proteins described herein below.
  • determining the expression values of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB proteins may indicate if a subject belongs to a pre-established population associated with negative lymph node metastatic status or with positive lymph node metastatic status.
  • the biomarker protein of the invention is the 40S ribosomal protein S24 (RPS24) Protein.
  • RPS24 as described herein, refers to the human RPS24 (Protein IDs P62847; E7ETK0; A0A087WUS0). This protein is required for processing of pre-rRNA and maturation of 40S ribosomal subunits.
  • the RPS24 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 1.
  • the biomarker protein of the invention is the U6 snRNA-associated Sm-like protein LSm4 (LSM4) protein.
  • LSM4 as described herein, refers to the human LSM4 (protein IDs V9GZ56; Q9Y4Z0; U3KQS7; U3KQK1; M0QXB0). This protein binds specifically to the 3'-terminal U-tract of U6 snRNA.
  • the LSM4 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 2.
  • the biomarker protein of the invention is the RNA-binding protein 12B (RBM12B) protein.
  • RBM12B as described herein refers to the human RBM12B (Protein IDs Q8IXT5; B9ZVT1; E5RHG1; E5RJ83; E5RJV8; E5RJW8). This protein contains several RNA- binding motifs, potential transmembrane domains, and proline-rich regions.
  • the RBM12B protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 3.
  • the biomarker protein of the invention is the 40S ribosomal protein S29 (RPS29) protein.
  • RPS29 ribosomal protein S29
  • this protein refers to the human RPS29 (Protein IDs P62273; A0A087WTT6).
  • This protein is a member of the S 14P family of ribosomal proteins that acts as a component of the 40S subunit and.
  • the protein which contains a C2-C2 zinc finger-like domain that can bind to zinc, can enhance the tumor suppressor activity of Ras-related protein 1A (KREV1). It is located in the cytoplasm.
  • the RPS29 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 4.
  • the biomarker protein of the invention is the Putative RNA-binding protein 3 (RBM3) protein.
  • RBM3 Putative RNA-binding protein 3
  • this protein refers to the human RBM3 (Protein IDs P98179; A0A024QYX3).
  • This protein is a cold-inducible mRNA binding protein that enhances global protein synthesis at both physiological and mild hypothermic temperatures. It reduces the relative abundance of microRNAs, when overexpressed and enhances phosphorylation of translation initiation factors and active polysome formation.
  • the RBM3 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 5.
  • the biomarker protein of the invention is the Purine nucleoside phosphorylase (PNP) protein.
  • PNP Purine nucleoside phosphorylase
  • this biomarker refers to the human PNP (Protein IDs P00491; V9HWH6; Q8N7G1; G3V5M2; G3V2H3; G3V393).
  • This protein catalyze the phosphorolytic breakdown of the N-glycosidic bond in the beta- (deoxy)ribonucleo side molecules, with the formation of the corresponding free purine bases and pentose- 1 -phosphate.
  • the PNP protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 6.
  • the biomarker protein of the invention is the Methionine aminopeptidase 2; (METAP2) protein.
  • this biomarker refers to the human METAP2 (Protein IDs P50579; B4DUX5; G3V1U3; B3KWL6; F8VSC4).
  • This protein co-translationally removes the N-terminal methionine from nascent proteins. It protects eukaryotic initiation factor EIF2S 1 from translation-inhibiting phosphorylation by inhibitory kinases such as EIF2AK2/PKR and EIF2AK1/HCR and plays a critical role in the regulation of protein synthesis.
  • the METAP2 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 7.
  • the biomarker protein of the invention is the Calcyphosin (CAPS) protein.
  • this biomarker refers to the human CAPS (Protein IDs Q13938; K7ES72). This protein is a calcium-binding protein.
  • the CAPS protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 8.
  • the biomarker protein of the invention is the Eukaryotic initiation factor 4A-III, (EIF4A3), protein.
  • EIF4A3 Eukaryotic initiation factor 4A-III
  • this biomarker refers to the human EIF4A3 (Protein IDs P38919; A0A024R8W0; I3L3H2).
  • This protein is an ATP-dependent RNA helicase.
  • the EIF4A3 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 9.
  • the biomarker protein of the invention is the Sorting nexin-12 (SNX12), Protein.
  • this biomarker refers to the human SNX12 (Protein IDs Q9UMY4; Q3SYF1; A0A087X0R6). This protein may be involved in several stages of intracellular trafficking.
  • the SNX12 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 10.
  • the biomarker protein of the invention is the Serine/arginine-rich splicing factor 1 (SRSF1) protein.
  • SRSF1 Serine/arginine-rich splicing factor 1
  • this biomarker refers to the human SRSF1 (Protein. IDs Q07955; J3KTL2; Q59FA2; A8K1L8; J3KSR8; J3QQV5; J3KSW7).
  • This protein plays a role in preventing exon skipping, ensuring the accuracy of splicing and regulating alternative splicing. It interacts with other spliceosomal components, via the RS domains, to form a bridge between the 5'- and 3'-splice site binding components, Ul snRNP and U2AF.
  • the SRSF1 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 11.
  • the biomarker protein of the invention is the Ras-related protein Rab-5C (RAB5C) protein.
  • RAB5C Ras-related protein Rab-5C
  • this biomarker refers to the human RAB5C (Protein IDs P51148; A0A024R1U4; K7ERI8; F8VVK3; K7ENY4; F8VWU4; F8VSF8; K7EIP6; F8VWZ7).
  • RAB5C protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 12.
  • the biomarker protein of the invention is the Transcription factor BTF3 homolog 4 (BTF3L4) protein.
  • BTF3L4 Transcription factor BTF3 homolog 4
  • this biomarker refers to the human BTF3L4 (Protein IDs Q96K17; Q6PJ77; E9PL10). This is a protein-coding gene.
  • the BTF3L4 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 13.
  • the biomarker protein of the invention is the U6 snRNA-associated Sm- like protein LSm2 (LSM2) protein.
  • LSM2 U6 snRNA-associated Sm- like protein LSm2
  • this biomarker refers to the human LSM2 (Protein ID Q9Y333). This protein binds specifically to the 3 '-terminal U-tract of U6 snRNA and may be involved in pre-mRNA splicing.
  • the LSM2 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 14.
  • the biomarker protein of the invention is the Cystatin-B (CSTB) protein.
  • this biomarker refers to the human CSTB (Protein IDs P04080; Q76LA1).
  • This protein is an intracellular thiol proteinase inhibitor and tightly binding reversible inhibitor of cathepsins L, H and B.
  • the CSTB protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 15.
  • the expression value of at least one biomarker protein is determined.
  • the methods of the invention may involve determination of the expression level of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins. It should be noted that the biomarker proteins of the invention are disclosed in Table 4 herein after.
  • the method of the invention may involve in step (a) determination of the expression level of at least five biomarker proteins in at least one biological sample of the examined subject, to obtain an expression value for each of the at least five biomarker proteins.
  • at least five biomarker proteins may be selected from RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins.
  • such at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise LSM2, METAP2, RPS24, RBM12B and CAPS.
  • the at least five biomarker proteins may comprise RPS29, RBM3, PNP, METAP2 and RAB5C.
  • the five biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten of the LSM4, RPS29, RBM3, PNP, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4 and CSTB biomarker proteins of the invention.
  • the selected biomarker proteins may further comprise at least one of RPS24, LSM4, RBM12B, CAPS, EIF4A3, SNX12, SRSFl, BTF3L4, LSM2 and CSTB.
  • the method of the invention may involve in step (a) determination of the expression level of at least four biomarker proteins.
  • biomarker proteins such at least four of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise RPS24, LSM4, RBM12B and RPS29.
  • the four biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten or at least eleven of the RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention.
  • the method of the invention may involve in step (a) determination of the expression level of at least six biomarker proteins.
  • biomarker proteins such at least six of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise RPS24, LSM4, RBM12B, RPS29, RBM3 and PNP.
  • the six biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight or at least nine of the METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention.
  • the method of the invention may involve in step (a) determination of the expression level of at least ten biomarker proteins.
  • biomarker proteins such at least ten of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3 and SNX12.
  • the ten biomarker proteins may further comprise at least one, at least two, at least three, at least four, or at least five of the SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention.
  • the method of the invention may provide and use detecting molecules specific for at least one, at least five, at least four, at least six or at least ten of the biomarkers of Table 4 and further, detecting molecule/s specific for at least one additional biomarker protein. It should be noted that each detecting molecule is specific for one biomarker.
  • the method as well as the kits of the invention described herein after may provide and use further detecting molecules specific for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more, specifically, 110, 120, 130, 140, 150, 160, 170, 180, 190
  • the methods, compositions and kits of the invention may provide and use in addition to detecting molecules specific for at least one of the biomarkers disclosed in Table 4, also detecting molecule/s specific for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 or 85 biomarker proteins disclosed in Table 2, and optionally, further detecting molecule/s specific for additional at least one biomarker protein/s, for example, any of the biomarkers presented in Table 3, or
  • the methods, as well as the compositions and kits of the invention may provide and use detecting molecules specific for at least one additional biomarker protein and at most, 499 additional marker protein/s.
  • the methods and kit/s of the invention may provide and use detecting molecules specific for at least one of the biomarker proteins of Table 4, and detecting molecules specific for at least one additional biomarkers, provided that detecting molecules specific for 100, 150, 200, 250, 300, 350, 384, 400, 450 and 500 at the most biomarker proteins are used.
  • the at least one additional biomarker protein may comprise any of the biomarker proteins presented in Figure 12B.
  • such additional biomarker proteins may comprise at least one of EDF1, PLEC, SNRPG, SRP9, RPL8, RPL23, RPL35A, RPS 15A, RPS23 and RPS28.
  • the methods of the invention as well as the compositions and kits described herein after may involve the determination of the expression levels of the biomarker proteins of the invention and/or the use of detecting molecules specific for said biomarker proteins. Specifically, at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen of the biomarker protein/s of the invention that may further comprise any additional biomarker proteins or control reference protein provided that 500 at the most biomarker proteins and control reference proteins are used.
  • the at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen of the biomarker protein/s of the invention may form at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or 100% of the biomarker proteins determined by the methods of the invention.
  • the detecting molecules specific for at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten at least eleven, at least twelve, at least thirteen, at least fourteen or at least fifteen of the biomarker protein/s of the invention, that are used by the methods of the invention and comprised within any of the compositions and kits of the invention may form at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or 100% of detecting molecules used in accordance with the invention. It should be appreciated that for each of the selected biomarker proteins at least one detecting molecules may be used. In case more than one detecting molecule is used for a certain biomarker protein, such detecting molecules may be either identible or different.
  • the proteins that were found to be down-regulated in primary tumor samples of patients suffering from lymph node negative breast tumors represent several cellular functions such as ribosomal proteins (RPS24 and RPS29), proteins involved in pre-mRNA splicing (LSM2 and LSM4) and RNA binding proteins (RBM3 and RBM12B), and also include the two possible invasiveness markers METAP2 and PNP.
  • the protein signature described above may predict the involvement of lymph node in cancer progression, based on expression levels in the primary tumors.
  • cancer is used herein interchangeably with the term “tumor” and denotes a mass of tissue found in or on the body that is made up of abnormal cells.
  • breast cancer refers to a cancer that develops from breast tissue. Development of breast cancer is often associated with a lump in the breast, a change in breast shape, dimpling of the skin, fluid coming from the nipple, or a red scaly patch of skin.
  • Breast cancer classification divides breast cancer into categories according to different schemes, each based on different criteria and serving a different purpose.
  • the major categories are the histopathological type, the grade of the tumor, the stage of the tumor, and the expression of proteins and genes.
  • breast cancers that tend to be aggressive and life-threatening should be treated with aggressive treatments that have major adverse effects.
  • Other breast cancers which are less aggressive can be treated with less aggressive treatments.
  • Breast cancers can be classified by criteria, each one influences treatment response and prognosis.
  • Classification includes at least one of the following parameters histopathological type, grade, stage,
  • Staging of breast cancer may be done by various methods for example using TNM staging which takes into account the size of the tumor (T), whether the cancer has spread to the lymph glands
  • lymph nodes (lymph nodes) (N), and whether the tumor has spread anywhere else in the body (M - for metastases).
  • staging can be expressed as a number on a scale of 0 through IV— with stage 0 describing non-invasive cancers that remain within their original location and stage IV describing invasive cancers that have spread outside the breast to other parts of the body.
  • the breast tumor is a non-invasive tumor. In some other embodiments, the breast tumor is an invasive tumor.
  • non-invasive cancer it should be noted as a cancer that do not grow into or invade normal tissues within or beyond the primary location, for example the breast. Non-invasive cancers are sometimes called carcinoma in situ ("in the same place") or pre-cancers. In connection with breast cancer, non invasive cancer stays in milk ducts or lobules in the breast.
  • invasive cancers it should be noted as caner that invade and grow in normal, healthy tissues to form metastasis.
  • metastatic cancer or “metastatic status” refers to a cancer that has spread from the place where it first started to another place in the body and specifically to the lymph node.
  • a tumor formed by metastatic cancer cells is called a metastatic tumor or a metastasis.
  • lymph node negative refers to a primary non-invasive breast tumor that remain within the breast.
  • lymph node positive refers to a primary invasive breast tumor that has spread outside the breast into the lymph node.
  • characterization of a breast tumor as non-invasive or invasive may depend on information collected from different methods and depends on the detection capability of each one of the methods. Therefore, when referring to LNN or LNP it should be understood as detection level of the method used.
  • Receptor status can also be used for classification of breast cancer into several molecular classes. The three most important receptors in the classification being: estrogen receptor (ER), progesterone receptor (PR), and HER2/neu.
  • ER+ cells expressing ER
  • Luminal B Breast cells characterized by being ER+ and but often high grade.
  • the diagnosed subject may be a subject suffering from a luminal A breast tumor or a luminal B breast tumor.
  • the method of the invention may be used as a diagnostic and prognostic tool by detecting the expression values of at least one of the marker proteins described herein.
  • determining the expression values of at least one of marker proteins described herein may differentiate metastatic breast cancer from non-metastatic breast cancer, namely at an early stage when a subject has a primary tumor, it may be possible to predict or prognose if a breast tumor will be LNN or LNP.
  • determining the expression values of at least one of the following biomarker proteins may indicate if a subject belongs to a pre-established population associated with negative lymph node metastatic status or positive lymph node metastatic status.
  • lymph node status of cancer patients and specifically breast cancer is considered to be most important variable in the management of the disease (Jatoi et al., 1999).
  • the methods of the invention may further comprise determining the expression level of at least one other biomarker protein in at least one biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein/s.
  • biomarker proteins may be at least one of Acyl-coenzyme A thioesterase 1 ; Acyl-coenzyme A thioesterase 2, mitochondrial (ACOT1; ACOT2), Small nuclear ribonucleoprotein G;Small nuclear ribonucleoprotein G-like protein (SNRPG), 40S ribosomal protein S27-like;40S ribosomal protein S27 (RPS27L), Nitrilase homolog 1 (NIT1), Protein RPS 10-NUDT3 (RPS 10-NUDT3), Protein PRRC1 (PRRC1), 39S ribosomal protein L43, mitochondrial (MRPL43), Ras suppressor protein 1 (RSU1), Nucleolysin TIAR (TIAL1), 28S ribosomal protein S 17, mitochondrial (MRPS 17), Replication protein A 32 kDa subunit (RPA2), Cytochrome c oxidase subunit 6B 1 (COX6B 1), Transformer-2
  • the method of the invention may involves the determination of the expression level of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 of the biomarker proteins of the invention, specifically, the proteins disclosed in Table 4, and optionally further at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 34, 35, 36, 37, 38, 39, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 or 85 of the biomarker proteins disclosed in Table 2.
  • the method of the invention may involve determination of the expression level of additional biomarker protein/s, specifically, additional at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96
  • the additional biomarker proteins may include any breast cancer-related proteins, such as estrogen receptor, progesterone receptor, ErbB2/Her2, EGF receptor, ki67, keratins, Mucin 1 (MUC1), Catapsin D and more.
  • further biomarker proteins may include at least one of ACOTl/2, SLC25A11, GLUL, POSTN, COL12A1, CDKN2A, LGALS 1, and any biomarker protein presented in the figures of the present invention, for example in Figure 9.
  • the methods of the invention may involve determination of the expression level of at least one control reference protein/s in at least one sample of the diagnosed subject. Control reference protein/s will be described in more details herein after.
  • the second step (b) involves calculating and determining if the expression value obtained in step (a) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample.
  • the methods of the invention may further comprise determining that a subject classified as belonging to a population having an LNP tumor, will develop, or has an increased probability to develop metastasis to the lymph node/s.
  • the inventors also determined the differences in the protein expression in a group of primary metastatic tumors and the corresponding lymph node metastases. As shown in Example 5, a set of ten proteins was found to be differently expressed in the primary tumor and the metastatic tissue.
  • the methods of the invention may be further used to differentiate between primary tumor and metastatic tissue.
  • the methods of the invention further comprise (a) determining the expression level of at least one biomarker protein in at least one biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein/s, wherein said biomarker proteins are selected from Periostin (POSTN), AP-2 complex subunit mu (AP2M1), Mimecan (OGN), Collagen alpha-l(XII) chain (COL12A1), Cyclin-dependent kinase inhibitor 2A, isoforms 1/2/3 (CDKN2A), Galectin-1 (LGALS 1), Synaptosomal-associated protein 23 (SNAP23), Adenine phosphoribosyltransferase (APRT) Protein-glutamine gamma-glutamyltransferase 2 (TGM2) and Ester hydrolase Cl lorf54 (Cl lorf54), or any combination thereof; and (b) determining if the expression value obtained in step (a) is any one of positive or negative with respect
  • POSTN
  • POSTN and LGALS 1 may facilitate migration to lymph nodes.
  • the invasive mechanisms of cells from the primary tumor may be 'shut off in the lymph node in favor of renewed proliferation, indicated by decreased expression of cell-cycle regulator CDKN2A in metastatic cells.
  • lymph node metastases As described herein, the comparison between primary breast tumors and their matched lymph node metastases showed a set of ten proteins that were differently expressed. Thus, it was concluded by the inventors that there is a similarity between the protein signature of tumors and the protein signature of lymph node metastases, despite the different microenvironment. This suggests that the lymph node metastases retain the expression patterns of their original primary breast tumor. This result is in agreement with several gene expression studies that examined both lymph node metastases and distant metastases.
  • the present invention also provides disease diagnosis and specifically methods of cancer diagnosis.
  • the inventors have analyzed protein expression in healthy tissue and tumor tissue.
  • the expression level of at least one of the biomarker proteins described herein is being determined.
  • level of expression or “expression level” are used interchangeably and generally refer to a numerical representation of the amount (quantity) of an amino acid product or polypeptide or protein in a biological sample.
  • level of expression or “expression level” refers to the numerical representation of the amount (quantity) of polynucleotide which may be gene in a biological sample.
  • “Expression” generally refers to the process by which gene-encoded information is converted into the structures present and operating in the cell.
  • gene expression values may be measured in the protein level, for example by MS methods or alternatively by immunological methods.
  • the expression may be measured in the nucleic acid level, for example using Real-Time Polymerase Chain Reaction, sometimes also referred to as RT-PCR or quantitative PCR (qPCR).
  • RT-PCR Real-Time Polymerase Chain Reaction
  • qPCR quantitative PCR
  • any gene encoding any of the biomarker proteins of the invention may refer to transcription into a polynucleotide and translation into a polypeptide. Fragments of the transcribed polynucleotide, the translated protein, or the post-translationally modified protein shall also be regarded as expressed whether they originate from a transcript generated by alternative splicing or a degraded transcript, or from a post-translational processing of the protein, e.g., by proteolysis. Methods for determining the level of expression of the biomarkers of the invention will be described in more detail herein after.
  • the methods of the invention refer to the level of the biomarker protein/s in the sample. It should be understood that the level of the protein reflects the level of expression but may also reflect the stability of the biomarker protein.
  • the method of the invention further comprises an additional and optional step of normalization.
  • the level of expression of at least one suitable control reference protein is being determined in the same sample.
  • a control reference protein may be any protein that is not differentially expressed in different tissues or different pathologic conditions.
  • appropriate control reference proteins in connection with the present invention are proteins that are expressed equally in primary tumors of patients diagnosed with LNN vs. LNP, and therefore cannot be used to distinguish between LNP and LNN patients based on their expression in primary tumors.
  • control reference proteins may include ARCN1 (Archain 1), MPZL1 (Myelin Protein Zero-Like 1), NSF (N-ethylmaleimide-sensitive factor), PRKCD (Protein Kinase C, Delta), CAT (catalase), actin, tubulin, or other cytoskeletal proteins.
  • the expression level of at least one of the biomarkers of the invention obtained in step (a) is normalized according to the expression level of said at least one reference control protein obtained in the additional optional step in said test sample, thereby obtaining a normalized expression value.
  • similar normalization is performed also in at least one control sample or a representing standard when applicable.
  • expression value refers to the result of a calculation, that uses as an input the "level of expression” or “expression level” obtained experimentally and by normalizing the "level of expression” or “expression level” by at least one normalization step as detailed herein, the calculated value termed herein "expression value” is obtained.
  • normalized values are the quotient of raw expression values of marker proteins, divided by the expression value of a control reference protein from the same sample. Any assayed sample may contain more or less biological material than is intended, due to human error and equipment failures. Importantly, the same error or deviation applies to both the marker protein of the invention and to the control reference protein, whose expression is essentially constant. Thus, division of the marker protein raw expression value by the control reference protein raw expression value yields a quotient which is essentially free from any technical failures or inaccuracies (except for major errors which destroy the sample for testing purposes) and constitutes a normalized expression value of said marker protein. This normalized expression value may then be compared with normalized cutoff values, i.e., cutoff values calculated from normalized expression values.
  • the control reference protein may be a protein that maintains stable in all samples analyzed.
  • Normalized biomarker protein expression level values that are higher (positive) or lower (negative) in comparison with a corresponding predetermined standard expression value or a cut-off value in a control sample predict to which population of patients the tested sample belongs or more specifically the disease stage, or the metastatic status of the subject.
  • an important step in the method of the inventions is determining whether the normalized expression value of any one of the biomarker proteins is changed compared to a pre-determined cut off, or is within the range of expression of such cutoff.
  • the next step of the method of the invention involves calculating and determining if the expression value obtained in step (a) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample.
  • Such step involves calculating and measuring the difference between the expression values of the examined sample and the cutoff value pre-determined for a certain population and determining whether the examined sample can be defined as positive or negative, with respect to said population.
  • the second step (b) of the method of the invention involves comparing the expression values determined for the tested sample with predetermined standard values or cutoff values, or alternatively, with expression values of at least one control sample.
  • comparing denotes any examination of the expression level and/or expression values obtained in the samples of the invention as detailed throughout in order to discover similarities or differences between at least two different samples. It should be noted that in some embodiments, comparing according to the present invention encompasses the possibility to use a computer based approach.
  • cutoff value is a value that meets the requirements for both high diagnostic sensitivity (true positive rate) and high diagnostic specificity (true negative rate).
  • sensitivity and “specificity” are used herein with respect to the ability of one or more markers, to correctly classify a sample as belonging to a pre-established population associated with negative lymph node metastatic status, or alternatively, to a pre- established population associated with positive lymph node metastatic status.
  • “Sensitivity” indicates the performance of the bio-marker of the invention, with respect to correctly classifying samples as belonging to pre-established populations that are likely to suffer from a disease or disorder or characterized at different stages of a disease to respond to therapy or to relapse, when applicable, wherein said bio-marker are consider here as any of the options provided herein.
  • Specificity indicates the performance of the bio-marker of the invention with respect to correctly classifying samples as belonging to pre-established populations of subjects suffering from the same disorder or populations of subjects that are likely to respond to a specific treatment or unlikely to relapse as will be discussed herein after.
  • sensitivity relates to the rate of correct identification of the patients (samples) as such out of a group of samples
  • specificity relates to the rate of correct identification of lymph node metastatic status samples as such out of a group of samples.
  • Cutoff values may be used as control sample/s or in addition to control sample/s, said cutoff values being the result of a statistical analysis of biomarker protein expression value/s (specifically the biomarker proteins of the invention) differences in pre-established populations healthy, metastatic, LNP or LNN.
  • a given population having specific clinical parameters will have a defined likelihood to have positive lymph node metastasis or negative lymph node metastasis based on the expression values of the marker proteins being above or below said cutoff values.
  • an individual having a positive expression value being up- regulated of least one of the following biomarker protein/s 40S ribosomal protein S24, U6 snRNA- associated Sm-like protein LSm4, RNA-binding protein 12B, 40S ribosomal protein S29; Putative RNA-binding protein 3, Purine nucleoside phosphorylase, Methionine aminopeptidase 2;Methionine aminopeptidase, Calcyphosin, Eukaryotic initiation factor 4A-III;Eukaryotic initiation factor 4A-III, N-terminally processed, Sorting nexin-12, (erine/arginine-rich splicing factor 1, Ras- related protein Rab-5C, Transcription factor BTF3 homolog 4, U6 snRNA-associated Sm-like protein LSm2 and Cystatin-B may be considered as belonging to a pre-established population associated with positive lymph node metastatic status.
  • a subject presenting a negative expression value that reflects down- reulation of at least one biomarker protein/s 40S ribosomal protein S24, U6 snRNA-associated Sm- like protein LSm4, RNA-binding protein 12B, 40S ribosomal protein S29; Putative RNA-binding protein 3, Purine nucleoside phosphorylase, Methionine aminopeptidase 2;Methionine aminopeptidase, Calcyphosin, Eukaryotic initiation factor 4A-III;Eukaryotic initiation factor 4A-III, N-terminally processed, Sorting nexin-12, (erine/arginine-rich splicing factor 1, Ras-related protein Rab-5C, Transcription factor BTF3 homolog 4, U6 snRNA-associated Sm-like protein LSm2 and Cystatin-B may be considered as belonging to a pre-established population associated with negative lymph node metastatic status.
  • a negative or positive determination of the expression value as compared to the predetermined cutoff values also encompass values that are within the range of said cutoff. More specifically, an expression value that is determined by the method of the invention as "positive" when compared to a predetermined cutoff of population of LNP patients, or for at least one known LNP patient, may indicate that the examined subject belongs to LNP population, in case that the expression value is either higher (positive) or within the range (the average values of the cutoff predetermined for LNP patient population). In a similar manner, a subject exhibiting an expression value that is "negative” (that is down-regulated) as compared to the cutoff patients, may be considered as belonging to LNN population. In more specific embodiments, the expression value of such subject should fall within the range of the cutoff value predetermined for LNN population. In some embodiments, "fall within the range” encompass values that differ from the cutoff value in about 1% to about 50% or more.
  • the nature of the invention is such that the accumulation of further patient data may improve the accuracy of the presently provided cutoff values, which are based on an ROC (Receiver Operating Characteristic) curve generated according to said patient data using analytical software program.
  • the biomarker protein expression values are selected along the ROC curve for optimal combination of prognostic sensitivity and prognostic specificity which are as close to 100 percent as possible, and the resulting values are used as the cutoff values that distinguish between patients who are diagnosed with positive lymph node metastasis at a certain rate, and those who will not (with said given sensitivity and specificity).
  • ROC curve may evolve as more and more data and related biomarker gene expression values are recorded and taken into consideration, modifying the optimal cutoff values and improving sensitivity and specificity.
  • the provided cutoff values should be viewed as a starting point that may shift as more data allows more accurate cutoff value calculation.
  • the presently provided values already provide good sensitivity and specificity, and are readily applicable in current clinical use, even in patients diagnosed with different cancer stages.
  • the expression value determined for the examined sample (or alternatively, the normalized expression value) is compared with a predetermined cutoff or a control sample. More specifically, in certain embodiments, the expression value obtained for the examined sample is compared with a predetermined standard or cutoff value.
  • the predetermined standard expression value, or cutoff value has been predetermined and calculated for a population comprising at least one of healthy subjects, subjects suffering from any disorder, subjects suffering from different stages of any disorder, subjects that respond to treatment, non-responder subjects, subjects in remission and subjects in relapse.
  • predetermined cutoff values may be calculated for a population of subject diagnosed with breast cancer, subjects diagnosed with metastatic breast cancer, subjects diagnosed with LNN and subjects diagnosed with LNP.
  • control sample is being used (instead of, or in addition to, pre-determined cutoff values)
  • the normalized expression values of the biomarker proteins used by the invention in the test sample are compared to the expression values in the control sample.
  • control sample may be obtained from at least one of a healthy subject, a subject suffering from a disorder at a specific stage, a subject suffering from a disorder at a different specific stage a subject that responds to treatment, a non-responder subject, a subject in remission and a subject in relapse.
  • predetermined cutoff values may be calculated for a population of subject diagnosed with breast cancer, subjects diagnosed with metastatic breast cancer, subjects diagnosed with LNN and subjects diagnosed with LNP.
  • Standard or a "predetermined standard” as used herein, denotes either a single standard value or a plurality of standards with which the level at least one of the biomarker protein expression from the tested sample is compared.
  • the standards may be provided, for example, in the form of discrete numeric values or is calorimetric in the form of a chart with different colors or shadings for different levels of expression; or they may be provided in the form of a comparative curve prepared on the basis of such standards (standard curve).
  • determining the level of expression of at least one of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins is performed by the step of contacting at least one detecting molecule or any combination or mixture of plurality of detecting molecules with a biological sample of said subject, or with any protein or nucleic acid product obtained therefrom, wherein each of said detecting molecules is specific for one of said biomarker proteins.
  • the methods of the invention may further comprise the step of providing at least one detecting molecule specific for determining the expression of at least on of said biomarker proteins of the invention.
  • detecting molecules may be provided as a mixture, as a composition or as a kit.
  • the at least one detecting molecules may be provided as a mixture of detecting molecules, wherein each detecting molecule is specific for one biomarker protein. It should be appreciated however, that for each biomarker protein, one or several specific detecting molecules may be used and provided.
  • the detecting molecules may be provided separately for each biomarker protein, e.g., in specific tube, containers, slots, spots, wells, and the like. It further alternative embodiments, the detecting molecules may be attached or immobilized to a solid support, specifically, in recorded location.
  • determining the level of expression of at least one of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins is performed by the step of contacting at least one detecting molecule or any combination or mixture of plurality of detecting molecules with a biological sample of said subject, or with any protein or nucleic acid product obtained therefrom, wherein each of said detecting molecules is specific for one of said biomarker proteins.
  • contacting mean to bring, put, incubates or mix together. As such, a first item is contacted with a second item when the two items are brought or put together, e.g., by touching them to each other or combining them.
  • the term "contacting” includes all measures or steps which allow interaction between the at least one of the detection molecules of at least one of the biomarker proteins, and optionally, for at least one suitable control reference protein of the tested sample.
  • the contacting is performed in a manner so that the at least one of detecting molecule of at least one of the biomarker proteins for example, can interact with or bind to the at least one of the biomarker proteins, in the tested sample.
  • the binding will preferably be non-covalent, reversible binding, e.g., binding via salt bridges, hydrogen bonds, hydrophobic interactions or a combination thereof.
  • the detection step further involves detecting a signal from the detecting molecules that correlates with the expression level of at least one of the biomarker proteins and in the sample from the subject, by a suitable means.
  • the signal detected from the sample by any one of the experimental methods detailed herein below reflects the expression level of at least one of the biomarker proteins. It should be noted that such signal-to- expression level data may be calculated and derived from a calibration curve.
  • the method of the invention may optionally further involve the use of a calibration curve created by detecting a signal for each one of increasing pre-determined concentrations of at least one of the biomarker proteins. Obtaining such a calibration curve may be indicative to evaluate the range at which the expression levels correlate linearly with the concentrations of at least one of the biomarker proteins. It should be noted in this connection that at times when no change in expression level of at least one of the biomarker proteins is observed, the calibration curve should be evaluated in order to rule out the possibility that the measured expression level is not exhibiting a saturation type curve, namely a range at which increasing concentrations exhibit the same signal.
  • the detecting molecules used for determining the expression levels at least one of the biomarker proteins are selected from isolated detecting amino acid molecules and isolated detecting nucleic acid molecules. It should be noted that the invention further encompasses any combination of nucleic and amino acids for use as detecting molecules for the methods of the invention. As noted above, in the first step of the method of the invention, the sample or any protein or nucleic acid obtained therefrom, is contacted with the detecting molecules of the invention.
  • a protein is composed of less than 200, less than 175, less than 150, less than 125, less than 100, less than 50, less than 45, less than 40, less than 35, less than 30, less than 25, less than 20, less than 15, less than 10, or less than 5 amino acids linked together by peptide bonds.
  • a protein is composed of at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500 or more amino acids linked together by peptide bonds.
  • peptide bond as described herein is a covalent amid bond formed between two amino acid residues.
  • the detecting molecules used by the methods of the invention may be recombinantly expressed or synthetically prepared.
  • the recombinantly or synthetically expressed and prepared detecting molecules may be labeled or tagged. It should be noted that in some embodiments, these detecting molecules may be isolated detecting molecules.
  • Recombinant proteins denotes proteins encoded by a recombinant DNA which is a genetically engineered DNA formed by laboratory methods of genetic recombination to bring together genetic material from multiple sources and thus creating variable sequences.
  • Recombinant proteins may be produced mainly, but not limited, by molecular cloning, namely incorporating the recombinant DNA into a living cell (e.g. bacteria or yeast) and using its system to express the DNA into mRNA and protein thereof.
  • MS Mass spectrometry
  • immunological techniques such as Western Blotting, Immunoprecipitation, ELISAs, protein microarray analysis, Flow cytometry and the like
  • the amino acid-based detecting molecules may comprise at least one of: (a) at least one labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof; optionally, such labeled proteins may be recombinant or synthetically produced proteins, (b) antibodies specific for said at least one of said biomarker proteins; (c) peptide aptamers specific for said at least one of said biomarker proteins; and (d) any combination of (a), (b) and (c).
  • the detecting molecules may be at least one labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragments, peptides or mixture thereof.
  • the term "labeled” or "tagged” may refer to direct labeling of the protein via, e.g., coupling (i.e., physically linking) or incorporating of a detectable substance to the protein.
  • Useful labels in the present invention may include but are not limited to include isotopes (e.g.
  • radiolabels e.g., 3 H, 125 I, 35 S, 14 C, or 32 P
  • magnetic beads e.g. DYNABEADS
  • fluorescent dyes e.g., fluorescein isothiocyanate, Texas red, rhodamine, green fluorescent protein, and the like
  • enzymes e.g., horseradish peroxidase, alkaline phosphatase and others commonly used in an ELISA and competitive ELISA, histochemistry and other similar methods known in the art
  • colorimetric labels such as colloidal gold or colored glass or plastic (e.g. polystyrene, polypropylene, latex, etc.) beads.
  • the protein may be tagged.
  • tags may be also used, for example, His, myc, HA, GFP, ABP, GST, biotin and the like, "tagged” as used herein may further include fusion or linking of the biomarker protein or any fragment or peptide thereof, that serves herein as a detecting molecule, a tag that in some embodiments may contain several amino acids or a peptide that may be recognized by affinity or immunologically, using specific antibodies.
  • the detecting molecules may be at least one, optionally, recombinant and/or isolated, labeled or tagged biomarker protein that may be any one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof. It should be noted that as indicated above, the invention encompasses the use of any biomarker protein, specifically, the biomarkers disclosed in the present invention, as detecting molecule/s.
  • the invention encompasses the use of any of the biomarker proteins of Table 4, as well as any of the biomarker proteins of Tables 2 and 3 as detecting molecule/s as described herein.
  • the biomarker proteins or any fragments or peptides thereof may be fluorescently labeled.
  • the biomarker proteins or any fragments or peptides thereof may be isotope labeled.
  • the term "recombinant isotope labeled" denotes a protein 'labeled' by replacing specific atoms by their isotope.
  • radiolabels may be detected using photographic film or scintillation counters
  • fluorescent markers may be detected using a photodetector to detect emitted illumination
  • Enzymatic labels are typically detected by providing the enzyme with a substrate and detecting the reaction product produced by the action of the enzyme on the substrate, and colorimetric labels are detected by simply visualizing the colored label.
  • the biomarker proteins of the invention or any fragment or peptide thereof when recombinantly expressed and labeled or tagged, may be used as detecting molecules for determining the quantity or level of expression of the biomarker proteins of the invention in the examined sample.
  • labeled form includes an isotope labeled form. Specifically, the labeled form is a chemically or metabolically isotope labeled, and more specifically a metabolically isotope labeled form of the biomarker proteins of the invention.
  • isotope labeled forms of the biomarker protein/s or any fragments or peptides thereof in accordance with the present invention are variants of naturally occurring molecules, in whose structure one or more atoms have been substituted with atom(s) of the same element having a different atomic weight, although isotope labeled forms in which the isotope has been covalently linked either directly or via a linker, or wherein the isotope has been complexed to the biomarker proteins are likewise contemplated. In either case, the isotope may be stable isotope.
  • a stable isotope as referred to herein is a non-radioactive isotopic form of an element having identical numbers of protons and electrons, but having one or more additional neutron(s), which increase(s) the molecular weight of the element.
  • the stable isotopes may be selected from the group consisting of 2 H, 13 C, 15 N, 170, 180, 33 P, 34 S and combinations thereof. Particularly specific examples include 13 C and 5 N, and combinations thereof.
  • a labeled reference biomarker (used as detecting molecule) can be synthesized using isotope labeled amino acids as precursor molecules, or chemically modified. Modification and labeling can be done on whole proteins or their fragments.
  • ICAT isotope-coded affinity tag
  • VICAT reagents label reference biomolecule such as proteins at the alkylation step of sample preparation (WO2004079370).
  • Visible ICAT reagents VIC AT reagents
  • VICAT-type reagent contains as a detectable moiety a fluorophore or radiolabel.
  • iTRAQ and similar methods may likewise be employed.
  • Metabolic labeling may also be used to produce the labeled reference biomarkers.
  • cells can be grown on media containing isotope labeled precursor molecules, such as isotope labeled amino acids, that are incorporated into proteins or peptides, which are thereby metabolically labeled.
  • the metabolic isotope labeling may be a stable isotope labeling with amino acids in cell culture (SILAC). If metabolic labeling is used, and the labeled form of the one or the plurality of reference biomarker protein/s is a SILAC labeled form of the reference biomarker protein/s, the standard mixture as defined above is also referred to as SUPER-SILAC mix.
  • the detecting amino acid molecules applicable for the invention may be isolated antibodies, with specific binding selectively to at least one of said biomarker proteins. More specifically, antibodies that specifically bind at least one of the biomarker proteins of the invention as listed in Table 4, and optionally, at least one of the biomarker proteins listed in Tables 2 and 3. It should be understood that each antibody specifically recognizes one biomarker protein.
  • the level of expression of at least one of the biomarker protein may be determined using an immunoassay which may be an assay that includes but not limited to FACS, a Western blot, an ELISA, a RIA, a slot blot, a dot blot, immune-histochemical assay and a radio-imaging assay. It should be noted that such assay may be performed using microarray protein arrays.
  • antibody as used in this invention includes whole antibody molecules as well as functional fragments thereof, such as Fab, F(ab')2, and Fv that are capable of binding with antigenic portions of the target polypeptide, i.e. at least one of the biomarker protein.
  • the antibody may be preferably monospecific, e.g., a monoclonal antibody, or antigen-binding fragment thereof.
  • monospecific antibody refers to an antibody that displays a single binding specificity and affinity for a particular target, e.g., epitope. This term includes a "monoclonal antibody” or “monoclonal antibody composition”, which as used herein refer to a preparation of antibodies or fragments thereof of single molecular composition.
  • the antibody can be a human antibody, a chimeric antibody, a recombinant antibody, a humanized antibody, a monoclonal antibody, or a polyclonal antibody.
  • the antibody can be an intact immuno globulin, e.g., an IgA, IgG, IgE, IgD, lgM or subtypes thereof.
  • the antibody can be conjugated to a labeling moiety as discussed above.
  • antibody also encompasses antigen-binding fragments of an antibody.
  • antigen -binding fragment of an antibody (or simply “antibody portion,” or “fragment”), as used herein, may be defined as follows:
  • Fab the fragment which contains a monovalent antigen-binding fragment of an antibody molecule, can be produced by digestion of whole antibody with the enzyme papain to yield an intact light chain and a portion of one heavy chain;
  • Fab' the fragment of an antibody molecule that can be obtained by treating whole antibody with pepsin, followed by reduction, to yield an intact light chain and a portion of the heavy chain; two Fab' fragments are obtained per antibody molecule;
  • Fv defined as a genetically engineered fragment containing the variable region of the light chain and the variable region of the heavy chain expressed as two chains
  • Single chain antibody (“SCA”, or ScFv), a genetically engineered molecule containing the variable region of the light chain and the variable region of the heavy chain, linked by a suitable polypeptide linker as a genetically fused single chain molecule.
  • Purification of serum immunoglobulin antibodies can be accomplished by a variety of methods known to those of skill in the art including, precipitation by ammonium sulfate or sodium sulfate followed by dialysis against saline, ion exchange chromatography, affinity or immuno-affinity chromatography as well as gel filtration, zone electrophoresis, etc.
  • the antibodies used by the present invention may optionally be covalently or non- covalently linked to a detectable label or tag.
  • the label and can also refer to indirect labeling of the protein by reactivity with another reagent that is directly labeled. Examples of indirect labeling include detection of at least one of the biomarker protein/s of the invention using a fluorescently labeled secondary antibody. More specifically, detectable labels suitable for such use include any composition detectable by spectroscopic, photochemical, biochemical, immunochemical, electrical, optical or chemical means.
  • each antibody is specific for one of the biomarker proteins of the invention, specifically, those disclosed in Table 4, and optionally, those disclosed in Tables 2 and 3. It should be appreciated that antibodies that may be used by the methods as well as the compositions and kits of the invention, may be antibodies directed not only against the biomarker proteins of the invention, but also in case the biomarkers are tagged, the antibodies may be directed against said tags.
  • binding specificity refers to a binding reaction which is determinative of the presence of the epitope in a heterogeneous population of proteins and other biologies.
  • "selectively bind" in the context of proteins encompassed by the invention refers to the specific interaction of a any two of a peptide, a protein, a polypeptide an antibody, wherein the interaction preferentially occurs as between any two of a peptide, protein, polypeptide and antibody preferentially as compared with any other peptide, protein, polypeptide and antibody.
  • the specified antibodies bind to a particular epitope at least two times the background and more typically more than 10 to 100 times background.
  • “Selective binding”, as the term is used herein, means that a molecule binds its specific binding partner with at least 2-fold greater affinity, and preferably at least 10-fold, 20-fold, 50-fold, 100-fold or higher affinity than it binds a non- specific molecule.
  • the antibodies used by the methods of the invention may be in some embodiments antibodies that are not naturally occurring antibodies. More specifically, the antibodies are not produced naturally in the body, and more specifically, it should be appreciated that production thereof involves immunological and recombinant techniques.
  • immunoassay formats may be used to select antibodies specifically immuno-reactive with a particular protein or carbohydrate.
  • solid-phase ELISA immunoassays are routinely used to select antibodies specifically immuno-reactive with a protein or carbohydrate.
  • epitope is meant to refer to that portion of any molecule capable of being bound by an antibody which can also be recognized by that antibody.
  • Epitopes or "antigenic determinants” usually consist of chemically active surface groupings of molecules such as amino acids or sugar side chains and have specific three dimensional structural characteristics as well as specific charge characteristics.
  • the detecting molecules are peptide aptamers specific for said at least one of said biomarker proteins.
  • eptide aptamers ⁇ as used herein refers to small peptides with a single variable loop region tied to a protein scaffold on both ends that binds to a specific molecular target (e.g. protein), and which are bind to their targets only with said variable loop region and usually with high specificity properties.
  • the expression level of the at least one of the biomarker protein, in the tested sample can be determined using different methods known in the art, specifically method disclosed herein below as non- limiting examples.
  • the detecting molecules may be at least one isolated, optionally recombinant or synthetic labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or fragments, peptides or mixture thereof, and optionally, at least one of the biomarkers of any one of the biomarker proteins disclosed in Tables 2 and 3.
  • the determination of the expression level of said at least one biomarker protein/s may be performed by mass spectrometry.
  • Mass spectrometry is used herein as an analytical chemistry technique to identify the amount and type of chemicals present in a sample by measuring the mass-to-charge ratio and abundance of gas-phase ions.
  • a mass spectrum is a plot of the ion signal as a function of the mass-to-charge ratio. The spectra are used to determine the elemental or isotopic signature of a sample, the masses of particles and of molecules, and to elucidate the chemical structures of molecules, such as peptides and other chemical compounds.
  • Mass spectrometry-based absolute quantification assays that generally require recombinant expression of full length, labeled protein standards.
  • Mass spectrometry is not inherently quantitative but many methods have been developed to overcome this limitation. Most of them are based on stable isotopes and introduce a mass shifted version of the peptides of interest, which are then quantified by their "heavy" to "light” ratio. Stable isotope labeling is either accomplished by chemical addition of labeled reagents, enzymatic isotope labeling, or metabolic labeling. Generally, these approaches are used to obtain relative quantitative information on protein expression levels in a light and a heavy labeled sample.
  • SILAC stable isotope labeling by amino acids in cell culture
  • Labeled protein can also be used as internal standards for determining expression levels of a cell or tissue protein of interest, such as in the spike-in SILAC approach.
  • absolute quantification AQUA
  • quantification concatamer QConCAT
  • PSAQ protein standard absolute quantification
  • SILAC absolute SILAC
  • FlexiQuant Several methods for absolute quantification have emerged over the last years and may be applicable for the present invention, including absolute quantification (AQUA), quantification concatamer (QConCAT), protein standard absolute quantification (PSAQ), absolute SILAC, and FlexiQuant. They all quantify the endogenous protein of interest by the heavy to light ratios to a defined amount of the labeled counterpart spiked into the sample and are chiefly distinguished by either spiking in heavy labeled peptides or heavy labeled full length proteins.
  • the AQUA strategy is convenient and streamlined: proteotypic peptides are chemically synthesized with heavy isotopes and spiked in after sample preparation.
  • the QconCAT approach is based on artificial proteins that are concatamers of proteotypic peptides. This artificial protein is recombinantly expressed in Escherichia coli and spiked into the sample before proteolysis.
  • QconCAT in principle allows efficient production of labeled peptides but does not automatically correct for protein fractionation effects or digestion efficiency in the native proteins versus the concatamers.
  • the PSAQ, absolute SILAC and FlexiQuant approaches sidestep these limitations by metabolically labeling full length proteins by heavy versions of the amino acids arginine and lysine.
  • the protein standard is added at an early stage, such as directly to cell lysate. Consequently, sample fractionation can be performed in parallel and the SILAC protein is digested together with the proteome under investigation.
  • Another quantitative approach applicable for the purpose of the present invention may be in some embodiments the SILAC-PrEST assay.
  • Protein Epitope Signature Tags are expressed recombinantly in E. coli and they consist of a short and unique region of the protein of interest as well as purification and solubility tags.
  • a highly purified, stable isotope labeling of amino acids in cell culture (SILAC)-labeled version of the solubility tag is first quantified and used to determine the precise amount of each PrEST by its SILAC ratios.
  • the PrESTs are then spiked into the examined sample (e.g., cell lysates) and the SILAC ratios of PrEST peptides to peptides from endogenous target proteins yield their cellular quantities.
  • the labeled or tagged biomarker/s of the invention or any labeled fragments or peptides thereof are mixed with the sample of with any protein extracted therefrom.
  • the resulting protein mixture may be then digested according to the FASP protocol (Wisniewski et al., 2009b) and the peptides are separated into fractions by anion exchange chromatography in a StageTip format (Wisniewski al., 2009a). Each fraction is analyzed by online reverse-phase chromatography coupled to high resolution, quantitative mass spectrometry analysis.
  • Mass analyzers with high mass accuracy, high sensitivity and high resolution include, but are not limited to, matrix-assisted laser desorption time-of-flight (MALDI-TOF) mass spectrometers, electrospray ionization time-of-flight (ESI-TOF) mass spectrometers, Fourier transform ion cyclotron mass analyzers (FT-ICR-MS), and Orbitrap analyzer instruments.
  • MALDI-TOF matrix-assisted laser desorption time-of-flight
  • EI-TOF electrospray ionization time-of-flight
  • FT-ICR-MS Fourier transform ion cyclotron mass analyzers
  • Orbitrap analyzer instruments include ion trap and triple quadrupole mass spectrometers.
  • ion trap MS In ion trap MS, analytes are ionized by electrospray ionization or MALDI and then put into an ion trap. Trapped ions can then be separately analyzed by MS upon selective release from the ion trap. Ion traps can also be combined with the other types of mass spectrometers described above.
  • Reference biomarker protein/s labeled with an ICAT or VICAT or iTRAQ type reagent, or SILAC labeled peptides can be analyzed, for example, by single stage mass spectrometry with a MALDI or ESI ionization and with TOF, quadrupole, iontrap, FT-ICR or Orbitrap analyzers.. Methods of mass spectrometry analysis are well known to those skilled in the art. For high resolution peptide fragment separation, liquid chromatography ESI- MS/MS or automated LC-MS/MS, can be used. MS analysis can be performed in a data-dependent manner or using targeted MS techniques such as selected reaction monitoring (SRM) or parallel reaction monitoring (PRM).
  • SRM selected reaction monitoring
  • PRM parallel reaction monitoring
  • the detecting molecules used are at least one of antibodies, nucleic acid, peptide aptamers or any combination thereof, specific for said at least one of said biomarker proteins
  • the determination of the expression level of said biomarker protein/s may be performed by an immunological assay.
  • ELISA Enzyme-Linked Immunosorbent Assay
  • a sample containing a protein substrate e.g., fixed cells or a protein solution
  • a substrate-specific antibody coupled to an enzyme is applied and allowed to bind to the substrate. Presence of the antibody is then detected and quantitated by a colorimetric reaction employing the enzyme coupled to the antibody.
  • Enzymes commonly employed in this method include horseradish peroxidase and alkaline phosphatase. If well calibrated and within the linear range of response, the amount of substrate present in the sample is proportional to the amount of color produced.
  • a substrate standard is generally employed to improve quantitative accuracy.
  • determination of the expression level of the biomarker may be performed using Western blot.
  • Western Blot as used herein involves separation of a substrate from other protein by means of an acryl amide gel followed by transfer of the substrate to a membrane (e.g., nitrocellulose, nylon, or PVDF). Presence of the substrate is then detected by antibodies specific to the substrate, which are in turn detected by antibody -binding reagents.
  • Antibody - binding reagents may be, for example, protein A or secondary antibodies.
  • Antibody -binding reagents may be radio labeled or enzyme-linked, as described hereinafter. Detection may be by autoradiography, colorimetric reaction, or chemiluminescence. This method allows both quantization of an amount of substrate and determination of its identity by a relative position on the membrane indicative of the protein's migration distance in the acryl amide gel during electrophoresis, resulting from the size and other characteristics of the protein.
  • Radioimmunoassay involves precipitation of the desired protein (i.e., the substrate) with a specific antibody and radio labeled antibody -binding protein (e.g., protein A labeled with I 125 ) immobilized on a perceptible carrier such as agars beads.
  • the radio-signal detected in the precipitated pellet is proportional to the amount of substrate bound.
  • a labeled substrate and an unlabelled antibody-binding protein are employed.
  • a sample containing an unknown amount of substrate is added in varying amounts.
  • the number of radio counts from the labeled substrate-bound precipitated pellet is proportional to the amount of substrate in the added sample.
  • determination of the expression level of the biomarker may be performed using FACS.
  • Fluorescence- Activated Cell Sorting involves detection of a substrate in situ in cells bound by substrate-specific, fluorescently labeled antibodies.
  • the substrate- specific antibodies are linked to fluorophore.
  • Detection is by means of a flow cytometry machine, which reads the wavelength of light emitted from each cell as it passes through a light beam. This method may employ two or more antibodies simultaneously, and is a reliable and reproducible procedure used by the present invention.
  • determination of the expression level of the biomarker may be performed using immunohistochemistry methods.
  • Immuno histochemical Analysis involves detection of a substrate in situ in fixed cells by substrate-specific antibodies.
  • the substrate specific antibodies may be enzyme-linked or linked to fluorophore. Detection is by microscopy, and is either subjective or by automatic evaluation. With enzyme-linked antibodies, a calorimetric reaction may be required. It will be appreciated that immunohistochemistry is often followed by counterstaining of the cell nuclei, using, for example, Hematoxyline or Giemsa stain.
  • isolated molecules when used in reference to a protein means that a naturally occurring sequence has been removed from its normal cellular environment or is synthesized in a non-natural environment (e.g., artificially synthesized). Thus, an "isolated” or “purified” sequence may be in a cell-free solution or placed in a different cellular environment.
  • purified does not imply that the sequence is the only nucleotide present, but that it is essentially free (about 90- 95% pure) of non-nucleotide material naturally associated with it, and thus is distinguished from isolated chromosomes.
  • isolated and purified in the context of a proteineous agent (e.g., a peptide, polypeptide, protein or antibody) refer to a proteineous agent which is substantially free of cellular material and in some embodiments, substantially free of heterologous proteineous agents (i.e. contaminating proteins) from the cell or tissue source from which it is derived, or substantially free of chemical precursors or other chemicals when chemically synthesized.
  • substantially free of cellular material includes preparations of a proteineous agent in which the proteineous agent is separated from cellular components of the cells from which it is isolated and/or recombinantly and/or synthetically produced.
  • a proteineous agent that is substantially free of cellular material includes preparations of a proteineous agent having less than about 30%, 20%, 10%, or 5% (by dry weight) of heterologous proteineous agent (e.g. protein, polypeptide, peptide, or antibody; also referred to as a "contaminating protein").
  • heterologous proteineous agent e.g. protein, polypeptide, peptide, or antibody; also referred to as a "contaminating protein”
  • the proteineous agent is recombinantly produced, it is also preferably substantially free of culture medium, i.e.
  • culture medium represents less than about 20%, 10%, or 5% of the volume of the protein preparation.
  • the proteinaceous agent is produced by chemical synthesis, it is preferably substantially free of chemical precursors or other chemicals, i.e., it is separated from chemical precursors or other chemicals which are involved in the synthesis of the proteinaceous agent. Accordingly, such preparations of a proteinaceous agent have less than about 30%, 20%, 10%, 5% (by dry weight) of chemical precursors or compounds other than the proteinaceous agent of interest.
  • proteinaceous agents disclosed herein are isolated.
  • nucleic acid detecting molecule may be used.
  • the nucleic acid detecting molecule/s of the invention may comprise at least one of: (a) nucleic acid aptamers specific for said at least one of said biomarker proteins; and (b) at least one isolated oligonucleotides, each oligonucleotide specifically hybridizes to a nucleic acid sequence encoding said at least one biomarker protein.
  • nucleic acid detecting molecules may comprise a nucleic acid aptamers specific for said at least one of the biomarker protein/s of the invention.
  • the nucleic acid detecting molecules may comprise at least one isolated oligonucleotide/s, each oligonucleotide specifically hybridizes to a nucleic acid sequence encoding one of said at least one biomarker protein.
  • the method of the invention may use nucleic acid detecting molecules specific for a nucleic acid sequence encoding the control reference protein/s.
  • nucleic acid molecules or “nucleic acid sequence” are interchangeable with the term “polynucleotide(s)” and it generally refers to any polyribonucleotide or poly- deoxyribonucleotide, which may be unmodified RNA or DNA or modified RNA or DNA or any combination thereof.
  • Nucleic acids include, without limitation, single- and double- stranded nucleic acids.
  • nucleic acid(s) also includes DNAs or RNAs as described above that contain one or more modified bases. Thus, DNAs or RNAs with backbones modified for stability or for other reasons are “nucleic acids”.
  • nucleic acids as it is used herein embraces such chemically, enzymatically or metabolically modified forms of nucleic acids, as well as the chemical forms of DNA and RNA characteristic of viruses and cells, including for example, simple and complex cells.
  • a "nucleic acid” or “nucleic acid sequence” may also include regions of single- or double- stranded RNA or DNA or any combinations.
  • oligonucleotide is defined as a molecule comprised of two or more deoxyribonucleotides and/or ribonucleotides, and preferably more than three. Its exact size will depend upon many factors which in turn, depend upon the ultimate function and use of the oligonucleotide.
  • the oligonucleotides may be from about 3 to about 1,000 nucleotides long.
  • oligonucleotides of 5 to 100 nucleotides are useful in the invention, preferred oligonucleotides range from about 5 to about 15 bases in length, from about 5 to about 20 bases in length, from about 5 to about 25 bases in length, from about 5 to about 30 bases in length, from about 5 to about 40 bases in length or from about 5 to about 50 bases in length. More specifically, the detecting oligonucleotides molecule used by the composition of the invention may comprise any one of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 bases in length.
  • oligonucleotide refers to a single stranded or double stranded oligomer or polymer of ribonucleic acid (RNA) or deoxyribonucleic acid (DNA) or mimetics thereof.
  • RNA ribonucleic acid
  • DNA deoxyribonucleic acid
  • oligonucleotides composed of naturally-occurring bases, sugars and covalent internucleoside linkages (e.g., backbone) as well as oligonucleotides having non-naturally-occurring portions which function similarly.
  • optional detecting molecule/s may be at least one nucleic acid aptamer specific for the at least one of said biomarker proteins.
  • aptamer or “specific aptamers” denotes single-stranded nucleic acid (DNA or RNA) molecules which specifically recognizes and binds to a target molecule.
  • the aptamers according to the invention may fold into a defined tertiary structure and can bind a specific target molecule with high specificities and affinities. Aptamers are usually obtained by selection from a large random sequence library, using methods well known in the art, such as SELEX and/or Molinex.
  • aptamers may include single- stranded, partially single- stranded, partially double-stranded or double-stranded nucleic acid sequences; sequences comprising nucleotides, ribonucleotides, deoxyribonucleotides, nucleotide analogs, modified nucleotides and nucleotides comprising backbone modifications, branch points and non-nucleotide residues, groups or bridges; synthetic RNA, DNA and chimeric nucleotides, hybrids, duplexes, heteroduplexes; and any ribonucleotide, deoxyribonucleotide or chimeric counterpart thereof and/or corresponding complementary sequence.
  • aptamers used by the invention are composed of deoxyribonucleotides.
  • the recognition between the aptamer and the antigen is specific and may be detected by the appearance of a detectable signal by using a colorimetric sensor or a fluorimetric/lumination sensor.
  • the aptamers as used according to some aspects of the invention may be biotinylated.
  • the aptamers may optionally include a chemically reactive group at the 3 and/or 5 termini.
  • the term reactive group is used herein to denote any functional group comprising a group of atoms which is found in a molecule and is involved in chemical reactions.
  • Some non-limiting examples for a reactive group include primary amines (NH 2 ), thiol (SH), carboxy group (COOH), phosphates (P04), Tosyl, and a photo-reactive group.
  • the aptamer as used herein may optionally comprise a spacer between the nucleic acid sequence and the reactive group.
  • the spacer may be an alkyl chain such as (CH 2 ) 6/12, namely comprising six to twelve carbon atoms.
  • the detection molecule may be at least one primer, at least one pair of primers, nucleotide probes and any combinations thereof.
  • compositions and kits of the invention may comprise, as an oligonucleo tide-based detection molecule, both primers and probes.
  • primer refers to an oligonucleotide, whether occurring naturally as in a purified restriction digest, or produced synthetically, which is capable of acting as a point of initiation of synthesis when placed under conditions in which synthesis of a primer extension product, which is complementary to a nucleic acid strand, is induced, i.e., in the presence of nucleotides and an inducing agent such as a DNA polymerase and at a suitable temperature and pH.
  • the primer may be single- stranded or double- stranded and must be sufficiently long to prime the synthesis of the desired extension product in the presence of the inducing agent.
  • the exact length of the primer will depend upon many factors, including temperature, source of primer and the method used.
  • the oligonucleotide primer typically contains 10-30 or more nucleotides, although it may contain fewer nucleotides. More specifically, the primer used by the methods, as well as the compositions and kits of the invention may comprise 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides or more.
  • such primers may comprise 30, 40, 50, 60, 70, 80, 90, 100 nucleotides or more.
  • the primers used by the method of the invention may have a stem and loop structure. The factors involved in determining the appropriate length of primer are known to one of ordinary skill in the art and information regarding them is readily available.
  • probe means oligonucleotides and analogs thereof and refers to a range of chemical species that recognize polynucleotide target sequences through hydrogen bonding interactions with the nucleotide bases of the target sequences.
  • the probe or the target sequences may be single- or double-stranded RNA or single- or double- stranded DNA or a combination of DNA and RNA bases.
  • a probe may be 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 and up to 30 nucleotides in length as long as it is less than the full length of the target mRNA or any gene encoding said mRNA.
  • Probes can include oligonucleotides modified so as to have a tag which is detectable by fluorescence, chemiluminescence and the like.
  • the probe can also be modified so as to have both a detectable tag and a quencher molecule, for example TaqMan(R) and Molecular Beacon(R) probes.
  • RNA or DNA may be RNA or DNA, or analogs of RNA or DNA, commonly referred to as antisense oligomers or antisense oligonucleotides.
  • RNA or DNA analogs comprise, but are not limited to, 2-'0-alkyl sugar modifications, methylphosphonate, phosphorothiate, phosphorodithioate, formacetal, 3-thioformacetal, sulfone, sulfamate, and nitroxide backbone modifications, and analogs, for example, LNA analogs, wherein the base moieties have been modified.
  • analogs of oligomers may be polymers in which the sugar moiety has been modified or replaced by another suitable moiety, resulting in polymers which include, but are not limited to, morpholino analogs and peptide nucleic acid (PNA) analogs.
  • Probes may also be mixtures of any of the oligonucleotide analog types together or in combination with native DNA or RNA.
  • the oligonucleotides and analogs thereof may be used alone or in combination with one or more additional oligonucleotides or analogs thereof.
  • the expression level may be determined using amplification assay.
  • amplification assay refers to methods that increase the representation of a population of nucleic acid sequences in a sample. Nucleic acid amplification methods, such as PCR, isothermal methods, rolling circle methods, etc., are well known to the skilled artisan. More specifically, as used herein, the term “amplified”, when applied to a nucleic acid sequence, refers to a process whereby one or more copies of a particular nucleic acid sequence is generated from a template nucleic acid, preferably by the method of polymerase chain reaction.
  • PCR Polymerase chain reaction
  • dNTPs each of the four deoxynucleotides dATP, dCTP, dGTP, and dTTP
  • primers primers
  • buffers DNA polymerase, and nucleic acid template.
  • the PCR reaction comprises providing a set of polynucleotide primers wherein a first primer contains a sequence complementary to a region in one strand of the nucleic acid template sequence and primes the synthesis of a complementary DNA strand, and a second primer contains a sequence complementary to a region in a second strand of the target nucleic acid sequence and primes the synthesis of a complementary DNA strand, and amplifying the nucleic acid template sequence employing a nucleic acid polymerase as a template-dependent polymerizing agent under conditions which are permissive for PCR cycling steps of (i) annealing of primers required for amplification to a target nucleic acid sequence contained within the template sequence, (ii) extending the primers wherein the nucleic acid polymerase synthesizes a primer extension product.
  • a set of polynucleotide primers "a set of PCR primers” or “pair of primers” can comprise two, three, four or more primers.
  • Real time nucleic acid amplification and detection methods are efficient for sequence identification and quantification of a target since no pre-hybridization amplification is required.
  • Amplification and hybridization are combined in a single step and can be performed in a fully automated, large- scale, closed-tube format.
  • hybridization-triggered fluorescent probes for real time PCR are based either on a quench-release fluorescence of a probe digested by DNA Polymerase (e.g., methods using TaqMan(R), MGB- TaqMan(R)), or on a hybridization- triggered fluorescence of intact probes (e.g., molecular beacons, and linear probes).
  • the probes are designed to hybridize to an internal region of a PCR product during annealing stage (also referred to as amplicon).
  • a "real time PCR” or “RT-PCT” assay provides dynamic fluorescence detection of amplified biomarker proteins of the invention or any control reference gene produced in a PCR amplification reaction.
  • the amplified products created using suitable primers hybridize to probe nucleic acids (TaqMan(R) probe, for example), which may be labeled according to some embodiments with both a reporter dye and a quencher dye.
  • the fluorescence of the reporter dye is suppressed.
  • a polymerase such as AmpliTaq GoldTM, having 5'-3' nuclease activity can be provided in the PCR reaction. This enzyme cleaves the fluorogenic probe if it is bound specifically to the target nucleic acid sequences between the priming sites.
  • the reporter dye and quencher dye are separated upon cleavage, permitting fluorescent detection of the reporter dye.
  • the fluorescent signal produced by the reporter dye is detected and/or quantified. The increase in fluorescence is a direct consequence of amplification of target nucleic acids during PCR.
  • QRT-PCR or "qPCR” which is quantitative in nature, can also be performed to provide a quantitative measure of gene expression levels.
  • QRT-PCR reverse transcription and PCR can be performed in two steps, or reverse transcription combined with PCR can be performed.
  • One of these techniques for which there are commercially available kits such as TaqMan(R) (Perkin Elmer, Foster City, CA), is performed with a transcript-specific antisense probe.
  • This probe is specific for the PCR product (e.g. a nucleic acid fragment derived from a gene) and is prepared with a quencher and fluorescent reporter probe attached to the 5' end of the oligonucleotide. Different fluorescent markers are attached to different reporters, allowing for measurement of at least two products in one reaction.
  • Taq DNA polymerase When Taq DNA polymerase is activated, it cleaves off the fluorescent reporters of the probe bound to the template by virtue of its 5-to-3' exonuclease activity. In the absence of the quenchers, the reporters now fluoresce. The color change in the reporters is proportional to the amount of each specific product and is measured by a fluorometer; therefore, the amount of each color is measured and the PCR product is quantified.
  • the PCR reactions can be performed in any solid support, for example, slides, microplates, 96 well plates, 384 well plates and the like so that samples derived from many individuals are processed and measured simultaneously.
  • the TaqMan(R) system has the additional advantage of not requiring gel electrophoresis and allows for quantification when used with a standard curve.
  • a second technique useful for detecting PCR products quantitatively without is to use an intercalating dye such as the commercially available QuantiTect SYBR Green PCR (Qiagen, Valencia California).
  • RT-PCR is performed using SYBR green as a fluorescent label which is incorporated into the PCR product during the PCR stage and produces fluorescence proportional to the amount of PCR product.
  • Both TaqMan(R) and QuantiTect SYBR systems can be used subsequent to reverse transcription of RNA.
  • Reverse transcription can either be performed in the same reaction mixture as the PCR step (one-step protocol) or reverse transcription can be performed first prior to amplification utilizing PCR (two-step protocol).
  • Molecular Beacons(R) which uses a probe having a fluorescent molecule and a quencher molecule, the probe capable of forming a hairpin structure such that when in the hairpin form, the fluorescence molecule is quenched, and when hybridized, the fluorescence increases giving a quantitative measurement of gene expression.
  • the detecting molecule may be in the form of probe corresponding and thereby hybridizing to any region or at least one of the biomarker protein or any control reference protein. More particularly, it is important to choose regions which will permit hybridization to the target nucleic acids. Factors such as the Tm of the oligonucleotide, the percent GC content, the degree of secondary structure and the length of nucleic acid are important factors. It should be further noted that a standard Northern blot assay can also be used to ascertain an RNA transcript size and the relative amounts of the biomarker proteins of the invention or any control gene product, in accordance with conventional Northern hybridization techniques known to those persons of ordinary skill in the art.
  • determining the level of expression of at least one or of at least five of the RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s of the invention may be performed by the step of subjecting a biological sample of the examined subject, or any protein product obtained therefrom to mass spectrometry analysis or assay.
  • the signature proteins may be also detected and quantified without the need for detection molecule/s.
  • Detection can be based on MS approaches using non-targeted or targeted methods such as selected reaction monitoring (SRM) or parallel reaction monitoring (PRM).
  • SRM selected reaction monitoring
  • PRM parallel reaction monitoring
  • SRM selected reaction monitoring
  • PRM parallel reaction monitoring
  • analyses can be performed with or without a reference heavy standard and provide quantitative measure of the peptide/protein amount.
  • the heavy reference can be a synthetic peptide, or a chemically labeled peptide/protein or metabolically labeled proteins.
  • the MS signal can provide the measure of peptide abundance.
  • the method of the invention may use as a sample any one of a biological sample of organ/s, cell/s or tissue/s and a blood sample.
  • a sample may be a primary tumor sample.
  • the methods of the invention may use a primary breast tumor sample.
  • sample refers to cells, sub-cellular compartments thereof, tissue or organs.
  • the tissue may be a whole tissue, or selected parts of a tissue. Tissue parts can be isolated by micro-dissection of a tissue, or by biopsy, or by enrichment of sub-cellular compartments.
  • sample further refers to healthy as well as diseased or pathologically changed cells or tissues.
  • the term further refers to a cell or a tissue associated with a disease, such a tumor, in particular carcinoma, breast cancer, and more specifically, Luminal A or B breast cancer.
  • a sample can be cells that are placed in or adapted to tissue culture.
  • a sample may also be a blood, body fluid such as plasma, lymph, urine, saliva, serum, cerebrospinal fluid, seminal plasma, pancreatic juice, breast milk, or lung lavage.
  • a sample can additionally be a cell or tissue from any species, including prokaryotic and eukaryotic species, specifically, humans.
  • a tissue sample can be further a fractionated or preselected sample, if desired, preselected or fractionated to contain or be enriched for particular cell types.
  • the sample can be fractionated or preselected by a number of known fractionation or pre selection techniques.
  • a sample can also be any extract of the above.
  • the term also encompasses protein fractions or alternatively, nucleic acid from cells or tissue.
  • the sample may be any one of a biological sample of organ/s, cell/s or tissue/s and a blood sample.
  • the sample may be a primary tumor sample.
  • the sample is obtained from a subject suffering from a luminal A or a luminal B breast tumor.
  • lymph node status of breast cancer patients is the most important variable in the management of the disease (Jatoi et al., 1999). Luminal tumors still confined to their original surroundings will mostly be treated with tamoxifen or aromatase inhibitors, while in the node-positive setting, chemotherapy will usually be applied and risk for recurrence rises substantially (Ellis and Perou, 2013). Currently, the lymph node status is mostly determined after the dissection of lymph nodes during surgery, but this may result in additional complications to the lymphatic system (Sakorafas et al., 2006).
  • the diagnostic and prognostic methods of the invention may provide an efficient tool for personalized treatment effective for specific subjects.
  • the methods of the inventions may be used for determining a treatment regimen for a subject suffering from a luminal A or a luminal B breast tumor. Accordingly, the method comprising in the first step, determining the expression level of at least one biomarker protein in at least one biological sample of a subject in need of such treatment, to obtain an expression value for each of said at least one biomarker protein.
  • the biomarker proteins may be at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten at least eleven, at least twelve, at least thirteen, at least fourteen or at least fifteen of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB or any combination thereof.
  • the biomarker proteins may be at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB or any combination thereof. In some further specific embodiments, the biomarker proteins may be at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s.
  • such five biomarker proteins may be LSM2, METAP2, RPS24, RBM12B and CAPS.
  • the at least five biomarker proteins may comprise RPS29, RBM3, PNP, METAP2 and RAB5C.
  • the biomarker proteins may be at least four of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s, specifically, RPS24, LSM4, RBM12B and RPS29.
  • the biomarker proteins of the invention may be at least six of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2, specifically, RPS24, LSM4, RBM12B, RPS29, RBM3 and PNP.
  • the biomarker proteins may be at least ten of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB. More specifically, RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3 and SNX12.
  • the second step of the method comprises determining if the expression value obtained in step (a) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein in at least one control sample.
  • the third step of the method involves providing an appropriate therapeutic regimen to a subject determined as exhibiting a negative expression value of said at least one, or alternatively, at least five, at least four, at least six or at least ten of the biomarker protein/s of the invention.
  • the therapy according with the present invention is any therapy applicable to cancer and specifically to breast cancer.
  • an endocrine therapy or any combination thereof with a biological therapy may be offered.
  • Endocrine therapy refers to a treatment that adds, blocks, or removes hormones.
  • endocrine therapy is provided to slow or stop the growth of breast cancers.
  • synthetic hormones or other drugs may be given to block the body's natural hormones.
  • therapy based on aromatase inhibitors may be offered.
  • Other therapeutic options may also include biological therapy (antibodies and the like) and cryotherapy.
  • chemotherapy, radiotherapy or any combinations thereof may be offered.
  • the method of the invention may be also applicable for evaluating or monitoring the responsiveness of a patient to treatment with any therapeutic agent or regimen. Accordingly, the patient may be evaluated in at least one time point after initiation of treatment in order to asses if the treatment protocol is efficient and appropriate. Determination can be carried out at an early time points such that a decision may be made regarding continuation of the treatment or alternatively readjusting the treatment protocol.
  • the present invention further provides the use of at least one of the biomarker proteins as markers for evaluating response of patients treated with a certain therapeutic agent or monitoring the efficacy of treatment with a certain therapeutic agent.
  • the method of the invention may be particularly suitable for monitoring and early diagnosis of response of the diagnosed disorder in the subject.
  • the invention provides a method for assessing responsiveness of a mammalian subject to treatment with a specific therapeutic agent or evaluating and/or monitoring the efficacy of treatment on a subject. This method is based on determining the expression values of the biomarkers of the invention before and any time after initiation of treatment, and calculating the ratio of the change in said values as a result of the treatment.
  • step (a) determining the expression level of at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen of the biomarker protein/s in a biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein, wherein said biomarker proteins are selected from RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB or any combination thereof.
  • the method of the invention may further encompass the use of at least one further additional detecting molecules, specifically, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more, specifically, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200,
  • the methods, compositions and kits of the invention may provide and use in addition to detecting molecules specific for at least one of the biomarkers disclosed in Table 4, also detecting molecule/s specific for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 or 85 biomarker proteins disclosed in Table 2, and optionally, further detecting molecule/s specific for additional at least one biomarker protein/s, for example, any of the biomarkers presented in Table 3, or
  • step (b) repeating step (a) in at least one other biological sample of said subject obtained after initiation of said treatment.
  • the third step (c) involves calculating the rate of change of the expression value of the biomarker proteins between said temporally separated samples, for example, samples obtained before and after initiation of said treatment.
  • step (d) concerns determining if the rate of change determined between at least two temporally separated samples or to the rate of change calculated for expression values in at least one control sample obtained from at least two temporally separated samples, wherein at least one sample of said at least two samples is obtained after the initiation of said treatment.
  • a negative rate of change of the expression value of at least one of said biomarker protein/s indicates that said subject exhibits a beneficial response to said treatment.
  • a positive rate of change is calculated for a subject, that means that the expression of the biomarker proteins of the invention is elevated in response to treatment and the subject may be thus classified as a non-responder to the particular treatment. Therefore, the invention provides a tool for monitoring the efficacy of a treatment with a therapeutic agent and the disease progression.
  • the methods of the invention involve in step (a) determination of the expression level of at least five biomarker proteins in at least one biological sample of the examined subject, to obtain an expression value for each of the at least five biomarker proteins.
  • at least five biomarker proteins may be selected from RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins. It should be understood that the at least five, at least four, at least six and at least tern biomarker proteins of the invention discussed herein before are also applicable for this method of the invention.
  • At least two "temporally- separated” test samples in order to assess the patient condition, or monitor the disease progression, as well as responsiveness to a certain treatment, at least two "temporally- separated” test samples must be collected from the examined patient and compared thereafter in order to obtain the rate of change in the expression value of at least one of the biomarker proteins between said samples.
  • at least two "temporally-separated" test samples and preferably more must be collected from the patient.
  • the expression value is then determined using the method of the invention, applied for each sample.
  • the rate of change in parameters is calculated by determining the ratio between at least two values of expression obtained from the same patient in different time-points or time intervals.
  • This period of time also referred to as "time interval", or the difference between time points (wherein each time point is the time when a specific sample was collected) may be any period deemed appropriate by medical staff and modified as needed according to the specific requirements of the patient and the clinical state he or she may be in.
  • this interval may be at least one day, at least three days, at least three days, at least one week, at least two weeks, at least three weeks, at least one month, at least two months, at least three months, at least four months, at least five months, at least one year, or even more.
  • one of the time points may correspond to a period in which a patient is experiencing a remission of the disease.
  • the rate of change When calculating the rate of change, one may use any two samples collected at different time points from the patient. To ensure more reliable results and reduce statistical deviations to a minimum, averaging the calculated rates of several sample pairs is preferable. A calculated or average value of a negative rate of change of the expression value of at least one of said biomarker protein/s indicates that said subject exhibits a beneficial response to said treatment; thereby monitoring the efficacy of a treatment with a therapeutic agent and the disease progression. It should be noted that in certain embodiments, where normalization step is being performed, the values referred to above, are normalized values.
  • the invention provides diagnostic and prognostic methods.
  • "Prognosis” is defined as a forecast of the future course of a disease or disorder, based on medical knowledge. This highlights the major advantage of the invention, namely, the ability to predict progression of the disease, based on the expression value of at least one of the biomarker proteins. More specifically, the ability to determine at early stage that the subject is suffering from a metastatic breast cancer, specifically, if a subject is classified as an LNN or alternatively as an LNP patient. This ability facilitates the selection of appropriate treatment regimen/s that may minimize side effects from unnecessary treatment, individually to each patient, as part of personalized medicine.
  • the prognostic method may be effective for predicting, monitoring and early diagnosing molecular alterations indicating response to treatment in said patient.
  • the prognostic method may be applicable for early, sub- symptomatic diagnosis of relapse when used for analysis of more than a single sample along the time-course of diagnosis, treatment and follow-up.
  • An “early diagnosis” provides diagnosis prior to appearance of clinical symptoms.
  • Prior as used herein is meant days, weeks, months or even years before the appearance of such symptoms. More specifically, at least 1 week, at least 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, or even few years before clinical symptoms appear.
  • the number of samples collected and used for evaluation of the subject may change according to the frequency with which they are collected.
  • the samples may be collected at least every day, every two days, every four days, every week, every two weeks, every three weeks, every month, every two months, every three months every four months, every 5 months, every 6 months, every 7 months, every 8 months, every 9 months, every 10 months, every 11 months, every year or even more.
  • the rate of change may be calculated as an average rate of change over at least three samples taken in different time points, or the rate may be calculated for every two samples collected at adjacent time points.
  • the sample may be obtained from the monitored patient in the indicated time intervals for a period of several months or several years. More specifically, for a period of 1 year, for a period of 2 years, for a period of 3 years, for a period of 4 years, for a period of 5 years, for a period of 6 years, for a period of 7 years, for a period of 8 years, for a period of 9 years, for a period of 10 years, for a period of 11 years, for a period of 12 years, for a period of 13 years, for a period of 14 years, for a period of 15 years or more.
  • the samples are taken from the monitored subject every two months for a period of 5 years.
  • the method for monitoring disease progression or early prognosis for disease relapse as detailed herein may be used for personalized medicine, by collecting at least two samples from the same patient at different stages of the disease.
  • the prediction obtained by the method of the invention made by comparing between the sample and the patient population may be dependent on the selection of population of patients to which the sample is compared to.
  • patient or “subject” it is meant any mammal that may be affected by the above-mentioned conditions, and to whom the treatment and diagnosis methods herein described is desired, including human, bovine, equine, canine, murine and feline subjects. Specifically, said patient is a human.
  • determining the level of expression of at least one or of at least five of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins may be performed by the step of contacting at least one detecting molecule or any combination or mixture of plurality of detecting molecules with a biological sample of the examined subject, or with any protein or nucleic acid product obtained therefrom. It should be noted that each of the detecting molecules is specific for one of the biomarker protein/s.
  • determining the level of expression of at least one or of at least five of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may be performed by the step of subjecting a biological sample of the examined subject, or any protein product obtained therefrom to mass spectrometry analysis or assay.
  • the determination of the expression level of the proteins can be achieved by quantification methods excluding the need of detection molecules. Label-free quantification of proteins can be conducted by liquid chromatography-mass spectrometry (LC-MS) with electrospray ionization.
  • LC-MS liquid chromatography-mass spectrometry
  • This method provides differential expression measurements and enables the discovery of biological markers.
  • Other methods for label-free quantification also can be used. Non-limiting examples of these methods include SRM or PRM. These analyses may be performed with or without a reference heavy standard and provide quantitative measure of the peptide/protein amount.
  • the invention relates to a diagnostic and/or prognostic composition
  • a diagnostic and/or prognostic composition comprising at least one detecting molecule or any combination or mixture of plurality of detecting molecules specific for determining the level of expression of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof.
  • each of said detecting molecules is specific for one of said biomarker proteins.
  • the composition of the invention may be at least one of diagnostic and prognostic composition.
  • the detecting molecules comprised within the composition of the invention may be attached to a solid support.
  • solid support that may be used as part of the diagnostic composition of the invention are described in more detail herein after, in connection with the kit of the invention. It should be appreciated that in some specific and non-limiting embodiments, the detecting molecules of the composition of the invention may be provided in a suitable medium or a buffer. In some alternative embodiments, the detecting molecules of the invention may be provided in a dried form.
  • compositions comprising detecting molecules specific for any combination of any of the marker protein used by the invention.
  • the composition of the invention may comprise at least one detecting molecule or any combination or mixture of plurality of detecting molecules specific for determining the level of expression of at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof. It should be noted that each of the detecting molecules is specific for one of said biomarker proteins.
  • such at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise LSM2, METAP2, RPS24, RBM12B and CAPS.
  • the at least five biomarker proteins may comprise RPS29, RBM3, PNP, METAP2 and RAB5C.
  • the five biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten of the LSM4, RPS29, RBM3, PNP, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4 and CSTB biomarker proteins of the invention.
  • the composition of the invention may comprise at least one detecting molecule specific for at least four of the biomarker proteins of the invention.
  • such at least four of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise RPS24, LSM4, RBM12B and RPS29.
  • the four biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten or at least eleven of the RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention.
  • the composition of the invention may comprise detecting molecules specific for at least six of the biomarker proteins of the invention.
  • such at least six of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise RPS24, LSM4, RBM12B, RPS29, RBM3 and PNP.
  • the six biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight or at least nine of the METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention. Still further embodiments relate to the compositions of the invention that may comprise detecting molecules specific for at least ten of the biomarker proteins of the invention, specifically, RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3 and SNX12.
  • the ten biomarker proteins may further comprise at least one, at least two, at least three, at least four, or at least five of the SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention.
  • composition of the invention may comprise detecting molecules specific for all fifteen biomarker proteins of the invention, specifically RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB. It should be appreciated that any of the combinations of at least five, at least four, at least six and at least ten of the biomarker proteins of the invention disclosed herein, are also applicable for any of the kits of the invention discussed herein after.
  • compositions of the invention may further comprise detecting molecules specific for control reference protein.
  • control reference protein may be used for normalizing the detected expression levels for the biomarker proteins used by the invention.
  • Non-limiting embodiments for control reference proteins may include ARCN1 (Archain 1), MPZL1 (Myelin Protein Zero-Like 1), NSF (N-ethylmaleimide-sensitive factor), PRKCD (Protein Kinase C, Delta), CAT (catalase), actin, tubulin, or other cytoskeletal proteins.
  • composition of the invention may comprise at least one detecting molecules specific for at least one biomarker of the invention, specifically, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 of the biomarkers of Table 4, specifically, RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB.
  • the composition of the invention may comprise detecting molecules specific for at least one further additional biomarker.
  • compositions of the invention may comprise also detecting molecule/s specific for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more, specifically, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, 300
  • the detecting molecules used by the methods, compositions and kits of the invention may be specific for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 or 85 biomarker proteins disclosed in Table 2, and optionally, further detecting molecule/s specific for additional at least one biomarker protein/s, for example, any of the biomarkers presented in Table 3, or any other biomarker/s.
  • the detection molecules of the invention may be in the form of isolated detecting amino acid molecules and isolated detecting nucleic acid molecules.
  • the composition of the invention may comprise amino acid detecting molecules. More specifically, such molecules may be at least one of: (a) at least one isolated recombinant labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof; (b) antibodies specific for said at least one of said biomarker protein/s; (c) peptide aptamers specific for said at least one biomarker protein/s; or (d) any combination of (a), (b) and (c).
  • the composition of the invention may comprise nucleic acid detecting molecules.
  • detecting molecules may include at least one of: (a) nucleic acid aptamers specific for said at least one biomarker proteins; (b) at least one isolated oligonucleotides, each oligonucleotide specifically hybridizes to a nucleic acid sequence encoding said at least one biomarker protein.
  • the detecting molecules of the composition/s of the invention may be at least one labeled or tagged optionally isolated and/or recombinant RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof.
  • the determination of the expression level of said at least one biomarker protein/s may be performed by mass spectrometry.
  • the detecting molecules may be at least one of antibodies, nucleic acid or peptide aptamers specific for said at least one of the at least one biomarker proteins, or any combination thereof.
  • the determination of the expression level of the at least one biomarker protein/s may be performed by an immunological assay.
  • the detecting molecules comprised within the composition of the invention may be attached to a solid support. More specifically, as defined herein, the detecting molecules are optionally attached to a support where each of the detecting molecules is attached to a support in a unique pre- selected and defined region. In some other embodiments, the detecting molecules may be provided in non-immobilized form, specifically, not attached to a solid support but separated in different vessels, tubes, wells and the like. Nevertheless, in yet some alternative embodiments, the detecting molecules may be provided in a mixture that contains variety of detecting molecules specific for at least one and at most 500 of the biomarker proteins of the invention.
  • composition of the invention may further comprise a biological sample. It should be appreciated that any of the biological samples described for the method of the invention are also applicable for the composition of the invention.
  • the invention may further comprise a composition comprising at least one of the detecting molecules specific for at least one biomarker protein/s of the invention, specifically, the biomarkers of Table 4, and a sample, specifically, a biological sample.
  • the composition of the invention may comprise detecting molecules specific for at least one further biomarker, provided that the detecting molecules of the compositions of the invention are specific for 500 biomarkers at the most.
  • such further biomarkers may be selected from the proteins listed in any one of Tables 2 and 3.
  • compositions of the invention may comprise detecting molecules specific for at least one additional biomarker protein, specifically, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more, specifically, 110, 120, 130, 140, 150, 160, 1
  • compositions of the invention may be used for predicting breast cancer progression, assessing the patient's condition and may be also used for monitoring responsiveness of a mammalian subject to treatment.
  • a third aspect of the invention relates to a kit comprising (a) detecting molecules specific for determining the level of expression of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof in a biological sample. It should be noted that each of said detecting molecules is specific for one of said biomarker proteins.
  • the kit of the invention may optionally further comprise at least one of: (b) pre-determined calibration curve/s or predetermined standard/s providing standard expression values of said at least one biomarker/s; and (c) at least one control sample.
  • the kit of the invention may comprise at least one detecting molecule specific for determining the level of expression of at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof in a biological sample.
  • the invention further encompass any kit comprising detecting molecules specific for at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen of the biomarker protein/s of the invention.
  • the kit of the invention may comprise detecting molecules specific for any combination of the biomarker proteins of the invention, specifically the combinations specified herein above in connection with the methods and compositions aspects. It should be appreciated that each of the detecting molecule/s is specific for one of said biomarker proteins.
  • the kit of the invention may optionally further comprises at least one of: pre-determined calibration curve/s or predetermined standard/s providing standard expression values of said at least one biomarker protein/s; and at least one control sample. It should be appreciated that all the combinations disclosed herein before in connection with the compositions of the invention are also applicable for any of the kits of the invention.
  • the detecting molecules comprised within the kit of the invention may be isolated detecting nucleic acid molecules, isolated detecting amino acid molecules or any combinations thereof.
  • kits of the invention may comprise amino acid detecting molecules, more specifically, at least one of: (a) at least one labeled or tagged, optionally, recombinant and/or isolated RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof; (b) antibodies specific for said at least one of said biomarker proteins; (c) peptide aptamers specific for said at least one of said biomarker protein/s; and (d) any combination of (a), (b) and (c).
  • the kit of the invention may comprise nucleic acid detecting molecule, for example, at least one of: (a) nucleic acid aptamers specific for said at least one biomarker proteins; (b) at least one isolated oligonucleotides, each oligonucleotide specifically hybridizes to a nucleic acid sequence encoding said at least one biomarker protein/s.
  • the detecting molecules comprised within the kit of the invention may be attached to a solid support.
  • detecting molecules of the invention were described in detailed in connection with the methods of the invention. It should be appreciated that all embodiments for detecting molecules mentioned therein are also applicable for the compositions and kits of the invention.
  • kit of the invention may further comprise instructions for use, wherein said instructions comprise at least one of:
  • the kit of the invention may further comprise at least one reagent for conducting a mass spectrometry assay.
  • reagents may include trypsin, buffers, filters and the like, for peptide purification.
  • the kit of the invention further comprising at least one reagent for conducting an immunological assay selected from protein microarray analysis, ELISA, RIA, slot blot, dot blot, FACS, western blot, immunohistochemical assay, immunofluorescent assay and a radio-imaging assay.
  • the kit of the invention may be used for predicting breast cancer, assessing the patient's condition and monitoring responsiveness of a mammalian subject to treatment.
  • the kit of the invention may be used in a method for determining the progression of breast cancer in a subject.
  • the subject is suffering from a luminal A or a luminal B breast tumor.
  • the kit of the invention may be applicable for early determination and diagnosis of tested subjects that has LNN or LNP.
  • the sample to be used is any one of a biopsy of organs or tissues and a blood sample.
  • the sample is a primary tumor sample.
  • the kits of the invention may use any appropriate biological sample.
  • biological sample in the present specification and claims is meant to include samples obtained from a mammalian subject.
  • the biological sample may be a bodily fluid, a tissue, a tissue biopsy, a skin swab, an isolated cell population or a cell preparation.
  • the population of cells comprises cancer cells. In another embodiment the population of cells is an in vitro cultured cell population.
  • the biological sample may be a bodily fluid selected from the group consisting of blood, serum, plasma, urine, cerebrospinal fluid, amniotic fluid, tear fluid, nasal wash, mucus, saliva, sputum, broncheoalveolar fluid, throat wash, vaginal fluid and semen.
  • the sample may be a tissue sample or blood sample which can be obtained using a syringe needle for example from a vein of the subject or from the tissue.
  • the cell may be isolated from the subject (e.g., for in vitro detection) or may optionally comprise a cell that has not been physically removed from the subject (e.g., in vivo detection).
  • the sample is any one of a biopsy of organ/s or tissue/s and a blood sample.
  • the sample is a primary tumor sample.
  • the sample may be lymph node tissue.
  • Primary Tumor may refer to the original, or first, tumor in the body. Cancer cells from a primary tumor may spread to other parts of the body and form new, or secondary, tumors (i.e. metastasis). In most cases, secondary tumors are the same type of cancer as the primary tumor. The term primary tumor may be interchanged with the term primary cancer. Still further, the inventors consider the kit of the invention in compartmental form.
  • the detecting molecules used for detecting the expression levels of the biomarker proteins may be provided in a kit attached to an array.
  • a "detecting molecule array” refers to a plurality of detection molecules that may be nucleic acids based or protein based detecting molecules, optionally attached to a support where each of the detecting molecules is attached to a support in a unique pre- selected and defined region.
  • the detecting molecules are attached to a solid support.
  • an array may contain different detecting molecules, such as specific antibodies, labeled or tagged proteins, peptides, aptamers, probes and/or primers.
  • the different detecting molecules for each target may be spatially arranged in a predetermined and separated location in an array.
  • an array may be a plurality of vessels (test tubes), plates, micro-wells in a micro-plate, each containing different detecting molecules, specifically, aptamers, primers and antibodies, specific for each marker protein used by the invention.
  • An array may also be any solid support holding in distinct regions (dots, lines, columns) different and known, predetermined detecting molecules.
  • solid support is defined as any surface to which molecules may be attached through either covalent or non-covalent bonds.
  • useful solid supports include solid and semisolid matrixes, such as aero gels and hydro gels, resins, beads, biochips (including thin film coated biochips), micro fluidic chip, a silicon chip, multi-well plates (also referred to as microtiter plates or microplates), membranes, filters, conducting and no conducting metals, glass (including microscope slides) and magnetic supports.
  • useful solid supports include silica gels, polymeric membranes, particles, derivative plastic films, glass beads, cotton, plastic beads, alumina gels, polysaccharides such as Sepharose, nylon, latex bead, magnetic bead, paramagnetic bead, super paramagnetic bead, starch and the like. This also includes, but is not limited to, microsphere particles such as Lumavidin.TM. Or LS-beads, magnetic beads, charged paper, Langmuir-Blodgett films, functionalized glass, germanium, silicon, PTFE, polystyrene, gallium arsenide, gold, and silver.
  • any other material known in the art that is capable of having functional groups such as amino, carboxyl, thiol or hydroxyl incorporated on its surface is also contemplated. This includes surfaces with any topology, including, but not limited to, spherical surfaces and grooved surfaces.
  • any of the reagents, substances or ingredients included in any of the methods and kits of the invention may be provided as reagents embedded, linked, connected, attached, placed or fused to any of the solid support materials described above.
  • the detecting molecules may be provided as molecules that are not attached to any solid support.
  • the non-attached detecting molecules may be provided in separate containers, wells, tube vessels and the like.
  • the attached or non-attached detecting molecules may be provided in a mixture that contains at least two detecting molecules specific for at least two biomarker protein/s of the invention.
  • the detecting molecules of the invention e.g., the recombinant (or synthetically produced) labeled or tagged biomarker protein/s or any fragment or peptide thereof may be provided as a mixture in a tube or any other vessel or container.
  • the invention provides a method for assessing expression status of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB in a biological sample.
  • the method comprising the step of: (a) providing at least one detecting molecule, any combination, mixture of plurality of detecting molecules or any composition of kit comprising the same, wherein each of said detecting molecules is specific for one of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB proteins.
  • step (b) contacting the at least one detecting molecule/s provided in (a) with the biological sample, or with any protein or nucleic acid product obtained therefrom and (c), performing a protein or nucleic acid detection assay to assess the expression status of said proteins.
  • the samples and detecting molecules described by the invention herein before are also applicable for this specific method.
  • Still further aspect of the invention relates to a diagnostic and/or prognostic method for determining the progression of breast cancer in a subject, the method comprising: (a) providing at least one detecting molecule specific for at least one of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins, and any additional biomarker proteins, for example any one of the biomarkers disclosed in Tables 2 and 3. It should be noted that in certain embodiments, the detecting molecules are specific for 500 biomarker/s at the most.
  • the method of the invention may be at least one of diagnostic and prognostic method.
  • the detecting molecules provided by the methods of the invention may be provided as an array, as a composition (specifically, any of the composition described herein above) or as a kit, as described herein before.
  • the method of the invention involves determining the expression level of at least one of the biomarker proteins of the invention in at least one biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein/s, using the detecting molecules provided in step (a).
  • step (c) determining if the expression value obtained in step (b) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample.
  • at least one of: (i) a positive expression value of said at least one biomarker protein/s in said sample, indicates that said subject belongs to a pre-established population associated with positive lymph node metastatic status; and (ii) a negative expression value of said at least one biomarker protein/s in said sample indicates that said subject belongs to a pre-established population associated with negative lymph node metastatic status.
  • composition or method may include additional ingredients and/or steps, and/or parts, but only if the additional ingredients and/or steps do not materially alter the basic and novel characteristics of the claimed composition or method.
  • word “comprise”, and variations such as “comprises” and “comprising” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. It should be noted that various embodiments of this invention may be presented in a range format.
  • range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub ranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed sub ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range. Whenever a numerical range is indicated herein, it is meant to include any cited numeral (fractional or integral) within the indicated range.
  • method refers to manners, means, techniques and procedures for accomplishing a given task including, but not limited to, those manners, means, techniques and procedures either known to, or readily developed from known manners, means, techniques and procedures by practitioners of the chemical, pharmacological, biological, biochemical and medical arts.
  • Formalin-fixed, paraffin-embedded (FFPE)-blocks were obtained from the department of pathology, Sheba Medical Center, Tel Hashomer, Israel. Included cases were ER-positive and Her-2 negative infiltrating ductal carcinoma of the breast.
  • FFPE blocks were sliced into twelve ⁇ -thick sections and mounted on histological slides. Slides were dried at 37°C overnight. Areas of high cellularity were marked on the slides to enrich for cancer cells (or healthy breast epithelia) and to avoid inclusion of connective, adipose or lymphatic tissue. ER staining was used as a guide for delineation of cancer cells. Proteins were denaturated, combined at a 1: 1 ratio with the super-SILAC mix and digested following the FFPE-FASP protocol (Ostasiewicz et al., 2010) was used. Peptides were fractionated by strong anion exchange (SAX) fractionation in a StageTip format.
  • SAX strong anion exchange
  • Super-SILAC mix was prepared previously by T. Geiger, as described (Geiger et al., 2010, and WO 2011/042467).
  • Four breast cancer cell lines that differ in origin, stage and receptor status HCC1599, MCF7, HCC1937 and HCC2218
  • HMEC normal mammary epithelial cells
  • Recombinant proteins are tagged (with HA, ABP, GST or other) to enable purification and absolute quantification. Proteins will purified with affinity chromatography, based on the specific tags. The absolute amount of each purified protein is determined by amino acid analysis and/or MS-based analysis relative to the purified tag. The heavy labeled recombinant proteins are combined with the unlabeled clinical samples to determine their absolute amounts. Sample preparation for MS analysis
  • proteins from the clinical samples were mixed with equal amounts of the super-SILAC mix and incubated with 0.1M dithiothreitol and 0.05M iodoacetamide.
  • the mixed lysates were digested with trypsin on top of 30kDa cutoff Amicon filters, as depicted in the Filter- Aided Sample Preparation (FASP) procedure (Wisniewski et al., 2009b).
  • FASP Filter- Aided Sample Preparation
  • the resulting peptides were fractionated using pH-based strong anion exchange (SAX) in StageTip format (Wisniewski et al., 2009a).
  • SAX pH-based strong anion exchange
  • Peptides were separated by nano-ultra high performance liquid chromatography (UHPLC) (EasynLClOOO, Thermo Fisher Scientific) coupled on-line to a Q-Exactive or Q-Exactive Plus mass spectrometers (Thermo Fisher Scientific) through the EASY-Spray ionization source. Peptides were loaded onto to a 50 cm EASY-Spray column with a buffer containing 0.1% formic acid (buffer A), and eluted with a buffer containing 80% acetonitrile and 0.1% formic acid (buffer B) using different gradients.
  • UHPLC nano-ultra high performance liquid chromatography
  • Buffer A buffer containing 0.1% formic acid
  • buffer B buffer containing 80% acetonitrile and 0.1% formic acid
  • Mass spectra were acquired in a data- dependent manner with the top- 10 precursor m/z values from each MS scan fragmented by higher energy collisional dissociation (HCD). MS-scans and MS/MS scans were performed with resolutions of 70,000 and 17,500, respectively.
  • MS raw files were analyzed by MaxQuant (version 1.5.0.36; Cox and Mann, 2008). MS/MS spectra were searched against the reference UNIPROT human proteome (published November 2014) by the Andromeda search engine (Cox et al., 2011). False Discovery Rate (FDR) of 0.01 was used in on both the peptide and protein levels and determined by a decoy database. Prior to bioinformatics analysis, the resulting protein list was filtered to eliminate common contaminants and decoy database hits.
  • FDR False Discovery Rate
  • Bioinformatic and statistical analyses were performed in the Perseus software and in MATLAB (version R2014a). For all analyses, we filtered the data to include only proteins quantified in >70% of the samples. Expression ratios towards internal standard were normalized by z-scoring and subtracting most frequent value on each sample, and missing data points were imputed by creating a normal distribution with a width of 0.3 and a downshift of 1.5. Correlation networks of proteins were constructed by generic k-means clustering of all protein pairs with a cutoff correlation of 0.3 for all samples and 0.5 for tumor samples. Additional network analyses were done using STRING database. All networks were visualized with Cytoscape. Welch's ttests for statistical significance were performed with permutation-based FDR correction threshold of 0.05.
  • TMAs activity assays of complex I and complex IV on breast cancer frozen tumor microarrays
  • TMAs BioChain institute, Inc. Newark, USA
  • complex IV activity assay TMAs were brought to room temperature, washed for 5 min with 25mM sodium phosphate buffer pH7.4, and then incubated for 90 min at 37°C with COX incubation mixture containing 1 mg/ml Cytochrome C (C7752, Sigma-Aldrich), 1 mg/ml 3,3'-diaminobenzidine tetrahydrochloride hydrate (D5637, Sigma-Aldrich) and 0.2 mg/ml catalase (C1345, Sigma-Aldrich) in 25 mM sodium phosphate buffer at pH7.2-7.4 (Whitaker- Menezes et al., 2011).
  • HMEC Human mammary epithelial cells
  • MCF7 Human mammary epithelial cells
  • HMEC Human mammary epithelial cells
  • MCF7 medium-heavy lysine and arginine
  • cell lysates were mixed with the same cells grown in light culture medium that serve as an internal standard. Lysates were digested overnight with trypsin in solution and were subjected to LC-MS/MS analysis and MaxQuant analysis as described above.
  • Breast cancer tumor microarrays were obtained from BioChain Institute, Inc. (Newark, USA), and stained with anti-ACOTl and anti-SLC25Al l (AbCam) or anti GLUL (Sigma- Aldrich/Prestige antibodies). Staining intensity of relevant cores (ER-positive, Her-2 negative invasive ductal carcinomas, and healthy ducts) was assessed by a pathologist on a 4-degree scale of 0 (no staining) to 3 (strong staining).
  • affiliated Lymph nodes with nodes score (>2, (>0.5,
  • FFPE formalin-fixed, paraffin-embedded
  • Samples were obtained from lumpectomies or mastectomies, from patients which have not received any treatment prior to surgery, to eliminate possible proteomic changes caused by the treatment.
  • the cohort included ER-positive, Her2-negative infiltrating ductal carcinomas (IDC, luminal), grade 2 or 3, based on immunohistochemical staining and pathologist review. Since for luminal tumors patient prognosis largely depends on the lymph node status, identification of the proteins that are altered in late stages can serve as prognostic markers, and understanding the processes that are altered upon cancer progression can lead to better understanding of the regulatory mechanisms of cancer invasion.
  • breast-cancer super-SILAC mix was used (Geiger et al., 2010) to serve as a common internal standard against which proteins from all samples are quantified in the MS analysis. It is a mixed lysate of four breast cancer cell lines and primary mammary epithelial cells, labeled metabolically with heavy lysine and arginine. As a mixture, it contains labeled counterparts of the vast majority of proteins expressed in the clinical samples at similar levels, and it serves as a platform for relative quantifications of these proteins across the entire cohort.
  • MaxQuant analysis identified overall 150,471 peptides and 10,124 proteins and quantified overall 10,043 of them ( Figure IB and in Figure 2A; false discovery rate of 1% on both peptide and protein levels).
  • the inventors found no overall differences in the total number of quantified proteins, between the groups of samples, with 9093, 9746 and 9450 proteins in the healthy tissue, tumor tissue (both LNN and LNP) and metastatic tissue, respectively, with an average of 5439 proteins identified and 4300 quantified in each tissue. From these, 1499 were identified in all 88 samples, and only 177 were found in less than 10 of the samples ( Figure 2B).
  • the combined dynamic range of protein expression encompassed eight orders of magnitude and the vast majority of proteins (97%) were expressed within four orders of magnitude (Figure 1C).
  • these proteins known luminal breast cancer-associated proteins such as GAT A3, FOXA1 and the estrogen receptor ESR were identified.
  • the inventors identified five histones, actin and tubulin, as well as ribosomal proteins.
  • the inventor found two clusters with high density (high intra-cluster protein associations) to be enriched for ribosomal proteins, translation and ribosome biogenesis, as well as oxidative phosphorylation, lysosomes and interferon signaling (clusters 6 and 7).
  • Other enriched and correlated processes include proteasome together with spliceosome (cluster 2), and glycolysis together with tRNA aminoacylation and focal adhesion (cluster 8).
  • the inventors constructed an additional network comprising of correlations between tumor samples only, with a correlation cutoff of 0.5 (Figure 5A).
  • the inventors found highest correlations within specific functions, mostly consisting of large protein complexes such as ribosomes (cluster 1), spliceosome (cluster 5), oxidative phosphorylation (cluster 7) and DNA replication (cluster 6). Beyond these, the inventors further found high correlations between distinct compartments, such as mitochondrial oxidative phosphorylation and the peroxisome (cluster 7), and distinct functions, such as translation and lysosomal degradation (cluster 5) or DNA replication and locomotion (cluster 6). These results highlight the potential of this proteomic resource to reveal fundamental cancer-related associations of cellular processes, which can serve as the basis for further functional research.
  • correlation matrix of primary tumors and healthy tissue showed major differences between the healthy tissue and tumors, and co-clustering of primary tumors and metastases, highlighting their similar protein expression patterns (Figure 4B).
  • the correlations between samples ranged from 0.06 to 0.87.
  • the median correlation between healthy and primary tumors was 0.38 (and 0.39 for matched samples), and the median correlation between primary and metastases was 0.58 (and 0.75 for matched samples, see also below).
  • the correlation within the tumor sample group both primary and metastases was significantly higher than the correlation between the healthy tissues (Figure 5B).
  • the inventors examined the differences between the healthy tissues and tumors.
  • the inventors found two known breast cancer markers, Mucin 1 (CA15-3) and Cathepsin D, to be higher in the breast cancer tissues compared to normal duct epithelia ( Figures 6C, 6D).
  • NMD nonsense-mediated mRNA decay
  • Lysosomal proteins were also significantly upregulated (Figure 4C) with the most prominent components belonging to the vacuolar-type proton ATPase, as well as several cathepsins (CTSA, CTSB, CTSD, and CTSZ). Taken together, these results suggest that protein homeostasis is impaired in tumor cells, and that tumors may over-produce improperly functioning proteins, which may interfere with proper cellular activities and thus facilitate tumorigenesis.
  • HMEC normal mammary epithelial cells
  • MCF7 ER-positive breast cancer cells
  • Oxidative phosphorylation proteins were significantly upregulated in the cancer samples, concurrently, key glycolytic enzymes such as HK2, GAPDH, ALDOA, LDHA and LDHB were downregulated (Figure 4C).
  • key glycolytic enzymes such as HK2, GAPDH, ALDOA, LDHA and LDHB were downregulated ( Figure 4C).
  • the increased activity of the electron transport chain in tumor cells was further validated by activity-based assays for mitochondrial complex I and complex IV, using breast cancer tumor (Duct carcinoma in situ) arrays ( Figures 6E, 6F).
  • Recon 1 pathways were shown to be down regulated in tumors: Alanine and Aspartate Metabolism, Arginine and Proline Metabolism, Ascorbate and Aldarate Metabolism, Cholesterol Metabolism, Fatty acid activation, Glycine, Serine, and Threonine Metabolism, Glycolysis/Gluconeogenesis, Glyoxylate and Dicarboxylate Metabolism, Histidine Metabolism, IMP Biosynthesis, Lysine Metabolism, Methionine Metabolism, Pentose Phosphate Pathway, Propanoate Metabolism, Selenoamino acid metabolism, Transport, Extracellular.
  • Recon 1 pathways were shown to be up regulated in tumors: Chondroitin sulfate degradation, Fatty Acid Metabolism, Galactose metabolism, Glutathione Metabolism, Heme Biosynthesis, Heme Degradation, Heparan sulfate degradation, Hyaluronan Metabolism, Keratan sulfate degradation, N-Glycan Degradation, Nucleotides, Oxidative Phosphorylation, Pentose and Glucuronate Interconversions, Pyrimidine Catabolism, ROS Detoxification, Sphingolipid Metabolism, Tetrahydrobiopterin, Transport, Lysosomal.
  • GLUL glutamine synthase
  • SLC1A5 the major importer for glutamine, SLC1A5, was significantly downregulated, as well as the bidirectional transporter SLC7A5/SLC3A2, which has been shown to control outward efflux of glutamine in exchange for essential amino acids.
  • ROS reactive-oxygen species
  • the inventors selected three metabolic enzymes that were higher in the cancer samples compared to the healthy tissue, and represent key regulated pathways: (i) GLUL, a key regulator of Gin production; (ii) SLC25A11, a member of the malate-aspartate shuttle; (iii) Acyl-CoA thioesterase (ACOT1), which operates in the ⁇ -oxidation of fatty acids and catalyzes the hydrolysis of acyl-CoA to coenzyme A and free fatty acid.
  • the inventors validated their overexpression in tumors using immunohistochemistry on commercial tumor microarrays.
  • FIGS 9B, 9C Ten of the downregulated proteins and twelve of the upregulated proteins showed a pattern of staining that matches the proteomic findings ( Figures 9D, 9E, respectively). Presumably, some of the discrepancies result from the semiquantitative nature of IHC and potential non-specific binding of some of the antibodies. Importantly, only one of the antibodies (against POSTN) showed strong extracellular staining and no staining of epithelial cells, and six additional antibodies showed mild involvement of extracellular staining. These results indicate that macrodis section of the tissues allowed capturing of the cancer-related proteome, with minor effects of the extracellular proteins.
  • the inventors next examined the metastatic tissue from lymph nodes in comparison to the primary tumors.
  • an unsupervised clustering twelve cases of matched tissues (from the same patient) clustered together, with a significantly higher correlation compared to all other unmatched tumor- metastasis couples (average 0.75 vs. 0.58) ( Figures 13B and 13C).
  • the 563 proteins that were significantly upregulated in tumor tissue compared to healthy tissue and the 406 downregulated proteins showed a similar median expression in lymph node metastases compared to the primary tumors (Figure 13D).
  • NIT1 Nitrilase homolog 1 Q86X76; B7Z410; B2R8D1 5.56E-05 1.0581136 Up in LNP
  • MRPS 17 S I 7, mitochondrial K7EQQ2; J3Q Y7; E9PNM0; 2.2E-05 0.7960314 Up in LNP F2Z2N8; 7ELI5; E9PSE6;
  • 60S ribosomal protein P83881 J3KQN4; H0Y5B4;
  • TIMM23 ' translocase subunit B4DDK6; B7ZB25; B4DI18;
  • RPS14 S14 A4D1M5 5.47E-05 0.477388 Up in LNP 40S ribosomal protein P42677; Q5T4L4; A4D1G5;
  • EPFP1 protein mitochondrial A0A024R3X7; B8ZZ54 0.001554 0.3817505 Up in LN
  • subunit 4 isoform 1, PI 3073; H3BN72; H3BNV9;
  • RPLPO ribosomal protein P0- F8VWV4; F8VS58; F8W1 K8;
  • Keratin type I Q13092; K7ENV3 ;Q7Z3Y9;
  • B4DFM1 B4DJB 1 ; C9JJ47;
  • a 15-protein signature predicts lymph node involvement
  • lymph node status is a crucial component in treatment decision making, but currently there are no reliable ways to determine node involvement without physically evaluating the lymph node.
  • RFE recursive feature elimination
  • the inventors plotted the false positive (FP; prediction of LNN as LNP) and false negative (FN; prediction of LNP as LNN) rates, together with the area under the receiver operating characteristic curve (AUC of ROC) generated by the classifier, using increasing number of features for classification ( Figure 14A and Table 4).
  • the optimal performance of the classifier using a minimum number of features, denoted by a low false incidents and high AUC, was reached using 15 proteins for classification (FP incidents: 3 out of 21; FN incidents: 4 out of 20; AUC 0.93, Figure 14B).
  • BM12B RNA-binding protein 12B E5RJ83; E5RJV8; E5RJW8 2 3
  • RJBM3 Putative RNA-binding protein 3 P98179; A0A024QYX3 4 5
  • Methionine aminopeptidase 2 P50579; B4DUX5; G3V1U3;

Landscapes

  • Chemical & Material Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Organic Chemistry (AREA)
  • Immunology (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Analytical Chemistry (AREA)
  • Zoology (AREA)
  • Molecular Biology (AREA)
  • Wood Science & Technology (AREA)
  • Pathology (AREA)
  • Physics & Mathematics (AREA)
  • General Health & Medical Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Biotechnology (AREA)
  • Microbiology (AREA)
  • Biochemistry (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Hematology (AREA)
  • Urology & Nephrology (AREA)
  • General Engineering & Computer Science (AREA)
  • Biophysics (AREA)
  • Biomedical Technology (AREA)
  • Cell Biology (AREA)
  • Hospice & Palliative Care (AREA)
  • Oncology (AREA)
  • Food Science & Technology (AREA)
  • Medicinal Chemistry (AREA)
  • General Physics & Mathematics (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
  • Investigating Or Analysing Biological Materials (AREA)

Abstract

The invention relates to diagnostic and prognostic methods, compositions and kits for determining the progression of breast cancer in a subject and for early detection of metastatic breast cancer.

Description

METHODS AND KITS FOR BREAST CANCER PROGNOSIS FIELD OF THE INVENTION
The invention relates to personalized medicine. More particularly, the invention relates to diagnostic and prognostic methods and kits for detecting metastatic breast cancer and for monitoring breast cancer progression.
BACKGROUND OF THE INVENTION
References considered to be relevant as background to the presently disclosed subject matter are listed below:
Barla, A., Jurman, G., Riccadonna, S., Merler, S., Chierici, M., and Furlanello, C. (2008). Machine learning methods for predictive proteomics. Briefings in Bioinformatics 9, 119-128.
Benjamini, Y., and Hochberg, Y. (1995). Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society Series B (Methodological) 57, 289-300.
Cox, J., and Mann, M. (2008). MaxQuant enables high peptide identification rates, individualized p.p.b. -range mass accuracies and proteome-wide protein quantification. Nat Biotech 26, 1367-1372. Cox, J., and Mann, M. (2012). ID and 2D annotation enrichment: a statistical method integrating quantitative proteomics with complementary high-throughput data. BMC Bioinformatics 13, S 12. Cox, J., Neuhauser, N., Michalski, A., Scheltema, R.A., Olsen, J.V., and Mann, M. (2011). Andromeda: A Peptide Search Engine Integrated into the MaxQuant Environment. Journal of Proteome Research 10, 1794-1805.
Duarte, N.C., Becker, S.A., Jamshidi, N., Thiele, I., Mo, M.L., Vo, T.D., Srivas, R., and Palsson, B.0. (2007). Global reconstruction of the human metabolic network based on genomic and bibliomic data. Proceedings of the National Academy of Sciences 104, 1777-1782.
Geiger, T., Cox, J., Ostasiewicz, P., Wisniewski, J.R., and Mann, M. (2010). Super-SILAC mix for quantitative proteomics of human tumor tissue. Nat Meth 7, 383-385.
Geiger, T., Madden, S.F., Gallagher, W.M., Cox, J., and Mann, M. (2012). Proteomic Portrait of Human Breast Cancer Progression Identifies Novel Prognostic Markers. Cancer Research 72, 2428- 2439.
Geiger, T., Wisniewski, J.R., Cox, J., Zanivan, S., Kruger, M., Ishihama, Y., and Mann, M. (2011). Use of stable isotope labeling by amino acids in cell culture as a spike-in standard in quantitative proteomics. Nat Protocols 6, 147-157.
Jatoi, I., Hilsenbeck, S.G., Clark, G.M., and Osborne, C.K. (1999). Significance of Axillary Lymph Node Metastasis in Primary Breast Cancer. Journal of Clinical Oncology 17, 2334. Li, Z., Huang, C, Bai, S., Pan, X., Zhou, R., Wei, Y., and Zhao, X. (2008). Prognostic evaluation of epidermal fatty acid-binding protein and calcyphosine, two proteins implicated in endometrial cancer using a proteomic approach. International Journal of Cancer 123, 2377-2383.
Liu, C.-W., and Jacobson, A.D. (2013). Functions of the 19S complex in proteasomal degradation. Trends in Biochemical Sciences 38, 103-110.
Ma, Y., and Hendershot, L.M. (2004). The role of the unfolded protein response in tumour development: friend or foe? Nat Rev Cancer 4, 966-977.
Ong, S.-E., Blagoev, B., Kratchmarova, I., Kristensen, D.B., Steen, H., Pandey, A., and Mann, M. (2002). Stable Isotope Labeling by Amino Acids in Cell Culture, SILAC, as a Simple and Accurate Approach to Expression Proteomics. Molecular & Cellular Proteomics 1, 376-386.
Ostasiewicz, P., Zielinska, D.F., Mann, M., and Wisniewski, J.R. (2010). Proteome, Phosphoproteome, and N-Glycoproteome Are Quantitatively Preserved in Formalin-Fixed Paraffin- Embedded Tissue and Analyzable by High-Resolution Mass Spectrometry. Journal of Proteome Research 9, 3688-3700.
Perou, CM., Sorlie, T., Eisen, M.B., van de Rijn, M., Jeffrey, S.S., Rees, C.A., Pollack, J.R., Ross, D.T., Johnsen, H., Akslen, L.A., et al. (2000). Molecular portraits of human breast tumours. Nature 406, 747-752.
Rappsilber, J., Mann, M., and Ishihama, Y. (2007). Protocol for micro-purification, enrichment, pre-fractionation and storage of peptides for proteomics using StageTips. Nat Protocols 2, 1896- 1906.
Sakorafas, G.H., Peros, G., Cataliotti, L., and Vlastos, G. (2006). Lymphedema following axillary lymph node dissection for breast cancer. Surgical Oncology 15, 153-165.
S0rlie, T., Tibshirani, R., Parker, J., Hastie, T., Marron, J.S., Nobel, A., Deng, S., Johnsen, H., Pesich, R., Geisler, S., et al. (2003). Repeated observation of breast tumor subtypes in independent gene expression data sets. Proceedings of the National Academy of Sciences 100, 8418-8423.
Sotiriou, C, and Pusztai, L. (2009). Gene-Expression Signatures in Breast Cancer. New England Journal of Medicine 360, 790-800.
Umar, A., Kang, H., Timmermans, A.M., Look, M.P., Meijer-van Gelder, M.E., den Bakker, M.A., Jaitly, N., Martens, J.W.M., Luider, T.M., Foekens, J.A., et al. (2009). Identification of a Putative Protein Profile Associated with Tamoxifen Therapy Resistance in Breast Cancer. Molecular & Cellular Proteomics 8, 1278-1294.
Whitaker-Menezes, D., Martinez-Outschoorn, U. E., Lin, Z., Ertel, A., Flomenberg, N., Witkiewicz, A. K., ... & Pestell, R. G. (2011). Evidence for a stromal-epithelial "lactate shuttle" in human tumors: MCT4 is a marker of oxidative stress in cancer-associated fibroblasts. Cell cycle, 10(11), 1772- 1783. Wisniewski, J.R., Zougman, A., and Mann, M. (2009a). Combination of FASP and StageTip-Based Fractionation Allows In-Depth Analysis of the Hippocampal Membrane Proteome. Journal of Proteome Research 8, 5674-5678.
Wisniewski, J.R., Zougman, A., Nagaraj, N., and Mann, M. (2009b). Universal sample preparation method for proteome analysis. Nat Meth 6, 359-362.
Wilhelm, M., Schlegl, J., Hahne, H., Gholami, A. M., Lieberenz, M., Savitski, M. M., ... & Mathieson, T. (2014). Mass-spectrometry-based draft of the human proteome. Nature, 509(7502), 582-587.Henras, A. K., Soudet, J., Gerus, M., Lebaron, S., Caizergues-Ferrer, M., Mougin, A., & Henry, Y. (2008). The post-transcriptional steps of eukaryotic ribosome biogenesis. Cellular and Molecular Life Sciences, 65(15), 2334-2359.
Ellis, M. J., & Perou, C. M. (2013). The genomic landscape of breast cancer as a therapeutic roadmap. Cancer discovery, 3(1), 27-34.
WO 2004/079370
WO 2011/042467
Acknowledgement of the above references herein is not to be inferred as meaning that these are in any way relevant to the patentability of the presently disclosed subject matter.
BACKGROUND OF THE INVENTION
Over the past decade there has been a tremendous progress in the characterization of fundamental molecular determinants of breast cancer. The development of high throughput technologies has made it possible to approach the disease in a global manner, and assess the contribution of the three players in the "central dogma" - namely DNA, RNA and protein - to the tumorigenic phenotype. While the genomic and transcriptomic levels have been extensively studied, due to technological challenges, the proteomic level has been mainly studied using cell lines or with low analytical depth. Seminal gene-expression studies defined molecular signatures that allowed the classification of breast tumors into four accepted "intrinsic" subtypes: luminal A and B, Her2-overexpressing and basal-like tumors (Perou et al., 2000). Further genomic and transcriptomic efforts expanded and refined the original signatures to slightly alter the classification. These efforts culminated in studies which make up the largest breast cancer genomic-profiling done to date, combining data from multiple platforms and utilizing next-generation techniques to study up to 2,000 breast tumors. These genomic and transcriptomic data serve as an invaluable resource of breast cancer associated mutations, chromosomal aberrations and further expanded the classification to additional subtypes. However, the actual manifestation of such genomic changes in the cancer phenotype is far from obvious. Proteomics makes a natural complement to the genomic and transcriptomic studies. As proteins convey the actual functional properties of cells, they represent the final combined effect of all genetic abnormalities, including mutations, copy-number variations, epigenetic and transcription- level regulation. Mass spectrometry (MS)-based proteomic analyses have undergone a revolution in the past decade, owing to improvements in instrumentation, sample preparation and quantification methods. The advanced MS instruments combine high resolution, high mass accuracy and high speed, and are capable of comprehensively cataloguing proteomes of yeast, mouse, and most recently, human. The implementation of proteomics to cancer studies is increasing; however, many of these studies are still limited in scope and quantification accuracy. A major improvement in quantification technique has been the introduction of Stable Isotope Labeling with Amino acids in Cell culture (SILAC)(Ong et al., 2002). In this metabolic labeling approach, cell lines incorporate heavy isotope versions of lysine and arginine into their proteome; when mixed with proteins from an unlabeled cell line, each protein will be present in the solution in two forms (labeled and unlabeled). This allows a very accurate relative quantification of the protein. A SILAC-labeled cell line can then be used as a common, 'spike-in', internal standard for comparing a theoretically unlimited number of samples, including samples that cannot be metabolically labeled such as clinical tumor samples (Geiger et al., 2011). Owing to the great complexity of clinical tumor samples and the wide repertoire of expressed proteins, a single SILAC cell line was found to be insufficient for accurate quantification. Rather, a mixture of cell lines, termed a super-SILAC mix, was found to dramatically increase the quantification accuracy of breast cancer proteomes (Geiger et al., 2010). The super-SILAC mix is further disclosed in WO 2011/042467 that is a previous application by one of the inventors. The combination of fast, high-resolution mass spectrometers for identification of proteins, and super-SILAC for their quantification, puts a deep view into breast tumors proteome at hand.
Clinical assessment of the intrinsic subtypes showed their relevance to the determination of breast cancer prognosis (S0rlie et al., 2003). Luminal tumors make up the vast majority of breast tumors and are characterized by expression of the estrogen receptor (ER). As such, they can be treated by endocrine therapy such as Tamoxifen. Luminal A tumors show overall favorable prognosis, while luminal B, which also express higher levels of the proliferation marker ki67 or Her2, have somewhat poorer prognosis. Despite the overall good prognosis of the luminal tumors, risk of recurrence rises substantially if the cancer had already metastasized to nearby lymph nodes by the time of diagnosis (lymph node positive, LNP). While lymph node negative patients (LNN) benefit from endocrine therapy alone, LNP patients are not likely to benefit from such therapy despite high ER expression levels. Moreover, increasing evidence shows that the added value from adjuvant chemotherapy is also questionable (Ellis and Perou, 2013), making the lymph node status a crucial component in treatment decision-making.
There is therefore a need for reliable, sensitive and rapid diagnostic methods for early determination of the patient's metastatic state using primary tumor samples. Such methods are specifically applicable in determining a suitable and efficient treatment regimen to the diagnosed subject. The above object is solved by the methods and kits of the invention that provide powerful tools for personalized medicine.
SUMMARY OF THE INVENTION
In a first aspect, the invention relates to a diagnostic and/or prognostic method for determining the progression of breast cancer in a subject. In certain embodiments, the method of the invention comprises the steps of: (a) determining the expression level of at least one biomarker protein in at least one biological sample of said subject, to obtain an expression value for each of the at least one biomarker protein/s. The biomarker proteins of the invention may be at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof; (b) calculating and determining if the expression value obtained in step (a) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample. In more specific embodiments, at least one of: (i) a positive expression value of said at least one biomarker protein/s in said sample, indicates that said subject belongs to a pre-established population associated with positive lymph node metastatic status; and (ii) a negative expression value of said at least one biomarker protein/s in said sample, indicates that said subject belongs to a pre-established population associated with negative lymph node metastatic status. According to a second aspect, the invention relates to a diagnostic and/or prognostic composition comprising at least one detecting molecule or any combination or mixture of plurality of detecting molecules specific for determining the level of expression of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof.
A third aspect of the invention relates to a kit comprising (a) detecting molecules specific for determining the level of expression of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof in a biological sample. BRIEF DESCRIPTION OF THE DRAWINGS
In order to better understand the subject matter that is disclosed herein and to exemplify how it may be carried out in practice, embodiments will now be described, by way of non-limiting example only, with reference to the accompanying drawings, in which:
Figures 1A to ID. General overview of the study
Fig. 1A. is a workflow depicting cohort assembly and sample preparation and analysis.
Fig. IB. is a bar graph showing the number of proteins quantified in this study and in each of the clinical groups.
Fig.lC. is a graph showing the distribution of expression intensities of the quantified proteins (filtered for proteins that were quantified in >10 samples) shows a large dynamic range of abundance, still the vast majority of the proteins are expressed within 4 orders of magnitude.
Fig. ID. shows a comparison to the breast- specific protein set (Wilhelm et al, 2014) shows a significant enrichment for these proteins (gray bars) in all groups of clinical samples (ID annotation enrichment, FDR=0.02. Median expression intensities are depicted in black (group overall) or gray (breast subset).
Figure 2A to 2C. Data distribution in clinical samples
Fig 2A. shows the number of proteins quantified in each sample.
Fig 2B. shows the protein occurrence by sample.
Fig 2C. is an intensity distribution of 'light' versus 'heavy' proteins shows that >90% of the protein intensities are within 5-fold from the heavy standard in all clinical groups (healthy tissue 90.9%, tumor tissue 91.7% and metastasic tissue 92.7% of the proteins with 5-fold (log 2.3) ratio).
Figure 3. Correlation networks of breast cancer clinical samples
Figure shows Pearson correlation for 2974 proteins that were present in >70% of the clinical samples. Network construction based on k-means clustering of the correlation (k=10) with a cutoff correlation of 0.3. Cluster 1 (DNA replication, rRNA processing, Pyrimidine and Purin metabolism, DNA repair), Cluster 2 (proteosome, protein export, spliceosome, regulation of cell cycle), Cluster 3 (TCA cyclate, Glutathion metabolism, protein glycosylation, RNA Splicing), Cluster 4 (Fatty acid β-oxidation, Peroxisome, Oxidative phosphorylation), Cluster 5 (RNA polyadenylation), Cluster 6 (Ribosome, Oxidative phosphorylation), Cluster 7 (Ribosome, Oxidative phosphorylation, Type-I interferon signaling, Antioxidant activity, Lysosome), Cluster 8 (Proteasome, tRNA aminoacylayion, Glycolysis/gluconeogenesis, Focal adhesion), Cluster 9 (Ribosome biogenesis, Glucose transport). Enriched pathways in each cluster are indicated (FDR=0.02). Figure 4A to 4D. Segregation between healthy tissue and breast cancer tumors
Fig 4A. shows a principal component analysis of healthy and cancer samples. Abbreviations: LNN, lymph node negative; LNP, lymph node positive; Pr, primary. Healthy samples are denoted by open circles, LN metastasis with black circles, LNP Pr. Tumor with filled circles, and LNN Pr with gray circles.
Fig 4B. is a hierarchical clustering of sample Pearson correlations.
Fig 4C. shows volcano plots demonstrating significantly changing proteins in primary tumors compared to healthy breast epithelia (black dots; FDR=0.05). Enriched GO and KEGG pathway annotations are presented (Fisher exact test, FDR=0.02. *p<0.001, †p<0.0001. Significantly changing proteins are colored in a brighter color.
Fig 4D. is a pulsed-SILAC experiment in normal mammary epithelial cells (HMEC) and ER- positive breast cancer cell line (MCF7) shows higher rates of synthesis, degradation and turnover in the cancer cell line. Gradient describes the change in H/L (left), M/L (middle) and H/M (right) ratios at Oh, 2h, 4h, 9, 12h and 24h after the pulse. Graphs represent mean of two biological replicates.
Figure 5A to 5E. Correlation between clinical groups.
Fig 5A. shows Pearson correlation calculated for proteins across tumor samples and a network construction by k-means (k=7) with cutoff correlation of 0.5. Cluster 1 (Ribosome, Translation), Cluster 5 (Ribosome, Nonsense-mediated decay, Protein targeting to ER, Translation, Lysosome, Oxidative phosporylation, Spliceosome), Cluster 6 (MCM complex, DNA replication, DNA repair, Protein folding, Locomotion, Nucleoulus, Cluster 7 (Peroxisome, Oxidative phosphorylation, Chemical homeostasis). Enriched pathways are indicated (FDR=0.0.2).
Fig 5B. shows that the correlation among the cancer samples (primary tumors and metastases is higher than among the healthy tissue samples. (Mann- Whitney p<0.0001).
Fig 5C. shows that the extracellular matrix (ECM) part proteins make a highly variably subgroup in both healthy and tumor samples (FDR=0.02).
Fig 5D. shows the fraction of the total intensity of ECM part proteins is significantly higher in healthy tissue compared to tumor and metastatic tissue.
Fig 5E. shows the fraction of the total intensity of ECM part proteins is significantly higher in primary LNN compared to primary LNP tumors.
Figure 6A to 6F. Proteomic differences between healthy breast duct epithelia and cancer cells
Fig 6A. shows that ratio distribution of the tumor downregulated and upregulated subsets (from right to left respectively) is significantly narrower than the overall distribution of proteins (the left measurements) among tumor samples (Bartlett's test, p=5.3e-21 and 2.3e-39, respectively). The box in the figure represents area of accurate quantification. Percentages denote the number of ratios within that range for every group.
Fig 6B. shows an enrichment and de-enrichment of the upregulated and downregulated subsets in comparison to the overall distribution of ratios among the tumors, respectively(lD annotation enrichment, FDR=0.02).
Fig 6C. shows that CTSD is upregulated in tumors (Welch's t-test FDR=0.05). (Welch's t-test FDR=0.05).
Fig 6D. shows that MUC1 is upregulated in tumors (Welch's t-test FDR=0.05). (Welch's t-test FDR=0.05).
Fig 6E. shows activity assay of mitochondrial complex I and complex IV, evaluated by histochemistry on frozen breast cancer tumor arrays (duct carcinoma in situ). Representative tissue cores are presented.
Fig 6F. showing staining intensity of complex I and IV, one and three normal cores, respectively, and fifteen ER-positive, Her2-negative breast cancer cores were evaluated. Bars represent average scores +SEM.
Figure 7A-7C. Metabolic remodeling in primary breast tumors
Fig. 7A. is a diagram depicting the changes in some major metabolic pathways in tumors compared to healthy breast epithelia. Solid black and dashed black represent significant upregulation or down regulation in tumors, respectively (Welch's t-test, FDR=0.05). Underlined font and Bold Italicized font represent upregulated or down regulated proteins, respectively. Gray lines show the reactions that are not significantly changing.
Fig. 7B. Expression changes of metabolic enzymes between healthy tissue and primary tumors, grouped by Recon 1 metabolic pathway. For pathway statistical analysis, an average expression all enzymes in the group was calculated for each sample (Welch's t-test, FDR=0.05 was used for both protein expression and average pathway expression.†p<0.02, *p<0.01, **p<0.001, ***p<0.0001, unmarked, not statistically significant.
Fig. 7C. The overexpression of three metabolic enzymes - glutamine synthethase (GLUL), acyl- CoA thioesterase (ACOT)-l/2 and oxoglutarate/malate carrier (SLC25A11) was validated using immunohistochemistry on tumor arrays.
Figure 8. RECON 1 Significantly Changing Pathways between Primary Tumors and Healthy Tissues
Figure shows upregulated and downregulated proteins of Recon 1 pathways. Each number represents a protein: 1 - Pentose and Glucuronate Interconversions, 2 - Heparan sulfate degradation, 3 - Tetrahydrobiopterin, 4 - Hyaluronan Metabolism, 5 - N-Glycan Degradation, 6 - Chondroitin sulfate degradation, 7 - Keratan sulfate degradation, 8 - Oxidative Phosphorylation, 9 - Pyrimidine Catabolism, 10 - Heme Degradation, 11 - ROS Detoxification, 12 - Sphingolipid Metabolism, 13 - Glutathione Metabolism, 14 - Transport, Lysosomal, 15 - Heme Biosynthesis, 16 - Galactose metabolism, 17 - Nucleotides, 18 - Fatty Acid Metabolism, 19 - Glycolysis/Gluconeogenesis, 20 - Arginine and Proline Metabolism, 21 - Pentose Phosphate Pathway, 22 - Selenoamino acid metabolism, 23 - Lysine Metabolism, 24 - Propanoate Metabolism, 25 - Glyoxylate and Dicarboxylate Metabolism, 26 - Methionine Metabolism, 27 - Histidine Metabolism, 28 - IMP Biosynthesis, 29 - Alanine and Aspartate Metabolism, 30 - Ascorbate and Aldarate Metabolism, 31 - Cholesterol Metabolism, 32 - Fatty acid activation, 33 - Glycine, Serine, and Threonine Metabolism, 34 - Transport, Extracellular.
Figure 9A to 9E. Immunohistochemistry for Significantly Changing Proteins
Fig 9A. shows staining intensity of Acyl-CoA thioesterase-1 (ACOT1), glutamine-synthetase (GLUL) and oxoglutarate carrier (SLC25A11) antibodies. The antibodies were used to stain FFPE tumor microarrays containing three duplicates of normal breast duct epithelia and 15 duplicate cores of ER-positive, Her2-negative invasive ductal carcinomas. Bars represent average scores +SEM (**p<0.001, ***p<0.0001).
Fig 9B and Fig 9C. show representative figures from the Human Protein Atlas database showing differential staining between healthy and tumor cores.
Fig 9D. shows validation of 12 downregulated proteins using the Human Protein Atlas database. Fig 9E. shows validation of 24 upregulated proteins using the Human Protein Atlas database.
Figure 10. Expression of electron transport chain components in breast tissue (healthy vs. primary tumor)
Figure shows expression changes in electron transport chain components between healthy tissue
(light gray) and primary tumors (dark gray) (Welch's t-test, FDR=0.05 was used for both protein expression and average pathway expression. *p<0.01, **p<0.001, ***p<0.0001).
Figure 11. Differential RNA and protein expression in healthy tissue and primary tumors
Figure shows a protein ranking according to their log fold change (healthy/tumor). The barcode plots show the ranking distribution of proteins in the category.
Figure 12A to 12B. Changes in protein expression during cancer progression
Fig. 12A. discloses that Ki67 does not discriminate between LNN and LNP tumors (Welch's t-test,
FDR=0.05)
Fig 12B. shows that fifteen proteins changed significantly between healthy tissue and LNN tumor and also between LNN and LNP Tumors. All expression differences are significant with p<0.01 (Welch's t-test, FDR=0.05).
Figure 13A-113E. Analysis of tumor progression Fig. 13A. Protein interaction network of proteins which were upregulated or downregulated (grey dashed nodes) in LNP tumors compared to LNN tumors shows a large network of upregulated ribosomal and translation proteins, as well as splicing factors and mitochondrial proteins (Welch's t-test, FDR=0.05).
Fig. 13B. Unsupervised clustering of 20 matched pairs of primary tumors (T) and lymph node metastases (N) shows a high incident of co-clustering.
Fig. 13C. Pearson correlation between matched pairs of tumors and metastases was significantly higher than the average tumor-metastases correlation.
Fig. 13D. significantly upregulated proteins in primary tumors compared to healthy samples (upper panel) and downregulated proteins (lower panel) show no change in the pattern of expression between primary tumors and metastases. White lines indicate z-scored median expression.
Fig. 13E. Four proteins were significantly upregulated in tumor and downregulated again in the lymph node (Welch's t-test, FDR=0.05). Bars represent medians +SEM.
Figure 14A-14C. A 15-protein signature predicts LNP primary tumors
Fig. 14A. A support vector machines (SVM)-based classifier was trained and tested on the 85 significantly changing proteins between LNN and LNP primary tumors, with different numbers of features (proteins) for classification; top- 15 ranked proteins were chosen as an optimal number for prediction.
Fig. 14B. AUC of ROC curve using 15 proteins for classification was 0.93.
Fig. 14C. The 15 proteins of the signature, grouped by cellular functions; all expression differences are statistically significant (Welch's t-test, FDR=0.05).
DETAILED DESCRIPTION OF THE INVENTION
The work described in the present invention presents the first genome- scale analysis of breast cancer progression, which is able to capture novel aspects of cancer development. In a three-level comparison, the inventors capture the functional difference between the healthy control tissue and the tumors, between the primary tumors and the metastases, and between pre-metastatic and metastatic breast cancer. While the inventor's data captured hundreds of regulated proteins in the comparison to the healthy tissues, a more challenging comparison was between the groups of tumor tissues. Surprisingly, despite the distinct microenvironment, the inventors found greater similarity between the primary tumors and the lymph node metastases than between two groups of primary tumors, associated with the tumor stage.
Breakdown of translational quality control processes
The most prominent network of upregulated proteins in tumors consisted of structural ribosomal proteins, with a concurrent down regulation of several of the most important co-players of the translational machinery - the tRNA aminoacyl synthetases (ARSs), and also of the auxiliary protein AEVIP2. The canonical role of the ARSs is to ligate an amino acid to its cognate tRNA, later to be added to the nascent polypeptide chain. Improper activity of the ARSs may impair the accuracy of protein synthesis, and not allow for appropriate folding of proteins. Furthermore, even if the proteins are correctly translated, the marked decrease in the expression of important chaperons may also have an adverse effect on their function. This in turn can inflict tumorigenesis if the misfolded proteins are tumor suppressors, as have been demonstrated for p53 and VHL. Alternatively, it can induce gain-of-function activities or interfere with protein localization. In recent years, increasing evidence point to additional, non-canonical roles that may make the ARSs important regulators of diverse cellular functions. These include activation of p53, interaction with transcription factors and regulation of angiogenesis. AEVIP2 mediates anti-proliferative and pro-apoptotic functions through regulation of ubiquitination, and as an outcome, mice lacking this protein died neonatally due to severe over-proliferation of lung epithelia, making AIMP2 a bona-fide tumor suppressor. The synthetase YARS has been shown to be secreted and cleaved into two fragments, one of which acts as a cytokine to induce angiogenesis. Secretion of YARS and possibly other ARSs by cancer cells may explain their reduced intracellular levels and may directly affect tumor progression through interaction with its microenvironment. Further efforts will be needed in order to elucidate the roles of tRNA aminoacyl synthetases in tumorigenesis.
The inventors propose that the elevated rate of protein production and degradation, together with down regulation of several quality control systems, impairs protein homeostasis in the cancer cells. Accumulation of DNA damage in the tumors, potentially due to high levels of ROS and impairment of repair mechanisms may lead to the accumulation of mutated transcripts; higher ribosome levels fail to produce functional proteins due to translation of damaged transcripts and reduced activity of ARSs; and the last line of defense against the accumulation of such proteins - the chaperones of the unfolded protein response - fail to launch a protective campaign. Damaged proteins may eventually be degraded by the proteasome or by lysosomal proteases, overall increasing protein turnover rates. Despite the tremendous energetic demand of such a mechanism, we speculate that it provides the system the necessary adaptability to changing conditions, and confers an evolutionary advantage to the cancer cells.
Metabolic alterations in cancer
The link between neoplastic transformation and cell metabolism has only recently received a deserved attention as an emerging hallmark of cancer. As tumor cells possess an ability to grow in changing environments, metabolic adaptations that support this growth must take place, as the cells need excess amounts of nucleotides, amino acids, fatty acids and ATP to carry out anabolic reactions. The most famous of these alterations, the Warburg effect, states that cancer cell produce ATP by increasing their glycolysis rates even in the presence of oxygen, despite the energetic inefficiency comparing to oxidative phosphorylation in the mitochondria. However, large-scale cancer studies from the last decade unveil enormous diversity even within cancers from the same origin; it is therefore hard to speculate that this phenomenon is common to all cancer subtypes. In contrast to the Warburg effect, significant upregulation of all four complexes of the ETC and the ATP synthase was found herein in cancer cells in comparison to normal cells, implying that the tumor cells rely heavily on cellular respiration for energy production. Increased cellular respiration requires increased availability of NADH. Based on the proteomic data, we propose that a major NADH source may be peroxisomal β-oxidation (based on the elevation of ECH1, HSD17B4 and ACAA1). This pathway ultimately leads to the production of acetyl-CoA and NADH, both of which are exported to the cytosol. NADH can then be shuttled in the form of reducing equivalents to the mitochondria, through the malate/aspartate shuttle. The inventors found that two components of this system - namely, the oxoglutarate/malate carrier (SLC25A11) and mitochondrial aspartate aminotransferase (GOT2) - are upregulated significantly in tumors, while the intermediate enzyme in this pathway, mitochondrial malate dehydrogenase (MDH2) was also upregulated, albeit to a lesser extent.
In agreement with the elevated beta-oxidation, the inventors found reduced levels of two major regulators of fatty acid synthesis, ATP citrate lyase (ACLY) and fatty acid synthase (FASN) as well as cholesterol biosynthesis enzymes (Figures 7 A and 7B).
Interestingly, it has been shown that knockdown of ACLY in differentiating adipocytes causes a significant decrease in glucose uptake and down regulation of several glycolytic enzymes, and this affect was attributed to global changes in histone acetylation, associated with glucose availability. The data of the present invention show a similar decrease in glycolytic enzymes such as hexokinase (HK)-2, glyceraldehyde 3 phosphate dehydrogenase (GAPDH) and pyruvate kinase (PKM), proposing the notion that ER-positive breast cancer presents an anti- Warburg phenotype. Without being bound by any theory, the inventors propose that the decrease in glycolysis rates together with the increase in oxidative phosphorylation stem from the proximity to adipose tissue within the breast and potentially, reduced glucose levels. Increased fatty acid catabolism may suggest an adaptation of the invading cancer cells to their surroundings and selection for cells that adequately change their metabolism.
Much like the Warburg dependence on glucose, some cancers are also known to require extracellular glutamine, up to the point of 'addiction'. The present invention reports herein increased glutamine production from glutamate by glutamine synthetase (GLUL), accompanied by a decrease in the opposite reaction, catalyzed by glutaminase (GLS), and down regulation of a primary glutamine importer, SLC1A5. These results indicate that the breast tumors in this study may be independent of extracellular glutamine. Glutamine addiction has been observed in basal-like breast cancer cell lines and primary tumors, but not in luminal cells.
An additional important metabolic alteration apparent through the proteomic data was the increased ROS -detoxification enzymes. A significant upregulation of several antioxidant molecules was observed, among which are glutathione peroxidases (GPX) and peroxiredoxin (PRDX) enzymes, which use reduced glutathione (GSSH) to protect the cells from oxidative stress. The formation of GSSH requires NADPH, which is thought to be produced mostly by the pentose phosphate pathway (PPP). Intriguingly, we found that the PPP is significantly downregulated in tumors, primarily due to a significant reduction in rate-limiting enzyme glucose-6-phosphate dehydrogenase (G6PD), potentially lowering the availability of cytosolic NADPH. A recent study examined the NADPH/NADP+ levels in A549 and MCF7 cell lines under stress conditions such as glucose deprivation and matrix detachment. The authors suggest that an AMPK-mediated inhibition of acetyl-CoA carboxylase alpha (AC AC A) and beta (ACACB) serves to maintain the NADPH/NADP+ balance through inhibiting fatty acids synthesis and increasing fatty acid oxidation, mirroring in our data by a decrease in expression of AC AC A in tumors and the overall decrease in fatty acid synthesis. Therefore, the decreased consumption of NAPDH in fatty acid synthesis is potentially supporting the elevated activity of ROS -detoxifying enzymes.
A 15 -protein signature predicts lymph node involvement based on the primary tumor
The unfolding of the Omics' era allowed for the use of large-scale data to assist in the prediction of several important clinical features of cancer, such as the risk for recurrence, tumor aggressiveness and response to therapy, and several commercial assays were developed as a result (reviewed in Sotiriou and Pusztai, 2009). The vast majority of these predictive signatures is based on gene expression profiling, and recently, proteomic studies are being increasingly implemented in prognostic and diagnostic context (Geiger et al., 2012; Umar et al., 2009). The present invention presents a 15-protein signature (shown in Table 4) that predicts the involvement of lymph node based on expression levels in the primary tumors, using supervised classification algorithms on proteins that significantly change in expression between LNN and LNP luminal tumors. To the inventor's knowledge, this is the first proteomic signature that allows such a classification. The proteomic signature of the invention combines biological importance with high AUC (0.93) and low error rates (14% for LNN primary tumors and 20% for LNP tumors). Several of the proteins found in the signature of the present invention have been reported to have a role in cancer progression, while others are novel indicators of aggressiveness. The highest-changing protein between LNN and LNP tumors, calcyphosine (CAPS), has recently been identified by a proteomic study as a marker of poor prognosis in endometrial cancer (Li et al., 2008). The splicing factor SRSF1 is overexpressed in lung cancer and is a transcriptional target of MYC, and the cathepsin inhibitor cystatin B (CSTB) is overexpressed in ovarian cancer. The inventors envision that this signature could be implemented into clinical applications in the future, to determine the aggressiveness of the tumor already by the initial biopsy, thereby potentially obliterating the need for intrusive surgical disruption of the lymph nodes, while also preventing unnecessary systemic treatment.
The present invention therefore presents a deep proteomic analysis of breast cancer progression and provide a novel, high-quality proteomic database. The invention shows that the regulation on protein production is severely impaired in tumors, and that luminal tumors display an anti-Warburg effect, characterized by high levels of cellular respiration and low glycolytic activity. The invention shows that the proteomic landscape of LNN and LNP primary tumors is similar, and that lymph node metastases retain the expression patterns of their original tumor. Finally, the present invention provides a 15 -protein signature that accurately predicts lymph node involvement based on the primary tumor.
Predicting the onset of a disease is highly valuable and clinically desired particularly for diseases that are often detected at advanced stages. Specifically, breast cancer diagnosis and prognosis are highly important and crucial in management patient's life quality. Providing tools and specifically non-invasive tools to accurately diagnose at early stage breast cancer and predict its progression may assist in reducing mortality and increase life quality of patients.
In addition, determining treatment regimen and monitoring patient response to treatment is highly valuable and clinically desired as it can provide information regarding suitable and successful treatment protocols enabling personalized medicine. This is appreciated in view of the fact that treatment protocols are often associated with some extent of undesired side effects.
Thus, predicting disease occurrence, progression and response to treatment may be considered as life saving and have the advantage of avoiding inadequate treatments thereby reducing side effects. As indicated above, in the present invention, the inventors used computational analysis to provide novel, unique, comprehensive and deep proteomics-scale analysis of breast cancer progression providing a high quality proteomic database. The inventors have identified an arsenal of proteins that was differently expressed in healthy tissues and at different stages of breast cancer development such as metastatic and non-metastatic breast cancer and successfully characterized different protein signatures that are unique for healthy tissue, tumor tissues at different stages and metastatic tissues. Using the method described herein, the inventors differentiated healthy breast tissue from breast tumors, pre-metastatic breast tumors from metastatic breast cancer and primary breast tumors from metastases.
Surprisingly, the inventors found significant differences in the expression of proteins in different stages of breast tumors. Specifically, as shown in Example 6 herein, the inventors identified a specific set of signature proteins, specifically, RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB that are differently expressed in primary tumor samples obtained from patients diagnosed as having lymph node negative breast tumors compared to patients diagnosed as having lymph node positive breast tumors.
The inventors have therefore suggested that the identified signatory proteins described herein are suitable as a powerful tool for early diagnosis and prognosis of breast cancer metastasis.
Thus, according to a first aspect, the invention relates to a diagnostic and prognostic method for determining the progression of breast cancer in a subject.
In certain embodiments, the method of the invention comprises the steps of:
First, step (a) involves determining the expression level of at least one biomarker protein in at least one biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein/s, wherein said biomarker proteins are selected from RPS24 (40S ribosomal protein S24), LSM4 (U6 snRNA-associated Sm-like protein LSm4), RBM12B (RNA-binding protein 12B), RPS29 (40S ribosomal protein S29); RBM3 (Putative RNA-binding protein 3), PNP (Purine nucleoside phosphorylase), METAP2 (Methionine aminopeptidase 2;Methionine aminopeptidase), CAPS (Calcyphosin), EIF4A3 (Eukaryotic initiation factor 4A-III;Eukaryotic initiation factor 4A-III, N-terminally processed), SNX12 (Sorting nexin-12), SRSF1 (Serine/arginine-rich splicing factor 1), RAB5C (Ras-related protein Rab-5C), BTF3L4 (Transcription factor BTF3 homolog 4), LSM2 (U6 snRNA-associated Sm-like protein LSm2) and CSTB (Cystatin-B), or any combination thereof. The second step (b) involves calculating and determining if the expression value obtained in step (a) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample.
In more specific embodiments, at least one of: (i) a positive expression value of said at least one biomarker protein/s in said sample, indicates that said subject belongs to a pre-established population associated with positive lymph node metastatic status; and (ii) a negative expression value of said at least one biomarker protein/s in said sample, indicates that said subject belongs to a pre-established population associated with negative lymph node metastatic status. In some further embodiments, the present invention provides a diagnostic and prognostic method for determining the progression of breast cancer in a subject, the method comprising the steps of: (a) providing at least one detecting molecule/s each specific for at least one biomarker protein, specifically, at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB. It should be noted that the detecting molecules may be provided in a diagnostic composition or in a kit either attached to a solid support or alternatively, in a mixture. Thus, the method of the invention encompasses in certain embodiments also the provision of a composition, kit, solid support or mixture comprising at least one detecting molecule specific for at least one of said biomarker proteins of the invention. The next step (b) requires determining the expression level of at least one of said biomarker protein/s in at least one biological sample of the diagnosed subject, to obtain an expression value for each of said at least one biomarker protein/s. The final step (c) requires determining if the expression value obtained in step (b) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample. It should be noted that at least one of: (i) a positive expression value of said at least one biomarker protein/s in said sample, indicates that said subject belongs to a pre-established population associated with positive lymph node metastatic status; and (ii) a negative expression value of said at least one biomarker protein/s in said sample, indicates that said subject belongs to a pre-established population associated with negative lymph node metastatic status.
More particularly, the method of the invention may use as diagnostic and prognostic tool, the expression values of any one of the marker proteins described herein below.
Specifically, determining the expression values of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB proteins may indicate if a subject belongs to a pre-established population associated with negative lymph node metastatic status or with positive lymph node metastatic status.
In some specific embodiments the biomarker protein of the invention is the 40S ribosomal protein S24 (RPS24) Protein. RPS24 as described herein, refers to the human RPS24 (Protein IDs P62847; E7ETK0; A0A087WUS0). This protein is required for processing of pre-rRNA and maturation of 40S ribosomal subunits. In more specific embodiments, the RPS24 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 1.
In other specific embodiments the biomarker protein of the invention is the U6 snRNA-associated Sm-like protein LSm4 (LSM4) protein. LSM4 as described herein, refers to the human LSM4 (protein IDs V9GZ56; Q9Y4Z0; U3KQS7; U3KQK1; M0QXB0). This protein binds specifically to the 3'-terminal U-tract of U6 snRNA. In more specific embodiments, the LSM4 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 2.
In certain embodiments, the biomarker protein of the invention is the RNA-binding protein 12B (RBM12B) protein. RBM12B as described herein refers to the human RBM12B (Protein IDs Q8IXT5; B9ZVT1; E5RHG1; E5RJ83; E5RJV8; E5RJW8). This protein contains several RNA- binding motifs, potential transmembrane domains, and proline-rich regions. In more specific embodiments, the RBM12B protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 3.
In certain embodiments, the biomarker protein of the invention is the 40S ribosomal protein S29 (RPS29) protein. As described herein, this protein refers to the human RPS29 (Protein IDs P62273; A0A087WTT6). This protein is a member of the S 14P family of ribosomal proteins that acts as a component of the 40S subunit and. The protein, which contains a C2-C2 zinc finger-like domain that can bind to zinc, can enhance the tumor suppressor activity of Ras-related protein 1A (KREV1). It is located in the cytoplasm. In more specific embodiments, the RPS29 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 4.
In certain embodiments, the biomarker protein of the invention is the Putative RNA-binding protein 3 (RBM3) protein. As described herein, this protein refers to the human RBM3 (Protein IDs P98179; A0A024QYX3). This protein is a cold-inducible mRNA binding protein that enhances global protein synthesis at both physiological and mild hypothermic temperatures. It reduces the relative abundance of microRNAs, when overexpressed and enhances phosphorylation of translation initiation factors and active polysome formation. In more specific embodiments, the RBM3 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 5. In certain embodiments, the biomarker protein of the invention is the Purine nucleoside phosphorylase (PNP) protein. As described herein this biomarker refers to the human PNP (Protein IDs P00491; V9HWH6; Q8N7G1; G3V5M2; G3V2H3; G3V393). This protein catalyze the phosphorolytic breakdown of the N-glycosidic bond in the beta- (deoxy)ribonucleo side molecules, with the formation of the corresponding free purine bases and pentose- 1 -phosphate. In more specific embodiments, the PNP protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 6.
In certain embodiments, the biomarker protein of the invention is the Methionine aminopeptidase 2; (METAP2) protein. As described herein, this biomarker refers to the human METAP2 (Protein IDs P50579; B4DUX5; G3V1U3; B3KWL6; F8VSC4). This protein co-translationally removes the N-terminal methionine from nascent proteins. It protects eukaryotic initiation factor EIF2S 1 from translation-inhibiting phosphorylation by inhibitory kinases such as EIF2AK2/PKR and EIF2AK1/HCR and plays a critical role in the regulation of protein synthesis. In more specific embodiments, the METAP2 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 7.
In certain embodiments, the biomarker protein of the invention is the Calcyphosin (CAPS) protein. As described herein, this biomarker refers to the human CAPS (Protein IDs Q13938; K7ES72). This protein is a calcium-binding protein. In more specific embodiments, the CAPS protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 8.
In certain embodiments, the biomarker protein of the invention is the Eukaryotic initiation factor 4A-III, (EIF4A3), protein. As described herein, this biomarker refers to the human EIF4A3 (Protein IDs P38919; A0A024R8W0; I3L3H2). This protein is an ATP-dependent RNA helicase. In more specific embodiments, the EIF4A3 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 9.
In certain embodiments, the biomarker protein of the invention is the Sorting nexin-12 (SNX12), Protein. As described herein, this biomarker refers to the human SNX12 (Protein IDs Q9UMY4; Q3SYF1; A0A087X0R6). This protein may be involved in several stages of intracellular trafficking. In more specific embodiments, the SNX12 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 10.
In certain embodiments, the biomarker protein of the invention is the Serine/arginine-rich splicing factor 1 (SRSF1) protein. As described herein, this biomarker refers to the human SRSF1 (Protein. IDs Q07955; J3KTL2; Q59FA2; A8K1L8; J3KSR8; J3QQV5; J3KSW7). This protein plays a role in preventing exon skipping, ensuring the accuracy of splicing and regulating alternative splicing. It interacts with other spliceosomal components, via the RS domains, to form a bridge between the 5'- and 3'-splice site binding components, Ul snRNP and U2AF. In more specific embodiments, the SRSF1 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 11. In certain embodiments, the biomarker protein of the invention is the Ras-related protein Rab-5C (RAB5C) protein. As described herein, this biomarker refers to the human RAB5C (Protein IDs P51148; A0A024R1U4; K7ERI8; F8VVK3; K7ENY4; F8VWU4; F8VSF8; K7EIP6; F8VWZ7). This is a protein transport and may be involved in vesicular traffic. In more specific embodiments, the RAB5C protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 12.
In certain embodiments, the biomarker protein of the invention is the Transcription factor BTF3 homolog 4 (BTF3L4) protein. As described herein, this biomarker refers to the human BTF3L4 (Protein IDs Q96K17; Q6PJ77; E9PL10). This is a protein-coding gene. In more specific embodiments, the BTF3L4 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 13.
In certain embodiments, the biomarker protein of the invention is the U6 snRNA-associated Sm- like protein LSm2 (LSM2) protein. As described herein, this biomarker refers to the human LSM2 (Protein ID Q9Y333). This protein binds specifically to the 3 '-terminal U-tract of U6 snRNA and may be involved in pre-mRNA splicing. In more specific embodiments, the LSM2 protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 14.
In certain embodiments, the biomarker protein of the invention is the Cystatin-B (CSTB) protein. As described herein, this biomarker refers to the human CSTB (Protein IDs P04080; Q76LA1). This protein is an intracellular thiol proteinase inhibitor and tightly binding reversible inhibitor of cathepsins L, H and B. In more specific embodiments, the CSTB protein as used herein comprises the amino acid sequence as denoted by SEQ ID NO. 15.
In some embodiments, the expression value of at least one biomarker protein, at times at least two proteins, at times at least three proteins, at times at least four proteins, at times at least five proteins, at times at least six proteins, at times at least seven proteins, at times at least eight proteins, at times at least nine proteins, at times at least ten proteins, at times at least eleven proteins, at times at least twelve proteins, at times at least thirteen proteins, at times at least fourteen proteins, at times at least fifteen proteins of any one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB is determined.
In certain embodiments, the methods of the invention may involve determination of the expression level of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins. It should be noted that the biomarker proteins of the invention are disclosed in Table 4 herein after.
In yet some particular and non-limiting embodiments, the method of the invention may involve in step (a) determination of the expression level of at least five biomarker proteins in at least one biological sample of the examined subject, to obtain an expression value for each of the at least five biomarker proteins. It should be noted that at least five biomarker proteins may be selected from RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins. In some particular and non-limiting embodiments of the invention, such at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise LSM2, METAP2, RPS24, RBM12B and CAPS. In yet some further alternative embodiments, the at least five biomarker proteins may comprise RPS29, RBM3, PNP, METAP2 and RAB5C. In some embodiments, the five biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten of the LSM4, RPS29, RBM3, PNP, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4 and CSTB biomarker proteins of the invention. In yet another embodiment, the selected biomarker proteins may further comprise at least one of RPS24, LSM4, RBM12B, CAPS, EIF4A3, SNX12, SRSFl, BTF3L4, LSM2 and CSTB.
In yet some further alternative embodiments, the method of the invention may involve in step (a) determination of the expression level of at least four biomarker proteins. In some particular and non-limiting embodiments of the invention, such at least four of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise RPS24, LSM4, RBM12B and RPS29. It should be appreciated that in some embodiments, the four biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten or at least eleven of the RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention.
Still further, the method of the invention may involve in step (a) determination of the expression level of at least six biomarker proteins. In some particular and non-limiting embodiments of the invention, such at least six of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise RPS24, LSM4, RBM12B, RPS29, RBM3 and PNP. It should be appreciated that in some embodiments, the six biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight or at least nine of the METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention.
In yet some further embodiments, the method of the invention may involve in step (a) determination of the expression level of at least ten biomarker proteins. In some particular and non-limiting embodiments of the invention, such at least ten of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3 and SNX12. It should be appreciated that in some embodiments, the ten biomarker proteins may further comprise at least one, at least two, at least three, at least four, or at least five of the SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention.
In certain embodiments, the method of the invention may provide and use detecting molecules specific for at least one, at least five, at least four, at least six or at least ten of the biomarkers of Table 4 and further, detecting molecule/s specific for at least one additional biomarker protein. It should be noted that each detecting molecule is specific for one biomarker. In some embodiments, the method as well as the kits of the invention described herein after may provide and use further detecting molecules specific for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more, specifically, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, 300, 350, 400, 450 and 500 at the most, additional biomarker proteins. In some embodiments, the methods, compositions and kits of the invention may provide and use in addition to detecting molecules specific for at least one of the biomarkers disclosed in Table 4, also detecting molecule/s specific for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 or 85 biomarker proteins disclosed in Table 2, and optionally, further detecting molecule/s specific for additional at least one biomarker protein/s, for example, any of the biomarkers presented in Table 3, or any other biomarker. In some embodiments, the methods, as well as the compositions and kits of the invention may provide and use detecting molecules specific for at least one additional biomarker protein and at most, 499 additional marker protein/s. In some specific embodiments, the methods and kit/s of the invention may provide and use detecting molecules specific for at least one of the biomarker proteins of Table 4, and detecting molecules specific for at least one additional biomarkers, provided that detecting molecules specific for 100, 150, 200, 250, 300, 350, 384, 400, 450 and 500 at the most biomarker proteins are used. In still further specific and non-limiting embodiments, the at least one additional biomarker protein may comprise any of the biomarker proteins presented in Figure 12B. In more specific embodiments, such additional biomarker proteins may comprise at least one of EDF1, PLEC, SNRPG, SRP9, RPL8, RPL23, RPL35A, RPS 15A, RPS23 and RPS28.
In yet some further embodiments, it should be understood that the methods of the invention as well as the compositions and kits described herein after may involve the determination of the expression levels of the biomarker proteins of the invention and/or the use of detecting molecules specific for said biomarker proteins. Specifically, at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen of the biomarker protein/s of the invention that may further comprise any additional biomarker proteins or control reference protein provided that 500 at the most biomarker proteins and control reference proteins are used. In some embodiments, the at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen of the biomarker protein/s of the invention may form at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or 100% of the biomarker proteins determined by the methods of the invention. In yet some further embodiments, the detecting molecules specific for at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten at least eleven, at least twelve, at least thirteen, at least fourteen or at least fifteen of the biomarker protein/s of the invention, that are used by the methods of the invention and comprised within any of the compositions and kits of the invention may form at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or 100% of detecting molecules used in accordance with the invention. It should be appreciated that for each of the selected biomarker proteins at least one detecting molecules may be used. In case more than one detecting molecule is used for a certain biomarker protein, such detecting molecules may be either identible or different.
As described herein below, the proteins that were found to be down-regulated in primary tumor samples of patients suffering from lymph node negative breast tumors (or up-regulated in primary tumor samples of patients suffering from lymph node positive breast tumors), represent several cellular functions such as ribosomal proteins (RPS24 and RPS29), proteins involved in pre-mRNA splicing (LSM2 and LSM4) and RNA binding proteins (RBM3 and RBM12B), and also include the two possible invasiveness markers METAP2 and PNP.
Without being bound by theory, it was suggested by the inventors that the protein signature described above may predict the involvement of lymph node in cancer progression, based on expression levels in the primary tumors. The inventors envision that this signature could be implemented into clinical applications in the future, to determine the aggressiveness of the tumor already by the initial biopsy, thereby potentially obliterating the need for intrusive surgical disruption of the lymph nodes, while also preventing unnecessary systemic treatment.
The term "cancer" is used herein interchangeably with the term "tumor" and denotes a mass of tissue found in or on the body that is made up of abnormal cells. As used herein, the term "breast cancer" refers to a cancer that develops from breast tissue. Development of breast cancer is often associated with a lump in the breast, a change in breast shape, dimpling of the skin, fluid coming from the nipple, or a red scaly patch of skin.
Breast cancer classification divides breast cancer into categories according to different schemes, each based on different criteria and serving a different purpose. The major categories are the histopathological type, the grade of the tumor, the stage of the tumor, and the expression of proteins and genes.
The purpose of classification among others is to select the appropriate treatment regimen. For example, breast cancers that tend to be aggressive and life-threatening should be treated with aggressive treatments that have major adverse effects. Other breast cancers which are less aggressive can be treated with less aggressive treatments.
Breast cancers can be classified by criteria, each one influences treatment response and prognosis.
Classification includes at least one of the following parameters histopathological type, grade, stage
(TNM), receptor status, and the presence or absence of certain receptors and markers:
Staging of breast cancer may be done by various methods for example using TNM staging which takes into account the size of the tumor (T), whether the cancer has spread to the lymph glands
(lymph nodes) (N), and whether the tumor has spread anywhere else in the body (M - for metastases).
Alternatively, staging can be expressed as a number on a scale of 0 through IV— with stage 0 describing non-invasive cancers that remain within their original location and stage IV describing invasive cancers that have spread outside the breast to other parts of the body.
In some embodiments, the breast tumor is a non-invasive tumor. In some other embodiments, the breast tumor is an invasive tumor. When referring to "non-invasive" cancer it should be noted as a cancer that do not grow into or invade normal tissues within or beyond the primary location, for example the breast. Non-invasive cancers are sometimes called carcinoma in situ ("in the same place") or pre-cancers. In connection with breast cancer, non invasive cancer stays in milk ducts or lobules in the breast. When referring to "invasive cancers" it should be noted as caner that invade and grow in normal, healthy tissues to form metastasis.
As used herein the term "metastatic cancer" or "metastatic status" refers to a cancer that has spread from the place where it first started to another place in the body and specifically to the lymph node. Such a tumor formed by metastatic cancer cells is called a metastatic tumor or a metastasis.
As used herein the term lymph node negative (LAW) refers to a primary non-invasive breast tumor that remain within the breast. The term lymph node positive (LNP) refers to a primary invasive breast tumor that has spread outside the breast into the lymph node.
It should be understood that characterization of a breast tumor as non-invasive or invasive may depend on information collected from different methods and depends on the detection capability of each one of the methods. Therefore, when referring to LNN or LNP it should be understood as detection level of the method used. Receptor status can also be used for classification of breast cancer into several molecular classes. The three most important receptors in the classification being: estrogen receptor (ER), progesterone receptor (PR), and HER2/neu. Breast cells characterized by being ER+ (cells expressing ER) and low grade are denoted Luminal A. Breast cells characterized by being ER+ and but often high grade are denoted Luminal B.
In accordance with some embodiments, the diagnosed subject may be a subject suffering from a luminal A breast tumor or a luminal B breast tumor.
As shown herein the method of the invention may be used as a diagnostic and prognostic tool by detecting the expression values of at least one of the marker proteins described herein. Particularly, determining the expression values of at least one of marker proteins described herein may differentiate metastatic breast cancer from non-metastatic breast cancer, namely at an early stage when a subject has a primary tumor, it may be possible to predict or prognose if a breast tumor will be LNN or LNP. In other words, determining the expression values of at least one of the following biomarker proteins may indicate if a subject belongs to a pre-established population associated with negative lymph node metastatic status or positive lymph node metastatic status.
From treatment decision making point of view, the lymph node status of cancer patients and specifically breast cancer is considered to be most important variable in the management of the disease (Jatoi et al., 1999).
In some other embodiments, in addition to determining the expression level of at least one of the fifteen biomarker proteins of the invention indicated above, the methods of the invention may further comprise determining the expression level of at least one other biomarker protein in at least one biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein/s. In more specific embodiments, such biomarker proteins may be at least one of Acyl-coenzyme A thioesterase 1 ; Acyl-coenzyme A thioesterase 2, mitochondrial (ACOT1; ACOT2), Small nuclear ribonucleoprotein G;Small nuclear ribonucleoprotein G-like protein (SNRPG), 40S ribosomal protein S27-like;40S ribosomal protein S27 (RPS27L), Nitrilase homolog 1 (NIT1), Protein RPS 10-NUDT3 (RPS 10-NUDT3), Protein PRRC1 (PRRC1), 39S ribosomal protein L43, mitochondrial (MRPL43), Ras suppressor protein 1 (RSU1), Nucleolysin TIAR (TIAL1), 28S ribosomal protein S 17, mitochondrial (MRPS 17), Replication protein A 32 kDa subunit (RPA2), Cytochrome c oxidase subunit 6B 1 (COX6B 1), Transformer-2 protein homolog alpha (TRA2A;HSU53209), Partner of Y14 and mago (WIBG), 40S ribosomal protein S23 (RPS23), l,2-dihydroxy-3-keto-5-methylthiopentene dioxygenase (ADI1), Mitochondrial import inner membrane translocase subunit Timl3 (TIMM13), 60S ribosomal protein L36a (RPL36A), Mitochondrial import inner membrane translocase subunit Tim23;Putative mitochondrial import inner membrane translocase subunit Tim23B (TIMM23;TIMM23B), 28S ribosomal protein S34, mitochondrial (MRPS34), Alpha-endosulfine (ENSA), SRA stem-loop-interacting RNA-binding protein, mitochondrial (SLIRP), Ribonuclease P protein subunit p30 (RPP30), THO complex subunit 4 (ALYREF), 60S ribosomal protein L37a (RPL37A), Selenide, water dikinase 1 (SEPHS 1), SWI/SNF-related matrix-associated actin-dependent regulator of chromatin subfamily D member 2 (SMARCD2), 60S ribosomal protein L23 (RPL23), Tubulin beta-8 chain (TUBB8), Proteasome subunit beta type-4 (PSMB4), Vesicle-associated membrane protein 8 (VAMP8), Signal recognition particle 9 kDa protein (SRP9), Multiple myeloma tumor-associated protein 2 (MMTAG2), PHD finger-like domain-containing protein 5A (PHF5A), Pre-mRN A- splicing factor SPF27 (BCAS2), 60S ribosomal protein L27 (RPL27), ER membrane protein complex subunit 3 (EMC3), 40S ribosomal protein S21 (RPS21), Cystatin-B (CSTB), 40S ribosomal protein S 14 (RPS 14), 40S ribosomal protein S27 (RPS27), 60S ribosomal protein L35a (RPL35A), Endothelial differentiation-related factor 1 (EDF1), 40S ribosomal protein S28 (RPS28), 60S ribosomal protein L31 (RPL31), 60S ribosomal protein L19;Ribosomal protein L19 (RPL19), 40S ribosomal protein S 15a (RPS 15A;hCG_1994130), 40S ribosomal protein SA (RPSA;LAMR1P15), 10 kDa heat shock protein, mitochondrial (HSPE1;EPFP1), Cytochrome c oxidase subunit 4 isoform 1, mitochondrial (COX4I1), 40S ribosomal protein S8 (RPS8), 60S ribosomal protein L30 (RPL30), 60S ribosomal protein L8 (RPL8), Malate dehydrogenase, mitochondrial;Malate dehydrogenase (MDH2), Eukaryotic translation initiation factor 4H (EIF4H; WBSCR1), 60S ribosomal protein L12 (RPL12;hCG_21173), Ubiquitin-conjugating enzyme E2 variant 1 (UBE2V1), Guanine nucleotide- binding protein subunit beta-2-like 1 (GNB2L1), Ubiquitin-associated protein 2-like (UBAP2L), 40S ribosomal protein S2 (RPS2;rps2; OK/KNS-cl.7), Proteasome subunit alpha type-7; Proteasome subunit alpha type (PSMA7; hCG_41772), 60S ribosomal protein L18a (RPL18A), 40S ribosomal protein S4, X isoform (RPS4X), 60S acidic ribosomal protein PO; 60S acidic ribosomal protein PO-like (RPLP0;RPLP0P6), Translational activator GCNl (PRIC295;GCN1L1), Plectin (PLEC), Fascin (FSCN1), Ribosome biogenesis regulatory protein homolog (RRS 1), ATP-binding cassette sub-family D member 3 (ABCD3), Major prion protein (PRNP), Keratin, type I cytoskeletal 14 (KRT14), or any combination thereof. It should be noted that said additional and optional biomarker proteins are disclosed in Tables 2 and 3 herein after.
In more specific embodiments the method of the invention may involves the determination of the expression level of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 of the biomarker proteins of the invention, specifically, the proteins disclosed in Table 4, and optionally further at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 34, 35, 36, 37, 38, 39, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 or 85 of the biomarker proteins disclosed in Table 2. It should be appreciated that further biomarkers may be used, for example, any of the biomarkers presented in Table 3, or any other biomarkers. In some further embodiments, the method of the invention may involve determination of the expression level of additional biomarker protein/s, specifically, additional at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more, specifically, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, 300, 350, 400, 450 and at most, 500 additional biomarker proteins. It should be understood that in certain embodiments, the additional biomarker proteins may include any breast cancer-related proteins, such as estrogen receptor, progesterone receptor, ErbB2/Her2, EGF receptor, ki67, keratins, Mucin 1 (MUC1), Catapsin D and more. In yet some further embodiments further biomarker proteins may include at least one of ACOTl/2, SLC25A11, GLUL, POSTN, COL12A1, CDKN2A, LGALS 1, and any biomarker protein presented in the figures of the present invention, for example in Figure 9.
It should be further appreciated that in certain embodiments, in addition to the biomarker protein/s of the invention, the methods of the invention may involve determination of the expression level of at least one control reference protein/s in at least one sample of the diagnosed subject. Control reference protein/s will be described in more details herein after.
Back to the method of the invention, the second step (b) involves calculating and determining if the expression value obtained in step (a) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample.
More specifically, wherein at least one of: (i) a positive expression value of said at least one biomarker protein/s selected from CAPS, ACOTl;ACOT2, SNRPG, RPS27L, LSM2, NIT1, RPS 10-NUDT3, PRRC1, MRPL43, RBM12B, RSU1, TIAL1, MRPS 17, RBM3, BTF3L4, RPA2, CCOX6B 1, TRA2A;HSU53209, LSM4, WIBG, METAP2, RPS23, ADI1, TIMM13, RPL36A, TIMM23;TIMM23B, MRPS34, SNX12, SLIRP, RPP30, RPL37A, SEPHS 1, SMARCD2, 60S RPL23, TUBB8, PSMB4, VAMP8, SRP9, MMTAG2, PHD PHF5A, BCAS2, RPL27, EMC3, RPS21, CSTB, RPS 14, RPS27, RPS29, RPL35A, EDF1, 4RPS28, SRSF1, RPL31, 60S RPL19, 4RPS 15A;hCG_1994130, RAB5C, RPSA;LAMR1P15, HSPE1;EPFP1, PNP, COX4I1, 40S RPS8, RPL30, RPL8, MDH2, EIF4H;WBSCR1, RPL12;hCG_21173, UBE2V1, GNB2L1, UBAP2L, RPS2;rps2;OK/KNS-cl.7, PSMA7;hCG_41772, RPL18A, RPS4X, RPLP0;RPLP0P6, in said sample, indicates that said biological tissue is from a lymph node positive tumor; and (ii) a negative expression value of said at least one biomarker protein/s selected from Translational activator GCN1 (PRIC295;GCN1L1), Plectin (PLEC), Fascin (FSCN1), Ribosome biogenesis regulatory protein homolog (RRS 1), ATP-binding cassette sub-family D member 3 (ABCD3), Major prion protein (PRNP), Keratin, type I cytoskeletal 14 (KRT14) in said sample, indicates that the subject belongs to a population associated with metastatic status and therefore, may be diagnosed as having a lymph node positive tumor.
In further embodiments, the methods of the invention may further comprise determining that a subject classified as belonging to a population having an LNP tumor, will develop, or has an increased probability to develop metastasis to the lymph node/s.
As described herein, the inventors also determined the differences in the protein expression in a group of primary metastatic tumors and the corresponding lymph node metastases. As shown in Example 5, a set of ten proteins was found to be differently expressed in the primary tumor and the metastatic tissue.
Thus, in some other embodiments, the methods of the invention may be further used to differentiate between primary tumor and metastatic tissue.
In some further embodiments, the methods of the invention further comprise (a) determining the expression level of at least one biomarker protein in at least one biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein/s, wherein said biomarker proteins are selected from Periostin (POSTN), AP-2 complex subunit mu (AP2M1), Mimecan (OGN), Collagen alpha-l(XII) chain (COL12A1), Cyclin-dependent kinase inhibitor 2A, isoforms 1/2/3 (CDKN2A), Galectin-1 (LGALS 1), Synaptosomal-associated protein 23 (SNAP23), Adenine phosphoribosyltransferase (APRT) Protein-glutamine gamma-glutamyltransferase 2 (TGM2) and Ester hydrolase Cl lorf54 (Cl lorf54), or any combination thereof; and (b) determining if the expression value obtained in step (a) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample; wherein at least one of: (i) a negative expression value of said at least one biomarker protein/s selected from SNAP23, APRT, TGM2 and Cl lorf54 in said sample, indicates that said biological tissue is a primary tissue; and (ii) a positive expression value of said at least one biomarker protein/s selected from POSTN, AP2M1, OGN, COL12A1, CDKN2A, LGALS 1 in said sample, indicates that said biological tissue is a primary tissue.
As shown in Table 3, four proteins - POSTN, LGALS 1, CDKN2A and COL12A1 - were upregulated in primary tumors and downregulated in metastases. Both POSTN and LGALS 1 have been shown to promote invasiveness in ovarian carcinoma and lung cancer, respectively, by interacting with integrins and several collagens were reported to be associated with high expression levels in tumors in comparison to metastases. CDKN2A is an important inhibitor of cell cycle progression and in the present invention was shown to be over expressed in breast tumors.
Without being bound by theory, it was suggested by the inventors that high expression of POSTN and LGALS 1 in tumor may facilitate migration to lymph nodes. Further, it was suggested by the inventors that the invasive mechanisms of cells from the primary tumor may be 'shut off in the lymph node in favor of renewed proliferation, indicated by decreased expression of cell-cycle regulator CDKN2A in metastatic cells.
As described herein, the comparison between primary breast tumors and their matched lymph node metastases showed a set of ten proteins that were differently expressed. Thus, it was concluded by the inventors that there is a similarity between the protein signature of tumors and the protein signature of lymph node metastases, despite the different microenvironment. This suggests that the lymph node metastases retain the expression patterns of their original primary breast tumor. This result is in agreement with several gene expression studies that examined both lymph node metastases and distant metastases.
It should be noted that the present invention also provides disease diagnosis and specifically methods of cancer diagnosis. To this end, the inventors have analyzed protein expression in healthy tissue and tumor tissue.
In accordance with some embodiments, in the first step (a) of the method of the invention, the expression level of at least one of the biomarker proteins described herein is being determined. The terms "level of expression" or "expression level" are used interchangeably and generally refer to a numerical representation of the amount (quantity) of an amino acid product or polypeptide or protein in a biological sample. In some embodiments, the "level of expression" or "expression level" refers to the numerical representation of the amount (quantity) of polynucleotide which may be gene in a biological sample.
"Expression" generally refers to the process by which gene-encoded information is converted into the structures present and operating in the cell. For example, gene expression values may be measured in the protein level, for example by MS methods or alternatively by immunological methods. Alternatively, the expression may be measured in the nucleic acid level, for example using Real-Time Polymerase Chain Reaction, sometimes also referred to as RT-PCR or quantitative PCR (qPCR). The luminosity in case of RT-PCR, or any other tag is captured by a detector that converts the signal intensity into a numerical representation which is said expression value, in terms of biomarker protein or a gene. Therefore, according to the invention "expression" of a gene, specifically, any gene encoding any of the biomarker proteins of the invention may refer to transcription into a polynucleotide and translation into a polypeptide. Fragments of the transcribed polynucleotide, the translated protein, or the post-translationally modified protein shall also be regarded as expressed whether they originate from a transcript generated by alternative splicing or a degraded transcript, or from a post-translational processing of the protein, e.g., by proteolysis. Methods for determining the level of expression of the biomarkers of the invention will be described in more detail herein after. It should be appreciated that the methods of the invention, as well as the compositions and kits disclosed herein after, refer to the level of the biomarker protein/s in the sample. It should be understood that the level of the protein reflects the level of expression but may also reflect the stability of the biomarker protein.
In certain and specific embodiments, the method of the invention further comprises an additional and optional step of normalization. According to this embodiment, in addition to determination of the level of expression of the biomarkers of the invention, the level of expression of at least one suitable control reference protein is being determined in the same sample. It should be noted that a control reference protein may be any protein that is not differentially expressed in different tissues or different pathologic conditions. In more specific embodiments, appropriate control reference proteins in connection with the present invention are proteins that are expressed equally in primary tumors of patients diagnosed with LNN vs. LNP, and therefore cannot be used to distinguish between LNP and LNN patients based on their expression in primary tumors. Non-limiting examples for such control reference proteins may include ARCN1 (Archain 1), MPZL1 (Myelin Protein Zero-Like 1), NSF (N-ethylmaleimide-sensitive factor), PRKCD (Protein Kinase C, Delta), CAT (catalase), actin, tubulin, or other cytoskeletal proteins. According to such embodiment, the expression level of at least one of the biomarkers of the invention obtained in step (a) is normalized according to the expression level of said at least one reference control protein obtained in the additional optional step in said test sample, thereby obtaining a normalized expression value. Optionally, similar normalization is performed also in at least one control sample or a representing standard when applicable.
The term "expression value" refers to the result of a calculation, that uses as an input the "level of expression" or "expression level" obtained experimentally and by normalizing the "level of expression" or "expression level" by at least one normalization step as detailed herein, the calculated value termed herein "expression value" is obtained.
More specifically, as used herein, "normalized values" are the quotient of raw expression values of marker proteins, divided by the expression value of a control reference protein from the same sample. Any assayed sample may contain more or less biological material than is intended, due to human error and equipment failures. Importantly, the same error or deviation applies to both the marker protein of the invention and to the control reference protein, whose expression is essentially constant. Thus, division of the marker protein raw expression value by the control reference protein raw expression value yields a quotient which is essentially free from any technical failures or inaccuracies (except for major errors which destroy the sample for testing purposes) and constitutes a normalized expression value of said marker protein. This normalized expression value may then be compared with normalized cutoff values, i.e., cutoff values calculated from normalized expression values. In certain embodiments, the control reference protein may be a protein that maintains stable in all samples analyzed.
Normalized biomarker protein expression level values that are higher (positive) or lower (negative) in comparison with a corresponding predetermined standard expression value or a cut-off value in a control sample predict to which population of patients the tested sample belongs or more specifically the disease stage, or the metastatic status of the subject.
It should be appreciated that an important step in the method of the inventions is determining whether the normalized expression value of any one of the biomarker proteins is changed compared to a pre-determined cut off, or is within the range of expression of such cutoff.
More specifically, as noted above, after determining the expression values of biomarker proteins of the invention, the next step of the method of the invention involves calculating and determining if the expression value obtained in step (a) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample. Such step involves calculating and measuring the difference between the expression values of the examined sample and the cutoff value pre-determined for a certain population and determining whether the examined sample can be defined as positive or negative, with respect to said population.
In yet more specific embodiments, the second step (b) of the method of the invention involves comparing the expression values determined for the tested sample with predetermined standard values or cutoff values, or alternatively, with expression values of at least one control sample. As used herein the term "comparing" denotes any examination of the expression level and/or expression values obtained in the samples of the invention as detailed throughout in order to discover similarities or differences between at least two different samples. It should be noted that in some embodiments, comparing according to the present invention encompasses the possibility to use a computer based approach.
As described hereinabove, the method of the invention refers to a predetermined cutoff value/s. It should be noted that a "cutoff value", sometimes referred to simply as "cutoff herein, is a value that meets the requirements for both high diagnostic sensitivity (true positive rate) and high diagnostic specificity (true negative rate).
It should be noted that the terms "sensitivity" and "specificity" are used herein with respect to the ability of one or more markers, to correctly classify a sample as belonging to a pre-established population associated with negative lymph node metastatic status, or alternatively, to a pre- established population associated with positive lymph node metastatic status.
"Sensitivity" indicates the performance of the bio-marker of the invention, with respect to correctly classifying samples as belonging to pre-established populations that are likely to suffer from a disease or disorder or characterized at different stages of a disease to respond to therapy or to relapse, when applicable, wherein said bio-marker are consider here as any of the options provided herein.
"Specificity" indicates the performance of the bio-marker of the invention with respect to correctly classifying samples as belonging to pre-established populations of subjects suffering from the same disorder or populations of subjects that are likely to respond to a specific treatment or unlikely to relapse as will be discussed herein after.
Simply put, "sensitivity" relates to the rate of correct identification of the patients (samples) as such out of a group of samples, whereas "specificity" relates to the rate of correct identification of lymph node metastatic status samples as such out of a group of samples. Cutoff values may be used as control sample/s or in addition to control sample/s, said cutoff values being the result of a statistical analysis of biomarker protein expression value/s (specifically the biomarker proteins of the invention) differences in pre-established populations healthy, metastatic, LNP or LNN.
Thus, a given population having specific clinical parameters will have a defined likelihood to have positive lymph node metastasis or negative lymph node metastasis based on the expression values of the marker proteins being above or below said cutoff values.
For example, an individual having a positive expression value, or in other words, being up- regulated of least one of the following biomarker protein/s 40S ribosomal protein S24, U6 snRNA- associated Sm-like protein LSm4, RNA-binding protein 12B, 40S ribosomal protein S29; Putative RNA-binding protein 3, Purine nucleoside phosphorylase, Methionine aminopeptidase 2;Methionine aminopeptidase, Calcyphosin, Eukaryotic initiation factor 4A-III;Eukaryotic initiation factor 4A-III, N-terminally processed, Sorting nexin-12, (erine/arginine-rich splicing factor 1, Ras- related protein Rab-5C, Transcription factor BTF3 homolog 4, U6 snRNA-associated Sm-like protein LSm2 and Cystatin-B may be considered as belonging to a pre-established population associated with positive lymph node metastatic status. In yet another example, a subject presenting a negative expression value, that reflects down- reulation of at least one biomarker protein/s 40S ribosomal protein S24, U6 snRNA-associated Sm- like protein LSm4, RNA-binding protein 12B, 40S ribosomal protein S29; Putative RNA-binding protein 3, Purine nucleoside phosphorylase, Methionine aminopeptidase 2;Methionine aminopeptidase, Calcyphosin, Eukaryotic initiation factor 4A-III;Eukaryotic initiation factor 4A-III, N-terminally processed, Sorting nexin-12, (erine/arginine-rich splicing factor 1, Ras-related protein Rab-5C, Transcription factor BTF3 homolog 4, U6 snRNA-associated Sm-like protein LSm2 and Cystatin-B may be considered as belonging to a pre-established population associated with negative lymph node metastatic status.
In yet some further embodiments, a negative or positive determination of the expression value as compared to the predetermined cutoff values, also encompass values that are within the range of said cutoff. More specifically, an expression value that is determined by the method of the invention as "positive" when compared to a predetermined cutoff of population of LNP patients, or for at least one known LNP patient, may indicate that the examined subject belongs to LNP population, in case that the expression value is either higher (positive) or within the range (the average values of the cutoff predetermined for LNP patient population). In a similar manner, a subject exhibiting an expression value that is "negative" (that is down-regulated) as compared to the cutoff patients, may be considered as belonging to LNN population. In more specific embodiments, the expression value of such subject should fall within the range of the cutoff value predetermined for LNN population. In some embodiments, "fall within the range" encompass values that differ from the cutoff value in about 1% to about 50% or more.
It should be emphasized that the nature of the invention is such that the accumulation of further patient data may improve the accuracy of the presently provided cutoff values, which are based on an ROC (Receiver Operating Characteristic) curve generated according to said patient data using analytical software program. The biomarker protein expression values are selected along the ROC curve for optimal combination of prognostic sensitivity and prognostic specificity which are as close to 100 percent as possible, and the resulting values are used as the cutoff values that distinguish between patients who are diagnosed with positive lymph node metastasis at a certain rate, and those who will not (with said given sensitivity and specificity). Similar analysis may be performed for example when diagnosis of cancer is being examined to distingue between healthy tissue and cancerous tissue or when responsiveness to treatment is being examined to distinguish between responsive and non-responsive subjects. The ROC curve may evolve as more and more data and related biomarker gene expression values are recorded and taken into consideration, modifying the optimal cutoff values and improving sensitivity and specificity. Thus, the provided cutoff values should be viewed as a starting point that may shift as more data allows more accurate cutoff value calculation. Although considered as initial cutoff values, the presently provided values already provide good sensitivity and specificity, and are readily applicable in current clinical use, even in patients diagnosed with different cancer stages.
As noted above, the expression value determined for the examined sample (or alternatively, the normalized expression value) is compared with a predetermined cutoff or a control sample. More specifically, in certain embodiments, the expression value obtained for the examined sample is compared with a predetermined standard or cutoff value.
In further embodiments, the predetermined standard expression value, or cutoff value has been predetermined and calculated for a population comprising at least one of healthy subjects, subjects suffering from any disorder, subjects suffering from different stages of any disorder, subjects that respond to treatment, non-responder subjects, subjects in remission and subjects in relapse. In more specific embodiments, predetermined cutoff values may be calculated for a population of subject diagnosed with breast cancer, subjects diagnosed with metastatic breast cancer, subjects diagnosed with LNN and subjects diagnosed with LNP.
Still further, in certain alternative embodiments where a control sample is being used (instead of, or in addition to, pre-determined cutoff values), the normalized expression values of the biomarker proteins used by the invention in the test sample are compared to the expression values in the control sample. In certain embodiments, such control sample may be obtained from at least one of a healthy subject, a subject suffering from a disorder at a specific stage, a subject suffering from a disorder at a different specific stage a subject that responds to treatment, a non-responder subject, a subject in remission and a subject in relapse. In more specific embodiments, predetermined cutoff values may be calculated for a population of subject diagnosed with breast cancer, subjects diagnosed with metastatic breast cancer, subjects diagnosed with LNN and subjects diagnosed with LNP.
It should be appreciated that "Standard" or a "predetermined standard" as used herein, denotes either a single standard value or a plurality of standards with which the level at least one of the biomarker protein expression from the tested sample is compared. The standards may be provided, for example, in the form of discrete numeric values or is calorimetric in the form of a chart with different colors or shadings for different levels of expression; or they may be provided in the form of a comparative curve prepared on the basis of such standards (standard curve).
In some particular embodiments, determining the level of expression of at least one of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins is performed by the step of contacting at least one detecting molecule or any combination or mixture of plurality of detecting molecules with a biological sample of said subject, or with any protein or nucleic acid product obtained therefrom, wherein each of said detecting molecules is specific for one of said biomarker proteins.
It should be noted that for determining the expression value/s of at least one of the biomarker proteins of the invention, the methods of the invention may further comprise the step of providing at least one detecting molecule specific for determining the expression of at least on of said biomarker proteins of the invention. In some embodiments, such detecting molecules may be provided as a mixture, as a composition or as a kit. Thus, in some embodiments, the at least one detecting molecules may be provided as a mixture of detecting molecules, wherein each detecting molecule is specific for one biomarker protein. It should be appreciated however, that for each biomarker protein, one or several specific detecting molecules may be used and provided. In yet some further alternative embodiments, the detecting molecules may be provided separately for each biomarker protein, e.g., in specific tube, containers, slots, spots, wells, and the like. It further alternative embodiments, the detecting molecules may be attached or immobilized to a solid support, specifically, in recorded location.
Still further, it should be noted that all steps for determining the different parameters indicated above, involve contacting the sample or any component thereof with a specific reagent (e.g., detecting molecules). In some specific embodiments, determining the level of expression of at least one of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins is performed by the step of contacting at least one detecting molecule or any combination or mixture of plurality of detecting molecules with a biological sample of said subject, or with any protein or nucleic acid product obtained therefrom, wherein each of said detecting molecules is specific for one of said biomarker proteins. The term "contacting" mean to bring, put, incubates or mix together. As such, a first item is contacted with a second item when the two items are brought or put together, e.g., by touching them to each other or combining them. In the context of the present invention, the term "contacting" includes all measures or steps which allow interaction between the at least one of the detection molecules of at least one of the biomarker proteins, and optionally, for at least one suitable control reference protein of the tested sample. The contacting is performed in a manner so that the at least one of detecting molecule of at least one of the biomarker proteins for example, can interact with or bind to the at least one of the biomarker proteins, in the tested sample. The binding will preferably be non-covalent, reversible binding, e.g., binding via salt bridges, hydrogen bonds, hydrophobic interactions or a combination thereof. In certain embodiments, the detection step further involves detecting a signal from the detecting molecules that correlates with the expression level of at least one of the biomarker proteins and in the sample from the subject, by a suitable means. According to some embodiments, the signal detected from the sample by any one of the experimental methods detailed herein below reflects the expression level of at least one of the biomarker proteins. It should be noted that such signal-to- expression level data may be calculated and derived from a calibration curve.
Thus, in certain embodiments, the method of the invention may optionally further involve the use of a calibration curve created by detecting a signal for each one of increasing pre-determined concentrations of at least one of the biomarker proteins. Obtaining such a calibration curve may be indicative to evaluate the range at which the expression levels correlate linearly with the concentrations of at least one of the biomarker proteins. It should be noted in this connection that at times when no change in expression level of at least one of the biomarker proteins is observed, the calibration curve should be evaluated in order to rule out the possibility that the measured expression level is not exhibiting a saturation type curve, namely a range at which increasing concentrations exhibit the same signal.
It must be appreciated that in certain embodiments such calibration curve as described above may by also part or component in any of the kits provided by the invention as described herein after. In other embodiments of the invention, the detecting molecules used for determining the expression levels at least one of the biomarker proteins are selected from isolated detecting amino acid molecules and isolated detecting nucleic acid molecules. It should be noted that the invention further encompasses any combination of nucleic and amino acids for use as detecting molecules for the methods of the invention. As noted above, in the first step of the method of the invention, the sample or any protein or nucleic acid obtained therefrom, is contacted with the detecting molecules of the invention.
The invention thus contemplates the use of amino acid based molecules such as proteins or polypeptides as detecting molecules disclosed herein and would be known by a person skilled in the art to measure the at least one biomarker protein. As used herein, the terms "protein" and "polypeptide" are used interchangeably to refer to a chain of amino acids linked together by peptide bonds. In a specific embodiment, a protein is composed of less than 200, less than 175, less than 150, less than 125, less than 100, less than 50, less than 45, less than 40, less than 35, less than 30, less than 25, less than 20, less than 15, less than 10, or less than 5 amino acids linked together by peptide bonds. In another embodiment, a protein is composed of at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500 or more amino acids linked together by peptide bonds. It should be noted that peptide bond as described herein is a covalent amid bond formed between two amino acid residues. In some embodiments, the detecting molecules used by the methods of the invention may be recombinantly expressed or synthetically prepared. In further embodiments, the recombinantly or synthetically expressed and prepared detecting molecules may be labeled or tagged. It should be noted that in some embodiments, these detecting molecules may be isolated detecting molecules. As used herein, "Recombinant proteins" denotes proteins encoded by a recombinant DNA which is a genetically engineered DNA formed by laboratory methods of genetic recombination to bring together genetic material from multiple sources and thus creating variable sequences. Recombinant proteins may be produced mainly, but not limited, by molecular cloning, namely incorporating the recombinant DNA into a living cell (e.g. bacteria or yeast) and using its system to express the DNA into mRNA and protein thereof.
Techniques for detection and quantification known to persons skilled in the art (for example, Mass spectrometry (MS) or different immunological techniques such as Western Blotting, Immunoprecipitation, ELISAs, protein microarray analysis, Flow cytometry and the like) can then be used to measure the level of protein products corresponding to the biomarker of the invention. In some embodiments, the amino acid-based detecting molecules may comprise at least one of: (a) at least one labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof; optionally, such labeled proteins may be recombinant or synthetically produced proteins, (b) antibodies specific for said at least one of said biomarker proteins; (c) peptide aptamers specific for said at least one of said biomarker proteins; and (d) any combination of (a), (b) and (c).
More specifically, in some embodiments, the detecting molecules may be at least one labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragments, peptides or mixture thereof. In some embodiments, the term "labeled" or "tagged" may refer to direct labeling of the protein via, e.g., coupling (i.e., physically linking) or incorporating of a detectable substance to the protein. Useful labels in the present invention may include but are not limited to include isotopes (e.g. 13C, 15N), or any other radiolabels (e.g., 3H, 125I, 35S, 14C, or 32P), magnetic beads (e.g. DYNABEADS), fluorescent dyes (e.g., fluorescein isothiocyanate, Texas red, rhodamine, green fluorescent protein, and the like), enzymes (e.g., horseradish peroxidase, alkaline phosphatase and others commonly used in an ELISA and competitive ELISA, histochemistry and other similar methods known in the art) and colorimetric labels such as colloidal gold or colored glass or plastic (e.g. polystyrene, polypropylene, latex, etc.) beads. In some embodiments, the protein may be tagged. Different tags may be also used, for example, His, myc, HA, GFP, ABP, GST, biotin and the like, "tagged" as used herein may further include fusion or linking of the biomarker protein or any fragment or peptide thereof, that serves herein as a detecting molecule, a tag that in some embodiments may contain several amino acids or a peptide that may be recognized by affinity or immunologically, using specific antibodies.
In some embodiments, the detecting molecules may be at least one, optionally, recombinant and/or isolated, labeled or tagged biomarker protein that may be any one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof. It should be noted that as indicated above, the invention encompasses the use of any biomarker protein, specifically, the biomarkers disclosed in the present invention, as detecting molecule/s. In more specific embodiments, the invention encompasses the use of any of the biomarker proteins of Table 4, as well as any of the biomarker proteins of Tables 2 and 3 as detecting molecule/s as described herein. In some other embodiments, the biomarker proteins or any fragments or peptides thereof may be fluorescently labeled. In another embodiment, the biomarker proteins or any fragments or peptides thereof may be isotope labeled. The term "recombinant isotope labeled" denotes a protein 'labeled' by replacing specific atoms by their isotope.
Means of detecting such labels are well known to those of skill in the art. Thus, for example, radiolabels may be detected using photographic film or scintillation counters, fluorescent markers may be detected using a photodetector to detect emitted illumination. Enzymatic labels are typically detected by providing the enzyme with a substrate and detecting the reaction product produced by the action of the enzyme on the substrate, and colorimetric labels are detected by simply visualizing the colored label.
More specifically, in certain embodiments the biomarker proteins of the invention or any fragment or peptide thereof, when recombinantly expressed and labeled or tagged, may be used as detecting molecules for determining the quantity or level of expression of the biomarker proteins of the invention in the examined sample. The term "labeled form" as used herein includes an isotope labeled form. Specifically, the labeled form is a chemically or metabolically isotope labeled, and more specifically a metabolically isotope labeled form of the biomarker proteins of the invention. Optional "isotope labeled forms" of the biomarker protein/s or any fragments or peptides thereof in accordance with the present invention are variants of naturally occurring molecules, in whose structure one or more atoms have been substituted with atom(s) of the same element having a different atomic weight, although isotope labeled forms in which the isotope has been covalently linked either directly or via a linker, or wherein the isotope has been complexed to the biomarker proteins are likewise contemplated. In either case, the isotope may be stable isotope. A stable isotope as referred to herein, is a non-radioactive isotopic form of an element having identical numbers of protons and electrons, but having one or more additional neutron(s), which increase(s) the molecular weight of the element. Specifically, the stable isotopes may be selected from the group consisting of 2 H, 13 C, 15 N, 170, 180, 33 P, 34 S and combinations thereof. Particularly specific examples include 13 C and 5 N, and combinations thereof.
The labeling can be effected by means known in the art. A labeled reference biomarker (used as detecting molecule) can be synthesized using isotope labeled amino acids as precursor molecules, or chemically modified. Modification and labeling can be done on whole proteins or their fragments. For example, isotope-coded affinity tag (ICAT) reagents label reference biomolecule such as proteins at the alkylation step of sample preparation (WO2004079370). Visible ICAT reagents (VIC AT reagents) may be likewise employed (WO2011042467), whereby the VICAT-type reagent contains as a detectable moiety a fluorophore or radiolabel. iTRAQ and similar methods may likewise be employed.
Metabolic labeling may also be used to produce the labeled reference biomarkers. For example, cells can be grown on media containing isotope labeled precursor molecules, such as isotope labeled amino acids, that are incorporated into proteins or peptides, which are thereby metabolically labeled. The metabolic isotope labeling may be a stable isotope labeling with amino acids in cell culture (SILAC). If metabolic labeling is used, and the labeled form of the one or the plurality of reference biomarker protein/s is a SILAC labeled form of the reference biomarker protein/s, the standard mixture as defined above is also referred to as SUPER-SILAC mix.
In specific embodiments, the detecting amino acid molecules applicable for the invention may be isolated antibodies, with specific binding selectively to at least one of said biomarker proteins. More specifically, antibodies that specifically bind at least one of the biomarker proteins of the invention as listed in Table 4, and optionally, at least one of the biomarker proteins listed in Tables 2 and 3. It should be understood that each antibody specifically recognizes one biomarker protein. Using these antibodies, the level of expression of at least one of the biomarker protein may be determined using an immunoassay which may be an assay that includes but not limited to FACS, a Western blot, an ELISA, a RIA, a slot blot, a dot blot, immune-histochemical assay and a radio-imaging assay. It should be noted that such assay may be performed using microarray protein arrays.
More specifically, he term "antibody" as used in this invention includes whole antibody molecules as well as functional fragments thereof, such as Fab, F(ab')2, and Fv that are capable of binding with antigenic portions of the target polypeptide, i.e. at least one of the biomarker protein. The antibody may be preferably monospecific, e.g., a monoclonal antibody, or antigen-binding fragment thereof. The term "monospecific antibody" refers to an antibody that displays a single binding specificity and affinity for a particular target, e.g., epitope. This term includes a "monoclonal antibody" or "monoclonal antibody composition", which as used herein refer to a preparation of antibodies or fragments thereof of single molecular composition.
It should be recognized that the antibody can be a human antibody, a chimeric antibody, a recombinant antibody, a humanized antibody, a monoclonal antibody, or a polyclonal antibody. The antibody can be an intact immuno globulin, e.g., an IgA, IgG, IgE, IgD, lgM or subtypes thereof. The antibody can be conjugated to a labeling moiety as discussed above.
As noted above, the term "antibody" also encompasses antigen-binding fragments of an antibody. The term "antigen -binding fragment" of an antibody (or simply "antibody portion," or "fragment"), as used herein, may be defined as follows:
(1) Fab, the fragment which contains a monovalent antigen-binding fragment of an antibody molecule, can be produced by digestion of whole antibody with the enzyme papain to yield an intact light chain and a portion of one heavy chain;
(2) Fab', the fragment of an antibody molecule that can be obtained by treating whole antibody with pepsin, followed by reduction, to yield an intact light chain and a portion of the heavy chain; two Fab' fragments are obtained per antibody molecule;
(3) (Fab')2, the fragment of the antibody that can be obtained by treating whole antibody with the enzyme pepsin without subsequent reduction; F(ab')2 is a dimer of two Fab' fragments held together by two disulfide bonds;
(4) Fv, defined as a genetically engineered fragment containing the variable region of the light chain and the variable region of the heavy chain expressed as two chains; and
(5) Single chain antibody ("SCA", or ScFv), a genetically engineered molecule containing the variable region of the light chain and the variable region of the heavy chain, linked by a suitable polypeptide linker as a genetically fused single chain molecule.
Methods of generating such antibody fragments are well known in the art (See for example, Harlow and Lane, Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, New York, 1988, incorporated herein by reference).
Purification of serum immunoglobulin antibodies (polyclonal antisera) or reactive portions thereof can be accomplished by a variety of methods known to those of skill in the art including, precipitation by ammonium sulfate or sodium sulfate followed by dialysis against saline, ion exchange chromatography, affinity or immuno-affinity chromatography as well as gel filtration, zone electrophoresis, etc.
Still further, the antibodies used by the present invention may optionally be covalently or non- covalently linked to a detectable label or tag. In addition, the label and can also refer to indirect labeling of the protein by reactivity with another reagent that is directly labeled. Examples of indirect labeling include detection of at least one of the biomarker protein/s of the invention using a fluorescently labeled secondary antibody. More specifically, detectable labels suitable for such use include any composition detectable by spectroscopic, photochemical, biochemical, immunochemical, electrical, optical or chemical means.
The antibody used as a detecting molecule according to the invention, specifically recognizes and binds at least one of the biomarker protein. It should be noted that in certain embodiments, each antibody is specific for one of the biomarker proteins of the invention, specifically, those disclosed in Table 4, and optionally, those disclosed in Tables 2 and 3. It should be appreciated that antibodies that may be used by the methods as well as the compositions and kits of the invention, may be antibodies directed not only against the biomarker proteins of the invention, but also in case the biomarkers are tagged, the antibodies may be directed against said tags. It should be therefore noted that the term "binding specificity", "specifically binds to an antigen", "specifically immuno- reactive with", "specifically directed against" or "specifically recognizes", when referring to an epitope, specifically, a recognized epitope within the at least one of the biomarker protein, refers to a binding reaction which is determinative of the presence of the epitope in a heterogeneous population of proteins and other biologies. More particularly, "selectively bind" in the context of proteins encompassed by the invention refers to the specific interaction of a any two of a peptide, a protein, a polypeptide an antibody, wherein the interaction preferentially occurs as between any two of a peptide, protein, polypeptide and antibody preferentially as compared with any other peptide, protein, polypeptide and antibody.
Thus, under designated immunoassay conditions, the specified antibodies bind to a particular epitope at least two times the background and more typically more than 10 to 100 times background. More specifically, "Selective binding", as the term is used herein, means that a molecule binds its specific binding partner with at least 2-fold greater affinity, and preferably at least 10-fold, 20-fold, 50-fold, 100-fold or higher affinity than it binds a non- specific molecule. It should be appreciated that the antibodies used by the methods of the invention, may be in some embodiments antibodies that are not naturally occurring antibodies. More specifically, the antibodies are not produced naturally in the body, and more specifically, it should be appreciated that production thereof involves immunological and recombinant techniques.
A variety of immunoassay formats may be used to select antibodies specifically immuno-reactive with a particular protein or carbohydrate. For example, solid-phase ELISA immunoassays are routinely used to select antibodies specifically immuno-reactive with a protein or carbohydrate. The term "epitope" is meant to refer to that portion of any molecule capable of being bound by an antibody which can also be recognized by that antibody. Epitopes or "antigenic determinants" usually consist of chemically active surface groupings of molecules such as amino acids or sugar side chains and have specific three dimensional structural characteristics as well as specific charge characteristics.
In some other embodiments, the detecting molecules are peptide aptamers specific for said at least one of said biomarker proteins. eptide aptamers^ as used herein refers to small peptides with a single variable loop region tied to a protein scaffold on both ends that binds to a specific molecular target (e.g. protein), and which are bind to their targets only with said variable loop region and usually with high specificity properties.
According to one embodiment, where amino acid-based detection molecules are used, the expression level of the at least one of the biomarker protein, in the tested sample can be determined using different methods known in the art, specifically method disclosed herein below as non- limiting examples.
In some embodiments, the detecting molecules may be at least one isolated, optionally recombinant or synthetic labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or fragments, peptides or mixture thereof, and optionally, at least one of the biomarkers of any one of the biomarker proteins disclosed in Tables 2 and 3. In such case the determination of the expression level of said at least one biomarker protein/s may be performed by mass spectrometry.
Mass spectrometry (MS) is used herein as an analytical chemistry technique to identify the amount and type of chemicals present in a sample by measuring the mass-to-charge ratio and abundance of gas-phase ions. A mass spectrum is a plot of the ion signal as a function of the mass-to-charge ratio. The spectra are used to determine the elemental or isotopic signature of a sample, the masses of particles and of molecules, and to elucidate the chemical structures of molecules, such as peptides and other chemical compounds.
As noted above, the invention contemplates the use of Mass spectrometry-based absolute quantification assays that generally require recombinant expression of full length, labeled protein standards. Mass spectrometry is not inherently quantitative but many methods have been developed to overcome this limitation. Most of them are based on stable isotopes and introduce a mass shifted version of the peptides of interest, which are then quantified by their "heavy" to "light" ratio. Stable isotope labeling is either accomplished by chemical addition of labeled reagents, enzymatic isotope labeling, or metabolic labeling. Generally, these approaches are used to obtain relative quantitative information on protein expression levels in a light and a heavy labeled sample. For example, stable isotope labeling by amino acids in cell culture (SILAC) is performed by metabolic incorporation of light or heavy labeled amino acids into the recombinant or synthetic protein. Labeled protein can also be used as internal standards for determining expression levels of a cell or tissue protein of interest, such as in the spike-in SILAC approach.
Several methods for absolute quantification have emerged over the last years and may be applicable for the present invention, including absolute quantification (AQUA), quantification concatamer (QConCAT), protein standard absolute quantification (PSAQ), absolute SILAC, and FlexiQuant. They all quantify the endogenous protein of interest by the heavy to light ratios to a defined amount of the labeled counterpart spiked into the sample and are chiefly distinguished by either spiking in heavy labeled peptides or heavy labeled full length proteins. The AQUA strategy is convenient and streamlined: proteotypic peptides are chemically synthesized with heavy isotopes and spiked in after sample preparation.
Still further, the QconCAT approach is based on artificial proteins that are concatamers of proteotypic peptides. This artificial protein is recombinantly expressed in Escherichia coli and spiked into the sample before proteolysis. QconCAT in principle allows efficient production of labeled peptides but does not automatically correct for protein fractionation effects or digestion efficiency in the native proteins versus the concatamers. The PSAQ, absolute SILAC and FlexiQuant approaches sidestep these limitations by metabolically labeling full length proteins by heavy versions of the amino acids arginine and lysine. PSAQ and FlexiQuant in vitro synthesize full-length proteins in wheat germ extracts or in bacterial cell extract, respectively, whereas absolute SILAC was described with recombinant protein expression in E. coli. The protein standard is added at an early stage, such as directly to cell lysate. Consequently, sample fractionation can be performed in parallel and the SILAC protein is digested together with the proteome under investigation. Another quantitative approach applicable for the purpose of the present invention may be in some embodiments the SILAC-PrEST assay. In this method, Protein Epitope Signature Tags (PrESTs) are expressed recombinantly in E. coli and they consist of a short and unique region of the protein of interest as well as purification and solubility tags. A highly purified, stable isotope labeling of amino acids in cell culture (SILAC)-labeled version of the solubility tag is first quantified and used to determine the precise amount of each PrEST by its SILAC ratios. The PrESTs are then spiked into the examined sample (e.g., cell lysates) and the SILAC ratios of PrEST peptides to peptides from endogenous target proteins yield their cellular quantities.
In some embodiments, in the context of the present invention, the labeled or tagged biomarker/s of the invention or any labeled fragments or peptides thereof are mixed with the sample of with any protein extracted therefrom. The resulting protein mixture may be then digested according to the FASP protocol (Wisniewski et al., 2009b) and the peptides are separated into fractions by anion exchange chromatography in a StageTip format (Wisniewski al., 2009a). Each fraction is analyzed by online reverse-phase chromatography coupled to high resolution, quantitative mass spectrometry analysis.
A variety of mass spectrometry systems can be employed in the methods of the invention for identifying and/or quantifying a biomarker protein of the invention in a sample or any fragment or peptide thereof. Mass analyzers with high mass accuracy, high sensitivity and high resolution include, but are not limited to, matrix-assisted laser desorption time-of-flight (MALDI-TOF) mass spectrometers, electrospray ionization time-of-flight (ESI-TOF) mass spectrometers, Fourier transform ion cyclotron mass analyzers (FT-ICR-MS), and Orbitrap analyzer instruments. Other modes of MS include ion trap and triple quadrupole mass spectrometers. In ion trap MS, analytes are ionized by electrospray ionization or MALDI and then put into an ion trap. Trapped ions can then be separately analyzed by MS upon selective release from the ion trap. Ion traps can also be combined with the other types of mass spectrometers described above.
Fragments can also be generated and analyzed. Reference biomarker protein/s labeled with an ICAT or VICAT or iTRAQ type reagent, or SILAC labeled peptides can be analyzed, for example, by single stage mass spectrometry with a MALDI or ESI ionization and with TOF, quadrupole, iontrap, FT-ICR or Orbitrap analyzers.. Methods of mass spectrometry analysis are well known to those skilled in the art. For high resolution peptide fragment separation, liquid chromatography ESI- MS/MS or automated LC-MS/MS, can be used. MS analysis can be performed in a data-dependent manner or using targeted MS techniques such as selected reaction monitoring (SRM) or parallel reaction monitoring (PRM).
In some other embodiments, when the detecting molecules used are at least one of antibodies, nucleic acid, peptide aptamers or any combination thereof, specific for said at least one of said biomarker proteins, the determination of the expression level of said biomarker protein/s may be performed by an immunological assay.
In some specific embodiments, determination of the expression level of the biomarker may be performed using ELISA. Enzyme-Linked Immunosorbent Assay (ELISA) is used herein involves fixation of a sample containing a protein substrate (e.g., fixed cells or a protein solution) to a surface such as a well of a microtiter plate. A substrate- specific antibody coupled to an enzyme is applied and allowed to bind to the substrate. Presence of the antibody is then detected and quantitated by a colorimetric reaction employing the enzyme coupled to the antibody. Enzymes commonly employed in this method include horseradish peroxidase and alkaline phosphatase. If well calibrated and within the linear range of response, the amount of substrate present in the sample is proportional to the amount of color produced. A substrate standard is generally employed to improve quantitative accuracy.
In some specific embodiments, determination of the expression level of the biomarker may be performed using Western blot. Western Blot as used herein involves separation of a substrate from other protein by means of an acryl amide gel followed by transfer of the substrate to a membrane (e.g., nitrocellulose, nylon, or PVDF). Presence of the substrate is then detected by antibodies specific to the substrate, which are in turn detected by antibody -binding reagents. Antibody - binding reagents may be, for example, protein A or secondary antibodies. Antibody -binding reagents may be radio labeled or enzyme-linked, as described hereinafter. Detection may be by autoradiography, colorimetric reaction, or chemiluminescence. This method allows both quantization of an amount of substrate and determination of its identity by a relative position on the membrane indicative of the protein's migration distance in the acryl amide gel during electrophoresis, resulting from the size and other characteristics of the protein.
In some specific embodiments, different RIA assays may be employed for determination of the expression level of the biomarker proteins of the invention. In one version, Radioimmunoassay (RIA) involves precipitation of the desired protein (i.e., the substrate) with a specific antibody and radio labeled antibody -binding protein (e.g., protein A labeled with I 125 ) immobilized on a perceptible carrier such as agars beads. The radio-signal detected in the precipitated pellet is proportional to the amount of substrate bound.
In an alternate version of RIA, a labeled substrate and an unlabelled antibody-binding protein are employed. A sample containing an unknown amount of substrate is added in varying amounts. The number of radio counts from the labeled substrate-bound precipitated pellet is proportional to the amount of substrate in the added sample.
Still further, in specific embodiments, determination of the expression level of the biomarker may be performed using FACS. Fluorescence- Activated Cell Sorting (FACS) involves detection of a substrate in situ in cells bound by substrate-specific, fluorescently labeled antibodies. The substrate- specific antibodies are linked to fluorophore. Detection is by means of a flow cytometry machine, which reads the wavelength of light emitted from each cell as it passes through a light beam. This method may employ two or more antibodies simultaneously, and is a reliable and reproducible procedure used by the present invention.
In some specific embodiments, determination of the expression level of the biomarker may be performed using immunohistochemistry methods. Immuno histochemical Analysis involves detection of a substrate in situ in fixed cells by substrate-specific antibodies. The substrate specific antibodies may be enzyme-linked or linked to fluorophore. Detection is by microscopy, and is either subjective or by automatic evaluation. With enzyme-linked antibodies, a calorimetric reaction may be required. It will be appreciated that immunohistochemistry is often followed by counterstaining of the cell nuclei, using, for example, Hematoxyline or Giemsa stain.
It should be appreciated that all the detecting molecules used by any of the methods, as well as the compositions and kits of the invention described herein after, are isolated and/or purified molecules. As used herein, "isolated" or "purified" when used in reference to a protein means that a naturally occurring sequence has been removed from its normal cellular environment or is synthesized in a non-natural environment (e.g., artificially synthesized). Thus, an "isolated" or "purified" sequence may be in a cell-free solution or placed in a different cellular environment. The term "purified" does not imply that the sequence is the only nucleotide present, but that it is essentially free (about 90- 95% pure) of non-nucleotide material naturally associated with it, and thus is distinguished from isolated chromosomes. As used herein, the terms "isolated" and "purified" in the context of a proteineous agent (e.g., a peptide, polypeptide, protein or antibody) refer to a proteineous agent which is substantially free of cellular material and in some embodiments, substantially free of heterologous proteineous agents (i.e. contaminating proteins) from the cell or tissue source from which it is derived, or substantially free of chemical precursors or other chemicals when chemically synthesized. The language "substantially free of cellular material" includes preparations of a proteineous agent in which the proteineous agent is separated from cellular components of the cells from which it is isolated and/or recombinantly and/or synthetically produced. Thus, a proteineous agent that is substantially free of cellular material includes preparations of a proteineous agent having less than about 30%, 20%, 10%, or 5% (by dry weight) of heterologous proteineous agent (e.g. protein, polypeptide, peptide, or antibody; also referred to as a "contaminating protein"). When the proteineous agent is recombinantly produced, it is also preferably substantially free of culture medium, i.e. culture medium represents less than about 20%, 10%, or 5% of the volume of the protein preparation. When the proteinaceous agent is produced by chemical synthesis, it is preferably substantially free of chemical precursors or other chemicals, i.e., it is separated from chemical precursors or other chemicals which are involved in the synthesis of the proteinaceous agent. Accordingly, such preparations of a proteinaceous agent have less than about 30%, 20%, 10%, 5% (by dry weight) of chemical precursors or compounds other than the proteinaceous agent of interest. Preferably, proteinaceous agents disclosed herein are isolated.
In some alternative embodiments, determination of the expression levels of the biomarker proteins of the invention may be performed in the nucleic acid level, specifically, the mRNA level. In such embodiments for determining the expression level of the biomarkers of the invention, nucleic acid detecting molecule may be used. In some embodiments, the nucleic acid detecting molecule/s of the invention may comprise at least one of: (a) nucleic acid aptamers specific for said at least one of said biomarker proteins; and (b) at least one isolated oligonucleotides, each oligonucleotide specifically hybridizes to a nucleic acid sequence encoding said at least one biomarker protein. More specifically, such nucleic acid detecting molecules may comprise a nucleic acid aptamers specific for said at least one of the biomarker protein/s of the invention. In some other embodiments, the nucleic acid detecting molecules may comprise at least one isolated oligonucleotide/s, each oligonucleotide specifically hybridizes to a nucleic acid sequence encoding one of said at least one biomarker protein. In an optional embodiment, where the expression levels of the biomarkers of the invention are normalized, the method of the invention may use nucleic acid detecting molecules specific for a nucleic acid sequence encoding the control reference protein/s. As used herein, "nucleic acid molecules" or "nucleic acid sequence" are interchangeable with the term "polynucleotide(s)" and it generally refers to any polyribonucleotide or poly- deoxyribonucleotide, which may be unmodified RNA or DNA or modified RNA or DNA or any combination thereof. "Nucleic acids" include, without limitation, single- and double- stranded nucleic acids. As used herein, the term "nucleic acid(s)" also includes DNAs or RNAs as described above that contain one or more modified bases. Thus, DNAs or RNAs with backbones modified for stability or for other reasons are "nucleic acids". The term "nucleic acids" as it is used herein embraces such chemically, enzymatically or metabolically modified forms of nucleic acids, as well as the chemical forms of DNA and RNA characteristic of viruses and cells, including for example, simple and complex cells. A "nucleic acid" or "nucleic acid sequence" may also include regions of single- or double- stranded RNA or DNA or any combinations.
As used herein, the term "oligonucleotide" is defined as a molecule comprised of two or more deoxyribonucleotides and/or ribonucleotides, and preferably more than three. Its exact size will depend upon many factors which in turn, depend upon the ultimate function and use of the oligonucleotide. The oligonucleotides may be from about 3 to about 1,000 nucleotides long. Although oligonucleotides of 5 to 100 nucleotides are useful in the invention, preferred oligonucleotides range from about 5 to about 15 bases in length, from about 5 to about 20 bases in length, from about 5 to about 25 bases in length, from about 5 to about 30 bases in length, from about 5 to about 40 bases in length or from about 5 to about 50 bases in length. More specifically, the detecting oligonucleotides molecule used by the composition of the invention may comprise any one of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 bases in length. It should be further noted that the term "oligonucleotide" refers to a single stranded or double stranded oligomer or polymer of ribonucleic acid (RNA) or deoxyribonucleic acid (DNA) or mimetics thereof. This term includes oligonucleotides composed of naturally-occurring bases, sugars and covalent internucleoside linkages (e.g., backbone) as well as oligonucleotides having non-naturally-occurring portions which function similarly.
In some specific embodiments, where the detecting molecules of the invention are nucleic acid based molecules, optional detecting molecule/s may be at least one nucleic acid aptamer specific for the at least one of said biomarker proteins.
As used herein the term "aptamer" or "specific aptamers " denotes single-stranded nucleic acid (DNA or RNA) molecules which specifically recognizes and binds to a target molecule. The aptamers according to the invention may fold into a defined tertiary structure and can bind a specific target molecule with high specificities and affinities. Aptamers are usually obtained by selection from a large random sequence library, using methods well known in the art, such as SELEX and/or Molinex. In various embodiments, aptamers may include single- stranded, partially single- stranded, partially double-stranded or double-stranded nucleic acid sequences; sequences comprising nucleotides, ribonucleotides, deoxyribonucleotides, nucleotide analogs, modified nucleotides and nucleotides comprising backbone modifications, branch points and non-nucleotide residues, groups or bridges; synthetic RNA, DNA and chimeric nucleotides, hybrids, duplexes, heteroduplexes; and any ribonucleotide, deoxyribonucleotide or chimeric counterpart thereof and/or corresponding complementary sequence. In certain specific embodiments, aptamers used by the invention are composed of deoxyribonucleotides.
According to the present invention and as appreciated in the art, the recognition between the aptamer and the antigen is specific and may be detected by the appearance of a detectable signal by using a colorimetric sensor or a fluorimetric/lumination sensor.
The aptamers as used according to some aspects of the invention may be biotinylated. The aptamers may optionally include a chemically reactive group at the 3 and/or 5 termini. The term reactive group is used herein to denote any functional group comprising a group of atoms which is found in a molecule and is involved in chemical reactions. Some non-limiting examples for a reactive group include primary amines (NH2), thiol (SH), carboxy group (COOH), phosphates (P04), Tosyl, and a photo-reactive group.
In some embodiments, the aptamer as used herein may optionally comprise a spacer between the nucleic acid sequence and the reactive group. The spacer may be an alkyl chain such as (CH2)6/12, namely comprising six to twelve carbon atoms.
In some other embodiments, the detection molecule may be at least one primer, at least one pair of primers, nucleotide probes and any combinations thereof.
Thus, it should be further appreciated that the methods, as well as the compositions and kits of the invention may comprise, as an oligonucleo tide-based detection molecule, both primers and probes. The term, "primer", as used herein refers to an oligonucleotide, whether occurring naturally as in a purified restriction digest, or produced synthetically, which is capable of acting as a point of initiation of synthesis when placed under conditions in which synthesis of a primer extension product, which is complementary to a nucleic acid strand, is induced, i.e., in the presence of nucleotides and an inducing agent such as a DNA polymerase and at a suitable temperature and pH. The primer may be single- stranded or double- stranded and must be sufficiently long to prime the synthesis of the desired extension product in the presence of the inducing agent. The exact length of the primer will depend upon many factors, including temperature, source of primer and the method used. For example, for diagnostic applications, depending on the complexity of the target sequence, the oligonucleotide primer typically contains 10-30 or more nucleotides, although it may contain fewer nucleotides. More specifically, the primer used by the methods, as well as the compositions and kits of the invention may comprise 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides or more. In certain embodiments, such primers may comprise 30, 40, 50, 60, 70, 80, 90, 100 nucleotides or more. In specific embodiments, the primers used by the method of the invention may have a stem and loop structure. The factors involved in determining the appropriate length of primer are known to one of ordinary skill in the art and information regarding them is readily available.
As used herein, the term "probe" means oligonucleotides and analogs thereof and refers to a range of chemical species that recognize polynucleotide target sequences through hydrogen bonding interactions with the nucleotide bases of the target sequences. The probe or the target sequences may be single- or double-stranded RNA or single- or double- stranded DNA or a combination of DNA and RNA bases. A probe may be 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 and up to 30 nucleotides in length as long as it is less than the full length of the target mRNA or any gene encoding said mRNA. Probes can include oligonucleotides modified so as to have a tag which is detectable by fluorescence, chemiluminescence and the like. The probe can also be modified so as to have both a detectable tag and a quencher molecule, for example TaqMan(R) and Molecular Beacon(R) probes.
The oligonucleotides and analogs thereof may be RNA or DNA, or analogs of RNA or DNA, commonly referred to as antisense oligomers or antisense oligonucleotides. Such RNA or DNA analogs comprise, but are not limited to, 2-'0-alkyl sugar modifications, methylphosphonate, phosphorothiate, phosphorodithioate, formacetal, 3-thioformacetal, sulfone, sulfamate, and nitroxide backbone modifications, and analogs, for example, LNA analogs, wherein the base moieties have been modified. In addition, analogs of oligomers may be polymers in which the sugar moiety has been modified or replaced by another suitable moiety, resulting in polymers which include, but are not limited to, morpholino analogs and peptide nucleic acid (PNA) analogs. Probes may also be mixtures of any of the oligonucleotide analog types together or in combination with native DNA or RNA. At the same time, the oligonucleotides and analogs thereof may be used alone or in combination with one or more additional oligonucleotides or analogs thereof.
According to this option, the expression level may be determined using amplification assay. The term "amplification assay", with respect to nucleic acid sequences, refers to methods that increase the representation of a population of nucleic acid sequences in a sample. Nucleic acid amplification methods, such as PCR, isothermal methods, rolling circle methods, etc., are well known to the skilled artisan. More specifically, as used herein, the term "amplified", when applied to a nucleic acid sequence, refers to a process whereby one or more copies of a particular nucleic acid sequence is generated from a template nucleic acid, preferably by the method of polymerase chain reaction. "Polymerase chain reaction" or "PCR" refers to an in vitro method for amplifying a specific nucleic acid template sequence. The PCR reaction involves a repetitive series of temperature cycles and is typically performed in a volume of 50-100 microliter. The reaction mix comprises dNTPs (each of the four deoxynucleotides dATP, dCTP, dGTP, and dTTP), primers, buffers, DNA polymerase, and nucleic acid template. The PCR reaction comprises providing a set of polynucleotide primers wherein a first primer contains a sequence complementary to a region in one strand of the nucleic acid template sequence and primes the synthesis of a complementary DNA strand, and a second primer contains a sequence complementary to a region in a second strand of the target nucleic acid sequence and primes the synthesis of a complementary DNA strand, and amplifying the nucleic acid template sequence employing a nucleic acid polymerase as a template-dependent polymerizing agent under conditions which are permissive for PCR cycling steps of (i) annealing of primers required for amplification to a target nucleic acid sequence contained within the template sequence, (ii) extending the primers wherein the nucleic acid polymerase synthesizes a primer extension product. "A set of polynucleotide primers", "a set of PCR primers" or "pair of primers" can comprise two, three, four or more primers.
Real time nucleic acid amplification and detection methods are efficient for sequence identification and quantification of a target since no pre-hybridization amplification is required. Amplification and hybridization are combined in a single step and can be performed in a fully automated, large- scale, closed-tube format.
Methods that use hybridization-triggered fluorescent probes for real time PCR are based either on a quench-release fluorescence of a probe digested by DNA Polymerase (e.g., methods using TaqMan(R), MGB- TaqMan(R)), or on a hybridization- triggered fluorescence of intact probes (e.g., molecular beacons, and linear probes). In general, the probes are designed to hybridize to an internal region of a PCR product during annealing stage (also referred to as amplicon). For those methods utilizing TaqMan(R) and MGB-TaqMan(R) the 5'-exonuclease activity of the approaching DNA Polymerase cleaves a probe between a fluorophore and a quencher, releasing fluorescence. Thus, a "real time PCR" or "RT-PCT" assay provides dynamic fluorescence detection of amplified biomarker proteins of the invention or any control reference gene produced in a PCR amplification reaction. During PCR, the amplified products created using suitable primers hybridize to probe nucleic acids (TaqMan(R) probe, for example), which may be labeled according to some embodiments with both a reporter dye and a quencher dye. When these two dyes are in close proximity, i.e. both are present in an intact probe oligonucleotide, the fluorescence of the reporter dye is suppressed. However, a polymerase, such as AmpliTaq GoldTM, having 5'-3' nuclease activity can be provided in the PCR reaction. This enzyme cleaves the fluorogenic probe if it is bound specifically to the target nucleic acid sequences between the priming sites. The reporter dye and quencher dye are separated upon cleavage, permitting fluorescent detection of the reporter dye. Upon excitation by a laser provided, e.g., by a sequencing apparatus, the fluorescent signal produced by the reporter dye is detected and/or quantified. The increase in fluorescence is a direct consequence of amplification of target nucleic acids during PCR.
More particularly, QRT-PCR or "qPCR" (Quantitative RT-PCR), which is quantitative in nature, can also be performed to provide a quantitative measure of gene expression levels. In QRT-PCR reverse transcription and PCR can be performed in two steps, or reverse transcription combined with PCR can be performed. One of these techniques, for which there are commercially available kits such as TaqMan(R) (Perkin Elmer, Foster City, CA), is performed with a transcript-specific antisense probe. This probe is specific for the PCR product (e.g. a nucleic acid fragment derived from a gene) and is prepared with a quencher and fluorescent reporter probe attached to the 5' end of the oligonucleotide. Different fluorescent markers are attached to different reporters, allowing for measurement of at least two products in one reaction.
When Taq DNA polymerase is activated, it cleaves off the fluorescent reporters of the probe bound to the template by virtue of its 5-to-3' exonuclease activity. In the absence of the quenchers, the reporters now fluoresce. The color change in the reporters is proportional to the amount of each specific product and is measured by a fluorometer; therefore, the amount of each color is measured and the PCR product is quantified. The PCR reactions can be performed in any solid support, for example, slides, microplates, 96 well plates, 384 well plates and the like so that samples derived from many individuals are processed and measured simultaneously. The TaqMan(R) system has the additional advantage of not requiring gel electrophoresis and allows for quantification when used with a standard curve. A second technique useful for detecting PCR products quantitatively without is to use an intercalating dye such as the commercially available QuantiTect SYBR Green PCR (Qiagen, Valencia California). RT-PCR is performed using SYBR green as a fluorescent label which is incorporated into the PCR product during the PCR stage and produces fluorescence proportional to the amount of PCR product.
Both TaqMan(R) and QuantiTect SYBR systems can be used subsequent to reverse transcription of RNA. Reverse transcription can either be performed in the same reaction mixture as the PCR step (one-step protocol) or reverse transcription can be performed first prior to amplification utilizing PCR (two-step protocol).
Additionally, other known systems to quantitatively measure mRNA expression products include Molecular Beacons(R) which uses a probe having a fluorescent molecule and a quencher molecule, the probe capable of forming a hairpin structure such that when in the hairpin form, the fluorescence molecule is quenched, and when hybridized, the fluorescence increases giving a quantitative measurement of gene expression.
According to this embodiment, the detecting molecule may be in the form of probe corresponding and thereby hybridizing to any region or at least one of the biomarker protein or any control reference protein. More particularly, it is important to choose regions which will permit hybridization to the target nucleic acids. Factors such as the Tm of the oligonucleotide, the percent GC content, the degree of secondary structure and the length of nucleic acid are important factors. It should be further noted that a standard Northern blot assay can also be used to ascertain an RNA transcript size and the relative amounts of the biomarker proteins of the invention or any control gene product, in accordance with conventional Northern hybridization techniques known to those persons of ordinary skill in the art.
In some alternative embodiments, determining the level of expression of at least one or of at least five of the RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s of the invention may be performed by the step of subjecting a biological sample of the examined subject, or any protein product obtained therefrom to mass spectrometry analysis or assay. Thus, it should be appreciated that in certain embodiments, the signature proteins, specifically, at least one, at least five, at least four, at least six or at least ten of the biomarker proteins of the invention or any protein-fragments thereof may be also detected and quantified without the need for detection molecule/s. Detection can be based on MS approaches using non-targeted or targeted methods such as selected reaction monitoring (SRM) or parallel reaction monitoring (PRM). These analyses can be performed with or without a reference heavy standard and provide quantitative measure of the peptide/protein amount. The heavy reference can be a synthetic peptide, or a chemically labeled peptide/protein or metabolically labeled proteins. In the absence of a standard, the MS signal can provide the measure of peptide abundance.
According to some embodiments, the method of the invention may use as a sample any one of a biological sample of organ/s, cell/s or tissue/s and a blood sample. In further specific embodiments a sample may be a primary tumor sample. In more particular embodiments, the methods of the invention may use a primary breast tumor sample.
As used herein, the term "sample" refers to cells, sub-cellular compartments thereof, tissue or organs. The tissue may be a whole tissue, or selected parts of a tissue. Tissue parts can be isolated by micro-dissection of a tissue, or by biopsy, or by enrichment of sub-cellular compartments. The term "sample" further refers to healthy as well as diseased or pathologically changed cells or tissues. Hence, the term further refers to a cell or a tissue associated with a disease, such a tumor, in particular carcinoma, breast cancer, and more specifically, Luminal A or B breast cancer. A sample can be cells that are placed in or adapted to tissue culture. In some embodiments, a sample may also be a blood, body fluid such as plasma, lymph, urine, saliva, serum, cerebrospinal fluid, seminal plasma, pancreatic juice, breast milk, or lung lavage. A sample can additionally be a cell or tissue from any species, including prokaryotic and eukaryotic species, specifically, humans. A tissue sample can be further a fractionated or preselected sample, if desired, preselected or fractionated to contain or be enriched for particular cell types. The sample can be fractionated or preselected by a number of known fractionation or pre selection techniques. A sample can also be any extract of the above. The term also encompasses protein fractions or alternatively, nucleic acid from cells or tissue. Thus, in some specific embodiments, the sample may be any one of a biological sample of organ/s, cell/s or tissue/s and a blood sample. In yet some other embodiments, the sample may be a primary tumor sample. In certain embodiments, the sample is obtained from a subject suffering from a luminal A or a luminal B breast tumor.
From treatment decision making point of view, the lymph node status of breast cancer patients is the most important variable in the management of the disease (Jatoi et al., 1999). Luminal tumors still confined to their original surroundings will mostly be treated with tamoxifen or aromatase inhibitors, while in the node-positive setting, chemotherapy will usually be applied and risk for recurrence rises substantially (Ellis and Perou, 2013). Currently, the lymph node status is mostly determined after the dissection of lymph nodes during surgery, but this may result in additional complications to the lymphatic system (Sakorafas et al., 2006). Other methods such as sentinel lymph node biopsy and imaging are evolving, but to date, there is no reliable way to determine lymph node involvement based on the primary tumor itself. Thus, in some embodiments, the diagnostic and prognostic methods of the invention may provide an efficient tool for personalized treatment effective for specific subjects. In accordance with some embodiments, the methods of the inventions may be used for determining a treatment regimen for a subject suffering from a luminal A or a luminal B breast tumor. Accordingly, the method comprising in the first step, determining the expression level of at least one biomarker protein in at least one biological sample of a subject in need of such treatment, to obtain an expression value for each of said at least one biomarker protein. The biomarker proteins may be at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten at least eleven, at least twelve, at least thirteen, at least fourteen or at least fifteen of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB or any combination thereof. In yet more specific embodiments, the biomarker proteins may be at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB or any combination thereof. In some further specific embodiments, the biomarker proteins may be at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s. Specifically, such five biomarker proteins may be LSM2, METAP2, RPS24, RBM12B and CAPS. In yet some further alternative embodiments, the at least five biomarker proteins may comprise RPS29, RBM3, PNP, METAP2 and RAB5C. In yet some further embodiments, the biomarker proteins may be at least four of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s, specifically, RPS24, LSM4, RBM12B and RPS29. In some particular and non-limiting embodiments of the invention, the biomarker proteins of the invention may be at least six of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2, specifically, RPS24, LSM4, RBM12B, RPS29, RBM3 and PNP. In some particular and non-limiting embodiments of the invention, the biomarker proteins may be at least ten of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB. More specifically, RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3 and SNX12. The second step of the method comprises determining if the expression value obtained in step (a) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein in at least one control sample. The third step of the method involves providing an appropriate therapeutic regimen to a subject determined as exhibiting a negative expression value of said at least one, or alternatively, at least five, at least four, at least six or at least ten of the biomarker protein/s of the invention.
The therapy according with the present invention is any therapy applicable to cancer and specifically to breast cancer. In some embodiments, for subjects classified as LNN patients by the methods of the invention, an endocrine therapy or any combination thereof with a biological therapy may be offered. Endocrine therapy refers to a treatment that adds, blocks, or removes hormones. In the context of the present disclosure, endocrine therapy is provided to slow or stop the growth of breast cancers. In this connection, synthetic hormones or other drugs may be given to block the body's natural hormones. In yet some further embodiments, therapy based on aromatase inhibitors may be offered. Other therapeutic options may also include biological therapy (antibodies and the like) and cryotherapy. In yet some other embodiments, where the subject is classified as an LNP patient, chemotherapy, radiotherapy or any combinations thereof may be offered.
As detailed herein, the method of the invention may be also applicable for evaluating or monitoring the responsiveness of a patient to treatment with any therapeutic agent or regimen. Accordingly, the patient may be evaluated in at least one time point after initiation of treatment in order to asses if the treatment protocol is efficient and appropriate. Determination can be carried out at an early time points such that a decision may be made regarding continuation of the treatment or alternatively readjusting the treatment protocol.
One of the challenges associated with cancer and specifically breast cancer treatment originates from non efficient treatments or resistance to treatment. Thus, the present invention further provides the use of at least one of the biomarker proteins as markers for evaluating response of patients treated with a certain therapeutic agent or monitoring the efficacy of treatment with a certain therapeutic agent. In some embodiments, the method of the invention may be particularly suitable for monitoring and early diagnosis of response of the diagnosed disorder in the subject.
Thus, in yet other embodiments, the invention provides a method for assessing responsiveness of a mammalian subject to treatment with a specific therapeutic agent or evaluating and/or monitoring the efficacy of treatment on a subject. This method is based on determining the expression values of the biomarkers of the invention before and any time after initiation of treatment, and calculating the ratio of the change in said values as a result of the treatment.
In more specific embodiments such method may comprise the step of:
First, in step (a), determining the expression level of at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen of the biomarker protein/s in a biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein, wherein said biomarker proteins are selected from RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB or any combination thereof. As noted above, the method of the invention may further encompass the use of at least one further additional detecting molecules, specifically, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more, specifically, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, 300, 350, 400, 450 and 500 at the most, additional biomarker proteins. In some embodiments, the methods, compositions and kits of the invention may provide and use in addition to detecting molecules specific for at least one of the biomarkers disclosed in Table 4, also detecting molecule/s specific for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 or 85 biomarker proteins disclosed in Table 2, and optionally, further detecting molecule/s specific for additional at least one biomarker protein/s, for example, any of the biomarkers presented in Table 3, or any other biomarker/s.
In the next step (b), repeating step (a) in at least one other biological sample of said subject obtained after initiation of said treatment.
The third step (c), involves calculating the rate of change of the expression value of the biomarker proteins between said temporally separated samples, for example, samples obtained before and after initiation of said treatment.
Finally, step (d) concerns determining if the rate of change determined between at least two temporally separated samples or to the rate of change calculated for expression values in at least one control sample obtained from at least two temporally separated samples, wherein at least one sample of said at least two samples is obtained after the initiation of said treatment.
In certain embodiments, a negative rate of change of the expression value of at least one of said biomarker protein/s (e.g., reduction in the expression of a least one of the biomarker protein/s of the invention in response to treatment) indicates that said subject exhibits a beneficial response to said treatment. In yet another embodiment, where a positive rate of change is calculated for a subject, that means that the expression of the biomarker proteins of the invention is elevated in response to treatment and the subject may be thus classified as a non-responder to the particular treatment. Therefore, the invention provides a tool for monitoring the efficacy of a treatment with a therapeutic agent and the disease progression. In some specific embodiments, the methods of the invention involve in step (a) determination of the expression level of at least five biomarker proteins in at least one biological sample of the examined subject, to obtain an expression value for each of the at least five biomarker proteins. It should be noted that at least five biomarker proteins may be selected from RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins. It should be understood that the at least five, at least four, at least six and at least tern biomarker proteins of the invention discussed herein before are also applicable for this method of the invention.
As indicated above, in accordance with some embodiments of the invention, in order to assess the patient condition, or monitor the disease progression, as well as responsiveness to a certain treatment, at least two "temporally- separated" test samples must be collected from the examined patient and compared thereafter in order to obtain the rate of change in the expression value of at least one of the biomarker proteins between said samples. In practice, to detect a change in at least one of these parameters between said samples, at least two "temporally-separated" test samples and preferably more must be collected from the patient.
The expression value is then determined using the method of the invention, applied for each sample. As detailed above, the rate of change in parameters is calculated by determining the ratio between at least two values of expression obtained from the same patient in different time-points or time intervals.
This period of time, also referred to as "time interval", or the difference between time points (wherein each time point is the time when a specific sample was collected) may be any period deemed appropriate by medical staff and modified as needed according to the specific requirements of the patient and the clinical state he or she may be in. For example, this interval may be at least one day, at least three days, at least three days, at least one week, at least two weeks, at least three weeks, at least one month, at least two months, at least three months, at least four months, at least five months, at least one year, or even more.
In some embodiments, one of the time points may correspond to a period in which a patient is experiencing a remission of the disease.
When calculating the rate of change, one may use any two samples collected at different time points from the patient. To ensure more reliable results and reduce statistical deviations to a minimum, averaging the calculated rates of several sample pairs is preferable. A calculated or average value of a negative rate of change of the expression value of at least one of said biomarker protein/s indicates that said subject exhibits a beneficial response to said treatment; thereby monitoring the efficacy of a treatment with a therapeutic agent and the disease progression. It should be noted that in certain embodiments, where normalization step is being performed, the values referred to above, are normalized values.
As indicated above, the invention provides diagnostic and prognostic methods. "Prognosis" is defined as a forecast of the future course of a disease or disorder, based on medical knowledge. This highlights the major advantage of the invention, namely, the ability to predict progression of the disease, based on the expression value of at least one of the biomarker proteins. More specifically, the ability to determine at early stage that the subject is suffering from a metastatic breast cancer, specifically, if a subject is classified as an LNN or alternatively as an LNP patient. This ability facilitates the selection of appropriate treatment regimen/s that may minimize side effects from unnecessary treatment, individually to each patient, as part of personalized medicine.
Still further, as indicated above, in order to execute the prognostic method of the invention, at least two different samples must be obtained from the subject in order to calculate the rate of change in the expression as detailed above. By obtaining at least two and preferably more biological samples from a subject and analyzing them according to the method of the invention, the prognostic method may be effective for predicting, monitoring and early diagnosing molecular alterations indicating response to treatment in said patient.
Thus, the prognostic method may be applicable for early, sub- symptomatic diagnosis of relapse when used for analysis of more than a single sample along the time-course of diagnosis, treatment and follow-up.
An "early diagnosis" provides diagnosis prior to appearance of clinical symptoms. Prior as used herein is meant days, weeks, months or even years before the appearance of such symptoms. More specifically, at least 1 week, at least 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, or even few years before clinical symptoms appear.
The number of samples collected and used for evaluation of the subject may change according to the frequency with which they are collected. For example, the samples may be collected at least every day, every two days, every four days, every week, every two weeks, every three weeks, every month, every two months, every three months every four months, every 5 months, every 6 months, every 7 months, every 8 months, every 9 months, every 10 months, every 11 months, every year or even more. Furthermore, to assess the trend in expression rates according to the invention, it is understood that the rate of change may be calculated as an average rate of change over at least three samples taken in different time points, or the rate may be calculated for every two samples collected at adjacent time points. It should be appreciated that the sample may be obtained from the monitored patient in the indicated time intervals for a period of several months or several years. More specifically, for a period of 1 year, for a period of 2 years, for a period of 3 years, for a period of 4 years, for a period of 5 years, for a period of 6 years, for a period of 7 years, for a period of 8 years, for a period of 9 years, for a period of 10 years, for a period of 11 years, for a period of 12 years, for a period of 13 years, for a period of 14 years, for a period of 15 years or more. In one particular example, the samples are taken from the monitored subject every two months for a period of 5 years.
The method for monitoring disease progression or early prognosis for disease relapse as detailed herein may be used for personalized medicine, by collecting at least two samples from the same patient at different stages of the disease.
As detailed above, the prediction obtained by the method of the invention made by comparing between the sample and the patient population may be dependent on the selection of population of patients to which the sample is compared to. A positive or higher expression value of the sample over a population of patients diagnosed as having an LNP breast tumor, indicates that the examined subject is suffering from a metastatic breast cancer, specifically, LNP.
By "patient" or "subject" it is meant any mammal that may be affected by the above-mentioned conditions, and to whom the treatment and diagnosis methods herein described is desired, including human, bovine, equine, canine, murine and feline subjects. Specifically, said patient is a human. In some embodiments, determining the level of expression of at least one or of at least five of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins may be performed by the step of contacting at least one detecting molecule or any combination or mixture of plurality of detecting molecules with a biological sample of the examined subject, or with any protein or nucleic acid product obtained therefrom. It should be noted that each of the detecting molecules is specific for one of the biomarker protein/s.
In some alternative embodiments, determining the level of expression of at least one or of at least five of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may be performed by the step of subjecting a biological sample of the examined subject, or any protein product obtained therefrom to mass spectrometry analysis or assay. The determination of the expression level of the proteins can be achieved by quantification methods excluding the need of detection molecules. Label-free quantification of proteins can be conducted by liquid chromatography-mass spectrometry (LC-MS) with electrospray ionization. This method provides differential expression measurements and enables the discovery of biological markers. Other methods for label-free quantification also can be used. Non-limiting examples of these methods include SRM or PRM. These analyses may be performed with or without a reference heavy standard and provide quantitative measure of the peptide/protein amount.
According to a second aspect, the invention relates to a diagnostic and/or prognostic composition comprising at least one detecting molecule or any combination or mixture of plurality of detecting molecules specific for determining the level of expression of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof. It should be noted that each of said detecting molecules is specific for one of said biomarker proteins. It should be appreciated that in certain embodiments, the composition of the invention may be at least one of diagnostic and prognostic composition. In certain embodiments, the detecting molecules comprised within the composition of the invention may be attached to a solid support. Definitions of solid support that may be used as part of the diagnostic composition of the invention are described in more detail herein after, in connection with the kit of the invention. It should be appreciated that in some specific and non-limiting embodiments, the detecting molecules of the composition of the invention may be provided in a suitable medium or a buffer. In some alternative embodiments, the detecting molecules of the invention may be provided in a dried form.
It should be appreciated that the invention encompasses compositions comprising detecting molecules specific for any combination of any of the marker protein used by the invention.
In some specific embodiments, the composition of the invention may comprise at least one detecting molecule or any combination or mixture of plurality of detecting molecules specific for determining the level of expression of at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof. It should be noted that each of the detecting molecules is specific for one of said biomarker proteins. In some particular and non-limiting embodiments of the invention, such at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise LSM2, METAP2, RPS24, RBM12B and CAPS. In yet some further alternative embodiments, the at least five biomarker proteins may comprise RPS29, RBM3, PNP, METAP2 and RAB5C. In some embodiments, the five biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten of the LSM4, RPS29, RBM3, PNP, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4 and CSTB biomarker proteins of the invention. In yet some further embodiments, the composition of the invention may comprise at least one detecting molecule specific for at least four of the biomarker proteins of the invention. In some particular and non-limiting embodiments of the invention, such at least four of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise RPS24, LSM4, RBM12B and RPS29. It should be appreciated that in some embodiments, the four biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten or at least eleven of the RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention. Still further, the composition of the invention may comprise detecting molecules specific for at least six of the biomarker proteins of the invention. In some embodiments, such at least six of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s may comprise RPS24, LSM4, RBM12B, RPS29, RBM3 and PNP. It should be appreciated that in some embodiments, the six biomarker proteins may further comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight or at least nine of the METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention. Still further embodiments relate to the compositions of the invention that may comprise detecting molecules specific for at least ten of the biomarker proteins of the invention, specifically, RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3 and SNX12. It should be appreciated that in some embodiments, the ten biomarker proteins may further comprise at least one, at least two, at least three, at least four, or at least five of the SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins of the invention.
Still further embodiments relate to the composition of the invention that may comprise detecting molecules specific for all fifteen biomarker proteins of the invention, specifically RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB. It should be appreciated that any of the combinations of at least five, at least four, at least six and at least ten of the biomarker proteins of the invention disclosed herein, are also applicable for any of the kits of the invention discussed herein after.
In certain embodiments, the compositions of the invention may further comprise detecting molecules specific for control reference protein. Such control reference protein may be used for normalizing the detected expression levels for the biomarker proteins used by the invention.
Non-limiting embodiments for control reference proteins may include ARCN1 (Archain 1), MPZL1 (Myelin Protein Zero-Like 1), NSF (N-ethylmaleimide- sensitive factor), PRKCD (Protein Kinase C, Delta), CAT (catalase), actin, tubulin, or other cytoskeletal proteins.
It should be appreciated that the composition of the invention may comprise at least one detecting molecules specific for at least one biomarker of the invention, specifically, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 of the biomarkers of Table 4, specifically, RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB. In some embodiments, the composition of the invention may comprise detecting molecules specific for at least one further additional biomarker. In more specific embodiments, the compositions of the invention may comprise also detecting molecule/s specific for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more, specifically, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, 300, 350, 384, 400, 450 and 500 at the most, additional biomarker proteins. In more specific embodiments, the detecting molecules used by the methods, compositions and kits of the invention may be specific for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 or 85 biomarker proteins disclosed in Table 2, and optionally, further detecting molecule/s specific for additional at least one biomarker protein/s, for example, any of the biomarkers presented in Table 3, or any other biomarker/s.
In one embodiment, the detection molecules of the invention may be in the form of isolated detecting amino acid molecules and isolated detecting nucleic acid molecules.
In some specific embodiments, the composition of the invention may comprise amino acid detecting molecules. More specifically, such molecules may be at least one of: (a) at least one isolated recombinant labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof; (b) antibodies specific for said at least one of said biomarker protein/s; (c) peptide aptamers specific for said at least one biomarker protein/s; or (d) any combination of (a), (b) and (c).
It should be noted that any of the amino acid based detecting molecules described herein before for the methods of the invention are also applicable for any of the compositions of the invention and are therefore encompassed by the present aspect as well.
In some alternative embodiments, the composition of the invention may comprise nucleic acid detecting molecules. Such detecting molecules may include at least one of: (a) nucleic acid aptamers specific for said at least one biomarker proteins; (b) at least one isolated oligonucleotides, each oligonucleotide specifically hybridizes to a nucleic acid sequence encoding said at least one biomarker protein.
In some embodiments, the detecting molecules of the composition/s of the invention may be at least one labeled or tagged optionally isolated and/or recombinant RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof. In such case, the determination of the expression level of said at least one biomarker protein/s may be performed by mass spectrometry. In yet some other embodiments, the detecting molecules may be at least one of antibodies, nucleic acid or peptide aptamers specific for said at least one of the at least one biomarker proteins, or any combination thereof. In some embodiments, the determination of the expression level of the at least one biomarker protein/s may be performed by an immunological assay.
In some specific embodiments, the detecting molecules comprised within the composition of the invention may be attached to a solid support. More specifically, as defined herein, the detecting molecules are optionally attached to a support where each of the detecting molecules is attached to a support in a unique pre- selected and defined region. In some other embodiments, the detecting molecules may be provided in non-immobilized form, specifically, not attached to a solid support but separated in different vessels, tubes, wells and the like. Nevertheless, in yet some alternative embodiments, the detecting molecules may be provided in a mixture that contains variety of detecting molecules specific for at least one and at most 500 of the biomarker proteins of the invention.
In some specific and optional embodiments, the composition of the invention may further comprise a biological sample. It should be appreciated that any of the biological samples described for the method of the invention are also applicable for the composition of the invention.
Thus, the invention may further comprise a composition comprising at least one of the detecting molecules specific for at least one biomarker protein/s of the invention, specifically, the biomarkers of Table 4, and a sample, specifically, a biological sample. It should be noted that in addition to the biomarker/s of Table 4, the composition of the invention may comprise detecting molecules specific for at least one further biomarker, provided that the detecting molecules of the compositions of the invention are specific for 500 biomarkers at the most. In some embodiments such further biomarkers may be selected from the proteins listed in any one of Tables 2 and 3. It should be appreciated that in more specific embodiments, the compositions of the invention may comprise detecting molecules specific for at least one additional biomarker protein, specifically, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more, specifically, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, 300, 350, 400, 450 and 500 at the most, additional biomarker proteins.
As noted above, it should be appreciated that any of the compositions of the invention may be used for predicting breast cancer progression, assessing the patient's condition and may be also used for monitoring responsiveness of a mammalian subject to treatment.
A third aspect of the invention relates to a kit comprising (a) detecting molecules specific for determining the level of expression of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof in a biological sample. It should be noted that each of said detecting molecules is specific for one of said biomarker proteins. In certain embodiments, the kit of the invention may optionally further comprise at least one of: (b) pre-determined calibration curve/s or predetermined standard/s providing standard expression values of said at least one biomarker/s; and (c) at least one control sample.
In some particular embodiments, the kit of the invention may comprise at least one detecting molecule specific for determining the level of expression of at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof in a biological sample. The invention further encompass any kit comprising detecting molecules specific for at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen of the biomarker protein/s of the invention. It should be further understood that the kit of the invention may comprise detecting molecules specific for any combination of the biomarker proteins of the invention, specifically the combinations specified herein above in connection with the methods and compositions aspects. It should be appreciated that each of the detecting molecule/s is specific for one of said biomarker proteins. In some embodiments, the kit of the invention may optionally further comprises at least one of: pre-determined calibration curve/s or predetermined standard/s providing standard expression values of said at least one biomarker protein/s; and at least one control sample. It should be appreciated that all the combinations disclosed herein before in connection with the compositions of the invention are also applicable for any of the kits of the invention. In further specific embodiments, the detecting molecules comprised within the kit of the invention may be isolated detecting nucleic acid molecules, isolated detecting amino acid molecules or any combinations thereof.
In more specific embodiments, the kits of the invention may comprise amino acid detecting molecules, more specifically, at least one of: (a) at least one labeled or tagged, optionally, recombinant and/or isolated RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof; (b) antibodies specific for said at least one of said biomarker proteins; (c) peptide aptamers specific for said at least one of said biomarker protein/s; and (d) any combination of (a), (b) and (c).
In yet some alternative embodiments, the kit of the invention may comprise nucleic acid detecting molecule, for example, at least one of: (a) nucleic acid aptamers specific for said at least one biomarker proteins; (b) at least one isolated oligonucleotides, each oligonucleotide specifically hybridizes to a nucleic acid sequence encoding said at least one biomarker protein/s.
In some further embodiments, the detecting molecules comprised within the kit of the invention may be attached to a solid support.
The detecting molecules of the invention were described in detailed in connection with the methods of the invention. It should be appreciated that all embodiments for detecting molecules mentioned therein are also applicable for the compositions and kits of the invention.
In yet another embodiment, the kit of the invention may further comprise instructions for use, wherein said instructions comprise at least one of:
(a) instructions for carrying out the detection and quantification of expression of said at least one of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s and optionally, of a control reference protein; and
(b) instructions for comparing the expression values of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB with a corresponding predetermined standard expression value or with expression value of at least one of said biomarker protein/s in said at least one control sample.
The components in the kit may depend on the method of detection and are not limited to any method. In some embodiments, the kit of the invention may further comprise at least one reagent for conducting a mass spectrometry assay. Such reagents may include trypsin, buffers, filters and the like, for peptide purification. In some other embodiments, the kit of the invention further comprising at least one reagent for conducting an immunological assay selected from protein microarray analysis, ELISA, RIA, slot blot, dot blot, FACS, western blot, immunohistochemical assay, immunofluorescent assay and a radio-imaging assay.
According to one specific embodiment, the kit of the invention may be used for predicting breast cancer, assessing the patient's condition and monitoring responsiveness of a mammalian subject to treatment. In accordance with some embodiments, the kit of the invention may be used in a method for determining the progression of breast cancer in a subject. In some other embodiments, the subject is suffering from a luminal A or a luminal B breast tumor. Thus, in yet some further specific embodiments, the kit of the invention may be applicable for early determination and diagnosis of tested subjects that has LNN or LNP.
In accordance with some other embodiments, the sample to be used is any one of a biopsy of organs or tissues and a blood sample. In some further embodiments, the sample is a primary tumor sample. Still further, according to certain embodiments, the kits of the invention may use any appropriate biological sample. The term "biological sample" in the present specification and claims is meant to include samples obtained from a mammalian subject.
In some embodiments, the biological sample may be a bodily fluid, a tissue, a tissue biopsy, a skin swab, an isolated cell population or a cell preparation.
In some specific embodiments, the population of cells comprises cancer cells. In another embodiment the population of cells is an in vitro cultured cell population.
In some embodiments, the biological sample may be a bodily fluid selected from the group consisting of blood, serum, plasma, urine, cerebrospinal fluid, amniotic fluid, tear fluid, nasal wash, mucus, saliva, sputum, broncheoalveolar fluid, throat wash, vaginal fluid and semen.
According to an embodiment of the invention, the sample may be a tissue sample or blood sample which can be obtained using a syringe needle for example from a vein of the subject or from the tissue. It should be noted that the cell may be isolated from the subject (e.g., for in vitro detection) or may optionally comprise a cell that has not been physically removed from the subject (e.g., in vivo detection).
In some embodiments, the sample is any one of a biopsy of organ/s or tissue/s and a blood sample. In some other embodiments, the sample is a primary tumor sample. In some other embodiments, the sample may be lymph node tissue._ Samples of "Primary Tumor" may refer to the original, or first, tumor in the body. Cancer cells from a primary tumor may spread to other parts of the body and form new, or secondary, tumors (i.e. metastasis). In most cases, secondary tumors are the same type of cancer as the primary tumor. The term primary tumor may be interchanged with the term primary cancer. Still further, the inventors consider the kit of the invention in compartmental form. It should be therefore noted that in certain embodiments the detecting molecules used for detecting the expression levels of the biomarker proteins may be provided in a kit attached to an array. As defined herein, a "detecting molecule array" refers to a plurality of detection molecules that may be nucleic acids based or protein based detecting molecules, optionally attached to a support where each of the detecting molecules is attached to a support in a unique pre- selected and defined region. In some embodiments, the detecting molecules are attached to a solid support.
For example, an array may contain different detecting molecules, such as specific antibodies, labeled or tagged proteins, peptides, aptamers, probes and/or primers. As indicated herein before, in case a combined detection of the biomarker proteins expression level, the different detecting molecules for each target may be spatially arranged in a predetermined and separated location in an array. For example, an array may be a plurality of vessels (test tubes), plates, micro-wells in a micro-plate, each containing different detecting molecules, specifically, aptamers, primers and antibodies, specific for each marker protein used by the invention. An array may also be any solid support holding in distinct regions (dots, lines, columns) different and known, predetermined detecting molecules.
As used herein, "solid support" is defined as any surface to which molecules may be attached through either covalent or non-covalent bonds. Thus, useful solid supports include solid and semisolid matrixes, such as aero gels and hydro gels, resins, beads, biochips (including thin film coated biochips), micro fluidic chip, a silicon chip, multi-well plates (also referred to as microtiter plates or microplates), membranes, filters, conducting and no conducting metals, glass (including microscope slides) and magnetic supports. More specific examples of useful solid supports include silica gels, polymeric membranes, particles, derivative plastic films, glass beads, cotton, plastic beads, alumina gels, polysaccharides such as Sepharose, nylon, latex bead, magnetic bead, paramagnetic bead, super paramagnetic bead, starch and the like. This also includes, but is not limited to, microsphere particles such as Lumavidin.TM. Or LS-beads, magnetic beads, charged paper, Langmuir-Blodgett films, functionalized glass, germanium, silicon, PTFE, polystyrene, gallium arsenide, gold, and silver. Any other material known in the art that is capable of having functional groups such as amino, carboxyl, thiol or hydroxyl incorporated on its surface, is also contemplated. This includes surfaces with any topology, including, but not limited to, spherical surfaces and grooved surfaces. It should be further appreciated that any of the reagents, substances or ingredients included in any of the methods and kits of the invention may be provided as reagents embedded, linked, connected, attached, placed or fused to any of the solid support materials described above. In some alternative embodiments, the detecting molecules may be provided as molecules that are not attached to any solid support. In some embodiments, the non-attached detecting molecules may be provided in separate containers, wells, tube vessels and the like. In some alternative embodiments, the attached or non-attached detecting molecules may be provided in a mixture that contains at least two detecting molecules specific for at least two biomarker protein/s of the invention.
Thus, in yet some further embodiments, when the compositions and kits of the invention are directed for using MS detection methods, the detecting molecules of the invention, e.g., the recombinant (or synthetically produced) labeled or tagged biomarker protein/s or any fragment or peptide thereof may be provided as a mixture in a tube or any other vessel or container.
Still further, the invention provides a method for assessing expression status of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB in a biological sample. In more specific embodiments the method comprising the step of: (a) providing at least one detecting molecule, any combination, mixture of plurality of detecting molecules or any composition of kit comprising the same, wherein each of said detecting molecules is specific for one of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB proteins. In the next step (b), contacting the at least one detecting molecule/s provided in (a) with the biological sample, or with any protein or nucleic acid product obtained therefrom and (c), performing a protein or nucleic acid detection assay to assess the expression status of said proteins. It should be appreciated that the samples and detecting molecules described by the invention herein before are also applicable for this specific method.
Still further aspect of the invention relates to a diagnostic and/or prognostic method for determining the progression of breast cancer in a subject, the method comprising: (a) providing at least one detecting molecule specific for at least one of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker proteins, and any additional biomarker proteins, for example any one of the biomarkers disclosed in Tables 2 and 3. It should be noted that in certain embodiments, the detecting molecules are specific for 500 biomarker/s at the most. It should be appreciated that in certain embodiments, the method of the invention may be at least one of diagnostic and prognostic method. It should be appreciated that the detecting molecules provided by the methods of the invention may be provided as an array, as a composition (specifically, any of the composition described herein above) or as a kit, as described herein before. In step (b), the method of the invention involves determining the expression level of at least one of the biomarker proteins of the invention in at least one biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein/s, using the detecting molecules provided in step (a). In the next step (c), determining if the expression value obtained in step (b) is any one of positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample. In certain embodiments, wherein at least one of: (i) a positive expression value of said at least one biomarker protein/s in said sample, indicates that said subject belongs to a pre-established population associated with positive lymph node metastatic status; and (ii) a negative expression value of said at least one biomarker protein/s in said sample, indicates that said subject belongs to a pre-established population associated with negative lymph node metastatic status.
It should be understood that any of the detecting molecules described by the invention are also applicable for this aspect.
All scientific and technical terms used herein have meanings commonly used in the art unless otherwise specified. The definitions provided herein are to facilitate understanding of certain terms used frequently herein and are not meant to limit the scope of the present disclosure.
The term "about" as used herein indicates values that may deviate up to 1%, more specifically 5%, more specifically 10%, more specifically 15%, and in some cases up to 20% higher or lower than the value referred to, the deviation range including integer values, and, if applicable, non-integer values as well, constituting a continuous range. As used herein the term "about" refers to + 10 %. The terms "comprises", "comprising", "includes", "including", "having" and their conjugates mean "including but not limited to". This term encompasses the terms "consisting of" and "consisting essentially of". The phrase "consisting essentially of" means that the composition or method may include additional ingredients and/or steps, and/or parts, but only if the additional ingredients and/or steps do not materially alter the basic and novel characteristics of the claimed composition or method. Throughout this specification and the Examples and claims which follow, unless the context requires otherwise, the word "comprise", and variations such as "comprises" and "comprising", will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. It should be noted that various embodiments of this invention may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub ranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed sub ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range. Whenever a numerical range is indicated herein, it is meant to include any cited numeral (fractional or integral) within the indicated range. The phrases "ranging/ranges between" a first indicate number and a second indicate number and "ranging/ranges from" a first indicate number "to" a second indicate number are used herein interchangeably and are meant to include the first and second indicated numbers and all the fractional and integral numerals there between.
As used herein the term "method" refers to manners, means, techniques and procedures for accomplishing a given task including, but not limited to, those manners, means, techniques and procedures either known to, or readily developed from known manners, means, techniques and procedures by practitioners of the chemical, pharmacological, biological, biochemical and medical arts.
It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub combination or as suitable in any other described embodiment of the invention. Certain features described in the context of various embodiments are not to be considered essential features of those embodiments, unless the embodiment is inoperative without those elements.
Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.
Disclosed and described, it is to be understood that this invention is not limited to the particular examples, methods steps, and compositions disclosed herein as such methods steps and compositions may vary somewhat. It is also to be understood that the terminology used herein is used for the purpose of describing particular embodiments only and not intended to be limiting since the scope of the present invention will be limited only by the appended claims and equivalents thereof.
It must be noted that, as used in this specification and the appended claims, the singular forms "a", "an" and "the" include plural referents unless the content clearly dictates otherwise. EXAMPLES
Experimental procedures
Cohort assembly
Formalin-fixed, paraffin-embedded (FFPE)-blocks were obtained from the department of pathology, Sheba Medical Center, Tel Hashomer, Israel. Included cases were ER-positive and Her-2 negative infiltrating ductal carcinoma of the breast.
Tumors with no lymph node metastases were classified as LNN primary tumors (n=21). Tumors with at least one macro metastasis in lymph nodes were classified as LNP primary tumors (n=20). LN metastatic tissue (n=25, 20 cases matched to LNP primary tumors) was obtained from lymph nodes dissected during breast surgery. Healthy breast duct epithelia (n=l l from LNN patients, n=l l from LNP patients) was obtained from the surgical borders of the specimen, defined by pathologists as "tumor free". All samples were pathologically examined. The full cohort can be seen in Table 1. Preparation of tissue samples
FFPE blocks were sliced into twelve ΙΟμπι-thick sections and mounted on histological slides. Slides were dried at 37°C overnight. Areas of high cellularity were marked on the slides to enrich for cancer cells (or healthy breast epithelia) and to avoid inclusion of connective, adipose or lymphatic tissue. ER staining was used as a guide for delineation of cancer cells. Proteins were denaturated, combined at a 1: 1 ratio with the super-SILAC mix and digested following the FFPE-FASP protocol (Ostasiewicz et al., 2010) was used. Peptides were fractionated by strong anion exchange (SAX) fractionation in a StageTip format. Briefly, sections were deparaffinized with xylene, and rehydrated using ethanol/water gradient. The marked tissue was scraped with 2-3μ1 of water into a tube containing lysis buffer (0.1M Tris-HCl pH-8 and 4% SDS). The samples were boiled for lh, following by centrifugation at 17000 rpm for 5 min to pellet tissue debris. Protein amounts were measured in the supernatant using BCA protein assay kit (Thermo Scientific Pierce).
Preparation of super-SILAC mix
Super-SILAC mix was prepared previously by T. Geiger, as described (Geiger et al., 2010, and WO 2011/042467). Four breast cancer cell lines that differ in origin, stage and receptor status (HCC1599, MCF7, HCC1937 and HCC2218) and normal mammary epithelial cells (HMEC) were cultured in a medium deprived of lysine and arginine, and supplemented with 'heavy' versions of
13 15 13 15
these amino acids C6 N2-Lys, 1JC6 N4-Arg) until they were fully incorporated into their proteomes. Cells were then lysed and combined in a ratio of 1: 1: 1: 1: 1 (protein amount) to create the super-SILAC mix.
Preparation of metabolically -labeled recombinant marker-proteins Recombinant proteins are expressed in E. Coli, which are heavy labeled with 13 C615 N2-Lys,
13 C 15 N4-Arg. Recombinant proteins are tagged (with HA, ABP, GST or other) to enable purification and absolute quantification. Proteins will purified with affinity chromatography, based on the specific tags. The absolute amount of each purified protein is determined by amino acid analysis and/or MS-based analysis relative to the purified tag. The heavy labeled recombinant proteins are combined with the unlabeled clinical samples to determine their absolute amounts. Sample preparation for MS analysis
Prior to protein digestion, proteins from the clinical samples were mixed with equal amounts of the super-SILAC mix and incubated with 0.1M dithiothreitol and 0.05M iodoacetamide. The mixed lysates were digested with trypsin on top of 30kDa cutoff Amicon filters, as depicted in the Filter- Aided Sample Preparation (FASP) procedure (Wisniewski et al., 2009b). In order to increase analytical depth, the resulting peptides were fractionated using pH-based strong anion exchange (SAX) in StageTip format (Wisniewski et al., 2009a). Each sample was separated into six fractions using buffers of different pH values followed by desalting and concentration on Cis StageTips (Rappsilber et al., 2007). Prior to MS analysis, peptides were eluted from StageTips using 80% acetonitrile, vacuum-concentrated and diluted in MS loading buffer (2% acetonitrile, 0.1% formic acid).
MS analysis
Peptides were separated by nano-ultra high performance liquid chromatography (UHPLC) (EasynLClOOO, Thermo Fisher Scientific) coupled on-line to a Q-Exactive or Q-Exactive Plus mass spectrometers (Thermo Fisher Scientific) through the EASY-Spray ionization source. Peptides were loaded onto to a 50 cm EASY-Spray column with a buffer containing 0.1% formic acid (buffer A), and eluted with a buffer containing 80% acetonitrile and 0.1% formic acid (buffer B) using different gradients. The pH=l l fraction was eluted with a linear 4-hour gradient of 5-25% buffer B; the pH=8, 6 and 5 fractions were eluted using a 7-28% gradient, and the pH=4 and 3 fractions were eluted with a 7-33% gradient. Mass spectra were acquired in a data- dependent manner with the top- 10 precursor m/z values from each MS scan fragmented by higher energy collisional dissociation (HCD). MS-scans and MS/MS scans were performed with resolutions of 70,000 and 17,500, respectively.
Data analysis
MS raw files were analyzed by MaxQuant (version 1.5.0.36; Cox and Mann, 2008). MS/MS spectra were searched against the reference UNIPROT human proteome (published November 2014) by the Andromeda search engine (Cox et al., 2011). False Discovery Rate (FDR) of 0.01 was used in on both the peptide and protein levels and determined by a decoy database. Prior to bioinformatics analysis, the resulting protein list was filtered to eliminate common contaminants and decoy database hits.
Bioinformatic and statistical analysis
Bioinformatic and statistical analyses were performed in the Perseus software and in MATLAB (version R2014a). For all analyses, we filtered the data to include only proteins quantified in >70% of the samples. Expression ratios towards internal standard were normalized by z-scoring and subtracting most frequent value on each sample, and missing data points were imputed by creating a normal distribution with a width of 0.3 and a downshift of 1.5. Correlation networks of proteins were constructed by generic k-means clustering of all protein pairs with a cutoff correlation of 0.3 for all samples and 0.5 for tumor samples. Additional network analyses were done using STRING database. All networks were visualized with Cytoscape. Welch's ttests for statistical significance were performed with permutation-based FDR correction threshold of 0.05. Categorical annotations of proteins were supplied from Uniprot in the form of Gene Ontology (GO) Biological Process, Molecular Function and Cellular Component, and KEGG pathway. Fisher exact tests for annotation enrichments were performed with FDR threshold of 0.02 against the human proteome. ID and 2D annotation enrichment were performed as described (Cox and Mann, 2012).
Mitochondrial activity assays
Functional examination of oxidative phosphorylation was performed by activity assays of complex I and complex IV on breast cancer frozen tumor microarrays (TMAs; BioChain institute, Inc. Newark, USA). For the complex IV activity assay, TMAs were brought to room temperature, washed for 5 min with 25mM sodium phosphate buffer pH7.4, and then incubated for 90 min at 37°C with COX incubation mixture containing 1 mg/ml Cytochrome C (C7752, Sigma-Aldrich), 1 mg/ml 3,3'-diaminobenzidine tetrahydrochloride hydrate (D5637, Sigma-Aldrich) and 0.2 mg/ml catalase (C1345, Sigma-Aldrich) in 25 mM sodium phosphate buffer at pH7.2-7.4 (Whitaker- Menezes et al., 2011). Counterstain was performed with hematoxylin (HHS32, Sigma-Aldrich) for 45 seconds. For complex I activity assay, the TMAs were brought to room temperature, washed for 5 min with 50mM Tris-HCl buffer pH7.4, and then incubated for 45 min at 37°C with NADH incubation mixture containing 2 mg/ml β-nicotinamide adenine dinucelotide (NAD; N7410, Sigma- Aldrich), 38 mM cobalt (II) chloride (C8661, Sigma-Aldrich) and 380 μg/ml nitrotetrazolium blue chloride (N6876, Sigma-Aldrich) in 1 ml of 77 mM Tris-HCl buffer pH7.4 (Whitaker-Menezes et al., 2011).
Pulsed-SILAC assay
Human mammary epithelial cells (HMEC) and the ER-positive breast cancer cell line MCF7 were fully labeled with medium-heavy lysine and arginine (Lys4 and Arg6) (Cambridge Isotopes Laboratory, MA, USA). Culture medium containing heavy versions of the same amino acids (Lys8 and ArglO) was pulsed into the medium-labeled cell lines at T=0, and harvested at Oh, 2h, 4h, 9h, 12h and 24h. At each time point, cell lysates were mixed with the same cells grown in light culture medium that serve as an internal standard. Lysates were digested overnight with trypsin in solution and were subjected to LC-MS/MS analysis and MaxQuant analysis as described above.
Immunohistochemistry
Breast cancer tumor microarrays were obtained from BioChain Institute, Inc. (Newark, USA), and stained with anti-ACOTl and anti-SLC25Al l (AbCam) or anti GLUL (Sigma- Aldrich/Prestige antibodies). Staining intensity of relevant cores (ER-positive, Her-2 negative invasive ductal carcinomas, and healthy ducts) was assessed by a pathologist on a 4-degree scale of 0 (no staining) to 3 (strong staining).
Table 1: Sample cohort
Lymph Her2 PR score
Affiliated Lymph nodes with nodes score (>2, (>0.5,
Case samples Grade macrometastases examined positive) positive)
A3 A3H, A3T 2 0 3 0 0 2.9
A5 A5T 2 0 6 0 2 3
A6 A6T 2 0 3 0 1 2.5
A7 A7T 2 0 4 0 2 3
A8 A8T 2 0 3 0 0 2.8
A9 A9T 2 0 8 0 1.5 3
All A11H, A11T 2 0 2 0 2 2.8
A12 A12T 2 0 2 0 0.3 2.8
A13 A13T 2 0 2 0 0.3 1.8
A14 A14H 3 0 4 0 0.8 2
A15 A15H, A15T 2 0 5 0 1.5 3
A16 A16H, A16T 2 0 3 0 0.8 1.5
A18 A18T 2 0 8 0 1.5 2.5
A19 A19H A19T 2 0 1 0 2.5 2.5
A30 A30H, A30T 2 0 1 0 0 3
A31 A31H, A3 IT 2 0 6 0 2 3
A32 A32T 2 0 2 0 2 3
A33 A33T 2 0 7 0 1 3
A34 A34H, A34T 2 0 1 0 3 3
A35 A35H, A35T 2 0 3 0 1 3
A36 A36H, A36T 2 0 7 0 0 3
A38 A38T 2-3 0 3 0 3 3
Bl B1M 3 2 5 0 0 2.8
B2 B2H, B2T, B2M 3 5 8 0 0.1 2.5
B3 B3H, B3T, B3M 2 1 2 0 0.5 3
B5 B5H, B5M 2 2 3 0 0.2 3
B6 B6T, B6M 2 4 10 0 0 3
B7 B7T, B7M 2 3 15 0 0.5 2.5
B8 B8H 2-3 1 1 0 1 2.5
B9 B9H, B9T, B9M 3 1 4 1 2.2 2.2
BIO B10T, B10M 2 1 2 1 1.5 3
Bll BUT, BUM 2 2 6 1 0.5 2.5
B12 B12T, B12M 2 1 5 1 3 3 B13 B13T, B13M 2-3 19 20 1 2 3
B14 B14T, B14M 2 1 2 0 0 3
B15 B15T, B15M 2 5 13 0 0 3
B16 B16T, B16M 2 1 1 0 2 3
B17 B17T, B17M 2 1 2 0 3 3
B18 B18T, B18M 2 4 10 0 0 3
B19H, B19T,
B19 B19M 3 2 17 0 0 1
B21H, B21T,
B21 B21M 3 1 7 0 0 3
B22H, B22T,
B22 B22M 3 3 3 0 0 3
B23 B23H, B23M 3 2 5 1 0 3
B24H, B24T,
B24 B24M 3 2 5 0 1 1.2
B25 B25T, B25M 3 10 14 0 2.5 2.8
B26 B26M 2 1 10 0 1.5 1.5
B27 B27T, B27M 3 1 1 0 0 2.8
B28 B28M 3 13 14 0 0 2.5
EXAMPLE 1
Proteomic profiling of breast cancer clinical samples
To gain insights into the proteomic changes that occur during breast cancer progression, the inventors assembled a cohort of formalin-fixed, paraffin-embedded (FFPE) breast clinical samples. Since for luminal tumors patient prognosis largely depends on the lymph node status (Paik et al., 2006), Identification of the processes that are altered upon cancer progression can lead to better understanding of the regulatory mechanisms of cancer. The cohort included 88 samples that represent different stages of breast cancer, with non-metastatic lesions (lymph-node negative; LNN, n=21), metastatic ones (lymph node positive; LNP, n=20), their matched lymph node metastases (n=25), and adjacent healthy ducts as controls (n=22) (Figure 1A).
Samples were obtained from lumpectomies or mastectomies, from patients which have not received any treatment prior to surgery, to eliminate possible proteomic changes caused by the treatment. The cohort included ER-positive, Her2-negative infiltrating ductal carcinomas (IDC, luminal), grade 2 or 3, based on immunohistochemical staining and pathologist review. Since for luminal tumors patient prognosis largely depends on the lymph node status, identification of the proteins that are altered in late stages can serve as prognostic markers, and understanding the processes that are altered upon cancer progression can lead to better understanding of the regulatory mechanisms of cancer invasion.
For accurate quantification the previously published breast-cancer super-SILAC mix was used (Geiger et al., 2010) to serve as a common internal standard against which proteins from all samples are quantified in the MS analysis. It is a mixed lysate of four breast cancer cell lines and primary mammary epithelial cells, labeled metabolically with heavy lysine and arginine. As a mixture, it contains labeled counterparts of the vast majority of proteins expressed in the clinical samples at similar levels, and it serves as a platform for relative quantifications of these proteins across the entire cohort.
MaxQuant analysis identified overall 150,471 peptides and 10,124 proteins and quantified overall 10,043 of them (Figure IB and in Figure 2A; false discovery rate of 1% on both peptide and protein levels). The inventors found no overall differences in the total number of quantified proteins, between the groups of samples, with 9093, 9746 and 9450 proteins in the healthy tissue, tumor tissue (both LNN and LNP) and metastatic tissue, respectively, with an average of 5439 proteins identified and 4300 quantified in each tissue. From these, 1499 were identified in all 88 samples, and only 177 were found in less than 10 of the samples (Figure 2B).
The combined dynamic range of protein expression encompassed eight orders of magnitude and the vast majority of proteins (97%) were expressed within four orders of magnitude (Figure 1C). Among these proteins known luminal breast cancer-associated proteins such as GAT A3, FOXA1 and the estrogen receptor ESR were identified. Among the most abundant proteins, the inventors identified five histones, actin and tubulin, as well as ribosomal proteins.
To examine the coverage of breast-related proteins, the inventors compared their dataset to the recently published draft of the human proteome (Wilhelm et al., 2014). The authors defined a "core" proteome comprising of 11,578 proteins, of which 8,596 (74%) were identified herein. Examination of the group of 1,021 breast tissue-specific proteins (as defined by Wilhelm et al), showed significantly higher median expression intensity (of 921 proteins identified in this study) compared to the overall protein intensity distribution, in all three of the clinical groups (Figure ID; annotation enrichment, Cox and Mann 2012; FDR=0.02).
These results show that the inventors were able to capture higher levels of the breast- specific proteome. To examine the ability of the super-SILAC mix to accurately quantify the proteins in each of the clinical groups the inventors determined the ratios between the samples and the super- SILAC mix proteins. For all groups of tumors the inventors found that more than 90% of the proteins are within a 5-fold difference from the super-SILAC standard, which provides accurate measurement of the ratios (Figure 2C).
These results endorse the inventor's data as providing the technical accuracy, and breast-cancer relevance of the identified proteins.
The large number of analyzed samples provides a global view of proteins and processes that act together and may be involved in tumor formation. To exploit that, the inventors constructed functional networks based on the correlations between pairs of proteins identified in this study. Pearson correlations with a cutoff of 0.3 were used as a basis for K-means generic clustering (Figure 3). The data was divided into 10 clusters followed by Fisher exact test to examine enriched processes within each cluster (FDR=0.02). This analysis revealed links between distinct cellular processes, which were previously not known to be related. For example, the inventor found two clusters with high density (high intra-cluster protein associations) to be enriched for ribosomal proteins, translation and ribosome biogenesis, as well as oxidative phosphorylation, lysosomes and interferon signaling (clusters 6 and 7). Other enriched and correlated processes include proteasome together with spliceosome (cluster 2), and glycolysis together with tRNA aminoacylation and focal adhesion (cluster 8). To exclude dominance of the differences between healthy tissue and the tumors, the inventors constructed an additional network comprising of correlations between tumor samples only, with a correlation cutoff of 0.5 (Figure 5A). In this analysis the inventors found highest correlations within specific functions, mostly consisting of large protein complexes such as ribosomes (cluster 1), spliceosome (cluster 5), oxidative phosphorylation (cluster 7) and DNA replication (cluster 6). Beyond these, the inventors further found high correlations between distinct compartments, such as mitochondrial oxidative phosphorylation and the peroxisome (cluster 7), and distinct functions, such as translation and lysosomal degradation (cluster 5) or DNA replication and locomotion (cluster 6). These results highlight the potential of this proteomic resource to reveal fundamental cancer-related associations of cellular processes, which can serve as the basis for further functional research.
EXAMPLE 2
Comparison of healthy and tumor tissue proteomes
As a first step in the analysis the inventors examined the data distribution in an unbiased manner. Principle component analysis revealed that the first component separates between the healthy tissues and the tumor samples, including the lymph node metastases (Figure 4A). All tumor samples (Primary LNN and LNP) and metastases were indistinguishable according to the first two components.
In agreement, correlation matrix of primary tumors and healthy tissue showed major differences between the healthy tissue and tumors, and co-clustering of primary tumors and metastases, highlighting their similar protein expression patterns (Figure 4B). The correlations between samples ranged from 0.06 to 0.87. The median correlation between healthy and primary tumors was 0.38 (and 0.39 for matched samples), and the median correlation between primary and metastases was 0.58 (and 0.75 for matched samples, see also below). Surprisingly, it was found that the correlation within the tumor sample group (both primary and metastases) was significantly higher than the correlation between the healthy tissues (Figure 5B). Examination of the variance of protein expression within each group of samples showed enrichment of extracellular matrix part in the proteins with highest standard deviation (2D annotation enrichment, Cox and Mann 2012; FDR=0.02; Figure 5C). These ECM proteins originate from imperfect dissection of the tissues, which is more prominent in the healthy ducts due to their small size (Figure 5D). While the ECM proteins contributed to the inter- tumor variance (Figure 5E), they constitute only 3-5% of the tissue mass, and upon data normalization these do not affect the overall ratio distribution (see Experimental Procedures).
Next, the inventors examined the differences between the healthy tissues and tumors. The inventors found overall 969 significantly different proteins between healthy ducts and the primary tumors, with 563 significantly upregulated proteins in the tumors and 406 proteins downregulated in tumors (Welch's t-test; FDR=0.05; list of 969 proteins not shown). More than 95% of these proteins are well within the 5-fold ratio towards the standard, and are therefore considered to be accurately determined (Figures 6A, 6B). Among these proteins, the inventors found two known breast cancer markers, Mucin 1 (CA15-3) and Cathepsin D, to be higher in the breast cancer tissues compared to normal duct epithelia (Figures 6C, 6D). Enrichment analysis of gene-ontology (GO) and KEGG pathways aimed at the identification of key processes that are up- or down-regulated in the cancer tissues vs. the healthy samples, showed significant changes in two major groups of cellular processes: (i) Protein homeostasis and quality control and (ii) Central metabolism.
EXAMPLE 3
Protein homeostasis alteration in tumors
Cellular protein levels are tightly controlled by a plethora of housekeeping functions that control protein synthesis, mainly transcription and translation, and post-translational control of protein folding and stability. At the basis of these, genomic instability, one of the central cancer hallmarks, increases the extent of aberrant proteins. In agreement, DNA repair-proteins were found to be significantly downregulated in tumors (Figure 4C). These include components of non-homologous end joining (NHEJ) complex Ku (XRCC5 and XRCC6), both components of MutS alpha mismatch repair system (MSH2 and MSH6), as well as condensin complex (SMC2 and SMC4) and PARP1, which have a role in single-strand DNA break repair. On the mRNA level, a key quality control mechanism, the nonsense-mediated mRNA decay (NMD), was downregulated. UPF1 and UPF2, two vital component of this system, were significantly reduced in tumors. Potentially, the reduced activity of the NMD can lead to translation of truncated proteins with deleterious gain of function or dominant negative activity.
On the protein level, the inventors found a highly significant enrichment and upregulation of structural ribosomal proteins, both cytosolic and mitochondrial (Figure 4C), implying an elevated rate of protein translation. While the demand for protein production is expected, the inventors found that fundamental translation auxiliary proteins, namely the aminoacyl tRNA synthetases (ARSs), were surprisingly downregulated in the tumors. The ARSs have a canonical role in translation regulation; of note, recent evidence also point to non-canonical roles which may also affect tumorigenesis). As translational regulators, improper activity of the ARSs can affect both the rate and accuracy of protein synthesis, making the ARSs important quality control mediators. An additional level of regulation involves proper protein folding.
In normal cells, accumulation of unfolded proteins would result in the activation of ER chaperones that collectively act under the unfolded protein response (UPR) (Ma and Hendershot, 2004). Interestingly, three major players in this response - namely, the chaperones GRP78 (HSPA5), GRP94 (HSP90B 1) and GRP170 (HYOU1) - were significantly downregulated in tumors relative to normal tissue (Figure 4C). Significant decrease is also evident in six out of eight core subunits of the TRiC/CCT chaperonin complex, as well as in the chaperone calnexin (Figure 4C). The substantial deregulation of protein folding machinery may lead to increased amounts of unfolded proteins, a state tightly linked with stress and disease (Ma and Hendershot, 2004).
The inventors reasoned that the increased mal-production of proteins may increase the need to degrade them, and potentially recycle the amino acids. Indeed, it was found that seven proteins from the 20S core proteolytic structure of the proteasome are significantly upregulated (PSMA1, PSMA5, PSMB 1, PSMB2, PSMB3, PS MB 4 and PSMB8) (Figure 4C) suggesting an increase in proteosomal activity. Interestingly, five members of the 19S regulatory complex (PSMC1, PSMC2, PSMDl, PSMD2 and PSMD7) were significantly downregulated, which may implicate alteration of substrate recognition and binding (Liu and Jacobson, 2013). Lysosomal proteins were also significantly upregulated (Figure 4C) with the most prominent components belonging to the vacuolar-type proton ATPase, as well as several cathepsins (CTSA, CTSB, CTSD, and CTSZ). Taken together, these results suggest that protein homeostasis is impaired in tumor cells, and that tumors may over-produce improperly functioning proteins, which may interfere with proper cellular activities and thus facilitate tumorigenesis.
In order to validate the higher protein turnover in tumor cells, a pulsed-SILAC experiment in normal mammary epithelial cells (HMEC) versus ER-positive breast cancer cells (MCF7) was conducted. To that end, cells were fully labeled with medium-heavy lysine and arginine (Lys4 and Arg6) followed by pulse-labeling with heavy amino acids (Lys8 and ArglO), and then mixed with light-labeled cells as a common standard. Proteomic analysis of the samples showed that in agreement with the tumor data, the H/L ratios (representing protein synthesis), M/L ratios (representing degradation) and H/M ratios (representing turnover), were all significantly higher in MCF7 cells than in HMEC (Figure 4D), supporting the functional output of the increased levels of ribosomal proteins, proteasome and lysosomes in the tumors.
EXAMPLE 4
Metabolic remodeling in tumor samples
Two central metabolic pathways were found to be enriched in the significantly changing proteins: Oxidative phosphorylation proteins were significantly upregulated in the cancer samples, concurrently, key glycolytic enzymes such as HK2, GAPDH, ALDOA, LDHA and LDHB were downregulated (Figure 4C). The increased activity of the electron transport chain in tumor cells was further validated by activity-based assays for mitochondrial complex I and complex IV, using breast cancer tumor (Duct carcinoma in situ) arrays (Figures 6E, 6F).
These surprising results contradict the known function of glycolysis as an energy production pathway and a carbon source for biosynthetic pathways in cancer. The inventors therefore zoomed- in on the entire metabolic network, to elucidate the individual regulated pathways. The inventors matched their data to the Recon 1 human metabolic network, which includes -3300 metabolic reactions assigned to over 2000 proteins (Duarte et al., 2007). In order to test for significantly changing metabolic pathways, the inventors averaged the expression values of all proteins which are assigned to a specific Recon 1 pathway in every sample and examined their statistical difference between the healthy ducts and the tumors. Overall, the inventors found 18 upregulated pathways and 16 downregulated pathways between healthy tissue and tumor tissue (Figures 7A-7B and 8). The following Recon 1 pathways were shown to be down regulated in tumors: Alanine and Aspartate Metabolism, Arginine and Proline Metabolism, Ascorbate and Aldarate Metabolism, Cholesterol Metabolism, Fatty acid activation, Glycine, Serine, and Threonine Metabolism, Glycolysis/Gluconeogenesis, Glyoxylate and Dicarboxylate Metabolism, Histidine Metabolism, IMP Biosynthesis, Lysine Metabolism, Methionine Metabolism, Pentose Phosphate Pathway, Propanoate Metabolism, Selenoamino acid metabolism, Transport, Extracellular.
The following Recon 1 pathways were shown to be up regulated in tumors: Chondroitin sulfate degradation, Fatty Acid Metabolism, Galactose metabolism, Glutathione Metabolism, Heme Biosynthesis, Heme Degradation, Heparan sulfate degradation, Hyaluronan Metabolism, Keratan sulfate degradation, N-Glycan Degradation, Nucleotides, Oxidative Phosphorylation, Pentose and Glucuronate Interconversions, Pyrimidine Catabolism, ROS Detoxification, Sphingolipid Metabolism, Tetrahydrobiopterin, Transport, Lysosomal.
Oxidative phosphorylation was highly upregulated (p=8e-13) alongside with increases in ROS detoxification enzymes (p=0.01), while glycolysis was downregulated (p=0.004) (Figure 7B). Two metabolic pathways branching from glycolytic intermediates, namely the pentose phosphate pathway (PPP) and the serine/glycine biosynthetic pathway, were also downregulated (p=0.005 and 2e-7, respectively), primarily due to the significant reduction in expression of branching enzymes glucose-6-phosphate dehydrogenase (G6PD) and phosphoglycerate dehydrogenase (PHGDH) (Figure 7B).
Interestingly, the inventors observed a trend towards reduced biosynthesis and increased breakdown of fatty acids. ATP-citrate lyase (ACLY), a key enzyme that produced acetyl-CoA for lipid and cholesterol biosynthesis, was significantly downregulated in tumors (p=8e-7; Figure 7B). Fatty acid synthase (FASN), the primary enzyme in lipid synthesis, was also lower in expression. The inventors observed a very prominent decrease in the biosynthesis of cholesterol (p=le-7), with five enzymes significantly downregulated (Figure 7B). Concurrently, several enzymes in the peroxisomal β-oxidation of fatty acids are significantly elevated, pointing to a utilization of this pathway to generate reducing power in the form of NADH, to be used in cellular respiration. Two components of the malate/aspartate shuttle, which transports the NADH from the cytosol to the mitochondria, were also significantly upregulated (Figure 7B).
The inventors found that tumor cells exhibit significantly elevated levels of glutamine synthase (GLUL), which uses ammonia and glutamate to produce glutamine. Concurrently, the reverse reaction, catalyzed by glutaminase (GLS) was significantly reduced, suggesting higher demand of 'in-house' produced glutamine (Figures 7A-7B). In support of this, the major importer for glutamine, SLC1A5, was significantly downregulated, as well as the bidirectional transporter SLC7A5/SLC3A2, which has been shown to control outward efflux of glutamine in exchange for essential amino acids. Glutamate was also increasingly catabolized to a-ketoglutarate by glutamate dehydrogenase (GLUD1), potentially to supply the TCA cycle by way of anaplerosis and increase NADH production. Finally, the inventors found that the reactive-oxygen species (ROS)- detoxification pathway is significantly upregulated (p=0.01), with increased activity of enzymes such as peroxiredoxin (PRDX)-3, 4 and 5, glutathione peroxidase (GPX)-l and 4 and superoxide dismutase (SOD)-l (Figure 7B).
As a validation of the proteomic results, the inventors selected three metabolic enzymes that were higher in the cancer samples compared to the healthy tissue, and represent key regulated pathways: (i) GLUL, a key regulator of Gin production; (ii) SLC25A11, a member of the malate-aspartate shuttle; (iii) Acyl-CoA thioesterase (ACOT1), which operates in the β-oxidation of fatty acids and catalyzes the hydrolysis of acyl-CoA to coenzyme A and free fatty acid. The inventors validated their overexpression in tumors using immunohistochemistry on commercial tumor microarrays. In agreement with the proteomic results, ACOT1 and GLUL were absent in the healthy tissue and showed a dramatic increase in staining intensity in the tumor samples (Figure 7C and Figure 9A). SLC25A11 was lowly expressed in healthy epithelia but was highly increased in tumor samples. Expression changes in electron transport chain components between healthy tissue and primary tumors are also shown in Figure 10. The results confirm the upregulation of these metabolic enzymes in primary breast tumors. In addition, the inventors selected 12 proteins from the downregulated set and 24 proteins from the upregulated set and examined their expression in the Human Protein Atlas database based on IHC on tissue microarrays (Uhlen et al., 2015). Nine ER- positive, Her2-negative cases were examined against up to three cases of healthy breast ducts. Representative figures are shown in Figures 9B, 9C. Ten of the downregulated proteins and twelve of the upregulated proteins showed a pattern of staining that matches the proteomic findings (Figures 9D, 9E, respectively). Presumably, some of the discrepancies result from the semiquantitative nature of IHC and potential non-specific binding of some of the antibodies. Importantly, only one of the antibodies (against POSTN) showed strong extracellular staining and no staining of epithelial cells, and six additional antibodies showed mild involvement of extracellular staining. These results indicate that macrodis section of the tissues allowed capturing of the cancer-related proteome, with minor effects of the extracellular proteins.
Finally, the inventors compared the protein expression data of five enriched categories to mRNA expression data published by the Cancer Genome Atlas (Network, 2012). The expression values of 22 healthy controls and 358 luminal tumors ware analyzed (Figure 11). Interestingly, it was found that the mRNA expression of ribosomal proteins is reduced in tumors while protein levels are significantly elevated, in agreement with known post-transcriptional regulation on ribosome synthesis (Henras et al., 2008). An opposite trend is seen in glycolytic proteins and DNA repair proteins. Oxidative phosphorylation and lysosome gene expression was elevated on both transcript and protein levels, however protein increases were significantly larger than mRNA. Taken together, these data demonstrate only partial protein-RNA concordance in both protein production and metabolic pathways.
EXAMPLE 5
Changes in protein expression during cancer progression
A much more challenging task, however, was to identify proteins and processes that discriminate between LNN and LNP primary tumors, and between LNP tumors and their matched metastases. To verify that the proteomic changes between those genuinely represent aggressiveness and not the existence of different luminal subtypes (luminal A and B), the inventors examined the levels of ki67, a marker of luminal B tumors, and found similar levels of Ki67 in LNN compared to LNP (Figure 12A). The inventors identified 78 proteins upregulated and 8 downregulated proteins in LNP tumors compared to LNN (Table 2, Welch t-test, FDR=0.05).
Network analysis of these significantly changing proteins showed that the processes associated with the elevated cancer stage mirror those of tumor formation (compared to the healthy tissues; Figure 13A). Similarly, 27 of the upregulated proteins were structural ribosomal proteins; 7 mitochondrial proteins were upregulated, including the mitochondrial malate dehydrogenase (MDH2). Splicing machinery was also enriched, including the spliceosome itself, (Fisher Exact test, FDR=0.02) and additional splicing proteins - particularly in the context of ribosomal RNA processing.
Interestingly, the expression of fifteen proteins was significantly altered both between healthy tissue and primary LNN tumor and between LNN and LNP primary lesions (Figure 12B). Twelve of them, mostly ribosomal proteins, were doubly upregulated; the cytoskeletal-linking protein plectin (PLEC) was doubly downregulated; and methionyl aminopeptidase (METAP)-2 and purine nucleoside phosphorylase (PNP) were downregulated in LNN tumors but upregulated in LNP tumors.
The inventors next examined the metastatic tissue from lymph nodes in comparison to the primary tumors. In an unsupervised clustering, twelve cases of matched tissues (from the same patient) clustered together, with a significantly higher correlation compared to all other unmatched tumor- metastasis couples (average 0.75 vs. 0.58) (Figures 13B and 13C). The 563 proteins that were significantly upregulated in tumor tissue compared to healthy tissue and the 406 downregulated proteins showed a similar median expression in lymph node metastases compared to the primary tumors (Figure 13D). In contrast to hundreds of proteins differentially expressed between healthy and tumor tissues, only four proteins were significantly upregulated and six were downregulated in the LN (Table 3). Four of the downregulated proteins - POSTN, COL12A1, CDKN2A and LGALS 1 - displayed a bidirectional trend (increased in tumor and decreased in metastases) (Figure 13E). Given the extracellular staining of POSTN and COL12A1 (based on the Human Protein Atlas database), reduced expression levels probably stems from the change in the microenvironment of the lymph nodes compared to the breast. Together, these data demonstrate that the proteomic profile of lymph node metastases is highly similar to that of the primary tumors as the primary site of cancer dissemination. Table 2: proteins that were significantly changed in expression between LNN primary tumors and LNP primary tumors (Welch's t-test, FDR=0.05)
Gene Welch test Welch test
names Protein names Protein IDs p-value Difference cluster
CAPS Calcyphosin Q13938;K7ES72 0.001249 1.9314532 Up in LNP
Acyl-coenzyme A Q86TX2;E9 L42; G3V4F2;
thioesterase 1 ; B7ZMC1 ;
Acyl-coenzyme A AOA087XOW7; P49753;
ACOT1 ; thioesterase 2, B4DV16; A1 L172;
ACOT2 mitochondrial A0A087WT95; B3KSA0 0.000343 1.2373475 Up in LNP
Small nuclear
ribonucleoprotein G;
mall nuclear P62308; A8MWD9; Q49AN9;
ribonucleoprotein G-like F5H013;
SNRPG protein C9JVQ0 0.001515 1.166724 Up in LNP
40S ribosomal protein
S27-like;40S ribosomal
RPS27L protein S27 Q71UM5; H0YMV8; C9JLI6 0.001737 1.16631 12 Up in LNP
U6 snR A-associated
LSM2 Sm-like protein LSm2 Q9Y333 1.89E-05 1.15861 17 Up in LNP
NIT1 Nitrilase homolog 1 Q86X76; B7Z410; B2R8D1 5.56E-05 1.0581136 Up in LNP
S4R435 0.000151 1.0532735 Up in LNP
PRRC1 Protein PPvRCl Q96M27 0.000354 0.9378307 Up in LNP
39S ribosomal protein Q8N983; B 1AL05; H0Y6Y8;
MRPL43 L43, mitochondrial H0YBU8; M0R051 0.001193 0.910031 1 Up in LNP
RNA -binding protein Q8IXT5; B9ZVT1 ; E5RHG1 ;
RBM12B 12B E5RJ83; E5RJV8; E5RJW8 0.000343 0.9065059 Up in LNP
RSU1 Ras suppressor protein 1 Q15404; Q32Q10; B0YJ73 0.001025 0.867636 Up in LNP
Q01085; Q2TSD2; 015187;
A8 5C4; E7ETJ9; E7ETC0;
A6NKZ9; B4DHS3; Q59G49;
TIAL1 Nucleolysin TIAR Q49AS9; F8WE 16 5.94E-05 0.8555785 Up in LNP
I3L0E3; Q9Y2R5; Q8IY71 ;
E9PE17; Q96SE7; Q03938;
Q 15929; Q5FWF6; F8WEN0;
C9JUE3; Q4G0J2; I3L137;
K7ENX6; E5RGC2; F8WBQ5;
Q4G 1C2; E9PPF8; E9PIN4;
E9PLW9; C9K0H2; M0QYA6;
M0QYM6; J3KQN0; F6W2C3;
M0QYV6; M0QZG9; E9PLF5;
K7ERS7; K7EL03; M0QXZ2;
7EM14; C9J5H1 ; F8WDJ7;
E9PMX5; E7EVQ0; C9J487;
28 S ribosomal protein B3KX01 ; F2Z3A3; Q0P6G 1 ;
MRPS 17 S I 7, mitochondrial K7EQQ2; J3Q Y7; E9PNM0; 2.2E-05 0.7960314 Up in LNP F2Z2N8; 7ELI5; E9PSE6;
K7ES26; M0QY27; E9PIT0;
E9PIY8; B2RBC6; A8 1S9;
A0A024R4L7
Putative RNA-binding
RBM3 protein 3 P98179; A0A024QYX3 0.000955 0.7845155 Up in LNP
Transcription factor
BTF3L4 BTF3 homolog 4 Q96 17; Q6PJ77; E9PL10 0.001782 0.7590517 Up in LNP
P15927; B2R7E8; B4DUL2;
Replication protein A B4DQD9; B4DL94; Q5TEJ0;
RPA2 32 kDa subunit Q5TEJ7 0.000784 0.7423794 Up in LNP
Cytochrome c oxidase
COX6B 1 subunit 6B1 P14854; K7EQD3 0.000314 0.7420914 Up in LNP
TRA2A;H Transformer-2 protein
U53209 homolog alpha Q13595; B4DQI6; Q549U1 0.000489 0.732815 Up in LNP
U6 snRNA-associated V9GZ56; Q9Y4Z0; U3 QS7;
LSM4 Sm-like protein LSm4 U3KQK1 ; M0QXB0 0.000699 0.720337 Up in LNP
Partner of Y14 and
WIBG mago Q9BRP8 0.000649 0.7169853 Up in LNP
Methionine
aminopeptidase P50579; B4DUX5; G3V1U3;
2;Methionine B3KWL6;
METAP2 aminopeptidase F8VSC4 0.000705 0.7160562 Up in LNP
RPS23 40S ribosomal protein P62266; A8 517; D6RD47;
S23 D6RDJ2; D6RIX0; D6R9I7 0.000234 0.7073038 Up in LNP
1 ,2-dihydroxy-3-keto-5- methy lthiop entene
ADI1 dioxygenase Q9BV57; H7C382 0.001961 0.6954838 Up in LNP
Mitochondrial import
inner membrane
translocase subunit
TIMM13 Tim 13 Q9Y5L4; K7E1T2 0.000946 0.68323 Up in LNP
60S ribosomal protein P83881 ; J3KQN4; H0Y5B4;
RPL36A L36a H7BZ11 ; H0Y3V9 0.000253 0.6821945 Up in LNP
Mitochondrial import
inner membrane
translocase subunit
Tim23;Putative
mitochondrial import
inner membrane 014925; Q5SRD1 ; B 1APJ0;
TIMM23;' translocase subunit B4DDK6; B7ZB25; B4DI18;
MM23B Tim23B B4DHQ8 0.001884 0.6775184 Up in LNP
40S ribosomal protein P62847; E7ETK0;
RPS24 S24 A0A087WUS0 1.77E-06 0.6714652 Up in LNP
28S ribosomal protein P82930; C9JJ19; A4UCR9;
MRPS34 S34, mitochondrial A0A087WUZ8 0.001282 0.6419934 Up in LNP
ENSA Alpha-endosulfme 043768; A6NMQ3; Q5T5H1 0.000167 0.6136546 Up in LNP Q9UMY4; Q3SYF1 ;
SNX12 Sorting nexin-12 AOAO87X0R6 0.000841 0.5972419 Up in LNP
Q9GZT3; A0A087WUN7;
SRA stem-loop-interH0YJ4O; G3V4X6; G3V2S9;
acting RNA-binding H0YJW7; H0YJ07; H0YJU7;
SLIRP protein, mitochondrial H0YJI1 5.86E-05 0.5904784 Up in LNP
Ribonuclease P protein P78346; Q5VU10; Q5VU11 ;
RPP30 subunit p30 B4DJR3 0.00054 0.5878007 Up in LNP
ALYREF THO complex subunit 4 Q86V81 2.89E-06 0.577972 Up in LNP
60S ribosomal protein P61513; C9J4Z3; Q6P4E4;
RPL37A L37a E9PEL3; G5E9R3; M0R0A1 0.001 166 0.5755252 Up in LNP
Selenide, water dikinase P49903; Q5T5U6; B4DLS1 ;
SEPHS1 1 Q5T5U7 0.000892 0.5697772 Up in LNP
SWI/SNF-related matrix- associated actin- dependent regulator of
chromatin subfamily D Q92925; J3KMX2; B9EGA3;
SMARCD member 2 J3QWB6 0.001293 0.5696361 Up in LNP
P62829; Q9BTQ7;
A0AO24R1Q8; J3KT29;
60S ribosomal protein J3KTJ3; C9JD32; B9ZVP7;
RPL23 L23 J3QQT9; A4D 142 4.83E-05 0.5590548 Up in LNP
Q3ZCM7; A0A075B736;
TUBB8 Tubulin beta-8 chain Q5SQY0; A0A075B724 0.002051 0.5583782 Up in LNP
Proteasome subunit beta
PSMB4 type-4 P28070; B4DFL3 2.81E-05 0.5524351 Up in LNP
Vesicle-associated
VAMP 8 membrane protein 8 Q9BV40; B8ZZT4; C9JXZ5 0.001562 0.5328731 Up in LNP
Signal recognition
SRP9 particle 9 kDa protein P49458; Q8WVW9 0.000932 0.5328677 Up in LNP
Multiple myeloma
MMTAG2 tumor-associated protein Q9BU76 0.001836 0.5294666 Up in LNP
PHD finger- like domain-
PHF5A containing protein 5A Q7RTV0; A0A024R1 U2 0.001415 0.5197709 Up in LNP
Pre-mRNA-splicing
BCAS2 factor SPF27 075934; B2R7W3; Q53HE3 0.001232 0.5191518 Up in LNP
P61353; E4W6B6;
A0A024R1V4; K7ELC7;
60S ribosomal protein B2R4D8 ; 7EQQ9; Q6LCU2;
RPL27 L27 K7ERY7 0.00174 0.5151735 Up in LNP
ER membrane protein
EMC3 complex subunit 3 Q9P0I2; S4R3U9 0.001 129 0.4913836 Up in LNP
40S ribosomal protein P63220; Q8WVC2; Q6FGH5;
RPS21 S21 Q9BYK1 0.000674 0.488004 Up in LNP
CSTB Cystatin-B P04080; Q76LA1 0.000353 0.4818774 Up in LNP
40S ribosomal protein P62263; H0YB22; E5RH77;
RPS14 S14 A4D1M5 5.47E-05 0.477388 Up in LNP 40S ribosomal protein P42677; Q5T4L4; A4D1G5;
RPS27 S27 C9J1C5 0.000426 0.4710156 Up in LNP
40S ribosomal protein
RPS29 S29 P62273; AOA087WTT6 0.000187 0.4664468 Up in LNP
60S ribosomal protein P18077; C9K025; F8WBS5;
RPL35A L35a F8WB72 0.000364 0.4491716 Up in LNP
Endothelial
differentiation-related
EDF1 factor 1 060869 0.001 107 0.4469682 Up in LNP
40S ribosomal protein
RPS28 S28 P62857;B 2R4R9 7.12E-05 0.4457975 Up in LNP
Q07955; J3KTL2; Q59FA2;
Serine/arginine-rich A8 1L8; J3KSR8; J3QQV5;
SRSF1 splicing factor 1 J3 SW7 0.000746 0.4187736 Up in LNP
P62899; B2R4C 1 ; H7C2W9;
60S ribosomal protein C9JU56; B7Z4E3; B7Z4C8;
RPL31 L31 B8ZZK4; Q76N53 0.001968 0.4160294 Up in LNP
60S ribosomal protein
L19;Ribosomal protein P84098; J3QR09; J3KTE4;
RPL19 L19 Q53G49; Q8IWR8; J3QL15 0.000254 0.4138308 Up in LNP
RPS15A; P62244; B2R4W8; I3L3P7;
hCG_ 40S ribosomal protein A8K7H3; I3L246; H3BV27;
1994130 S15a H3BT37; I3L303; H3BVC7 0.000269 0.41 10285 Up in LNP
P51 148; A0A024R1U4;
K7ERI8; F8VV 3; 7ENY4;
Ras-related protein Rab- F8VWU4; F8VSF8; K7EIP6;
RAB5C 5C F8VWZ7 0.000695 0.4055087 Up in LNP
RPSA; P08865; C9J9 3;
LAMR 40S ribosomal protein A0A024R2P0; Q96RS2;
1P15 SA A0A024R7P5; A0A024RCF2 5.7E-05 0.3915079 Up in LNP
HSPE1 ; 10 kDa heat shock P61604; Q9UNM1 ; B8ZZL8;
EPFP1 protein, mitochondrial A0A024R3X7; B8ZZ54 0.001554 0.3817505 Up in LN
Purine nucleoside P00491 ; V9HWH6; Q8N7G1 ;
PNP phosphorylase G3V5M2; G3V2H3; G3V393 0.000577 0.3809596 Up in LNP
Cytochrome c oxidase
subunit 4 isoform 1, PI 3073; H3BN72; H3BNV9;
COX4I 1 mitochondrial H3BPG0; Q86WV2; H3BNI5 0.000856 0.369895 Up in LNP
40S ribosomal protein P62241 ; Q5JR94; Q5JR95;
RPS8 S8 Q9BS 10 3.49E-05 0.3559486 Up in LNP
60S ribosomal protein P62888; A0A024R9D3;
RPL30 L30 E5RI99; E5RJH3 0.001539 0.3482403 Up in LNP
60S ribosomal protein P62917; E9PKZ0; E9P U4;
RPL8 L8 G3V1A1 ; B4DVG7; E9PP36 0.001 128 0.3372503 Up in LNP
Malate dehydrogenase, P40926; Q75MT9; Q6FHZ0;
mitochondrial;Malate A0A024R4K3; Q0QF37;
MDH2 dehydrogenase G3XAL0; B3KTM1 0.001902 0.3348827 Up in LNP EIF4H; Eukaryotic translation Q15056; Q75MU1 ; B4DMV6;
WBSCR1 initiation factor 4H Q75MU2; Q75MT8 0.00165 0.3300902 Up in LNP
RPL12;
hCG_ 60S ribosomal protein
21 173 L12 P30050; Q59FI9; D3DS95 0.001253 0.3283965 Up in LNP
Ubiquitin-conjugating
UBE2V1 enzyme E2 variant 1 Q 13404 0.00033 0.3263813 Up in LNP
P63244; E9KL35; J3KPE3;
D6RAC2; D6REE5; D6RHH4;
H0Y8W2; HO YAF8 ;H0 YAM7 ;
D6R9Z1 ; D6R9L0; D6RFX4;
D6RAU2; E9PD14; D6RBD0;
D6RFZ9; B4DVD2; B4E0C3;
Guanine nucleotide- H0Y8R5; D6R909; B4DWC6;
binding protein subunit D6RGK8; D6RHJ5; H0Y9P0;
GNB2L1 beta-2-like 1 D6RDI0; I3QNU9 0.000154 0.3186193 Up in LNP
Q 14157;B4DYY5;F8 W726;
Ubiquitin-associated Q5VU77;Q5VU79;Q5VU78;
UBAP2L protein 2-like Q5VU81 ;Q5 VU80;H0Y5H6 0.001808 0.3159722 Up in LNP
P15880; Q6IPX5; Q3KQT6;
Q8N5L9; H0YEN5; Q8J014;
E9PQD7; E9PMM9; I3L404;
RPS2; E9PM36; E9PPT0; Q8NI61 ;
rps2; Q9BSW5; 060249; D3DU83;
OK/KNS- 40S ribosomal protein A4D0Y7; H3BNG3; H0YE27;
cl.7 S2 Q6Z N8 0.00193 0.3032843 Up in LNP
Proteasome subunit
PSMA7; alpha type-7;
hCG_ Proteasome subunit 014818; Q05DH1 ; H0UI83;
41772 alpha type H0Y586; F5GY34 0.00051 1 0.2924873 Up in LNP
Q02543; M0R3D6; MORI A7;
MORI 17; B4DM74; B2R4C0;
60S ribosomal protein M0R0P7; B4DM94; B4DUV3;
RPL18A L18a Q76N54 0.001912 0.278734 Up in LNP
P62701 ; Q96IR1 ; B2R491 ;
Q53HV1 ; Q8TD47; A6NH36;
40S ribosomal protein P22090; C9JEH7; A4FU1 1 ;
RPS4X S4, X isoform Q496E4; B7Z1M6; Q53H16 0.000556 0.277622 Up in LNP
P05388; A8K4Z4; A0A024RBS
Q6NSF2; F8VWS0; Q53HK9;
Q53HW2; F8VU65; F8VPE8;
60S acidic ribosomal G3V210; F8VW21 ; F8VQY6;
protein P0;60S acidic F8VRK7; Q8NHW5; B4E3D5;
RPLPO; ribosomal protein P0- F8VWV4; F8VS58; F8W1 K8;
RPLP0P6 like Q3MHV2 0.000576 0.2730991 Up in LNP Eukaryotic initiation fact
4A-III;Eukaryotic
initiation factor 4A-III, P38919; A0A024R8W0;
EIF4A3 N-terminally processed I3L3H2 0.00044 0.2634808 Up in LNP
PRIC295;< Translational activator E1NZA1 ; Q92616; Down in
CN1L1 GCN1 AOA024RBS 1 ; B4DM32 0.001847 -0.3358287 LNP
Q15149; D3DWL0; Q96IE3:
PLEC Plectin E9PIA2; E9PQ28 0.000967 -0.4230442 Down in LNP
Q16658; B3KTA3; B3KTM9;
FSCN1 Fascin C9JFC0; C9JPH9 0.001916 -0.4929755 Down in LNP
Ribosome biogenesis
regulatory protein
RRS1 homolog Q15050 0.001964 -0.5547892 Down in LNP
ATP-binding cassette
ABCD3 sub-family D member 3 P28288; B4DL07; B4DZ22 0.001412 -0.6309443 Down in LNP
B4DJ65; Q86XR1 ; A2A2V1;
Q6FGR8; Q6FGN5; Q53YK7;
P04156; Al YVW6; Q6SES1;
075942; D4P3Q7; B2R5Q9;
PRNP Major prion protein B3KQX7; B4DI53; B4DDS 1 0.001985 -0.8339759 Down in LNP
P02533; CON_P02533;
Keratin, type I Q13092; K7ENV3 ;Q7Z3Y9;
RT14 cytoskeletal 14 CON_Q7Z3Y9 0.002037 -1.0550838 Down in LNP
Table 3: proteins that were significantly changed in expression between primary tumors and lymph node metastases (Welch's t-test, FDR=0.05)
Gene Welch test Welch test
names Protein names Protein IDs Difference Difference Cluster
Down in
POSTN Periostin Q 15063; AO A024RDS2 -1.58284 -1.58284 LN metastases
Q96CW1 ; A0A087WY71 ;
E9PFW3; B4DNB9; B4E304;
B4DFM1 ; B4DJB 1 ; C9JJ47;
Down in
AP-2 complex subunit B4DTI4; B7Z4N2; C9JJD3;
AP2M1 mu C9JTK4; H7C4C3 -0.469366 -0.469366 LN metastases
Down in
P20774; Q7Z532; A8KOR3;
OGN Mimecan B4DI63; Q5TBF5 -3.46606 -3.46606 LN metastases
Q99715; D6RGG3;
Down in
Collagen alpha- 1 (XII) A0A087X0A8; B9EJB8;
COL12A1 chain H0Y4P7; H0Y991 -1.70424 -1.70424 LN metastases
P42771 ; K7PML8; J3QRG6
Cyclin-dependent kinase L8E941 ;D1 LYX3; 7ES20
Down in inhibitor 2A, isoforms K7ENC6; R9S252; Q208B5
CDKN2A 1/2/3 G3XAG3 ; Q9UPB7; Q2MJE o -0.761955 -0.761955 LN metastases
Down in
LGALS 1 Galectin-1 P09382; F8WEI7 -0.702194 -0.702194 LN metastases
Synaptosomal-associatec
protein 23; 000161 ; A8K287;
Up in
Synaptosomal-associatec A0A024R9R8; H3BM38;
SNAP23 protein H3BP15; H3BQY9 0.885205 0.885205 LN metastases
Up in
Adenine phosphoribosyl- P07741 ; H3BQZ9; H3BQB1 ;
APRT transferase H3BQF1 ; H3BSW3; Q12898 0.361784 0.361784 LN metastases
P21980; V9HWG3; B4DIT7;
Protein-glutamine B4DTN7; A2A299; Q6DKH2;
Up in gamma- A2A2A0; Q9H035; B4YUQ3;
TGM2 glutamyltransferase 2 B4YUQ2 0.580455 0.580455 LN metastases
Q9H0W9; A8K718;
A0A024R396; A0A087WT99;
A0A024R3B0; E9PQS 1 ;
E9PPB5; E9PIP1 ; E9PSC3;
Up in E9PLC5; E9PLB3; E9PJU8;
Cl lorf54 Ester hydrolase CI lorf5< E9PR95 0.863908 0.863908 LN metastases EXAMPLE 6
A 15-protein signature predicts lymph node involvement
Finally, the inventors tested if the 85 significantly changing proteins between LNN and LNP primary tumors, or a subset thereof, can also predict lymph node involvement. Lymph node status is a crucial component in treatment decision making, but currently there are no reliable ways to determine node involvement without physically evaluating the lymph node.
The inventors trained a classifier using support vector machine (SVM) algorithm on all primary tumor samples (n=21 for LNN, n=20 for LNP). For sample cross-validation, the inventors used random subsampling of the data. 85 percent of the samples were used as a training set, and classification was subsequently tested on the remaining 15 percent, done recursively for 250 iterations. Feature selection was done by recursive feature elimination (RFE), where features with low weights were removed after every iteration (Barla et al., 2008).
In order to test the classifier performance, the inventors plotted the false positive (FP; prediction of LNN as LNP) and false negative (FN; prediction of LNP as LNN) rates, together with the area under the receiver operating characteristic curve (AUC of ROC) generated by the classifier, using increasing number of features for classification (Figure 14A and Table 4). The optimal performance of the classifier using a minimum number of features, denoted by a low false incidents and high AUC, was reached using 15 proteins for classification (FP incidents: 3 out of 21; FN incidents: 4 out of 20; AUC=0.93, Figure 14B). These 15 proteins were upregulated in LNP tumors and represent several cellular functions (Figure 14C), including ribosomal proteins (RPS24 and RPS29), proteins involved in pre-mRNA splicing (LSM2 and LSM4) and RNA binding proteins (RBM3 and RBM12B), and also include the two possible invasiveness markers METAP2 and PNP. This proteomic signature may be further translated to clinical tests for early identification of aggressive luminal breast cancer. Table 4: protein signature that segregates LNN from LNP primary tumors. All proteins are significantly changing in expression between LNN and LNP primary tumors (Welch's t-test, FDR=0.05)
Gene
names Protein names Protein IDs Ranks SEQ ID NO:
RPS24 40S ribosomal protein S24 P62847; E7ET 0; A0A087WUS0 0 1
U6 snRNA-associated V9GZ56; Q9Y4Z0; U3KQS7;
LSM4 Sm-like protein LSm4 U3 QK1 ; M0QXBO 1 2
Q8IXT5; B9ZVT1 ; E5RHG1 ;
BM12B RNA-binding protein 12B E5RJ83; E5RJV8; E5RJW8 2 3
RPS29 40S ribosomal protein S29 P62273; A0A087WTT6 J 4
RJBM3 Putative RNA-binding protein 3 P98179; A0A024QYX3 4 5
P00491 ; V9HWH6; Q8N7G1 ;
PNP Purine nucleoside phosphorylase G3V5M2; G3V2H3; G3V393 5 6
Methionine aminopeptidase 2; P50579; B4DUX5; G3V1U3;
METAP2 Methionine aminopeptidase B3KWL6; F8VSC4 6 7
CAPS Calcyphosin Q 13938; K7ES72 7 8
Eukaryotic initiation factor 4A-III;
Eukaryotic initiation factor 4A-III,
EIF4A3 N-terminally processed P38919; A0A024R8W0; I3L3H2 8 9
SNX12 Sorting nexin-12 Q9UMY4; Q3SYF1; AOA087XOR6 9 10
Q07955;J 3KTL2; Q59FA2;
SRSF1 Serine/arginine-rich splicing factor 1 A8K1 L8; J3KSR8; J3QQV5; J3 SW7 10 1 1
P51148; A0A024R1U4; 7ERI8;
F8VVK3; 7ENY4; F8VWU4;
RAB5C Ras-related protein Rab-5C F8VSF8; 7EIP6; F8VWZ7 1 1 12
BTF3L4 Transcription factor BTF3 homolog 4 Q96K17; Q6PJ77; E9PL10 12 13
U6 snRNA-associated
LSM2 Sm-like protein LSm2 Q9Y333 13 14
CSTB Cystatin-B P04080; Q76LA1 14 15

Claims

CLAIMS:
1. A diagnostic or prognostic method for determining the progression of breast cancer in a subject, the method comprising:
a. determining the expression level of at least one biomarker protein in at least one biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein/s, wherein said biomarker proteins are selected from RPS24 (40S ribosomal protein S24), LSM4 (U6 snRNA-associated Sm-like protein LSm4), RBM12B (RNA-binding protein 12B), RPS29 (40S ribosomal protein S29); RBM3 (Putative RNA-binding protein 3), PNP (Purine nucleoside phosphorylase), METAP2 (Methionine aminopeptidase 2; Methionine aminopeptidase), CAPS (Calcyphosin), EIF4A3 (Eukaryotic initiation factor 4A-III;Eukaryotic initiation factor 4A- III, N-terminally processed), SNX12 (Sorting nexin-12), SRSF1 (Serine/arginine-rich splicing factor 1), RAB5C (Ras-related protein Rab-5C), BTF3L4 (Transcription factor BTF3 homolog 4), LSM2 (U6 snRNA-associated Sm-like protein LSm2) and CSTB (Cystatin-B), or any combination thereof; and
b. determining if the expression value obtained in step (a) is_positive or negative with respect to a predetermined standard expression value or to an expression value of said biomarker protein/s in at least one control sample;
wherein at least one of: (i) a positive expression value of said at least one biomarker protein/s in said sample, indicates that said subject belongs to a pre-established population associated with positive lymph node metastatic status; and (ii) a negative expression value of said at least one biomarker protein/s in said sample, indicates that said subject belongs to a pre-established population associated with negative lymph node metastatic status.
2. The method according to claim 1, wherein in step (a) the expression level of at least five biomarker proteins is determined in at least one biological sample of said subject, to obtain an expression value for each of said at least five biomarker protein/s, wherein said at least five biomarker proteins are selected from RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB.
3. The method according to any one of claims 1 and 2, wherein determining the level of expression of at least one or of at least five of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s is performed by the step of contacting at least one detecting molecule or any combination or mixture of plurality of detecting molecules with a biological sample of said subject, or with any protein or nucleic acid product obtained therefrom, wherein each of said detecting molecules is specific for one of said biomarker proteins.
4. The method according to claim 3, wherein said detecting molecule/s is selected from amino acid detecting molecules and nucleic acid detecting molecules.
5. The method according to claim 4, wherein said amino acid detecting molecule/s comprise at least one of:
a. at least one labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof;
b. at least one antibody specific for said at least one of said biomarker proteins;
c. at least one peptide aptamer/s specific for said at least one of said biomarker proteins;
d. any combination of (a), (b) and (c).
6. The method according to claim 4, wherein said nucleic acid detecting molecule/s comprise at least one of:
a. at least one nucleic acid aptamer/s specific for said at least one of said biomarker proteins; b. at least one oligonucleotide/s, each oligonucleotide specifically hybridizes to a nucleic acid sequence encoding said at least one biomarker protein.
7. The method according to claim 5, wherein said detecting molecule/s are at least one labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof, and wherein determination of the expression level of said at least one biomarker protein/s is performed by mass spectrometry.
8. The method according to claim 5, wherein said detecting molecule/s are at least one of antibodies, nucleic acid or peptide aptamer/s specific for said at least one of said biomarker protein/s, or any combination thereof, and wherein determination of the expression level of said at least one biomarker protein/s is performed by an immunological assay.
9. The method according to any one of claims 1 and 2, wherein determining the level of expression of at least one or of at least five of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s is performed by the step of subjecting a biological sample of said subject, or any protein product obtained therefrom to a mass spectrometry assay.
10. The method according to any one of claims 1 and 2, wherein said sample is any one of a biological sample of organ/s, cell/s or tissue/s and a blood sample.
11. The method according to claim 10, wherein said sample is a primary tumor sample.
12. The method according to any one of claims 1 and 2, wherein said subject is suffering from a luminal A or a luminal B breast tumor.
13. The method according to any one of claims 1 and 2, for monitoring the efficacy of a treatment with a therapeutic agent and the disease progression, said method comprises the steps of: a. determining the expression level of at least one biomarker protein in a biological sample of said subject, to obtain an expression value for each of said at least one biomarker protein/s, wherein said biomarker protein/s are selected from RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB or any combination thereof;
b. repeating step (a) to obtain expression values of said at least one biomarker protein/s, for at least one more temporally- separated test sample; wherein at least one of said temporally separated samples is obtained after the initiation of said treatment;
c. calculating the rate of change of said expression values of said at least one biomarker protein between said temporally-separated test samples;
d. determining if the rate of change obtained in step (c) is positive or negative with respect to a predetermined standard rate of change determined between at least two temporally separated samples or to the rate of change calculated for expression values in at least one control sample obtained from at least two temporally separated samples, wherein at least one sample of said at least two samples is obtained after the initiation of said treatment;
Wherein a negative rate of change of the expression value of at least one of said biomarker protein/s indicates that said subject exhibits a beneficial response to said treatment; thereby monitoring the efficacy of a treatment with a therapeutic agent and the disease progression.
14. The method according to claim 13, wherein in step (a) the expression level of at least five biomarker proteins is determined in at least one biological sample of said subject, to obtain an expression value for each of said at least five biomarker protein/s, wherein said at least five biomarker proteins are selected from RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB.
15. The method according to any one of claims 13 and 14, wherein determining the level of expression of at least one or of at least five of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s is performed by the step of contacting at least one detecting molecule or any combination or mixture of plurality of detecting molecules with a biological sample of said subject, or with any protein or nucleic acid product obtained therefrom, wherein each of said detecting molecules is specific for one of said biomarker proteins.
16. The method according to any one of claims 13 and 14, wherein determining the level of expression of at least one or of at least five of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s is performed by the step of subjecting a biological sample of said subject, or any protein product obtained therefrom to a mass spectrometry assay.
17. A diagnostic or prognostic composition comprising at least one detecting molecule or any combination or mixture of plurality of detecting molecules specific for determining the level of expression of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof, wherein each of said detecting molecules is specific for one of said biomarker protein/s.
18. The composition according to claim 17, wherein said composition comprises at least one detecting molecule or any combination or mixture of plurality of detecting molecules specific for determining the level of expression of at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof, wherein each of said detecting molecules is specific for one of said biomarker proteins.
19. The composition according to any one of claims 17 and 18, wherein said detecting molecules are selected from amino acid detecting molecules and nucleic acid detecting molecules.
20. The composition according to claim 19, wherein said amino acid detecting molecules comprise at least one of:
a. at least one labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof;
b. at least one antibody specific for said at least one of said biomarker protein/s;
c. at least one peptide aptamer/s specific for said at least one biomarker protein/s;
d. any combination of (a), (b) and (c).
21. The composition according to claim 19, wherein said nucleic acid detecting molecule comprise at least one of:
a. at least one nucleic acid aptamer/s specific for said at least one biomarker proteins;
b. at least one oligonucleotide/s, each oligonucleotide specifically hybridizes to a nucleic acid sequence encoding said at least one biomarker protein/s.
22. The composition according to any one of claims 17 and 18, wherein said detecting molecules are attached to a solid support.
23. The composition according to any one of claims 17 and 18, wherein said composition further comprises a biological sample.
24. A kit comprising:
a. at least one detecting molecule specific for determining the level of expression of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof in a biological sample, wherein each of said detecting molecule/s is specific for one of said biomarker proteins; said kit optionally further comprises at least one of:
b. pre-determined calibration curve/s or predetermined standard/s providing standard expression values of said at least one biomarker/s;
c. at least one control sample.
25. The kit according to claim 24, comprising at least one detecting molecule specific for determining the level of expression of at least five of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any combination thereof in a biological sample, wherein each of said detecting molecule/s is specific for one of said biomarker proteins; said kit optionally further comprises at least one of: b. pre-determined calibration curve/s or predetermined standard/s providing standard expression values of said at least one biomarker protein/s;
c. at least one control sample.
26. The kit according to any one of claims 24 and 25, wherein said detecting molecules are selected from amino acid detecting molecule/s and nucleic acid detecting molecule/s.
27. The kit according to claim 26, wherein said amino acid detecting molecules comprise at least one of:
a. at least one labeled or tagged RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB protein/s or any fragment/s, peptide/s or mixture thereof;
b. at least one antibody specific for said at least one of said biomarker proteins;
c. at least one peptide aptamer/s specific for said at least one of said biomarker protein/s;
d. any combination of (a), (b) and (c).
28. The kit according to claim 26, wherein said nucleic acid detecting molecule comprise at least one of: a. at least one nucleic acid aptamer/s specific for said at least one biomarker proteins;
b. at least one oligonucleotides, each oligonucleotide specifically hybridizes to a nucleic acid sequence encoding said at least one biomarker protein/s.
29. The kit according to any one of claims 24 and 25, wherein said detecting molecule/s are attached to a solid support.
30. The kit according to any one of claims 24 and 25, wherein said detecting molecule/s is provided in a mixture.
31. The kit according to any one of claims 24 and 25, further comprising instructions for use, wherein said instructions comprise at least one of:
a. instructions for carrying out the detection and quantification of the expression of said at least one of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s and optionally, of a control reference protein; and
b. instructions for comparing the expression values of at least one of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB with a corresponding predetermined standard expression value or with expression value of at least one of said biomarker protein/s in said at least one control sample.
32. The kit according to any one of claims 24 and 25, further comprising at least one reagent for conducting a mass spectrometry assay.
33. The kit according to any one of claims 24 and 25, further comprising at least one reagent for conducting an immunological assay selected from protein microarray analysis, ELISA, RIA, slot blot, dot blot, FACS, western blot, immunohistochemical assay, immunofluorescent assay and a radio-imaging assay.
34. The kit according to any one of claims 24 and 25, for use in a method for determining the progression of breast cancer in a subject.
35. The kit according to claim 34, wherein said subject is suffering from a luminal A or a luminal B breast tumor.
36. The kit according to any one of claims 24 and 25, wherein said sample is any one of a biological sample of organ/s, cell/s or tissue/s and a blood sample.
37. The kit according to claim 36, wherein said sample is a primary tumor sample.
38. A method for assessing expression status of RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSFl, RAB5C, BTF3L4, LSM2 and CSTB in a biological sample, the method comprising the step of : a. providing at least one detecting molecule, any combination, mixture of plurality of detecting molecules or any composition of kit comprising the same, wherein each of said detecting molecules is specific for one of said RPS24, LSM4, RBM12B, RPS29, RBM3, PNP, METAP2, CAPS, EIF4A3, SNX12, SRSF1, RAB5C, BTF3L4, LSM2 and CSTB biomarker protein/s;
b. contacting said at least one detecting molecule/s provided in (a) with a biological sample or with any protein or nucleic acid product obtained therefrom;
c. performing a protein or nucleic acid detection assay.
PCT/IL2016/050480 2015-05-06 2016-05-05 Methods and kits for breast cancer prognosis Ceased WO2016178236A1 (en)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
US201562157723P 2015-05-06 2015-05-06
US62/157,723 2015-05-06
US201662289525P 2016-02-01 2016-02-01
US62/289,525 2016-02-01

Publications (1)

Publication Number Publication Date
WO2016178236A1 true WO2016178236A1 (en) 2016-11-10

Family

ID=57218183

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/IL2016/050480 Ceased WO2016178236A1 (en) 2015-05-06 2016-05-05 Methods and kits for breast cancer prognosis

Country Status (1)

Country Link
WO (1) WO2016178236A1 (en)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108753982A (en) * 2018-08-01 2018-11-06 徐建震 A kind of cancer of the esophagus prognosis biomarker and its detection method and application
CN113252895A (en) * 2020-02-10 2021-08-13 首都医科大学附属北京世纪坛医院 Application of serum cathepsin D in lymphedema diseases
CN113667748A (en) * 2021-07-19 2021-11-19 中山大学肿瘤防治中心(中山大学附属肿瘤医院、中山大学肿瘤研究所) Inhibitor of circIKB and application of detection reagent thereof in kit for diagnosis, treatment and prognosis of breast cancer bone metastasis

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2014205293A1 (en) * 2013-06-19 2014-12-24 Memorial Sloan-Kettering Cancer Center Methods and compositions for the diagnosis, prognosis and treatment of brain metastasis

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2014205293A1 (en) * 2013-06-19 2014-12-24 Memorial Sloan-Kettering Cancer Center Methods and compositions for the diagnosis, prognosis and treatment of brain metastasis

Non-Patent Citations (5)

* Cited by examiner, † Cited by third party
Title
BERTUCCI, FRANCOIS ET AL.: "Gene expression profiling of primary breast carcinomas using arrays of candidate genes.", HUMAN MOLECULAR GENETICS, vol. 9.20, 12 December 2000 (2000-12-12), pages 2981 - 2991, XP002225994, Retrieved from the Internet <URL:https://jeccr.biomedcentral.com/articles/10.1186/s13046-014-0097-2> *
CAO, YU-WEN ET AL.: "Correlation and prognostic value of SIRT and Notchl signaling in breast cancer.", JOURNAL OF EXPERIMENTAL & CLINICAL CANCER RESEARCH, vol. 33.1 : 1, 25 November 2014 (2014-11-25), XP055326179, Retrieved from the Internet <URL:https://jeccr.biomedcentral.com/articles/10.1186/s13046-014-0097-2> *
HUANG, ERICH ET AL.: "Gene expression predictors of breast cancer outcomes.", THE LANCET, vol. 361.9369, 2003, pages 1590 - 1596, XP004782813 *
SCHMIDT, MARCUS ET AL.: "Prediction of late metastasis in node-negative breast cancer.", ASCO ANNUAL MEETING PROCEEDINGS., 5 June 2012 (2012-06-05), pages 10551, Retrieved from the Internet <URL:http://meetinglibrary.asco.org/content/97502-114> *
YANG, YANG ET AL.: "Clinical implications of high NQO expression in breast cancers.", JOURNAL OF EXPERIMENTAL & CLINICAL CANCER RESEARCH, vol. 33.1, no. 1., 5 February 2014 (2014-02-05), pages 4, Retrieved from the Internet <URL:http://jeccr.biomedcentral.com/articles/10.1186/1756-9966-33-14> *

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108753982A (en) * 2018-08-01 2018-11-06 徐建震 A kind of cancer of the esophagus prognosis biomarker and its detection method and application
CN108753982B (en) * 2018-08-01 2022-05-06 徐建震 Esophageal cancer prognosis biomarker and detection method and application thereof
CN113252895A (en) * 2020-02-10 2021-08-13 首都医科大学附属北京世纪坛医院 Application of serum cathepsin D in lymphedema diseases
CN113667748A (en) * 2021-07-19 2021-11-19 中山大学肿瘤防治中心(中山大学附属肿瘤医院、中山大学肿瘤研究所) Inhibitor of circIKB and application of detection reagent thereof in kit for diagnosis, treatment and prognosis of breast cancer bone metastasis
WO2023000282A1 (en) * 2021-07-19 2023-01-26 中山大学肿瘤防治中心(中山大学附属肿瘤医院 中山大学肿瘤研究所) Circikbkb inhibitor and use of circlkbkb detection reagent in diagnosis, treatment and prognosis kit for breast cancer bone metastasis

Similar Documents

Publication Publication Date Title
US11079384B2 (en) Biomarkers and methods for diagnosis of early stage pancreatic ductal adenocarcinoma
EP2430193B1 (en) Markers for detection of gastric cancer
US10689711B2 (en) Test kits and methods for their use to detect genetic markers for urothelial carcinoma of the bladder and treatment thereof
US20150079078A1 (en) Biomarkers for triple negative breast cancer
DK2638398T3 (en) Hitherto UNKNOWN MARKET FOR THE DETECTION OF BLADE CANCER AND / OR INFLAMMATORY CONDITIONS IN THE BLADE.
WO2012125411A1 (en) Methods of predicting prognosis in cancer
WO2012009382A2 (en) Molecular indicators of bladder cancer prognosis and prediction of treatment response
CN111065925A (en) CTNB1 as a marker for endometrial cancer
WO2016178236A1 (en) Methods and kits for breast cancer prognosis
JP5403534B2 (en) Methods to provide information for predicting prognosis of esophageal cancer
WO2024196872A1 (en) Methods and compositions for diagnosis of ectopic pregnancy
US20240402177A1 (en) Protein markers for estrogen receptor (er)-positive luminal a (la)-like and luminal b1 (lb1)-like breast cancer
JP2007263896A (en) Biomarker and method for predicting postoperative prognosis of lung cancer patients
US20150011411A1 (en) Biomarkers of cancer
US20230059578A1 (en) Protein markers for estrogen receptor (er)-positive-like and estrogen receptor (er)-negative-like breast cancer
EP2607494A1 (en) Biomarkers for lung cancer risk assessment
KR20260044139A (en) Proteomic heterogeneity of the extracellular matrix identifies histologic subtype-specific fibroblast in gastric cancer
KR20240126130A (en) Biomarker for predicting colorectal cancer prognosis, and prognosis prediction method using thereof
Kass et al. Oncogenomics/Proteomics of Head and Neck Cancers
Moskowitz et al. Oncogenomics/Proteomics of Head and Neck Cancers
NZ577012A (en) Human zymogen granule protein 16 as a marker for detection of gastric cancer
WO2009047274A2 (en) Ptpl1 as a biomaker of survival in breast cancer

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16789424

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16789424

Country of ref document: EP

Kind code of ref document: A1