EP4721073A1 - Molecular diagnostic method for classifying a biological sample on the basis of cellularity - Google Patents
Molecular diagnostic method for classifying a biological sample on the basis of cellularityInfo
- Publication number
- EP4721073A1 EP4721073A1 EP24731083.2A EP24731083A EP4721073A1 EP 4721073 A1 EP4721073 A1 EP 4721073A1 EP 24731083 A EP24731083 A EP 24731083A EP 4721073 A1 EP4721073 A1 EP 4721073A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- tumour
- sample
- genes
- training
- cellularity
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B25/00—ICT specially adapted for hybridisation; ICT specially adapted for gene or protein expression
- G16B25/10—Gene or protein expression profiling; Expression-ratio estimation or normalisation
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/20—Supervised data analysis
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/30—Unsupervised data analysis
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/20—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Physics & Mathematics (AREA)
- Data Mining & Analysis (AREA)
- General Health & Medical Sciences (AREA)
- Public Health (AREA)
- Biotechnology (AREA)
- Epidemiology (AREA)
- Theoretical Computer Science (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Databases & Information Systems (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biophysics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Evolutionary Computation (AREA)
- Software Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Bioethics (AREA)
- Biomedical Technology (AREA)
- Artificial Intelligence (AREA)
- Genetics & Genomics (AREA)
- Pathology (AREA)
- Primary Health Care (AREA)
- Molecular Biology (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
The method enables, by analysing the gene expression of a predefined list of genes specific to the cellular tumour condition with a trained binary classifier algorithm, the estimation of the tumours of a biological sample.
Description
MOLECULAR DIAGNOSTIC METHOD FOR CLASSIFYING A BIOLOGICAL SAMPLE ON THE BASIS OF CELLULARITY TECHNICAL FIELD The present invention relates to a method to support in vitro molecular diagnosis based on the use of a digital RNA-seq method. In particular, the invention relates to a method for analysing and classifying biological samples obtained from subjects with suspected cancer, for example breast cancer. Prior Art Cancer is a leading cause of death worldwide, and approximately 18 million new cases worldwide in only 2020 (the Global Cancer Observatory). The most common cancers in both developed and developing countries are breast, lung, colorectal, rectal and prostate. Although the etiology of various types of cancer remains unknown, several factors are responsible for the abnormal cell development. A solid malignant tumor is an abnormal mass of tissue that usually does not contain cysts or areas of fluid. These solid tumors vary according to their site of origin, such as sarcomas, carcinomas, and lymphomas. In the last decade, important scientific and technological advances have allowed us to witness a steady improvement in the treatment and survival of many tumors, including breast cancer patients. The clinical validation and approval of new targeted therapies for oncological diseases in the context of precision medicine is bringing about a significant change in the reduction of all cancer mortality rates, with the prospect of further improvements in patient survival over the next 10 years. In fact, the medical community is facing an unprecedented revolution in the field of oncology with the introduction of various types of agnostic drugs aimed at new therapeutic approaches, including immunotherapy and molecularly targeted drugs. Specifically, the eligibility for treatment of a molecularly targeted agnostic drug is based solely on the molecular characteristics identified in the patient's tumour. In particular, it has been shown that solid tumours are characterised by great molecular heterogeneity, which means that each patient must be considered as a carrier of a unique disease requiring personalised treatment. Thanks to this evidence, a more correct diagnosis of the disease state is now made, directing the physician towards planning a better treatment plan. Tumor-agnostic therapies based on the cancer’s genetic and molecular features without regard to the cancer type or where the cancer started in the body are gaining significant advantages for the treatment of many solid tumors. Recently, HER2 drugs (i.e trastuzumab deruxtecan) have been approved by the Food and Drug Administration (FDA) and European Medicines Agency (EMA) for the treatment of
specific HER2 clinical and molecular features, allowing to specifically target the protein underlying tumour development. In addition to the example reported, there are currently more than 1,000 ongoing studies for breast cancer, identified under the terms 'Targeted Therapy' and 'Breast Cancer', in the ClinicalTrials.gov database and of the 'Interventional Study' type, which use targeted therapies according to tumour characteristics. The introduction of molecular precision medicine technologies, aimed at identifying the 'disease' more accurately, proves to be essential for a correct molecular stratification of the tumour and for framing the correct diagnostic and therapeutic protocols. In general, with the advancement of research, it is possible to identify, through molecular analysis of tumour samples, a plurality of genes whose expression occurs specifically in a tumour e.g. ovarian, prostate, etc. so as to be able to diagnose and prognosticate, sometimes even early, the specific tumour and more generally estimate whether a biological sample is tumoral, and thus needs to be checked more thoroughly, or non-tumoral. To improve the accessibility of targeted therapies in clinical practice, it is necessary to promote the use of personalized medicine diagnostic tools and technologies suitable for future scenario of precision medicine where drugs can be prescribed on the basis of molecular features (companion diagnostics). Among these technologies, multigenic tests represent the only possible alternative and especially those based on Next Generation Sequencing (NGS) are the most promising. Indeed, NGS is currently considered the most comprehensive and reliable technology that has contributed substantially to the advancement of precision medicine in scientific and clinical settings, particularly for the study and characterisation of tumours (Hussen BM, Abdullah ST, Salihi A, Sabir DK, Sidiq KR, Rasul MF, Hidayat HJ, Ghafouri-Fard S, Taheri M, Jamali E. The emerging roles of NGS in clinical oncology and personalised medicine. Pathol Res Pract. 2022 Feb;230:153760. doi: 10.1016/j.prp.2022.153760. Epub 2022 Jan 10. PMID: 35033746.). However, the lack of highly processive NGS-based diagnostic tests stems from some major limitations, including very high costs and the difficulty in accurately measuring degraded nucleic acids derived from molecular pathology preparations. The FFPE (Formalin-Fixed Paraffin-Embedded) tissue preservation method, commonly used in pathological anatomy, allows for the preservation of tissue architecture, morphology, and cellular processes, but causes degradation of nucleic acids, which is physiologically greater for RNA than for DNA, limiting the application of NGS sequencing technologies. These technical limitations make the current NGS-based multigenic solutions unable to achieve a level of accuracy and reproducibility suitable for diagnosis and clinical standards. Moreover, current NGS-based transcriptomic solutions are offered at a higher cost than immunohistochemistry and generate results in a format that is still difficult to interpret. In fact, NGS results measure thousands of genes simultaneously, and these results, if not properly analysed and contextualized, can be difficult to interpret in a way that is clinically meaningful and actionable.
Even if several methodologies have been developed based on other methods such as microarrays, qRT-PCR or targeted-NGS library preparations, there is no method that shows a complete cost/accuracy improvement over immunohistochemistry and no NGS tools has been developed able to substitute immunohistochemistry in clinical practice. A fundamental problem in the molecular analysis of clinical samples is the heterogeneity of the tissue to be analyzed, which is often a mixture of tumor and normal adjacent cell type, including extracellular matrix, epithelial tissue, connective tissue, immune cells, etc. For this reason, all tissue biopsies must be evaluated for tumor content in order to be analyzed by following molecular techniques for tumor characterization. Pathologists rely on the gold standard of color staining of the bioptic tissue by Hematoxylin and Eosin staining (H&E staining), which allows quantitative visual measurement of tumor cellularity. Histopathological diagnosis of tumors from tissue biopsies using hematoxylin and eosin (H&E) staining is the standard of care in oncology. To date, no molecular analysis can be performed without such control of the tumour content of the sample. However, H&E staining is a methodology not without limitations, requires trained operators, dyes and reagents, and valuable tissue samples that cannot be reused. Also, analytical reproducibility is an issue derived mainly from heterogeneity of the tissue cellular types, influencing the coloring reaction and subjectivity of the operator visual estimation method (Smits AJ, Kummer JA, de Bruin PC, Bol M, van den Tweel JG, Seldenrijk KA, Willems SM, Offerhaus GJ, de Weger RA, van Diest PJ, Vink A. The estimation of tumor cell percentage for molecular testing by pathologists is not accurate. Mod Pathol.2014 Feb;27(2):168-74. doi: 10.1038/modpathol.2013.134. Epub 2013 Jul 26. PMID: 23887293.). For this reason several groups have tried to apply H&E together with more reproducible and robust analytical method based on gene expression for tissue analysis (Wai jin TAN et al, Braest Cancer Res. 2016; Y. Yuan et al., Science Translational Medicine, 2012; Marko NF et al, Genomic Academic Press et al 210; DW Scott et al Journal of Clinical Oncology 2012). In particular, Wai Jin Tan et al. (A five-gene reverse transcription-PCR assay for pre- operative classification of breast fibroepithelial lesions - Breast Cancer Res. 2016) developed a 5-gene assay to distinguish fibroadenomas and phyllodes tumors in bioptic tissue. The authors initially profiled the transcriptome of a training set of formalin-fixed paraffin embedded (FFPE) fibroadenomas and phyllodes tumours by a cDNA-mediated Annealing, Selection, extension and Ligation assay, performed with a package specifically designed for microarray analysis (samr: significance analysis of microarray r). They identified and then validated, using qPCR, 43 genes and obtained a final 5 gene model useful for separating fibroadenomas and phyllodes tumours. The limits of said assay relied again on the difficulty to analyze FFPE tissue, extremely degraded on the ribonucleoprotein acid component, recurring on two parallel methods
such as microarray and qRT-PCR. The authors of the paper conclude that their assay should not replace the current diagnostic framework for said tumours. There is therefore still a need for a methodology to support molecular pathologist to assess tumour content of the tissues to be molecularly analysed. Innovative NGS based diagnostics tool that can overcome the limitations of the known art, that can be used on heterogeneous samples, taking into consideration the complexity of each sample, which may contain varying percentages of cancer cells and that is developed with a technique that could analyse the RNA and be superior in terms of sensitivity, resolution, genomic coverage, reproducibility than microarray. PURPOSES AND SUMMARY OF THE INVENTION The aims of the invention are achieved by means of a NGS-derived computer-based method to classify a biological sample comprising the steps of: - Receiving the biological test sample of a subject with a suspected cancer - Receive an indication of the tumour to be tested in the sample biological test - Extract RNA and generate a library of fragments of PolyA+ sequences from the sample RNA - Processing through a sequencer and subsequent alignment to the reference genome and read counts, to obtain an output file of the biological test sample including a plurality of counts of occurrences of corresponding genes represented by the PoliA+ fragments in the sample, each count defining an expression level of the corresponding gene - Apply to the output file a binary classification algorithm trained with biological training samples whose classification as non-tumour or tumour cancer is know to be verified by histological cellularity examinations and whose gene expression levels are known - In which the said trained binary classification algorithm is configured to generate, based on the expression levels in the output file of the biological test sample of a first predefined list of genes (list 1) whose expression levels are indicative of the cellularity of the tumour to be tested, a binary estimate of whether the biological test sample is tumour-like of the tumour to be tested or non-cancerous. Using the binary classifier, it is possible to discriminate gene expression levels of the test sample to estimate tumours or non-tumours in a quantitative and highly repeatable manner. The genes on the list, identified according to the procedure of claim 2 applicable to each tumour whose cellularity classification e.g. by histological examinations is known, show high levels of accuracy. This accuracy is influenced by the value of the
correlation factor e.g. Bayes factor in the sense that the greater the correlation of the gene expression level with tumour cellularity, the more accurate the classification. When training samples are divided into at least two different levels of cellularity, the process is simplified to identify the most specific genes whose variation in expression with cellularity allows high classification accuracy. The use of clustering, in particular hierarchical clustering, reduces the number of genes to be identified during the testing phase and at the same time identifies the most specific genes for the tumour cell condition, making classification faster and more accurate, and also requiring both widely available and inexpensive computational devices and a relatively small number of training samples to provide satisfactory estimates. When the classification algorithm is a GLM, it has the dual benefit of adhering more closely to current molecular classification guidelines based on histopathology of biological samples and of preserving accuracy from training on biological samples of known classification to testing on a biological sample to be examined. The inventors developed a solution using advanced RNA processing technology to precisely measure polyadenylated RNA expression from fresh and pathology inclusion tissues, including FFPE and heavily degraded tissues, exclusively using NGS. This technology has a similar cost per sample to current clinical methods, such as immunohistochemistry, and is seen as promising for overhauling tumour assessment, which is currently based solely on H&E. The inventors have focused their algorithm and study on identifying the features that characterize a heterogeneous tissue derived from molecular pathology biopsies. Such tissue can presumably be composed of tumour cells and other cellular components represented in different abundances (immune cells, epithelial cells, adipose cells, etc.). The tumour must be properly detected to be eligible for following molecular assays, including the analysis of specific tumor biomarkers. For this reason, the described invention will enable to verify whether a tissue obtained from molecular pathology biopsies contains enough abundance of tumoral cells and is suitable for subsequent analysis of tumour biomarkers. The inventors created a highly processive solution, eligible to be integrated with the clinical standard for pathology and oncology. The diagnostic solution is based on NGS technology, specifically 3'Digital Gene Expression (3'DGE) Sequencing, which enables total RNA extraction and sequencing of the polyA near-tail region. This makes sequence analysis rapid and reproducible, demonstrating quantitative accuracy comparable to traditional whole RNA fragment sequencing technologies on intact tissues, while on degraded tissues 3'DGE shows higher mRNA measurement accuracy than traditional NGS technologies. Thanks to this technology, the laboratory protocol and algorithm developed, it is possible to analyse the degree of neoplastic cellularity digitally for each tumour without any need of H&E staining for assessment of tumour cellularity. The same 3’DGE sequencing results may allow subsequent analysis of the tumour tissue, without the need to repeat any other molecular assay. Therefore, 3'DGE sequencing, and the algorithm developed for the analysis, have a very high potential as an alternative to histology
methods and in particular IHC for the processing of FFPE tissues with degraded RNA, typical of pathological anatomy. In fact, the inventors undertook a clinical study, which demonstrated a very high accuracy of the parameters detected for the definition of molecular subtypes as defined by the recognised and globally used standard of pathological anatomy. The invention is characterised by a preliminary laboratory sample processing phase and three computational phases consisting of three algorithms. The algorithm of step 1 allows the calculation of the number of counts for all genes analysed by sequencing; the algorithm of step 2 consists of a model training for the definition of classification parameters on the basis of biological samples of known H&E classification; the algorithm of step 3 allows these parameters to be applied to classify a sample. With the method according to the invention, it is possible in step 3, from a biological sample obtained from a patient with an unknown classification, to obtain a quantification of the neoplastic cellularity of the sample itself, possibly identifying the sample as tumoral by step 3.2. The method according to the invention overcomes the limitations of the known art and, in particular, makes it possible to obtain the results described above by lowering the very high cost and time associated with multigenic testing technologies currently used in the diagnostic field and the technical limitations of current H&E staining and IHC technology; the method according to the invention overcomes the limitations of traditional NGS related to the inaccurate measurement of degraded nucleic acids (RNA) extracted from tissues derived from pathological anatomy, and to the difficulty in understanding the results, which are very extensive and often difficult to process and consult, allowing to classify tumour cellularity and monitor statistical robustness of the classification. This method presented in the invention is in continuity with the guidelines of pathological anatomy (ESMO, AIOM, ESCO/CAP), as well as guaranteeing an immediate understanding of the data by specialists. The method of the invention, which is in line with existing multigenic solutions, overcomes the limitations between those identified for NGS and other existing technologies (e.g. RT-PCR, microarrays). The method of the invention, thanks to its competitive cost compared to other NGS technologies and its robust standards of accuracy and interpretation of the results provided, is able to support the anatomo- pathologist in daily clinical practice, introducing a paradigm shift in molecular diagnostics, shifting for the first-time clinical practice towards digital methods of analysis. In particular, a preliminary clinical study carried out using the method of the invention, demonstrated a very high accuracy of the parameters detected for the definition of molecular subtypes in breast cancer as defined by the recognised and globally used standard of pathological anatomy and Immunohistochemistry (ESMO; 2019; AIOM, Edition 2021; ASCO/CAP Guidelines). Thus, the anatomo-pathologist can work in continuity with the guidelines and diagnostic protocols used in recent decades, adopting an innovative and extremely reliable precision medicine tool.
BRIEF DESCRIPTION OF THE FIGURES The present invention is described below on the basis of the following illustrative and non-limiting figures: - Figure 1: A block diagram showing the general structure of the steps of the invention - Figure 2: A block diagram describing the individual steps of steps 1, 2 and 3 of the bioinformatics analysis and algorithm - Figure 3: Accuracy plot for number of genes to be selected. - Figure 4: Confusion matrix of the testing dataset - Figure 5: Confusion matrix on trining dataset based on method from prior art as described in Wai jin Tan et al. - Figure 6: Confusion matrix of the testing dataset based on method from prior art as described in Wai jin Tan et al. Table 2: List of genes detected upon cellularity feature selection DESCRIPTION OF THE INVENTION Unless otherwise defined, technical and scientific terms used herein have the same meaning commonly understood by a person skilled in the technical field of the present invention. In this description, the term 'method' expressly refers to an in vitro method that is not performed directly on the human body. In this description, 'sequencing' means massive sequencing, next-generation sequencing (NGS), next-generation genetic sequencing, second-generation sequencing, parallel sequencing, third-generation sequencing, pyrosequencing. In this description, "3'DGE" or "digital RNA-seq," refers to a next-generation sequencing-based method as described in Xiong, Y., Soumillon, M., Wu, J. et al.2017, for estimating gene expression by amplifying and quantifying the poly-A tail regions of all expressed transcripts. This system is based on a precise measurement of the number of fragments derived from the RNA extracted from the sample, starting with the amplification of fragments at one end, defined 3'. This end is the same for messenger RNAs and some non-coding RNAs and contains a conserved and repeated sequence of A nucleotides (adenine), defined as the polyadenylation or poly- A tail. When 3′DGE is performed, cDNA is generated directly on the extracted sample, without performing the steps of pre-capture of polyadenylated RNAs in the sample or fragmentation steps; 3'DGE analysis is performed on the extracted RNA and involves as initial step of library preparation, pairing with dT oligos directed to the 3′ of the transcript followed by subsequent retrotranscription and amplification steps. According to the present invention 3' DGE can also be performed using specially designed kits, by way of non-limiting example YourSeq FT & 3'DGE (Active Motif), 3' DE kit (Takara Bio), Quantseq 3' (Lexogen). According to the present invention
sequencing can be carried out using machines suitable for next generation sequencing, by way of non-limiting example MiSeq, NextSeq series, NovaSeq series (Illumina), MGI- G-50 and MGI G-400 (MGI Tech), Elemnt AVITI System (Element Biosciences), UG100 (Ultima Genomics), ONSO (Pacific Biosciences). The term 'expression level' refers to a specific level of gene expression. This can be a determined expression level as an absolute or normalised value for each sample and is based on the count of occurrences of a specific gene in the biological sample examined. The term 'oligonucleotides' refers to a DNA fragment composed of a short number of nucleotides. The term 'tumour', as used here, refers to a dysregulated state of neoplastic cell growth and proliferation, both malignant and benign, and to all precancerous and cancerous cells and tissues. In this description, 'cellularity' refers to the composition of cells in a tissue; 'tumour cellularity' represents the fraction of tumour cells compared to normal cells in a certain tissue. In this description 'sample' or 'biological sample' means a sample obtained from a subject and more specifically a portion of body tissue or fluid obtained from a subject; by way of non-limiting example breast tissue, tissue obtained from biopsies or needle aspirates, blood, plasma, serum, spontaneous or induced tumour exudates, tissue lysates and fractions thereof. The term 'sample' or 'biological sample' also includes any preparation derived from the same; by way of non-limiting example, lysates or extracts prepared from the tissue, paraffin embedding, formalin fixation, fresh frozen tissue, fresh tissue. In a preferred embodiment of the present invention, 'sample' means a tumour sample, i.e. comprising one or more tumour cells. Marker' refers to the expression of an RNA or protein detected by analytical measurement of gene expression level or protein expression. In an embodiment, the invention relates to an in vitro method of determining the tumour cellularity of a sample. The method according to the invention is therefore based on the use of digital 3' RNA- seq (3'DGE) (Xiong, Y., Soumillon, M., Wu, J. et al. 2017, https://doi.org/10.1038/s41598-017-14892-x) to analyse gene expression of specific transcripts. By applying the 3'DGE to nucleic acids derived from sample tissues, particularly those of suspected tumour origin, the inventors are able to simultaneously quantify the expression of more than 10000 genes. This method is used for the analysis and classification of nucleic acids extracted from tissues and biological fluids derived from patients with suspected cancer; the sample may be any isolated human biological material, e.g. tissue or biological fluids or both; the technique is applicable on any biological sample and, in a preferred embodiment, such sample is obtained from biopsies of tumour tissue and the use of paraffin-fixed tissue (FFPE) is particularly preferred.
The inventors are in fact able, even with degraded RNA, with the method according to the invention, to determine the presence of gene expression profiles associated with tumour versus non-tumour cellular types, which are useful in the practice of pathological anatomy. In an embodiment, the method according to the invention provides for the application of the 3'DGE method for preparation of libraries through oligo dT, associated with an innovative computational analysis method that allows a classification of the sample taken from patients according to the main indications of national/international molecular pathology guidelines (AIOM, ESMO; ASCO/CAP). Figure 1 shows the overall flow and related data analysis steps, which characterise the diagnostic solution described here. Figure 2 shows the 3 phases of data analysis and thus the training and testing algorithms. In particular, the method according to the invention provides a training algorithm, created on the basis of data derived from a group of samples with known classification and based on Hematoxylin and Eosin (H&E) staining, according to the guidelines and protocols of the AIOM, ESMO and ASCO/CAP guidelines. Samples and metadata were obtained from: PrecisionMed Inc (study number 1179119 and IRB approval no.20172130); Cutting Edge Inc (study number 1302233 and IRB approval no.20210553); European Oncology Institute (IEO) authorised on the basis of IEO procedure DSC.RE. 7173.A (reference number UID 3738); Azienda Ospedaliero-Universitaria di Modena authorised by Prot. AOU 0004895/22 of 17/02/2022 with reference to Practice C.E.N 79/2021 SIRER ID 2123 Study Itanet_Registry. As illustrated in Figure 1, e.g. the 3'DGE method is applied to estimate gene expression (gene count) either in training samples (whose subtypes and characteristics or metadata are known e.g. from previous histology staining analysis) or in testing samples e.g. tissues from patients with suspected breast cancer e.g. (phase 1). The Algorithm in Phase 2 of the initial training is used on samples with known tumour cellularity (measured by a certified pathologist using Haematoxylin-Eosin) to identify a list of genes associated with tumours in the sample. Once step 2 has been performed, the determined thresholds and/or numerical parameters remain unchanged and can be updated following the execution of new steps 1 and 2 applied to tissue samples with known subtype, different from or in addition to those used to obtain the previous classification thresholds. Once the classification thresholds and/or numerical parameters have been determined, they are applied (step 3) to an output file obtained during step 1 performed on a biological sample taken from a patient with suspected cancer, such as breast cancer, whose classification is not yet known. In particular, both phase 2 and phase 3 allow the results of e.g.3'DGE to be classified in tumour tissue using a classification algorithm. The training dataset (phase 1) comprises known samples, analysed according to the clinical standard of pathological anatomy and according to guidelines e.g. by an
anatomo-pathologist. These samples are used as a reference dataset for the calculation of numerical parameters and for the definition of classification thresholds e.g. by means of logistic regression, which is e.g. used in the first algorithm. In particular, these samples are selected according to certain analytical quality parameters derived from the sequencing output with 3'DGE. Quality standard refers to objective metrics produced by the NGS technique such as: number of sequences produced, and number of genes identified per sample. Only samples that meet minimum quality specifications can become part of the training dataset ( e.g. preferably between 6M and 12 M sequences produced; e.g. preferably between 10000 and 12000 genes identified). Within the training dataset, two types of tissues have been collected: a tumour sample, of neoplastic type, identified by the term 'TUM' that may contain a mixture of the tumor and adjacent tissue with tumor represented at different abundances, and a healthy tissue sample taken from a distal region of the same tissue of origin of the tumor with no tumoral cellularity detected by H&E staining, identified by the term 'NAT'. In addition, samples must have certain associated characteristics (also referred to as metadata) in order to be included in the training dataset: -tumour cellularity of the tissue, obtained by staining with haematoxylin-eosin and measured by an anatomo-pathologist. In future, we will refer to this cellularity as 'real cellularity'. The testing dataset consists of biopsy samples with suspected cancer, the classification of which is unknown and have to be classified. Only one sample per patient is required for the testing dataset with a minimum quality of 6-8M sequences produced and at least 9000-12000 genes identified. The method according to the present invention involves a step of determining the tumour cellularity of a sample. This analysis will be explained in detail in the following paragraphs. Preliminary phase. Analytical preparation of the library (3'DGE) The method according to the invention preferably comprises the processing of tissue biopsies, fresh, fresh frozen or embedded in FFPE. In a preferred embodiment, the tissue is in FFPE from which, by way of non-limiting example, about 5-10 microns of tissue with dimensions ranging from 20mm2 to 200mm2 and a total of up to 2.0mm3 are taken. RNA is extracted from the sample using methods known to the expert in the field and suitable for subsequent application of NGS techniques. For example, commercial kits with protocols established by the manufacturer can be used for this purpose. RNA quality can be checked, for example, using screentape RNA (Agilent) and the DV200 quality index.
Samples must reach a concentration of at least 1ng/ul, preferably samples with a concentration of 50ng/μl are selected. According to the invention, extracted RNA samples are used for the preparation of libraries (analytical step - Figure 1). This preparation consists of several steps (i) mRNA retrotranscription with oligo-dT oligonucleotides that anneal to the poly-A tail and trigger the polymerase for retrotranscription of the mRNA to cDNA; (ii) enzymatic degradation of the mRNA template; (iii) amplification of the single-stranded cDNA by priming with random oligonucleotides to obtain a double-stranded cDNA (iiii) purification of the fragments obtained and amplification by Polymerase Chain Reaction (PCR) to insert the sequences required for the sequencing reaction. The oligonucleotides used as primers in (i), (iii) (poly-A and Random) have the sequences to trigger the synthesis of read 1 and possibly also read 2 for the sequencing reaction. Depending on the technology and the commercial kit or preparation protocol of the 3'DGE, there are specific adapters and the choice of these is up to the expert in the field. Sequencing may be performed through next generation technology and by any platform based on sequencing by synthesis (SBS) . Phase 1. Determination of expression values per gene In this section, a detailed example of the process of handling biological samples in order to obtain gene expression values is given. In particular, selected samples, for which detailed information on percentage (%) tumor cellularity by H&E staining and metadata are available, will be used to train the decision-making model during the training phase and to identify threshold values for the evaluation of gene markers. The results obtained will then be applied to assess gene expression values in samples for which the classification for H&E is unknown. After sequencing, the resulting information is saved and converted into sample- specific files using the Fastq format. In a preferred implementation, samples that have produced more than 12 million sequences are subsampled so as not to exceed this threshold (phase 1.1 Downsampling – Figure 2), in order to avoid excessive imbalance between samples. The resulting sequences are then processed with a specific tool, such as bbduk (BBMap - Bushnell B. - sourceforge.net/projects/bbmap/), to remove any adapters added during the library preparation phase. In NGS, the read accuracy of a nucleotide progressively decreases with each successive read. This means that the longer the sequence to be read, the greater the probability of reading errors towards the end of the sequence. To remedy this problem, a removal of the end parts of sequences with low read quality (phase 1.2 Trimming – Figure 2) is carried out. Furthermore, reads that are shorter than 20 bases after this operation may not align correctly and are therefore deleted. In phase 1.3 (Figure 2), the sequences obtained are aligned to a reference genome. In the present study, the human reference genome used for the described steps was
GRCh38.p5 (https://www.ncbi.nlm.nih.gov/assembly/GCF_000001405.31/). The alignment is performed using specialised software known to experts in the field. Examples include, but are not limited to, the STAR software (Dobin A, Davis CA, Schlesinger F, et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics. 2013;29(1):15-21. doi:10.1093/bioinformatics/bts635). It is recommended to preferably discard sequences whose alignment on the reference genome covers less than 60% of their length. To quantify the expression of each gene, the aligned reads are counted using specific software such as HTseq-counts, which is known to the expert in the field (step 1.4 - figure 2). The reference genes are listed in a General Transfer Format (GTF) file and can be easily obtained from the standard human GTF available on the National Centre for Biotechnology Information (NCBI) platform. The count file, obtained by measuring the expression of all sequenced genes that followed analysis steps 1.1 to 1.4, is stored for subsequent analysis steps. Phase 2. Feature selection, training of the model and parameters definition. Phase 2.1 Data aggregation. The steps to define thresholds and analysis criteria are performed only on a subset of samples defined as the training set, that has been previously characterized by H&E staining and classified by a certified molecular pathologist. These samples are selected on the basis of certain analytical quality parameters (PQA): only those samples that after sequencing analysis exceed 8 million fragments produced and a number of identified genes greater than or equal to 12 thousand (with at least 5 assigned sequences per gene) are considered in the subsequent steps of phase 2. Training set samples may be identified as TUM when they have tumoral cells and NAT when the tissue does not present tumoral cells. Every sample is associated with metadata that describes the sample as tumoral or normal and data on % tumor cellularity in case of classified tumoral tissues (Figure 1, 2). Dataset consists of 419 samples, of which 56 NAT and the other sample distributed upon different level of cellularity as shown in Table 1. Classification NAT TUM Cellularity % 11- 21- 31- 41- 51- 61- 71- BIN 0 1-10 20 30 40 50 60 70 80 81-90
Num. of samples 56 13 15 59 60 39 63 63 35 16 Table 1. Number of samples and metadata and cellularity bins In step 2.1, the metadata associated with the training samples (e.g. tumour or non- tumour classification, % tumour cells, values for IHC of markers) is aggregated into a single database (training database) together with the raw counts derived from sequencing and step 1 (gtf files). In the example case of breast cancer for TUM samples, metadata related to tumour cellularity (1% to 100% cellularity) measured by haematoxylin-eosin (H&E) is taken into account. For the analysis of NAT samples, the relevant metadata is the information on the absence of tumour cells (0% cellularity), in contrast to TUM samples where the presence of tumour cells is crucial. No collected tissue from biopsy has been evaluated as 100% cellularity, with 90% cellularity being the highest amount of tumour cellular component over all non-tumour cellular components. 2.2 Tumor cellularity assessment criterion by feature selection of genes variables As already mentioned, in addition to the analytical quality parameters, the training dataset comprises, preferably consists of samples derived from the pathology department where two types of tissues have been collected: tissue from tumour biopsies and biopsies of normal adjacent areas of the same tissue. Both biopsies have been formalin-fixed as FFPE and stained with hematoxylin-eosin (H&E). A specialized and certified pathologist performed H&E-based evaluation of tumor cellularity for each biopsy obtained. Tissue with 0% cellularity were identified 'NAT', and bopsies whose tumor cellularity is greater than 0% were considered 'TUM' samples. Such classification refers to standard clinical practice and international guidelines as reported for solid tumors, e.g. breast cancer, ovarian cancer, prostate cancer etc. (ESMO, 2019). . As will be indicated in more detail below, the 'TUM' samples are analysed for a plurality of expressed genes, this plurality being, as a whole, representative of the specific tumour to which the TUM samples refer (step 2.2, Figure 2). The method used in the invention is based on the evidence that bioptic tissue is a heterogeneous mixture of tumor cells and adjacent non-tumor cells (e.g. immune, epithelial, connective, adipose, etc.). The clinical need is based on the importance of assessing the presence of tumor cellularity, which can ensure the accuracy of subsequent molecular analysis, like measurments of other tumor biomarkers. Based on this clinical need and the evidence of tissue biopsy heterogeneity, the invention is focused on finding an accurate method for detecting presence of tumor celluarity in a tissue.
The invention is based on the assumption that if the amount of tumour cells relative to non-tumoral cells in a biological tissue increases, and thus the tumour cellularity, the expression level of characteristic tumour genes can steadily increase or decrease. These characteristic genes and their expression level are identified on a specific method for feature selection and a training algorithm based on samples with known tumour cellularity (TUM) or no tumor cellularity (NAT) obtained from H&E staining. The training NAT and TUM samples are divided into 10 groups with different and increasing tumour cellularity. These groups are treated as time points in a statistical regression method based on Gaussian processes. Specifically, the increasing variation of gene expression as a function of known H&E tumour cellularity was assessed for each measured gene. In this step, to fit known statistical models, cellularity is 'labelled' as a time-point variable. For example, 0% cellularity will correspond to time 0 and will refer to NAT samples, while 10% cellularity will correspond to time 1 and so on up to time 10 for a total of 10 groups, consisting of 1 NAT group and 9 groups of TUM samples with increasing tumour cellularity. For each sample of the 9 TUM groups, cellularity- associated genes were selected using a Bayes Factor. The so-called time-point series analysis is repeated 20 times, in which, for each repetition, only a random subset of 10 samples per time-point is used to calculate the Bayes Factor. The purpose of these repetitions is not to have a single defined ranking score, but to make a more robust and accurate choice of genes that are most representative of the condition to be classified. After the 20 repetitions, the different lists of ranked genes are aggregated to identify an overall ranked list according to the Borda score calculation. It should be noted that the identification of genes representative of cellularity is also expressed at measurable levels (mean expression >= 5 raw counts) across all the samples. The identified ranked genes show a statistically robust association between tumour cellularity and expression level. It is possible, at this point, to select a group of top genes, i.e. 100, and inspect them to identify a subset that is highly specific for the tumour cellularity trend by applying an unsupervised learning model, i.e. kmeans or hierarchical clustering on all TUM and NAT samples to select the cluster with a increasing or decreasing expression level correlated with the increasing tumour cellularity from the set of training samples. The choice to add the information obtained through an unsupervised machine learning method, i.e., clustering, to the ranking score information obtained through Borda, allows us to perform an initial filter for choosing that group of genes that is identifiable with an increasing and decreasing trend in cellularity across the various bins. The trend in the average expression of genes within a cluster shows, in fact, how some genes are characteristic of an increasing trend over time or decreasing trend or genes which trend is not linear. A first filter on genes is performed by eliminating those genes exhibiting fluctuating mean expression across the bins. To make the selection of genes to be used to train the machine learning model even more robust, the remaining genes are analyzed over the entire dataset by calculating the importance of the variables using a model-based approach. The advantage of using a model-based approach is that is more closely tied to the model performance and that it may be able to incorporate the correlation structure between the predictors into the importance calculation. A cumulative
variance calculation step allows us to identify key genes in the classification model. (Figure 3). The cumulative importance represents the cumulative sum of gene implication in describing the data, indicating the proportion of total importance in the dataset that can be attributed to a combination of genes selected. With this third filter and a threshold value, i.e., 80 percent, it is possible to determine the point at which to stop gene selection, ensuring that only that amount of importance that is deemed significant or useful for analysis is taken into account. As shown, at lest 40 genes or more provide a reasonable number of genes for an accurate estimation. Therefore, as the number of training samples changes, the genes that are part of the group may change as exemplified in table 2 and 3. The selected genes are used as a specific category of genes to investigate their expression in a statistical logistic regression model, as an example of a generalised linear model (GLM). These models allow the estimation of the probability of occurrence of a binary event (true/false, yes/no, 0/1), which define the classes of the model, i.e. tumour or non-tumour sample, based on one or more independent variables, such as the level of gene expression. In the case of a generic GLM model, the probability threshold chosen to distinguish the two classes is, by way of example, 0.5. This choice is based on the objective of minimising the number of false positives and false negatives in the calculation of model accuracy (Fluss R, Faraggi D, Reiser B. Estimation of the Youden Index and its associated cutoff point. Biom J.2005 Aug;47(4):458-72. doi: 10.1002/bimj.200410135. PMID: 16161804.). Gene Ensembl chromosome start end RNF223 ENSG00000237330 1 1070966 1074307 ESPN ENSG00000187017 1 6424788 6461370 ENSA ENSG00000143420 1 150600851 150629612 TUFT1 ENSG00000143367 1 151540305 151583583 RAB13 ENSG00000143545 1 153981617 153986358 LAMTOR2 ENSG00000116586 1 156054752 156058510 DNAH14 ENSG00000185842 1 224896262 225399292 AKT3 ENSG00000117020 1 243488233 243851079 MAP4K4 ENSG00000071054 2 101696850 101894689 LIMS2 ENSG00000072163 2 127638381 127681786 RBMS1 ENSG00000153250 2 160272151 160493794 ARL4C ENSG00000188042 2 234493041 234497053 CDS1 ENSG00000163624 4 84582979 84651338 NR3C1 ENSG00000113580 5 143277931 143435512 TFAP2A ENSG00000137203 6 10393186 10419659 HIST1H2AC ENSG00000180573 6 26124145 26139116 CFB ENSG00000243649 6 31945650 31952084 C6orf132 ENSG00000188112 6 42101118 42142619 EPHB6 ENSG00000106123 7 142855061 142871094
SYTL5 ENSG00000147041 x 38006582 38128819 IL33 ENSG00000137033 9 6215786 6257983 CELF2 ENSG00000048740 10 11005321 11336675 NMT2 ENSG00000152465 10 15102584 15168693 C1RL-AS1 ENSG00000205885 12 7108052 7122501 RASA3 ENSG00000185989 13 113977783 114132611 ZBTB42 ENSG00000179627 14 104800596 104804712 CCDC64B ENSG00000162069 16 3027682 3036926 RP11-1136G4.2 ENSG00000273615 16 54851640 54852733 CRNDE ENSG00000245694 16 54918863 54929189 KRT19 ENSG00000171345 17 41523617 41528308 PRKCA ENSG00000154229 17 66302636 66810743 TJP3 ENSG00000105289 19 3708109 3750813 CD209 ENSG00000090659 19 7739994 7747564 CTXN1 ENSG00000178531 19 7924485 7926166 PRDX2 ENSG00000167815 19 12796820 12801859 TMEM238 ENSG00000233493 19 55379245 55384598 Table 2. Genes for calculating digital cellularity, with reference to breast cancer (Genome Build GRCh38.p5; Genome V. GRCh38; Genome Date 2013-12; Genome Build accession NCBI:GCA_000001405.20; Gene build last updated 2015-10) Ensembl Gene ID Gene Symbol CHR Start End ENSG00000028277 POU2F2 chr19 42086110 42196585 ENSG00000065357 DGKA chr12 55927319 55954027 ENSG00000075426 FOSL2 chr2 28392448 28417317 ENSG00000079950 STX7 chr6 132445867 132513198 ENSG00000087245 MMP2 chr16 55389700 55506691 ENSG00000088367 EPB41L1 chr20 36091504 36232799 ENSG00000089820 ARHGAP4 chrX 153907367 153934999 ENSG00000091592 NLRP1 chr17 5499415 5619424 ENSG00000099814 CEP170B chr14 104865268 104896747 ENSG00000100055 CYTH4 chr22 37282027 37315341 ENSG00000101825 MXRA5 chrX 3308565 3346652 ENSG00000103855 CD276 chr15 73683966 73714514 ENSG00000105639 JAK3 chr19 17824780 17848071 ENSG00000115977 AAK1 chr2 69457997 69674349 ENSG00000118473 SGIP1 chr1 66533267 66751139 ENSG00000123338 NCKAP1L chr12 54497752 54548243
ENSG00000123989 CHPF chr2 219538954 219543809 ENSG00000124098 FAM210B chr20 56358974 56368663 ENSG00000125354 SEPT6 chrX 119615724 119693370 ENSG00000130508 PXDN chr2 1631887 1744852 ENSG00000130635 COL5A1 chr9 134641803 134844843 ENSG00000134954 ETS1 chr11 128458761 128587558 ENSG00000136490 LIMD2 chr17 63695888 63701172 ENSG00000139329 LUM chr12 91102629 91111494 ENSG00000142949 PTPRF chr1 43525187 43623666 ENSG00000152601 MBNL1 chr3 152243828 152465780 ENSG00000157227 MMP14 chr14 22836560 22849041 ENSG00000158985 CDC42SE2 chr5 131245493 131398447 ENSG00000160113 NR2F6 chr19 17231883 17245919 ENSG00000164050 PLXNB1 chr3 48403854 48430086 ENSG00000166033 HTRA1 chr10 122458551 122514907 ENSG00000166741 NNMT chr11 114257787 114313536 ENSG00000167291 TBC1D16 chr17 79932343 80035872 ENSG00000168528 SERINC2 chr1 31409565 31434680 ENSG00000169442 CD52 chr1 26317958 26320523 ENSG00000170421 KRT8 chr12 52897187 52949954 ENSG00000172349 IL16 chr15 81159575 81314058 ENSG00000177508 IRX3 chr16 54283304 54286787 ENSG00000180096 SEPT1 chr16 30378135 30395991 ENSG00000184922 FMNL1 chr17 45221444 45247319 ENSG00000186340 THBS2 chr6 169215780 169254050 ENSG00000188157 AGRN chr1 1020120 1056118 ENSG00000196199 MPHOSPH8 chr13 19633659 19673441 ENSG00000214548 MEG3 chr14 100779410 100861031 ENSG00000213654 GPSM3 chr6 32190766 32195523 Table 3 Genes for calculating digital cellularity, with reference to breast cancer (Genome Build GRCh38.p5; Genome V. GRCh38; Genome Date 2013-12; Genome Build accession NCBI:GCA_000001405.20; Gene build last updated 2015-10). Based on what has been described above, it is possible to select for the training phase of the algorithm TUM and NAT testing samples from other neoplasms, in order to identify through the method of step 2.2 an additional set of genes whose expression correlates with tumour cellularity values, and which fulfil the requirements for tumour/non-tumour classification. It is possible to use as a correlation parameter not just the Bayes Factor.
In the example above, 10 different expression levels of the selected genes corresponding to 10 different levels of cellularity were identified, but it is possible to increase or decrease these levels. 2.3 Database Creation As a final step in phase 2, a Reference Database is created in which they are stored: - the gene expression count matrix of the training samples. - the metadata associated with the training samples; - all threshold values and parameters derived from the training samples (as described above for step 2.2). - The predictive models obtained in phase 2.2 Step 3. Samples classification Following phase 2 training, after identifying the characteristic cellularity genes (table 2), it is possible to analyse a biological sample derived from a patient with a suspected neoplasia, the classification of which is not known to provide useful information for the molecular diagnosis of a carcinoma e.g. breast cancer. In particular, with evidence for stage 3 shown in Figure 2. Step 3.1 Sample QC on PQA The sample to be analysed (testing) is subjected to an analytical quality control (QC): for example, at least 6 million sequenced fragments and a number of identified genes exceeding 10 thousand. Only samples that pass QC can continue the analysis of phase 3 testing. For each sample entering testing, the expression values (cpm) obtained for all identified genes are compared to the expression values (cpm) within the training dataset. For each individual gene is reported: ^ the level of expression in each sample; ^ the distributions of expression values within the entire training dataset. Step 3.2 Classification for tumor cellularity based on the model On the sample passing the QC of step 3.1, an analysis is performed to identify the presence of tumour cells, based on the method described in step 2.2, which is derived from the statistical study of the expression levels of the list of genes associated with tumour cellularity (step 2.2, Table 2 or 3) The logistic regression model obtained in step 2.2 is applied to the test sample, producing a numerical value representing the probability of the sample being cancerous or not, preferably producing the numerical value as direct output of the computer implemented trained probabilistic model after the elaboration of the extraction from output file with the counts of the biological sample to be tested, of the expression of the genes in Table 2 or 3. The sample with a probability of being cancerous below a threshold defined by the linear regression model (step 2.2), is considered non-tumor and is therefore excluded from further analysis. In fact, a non- tumor sample has a gene expression inconsistent with the analytical approach
followed during the subsequent classification steps. Such samples are therefore considered unsuitable for further testing. Validation of the method The analytical validation results obtained refer to the results derived from the different classification steps indicated in step 3 and compared to the results from the current clinical standard, based solely on histopathology H&E results. The validation consists of using blinded samples of known classification (classified according to ESMO, ASCO/CAP, AIOM clinical practice guidelines) and comparing the results obtained by the methodology described in the invention. The samples in table 1 were used for training and a k-Fold Cross Validation method. K-fold cross validation divides the data into k subsets and performs the holdout method k times. Specifically, one of the k subsets is used as the test set and the other k−1 subsets are put together to form a training set. The k-fold cross-validation technique in caret, implemented in R, partitions the dataset into k subsets. It then iteratively trains the model on k-1 subsets and evaluates it on the remaining subset. This process is repeated K times, providing more reliable estimates of model performance, reducing the risk of overfitting or underfitting and making the best use of all available data. Finally, given these results the overall accuracy of the method is calculated using the R caret package (https://CRAN.R-project.org/package=caret) using the formula: ^^ + ^^ ^^ + ^^ + ^^ + ^^ Where to: VP (true positives) and VN (true negatives) we mean the samples whose classification result from step 3 of the algorithm corresponds with the results derived from the IHC. FP (false positives) and FN (false negatives) mean the samples whose classification result from step 3 of the algorithm does not match the results derived from the IHC. The preliminary results obtained show an accuracy of the method of 0.99%. The results are shown by the confusion matrix, of the tested dataset shown below.
Comparison with the prior art In Wai jin Tan et al. the first step performed is a gene expression analysis by Differential expression (DE) between two conditions, performed with a specific package for microarray analysis (samr: significance analysis of microarray r). The conditions examined are two fibroadenoma (considered normal) and phyllodes tumor (tumor/malignant counterpart). The first step is to choose a list of biologically significant genes from among the differentially expressed genes. To do this, the authors use some filters on the values of R-fold, a parameter that is not directly related to statistical significance. At this point, the authors move the analysis of gene expression from microarray data to qRT-PCR, since this technology is most likely not suitable for future clinical practice and only a few genes may actually be used to develop a potential test to be applied. To this end, the authors tested oligonucleotides presumably for a list of 495 genes (Table S3) and validated 43 of those genes with primers “*successfully designed for qPCR assays” (*at the bottom of Table S3). The 43 selected genes are not ranked based on differential expression, but rather on the success of the assay to be used. Having chosen the genes based on qRT-PCR assay results, the authors transform ^Ct measurments to apply machine learning techniques to: 1) a random forest model for feature selection of the genes selected by DE and qRT- PCR from the previous steps 2) glm method for binary classification of the output.
About the Random Forest approach it is not clarified by the authors whether the model uses cross-validation or bootstrap techniques to obtain a stable and robust model. Anyway, from the random forest 7 genes are identified prioritized through the highest score resulting from the method: it is not specified why 7 and not more or less, very likely the choice is to obtain a number of genes feasible for qRT-PCR assay. Starting from the 7 genes, a brute force method with the glmulti package is used to select single genes or in combination that allows to obtain through AIC (Akaike Information Criteria) the most significant value. Thus 5 genes are identified, that also in combination with each other, give the best model based on AIC. In order to make a comparison as complete as possible with our data, the method described in Wai jin Tan et al. was performed with the samples described in our invention (Table 1) and relative 2 conditions TUM vs NAT. For parameters not specified in the method, default parameters were utilized. Regarding the training dataset, a group of samples of the size and characteristics to that used in the method described in Wai jin Tan et al. was selected: 20 matched samples (10 Tumor and 10 normals, adjacent) and another 28 unmatched samples (14 tumour and 14 normal taken randomly and not matched to tumours). The first step was to perform differential expression (DE). Since the data according to the invention were produced by NGS, it was not possible to use the paper’s samr tool, and the data were processed with a tool considered the gold standard in the field of gene expression analysis for NGS: DESeq2. From this analysis, 3070 significantly expressed genes were identified. In contrast to Wai jin Tan et al., our test and the entire workflow is uniquely based on next-generation sequencing data, which is used as input for feature selection (genes), machine learning, and for the classification of the sample in testing phase. The advantage is to use only one comprehensive test and analysis for all molecular biomarkers and classification of tumor to be analyzed. After the DE analysis, the machine learning phase began: - Random Forest was applied on the 3070 genes and from this analysis, sorting the genes according to importance the first 7 were chosen, to simulate the choice performed in in Wai jin Tan et al. paper. Since there is no detail on the parameters of the random forest model used in the paper, a standard bootstrap parameter was used. - By bootstrapping glmulti with the default values using the AIC method we obtained 4 genes: FGFR, SEMA3F, WNT16, ICA1. We could not obtain the combinations of the genes, since the computational resources available were insufficient for successfully run the model (cloud virtual machine n2-highmem-4 with 4 cpu e 32GB of ram). Unfortunately, no specific parameters were specified by Wai jin Tan et al. to improve the success of the run. At this point the the 5 genes obtained was used by glmulti as training for the predictive model.
To conclude, all the adjustments were made to be able to replicate the method described in Wai jin Tan et al. and the obtained model was applied on the training dataset (40 total samples) obtaining a classification accuracy of 100% on the training, validating that the correct execution of the training. The Confusion matrix on training dataset based on method from prior art as described in Wai jin Tan et al is shown below.
Next for validation of Wai jin Tan et al. on our samples a testing analysis was performed on a testing dataset, including 230 samples and thus comparable with the method dataset presented in in Wai jin Tan et al. In this case, the accuracy is 76.09%. However, it is important to note that, all but 1 normal sample in the testing dataset (n=55) are misclassified, with a balanced accuracy between the two categories normal (NAT) vs tumor (TUM) of 50% as shown in the validation statistics. The confusion matrix of the testing dataset based on method from prior art as described in Wai jin Tan et al. is shown below.
Claims
CLAIMS 1. Computer-based method to classify a biological sample comprising the steps of: - Receiving the biological test sample of a subject with suspected cancer - Receive an indication of the tumour to be verified in the biological test sample - Extract RNA and generate a library of fragments of PolyA+ sequences from the RNA of the sample - Processing through a sequencer and subsequent alignment to the reference genome and counts of fragment or reads, to obtain an output file of the biological test sample including a plurality of counts of occurrences of corresponding genes represented by the PoliA+ fragments in the sample, each count defining an expression level of the corresponding gene - Apply to the output file a binary classification algorithm trained with biological training samples whose classification as non-tumour or tumour cancer to be verified by histological cellularity examinations and whose gene expression levels are known - In which said trained binary classification algorithm is configured to generate, on the basis of the expression levels in the output file of the biological test sample of a predefined list of genes (list 1) whose expression levels are indicative of the cellularity of the tumour to be tested, a binary estimate as to whether the biological test sample is tumour-like of the tumour to be tested or non-tumour. 2. Method according to claim 1, comprising the following training steps of the first binary classification algorithm: - Extract the RNA, for each of said biological training samples of tumour and nontumour tissues, and generate a library of fragments of PolyA+ sequences from the RNA for each of a plurality of first training samples classified as non-tumour (NAT) and tumour (TUM), each first tumour training sample having a known level of cellularity preferably by histological examination, and each non-tumour training sample (NAT) being an adjacent, non-tumour tissue to the tissue of a corresponding tumour training sample (TUM) - Process through a sequencer and subsequent alignment to the reference genome and read counts, to obtain a first training output file for each first training sample including a plurality of counts of occurrences of corresponding genes represented by the PoliA+ fragments in the sample, each count defining an expression level of the corresponding gene
- Calculate for each gene in the first training output file of each training tumour sample a value of a correlation index, preferably the Bayes Factor of the expression level with the known cellularity of the corresponding training tumour sample - Include in the list a gene based on the corresponding correlation index value with the known cellularity of the corresponding training samples - to train said classification algorithm in such a way that, for each training sample (NAT, TUM), the expression level of each of the genes of said list (list 1) is received as input and the classification known by histological examinations as tumour or non- tumour of the training sample is received as output, where said list of genes (list 1) is such that an accuracy of the estimation of the classification algorithm applied to the training samples is greater than or equal to 95%. 3. Method according to claim 2, wherein the known cellularity levels of the biological training samples are at least 2 with corresponding non-zero known cellularity levels associated with corresponding training output files containing gene expression levels further comprising the following steps: - identify a first and second plurality of genes on the basis of the said correlation index for each of the training samples and each non-zero cellularity level and in which the list (List 1) includes genes from the first and second plurality of genes whose expression levels vary with variation in cellularity. 4. Method according to any one of claims 2 or 3, further comprising the steps of: - apply an unsupervised learning algorithm receiving as input the expression levels of genes whose correlation index value is above the said threshold and divide the corresponding training samples into at least two groups; and - when in each of the said groups the splitting accuracy between NAT and TUM samples is below a predefined threshold, e.g.100%, increase a threshold of the correlation index value to increase the specificity of the genes to the tumour cell condition. 5. Method according to any one of claims 1 to 4, wherein the first classifier is a general linearised model (GLM). 6. Method according to any one of claims 2 to 5, wherein the Bayes Factor is greater than or equal to 8, preferably greater than or equal to 10. 7. Method according to any one of the preceding claims in which the library preparation and analysis steps are carried out using the 3'DGE technique. 8. Method according to any one of the preceding claims wherein the sample is obtained
from a breast tissue biopsy. 9. Method according to any one of the preceding claims wherein the sample is a formalin-fixed, paraffin-embedded (FFPE) biological tissue sample or an untreated biological tissue sample.
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| IT102023000010068A IT202300010068A1 (en) | 2023-05-24 | 2023-05-24 | MOLECULAR DIAGNOSIS METHOD TO CLASSIFY A BIOLOGICAL SAMPLE BASED ON CELLULARITY |
| IT102023000019611A IT202300019611A1 (en) | 2023-09-22 | 2023-09-22 | MOLECULAR DIAGNOSIS METHOD FOR CLASSIFYING A BIOLOGICAL SAMPLE BASED ON ITS CELLULARITY |
| PCT/IB2024/054909 WO2024241204A1 (en) | 2023-05-24 | 2024-05-20 | Molecular diagnostic method for classifying a biological sample on the basis of cellularity |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4721073A1 true EP4721073A1 (en) | 2026-04-08 |
Family
ID=93588999
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24731083.2A Pending EP4721073A1 (en) | 2023-05-24 | 2024-05-20 | Molecular diagnostic method for classifying a biological sample on the basis of cellularity |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4721073A1 (en) |
| CN (1) | CN121127921A (en) |
| AU (1) | AU2024277269A1 (en) |
| WO (1) | WO2024241204A1 (en) |
-
2024
- 2024-05-20 WO PCT/IB2024/054909 patent/WO2024241204A1/en not_active Ceased
- 2024-05-20 AU AU2024277269A patent/AU2024277269A1/en active Pending
- 2024-05-20 CN CN202480032819.8A patent/CN121127921A/en active Pending
- 2024-05-20 EP EP24731083.2A patent/EP4721073A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| CN121127921A (en) | 2025-12-12 |
| WO2024241204A1 (en) | 2024-11-28 |
| AU2024277269A1 (en) | 2025-10-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11098372B2 (en) | Gene expression panel for prognosis of prostate cancer recurrence | |
| EP2852689B1 (en) | Nano46 genes and methods to predict breast cancer outcome | |
| US20260043085A1 (en) | Compositions, methods and kits for diagnosis of a gastroenteropancreatic neuroendocrine neoplasm | |
| US20140154681A1 (en) | Methods to Predict Breast Cancer Outcome | |
| US20060265138A1 (en) | Expression profiling of tumours | |
| KR101672531B1 (en) | Genetic markers for prognosing or predicting early stage breast cancer and uses thereof | |
| WO2017223216A1 (en) | Compositions and methods for diagnosing lung cancers using gene expression profiles | |
| CN108588230B (en) | Marker for breast cancer diagnosis and screening method thereof | |
| US20160115551A1 (en) | Methods to predict risk of recurrence in node-positive early breast cancer | |
| US20250137066A1 (en) | Compostions and methods for diagnosing lung cancers using gene expression profiles | |
| EP4721073A1 (en) | Molecular diagnostic method for classifying a biological sample on the basis of cellularity | |
| PL241607B1 (en) | Panel of miRNA biomarkers for differential diagnosis of histopathological subtypes of non-small cell lung cancer | |
| IT202300010068A1 (en) | MOLECULAR DIAGNOSIS METHOD TO CLASSIFY A BIOLOGICAL SAMPLE BASED ON CELLULARITY | |
| CA3239042A1 (en) | Compositions and methods for identifying transplant rejection or the risk thereof | |
| IT202300019611A1 (en) | MOLECULAR DIAGNOSIS METHOD FOR CLASSIFYING A BIOLOGICAL SAMPLE BASED ON ITS CELLULARITY | |
| AU2004219989B2 (en) | Expression profiling of tumours | |
| WO2019015549A1 (en) | Cell type identification method and system thereof | |
| HK1231515B (en) | Gene expression panel for prognosis of prostate cancer recurrence |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251125 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |