WO2024252405A1 - Non-invasive bone marrow diagnostics - Google Patents
Non-invasive bone marrow diagnostics Download PDFInfo
- Publication number
- WO2024252405A1 WO2024252405A1 PCT/IL2024/050568 IL2024050568W WO2024252405A1 WO 2024252405 A1 WO2024252405 A1 WO 2024252405A1 IL 2024050568 W IL2024050568 W IL 2024050568W WO 2024252405 A1 WO2024252405 A1 WO 2024252405A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- subject
- bone marrow
- dataset
- cellular
- cells
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6806—Preparing nucleic acids for analysis, e.g. for polymerase chain reaction [PCR] assay
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61K—PREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
- A61K39/00—Medicinal preparations containing antigens or antibodies
- A61K39/0005—Vertebrate antigens
- A61K39/0011—Cancer antigens
- A61K39/001196—Fusion proteins originating from gene translocation in cancer cells
- A61K39/001197—Breakpoint cluster region-abelson tyrosine kinase [BCR-ABL]
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61P—SPECIFIC THERAPEUTIC ACTIVITY OF CHEMICAL COMPOUNDS OR MEDICINAL PREPARATIONS
- A61P35/00—Antineoplastic agents
- A61P35/02—Antineoplastic agents specific for leukemia
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6876—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
- C12Q1/6881—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for tissue or cell typing, e.g. human leukocyte antigen [HLA] probes
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6876—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
- C12Q1/6883—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
- C12Q1/6886—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material for cancer
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
- G16B30/10—Sequence alignment; Homology search
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/20—Supervised data analysis
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H20/00—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance
- G16H20/10—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance relating to drugs or medications, e.g. for ensuring correct administration to patients
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/20—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/30—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for calculating health indices; for individual health risk assessment
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2600/00—Oligonucleotides characterized by their use
- C12Q2600/156—Polymorphic or mutational markers
Definitions
- the present invention is in the field of bone marrow diagnostics.
- HSPCs acquire somatic mutations, however, certain mutations in leukemia-related genes, namely pre-leukemic mutations - pLMs, can lead to clonal expansion of HSPCs, a phenomenon termed clonal hematopoiesis (CH). While CH is quite common among the elderly, it remains poorly understood why pLMs lead to clonal expansion, and how CH and other age-related blood phenomena are related to each other.
- CH clonal hematopoiesis
- PB HSPCs can be a good surrogate for studying inter-individual HSPC transcriptional heterogeneity.
- a new accurate, non-invasive test for assessing MSPCs of the bone marrow by examining HSPCs in PB therefore greatly needed.
- the present invention provides non-invasive methods of detecting pathology of the bone marrow comprising receiving a subject cellular dataset based on single cell RNA sequencing (scRNA-seq) of CD34 positive cells from peripheral blood of the subject and analyzing the received subject cellular dataset in relation to a control dataset comprising a plurality of cellular datasets wherein each cellular dataset of the plurality is based on scRNAseq of CD34 positive cells from peripheral blood of a healthy subject.
- scRNA-seq single cell RNA sequencing
- a non-invasive method of detecting pathology of the bone marrow in a subject in need thereof comprising: a. receiving a subject cellular dataset based on single cell RNA sequencing (scRNA-seq) of CD34 positive cells from peripheral blood of the subject; and b. analyzing the received subject cellular dataset in relation to a control dataset comprising a plurality of cellular datasets wherein each cellular dataset of the plurality is based on scRNA-seq of CD34 positive cells from peripheral blood of a healthy subject, wherein a deviation of the subject cellular dataset from the control dataset indicates a bone marrow pathology; thereby detecting pathology of the bone marrow.
- scRNA-seq single cell RNA sequencing
- the cellular dataset comprises statistical data of the totality of CD34 positive cells in a peripheral blood sample.
- the analyzing comprises producing a feature vector representing deviation of the subject’s cellular data from the control cellular data.
- the analyzing comprises applying a trained machine learning model to the received dataset, wherein the machine learning model is trained on a training set comprising the plurality of cellular datasets and wherein the machine learning model classifies the subject’s bone marrow as being a healthy or not.
- the training set further comprises cellular datasets based on scRNA-seq of CD34 positive cells from peripheral blood of subjects suffering from pathology of the bone marrow and labels indicating a cellular dataset is from a healthy subject or a subject with pathology of the bone marrow; and wherein the machine learning model classifies the subject as being heathy or suffering from a pathology of the bone marrow.
- the analyzing comprises applying a trained machine learning model to the feature vector, wherein the machine learning model is trained on a training set comprising: feature vectors from healthy subjects and subjects suffering from pathology of the bone marrow and labels indicating a feature vector is from a healthy subject or a subject with pathology of the bone marrow; and wherein the machine learning model classifies the subject as being heathy or suffering from a pathology of the bone marrow.
- the analyzing comprises applying a trained machine learning model to a parameter extracted from the cellular dataset, wherein the machine learning model is trained on a training set comprising: the parameter extracted from cellular datasets of healthy subjects and optionally subjects suffering from a bone marrow pathology and wherein the machine learning model classifies the subject as being a healthy subject or not.
- the cellular dataset is selected from: a metacell model of the totality of CD34 positive cells in a peripheral blood sample, a transcriptome of each of the CD34 positive cells in a peripheral blood sample, an annotated cell atlas of CD34 positive cell types present in a peripheral blood sample.
- the pathology of the bone marrow is selected from myelodysplastic syndrome (MDS), Chronic myelomonocytic leukemia (CMML), Acute myeloid leukemia (AML), polycythemia vera (PV), essential thrombocythemia (ET), Mastocytosis, chronic eosinophilic leukemia, myelofibrosis (MF), acute lymphoblastic leukemia (ALL), acute leukemia of ambiguous lineage, multiple myeloma (MM), myeloproliferative neoplasm (MPN) and blastic plasmacytoid dendritic cell leukemia.
- MDS myelodysplastic syndrome
- CMML Chronic myelomonocytic leukemia
- AML Acute myeloid leukemia
- PV polycythemia vera
- ET essential thrombocythemia
- Mastocytosis chronic eosinophilic leukemia
- the method is a method of detecting MDS and wherein deviation in the frequency of erythrocyte progenitor cells (ERYP), basophil/eosinophil/mast progenitor cells (BEMP), and/or megakaryocyte/erythrocyte/basophil/eosinophil/mast progenitor cells (MEBEMP) indicates the presences of MDS .
- ERP erythrocyte progenitor cells
- BEMP basophil/eosinophil/mast progenitor cells
- MEBEMP megakaryocyte/erythrocyte/basophil/eosinophil/mast progenitor cells
- the method is a method of detecting CMML and wherein deviation in the frequency of early granulocyte-monocyte progenitor cells (GMP- E) indicates the presence of CMML.
- GMP- E early granulocyte-monocyte progenitor cells
- the method is a method of detecting AML and wherein deviation in the frequency of common lymphoid progenitor cells (CLP) and/or natural killer/T/dendritic cell progenitor cells (NKTDP) indicates the presence of AML.
- CLP common lymphoid progenitor cells
- NKTDP natural killer/T/dendritic cell progenitor cells
- the deviation is higher or lower levels of a cell types than is present in the healthy subjects.
- deviation in the frequency of CLPs is also indicative of MDS and wherein the deviation is lower levels of the CLPs than is present in the healthy subjects.
- deviation in the frequency of CLPs is also indicative of CMML, MF or MPN and wherein the deviation is lower levels of the CLPs than is present in the healthy subjects.
- the method is a method of detecting MDS and wherein a decrease in the frequence of CLP, NKTDP or both as compared to healthy subjects is indicative of MDS.
- a decrease in the frequency of both CLP and NKTDP as compared to healthy subjects is indicative of MDS.
- the pathology of the bone marrow comprises an increased percentage of blasts, wherein deviation is an increase and wherein a deviation in the frequency of early common lymphoid progenitor cells (CLP-E) indicates the presence of an increased percentage of blasts.
- CLP-E early common lymphoid progenitor cells
- the method further comprises administering at least one therapeutic agent to a subject determined to suffer from a bone marrow pathology.
- a non-invasive method of predicting the percentage of blasts in the bone marrow of a subject in need thereof comprising receiving a measure of the CLP-E cells in the peripheral blood of the subject wherein the measure is proportional to the percentage of blasts in the bone marrow of the subject, thereby predicting the percentage of blasts in the bone marrow of a subject.
- the method further comprises analyzing the received measure in relation to a control dataset comprising a plurality of measures of CLP- E cells in the peripheral blood of healthy subjects and subjects suffering from pathology of the bone marrow, wherein the percentage of blasts in the bone marrow is known for each subject of the control dataset.
- a non-invasive method of predicting the percentage of blasts in the bone marrow of a subject in need thereof comprising: a. receiving a subject cellular dataset based on single cell RNA sequencing (scRNA-seq) of CD34 positive cells from peripheral blood of the subject; and b.
- scRNA-seq single cell RNA sequencing
- the machine learning model is trained on a training set comprising a plurality of cellular datasets wherein each cellular dataset of the plurality is based on scRNA-seq of CD34 positive cells from peripheral blood of a control subject and labels indicating the percentage of blasts in the bone marrow of the control subjects that provided each cellular dataset of the plurality of cellular datasets; and wherein the machine learning model outputs a predicted percentage of blasts in the bone marrow of the subject; thereby predicting the percentage of blasts in the bone marrow of a subject.
- the subject suffers from leukemia.
- control subjects comprise subjects suffering from leukemia and non-leukemic subjects.
- the cellular dataset is selected from: a metacell model of the totality of CD34 positive cells in a peripheral blood sample, a transcriptome of each of the CD34 positive cells in a peripheral blood sample, an annotated cell atlas of CD34 positive cell types present in a peripheral blood sample.
- the cellular data set is a metacell model and is produced by a method comprising: a. receiving a peripheral blood sample from a subject; b. isolating CD34 positive hematopoietic stem and progenitor cells (HSPCs) from the peripheral blood sample; c. performing scRNA-seq of the isolated HSPCs to produce a transcriptome for each isolated HSPC; and d. producing a metacell model of the HSPCs based on their transcriptomes.
- HSPCs hematopoietic stem and progenitor cells
- a metacell is a cluster of cells with a similar transcriptome.
- a cellular dataset comprises groupings of cells into cell types that share a common differentiation within the HSPC spectrum of differentiation.
- the cell types are selected from: BEMP, ERYP, MEBEMP-L, MEBEMP-E, GMP-E, multipotent progenitor cells (MPP), hematopoietic stem cells (HSC), CLP-E, CLP-M, CLP-L and NKTDP.
- the method is a method of detecting MDS and/or leukemia and wherein a percentage of blasts above a predetermined threshold indicates the subject suffers from MDS and/or leukemia.
- the method further comprises administering to a subject suffering from MDS and/or leukemia at least one anticancer therapy.
- a non-invasive method of calculating a Molecular International Prognostic Scoring System (IPSS-M) risk score for a subject suffering from a bone marrow malignancy comprising: a. predicting the percentage of blasts in the bone marrow of the subject by a method of the invention; b. detecting the presence of bone marrow mutations and karyotype abnormalities based on scRNA-seq reads from CD34 positive cells from peripheral blood of the subject; c. receiving hemoglobin levels, and platelet counts in peripheral blood from the subject; and d. calculating the IPSS-M risk score based on the predicted blast percentage, detected mutations and karyotyping and received hemoglobin levels and platelet counts; thereby calculating an IPSS-M risk score.
- IPS-M Molecular International Prognostic Scoring System
- the method further comprises administering to the subject a treatment regimen based on the IPSS-M risk score, where in a subject with a higher score is administered a more intense treatment regimen and a subject with a lower score is administered a reduced treatment regimen.
- a system for evaluating bone marrow health in a subject comprising: a scRNA sequencing device; a non-transitory memory device, wherein modules of instruction code are stored; and at least one processor associated with the memory device, and configured to execute the modules of instruction code, whereupon execution of the modules of instruction code, the at least one processor is configured to: obtain from the scRNA sequencing device single cell transcriptomes from CD34 positive cells from peripheral blood of the subject; produce a cellular dataset based on the obtained single cell transcriptomes; analyze the produced cellular dataset in relation to a control dataset comprising a plurality of cellular datasets wherein each cellular dataset of the plurality is based on scRNA-seq of CD34 positive cells from peripheral blood of a healthy subject and output a finding of healthy bone marrow or pathology of the bone marrow in the subject based on deviation of the subject cellular dataset from the control dataset.
- the cellular dataset is a metacell model with similar transcriptomes from the obtained single cell transcriptomes clustered into metacells.
- Figures 1A-1H (1A) experimental design.
- (IB) annotated 2D UMAP projection of the metacell manifold following filtration of metacells with low CD34 expression.
- (1C-D) (1C) Symmetric and (ID) asymmetric regulation of specific HSC transcription factors upon bifurcation to the CLP (right) and MEBEMP (left) lineages.
- Each panel shows the expression of one gene (Y axis).
- Metacells in all panels are ordered (left to right) by increasing A VP expression in the MEBEMP lineage and decreasing AVP expression in the CLP lineage. Units for gene expression in all the figure panels are log2 of each gene’s fractional expression.
- IE the metacell population of interest (dotted line) linking BEMPs to their MEBEMP-L precursors.
- IF positively and negatively regulated TFs involved in early BEMP differentiation.
- (1G gene-gene plot of IRF8 against TCF7 expression as hallmark markers of DC and T cell differentiation respectively. The high ACY3 NKTDP metacell population of interest is depicted (dotted line).
- This population exhibits high expression of both T and dendritic cell regulators, forming a gradient consisting of NK/T cell-like progenitors exhibiting a high TCF7/IRF8 expression ratio along with high expression of other T cell hallmarks such as CD7, MAF, IL7R, TRBC2, and DC-like progenitors exhibiting a low TCF7/IRF8 expression ratio, along with high expression of other DC hallmarks, such as the myeloid TF PU.l and the MHC class II gene CD74.
- FIGS. 2A-2H (2A) characterization of inter-individual HSPC compositional state variation (scheme). (2B) boxplots of cell state frequency distributions across individuals (logarithmic scale). Percents calculated out of CD34+ population. Boxplot centers, hinges and whiskers represent median, first and third quartiles and 1.5x interquartile range, respectively. Numbers represent mean +/- SD for each distribution. (2C) correlation of cell state frequencies between 20 biological replicates and their original samples, for CLP (CLP- E, CLP-M, CLP-L, NKTDP - top) and MEBEMP (MEBEMP-E, MEBEMP-L, ERYP, BEMP - bottom) populations.
- CLP CLP- E, CLP-M, CLP-L, NKTDP - top
- MEBEMP MEBEMP-E, MEBEMP-L, ERYP, BEMP - bottom
- FIG. 3A-3K (3A) Analysis of age-linked compositional differences in MEBEMP (MEBEMP-E, MEBEMP- L, ERYP, BEMP, left) and CLP (CLP-E, CLP-M, CLP-L, right) populations, comparing specific cell state frequencies (out of total CD34+ population) in young ( ⁇ 50 years) vs. old (>60-years) individuals without clonal hematopoiesis, performed for males (blue) and females (red) separately. Kruskal- Wallis p values for group differences are denoted on top.
- Y axis denotes log2(observed/expected expression) normalized for composition. Boxplot centers, hinges and whiskers represent median, first and third quartiles and 1.5x interquartile range, respectively.
- (31) individual heatmaps of single cell counts over 20 bins of sternness (A VP signature, y axis) and MEBEMP differentiation (GATA1 signature, x axis). Individual identifier, as well as his/her RBC, and MCV are denoted on top. (3J) comparison between individual sync scores and clinical parameters (RBC/MCV) across males. High and low sync scores (denoted by red and black dots respectively) define clinically distinct populations. (3K) correlation between individual sync scores and cell state compositions. Permutation test p value denoted on top.
- Figures 4A-4G (4A) composition bias score variation with age. (4B) cell typespecific comparison of S-phase signatures in circulating (left) vs. BM (right) HSPCs. (4C) S -phase signature variation with age in the late MEBEMP trajectory. (4D) corresponding individual S-phase signatures (X axis) and composition bias scores (Y) for individuals younger (left) and older (right) than 65 years. (4E-4G) like 4D, but showing the (4E) LMNA signature, (4F) sync scores, and (4G) RDW instead of S-phase, respectively.
- Figures 5A-5F (5A) diagnostic approach to leukemia analysis using the cHSPC reference atlas (scheme): 1. scRNA-seq on CD34-enriched PB and construction of a patientspecific metacell model, 2. Projection of patient derived metacells on the healthy reference atlas. 3. Compositional (relative cell state frequency) analysis and 4. Composition-controlled differential gene expression analysis. 5. Mutational and CNV analysis using targeted DNA sequencing and RNA-based karyotyping.
- (5E) CompositionaL controlled differential gene expression of 2 healthy and 7 patient samples (1 of them prior to and following treatment initiation) to the normal reference, quantifying the number of differentially expressed genes and identifying specific genes recurrently induced or repressed in disease.
- (5F) scRNA-seq karyotyping for 2 AML patients and 1 MDS patient prior to and following treatment initiation. Metacell models were created for each MDS/AML patient and projected over our healthy reference map. Coupled reference and projected (patient) metacells were then used for calculating expression ratios over all expressed genes in all chromosomes. Log2 fold-change expression (patient/healthy) for all expressed genes across all chromosomes is shown. Red lines represent the median of each chromosomal fold-change distribution.
- FIG. 6 A block diagram, depicting a computing device which may be included in a system for determining a Hematopoietic Stem Cells (HSC) condition in a subject, according to some embodiments of the invention.
- HSC Hematopoietic Stem Cells
- Figures 7A-7B Block diagrams, depicting systems for determining (7A) and indication or (7B) an IPSS-M score in a subject according to some embodiments of the invention.
- Figure 8 A flow diagram, depicting a method of determining an HSC condition in a subject according to some embodiments of the invention.
- Figure 9 Factors involved in BEMP and NKTDP differentiation. Factors positively (CNRIP1, HPGDS, TET2, TNFSF10) and negatively (CD34, HBD, CD74 and BLVRB) regulated in the early stages of BEMP specification.
- Figure 10 True age (x) vs age predicted based on composition-controlled CLP expression (y).
- Figure 11 gene-gene correlation heatmap, calculated over individual-level CLP gene expression controlled for CLP composition.
- Figures 12A-12C (12A) heatmap of individual LMNA signature expression across the MEBEMP trajectory. Individual age and sex are color-coded on top. (12B) LMNA signature expression correlations between 39 technical & 20 biological replicates and their original samples. (12C) sync score correlations between 39 technical & 20 biological replicates and their original samples. All biological replicates were sampled 1 year following original blood draw.
- FIG. 13A-13C (13A) each of 4 panels refers to a different cell state gene signature as noted on the x-axis.
- Panel top boxplots of gene module expression distributions for different cell states in our reference atlas.
- Panel bottom Gene signature expression density plots for each of the AML subclones. Reference gene signature distributions (panel top) were used to identify subpopulations of AML cells with CLP, MEBEMP, HSC and NKTDP characteristics (panel bottom). Dashed lines represent the threshold for expressing a gene signature, and the fraction of cells expressing a signature per AML clone is listed.
- the malignant state is characterized by multiple novel gene signatures in addition to aberrant expression of "healthy" differentiation-related modules, right - UMAP projection of the metacell models of AML-1 (top) and AML-2 (bottom), colored by relative expression of differentially expressed genes.
- Overexpression of BCL2 in AML-1-2 compared to AML-1-1 can be seen.
- AML-1 gene signature BCL2, VPREB1, RUNX3, GATA2, SELL, LMNA, ID2.
- AML-2 gene signature CCL4, MPO, LMNA, MME, JCHAIN, ACY3, DNTT, GATA2.
- the present invention provides non-invasive methods of detecting pathology of the bone marrow comprising receiving a subject cellular dataset based on single cell RNA sequencing (scRNA-seq) of CD34 positive cells from peripheral blood of the subject and analyzing the received subject cellular dataset in relation to a control dataset.
- Non-invasive methods of predicting the percentage of blasts in the bone marrow comprising applying a trained machine learning model to a received subject cellular dataset based on single cell RNA sequencing (scRNA-seq) of CD34 positive cells from peripheral blood are also provided.
- Non-invasive methods of calculating an IPSS-M risk score are also provided.
- Systems for performing the methods of the invention are also provided.
- the present invention is based, at least in part, on the surprising finding that single cell RNA- sequencing (scRNA-Seq) of HSPCs in the blood can be used to recapitulate the status of HSPCs in the bone marrow and thereby detect bone marrow pathology, detect the presence and percentage of bone marrow blasts and predicts clinical outcome and treatment based on a divergence from what is observed in healthy controls.
- scRNA-Seq single cell RNA- sequencing
- the magnitude of the cohort allowed the inventors to characterize in detail the transcriptional programs of diverse, sometimes rare (NKTDP, BEMP), HSPC sub-populations, refining and augmenting previous findings from much smaller cohorts (Fig. 1).
- the disclosure defines a normal reference range for cHSPC subpopulation frequencies within a large age- and sex- diverse healthy population and shows that cHSPC subtype compositions are highly variable between individuals, while the cell states themselves are remarkably universal (Fig. 2). These compositions remained stable over a one-year follow-up period, placing them as a strong individual characteristic.
- a significant correlation between low CLP frequencies, CH and increased RDW was discovered (Fig.
- RNA expression clock which correlated with age
- LMNA novel gene module
- a method of analyzing the bone marrow of a subject comprising: a. receiving a dataset based on CD34 positive cells from blood of the subject; and b. analyzing the received subject dataset in relation to a control dataset, thereby analyzing bone marrow of a subject
- the method is a diagnostic method. In some embodiments, the method is a prognostic method. In some embodiments, the method is a non-invasive method. In some embodiments, the method is an in vitro method. In some embodiments, the method is an ex vivo method. In some embodiments, the method is a method of treatment. In some embodiments, the method is a computerized method. In some embodiments, the method is performed by at least one processor. In some embodiments, the method requires analyzing data that is beyond the capability of the human mind.
- non-invasive refers to a method that does not require extraction of a sample from the bone marrow.
- Bone marrow biopsies and aspirations are invasive, painful and expensive procedures that provide a diagnostician with a sample of cells in the bone marrow.
- the instant method circumvents the drawbacks of invasive bone marrow samples by analyzing the bone marrow via the circulating CD34 positive cells found in blood.
- the instant method is highly beneficial as it is non-invasive.
- blood is peripheral blood.
- blood is venous blood.
- blood is circulating blood.
- blood is not from an organ.
- blood is not from tissue.
- blood is not from the bone marrow.
- blood is a blood sample.
- the CD34 positive cells are hematopoietic stem progenitor cells (HSPCs).
- HSPCs hematopoietic stem progenitor cells
- CD34 is a transmembrane cell surface protein that marks hematopoietic stem cells (HSCs) as well as early progenitor cells that have differentiated from HSCs.
- CD34 positive cells run the gamut from fully stem cells (HSCs) to cells that have begun to differentiate toward one of two lineage programs: common lymphoid progenitor (CLP) lineage or megakaryocyte/erythrocyte/basophil/eosinophil/mast progenitors (MEBEM-P) lineage.
- CLP common lymphoid progenitor
- MEBEM-P megakaryocyte/erythrocyte/basophil/eosinophil/mast progenitors
- CD34 Positive Isolation Kit ThermoFisher
- 1-0 Human CD34+ Cell Isolation Kit Creative Biolabs
- EasySep Human CD34 Positive Selection Kit Stemcell Technologies
- CD34 MicroBead Kit human (Miltenyi Biotec).
- the dataset is based on CD34 positive cells from a blood sample from the subject. In some embodiments, the dataset is based on all CD34 positive cells in the sample. In some embodiments, the dataset is a cellular dataset. In some embodiments, the dataset is an ensemble of the CD34 positive cells in the blood. In some embodiments, the dataset is a per cell dataset. In some embodiments, the dataset contains an entry for each CD34 positive cell. In some embodiments, the data is data on the totality of CD34 positive cells in the blood. In some embodiments, the dataset is statistical data. In some embodiments, statistical data is statistical data is a data transformation of the cellular data. In some embodiments, the dataset is based on single cell data.
- the dataset comprises single cell data. In some embodiments, the dataset consists of single cell data. In some embodiments, the single cell data is single cell RNA data. In some embodiments, the single cell RNA data is single cell RNA sequencing (scRNA-seq) data. In some embodiments, the data is reads. In some embodiments, reads are sequencing reads. In some embodiments, the data is transcriptome data. In some embodiments, the single cell data is protein data. In some embodiments, the single cell data is proteome data. In some embodiments, the dataset comprises a transcriptome of each of the CD34 positive cells. In some embodiments, the dataset comprises the proteome of each of the CD34 positive cells. In some embodiments, the dataset is a cell atlas. In some embodiments, the cell atlas is annotated. In some embodiments, the annotation is the cell type.
- the method further comprises receiving a blood sample from the subject. In some embodiments, the method further comprises extracting a blood sample from the subject. In some embodiments, a blood sample is a peripheral blood sample. In some embodiments, the method further comprises producing a dataset from the sample. In some embodiments, the method further comprises isolating CD34 positive cells from the sample. In some embodiments, isolating comprises extracting. In some embodiments, isolating is positive selection. In some embodiments, isolating is negative selection.
- the method comprises sequencing the CD34 positive cells.
- sequencing is single cell sequencing.
- the sequencing is next generation sequencing.
- the sequencing is high throughput sequencing.
- the sequencing is massively parallel sequencing.
- the dataset is a dataset of sequences.
- the dataset is a dataset of expression.
- expression is gene expression.
- CD34 cells are clustered into cell types.
- cell types are defined by their transcriptional profile.
- cell types are defined by their transcriptome.
- cell types are defined by their proteome.
- cell types are defined by their level of differentiation.
- cell types are defined by their differentiation status.
- cell types are defined by how similarly they have differentiated.
- the dataset is a metacell model of the CD34 positive cells.
- the model is of the totality of CD34 positive cells. Metacell modeling computes partitions of cells by similarity to produce mostly homogenous groups (e.g., cell types) which are defined as metacells.
- a cell type comprises a plurality of metacells.
- the cell type comprises metacells with similar differentiation.
- Methods of producing metacells from single cell data are well known and are described hereinbelow as well as for example in Baran, et al., “MetaCell: analysis of single-cell RNA-seq data using K-nn graph partitions”, Genome Biol. 2019 Oct 11 ;20(l):206 and Ben-Kiki et al., “Metacell-s: a divide and conquer metacell algorithm for scalable scRNA-seq analysis”, Genome Biol. 2022 Apr 19;23(l):100 the contents of which are hereby incorporated herein by reference in their entirety.
- the metacell program is freely available at github.com/tanaylab/metacells.
- the method comprises generating metacells from the scRNA-seq data.
- control dataset comprises the same type of data as the subject dataset.
- the control dataset comprises a plurality of subject datasets.
- the control dataset comprises a plurality of datasets.
- each of the plurality of datasets in the control dataset is from a different control subject.
- the control dataset comprises a plurality of control subject datasets.
- each dataset of the plurality is based on scRNA-seq of CD34 positive cells.
- the CD34 positive cells are from blood.
- control subjects are from control subjects.
- control subjects are healthy subjects.
- control subjects are subjects with a pathology of the bone marrow.
- control subjects are both healthy subjects and subjects with a pathology of the bone marrow.
- the control dataset is an atlas of control cells.
- the control dataset is an atlas of metacells from control subjects.
- the atlas is an atlas of datasets.
- a dataset comprises grouping of the cells into cell types.
- the metacells are grouped into cell types.
- cell types share a common transcription profile.
- cell types share a common differentiation state.
- the differentiation state is within the HSPC spectrum of differentiation.
- the control dataset comprises amounts of cell types in control subjects.
- amounts are ranges.
- cell types are types of metacells.
- cell types are differentiation states.
- amounts are relative amounts.
- amounts are amounts of all cell types in a control subject.
- ranges are ranges of all cell types in control subjects.
- the cell types are selected from different differentiation states of the CD34 positive cells.
- the cell types are selected from hematopoietic stem cells (HSC), common lymphoid progenitor cells (CLP), natural killer/T/dendritic cell progenitor cells (NKTDP), multipotent progenitor cells (MPP), early granulocyte- monocyte progenitor cells (GMP-E), megakaryocyte/erythrocyte/basophil/eosinophil/mast progenitor cells (MEBEMP), erythrocyte progenitor cells (ERYP) and basophil/eosinophil/mast progenitor cells (BEMP).
- HSC hematopoietic stem cells
- CLP common lymphoid progenitor cells
- NKTDP natural killer/T/dendritic cell progenitor cells
- MPP multipotent progenitor cells
- GMP-E early granulocyte- monocyte progenitor cells
- MEBEMP megak
- CLPs comprise early CLPs (CLP-E), intermediate CLPs (CLP-M) and late CLPs (CLP-L).
- MEBEMPs comprise early MEBEMPs (MEBEMP-E) and late MEBEMPs (MEBEMP-L).
- the cell types are selected from BEMP, ERYP, MEBEMP-L, MEBEMP-E, GMP-E, MPP, HSC, CLP-E, CLP-M, CLP-L and NKTDP.
- CLP comprises NKTDP.
- CLP comprises CLP-E, CLP-M, CLP-L and NKTDP.
- MEBEMP comprises BEMP.
- MEBEMP comprises ERYP.
- MEBEMP comprises BEMP, ERYP and MEBEMP-L.
- control dataset comprises control ranges for each cell type.
- control ranges are relative ranges.
- relative ranges are relative abundance.
- control relative ranges are relative percentage of all CD34 positive cells.
- percentage is percent of CD34 positive cells in a sample.
- the control ranges are provided in Figure 2B.
- the control range for BEMP is about 4.4% of CD34 positive cells.
- about 4.4% is 4.4 +/- 4.1%.
- the control range for ERYP is about 1.4% of CD34 positive cells. In some embodiments, about 1.4% is 1.4 +/- 0.7%.
- the control range for MEMBEMP-L is about 8.2% of CD34 positive cells. In some embodiments, about 8.2% is 8.2 +/- 2.2%. In some embodiments, the control range for MEMBEMP-E is about 38.0% of CD34 positive cells. In some embodiments, about 38.0% is 38.0 +/- 6.5%. In some embodiments, the control range for GMP-E is about 3.0% of CD34 positive cells. In some embodiments, about 3.0% is 3.0 +/- 0.9%. In some embodiments, the control range for MPP is about 21.6% of CD34 positive cells. In some embodiments, about 31.6% is 31.6 +/- 4.7%. In some embodiments, the control range for HSC is about 1.8% of CD34 positive cells.
- the control range for CLP-E is about 2.5% of CD34 positive cells. In some embodiments, about 2.5% is 2.5 +/- 0.8%. In some embodiments, the control range for CLP-M is about 7.9% of CD34 positive cells. In some embodiments, about 7.9% is 7.9 +/- 5.2%. In some embodiments, the control range for CLP- L is about 5.7% of CD34 positive cells. In some embodiments, about 5.7% is 5.7 +/- 3.6%. In some embodiments, the control range for NKTDP is about 5.1% of CD34 positive cells. In some embodiments, about 45.1% is 5.1 +/- 3.0%.
- analyzing is comparing. In some embodiments, analyzing comprises projecting the dataset onto the control dataset. In some embodiments, the analyzing is determining cell type differences between the subject dataset and the control dataset. In some embodiments, changes are loss of cells of a cell type. In some embodiments, changes are gains of cells of a cell type. In some embodiments, cells are metacells. In some embodiments, analyzing is analyzing the totality of the subject dataset. In some embodiments, analyzing is analyzing the subject dataset in relation to all of the plurality of datasets within the control dataset.
- analyzing bone marrow comprises detecting a pathology of the bone marrow. In some embodiments, detecting comprises determining the pathology of the bone marrow. In some embodiments, analyzing comprises diagnosing a pathology of the bone marrow. In some embodiments, analyzing comprises prognosing a pathology of the bone marrow. In some embodiments, analyzing comprises determining the proper treatment of a pathology of the bone marrow. In some embodiments, analyzing comprises determining the amount of blasts in the bone marrow. In some embodiments, determining is predicting. In some embodiments, determining is estimating. In some embodiments, determining is approximating. In some embodiments, the determining is without actually counting blasts in the bone marrow.
- deviation of the subject dataset from the control dataset indicates a bone marrow pathology. In some embodiments, deviation of the subject dataset from the control dataset indicates a specific bone marrow pathology. In some embodiments, deviation of the subject dataset from the control dataset indicates a disease of the bone marrow. In some embodiments, deviation comprises a difference when the subject dataset is projected onto the control dataset. In some embodiments, deviation is higher levels/amounts of a cell type being present in the subject than the healthy controls. In some embodiments, deviation is a higher frequency of a cell type in the subject than the healthy controls. In some embodiments, deviation is lower levels/amounts of a cell types being present in the subject than the healthy controls. In some embodiments, deviation is a lower frequency of a cell type in the subject than the healthy controls. In some embodiments, lower amounts is the absence of a cell type. In some embodiments, higher amounts is the presence of new cell type.
- a pathology of the bone marrow refers to any disease or condition affecting the bone marrow of humans.
- a pathology is a disease.
- a pathology is an abnormality of the bone marrow.
- bone marrow pathologies include but are not limited to: myelodysplastic syndrome (MDS), Chronic myelomonocytic leukemia (CMML), Chronic myeloid leukemia (CML), Acute myeloid leukemia (AML), polycythemia vera (PV), essential thrombocythemia (ET), Mastocytosis, chronic eosinophilic leukemia, primary myelofibrosis (MF), post-ET myelofibrosis, post PV myelofibrosis, acute lymphoblastic leukemia (ALL), acute leukemia of ambiguous lineage, multiple myeloma (MM), myeloproliferative neoplasm (MPN) and blastic plasmacytoid dendritic cell leukemia.
- MDS myelodysplastic syndrome
- CMML Chronic myelomonocytic leukemia
- CML Chronic myeloid leukemia
- AML Acute myeloid leukemia
- the pathology is cancer. In some embodiments, the cancer is a hematopoietic cancer. In some embodiments, the cancer is leukemia. In some embodiments, the pathology is MDS. MDS is a well-known group of cancers in which immature blood cells (HSPCs) within the bone marrow do not mature to become healthy blood cells. In some embodiments, the pathology is CMML. In some embodiments, the pathology is MF. In some embodiments, MF is selected from primary MF, post-ET MF and post PV MF. In some embodiments, the pathology is MPN. In some embodiments, the pathology is MDS/MPN. In some embodiments, the pathology is AML.
- the pathology is not AML. In some embodiments, the pathology is selected from the group consisting of: MDS, CMML, MF, and MPN. In some embodiments, the pathology is selected from the group consisting of: MDS, CMML, MF, and MDS/MPN. In some embodiments, the pathology is selected from the group consisting of: MDS, CMML, MF, MPN and AML. In some embodiments, the pathology is selected from the group consisting of: MDS, CMML, MF, MDS/MPN and AML. In some embodiments, MDS is MDS with a del5q mutation.
- the method is a method of detecting MDS.
- deviation in the amount or frequency of ERYP cells indicates the presence of MDS.
- deviation in the amount or frequency of BEMP cells indicates the presence of MDS.
- deviation in the amount or frequency of MEBEMP cells indicates the presence of MDS.
- deviation in the amount or frequency of any one of ERYP, BEMP and MEBEMP cells indicates the presence of MDS.
- deviation in the amount or frequency of all of ERYP, BEMP and MEBEMP cells indicates the presence of MDS.
- MEBEMP is MEBEMP-L or MEBEMP-E.
- MEBEMP is MEBEMP-L and MEBEMP-E.
- the deviation is an increase.
- deviation in the amount or frequency of CLP cells indicates the presence of MDS.
- the deviation is a decrease.
- CLP is CLP-L, CLP-M or CLP-E.
- CLP is any two of CLP-L, CLP-M and CLP-E.
- CLP is CLP-L, CLP-M and CLP-E.
- decrease in CLP amount or frequence indicates the presence of MDS.
- MDS is MDS/MPN.
- decrease in CLP amount or frequence indicates the presence of MDS or MPN. In some embodiments, deviation in the amount or frequency of NKTDP cells indicates the presence of MDS. In some embodiments, decrease in NKTDP amount or frequence indicates the presence of MDS. In some embodiments, MDS is MDS/MPN. In some embodiments, decrease in NKTDP amount or frequence indicates the presence of MDS or MPN.
- the method is a method of detecting CMML.
- deviation in the amount or frequency of GMP cells indicates the presence of CMML.
- GMP is GMP-E.
- the deviation is an increase.
- deviation in the amount or frequency of CLP cells indicates the presence of CMML.
- the deviation is a decrease.
- CLP is CLP-L, CLP-M or CLP-E.
- CLP is any two of CLP-L, CLP-M and CLP-E.
- CLP is CLP-L, CLP-M and CLP-E.
- the method is a method of detecting MF.
- deviation in the amount or frequency of CLP cells indicates the presence of MF.
- the deviation is a decrease.
- deviation in the amount or frequency of NKTDP cells indicates the presence of MF.
- decrease in NKTDP amount or frequence indicates the presence of MF.
- the method is a method of detecting MPN.
- deviation in the amount or frequency of CLP cells indicates the presence of MPN.
- the deviation is a decrease.
- CLP is CLP-L, CLP-M or CLP-E.
- CLP is any two of CLP-L, CLP-M and CLP-E.
- CLP is CLP-L, CLP-M and CLP-E.
- decrease in CLP amount or frequence indicates the presence of MPN.
- MPN is MDS/MPN.
- decrease in CLP amount or frequence indicates the presence of MDS or MPN.
- deviation in the amount or frequency of NKTDP cells indicates the presence of MPN.
- decrease in NKTDP amount or frequence indicates the presence of MPN.
- MPN is MDS/MPN.
- decrease in NKTDP amount or frequence indicates the presence of MDS or MPN.
- the method is a method of detecting AML.
- deviation in the amount or frequency of CLP cells indicates the presence of AML.
- CLP is CLP-L, CLP-M or CLP-E.
- CLP is any two of CLP-L, CLP-M and CLP-E.
- CLP is CLP-L, CLP-M and CLP-E.
- deviation in the amount or frequency of NKTDP cells indicates the presence of AML.
- the deviation is an increase.
- an increase in the amount or frequency of NKTDP cells indicates the presence of AML.
- the pathology of the bone marrow comprises an increased percentage of blasts. In some embodiments, the pathology of the bone marrow is characterized by an increased percentage of blasts. In some embodiments, the pathology of the bone marrow is selected from AML and MDS. In some embodiments, MDS is MDS/MPN. In some embodiments, the pathology of the bone marrow is selected from AML, MPN and MDS. In some embodiments, AML and MDS are characterized by an increased percentage of blasts. In some embodiments, a deviation in the frequency of CLP-E indicates the presence of an increased amount of blasts. In some embodiments, a deviation is an increase. In some embodiments, an increase in CLP-E is the deviation.
- the magnitude of the increase is proportionate to the increase in the amount of blasts.
- an increase in blasts is as compared to the amount of blasts in a healthy control.
- a healthy control is a healthy cohort.
- the healthy cohort is the subjects that make up the control dataset.
- a linear regression predicts the amount of blasts from the amount of CLP-E.
- analyzing comprises producing a feature vector representing deviation of the subject’s cellular data from the control cellular data.
- the feature vector comprises a plurality of entries.
- each entry corresponds to a specific cell type.
- each entry corresponds to an amount of each cell type.
- the amount is the number.
- the amount is the frequency.
- the frequency is the percentage of all CD34 positive cells.
- each entry represents or corresponds to the deviation from a reference value.
- the deviation is the magnitude of deviation.
- the reference value is the values from the control dataset.
- the reference value is a range of the amount of a cell type.
- a cell type is a cell population.
- the range is the control range.
- the range is the healthy range.
- analyzing comprises applying a trained machine learning model to the received dataset.
- the machine learning model is trained on a training set.
- the training set comprises the control dataset.
- the training set comprises the plurality of cellular datasets.
- the machine learning model outputs a classification of the subject’s bone marrow.
- the machine learning model outputs a classification of the subject.
- the machine learning model outputs an analysis of the subject’s bone marrow.
- the classification is healthy or not.
- the training set comprises datasets from healthy subjects. In some embodiments, training set comprises datasets from subjects suffering from pathology of the bone marrow. In some embodiments, the training set comprises datasets from subjects suffering from a plurality of pathologies of the bone marrow. In some embodiments, the training set further comprises labels. In some embodiments, the labels label the datasets. In some embodiments, the labels indicate if the dataset is from a healthy subject or subject with a pathology of the bone marrow. In some embodiments, the label indicates the pathology of the bone marrow. In some embodiments, the label indicates the type of pathology. In some embodiments, classification is healthy or suffering from a pathology of the more marrow. In some embodiments, classification comprises classifying what the pathology is. In some embodiments, classification comprises classifying the type of pathology of the bone marrow.
- analyzing comprises applying a trained machine learning model to a parameter extracted from the dataset. In some embodiments, analyzing comprises applying a trained machine learning model to the feature vector.
- the feature vector is a vector of the amounts of cell types. In some embodiments, cell types are all cell types of the CD34 positive cells in a sample. In some embodiments, the cell types are the full ensemble of CD34 positive cells in a sample.
- the machine learning model is trained on a training set. In some embodiments, the training set comprises feature vectors from healthy subject. In some embodiments, the training set comprises parameters extracted from datasets from healthy subjects. In some embodiments, the training set comprises feature vectors from subject suffering from a bone marrow pathology.
- the training set comprises parameters extracted from datasets from subjects suffering from a bone marrow pathology.
- the training set comprises labels.
- the labels indicate a feature vector is from a healthy subject or subject with a bone marrow pathology.
- the labels indicate an extracted parameter is from a healthy subject or subject with a bone marrow pathology.
- analyzing further comprises applying a trained machine learning model to at least one clinical parameter.
- the clinical parameter is a clinical parameter of the subject.
- the clinical parameter is age.
- the clinical parameter is sex.
- the clinical parameter is sex and age.
- the machine learning model is trained on a training set comprises at least one clinical parameter.
- a method of predicting the amount of blasts in the bone marrow of a subject comprising receiving a measure of the CLP-E cells in peripheral blood from the subject, thereby predicting the amount of blasts in the bone marrow of a subject.
- the measure of CLP-E cells is proportional to the amount of blasts in the bone marrow of the subject. In some embodiments, proportional is linearly proportional. In some embodiments, a linear regression indicates the amount of blasts from the measure of CLP-E. In some embodiments, indicates is predicts. In some embodiments, a measure above a predetermined threshold indicates blasts above a predetermined threshold. In some embodiments, the measure of CLP-E cells is the amount of CLP-E cells. In some embodiments, the measure of CLP-E cells is the number of CLP-E cells. In some embodiments, the measure of CLP-E cells is the proportion of CLP-E cells in the CD34 positive cells in the peripheral blood.
- CLP-E cells can be measured by any method known in the art, comprising flow cytometry, immunostaining, sequencing, producing of metacells from scRNA-seq and the like. Methods of identifying these cells in a sample, including a blood sample, are known in the art and any such method may be used. Methods of identifying CLP-E cells for example, are provided hereinbelow and in Ding and Morrison, “Haematopoietic stem cells and early lymphoid progenitors occupy distinct bone marrow niches”, Nature. 2013, Mar 14; 495(7440): 231-235, the contents of which are herein incorporated by reference in their entirety.
- the method further comprises receiving a peripheral blood sample. In some embodiments, the method further comprises measuring CLP-E cells in the sample. In some embodiments, measuring is counting. In some embodiments, the method further comprises receiving scRNA-seq data from CD34 positive cells in the blood and calculating the number/amount/percentage of CLP-E cells in the blood. In some embodiments, in the blood is in the sample. In some embodiments, the method further comprises analyzing the received measure in relation to a control dataset.
- a method of predicting the amount of blasts in the bone marrow of a subject comprising: a. receiving a dataset based on CD34 positive cells from blood of the subject; and b. applying a trained machine learning model to the received dataset, wherein the machine learning model outputs a predicted amount of blasts in the bone marrow of the subject; thereby predicting the amount of blasts in the bone marrow of a subject.
- the subject is a mammal. In some embodiments, the mammal is a human. In some embodiments, the subject is in need of a method of the invention. In some embodiments, the subject is male. In some embodiments, the subject is female. In some embodiments, the subject suffers from a pathology of the bone marrow. In some embodiments, a bone marrow pathology is a bone marrow malignancy. In some embodiments, the subject suffers from leukemia.
- leukemia is selected from AML, CMML, CML, Mastocytosis, chronic eosinophilic leukemia, acute leukemia of ambiguous lineage and blastic plasmacytoid dendritic cell leukemia.
- the amount of blasts is the number of blasts. In some embodiments, the amount of blasts is the frequency of blasts. In some embodiments, the amount of blasts is the percentage of blasts in the bone marrow. In some embodiments, percentage is relative to all cells in the bone marrow. In some embodiments, all cells are all CD34 positive cells.
- the training set comprises subjects suffering from MDS. In some embodiments, the training set comprises non-MDS subjects. In some embodiments, the training set comprises leukemic subject. In some embodiments, the training set comprises leukemic and non-leukemic subjects. In some embodiments, the training set further comprises labels. In some embodiments, the labels label the datasets. In some embodiments, the labels indicate the amount of blasts in the subject that provided the dataset. In some embodiments, the percentage of blasts in the bone marrow is known for each subject of the control dataset. In some embodiments, a subject of the control dataset is a subject that provided data for the control dataset. In some embodiments, the dataset is a dataset of the plurality of datasets. In some embodiments, the dataset is a control dataset. In some embodiments, the machine learning model outputs the amount of blasts in the subject.
- the method is a method of detecting MDS and an amount of blasts above a predetermined threshold indicates the subject suffers from MDS. In some embodiments, the method is a method of detecting leukemia and an amount of blasts above a predetermined threshold indicates the subject suffers from leukemia. In some embodiments, the threshold is 0%. In some embodiments, the threshold is 5%. In some embodiments, the threshold is 9%. In some embodiments, the threshold is 10%. In some embodiments, the threshold is 15%.
- the method further comprises not administering a therapeutic agent to a subject with amounts of blasts below the predetermined threshold. In some embodiments, the method further comprises administering a therapeutic agent to a subject determined to suffer from a pathology of the bone marrow. In some embodiments, the method further comprises administering a therapeutic agent to a subject with amounts of blasts above the predetermined threshold.
- the agent is an anticancer agent and the subject suffers from cancer. In some embodiments, the cancer is MDS. In some embodiments, the agent is an anti-MDS agent. In some embodiments, the anti-MDS agent is lenalidomide. In some embodiments, the agent is an anti-leukemia agent.
- Anticancer agents are well known in the art and any such agent may be used, this includes, but is not limited to, chemotherapy, radiation therapy, immunotherapy, and targeted therapy.
- the agent is a chemotherapy.
- the agent is radiation therapy.
- the agent is an immunotherapy.
- the immunotherapy is immune checkpoint inhibition.
- the checkpoint is PD-1/PD-L1.
- the immunotherapy is CAR-T or CAR-NK therapy.
- the anticancer agent is a hypomethylating agent.
- the hypomethylating agent is azacytidine.
- the hypomethylating agent is decitabine.
- the anticancer agent is azacytidine in combination with venetoclax.
- the subject suffers from leukemia and the anticancer agent is venetoclax.
- the subject suffers from MDS and the agent is azacytidine.
- the subject suffers from MDS and the agent is azacytidine in combination with venetoclax.
- the leukemia is chronic lymphocytic leukemia, small lymphocytic lymphoma, or acute myeloid leukemia.
- the method further comprises performing a bone marrow transplant on a subject determined to suffer from a pathology of the bone marrow.
- the method further comprises performing a bone marrow transplant on a subject with an amount of blasts above a predetermined threshold.
- the subject suffers from MPN and the agent is an interferon.
- the subject suffers from MPN and the method further comprises administering interferon therapy.
- the interferon is interferon alpha.
- interferon is a type I interferon.
- interferon is interferon beta.
- interferon beta is interferon beta 1 (IFNB1).
- interferon is interferon alpha.
- interferon alpha is selected from interferon alpha 1, 2, 4, 5, 6, 7, 8, 10, 13, 14, 16, 17 and 21. In some embodiments, interferon is interferon alpha-2b. In some embodiments, the agent is Ropeginterferon alfa-2b (Besremi).
- a method of calculating a Molecular International Prognostic Scoring System (IPSS-M) risk score for a subject comprising: a. predicting the percentage of blasts in the bone marrow to the subject by a method of the invention; b. receiving data as to the presence of bone marrow mutations and/or karyotype abnormalities in the subject; c. receiving hemoglobin levels and/or platelet counts in peripheral blood from the subject; and d. calculating the IPSS-M risk score based on the predicted blast percentage, received mutations and/or karyotyping data and received hemoglobin levels and/or platelet counts; thereby calculating an IPSS-M risk score.
- IPS-M Molecular International Prognostic Scoring System
- the method further comprises detecting the presence of bone marrow mutations. In some embodiments, the method further comprises detecting karyotype abnormalities. In some embodiments, the detecting is in the scRNA data. In some embodiments, the detecting is a non-invasive detecting. In some embodiments, the detecting does not comprise detecting within the bone marrow. It will be understood that all steps of the method can be performed non-invasively and one of the major benefits of the method of the invention is that is does not require a bone marrow sample in order to learn important information (e.g., IPSS-score) about the bone marrow.
- important information e.g., IPSS-score
- the mutation or karyotype abnormality is del(5q). In some embodiments, the mutation or karyotype abnormality is -7/del(7q). In some embodiments, the mutation or karyotype abnormality is -17/del(17p). In some embodiments, the mutation or karyotype abnormality is a complex karyotype. In some embodiments, the mutation or karyotype abnormality is del(llq). In some embodiments, the mutation or karyotype abnormality is del(5q). In some embodiments, the mutation or karyotype abnormality is del(12p). In some embodiments, the mutation or karyotype abnormality is del (20q).
- the mutation or karyotype abnormality is del (7q). In some embodiments, the mutation or karyotype abnormality is +8. In some embodiments, the mutation or karyotype abnormality is +19. In some embodiments, the mutation or karyotype abnormality is i(17q). In some embodiments, the mutation or karyotype abnormality is -Y. In some embodiments, the mutation or karyotype abnormality is -7. In some embodiments, the mutation or karyotype abnormality is (inv)3/t(3q)/del(3q).
- the mutation is a variant allele.
- the mutation is mutation within tumor protein p53 (TP53).
- mutation is the number of mutations.
- the mutation or karyotype abnormality is loss of heterozygosity of the TP53 locus.
- the mutation is MLL (lysine methyltransferase 2A (KMT2A)) mutation.
- the mutation is fms related receptor tyrosine kinase 3 (FLT3) mutation.
- the mutation is ASXL transcriptional regulator 1 (ASXL1) mutation.
- the mutation or karyotype abnormality is Cbl proto-oncogene (CBL) mutation.
- the mutation is DNA methyltransferase 3 alpha (DNMT3A) mutation.
- the mutation is ETS variant transcription factor 6 (ETV6) mutation.
- the mutation is Enhancer Of Zeste 2 Polycomb Repressive Complex 2 Subunit (EZH2) mutation.
- the mutation is isocitrate dehydrogenase (NADP(+)) 2 (IDH2) mutation.
- the mutation is KRAS proto-oncogene, GTPase
- the mutation is nucleophosmin 1 (NPM1) mutation.
- the mutation is NRAS proto-oncogene, GTPase (NRAS) mutation.
- the mutation is RUNX family transcription factor 1 (RUNX1) mutation.
- the mutation is splicing factor 3b subunit 1 (SF3B1) mutation.
- the mutation is serine and arginine rich splicing factor 2 (SRSF2) mutation.
- the mutation is U2 small nuclear RNA auxiliary factor 1 (U2AF1) mutation.
- the mutation is BCL6 corepressor (BCOR) mutation.
- the mutation is BCL6 corepressor like
- the mutation is CCAAT enhancer binding protein alpha (CEBPA) mutation.
- the mutation is ethanolamine kinase 1 (ETNK1) mutation.
- the mutation is GATA binding protein
- the mutation is G protein subunit beta 1 (GNB 1) mutation. In some embodiments, the mutation is isocitrate dehydrogenase (NADP(+)) 1 (IDH1) mutation. In some embodiments, the mutation is neurofibromin 1 (NF1) mutation. In some embodiments, the mutation is PHD finger protein 6 (PHF6) mutation. In some embodiments, the mutation is protein phosphatase, Mg2+/Mn2+ dependent ID (PPM ID) mutation. In some embodiments, the mutation is pre-mRNA processing factor 8 (PRPF8) mutation. In some embodiments, the mutation is protein tyrosine phosphatase non-receptor type 11 (PTPN11) mutation. In some embodiments, the mutation is SET binding protein 1 (SETBP 1) mutation. In some embodiments, the mutation is STAG2 cohesin complex component (STAG2) mutation. In some embodiments, the mutation is WT1 transcription factor (WT1) mutation.
- GNB1 G protein subunit beta 1
- the mutation is isocitrate de
- hemoglobin levels are received. In some embodiments, the method further comprises measuring hemoglobin levels. In some embodiments, the method comprises receiving a blood sample from the subject. In some embodiments, the hemoglobin levels are calculated in the blood sample. In some embodiments, platelet counts are received. In some embodiments, the method further comprises counting platelets. In some embodiments, the platelets are in the blood sample. In some embodiments, the method further comprises receiving neutrophil counts. In some embodiments, the method further comprises counting neutrophils. In some embodiments, neutrophils in the sample are counted. In some embodiments, the subject’s age is also received. In some embodiments, the subject’s sex/gender is also received.
- the IPSS-M risk score is calculated based on any combination of received data. In some embodiments, the IPSS-M risk score is calculated based on the predicted blast percentage. In some embodiments, the IPSS-M risk score is calculated based on the predicted blast percentage and received mutations and karyotyping. In some embodiments, the IPSS-M risk score is calculated based on the predicted blast percentage and the received hemoglobin levels and platelet counts. In some embodiments, the IPSS-M risk score is calculated based on the predicted blast percentage, received mutations and karyotyping and received hemoglobin levels and platelet counts. In some embodiments, the IPSS-M risk score is calculated further based on the neutrophil counts and/or the patient’s age.
- the IPSS-M score is well known in the art. It ranges from 0 to 16. The scores are divided into six risk possibilities: Very Low (VL) risk, Low (L) risk, Medium Low (ML) risk, Medium High (MH) risk, High risk (H) and Very High (VH) risk.
- Subjects with low risk may receive no treatment or treatment to manage symptoms such as Erythropoiesisstimulating agents (ESA) to treat anemia. Patients with thrombocytopenia may receive romiplostim or eltrombopag. Similarly, Luspatercept can be administered if ESA is ineffective (and/or there is a mutation in SF3B 1 or ring sideroblasts are present).
- Subjects with high risk may receive hypomethylating agents, or other anticancer treatments. High risk subjects may have a bone marrow transplant.
- the method further comprises administering to a subject a treatment regimen based on the calculated IPSS-M score.
- a subject with a higher score is administered a more intense treatment regimen.
- a subject with a lower score is administered a reduced treatment regimen.
- more intense is increased.
- reduced is less intense.
- a method of detecting AML in a subject comprising detecting the presence of an R353K mutation within GAT A3 in a sample from the subject, thereby detecting AML in a subject.
- the sample comprises cells.
- the cells are hematopoietic cells.
- the cells are blasts.
- the cells are CD34 positive cells.
- mutation is a mutation of arginine 353 in GATA3.
- the arginine is mutated to lysine.
- the mutation is indicative of AML.
- Computing device 1 may include a processor or controller 2 that may be, for example, a central processing unit (CPU) processor, a chip or any suitable computing or computational device, an operating system 3, a memory 4, executable code 5, a storage system 6, input devices 7 and output devices 8.
- processor 2 or one or more controllers or processors, possibly across multiple units or devices
- More than one computing device 1 may be included in, and one or more computing devices 1 may act as the components of, a system according to embodiments of the invention.
- Operating system 3 may be or may include any code segment (e.g., one similar to executable code 5 described herein) designed and/or configured to perform tasks involving coordination, scheduling, arbitration, supervising, controlling or otherwise managing operation of computing device 1, for example, scheduling execution of software programs or tasks or enabling software programs or other modules or units to communicate.
- Operating system 3 may be a commercial operating system. It will be noted that an operating system 3 may be an optional component, e.g., in some embodiments, a system may include a computing device that does not require or include an operating system 3.
- Memory 4 may be or may include, for example, a Random- Access Memory (RAM), a read only memory (ROM), a Dynamic RAM (DRAM), a Synchronous DRAM (SD-RAM), a double data rate (DDR) memory chip, a Flash memory, a volatile memory, a non-volatile memory, a cache memory, a buffer, a short term memory unit, a long term memory unit, or other suitable memory units or storage units.
- Memory 4 may be or may include a plurality of possibly different memory units.
- Memory 4 may be a computer or processor non- transitory readable medium, or a computer non-transitory storage medium, e.g., a RAM.
- a non-transitory storage medium such as memory 4, a hard disk drive, another storage device, etc. may store instructions or code which when executed by a processor may cause the processor to carry out methods as described herein.
- Executable code 5 may be any executable code, e.g., an application, a program, a process, task, or script. Executable code 5 may be executed by processor or controller 2 possibly under control of operating system 3. For example, executable code 5 may be an application that may calculate an IPSS-M score for a subject as further described herein. Although, for the sake of clarity, a single item of executable code 5 is shown in Figure 6, a system according to some embodiments of the invention may include a plurality of executable code segments similar to executable code 5 that may be loaded into memory 4 and cause processor 2 to carry out methods described herein.
- Storage system 6 may be or may include, for example, a flash memory as known in the art, a memory that is internal to, or embedded in, a micro controller or chip as known in the art, a hard disk drive, a CD-Recordable (CD-R) drive, a Blu-ray disk (BD), a universal serial bus (USB) device or other suitable removable and/or fixed storage unit.
- Data pertaining to single cell RNA sequencing (scRNA-seq) reads may be stored in storage system 6 and may be loaded from storage system 6 into memory 4 where it may be processed by processor or controller 2.
- scRNA-seq single cell RNA sequencing
- memory 4 may be a non-volatile memory having the storage capacity of storage system 6. Accordingly, although shown as a separate component, storage system 6 may be embedded or included in memory 4.
- Input devices 7 may be or may include any suitable input devices, components, or systems, e.g., a detachable keyboard or keypad, a mouse and the like.
- Output devices 8 may include one or more (possibly detachable) displays or monitors, speakers and/or any other suitable output devices.
- Any applicable input/output (I/O) devices may be connected to Computing device 1 as shown by blocks 7 and 8.
- NIC network interface card
- USB universal serial bus
- any suitable number of input devices 7 and output device 8 may be operatively connected to Computing device 1 as shown by blocks 7 and 8.
- a system may include components such as, but not limited to, a plurality of central processing units (CPU) or any other suitable multi-purpose or specific processors or controllers (e.g., similar to element 2), a plurality of input units, a plurality of output units, a plurality of memory units, and a plurality of storage units.
- CPU central processing units
- controllers e.g., similar to element 2
- NN neural network
- ANN artificial neural network
- ML machine learning
- Al artificial intelligence
- NN neural network
- ANN artificial neural network
- a NN may be configured or trained for a specific task, e.g., pattern recognition or classification. Training a NN for the specific task may involve adjusting these weights based on examples.
- Each neuron of an intermediate or last layer may receive an input signal, e.g., a weighted sum of output signals from other neurons, and may process the input signal using a linear or nonlinear function (e.g., an activation function).
- the results of the input and intermediate layers may be transferred to other neurons and the results of the output layer may be provided as the output of the NN.
- the neurons and links within a NN are represented by mathematical constructs, such as activation functions and matrices of data elements and weights.
- At least one processor e.g., processor 2 of Fig. 6) such as one or more CPUs or graphics processing units (GPUs), or a dedicated hardware device may perform the relevant calculations.
- Figure 7 depicts a system 10 for analyzing bone marrow in a subject, according to some embodiments.
- system 10 may be implemented as a software module, a hardware module, or any combination thereof.
- system may be or may include a computing device such as element 1 of Figure 6 and may be adapted to execute one or more modules of executable code (e.g., element 5 of Fig. 6) to analyze bone marrow in a subject, as further described herein.
- modules of executable code e.g., element 5 of Fig. 6
- arrows may represent flow of one or more data elements to and from system 10 and/or among modules or elements of system 10. Some arrows have been omitted in Figure 7 for the purpose of clarity.
- analyzing comprises producing a feature vector representing deviation of the subject’s cellular data from the control cellular data.
- system 10 may include, or may be communicatively connected to a single cell RNA sequencing (scRNA-seq) module or device 20, which may be configured to produce scRNA-seq data 20S (or “data 20S”, for short) as elaborated herein.
- scRNA-seq single cell RNA sequencing
- An analysis module 100 of system 10 may be configured to analyze data 20S, to extract a feature vector 150F.
- feature vector 150F may include one or more values indicative of a CD34 positive population in a peripheral blood sample of a subject (e.g., patient) of interest.
- feature vector 150F may include a plurality of entries, each corresponding to a specific cell type.
- the value of each entry of feature vector 150F may represent a relation to, or deviation from a reference value, or a range of cell populations.
- entries of a feature vector 150F pertaining to a specific subject may include values of stem cell population of that subject.
- entries of a feature vector 150F may include statistical numerical values representing deviation from a reference.
- Such a reference may include a mean (black, dashed line) and/or normal range (grey lines) of stem cell population in a cohort of subjects.
- analyzing comprises applying a trained Machine Learning (ML) based module 200, also referred to herein as a classifier 200, to the received dataset 20S. Additionally, or alternatively, analyzing may include applying ML 200 on feature vector 150F.
- the ML module is trained on a training set.
- the training set comprises the control dataset. In some embodiments, the training set comprises the plurality of cellular datasets.
- the ML 200 may output (e.g., via output device 8 of Fig. 6) an indication 30.
- Indication 30 may be, for example, a classification of the subject’s bone marrow.
- indication 30 may include a classification of the subject, an analysis of the subject’s bone marrow.
- indication 30 may include a notification regarding a health condition of the subject (e.g., healthy, or not), a diagnosis of the subject (e.g., a suspected pathology of the bone marrow), a prognosis of a subject’s condition, and the like.
- analyzing may include applying ML 200 to a parameter extracted from the dataset. In some embodiments, analyzing comprises applying a trained machine learning model to feature vector 150F.
- FIG. 7B is a block diagram depicting a nonlimiting example for implementation of system 10, according to some embodiments of the invention.
- System 10 of Figure 7B may be the same as system 10 of Figure7A.
- analysis module 100 may include a feature extraction module 110.
- feature extraction module 110 may be configured to extract, from data 20, a plurality of features 110F, or parameters pertaining to, or characterizing of a plurality of specific cells in peripheral blood samples. These features may be expression profiles of informative genes or other transcriptional data extracted from the scRNA-seq data. The features may be the whole transcriptome or informative parts of the transcriptome of the cells.
- Analysis module 100 may use features 110F to bin, or cluster features 110F to form high-level representations of cell population in the peripheral blood samples.
- a subject module 130 of analysis module 100 may be configured to produce at least one subject- specific model 130M.
- Subject- specific model 130M may pertain to a specific peripheral blood test, taken from a specific subject.
- subject- specific model 130M may include a plurality of metacell entities, each representing an abstraction of cell population data pertaining to that subject, as elaborated herein.
- a cohort reference generator module 120 of analysis module 100 may be configured to produce a reference data 120M, or cohort data model 120M, also referred to herein as an HSPC atlas 120M.
- reference data 120M may include a plurality of metacell entities, each representing an abstraction of cell population data pertaining to a cohort of subjects, as elaborated herein.
- analysis module 100 may include a projection module 150, configured to project, or compare features 110F of a specific subject of interest, as manifested by subject- specific model 130M, onto features of the cohort of subjects, as manifested by reference data (e.g., HSPC atlas) 120M.
- a projection module 150 configured to project, or compare features 110F of a specific subject of interest, as manifested by subject- specific model 130M, onto features of the cohort of subjects, as manifested by reference data (e.g., HSPC atlas) 120M.
- projection module 150 may produce a feature vector 150F, also denoted herein as a “normalcy vector” 150F. Normalcy vector 150F may be indicative of the specific subject’s condition.
- system 10 may infer classifier 200 on feature vector (e.g., normalcy vector) 150F, to produce indication 30 of Figure 7A.
- classifier 200 may be, or may include an ML-based classification model, that may be trained on a training dataset, that includes a plurality of labeled, or annotated normalcy vectors 150F.
- Annotation of normalcy vectors 150F may include, for example, an expert indication 30 (e.g., diagnosis) of corresponding peripheral blood samples.
- ML-based classification model 200 may be trained to produce indication 30 of incident normalcy vectors 150F by a supervised training scheme, using the annotations as supervisory data.
- system 10 may infer classifier 200 on subject- specific model 130M data, to produce indication 30.
- classifier 200 may be, or may include an ML-based classification model, that may be trained on a training dataset, that includes a plurality of labeled, or annotated subject- specific model 130M data entities.
- Annotations of subject-specific models 130M of the dataset may include, for example, expert indications 30 (e.g., diagnosis) of corresponding peripheral blood samples.
- ML-based classification model 200 may thus be trained to produce indication 30 by a supervised training scheme, using the annotations as supervisory data.
- system 10 may infer classifier 200 on feature vector (e.g., normalcy vector) 150F, to produce a prediction of blast level 210B in bone marrow.
- classifier 200 may be, or may include an ML-based classification model 210, that may be trained on a training dataset, that includes a plurality of labeled, or annotated normalcy vectors 150F.
- Annotation of normalcy vectors 150F may include levels of blasts 210B in bone marrows, corresponding to respective patient peripheral blood samples.
- ML-based classification model 210 may be trained to predict bone marrow blast levels 210B by a supervised training scheme, using the annotations as supervisory data.
- system 10 may include an auxiliary data extraction module 140 (or “auxiliary module 140” for short).
- auxiliary module 140 may be configured to produce, from data 20, auxiliary information 140A such as karyotype data 140A or mutational data, as known in the art.
- classifier module 200 may include an IPSS-M risk score calculation module 220, configured to calculate an IPSS- M risk score 220S based on the predicted bone-marrow blast level 210B, the calculated karyotype data 140A, mutational data and other clinical blood measurements, as known in the art.
- Figure 8 is a flow diagram depicting a method of analyzing bone marrow in a subject, by at least one processor, according to some embodiments.
- the at least one processor may receive a subject cellular dataset 20S based on single cell RNA sequencing (scRNA-seq) 20 of CD34 positive cells from peripheral blood of the subject.
- scRNA-seq single cell RNA sequencing
- the at least one processor may employ an analysis module 100 (e.g., as elaborated herein in relation to Figs. 7A, 7B), to analyze said received subject cellular dataset (e.g., 130M) in relation to a control dataset (e.g., 120M) comprising a plurality of cellular datasets.
- a control dataset e.g., 120M
- Each cellular dataset of said plurality may be based on scRNA- seq 20S of CD34 positive cells from peripheral blood of a healthy subject, wherein a deviation (e.g., feature vector, or normalcy vector 150F) of said subject cellular dataset from said control dataset may indicate a bone marrow pathology.
- Embodiments of the invention may thereby produce an indication 30, representing, or notifying detection of pathology of the bone marrow in the subject.
- the system is for evaluating bone marrow healthy. In some embodiments, the system is for measuring blast number in the bone marrow. In some embodiments, the system is a non-invasive system.
- the system comprises a scRNA sequencing device.
- sequencing device is a scRNA sequencer.
- the system comprises a non-transitory memory device, wherein modules of instruction code are stored.
- the system comprises at least one processor.
- the processor is associated with the memory device.
- the processor is configured to perform a method of the invention.
- the processor is configured to execute the modules of instruction code, whereupon execution of said modules of instruction code, the at least one processor is configured to perform a method of the invention.
- the method comprises obtaining from the scRNA sequencing device single cell transcriptomes from CD34 positive cells from peripheral blood. In some embodiments, the peripheral blood is from the subject. In some embodiments, the method comprises producing a cellular dataset from the obtained single cell transcriptomes. In some embodiments, the method comprises producing a cellular dataset based on the obtained single cell transcriptomes. In some embodiments, the method comprises producing a cellular dataset derived from the obtained single cell transcriptomes. In some embodiments, the method comprises analyzing the produced dataset. In some embodiments, the analyzing is in relation to a control dataset. In some embodiments, the method comprises accessing a control dataset. In some embodiments, the control dataset is a control database.
- the control dataset is a plurality of datasets.
- the method comprises outputting a finding.
- the finding is the health of the subject.
- the finding is the health of the bone marrow.
- the finding is healthy.
- the finding is the presence of bone marrow pathology.
- the finding is what the bone marrow pathology is.
- the finding is based on deviation or lack thereof of the subject dataset from the control dataset.
- “and/or” is to be taken as specific disclosure of each of the two specified features or components with or without the other.
- the term “and/or” as used in a phrase such as “A and/or B” is intended to include A and B, A or B, A (alone), and B (alone).
- the term “and/or” as used in a phrase such as “A, B, and/or C” is intended to include A, B, and C; A, B, or C; A or B; A or C; B or C; A and B; A and C; B and C; A (alone); B (alone); and C (alone).
- PB peripheral blood
- DNA production 50 ml of peripheral blood (PB) were drawn from each individual into lithiumheparin tubes. 1 ml of blood was used for DNA production, and the remaining volume was used for PBMC isolation via Ficoll, using Lymphoprep filled Sepmate tubes (StemCell technologies), followed by CD34 magnetic bead-based enrichment using the EasySep human CD34 positive selection kit II (StemCell technologies).
- This enrichment strategy was found to be simple and reproducible and it was chosen for several reasons: 1) RNA-seq data was most reproducible when cells were not sorted, but rather enriched-for using beads (lower mitochondrial gene fraction).
- CD34 purity could be highly regulated by this method, to achieve anywhere between 50-95% enrichment of CD34-positive cells, which could later be easily distinguished based on their single cell expression data.
- cell numbers - 50 ml of blood would yield anywhere between 50 to 100 million PBMCs following Ficoll, 1/1000 of which are expected to be CD34+, such that this population’s representation was increased from 0.1% in the periphery to at least 50% of cells loaded for analysis.
- scRNA-seq of CD34+ PBMCs Single cell RNA libraries were generated using the lOx genomics scRNA-seq platform (Chromium Next Gem single cell 3’ reagent kit V3.1). Chip loading was preceded by flow-cytometry to verify that enrichment was successful, and that enough CD34+CD45 int live cells were gathered. All blood samples were freshly drawn at the Weizmann Institute of Science on the morning of each experiment day, and time from blood draw to lOx loading was restricted to 5 hours. The motivation for working with fresh samples was based on previous experience with PB CD34+ cells being vulnerable to freezing/thawing rounds and long manipulation times.
- Genotype-based demultiplexing All cells were traced back to their sample of origin using genotype-based de-multiplexing. This method allowed pooling of blood samples immediately following extraction of the DNA aliquot, such that CD34-enrichment was performed on the entire pool of PBMCs produced.
- SNP-based multiplexing has several advantages to alternative antibody-based cell hashing methods: 1) it is extremely cost effective, such that the cost of sequencing a single individual on a 2000 SNP Molecular Inversion Probe (MIP) panel at a depth of 1000X per SNP (adequate for de-multiplexing purposes) is several folds cheaper than antibody staining, 2) genotyping eliminates the need to keep samples separated prior to loading, it entails shorter handling times and less cell manipulation, as it does not require antibody incubation periods and multiple wash centrifuges. This was very evident in cell viability prior to chip loading. As with other methods of sample multiplexing, genotype-based multiplexing allows for robust doublet detection during data analysis, which enabled loading of 30-40K cells from between 4-6 individuals on each Chromium Chip lane, yielding 15-25k cells per library.
- MIP Molecular Inversion Probe
- Molecular inversion probe (MIP) panels Both the CH and genotyping panels are Molecular inversion probes (MlP)-based panels described in detail previously in Biezuner, T. et al., “An improved molecular inversion probe based targeted sequencing approach for low variant allele frequency.” NAR Genom Bioinform 4, (2022) herein incorporated by reference in its entirety.
- the CH panel contains 705 probes, covering pre-leukemic SNVs and Indels in 47 genes, and is complemented by 2 amplicon sequencing reactions to cover GC rich regions in SRSF2 and ASXL1. As MIP sequencing is cost-effective yet noisy, an in-house variant-calling method was designed to identify low VAF CH events.
- the genotyping panel allows for the simultaneous detection of >2000 common genetic variants, all of which are extensively covered in all cell types in the data. It includes heterozygous sites with at least 5% minor allele frequency from the IK genomes project, which were highly covered by RNA molecules in the data (at least 80 UMIs across all cells in a test lOx library), excluding sites in repetitive elements and in sex chromosomes. Both panels were designed using MIPgen to ensure capture uniformity and specificity.
- CH sequencing of high RDW samples and controls In order to compare propensity for CH and high risk CH mutations in high RDW cases and normal RDW controls, deep targeted sequencing was performed on DNA samples from 602 high RDW (>15%) individuals, who did not show signs of anemia and whose blood count did not meet MDS criteria (11.5 g/dL ⁇ Hg ⁇ 15.5 g/dL [F], 13g/dL ⁇ Hg ⁇ 17g/dL [M], 80fL ⁇ MCV ⁇ 96fL, PLT>100xl09/L, Abs Neut>1.8xl09/L), and 602 normal RDW (11.5g/dL ⁇ Hg ⁇ 15.5g/dL [F], 13g/dL ⁇ Hg ⁇ 17g/dL[M], 80fL ⁇ MCV ⁇ 96fL, PLT>100xl09/L, Abs Neut>1.8*109/L), age and gender-matched controls.
- scRNA-seq processing fastq files were processed by executing cellranger- with an hg-38 reference genome. Cells were filtered with at least 20% mitochondrial expression and ⁇ 500 UMIs from unfiltered genes.
- Doublet calling Several steps were performed to assign cells to their individuals and to detect doublets. The pipeline is made of several steps:
- a metacell model was built with cells from all libraries. This model included cells that were already identified as doublets.
- the model was built with metacell (see Lee-Six, H. et al., “Population dynamics of normal human blood inferred from somatic mutations.” Nature 561, 473-478 (2016), herein incorporated by reference in its entirety), with a target metacell size of 200 cells. All metacells where at least 35% of the cells were already marked as doublets were then marked, and all metacells that expressed key markers of distinct cell types, as doublet metacells. All cells that belonged to a doublet metacell were then marked as doublets. An additional metacell model (see below) was then built, without cells that were marked as doublets.
- the metacell model was built with metacell 2, with a target metacell size of 200 cells. Histone, cell cycle, ribosomal, sex-linked, and stress response genes (including FOS, JUN) were marked as forbidden genes, as were genes with high technical variation, such as those with high or inconsistent differences between Illumina- and Ultima-sequenced technical replicates. These genes were not used for calculating gene-gene similarities but were included in downstream analyses. Metacells were annotated using known markers. Metacells with low CD34 expression, such as mature monocytes, B cells, T cells, NK cells, DCs, and endothelial cells were excluded from most downstream analyses. UMAP projections of the metacell expression vector over genes with specific enrichment over cell types were used for visualization of the metacell manifold.
- BM comparisons and projections Three BM datasets were used for comparison purposes: a dataset including CD34-enriched cells from 2 individual BMs collected by us and processed similarly to PB (Fig 1A); the Human Cell Atlas (HCA) dataset and a CD34+ bead-enriched BM datasets from Setty, et al. “Characterization of cell fate probabilities in single-cell data with Palantir”, Nat Biotechnol 37, 451-460 (2019), the contents of which are hereby incorporated by reference in their entirety.. The HCA datasets was previously processed and annotated in a metacell model.
- a metacell model for the 2 CD34-enriched BM samples collected was constructed using metacell, in a similar fashion to that described previously, and the Setty et al. sequencing data was downloaded and processed by running cellranger and a third BM metacell model was created from that data.
- the HCA metacells and the Setty and PB metacells were correlated over genes showing high variance over the HCA metacell model.
- Each Setty metacell was annotated using the mode of the 5 most correlated HCA metacells and expression of gene markers. Metacells from each of these models was projected on the HCA UMAP using the mean x and y values of its 5 most correlated HCA metacells.
- HSC differentiation gene programs To visualize transcriptional dynamics in HSC cells, MEBEMP and CLP metacells were sorted based on their AVP expression. To calculate differential expression (DE) between HSC and neighboring cell types, the geometric mean of each gene was calculated across HSCs, CLP-E and MPP metacells, and the difference between HSC and MPP, and between HSC and CLP-E was selected.
- DE differential expression
- each metacell was downsampled to have 90K UMIs and all UMIs of the metacell each cell belongs to were summed for each individual. This matched expression was normalized to sum to 1 and log2 was calculated. All genes that were expressed in either the observed or matched expression in any individual (log2 expression > 2 A -14.5), with at least a 2-fold change between observed and matched in at least one individual were plotted. Genes exhibiting strong batch effects were excluded.
- HSPC compositional analysis To explore variance in cell type composition between individuals, first the distribution of each individual’s cells across the CD34+ cell states were calculated. Further, cells from CD34+ states were partitioned into finer grained bins using one HSC bin, four CLP bins, and ten MEBEMP / MPP bins, for a total of 15 bins. HSC cells were assigned to bin 0, CLP-E cells to CLP bin 1, and CLP-M/L cells to CLP bins 2-4 based on an AVP expression gradient, such that each of these bins consisted of an equal number of cells. Similarly, MPP and MEBEMP-E/L cells were assigned into equal size MPP/MEBEMP bins 1-10 based on decreasing AVP expression.
- the bottom panel of Figure 2D shows individual enrichment across bins (log2 of the ratio each individual’s cell frequency in each bind to the median cell frequency in that bin across individuals). Individuals were partitioned into three group based on their mean enrichment across CLP bins 2-4 - those with mean enrichment > 0.5 are high CLP, those with ⁇ -0.5 are low, and the rest are intermediate.
- the sternness score was defined as the ratio between the number of cells in MPP / MEBEMP bins 1-5 and the total MPP / MEBEMP number (cells in bins 1-10). Individuals with sternness score > 0.5 had enriched sternness. Individuals within each cluster were further sorted based on their sternness score. The combination of CLP enrichment and sternness defines the six classes shown in the figure.
- Test for association between cell state compositions and a numerical label Permutation tests were used to test the association between cell state distribution and a label, such as CBC indices or sync-scores.
- a label such as CBC indices or sync-scores.
- Variably expressed gene modules Genes modules with high variance were detected across individuals while controlling or compositional variant. This was performed, separately for myeloid and lymphoid states, in the following manner:
- CLPs Fig. 11
- the analysis included cells from CLP-M metacells. The cells were partitioned into 6 equal size bins, and partitioning was based on the average of their DNTT and VPREB 1 expression. Genes with high absolute correlation to the average DNTT and VPREB 1 expression were excluded. This was followed by hierarchical clustering of the gene-gene correlation profiles, and removal of genes as described for MEBEMPs.
- Age regression models were developed for MEBEMP and CLP expression separately. To predict age, the difference between an individual’s observed and expected gene expression was used as described above. Genes with minimal expression >2 A - 14.5 for MEBEMPs and >2 A -15.5 for CLPs across individuals were used. A LASSO model was trained using nested leave-one-out cross validation. For each left-out sample cross validation was performed on the remaining samples to select LASSO’s A parameter, a model was trained using the selected A and a prediction was made on the left-out sample.
- LMNA signature The difference between an individual’s observed and expected gene expression was used and this difference was correlated to ALMNA separately for MEBEMPs and CLPs. The MEBEMP and CLP correlation values were then summed and genes whose summed correlation was >0.7 were kept. Further, genes with high technical variance were removed, resulting in retaining 17 genes in the LMNA signature. To calculate individual LMNA signatures, the average value of these 17 genes in the observed-expected matrix of each individual for MEBEMPs and CLPs were selected separately. To plot Figure 12A, the geometric mean of LMNA signature gene expression were calculated for each individual in each one of the 10 MPP/MEBEMP bins described earlier (Fig 2D).
- GoT 20 performed on sample N122 allowed the marking of this individual's cells as wild-type or mutated. Due to the low VAF of N122’s DNMT3A mutation, and in order to increase power, cells whose DNMT3A mutation status could not be determined by GoT were marked as wild-type cells. For Figure 2G, N122’s cells’ distribution across cell states was examined.
- Sync-score The AVP signature was defined to include genes with high correlation (> 0.6) to AVP across HSC, MPP and MEBEMP metacells, and the GATA1 signature to include those with high correlation (> 0.7) to GATA1. Genes with mean relative expression > 2 A -10 were filtered in these metacells, to preclude a small number of genes from dominating the signatures. All HSC, MPP, MEBEMP-E and MEBEMP-L cells was then scored by their fraction of its UMIs from the AVP and GATA1 signatures and all cells were partitioned into 20 equal-size bins of AVP signature expression and 20 equal-size bins of GATA1 signature expression. The sync-score is then defined as the fraction of cells in GATA1 bins 13 and above (upper two quintiles of GATA1) that are in AVP bins 9 and above (upper three quintiles of AVP expression).
- this 20 x 20 bin matrix was normalized to sum to 1, the obtained matrix was smoothed by averaging cells using a running window of length 3, and log2 calculated.
- Differential gene expression with respect to age and CBC Differential expression was performed separately for MPP / MEBEMP and CLP cells as well as for males and females.
- Individual gene expressions were correlated with age, max VAF of CH mutations and 20 CBC indices using Spearman correlation, and the correlation was tested for significance, p-values were FDR-corrected (Benjamini-Hochberg) for each label separately.
- max VAF a Mann- Whitney test comparing individuals with and without detected mutations was additionally performed.
- Differential expression between males and females was performed using a Mann-Whitney test on the same expression matrices.
- Patient scRNA-seq initial processing All patient- including lOx libraries were multiplexed with additional healthy samples. These were processed using cellranger as described previously. Doublets were detected using Vireo and Souporcell and cells were assigned to individuals as described above. All patient data was sequenced on the Ultima platform and was corrected by downsampling and resampling of UMIs as described above. A metacell model was then created for each of 12 samples separately: 2 healthy individuals, 2 MDS patients (one of which was a del5q patient sampled twice-before and after treatment initiation), 3 CMML patients, 1 MDS/MPN overlap patient, 1 myelofibrosis patient and 2 AML patients.
- Projection of disease data on the HSPC model To project patient metacells on the healthy reference, patient (query) metacells were correlated with reference metacells. Due to sequencing depth variability, query metacells were first downsampled to 150K UMIs per metacell. The correlation was performed in log2 scale using variable genes from the reference. Query metacells were then annotated using the mode (most common cell state) of the 5 reference metacells they were most correlated to. Query metacells that mapped to CD34-negative reference metacells were discarded from downstream analyses. Figure 4B shows the mean correlation between each query metacell and its 5 most correlated reference metacells and places each query metacell on the mean x and y coordinates of these metacells on the reference UMAP.
- Figures 4C-D are based on projection of single cells. Patient query cells were correlated to reference metacells using raw UMI vectors. Figure 4C then shows the distribution of projected metacell annotations (determined by the mode of each query cell’s 5 most correlated reference metacells). To create Figure 4D, each query cell was assigned to the bin (as in Fig. 2D) most common among cells in the metacell to which it mapped. To create Figure 4E, each patient’s observed expression was calculated as the geometric mean expression across all his metacells annotated as either MPP/MEBEMP- E/MEBEMP-L. To calculate expected expression, the geometric mean expression of the 5 most correlated reference metacells of each query metacell were selected.
- expected values were normalized by sorting genes based on their mean expression in the observed and expected and subtracting from the expected value a rolling mean of the difference between the expected and observed across 500 consecutive genes.
- Differentially expressed (DE) genes were defined as having >2-fold difference between observed and expected.
- AML-1 patient N 186 metacells were separated into AML-1-1 and AML-1-2 by their karyotype. Lists of differentially (over)expressed genes in healthy HSC, MEBEMP, NKTDP and CLP metacells were created. Each AML metacell was scored by the geometric mean of all genes in each of these cellstate specific gene lists, and of the LMNA signature genes. State-specific expression thresholds were then set by observing the expression of each state-specific gene program across all reference metacells belonging to the relevant cell state (e.g., NKTDP metacells for the NKTDP gene program, see dashed lines in FIG. 13A). To discover de-novo gene programs in the AML samples (FIG.
- Example 1 Universal stem and progenitor states observed across humans in CD34+ peripheral blood [0189] To evaluate interpersonal diversity in subtype distribution and regulation of circulating HSPCs (cHSPCs)from healthy humans, multiplexed scRNA-seq was combined with genotyping, and integrated clinical data. Multiplexing was resolved using SNPs identified in the 3’ UTR of cHSPC RNA facilitating precise matching of cells to individuals, and improving control for batch effects and doublets (Fig. 1A). Altogether, HSPCs from 79 males and 69 females between the ages of 23 and 91 years (median 61.5) were collected.
- cHSPCs circulating HSPCs
- BM bone marrow
- cHSPCs can serve as a highly accessible proxy for hematopoietic dynamics, both within and between individuals.
- Example 2 High resolution circulating HSC map shows HLF, GATA3, HOXB5 and TLE4 as distinct HSC TFs
- One of the hallmarks of this cHSPC model is a distinct HSC state that is transcriptionally linked with two major differentiation gradients: the first representing a continuum of common lymphoid progenitor (CLP) programs; the second, and more common branch, representing multipotent progenitor (MPP) states and their differentiation toward granulocyte-monocyte progenitors (GMPs), erythrocyte progenitors (ERYPs) and basophil/eosinophil/mast progenitors (BEMPs).
- CLP common lymphoid progenitor
- MPP multipotent progenitor
- GFP granulocyte-monocyte progenitors
- ERPs erythrocyte progenitors
- BEMPs basophil/eosinophil/mast progenitors
- MEBEMP megakaryocyte/erythrocyte/basophil/eosinophil/mast progenitors
- HSCs are marked by high AVP and HLF expression and were previously shown to represent a rare cell population with self-renewal capacity in BM and cord blood.
- This model included data on -14,440 HLF/ A VP HSCs that could be matched with cells from independent BM atlases, suggesting that under steady-state, HSCs with potential selfrenewal capacity are present in the peripheral blood.
- 14 genes were discovered that were expressed at least 1.75-fold higher in HSCs as compared to their two immediate differentiation branches.
- TFs transcription factors
- HSC state is defined by unique markers that are symmetrically down-regulated upon exit of the CLP and MEBEMP trajectories (Fig. 1C)
- Fig. ID the HSC state is defined by unique markers that are symmetrically down-regulated upon exit of the CLP and MEBEMP trajectories
- Fig. ID the HSC state is defined by unique markers that are symmetrically down-regulated upon exit of the CLP and MEBEMP trajectories
- Fig. ID lineage-specific regulators at intermediate levels, which are bifurcating anti-symmetrically upon exit from the HSC state to the CLP and MEBEMP trajectories
- Example 3 NK-T-dendritic and basophil-eosinophil-mast progenitors are enriched in circulating HSPCs
- the cHSPC atlas was enriched for basophil-eosinophil-mast progenitors (BEMP), mapped as one possible terminus of the HSC differentiation. While classical studies linked these cells with a granulocyte/monocyte progenitor (GMP) origin, more recent studies suggested these emerge, at least in part, from erythroid progenitors in mice and humans. This analysis allowed for focusing on a small population of metacells linking BEMPs with their MEBEMP-L precursors (Fig. IE). This highlighted TFs (Fig. IF) and other factors (Fig. 9) positively or negatively regulated in this postulated early stage of BEMP specification.
- BEMP basophil-eosinophil-mast progenitors
- NK/T/DC progenitors This subpopulation was therefore termed NK/T/DC progenitors (NKTDP).
- NKTDP NK/T/DC progenitors
- Example 4 Inter-individual variation in cHSPC sternness and in lymphoid/myeloid differentiation bias
- lymphoid progenitor frequencies CLP-M, CLP-L, NKTDP
- MEBEMP-E MEBEMP-L, ERYP, BEMP
- each individual’s enrichment was profiled over the MEBEMP and CLP trajectories. Clustering of these enrichment profiles yielded six archetypes of cHSPC composition within the healthy population (classes I - VI) (Fig. 2D) These were composed of individuals with relative lymphoid enrichment (class I- III) or depletion (class V-VI) further subdivided by a sternness gradient, enriched in classes II, IV and VI, and depleted in class I, III and V. Analysis of technical and biological replicates confirmed this variation to be robust and patient specific. To summarize, the instant analysis provides the first cHSPC subpopulation normal reference range (Fig. 2B), characterized by extensive variation among healthy individuals, and show these compositional differences are a true individual characteristic, with potential clinical implications.
- RBC indices including RBC, hematocrit (HCT), mean corpuscular volume (MCV), and RDW, were analyzed separately in males and females.
- a significant negative correlation (P ⁇ 0.02) was observed between CLP frequencies and HCT (males, Fig. E- middle).
- P ⁇ 0.01 a significant positive correlation between increased RDW- a cHSPC myeloid bias- and a relative CLP depletion (males, Fig. 2E-right).
- Example 7 Composition-controlled HSPC expression correlates with age
- cHSPC composition provides an initial blueprint of hematopoietic dynamics along the sternness and CLP/MEBEMP axes. Further analysis of transcriptional variation could now be carried out, while controlling for the dominant effect of cHSPC composition, in order to characterize additional gene expression signatures that could distinguish between individuals. Composition-controlled individual expression profiles showed high information content when correlated with age, enabling age prediction based on normalized expression alone (Fig 3E, 10). Next, a search for gene groups (signatures) that co-variate between individuals was performed, filtering out sex-linked signatures and those showing strong batch effects.
- LMNA Lamin-A
- ANXA1 annexin Al
- AHNAK AHNAK nucleoprotein
- MY ADM myeloid associated differentiation marker
- TSPAN2 tetraspanin 2
- VIM vimentin
- Individual LMNA signature expression varied across a range of more than 2-fold (Fig. 12A), exhibiting high expression variability in HSCs and early myeloid and lymphoid cell states, and a homogeneously low expression in late MEBEMPs and CLPs.
- Individual LMNA signature expression was consistent in myeloid and lymphoid cell states (Fig. 3G) and was stable in a follow-up cohort (Fig. 12B).
- Example 8 Rapid repression of sternness signatures in MEBEMPs is linked with lower red cell counts and higher red cell volumes [0199]
- the differentiation of HSPCs toward MEBEMP and CLP fates involves coordinated activation of specific transcriptional programs that were generally universal among individuals. Yet, the screen for individual-specific gene signatures suggested that individuals differed in the way they synchronized the opposing effects of these sternness and differentiation programs.
- AVP sternness
- GATA1 MEBEMP differentiation
- Example 9 Age-related perturbation of HSPC composition and transcriptional signatures
- Example 10 Using the cHSPC atlas for mapping, dissecting and annotating myeloid malignancies
- FIG. 5A there is shown a new stepwise approach for analysis of myeloid disorders based on sampling of cHSPCs and comparison to their compositions, normalized expression and copy number variations (karyotyping) to the new normal reference.
- CMML chronic myelomonocytic leukemia
- MDS myeloproliferative neoplasm
- MF myelofibrosis
- Compositional- controlled gene expression comparisons of patient samples to the normal model identified specific genes that were recurrently induced or repressed in disease (Fig. 5E). While both healthy individuals and the treated MDS case showed a minimal number of DE genes, all leukemic cases exhibited a substantial increase in the number of DE genes when compared to the healthy population model (Fig. 5E).
Landscapes
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Medical Informatics (AREA)
- Physics & Mathematics (AREA)
- General Health & Medical Sciences (AREA)
- Public Health (AREA)
- Organic Chemistry (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Data Mining & Analysis (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Analytical Chemistry (AREA)
- Biophysics (AREA)
- Biotechnology (AREA)
- Epidemiology (AREA)
- Theoretical Computer Science (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Genetics & Genomics (AREA)
- Immunology (AREA)
- Databases & Information Systems (AREA)
- Pathology (AREA)
- Software Systems (AREA)
- Biomedical Technology (AREA)
- Molecular Biology (AREA)
- General Engineering & Computer Science (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Microbiology (AREA)
- Primary Health Care (AREA)
- Biochemistry (AREA)
- Bioethics (AREA)
- Oncology (AREA)
- Medicinal Chemistry (AREA)
- Chemical Kinetics & Catalysis (AREA)
Abstract
Description
Claims
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24736111.6A EP4724599A1 (en) | 2023-06-08 | 2024-06-09 | Non-invasive bone marrow diagnostics |
| IL325192A IL325192A (en) | 2023-06-08 | 2024-06-09 | Non-invasive bone marrow diagnostics |
| US19/411,342 US20260094717A1 (en) | 2023-06-08 | 2025-12-07 | Non-invasive bone marrow diagnostics |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| IL303582A IL303582A (en) | 2023-06-08 | 2023-06-08 | Non-invasive bone marrow diagnostics |
| IL303582 | 2023-06-08 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US19/411,342 Continuation US20260094717A1 (en) | 2023-06-08 | 2025-12-07 | Non-invasive bone marrow diagnostics |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024252405A1 true WO2024252405A1 (en) | 2024-12-12 |
Family
ID=91664804
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/IL2024/050568 Ceased WO2024252405A1 (en) | 2023-06-08 | 2024-06-09 | Non-invasive bone marrow diagnostics |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20260094717A1 (en) |
| EP (1) | EP4724599A1 (en) |
| IL (2) | IL303582A (en) |
| WO (1) | WO2024252405A1 (en) |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US4666828A (en) | 1984-08-15 | 1987-05-19 | The General Hospital Corporation | Test for Huntington's disease |
| US4683202A (en) | 1985-03-28 | 1987-07-28 | Cetus Corporation | Process for amplifying nucleic acid sequences |
| US4801531A (en) | 1985-04-17 | 1989-01-31 | Biotechnology Research Partners, Ltd. | Apo AI/CIII genomic polymorphisms predictive of atherosclerosis |
| US5192659A (en) | 1989-08-25 | 1993-03-09 | Genetype Ag | Intron sequence analysis method for detection of adjacent and remote locus alleles as haplotypes |
| US5272057A (en) | 1988-10-14 | 1993-12-21 | Georgetown University | Method of detecting a predisposition to cancer by the use of restriction fragment length polymorphism of the gene for human poly (ADP-ribose) polymerase |
| WO2018016656A1 (en) * | 2016-07-20 | 2018-01-25 | 一般社団法人 東京血液疾患研究所 | Method for diagnosing myelodysplastic syndrome |
| US20210246503A1 (en) * | 2018-06-11 | 2021-08-12 | The Broad Institute, Inc. | Lineage tracing using mitochondrial genome mutations and single cell genomics |
-
2023
- 2023-06-08 IL IL303582A patent/IL303582A/en unknown
-
2024
- 2024-06-09 WO PCT/IL2024/050568 patent/WO2024252405A1/en not_active Ceased
- 2024-06-09 EP EP24736111.6A patent/EP4724599A1/en active Pending
- 2024-06-09 IL IL325192A patent/IL325192A/en unknown
-
2025
- 2025-12-07 US US19/411,342 patent/US20260094717A1/en active Pending
Patent Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US4666828A (en) | 1984-08-15 | 1987-05-19 | The General Hospital Corporation | Test for Huntington's disease |
| US4683202A (en) | 1985-03-28 | 1987-07-28 | Cetus Corporation | Process for amplifying nucleic acid sequences |
| US4683202B1 (en) | 1985-03-28 | 1990-11-27 | Cetus Corp | |
| US4801531A (en) | 1985-04-17 | 1989-01-31 | Biotechnology Research Partners, Ltd. | Apo AI/CIII genomic polymorphisms predictive of atherosclerosis |
| US5272057A (en) | 1988-10-14 | 1993-12-21 | Georgetown University | Method of detecting a predisposition to cancer by the use of restriction fragment length polymorphism of the gene for human poly (ADP-ribose) polymerase |
| US5192659A (en) | 1989-08-25 | 1993-03-09 | Genetype Ag | Intron sequence analysis method for detection of adjacent and remote locus alleles as haplotypes |
| WO2018016656A1 (en) * | 2016-07-20 | 2018-01-25 | 一般社団法人 東京血液疾患研究所 | Method for diagnosing myelodysplastic syndrome |
| US20210246503A1 (en) * | 2018-06-11 | 2021-08-12 | The Broad Institute, Inc. | Lineage tracing using mitochondrial genome mutations and single cell genomics |
Non-Patent Citations (16)
| Title |
|---|
| "Cell Biology: A Laboratory Handbook'', Volumes I- III Cellis", 1994, APPLETON & LANGE |
| "Strategies for Protein Purification and Characterization - A Laboratory Course Manual", 1996, CSHL PRESS |
| BARAN ET AL.: "MetaCell: analysis of single-cell RNA-seq data using K-nn graph partitions", GENOME BIOL., vol. 20, no. 1, 11 October 2019 (2019-10-11), pages 206 |
| BARYAWNO NINIB ET AL: "A Cellular Taxonomy of the Bone Marrow Stroma in Homeostasis and Leukemia", CELL, vol. 177, no. 7, 13 June 2019 (2019-06-13), pages 1915, XP085712698, ISSN: 0092-8674, DOI: 10.1016/J.CELL.2019.04.040 * |
| BEN-KIKI ET AL.: "Metacell-s: a divide and conquer metacell algorithm for scalable scRNA-seq analysis", GENOME BIOL., vol. 23, no. 1, 19 April 2022 (2022-04-19), pages 100 |
| BIRREN ET AL.: "Genome Analysis: A Laboratory Manual Series", 1998, COLD SPRING HARBOR LABORATORY PRESS, pages: 1 - 4 |
| DE BIE JOLIEN ET AL: "Single-cell sequencing reveals the origin and the order of mutation acquisition in T-cell acute lymphoblastic leukemia", BLOOD CANCER JOURNAL, NATURE PUBLISHING GROUP UK, LONDON, vol. 32, no. 6, 18 April 2018 (2018-04-18), pages 1358 - 1369, XP036519479, ISSN: 0887-6924, [retrieved on 20180418], DOI: 10.1038/S41375-018-0127-8 * |
| DINGMORRISON: "Haematopoietic stem cells and early lymphoid progenitors occupy distinct bone marrow niches", NATURE, vol. 495, no. 7440, 14 March 2013 (2013-03-14), pages 231 - 235 |
| HERBIG MAIK ET AL: "Machine learning assisted real-time deformability cytometry of CD34+ cells allows to identify patients with myelodysplastic syndromes", SCIENTIFIC REPORTS, vol. 12, no. 1, 18 January 2022 (2022-01-18), US, XP093203035, ISSN: 2045-2322, Retrieved from the Internet <URL:https://www.nature.com/articles/s41598-022-04939-z> DOI: 10.1038/s41598-022-04939-z * |
| LEE-SIX, H. ET AL.: "Population dynamics of normal human blood inferred from somatic mutations", NATURE, vol. 561, 2018, pages 473 - 478, XP036600603, DOI: 10.1038/s41586-018-0497-0 |
| PERBAL: "A Practical Guide to Molecular Cloning", 1988, JOHN WILEY & SONS |
| PETTI ET AL.: "A general approach for detecting expressed mutations in AML cells using single cell RNA sequencing", NATURE COMMUNICATIONS, vol. 10, no. 3660, 2019 |
| SAMBROOK ET AL.: "Current Protocols in Molecular Biology", 1989, JOHN WILEY AND SONS |
| SETTY ET AL.: "Characterization of cell fate probabilities in single-cell data with Palantir", NAT BIOTECHNOL, vol. 37, 2019, pages 451 - 460, XP036888454, DOI: 10.1038/s41587-019-0068-4 |
| WEISSBEIN ET AL.: "Analysis of chromosomal aberrations and recombination by allelic bias in RNA-Seq", NATURE COMMUNICATIONS, vol. 7, no. 12144, 2016 |
| ZHAO XIN ET AL: "Comprehensive analysis of single-cell RNA sequencing data from healthy human marrow hematopoietic cells", BMC RESEARCH NOTES, vol. 13, no. 1, 1 December 2020 (2020-12-01), GB, XP093202001, ISSN: 1756-0500, Retrieved from the Internet <URL:https://epo.summon.serialssolutions.com/2.0.0/link/0/eLvHCXMwrV1db9MwFLXYJgQvfA8KAwVeeICsieOPRJqQysQGLwNNGxJPlu3Y27Q2qdpVqP-In8m9TlKawQNIvLXxiWq719f3JPceE5LR3SS-5hNYCWRZaKalTYT3VmABZWGAb3CDku74iHefnRzK40PWpUNhaYyZQOw0r2oIu3bX69HHwYnDB3s5nJa-Wfu5GM5TCFdYjJwIhUtkvNwgWxChczD4rdOjL6NvoUCSA4vmSdIV0fzxxt5> DOI: 10.1186/s13104-020-05357-y * |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4724599A1 (en) | 2026-04-15 |
| IL325192A (en) | 2026-02-01 |
| IL303582A (en) | 2025-01-01 |
| US20260094717A1 (en) | 2026-04-02 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Miller et al. | Programs, origins and immunomodulatory functions of myeloid cells in glioma | |
| Kwok et al. | Neutrophils and emergency granulopoiesis drive immune suppression and an extreme response endotype during sepsis | |
| Leday et al. | Replicable and coupled changes in innate and adaptive immune gene expression in two case-control studies of blood microarrays in major depressive disorder | |
| Bruno et al. | Molecular signatures of self-renewal, differentiation, and lineage choice in multipotential hemopoietic progenitor cells in vitro | |
| Beretta et al. | Genome-wide whole blood transcriptome profiling in a large European cohort of systemic sclerosis patients | |
| Dussiau et al. | Hematopoietic differentiation is characterized by a transient peak of entropy at a single-cell level | |
| Wu et al. | Long noncoding RNAs of single hematopoietic stem and progenitor cells in healthy and dysplastic human bone marrow | |
| Zhang et al. | Single-cell analysis highlights a population of Th17-polarized CD4+ naïve T cells showing IL6/JAK3/STAT3 activation in pediatric severe aplastic anemia | |
| Liu et al. | A distinct glycerophospholipid metabolism signature of acute graft versus host disease with predictive value | |
| Bandyopadhyay et al. | Identification of functionally primitive and immunophenotypically distinct subpopulations in secondary acute myeloid leukemia by mass cytometry | |
| Furer et al. | A reference model of circulating hematopoietic stem cells across the lifespan with applications to diagnostics | |
| Myers et al. | Integrated single-cell genotyping and chromatin accessibility charts JAK2V617F human hematopoietic differentiation | |
| Chen et al. | Single-cell transcriptomic analysis of immune cell dynamics in the healthy human endometrium | |
| Wu et al. | Human autoimmunity at single cell resolution in aplastic anemia before and after effective immunotherapy | |
| Li et al. | Basal-to-inflammatory transition and tumor resistance via crosstalk with a pro-inflammatory stromal niche | |
| Sun et al. | Minimal residual disease monitoring in acute myeloid leukemia: Focus on MFC‐MRD and treatment guidance for elderly patients | |
| Cheng et al. | The role of CD101 and Tim3 in the immune microenvironment of gastric cancer and their potential as prognostic biomarkers | |
| US20260094717A1 (en) | Non-invasive bone marrow diagnostics | |
| CN120870559A (en) | Biomarkers for prognosis and treatment selection of acute myeloid leukemia and uses thereof | |
| Sui et al. | Single‐Cell Multiomics Reveals TCR Clonotype‐Specific Phenotype and Stemness Heterogeneity of T‐ALL Cells | |
| Ni et al. | Transcriptome and single-cell analysis reveal the contribution of immunosuppressive microenvironment for promoting glioblastoma progression | |
| Bruins et al. | Immune signatures in older patients with newly diagnosed multiple myeloma are associated with survival outcomes of first‐line therapy irrespective of frailty levels | |
| Latorre-Crespo et al. | Clinical progression of clonal hematopoiesis is determined by a combination of mutation timing, fitness, and clonal structure | |
| WO2025257825A1 (en) | Non-invasive myelodysplastic syndrome diagnostics | |
| Roostaei et al. | Cell type-and state-resolved immune transcriptomic profiling identifies glucocorticoid-responsive molecular defects in multiple sclerosis T cells |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24736111 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 325192 Country of ref document: IL |
|
| ENP | Entry into the national phase |
Ref document number: 2025571482 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2025571482 Country of ref document: JP |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024736111 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2024736111 Country of ref document: EP Effective date: 20260108 |
|
| ENP | Entry into the national phase |
Ref document number: 2024736111 Country of ref document: EP Effective date: 20260108 |
|
| WWP | Wipo information: published in national office |
Ref document number: 325192 Country of ref document: IL |
|
| ENP | Entry into the national phase |
Ref document number: 2024736111 Country of ref document: EP Effective date: 20260108 |
|
| WWP | Wipo information: published in national office |
Ref document number: 2024736111 Country of ref document: EP |