WO2024246152A1 - Method for determining proportions of populations in an ensemble of biological objects - Google Patents

Method for determining proportions of populations in an ensemble of biological objects Download PDF

Info

Publication number
WO2024246152A1
WO2024246152A1 PCT/EP2024/064824 EP2024064824W WO2024246152A1 WO 2024246152 A1 WO2024246152 A1 WO 2024246152A1 EP 2024064824 W EP2024064824 W EP 2024064824W WO 2024246152 A1 WO2024246152 A1 WO 2024246152A1
Authority
WO
WIPO (PCT)
Prior art keywords
vector
matrix
data
label
cell
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/EP2024/064824
Other languages
French (fr)
Inventor
Bastien DUSSAP
Gilles BLANCHARD
Baptiste LABARTHE
Badr-Eddine CHERIEF-ABDELLATIF
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Centre National de la Recherche Scientifique CNRS
Metafora Biosystems SAS
Universite Paris Saclay
Original Assignee
Centre National de la Recherche Scientifique CNRS
Metafora Biosystems SAS
Universite Paris Saclay
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Centre National de la Recherche Scientifique CNRS, Metafora Biosystems SAS, Universite Paris Saclay filed Critical Centre National de la Recherche Scientifique CNRS
Publication of WO2024246152A1 publication Critical patent/WO2024246152A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B40/00ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding

Definitions

  • the present invention relates to the field of cytometric data analysis, including single cell data analysis.
  • BACKGROUND OF INVENTION [0002]
  • the classical descriptive and reductionist approach (one gene, one messenger RNA, one protein) has been replaced by a more global understanding of biological systems based on the analysis of sets of biological elements ("-omes").
  • the basic idea associated with "omics” approaches is to apprehend the complexity of living organisms as a whole, using methodologies that are the least restrictive possible in terms of description.
  • genomics projecting of genes
  • transcriptomics analysis of gene expression and its regulation
  • proteomics projecting of proteins
  • metabolomics analysis of metabolites
  • metabonomics projecting of metabolic profiles in vivo.
  • Genomics is divided into two branches: structural genomics, which relates to the sequencing of the entire genome, and functional genomics, which aims to determine the function and expression of the sequenced genes.
  • functional genomics techniques are applied to a large number of genes in parallel: for example, the phenotype of mutants for a whole family of genes, or the expression of all the genes of an entire organism, can thus be analyzed.
  • Transcriptomics is the study of all the messenger RNAs produced during the genetic transcription process.
  • Proteomics is the analysis of all the proteins of an organelle, a cell, a tissue, an organ or an organism under given conditions. Proteomics attempts to identify in a global way the proteins extracted from a cell culture, a tissue or a biological fluid, their localization in the cellular compartments, their possible post-translational modifications, as well as their quantity. It makes it possible to quantify the variations in their level of expression, for example as a function of time, of their environment, of their state of development, of their physiological and pathological state, of the species of origin, etc.
  • Metabolomics studies all metabolites (sugars, amino acids, fatty acids, etc.), such as metabolic substrates, intermediates, products, as well as hormones, other signaling molecules and secondary metabolites, present in a cell, an organ, an organism.
  • Metabonomics consists of monitoring metabolic profiles in vivo, which makes it possible to provide information on the toxicity of drugs, on pathological processes, and on the function of genes.
  • Cytomics corresponds to the analysis of the cytome, the set of cellular constituents, in particular morphological, antigenic and functional, which make it possible to define or model the state and the functioning of a cell at a given time.
  • the analysis of the cytome makes it possible to describe the structural and functional heterogeneity of the various cells of an organism.
  • the preceding approaches make it possible to obtain a great amount of information on the cellular and/or tissue response to an in vitro or in vivo exposure.
  • Flow cytometry is a technique for analyzing a biological fluid comprising a collection of biological objects, such as cells, vesicles or particles, in suspension. Flow cytometry is an essential analytical technique for the identification and characterization of populations of cells, vesicles or particles.
  • individual cell heterogeneity within a cell population can be critical to its specific function and fate.
  • Cell-to-cell variations for instance in RNA transcripts or protein expression, may be key in research and development studies in the field of cancer, neurobiology, stem cell biology, immunology and developmental biology.
  • Conventional cell-based assays mainly analyze the average responses from a population of cells, without providing individual cell phenotypes.
  • Single-cell analysis is the study of genomics, transcriptomics, proteomics, metabolomics and cell–cell interactions at the single cell level. Analyzing a single cell makes it possible to discover mechanisms not seen when studying a cell population. High- throughput and multiparameter approaches may be combined in single cell analysis in order to reflect cell-to-cell variability and heterogeneous differences in the individual cells.
  • Technologies such as fluorescence-activated cell sorting (FACS) allow the precise isolation of selected single cells from complex samples, while high throughput single cell partitioning technologies, enable the simultaneous molecular analysis of hundreds or thousands of single unsorted cells.
  • FACS fluorescence-activated cell sorting
  • flow cytometry can address the issue of cellular processes and intercellular interactions because cytometry provides a single-cell view of more complex cellular conglomerates, such as tissues in multicellular organisms or communities in the case of single-cell bacteria, yeasts, algae, etc.
  • cytometry provides a single-cell view of more complex cellular conglomerates, such as tissues in multicellular organisms or communities in the case of single-cell bacteria, yeasts, algae, etc.
  • RNA sequencing Single ⁇ cell RNA sequencing (scRNA ⁇ seq) technology has become a state ⁇ of ⁇ the ⁇ art approach for unravelling the heterogeneity and complexity of RNA transcripts within individual cells, as well as revealing the composition of different cell types and functions within tissues, organs or organisms.
  • In situ sequencing and fluorescence in situ hybridization do not require that cells be isolated and may be used for analysis of tissues.
  • Other technologies enable quantifying thousands of proteins across hundreds of single cells. For example, mass spectrometry techniques have become important analytical tools for proteomic and metabolomic analysis of single cells.
  • Individual cells differentially use their genes, RNA, and proteins across tissue types. Spatial biology designates a series of techniques providing detailed cellular information according to the positional context of cells in a tissue.
  • Spatial transcriptomics is a technology that allows measuring all the gene activity in a tissue sample and spatially map where the activity is occurring. It combines unbiased, high-throughput total mRNA analysis for intact tissue sections with morphological context. Spatial proteomics combines microscopy with detection of multiple proteins, usually through the use of specific antibodies. It allows the integration of morphological data and protein expression data at the single-cell level. Detection methods may be based for example on mass spectrometry, on chromogenic staining, or on oligonucleotide barcodes.
  • the quantification of the different populations of biological objects contained in this ensemble is of particular interest.
  • the quantification of the different populations of biological objects contained in an ensemble of analyzed biological objects is generally one of the first assessed parameters. Indeed, the quantification of the different populations of biological objects may for instance provide information on the validity of the data obtained during a specific assay, or on the reproducibility of this kind of assay.
  • This invention thus relates to computer-implemented method for determining proportions ⁇ i of C reference populations Pi in an ensemble of biological objects being chosen among cells, vesicles of cellular origin, acellular microorganisms, and/or biofunctionalized materials, C being an integer number, said ensemble of biological objects being associated with a set of test data, the method comprising: - receiving at least one set of reference data associated with a reference ensemble of biological objects, wherein each reference data is labelled with a label Li belonging to a plurality of C labels Lj, each label Lj among the plurality of C labels Lj being associated with one corresponding reference population Pj among the C reference populations, - determining said proportions ⁇ i by minimizing a distance between a first vector and a second vector, said first vector being a weighted sum of C transformed vectors ⁇ ⁇ ⁇ ⁇ for
  • a matrix comprising a sample composed of the reference/test data associated with a reference ensemble of biological objects
  • said matrix comprises respective data/characteristics of said (biological) sample.
  • said matrix may comprise rows and columns when each row corresponds to one biological object of said ensemble and each column corresponds to a characteristic of corresponding biological objects.
  • a minimization problem is a quadratic programming (QP) in dimension c and can be solved efficiency by any known method in the literature.
  • the predetermined function z is based on a kernel function k.
  • computing P involves a computation of nxn values k(xil,xjt), wherein xil denotes an i th vector with the label Ll in a matrix P0 and xjt denotes a j th vector with the label Lt of the matrix P0, the matrix P0 being a concatenation of the ⁇ ⁇ matrices ⁇ ⁇ , n being equal to ⁇ ⁇ ⁇ ⁇ , and, Q comprising m rows y j , computing q involves a computation of n x m values k(x i ,y j ).
  • the kernel function k is a Gaussian function.
  • the number of calculations performed in the method is advantageously reduced. Therefore, as a consequence, the computation time is reduced in comparison to prior art methods, as will be illustrated further. Furthermore, in comparison to machine learning models, no hyperparameter is when two sets of test data have to be compared, once the C transformed and the second vector have been computed, the computation of an Euclidean distance is sufficient to perform the comparison, as compared to prior art methods as will be illustrated further. The comparison is therefore simpler, and thus faster.
  • the kernel function k is a Gaussian function. Having a kernel function k which is a Gaussian allows for easy computation of the distribution ⁇ k.
  • the method further comprises: - receiving a set of K clusters G k associated with the set of test data, k being integer varying from 1 to K, and G k denoting the k-th cluster, wherein each cluster Gk in the set of K clusters Gk is associated with a population and is represented with a matrix Xk, - computing, for each cluster Gk and for each vector ⁇ ⁇ , a distance d_ki between z(Xk) and ⁇ , - determining, for each cluster Gk, the reference population Pp k for which the distance dkp is minimal, wherein the reference population Pp k for is associated with the label Lp k , - assigning the label Lpk to the k-th cluster Gk.
  • the method advantageously allows for match labeling.
  • the method allows, for instance, using the predetermined function defined as a vectorial function previously described, to label unlabelled clusters obtained from a segmentation method of cytometric data of an ensemble of biological objects.
  • the biological objects are: - cells of animal, plant, fungal, protist, bacterial or archaebacterial origin, - vesicles of cellular origin chosen from exosomes, ectosomes, microvesicles, microparticles, prostasomes, oncosomes, matrix/calcification vesicles or apoptotic bodies, - acellular microorganisms chosen from viruses, viroids and prions, and/or - biofunctionalized materials comprising a material of synthetic or biological origin chosen from a nanoparticle (such as a nanobead, a nanosphere or a nanocapsule), a microparticle (such as a microbead, a microsphere or a microcapsule), a lipid vesicle (such as a unilamellar vesicle, a multilamellar vesicle, a lipoplex, a polyplex, a lipopolyplex, a lip
  • the reference and test data are obtained by flow cytometry (or FACS for "fluorescence activated cell sorting"), by PCR-activated cell sorting (PACS), by microsphere affinity proteomics (MAP), by mass spectrometry, by chromatography, by CYTOF, by spectral cytometry, by mass cytometry, by image cytometry, by gene expression on chips (microarray), by sequencing (for example DNA-seq, RNA-seq, scDNA-seq, or scRNA-seq), by in situ hybridization, and/or by microscopy.
  • the reference and test data are obtained by single-cell analysis technologies or spatial biology analysis technologies.
  • Another aspect of the invention pertains to a device comprising a processor configured to carry out the method previously described.
  • Another aspect of the invention pertains to a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method previously described.
  • Another aspect of the invention relates to a non-transitory computer-readable recording medium comprising instructions which, when executed by a computer, cause the computer to carry out steps of the method previously described.
  • processor or “processing unit” should not be construed as being limited to hardware capable of executing software, and refers generally to a processing device, which may include, for example, a computer, microprocessor, integrated circuit, or programmable logic device (PLD).
  • the processor may also include one or more graphics processing units (GPUs), whether they are used for computer graphics and image processing or other functions.
  • GPUs graphics processing units
  • the instructions and/or data for executing the associated and/or resulting functionality may be stored on any media readable by the processor such as, for example, an integrated circuit, a hard disk, a CD (Compact Disc), an optical disk such as a DVD (Digital Versatile Disc), a RAM (Random-Access Memory) or a ROM (Read-Only Memory).
  • instructions may be stored in hardware, software, firmware or any combination thereof.
  • Cytometry generally refers to the detection and/or measurement of characteristics of a cell, which is the basic structural, functional and biological unit of living organisms.
  • cytometry also refers to the detection and/or measurement of characteristics of vesicles of cellular origin (such as extracellular or intracellular vesicles), or of particles of nano- or micrometer size such as acellular microorganisms (e.g., viruses), or bio-functionalized materials (e.g., microspheres coated with biological molecules such as proteins, or microspheres coated with bio-active molecules such as immunomodulatory molecules or drugs).
  • acellular microorganisms e.g., viruses
  • bio-functionalized materials e.g., microspheres coated with biological molecules such as proteins, or microspheres coated with bio-active molecules such as immunomodulatory molecules or drugs.
  • the "cells” may come from a unicellular or pluricellular organism, of prokaryotic or eukaryotic origin, of animal, vegetable, fungal, protist, bacterial or archaebacterial origin. They can be living, dead or fixed.
  • the cells may, for example, come from solid or liquid tissues (such as bone marrow, blood or lymph, etc), or from body fluids (such as cerebrospinal fluid, urine, bronchoalveolar fluid, etc).
  • the cells, vesicles or particles may be in suspension (aqueous or non-aqueous, biological or synthetic suspension), or immobilized on a solid support (e.g., a flask, a cell culture dish with one or more wells, a plate or microplate, a glass slide, a chip, etc.). They may have been pre-treated with biological or chemical agent(s).
  • the cells, vesicles or particles to be analyzed may be comprised within one or more populations.
  • a "population” is an ensemble of cells, vesicles, or particles with same or similar characteristics. A population can be subdivided into different subpopulations with secondary characteristics different from each other.
  • Cell-derived vesicles are vesicles composed of a membrane of cellular origin. It can be intracellular or extracellular vesicles. Extracellular vesicles have generally been secreted or excreted by a cell.
  • Cell-derived vesicles include exosomes, ectosomes, microvesicles, microparticles, prostasomes, oncosomes, matrix/calcification vesicles, apoptotic bodies and other subsets of cellular vesicles.
  • Acellular microorganisms designate organisms of microscopic size whose structure is not cellular.
  • Acellular microorganisms include viruses, viroids and prions.
  • bio-functionalized materials or “functionalized biomaterials” or “bio- functionalized devices” or “bio-functionalized systems” designate any type of object comprising a material of synthetic or organic origin coated on its surface, coupled or linked to one or more molecule(s) or functional group(s) of biological origin.
  • the material of synthetic or biological origin can be a nanoparticle (such as a nanobead, a nanosphere or a nanocapsule), a microparticle (such as a microbead, a microsphere or a microcapsule), a lipid vesicle such as a unilamellar vesicle, a multilamellar vesicle, a lipoplex, a polyplex, a lipopolyplex a liposome, a niosome, a cochleate, a virosome, an immune-stimulating complex (ISCOM®), etc.
  • a nanoparticle such as a nanobead, a nanosphere or a nanocapsule
  • a microparticle such as a microbead, a microsphere or a microcapsule
  • a lipid vesicle such as a unilamellar vesicle, a multilamellar vesicle, a lipoplex
  • Nanoparticles and microparticles can be organic, inorganic, magnetic or radioactive.
  • nanoparticles or microparticles can be made of metal, silica, aluminium, titanium, glass, ceramic, polystyrene, poly(methyl methacrylate), melamine, polylactide, latex, dextran, oxide, graphene, magnetic material, radioactive material, or a combination thereof.
  • the material can be combined with, or coated with, one or more peptide(s), protein(s), antibody(ies), fragment(s) of antibody, receptor(s), cytokine(s), chemokine(s), toxin(s), oligonucleotide(s), colored or fluorescent molecule(s), amine, carboxyl or hydroxyl group(s), bioactive molecule(s) (such as an immunomodulatory molecule, a small chemical molecule, a peptido-mimetic, a drug), biotin, avidin or streptavidin molecule(s), or a combination thereof.
  • bioactive molecule(s) such as an immunomodulatory molecule, a small chemical molecule, a peptido-mimetic, a drug
  • biotin avidin or streptavidin molecule(s), or a combination thereof.
  • cytometric when qualifying data, parameters, measurements, events, may refer to cells, vesicles of cellular origin, or to particles such as acellular microorganisms or bio-functionalized materials.
  • a “statistical sample” refers to a subset of items from within a statistical population to estimate characteristics of the whole statistical population.
  • Figure 1 illustrates a device for determining proportions ⁇ i of C reference populations P i in an ensemble of biological objects.
  • Figure 2 is an example of a flow-chart representing a method for determining proportions ⁇ i of C reference populations Pi in an ensemble of biological objects.
  • Figure 3 shows a comparison of performances of the method illustrated in Figure 2 and a method for quantification of biological objects of the prior art.
  • Figure 4 is an example of a flow-chart representing a method for match labeling.
  • DETAILED DESCRIPTION [0062] This invention relates to a computer-implemented method 100 (see figure 2) for determining proportions ⁇ i of C reference populations Pi in an ensemble of biological objects being chosen among cells, vesicles of cellular origin, acellular microorganisms, and/or biofunctionalized materials.
  • the biological objects are: - cells of animal, plant, fungal, protist, bacterial or archaebacterial origin, - vesicles of cellular origin chosen from exosomes, ectosomes, microvesicles, microparticles, prostasomes, oncosomes, matrix/calcification vesicles or apoptotic bodies, - acellular microorganisms chosen from viruses, viroids and prions, and/or - biofunctionalized materials comprising a material of synthetic or biological origin chosen from a nanoparticle (such as a nanobead, a nanosphere or a nanocapsule), a microparticle (such as a microbead, a microsphere or a microcapsule), a lipid vesicle (such as a unilamellar vesicle, a multilamellar vesicle, a lipoplex, a polyplex, a lipopolyplex, a liposome
  • Cytometric data can be obtained by any process allowing to analyze, determine or measure characteristics of a cell, vesicle or particle.
  • the analysis can relate to a sample including a plurality of cells, vesicles and/or particles.
  • cytometric data can be obtained from a biological sample of an individual, such as one or more organ(s), tissue(s), cell(s) or cell fragment(s) of the individual.
  • the cytometric data acquisition process may be preceded by a sample preparation process, including for example a stage of isolation of cells, vesicles and/or particles, or of preliminary isolation of components of said cells.
  • the sample preparation process comprising the biological objects with which cytometric data is associated may include, as non-limiting examples, one or more step(s) of magnetic-activated cell sorting (MACS), laser capture microdissection (LCM), manual cell or micro-manipulation harvesting, micro-fluidic, limit dilution, separation by optical tools, electrophoresis, separation by beads, immunoprecipitation, immunopanning, labeling, immunostaining, immunofluorescence, multiplexing and/or use of biochips.
  • Cytometric data can be data of "omics” or "meta-omics” technologies, for example genomics, epigenomics, transcriptomics, proteomics, metabolomics, lipidomics, etc.
  • cytometric data can for example be relating to the detection and/or quantification of DNA (for example genes), RNA, proteins (for example cytokines, receptors, etc.), sugars, amino acids and/or fatty acids present in or on the surface of the cell, vesicle or particle, or any other molecule contained in the cell, vesicle or particle.
  • DNA for example genes
  • proteins for example cytokines, receptors, etc.
  • sugars amino acids and/or fatty acids present in or on the surface of the cell, vesicle or particle, or any other molecule contained in the cell, vesicle or particle.
  • Cytometric data can also be relating to the detection of the presence or absence of a particular molecule or a particular assembly in or on the surface of the cell, vesicle or particle, or to the detection of interactions between molecules present in or on the surface of the cell, vesicle or particle, or to the detection of interactions between cells and/or objects such as vesicles, microorganisms, or bio-functionalized materials.
  • the cytometric parameters measured or analyzed can be the level of expression of one or more intracellular or extracellular protein(s), or the level of expression of one or more receptor(s) or marker(s) on the surface of the cell, vesicle or particle.
  • cytometric parameters include the size of cells, vesicles or particles, their density, their granularity, their morphology/form, their refractive index, the composition of their membrane, their molecular content (such as for example the presence or the absence of an intracellular or surface molecule) or their content in this molecule (such as intracellular content, or ions, DNA, or RNA content %) or the state of oxydo-reduction of the cell, its status in the cell cycle, its apoptotic state, or its level of phosphorylation ...
  • the cytometric data associated to the ensemble of biological objects may be obtained by flow cytometry (or FACS for "fluorescence activated cell sorting"), by PCR-activated cell sorting (PACS), by microsphere affinity proteomics (MAP), by mass spectrometry, by chromatography, by CYTOF, by spectral cytometry, by mass cytometry, by image cytometry, by gene expression on chips (microarray), by sequencing (for example DNA-seq, RNA-seq, scDNA-seq, or scRNA-seq), by in situ hybridization, and/or by microscopy.
  • flow cytometry or FACS for "fluorescence activated cell sorting"
  • PES PCR-activated cell sorting
  • MAP microsphere affinity proteomics
  • mass spectrometry by mass spectrometry
  • mass chromatography by CYTOF
  • spectral cytometry by mass cytometry
  • image cytometry by gene expression on chips (microarray)
  • sequencing
  • microscopy notably designates atomic force microscopy (AFM), electrochemical detection (EC), scanning electron microscopy (SEM), transmission electron microscopy (TEM), surface plasmon resonance (SPR), Raman micro-spectroscopy, etc.
  • the cytometric data can also be obtained by combining several of these techniques.
  • multiplexed profiling of RNA and protein expression at the single cell level can be obtained simultaneously by combining the PLAYR (Proximity Ligation Assay for RNA) technique, which makes it possible to measure the expression level of a large number of RNAs by flow or mass cytometry, with a technique for detecting surface or internal proteins labeled with antibodies.
  • PLAYR Proximity Ligation Assay for RNA
  • the cytometric data are obtained by single-cell analysis technologies.
  • the cytometric data are obtained by spatial biology analysis technologies, such as e.g. spatial genomics, transcriptomics or proteomics analysis technologies.
  • the cytometric data may be obtained by single cell imaging-based technologies, such as e.g. flow, spectral or mass cytometry, or microscopy; or by sequencing-based technologies, such as e.g.
  • RNAseq single cell sequencing platform including for instance single-cell RNAseq, spatial transcriptomics, single-cell CITE-seq...
  • Another purely illustrative example of a technique for obtaining cytometric data is the profiling of single cells isolated in liquid droplets, enveloped by a thin semi- permeable membrane (microcapsules), by high-throughput RNA cytometry, for example by multiplexed RT-PCR.
  • the analysis of the ensemble of biological objects outputted a set of test data 10.
  • the set of test data 10 comprises statistical data about individual cell populations.
  • the set of test data 10 will be sometimes referred to as “test statistical sample”, where the expression “statistical sample” has been defined in the Definitions section.
  • the set of test data 10 is stored in a matrix Q.
  • the matrix Q comprises rows, each row corresponding to one biological object, and columns, each column corresponding to a characteristic of the corresponding biological object.
  • the method 100 may be implemented with a device 200 such as illustrated on Figure 1.
  • Figure 1 illustrates a device 200 for determining proportions ⁇ i of C reference populations P i in an ensemble of biological objects according to embodiments of the invention.
  • the device 200 comprises a computer, this computer comprising a memory 201 to store program instructions loadable into a circuit and adapted to cause a circuit 202 to carry out steps of the methods of Figures 2 and 3 when the program instructions are run by the circuit 202.
  • the memory 201 may also store data and useful information for carrying steps of the present invention as described above.
  • the circuit 202 may be for instance: - a processor or a processing unit adapted to interpret instructions in a computer language, the processor or the processing unit may comprise, may be associated with or be attached to a memory comprising the instructions, or - the association of a processor / processing unit and a memory, the processor or the processing unit adapted to interpret instructions in a computer language, the memory comprising said instructions, or - an electronic card wherein the steps of the invention are described within silicon, or - a programmable electronic chip such as a FPGA chip (for « Field-Programmable Gate Array »).
  • the computer may also comprise an input interface 203 for the reception of input data and an output interface 204 to provide output data. Examples of input data and output data will be provided further.
  • a screen 205 and a keyboard 206 may be provided and connected to the computer circuit 202.
  • a database comprising at least one set of reference data is preliminarily collected.
  • the at least one set of reference data comprises N sets of reference data, with N being an integer number higher than or equal to 1.
  • Each set of reference data comprises cytometric data corresponding to a given reference ensemble of biological objects.
  • the cytometric data associated to a plurality of given reference biological objects have been obtained by flow cytometry (or FACS for "fluorescence activated cell sorting"), by PCR-activated cell sorting (PACS), by microsphere affinity proteomics (MAP), by mass spectrometry, by chromatography, by CYTOF, by spectral cytometry, by mass cytometry, by image cytometry, by gene expression on chips (microarray), by sequencing (for example DNA-seq, RNA-seq, scDNA-seq, or scRNA-seq), by in situ hybridization, and/or by microscopy.
  • flow cytometry or FACS for "fluorescence activated cell sorting"
  • PES PCR-activated cell sorting
  • MAP microsphere affinity proteomics
  • mass spectrometry by mass spectrometry
  • mass chromatography by CYTOF
  • spectral cytometry by mass cytometry
  • image cytometry by gene expression on chips (microarray
  • Each given reference ensemble of biological objects originates from one among the given reference ensemble of biological objects, each of which containing different reference populations D i of biological objects.
  • Each reference population D i is assigned a label Li belonging to a plurality of C labels Lj, C being an integer number higher than or equal to 2 and j varying from 1 to C.
  • the cytometric data of each set of reference data will often be referred to as “reference statistical sample”, where the expression “statistical sample” has been defined in the Definitions section.
  • each reference data in a set of reference data is labelled with a label Li belonging to a plurality of C labels Lj, j being an integer varying between 1 and C.
  • the data contained in the database take the form of an electronic file comprising C matrices ⁇ ⁇ + .
  • Each matrix ⁇ ⁇ + contains a concatenation of the reference data labelled with the label Lj in the N sets of reference data. The higher the number N of sets of reference data, the more precise the result of the quantification obtained with the method 100 will be.
  • each matrix ⁇ ⁇ + comprises n j rows, each row corresponding to data pertaining to a biological object. All biological objects represented in the matrix ⁇ ⁇ + present the label L j .
  • the electronic file can be received by the device 200 via the input interface 203 and subsequently stored in the memory 201 of the device 200.
  • the computer-implemented method 100 aims at determining the proportions ⁇ j of each reference population Dj among the C reference populations Dj, j being an integer varying between 1 and C.
  • the proportions ⁇ j can be gathered in a single vector ⁇ .
  • Figure 2 is a flowchart illustrating an example of steps to be executed to implement the computer-implemented method 100.
  • step S1 the at least one set of reference data is received.
  • the at least one set of reference data can take the form of an electronic file comprising C matrices ⁇ ⁇ + as previously described.
  • transformed vectors M j are computed.
  • the transformed vectors result from the application of a predetermined function z on the matrices ⁇ ⁇ + .
  • the predetermined function z will be described further.
  • a transformed vector R is computed.
  • the test transformed vector R results from the application of the predetermined function z on a matrix Q comprising the set of test data 10 as previously described.
  • the test transformed vector R is defined by: .
  • Step S4 the proportions ⁇ j, for j varying between 1 and C, by minimizing a distance between the weighted sum of the transformed # ⁇ and the test transformed vector R.
  • the proportions ⁇ js are the argument of the minima (i.e., arg min) of the minimization of said distance between the weighted sum of the transformed vectors ⁇ ⁇ # ⁇ ⁇ # ⁇ ⁇ ⁇ ⁇ and the test transformed vector R.
  • said minimization problem is solved with respect to the proportions ⁇ j , for j varying between 1 and C.
  • Steps S1 to S4 may be implemented by the processor of the circuit 202.
  • ⁇ ′ ⁇ denotes the transposed matrix of the matrix P’. The minimization is carried out under the conditions that ⁇ # ⁇ ⁇ # is equal to 1 and all the proportions ⁇ i are positive.
  • the predetermined function z is a kernel function k, such as functions defined with respect to reproducing kernel Hilbert spaces (RKHS) and used in techniques known as kernel methods.
  • the concept of kernel mean embedding is used.
  • the kernel function k is Gaussian function.
  • the notation of matrix P 0 is the concatenation of the matrices ⁇ ⁇ + , in other words, 3.
  • Each matrix ⁇ ⁇ + for j taking values between 1 and C, comprises nj rows.
  • the matrix P 0 comprises n rows, where n is equal to the sum of the numbers n j for j comprised between 1 and C.
  • the rows of the matrix P0 are denoted xil, where i denotes the i th row of the matrix ⁇ ⁇ 4 included in the matrix P0. It can be shown that the computation of the quantity P in the definition of the quantity D involves the computation of nxn values k(x il ,x jt ). Similarly, the matrix Q comprises m rows y i . In this case, it can be shown that the computation of the quantity q used in the definition of the quantity D involves a computation of n x m values k(x il ,y j ).
  • the translation invariant kernel function k is a Gaussian function.
  • step S4 is simpler as it can be shown that no further calculation, except for the minimization of the quantity D, is necessary.
  • the computation of the transformed vectors M j ( ⁇ ⁇ ⁇ ⁇ ), and of the test transformed vector R ( ⁇ is sufficient.
  • the matrix containing the concatenated N sets of reference data is multiplied by the matrix containing the values of the statistical sample ( ⁇ i) i ⁇ 1 .. M/2 of the distribution ⁇ k , Then, the cosine and the sine functions are applied to the results.
  • a solver program can be used.
  • FIG. 3 shows a comparison of computation times (for estimation of the parameters) between the method 100 (referred to on Figure 3 as KQuant) and a method for quantification of biological objects of the prior art referred to on Figure 3 as GMM and described in the paper “A Machine Learning Approach to the Classification of Acute Leukemias and Distinction From Nonneoplastic Cytopenias Using Flow Cytometric Data”, by Monaghan et al., in American Society for Clinical Pathology, 2021, which is incorporated herein by reference.
  • the method GMM (which stands for Gaussian Mixture Model) is a machine learning based method for classifying multidimensional data obtained by flow cytometry.
  • a GMM makes it possible to define a descriptor of a distribution (for example the distribution representing a cell population in a flow cytometry experiment) and an estimate of the proportion.
  • GMMs as a descriptor of a cell population in a flow cytometry can be realized according to any known prior art, for example the above-cited article of Monaghan et al., in American Society for Clinical Pathology, 2021, or according to the article of Ko et al., “Clinically validated machine learning algorithm for detecting residual diseases with multicolor flow cytometry analysis in acute myeloid leukemia and myelodysplastic syndrome”, EBioMedicine 37, 2018, which is incorporated herein by reference. In the results presented in Figure 3, both methods 100 and GMM were applied on a series of test data of increasing size (from 25000 points, i.e. biological objects in the test data, to 1,000,000 points).
  • the computation time is shorter with the method 100 than with the method GMM from the prior art for all sizes.
  • the method 100 is more efficient as a descriptor as it is agnostic to the type of overall distribution to be described.
  • the method 100 is more efficient in computation time.
  • a specific library can be used.
  • the KeOps library that allows computing reductions of large arrays whose entries are given by a mathematical formula or a neural network, can be used.
  • This library can be used with Python (NumPy, PyTorch), Matlab and R.
  • the method 100 further comprises additional steps S5, S6, S7 and S8, that will be described below, and allowing to perform match labeling.
  • additional steps S5, S6, S7 and S8, that will be described below, and allowing to perform match labeling.
  • S5, S6, S7 and S8 that will be described below.
  • the predetermined function z is the one as used in the vectorization embodiments previously described.
  • Figure 4 is a flowchart illustrating an example of further steps of the method 100 as illustrated in Figure 2 to be executed to implement a computer-implemented method for match labeling.
  • a set of clusters G k associated with the set of test data 10 is received.
  • the set of clusters Gk may have been preliminary obtained by the application of a segmentation method on the set of test data 10.
  • the set of clusters Gk are received via the input interface 203 of the device 200.
  • each cluster G k in the set of clusters G k is associated with a population.
  • each cluster G k is received in the form of a matrix Xk.
  • a distance dki is computed between z(Xk) for each cluster Gk and for each matrix ⁇ ⁇ ⁇ as defined above, for i taking values between 1 and C.
  • a step S7 for each cluster Gk, the reference population Ppk for which the distance d_kp k is minimal is determined, p k being an integer comprised between 1 and C.
  • p k being an integer comprised between 1 and C.
  • a minimum of computed d ki in step S6 for i taking values between 1 and C is determined as d_kpk and the respective population Ppk to this minimum distance is determined as the reference population.
  • each label Lp k corresponding to the reference population Pp k determined at step S7 is assigned to the corresponding cluster Gk.
  • Steps S6, S7 and S8 may be implemented by the processor of the circuit 201 of the device 201.

Landscapes

  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Medical Informatics (AREA)
  • Epidemiology (AREA)
  • Software Systems (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Bioethics (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Public Health (AREA)
  • Biophysics (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Biotechnology (AREA)
  • Evolutionary Biology (AREA)
  • General Health & Medical Sciences (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Theoretical Computer Science (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

The present invention relates to a computer-implemented method (100) for determining proportions αi of C reference populations (Pi) in an ensemble of biological objects being chosen among cells, vesicles of cellular origin, acellular microorganisms, and/or biofunctionalized materials.

Description

METHOD FOR DETERMINING PROPORTIONS OF POPULATIONS IN AN ENSEMBLE OF BIOLOGICAL OBJECTS FIELD OF INVENTION [0001] The present invention relates to the field of cytometric data analysis, including single cell data analysis. BACKGROUND OF INVENTION [0002] The emergence of protein sequencing, and then DNA sequencing, and the development of automatic sequencers have revolutionized biology. The classical descriptive and reductionist approach (one gene, one messenger RNA, one protein) has been replaced by a more global understanding of biological systems based on the analysis of sets of biological elements ("-omes"). The basic idea associated with "omics" approaches is to apprehend the complexity of living organisms as a whole, using methodologies that are the least restrictive possible in terms of description. [0003] Such approaches mainly include: genomics (study of genes), transcriptomics (analysis of gene expression and its regulation), proteomics (study of proteins), metabolomics (analysis of metabolites) and metabonomics (study of metabolic profiles in vivo). [0004] Genomics is divided into two branches: structural genomics, which relates to the sequencing of the entire genome, and functional genomics, which aims to determine the function and expression of the sequenced genes. In functional genomics, techniques are applied to a large number of genes in parallel: for example, the phenotype of mutants for a whole family of genes, or the expression of all the genes of an entire organism, can thus be analyzed. [0005] Transcriptomics is the study of all the messenger RNAs produced during the genetic transcription process. It is based on the quantification of all of these messenger RNAs, which provides a relative indication of the transcription rate of different genes under given conditions. [0006] Proteomics is the analysis of all the proteins of an organelle, a cell, a tissue, an organ or an organism under given conditions. Proteomics attempts to identify in a global way the proteins extracted from a cell culture, a tissue or a biological fluid, their localization in the cellular compartments, their possible post-translational modifications, as well as their quantity. It makes it possible to quantify the variations in their level of expression, for example as a function of time, of their environment, of their state of development, of their physiological and pathological state, of the species of origin, etc. It also studies the interactions that proteins have with other proteins, with DNA or RNA, or with other substances. [0007] Metabolomics studies all metabolites (sugars, amino acids, fatty acids, etc.), such as metabolic substrates, intermediates, products, as well as hormones, other signaling molecules and secondary metabolites, present in a cell, an organ, an organism. [0008] Metabonomics consists of monitoring metabolic profiles in vivo, which makes it possible to provide information on the toxicity of drugs, on pathological processes, and on the function of genes. [0009] Cytomics corresponds to the analysis of the cytome, the set of cellular constituents, in particular morphological, antigenic and functional, which make it possible to define or model the state and the functioning of a cell at a given time. The analysis of the cytome makes it possible to describe the structural and functional heterogeneity of the various cells of an organism. [0010] The preceding approaches make it possible to obtain a great amount of information on the cellular and/or tissue response to an in vitro or in vivo exposure. They can in particular be useful for highlighting and identifying new biomarkers (of diagnosis, susceptibility, prognosis, exposure, effect), generating new knowledge on the mechanistic level (modes of action), or further develop new efficacy or predictive toxicology tools to help identify new therapeutic targets or new candidate drugs or candidate vaccines. [0011] The automation of sequencing techniques and the development of high- throughput techniques, in particular made possible by the emergence of specialized technological platforms, have allowed the industrialization of data production and the simultaneous analysis of a large number of parameters. [0012] The result is a very large amount of data to be processed, analyzed, visualized and interpreted in the most informative way possible in order to extract the maximum amount of information on the biological process or on the biological system studied. [0013] From the biostatistic point of view, the data obtained by "omics" approaches relate to very numerous variables that should be analyzed jointly. For example, transcriptomic analyzes make it possible to simultaneously study the expression of several thousand genes. [0014] It is therefore desirable to have powerful biostatistical and bioinformatics means for processing, analyzing and interpreting the mass of data generated by "omics" approaches. [0015] There are many techniques for acquiring "omics" data, such as, as non-limiting examples, mass spectrometry, chromatography, sequencing, spectral cytometry, mass cytometry, flow cytometry, etc. [0016] Flow cytometry is a technique for analyzing a biological fluid comprising a collection of biological objects, such as cells, vesicles or particles, in suspension. Flow cytometry is an essential analytical technique for the identification and characterization of populations of cells, vesicles or particles. [0017] Besides, individual cell heterogeneity within a cell population can be critical to its specific function and fate. Cell-to-cell variations, for instance in RNA transcripts or protein expression, may be key in research and development studies in the field of cancer, neurobiology, stem cell biology, immunology and developmental biology. Conventional cell-based assays mainly analyze the average responses from a population of cells, without providing individual cell phenotypes. To better understand cell-to-cell differences, detailed information may be provided by single cell analyses. [0018] Single-cell analysis is the study of genomics, transcriptomics, proteomics, metabolomics and cell–cell interactions at the single cell level. Analyzing a single cell makes it possible to discover mechanisms not seen when studying a cell population. High- throughput and multiparameter approaches may be combined in single cell analysis in order to reflect cell-to-cell variability and heterogeneous differences in the individual cells. [0019] Technologies such as fluorescence-activated cell sorting (FACS) allow the precise isolation of selected single cells from complex samples, while high throughput single cell partitioning technologies, enable the simultaneous molecular analysis of hundreds or thousands of single unsorted cells. This is particularly useful for the analysis of transcriptome variation in genotypically identical cells, allowing the definition of otherwise undetectable cell subtypes. Thus, flow cytometry can address the issue of cellular processes and intercellular interactions because cytometry provides a single-cell view of more complex cellular conglomerates, such as tissues in multicellular organisms or communities in the case of single-cell bacteria, yeasts, algae, etc. [0020] The development of new technologies nowadays enables to analyze the genome and transcriptome of single cells, as well as to quantify their proteome and metabolome. Single‐cell RNA sequencing (scRNA‐seq) technology has become a state‐of‐the‐art approach for unravelling the heterogeneity and complexity of RNA transcripts within individual cells, as well as revealing the composition of different cell types and functions within tissues, organs or organisms. In situ sequencing and fluorescence in situ hybridization (FISH) do not require that cells be isolated and may be used for analysis of tissues. Other technologies enable quantifying thousands of proteins across hundreds of single cells. For example, mass spectrometry techniques have become important analytical tools for proteomic and metabolomic analysis of single cells. [0021] Individual cells differentially use their genes, RNA, and proteins across tissue types. Spatial biology designates a series of techniques providing detailed cellular information according to the positional context of cells in a tissue. Thus, spatial biology studies the cell types, their location in tissues, their biomarker co-expression patterns, and how cells organize, interact and influence the tissue microenvironment. [0022] Spatial transcriptomics is a technology that allows measuring all the gene activity in a tissue sample and spatially map where the activity is occurring. It combines unbiased, high-throughput total mRNA analysis for intact tissue sections with morphological context. Spatial proteomics combines microscopy with detection of multiple proteins, usually through the use of specific antibodies. It allows the integration of morphological data and protein expression data at the single-cell level. Detection methods may be based for example on mass spectrometry, on chromogenic staining, or on oligonucleotide barcodes. [0023] Among the various parameters that may be analyzed in a set of data obtained by flow cytometry or other assays, and corresponding to a particular ensemble of biological objects, the quantification of the different populations of biological objects contained in this ensemble is of particular interest. The quantification of the different populations of biological objects contained in an ensemble of analyzed biological objects is generally one of the first assessed parameters. Indeed, the quantification of the different populations of biological objects may for instance provide information on the validity of the data obtained during a specific assay, or on the reproducibility of this kind of assay. [0024] It is therefore desirable to have powerful biostatistical and bioinformatics means and methods for rapidly determining quantification of different populations of biological objects for which data have been obtained by an assay. These methods must be easy to implement and robust, and provide reliable and reproducible results. [0025] The invention proposes a solution to this need. SUMMARY [0026] This invention thus relates to computer-implemented method for determining proportions αi of C reference populations Pi in an ensemble of biological objects being chosen among cells, vesicles of cellular origin, acellular microorganisms, and/or biofunctionalized materials, C being an integer number, said ensemble of biological objects being associated with a set of test data, the method comprising: - receiving at least one set of reference data associated with a reference ensemble of biological objects, wherein each reference data is labelled with a label Li belonging to a plurality of C labels Lj, each label Lj among the plurality of C labels Lj being associated with one corresponding reference population Pj among the C reference populations, - determining said proportions αi by minimizing a distance between a first vector and a second vector, said first vector being a weighted sum of C transformed vectors ^^^^ ^^ for i an integer varying from 1 to C, where ^^ ^ represents a matrix with ni rows comprising a sample composed of the reference data labelled with the label Li, z denotes a predetermined function, and the weight of each transformed vector ^^^^ ^^ is the proportion αi, said second vector resulting from the application of the predetermined function z to a second matrix Q comprising a sample composed of the test data. [0027] By “a matrix comprising a sample composed of the reference/test data associated with a reference ensemble of biological objects” it is meant that said matrix comprises respective data/characteristics of said (biological) sample. In examples, said matrix may comprise rows and columns when each row corresponds to one biological object of said ensemble and each column corresponds to a characteristic of corresponding biological objects. [0028] The proposed advantageously method allows for quick quantification of populations in an ensemble of biological objects starting from labelled data, with input cytometric data and by only performing a minimization. [0029] Advantageously, minimizing the distance between the first vector and the second vector comprises minimizing a quantity D defined by: ^ = ^ ^^ ^ ^ ^^ + ^ ^, wherein α represents a vector comprising the proportions ^ ^
Figure imgf000009_0001
^ ^ = ^′ ^′, where P’ is a matrix composed of the C vectors ^ = −2 ^′^^^^^. [0030] Such a minimization problem is a quadratic programming (QP) in dimension c and can be solved efficiency by any known method in the literature. [0031] In some embodiments, the predetermined function z is based on a kernel function k. [0032] In some embodiments, computing P involves a computation of nxn values k(xil,xjt), wherein xil denotes an ith vector with the label Ll in a matrix P0 and xjt denotes a jth vector with the label Lt of the matrix P0, the matrix P0 being a concatenation of the ^ ^
Figure imgf000009_0002
matrices ^^ , n being equal to ∑^^^ ^^ , and, Q comprising m rows yj, computing q involves a computation of n x m values k(xi,yj). [0033] Advantageously, the kernel function k is a Gaussian function. [0034] In some other embodiments, the predetermined function is a vectorial function taking values in ℝM and defined, for a given vector x in ℝs by: ^^^^ = ^ ^ ^ ^cos^"# ^^^ , sin^"# ^ ^ ^^' ( #^^ ^ , X comprising r rows denoted x1, x2,…xr, by:
Figure imgf000009_0003
* , where (ωi)i ^ 1 .. M/2 is a statistical sample of a distribution Λk, Λk being the Fourier transform of a translation invariant kernel function k. In those embodiments, the number of calculations performed in the method is advantageously reduced. Therefore, as a consequence, the computation time is reduced in comparison to prior art methods, as will be illustrated further. Furthermore, in comparison to machine learning models, no hyperparameter is when two sets of test data have to be compared, once the
Figure imgf000009_0004
C transformed and the second vector have been computed, the computation of an Euclidean distance is sufficient to perform the comparison, as compared to prior art methods as will be illustrated further. The comparison is therefore simpler, and thus faster. [0035] Advantageously, the kernel function k is a Gaussian function. Having a kernel function k which is a Gaussian allows for easy computation of the distribution Λk. [0036] In some embodiments, the method further comprises: - receiving a set of K clusters Gk associated with the set of test data, k being integer varying from 1 to K, and Gk denoting the k-th cluster, wherein each cluster Gk in the set of K clusters Gk is associated with a population and is represented with a matrix Xk, - computing, for each cluster Gk and for each vector ^^^ , a distance d_ki between z(Xk) and ^^^^^^, - determining, for each cluster Gk, the reference population Ppk for which the distance dkp is minimal, wherein the reference population Ppk for is associated with the label Lpk, - assigning the label Lpk to the k-th cluster Gk. Therefore, the method advantageously allows for match labeling. In other words, the method allows, for instance, using the predetermined function defined as a vectorial function previously described, to label unlabelled clusters obtained from a segmentation method of cytometric data of an ensemble of biological objects. [0037] In some embodiments, the biological objects are: - cells of animal, plant, fungal, protist, bacterial or archaebacterial origin, - vesicles of cellular origin chosen from exosomes, ectosomes, microvesicles, microparticles, prostasomes, oncosomes, matrix/calcification vesicles or apoptotic bodies, - acellular microorganisms chosen from viruses, viroids and prions, and/or - biofunctionalized materials comprising a material of synthetic or biological origin chosen from a nanoparticle (such as a nanobead, a nanosphere or a nanocapsule), a microparticle (such as a microbead, a microsphere or a microcapsule), a lipid vesicle (such as a unilamellar vesicle, a multilamellar vesicle, a lipoplex, a polyplex, a lipopolyplex, a liposome, a niosome, a cochleate, a virosome, an immunostimulating complex (ISCOM®)), said material of synthetic or biological origin being coupled to, or coated with, one or more peptide(s), protein(s), antibody(ies), antibody fragment(s), receptor(s), cytokine(s), chemokine(s), toxin(s), oligonucleotide(s), colored or fluorescent molecule(s), amine, carboxyl or hydroxyl group(s), bioactive molecule(s) (such as an immunomodulatory molecule, a small chemical molecule, a peptido-mimetic, a drug), molecule(s) of biotin, avidin or streptavidin, or a combination thereof. [0038] Advantageously, the reference and test data are obtained by flow cytometry (or FACS for "fluorescence activated cell sorting"), by PCR-activated cell sorting (PACS), by microsphere affinity proteomics (MAP), by mass spectrometry, by chromatography, by CYTOF, by spectral cytometry, by mass cytometry, by image cytometry, by gene expression on chips (microarray), by sequencing (for example DNA-seq, RNA-seq, scDNA-seq, or scRNA-seq), by in situ hybridization, and/or by microscopy. [0039] Advantageously, the reference and test data are obtained by single-cell analysis technologies or spatial biology analysis technologies. [0040] Another aspect of the invention pertains to a device comprising a processor configured to carry out the method previously described. [0041] Another aspect of the invention pertains to a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method previously described. [0042] Another aspect of the invention relates to a non-transitory computer-readable recording medium comprising instructions which, when executed by a computer, cause the computer to carry out steps of the method previously described. DEFINITIONS [0043] In the present invention, the following terms have the following meanings: [0044] The term "processor" or “processing unit” should not be construed as being limited to hardware capable of executing software, and refers generally to a processing device, which may include, for example, a computer, microprocessor, integrated circuit, or programmable logic device (PLD). The processor may also include one or more graphics processing units (GPUs), whether they are used for computer graphics and image processing or other functions. In addition, the instructions and/or data for executing the associated and/or resulting functionality may be stored on any media readable by the processor such as, for example, an integrated circuit, a hard disk, a CD (Compact Disc), an optical disk such as a DVD (Digital Versatile Disc), a RAM (Random-Access Memory) or a ROM (Read-Only Memory). In particular, instructions may be stored in hardware, software, firmware or any combination thereof. [0045] "Cytometry" generally refers to the detection and/or measurement of characteristics of a cell, which is the basic structural, functional and biological unit of living organisms. By extension, cytometry also refers to the detection and/or measurement of characteristics of vesicles of cellular origin (such as extracellular or intracellular vesicles), or of particles of nano- or micrometer size such as acellular microorganisms (e.g., viruses), or bio-functionalized materials (e.g., microspheres coated with biological molecules such as proteins, or microspheres coated with bio-active molecules such as immunomodulatory molecules or drugs). [0046] The "cells" may come from a unicellular or pluricellular organism, of prokaryotic or eukaryotic origin, of animal, vegetable, fungal, protist, bacterial or archaebacterial origin. They can be living, dead or fixed. The cells may, for example, come from solid or liquid tissues (such as bone marrow, blood or lymph, etc), or from body fluids (such as cerebrospinal fluid, urine, bronchoalveolar fluid, etc). [0047] During the cytometric data acquisition process, the cells, vesicles or particles may be in suspension (aqueous or non-aqueous, biological or synthetic suspension), or immobilized on a solid support (e.g., a flask, a cell culture dish with one or more wells, a plate or microplate, a glass slide, a chip, etc.). They may have been pre-treated with biological or chemical agent(s). [0048] The cells, vesicles or particles to be analyzed may be comprised within one or more populations. [0049] A "population" is an ensemble of cells, vesicles, or particles with same or similar characteristics. A population can be subdivided into different subpopulations with secondary characteristics different from each other. [0050] "Cell-derived vesicles" are vesicles composed of a membrane of cellular origin. It can be intracellular or extracellular vesicles. Extracellular vesicles have generally been secreted or excreted by a cell. "Cell-derived vesicles" include exosomes, ectosomes, microvesicles, microparticles, prostasomes, oncosomes, matrix/calcification vesicles, apoptotic bodies and other subsets of cellular vesicles. [0051] "Acellular microorganisms" designate organisms of microscopic size whose structure is not cellular. "Acellular microorganisms" include viruses, viroids and prions. [0052] The "bio-functionalized materials" or "functionalized biomaterials" or "bio- functionalized devices" or "bio-functionalized systems" designate any type of object comprising a material of synthetic or organic origin coated on its surface, coupled or linked to one or more molecule(s) or functional group(s) of biological origin. [0053] As non-limitative examples, the material of synthetic or biological origin can be a nanoparticle (such as a nanobead, a nanosphere or a nanocapsule), a microparticle (such as a microbead, a microsphere or a microcapsule), a lipid vesicle such as a unilamellar vesicle, a multilamellar vesicle, a lipoplex, a polyplex, a lipopolyplex a liposome, a niosome, a cochleate, a virosome, an immune-stimulating complex (ISCOM®), etc. [0054] Nanoparticles and microparticles, such as beads, spheres or capsules, can be organic, inorganic, magnetic or radioactive. For example, nanoparticles or microparticles can be made of metal, silica, aluminium, titanium, glass, ceramic, polystyrene, poly(methyl methacrylate), melamine, polylactide, latex, dextran, oxide, graphene, magnetic material, radioactive material, or a combination thereof. [0055] As non-limited examples, to be bio-functionalized, the material can be combined with, or coated with, one or more peptide(s), protein(s), antibody(ies), fragment(s) of antibody, receptor(s), cytokine(s), chemokine(s), toxin(s), oligonucleotide(s), colored or fluorescent molecule(s), amine, carboxyl or hydroxyl group(s), bioactive molecule(s) (such as an immunomodulatory molecule, a small chemical molecule, a peptido-mimetic, a drug), biotin, avidin or streptavidin molecule(s), or a combination thereof. [0056] In the context of the present invention, the term "cytometric", when qualifying data, parameters, measurements, events, may refer to cells, vesicles of cellular origin, or to particles such as acellular microorganisms or bio-functionalized materials. [0057] A “statistical sample” refers to a subset of items from within a statistical population to estimate characteristics of the whole statistical population. BRIEF DESCRIPTION OF THE DRAWINGS [0058] Figure 1 illustrates a device for determining proportions αi of C reference populations Pi in an ensemble of biological objects. [0059] Figure 2 is an example of a flow-chart representing a method for determining proportions αi of C reference populations Pi in an ensemble of biological objects. [0060] Figure 3 shows a comparison of performances of the method illustrated in Figure 2 and a method for quantification of biological objects of the prior art. [0061] Figure 4 is an example of a flow-chart representing a method for match labeling. DETAILED DESCRIPTION [0062] This invention relates to a computer-implemented method 100 (see figure 2) for determining proportions αi of C reference populations Pi in an ensemble of biological objects being chosen among cells, vesicles of cellular origin, acellular microorganisms, and/or biofunctionalized materials. [0063] For instance, the biological objects are: - cells of animal, plant, fungal, protist, bacterial or archaebacterial origin, - vesicles of cellular origin chosen from exosomes, ectosomes, microvesicles, microparticles, prostasomes, oncosomes, matrix/calcification vesicles or apoptotic bodies, - acellular microorganisms chosen from viruses, viroids and prions, and/or - biofunctionalized materials comprising a material of synthetic or biological origin chosen from a nanoparticle (such as a nanobead, a nanosphere or a nanocapsule), a microparticle (such as a microbead, a microsphere or a microcapsule), a lipid vesicle (such as a unilamellar vesicle, a multilamellar vesicle, a lipoplex, a polyplex, a lipopolyplex, a liposome, a niosome, a cochleate, a virosome, an immunostimulating complex (ISCOM®)), said material of synthetic or biological origin being coupled to, or coated with, one or more peptide(s), protein(s), antibody(ies), antibody fragment(s), receptor(s), cytokine(s), chemokine(s), toxin(s), oligonucleotide(s), colored or fluorescent molecule(s), amine, carboxyl or hydroxyl group(s), bioactive molecule(s) (such as an immunomodulatory molecule, a small chemical molecule, a peptido-mimetic, a drug), molecule(s) of biotin, avidin or streptavidin, or a combination thereof. [0064] Cytometric data can be obtained by any process allowing to analyze, determine or measure characteristics of a cell, vesicle or particle. The analysis can relate to a sample including a plurality of cells, vesicles and/or particles. For example, cytometric data can be obtained from a biological sample of an individual, such as one or more organ(s), tissue(s), cell(s) or cell fragment(s) of the individual. [0065] In some cases, the cytometric data acquisition process may be preceded by a sample preparation process, including for example a stage of isolation of cells, vesicles and/or particles, or of preliminary isolation of components of said cells. [0066] The sample preparation process comprising the biological objects with which cytometric data is associated may include, as non-limiting examples, one or more step(s) of magnetic-activated cell sorting (MACS), laser capture microdissection (LCM), manual cell or micro-manipulation harvesting, micro-fluidic, limit dilution, separation by optical tools, electrophoresis, separation by beads, immunoprecipitation, immunopanning, labeling, immunostaining, immunofluorescence, multiplexing and/or use of biochips. [0067] Cytometric data can be data of "omics" or "meta-omics" technologies, for example genomics, epigenomics, transcriptomics, proteomics, metabolomics, lipidomics, etc. Thus, cytometric data can for example be relating to the detection and/or quantification of DNA (for example genes), RNA, proteins (for example cytokines, receptors, etc.), sugars, amino acids and/or fatty acids present in or on the surface of the cell, vesicle or particle, or any other molecule contained in the cell, vesicle or particle. Cytometric data can also be relating to the detection of the presence or absence of a particular molecule or a particular assembly in or on the surface of the cell, vesicle or particle, or to the detection of interactions between molecules present in or on the surface of the cell, vesicle or particle, or to the detection of interactions between cells and/or objects such as vesicles, microorganisms, or bio-functionalized materials. [0068] By way of non-limitative and purely illustrative examples, when cytometric data are of the proteomic type, the cytometric parameters measured or analyzed can be the level of expression of one or more intracellular or extracellular protein(s), or the level of expression of one or more receptor(s) or marker(s) on the surface of the cell, vesicle or particle. [0069] Other examples of cytometric parameters include the size of cells, vesicles or particles, their density, their granularity, their morphology/form, their refractive index, the composition of their membrane, their molecular content (such as for example the presence or the absence of an intracellular or surface molecule) or their content in this molecule (such as intracellular content, or ions, DNA, or RNA content ...) or the state of oxydo-reduction of the cell, its status in the cell cycle, its apoptotic state, or its level of phosphorylation ... [0070] For instance, the cytometric data associated to the ensemble of biological objects may be obtained by flow cytometry (or FACS for "fluorescence activated cell sorting"), by PCR-activated cell sorting (PACS), by microsphere affinity proteomics (MAP), by mass spectrometry, by chromatography, by CYTOF, by spectral cytometry, by mass cytometry, by image cytometry, by gene expression on chips (microarray), by sequencing (for example DNA-seq, RNA-seq, scDNA-seq, or scRNA-seq), by in situ hybridization, and/or by microscopy. The term "microscopy" notably designates atomic force microscopy (AFM), electrochemical detection (EC), scanning electron microscopy (SEM), transmission electron microscopy (TEM), surface plasmon resonance (SPR), Raman micro-spectroscopy, etc. [0071] The cytometric data can also be obtained by combining several of these techniques. By way of purely illustrative example, multiplexed profiling of RNA and protein expression at the single cell level can be obtained simultaneously by combining the PLAYR (Proximity Ligation Assay for RNA) technique, which makes it possible to measure the expression level of a large number of RNAs by flow or mass cytometry, with a technique for detecting surface or internal proteins labeled with antibodies. [0072] In some embodiments, the cytometric data are obtained by single-cell analysis technologies. In some embodiments, the cytometric data are obtained by spatial biology analysis technologies, such as e.g. spatial genomics, transcriptomics or proteomics analysis technologies. [0073] For example, the cytometric data may be obtained by single cell imaging-based technologies, such as e.g. flow, spectral or mass cytometry, or microscopy; or by sequencing-based technologies, such as e.g. single cell sequencing platform including for instance single-cell RNAseq, spatial transcriptomics, single-cell CITE-seq… [0074] Another purely illustrative example of a technique for obtaining cytometric data is the profiling of single cells isolated in liquid droplets, enveloped by a thin semi- permeable membrane (microcapsules), by high-throughput RNA cytometry, for example by multiplexed RT-PCR. [0075] The analysis of the ensemble of biological objects outputted a set of test data 10. The set of test data 10 comprises statistical data about individual cell populations. The set of test data 10 will be sometimes referred to as “test statistical sample”, where the expression “statistical sample” has been defined in the Definitions section. For instance, the set of test data 10 is stored in a matrix Q. The matrix Q comprises rows, each row corresponding to one biological object, and columns, each column corresponding to a characteristic of the corresponding biological object. [0076] The method 100 may be implemented with a device 200 such as illustrated on Figure 1. [0077] Figure 1 illustrates a device 200 for determining proportions αi of C reference populations Pi in an ensemble of biological objects according to embodiments of the invention. [0078] In these embodiments, the device 200 comprises a computer, this computer comprising a memory 201 to store program instructions loadable into a circuit and adapted to cause a circuit 202 to carry out steps of the methods of Figures 2 and 3 when the program instructions are run by the circuit 202. The memory 201 may also store data and useful information for carrying steps of the present invention as described above. [0079] The circuit 202 may be for instance: - a processor or a processing unit adapted to interpret instructions in a computer language, the processor or the processing unit may comprise, may be associated with or be attached to a memory comprising the instructions, or - the association of a processor / processing unit and a memory, the processor or the processing unit adapted to interpret instructions in a computer language, the memory comprising said instructions, or - an electronic card wherein the steps of the invention are described within silicon, or - a programmable electronic chip such as a FPGA chip (for « Field-Programmable Gate Array »). [0080] The computer may also comprise an input interface 203 for the reception of input data and an output interface 204 to provide output data. Examples of input data and output data will be provided further. [0081] To ease the interaction with the computer, a screen 205 and a keyboard 206 may be provided and connected to the computer circuit 202. [0082] A database comprising at least one set of reference data is preliminarily collected. [0083] In some embodiments, the at least one set of reference data comprises N sets of reference data, with N being an integer number higher than or equal to 1. Each set of reference data comprises cytometric data corresponding to a given reference ensemble of biological objects. For instance, the cytometric data associated to a plurality of given reference biological objects have been obtained by flow cytometry (or FACS for "fluorescence activated cell sorting"), by PCR-activated cell sorting (PACS), by microsphere affinity proteomics (MAP), by mass spectrometry, by chromatography, by CYTOF, by spectral cytometry, by mass cytometry, by image cytometry, by gene expression on chips (microarray), by sequencing (for example DNA-seq, RNA-seq, scDNA-seq, or scRNA-seq), by in situ hybridization, and/or by microscopy. [0084] Each given reference ensemble of biological objects originates from one among the given reference ensemble of biological objects, each of which containing different reference populations Di of biological objects. Each reference population Di is assigned a label Li belonging to a plurality of C labels Lj, C being an integer number higher than or equal to 2 and j varying from 1 to C. The cytometric data of each set of reference data will often be referred to as “reference statistical sample”, where the expression “statistical sample” has been defined in the Definitions section. [0085] In those embodiments, each reference data in a set of reference data is labelled with a label Li belonging to a plurality of C labels Lj, j being an integer varying between 1 and C. [0086] In some embodiments, the data contained in the database take the form of an electronic file comprising C matrices ^^ + . Each matrix ^^ + contains a concatenation of the reference data labelled with the label Lj in the N sets of reference data. The higher the number N of sets of reference data, the more precise the result of the quantification obtained with the method 100 will be. For instance, each matrix ^^ + comprises nj rows, each row corresponding to data pertaining to a biological object. All biological objects represented in the matrix ^^ + present the label Lj. The electronic file can be received by the device 200 via the input interface 203 and subsequently stored in the memory 201 of the device 200. [0087] The computer-implemented method 100 aims at determining the proportions αj of each reference population Dj among the C reference populations Dj, j being an integer varying between 1 and C. The proportions αj can be gathered in a single vector α. [0088] Figure 2 is a flowchart illustrating an example of steps to be executed to implement the computer-implemented method 100. [0089] In a step S1, the at least one set of reference data is received. As explained above, the at least one set of reference data can take the form of an electronic file comprising C matrices ^^ + as previously described. [0090] In a step S2, transformed vectors Mj are computed. The transformed vectors result from the application of a predetermined function z on the matrices ^^ + . The predetermined function z will be described further. In other words, a transformed vector Mj is defined by: ,- = ^^^^ +^. [0091] In a step S3, a transformed vector R, referred to as test transformed vector, is computed. The test transformed vector R results from the application of the predetermined function z on a matrix Q comprising the set of test data 10 as previously described. In other words, the test transformed vector R is defined by: . = ^^^^ [0092] In a step S4, the proportions αj, for j varying between 1 and C, by
Figure imgf000020_0001
minimizing a distance between the weighted sum of the transformed #^^ and the test transformed vector R. In other words, the proportions αjs are the argument of the minima (i.e., arg min) of the minimization of said distance between the weighted sum of the transformed vectors ∑^ #^^ ^#^^^^ ^^ and the test transformed vector R. In yet other words, said minimization problem is solved with respect to the proportions αj, for j varying between 1 and C. [0093] Steps S1 to S4 may be implemented by the processor of the circuit 202. Advantageously, the processor includes one or more graphics processing units (GPUs). [0094] In some embodiments, step S4 is implemented by minimizing the quantity D defined by: ^ = ^ ^^^^ + ^^ ^ ^, where: ^ ^ = ^′ ^ ^′, P’ representing a matrix comp
Figure imgf000021_0001
^ osed of the C ^ = −2 ^′^^^^^. [0095] The expression ^′^ denotes the transposed matrix of the matrix P’. The minimization is carried out under the conditions that ∑^ #^^ ^# is equal to 1 and all the proportions αi are positive. [0096] In some embodiments referred to as kernel embodiments, the predetermined function z is a kernel function k, such as functions defined with respect to reproducing kernel Hilbert spaces (RKHS) and used in techniques known as kernel methods. In those embodiments, the concept of kernel mean embedding is used. For instance, the kernel function k is Gaussian function. The notation of matrix P0is the concatenation of the matrices ^^ + , in other words, 3.
Figure imgf000021_0002
[0097] Each matrix ^^ + , for j taking values between 1 and C, comprises nj rows. In this case, the matrix P0 comprises n rows, where n is equal to the sum of the numbers nj for j comprised between 1 and C. The rows of the matrix P0 are denoted xil, where i denotes the ith row of the matrix ^^ 4 included in the matrix P0. It can be shown that the computation of the quantity P in the definition of the quantity D involves the computation of nxn values k(xil,xjt). Similarly, the matrix Q comprises m rows yi. In this case, it can be shown that the computation of the quantity q used in the definition of the quantity D involves a computation of n x m values k(xil,yj). [0098] In some other embodiments referred to as vectorization embodiments, the predetermined function z is a vectorial function taking values in ℝM and defined, for a given vector x in ℝs by: ^^^^ = ^ ^ ^cos^"^^^ , sin^"^ ^ ^' ( ^ ^ # # ^ #^^ , X comprising r rows denoted x1, x2,…xr, by:
Figure imgf000022_0001
* , where (ωi)i ^ 1 .. M/2 is a statistical sample of a distribution Λk, Λk being the Fourier transform of a translation invariant kernel function k. For instance, the translation invariant kernel function k is a Gaussian function. [0099] In those embodiments, the concept of Random Fourier feature matching is used. [0100] In the vectorization embodiments, step S4 is simpler as it can be shown that no further calculation, except for the minimization of the quantity D, is necessary. The computation of the transformed vectors Mj (^^^^ ^^), and of the test transformed vector R (^^^^^ is sufficient. In other words, to compute ^^^5^, the matrix containing the concatenated N sets of reference data is multiplied by the matrix containing the values of the statistical sample (ωi)i ^ 1 .. M/2 of the distribution Λk, Then, the cosine and the sine functions are applied to the results. For instance, to minimize the quantity D, a solver program can be used. An example of software package that can be used is CVXOPT. [0101] This reduced number of computations provides one of the advantages of the method 100 in the vectorization embodiments, which is the computational speed, as compared to prior art methods. [0102] Figure 3 shows a comparison of computation times (for estimation of the parameters) between the method 100 (referred to on Figure 3 as KQuant) and a method for quantification of biological objects of the prior art referred to on Figure 3 as GMM and described in the paper “A Machine Learning Approach to the Classification of Acute Leukemias and Distinction From Nonneoplastic Cytopenias Using Flow Cytometric Data”, by Monaghan et al., in American Society for Clinical Pathology, 2021, which is incorporated herein by reference. The method GMM (which stands for Gaussian Mixture Model) is a machine learning based method for classifying multidimensional data obtained by flow cytometry. A GMM makes it possible to define a descriptor of a distribution (for example the distribution representing a cell population in a flow cytometry experiment) and an estimate of the proportion. [0103] The use of GMMs as a descriptor of a cell population in a flow cytometry can be realized according to any known prior art, for example the above-cited article of Monaghan et al., in American Society for Clinical Pathology, 2021, or according to the article of Ko et al., “Clinically validated machine learning algorithm for detecting residual diseases with multicolor flow cytometry analysis in acute myeloid leukemia and myelodysplastic syndrome”, EBioMedicine 37, 2018, which is incorporated herein by reference. In the results presented in Figure 3, both methods 100 and GMM were applied on a series of test data of increasing size (from 25000 points, i.e. biological objects in the test data, to 1,000,000 points). It can be observed that the computation time is shorter with the method 100 than with the method GMM from the prior art for all sizes. This shows that the method 100 is more efficient as a descriptor as it is agnostic to the type of overall distribution to be described. Furthermore, the method 100 is more efficient in computation time. [0104] Advantageously, in the vectorization embodiments, a specific library can be used. For instance, the KeOps library, that allows computing reductions of large arrays whose entries are given by a mathematical formula or a neural network, can be used. This library can be used with Python (NumPy, PyTorch), Matlab and R. [0105] In some embodiments, referred to as match labeling embodiments, the method 100 further comprises additional steps S5, S6, S7 and S8, that will be described below, and allowing to perform match labeling. [0106] Indeed, there exists some methods to segment a set of cytometric data into clusters corresponding to populations or subpopulations. However, often, those methods do not provide labels of the clusters obtained via the segmentation. [0107] In the match labeling embodiments, the predetermined function z is the one as used in the vectorization embodiments previously described. [0108] Figure 4 is a flowchart illustrating an example of further steps of the method 100 as illustrated in Figure 2 to be executed to implement a computer-implemented method for match labeling. [0109] In a step S5, a set of clusters Gk associated with the set of test data 10 is received. The set of clusters Gk may have been preliminary obtained by the application of a segmentation method on the set of test data 10. For instance, the set of clusters Gk are received via the input interface 203 of the device 200. By definition, each cluster Gk in the set of clusters Gk is associated with a population. Typically, each cluster Gk is received in the form of a matrix Xk.
Figure imgf000024_0001
[0110] In a step S6, a distance dki is computed between z(Xk) for each cluster Gk and for each matrix ^^ ^ as defined above, for i taking values between 1 and C. [0111] In a step S7, for each cluster Gk, the reference population Ppk for which the distance d_kpk is minimal is determined, pk being an integer comprised between 1 and C. In other words, in step S7, for each cluster Gk, a minimum of computed dki in step S6 for i taking values between 1 and C, is determined as d_kpk and the respective population Ppk to this minimum distance is determined as the reference population. [0112] In a step S8, each label Lpk corresponding to the reference population Ppk determined at step S7 is assigned to the corresponding cluster Gk. [0113] Steps S6, S7 and S8 may be implemented by the processor of the circuit 201 of the device 201.

Claims

CLAIMS 1. A computer-implemented method (100) for determining proportions αi of C reference populations (Pi) in an ensemble of biological objects being chosen among cells, vesicles of cellular origin, acellular microorganisms, and/or biofunctionalized materials, C being an integer number, said ensemble of biological objects being associated with a set of test data (10), said method comprising: - receiving at least one set of reference data associated with a reference ensemble of biological objects, wherein each reference data is labelled with a label Li belonging to a plurality of C labels (Lj), each label (Lj) among the plurality of C labels (Lj) being associated with one corresponding reference population (Pj) among the C reference populations, - determining said proportions αi by minimizing a distance between a first vector and a second vector, said first vector being defined weighted sum over i, with proportions αi as
Figure imgf000025_0001
weights, of C transformed for i an integer varying from 1 to C, where each ^^ ^ represents a matrix with ni rows, nj for j comprised between 1 and C, wherein each row comprises a sample composed of the reference data labelled with the label denotes a predetermined function, and the weight of each transformed vector is the proportion αi, said second vector resulting from the application of said predetermined function z to a second matrix Q comprising a sample comprising the set of test data (10). 2. The method according to claim 1, wherein minimizing the distance between the first vector and the second vector comprises minimizing a quantity D defined by: ^ = ^ ^^^^ + ^ ^ ^ ^, wherein α represents a vector having as coefficients the αi, ^
Figure imgf000025_0002
^ ^ = ^′^^′, where P’ is a matrix composed of the C vectors ^ = −2 ^′^^^^^. 3. The method according to claim 1 or 2, wherein the predetermined function z is based on a kernel function k. 4. The method according to claim 3, wherein, -computing P involves a computation of nxn values k(xil,xjt), wherein xil denotes an ith vector with the label Ll in a matrix P0 and xjt denotes a jth vector with the label Lt of P0, the matrix P0 being a concatenation of the matrices ^^ ^ , n being
Figure imgf000026_0001
^^^ , - Q comprises m vectors yj, and computing q involves a computation of n x m values k(xil,yj). 5. The method according to claim 3 or 4, wherein the kernel function k is a Gaussian function. 6. The method according to claim 1 or 2, wherein the predetermined function z is a vectorial function taking values in ℝM and defined, vector x in ℝs by:
Figure imgf000026_0002
^ ^cos^"^ ^ ^^ , sin^ ^ ^' ( ^ ^ # "# ^ #^^ , X comprising r rows denoted x1, x2,…xr, by:
Figure imgf000026_0003
* , where (ωi)i ^ 1.. M/2 is a statistical sample of a distribution Λk, Λk being the Fourier transform of a translation invariant kernel function k. 7. The method according to claim 6, wherein the invariant kernel function k is a Gaussian function. 8. The method according to claim 6 or 7, further comprising: - receiving a set of K clusters (Gk) associated with the set of test data (10), k being integer varying from 1 to K, and (Gk) denoting the k-th cluster, wherein each cluster (Gk) in the set of K clusters (Gk) is associated with a population and is represented with a matrix Xk, - computing, for each cluster (Gk) and for each matrix ^^ ^ , a distance d_ki between z(Xk) and ^^^^ ^^, - determining, for each cluster (Gk), the reference population Ppk for which the distance d_kpk is minimal, wherein the reference population Ppk is associated with the label (Lpk), - assigning the label (Lpk) to the k-th cluster (Gk). 9. The method according to any one of claims 1 to 8, wherein the biological objects are: - cells of animal, plant, fungal, protist, bacterial or archaebacterial origin, - vesicles of cellular origin chosen from exosomes, ectosomes, microvesicles, microparticles, prostasomes, oncosomes, matrix/calcification vesicles or apoptotic bodies, - acellular microorganisms chosen from viruses, viroids and prions, and/or - biofunctionalized materials comprising a material of synthetic or biological origin chosen from a nanoparticle (such as a nanobead, a nanosphere or a nanocapsule), a microparticle (such as a microbead, a microsphere or a microcapsule), a lipid vesicle (such as a unilamellar vesicle, a multilamellar vesicle, a lipoplex, a polyplex, a lipopolyplex, a liposome, a niosome, a cochleate, a virosome, an immunostimulating complex), said material of synthetic or biological origin being coupled to, or coated with, one or more peptide(s), protein(s), antibody(ies), antibody fragment(s), receptor(s), cytokine(s), chemokine(s), toxin(s), oligonucleotide(s), colored or fluorescent molecule(s), amine, carboxyl or hydroxyl group(s), bioactive molecule(s) (such as an immunomodulatory molecule, a small chemical molecule, a peptido- mimetic, a drug), molecule(s) of biotin, avidin or streptavidin, or a combination thereof. 10. The method according to any one of claims 1 to 9, wherein the reference and test data are obtained by flow cytometry (or FACS for "fluorescence activated cell sorting"), by PCR-activated cell sorting (PACS), by microsphere affinity proteomics (MAP), by mass spectrometry, by chromatography, by CYTOF, by spectral cytometry, by mass cytometry, by image cytometry, by gene expression on chips (microarray), by sequencing, by in situ hybridization, and/or by microscopy. 11. The method according to any one of claims 1 to 10, wherein the reference and test data are obtained by single-cell analysis technologies or spatial biology analysis technologies. 12. A device comprising a processor configured to carry out a method according to any one of claims 1 to 11. 13. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any one of claims 1 to 11.
PCT/EP2024/064824 2023-05-29 2024-05-29 Method for determining proportions of populations in an ensemble of biological objects Ceased WO2024246152A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP23305847.8A EP4471787A1 (en) 2023-05-29 2023-05-29 Method for determining proportions of populations in an ensemble of biological objects
EP23305847.8 2023-05-29

Publications (1)

Publication Number Publication Date
WO2024246152A1 true WO2024246152A1 (en) 2024-12-05

Family

ID=86776433

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/EP2024/064824 Ceased WO2024246152A1 (en) 2023-05-29 2024-05-29 Method for determining proportions of populations in an ensemble of biological objects

Country Status (2)

Country Link
EP (1) EP4471787A1 (en)
WO (1) WO2024246152A1 (en)

Non-Patent Citations (5)

* Cited by examiner, † Cited by third party
Title
KO ET AL.: "Clinically validated machine learning algorithm for detecting residual diseases with multicolor flow cytometry analysis in acute myeloid leukemia and myelodysplastic syndrome", EBIOMEDICINE, vol. 37, 2018
MONAGHAN ET AL., AMERICAN SOCIETY FOR CLINICAL PATHOLOGY, 2021
MONAGHAN ET AL.: "A Machine Learning Approach to the Classification of Acute Leukemias and Distinction From Nonneoplastic Cytopenias Using Flow Cytometric Data", AMERICAN SOCIETY FOR CLINICAL PATHOLOGY, 2021
MONAGHAN SARA A ET AL: "A Machine Learning Approach to the Classification of Acute Leukemias and Distinction From Nonneoplastic Cytopenias Using Flow Cytometry Data", vol. 157, no. 4, 1 April 2022 (2022-04-01), US, pages 546 - 553, XP093093510, ISSN: 0002-9173, Retrieved from the Internet <URL:https://watermark.silverchair.com/aqab148.pdf?token=AQECAHi208BE49Ooan9kkhW_Ercy7Dm3ZL_9Cf3qfKAc485ysgAAA2cwggNjBgkqhkiG9w0BBwagggNUMIIDUAIBADCCA0kGCSqGSIb3DQEHATAeBglghkgBZQMEAS4wEQQMxrOFrzBViwhB8boqAgEQgIIDGlrvTq23EmLaKekUG3B_5EE__sp81_fS0_eUTCUczS9sMw_-6Gf4zlUJoB3vIvRCsf146OdWnUoK9AWDgum-5IfwQ1Tj> DOI: 10.1093/ajcp/aqab148 *
WIKIPEDIA: "Support vector machine", 22 May 2023 (2023-05-22), pages 1 - 11, XP093093620, Retrieved from the Internet <URL:https://en.wikipedia.org/w/index.php?title=Support_vector_machine&oldid=1156332283> [retrieved on 20231020] *

Also Published As

Publication number Publication date
EP4471787A1 (en) 2024-12-04

Similar Documents

Publication Publication Date Title
Davis-Marcisak et al. From bench to bedside: Single-cell analysis for cancer immunotherapy
Su et al. Single cell proteomics in biomedicine: High‐dimensional data acquisition, visualization, and analysis
Bendall et al. From single cells to deep phenotypes in cancer
JP5888521B2 (en) Method for automated tissue analysis
Porter et al. A motif-based analysis of glycan array data to determine the specificities of glycan-binding proteins
Di Carlo et al. Introduction: why analyze single cells?
Goto-Silva et al. Single-cell proteomics: A treasure trove in neurobiology
JP2009526519A (en) Methods for predicting biological system response
JP7126337B2 (en) Program, apparatus and method for predicting biological activity of compounds
O’Connor et al. Systems Biology and immune aging
WO2022207125A1 (en) Method and device for optically recognising a discrete entity from a plurality of discrete entities
Liu et al. Scalable, compressed phenotypic screening using pooled perturbations
EP4029019A1 (en) Systems and methods for pairwise inference of drug-gene interaction networks
Papp et al. Life on a microarray: assessing live cell functions in a microarray format
Filippov et al. Comparative transcriptomic analyses of thymocytes using 10x Genomics and Parse scRNA-seq technologies
Demangel et al. Host-pathogen interactions from a metabolic perspective: methods of investigation
EP4471787A1 (en) Method for determining proportions of populations in an ensemble of biological objects
Bickle High-content screening: a new primary screening tool?
WO2002072871A2 (en) Method for association of genomic and proteomic pathways associated with physiological or pathophysiological processes
US20250014674A1 (en) Method for cytometric analysis
Gambe et al. Development of a multistage classifier for a monitoring system of cell activity based on imaging of chromosomal dynamics
US20240159647A1 (en) Method and device for optically recognizing a discrete entity from a plurality of discrete entities
Schauer et al. A novel organelle map framework for high-content cell morphology analysis in high throughput
Kuo et al. Analytical Technology for Single-Cancer-Cell Analysis
Song et al. Analysis of image-based phenotypic parameters for high throughput gene perturbation assays

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24730926

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE