WO2025230992A1 - Analyzing multiplexed and multimodal imaging data - Google Patents
Analyzing multiplexed and multimodal imaging dataInfo
- Publication number
- WO2025230992A1 WO2025230992A1 PCT/US2025/026820 US2025026820W WO2025230992A1 WO 2025230992 A1 WO2025230992 A1 WO 2025230992A1 US 2025026820 W US2025026820 W US 2025026820W WO 2025230992 A1 WO2025230992 A1 WO 2025230992A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- cell type
- cell
- computing device
- executed
- segments
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/762—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using clustering, e.g. of similar faces in social networks
- G06V10/763—Non-hierarchical techniques, e.g. based on statistics of modelling distributions
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/766—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using regression, e.g. by projecting features on hyperplanes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/60—Type of objects
- G06V20/69—Microscopic objects, e.g. biological cells or cellular parts
- G06V20/698—Matching; Classification
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/30—Unsupervised data analysis
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B5/00—ICT specially adapted for modelling or simulations in systems biology, e.g. gene-regulatory networks, protein interaction networks or metabolic networks
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H10/00—ICT specially adapted for the handling or processing of patient-related medical or healthcare data
- G16H10/40—ICT specially adapted for the handling or processing of patient-related medical or healthcare data for data related to laboratory analysis, e.g. patient specimen analysis
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H30/00—ICT specially adapted for the handling or processing of medical images
- G16H30/20—ICT specially adapted for the handling or processing of medical images for handling medical images, e.g. DICOM, HL7 or PACS
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H30/00—ICT specially adapted for the handling or processing of medical images
- G16H30/40—ICT specially adapted for the handling or processing of medical images for processing medical images, e.g. editing
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/20—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/30—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for calculating health indices; for individual health risk assessment
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/70—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for mining of medical data, e.g. analysing previous cases of other patients
Definitions
- TITLE ANALYZING MULTIPLEXED AND MULTIMODAL IMAGING DATA
- Single-cell and spatial multiomics have revolutionized biology and biomedical research by enabling analysis at near single molecular resolution.
- Single-cell and spatial multiomics can allow for the integrated measurement of various cell features while preserving their spatial context within biological samples such as human tissues, human cells, and human biofluids. While the promise is great for singlecell and spatial approaches in science, many challenges remain.
- Some of the problems related to manual feature annotation include subjectivity and variability, time-intensive, limited scalability, inter-observer discrepancies, annotation bias, difficulty in integrating multimodal data, and other reproducibility concerns.
- FIG. 1 is a drawing of a network environment according to various embodiments of the present disclosure.
- FIGS. 2A and 2B show a flowchart illustrating one example of functionality implemented as portions of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
- FIG. 3 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
- FIG. 4 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
- FIG. 5 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
- FIG. 6 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
- FIG. 7 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
- FIG. 8 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
- FIG. 9 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
- Single-cell and spatial multiomics can allow for the integrated measurement of various cell features while preserving their spatial context within biological samples such as human tissues, human cells, and human biofluids. Whilst biological samples may or may not also include individual microbes or microbes within complex communities.
- the present disclosure provides for solutions capable of analyzing highly complex data sets (images), identifying data associated with different cell types and states, and providing annotation or metadata associated with the different cell types and states.
- the present disclosure provides for a computational algorithm and associated software application that can automatically determine cell type identities, cell states, and other cell and molecular features in highly multiplexed and multimodal imaging data. It is an automated, scalable, and platform -agnostic computer-implemented method (or algorithm) which can be deployed as a web-based application (or software).
- the application can classify large quantities of spatially resolved cells from multiplexed imaging data without reliance on pre-labeled training datasets.
- the application can employ a hierarchical, multi-level classification framework that incorporates both broad and fine-grained ontologies, allowing for scalable assignment of structural, immune, and specialized subtypes.
- the application begins with predefined “cell recipes” (pre-defined set of markers used to specify cell identities and states by combining proteins or transcripts) and applies cluster-aware micro segmentation to calculate relevance scores for each marker in the context of local cellular neighborhoods. Thresholds for marker positivity can be learned dynamically using segmental regression, enabling the detection of both high-confidence and low-abundance or transitional cell states.
- a user can modify the threshold for marker positivity of one or more cell types, markers, genes, proteins, etc.
- Classification outputs can be refined using k-nearest-neighbor- based deconvolution, local enrichment scoring, and outlier filtering, resulting in accurate cell identity assignments that are robust to noise and applicable across tissue types, disease conditions, and imaging modalities.
- the software inputs are feature matrix and multi-tiered cell type signature matrix, where each cell type is required to be positive in the individual marker/gene.
- the software output is the assigned cell type identity for each cell.
- Linking the algorithm and the software provides a visualization tool that allows for, but is not limited to, 1 ) boundary identification from the associated feature input such as cells and microbes, 2) spatial neighborhood relationships, 3) cell motif analysis, and 4) the mapping of identified cells and microbes in their relevant feature space and their exact locations.
- the visualization tool provides a user interface that allows for interactive visualization and analysis of spatial cell type data including spatial plots and uniform manifold approximation and projection (UMAP) visuals featuring annotations, marker expression thresholds, and weighted cell type calculations. Users can also modify these thresholds for one or more cell types, markers, genes, proteins, etc. within the visualization tool and run the cell type application again with the modified thresholds.
- UMAP uniform manifold approximation and projection
- Users can navigate through complex tissue maps in real-time, including zooming, panning, and dynamic toggling between classification tiers. Additionally, users can overlay single or multiple markers from protein or transcriptom ic data and examine cell identity, expression intensity, spatial coordinates, and relationships within the broader tissue context.
- the visualization tool can allow users to filter cell types, cell states, markers, or a combination of any thereof. Users can also access color annotations, spatial neighborhood connections between cell types across the whole tissue or a user-defined region of interest (ROI), and statistical summaries of marker expression and cell type composition within regions to identify spatial autocorrelation. Additional tools can include annotated mean heatmaps, Voronoi plots, Delaunay triangulation, and proportions of cell types and cell state markers.
- Users can also customize cell type annotations with user- defined colors and select specific tissues or regions for focused visualization.
- the generation of heatmaps that display cell type distributions based at least in part on marker expression can be tailored by the user to reflect either cell type-specific markers or cell state markers.
- the visualization tool can also facilitate the exploration of cell states, providing insights into the proportions of various cell types and states across tissues or regions. [0019] Laborious and time consuming preprocessing steps, such as cell segmentation, can be combined in the present disclosure with whole-slide cell identification and, in some examples, metacellular annotation (such as identification of multicellular arrangements, or “cell neighborhoods”).
- This integration can collapse what were previously separate workflows into a single, streamlined process, which can allow for faster analysis, reduce redundancy, and provide an interactive environment where users can check and refine both cell type annotations and segmentation borders (for example, with human-in-the-loop validation).
- the present disclosure can offer an innovative framework for threshold-based assignment of cell types and cell states from multiplexed and multimodal imaging data.
- the present disclosure can systematically and statistically enhance setting marker positivity thresholds, which can ensure a robust cell-type assignment.
- the accuracy of the present disclosure can be further augmented by capitalizing on the importance assigned to each marker within a multi-tiered cell-type taxonomy.
- the present disclosure can take two inputs: a matrix of marker intensities in each column, cells in rows, and the cell type signature matrix with cell types in rows and markers in columns (1 indicating marker positivity for that cell type, 0 otherwise).
- the present disclosure can calculate Cell Type Relevance scores (CTR), a score that measures the relevance of a molecular profile of a cell to a particular cell type.
- CTR Cell Type Relevance scores
- cell identifies can be further assessed for cell state changes within the tiered cell identity assignments, including high-level cell assignments (e.g., epithelia versus immune cell) or detailed assignments (e.g., oral keratinocytes versus regulatory T cell).
- Benchmarked datasets have been used for current single modality testing, in which a trained pathologist established cell-type assignments. Comparative evaluations have been conducted against CELESTA and SCINA tools, and the Seurat version 5 pipeline. The present disclosure consistently demonstrated superior performance, exhibiting the highest agreement with the reference across datasets and averaging a significant 10% accuracy gain over alternative methods. Notably, challenges were observed in CELESTA and SCINA, which are partially attributable to inaccurate assumptions regarding marker distribution. Furthermore, all methods, except the present disclosure, encountered difficulties with the substantial ⁇ 2.5 million-cell intestine dataset. The present disclosure’s robustness and efficacy can make it a highly reliable solution across varying dataset sizes and complexities.
- the network environment 100 can include a computing environment 103 and a client device 106, which can be in data communication with each other via a network 109.
- the network 109 can include wide area networks (WANs), local area networks (LANs), personal area networks (PANs), or a combination thereof. These networks can include wired or wireless components or a combination thereof. Wired networks can include Ethernet networks, cable networks, fiber optic networks, and telephone networks such as dial-up, digital subscriber line (DSL), and integrated services digital network (ISDN) networks. Wireless networks can include cellular networks, satellite networks, Institute of Electrical and Electronic Engineers (IEEE) 802.11 wireless networks (/.e., WIFI®), BLUETOOTH® networks, microwave transmission networks, as well as other networks relying on radio broadcasts. The network 109 can also include a combination of two or more networks 109. Examples of networks 109 can include the Internet, intranets, extranets, virtual private networks (VPNs), and similar networks.
- VPNs virtual private networks
- the computing environment 103 can include one or more computing devices that include a processor, a memory, and/or a network interface.
- the computing devices can be configured to perform computations on behalf of other computing devices or applications.
- such computing devices can host and/or provide content to other computing devices in response to requests for content.
- the computing environment 103 can employ a plurality of computing devices that can be arranged in one or more server banks or computer banks or other arrangements. Such computing devices can be located in a single installation or can be distributed among many different geographical locations.
- the computing environment 103 can include a plurality of computing devices that together can include a hosted computing resource, a grid computing resource or any other distributed computing arrangement.
- the computing environment 103 can correspond to an elastic computing resource where the allotted capacity of processing, network, storage, or other computing-related resources can vary over time.
- Various applications or other functionality can be executed in the computing environment 103.
- the components executed on the computing environment 103 include cell type application 113, a machine learning model 116, a visualization service 118, and other applications, services, processes, systems, engines, or functionality not discussed in detail herein.
- the cell type application 113 can be executed to identify cell type, cell state, clusters of cells, and other various information from matrix representations of features collected from biological samples, such as, for example, images.
- the cell type application 113 is an algorithm using feature matrices 126 and cell type signatures 129 to identify various cell types from an image.
- the cell type application 113 weighs the importance of the extracted features 133 in order to determine which cell type applies to the identified cell.
- the cell type application 113 can identify more cell types than those provided in the signature matrix and determine the respective weights of those new cell types as well.
- the cell type application 113 can identify pure positive cells having only one cell type and compare the mixed cell having multiple cell types to the pure cell in order to determine the correct cell type. In some embodiments, the cell type application 113 can further communicate with a machine learning model 116 to determine various aspects throughout the process.
- the machine learning model 116 can be executed to perform segmentation of various images to identify cell boundaries. Additionally, the machine learning model 116 can be executed to receive data from the cell type application 113 such as, for example, cell type relevance scores for individual segments and identify micro clusters or other micro architectural features of the cells. In some examples, the machine learning model 116 can be used to identify the cell features from the images.
- various data is stored in a data store 119 that is accessible to the computing environment 103.
- the data store 119 can be representative of a plurality of data stores 119, which can include relational databases or non-relational databases such as object-oriented databases, hierarchical databases, hash tables or similar key-value data stores, as well as other data storage applications or data structures. Moreover, combinations of these databases, data storage applications, and/or data structures may be used together to provide a single, logical, data store.
- the data stored in the data store 119 is associated with the operation of the various applications or functional entities described below. This data can include images 123, feature matrices 126, cell type signatures 129, cell features 133, signature matrices 134, and potentially other data.
- the images 123 can represent image files from a variety of modalities, such as cellular imaging (e.g., phase-contrast microscopy, fluorescent imaging, brightfield microscopy, confocal microscopy, time-lapse microscopy, etc.), tissue imaging (e.g., multiplex colorimetric immunohistochemistry (rnCIHC), multiplex immunofluorescence (mIF), cyclic immunofluorescence (CycIF), multiplexed ion beam imaging (MIBI), codetection by indexing (CODEX)/Phenocycler (PCF), digital spatial profiling (DSP), etc., or other form of imaging.
- cellular imaging e.g., phase-contrast microscopy, fluorescent imaging, brightfield microscopy, confocal microscopy, time-lapse microscopy, etc.
- tissue imaging e.g., multiplex colorimetric immunohistochemistry (rnCIHC), multiplex immunofluorescence (mIF), cyclic immunofluorescence (Cyc
- the images 123 can represent segmented images, or images which have been preprocessed and segmented into various segments corresponding to individual cells, cell clusters, tissues, or other features.
- the images 123 are representative of any matrix representations of features collected from biological samples.
- the feature matrices 126 can represent matrices of cell features 133 and associated intensities/count in each column and cell type signatures 129 in rows, as well as cell type signature matrices with cell type signatures 129 in columns and cell features 133 in rows.
- feature matrices 126 could include z-normalized intensity values indicating an intensity level of specific markers within a set of cells.
- feature matrices 126 could include log-normalized counts of transcripts for a variety of genes within a set of cells.
- the feature matrices 126 can be generated by the machine learning model 116 from segmented images 123 that could be trained at least in part on existing cell type segmentation and/or morphological features.
- the feature matrices 126 can further include data regarding the respective importance, or weights, associated with each cell feature 133.
- the cell type signatures 129 can represent various broad cell types (e.g., epithelial, immune cells, stem cells, blood cells, bone cells, muscle cells, gametes, neurons, etc.), specific cell types (e.g., oral keratinocytes, regulatory T cell, erythrocytes, odontoblasts, etc.), cell states (e.g., activated, resting, diseased, etc. , or other cellular or molecular features which can be identified by the cell type application 113.
- the cell type signatures 129 can be stored in the data store 119 and associated with one or more cell features 133.
- the cell features 133 can represent various observable characteristics of a particular cell type or cell state which can be identified from the images 123.
- cell features 133 can include cellular morphology, nuclei characteristics, cytoplasm, cellular borders, inclusion bodies and vacuoles, cellular arrangement, mitotic activity, cellular diversity, etc. as well as associated and unassociated biomolecules and markers, including DNA, RNA, proteins, metabolites, etc.
- the data store 119 can also contain cell features 133 that did not originate from the segmented image 123.
- the signature matrices 134 can be representative of matrices having various cell types listed in rows and various cell markers, or cell features 133, in columns.
- the signature matrices 134 can include a value of “1 ,” indicating marker positivity for that cell type, or a value of “0” to indicate otherwise.
- the client device 106 is representative of a plurality of client devices that can be coupled to the network 109.
- the client device 106 can include a processor-based system such as a computer system.
- a computer system can be embodied in the form of a personal computer (e.g., a desktop computer, a laptop computer, or similar device), a mobile computing device (e.g., personal digital assistants, cellular telephones, smartphones, web pads, tablet computer systems, music players, portable game consoles, electronic book readers, and similar devices), media playback devices (e.g., media streaming devices, BluRay® players, digital video disc (DVD) players, set-top boxes, and similar devices), a videogame console, or other devices with like capability.
- a personal computer e.g., a desktop computer, a laptop computer, or similar device
- a mobile computing device e.g., personal digital assistants, cellular telephones, smartphones, web pads, tablet computer systems, music players, portable game consoles, electronic book readers, and similar devices
- the client device 106 can include one or more displays 136, such as liquid crystal displays
- the display 136 can be a component of the client device 106 or can be connected to the client device 106 through a wired or wireless connection.
- the client device 106 can be configured to execute various applications such as a client application 139 or other applications.
- the client application 139 can be executed in a client device 106 to access network content served up by the computing environment 103 or other servers, thereby rendering a user interface 143 on the display 136.
- the client application 139 can include a browser, a dedicated application, or other executable, and the user interface 143 can include a network page, an application screen, or other user mechanism for obtaining user input.
- the client device 106 can be configured to execute applications beyond the client application 139 such as email applications, social networking applications, word processors, spreadsheets, or other applications.
- a user can initiate a process through the client application 139 on a client device 106.
- the user can provide images 123 via the client device 106.
- the images 123 can be segmented by a machine learning model 116 to produce segmented images 123, which can then be used by the cell type application 113 for further processing.
- the cell type application 113 can identify various cell features 133 obtained from within the segmented images 123.
- the cell type application 113 can use the feature matrices 126 (consisting of cell features 133 and cell type signatures 129) to assign a cell type relevance score to each cell based at least in part on the data in the feature matrices 126 and based at least in part on the cell features 133 identified from the images 123.
- cell type relevance scores can be assigned by multiplying the features matrices 126 by the signature matrices 134.
- the cell type application 113 can use clustering methods with the machine learning model 116 to subcluster the feature matrices 126. After that, the cell type application 113 can use the subclusters to determine a distribution of cell type relevance scores.
- Piece-wise linear regression can be applied to detect points indicating drastic changes along the growing trend of cell type relevance scores. Those points divide the subclusters into a low-relevance group and a high- relevance group. A threshold is then determined by minimizing the misplacement of cells between low and high relevance groups. Based at least in part on the positive identification threshold and the cell type relevance score, the cell type application 113 can determine whether to assign a cell type provided in cell type signatures 129 to a particular cell. In some examples, the threshold can be modified by a user for one or more cell type signatures 129, genes, markers, or proteins.
- the cell type application 113 can be executed to call a function to assign a cell type provided in cell type signatures 129 and provide the threshold set by a user and the cell type relevance scores as arguments to the function. Finally, for cells or segments which are positive for multiple cell type signatures 129, the cell type application 113 can identify “pure cells” having only one assigned cell type signature 129, and “mixed cells” having multiple assigned cell types 129. The cell type application 113 can compare the “pure cells” to the “mixed cells” to determine a correct cell type signature 129 based at least in part on the similarity. In some examples, the cell type application 113 can identify “unknown cells” that are not classified as positive for any cell type signature 129.
- FIG. 2 shown is a flowchart that provides one example of the operation of a portion of the cell type application 113.
- the flowchart of FIG. 2 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the cell type application 113.
- the flowchart of FIG. 2 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
- the cell type application 113 can be executed to receive a segmented image 123.
- the cell type application 113 can receive the segmented image 123 from a client application 139 on a client device 106.
- the client application 139 can obtain the segmented image 123 from a data store 119 in response to a user interaction.
- the client application 139 can then send the segmented image 123 to the cell type application 113 to initiate the processing of the image 123.
- the cell type application 113 receives the segmented image 123 from the machine learning model 116, from the data store 119, or from another system, service, or application in the network environment 100.
- the cell type application 113 can be executed to identify one or more cell features 133 from the segmented image 123.
- the cell type application 113 can use imaging processing techniques to extract various data from the segmented images 123 received at block 200. For example, the cell type application 113 can determine the fluorescence of a particular segment in order to determine which protein is being expressed by the corresponding cell and how much of that protein.
- the cell type application 113 can identify a plurality of cell features 133 from the segmented image 123.
- the cell type application 113 can be executed to determine weights for the cell features 133.
- the cell type application 113 can determine a respective weight for each of the cell features 133 identified at block 203.
- the weights determined by the cell type application 113 can be representative of the importance of the cell feature 133 in determining the cell type as provided in cell type signatures 129.
- the cell type application 113 can determine a respective weight for a cell feature 133 based at least in part on a feature matrix 126.
- the cell type application 113 can compare the cell features 133 identified in a particular segment of the segmented image 123 to cell features 133 associated with a particular cell type provided in cell type signatures 129 in a feature matrix 126 and use the stored importances or weights from the feature matrix 126 to determine the weight of the respective cell feature 133.
- the cell type application 113 can determine a respective weight for a cell feature 133 based at least in part on other data.
- the cell type application 113 can be executed to assign feature scores to cell features 133.
- the cell type application 113 can assign a feature score to each respective cell feature 133 based at least in part on the weights determined at block 206.
- the cell type application 113 can use a feature matrix 126 from the data store 119 to determine a feature score for each respective cell feature 133 identified at block 203.
- the feature score given to each cell feature 133 by the cell type application 113 can represent a score of the similarity between the cell features 133 identified at block 203 and the cell features 133 listed in the feature matrix 126 and associated with a cell type signature 129.
- the cell type application 113 can assign a low feature score for that particular cell type signature 129.
- the cell type application 113 can be executed to calculate the cell type relevance (CTR) score.
- CTR cell type relevance
- the cell type application 113 can calculate a CTR score for each cell of the segmented image received at block 200.
- the cell type application 113 can calculate the CTR score based at least in part on the cell features 133 identified at block 203, the weights determined at block 206, and/or the feature scores assigned at block 209.
- the cell type application 113 calculates the CTR score by performing a weighted linear combination of the feature scores assigned at block 209 using the weights determined at block 206.
- the cell type application 113 can calculate a CTR score for each cell type stored in cell type signatures 129.
- the cell type application 113 can calculate the CTR scores by multiplying a matrix of the feature scores determined at block 209 with a signature matrix 134.
- the cell type application 113 can be executed to assign a cell type signature 129.
- the cell type application 113 can assign a cell type signature 129 to each segment of the segmented images 123 based at least in part on the CTR score calculated at block 213.
- the cell type application 113 can assign a cell type signature 129 which corresponds to the highest CTR score associated with the segment.
- the cell type application 113 can indicate one or more cell types not present in cell type signatures 129 as a potential new cell type and flag the potential new cell type for review.
- the cell type application 113 can be executed to identify clusters using machine learning (ML) or other methods.
- the cell type application 113 can use the machine learning model 116 to identify one or more clusters of cell type signatures 129 in the segmented image 123.
- the cell type application 113 can send the segmented image 123 with the assigned CTR scores to the machine learning model 116 with an instruction to identify clusters based at least in part on the CTR scores.
- the cell type application 113 can receive the identified clusters from the machine learning model 116.
- the cell type application 113 can be executed to determine the CTR score distribution.
- the cell type application 113 can determine the CTR score distribution across the clusters identified at block 219.
- the cell type application 113 can determine the CTR score distribution for each cluster and/or across all clusters in the entire feature matrix 126.
- the cell type application 113 can be executed to determine a positive identification threshold.
- the positive identification threshold can represent a threshold which can be used to determine the difference between a positively identified cell and the background of cell features 133 collected from, but not limited to, the image 123.
- the positive identification threshold can be a minimum CTR score which is required to assign a cell type signature 129 to a cell.
- the cell type application 113 can determine the positive identification threshold based at least in part on the CTR score distribution determined at block 223. In some examples, the cell type application 113 can determine the positive identification threshold using piecewise linear regression.
- the cell type application 113 could calculate and store as a vector the median CTR scores across all clusters identified at block 219 for any given cell type and then fit a segmented regression model to identify breakpoints that divide the data into distinct linear segments. In some examples, a maximum of three breakpoints can be identified. In some examples, the highest identified breakpoint is the positive identification threshold. In some examples, the positive identification threshold can be set by a user through the user interface 143 for one or more cell types, markers, genes, proteins, or a combination of any thereof. In some examples, the cell type application 113 can be executed to call a function to determine a positive identification threshold and provide the threshold set by the user as an argument to the function.
- the cell type application 113 can be executed to determine a cell identification.
- the cell type application 113 can determine a positive or a negative cell identification (e.g., a measure of whether a cell type signature 129 can be assigned to a segment) based at least in part on the positive identification threshold determined at block 226.
- the cell type application 113 can determine a positive cell identification for an individual segment based at least in part on the corresponding CTR score exceeding the positive identification threshold determined at block 226.
- the cell type application 113 can determine a negative cell identification for an individual segment based at least in part on the corresponding CTR score failing to exceed the positive identification threshold determined at block 226.
- the cell type application 113 can be executed to identify “mixed” cells.
- “Mixed” cells can represent cells of the segmented image 123, or other data source, which have a positive cell identification determined at block 229 for more than one cell type signature 129, or which is assigned multiple cell type signatures 129. For example, if a particular segment representing a cell has more than one CTR score exceeding the respective positive identification threshold, the cell type application 113 can determine that the particular segment represents a “mixed” cell. In some examples, the cell type application 113 can identify all of the “mixed” cells which occur in an image 123.
- the cell type application 113 can be executed to identify “pure” cells.
- “Pure” cells can represent cells of the segmented image 123, or other data source, which have a positive cell identification determined at block 229 for only a single cell type signature 129, or which only has a single assigned cell type signature 129. For example, if a particular segment representing a cell has only one CTR score exceeding the respective positive identification threshold, the cell type application 113 can determine that the particular segment represents a “pure” cell. In some examples, the cell type application 113 can identify all of the “pure” cells which occur in an image 123.
- the cell type application 113 can be executed to assign the correct cell type signature 129.
- the cell type application 113 can assign the correct cell type signatures 129 to “mixed” cells identified at block 233 based at least in part on a comparison of the “mixed” cell to the “pure” cells identified at block 236. For example, the cell type application 113 can take each “mixed” cell identified at block 233 and determine which cell type signatures 129 are positively identified or assigned to the cell. Then, the cell type application 113 can compare the “mixed” cell to a “pure” cell of each positively identified cell type 129 which applies to the “mixed” cell.
- the cell type application 113 can determine which of the “pure” cell types 129 is most similar to the “mixed” cell and assign the correct cell type signature 129 accordingly.
- the machine learning model 116 can be executed to validate the correct cell type signatures 129 assigned by the cell type application 113.
- a user could validate the correct cell type signatures 129 assigned by the cell type application 113.
- FIG. 3 shown is a flowchart that provides one example of the operation of a portion of the cell type application 113.
- the flowchart of FIG. 3 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the cell type application 113.
- the flowchart of FIG. 3 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
- the cell type application 113 can be executed to assign one or more cell type signatures 129 to one or more cells within a segmented image 123.
- the cell type application 113 can identify various cell features 133 obtained from within, but not limited to, the segmented images 123. Having collected the cell features 133, the cell type application 113 can use the feature matrices 126 (consisting of cell features 133 and cell type signatures 129) to assign a cell type relevance score to each cell based at least in part on the data in the feature matrices 126 and based at least in part on the cell features 133 identified from, but not limited to, the images 123.
- cell type relevance scores can be assigned by multiplying the features matrices 126 by the signature matrices 134.
- the cell type application 113 can use clustering methods with the machine learning model 116 to subcluster the feature matrices 126. After that, the cell type application 113 can use the subclusters to determine a distribution of cell type relevance scores.
- Piece-wise linear regression can be applied to detect points indicating drastic changes along the growing trend of cell type relevance scores. Those points divide the subclusters into a low-relevance group and a high-relevance group.
- a threshold is then determined by minimizing the misplacement of cells between low and high relevance groups. In some examples, the threshold can be set by a user.
- the cell type application 113 can determine whether to assign a cell type provided in cell type signatures 129 to a particular cell. Finally, for cells or segments which are positive for multiple cell type signatures 129, the cell type application 113 can identify “pure cells” having only one assigned cell type signature 129, and “mixed cells” having multiple assigned cell types 129. The cell type application 113 can compare the “pure cells” to the “mixed cells” to determine a correct cell type signature 129 based at least in part on the similarity. In some examples, the cell type application 113 can identify “unknown cells” or potential new cell types that are not classified as positive for any cell type signature 129.
- the cell type application 113 can be executed to identify one or more tiers for the cell type signatures 129 assigned at block 300.
- the tiers can be defined in the cell type signatures 129.
- the tiers can be defined by a user.
- the cell type application 113 can be executed to receive defined tiers from another application within the network environment 100.
- the cell type application 113 can be executed to call a function to identify one or tiers and provide the cell type signatures 129 and the cell type signatures 129 or defined tiers as arguments to the function.
- a tier can be defined as broad categorizations of cell type signatures 129.
- a Tier zero (0) cell type signature 129 could be defined as structural and immune cell type signatures 129.
- a tier can further divide an existing tier into more specific cell types.
- a Tier one (1 ) cell type signature 129 could further define structural cell type signatures 129 as epithelial cells, vasculature, muscle, etc. and define immune cell type signatures 129 as fibroblasts, immune myeloid cells, lymphoid-derived cells, etc.
- the cell type application 113 can be executed to annotate the cell type signatures 129 based at least in part on tiers identified at block 303.
- the cell type application 113 can be executed to call a function to label the cell type signatures 129 assigned at block 300 with one or more respective tiers identified at block 303 and provide the cell type signatures 129 and the tiers as arguments to the function.
- the cell type application 113 could label a fibroblast identified at block 300 as a Tier zero (0) immune cell type signature 129 and a Tier one (1 ) fibroblast based at least in part on the tiers identified at block 303.
- the cell type application 113 can be further executed to send the annotations to the visualization service 118 for display to a user on a client device 106.
- FIG. 4 shown is a flowchart that provides one example of the operation of a portion of the machine learning model 116.
- the flowchart of FIG. 4 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the machine learning model 116.
- the flowchart of FIG. 4 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
- the machine learning model 116 can be executed to receive a segmented image 123, associated one or more cell type signatures 129, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113.
- the image coordinates can represent the x and y coordinates of a cell within the segmented image 123.
- the machine learning model 116 can be executed to call a function to retrieve a segmented image 123, associated one or more cell type signatures 129, and image coordinates assigned to one or more cells from the cell type application 113 and include a file path as an argument to the function.
- the cell type application 113 can automatically send the segmented image 123, associated cell type signatures 129, and image coordinates to the machine learning model 116 after assigning cell type signatures 129 to the segmented image 123.
- the machine learning model 116 can be executed to partition a segmented image 123 into one or more subsets.
- the machine learning model 116 can be executed to partition the segmented image 123 into one or more vertical subsets based at least in part on defined x coordinate boundaries.
- the machine learning model 116 can be executed to partition the segmented image 123 into one or more horizontal subsets based at least in part on defined y coordinate boundaries.
- the machine learning model 116 can be executed to partition the segmented image 123 into one or more square grid subsets by dividing coordinates along both axes.
- the x coordinate boundaries and/or y coordinate boundaries are defined by a user.
- the partitions can be recursively refined to ensure the partitions contain a minimum and/or maximum number of cells.
- the partitions exceeding a maximum cell count can be subdivided and partitions containing a number of cells below the minimum cell count can be merged.
- the machine learning model 116 can be executed to construct one or more sub-graphs based at least in part on the subsets created at block 403.
- the machine learning model 116 can be executed to call a function to construct one or more sub-graphs and provide the subsets created at block 403 as arguments to the function.
- the nodes of the sub-graphs can represent cells having one or more cell type signatures 129 and the edges can capture one or more spatial relationships based at least in part on the k-nearest neighbors of the cells.
- the machine learning model 116 can be executed to assign one or more cell neighborhoods.
- the machine learning model 116 can be executed to call a minimum cut-based loss function and provide the nodes and the edges from the one or more sub-graphs generated at block 406.
- the one or more sub-graphs can be randomly shuffled in every epoch.
- the one or more sub-graphs can be processed in a graph neural network (GNN).
- the one or more sub-graphs can be treated as independent training samples and can be processed in mini batches with a gradient accumulation technique.
- the machine learning model 116 can be executed to perform one or more consensus-based tissue cell neighborhood assignments.
- the machine learning model 116 can be executed to call a function to assign consensusbased tissue cell neighborhoods and provide the cell neighborhoods from block 409 as arguments to the function.
- the one or more consensus-based tissue cell neighborhood assignments can be derived through majority voting over the tissue cell neighborhood assignments from the one or more partitioning methods.
- cells without consensus across the one or more partitioning methods can be removed.
- FIG. 5 shown is a flowchart that provides one example of the operation of a portion of the visualization service 118.
- the flowchart of FIG. 5 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the visualization service 118.
- the flowchart of FIG. 5 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
- the visualization service 118 can be executed to receive a segmented image 123, one or more associated cell type signatures 129, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113.
- the image coordinates can represent the x and y coordinates of a cell within a segmented image 123.
- the visualization service 118 can be executed to call a function to retrieve a segmented image 123 from the data store 119, the cell type application 113, or another program within the network environment 100.
- the visualization service 118 can be executed to call a function to retrieve one or more cell type signatures 129 and image coordinates assigned to one or more cells from the cell type application 113 and include a file path as an argument to the function.
- the cell type application 113 can automatically send the cell type signatures 129 and image coordinates to the visualization service 118 after assigning cell type signatures 129 to a segmented image 123.
- the visualization service 118 can be executed to partition a segmented image 123 into one or more subsets.
- the visualization service 118 can be executed to partition the segmented image 123 into one or more vertical subsets based at least in part on defined x coordinate boundaries.
- the visualization service 118 can be executed to partition the segmented image 123 into one or more horizontal subsets based at least in part on defined y coordinate boundaries.
- the visualization service 118 can be executed to partition the segmented image 123 into one or more square grid subsets by dividing coordinates along both axes.
- the x coordinate boundaries and/or y coordinate boundaries are defined by a user.
- the partitions can be recursively refined to ensure the partitions contain a minimum and/or maximum number of cells. In some examples, the partitions exceeding a maximum cell count can be subdivided and partitions containing a number of cells below the minimum cell count can be merged. In some examples, the visualization service 118 can be executed to partition the segmented image 123 based at least in part on one or more partitions defined by a user.
- the visualization service 118 can be executed to locate one or more cell type signatures 129 within one or more partitions created at block 503.
- the visualization service 118 can be executed to call a function to locate one or more cell type signatures 129 and provide the cell type signatures 129 and image coordinates received at block 500 as arguments to the function.
- the one or more cell type signatures 129 to be located by the visualization service 118 within the partitions are defined by a user.
- the visualization service 118 can be executed to determine a cooccurrence between two or more cell type signatures 129 located in the partitions at block 506.
- the visualization service 118 can be executed to call a function to determine one or more cell type signatures 129 located within a defined distance of one or more different cell type signatures 129 and provide the cell type signatures 129, image coordinates, and the defined distance as arguments to the function.
- the visualization service 118 can be executed to identify a first cell type signature 129 located within a defined distance from a second cell type signature 129.
- the defined distance can be modified by a user.
- the visualization service 118 can be executed to determine cells having a cell type A that are located within fifty (50) microns of cells having a cell type B within a two-hundred (200) micron partition of a segmented image 123.
- the visualization service 118 can be further executed to send the co-occurrences to a user interface 143 for display to a user. After block 509, the process depicted by the flowchart of FIG. 5 can come to an end.
- FIG. 6 shown is a flowchart that provides one example of the operation of a portion of the visualization service 118.
- the flowchart of FIG. 6 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the visualization service 118.
- the flowchart of FIG. 6 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
- the visualization service 118 can be executed to receive a segmented image 123, one or more associated cell type signatures 129, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113.
- the image coordinates can represent the x and y coordinates of a cell within a segmented image 123.
- the visualization service 118 can be executed to call a function to retrieve a segmented image 123 from the data store 119, the cell type application 113, or another program within the network environment 100.
- the visualization service 118 can be executed to call a function to retrieve one or more cell type signatures 129 and image coordinates assigned to one or more cells from the cell type application 113 and include a file path as an argument to the function.
- the cell type application 113 can automatically send the cell type signatures 129 and image coordinates to the visualization service 118 after assigning cell type signatures 129 to a segmented image 123.
- the visualization service 118 can be executed to compare the cells within the plurality of segmented images 123.
- the visualization service 118 can be executed to call a function to compare the cells within a first segmented image 123 with the cells within a second segmented image 123 and provide the cell type signatures 129 and coordinates received at block 600 as arguments to the function.
- the visualization service 118 can be executed to compare one or more cell states, a proportion of one or more cell type signatures 129, one or more co-occurrences, one or more biomarkers, and/or one or more gene expressions.
- the visualization service 118 can be executed to compare the cells within a segmented image 123 associated with one condition with the cells within one or more segmented images 123 associated with other conditions. For example, the visualization service 118 can compare the cells from a segmented image 123 associated with a healthy state with the cells from a segmented image 123 associated with a diseased state. In another example, the visualization service 118 can compare the cells from a segmented image 123 associated with a first condition of a diseased state with the cells from segmented images 123 associated with other conditions of the diseased state. In some examples, the visualization service 118 could compare a region of interest between different images associated with different conditions of a diseased state or between a healthy state and a diseased state.
- the visualization service 118 can be executed to identify one or more patterns based at least in part on the comparisons made at block 603.
- the visualization service 118 can be executed to call a function to estimate a survival outcome based at least in part on the cells within a segmented image 123 and provide a time to event associated with comparisons made at block 603 as an argument to the function.
- a co-occurrence between T-cells and B-cells could be a biomarker for a cancer type and greater percentages of the total cell distribution of this biomarker can correlate to lower survival outcomes.
- the visualization service 118 could identify this pattern by comparing survival outcomes (or times to event) associated with varying levels of this biomarker and estimate a survival outcome for a patient associated with a segmented image 123 having the biomarker present in 25% of the total cell distribution.
- the visualization service 118 can be executed to validate the one or more patterns identified at block 606.
- the visualization service 118 can be executed to receive user feedback associated with the one or more patterns. For example, a user could provide feedback that a biomarker present in 20% or greater of the total cell distribution is associated with a 50% survival outcome. In another example, a user could provide feedback that a patient associated with a segmented image 123 used to identify a pattern at block 606 was diagnosed with a cancer type a specified time after the pattern was identified. In another example, a user could provide feedback that a biomarker present in 20% or greater of the total cell distribution is associated with decreased response to a drug or with increased probability of cancer recurrence.
- the visualization service 118 can be executed to receive feedback from the machine learning model 116 associated with the one or more patterns.
- the visualization service 118 could be executed to send the one or more patterns to the machine learning model 116 and receive feedback from the machine learning model 116.
- the visualization service 118 could then be executed to modify the pattern based at least in part on the feedback.
- FIG. 7 shown is a flowchart that provides one example of the operation of a portion of the visualization service 118.
- the flowchart of FIG. 7 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the visualization service 118.
- the flowchart of FIG. 7 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
- the visualization service 118 can be executed to receive a segmented image 123, one or more associated cell features 133, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113.
- the image coordinates can represent the x and y coordinates of a cell within a segmented image 123.
- the visualization service 118 can be executed to call a function to retrieve a segmented image 123 from the data store 119, the cell type application 113, or another program within the network environment 100.
- the visualization service 118 can be executed to call a function to retrieve one or more cell features 133 and image coordinates assigned to one or more cells from the cell type application 113 and include a file path as an argument to the function.
- the cell type application 113 can automatically send the cell features 133 and image coordinates to the visualization service 118 after assigning cell type signatures 129 to a segmented image 123.
- the visualization service 118 can be executed to identify one or more biomarkers from the cell features 133 within the cells of the segmented image 123.
- the visualization service 118 can be executed to identify one or more biomarkers specified by a user.
- the visualization service 118 can be executed to identify one or more new combinations of biomarkers present in the segmented image 123. For example, cells that have biomarker A and biomarker B could be identified and flagged as having a new combination of biomarkers if they also have biomarker C.
- the visualization service 118 can identify and flag a new combination of biomarkers if the combination occurs greater than a defined threshold within a segmented image 123.
- the defined threshold can be selected by a user. In other examples, the defined threshold can be selected by the cell type application 113 or machine learning model 116.
- the visualization service 118 can be executed to add one or more new cell type signatures 129 based at least in part on the one or more biomarkers identified at block 703.
- the visualization service 118 could identify a biomarker present in unidentified cells and add a new cell type signature 129 based at least in part on the biomarker.
- the visualization service 118 could add the new cell type signature 129 if the number of unidentified cells having the biomarker exceeds a defined threshold.
- the visualization service 118 could add a new cell type signature 129 based at least in part on a new combination of biomarkers identified at block 703.
- the visualization service 118 could add a new cell type signature 129 for cells having biomarkers A, B, and C. In some examples, the visualization service 118 could add the new cell type signature 129 if the number of cells having the combination of biomarkers exceeds a defined threshold. In some examples, the defined threshold can be selected by a user. In other examples, the defined threshold can be selected by the cell type application 113 or the machine learning model 116.
- the visualization service 118 can be executed to validate the one or more new cell type signatures 129 added at block 706.
- the visualization service 118 can be executed to flag new cell type signatures 129 added at block 706, send the new cell type signatures 129 to a user interface 143, and receive user feedback entered in the user interface 143.
- the user feedback can confirm the one or more new cell type signatures 129.
- the user feedback can cause the visualization service 118 to be executed to modify or delete the one or more new cell type signatures 129.
- the visualization service 118 can be executed to send the new cell type signatures 129 to the machine learning model 116 and receive feedback from the machine learning model 116.
- the feedback from the machine learning model 116 can confirm the one or more new cell type signatures 129. In other examples, the feedback from the machine learning model 116 can cause the visualization service 118 to be executed to modify or delete the one or more new cell type signatures 129.
- the process depicted by the flowchart of FIG. 7 can end.
- FIG. 8 shown is a flowchart that provides one example of the operation of a portion of the visualization service 118.
- the flowchart of FIG. 8 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the visualization service 118.
- the flowchart of FIG. 8 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
- the visualization service 118 can be executed to receive a segmented image 123, one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113.
- the image coordinates can represent the x and y coordinates of a cell within a segmented image 123.
- the visualization service 118 can be executed to call a function to retrieve a segmented image 123 and/or a list of ligand-receptor pairs from the data store 119, the cell type application 113, or another program within the network environment 100.
- the visualization service 118 can be executed to call a function to retrieve one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113 and include a file path as an argument to the function.
- the cell type application 113 can automatically send the one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 to the visualization service 118 after assigning cell type signatures 129 to a segmented image 123.
- the visualization service 118 can be executed to identify one or more cells within a defined radius for each cell in the segmented image 123.
- the defined radius can be modified by a user.
- the defined radius can be set by the cell type application 113 or the machine learning model 116.
- the visualization service 118 can be executed to establish one or more connections between cells expressing a ligand gene and other cells within the defined radius expressing a corresponding receptor gene.
- the visualization service 118 can be executed to divide the segmented image 123 into a grid.
- the visualization service 118 can then be executed to calculate a spatial kernal density score for each connection.
- the visualization service 118 can then be executed to identify one or more regions with dense ligandreceptor interactions based at least in part on the spatial kernal density scores.
- the visualization service 118 can be executed to identify one or more clusters of ligand-receptor pairs.
- the visualization service 118 can be executed to apply a Louvain clustering algorithm to the density scores calculated at block 806.
- the visualization service 118 can be executed to send the density scores to the machine learning model 116 and receive clusters from the machine learning model 116.
- FIG. 9 shown is a flowchart that provides one example of the operation of a portion of the visualization service 118.
- the flowchart of FIG. 9 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the visualization service 118.
- the flowchart of FIG. 9 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
- the visualization service 118 can be executed to receive a segmented image 123, one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113.
- the image coordinates can represent the x and y coordinates of a cell within a segmented image 123.
- the visualization service 118 can be executed to call a function to retrieve a segmented image 123 and/or a list of ligand-receptor pairs from the data store 119, the cell type application 113, or another program within the network environment 100.
- the visualization service 118 can be executed to call a function to retrieve one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113 and include a file path as an argument to the function.
- the cell type application 113 can automatically send the one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 to the visualization service 118 after assigning cell type signatures 129 to a segmented image 123.
- the visualization service 118 can be executed to establish one or more connections between cells expressing a ligand gene and other cells within a defined radius expressing a corresponding receptor gene.
- the visualization service 118 can be executed to identify one or more cells within a defined radius for each cell in the segmented image 123.
- the defined radius can be modified by a user.
- the defined radius can be set by the cell type application 113 or the machine learning model 116.
- the visualization service 118 can be executed to call a function to establish a connection between a cell expressing a ligand gene and a cell expressing a corresponding receptor gene within a defined radius and provide data received at block 900 and a defined radius as arguments to the function.
- the visualization service 118 could establish a connection between a cell expressing a CXCL12 gene and a nearby cell expressing a CXCR4 gene.
- the visualization service 118 can be executed to create one or more receptor-ligand modules based at least in part on ligand-receptor connections established at block 903.
- the visualization service 118 can be executed to divide the segmented image 123 into a grid.
- the visualization service 118 can then be executed to calculate a spatial kernal density score for each connection.
- the visualization service 118 can then be executed to identify one or more regions with dense ligand-receptor interactions based at least in part on the spatial kernal density scores.
- the visualization service 118 could create a module defined by CXCL12-CXCR4, CCL14-CCR1 , and CXCL16- CXCR6 pairs.
- the visualization service 118 can be executed to assess patterns based at least in part on the receptor-ligand modules and data identified at block 900.
- the visualization service 118 can be executed to identify one or more clusters of ligand-receptor pairs.
- the visualization service 118 can be executed to apply a Louvain clustering algorithm to the calculated density scores.
- the visualization service 118 can be executed to send the density scores to the machine learning model 116 and receive clusters from the machine learning model 116.
- the visualization service 118 could identify cell type signatures 129 associated with one or more modules identified at block 906.
- the visualization service 118 could identify that a module defined by CXCL12-CXCR4, CXCL16-CXCR6, and CCL14-CCR1 pairs was both a hub sharing spatial distribution with at least three modules in both gland and mucosa tissues and was also composed of antigen-presenting cells along with lymphatic endothelial and immune cell type signatures 129.
- the visualization service 118 can be executed to add one or more cell type signatures 129 based at least in part on the patterns identified at block 909.
- the visualization service 118 can be executed to call a function to add a new cell type signature 129 and provide one or more ligand-receptor pairs as arguments to the function.
- the visualization service 118 could look for cell type signatures 129 that are associated with ligand-receptor modules identified at block 906 and create a new cell type signature 129 based at least in part on the patterns identified at block 909.
- the visualization service 118 could create a new cell type signature 129 for unique receptor-ligand pairs across different tissue sites.
- executable means a program file that is in a form that can ultimately be run by the processor.
- executable programs can be a compiled program that can be translated into machine code in a format that can be loaded into a random access portion of the memory and run by the processor, source code that can be expressed in proper format such as object code that is capable of being loaded into a random access portion of the memory and executed by the processor, or source code that can be interpreted by another executable program to generate instructions in a random access portion of the memory to be executed by the processor.
- An executable program can be stored in any portion or component of the memory, including random access memory (RAM), read-only memory (ROM), hard drive, solid-state drive, Universal Serial Bus (USB) flash drive, memory card, optical disc such as compact disc (CD) or digital versatile disc (DVD), floppy disk, magnetic tape, or other memory components.
- RAM random access memory
- ROM read-only memory
- USB Universal Serial Bus
- CD compact disc
- DVD digital versatile disc
- floppy disk magnetic tape, or other memory components.
- the memory includes both volatile and nonvolatile memory and data storage components. Volatile components are those that do not retain data values upon loss of power. Nonvolatile components are those that retain data upon a loss of power.
- the memory can include random access memory (RAM), read-only memory (ROM), hard disk drives, solid-state drives, USB flash drives, memory cards accessed via a memory card reader, floppy disks accessed via an associated floppy disk drive, optical discs accessed via an optical disc drive, magnetic tapes accessed via an appropriate tape drive, or other memory components, or a combination of any two or more of these memory components.
- the RAM can include static random access memory (SRAM), dynamic random access memory (DRAM), or magnetic random access memory (MRAM) and other such devices.
- the ROM can include a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable readonly memory (EEPROM), or other like memory device.
- each block can represent a module, segment, or portion of code that includes program instructions to implement the specified logical function(s).
- the program instructions can be embodied in the form of source code that includes human-readable statements written in a programming language or machine code that includes numerical instructions recognizable by a suitable execution system such as a processor in a computer system.
- the machine code can be converted from the source code through various processes. For example, the machine code can be generated from the source code with a compiler prior to execution of the corresponding application. As another example, the machine code can be generated from the source code concurrently with execution with an interpreter. Other approaches can also be used.
- each block can represent a circuit or a number of interconnected circuits to implement the specified logical function or functions.
- the flowcharts show a specific order of execution, it is understood that the order of execution can differ from that which is depicted. For example, the order of execution of two or more blocks can be scrambled relative to the order shown. Also, two or more blocks shown in succession can be executed concurrently or with partial concurrence. Further, in some embodiments, one or more of the blocks shown in the flowcharts can be skipped or omitted.
- any number of counters, state variables, warning semaphores, or messages might be added to the logical flow described herein, for purposes of enhanced utility, accounting, performance measurement, or providing troubleshooting aids, etc. It is understood that all such variations are within the scope of the present disclosure.
- any logic or application described herein that includes software or code can be embodied in any non-transitory computer-readable medium for use by or in connection with an instruction execution system such as a processor in a computer system or other system.
- the logic can include statements including instructions and declarations that can be fetched from the computer-readable medium and executed by the instruction execution system .
- a "computer-readable medium" can be any medium that can contain, store, or maintain the logic or application described herein for use by or in connection with the instruction execution system.
- the computer-readable medium can include any one of many physical media such as magnetic, optical, or semiconductor media. More specific examples of a suitable computer-readable medium would include, but are not limited to, magnetic tapes, magnetic floppy diskettes, magnetic hard drives, memory cards, solid-state drives, USB flash drives, or optical discs. Also, the computer-readable medium can be a random access memory (RAM) including static random access memory (SRAM) and dynamic random access memory (DRAM), or magnetic random access memory (MRAM).
- RAM random access memory
- SRAM static random access memory
- DRAM dynamic random access memory
- MRAM magnetic random access memory
- the computer-readable medium can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or other type of memory device.
- ROM read-only memory
- PROM programmable read-only memory
- EPROM erasable programmable read-only memory
- EEPROM electrically erasable programmable read-only memory
- any logic or application described herein can be implemented and structured in a variety of ways.
- one or more applications described can be implemented as modules or components of a single application.
- one or more applications described herein can be executed in shared or separate computing devices or a combination thereof.
- a plurality of the applications described herein can execute in the same computing device, or in multiple computing devices in the same computing environment 103.
- Disjunctive language such as the phrase “at least one of X, Y, or Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., can be either X, Y, or Z, or any combination thereof (e.g., X; Y; Z; X or Y; X or Z; Y or Z; X, Y, or Z; etc.).
- X Y
- Z X or Y
- Y or Z X or Z
- a system comprising: a computing device comprising a processor and a memory; and machine-readable instructions stored in the memory that, when executed by the processor, cause the computing device to at least: receive a segmented image, the segmented image comprising a plurality of segments; identify a plurality of cell features from each of the plurality of segments; calculate a cell type relevance score for each of the plurality of segments based at least in part on the plurality of cell features; and assign a cell type signature for each segment based at least in part on the cell type relevance score.
- Clause 2 The system of clause 1 , wherein the machine-readable instructions which, when executed by the processor, cause the computing device to calculate the cell type relevance score further cause the computing device to at least: determine a weight for each of the plurality of cell features; assign a feature score to each of the plurality of cell features; and perform a weighted linear combination of each of the feature scores to obtain the cell type relevance score.
- Clause 3 The system of clause 1 or 2, wherein each of the plurality of segments represents an individual cell in the segmented image.
- Clause 4 The system of any of clauses 1 -3, wherein the machine-readable instructions, when executed by the processor, further cause the computing device to at least: identify with a machine learning model, one or more clusters of individual segments of the plurality of segments based at least in part on the cell type relevance score for each of the plurality of segments; determine a distribution of cell type relevance scores across the one or more clusters; determine a positive identification threshold based at least in part on the distribution; and determine for individual segments of the plurality of segments a cell identification based at least in part on a comparison of the corresponding cell type relevance score to the positive identification threshold.
- Clause 8 The system of any of clauses 1 -7, wherein the machine-readable instructions which, when executed by the processor, cause the computing device to assign the cell type signature, further cause the computing device to at least: identify a first individual segment having multiple assigned cell type signatures; identify a second individual segment having a single assigned cell type signature; compare the first individual segment with the second individual segment to determine a correct cell type signature; and assign the correct cell type signature to the first individual segment.
- Clause 9 - A method comprising: receiving, by a computing device, a segmented image, the segmented image comprising a plurality of segments; identifying, by the computing device, a plurality of cell features from each of the plurality of segments; calculating, by the computing device, a cell type relevance score for each of the plurality of segments based at least in part on the plurality of cell features; and assigning, by the computing device, a cell type signature for each segment based at least in part on the cell type relevance score.
- Clause 10 The method of clause 9, further comprising: identifying, with a machine learning model, one or more clusters of individual segments of the plurality of segments based at least in part on the cell type relevance score for each of the plurality of segments; determining, by a computing device, a distribution of cell type relevance scores across the one or more clusters; determining, by the computing device, a positive identification threshold based at least in part on the distribution; and determining, by the computing device, a cell identification for individual segments of the plurality of segments based at least in part on a comparison of the corresponding cell type relevance score to the positive identification threshold.
- determining a positive identification threshold further comprises at least: determining, by the computing device, a distribution of cell type relevance scores across the one or more clusters; and using piecewise linear regression, determining, by the computing device, the positive identification threshold based at least in part on the distribution of cell type relevance scores.
- Clause 12 The method of clause 10 or 11 , further comprising at least determining, by the computing device, a positive cell identification for an individual segment of the plurality of segments based at least in part on the corresponding cell type relevance score exceeding the positive identification threshold.
- Clause 13 The method of clause 10 or 11 , further comprising at least determining, by the computing device, a negative cell identification for an individual segment of the plurality of segments based at least in part on the corresponding cell type relevance score failing to exceed the positive identification threshold.
- calculating the cell type relevance score further comprises at least: determining, by the computing device, a weight for each of the plurality of cell features; assigning, by the computing device, a feature score to each of the plurality of cell features; and performing, by the computing device, a weighted linear combination of each of the feature scores to obtain the cell type relevance score.
- assigning the cell type further comprises at least: identifying, by the computing device, a first individual segment having multiple assigned cell type signatures; identifying, by the computing device, a second individual segment having a single assigned cell type signature; comparing, by the computing device, the first individual segment with the second individual segment to determine a correct cell type signature; and assigning, by the computing device, the correct cell type signature to the first individual segment.
- comparing the first individual segment with the second individual segment further comprises: using K-nearest neighbor, determining, by the computing device, a similarity between the first individual segment and the second individual segment; and determining the correct cell type signature based at least in part on the similarity.
- Clause 17 -A system comprising: a computing device comprising a processor and a memory; and machine-readable instructions stored in the memory that, when executed by the processor, cause the computing device to at least: receive a segmented image, individual segments of the segmented image corresponding to individual cells; receive a feature matrix associated with a cell type signature; identify a plurality of cell features from the individual segments; compare the plurality of cell features with the feature matrix to generate a cell type relevance score for individual segments; and assign a cell type signature for individual segments based at least in part on the cell type relevance score.
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- General Health & Medical Sciences (AREA)
- Public Health (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Epidemiology (AREA)
- Databases & Information Systems (AREA)
- Primary Health Care (AREA)
- Software Systems (AREA)
- Life Sciences & Earth Sciences (AREA)
- Biomedical Technology (AREA)
- General Physics & Mathematics (AREA)
- Artificial Intelligence (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Evolutionary Computation (AREA)
- Computing Systems (AREA)
- Pathology (AREA)
- Multimedia (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Biotechnology (AREA)
- Radiology & Medical Imaging (AREA)
- Biophysics (AREA)
- Molecular Biology (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Nuclear Medicine, Radiotherapy & Molecular Imaging (AREA)
- Physiology (AREA)
- General Engineering & Computer Science (AREA)
- Bioethics (AREA)
- Mathematical Physics (AREA)
- Probability & Statistics with Applications (AREA)
- Image Analysis (AREA)
Abstract
Disclosed are various embodiments for analyzing multiplexed and multimodal imaging data. First, a segmented image is received. Then, a plurality of cell features are identified from each of a plurality of segments within the segmented image. Next, a cell type relevance score is calculated for each of the plurality of segments based at least in part on the plurality of cell features. Later, a cell type signature is assigned for each segment based at least in part on the cell type relevance score.
Description
TITLE: ANALYZING MULTIPLEXED AND MULTIMODAL IMAGING DATA
Inventors: Jinze Liu, Katarzyna Marta Tyc, Khoa Le Anh Huynh, Kevin Matthew Byrd, and Bruno F. Matuck
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to, and the benefit of, U.S. Provisional Patent Application No. 63/640,148, entitled “Analyzing Multiplexed and Multimodal Imaging Data” and filed on April 29, 2024, which is incorporated by reference as if set forth herein in its entirety.
BACKGROUND
[0002] The development of single-cell and spatial multiomics has revolutionized biology and biomedical research by enabling analysis at near single molecular resolution. Single-cell and spatial multiomics can allow for the integrated measurement of various cell features while preserving their spatial context within biological samples such as human tissues, human cells, and human biofluids. While the promise is great for singlecell and spatial approaches in science, many challenges remain. Some of the problems related to manual feature annotation include subjectivity and variability, time-intensive, limited scalability, inter-observer discrepancies, annotation bias, difficulty in integrating multimodal data, and other reproducibility concerns.
BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Many aspects of the present disclosure can be better understood with reference to the following drawings. The components in the drawings are not necessarily to scale, with emphasis instead being placed upon clearly illustrating the principles of the disclosure. Moreover, in the drawings, like reference numerals designate corresponding parts throughout the several views.
[0004] FIG. 1 is a drawing of a network environment according to various embodiments of the present disclosure.
[0005] FIGS. 2A and 2B show a flowchart illustrating one example of functionality implemented as portions of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
[0006] FIG. 3 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
[0007] FIG. 4 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
[0008] FIG. 5 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
[0009] FIG. 6 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
[0010] FIG. 7 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
[0011] FIG. 8 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
[0012] FIG. 9 shows a flowchart illustrating one example of functionality implemented as a portion of an application executed in a computing environment in the network environment of FIG. 1 according to various embodiments of the present disclosure.
DETAILED DESCRIPTION
[0013] Disclosed are various approaches for a computational algorithm and associated software application that can automatically determine cell type identities, cell states, and other cell and molecular features in highly multiplexed and multimodal imaging data. The development of single-cell and spatial multiomics has revolutionized biology and biomedical research by enabling analysis at near single molecular resolution, including transcriptom ics, genomics, proteomics, metabolomics, epigenomics, metagenomics, lipidomics, glycomics, microbiomics, and interactomics. Single-cell and spatial multiomics can allow for the integrated measurement of various cell features while preserving their spatial context within biological samples such as human tissues, human
cells, and human biofluids. Whilst biological samples may or may not also include individual microbes or microbes within complex communities.
[0014] There are times when these assays can focus on model or non-traditional model organisms such as animals, plants, and/or microbes, measuring samples in two, three, or four dimensions (e.g., time). There are times in which these assays can combine widely used histology approaches, such as staining, in combination with other imaging modalities to measure relationships between molecular composition and spatial organization within tissues, offering insights into cellular interactions, signaling pathways, and the heterogeneity of biological systems.
[0015] These technologies can offer insights into development, disease mechanisms, and the intricate interplay of cells within their microenvironments. In the near future, many fields of research can benefit from these technologies, such as cancer research, neuroscience, developmental biology, immunology, stem cell research, autoimmune and infectious diseases, aging, regenerative medicine, drug development, ecology and environmental sciences, and other biotechnology. While the promise is great for singlecell and spatial approaches in science, many challenges remain related to data analysis complexity, current spatial resolution limitations, cellular context preservation, standardization, scalability, reproducibility, cost, accessibility, and multimodal integration of datasets. One of the most significant rate-limiting steps in single-cell and spatial assays can be the analysis, and these assays can continue to scale in terms of size. The majority of technologies in single-cell spatial biology are currently focused on multimodal imaging, and the application of image-based technologies, with or without the use of artificial
intelligence assisted output, that introduce technical hurdles not encountered in singlecell sequencing alone.
[0016] While not limited to this example, manual examination, analysis, and feature extraction is a last resort for whole tissue, cell (including cell type, state, patterning, and/or motifs), and molecular, collectively multi-layer feature annotation of highly multiplexed imaging data. Some of the problems related to manual feature annotation include subjectivity and variability, time-intensive, limited scalability, inter-observer discrepancies, annotation bias, difficulty in integrating multimodal data, and other reproducibility concerns. Automated cell-type and cell-state identification methods could be asked to handle the following challenges present across assays and modalities, including erroneous measurements stemming from signal intensity quantifications and challenges associated with boundary segmentation to quantify features accurately. Other challenges to be addressed for single-cell and spatial biology include the number of cells captured in a single experiment, which is anticipated to surpass one million cells in the coming year. This represents a magnitude that is a hundredfold greater than what the current algorithms can be equipped to handle.
[0017] Accordingly, the present disclosure provides for solutions capable of analyzing highly complex data sets (images), identifying data associated with different cell types and states, and providing annotation or metadata associated with the different cell types and states. The present disclosure provides for a computational algorithm and associated software application that can automatically determine cell type identities, cell states, and other cell and molecular features in highly multiplexed and multimodal imaging data. It is an automated, scalable, and platform -agnostic computer-implemented method (or
algorithm) which can be deployed as a web-based application (or software). The application can classify large quantities of spatially resolved cells from multiplexed imaging data without reliance on pre-labeled training datasets. The application can employ a hierarchical, multi-level classification framework that incorporates both broad and fine-grained ontologies, allowing for scalable assignment of structural, immune, and specialized subtypes. The application begins with predefined “cell recipes” (pre-defined set of markers used to specify cell identities and states by combining proteins or transcripts) and applies cluster-aware micro segmentation to calculate relevance scores for each marker in the context of local cellular neighborhoods. Thresholds for marker positivity can be learned dynamically using segmental regression, enabling the detection of both high-confidence and low-abundance or transitional cell states. In some examples, a user can modify the threshold for marker positivity of one or more cell types, markers, genes, proteins, etc. Classification outputs can be refined using k-nearest-neighbor- based deconvolution, local enrichment scoring, and outlier filtering, resulting in accurate cell identity assignments that are robust to noise and applicable across tissue types, disease conditions, and imaging modalities. The software inputs are feature matrix and multi-tiered cell type signature matrix, where each cell type is required to be positive in the individual marker/gene. The software output is the assigned cell type identity for each cell.
[0018] Linking the algorithm and the software provides a visualization tool that allows for, but is not limited to, 1 ) boundary identification from the associated feature input such as cells and microbes, 2) spatial neighborhood relationships, 3) cell motif analysis, and 4) the mapping of identified cells and microbes in their relevant feature space and their
exact locations. The visualization tool provides a user interface that allows for interactive visualization and analysis of spatial cell type data including spatial plots and uniform manifold approximation and projection (UMAP) visuals featuring annotations, marker expression thresholds, and weighted cell type calculations. Users can also modify these thresholds for one or more cell types, markers, genes, proteins, etc. within the visualization tool and run the cell type application again with the modified thresholds. Users can navigate through complex tissue maps in real-time, including zooming, panning, and dynamic toggling between classification tiers. Additionally, users can overlay single or multiple markers from protein or transcriptom ic data and examine cell identity, expression intensity, spatial coordinates, and relationships within the broader tissue context. The visualization tool can allow users to filter cell types, cell states, markers, or a combination of any thereof. Users can also access color annotations, spatial neighborhood connections between cell types across the whole tissue or a user-defined region of interest (ROI), and statistical summaries of marker expression and cell type composition within regions to identify spatial autocorrelation. Additional tools can include annotated mean heatmaps, Voronoi plots, Delaunay triangulation, and proportions of cell types and cell state markers. Users can also customize cell type annotations with user- defined colors and select specific tissues or regions for focused visualization. The generation of heatmaps that display cell type distributions based at least in part on marker expression can be tailored by the user to reflect either cell type-specific markers or cell state markers. The visualization tool can also facilitate the exploration of cell states, providing insights into the proportions of various cell types and states across tissues or regions.
[0019] Laborious and time consuming preprocessing steps, such as cell segmentation, can be combined in the present disclosure with whole-slide cell identification and, in some examples, metacellular annotation (such as identification of multicellular arrangements, or “cell neighborhoods”). This integration can collapse what were previously separate workflows into a single, streamlined process, which can allow for faster analysis, reduce redundancy, and provide an interactive environment where users can check and refine both cell type annotations and segmentation borders (for example, with human-in-the-loop validation).
[0020] In the following discussion, a general description of the system and its components is provided, followed by a discussion of the operation of the same. Although the following discussion provides illustrative examples of the operation of various components of the present disclosure, the use of the following illustrative examples does not exclude other implementations that are consistent with the principles disclosed by the following illustrative examples.
[0021] By leveraging well-established principles from diverse fields such as histology, pathology, cell biology, genomics, biotechnology, bioinformatics, microscopy, immunology, statistics, machine learning, and engineering, the present disclosure can offer an innovative framework for threshold-based assignment of cell types and cell states from multiplexed and multimodal imaging data.
[0022] In an exemplary embodiment of the present disclosure and the present disclosure’s utility, the present disclosure can systematically and statistically enhance setting marker positivity thresholds, which can ensure a robust cell-type assignment. The accuracy of the present disclosure can be further augmented by capitalizing on the
importance assigned to each marker within a multi-tiered cell-type taxonomy. The present disclosure can take two inputs: a matrix of marker intensities in each column, cells in rows, and the cell type signature matrix with cell types in rows and markers in columns (1 indicating marker positivity for that cell type, 0 otherwise). The present disclosure can calculate Cell Type Relevance scores (CTR), a score that measures the relevance of a molecular profile of a cell to a particular cell type. The higher the CTR score, the more substantial the evidence that the cell represents a given cell type. The present disclosure aims to identify salient changes in cell type relevance scores exhibited by fine-grained subpopulations, where the sharp increase of cell type relevance scores between subpopulations is assumed to indicate the emergence of corresponding cell types. Once assigned, cell identifies can be further assessed for cell state changes within the tiered cell identity assignments, including high-level cell assignments (e.g., epithelia versus immune cell) or detailed assignments (e.g., oral keratinocytes versus regulatory T cell).
[0023] Benchmarked datasets have been used for current single modality testing, in which a trained pathologist established cell-type assignments. Comparative evaluations have been conducted against CELESTA and SCINA tools, and the Seurat version 5 pipeline. The present disclosure consistently demonstrated superior performance, exhibiting the highest agreement with the reference across datasets and averaging a significant 10% accuracy gain over alternative methods. Notably, challenges were observed in CELESTA and SCINA, which are partially attributable to inaccurate assumptions regarding marker distribution. Furthermore, all methods, except the present disclosure, encountered difficulties with the substantial ~2.5 million-cell intestine dataset.
The present disclosure’s robustness and efficacy can make it a highly reliable solution across varying dataset sizes and complexities.
[0024] With reference to FIG. 1 , shown is a network environment 100 according to various embodiments. The network environment 100 can include a computing environment 103 and a client device 106, which can be in data communication with each other via a network 109.
[0025] The network 109 can include wide area networks (WANs), local area networks (LANs), personal area networks (PANs), or a combination thereof. These networks can include wired or wireless components or a combination thereof. Wired networks can include Ethernet networks, cable networks, fiber optic networks, and telephone networks such as dial-up, digital subscriber line (DSL), and integrated services digital network (ISDN) networks. Wireless networks can include cellular networks, satellite networks, Institute of Electrical and Electronic Engineers (IEEE) 802.11 wireless networks (/.e., WIFI®), BLUETOOTH® networks, microwave transmission networks, as well as other networks relying on radio broadcasts. The network 109 can also include a combination of two or more networks 109. Examples of networks 109 can include the Internet, intranets, extranets, virtual private networks (VPNs), and similar networks.
[0026] The computing environment 103 can include one or more computing devices that include a processor, a memory, and/or a network interface. For example, the computing devices can be configured to perform computations on behalf of other computing devices or applications. As another example, such computing devices can host and/or provide content to other computing devices in response to requests for content.
[0027] Moreover, the computing environment 103 can employ a plurality of computing devices that can be arranged in one or more server banks or computer banks or other arrangements. Such computing devices can be located in a single installation or can be distributed among many different geographical locations. For example, the computing environment 103 can include a plurality of computing devices that together can include a hosted computing resource, a grid computing resource or any other distributed computing arrangement. In some cases, the computing environment 103 can correspond to an elastic computing resource where the allotted capacity of processing, network, storage, or other computing-related resources can vary over time.
[0028] Various applications or other functionality can be executed in the computing environment 103. The components executed on the computing environment 103 include cell type application 113, a machine learning model 116, a visualization service 118, and other applications, services, processes, systems, engines, or functionality not discussed in detail herein.
[0029] The cell type application 113 can be executed to identify cell type, cell state, clusters of cells, and other various information from matrix representations of features collected from biological samples, such as, for example, images. The cell type application 113 is an algorithm using feature matrices 126 and cell type signatures 129 to identify various cell types from an image. The cell type application 113 weighs the importance of the extracted features 133 in order to determine which cell type applies to the identified cell. In some examples, the cell type application 113 can identify more cell types than those provided in the signature matrix and determine the respective weights of those new cell types as well. When a cell has been identified as positive for more than one cell type,
the cell type application 113 can identify pure positive cells having only one cell type and compare the mixed cell having multiple cell types to the pure cell in order to determine the correct cell type. In some embodiments, the cell type application 113 can further communicate with a machine learning model 116 to determine various aspects throughout the process.
[0030] The machine learning model 116 can be executed to perform segmentation of various images to identify cell boundaries. Additionally, the machine learning model 116 can be executed to receive data from the cell type application 113 such as, for example, cell type relevance scores for individual segments and identify micro clusters or other micro architectural features of the cells. In some examples, the machine learning model 116 can be used to identify the cell features from the images.
[0031] Also, various data is stored in a data store 119 that is accessible to the computing environment 103. The data store 119 can be representative of a plurality of data stores 119, which can include relational databases or non-relational databases such as object-oriented databases, hierarchical databases, hash tables or similar key-value data stores, as well as other data storage applications or data structures. Moreover, combinations of these databases, data storage applications, and/or data structures may be used together to provide a single, logical, data store. The data stored in the data store 119 is associated with the operation of the various applications or functional entities described below. This data can include images 123, feature matrices 126, cell type signatures 129, cell features 133, signature matrices 134, and potentially other data.
[0032] The images 123 can represent image files from a variety of modalities, such as cellular imaging (e.g., phase-contrast microscopy, fluorescent imaging, brightfield
microscopy, confocal microscopy, time-lapse microscopy, etc.), tissue imaging (e.g., multiplex colorimetric immunohistochemistry (rnCIHC), multiplex immunofluorescence (mIF), cyclic immunofluorescence (CycIF), multiplexed ion beam imaging (MIBI), codetection by indexing (CODEX)/Phenocycler (PCF), digital spatial profiling (DSP), etc., or other form of imaging. The images 123 can represent segmented images, or images which have been preprocessed and segmented into various segments corresponding to individual cells, cell clusters, tissues, or other features. In some examples, the images 123 are representative of any matrix representations of features collected from biological samples.
[0033] The feature matrices 126 can represent matrices of cell features 133 and associated intensities/count in each column and cell type signatures 129 in rows, as well as cell type signature matrices with cell type signatures 129 in columns and cell features 133 in rows. For example, feature matrices 126 could include z-normalized intensity values indicating an intensity level of specific markers within a set of cells. In another example, feature matrices 126 could include log-normalized counts of transcripts for a variety of genes within a set of cells. In some embodiments the feature matrices 126 can be generated by the machine learning model 116 from segmented images 123 that could be trained at least in part on existing cell type segmentation and/or morphological features. In some embodiments, the feature matrices 126 can further include data regarding the respective importance, or weights, associated with each cell feature 133.
[0034] The cell type signatures 129 can represent various broad cell types (e.g., epithelial, immune cells, stem cells, blood cells, bone cells, muscle cells, gametes, neurons, etc.), specific cell types (e.g., oral keratinocytes, regulatory T cell, erythrocytes,
odontoblasts, etc.), cell states (e.g., activated, resting, diseased, etc. , or other cellular or molecular features which can be identified by the cell type application 113. The cell type signatures 129 can be stored in the data store 119 and associated with one or more cell features 133.
[0035] The cell features 133 can represent various observable characteristics of a particular cell type or cell state which can be identified from the images 123. In some examples, cell features 133 can include cellular morphology, nuclei characteristics, cytoplasm, cellular borders, inclusion bodies and vacuoles, cellular arrangement, mitotic activity, cellular diversity, etc. as well as associated and unassociated biomolecules and markers, including DNA, RNA, proteins, metabolites, etc. The data store 119 can also contain cell features 133 that did not originate from the segmented image 123.
[0036] The signature matrices 134 can be representative of matrices having various cell types listed in rows and various cell markers, or cell features 133, in columns. In some examples, the signature matrices 134 can include a value of “1 ,” indicating marker positivity for that cell type, or a value of “0” to indicate otherwise.
[0037] The client device 106 is representative of a plurality of client devices that can be coupled to the network 109. The client device 106 can include a processor-based system such as a computer system. Such a computer system can be embodied in the form of a personal computer (e.g., a desktop computer, a laptop computer, or similar device), a mobile computing device (e.g., personal digital assistants, cellular telephones, smartphones, web pads, tablet computer systems, music players, portable game consoles, electronic book readers, and similar devices), media playback devices (e.g., media streaming devices, BluRay® players, digital video disc (DVD) players, set-top
boxes, and similar devices), a videogame console, or other devices with like capability.
The client device 106 can include one or more displays 136, such as liquid crystal displays
(LCDs), gas plasma-based flat panel displays, organic light emitting diode (OLED) displays, electrophoretic ink (“E-ink”) displays, projectors, or other types of display devices. In some instances, the display 136 can be a component of the client device 106 or can be connected to the client device 106 through a wired or wireless connection.
[0038] The client device 106 can be configured to execute various applications such as a client application 139 or other applications. The client application 139 can be executed in a client device 106 to access network content served up by the computing environment 103 or other servers, thereby rendering a user interface 143 on the display 136. To this end, the client application 139 can include a browser, a dedicated application, or other executable, and the user interface 143 can include a network page, an application screen, or other user mechanism for obtaining user input. The client device 106 can be configured to execute applications beyond the client application 139 such as email applications, social networking applications, word processors, spreadsheets, or other applications.
[0039] Next, a general description of the operation of the various components of the network environment 100 is provided. To begin, a user can initiate a process through the client application 139 on a client device 106. In some examples, the user can provide images 123 via the client device 106. The images 123 can be segmented by a machine learning model 116 to produce segmented images 123, which can then be used by the cell type application 113 for further processing. The cell type application 113 can identify various cell features 133 obtained from within the segmented images 123. Having
collected the cell features 133, the cell type application 113 can use the feature matrices 126 (consisting of cell features 133 and cell type signatures 129) to assign a cell type relevance score to each cell based at least in part on the data in the feature matrices 126 and based at least in part on the cell features 133 identified from the images 123. In some examples, cell type relevance scores can be assigned by multiplying the features matrices 126 by the signature matrices 134. The cell type application 113 can use clustering methods with the machine learning model 116 to subcluster the feature matrices 126. After that, the cell type application 113 can use the subclusters to determine a distribution of cell type relevance scores. Piece-wise linear regression can be applied to detect points indicating drastic changes along the growing trend of cell type relevance scores. Those points divide the subclusters into a low-relevance group and a high- relevance group. A threshold is then determined by minimizing the misplacement of cells between low and high relevance groups. Based at least in part on the positive identification threshold and the cell type relevance score, the cell type application 113 can determine whether to assign a cell type provided in cell type signatures 129 to a particular cell. In some examples, the threshold can be modified by a user for one or more cell type signatures 129, genes, markers, or proteins. In some examples, the cell type application 113 can be executed to call a function to assign a cell type provided in cell type signatures 129 and provide the threshold set by a user and the cell type relevance scores as arguments to the function. Finally, for cells or segments which are positive for multiple cell type signatures 129, the cell type application 113 can identify “pure cells” having only one assigned cell type signature 129, and “mixed cells” having multiple assigned cell types 129. The cell type application 113 can compare the “pure cells” to the “mixed cells”
to determine a correct cell type signature 129 based at least in part on the similarity. In some examples, the cell type application 113 can identify “unknown cells” that are not classified as positive for any cell type signature 129.
[0040] Referring next to FIG. 2, shown is a flowchart that provides one example of the operation of a portion of the cell type application 113. The flowchart of FIG. 2 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the cell type application 113. As an alternative, the flowchart of FIG. 2 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
[0041] Beginning with block 200 of FIG. 2A, the cell type application 113 can be executed to receive a segmented image 123. In some examples, the cell type application 113 can receive the segmented image 123 from a client application 139 on a client device 106. For example, the client application 139 can obtain the segmented image 123 from a data store 119 in response to a user interaction. The client application 139 can then send the segmented image 123 to the cell type application 113 to initiate the processing of the image 123. In some examples, the cell type application 113 receives the segmented image 123 from the machine learning model 116, from the data store 119, or from another system, service, or application in the network environment 100.
[0042] Next, at block 203, the cell type application 113 can be executed to identify one or more cell features 133 from the segmented image 123. In some examples, the cell type application 113 can use imaging processing techniques to extract various data from the segmented images 123 received at block 200. For example, the cell type application 113 can determine the fluorescence of a particular segment in order to determine which
protein is being expressed by the corresponding cell and how much of that protein. In some examples, the cell type application 113 can identify a plurality of cell features 133 from the segmented image 123.
[0043] At block 206, the cell type application 113 can be executed to determine weights for the cell features 133. The cell type application 113 can determine a respective weight for each of the cell features 133 identified at block 203. The weights determined by the cell type application 113 can be representative of the importance of the cell feature 133 in determining the cell type as provided in cell type signatures 129. In some examples, the cell type application 113 can determine a respective weight for a cell feature 133 based at least in part on a feature matrix 126. For example, the cell type application 113 can compare the cell features 133 identified in a particular segment of the segmented image 123 to cell features 133 associated with a particular cell type provided in cell type signatures 129 in a feature matrix 126 and use the stored importances or weights from the feature matrix 126 to determine the weight of the respective cell feature 133. In some examples, the cell type application 113 can determine a respective weight for a cell feature 133 based at least in part on other data.
[0044] At block 209, the cell type application 113 can be executed to assign feature scores to cell features 133. In some examples, the cell type application 113 can assign a feature score to each respective cell feature 133 based at least in part on the weights determined at block 206. In some examples, the cell type application 113 can use a feature matrix 126 from the data store 119 to determine a feature score for each respective cell feature 133 identified at block 203. The feature score given to each cell feature 133 by the cell type application 113 can represent a score of the similarity between
the cell features 133 identified at block 203 and the cell features 133 listed in the feature matrix 126 and associated with a cell type signature 129. For example, if a feature matrix 126 associates a particular cell type signature 129 with a particular cell feature 133 at a high occurrence, but the particular cell feature 133 identified at block 213 has a low occurrence for the segment in question, the cell type application 113 can assign a low feature score for that particular cell type signature 129.
[0045] Next, at block 213, the cell type application 113 can be executed to calculate the cell type relevance (CTR) score. In some examples, the cell type application 113 can calculate a CTR score for each cell of the segmented image received at block 200. The cell type application 113 can calculate the CTR score based at least in part on the cell features 133 identified at block 203, the weights determined at block 206, and/or the feature scores assigned at block 209. In some examples, the cell type application 113 calculates the CTR score by performing a weighted linear combination of the feature scores assigned at block 209 using the weights determined at block 206. In some examples, the cell type application 113 can calculate a CTR score for each cell type stored in cell type signatures 129. In some examples, the cell type application 113 can calculate the CTR scores by multiplying a matrix of the feature scores determined at block 209 with a signature matrix 134.
[0046] At block 216, the cell type application 113 can be executed to assign a cell type signature 129. In some examples, the cell type application 113 can assign a cell type signature 129 to each segment of the segmented images 123 based at least in part on the CTR score calculated at block 213. For example, the cell type application 113 can assign a cell type signature 129 which corresponds to the highest CTR score associated
with the segment. In some other examples, the cell type application 113 can indicate one or more cell types not present in cell type signatures 129 as a potential new cell type and flag the potential new cell type for review.
[0047] Moving to block 219 in FIG. 2B, the cell type application 113 can be executed to identify clusters using machine learning (ML) or other methods. In some examples, the cell type application 113 can use the machine learning model 116 to identify one or more clusters of cell type signatures 129 in the segmented image 123. For example, the cell type application 113 can send the segmented image 123 with the assigned CTR scores to the machine learning model 116 with an instruction to identify clusters based at least in part on the CTR scores. In some examples, the cell type application 113 can receive the identified clusters from the machine learning model 116.
[0048] Then, at block 223, the cell type application 113 can be executed to determine the CTR score distribution. In some examples, the cell type application 113 can determine the CTR score distribution across the clusters identified at block 219. In some embodiments, the cell type application 113 can determine the CTR score distribution for each cluster and/or across all clusters in the entire feature matrix 126.
[0049] At block 226, the cell type application 113 can be executed to determine a positive identification threshold. The positive identification threshold can represent a threshold which can be used to determine the difference between a positively identified cell and the background of cell features 133 collected from, but not limited to, the image 123. For example, the positive identification threshold can be a minimum CTR score which is required to assign a cell type signature 129 to a cell. The cell type application 113 can determine the positive identification threshold based at least in part on the CTR
score distribution determined at block 223. In some examples, the cell type application 113 can determine the positive identification threshold using piecewise linear regression. For example, the cell type application 113 could calculate and store as a vector the median CTR scores across all clusters identified at block 219 for any given cell type and then fit a segmented regression model to identify breakpoints that divide the data into distinct linear segments. In some examples, a maximum of three breakpoints can be identified. In some examples, the highest identified breakpoint is the positive identification threshold. In some examples, the positive identification threshold can be set by a user through the user interface 143 for one or more cell types, markers, genes, proteins, or a combination of any thereof. In some examples, the cell type application 113 can be executed to call a function to determine a positive identification threshold and provide the threshold set by the user as an argument to the function.
[0050] At block 229, the cell type application 113 can be executed to determine a cell identification. The cell type application 113 can determine a positive or a negative cell identification (e.g., a measure of whether a cell type signature 129 can be assigned to a segment) based at least in part on the positive identification threshold determined at block 226. In some embodiments, the cell type application 113 can determine a positive cell identification for an individual segment based at least in part on the corresponding CTR score exceeding the positive identification threshold determined at block 226. In some examples, the cell type application 113 can determine a negative cell identification for an individual segment based at least in part on the corresponding CTR score failing to exceed the positive identification threshold determined at block 226.
[0051] Next, at block 233, the cell type application 113 can be executed to identify “mixed” cells. “Mixed” cells can represent cells of the segmented image 123, or other data source, which have a positive cell identification determined at block 229 for more than one cell type signature 129, or which is assigned multiple cell type signatures 129. For example, if a particular segment representing a cell has more than one CTR score exceeding the respective positive identification threshold, the cell type application 113 can determine that the particular segment represents a “mixed” cell. In some examples, the cell type application 113 can identify all of the “mixed” cells which occur in an image 123.
[0052] At block 236, the cell type application 113 can be executed to identify “pure” cells. “Pure” cells can represent cells of the segmented image 123, or other data source, which have a positive cell identification determined at block 229 for only a single cell type signature 129, or which only has a single assigned cell type signature 129. For example, if a particular segment representing a cell has only one CTR score exceeding the respective positive identification threshold, the cell type application 113 can determine that the particular segment represents a “pure” cell. In some examples, the cell type application 113 can identify all of the “pure” cells which occur in an image 123.
[0053] Finally, at block 239, the cell type application 113 can be executed to assign the correct cell type signature 129. The cell type application 113 can assign the correct cell type signatures 129 to “mixed” cells identified at block 233 based at least in part on a comparison of the “mixed” cell to the “pure” cells identified at block 236. For example, the cell type application 113 can take each “mixed” cell identified at block 233 and determine which cell type signatures 129 are positively identified or assigned to the cell. Then, the
cell type application 113 can compare the “mixed” cell to a “pure” cell of each positively identified cell type 129 which applies to the “mixed” cell. Using a similarity calculation under relevant cell type signatures 129 and using a classification algorithm, such as, for example, k-nearest neighbor (KNN), the cell type application 113 can determine which of the “pure” cell types 129 is most similar to the “mixed” cell and assign the correct cell type signature 129 accordingly. In some examples, the machine learning model 116 can be executed to validate the correct cell type signatures 129 assigned by the cell type application 113. In some examples, a user could validate the correct cell type signatures 129 assigned by the cell type application 113. After block 239, the process depicted by the flowchart of FIG. 2 can come to an end.
[0054] Referring now to FIG. 3, shown is a flowchart that provides one example of the operation of a portion of the cell type application 113. The flowchart of FIG. 3 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the cell type application 113. As an alternative, the flowchart of FIG. 3 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
[0055] At block 300, the cell type application 113 can be executed to assign one or more cell type signatures 129 to one or more cells within a segmented image 123. The cell type application 113 can identify various cell features 133 obtained from within, but not limited to, the segmented images 123. Having collected the cell features 133, the cell type application 113 can use the feature matrices 126 (consisting of cell features 133 and cell type signatures 129) to assign a cell type relevance score to each cell based at least in part on the data in the feature matrices 126 and based at least in part on the cell
features 133 identified from, but not limited to, the images 123. In some examples, cell type relevance scores can be assigned by multiplying the features matrices 126 by the signature matrices 134. The cell type application 113 can use clustering methods with the machine learning model 116 to subcluster the feature matrices 126. After that, the cell type application 113 can use the subclusters to determine a distribution of cell type relevance scores. Piece-wise linear regression can be applied to detect points indicating drastic changes along the growing trend of cell type relevance scores. Those points divide the subclusters into a low-relevance group and a high-relevance group. A threshold is then determined by minimizing the misplacement of cells between low and high relevance groups. In some examples, the threshold can be set by a user. Based at least in part on the positive identification threshold and the cell type relevance score, the cell type application 113 can determine whether to assign a cell type provided in cell type signatures 129 to a particular cell. Finally, for cells or segments which are positive for multiple cell type signatures 129, the cell type application 113 can identify “pure cells” having only one assigned cell type signature 129, and “mixed cells” having multiple assigned cell types 129. The cell type application 113 can compare the “pure cells” to the “mixed cells” to determine a correct cell type signature 129 based at least in part on the similarity. In some examples, the cell type application 113 can identify “unknown cells” or potential new cell types that are not classified as positive for any cell type signature 129. [0056] Next, at block 303, the cell type application 113 can be executed to identify one or more tiers for the cell type signatures 129 assigned at block 300. In some examples, the tiers can be defined in the cell type signatures 129. In some examples, the tiers can be defined by a user. In some examples, the cell type application 113 can be
executed to receive defined tiers from another application within the network environment 100. In some examples, the cell type application 113 can be executed to call a function to identify one or tiers and provide the cell type signatures 129 and the cell type signatures 129 or defined tiers as arguments to the function. In some examples, a tier can be defined as broad categorizations of cell type signatures 129. For example, a Tier zero (0) cell type signature 129 could be defined as structural and immune cell type signatures 129. In some examples, a tier can further divide an existing tier into more specific cell types. For example, a Tier one (1 ) cell type signature 129 could further define structural cell type signatures 129 as epithelial cells, vasculature, muscle, etc. and define immune cell type signatures 129 as fibroblasts, immune myeloid cells, lymphoid-derived cells, etc.
[0057] At block 306, the cell type application 113 can be executed to annotate the cell type signatures 129 based at least in part on tiers identified at block 303. The cell type application 113 can be executed to call a function to label the cell type signatures 129 assigned at block 300 with one or more respective tiers identified at block 303 and provide the cell type signatures 129 and the tiers as arguments to the function. For example, the cell type application 113 could label a fibroblast identified at block 300 as a Tier zero (0) immune cell type signature 129 and a Tier one (1 ) fibroblast based at least in part on the tiers identified at block 303. In some examples, the cell type application 113 can be further executed to send the annotations to the visualization service 118 for display to a user on a client device 106. After block 306, the process depicted by the flowchart of FIG. 3 can come to an end.
[0058] Referring now to FIG. 4, shown is a flowchart that provides one example of the operation of a portion of the machine learning model 116. The flowchart of FIG. 4 provides
merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the machine learning model 116. As an alternative, the flowchart of FIG. 4 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
[0059] At block 400, the machine learning model 116 can be executed to receive a segmented image 123, associated one or more cell type signatures 129, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113. In some examples, the image coordinates can represent the x and y coordinates of a cell within the segmented image 123. In some examples, the machine learning model 116 can be executed to call a function to retrieve a segmented image 123, associated one or more cell type signatures 129, and image coordinates assigned to one or more cells from the cell type application 113 and include a file path as an argument to the function. In some examples, the cell type application 113 can automatically send the segmented image 123, associated cell type signatures 129, and image coordinates to the machine learning model 116 after assigning cell type signatures 129 to the segmented image 123.
[0060] Next, at block 403, the machine learning model 116 can be executed to partition a segmented image 123 into one or more subsets. In some examples, the machine learning model 116 can be executed to partition the segmented image 123 into one or more vertical subsets based at least in part on defined x coordinate boundaries. In some examples, the machine learning model 116 can be executed to partition the segmented image 123 into one or more horizontal subsets based at least in part on defined y coordinate boundaries. In some examples, the machine learning model 116 can
be executed to partition the segmented image 123 into one or more square grid subsets by dividing coordinates along both axes. In some examples, the x coordinate boundaries and/or y coordinate boundaries are defined by a user. In some examples, the partitions can be recursively refined to ensure the partitions contain a minimum and/or maximum number of cells. In some examples, the partitions exceeding a maximum cell count can be subdivided and partitions containing a number of cells below the minimum cell count can be merged.
[0061] At block 406, the machine learning model 116 can be executed to construct one or more sub-graphs based at least in part on the subsets created at block 403. In some examples, the machine learning model 116 can be executed to call a function to construct one or more sub-graphs and provide the subsets created at block 403 as arguments to the function. In some examples, the nodes of the sub-graphs can represent cells having one or more cell type signatures 129 and the edges can capture one or more spatial relationships based at least in part on the k-nearest neighbors of the cells.
[0062] Then, at block 409, the machine learning model 116 can be executed to assign one or more cell neighborhoods. In some examples, the machine learning model 116 can be executed to call a minimum cut-based loss function and provide the nodes and the edges from the one or more sub-graphs generated at block 406. In some examples, the one or more sub-graphs can be randomly shuffled in every epoch. In some examples, the one or more sub-graphs can be processed in a graph neural network (GNN). In some examples, the one or more sub-graphs can be treated as independent training samples and can be processed in mini batches with a gradient accumulation technique.
[0063] At block 412, the machine learning model 116 can be executed to perform one or more consensus-based tissue cell neighborhood assignments. In some examples, the machine learning model 116 can be executed to call a function to assign consensusbased tissue cell neighborhoods and provide the cell neighborhoods from block 409 as arguments to the function. In some examples, the one or more consensus-based tissue cell neighborhood assignments can be derived through majority voting over the tissue cell neighborhood assignments from the one or more partitioning methods. In some examples, cells without consensus across the one or more partitioning methods can be removed. After block 412, the process depicted by the flowchart of FIG. 4 can come to an end.
[0064] Referring now to FIG. 5, shown is a flowchart that provides one example of the operation of a portion of the visualization service 118. The flowchart of FIG. 5 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the visualization service 118. As an alternative, the flowchart of FIG. 5 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
[0065] At block 500, the visualization service 118 can be executed to receive a segmented image 123, one or more associated cell type signatures 129, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113. In some examples, the image coordinates can represent the x and y coordinates of a cell within a segmented image 123. In some examples, the visualization service 118 can be executed to call a function to retrieve a segmented image 123 from the data store 119, the cell type application 113, or another program within the network
environment 100. In some examples, the visualization service 118 can be executed to call a function to retrieve one or more cell type signatures 129 and image coordinates assigned to one or more cells from the cell type application 113 and include a file path as an argument to the function. In some examples, the cell type application 113 can automatically send the cell type signatures 129 and image coordinates to the visualization service 118 after assigning cell type signatures 129 to a segmented image 123.
[0066] Next, at block 503, the visualization service 118 can be executed to partition a segmented image 123 into one or more subsets. In some examples, the visualization service 118 can be executed to partition the segmented image 123 into one or more vertical subsets based at least in part on defined x coordinate boundaries. In some examples, the visualization service 118 can be executed to partition the segmented image 123 into one or more horizontal subsets based at least in part on defined y coordinate boundaries. In some examples, the visualization service 118 can be executed to partition the segmented image 123 into one or more square grid subsets by dividing coordinates along both axes. In some examples, the x coordinate boundaries and/or y coordinate boundaries are defined by a user. In some examples, the partitions can be recursively refined to ensure the partitions contain a minimum and/or maximum number of cells. In some examples, the partitions exceeding a maximum cell count can be subdivided and partitions containing a number of cells below the minimum cell count can be merged. In some examples, the visualization service 118 can be executed to partition the segmented image 123 based at least in part on one or more partitions defined by a user.
[0067] At block 506, the visualization service 118 can be executed to locate one or more cell type signatures 129 within one or more partitions created at block 503. In some
examples, the visualization service 118 can be executed to call a function to locate one or more cell type signatures 129 and provide the cell type signatures 129 and image coordinates received at block 500 as arguments to the function. In some examples, the one or more cell type signatures 129 to be located by the visualization service 118 within the partitions are defined by a user.
[0068] At block 509, the visualization service 118 can be executed to determine a cooccurrence between two or more cell type signatures 129 located in the partitions at block 506. In some examples, the visualization service 118 can be executed to call a function to determine one or more cell type signatures 129 located within a defined distance of one or more different cell type signatures 129 and provide the cell type signatures 129, image coordinates, and the defined distance as arguments to the function. In some examples, the visualization service 118 can be executed to identify a first cell type signature 129 located within a defined distance from a second cell type signature 129. In some examples, the defined distance can be modified by a user. For example, the visualization service 118 can be executed to determine cells having a cell type A that are located within fifty (50) microns of cells having a cell type B within a two-hundred (200) micron partition of a segmented image 123. In some examples, the visualization service 118 can be further executed to send the co-occurrences to a user interface 143 for display to a user. After block 509, the process depicted by the flowchart of FIG. 5 can come to an end.
[0069] Referring now to FIG. 6, shown is a flowchart that provides one example of the operation of a portion of the visualization service 118. The flowchart of FIG. 6 provides merely an example of the many different types of functional arrangements that can be
employed to implement the operation of the depicted portion of the visualization service 118. As an alternative, the flowchart of FIG. 6 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
[0070] At block 600, the visualization service 118 can be executed to receive a segmented image 123, one or more associated cell type signatures 129, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113. In some examples, the image coordinates can represent the x and y coordinates of a cell within a segmented image 123. In some examples, the visualization service 118 can be executed to call a function to retrieve a segmented image 123 from the data store 119, the cell type application 113, or another program within the network environment 100. In some examples, the visualization service 118 can be executed to call a function to retrieve one or more cell type signatures 129 and image coordinates assigned to one or more cells from the cell type application 113 and include a file path as an argument to the function. In some examples, the cell type application 113 can automatically send the cell type signatures 129 and image coordinates to the visualization service 118 after assigning cell type signatures 129 to a segmented image 123.
[0071] Next, at block 603, the visualization service 118 can be executed to compare the cells within the plurality of segmented images 123. In some examples, the visualization service 118 can be executed to call a function to compare the cells within a first segmented image 123 with the cells within a second segmented image 123 and provide the cell type signatures 129 and coordinates received at block 600 as arguments to the function. In some examples, the visualization service 118 can be executed to compare one or more cell states, a proportion of one or more cell type signatures 129,
one or more co-occurrences, one or more biomarkers, and/or one or more gene expressions. In some examples, the visualization service 118 can be executed to compare the cells within a segmented image 123 associated with one condition with the cells within one or more segmented images 123 associated with other conditions. For example, the visualization service 118 can compare the cells from a segmented image 123 associated with a healthy state with the cells from a segmented image 123 associated with a diseased state. In another example, the visualization service 118 can compare the cells from a segmented image 123 associated with a first condition of a diseased state with the cells from segmented images 123 associated with other conditions of the diseased state. In some examples, the visualization service 118 could compare a region of interest between different images associated with different conditions of a diseased state or between a healthy state and a diseased state.
[0072] At block 606, the visualization service 118 can be executed to identify one or more patterns based at least in part on the comparisons made at block 603. In some examples, the visualization service 118 can be executed to call a function to estimate a survival outcome based at least in part on the cells within a segmented image 123 and provide a time to event associated with comparisons made at block 603 as an argument to the function. For example, a co-occurrence between T-cells and B-cells could be a biomarker for a cancer type and greater percentages of the total cell distribution of this biomarker can correlate to lower survival outcomes. In this example, the visualization service 118 could identify this pattern by comparing survival outcomes (or times to event) associated with varying levels of this biomarker and estimate a survival outcome for a
patient associated with a segmented image 123 having the biomarker present in 25% of the total cell distribution.
[0073] Next, at block 609, the visualization service 118 can be executed to validate the one or more patterns identified at block 606. In some examples, the visualization service 118 can be executed to receive user feedback associated with the one or more patterns. For example, a user could provide feedback that a biomarker present in 20% or greater of the total cell distribution is associated with a 50% survival outcome. In another example, a user could provide feedback that a patient associated with a segmented image 123 used to identify a pattern at block 606 was diagnosed with a cancer type a specified time after the pattern was identified. In another example, a user could provide feedback that a biomarker present in 20% or greater of the total cell distribution is associated with decreased response to a drug or with increased probability of cancer recurrence. In some examples, the visualization service 118 can be executed to receive feedback from the machine learning model 116 associated with the one or more patterns. For example, the visualization service 118 could be executed to send the one or more patterns to the machine learning model 116 and receive feedback from the machine learning model 116. In some examples, the visualization service 118 could then be executed to modify the pattern based at least in part on the feedback. After block 609, the process depicted by the flowchart of FIG. 6 can end.
[0074] Referring now to FIG. 7, shown is a flowchart that provides one example of the operation of a portion of the visualization service 118. The flowchart of FIG. 7 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the visualization service
118. As an alternative, the flowchart of FIG. 7 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
[0075] At block 700, the visualization service 118 can be executed to receive a segmented image 123, one or more associated cell features 133, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113. In some examples, the image coordinates can represent the x and y coordinates of a cell within a segmented image 123. In some examples, the visualization service 118 can be executed to call a function to retrieve a segmented image 123 from the data store 119, the cell type application 113, or another program within the network environment 100. In some examples, the visualization service 118 can be executed to call a function to retrieve one or more cell features 133 and image coordinates assigned to one or more cells from the cell type application 113 and include a file path as an argument to the function. In some examples, the cell type application 113 can automatically send the cell features 133 and image coordinates to the visualization service 118 after assigning cell type signatures 129 to a segmented image 123.
[0076] Next, at block 703, the visualization service 118 can be executed to identify one or more biomarkers from the cell features 133 within the cells of the segmented image 123. In some examples, the visualization service 118 can be executed to identify one or more biomarkers specified by a user. In some examples, the visualization service 118 can be executed to identify one or more new combinations of biomarkers present in the segmented image 123. For example, cells that have biomarker A and biomarker B could be identified and flagged as having a new combination of biomarkers if they also have biomarker C. In some examples, the visualization service 118 can identify and flag a new
combination of biomarkers if the combination occurs greater than a defined threshold within a segmented image 123. In some examples, the defined threshold can be selected by a user. In other examples, the defined threshold can be selected by the cell type application 113 or machine learning model 116.
[0077] At block 706, the visualization service 118 can be executed to add one or more new cell type signatures 129 based at least in part on the one or more biomarkers identified at block 703. For example, the visualization service 118 could identify a biomarker present in unidentified cells and add a new cell type signature 129 based at least in part on the biomarker. In some examples, the visualization service 118 could add the new cell type signature 129 if the number of unidentified cells having the biomarker exceeds a defined threshold. In other examples, the visualization service 118 could add a new cell type signature 129 based at least in part on a new combination of biomarkers identified at block 703. For example, the visualization service 118 could add a new cell type signature 129 for cells having biomarkers A, B, and C. In some examples, the visualization service 118 could add the new cell type signature 129 if the number of cells having the combination of biomarkers exceeds a defined threshold. In some examples, the defined threshold can be selected by a user. In other examples, the defined threshold can be selected by the cell type application 113 or the machine learning model 116.
[0078] At block 709, the visualization service 118 can be executed to validate the one or more new cell type signatures 129 added at block 706. In some examples, the visualization service 118 can be executed to flag new cell type signatures 129 added at block 706, send the new cell type signatures 129 to a user interface 143, and receive user feedback entered in the user interface 143. In some examples, the user feedback can
confirm the one or more new cell type signatures 129. In other examples, the user feedback can cause the visualization service 118 to be executed to modify or delete the one or more new cell type signatures 129. In some examples, the visualization service 118 can be executed to send the new cell type signatures 129 to the machine learning model 116 and receive feedback from the machine learning model 116. In some examples, the feedback from the machine learning model 116 can confirm the one or more new cell type signatures 129. In other examples, the feedback from the machine learning model 116 can cause the visualization service 118 to be executed to modify or delete the one or more new cell type signatures 129. After block 709, the process depicted by the flowchart of FIG. 7 can end.
[0079] Referring now to FIG. 8, shown is a flowchart that provides one example of the operation of a portion of the visualization service 118. The flowchart of FIG. 8 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the visualization service 118. As an alternative, the flowchart of FIG. 8 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
[0080] At block 800, the visualization service 118 can be executed to receive a segmented image 123, one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113. In some examples, the image coordinates can represent the x and y coordinates of a cell within a segmented image 123. In some examples, the visualization
service 118 can be executed to call a function to retrieve a segmented image 123 and/or a list of ligand-receptor pairs from the data store 119, the cell type application 113, or another program within the network environment 100. In some examples, the visualization service 118 can be executed to call a function to retrieve one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113 and include a file path as an argument to the function. In some examples, the cell type application 113 can automatically send the one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 to the visualization service 118 after assigning cell type signatures 129 to a segmented image 123.
[0081] Next, at block 803, the visualization service 118 can be executed to identify one or more cells within a defined radius for each cell in the segmented image 123. In some examples, the defined radius can be modified by a user. In some examples, the defined radius can be set by the cell type application 113 or the machine learning model 116.
[0082] At block 806, the visualization service 118 can be executed to establish one or more connections between cells expressing a ligand gene and other cells within the defined radius expressing a corresponding receptor gene. In some examples, the visualization service 118 can be executed to divide the segmented image 123 into a grid.
In some examples, the visualization service 118 can then be executed to calculate a spatial kernal density score for each connection. In some examples, the visualization service 118 can then be executed to identify one or more regions with dense ligandreceptor interactions based at least in part on the spatial kernal density scores.
[0083] At block 809, the visualization service 118 can be executed to identify one or more clusters of ligand-receptor pairs. In some examples, the visualization service 118 can be executed to apply a Louvain clustering algorithm to the density scores calculated at block 806. In some examples, the visualization service 118 can be executed to send the density scores to the machine learning model 116 and receive clusters from the machine learning model 116. After block 809, the process depicted by the flowchart of FIG. 8 can then end.
[0084] Referring now to FIG. 9, shown is a flowchart that provides one example of the operation of a portion of the visualization service 118. The flowchart of FIG. 9 provides merely an example of the many different types of functional arrangements that can be employed to implement the operation of the depicted portion of the visualization service 118. As an alternative, the flowchart of FIG. 9 can be viewed as depicting an example of elements of a method implemented within the network environment 100.
[0085] At block 900, the visualization service 118 can be executed to receive a segmented image 123, one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113. In some examples, the image coordinates can represent the x and y
coordinates of a cell within a segmented image 123. In some examples, the visualization service 118 can be executed to call a function to retrieve a segmented image 123 and/or a list of ligand-receptor pairs from the data store 119, the cell type application 113, or another program within the network environment 100. In some examples, the visualization service 118 can be executed to call a function to retrieve one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 from the cell type application 113 and include a file path as an argument to the function. In some examples, the cell type application 113 can automatically send the one or more associated cell features 133, one or more associated cell type signatures 129, a data matrix containing expression data for one or more ligands and one or more receptors per cell, a list of ligand-receptor pairs, and image coordinates assigned to one or more cells within the segmented image 123 to the visualization service 118 after assigning cell type signatures 129 to a segmented image 123.
[0086] Next, at block 903, the visualization service 118 can be executed to establish one or more connections between cells expressing a ligand gene and other cells within a defined radius expressing a corresponding receptor gene. In some examples, the visualization service 118 can be executed to identify one or more cells within a defined radius for each cell in the segmented image 123. In some examples, the defined radius can be modified by a user. In some examples, the defined radius can be set by the cell type application 113 or the machine learning model 116. In some examples, the visualization service 118 can be executed to call a function to establish a connection
between a cell expressing a ligand gene and a cell expressing a corresponding receptor gene within a defined radius and provide data received at block 900 and a defined radius as arguments to the function. For example, the visualization service 118 could establish a connection between a cell expressing a CXCL12 gene and a nearby cell expressing a CXCR4 gene.
[0087] At block 906, the visualization service 118 can be executed to create one or more receptor-ligand modules based at least in part on ligand-receptor connections established at block 903. In some examples, the visualization service 118 can be executed to divide the segmented image 123 into a grid. In some examples, the visualization service 118 can then be executed to calculate a spatial kernal density score for each connection. In some examples, the visualization service 118 can then be executed to identify one or more regions with dense ligand-receptor interactions based at least in part on the spatial kernal density scores. For example, the visualization service 118 could create a module defined by CXCL12-CXCR4, CCL14-CCR1 , and CXCL16- CXCR6 pairs.
[0088] At block 909, the visualization service 118 can be executed to assess patterns based at least in part on the receptor-ligand modules and data identified at block 900. In some examples, the visualization service 118 can be executed to identify one or more clusters of ligand-receptor pairs. In some examples, the visualization service 118 can be executed to apply a Louvain clustering algorithm to the calculated density scores. In some examples, the visualization service 118 can be executed to send the density scores to the machine learning model 116 and receive clusters from the machine learning model 116. In some examples, the visualization service 118 could identify cell type signatures
129 associated with one or more modules identified at block 906. For example, the visualization service 118 could identify that a module defined by CXCL12-CXCR4, CXCL16-CXCR6, and CCL14-CCR1 pairs was both a hub sharing spatial distribution with at least three modules in both gland and mucosa tissues and was also composed of antigen-presenting cells along with lymphatic endothelial and immune cell type signatures 129.
[0089] At block 912, the visualization service 118 can be executed to add one or more cell type signatures 129 based at least in part on the patterns identified at block 909. In some examples, the visualization service 118 can be executed to call a function to add a new cell type signature 129 and provide one or more ligand-receptor pairs as arguments to the function. For example, the visualization service 118 could look for cell type signatures 129 that are associated with ligand-receptor modules identified at block 906 and create a new cell type signature 129 based at least in part on the patterns identified at block 909. For example, the visualization service 118 could create a new cell type signature 129 for unique receptor-ligand pairs across different tissue sites. After block 912, the process depicted by the flowchart of FIG. 9 can then end.
[0090] A number of software components previously discussed are stored in the memory of the respective computing devices and are executable by the processor of the respective computing devices. In this respect, the term "executable" means a program file that is in a form that can ultimately be run by the processor. Examples of executable programs can be a compiled program that can be translated into machine code in a format that can be loaded into a random access portion of the memory and run by the processor, source code that can be expressed in proper format such as object code that is capable
of being loaded into a random access portion of the memory and executed by the processor, or source code that can be interpreted by another executable program to generate instructions in a random access portion of the memory to be executed by the processor. An executable program can be stored in any portion or component of the memory, including random access memory (RAM), read-only memory (ROM), hard drive, solid-state drive, Universal Serial Bus (USB) flash drive, memory card, optical disc such as compact disc (CD) or digital versatile disc (DVD), floppy disk, magnetic tape, or other memory components.
[0091] The memory includes both volatile and nonvolatile memory and data storage components. Volatile components are those that do not retain data values upon loss of power. Nonvolatile components are those that retain data upon a loss of power. Thus, the memory can include random access memory (RAM), read-only memory (ROM), hard disk drives, solid-state drives, USB flash drives, memory cards accessed via a memory card reader, floppy disks accessed via an associated floppy disk drive, optical discs accessed via an optical disc drive, magnetic tapes accessed via an appropriate tape drive, or other memory components, or a combination of any two or more of these memory components. In addition, the RAM can include static random access memory (SRAM), dynamic random access memory (DRAM), or magnetic random access memory (MRAM) and other such devices. The ROM can include a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable readonly memory (EEPROM), or other like memory device.
[0092] Although the applications and systems described herein can be embodied in software or code executed by general purpose hardware as discussed above, as an
alternative the same can also be embodied in dedicated hardware or a combination of software/general purpose hardware and dedicated hardware. If embodied in dedicated hardware, each can be implemented as a circuit or state machine that employs any one of or a combination of a number of technologies. These technologies can include, but are not limited to, discrete logic circuits having logic gates for implementing various logic functions upon an application of one or more data signals, application specific integrated circuits (ASICs) having appropriate logic gates, field-programmable gate arrays (FPGAs), or other components, etc. Such technologies are generally well known by those skilled in the art and, consequently, are not described in detail herein.
[0093] The flowcharts show the functionality and operation of an implementation of portions of the various embodiments of the present disclosure. If embodied in software, each block can represent a module, segment, or portion of code that includes program instructions to implement the specified logical function(s). The program instructions can be embodied in the form of source code that includes human-readable statements written in a programming language or machine code that includes numerical instructions recognizable by a suitable execution system such as a processor in a computer system. The machine code can be converted from the source code through various processes. For example, the machine code can be generated from the source code with a compiler prior to execution of the corresponding application. As another example, the machine code can be generated from the source code concurrently with execution with an interpreter. Other approaches can also be used. If embodied in hardware, each block can represent a circuit or a number of interconnected circuits to implement the specified logical function or functions.
[0094] Although the flowcharts show a specific order of execution, it is understood that the order of execution can differ from that which is depicted. For example, the order of execution of two or more blocks can be scrambled relative to the order shown. Also, two or more blocks shown in succession can be executed concurrently or with partial concurrence. Further, in some embodiments, one or more of the blocks shown in the flowcharts can be skipped or omitted. In addition, any number of counters, state variables, warning semaphores, or messages might be added to the logical flow described herein, for purposes of enhanced utility, accounting, performance measurement, or providing troubleshooting aids, etc. It is understood that all such variations are within the scope of the present disclosure.
[0095] Also, any logic or application described herein that includes software or code can be embodied in any non-transitory computer-readable medium for use by or in connection with an instruction execution system such as a processor in a computer system or other system. In this sense, the logic can include statements including instructions and declarations that can be fetched from the computer-readable medium and executed by the instruction execution system . In the context of the present disclosure, a "computer-readable medium" can be any medium that can contain, store, or maintain the logic or application described herein for use by or in connection with the instruction execution system. Moreover, a collection of distributed computer-readable media located across a plurality of computing devices (e.g., storage area networks or distributed or clustered filesystems or databases) may also be collectively considered as a single non- transitory computer-readable medium.
[0096] The computer-readable medium can include any one of many physical media such as magnetic, optical, or semiconductor media. More specific examples of a suitable computer-readable medium would include, but are not limited to, magnetic tapes, magnetic floppy diskettes, magnetic hard drives, memory cards, solid-state drives, USB flash drives, or optical discs. Also, the computer-readable medium can be a random access memory (RAM) including static random access memory (SRAM) and dynamic random access memory (DRAM), or magnetic random access memory (MRAM). In addition, the computer-readable medium can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or other type of memory device.
[0097] Further, any logic or application described herein can be implemented and structured in a variety of ways. For example, one or more applications described can be implemented as modules or components of a single application. Further, one or more applications described herein can be executed in shared or separate computing devices or a combination thereof. For example, a plurality of the applications described herein can execute in the same computing device, or in multiple computing devices in the same computing environment 103.
[0098] Disjunctive language such as the phrase “at least one of X, Y, or Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., can be either X, Y, or Z, or any combination thereof (e.g., X; Y; Z; X or Y; X or Z; Y or Z; X, Y, or Z; etc.). Thus, such disjunctive language is
not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y, or at least one of Z to each be present.
[0099] It should be emphasized that the above-described embodiments of the present disclosure are merely possible examples of implementations set forth for a clear understanding of the principles of the disclosure. Many variations and modifications can be made to the above-described embodiments without departing substantially from the spirit and principles of the disclosure. All such modifications and variations are intended to be included herein within the scope of this disclosure and protected by the following claims.
[0100] Clause 1 - A system, comprising: a computing device comprising a processor and a memory; and machine-readable instructions stored in the memory that, when executed by the processor, cause the computing device to at least: receive a segmented image, the segmented image comprising a plurality of segments; identify a plurality of cell features from each of the plurality of segments; calculate a cell type relevance score for each of the plurality of segments based at least in part on the plurality of cell features; and assign a cell type signature for each segment based at least in part on the cell type relevance score.
[0101] Clause 2 - The system of clause 1 , wherein the machine-readable instructions which, when executed by the processor, cause the computing device to calculate the cell type relevance score further cause the computing device to at least: determine a weight for each of the plurality of cell features; assign a feature score to each of the plurality of cell features; and perform a weighted linear combination of each of the feature scores to obtain the cell type relevance score.
[0102] Clause 3 - The system of clause 1 or 2, wherein each of the plurality of segments represents an individual cell in the segmented image.
[0103] Clause 4 - The system of any of clauses 1 -3, wherein the machine-readable instructions, when executed by the processor, further cause the computing device to at least: identify with a machine learning model, one or more clusters of individual segments of the plurality of segments based at least in part on the cell type relevance score for each of the plurality of segments; determine a distribution of cell type relevance scores across the one or more clusters; determine a positive identification threshold based at least in part on the distribution; and determine for individual segments of the plurality of segments a cell identification based at least in part on a comparison of the corresponding cell type relevance score to the positive identification threshold.
[0104] Clause 5 - The system of clause 4, wherein the machine-readable instructions which, when executed by the processor, cause the computing device to determine a positive identification threshold, further cause the computing device to at least determine, using piecewise linear regression, the positive identification threshold based at least in part on the distribution of cell type relevance scores.
[0105] Clause 6 - The system of clause 4 or 5, wherein the machine-readable instructions, when executed by the processor, further cause the computing device to at least determine a positive cell identification for an individual segment of the plurality of segments based at least in part on the corresponding cell type relevance score exceeding the positive identification threshold.
[0106] Clause 7 - The system of clause 4 or 5, wherein the machine-readable instructions, when executed by the processor, further cause the computing device to at
least determine a negative cell identification for an individual segment of the plurality of segments based at least in part on the corresponding cell type relevance score failing to exceed the positive identification threshold.
[0107] Clause 8 - The system of any of clauses 1 -7, wherein the machine-readable instructions which, when executed by the processor, cause the computing device to assign the cell type signature, further cause the computing device to at least: identify a first individual segment having multiple assigned cell type signatures; identify a second individual segment having a single assigned cell type signature; compare the first individual segment with the second individual segment to determine a correct cell type signature; and assign the correct cell type signature to the first individual segment.
[0108] Clause 9 - A method, comprising: receiving, by a computing device, a segmented image, the segmented image comprising a plurality of segments; identifying, by the computing device, a plurality of cell features from each of the plurality of segments; calculating, by the computing device, a cell type relevance score for each of the plurality of segments based at least in part on the plurality of cell features; and assigning, by the computing device, a cell type signature for each segment based at least in part on the cell type relevance score.
[0109] Clause 10 - The method of clause 9, further comprising: identifying, with a machine learning model, one or more clusters of individual segments of the plurality of segments based at least in part on the cell type relevance score for each of the plurality of segments; determining, by a computing device, a distribution of cell type relevance scores across the one or more clusters; determining, by the computing device, a positive identification threshold based at least in part on the distribution; and determining, by the
computing device, a cell identification for individual segments of the plurality of segments based at least in part on a comparison of the corresponding cell type relevance score to the positive identification threshold.
[0110] Clause 11 - The method of clause 10, wherein determining a positive identification threshold further comprises at least: determining, by the computing device, a distribution of cell type relevance scores across the one or more clusters; and using piecewise linear regression, determining, by the computing device, the positive identification threshold based at least in part on the distribution of cell type relevance scores.
[0111] Clause 12 - The method of clause 10 or 11 , further comprising at least determining, by the computing device, a positive cell identification for an individual segment of the plurality of segments based at least in part on the corresponding cell type relevance score exceeding the positive identification threshold.
[0112] Clause 13 - The method of clause 10 or 11 , further comprising at least determining, by the computing device, a negative cell identification for an individual segment of the plurality of segments based at least in part on the corresponding cell type relevance score failing to exceed the positive identification threshold.
[0113] Clause 14 - The method of any of clauses 9-13, wherein calculating the cell type relevance score further comprises at least: determining, by the computing device, a weight for each of the plurality of cell features; assigning, by the computing device, a feature score to each of the plurality of cell features; and performing, by the computing device, a weighted linear combination of each of the feature scores to obtain the cell type relevance score.
[0114] Clause 15 - The method of any of clauses 9-14, wherein assigning the cell type further comprises at least: identifying, by the computing device, a first individual segment having multiple assigned cell type signatures; identifying, by the computing device, a second individual segment having a single assigned cell type signature; comparing, by the computing device, the first individual segment with the second individual segment to determine a correct cell type signature; and assigning, by the computing device, the correct cell type signature to the first individual segment.
[0115] Clause 16 - The method of clause 15, wherein comparing the first individual segment with the second individual segment further comprises: using K-nearest neighbor, determining, by the computing device, a similarity between the first individual segment and the second individual segment; and determining the correct cell type signature based at least in part on the similarity.
[0116] Clause 17 -A system, comprising: a computing device comprising a processor and a memory; and machine-readable instructions stored in the memory that, when executed by the processor, cause the computing device to at least: receive a segmented image, individual segments of the segmented image corresponding to individual cells; receive a feature matrix associated with a cell type signature; identify a plurality of cell features from the individual segments; compare the plurality of cell features with the feature matrix to generate a cell type relevance score for individual segments; and assign a cell type signature for individual segments based at least in part on the cell type relevance score.
[0117] Clause 18 - The system of clause 17, wherein the machine-readable instructions which, when executed by the processor, cause the computing device to
compare the plurality of cell features with the feature matrix, further cause the computing device to at least: determine a weight for each of the plurality of cell features based at least in part on the feature matrix; assign a feature score to each of the plurality of cell features based at least in part on the individual segments; and calculate a cell type relevance score for each of the plurality of cell features based at least in part on the weight and the feature scores.
[0118] Clause 19 - The system of clause 17 or 18, wherein the machine-readable instructions, when executed by the processor, further cause the computing device to at least: identify with a machine learning model, one or more clusters of individual cells of based at least in part on the cell type relevance score for each of the plurality of cell features; determine a distribution of cell type relevance scores across the one or more clusters; determine a positive identification threshold based at least in part on the distribution; and determine a cell identification for individual cells based at least in part on a comparison of the corresponding cell type relevance score to the positive identification threshold.
[0119] Clause 20 - The system of any of claims 17-19, wherein the machine-readable instructions which, when executed by the processor, cause the computing device to assign the cell type, further cause the computing device to at least: identify a first individual segment having multiple assigned cell type signatures; identify a second individual segment having a single assigned cell type signature; compare the first individual segment with the second individual segment to determine a correct cell type signature; and assign the correct cell type signature to the first individual segment.
Claims
1. A system, comprising: a computing device comprising a processor and a memory; and machine-readable instructions stored in the memory that, when executed by the processor, cause the computing device to at least: receive a segmented image, the segmented image comprising a plurality of segments; identify a plurality of cell features from each of the plurality of segments; calculate a cell type relevance score for each of the plurality of segments based at least in part on the plurality of cell features; and assign a cell type signature for each segment based at least in part on the cell type relevance score.
2. The system of claim 1 , wherein the machine-readable instructions which, when executed by the processor, cause the computing device to calculate the cell type relevance score further cause the computing device to at least: determine a weight for each of the plurality of cell features; assign a feature score to each of the plurality of cell features; and perform a weighted linear combination of each of the feature scores to obtain the cell type relevance score.
3. The system of claim 1 or 2, wherein each of the plurality of segments represents an individual cell in the segmented image.
4. The system of any of claims 1 -3, wherein the machine-readable instructions, when executed by the processor, further cause the computing device to at least: identify with a machine learning model, one or more clusters of individual segments of the plurality of segments based at least in part on the cell type relevance score for each of the plurality of segments; determine a distribution of cell type relevance scores across the one or more clusters; determine a positive identification threshold based at least in part on the distribution; and determine for individual segments of the plurality of segments a cell identification based at least in part on a comparison of the corresponding cell type relevance score to the positive identification threshold.
5. The system of claim 4, wherein the machine-readable instructions which, when executed by the processor, cause the computing device to determine a positive identification threshold, further cause the computing device to at least determine, using piecewise linear regression, the positive identification threshold based at least in part on the distribution of cell type relevance scores.
6. The system of claim 4 or 5, wherein the machine-readable instructions, when executed by the processor, further cause the computing device to at least determine a positive cell identification for an individual segment of the plurality of segments based at least in part on the corresponding cell type relevance score exceeding the positive identification threshold.
7. The system of claim 4 or 5, wherein the machine-readable instructions, when executed by the processor, further cause the computing device to at least determine a negative cell identification for an individual segment of the plurality of segments based at least in part on the corresponding cell type relevance score failing to exceed the positive identification threshold.
8. The system of any of claims 1 -7, wherein the machine-readable instructions which, when executed by the processor, cause the computing device to assign the cell type signature, further cause the computing device to at least: identify a first individual segment having multiple assigned cell type signatures; identify a second individual segment having a single assigned cell type signature; compare the first individual segment with the second individual segment to determine a correct cell type signature; and assign the correct cell type signature to the first individual segment.
9. A method, comprising: receiving, by a computing device, a segmented image, the segmented image comprising a plurality of segments; identifying, by the computing device, a plurality of cell features from each of the plurality of segments; calculating, by the computing device, a cell type relevance score for each of the plurality of segments based at least in part on the plurality of cell features; and assigning, by the computing device, a cell type signature for each segment based at least in part on the cell type relevance score.
10. The method of claim 9, further comprising: identifying, with a machine learning model, one or more clusters of individual segments of the plurality of segments based at least in part on the cell type relevance score for each of the plurality of segments; determining, by a computing device, a distribution of cell type relevance scores across the one or more clusters; determining, by the computing device, a positive identification threshold based at least in part on the distribution; and determining, by the computing device, a cell identification for individual segments of the plurality of segments based at least in part on a comparison of the corresponding cell type relevance score to the positive identification threshold.
11. The method of claim 10, wherein determining a positive identification threshold further comprises at least: determining, by the computing device, a distribution of cell type relevance scores across the one or more clusters; and using piecewise linear regression, determining, by the computing device, the positive identification threshold based at least in part on the distribution of cell type relevance scores.
12. The method of claim 10 or 11 , further comprising at least determining, by the computing device, a positive cell identification for an individual segment of the plurality of segments based at least in part on the corresponding cell type relevance score exceeding the positive identification threshold.
13. The method of claim 10 or 11 , further comprising at least determining, by the computing device, a negative cell identification for an individual segment of the plurality of segments based at least in part on the corresponding cell type relevance score failing to exceed the positive identification threshold.
14. The method of any of claims 9-13, wherein calculating the cell type relevance score further comprises at least: determining, by the computing device, a weight for each of the plurality of cell features; assigning, by the computing device, a feature score to each of the plurality of cell features; and performing, by the computing device, a weighted linear combination of each of the feature scores to obtain the cell type relevance score.
15. The method of any of claims 9-14, wherein assigning the cell type further comprises at least: identifying, by the computing device, a first individual segment having multiple assigned cell type signatures; identifying, by the computing device, a second individual segment having a single assigned cell type signature; comparing, by the computing device, the first individual segment with the second individual segment to determine a correct cell type signature; and assigning, by the computing device, the correct cell type signature to the first individual segment.
16. The method of claim 15, wherein comparing the first individual segment with the second individual segment further comprises: using K-nearest neighbor, determining, by the computing device, a similarity between the first individual segment and the second individual segment; and determining the correct cell type signature based at least in part on the similarity.
17. A system, comprising: a computing device comprising a processor and a memory; and machine-readable instructions stored in the memory that, when executed by the processor, cause the computing device to at least: receive a segmented image, individual segments of the segmented image corresponding to individual cells; receive a feature matrix associated with a cell type signature; identify a plurality of cell features from the individual segments; compare the plurality of cell features with the feature matrix to generate a cell type relevance score for individual segments; and assign a cell type signature for individual segments based at least in part on the cell type relevance score.
18. The system of claim 17, wherein the machine-readable instructions which, when executed by the processor, cause the computing device to compare the plurality of cell features with the feature matrix, further cause the computing device to at least: determine a weight for each of the plurality of cell features based at least in part on the feature matrix; assign a feature score to each of the plurality of cell features based at least in part on the individual segments; and calculate a cell type relevance score for each of the plurality of cell features based at least in part on the weight and the feature scores.
19. The system of claim 17 or 18, wherein the machine-readable instructions, when executed by the processor, further cause the computing device to at least: identify with a machine learning model, one or more clusters of individual cells of based at least in part on the cell type relevance score for each of the plurality of cell features; determine a distribution of cell type relevance scores across the one or more clusters; determine a positive identification threshold based at least in part on the distribution; and determine a cell identification for individual cells based at least in part on a comparison of the corresponding cell type relevance score to the positive identification threshold.
20. The system of any of claims 17-19, wherein the machine-readable instructions which, when executed by the processor, cause the computing device to assign the cell type, further cause the computing device to at least: identify a first individual segment having multiple assigned cell type signatures; identify a second individual segment having a single assigned cell type signature; compare the first individual segment with the second individual segment to determine a correct cell type signature; and assign the correct cell type signature to the first individual segment.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463640148P | 2024-04-29 | 2024-04-29 | |
| US63/640,148 | 2024-04-29 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025230992A1 true WO2025230992A1 (en) | 2025-11-06 |
Family
ID=97562189
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2025/026820 Pending WO2025230992A1 (en) | 2024-04-29 | 2025-04-29 | Analyzing multiplexed and multimodal imaging data |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025230992A1 (en) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20170212028A1 (en) * | 2014-09-29 | 2017-07-27 | Biosurfit S.A. | Cell counting |
| US20170213067A1 (en) * | 2016-01-26 | 2017-07-27 | Ge Healthcare Bio-Sciences Corp. | Automated cell segmentation quality control |
| US20210166785A1 (en) * | 2018-05-14 | 2021-06-03 | Tempus Labs, Inc. | Predicting total nucleic acid yield and dissection boundaries for histology slides |
-
2025
- 2025-04-29 WO PCT/US2025/026820 patent/WO2025230992A1/en active Pending
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20170212028A1 (en) * | 2014-09-29 | 2017-07-27 | Biosurfit S.A. | Cell counting |
| US20170213067A1 (en) * | 2016-01-26 | 2017-07-27 | Ge Healthcare Bio-Sciences Corp. | Automated cell segmentation quality control |
| US20210166785A1 (en) * | 2018-05-14 | 2021-06-03 | Tempus Labs, Inc. | Predicting total nucleic acid yield and dissection boundaries for histology slides |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Narayan et al. | Assessing single-cell transcriptomic variability through density-preserving data visualization | |
| Li et al. | Benchmarking spatial and single-cell transcriptomics integration methods for transcript distribution prediction and cell type deconvolution | |
| Greenwald et al. | Whole-cell segmentation of tissue images with human-level performance using large-scale data annotation and deep learning | |
| Wu et al. | Graph deep learning for the characterization of tumour microenvironments from spatial protein profiles in tissue specimens | |
| Zormpas et al. | Mapping the transcriptome: Realizing the full potential of spatial data analysis | |
| Dorkenwald et al. | Automated synaptic connectivity inference for volume electron microscopy | |
| Yao et al. | Cell type classification and unsupervised morphological phenotyping from low-resolution images using deep learning | |
| Yang et al. | Spatial integration of multi-omics single-cell data with SIMO | |
| US20170076448A1 (en) | Identification of inflammation in tissue images | |
| Warchal et al. | Evaluation of machine learning classifiers to predict compound mechanism of action when transferred across distinct cell lines | |
| Schoenauer Sebag et al. | A generic methodological framework for studying single cell motility in high-throughput time-lapse data | |
| Janeiro et al. | Spatially resolved tissue imaging to analyze the tumor immune microenvironment: beyond cell-type densities | |
| Blampey et al. | Novae: a graph-based foundation model for spatial transcriptomics data | |
| Wang et al. | CW-NET for multitype cell detection and classification in bone marrow examination and mitotic figure examination | |
| Zhu et al. | CellLENS enables cross-domain information fusion for enhanced cell population delineation in single-cell spatial omics data | |
| Yuan et al. | SPANN: annotating single-cell resolution spatial transcriptome data with scRNA-seq data | |
| Hu et al. | Multisite assessment of reproducibility in high‐content cell migration imaging data | |
| Wang et al. | A generalizable Hi-C foundation model for chromatin architecture, single-cell and multi-omics analysis across species | |
| Baker et al. | emObject: domain specific data abstraction for spatial omics | |
| Tao et al. | Benchmarking mapping algorithms for cell-type annotating in mouse brain by integrating single-nucleus RNA-seq and Stereo-seq data | |
| Shah et al. | eLIMS: Ensemble Learning-Based Spatial Segmentation of Mass Spectrometry Imaging to Explore Metabolic Heterogeneity | |
| Wang et al. | stHGC: a self-supervised graph representation learning for spatial domain recognition with hybrid graph and spatial regularization | |
| WO2025230992A1 (en) | Analyzing multiplexed and multimodal imaging data | |
| Jiang et al. | Reconstructing spatial transcriptomics at the single-cell resolution with bayesDeep | |
| Li et al. | SC-Track: a robust cell-tracking algorithm for generating accurate single-cell lineages from diverse cell segmentations |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25798524 Country of ref document: EP Kind code of ref document: A1 |