WO2025199125A1 - Methods and systems for determining cellular and tissue images from gene expression - Google Patents
Methods and systems for determining cellular and tissue images from gene expressionInfo
- Publication number
- WO2025199125A1 WO2025199125A1 PCT/US2025/020407 US2025020407W WO2025199125A1 WO 2025199125 A1 WO2025199125 A1 WO 2025199125A1 US 2025020407 W US2025020407 W US 2025020407W WO 2025199125 A1 WO2025199125 A1 WO 2025199125A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- cell
- computer
- gene expression
- image
- images
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
- G06N3/0455—Auto-encoder networks; Encoder-decoder networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/12—Computing arrangements based on biological models using genetic models
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/12—Computing arrangements based on biological models using genetic models
- G06N3/126—Evolutionary algorithms, e.g. genetic algorithms or genetic programming
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B25/00—ICT specially adapted for hybridisation; ICT specially adapted for gene or protein expression
- G16B25/10—Gene or protein expression profiling; Expression-ratio estimation or normalisation
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/30—Unsupervised data analysis
Definitions
- RNA-seq single-cell RNA-seq
- TCR T cell receptor
- BCR B cell receptor
- a generative model based on transformer and diffusion models that learns a multi-modal representation of cells and tissues and reconstructs morphological information from transcriptome profiles to generate cellular and tissue images from single-cell expression profiles during inference.
- featured is a computer implemented method of generating a multi-dimensional (e.g., two-dimensional or three-dimensional) image of a cell or organelle from single cell sequencing data.
- the method includes providing a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle and using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
- the computer is programmed with a plurality of modules that generate the multi-dimensional image.
- the computer may be programed with three modules: an encoder model (e.g., a Bidirectional Encoder Representations from Transformers (BERT)-based single-cell language model), a diffusion model (e.g., a Latent Diffusion Model (LDM)), and a pretraining model (e.g., a biomedical Contrastive Language-Image Pretraining (CLIP) model).
- an encoder model e.g., a Bidirectional Encoder Representations from Transformers (BERT)-based single-cell language model
- a diffusion model e.g., a Latent Diffusion Model (LDM)
- LDM Latent Diffusion Model
- pretraining model e.g., a biomedical Contrastive Language-Image Pretraining (CLIP) model
- the method further includes providing a raw image of the cell or organelle.
- the computer is programmed to normalize and bin the gene expression of each of the plurality of genes.
- the computer is programmed to rank each of the plurality of genes by the normalized binned gene expression of each of the plurality of genes.
- the computer is programmed to transform the binned expression of each of the plurality of genes to produce the image.
- the computer is programmed to weight the binned expression or an identification thereof of each of the plurality of genes.
- the computer is programmed to mask a subset of the binned expression or the identification thereof of the plurality of genes.
- the computer is programmed to predict a remainder of the genes or binned gene expression that are not part of the masked subset.
- the computer is programmed to compress the image to reduce a dimensionality and add or reduce noise of the image.
- the computer is programmed to iteratively add or reduce noise of the image.
- the computer is programmed with a diffusion model.
- the machine learning algorithm model is trained with an empirical database of single cell gene expression profiles.
- the computer is programmed to optimize the image based on the empirical database to produce the multi-dimensional image of the cell or organelle.
- the computer is programmed to optimize the image based on the empirical database including a plurality of sets including matched gene expression profiles, images, and/or text annotation.
- the computer is programmed to correlate the set of gene expression profiles with the image of the cell or organelle.
- the computer is programmed to align the plurality of gene expression profiles with the plurality of sets including the matched gene expression profiles, images, and/or text annotations to produce a plurality of gene expression matrices.
- the computer is programmed to resize each image from the plurality of sets to match the size of the image.
- the computer is programmed to translate the plurality of gene expression matrices into the multi-dimensional image.
- the computer is programmed to rank the multi-dimensional image based on a similarity with the plurality of sets including the matched gene expression profiles, images, and/or text annotations. In some embodiments, the computer is programmed to produce a plurality of multi-dimensional images, and the method further includes selecting an image with the highest rank.
- the gene expression profiled is obtained by performing a spatial transcriptomic analysis.
- the spatial transcriptomic analysis includes single nucleus RNA (snRNA) sequencing, single cell RNA (scRNA) sequencing, patch sequencing, Visium, Xenium, NanoString GeoMx, STARmap, Slide-Seq, Slide-Tags, fluorescence in situ hybridization (FISH), multiplexed error- robust fluorescence in situ hybridization (MERFISH), electron microscopy, hematoxylin and eosin (H&E) imaging, fluorescence microscopy, 4',6-diamidino-2-phenylindole (DAPI), immunofluorescence (IF), or a combination thereof.
- snRNA single nucleus RNA
- scRNA single cell RNA
- patch sequencing Visium, Xenium, NanoString GeoMx, STARmap, Slide-Seq, Slide-Tags
- FISH fluorescence in situ hybridization
- MEFISH multiplexed error- robust fluorescence in situ hybridization
- H&E hematoxyl
- the organelle is a nucleus, nucleolus, ribosome, vesicle, endoplasmic reticulum (ER; e.g., rough ER or smooth ER), Golgi apparatus, vacuole, centriole, lysosome, or mitochondria.
- ER endoplasmic reticulum
- Golgi apparatus vacuole, centriole, lysosome, or mitochondria.
- the cell is a neuron
- the method shows a morphology of the neuron
- the method generates a multi-dimensional image of a plurality of cells or organelles.
- a tissue includes the plurality of cells, and the method produces a multidimensional image of the tissue.
- the method produces an electron microscopy, hematoxylin and eosin, or a fluorescence microscopy multi-dimensional image.
- the method produces fluorescence cell morphology.
- a computer implemented system including, for example, a storage device and a processor communicatively coupled to the storage device.
- the processor may be configured to execute code instructions stored on the storage device to cause the system to perform the computer implemented methods as described herein, e.g., of receiving a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multi-dimensional (e.g., two-dimensional or three-dimensional) image of the cell or organelle.
- a gene expression profile e.g., from sequencing and/or hybridization data
- the computer program product may include a non-transitory computer-readable storage device having computer-executable program instructions embodied thereon that when executed by a computer cause the computer to determine a multi-dimensional (e.g., two-dimensional or three-dimensional) image of a cell or an organelle.
- the computer-executable program instructions may include, for example, computer-executable program instructions to receive a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and computer-executable program instructions to determine the multi-dimensional image of the cell or organelle using the gene expression profile.
- the computer program product may be programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
- a software program product for determining a multi-dimensional (e.g., two-dimensional or three-dimensional) image of a cell or organelle.
- the software program product may include, for example, a computer-executable software program that includes instructions that when executed by a computer cause the computer to determine a multi-dimensional image of a cell or an organelle.
- the computer-executable program instructions may include computer-executable software program instructions to receive a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and computer-executable software program instructions to determine the multi-dimensional image of the cell or organelle using the gene expression profile.
- the software program product may be programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
- a computer may refer to device embodied in any of a number of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer, as nonlimiting examples. Additionally, a computer may be embedded in a device not generally regarded as a computer but with suitable processing capabilities, including a Personal Digital Assistant (PDA), a smartphone or any other suitable portable or fixed electronic device.
- PDA Personal Digital Assistant
- a computer may have one or more communication devices, which may be used to interconnect the computer to one or more other devices and/or systems, such as, for example, one or more networks in any suitable form, including a local area network or a wide area network, such as an enterprise network, and intelligent network (IN) or the Internet.
- Such networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks or wired networks.
- a computer may have one or more input devices and/or one or more output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that may be used to provide a user interface include printers or display screens for visual presentation of output and speakers or other sound generating devices for audible presentation of output. Examples of input devices that may be used for a user interface include keyboards, and pointing devices, such as mice, touch pads, and digitizing tablets. As another example, a computer may receive input information through speech recognition or in other audible formats. Computer-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices.
- program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types.
- the functionality of the program modules may be combined or distributed as desired in various embodiments.
- Databases if employed in the methods or devices or systems herein, may include computer readable memory (also referred to as “memory”).
- data storage space 3memlN may be and/or include computer readable memory, used to store data as described in the disclosure.
- Memory may be embodied by suitable hardware, including but not limited to the following: hard disk drives, serial advanced technology attachment (SATA) hard drives, SATA solid state drives (SSDs), non-volatile memory express (NVMe) SSDs, tape drives.
- SATA serial advanced technology attachment
- SSDs SATA solid state drives
- NVMe non-volatile memory express
- program “program,” “app,” and “software” are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that may be employed to program a computer or other processor to implement various embodiments described herein. Additionally, it should be appreciated that, according to one aspect, one or more computer programs that when executed perform methods of this application need not reside on a single computer or processor but may be distributed in a modular fashion among a number of different computers or processors to implement various embodiments of this application.
- FIG. 1A-FIG. 1D are a schematic flow chart showing the GenVinci workflow.
- FIG 1A Overview of GenVinci.
- [Upper] Input of GenVinci, which can be the gene expression data from single-nucleus RNA sequencing (snRNA-seq), single-cell RNA sequencing (scRNA-seq), spatially resolved transcriptomics (SRT) across different resolutions (nucleus/cell/spot), and Patch-seq.
- snRNA-seq single-nucleus RNA sequencing
- scRNA-seq single-cell RNA sequencing
- SRT spatially resolved transcriptomics
- [Middle] Representative deep learning architectures in GenVinci, which include Transformer, U-Net, and Autoencoder.
- GenVinci Output of GenVinci, which can be different imaging modalities at single-cell resolution (Hematoxylin-and-Eosin (H&E), 4',6-diamidino-2-phenylindole (DAPI), immunofluorescence (IF), neuron morphology, and electron microscopy) and multi-cell resolution (spot (1 -10 cells), tissue (5- 100 cells), and whole slide), and various downstream tasks (multi-modal clustering, zero-shot annotation, and data augmentation).
- FIG. 1 B GenVinci collects over 210 million gene expression data points from more than 20 data portals, including CELLxGENE, Single Cell PORTAL, and 10x Genomics to build a large-scale single-cell corpus (GenCorpus-210M).
- FIG. 1C The first stage (pretraining) of GenVinci. Gene expression data is sent into a newly-designed transformer framework for pretraining (masked language modeling). This step leverages prior biological knowledge from the collected large- scale single-cell corpus (Gencorpus-210M) to learn better gene expression embeddings specific to the utilized sequencing techniques.
- the utilized transformer framework employs a binning-based rank value encoding method to encode gene expression data and consists of several transformer encoders and one Mixture-of-Experts (MoE) decoder (FIG. 7A).
- FIG. 7A The utilized transformer framework employs a binning-based rank value encoding method to encode gene expression data and consists of several transformer encoders and one Mixture-of-Experts (MoE) decoder
- 1D The second stage (fine-tuning) of GenVinci.
- Gene expression data is sent to the previously trained GenVinci Large Language Model (LLM) and passed through a projector to obtain the transcriptomic prompt as the input of the generative model.
- LLM GenVinci Large Language Model
- a pre-trained diffusion model is employed as the backbone of the generative model to generate cellular images based on the transcriptomic prompt (FIG. 7B).
- the forward diffusion process generates the paired noisy images and the amount of noise added at each step for training (FIG. 7C).
- a vision-language model (VLM) is introduced for weak supervision, aiming to align transcriptomes, images, and potential text information within a shared latent space (FIG. 7D).
- FIG. 2A-FIG. 2H are H&E images produced by GenVinci from gene expression data across different sequencing resolutions. For each cell/spot/slide, 5 images were generated from the same gene expression data that were randomly sampled five times. The generated images were ranked based on the relative (normalized) cosine similarity scores (the best one with highest normalized cosine similarity score is 1 .0). The ground-truth (GT) images are listed in the first column.
- scH&E single-cell H&E
- FIG. 2B GenVinci generates multi-cell H&E (mcH&E) images from MERFISH single-cell gene expression (scGE) data.
- FIG. 2C GenVinci generates multi-cell H&E images (mcH&E) from 10x Visium multi-cell gene expression (mcGE) data (1 -10 cells/spot).
- FIG. 2D GenVinci generates multi-cell H&E images (mcH&E) from Spatial Transcriptomics multi-cell gene expression (mcGE) data (5-100 cells/spot).
- FIG. 2E GenVinci generates whole-slide H&E (wsH&E) images from Spatial Transcriptomics whole-slide gene expression matrices (wsGEM).
- FIG. 2F GenVinci-generated images are compared with ground-truth images and images generated by a Convolutional Neural Network (CNN) baseline. We selected the images generated by GenVinci with the highest similarity scores for comparison.
- FIG. 2G Comparisons of the Frechet Inception Distance (FID) scores (lower is better) and cosine similarity for the CNN baseline and GenVinci. The FID scores are calculated between the groundtruth and generated images.
- FID Frechet Inception Distance
- the cosine similarity (higher is better) is calculated based on the image embeddings extracted from both ground-truth and generated images using an ImageNet-pretrained ResNet-50 model.
- FIG. 2H Visualization of the ground-truth, CNN baseline, and GenVinci image embeddings using Uniform Manifold Approximation and Projection (UMAP).
- UMAP Uniform Manifold Approximation and Projection
- FIG. 3A-FIG. 3E are fluorescence cell morphology images produced by GenVinci from gene expression data across different sequencing resolutions.
- FIG. 3A GenVinci generates single-nucleus 4',6-diamidino-2-phenylindole (snDAPI) images from CosMx SMI single-nucleus gene expression (snGE) data.
- FIG. 3B GenVinci generates single-cell immunofluorescence (scIF) images from CosMx SMI single-cell gene expression (scGE) data.
- One channel represents DAPI
- another channel represents PanCK
- one channel represents CD45
- CD3 CD3.
- 3C GenVinci-generated best DAPI and IF-stained images are compared with ground-truth images and images generated by a CNN baseline.
- FIG. 3D Comparisons of the FID scores and cosine similarity for the CNN baseline and GenVinci.
- FIG. 3E Visualization of the ground-truth, CNN baseline, and GenVinci image embeddings using UMAP.
- FIG. 4A-FIG. 4D are single cell neuronon morphology images and graphs showing similarity with ground truth images.
- FIG. 4A GenVinci generates realistic single-cell neuron morphology (scNM) images from Patch-seq single-cell gene expression (scGE) data.
- FIG. 4B GenVinci-generated best neuron images are compared with ground-truth images and images generated by a CNN baseline.
- FIG. 4C Comparisons of the FID scores and cosine similarity for the CNN baseline and GenVinci.
- FIG. 4D Visualization of the ground-truth, CNN baseline, and GenVinci image embeddings using UMAP.
- FIG. 5D are single cell electron microscopy images (EM) and graphs showing similarly with ground truth images.
- FIG. 5A GenVinci generates realistic single-cell electron microscopy (scEM) images from MERFISH single-cell gene expression (scGE) data.
- FIG. 5B GenVinci-generated best EM images are compared with ground-truth images and images generated by a CNN baseline.
- FIG. 5C Comparisons of the FID scores and cosine similarity for the CNN baseline and GenVinci.
- FIG. 5D Visualization of the ground-truth, CNN baseline, and GenVinci image embeddings using UMAP.
- FIG. 6A-FIG. 6F are images and graphs showing how GenVinci can be flexibly applied to a diverse panel of downstream tasks, such as multi-modal clustering (FIG. 6A), zero-shot classification (FIGS. 6B and 6C), and data augmentation (FIGS. 6D-6F).
- FIG. 6A The clustering map of gene expression embeddings extracted by Geneformer (a recently published single-cell foundation model), continued-pre-trained Geneformer, and fine-tuned GenVinci. Cells/spots are colored with the cell- type/pathological annotations.
- FIG. 6B and 6C The classification performance of the cell- type/pathological annotation task (cancer/non-cancer) is evaluated on the st-HBC dataset and the sn- MBC dataset.
- cancer-related cells include MBC_neuronal
- non-cancer-related cells include endothelial_angiogenic, NK, and smooth muscle_vascular types. This classification could be useful for exploring the differences between cancerous and non-cancerous cells within the tumor microenvironment or metastatic niches.
- FIG. 6D The performance of the classification model trained solely with ground-truth (GT) images and GenVinci-generated (Gen) images with different augmentation ratios.
- GT ground-truth
- Gen GenVinci-generated
- FIG. 6E The performance of the classification model trained on GT images combined with GenVinci-generated images from the MERFISH dataset.
- FIG. 6F The performance of the classification model trained on GT images combined with GenVinci-generated images from an external scRNA-seq dataset. We set different ratios for the size of the augmented data.
- FIG. 7A-FIG. 7E are a schematic flow chart showing the GenVinci workflow
- FIG. 7A Transformer is utilized as the backbone of the sc-LLM model of GenVinci.
- GenVinci employs a binning-based rank value encoding method to encode gene expression data and consists of several transformer and MoE blocks to extract the gene expression embeddings.
- FIG. 7B Stable Diffusion is utilized as the backbone of the generative model of GenVinci.
- the transcriptomic prompt from the projector is concatenated with the Gaussian noise to generate cellular images by the reverse diffusion process.
- FIG. 7C The forward diffusion process is applied to the latent image embeddings to produce a series of noisy image embeddings.
- FIG. 7D The VLM BiomedCLIP, trained on the PMC-15M corpus, includes 15 million biomedical figure-caption pairs extracted from over 3 million biomedical research articles. BiomedCLIP aims to align images and text descriptions within a shared latent space.
- FIG. 7E The rank-based inference of GenVinci. During the inference, GenVinci can generate multiple images from the same gene expression data and select the ‘best’ one with the highest normalized cosine similarity score between the input gene expression and generated images.
- FIG. 8 is a schematic drawing showing an example of the CNN baseline architecture (st-HBC dataset).
- the CNN baseline model can predict cellular images from the original gene expression data.
- the kernel size in the convolutional layers and the number of nodes in the fully connected layers are adjusted according to the input dimension (number of genes) and the desired output image size across different datasets.
- FIG. 9 is a schematic workflow of the matrix-to-slide GenVinci.
- the cell-to-tile GenVinci is extended to a matrix-to-slide version by transferring the weights from both pre-trained LLM and generative model.
- the pre-trained LLM is frozen to extract gene expression embeddings from the matrix.
- a new projector is implemented to transform the matrix embeddings into a slide-level transcriptomic prompt, which is sent to the pre-trained diffusion model to generate whole-slide images.
- Other modules remain the same as in the cell-to-tile GenVinci.
- FIG. 10A-FIG. 10C are fluorescent and electron microscopy images and graphs showing paired images.
- FIG. 10A The adjacent DAPI images (from MERFISH) and EM images are aligned into the common coordinate system using STcEM.
- FIG. 10B The cells in MERFISH and EM images are paired using the nearest-neighbor searching method; paired MERFISH and EM images are connected with a grey line.
- FIG. 10C The cells in MERFISH and EM images are paired using the Hungarian algorithm.
- FIG. 11 A and FIG. 11B are clustering maps and comparisons of clustering performance graphs.
- FIG. 11 A Left. The clustering map of the original image embeddings extracted using an ImageNet- pretrained ResNet-50 model. Middle. The combination of the original image embeddings with paired gene expression embeddings. Right. The concatenation of the GenVinci-refined gene expression embeddings with the GenVinci-generated image embeddings.
- FIG. 11 B The comparisons of clustering performance on normalized mutual information (NMI), average silhouette width (ASW), and adjusted Rand index (ARI).
- NMI normalized mutual information
- ASW average silhouette width
- ARI adjusted Rand index
- FIG. 12 is a schematic drawing showing a pipeline for the zero-shot annotation task.
- the text prompts are first created based on the categories of interest. Text embeddings are then extracted from the VLM text encoder. At the same time, the gene expression data is sent into GenVinci to extract cell embeddings to compare the similarities with text embeddings.
- the cell embeddings can be extracted from the output of the projector in the VLM supervision module or the generated cellular image embeddings (extracted by the VLM image encoder).
- FIG. 13 is a schematic drawing showing a pipeline for data augmentation.
- GenVinci was initially trained on paired single-cell gene expression data (MERFISH) and H&E images.
- MEFISH paired single-cell gene expression data
- an external, well- annotated scRNA-seq dataset is then fed into the well-trained GenVinci to generate H&E images.
- These generated H&E images are paired with well-annotated cell types from the scRNA-seq dataset and added as data augmentation for the downstream annotation task.
- Described herein are computer implemented methods of generating a multi-dimensional (e.g., two-dimensional or three-dimensional) image of a cell or organelle from single cell sequencing data.
- the methods described herein generally include providing a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle and using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multidimensional image of the cell or organelle (e.g., a nucleus, nucleolus, ribosome, vesicle, endoplasmic reticulum (ER; e.g., rough ER or smooth ER), Golgi apparatus, vacuole, centriole, lysosome, or mitochondria).
- a gene expression profile e.g., from sequencing and/or hybridization data
- a machine learning algorithm model e.g., a machine learning algorithm model to determine the multidimensional image of the cell or organelle (e.g.
- the methods described herein may also be used to generate multi-dimensional images of a plurality of cells, e.g., a tissue.
- the methods and systems are particularly useful for showing complex cell morphologies, particularly of dynamic cell types (e.g. neurons).
- the methods are highly adaptable to work with a variety of cellular imaging techniques, such as electron microscopy, hematoxylin and eosin (H&E) imaging, fluorescence microscopy, 4',6-diamidino-2-phenylindole (DAPI), and immunofluorescence (IF). This is particularly advantageous for streamlining typically expensive and low throughput imaging techniques, such as electron microscopy.
- H&E hematoxylin and eosin
- DAPI 4',6-diamidino-2-phenylindole
- IF immunofluorescence
- Also featured herein are computer implemented systems including, for example, a storage device and a processor communicatively coupled to the storage device.
- the processor may be configured to execute code instructions stored on the storage device to cause the system to perform the computer implemented methods as described herein, e.g., of receiving a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multi-dimensional (e.g., two-dimensional or three-dimensional) image of the cell or organelle.
- a gene expression profile e.g., from sequencing and/or hybridization data
- the computer program product may include a non-transitory computer-readable storage device having computer-executable program instructions embodied thereon that when executed by a computer cause the computer to determine a multi-dimensional (e.g., two-dimensional or three-dimensional) image of a cell or an organelle.
- the computer-executable program instructions may include, for example, computer-executable program instructions to receive a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and computer-executable program instructions to determine the multi-dimensional image of the cell or organelle using the gene expression profile.
- the computer program product may be programmed with a machine learning algorithm model to determine the multidimensional image of the cell or organelle.
- the software program product may include, for example, a computer-executable software program that includes instructions that when executed by a computer cause the computer to determine a multi-dimensional image of a cell or an organelle.
- the computer-executable program instructions may include computer-executable software program instructions to receive a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and computer-executable software program instructions to determine the multi-dimensional image of the cell or organelle using the gene expression profile.
- the software program product may be programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
- the methodology described herein includes a generative model (in some instances herein referred to as “GenVinci”) with a machine learning algorithm that learns a multi-modal representation of morphological and molecular profiles and uses it for gene-to-image translation across various profiling platforms and imaging techniques.
- the model s cascaded architecture allows one to flexibly apply it across a range of downstream tasks, including image generation, multi-modal clustering, zero-shot classification, and data augmentation.
- the model can be trained in two stages: an initial general-purpose pretraining on expression data, followed by fine-tuning with paired imaging data for the gene-to-image translation task. In the pretraining step, a single-cell large language model can be trained to encode high-dimensional expression data into robust latent embeddings.
- GenVinci was pre-trained on GenCorpus-21 OM, a large- scale scRNA-seq corpus of over 210 million scRNA-seq and spatial transcriptomics profiles from a broad range of tissues across human and mouse samples.
- the pre-training module can be further pre-trained on additional gene expression data, e.g., using the masked language modeling strategy, where some genes or gene expression values are randomly masked in a cell and then predicted.
- the pre-trained model can be cascaded with a diffusion model (e.g., Stable Diffusion (SD)), a diffusion model where the extracted gene expression embeddings can serve as transcriptomic prompts to generate cell images.
- a diffusion model e.g., Stable Diffusion (SD)
- SD Stable Diffusion
- the diffusion model is specifically designed for text-to-image generation and may be trained on billions of image-text pairs, it may not generalize well with limited gene expression-image pairs.
- another biomedical vision-language model such as BiomedCLIP, may be employed to provide weak supervision to reduce modality discrepancies for better aligning the transcriptomic prompts with image and text embeddings.
- BiomedCLIP can extract multi-modal cell image embeddings that are well-matched with their potential text embeddings.
- the enhanced cell embeddings align more closely with both image and potential text embeddings, rendering them more effective for fine-tuning the diffusion model for cell image generation, thereby improving the quality of the resulting cell images.
- the method can be programmed to generate multiple cell images from a single expression profile or signature, reflecting various perspectives of the cell morphology, which is of great significance even when only partial or individual sampling of gene expression distributions is available, rather than a complete profile (e.g., with methods like multiplexed error-robust fluorescence in situ hybridization (MERFISH)) or full distributions representing cell states. It can then rank the generated images that are closest to the ground truth based on similarity scores between the input expression profile and the generated images. This flexible, rank-based approach to inference can be highly effective in real-world clinical settings when pathologists or physicians want to avoid incorrect predictions and choose the most confident generation.
- MEFISH multiplexed error-robust fluorescence in situ hybridization
- the cascaded architecture of the methods described herein overcomes the limitation of having few expression-image pairs for fine-tuning and enhances the cell representations obtained from the pretraining step for more effective image generation. Furthermore, by including large language models, GenVinci can be flexibly applied to diverse downstream tasks, leveraging the models’ capabilities in a multi-modal manner.
- GenVinci is a deep generative model that can generate cellular morphology from single-cell gene expression. It can include three modules: an encoder model (e.g., a Bidirectional Encoder Representations from Transformers (BERT)-based single-cell language model), a diffusion model (e.g., a Latent Diffusion Model (LDM)), and a pretraining model (e.g., a biomedical Contrastive Language-Image Pretraining (CLIP) model).
- an encoder model e.g., a Bidirectional Encoder Representations from Transformers (BERT)-based single-cell language model
- a diffusion model e.g., a Latent Diffusion Model (LDM)
- LDM Latent Diffusion Model
- pretraining model e.g., a biomedical Contrastive Language-Image Pretraining (CLIP) model.
- LDM module maps single-cell gene expression embeddings to cellular images. LDM approaches typically accept natural language embeddings, which are
- GenVinci includes a single-cell language model that produces gene expression embeddings in a manner similar to natural language models.
- the representations learned by GenVinci are unaligned to the pretraining of the LDM, BiomedCLIP, a model that learns to align biomedical figure captions to their corresponding images, was employed to supervise the fine-tuning of the entire approach.
- the integration of three modules allows GenVinci not only to generate realistic cellular images but also to be easily applied for diverse downstream tasks. These modules were introduced along with their tailored pretraining and fine-tuning strategies for GenVinci, as described in the subsequent section.
- a singlecell language model can be trained to encode gene expression data into a low-dimensional latent space.
- GenVinci is a BERT-like transformer model pre-trained on a large-scale corpus (GenCorpus-21 OM) of approximately 210 million single-cell and spatial transcriptome to allow context-aware predictions in settings with limited data in network biology.
- the single-cell language model can accept single-cell gene expression data from different sequencing measurements without the need to standardize into uniform gene lists.
- domain knowledge learned from an extensive single-cell transcriptomics corpus, particularly concerning genegene interactions can enhance cellular representation across a spectrum of downstream tasks.
- GenVinci is a transformer-based approach that expects inputs to be sequences
- a binning- based rank value encoding method can be employed to induce a sequence from the (tabular) gene expression observations.
- Rank value encoding is a non-parametric approach that ranks genes by their expression levels within a single cell. Initially, given a cell-by-gene matrix, only protein-coding genes may be retained based on a predefined vocabulary by GenCorpus-21 OM. Subsequently, the gene expression data can first be normalized using a scale factor, followed by a log transformation to bring them to a comparable scale. The bin size is set to be 50.
- x ⁇ represent the expression value for a gene i in cell c
- genes within each specific cell can be assigned into binning values by their normalized expression values using a quantile binning method and ranked in descending order and subsequently tokenized according to the predefined vocabulary, where each gene g k is assigned an integer identifier ic/( ⁇ fc ).
- the vocabulary can also include additional (e.g., two or more) special tokens for padding and masking and context tokens to include the cell context information (e.g., modality and assay).
- the binning-based rank value encoding method prioritizes genes that uniquely characterize cell states by using expression data across the GenCorpus- 210M, thereby downgrading common housekeeping genes and elevating critical transcription factors that define cell identity. It can offer robustness against technical biases, maintaining the relative gene ranking within each cell despite potential variations in absolute transcript counts. Furthermore, top m expressed genes with the normalized expression value x' > 0 for can be selected for each cell. For the cells with the number of expressed genes less than m, padding tokens can be added to meet the required length. In one example, m can be set to be 1024 (which fully represents more than 90% of rank value encodings in GenCorpus-210M).
- the final embeddings of the tokenized genes for cell c can be denoted as: where the ranked genes are represented by their integer identifiers, and Emb g is a learnable conventional embedding layer for the gene tokens. Similar to the position embedding in BERT, trainable binning-based expression embeddings denoted as EE C can be used for each gene token within the cell, which allows the modeling of the relationship between the binned gene expression values. GenVinci also includes the conditional embedding to distinguish the context tokens and gene tokens which is denoted as CE c .
- the final input embedding for each cell can be represented as the element-wise sum of the gene and binning- based expression embeddings:
- GeneVinci can be composed of six identical transformer encoders, each featuring, for example, two sub-layers.
- the first sub-layer can be a multi-head self-attention layer
- the second sub-layer can be a simple position-wise feed-forward neural network layer.
- the scaled dot-product attention can be implemented in the self-attention layer as:
- W° are the parameter matrices of projections.
- the linear layers may share the same architecture but have different parameters across various positions.
- 15% of the genes or gene expression values within each specific cell were randomly masked in spatial transcriptomics data, and the model can be trained to predict the masked genes or gene expression values based on the context of the remaining unmasked spatial genes or gene expression values.
- the cross-entropy loss can be employed as the masked learning objective function, which for a single masked token can be formulated as: where V is the size of the vocabulary. y t and ft are the ground truth and predicted probabilities of the i-th gene token, respectively.
- masked language modeling loss can be computed for all masked gene or gene expression tokens within a batch, and this loss can then be averaged to produce the final loss against which the model parameters can be updated.
- GenVinci Transformers instead of a single feedforward network like in standard Transformers, GenVinci Transformers use Mixture of Experts which consists of multiple expert FFNs, with a gating network that dynamically selects a few experts for each input token. This approach allows the model to scale efficiently, increasing capacity while keeping computational costs lower than a fully dense model.
- SD may include two modules: (1 ) a regularized autoencoder that compresses images into lowdimensional latent features and reconstructs them, and (2) a U-Net-based denoising model with attention mechanisms for more flexible conditional image generation.
- the t-th noisy embedding z t is sampled from the following distribution: where I is the identity matrix, is a variance schedule parameter that determines how much noises are added at each step. These noisy embeddings can then serve as the training data for the reverse diffusion process of SD.
- Conditional information can be applied through cross-attention heads in the attentionbased U-Net. For example, to match the dimensions of the SD inputs, the outputs y of the gene expression encoder can be further projected with a projector x e into a transcriptomic prompt r e (y) e aim t 0
- the gene expression encoder and the U-Net can be optimized together.
- the following loss function may be used: where e g is the denoising function implemented as a U-Net.
- the diffusion model can take the Gaussian noise as the initial input and denoise it iteratively through T steps.
- the reverted image latent embedding z 0 can be obtained and input into the decoder ⁇ (•) to generate the cellular images.
- GenVinci gene expression encoder By jointly training GenVinci gene expression encoder and the SD model, one can fine-tune the projector to get better transcriptomic prompts for cellular image generation tasks.
- the original SD model may be tailored for the text-to-image generation task, the unique properties of gene expression data result in latent spaces that are quite different from those of texts. This means that simply fine-tuning the SD model end-to-end with a limited set of gene expression-image pairs may make it difficult to yield an effective alignment between the gene expression and text embeddings.
- BiomedCLIP a biomedical vision-language foundation model trained on the PMC-15M corpus
- the PMC-15M corpus includes 15 million figure-caption pairs extracted from over 3 million biomedical research articles in PubMed Central, which includes a diverse range of biomedical images.
- previous gene expression projector with an additional alignment projector can be followed to refine the transcriptomic embeddings into multi-modal pseudo vision-language embeddings. These embeddings can be adjusted to be either dimensionally or distributionally compatible with those used in BiomedCLIP.
- a weak supervised constraint may be applied to minimize the similarity between the transcriptomic embeddings and the multi-modal embeddings produced by the BiomedCLIP image encoder, which is defined as follows: where h represents the alignment projector, I represents the cellular image sent to the BiomedCLIP image encoder E,.
- the loss function can serve to encourage closer alignment between the transcriptomic embeddings and the cellular image embeddings, and by extension, it also can align with the potential textual description features of the images. As a result, the spaces for transcriptome, images, and texts can be unified, and the optimized transcriptomic representation can become more suitable for SD image generation, thereby improving the quality of the generated images.
- the parameters of the pre-trained BiomedCLIP model can remain fixed.
- the entire loss function can be formulated as:
- the goal can be to generalize GenVinci from a cell-to-tile level to a matrix-to-slide level.
- the process may begin with resizing whole-slide images, e.g., to n x n pixel images, e.g., 512 x 512-pixel images, to match the output dimensions of the SD model.
- cells with insufficient information e.g., those with a total count of fewer than 200, can be removed.
- n (e.g., 512) cells can be randomly sampled to represent the tissue transcriptome signals.
- matrices contain fewer than n (e.g., 512) cells, they can be padded with zero- expressed genes to maintain consistent dimensions. This approach is adaptable across various imaging and sequencing technologies, accommodating variations in the number of captured cells.
- the pretrained cell embeddings can be extracted for each cell within the gene expression matrix by using the gene expression encoder of cell-to-tile pre-trained GenVinci. These embeddings can be passed through a new projector to reshape the matrix to fit the requirements of the SD model. The projector can be further fine-tuned, e.g., in conjunction with the U-Net component of the SD model. All other modules and parameters may remain consistent with those used in the cell-to-tile GenVinci.
- Various evaluation metrics may be employed to assess the accuracy of the GenVinci model as described below, e.g., in order to determine if the output image is consistent with the input gene expression profile and empirical databases of two-dimensional and three-dimensional images.
- a Convolutional Neural Network can be implemented as the baseline for the gene-to- image translation task to compare with GenVinci.
- the CNN baseline is inspired by the decoder network of U-Net, which consists of fully connected layers and transpose convolutional layers to increase the spatial dimension, and batch normalization and ReLU activations to enhance the model's ability to learn nonlinear functions.
- the CNN baseline can be trained to predict cellular images from the normalized gene expression using MSE loss.
- the top 2000 highly variable genes (HVGs) can be selected as input for the datasets with whole transcriptomic sequencing.
- the model can be trained for 500 epochs with the Adam optimizer, using a learning rate of 0.001 and a batch size of 32.
- FID measures the difference between the distributions of real and generated images.
- the ImageNet-pre-trained lnception-V3 convolutional neural network can be adopted to extract 2048- dimensional feature vectors for both real and generated images. These feature vectors capture the essential visual information of the images, as learned by the network during its training on ImageNet dataset. Then the mean and covariance can be computed for each set of features to calculate the Wasserstein-2 distance. A lower FID score indicates that the distributions of generated images are closer to the real images, suggesting better quality and realism of the generated images.
- FID is expressed as: where x denotes the feature vectors of the real images, g denotes the feature vectors of the generated images. Tr denotes the trace of a matrix, which is the sum of all the diagonal elements.
- Cosine similarity can also be used to measure the similarity between the real and generated images.
- Cosine similarity is a mathematical measure used to determine the similarity between two vectors in a multi-dimensional space.
- the cosine similarity can be calculated by the extracted 2048- dimensional feature vectors of the real and generated images from an ImageNet-pre-trained ResNet-50 model.
- classifiers can be trained based on the real and generated images for cell-type or pathological annotations.
- the image features can be extracted from an ImageNet-pre-trained ResNet-50 model with classification heads tailored for different annotation tasks. Both Accuracy and F1 scores can be reported as the evaluation metrics. The higher Accuracy and F1 score may indicate better classification performance.
- one or more commonly used metrics can be adopted, such as normalized mutual information (NMI), adjusted Rand index (ARI), and average silhouette width (ASW).
- NMI score can be computed to measure the concurrence between ground truth cell type labels and Louvain cluster labels from integrated cell embeddings, with clustering resolutions ranging from 0.1 to 2.
- the ARI can be used to assess the agreement between annotated labels and NMI-optimized Louvain clusters, adjusting for randomly correct labels.
- the ASW score ranging from -1 to 1 , can be calculated to evaluate the relationship between a cell's within-cluster distances and its distances to the closest cluster boundaries. Higher scores for NMI and ARI may indicate better matches, while an ASW score of 1 may indicate well- defined clusters.
- the same hyperparameters as the Geneformer can be used during the pretraining of GenVinci.
- a linear learning rate scheduler with warmup can be employed, setting the maximum learning rate, e.g., to 5 x 10 -5 .
- An Adam optimizer can be used, e.g., with a weight decay of 0.001 and a batch size of 16.
- the dimensionality of the input and output can be, for example, 256, and the inner layer can have a dimensionality of, e.g., 512.
- a nonlinear activation function may be utilized in the form of a Rectified Linear Unit (ReLU); a dropout probability of 0.02; a standard deviation of 0.02 for the initializer of weight matrices; and an epsilon value of 1 x 10 -12 for layer normalization layers.
- ReLU Rectified Linear Unit
- SD version 1 .5 can be utilized for image generation and denoising diffusion implicit model (DDIM) sampling strategy, with a batch size of, e.g., 4.
- the Adam optimizer may be employed, setting the learning rate to, e.g., 5 x 10 -5 , for training.
- Each model can be fine-tuned for 500 epochs, and the number of sampling steps can be set to 250.
- the parameters of the BiomedCLIP model may always remain fixed.
- the methods described herein include providing a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from a cell or organelle.
- the gene expression profile may be obtained from one or more known sequencing and/or imaging methodologies known in the art, e.g., as described in more detail below.
- the spatial transcriptomic analysis includes single nucleus RNA (snRNA) sequencing, single cell RNA (scRNA) sequencing, patch sequencing, Visium, Xenium, NanoString GeoMx, STARmap, Slide-Seq, Slide-Tags, fluorescence in situ hybridization (FISH), multiplexed error-robust fluorescence in situ hybridization (MERFISH), electron microscopy, hematoxylin and eosin (H&E) imaging, fluorescence microscopy, 4',6-diamidino-2-phenylindole (DAPI), immunofluorescence (IF), or a combination thereof.
- snRNA single nucleus RNA
- scRNA single cell RNA
- patch sequencing Visium, Xenium, NanoString GeoMx, STARmap, Slide-Seq, Slide-Tags
- FISH fluorescence in situ hybridization
- MEFISH multiplexed error-robust fluorescence in situ hybridization
- H&E
- Example 1 Generation of cellular and tissue images from single-cell expression profiles with GenVinci
- GenVinci a generative model based on transformer and diffusion models that learns a multi-modal representation of cells and tissues and reconstructs morphological information from transcriptome profiles to generate cellular and tissue images from single-cell expression profiles during inference. It was demonstrated how GenVinci can generate diverse types of imaging data, including histological (hematoxylin and eosin (H&E)) stains, fluorescence microscopy images, neuron morphology, and electron microscopy, from various molecular profiles at different resolutions.
- H&E histological hematoxylin and eosin
- GenVinci generates accurate cellular and tissue images compared to ground-truth morphologies, cell types, and pathological annotations. It outperforms a convolutional neural network (CNN) baseline in both generation quality and semantic mapping. Thanks to its cascaded architecture, GenVinci can be flexibly applied to a range of downstream tasks, including image generation, multi-modal clustering, zero-shot classification, and data augmentation. GenVinci can unify different views of cell and tissue biology, while greatly reducing the need for multiple measurements, towards an ultimate goal of virtual cell and tissue simulators.
- CNN convolutional neural network
- complex neuronal morphologies are highly cell-type-specific, essential for their function, and are associated with genomic mutations, gene expression levels, and methylation patterns. Therefore, linking molecular information with cell morphology would provide key insights for further studying the important functions of cells such as the formation of neural circuits, stem cell differentiation, and cancer metastasis.
- Generative artificial intelligence (Al) methods offer a compelling new approach to tackle this challenge. Al-generated content has made a profound impact across a wide range of fields.
- Text-to- image generation models like DALL-E 2, Stable Diffusion, and Midjourney have created complex and realistic images based on textual input, surpassing all previous generative adversarial network (GAN) and variational autoencoder (VAE)-based methods in quality metrics and synthesis capabilities.
- GAN generative adversarial network
- VAE variational autoencoder
- Sora also creates realistic and imaginative scenes from text instructions and serves as a foundation to understand and simulate the physical world. All these methods rely on diffusion models, a novel paradigm that is different from previous GAN and VAE-based approaches and generates images through the systematic introduction of noise with repeating steps.
- GenVinci a single-cell gene-to-image generative model, that bridges the gap between expression profiles and cell morphology.
- GenVinci pretrains a state-of-the- art single-cell language model to encode gene expression into a robust latent representation and then fine-tunes a text-to-image diffusion model to generate cellular images based on the latent transcriptomic prompt.
- GenVinci further employs a biomedical vision-language model to provide supervision, aligning transcriptome, images, and texts into the shared latent space. This enhances our understanding of cell states and transitions and should help translate molecular profiles into physiologically and clinically interpretable results.
- GenVinci is a generative model that learns a multi-modal representation of morphological and molecular profiles and uses it for gene-to-image translation across various profiling platforms and imaging techniques (FIG. 1A). GenVinci’s cascaded architecture allows one to flexibly apply it across a range of downstream tasks, including image generation, multi-modal clustering, zero-shot classification, and data augmentation (FIGS. 1A, 1C, and 1D).
- GenVinci model was trained in two stages: an initial general-purpose pretraining on expression data, followed by fine-tuning with paired imaging data for the gene-to-image translation task (FIGS. 1B, 1C, and 1D).
- GenVinci trains a single-cell large language model to encode high-dimensional expression data into robust latent embeddings.
- GenVinci was pre-trained on GenCorpus-210M, a large-scale scRNA-seq corpus of 210 million human and mouse single-cell and spatial RNA-seq profiles from a broad range of tissues.
- GenVinci can be further pre-trained on additional relevant gene expression data, using the original masked language modeling strategy, where some genes or gene expression values are randomly masked in a cell and then predicted (Methods, FIGS. 1C and 7A).
- the pre-trained Transformer was cascaded with Stable Diffusion (SD), a diffusion model where the extracted gene expression embeddings serve as transcriptomic prompts to generate cell images (Methods, FIGS. 1D, 7B, and 7C). Since the SD model is specifically designed for text-to-image generation and was trained on billions of image-text pairs, it may not generalize well with limited gene expression-image pairs. Thus, another biomedical vision-language model, BiomedCLIP, was employed to provide weak supervision to reduce modality discrepancies for better aligning the transcriptomic prompts with image and text embeddings (Methods, FIGS. 1D and 7D).
- BiomedCLIP biomedical vision-language model
- BiomedCLIP extracts multi-modal cell image embeddings that are well-matched with their potential text embeddings. These are then used to refine the cell embeddings from GenVinci. Consequently, the enhanced cell embeddings align more closely with both image and potential text embeddings, rendering them more effective for fine-tuning the SD model for cell image generation, thereby improving the quality of the resulting cell images.
- GenVinci generates multiple cell images from a single expression profile or signature, reflecting various perspectives of the cell morphology, which is of great significance even when only partial or individual sampling of gene expression distributions is available, rather than a complete profile (e.g., with methods like multiplexed error-robust fluorescence in situ hybridization (MERFISH)) or full distributions representing cell states. It then ranks the generated images that are closest to the ground truth based on similarity scores between the input expression profile and the generated images. This flexible, rank-based approach to inference can be highly effective in real- world clinical settings when pathologists or physicians want to avoid incorrect predictions and choose the most confident generation.
- MEFISH multiplexed error-robust fluorescence in situ hybridization
- GenVinci s cascaded architecture overcomes the limitation of having few expression-image pairs for fine-tuning and enhances the cell representations obtained from the pretraining step for more effective image generation. Furthermore, by including large language models, GenVinci can be flexibly applied to diverse downstream tasks, leveraging the models’ capabilities in a multi-modal manner.
- GenVinci s ability to generate Hematoxylin-and-Eosin (H&E) histology images from single cell (sc)/ or single nucleus (sn) RNA-seq of metastatic breast cancer (MBC) tumors from the Human Tumor Atlas Project Pilot (HTAPP) (Datasets).
- sc/snRNA-seq single cell tumor
- HTAPP Human Tumor Atlas Project Pilot
- a single tumor specimen was split such that one part was profiled by sc/snRNA-seq and another was sectioned serially, with consecutive sections spatially profiled by MERFISH (212 genes) or stained by H&E.
- GenVinci’s generalization ability was evaluated at two imaging resolutions.
- GenVinci was compared to a “CNN baseline” model consisting of several transposed convolutional layers inspired by U-Net decoders (Methods, FIG. 8), trained to predict cell images from expression data using mean squared error (MSE) loss. The image with the highest similarity score was selected among the five generated images for benchmarking comparison.
- CNN baseline consisting of several transposed convolutional layers inspired by U-Net decoders (Methods, FIG. 8), trained to predict cell images from expression data using mean squared error (MSE) loss.
- MSE mean squared error
- GenVinci-generated images were qualitatively (by visual inspection) and quantitatively very similar to the ground-truth images (FIGS. 2A-2F). Qualitatively, GenVinci images reflected nuances of complex shape, cell intensity patterns, color distribution, and even tissue profile across all resolutions (FIGS. 2A-2E, in each figure, the first column on the left contains the ground-truth images), while CNN baseline could not generate realistic H&E images, especially multi-cellular (tile level) ones, and single-cell images (on sn-MBC) were blurred (FIG. 2F).
- GenVinci images had much smaller distribution differences with ground-truth images than CNN baseline by the Frechet Inception Distance (FID; a standard metric for GAN models that measures the distance between the distribution of generated and ground-truth images) (FIG. 2G, GenVinci FIDs less than 100 in most cases, vs. more than 200 for CNN baseline), and by average cosine similarity (between the ground-truth and generated images (Methods), FIG. 2G, with similarity score > 0.8 in all cases, vs. less than 0.5 for the CNN baseline).
- FID Frechet Inception Distance
- Uniform Manifold Approximation and Projection of the image embeddings from the ground-truth, GenVinci, and CNN baseline images (FIG. 2H) also show good agreement between ground truth and GenVinci (but not CNN baseline) distributions for Visium-HBC and st-HBC datasets, and less so for sc-MBC (but still better than CNN baseline).
- GenVinci generates fluorescence cell morphology
- GenVinci was applied on a CosMx SMI dataset (with each cell profiled by 960 genes) from one non-small cell lung cancer (NSCLC) tissue sample to generate fluorescence microscopy images from single-cell gene expression data (Datasets).
- NSCLC non-small cell lung cancer
- Datasets single-cell gene expression data
- snGE single-nucleus DAPI images from the genes measured in those segmented nuclei
- IF immunofluorescence
- one channel represents DAPI
- one channel represents PanCK (an epithelial cell marker)
- one channel represents CD45 (a marker of all nucleated hematopoietic cells)
- one channel represents CD3 (a T cell marker). Since there is only one tissue sample in this dataset, in both cases, the gene expression-image pairs were randomly split from two fields of view (FOVs) for testing and retained the remaining 28 FOVs for training.
- FOVs fields of view
- GenVinci recovered the fluorescence intensity patterns and generated realistic single-nucleus DAPI images from single-nucleus gene expression data and even recovered the distribution of different marker stains in single-cell IF images from the gene expression information (FIGS. 3A and 3B).
- GenVinci showed superior generation quality and distribution similarity (by FID, FIG. 3D, top) and higher cosine similarity between generated and ground-truth images (FIG. 3D, bottom), and high overlap of image embeddings (FIG. 3E), demonstrate GenVinci’s ability to generate well across different data modalities. Note that across all metrics, GenVinci performed less well on the DAPI image generation task versus the IF-stained image generation task.
- GenVinci generates whole-neuron morphology
- GenVinci generated highly realistic neuron morphologies that reflect the provided Patch-seq gene expression and ground-truth images qualitatively and quantitatively (FIGS. 4A-4C), while CNN baseline could not (FIGS. 4B-4D).
- GenVinci s performance in generating electron microscopy (EM) images, a specialized data modality that is challenging to generate experimentally, was next tested.
- EM imaging provides a nanometer-resolution view of tissue ultrastructure, but its acquisition is both expensive and timeconsuming, and generating it through simpler profiling modalities could be transformative.
- To fine-tune GenVinci recently collected data was leveraged, where large area scanning EM and MERFISH were collected in two consecutive sections from the brain of a mouse with locally induced injury, followed by spatial alignment and projection into a common coordinate system (FIG. 10A).
- the EM images and MERFISH profiles do not match perfectly in the coordinate space and when aligned with a nearest neighbor search, different EM images were occasionally mapped to the same MERFISH profile, leading to loss of more than half of the gene expression-image pairs (FIG. 10B).
- the Hungarian algorithm was employed, thus ensuring an optimal solution to the assignment problem by finding a one-to-one correspondence between points in two sets by minimizing the total cost across all pairings. This approach allowed achievement of a one-to-one pairing, although some EM images did not structurally match their assigned profiles (FIG. 10C).
- GenVinci still performed well (FIGS. 5A and 5D). This may be beneficial from the supervision of BiomedCLIP, which incorporated EM-text pairs in its training corpus. Furthermore, GenVinci generates far more accurate EM images compared with the CNN baseline (FIGS. 5B-5D).
- GenVinci uncovers multi-modal information from unseen embeddings
- GenVinci learns multi-modal embeddings through generative training
- the cell/spot embeddings extracted from Geneformer (a recently published single-cell foundation model), continued pre-trained Geneformer, and gene expression-image fine-tuned GenVinci were visualized and compared (FIG. 6A).
- Geneformer a recently published single-cell foundation model
- Geneformer continued pre-trained Geneformer
- gene expression-image fine-tuned GenVinci were visualized and compared (FIG. 6A).
- st-HBC, sn-MBC, and sn-NSCLC datasets were investigated.
- GenVinci improved the identification of cell types compared to the original embeddings (FIGS. 6A and 11 B).
- NMI normalized mutual information
- ASW average silhouette width
- ARI adjusted Rand index
- GenVinci learned multi-modal embeddings from the gene-to-image task that helped distinguish tissue pathology and cell types. For example, tumor and normal spots were separated in the st-HBC dataset, while tumor, endothelial fibroblast, and plasmablast cells were separated in the sn- NSCLC dataset (FIG. 6A, right, FIG. 11B, upper). However, GenVinci also lost some cell-type information when training to translate the gene expression data into cellular images (e.g., hepatocytes in the sn-MBC dataset were clustered in GenVinci but scattered in the GenVinci embedding).
- GenVinci's potential in integrating single-cell expression data and image features was also explored. Specifically, the original gene expression embeddings were concatenated with ground-truth image embeddings and compared with GenVinci’s refined gene expression embeddings concatenated with its generated image embeddings (FIGS. 11 A, middle and right, FIG. 11B, bottom). Compared to the original gene expression and image features, the refined gene expression embeddings paired with generated image features captured more heterogeneity within the cell types or pathological annotations (FIG. 11 A, right and FIG. 11B, bottom).
- malignant cell spots were clustered together in the st-HBC dataset, and endothelial cells and hepatocytes were separated in the sn-MBC dataset.
- the clusters of endothelial, epithelial, and fibroblast cell types were better delineated in GenVinci embeddings of the sn-NSCLC dataset.
- GenVinci infers its annotation by calculating the similarity between cell embeddings and textual prompt embeddings without additional training.
- the cell embeddings can be represented by gene expression embeddings extracted from the projector or by the generated image embeddings from the vision-language model (VLM) image encoder (FIGS. 1D and 12). This is similar to the zero-shot image classification in the CLIP model.
- VLM vision-language model
- BiomedCLIP is trained with billions of biomedical image-text pairs in the PMC-15M corpus, this was solely focused on general annotation tasks evaluated in the BiomedCLIP paper, i.e. , tjhezero-shot-based pathological annotation task on datasets with H&E images were exclusively tested.
- GenVinci was used to generate the same number of training images for comparison.
- a ResNet-50 model was used to extract image features and implement a linear probing classifier for the pathological annotation task. The sampling frequency was varied to generate different numbers of images by GenVinci and trained the model to compare their performance to that of real images (FIGS. 6D-6F).
- Pathological annotations were obtained from the single-cell gene expression data and used as labels (at single cell level) for both ground-truth and generated images.
- GenVinci the first generative model designed to translate single-cell gene expression into various cellular and tissue images.
- GenVinci represents a significant leap in realizing digital biology, aiming to bridge the gap between the complexity of gene expression data and the interpretability offered by traditional imaging techniques.
- GenVinci addresses this issue by generating visual representations of cellular and tissue structures from transcriptomic data, thereby enhancing its interpretability and potential utility in clinical diagnostics and research. Its performance was evaluated from a range of sequencing and imaging measurements at different resolutions. GenVinci shows strong generalization and robustness.
- GenVinci offers a novel conceptual and technological framework for visualizing unseen information in gene expression data, facilitating a more integrated understanding of cellular function. GenVinci introduced a conceptual shift in generating new and hidden multi-modal biological information without measuring all modalities.
- GenVinci is a deep generative model that can generate cellular morphology from single-cell gene expression. It consists of three modules: a Bidirectional Encoder Representations from Transformers (BERT)-based single-cell language model, a Latent Diffusion Model (LDM), and a biomedical Contrastive Language-Image Pretraining (CLIP) model.
- LDM Low Density Multimedia Subsystem
- LDM approaches typically accept natural language embeddings, which are then translated into corresponding images. Because of this, GenVinci trains a single-cell language model that produces gene expression embeddings in a manner similar to natural language models.
- GenVinci trains a single-cell language model to encode gene expression data into a low-dimensional latent space.
- GenVinci is a BERT-like transformer model pre-trained on a large-scale corpus (GenCorpus-210M) of approximately 210 million single-cell and spatial transcriptome to allow context- aware predictions in settings with limited data in network biology.
- GenVinci can accept single-cell gene expression data from different sequencing measurements without the need to standardize into uniform gene lists.
- domain knowledge learned from an extensive single-cell transcriptomics corpus particularly concerning gene-gene interactions, enhances cellular representation across a spectrum of downstream tasks.
- GenVinci is a transformer-based approach that expects inputs to be sequences
- a binning- based rank value encoding method is employed to induce a sequence from the (tabular) gene expression observations.
- Rank value encoding is a non-parametric approach that ranks genes by their expression levels within a single cell. Initially, given a cell-by-gene matrix, only protein-coding genes was retained based on a predefined vocabulary by GenCorpus-210M. Subsequently, the gene expression data is first normalized using a scale factor, followed by a log transformation to bring them to a comparable scale. The bin size is set to be 50.
- x ⁇ represent the expression value for a gene i in cell c
- genes within each specific cell were assigned into bins using a quantile binning method and ranked by their binned values in descending order and subsequently tokenized according to the predefined vocabulary, where each gene g k is assigned an integer identifier id g k ).
- the vocabulary also includes two more special tokens for padding and masking and context tokens to include the cell context information (e.g., modality and assay).
- the rank value encoding method prioritizes genes that uniquely characterize cell states by using expression data across the GenCorpus-21 OM, thereby downgrading common housekeeping genes and elevating critical transcription factors that define cell identity. It offers robustness against technical biases, maintaining the relative gene ranking within each cell despite potential variations in absolute transcript counts. Furthermore, top m expressed genes with the normalized expression value x' > 0 for were selected for each cell. For the cells with the number of expressed genes less than m, padding tokens are added to meet the required length. In this study, m is set to be 1024 (which fully represents 90% of rank value encodings in GenCorpus-21 OM).
- the final embeddings of the tokenized genes for cell c are denoted as: where the ranked genes are represented by their integer identifiers, and Emb g is a learnable conventional embedding layer for the gene tokens. Similar to the position embedding in BERT, trainable binning-based expression embeddings denoted as EE C can be used for each gene token within the cell, which allows the modeling of the relationship between the binned gene expression values. GenVinci also includes the conditional embedding to distinguish the context tokens and gene tokens which is denoted as CE c .
- the final input embedding for each cell can be represented as the element-wise sum of the gene and binning- based expression embeddings:
- GenVinci is composed of six identical transformer encoders, each featuring two sub-layers.
- the first one is a multi-head self-attention layer, followed by a simple position-wise feed-forward neural network layer.
- the scaled dot-product attention is implemented in the self-attention layer as:
- V“/c the magnitude of attention weights.
- the output is computed as a weighted sum of the values, where the weight is calculated by a compatibility function of the query with the corresponding key.
- W° are the parameter matrices of projections.
- Two feed-forward neural network layers are followed by the self-attention layers, which are applied to each position separately and denoted as:
- FFN(x)' max (0, x W 1 + b W 2 + b 2
- the linear layers share the same architecture but have different parameters across various positions.
- 15% of the genes or gene expression values within each specific cell were randomly masked in spatial transcriptomics data, and the model was trained to predict the masked genes or gene expression values based on the context of the remaining unmasked spatial genes or gene expression values.
- the cross-entropy loss is employed as the masked learning objective function, which for a single masked token can be formulated as: where V is the size of the vocabulary.
- y t and ft are the ground truth and predicted probabilities of the i-th gene token, respectively.
- masked language modeling loss was computed for all masked gene or gene expression tokens within a batch, and this loss was then averaged to produce the final loss against which the model parameters are updated.
- GenVinci Transformers instead of a single feedforward network like in standard Transformers, GenVinci Transformers use Mixture of Experts which consists of multiple expert FFNs, with a gating network that dynamically selects a few experts for each input token. This approach allows the model to scale efficiently, increasing capacity while keeping computational costs lower than a fully dense model.
- SD consists of two modules: (1 ) a regularized autoencoder that compresses images into lowdimensional latent features and reconstructs them, and (2) a U-Net-based denoising model with attention mechanisms for more flexible conditional image generation.
- the t-th noisy embedding z t is sampled from the following distribution: where I is the identity matrix, is a variance schedule parameter that determines how much noises are added at each step. These noisy embeddings were then served as the training data for the reverse diffusion process of SD. According to previous studies, conditional information was applied through cross- attention heads in the attention-based U-Net. In this study, to match the dimensions of the SD inputs, the outputs y of the gene expression encoder is further projected with a projector T 8 into a transcriptomic prompt r e (y) w jth the aim to learn the reverse diffusion process formulated by During the fine-tuning process, the gene expression encoder and the U-Net were optimized together.
- e 8 is the denoising function implemented as a U-Net.
- the diffusion model took the Gaussian noise as the initial input and denoises it iteratively through T steps.
- the reverted image latent embedding z 0 was obtained and input into the decoder ⁇ (•) to generate the cellular images.
- BiomedCLIP a biomedical vision-language foundation model trained on the PMC-15M corpus, was incorporated to enhance the alignment of spaces related to the transcriptome, images, and text.
- the PMC-15M corpus includes 15 million figure-caption pairs extracted from over 3 million biomedical research articles in PubMed Central, which includes a diverse range of biomedical images.
- previous gene expression projector with an additional alignment projector was followed to refine the transcriptomic embeddings into multi-modal pseudo vision-language embeddings. These embeddings were adjusted to be either dimensionally or distributional ⁇ compatible with those used in BiomedCLIP.
- a weak supervised constraint is applied to minimize the similarity between the transcriptomic embeddings and the multi-modal embeddings produced by the BiomedCLIP image encoder, which is defined as follows: where h represents the alignment projector, I represents the cellular image sent to the BiomedCLIP image encoder E,.
- the loss function serves to encourage closer alignment between the transcriptomic embeddings and the cellular image embeddings, and by extension, it also aligns with the potential textual description features of the images. As a result, the spaces for transcriptome, images, and texts are unified, and the optimized transcriptomic representation became more suitable for SD image generation, thereby improving the quality of the generated images.
- GenVinci to translate gene expression matrices into whole-slide images.
- numerous gene expression-image pairs at the single-cell level there are only a few paired gene expression matrices and whole-slide images available.
- the varying number of cells captured across different tissue slides presents a challenge and limits the feasibility of generating whole-slide images.
- a matrix-to-slide GenVinci was successfully implemented that leverages the previous pre-trained GenVinci by gene expression-image pairs.
- the goal was to generalize GenVinci from a cell-to-tile level to a matrix-to-slide level.
- the process begins with resizing whole-slide images to 512 x 512-pixel images to match the output dimensions of the SD model.
- cells with insufficient information i.e., those with a total count of fewer than 200, were removed.
- 512 cells were randomly sampled to represent the tissue transcriptome signals.
- matrices contain fewer than 512 cells, they were padded with zero-expressed genes to maintain consistent dimensions. This approach is adaptable across various imaging and sequencing technologies, accommodating variations in the number of captured cells.
- the pre-trained cell embeddings for each cell within the gene expression matrix were extracted by using the gene expression encoder of cell-to-tile pre-trained GenVinci. These embeddings were passed through a new projector to reshape the matrix to fit the requirements of the SD model. The projector was further fine tuned in conjunction with the U-Net component of the SD model. All other modules and parameters remained consistent with those used in the cell-to-tile GenVinci.
- CNN Convolutional Neural Network
- FID measures the difference between the distributions of real and generated images.
- the ImageNet-pre-trained lnception-V3 convolutional neural network was adopted to extract 2048- dimensional feature vectors for both real and generated images. These feature vectors capture the essential visual information of the images, as learned by the network during its training on ImageNet dataset. Then the mean and covariance are computed for each set of features to calculate the Wasserstein-2 distance. A lower FID score indicates that the distributions of generated images are closer to the real images, suggesting better quality and realism of the generated images.
- FID is expressed as: where x denotes the feature vectors of the real images, g denotes the feature vectors of the generated images. Tr denotes the trace of a matrix, which is the sum of all the diagonal elements.
- Cosine similarity was also used to measure the similarity between the real and generated images.
- Cosine similarity is a mathematical measure used to determine the similarity between two vectors in a multi-dimensional space.
- the cosine similarity was calculated by the extracted 2048-dimensional feature vectors of the real and generated images from an ImageNet-pre-trained ResNet-50 model.
- classifiers were trained based on the real and generated images for cell-type or pathological annotations. Specifically, the image features were extracted from an ImageNet-pre-trained ResNet-50 model with classification heads tailored for different annotation tasks. Both Accuracy and F1 scores were reported as the evaluation metrics. The higher Accuracy and F1 score indicated better classification performance. For clustering performance evaluation, three commonly used metrics were adopted: normalized mutual information (NMI), adjusted Rand index (ARI), and average silhouette width (ASW). The NMI score was computed to measure the concurrence between ground truth cell type labels and Louvain cluster labels from integrated cell embeddings, with clustering resolutions ranging from 0.1 to 2.
- NMI normalized mutual information
- ARI adjusted Rand index
- ASW average silhouette width
- the ARI was used to assess the agreement between annotated labels and NMI-optimized Louvain clusters, adjusting for randomly correct labels.
- the ASW score ranging from -1 to 1 , was calculated to evaluate the relationship between a cell's within-cluster distances and its distances to the closest cluster boundaries. Higher scores for NMI and ARI indicate better matches, while an ASW score of 1 indicates well-defined clusters.
- the same hyperparameters as the Geneformer were used during the pretraining of GenVinci. Specifically, a linear learning rate scheduler with warmup was employed, setting the maximum learning rate to 5 x 10 -5 , and an Adam optimizer was used with a weight decay of 0.001 and a batch size of 16. The dimensionality of the input and output was 256, and the inner layer had a dimensionality of 512. For all fully connected layers, a nonlinear activation function was utilized in the form of a Rectified Linear Unit (ReLU); a dropout probability of 0.02; a standard deviation of 0.02 for the initializer of weight matrices; and an epsilon value of 1 x 10 -12 for layer normalization layers.
- ReLU Rectified Linear Unit
- SD version 1 .5 was utilized for image generation and denoising diffusion implicit model (DDIM) sampling strategy, with a batch size of 4.
- the Adam optimizer was employed, setting the learning rate to 5 x 10 -5 for training.
- Each model was fine-tuned for 500 epochs, and the number of sampling steps was set to 250.
- the parameters of the BiomedCLIP model are always fixed.
- GenVinci collects more than 210 million dissociated single-cell and targeted spatial transcriptomics data points for pretraining, covering 82 general tissues and 74 sequencing technologies from both human and mouse samples. Among them, more than 60 million dissociated single-cell data points were downloaded from CELLxGENE, a big data platform that provides curated and interoperable single-cell datasets. Additionally, we downloaded 1 million single-cell datasets from SenNet (focusing on aging and senescent cells), the Gene Expression Omnibus, and other public sources. For spatial transcriptomics datasets, we retrieved approximately 50 million data points from several databases: 10x Genomics (10 million), SODB (30 million), HEST-1 K (2 million), and STimage-1 K4M (4 million).
- the Spatial Transcriptomics dataset consists of sections from 23 breast cancer patients with Luminal A, Luminal B, triple-negative, and HER2-positive subtypes. Each patient includes 2 to 3 sections with both H&E images (scanned at x20 magnification) and spatial gene expression data.
- the number of spots in each section ranged from 256 to 712, and the number of cells in each spot varied from 5 to several tens.
- Each spot had a diameter of 100 m and was arranged in a grid with a center-to-center distance of 200 pm.
- 30,612 spots from 68 tissue sections were obtained, and each spot was measured by 26,949 gene expression.
- Detailed steps for collecting and handling breast cancer biopsies are known in the art, and the RNA sequencing of the barcoded libraries and subsequent analysis of tissue sections were performed using standardized protocols.
- the Visium dataset included two sections of invasive ductal carcinoma breast tissue.
- the tissue was embedded and cryosectioned, as described in the Visium Spatial Protocols (CG000240).
- Tissue sections, with a thickness of 10 pm, were placed on Visium Gene Expression Slides for sequencing.
- the sections were scanned at x 20 magnification, and the H&E images were acquired using a Nikon Ti2-E microscope.
- the Visium Spatial Gene Expression library was prepared following the instructions provided in the Visium Spatial Gene Expression Reagent Kits User Guide (CG000239). Previous studies were followed to segment patches of 224 x 224 pixels centered on the Visium spots. The standard quality control steps in Scanpy were performed for the gene expression data. Finally, obtained were 3,798 spots for Section 1 and 3,987 spots for Section 2, both including the measurement of 30,612 spatially resolved gene expression.
- the MERFISH dataset was collected by the Human Tumor Atlas Project Pilot (HTAPP). It consists of two sets of 4 and 5 metastatic breast cancer (MBC) tissue samples, along with single-nucleus gene expression data (sn-MBC) and single-cell gene expression data (sc-MBC), respectively. Both H&E images and gene expression data for each patient sample were derived from different biopsies of the same metastasis. Multiplex FISH measurements of 212 genes were performed on consecutive sections using MERFISH and included regional annotations on the H&E sections provided by expert pathologists for both datasets. The H&E images were segmented and padded to a size of 64 x 64 pixels.
- the NanoString dataset was generated using a 960-plex CosMx RNA panel run on a CosMx SMI prototype which contains eight different samples from five non-small cell lung cancer (NSCLC) tissues.
- TheLung-5-1 data files were selected with the corresponding raw morphological image files for analysis.
- the selected dataset consisted of fluorescence imaging data from 30 FOVs, each stained with nucleus (DAPI) and membrane markers (CD45, PanCK, and CD3) to facilitate morphology-based cell segmentation.
- the single-nucleus DAPI images were segmented and padded to a size of 96 x 96 pixels, and the whole-cell IF-stained images are segmented and padded to a size of 128 x 128 pixels.
- the STcEM MERFISH dataset was retrieved from a recently published study, which collected paired MERFISH and EM data from a mouse model of LPC-induced demyelinating injury. This dataset correlates large-area electron microscopy scanning with MERFISH on adjacent tissue sections.
- the adjacent sections from a mouse brain were processed in parallel using harmonized MERFISH and EM protocols and spatially aligned to connect transcriptional profiles with the ultrastructure of the regions of interest. Aligning MERFISH and EM imaging data allowed for projection into a common coordinate system for analysis, resulting in 934 EM images and 1 ,502 MERFISH profiles.
- RNA-EM pairs were then segmented and padded to a size of 256 x 256 pixels, and Hungarian algorithm was utilized to complete a one-to-one pairing for cells in MERFISH and EM sections. This resulted in 934 RNA-EM pairs, each cell measured for 286 genes.
- genes detected in fewer than 3 cells were excluded from all the datasets.
- the imaging data are segmented as described in the previous section.
- we augmented the dataset by randomly rotating the image by 0, 90, 180, or 270 degrees and flipping the image horizontally 50% of the time.
- a computer implemented method of generating a multi-dimensional image of a cell or organelle comprising:
- the computer is programmed to mask a subset of the weighted binned expression or the identification thereof of the plurality of genes.
- the spatial transcriptomic analysis comprises single nucleus RNA (snRNA) sequencing, single cell RNA (scRNA) sequencing, patch sequencing, Visium, Xenium, NanoString GeoMx, STARmap, Slide-Seq, Slide-Tags, fluorescence in situ hybridization (FISH), multiplexed error-robust fluorescence in situ hybridization (MERFISH), electron microscopy, hematoxylin and eosin (H&E) imaging, fluorescence microscopy, 4',6-diamidino-2-phenylindole (DAPI), immunofluorescence (IF), or a combination thereof. 25.
- organelle is a nucleus, nucleolus, ribosome, vesicle, endoplasmic reticulum, Golgi apparatus, vacuole, centriole, lysosome, or mitochondria.
- a system to determine a multi-dimensional image of a cell or organelle comprising: a storage device; and a processor communicatively coupled to the storage device, wherein the processor executes application code instructions that are stored in the storage device to cause the system to:
- a computer program product for determining a multi-dimensional image of a cell or organelle comprising: a non-transitory computer-readable storage device having computer-executable program instructions embodied thereon that when executed by a computer cause the computer to determine a multi-dimensional image of a cell or an organelle, the computer-executable program instructions comprising:
- a software program product for determining a multi-dimensional image of a cell or organelle comprising: a computer-executable software program comprising instructions that when executed by a computer cause the computer to determine a multi-dimensional image of a cell or an organelle, the computer-executable program instructions comprising:
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Theoretical Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Biophysics (AREA)
- Evolutionary Computation (AREA)
- General Health & Medical Sciences (AREA)
- Software Systems (AREA)
- Data Mining & Analysis (AREA)
- Artificial Intelligence (AREA)
- Molecular Biology (AREA)
- Computing Systems (AREA)
- General Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- General Engineering & Computer Science (AREA)
- Mathematical Physics (AREA)
- Biomedical Technology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Genetics & Genomics (AREA)
- Medical Informatics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Databases & Information Systems (AREA)
- Biotechnology (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Physiology (AREA)
- Multimedia (AREA)
- Bioethics (AREA)
- Epidemiology (AREA)
- Public Health (AREA)
- Image Analysis (AREA)
Abstract
Provided herein is a generative model based on transformer and diffusion models that learns a multi-modal representation of cells and tissues and reconstructs morphological information from transcriptome profiles to generate cellular and tissue images from single-cell expression profiles during inference.
Description
METHODS AND SYSTEMS FOR DETERMINING CELLULAR AND TISSUE IMAGES FROM GENE EXPRESSION
Background of the Invention
Cells were discovered through microscopy, and their subsequent characterization relied on imaging. In the past decades, molecular profiling, most recently single-cell profiling has opened the way to cell characterization at unprecedented scale, resolution, and interpretability, in both dissociated and tissue contexts. First, single-cell RNA-seq (scRNA-seq) has spurred in-depth analyses of cells, during development, differentiation, homeostasis, and pathogenesis, followed by methods of other molecular profiles, including cell surface and intracellular proteins, chromatin accessibility and modifications, T cell receptor (TCR) and B cell receptor (BCR) repertoires, genomic and mitochondrial somatic genetic variation, spatial location, cell morphology, and some of their multi-modal combinations, increasingly generating a multi-view perspective of cells.
However, current studies largely rely on one measurement modality, and if a multi-modal view is desired it requires direct measurement, ideally in the same cell or tissue. This poses several challenges. First, many measurement modalities are costly, laborious, and time-consuming, requiring specialized assays and expertise. As the number of experimental methods grows at a staggering pace, so does the experimental complexity. Second, independent methods often require separate specimens, which is especially difficult in a clinical setting, when samples are often limited in quantity (e.g., sections from miniscule needle biopsies) or can only be collected in certain ways. Finally, different modalities may have no immediate correspondence in their nominal variables (e.g., gene levels vs. non-molecular imaging features), and thus hard to unify to a coherent view. No approach to date has focused on revealing hidden morphological information from expression profiles to generate cell and tissue images from singlecell expression profiles. For example, complex neuronal morphologies are highly cell-type-specific, essential for their function, and are associated with genomic mutations, gene expression levels, and methylation patterns. Therefore, improved imaging modalities that link molecular information with cell morphology are needed for further studying the important functions of cells.
Summary of the Invention
Provided herein a generative model based on transformer and diffusion models that learns a multi-modal representation of cells and tissues and reconstructs morphological information from transcriptome profiles to generate cellular and tissue images from single-cell expression profiles during inference.
In one aspect, featured is a computer implemented method of generating a multi-dimensional (e.g., two-dimensional or three-dimensional) image of a cell or organelle from single cell sequencing data. The method includes providing a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle and using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
In some embodiments, the computer is programmed with a plurality of modules that generate the multi-dimensional image. For example, the computer may be programed with three modules: an encoder
model (e.g., a Bidirectional Encoder Representations from Transformers (BERT)-based single-cell language model), a diffusion model (e.g., a Latent Diffusion Model (LDM)), and a pretraining model (e.g., a biomedical Contrastive Language-Image Pretraining (CLIP) model).
In some embodiments, the method further includes providing a raw image of the cell or organelle.
In some embodiments, the computer is programmed to normalize and bin the gene expression of each of the plurality of genes.
In some embodiments, the computer is programmed to rank each of the plurality of genes by the normalized binned gene expression of each of the plurality of genes.
In some embodiments, the computer is programmed to transform the binned expression of each of the plurality of genes to produce the image.
In some embodiments, the computer is programmed to weight the binned expression or an identification thereof of each of the plurality of genes.
In some embodiments, the computer is programmed to mask a subset of the binned expression or the identification thereof of the plurality of genes.
In some embodiments, the computer is programmed to predict a remainder of the genes or binned gene expression that are not part of the masked subset.
In some embodiments, the computer is programmed to compress the image to reduce a dimensionality and add or reduce noise of the image.
In some embodiments, the computer is programmed to iteratively add or reduce noise of the image.
In some embodiments, the computer is programmed with a diffusion model.
In some embodiments, the machine learning algorithm model is trained with an empirical database of single cell gene expression profiles.
In some embodiments, the computer is programmed to optimize the image based on the empirical database to produce the multi-dimensional image of the cell or organelle.
In some embodiments, the computer is programmed to optimize the image based on the empirical database including a plurality of sets including matched gene expression profiles, images, and/or text annotation.
In some embodiments, the computer is programmed to correlate the set of gene expression profiles with the image of the cell or organelle.
In some embodiments, the computer is programmed to align the plurality of gene expression profiles with the plurality of sets including the matched gene expression profiles, images, and/or text annotations to produce a plurality of gene expression matrices.
In some embodiments, the computer is programmed to resize each image from the plurality of sets to match the size of the image.
In some embodiments, the computer is programmed to translate the plurality of gene expression matrices into the multi-dimensional image.
In some embodiments, the computer is programmed to rank the multi-dimensional image based on a similarity with the plurality of sets including the matched gene expression profiles, images, and/or text annotations.
In some embodiments, the computer is programmed to produce a plurality of multi-dimensional images, and the method further includes selecting an image with the highest rank.
In some embodiments, the gene expression profiled is obtained by performing a spatial transcriptomic analysis.
In some embodiments, the spatial transcriptomic analysis includes single nucleus RNA (snRNA) sequencing, single cell RNA (scRNA) sequencing, patch sequencing, Visium, Xenium, NanoString GeoMx, STARmap, Slide-Seq, Slide-Tags, fluorescence in situ hybridization (FISH), multiplexed error- robust fluorescence in situ hybridization (MERFISH), electron microscopy, hematoxylin and eosin (H&E) imaging, fluorescence microscopy, 4',6-diamidino-2-phenylindole (DAPI), immunofluorescence (IF), or a combination thereof.
In some embodiments, the organelle is a nucleus, nucleolus, ribosome, vesicle, endoplasmic reticulum (ER; e.g., rough ER or smooth ER), Golgi apparatus, vacuole, centriole, lysosome, or mitochondria.
In some embodiments, the cell is a neuron, and the method shows a morphology of the neuron.
In some embodiments, the method generates a multi-dimensional image of a plurality of cells or organelles.
In some embodiments, a tissue includes the plurality of cells, and the method produces a multidimensional image of the tissue.
In some embodiments, the method produces an electron microscopy, hematoxylin and eosin, or a fluorescence microscopy multi-dimensional image.
In some embodiments, the method produces fluorescence cell morphology.
In another aspect is featured a computer implemented system including, for example, a storage device and a processor communicatively coupled to the storage device. The processor may be configured to execute code instructions stored on the storage device to cause the system to perform the computer implemented methods as described herein, e.g., of receiving a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multi-dimensional (e.g., two-dimensional or three-dimensional) image of the cell or organelle.
In another aspect is featured a computer program configured to carry out executable program instructions to perform the computer implemented methods as described herein. The computer program product may include a non-transitory computer-readable storage device having computer-executable program instructions embodied thereon that when executed by a computer cause the computer to determine a multi-dimensional (e.g., two-dimensional or three-dimensional) image of a cell or an organelle. The computer-executable program instructions may include, for example, computer-executable program instructions to receive a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and computer-executable program instructions to determine the multi-dimensional image of the cell or organelle using the gene expression profile. The computer program product may be programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
In another aspect is featured a software program product for determining a multi-dimensional (e.g., two-dimensional or three-dimensional) image of a cell or organelle. The software program product
may include, for example, a computer-executable software program that includes instructions that when executed by a computer cause the computer to determine a multi-dimensional image of a cell or an organelle. The computer-executable program instructions may include computer-executable software program instructions to receive a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and computer-executable software program instructions to determine the multi-dimensional image of the cell or organelle using the gene expression profile. The software program product may be programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
Definitions
It is to be understood that aspects and embodiments of the invention described herein include “comprising,” “consisting,” and “consisting essentially of” aspects and embodiments.
As used herein, the singular form “a,” “an,” and “the” includes plural references unless indicated otherwise.
The term “about” as used herein refers to the usual error range for the respective value readily known to the skilled person in this technical field. Reference to “about” a value or parameter herein includes (and describes) embodiments that are directed to that value or parameter per se. In some instances, reference to “about” a value or parameter herein indicates the value or parameter ± 10%.
The term “computer” as used herein may refer to device embodied in any of a number of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer, as nonlimiting examples. Additionally, a computer may be embedded in a device not generally regarded as a computer but with suitable processing capabilities, including a Personal Digital Assistant (PDA), a smartphone or any other suitable portable or fixed electronic device. A computer may have one or more communication devices, which may be used to interconnect the computer to one or more other devices and/or systems, such as, for example, one or more networks in any suitable form, including a local area network or a wide area network, such as an enterprise network, and intelligent network (IN) or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks or wired networks. A computer may have one or more input devices and/or one or more output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that may be used to provide a user interface include printers or display screens for visual presentation of output and speakers or other sound generating devices for audible presentation of output. Examples of input devices that may be used for a user interface include keyboards, and pointing devices, such as mice, touch pads, and digitizing tablets. As another example, a computer may receive input information through speech recognition or in other audible formats. Computer-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or distributed as desired in various embodiments. Databases, if employed in the methods or devices or systems herein, may include computer readable memory (also referred to as “memory”). For example, data storage space 3memlN may be and/or include computer readable memory, used to store data as described in the
disclosure. Memory may be embodied by suitable hardware, including but not limited to the following: hard disk drives, serial advanced technology attachment (SATA) hard drives, SATA solid state drives (SSDs), non-volatile memory express (NVMe) SSDs, tape drives.
The terms “program,” “app,” and “software” are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that may be employed to program a computer or other processor to implement various embodiments described herein. Additionally, it should be appreciated that, according to one aspect, one or more computer programs that when executed perform methods of this application need not reside on a single computer or processor but may be distributed in a modular fashion among a number of different computers or processors to implement various embodiments of this application.
Brief Description of the Drawings
The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
FIG. 1A-FIG. 1D are a schematic flow chart showing the GenVinci workflow. FIG 1A: Overview of GenVinci. [Upper]: Input of GenVinci, which can be the gene expression data from single-nucleus RNA sequencing (snRNA-seq), single-cell RNA sequencing (scRNA-seq), spatially resolved transcriptomics (SRT) across different resolutions (nucleus/cell/spot), and Patch-seq. [Middle]: Representative deep learning architectures in GenVinci, which include Transformer, U-Net, and Autoencoder. [Bottom]: Output of GenVinci, which can be different imaging modalities at single-cell resolution (Hematoxylin-and-Eosin (H&E), 4',6-diamidino-2-phenylindole (DAPI), immunofluorescence (IF), neuron morphology, and electron microscopy) and multi-cell resolution (spot (1 -10 cells), tissue (5- 100 cells), and whole slide), and various downstream tasks (multi-modal clustering, zero-shot annotation, and data augmentation). FIG. 1 B: GenVinci collects over 210 million gene expression data points from more than 20 data portals, including CELLxGENE, Single Cell PORTAL, and 10x Genomics to build a large-scale single-cell corpus (GenCorpus-210M). The database encompass both human and mouse species, covering 82 general tissue types, and the technology spans 30 single-cell RNA sequencing (scRNA-seq) platforms and 44 spatial transcriptomics technologies. FIG. 1C: The first stage (pretraining) of GenVinci. Gene expression data is sent into a newly-designed transformer framework for pretraining (masked language modeling). This step leverages prior biological knowledge from the collected large- scale single-cell corpus (Gencorpus-210M) to learn better gene expression embeddings specific to the utilized sequencing techniques. The utilized transformer framework employs a binning-based rank value encoding method to encode gene expression data and consists of several transformer encoders and one Mixture-of-Experts (MoE) decoder (FIG. 7A). FIG. 1D: The second stage (fine-tuning) of GenVinci. Gene expression data is sent to the previously trained GenVinci Large Language Model (LLM) and passed through a projector to obtain the transcriptomic prompt as the input of the generative model. A pre-trained diffusion model is employed as the backbone of the generative model to generate cellular images based on the transcriptomic prompt (FIG. 7B). The forward diffusion process generates the paired noisy images and the amount of noise added at each step for training (FIG. 7C). Additionally, a vision-language model
(VLM) is introduced for weak supervision, aiming to align transcriptomes, images, and potential text information within a shared latent space (FIG. 7D).
FIG. 2A-FIG. 2H are H&E images produced by GenVinci from gene expression data across different sequencing resolutions. For each cell/spot/slide, 5 images were generated from the same gene expression data that were randomly sampled five times. The generated images were ranked based on the relative (normalized) cosine similarity scores (the best one with highest normalized cosine similarity score is 1 .0). The ground-truth (GT) images are listed in the first column. FIG. 2A: GenVinci generates single-cell H&E (scH&E) images from multiplexed error-robust fluorescence in situ hybridization (MERFISH) single-nucleus gene expression (snGE) data. FIG. 2B: GenVinci generates multi-cell H&E (mcH&E) images from MERFISH single-cell gene expression (scGE) data. FIG. 2C: GenVinci generates multi-cell H&E images (mcH&E) from 10x Visium multi-cell gene expression (mcGE) data (1 -10 cells/spot). FIG. 2D: GenVinci generates multi-cell H&E images (mcH&E) from Spatial Transcriptomics multi-cell gene expression (mcGE) data (5-100 cells/spot). FIG. 2E: GenVinci generates whole-slide H&E (wsH&E) images from Spatial Transcriptomics whole-slide gene expression matrices (wsGEM). Due to the limited size of the whole-slide level samples (several tens), we only generated a few images as an initial proof of concept without further quantitative evaluation. FIG. 2F: GenVinci-generated images are compared with ground-truth images and images generated by a Convolutional Neural Network (CNN) baseline. We selected the images generated by GenVinci with the highest similarity scores for comparison. FIG. 2G: Comparisons of the Frechet Inception Distance (FID) scores (lower is better) and cosine similarity for the CNN baseline and GenVinci. The FID scores are calculated between the groundtruth and generated images. The cosine similarity (higher is better) is calculated based on the image embeddings extracted from both ground-truth and generated images using an ImageNet-pretrained ResNet-50 model. FIG. 2H: Visualization of the ground-truth, CNN baseline, and GenVinci image embeddings using Uniform Manifold Approximation and Projection (UMAP).
FIG. 3A-FIG. 3E are fluorescence cell morphology images produced by GenVinci from gene expression data across different sequencing resolutions. FIG. 3A: GenVinci generates single-nucleus 4',6-diamidino-2-phenylindole (snDAPI) images from CosMx SMI single-nucleus gene expression (snGE) data. FIG. 3B: GenVinci generates single-cell immunofluorescence (scIF) images from CosMx SMI single-cell gene expression (scGE) data. One channel represents DAPI, another channel represents PanCK, one channel represents CD45, and one channel represents CD3. FIG. 3C: GenVinci-generated best DAPI and IF-stained images are compared with ground-truth images and images generated by a CNN baseline. FIG. 3D: Comparisons of the FID scores and cosine similarity for the CNN baseline and GenVinci. FIG. 3E: Visualization of the ground-truth, CNN baseline, and GenVinci image embeddings using UMAP.
FIG. 4A-FIG. 4D are single cell neuronon morphology images and graphs showing similarity with ground truth images. FIG. 4A: GenVinci generates realistic single-cell neuron morphology (scNM) images from Patch-seq single-cell gene expression (scGE) data. FIG. 4B: GenVinci-generated best neuron images are compared with ground-truth images and images generated by a CNN baseline. FIG. 4C: Comparisons of the FID scores and cosine similarity for the CNN baseline and GenVinci. FIG. 4D: Visualization of the ground-truth, CNN baseline, and GenVinci image embeddings using UMAP.
FIG. 5A-FIG. 5D are single cell electron microscopy images (EM) and graphs showing similarly with ground truth images. FIG. 5A: GenVinci generates realistic single-cell electron microscopy (scEM) images from MERFISH single-cell gene expression (scGE) data. FIG. 5B: GenVinci-generated best EM images are compared with ground-truth images and images generated by a CNN baseline. FIG. 5C: Comparisons of the FID scores and cosine similarity for the CNN baseline and GenVinci. FIG. 5D: Visualization of the ground-truth, CNN baseline, and GenVinci image embeddings using UMAP.
FIG. 6A-FIG. 6F are images and graphs showing how GenVinci can be flexibly applied to a diverse panel of downstream tasks, such as multi-modal clustering (FIG. 6A), zero-shot classification (FIGS. 6B and 6C), and data augmentation (FIGS. 6D-6F). FIG. 6A: The clustering map of gene expression embeddings extracted by Geneformer (a recently published single-cell foundation model), continued-pre-trained Geneformer, and fine-tuned GenVinci. Cells/spots are colored with the cell- type/pathological annotations. FIGS. 6B and 6C: The classification performance of the cell- type/pathological annotation task (cancer/non-cancer) is evaluated on the st-HBC dataset and the sn- MBC dataset. For the sn-MBC dataset, we divided the original cell types into cancer-related and non- cancer-related categories. Cancer-related cells include MBC_neuronal, while non-cancer-related cells include endothelial_angiogenic, NK, and smooth muscle_vascular types. This classification could be useful for exploring the differences between cancerous and non-cancerous cells within the tumor microenvironment or metastatic niches. FIG. 6D: The performance of the classification model trained solely with ground-truth (GT) images and GenVinci-generated (Gen) images with different augmentation ratios. FIG. 6E: The performance of the classification model trained on GT images combined with GenVinci-generated images from the MERFISH dataset. FIG. 6F: The performance of the classification model trained on GT images combined with GenVinci-generated images from an external scRNA-seq dataset. We set different ratios for the size of the augmented data.
FIG. 7A-FIG. 7E are a schematic flow chart showing the GenVinci workflow FIG. 7A: Transformer is utilized as the backbone of the sc-LLM model of GenVinci. GenVinci employs a binning-based rank value encoding method to encode gene expression data and consists of several transformer and MoE blocks to extract the gene expression embeddings. FIG. 7B: Stable Diffusion is utilized as the backbone of the generative model of GenVinci. The transcriptomic prompt from the projector is concatenated with the Gaussian noise to generate cellular images by the reverse diffusion process. FIG. 7C: The forward diffusion process is applied to the latent image embeddings to produce a series of noisy image embeddings. Noises are iteratively added until the image is fully destroyed to a standard Gaussian noise. These noisy embeddings are served as the training data for the reverse diffusion process of Stable Diffusion. FIG. 7D: The VLM BiomedCLIP, trained on the PMC-15M corpus, includes 15 million biomedical figure-caption pairs extracted from over 3 million biomedical research articles. BiomedCLIP aims to align images and text descriptions within a shared latent space. FIG. 7E: The rank-based inference of GenVinci. During the inference, GenVinci can generate multiple images from the same gene expression data and select the ‘best’ one with the highest normalized cosine similarity score between the input gene expression and generated images.
FIG. 8 is a schematic drawing showing an example of the CNN baseline architecture (st-HBC dataset). The CNN baseline model can predict cellular images from the original gene expression data. The kernel size in the convolutional layers and the number of nodes in the fully connected layers are
adjusted according to the input dimension (number of genes) and the desired output image size across different datasets.
FIG. 9 is a schematic workflow of the matrix-to-slide GenVinci. The cell-to-tile GenVinci is extended to a matrix-to-slide version by transferring the weights from both pre-trained LLM and generative model. The pre-trained LLM is frozen to extract gene expression embeddings from the matrix. A new projector is implemented to transform the matrix embeddings into a slide-level transcriptomic prompt, which is sent to the pre-trained diffusion model to generate whole-slide images. Other modules remain the same as in the cell-to-tile GenVinci.
FIG. 10A-FIG. 10C are fluorescent and electron microscopy images and graphs showing paired images. FIG. 10A: The adjacent DAPI images (from MERFISH) and EM images are aligned into the common coordinate system using STcEM. FIG. 10B: The cells in MERFISH and EM images are paired using the nearest-neighbor searching method; paired MERFISH and EM images are connected with a grey line. FIG. 10C: The cells in MERFISH and EM images are paired using the Hungarian algorithm.
FIG. 11 A and FIG. 11B are clustering maps and comparisons of clustering performance graphs. FIG. 11 A: Left. The clustering map of the original image embeddings extracted using an ImageNet- pretrained ResNet-50 model. Middle. The combination of the original image embeddings with paired gene expression embeddings. Right. The concatenation of the GenVinci-refined gene expression embeddings with the GenVinci-generated image embeddings. FIG. 11 B: The comparisons of clustering performance on normalized mutual information (NMI), average silhouette width (ASW), and adjusted Rand index (ARI).
FIG. 12 is a schematic drawing showing a pipeline for the zero-shot annotation task. The text prompts are first created based on the categories of interest. Text embeddings are then extracted from the VLM text encoder. At the same time, the gene expression data is sent into GenVinci to extract cell embeddings to compare the similarities with text embeddings. The cell embeddings can be extracted from the output of the projector in the VLM supervision module or the generated cellular image embeddings (extracted by the VLM image encoder).
FIG. 13 is a schematic drawing showing a pipeline for data augmentation. GenVinci was initially trained on paired single-cell gene expression data (MERFISH) and H&E images. Next, an external, well- annotated scRNA-seq dataset is then fed into the well-trained GenVinci to generate H&E images. These generated H&E images are paired with well-annotated cell types from the scRNA-seq dataset and added as data augmentation for the downstream annotation task.
Detailed Description
Described herein are computer implemented methods of generating a multi-dimensional (e.g., two-dimensional or three-dimensional) image of a cell or organelle from single cell sequencing data. The methods described herein generally include providing a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle and using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multidimensional image of the cell or organelle (e.g., a nucleus, nucleolus, ribosome, vesicle, endoplasmic reticulum (ER; e.g., rough ER or smooth ER), Golgi apparatus, vacuole, centriole, lysosome, or mitochondria). The methods described herein may also be used to generate multi-dimensional images of
a plurality of cells, e.g., a tissue. The methods and systems are particularly useful for showing complex cell morphologies, particularly of dynamic cell types (e.g. neurons). The methods are highly adaptable to work with a variety of cellular imaging techniques, such as electron microscopy, hematoxylin and eosin (H&E) imaging, fluorescence microscopy, 4',6-diamidino-2-phenylindole (DAPI), and immunofluorescence (IF). This is particularly advantageous for streamlining typically expensive and low throughput imaging techniques, such as electron microscopy.
Also featured herein are computer implemented systems including, for example, a storage device and a processor communicatively coupled to the storage device. The processor may be configured to execute code instructions stored on the storage device to cause the system to perform the computer implemented methods as described herein, e.g., of receiving a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multi-dimensional (e.g., two-dimensional or three-dimensional) image of the cell or organelle.
Also featured are computer programs configured to carry out executable program instructions to perform the computer implemented methods as described herein. The computer program product may include a non-transitory computer-readable storage device having computer-executable program instructions embodied thereon that when executed by a computer cause the computer to determine a multi-dimensional (e.g., two-dimensional or three-dimensional) image of a cell or an organelle. The computer-executable program instructions may include, for example, computer-executable program instructions to receive a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and computer-executable program instructions to determine the multi-dimensional image of the cell or organelle using the gene expression profile. The computer program product may be programmed with a machine learning algorithm model to determine the multidimensional image of the cell or organelle.
Also featured herein are software program products for determining a multi-dimensional (e.g., two-dimensional or three-dimensional) image of a cell or organelle. The software program product may include, for example, a computer-executable software program that includes instructions that when executed by a computer cause the computer to determine a multi-dimensional image of a cell or an organelle. The computer-executable program instructions may include computer-executable software program instructions to receive a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from the cell or organelle; and computer-executable software program instructions to determine the multi-dimensional image of the cell or organelle using the gene expression profile. The software program product may be programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
The methodology described herein includes a generative model (in some instances herein referred to as “GenVinci”) with a machine learning algorithm that learns a multi-modal representation of morphological and molecular profiles and uses it for gene-to-image translation across various profiling platforms and imaging techniques. The model’s cascaded architecture allows one to flexibly apply it across a range of downstream tasks, including image generation, multi-modal clustering, zero-shot classification, and data augmentation.
The model can be trained in two stages: an initial general-purpose pretraining on expression data, followed by fine-tuning with paired imaging data for the gene-to-image translation task. In the pretraining step, a single-cell large language model can be trained to encode high-dimensional expression data into robust latent embeddings. GenVinci was pre-trained on GenCorpus-21 OM, a large- scale scRNA-seq corpus of over 210 million scRNA-seq and spatial transcriptomics profiles from a broad range of tissues across human and mouse samples. To obtain better cell embeddings specific to the utilized profiling techniques, the pre-training module can be further pre-trained on additional gene expression data, e.g., using the masked language modeling strategy, where some genes or gene expression values are randomly masked in a cell and then predicted.
In a subsequent fine-tuning stage, the pre-trained model can be cascaded with a diffusion model (e.g., Stable Diffusion (SD)), a diffusion model where the extracted gene expression embeddings can serve as transcriptomic prompts to generate cell images. Since the diffusion model is specifically designed for text-to-image generation and may be trained on billions of image-text pairs, it may not generalize well with limited gene expression-image pairs. Thus, another biomedical vision-language model, such as BiomedCLIP, may be employed to provide weak supervision to reduce modality discrepancies for better aligning the transcriptomic prompts with image and text embeddings. BiomedCLIP can extract multi-modal cell image embeddings that are well-matched with their potential text embeddings. These are then used to refine the cell embeddings from GenVinci gene expression encoder. Consequently, the enhanced cell embeddings align more closely with both image and potential text embeddings, rendering them more effective for fine-tuning the diffusion model for cell image generation, thereby improving the quality of the resulting cell images.
During the inference stage, the method can be programmed to generate multiple cell images from a single expression profile or signature, reflecting various perspectives of the cell morphology, which is of great significance even when only partial or individual sampling of gene expression distributions is available, rather than a complete profile (e.g., with methods like multiplexed error-robust fluorescence in situ hybridization (MERFISH)) or full distributions representing cell states. It can then rank the generated images that are closest to the ground truth based on similarity scores between the input expression profile and the generated images. This flexible, rank-based approach to inference can be highly effective in real-world clinical settings when pathologists or physicians want to avoid incorrect predictions and choose the most confident generation.
The cascaded architecture of the methods described herein overcomes the limitation of having few expression-image pairs for fine-tuning and enhances the cell representations obtained from the pretraining step for more effective image generation. Furthermore, by including large language models, GenVinci can be flexibly applied to diverse downstream tasks, leveraging the models’ capabilities in a multi-modal manner.
Methods
GenVinci is a deep generative model that can generate cellular morphology from single-cell gene expression. It can include three modules: an encoder model (e.g., a Bidirectional Encoder Representations from Transformers (BERT)-based single-cell language model), a diffusion model (e.g., a Latent Diffusion Model (LDM)), and a pretraining model (e.g., a biomedical Contrastive Language-Image
Pretraining (CLIP) model). On aspect of the approach includes the LDM module, which maps single-cell gene expression embeddings to cellular images. LDM approaches typically accept natural language embeddings, which are then translated into corresponding images. Because of this, GenVinci includes a single-cell language model that produces gene expression embeddings in a manner similar to natural language models. However, since the representations learned by GenVinci are unaligned to the pretraining of the LDM, BiomedCLIP, a model that learns to align biomedical figure captions to their corresponding images, was employed to supervise the fine-tuning of the entire approach. The integration of three modules allows GenVinci not only to generate realistic cellular images but also to be easily applied for diverse downstream tasks. These modules were introduced along with their tailored pretraining and fine-tuning strategies for GenVinci, as described in the subsequent section.
Unified transcriptomic representation
Given that the raw gene expression data is usually sparse, high-dimensional, and noisy, a singlecell language model can be trained to encode gene expression data into a low-dimensional latent space. GenVinci is a BERT-like transformer model pre-trained on a large-scale corpus (GenCorpus-21 OM) of approximately 210 million single-cell and spatial transcriptome to allow context-aware predictions in settings with limited data in network biology. Benefiting from the flexible nature of the attention mechanism, the single-cell language model can accept single-cell gene expression data from different sequencing measurements without the need to standardize into uniform gene lists. In addition, domain knowledge learned from an extensive single-cell transcriptomics corpus, particularly concerning genegene interactions, can enhance cellular representation across a spectrum of downstream tasks.
Since GenVinci is a transformer-based approach that expects inputs to be sequences, a binning- based rank value encoding method can be employed to induce a sequence from the (tabular) gene expression observations. Rank value encoding is a non-parametric approach that ranks genes by their expression levels within a single cell. Initially, given a cell-by-gene matrix, only protein-coding genes may be retained based on a predefined vocabulary by GenCorpus-21 OM. Subsequently, the gene expression data can first be normalized using a scale factor, followed by a log transformation to bring them to a comparable scale. The bin size is set to be 50. Specifically, let x^ represent the expression value for a gene i in cell c, the normalization gene expression x' is given by: x^ = log (
where n is the total number of genes within the cell. After normalization, genes within each specific cell can be assigned into binning values by their normalized expression values using a quantile binning method and ranked in descending order and subsequently tokenized according to the predefined vocabulary, where each gene gk is assigned an integer identifier ic/(^fc). The vocabulary can also include additional (e.g., two or more) special tokens for padding and masking and context tokens to include the cell context information (e.g., modality and assay). The binning-based rank value encoding method prioritizes genes that uniquely characterize cell states by using expression data across the GenCorpus- 210M, thereby downgrading common housekeeping genes and elevating critical transcription factors that define cell identity. It can offer robustness against technical biases, maintaining the relative gene ranking within each cell despite potential variations in absolute transcript counts. Furthermore, top m expressed
genes with the normalized expression value x' > 0 for can be selected for each cell. For the cells with the number of expressed genes less than m, padding tokens can be added to meet the required length. In one example, m can be set to be 1024 (which fully represents more than 90% of rank value encodings in GenCorpus-210M). The final embeddings of the tokenized genes for cell c can be denoted as:
where the ranked genes are represented by their integer identifiers, and Embg is a learnable conventional embedding layer for the gene tokens. Similar to the position embedding in BERT, trainable binning-based expression embeddings denoted as EEC can be used for each gene token within the cell, which allows the modeling of the relationship between the binned gene expression values. GenVinci also includes the conditional embedding to distinguish the context tokens and gene tokens which is denoted as CEc. The final input embedding for each cell can be represented as the element-wise sum of the gene and binning- based expression embeddings:
GEC — GEC (J EEC (J CEc where GE' e 7?mxd, d is the dimension of each token embedding vector. The embeddings can then sent into the GenVinci transformer blocks.
Spatially resolved transcriptomics pre-trained transformers
GeneVinci can be composed of six identical transformer encoders, each featuring, for example, two sub-layers. The first sub-layer can be a multi-head self-attention layer, and the second sub-layer can be a simple position-wise feed-forward neural network layer. The scaled dot-product attention can be implemented in the self-attention layer as:
Where Q,K, V represent the query, keys, and values, respectively. These are sets of vectors derived from the input embeddings. dk is the dimension of the queries and keys, and - = is a scaling factor to control -k the magnitude of attention weights. The output is computed as a weighted sum of the values, where the weight is calculated by a compatibility function of the query with the corresponding key. Four attention heads can be used per transformer encoder to jointly attend to information from different representation subspaces:
and W° are the parameter matrices of projections. Two feed-forward neural network layers can be followed by the self-attention layers, which can be applied to each position separately and denoted as:
FFN(xj = max (0, x W, + b^ W2 + b2
The linear layers may share the same architecture but have different parameters across various positions. During pretraining, 15% of the genes or gene expression values within each specific cell were randomly masked in spatial transcriptomics data, and the model can be trained to predict the masked genes or gene expression values based on the context of the remaining unmasked spatial genes or gene
expression values. The cross-entropy loss can be employed as the masked learning objective function, which for a single masked token can be formulated as:
where V is the size of the vocabulary. yt and ft are the ground truth and predicted probabilities of the i-th gene token, respectively. Herein, masked language modeling loss can be computed for all masked gene or gene expression tokens within a batch, and this loss can then be averaged to produce the final loss against which the model parameters can be updated. In addition, instead of a single feedforward network like in standard Transformers, GenVinci Transformers use Mixture of Experts which consists of multiple expert FFNs, with a gating network that dynamically selects a few experts for each input token. This approach allows the model to scale efficiently, increasing capacity while keeping computational costs lower than a fully dense model.
Fine-tuning with latent diffusion models
After obtaining a robust representation of gene expression data through pretraining, it can be employed as the transcriptomic prompt for cellular image generation using a latent diffusion model. Latent diffusion models have been shown to achieve highly competitive performance on various tasks, including image inpainting, image generation, super-resolution, and semantic scene synthesis. Given the limited number of gene expression-image pairs in spatial transcriptomics datasets, we opted to fine-tune a latent diffusion model for the text-to-image generation task, specifically Stable Diffusion (SD) to generate cellular images from transcriptomic prompts. As opposed to traditional diffusion models, which generate samples in pixel space, SD operates within a latent feature space, which has been shown to significantly reduce computational requirements and introduce fewer spatial down-sampling, resulting in better image synthesis quality.
SD may include two modules: (1 ) a regularized autoencoder that compresses images into lowdimensional latent features and reconstructs them, and (2) a U-Net-based denoising model with attention mechanisms for more flexible conditional image generation. At the training stage of SD, the forward diffusion process can first be applied to image latent embeddings derived by the autoencoder to produce a series of noisy embeddings. Specifically, given a cellular image x in pixel space, its corresponding latent embedding z is obtained via an autoencoder £(•), i.e. , z = £(x). Noises can then be iteratively added to z for T time steps until it is fully destroyed to a standard Gaussian noise. The t-th noisy embedding zt is sampled from the following distribution:
where I is the identity matrix,
is a variance schedule parameter that determines how much noises are added at each step. These noisy embeddings can then serve as the training data for the reverse diffusion process of SD. Conditional information can be applied through cross-attention heads in the attentionbased U-Net. For example, to match the dimensions of the SD inputs, the outputs y of the gene expression encoder can be further projected with a projector xe into a transcriptomic prompt re(y) e
aim t0 |earn toe reverse diffusion process formulated by q(zt-i\zt,T0(y ). During the fine-
tuning process, the gene expression encoder and the U-Net can be optimized together. The following loss function may be used:
where eg is the denoising function implemented as a U-Net. In the sampling process, the diffusion model can take the Gaussian noise as the initial input and denoise it iteratively through T steps. Finally, the reverted image latent embedding z0 can be obtained and input into the decoder ©(•) to generate the cellular images.
Aligning transcriptome, images, and text
By jointly training GenVinci gene expression encoder and the SD model, one can fine-tune the projector to get better transcriptomic prompts for cellular image generation tasks. However, the original SD model may be tailored for the text-to-image generation task, the unique properties of gene expression data result in latent spaces that are quite different from those of texts. This means that simply fine-tuning the SD model end-to-end with a limited set of gene expression-image pairs may make it difficult to yield an effective alignment between the gene expression and text embeddings.
Fortunately, the SD model may already exhibit well-aligned text and image spaces by training on billions of image-text pairs. Therefore, BiomedCLIP, a biomedical vision-language foundation model trained on the PMC-15M corpus, can be incorporated to enhance the alignment of spaces related to the transcriptome, images, and text. The PMC-15M corpus includes 15 million figure-caption pairs extracted from over 3 million biomedical research articles in PubMed Central, which includes a diverse range of biomedical images. Specifically, previous gene expression projector with an additional alignment projector can be followed to refine the transcriptomic embeddings into multi-modal pseudo vision-language embeddings. These embeddings can be adjusted to be either dimensionally or distributionally compatible with those used in BiomedCLIP. During training, a weak supervised constraint may be applied to minimize the similarity between the transcriptomic embeddings and the multi-modal embeddings produced by the BiomedCLIP image encoder, which is defined as follows:
where h represents the alignment projector, I represents the cellular image sent to the BiomedCLIP image encoder E,. The loss function can serve to encourage closer alignment between the transcriptomic embeddings and the cellular image embeddings, and by extension, it also can align with the potential textual description features of the images. As a result, the spaces for transcriptome, images, and texts can be unified, and the optimized transcriptomic representation can become more suitable for SD image generation, thereby improving the quality of the generated images. Throughout the fine-tuning phase, the parameters of the pre-trained BiomedCLIP model can remain fixed. Finally, the entire loss function can be formulated as:
^All = SD + oup where the weight A is a hyperparameter that controls the degree of supervision from BiomedCLIP.
T anslating gene expression matrix to whole-slide image
The potential of using the methods described herein to translate gene expression matrices into whole-slide images can be implemented. However, unlike the numerous gene expression-image pairs at the single-cell level, there may be only a few paired gene expression matrices and whole-slide images available. In addition, the varying number of cells captured across different tissue slides presents a challenge and limits the feasibility of generating whole-slide images. Despite these constraints, a matrix- to-slide GenVinci can be successfully implemented to leverage the previous pre-trained GenVinci by gene expression-image pairs.
Specifically, the goal can be to generalize GenVinci from a cell-to-tile level to a matrix-to-slide level. The process may begin with resizing whole-slide images, e.g., to n x n pixel images, e.g., 512 x 512-pixel images, to match the output dimensions of the SD model. Next, for each gene expression matrix, cells with insufficient information, e.g., those with a total count of fewer than 200, can be removed. Then, for example, n (e.g., 512) cells can be randomly sampled to represent the tissue transcriptome signals. In cases where matrices contain fewer than n (e.g., 512) cells, they can be padded with zero- expressed genes to maintain consistent dimensions. This approach is adaptable across various imaging and sequencing technologies, accommodating variations in the number of captured cells. Next, the pretrained cell embeddings can be extracted for each cell within the gene expression matrix by using the gene expression encoder of cell-to-tile pre-trained GenVinci. These embeddings can be passed through a new projector to reshape the matrix to fit the requirements of the SD model. The projector can be further fine-tuned, e.g., in conjunction with the U-Net component of the SD model. All other modules and parameters may remain consistent with those used in the cell-to-tile GenVinci.
Evaluation metrics
Various evaluation metrics may be employed to assess the accuracy of the GenVinci model as described below, e.g., in order to determine if the output image is consistent with the input gene expression profile and empirical databases of two-dimensional and three-dimensional images.
CNN baseline
A Convolutional Neural Network (CNN) can be implemented as the baseline for the gene-to- image translation task to compare with GenVinci. The CNN baseline is inspired by the decoder network of U-Net, which consists of fully connected layers and transpose convolutional layers to increase the spatial dimension, and batch normalization and ReLU activations to enhance the model's ability to learn nonlinear functions. The CNN baseline can be trained to predict cellular images from the normalized gene expression using MSE loss. The top 2000 highly variable genes (HVGs) can be selected as input for the datasets with whole transcriptomic sequencing. The model can be trained for 500 epochs with the Adam optimizer, using a learning rate of 0.001 and a batch size of 32.
Frechet Inception Distance (FID)
FID measures the difference between the distributions of real and generated images. Specifically, the ImageNet-pre-trained lnception-V3 convolutional neural network can be adopted to extract 2048- dimensional feature vectors for both real and generated images. These feature vectors capture the
essential visual information of the images, as learned by the network during its training on ImageNet dataset. Then the mean and covariance can be computed for each set of features to calculate the Wasserstein-2 distance. A lower FID score indicates that the distributions of generated images are closer to the real images, suggesting better quality and realism of the generated images. FID is expressed as:
where x denotes the feature vectors of the real images, g denotes the feature vectors of the generated images. Tr denotes the trace of a matrix, which is the sum of all the diagonal elements.
Cosine similarity
Cosine similarity can also be used to measure the similarity between the real and generated images. Cosine similarity is a mathematical measure used to determine the similarity between two vectors in a multi-dimensional space. The cosine similarity can be calculated by the extracted 2048- dimensional feature vectors of the real and generated images from an ImageNet-pre-trained ResNet-50 model.
Accuracy of annotation and clustering tasks
To further determine whether the generated cellular images not only reflect gene expression constraints but also preserve pathological signals, classifiers can be trained based on the real and generated images for cell-type or pathological annotations. Specifically, the image features can be extracted from an ImageNet-pre-trained ResNet-50 model with classification heads tailored for different annotation tasks. Both Accuracy and F1 scores can be reported as the evaluation metrics. The higher Accuracy and F1 score may indicate better classification performance. For clustering performance evaluation, one or more commonly used metrics can be adopted, such as normalized mutual information (NMI), adjusted Rand index (ARI), and average silhouette width (ASW). The NMI score can be computed to measure the concurrence between ground truth cell type labels and Louvain cluster labels from integrated cell embeddings, with clustering resolutions ranging from 0.1 to 2. The ARI can be used to assess the agreement between annotated labels and NMI-optimized Louvain clusters, adjusting for randomly correct labels. The ASW score, ranging from -1 to 1 , can be calculated to evaluate the relationship between a cell's within-cluster distances and its distances to the closest cluster boundaries. Higher scores for NMI and ARI may indicate better matches, while an ASW score of 1 may indicate well- defined clusters.
Implementation details
The same hyperparameters as the Geneformer can be used during the pretraining of GenVinci. Specifically, a linear learning rate scheduler with warmup can be employed, setting the maximum learning rate, e.g., to 5 x 10-5. An Adam optimizer can be used, e.g., with a weight decay of 0.001 and a batch size of 16. The dimensionality of the input and output can be, for example, 256, and the inner layer can have a dimensionality of, e.g., 512. For all fully connected layers, a nonlinear activation function may be utilized in the form of a Rectified Linear Unit (ReLU); a dropout probability of 0.02; a standard deviation of 0.02 for the initializer of weight matrices; and an epsilon value of 1 x 10-12 for layer normalization layers.
For the fine-tuning of the SD model, SD version 1 .5 can be utilized for image generation and denoising diffusion implicit model (DDIM) sampling strategy, with a batch size of, e.g., 4. The Adam optimizer may be employed, setting the learning rate to, e.g., 5 x 10-5, for training. Each model can be fine-tuned for 500 epochs, and the number of sampling steps can be set to 250. The parameters of the BiomedCLIP model may always remain fixed.
Datasets
The methods described herein include providing a gene expression profile (e.g., from sequencing and/or hybridization data) of a plurality of genes from a cell or organelle. The gene expression profile may be obtained from one or more known sequencing and/or imaging methodologies known in the art, e.g., as described in more detail below.
For example, in some embodiments, the spatial transcriptomic analysis includes single nucleus RNA (snRNA) sequencing, single cell RNA (scRNA) sequencing, patch sequencing, Visium, Xenium, NanoString GeoMx, STARmap, Slide-Seq, Slide-Tags, fluorescence in situ hybridization (FISH), multiplexed error-robust fluorescence in situ hybridization (MERFISH), electron microscopy, hematoxylin and eosin (H&E) imaging, fluorescence microscopy, 4',6-diamidino-2-phenylindole (DAPI), immunofluorescence (IF), or a combination thereof.
Examples
The following example is put forth so as to provide those of ordinary skill in the art with a description of how the systems and methods described herein may be used and evaluated and is intended to be purely exemplary and is not intended to limit the scope of the disclosure.
Example 1 : Generation of cellular and tissue images from single-cell expression profiles with GenVinci
Recent advances in generative Al and deep learning have opened the way to generate multiple views through non-linear transformation and may offer a way to unify different aspects of cells and tissues. Here, is introduced GenVinci, a generative model based on transformer and diffusion models that learns a multi-modal representation of cells and tissues and reconstructs morphological information from transcriptome profiles to generate cellular and tissue images from single-cell expression profiles during inference. It was demonstrated how GenVinci can generate diverse types of imaging data, including histological (hematoxylin and eosin (H&E)) stains, fluorescence microscopy images, neuron morphology, and electron microscopy, from various molecular profiles at different resolutions. GenVinci generates accurate cellular and tissue images compared to ground-truth morphologies, cell types, and pathological annotations. It outperforms a convolutional neural network (CNN) baseline in both generation quality and semantic mapping. Thanks to its cascaded architecture, GenVinci can be flexibly applied to a range of downstream tasks, including image generation, multi-modal clustering, zero-shot classification, and data
augmentation. GenVinci can unify different views of cell and tissue biology, while greatly reducing the need for multiple measurements, towards an ultimate goal of virtual cell and tissue simulators.
Currently, studies largely rely on one measurement modality, and, if a multi-modal view is desired it requires direct measurement, ideally in the same cell or tissue. This poses several challenges. First, many measurement modalities are costly, laborious, and time-consuming, requiring specialized assays and expertise. As the number of experimental methods grows at a staggering pace, so does the experimental complexity. Second, independent methods often require separate specimens, which is especially difficult in a clinical setting, when samples are often limited in quantity (e.g., sections from miniscule needle biopsies) or can only be collected in certain ways. Finally, different modalities may have no immediate correspondence in their nominal variables (e.g., gene levels vs. non-molecular imaging features), and thus hard to unify to a coherent view. Recent studies have begun to tackle these challenges with an increasing number of multi-modal measurements in the same cell or tissue section, or by pioneering computational frameworks to infer data from imaging to molecular modalities. However, no approach focused on revealing hidden morphological information from expression profiles to generate cell and tissue images from single-cell expression profiles. Such methods would be especially beneficial, because cell and tissue imaging, especially histology, are a bedrock of clinical diagnostics and research, offering clear, tangible, and clinically interpretable results. Thus, equipping expression profiles with rich imaging would facilitate their physiological, functional, and clinical interpretation. For example, complex neuronal morphologies are highly cell-type-specific, essential for their function, and are associated with genomic mutations, gene expression levels, and methylation patterns. Therefore, linking molecular information with cell morphology would provide key insights for further studying the important functions of cells such as the formation of neural circuits, stem cell differentiation, and cancer metastasis.
Generative artificial intelligence (Al) methods offer a compelling new approach to tackle this challenge. Al-generated content has made a profound impact across a wide range of fields. Text-to- image generation models like DALL-E 2, Stable Diffusion, and Midjourney have created complex and realistic images based on textual input, surpassing all previous generative adversarial network (GAN) and variational autoencoder (VAE)-based methods in quality metrics and synthesis capabilities. The latest video-generation model, Sora, also creates realistic and imaginative scenes from text instructions and serves as a foundation to understand and simulate the physical world. All these methods rely on diffusion models, a novel paradigm that is different from previous GAN and VAE-based approaches and generates images through the systematic introduction of noise with repeating steps.
Here, advances in text-to-image diffusion models and large-scale biomedical and single-cell language models were leveraged to introduce GenVinci, a single-cell gene-to-image generative model, that bridges the gap between expression profiles and cell morphology. GenVinci pretrains a state-of-the- art single-cell language model to encode gene expression into a robust latent representation and then fine-tunes a text-to-image diffusion model to generate cellular images based on the latent transcriptomic prompt. GenVinci further employs a biomedical vision-language model to provide supervision, aligning transcriptome, images, and texts into the shared latent space. This enhances our understanding of cell states and transitions and should help translate molecular profiles into physiologically and clinically interpretable results.
Results
The GenVinci model
GenVinci is a generative model that learns a multi-modal representation of morphological and molecular profiles and uses it for gene-to-image translation across various profiling platforms and imaging techniques (FIG. 1A). GenVinci’s cascaded architecture allows one to flexibly apply it across a range of downstream tasks, including image generation, multi-modal clustering, zero-shot classification, and data augmentation (FIGS. 1A, 1C, and 1D).
Briefly, the GenVinci model was trained in two stages: an initial general-purpose pretraining on expression data, followed by fine-tuning with paired imaging data for the gene-to-image translation task (FIGS. 1B, 1C, and 1D). In the pretraining step, GenVinci trains a single-cell large language model to encode high-dimensional expression data into robust latent embeddings. GenVinci was pre-trained on GenCorpus-210M, a large-scale scRNA-seq corpus of 210 million human and mouse single-cell and spatial RNA-seq profiles from a broad range of tissues. To obtain better cell embeddings specific to the utilized profiling techniques, GenVinci can be further pre-trained on additional relevant gene expression data, using the original masked language modeling strategy, where some genes or gene expression values are randomly masked in a cell and then predicted (Methods, FIGS. 1C and 7A).
In the fine-tuning stage, the pre-trained Transformer was cascaded with Stable Diffusion (SD), a diffusion model where the extracted gene expression embeddings serve as transcriptomic prompts to generate cell images (Methods, FIGS. 1D, 7B, and 7C). Since the SD model is specifically designed for text-to-image generation and was trained on billions of image-text pairs, it may not generalize well with limited gene expression-image pairs. Thus, another biomedical vision-language model, BiomedCLIP, was employed to provide weak supervision to reduce modality discrepancies for better aligning the transcriptomic prompts with image and text embeddings (Methods, FIGS. 1D and 7D). BiomedCLIP extracts multi-modal cell image embeddings that are well-matched with their potential text embeddings. These are then used to refine the cell embeddings from GenVinci. Consequently, the enhanced cell embeddings align more closely with both image and potential text embeddings, rendering them more effective for fine-tuning the SD model for cell image generation, thereby improving the quality of the resulting cell images.
During the inference stage (FIG. 7E), GenVinci generates multiple cell images from a single expression profile or signature, reflecting various perspectives of the cell morphology, which is of great significance even when only partial or individual sampling of gene expression distributions is available, rather than a complete profile (e.g., with methods like multiplexed error-robust fluorescence in situ hybridization (MERFISH)) or full distributions representing cell states. It then ranks the generated images that are closest to the ground truth based on similarity scores between the input expression profile and the generated images. This flexible, rank-based approach to inference can be highly effective in real- world clinical settings when pathologists or physicians want to avoid incorrect predictions and choose the most confident generation.
GenVinci’s cascaded architecture overcomes the limitation of having few expression-image pairs for fine-tuning and enhances the cell representations obtained from the pretraining step for more effective
image generation. Furthermore, by including large language models, GenVinci can be flexibly applied to diverse downstream tasks, leveraging the models’ capabilities in a multi-modal manner.
GenVinci generates H&E histology images
First was evaluated GenVinci’s ability to generate Hematoxylin-and-Eosin (H&E) histology images from single cell (sc)/ or single nucleus (sn) RNA-seq of metastatic breast cancer (MBC) tumors from the Human Tumor Atlas Project Pilot (HTAPP) (Datasets). In this study, a single tumor specimen was split such that one part was profiled by sc/snRNA-seq and another was sectioned serially, with consecutive sections spatially profiled by MERFISH (212 genes) or stained by H&E. GenVinci’s generalization ability was evaluated at two imaging resolutions. For snRNA-seq data (sn-MBC), H&E images were segmented to single cells and paired cell images with snRNA-seq profiles to create the task of generating single-cell H&E images from snRNA-seq profiles (snGE => scH&E). For scRNA-seq (sc- MBC), 1 -10 cell tile was segmented around each single cell and aimed to generate multi-cell H&E images (tile) from one scRNA-seq gene expression profile (scGE => mcH&E). In addition, generating multi-cell tiles will be more meaningful in real-world clinical settings, as they are more informative for pathological annotation and better reflect the cellular microenvironment compared to single-cell images. GenVinci was further tested on two other human breast cancer (HBC) datasets (Visium-HBC and st-HBC) obtained by Visium (1 -10 cells/spot) and Spatial Transcriptomics (5-100 cells/spot) (Datasets), which measure whole transcriptome but at less than single-cell resolution. In this case, the H&E tiles were segmented by spot coordinates and paired them with the corresponding spot’s expression profile to create the multi-cell gene expression data to multi-cell H&E task (mcGE => mcH&E). The st-HBC dataset was also used to evaluate the direct generation of whole-slide images from gene expression matrices (wsGEM => wsH&E). Given the comprehensive spot-based gene expression dataset, which lacks spatial information, the capabilities of the single-profile level GenVinci were extended to generate whole-slide images from the corresponding paired gene expression matrix.
GenVinci’s robustness was evaluated in an out-of-sample setting. Specifically, for the sn-MBC and sc-MBC datasets, one patient was randomly held out for testing and used the remaining (n=3 and n=4) for training. For the Visium-HBC dataset, one section was used for training and one for testing. For the st-HBC dataset, samples from four patients were held out for testing and n=19 samples were used for training. For each nucleus, cell, or spot profile in the test data, 5 cell images were randomly generated and ranked them by the normalized cosine similarity scores between the input expression and generated image embeddings to compare with the ground-truth images (FIGS. 2A-2E). GenVinci was compared to a “CNN baseline” model consisting of several transposed convolutional layers inspired by U-Net decoders (Methods, FIG. 8), trained to predict cell images from expression data using mean squared error (MSE) loss. The image with the highest similarity score was selected among the five generated images for benchmarking comparison.
GenVinci-generated images were qualitatively (by visual inspection) and quantitatively very similar to the ground-truth images (FIGS. 2A-2F). Qualitatively, GenVinci images reflected nuances of complex shape, cell intensity patterns, color distribution, and even tissue profile across all resolutions (FIGS. 2A-2E, in each figure, the first column on the left contains the ground-truth images), while CNN baseline could not generate realistic H&E images, especially multi-cellular (tile level) ones, and single-cell
images (on sn-MBC) were blurred (FIG. 2F). Quantitatively, GenVinci images had much smaller distribution differences with ground-truth images than CNN baseline by the Frechet Inception Distance (FID; a standard metric for GAN models that measures the distance between the distribution of generated and ground-truth images) (FIG. 2G, GenVinci FIDs less than 100 in most cases, vs. more than 200 for CNN baseline), and by average cosine similarity (between the ground-truth and generated images (Methods), FIG. 2G, with similarity score > 0.8 in all cases, vs. less than 0.5 for the CNN baseline). Additionally, as a visualization, Uniform Manifold Approximation and Projection (UMAP) of the image embeddings from the ground-truth, GenVinci, and CNN baseline images (FIG. 2H) also show good agreement between ground truth and GenVinci (but not CNN baseline) distributions for Visium-HBC and st-HBC datasets, and less so for sc-MBC (but still better than CNN baseline).
Next were generated whole-slide images from gene expression matrices (Methods, FIG. 2E). Pathologists often require whole-slide information for accurate cancer diagnosis, while annotating individual small tiles is difficult. Because slide-based spatial transcriptomics data are still limited and varied, it is challenging to effectively train a deep learning model. It was sought to extend the capabilities of the current GenVinci model from translating individual cells/spots to tiles to translating entire gene expression matrices to whole-slide images (FIG. 9). Performance on the st-HBC dataset, which contains 68 paired slides and gene expression matrices, was evaluated. Given the limited size of the training data, the generation results were not as refined as those at the single-profile level; however, some consistency in morphology and structure in the generated whole slides was still observed (FIG. 2E).
GenVinci generates fluorescence cell morphology
GenVinci was applied on a CosMx SMI dataset (with each cell profiled by 960 genes) from one non-small cell lung cancer (NSCLC) tissue sample to generate fluorescence microscopy images from single-cell gene expression data (Datasets). To evaluate the capability of GenVinci at a multi-scale level, two datasets were derived based on different resolutions (sn-NSCLC and sc-NSCLC). For the sn-NSCLC dataset, multiple z-stack DAPI images were merged using maximum intensity projection to create two- dimensional DAPI images and collected gene expression in segmented nuclei, to construct the task of generating single-nucleus DAPI images from the genes measured in those segmented nuclei (snGE => snDAPI). For the sc-NSCLC dataset, composite images of immunofluorescence (IF) intensity were used from membrane markers and gene expression within cell boundaries were collected to create a task that generates whole-cell IF-stained images from the genes measured in those segmented cells (scGE => scIF). In the composite IF-stained images, one channel represents DAPI, one channel represents PanCK (an epithelial cell marker), one channel represents CD45 (a marker of all nucleated hematopoietic cells), and one channel represents CD3 (a T cell marker). Since there is only one tissue sample in this dataset, in both cases, the gene expression-image pairs were randomly split from two fields of view (FOVs) for testing and retained the remaining 28 FOVs for training.
GenVinci recovered the fluorescence intensity patterns and generated realistic single-nucleus DAPI images from single-nucleus gene expression data and even recovered the distribution of different marker stains in single-cell IF images from the gene expression information (FIGS. 3A and 3B). Compared to the CNN baseline (FIG. 3C), GenVinci showed superior generation quality and distribution similarity (by FID, FIG. 3D, top) and higher cosine similarity between generated and ground-truth images
(FIG. 3D, bottom), and high overlap of image embeddings (FIG. 3E), demonstrate GenVinci’s ability to generate well across different data modalities. Note that across all metrics, GenVinci performed less well on the DAPI image generation task versus the IF-stained image generation task. It was posited that this may be due to single-channel DAPI images, which are not specific to genetic or molecular markers, providing a weaker correlative signal with gene expression data. In contrast, multi-channel IF-stained images, which profile multiple proteins, offer a closer correspondence to RNA expression, encompassing a broader and more detailed spectrum of morphological features that reflect molecular signals.
GenVinci generates whole-neuron morphology
Next, GenVinci was fine tuned with Patch-seq data to generate whole-neuron morphology from expression profiles. Patch-seq combines electrophysiological recordings with single-cell whole- transcriptome profiling in individual neurons. Whole-neuron morphology is very complex, with intricate dendrite and axon branching structures, and this morphology has functional implications. Two Patch-seq datasets were used with both single-cell expression profiles and morphological reconstruction from the mouse visual and motor cortex. Mouse genes were mapped to human orthologs with mousipy package (github.com/stefanpeidli/mousipy) (Methods), projected the reconstructed neuron morphologies onto the xy-plane and color-coded the pixels by neuronal component using neuroM, and merged the two datasets to achieve a sufficient number of cells (k=1 ,245) with cellular morphology images. The task of generating neuron morphology was constructed from single-cell gene expression data (scGE => scNM), by randomly sampling 90% of cells for training and 10% held out for testing across all the cell types.
Despite the limited number of cells in the training, GenVinci generated highly realistic neuron morphologies that reflect the provided Patch-seq gene expression and ground-truth images qualitatively and quantitatively (FIGS. 4A-4C), while CNN baseline could not (FIGS. 4B-4D). In particular, GenVinci generated accurate soma and neuron synapses and can distinguish between the dendrite and axon structures from different channels (FIG. 4B), with image embeddings that closely match those of the ground truth (FIGS. 4C and 4D).
GenVinci generates electron microscopy images
GenVinci’s performance in generating electron microscopy (EM) images, a specialized data modality that is challenging to generate experimentally, was next tested. EM imaging provides a nanometer-resolution view of tissue ultrastructure, but its acquisition is both expensive and timeconsuming, and generating it through simpler profiling modalities could be transformative. To fine-tune GenVinci, recently collected data was leveraged, where large area scanning EM and MERFISH were collected in two consecutive sections from the brain of a mouse with locally induced injury, followed by spatial alignment and projection into a common coordinate system (FIG. 10A). Notably, the EM images and MERFISH profiles do not match perfectly in the coordinate space and when aligned with a nearest neighbor search, different EM images were occasionally mapped to the same MERFISH profile, leading to loss of more than half of the gene expression-image pairs (FIG. 10B). The Hungarian algorithm was employed, thus ensuring an optimal solution to the assignment problem by finding a one-to-one correspondence between points in two sets by minimizing the total cost across all pairings. This approach allowed achievement of a one-to-one pairing, although some EM images did not structurally match their
assigned profiles (FIG. 10C). A task was then implemented to generate a single-cell EM image from a MERFISH gene expression signature (scGE => scEM).
Despite the lingering inconsistencies in structural matching between the EM image and its neighboring profiles and the limited training samples in this dataset (i.e., less than 1 ,000 cells), GenVinci still performed well (FIGS. 5A and 5D). This may be beneficial from the supervision of BiomedCLIP, which incorporated EM-text pairs in its training corpus. Furthermore, GenVinci generates far more accurate EM images compared with the CNN baseline (FIGS. 5B-5D).
GenVinci uncovers multi-modal information from unseen embeddings
To further investigate whether GenVinci learns multi-modal embeddings through generative training, the cell/spot embeddings extracted from Geneformer (a recently published single-cell foundation model), continued pre-trained Geneformer, and gene expression-image fine-tuned GenVinci were visualized and compared (FIG. 6A). Considering the limited available cell-type annotations and dataset sizes that may not generate meaningful clusters in some of our datasets, only the st-HBC, sn-MBC, and sn-NSCLC datasets were investigated. For st-HBC, theretrieved the spot-level pathological annotations for visualization wree retrieved, as each spot may contain one to tens of cells. Three cell type-clustering metrics, normalized mutual information (NMI), average silhouette width (ASW), and adjusted Rand index (ARI) were adopted to further evaluate the clustering performance (Methods). GenVinci improved the identification of cell types compared to the original embeddings (FIGS. 6A and 11 B). For both the sn- MBC and sn-NSCLC datasets, cells with the same annotations are well-clustered together (e.g., endothelial cells), demonstrating the efficiency of continuing GenVinci to generalize it to spatial transcriptomics data (FIG. 6A, left and middle, FIG. 11B, upper).
Moreover, GenVinci learned multi-modal embeddings from the gene-to-image task that helped distinguish tissue pathology and cell types. For example, tumor and normal spots were separated in the st-HBC dataset, while tumor, endothelial fibroblast, and plasmablast cells were separated in the sn- NSCLC dataset (FIG. 6A, right, FIG. 11B, upper). However, GenVinci also lost some cell-type information when training to translate the gene expression data into cellular images (e.g., hepatocytes in the sn-MBC dataset were clustered in GenVinci but scattered in the GenVinci embedding). This was attributed to the discrepancy between gene expression data and cell images, as shown by the UMAP visualization of the original image embeddings (FIG. 11 A, left, FIG. 11B, bottom). In addition, the DAPI images of nuclei from all cell types in the sn-NSCLC dataset were grouped together, making it challenging to discern distinct groups within the original DAPI image features (FIG. 11 A, bottom). This suggests that the DAPI images may encode less information compared to the gene expression data. Therefore, during the training of the gene-to-image task, the model might learn more general or even noise information rather than the specific cell types.
GenVinci's potential in integrating single-cell expression data and image features was also explored. Specifically, the original gene expression embeddings were concatenated with ground-truth image embeddings and compared with GenVinci’s refined gene expression embeddings concatenated with its generated image embeddings (FIGS. 11 A, middle and right, FIG. 11B, bottom). Compared to the original gene expression and image features, the refined gene expression embeddings paired with generated image features captured more heterogeneity within the cell types or pathological annotations
(FIG. 11 A, right and FIG. 11B, bottom). For example, malignant cell spots were clustered together in the st-HBC dataset, and endothelial cells and hepatocytes were separated in the sn-MBC dataset. Besides, the clusters of endothelial, epithelial, and fibroblast cell types were better delineated in GenVinci embeddings of the sn-NSCLC dataset.
Zero-shot annotation with GenVinci
The incorporation of BiomedCLIP and SD for the gene-to-image translation task refines the original cell embeddings into transcriptomic prompts with multi-modal image and textual information. Since GenVinci leverages BiomedCLIP to align transcriptome and images into a shared latent space, it is feasible to implement zero-shot-based cell-type or pathological annotation tasks based on the gene expression data. Zero-shot annotation involves classifying objects into different categories using a model that was not explicitly trained on data containing labeled examples from those specific categories. This approach is especially valuable in biological fields, where labels may be scarce or difficult to collect. It allows for rapid adaptation of GenVinci to new biological tasks, such as identifying rare cell types or novel disease markers, without the need for extensive retraining.
Specifically, given a cell’s gene expression profile, GenVinci infers its annotation by calculating the similarity between cell embeddings and textual prompt embeddings without additional training. The cell embeddings can be represented by gene expression embeddings extracted from the projector or by the generated image embeddings from the vision-language model (VLM) image encoder (FIGS. 1D and 12). This is similar to the zero-shot image classification in the CLIP model. As BiomedCLIP is trained with billions of biomedical image-text pairs in the PMC-15M corpus, this was solely focused on general annotation tasks evaluated in the BiomedCLIP paper, i.e. , tjhezero-shot-based pathological annotation task on datasets with H&E images were exclusively tested. It was reasoned that if the model is well- trained, then the extracted cell embeddings should exhibit high similarity with their paired image embeddings (as well as any potential text embeddings that can describe the images). Based on this insight, prompts were created similar to those used in BiomedCLIP to perform the zero-shot annotation task. Specifically, the well-trained GenVinci was used to extract cell/spot embeddings from the test gene expression data of the st-HBC and sn-MBC datasets. Two text prompts were then created relevant to these datasets: 'This is a photo of breast cancer tissue' and 'This is a photo of normal breast tissue.' Next, text embeddings of these prompts were extracted and their similarities calculated with the extracted cell/spot embeddings, and aligned the class mentioned in the prompt with the cells/spots exhibiting the highest similarity (FIG. 12). In this non-training paradigm, annotating gene expression is rapid and does not require an expert or annotation tool. The performance of using cell embeddings (both from the generated cellular images and from gene expression data) was evaluated for zero-shot annotation and compared it with ground-truth image embeddings (FIG. 6B). Both gene expression embeddings and the generated image embeddings by GenVinci consistently achieved comparable or better annotation performance compared to the ground-truth images. This underscored the potential of GenVinci for clinically relevant annotations, particularly in cases with limited examples of rare conditions.
Data augmentation with GenVinci
Given the need for expertise in pathological annotations in spatial transcriptomics and the limited availability of publicly annotated histology databases at the cellular level, the feasibility of using GenVinci to generate H&E images for data augmentation (from well-annotated scRNA-seq) to enhance current H&E-based pathological annotation models was explored (FIG. 13). Specifically, classification models were trained using either ground-truth or GenVinci-generated H&E images and then assessed their performance in classifying real H&E images. A MERFISH sn-MBC dataset (considering its single-cell resolution for both H&E and gene expression data) was used for evaluation. The same as before, one patient was randomly held out for testing and used the remaining (n=3) to train GenVinci. Then the held- out patient was split into 90% training and 10% test. For benchmarking, GenVinci was used to generate the same number of training images for comparison. A ResNet-50 model was used to extract image features and implement a linear probing classifier for the pathological annotation task. The sampling frequency was varied to generate different numbers of images by GenVinci and trained the model to compare their performance to that of real images (FIGS. 6D-6F). Pathological annotations were obtained from the single-cell gene expression data and used as labels (at single cell level) for both ground-truth and generated images.
While the classifier trained only on generated images performs worse than one trained on real images, increasing the number of generated images improved performance (FIG. 6D), as is commonly observed with deep learning applications. Moreover, when real training data was augmented with generated images from both MERFISH sn-MBC gene expression data (FIG. 6E) and from an external scRNA-seq dataset (FIG. 6F), the augmented models achieved a 4-10% compared to a classifier trained on real data only. Specifically, the model augmented with scRNA-seq data showed superior performance compared to that augmented on MERFISH data. This supports the conclusion that the limited number of genes in the MERFISH panel may provide inadequate conditional information, leading to relatively poor performance in cellular image generation.
Discussion
This study shows the capabilities of GenVinci, the first generative model designed to translate single-cell gene expression into various cellular and tissue images. GenVinci represents a significant leap in realizing digital biology, aiming to bridge the gap between the complexity of gene expression data and the interpretability offered by traditional imaging techniques. These findings support that GenVinci not only enhances the accessibility of gene expression data for clinical and research applications but also paves the way for innovative approaches to understanding cellular biology.
Specifically, GenVinci addresses this issue by generating visual representations of cellular and tissue structures from transcriptomic data, thereby enhancing its interpretability and potential utility in clinical diagnostics and research. Its performance was evaluated from a range of sequencing and imaging measurements at different resolutions. GenVinci shows strong generalization and robustness.
GenVinci offers a novel conceptual and technological framework for visualizing unseen information in gene expression data, facilitating a more integrated understanding of cellular function. GenVinci introduced a conceptual shift in generating new and hidden multi-modal biological information without measuring all modalities.
Methods
GenVinci is a deep generative model that can generate cellular morphology from single-cell gene expression. It consists of three modules: a Bidirectional Encoder Representations from Transformers (BERT)-based single-cell language model, a Latent Diffusion Model (LDM), and a biomedical Contrastive Language-Image Pretraining (CLIP) model. At the core of the approach is the LDM module, which maps single-cell gene expression embeddings to cellular images. LDM approaches typically accept natural language embeddings, which are then translated into corresponding images. Because of this, GenVinci trains a single-cell language model that produces gene expression embeddings in a manner similar to natural language models. However, since the representations learned by GenVinci are unaligned to the pretraining of the LDM, BiomedCLIP, a model that learns to align biomedical figure captions to their corresponding images, was employed to supervise the fine-tuning of the entire approach. The integration of three modules allows GenVinci not only to generate realistic cellular images but also to be easily applied for diverse downstream tasks. These modules were introduced along with their tailored pretraining and fine-tuning strategies for GenVinci, as described in the subsequent section.
Unified transcriptomic representation
Given that the raw gene expression data is usually sparse, high-dimensional, and noisy, GenVinci trains a single-cell language model to encode gene expression data into a low-dimensional latent space. GenVinci is a BERT-like transformer model pre-trained on a large-scale corpus (GenCorpus-210M) of approximately 210 million single-cell and spatial transcriptome to allow context- aware predictions in settings with limited data in network biology. Benefiting from the flexible nature of the attention mechanism, GenVinci can accept single-cell gene expression data from different sequencing measurements without the need to standardize into uniform gene lists. In addition, domain knowledge learned from an extensive single-cell transcriptomics corpus, particularly concerning gene-gene interactions, enhances cellular representation across a spectrum of downstream tasks.
Since GenVinci is a transformer-based approach that expects inputs to be sequences, a binning- based rank value encoding method is employed to induce a sequence from the (tabular) gene expression observations. Rank value encoding is a non-parametric approach that ranks genes by their expression levels within a single cell. Initially, given a cell-by-gene matrix, only protein-coding genes was retained based on a predefined vocabulary by GenCorpus-210M. Subsequently, the gene expression data is first normalized using a scale factor, followed by a log transformation to bring them to a comparable scale. The bin size is set to be 50. Specifically, let x^ represent the expression value for a gene i in cell c, the normalization gene expression x' is given by: x^ = log (
where n is the total number of genes within the cell. After normalization, genes within each specific cell were assigned into bins using a quantile binning method and ranked by their binned values in descending order and subsequently tokenized according to the predefined vocabulary, where each gene gk is assigned an integer identifier id gk). The vocabulary also includes two more special tokens for padding and masking and context tokens to include the cell context information (e.g., modality and assay). The
rank value encoding method prioritizes genes that uniquely characterize cell states by using expression data across the GenCorpus-21 OM, thereby downgrading common housekeeping genes and elevating critical transcription factors that define cell identity. It offers robustness against technical biases, maintaining the relative gene ranking within each cell despite potential variations in absolute transcript counts. Furthermore, top m expressed genes with the normalized expression value x' > 0 for were selected for each cell. For the cells with the number of expressed genes less than m, padding tokens are added to meet the required length. In this study, m is set to be 1024 (which fully represents 90% of rank value encodings in GenCorpus-21 OM). The final embeddings of the tokenized genes for cell c are denoted as:
where the ranked genes are represented by their integer identifiers, and Embg is a learnable conventional embedding layer for the gene tokens. Similar to the position embedding in BERT, trainable binning-based expression embeddings denoted as EEC can be used for each gene token within the cell, which allows the modeling of the relationship between the binned gene expression values. GenVinci also includes the conditional embedding to distinguish the context tokens and gene tokens which is denoted as CEc. The final input embedding for each cell can be represented as the element-wise sum of the gene and binning- based expression embeddings:
GEC — GEC (J EEC (J CEc where GE' e 7?mxd, d is the dimension of each token embedding vector. The embeddings were then sent into the transformer blocks.
Spatially resolved transcriptomics pre-trained transformers
GenVinci is composed of six identical transformer encoders, each featuring two sub-layers. The first one is a multi-head self-attention layer, followed by a simple position-wise feed-forward neural network layer. The scaled dot-product attention is implemented in the self-attention layer as:
Where Q,K, V represent the query, keys, and values, respectively. These are sets of vectors derived from the input embeddings. dk is the dimension of the queries and keys, and -5= is a scaling factor to control
V“/c the magnitude of attention weights. The output is computed as a weighted sum of the values, where the weight is calculated by a compatibility function of the query with the corresponding key. Four attention heads were used per transformer encoder to jointly attend to information from different representation subspaces:
and W° are the parameter matrices of projections. Two feed-forward neural network layers are followed by the self-attention layers, which are applied to each position separately and denoted as:
FFN(x)' = max (0, x W1 + b W2 + b2
The linear layers share the same architecture but have different parameters across various positions. During pretraining, 15% of the genes or gene expression values within each specific cell were randomly masked in spatial transcriptomics data, and the model was trained to predict the masked genes or gene expression values based on the context of the remaining unmasked spatial genes or gene expression values. The cross-entropy loss is employed as the masked learning objective function, which for a single masked token can be formulated as:
where V is the size of the vocabulary. yt and ft are the ground truth and predicted probabilities of the i-th gene token, respectively. Herein, masked language modeling loss was computed for all masked gene or gene expression tokens within a batch, and this loss was then averaged to produce the final loss against which the model parameters are updated. In addition, instead of a single feedforward network like in standard Transformers, GenVinci Transformers use Mixture of Experts which consists of multiple expert FFNs, with a gating network that dynamically selects a few experts for each input token. This approach allows the model to scale efficiently, increasing capacity while keeping computational costs lower than a fully dense model.
Fine-tuning with latent diffusion models
After obtaining a robust representation of gene expression data through continued pretraining, it was employed as the transcriptomic prompt for cellular image generation using a latent diffusion model. Latent diffusion models have been shown to achieve highly competitive performance on various tasks, including image inpainting, image generation, super-resolution, and semantic scene synthesis. Given the limited number of gene expression-image pairs in spatial transcriptomics datasets, we opted to fine-tune a latent diffusion model for the text-to-image generation task, specifically Stable Diffusion (SD) to generate cellular images from transcriptomic prompts. As opposed to traditional diffusion models, which generate samples in pixel space, SD operates within a latent feature space, which has been shown to significantly reduce computational requirements and introduce fewer spatial down-sampling, resulting in better image synthesis quality.
SD consists of two modules: (1 ) a regularized autoencoder that compresses images into lowdimensional latent features and reconstructs them, and (2) a U-Net-based denoising model with attention mechanisms for more flexible conditional image generation. At the training stage of SD, the forward diffusion process was first applied to image latent embeddings derived by the autoencoder and produces a series of noisy embeddings. Specifically, given a cellular image x in pixel space, its corresponding latent embedding z is obtained via an autoencoder £(•), i.e. , z = £(x). Noises were then iteratively added to z for T time steps until it is fully destroyed to a standard Gaussian noise. The t-th noisy embedding zt is sampled from the following distribution:
where I is the identity matrix,
is a variance schedule parameter that determines how much noises are added at each step. These noisy embeddings were then served as the training data for the reverse diffusion process of SD. According to previous studies, conditional information was applied through cross-
attention heads in the attention-based U-Net. In this study, to match the dimensions of the SD inputs, the outputs y of the gene expression encoder is further projected with a projector T8 into a transcriptomic prompt re(y) wjth the aim to learn the reverse diffusion process formulated by
During the fine-tuning process, the gene expression encoder and the U-Net were optimized together. The following loss function was used:
where e8 is the denoising function implemented as a U-Net. In the sampling process, the diffusion model took the Gaussian noise as the initial input and denoises it iteratively through T steps. Finally, the reverted image latent embedding z0 was obtained and input into the decoder ©(•) to generate the cellular images.
Aligning transcriptome, images, and text with BiomedCLIP
By jointly training GenVinci and the SD model, one can fine-tune the projector to get better transcriptomic prompts for cellular image generation tasks. However, the original SD model is tailored for the text-to-image generation task, the unique properties of gene expression data result in latent spaces that are quite different from those of texts. This means that simply fine-tuning the SD model end-to-end with a limited set of gene expression-image pairs makes it difficult to yield an effective alignment between the gene expression and text embeddings.
Fortunately, the SD model already exhibits well-aligned text and image spaces by training on billions of image-text pairs. Therefore, BiomedCLIP, a biomedical vision-language foundation model trained on the PMC-15M corpus, was incorporated to enhance the alignment of spaces related to the transcriptome, images, and text. The PMC-15M corpus includes 15 million figure-caption pairs extracted from over 3 million biomedical research articles in PubMed Central, which includes a diverse range of biomedical images. Specifically, previous gene expression projector with an additional alignment projector was followed to refine the transcriptomic embeddings into multi-modal pseudo vision-language embeddings. These embeddings were adjusted to be either dimensionally or distributional^ compatible with those used in BiomedCLIP. During training, a weak supervised constraint is applied to minimize the similarity between the transcriptomic embeddings and the multi-modal embeddings produced by the BiomedCLIP image encoder, which is defined as follows:
where h represents the alignment projector, I represents the cellular image sent to the BiomedCLIP image encoder E,. The loss function serves to encourage closer alignment between the transcriptomic embeddings and the cellular image embeddings, and by extension, it also aligns with the potential textual description features of the images. As a result, the spaces for transcriptome, images, and texts are unified, and the optimized transcriptomic representation became more suitable for SD image generation, thereby improving the quality of the generated images. Throughout the fine-tuning phase, the parameters of the pre-trained BiomedCLIP model remain fixed. Finally, the entire loss function can be formulated as: L-AH = LSD + LCLIP where the weight A is a hyperparameter that controls the degree of supervision from BiomedCLIP.
T anslating gene expression matrix to whole-slide image
The potential of using GenVinci to translate gene expression matrices into whole-slide images was also explored. However, unlike the numerous gene expression-image pairs at the single-cell level, there are only a few paired gene expression matrices and whole-slide images available. In addition, the varying number of cells captured across different tissue slides presents a challenge and limits the feasibility of generating whole-slide images. Despite these constraints, a matrix-to-slide GenVinci was successfully implemented that leverages the previous pre-trained GenVinci by gene expression-image pairs.
Specifically, the goal was to generalize GenVinci from a cell-to-tile level to a matrix-to-slide level. The process begins with resizing whole-slide images to 512 x 512-pixel images to match the output dimensions of the SD model. Next, for each gene expression matrix, cells with insufficient information, i.e., those with a total count of fewer than 200, were removed. Then, 512 cells were randomly sampled to represent the tissue transcriptome signals. In cases where matrices contain fewer than 512 cells, they were padded with zero-expressed genes to maintain consistent dimensions. This approach is adaptable across various imaging and sequencing technologies, accommodating variations in the number of captured cells. Next, the pre-trained cell embeddings for each cell within the gene expression matrix were extracted by using the gene expression encoder of cell-to-tile pre-trained GenVinci. These embeddings were passed through a new projector to reshape the matrix to fit the requirements of the SD model. The projector was further fine tuned in conjunction with the U-Net component of the SD model. All other modules and parameters remained consistent with those used in the cell-to-tile GenVinci.
CNN baseline
First was implemented a Convolutional Neural Network (CNN) as the baseline for the gene-to- image translation task to compare with GenVinci. The CNN baseline is inspired by the decoder network of U-Net, which consists of fully connected layers and transpose convolutional layers to increase the spatial dimension, and batch normalization and ReLU activations to enhance the model's ability to learn nonlinear functions. The CNN baseline was trained to predict cellular images from the normalized gene expression using MSE loss. The top 2000 highly variable genes (HVGs) were selected as input for the datasets with whole transcriptomic sequencing. The model was trained for 500 epochs with the Adam optimizer, using a learning rate of 0.001 and a batch size of 32.
Evaluation metrics
Frechet Inception Distance (FID)
FID measures the difference between the distributions of real and generated images. Specifically, the ImageNet-pre-trained lnception-V3 convolutional neural network was adopted to extract 2048- dimensional feature vectors for both real and generated images. These feature vectors capture the essential visual information of the images, as learned by the network during its training on ImageNet dataset. Then the mean and covariance are computed for each set of features to calculate the Wasserstein-2 distance. A lower FID score indicates that the distributions of generated images are closer to the real images, suggesting better quality and realism of the generated images. FID is expressed as:
where x denotes the feature vectors of the real images, g denotes the feature vectors of the generated images. Tr denotes the trace of a matrix, which is the sum of all the diagonal elements.
Cosine similarity
Cosine similarity was also used to measure the similarity between the real and generated images. Cosine similarity is a mathematical measure used to determine the similarity between two vectors in a multi-dimensional space. In this study, the cosine similarity was calculated by the extracted 2048-dimensional feature vectors of the real and generated images from an ImageNet-pre-trained ResNet-50 model.
Accuracy of annotation and clustering tasks
To further determine whether the generated cellular images not only reflect gene expression constraints but also preserve pathological signals, classifiers were trained based on the real and generated images for cell-type or pathological annotations. Specifically, the image features were extracted from an ImageNet-pre-trained ResNet-50 model with classification heads tailored for different annotation tasks. Both Accuracy and F1 scores were reported as the evaluation metrics. The higher Accuracy and F1 score indicated better classification performance. For clustering performance evaluation, three commonly used metrics were adopted: normalized mutual information (NMI), adjusted Rand index (ARI), and average silhouette width (ASW). The NMI score was computed to measure the concurrence between ground truth cell type labels and Louvain cluster labels from integrated cell embeddings, with clustering resolutions ranging from 0.1 to 2. The ARI was used to assess the agreement between annotated labels and NMI-optimized Louvain clusters, adjusting for randomly correct labels. The ASW score, ranging from -1 to 1 , was calculated to evaluate the relationship between a cell's within-cluster distances and its distances to the closest cluster boundaries. Higher scores for NMI and ARI indicate better matches, while an ASW score of 1 indicates well-defined clusters.
Implementation details
The same hyperparameters as the Geneformer were used during the pretraining of GenVinci. Specifically, a linear learning rate scheduler with warmup was employed, setting the maximum learning rate to 5 x 10-5, and an Adam optimizer was used with a weight decay of 0.001 and a batch size of 16. The dimensionality of the input and output was 256, and the inner layer had a dimensionality of 512. For all fully connected layers, a nonlinear activation function was utilized in the form of a Rectified Linear Unit (ReLU); a dropout probability of 0.02; a standard deviation of 0.02 for the initializer of weight matrices; and an epsilon value of 1 x 10-12 for layer normalization layers.
For the fine-tuning of the SD model, SD version 1 .5 was utilized for image generation and denoising diffusion implicit model (DDIM) sampling strategy, with a batch size of 4. The Adam optimizer was employed, setting the learning rate to 5 x 10-5 for training. Each model was fine-tuned for 500 epochs, and the number of sampling steps was set to 250. The parameters of the BiomedCLIP model are always fixed.
Datasets
Collection of the GenCorpus-210M
GenVinci collects more than 210 million dissociated single-cell and targeted spatial transcriptomics data points for pretraining, covering 82 general tissues and 74 sequencing technologies from both human and mouse samples. Among them, more than 60 million dissociated single-cell data points were downloaded from CELLxGENE, a big data platform that provides curated and interoperable single-cell datasets. Additionally, we downloaded 1 million single-cell datasets from SenNet (focusing on aging and senescent cells), the Gene Expression Omnibus, and other public sources. For spatial transcriptomics datasets, we retrieved approximately 50 million data points from several databases: 10x Genomics (10 million), SODB (30 million), HEST-1 K (2 million), and STimage-1 K4M (4 million). Furthermore, we manually downloaded over 100 million spatial transcriptomics datasets from sources including Vizgen, NanoString, Single Cell Portal, Gene Expression Omnibus, BICAN, HubMAP, Allen Brain Map, Spatial Research, SenNet, and recently published HCA and HTAN paper bundles. Most of these data are not publicly available or previously used and will contribute novel biological insights for model training.
Datasets used for downstream tasks and evaluations Spatial Transcriptomics with H&E images
The Spatial Transcriptomics dataset consists of sections from 23 breast cancer patients with Luminal A, Luminal B, triple-negative, and HER2-positive subtypes. Each patient includes 2 to 3 sections with both H&E images (scanned at x20 magnification) and spatial gene expression data. The number of spots in each section ranged from 256 to 712, and the number of cells in each spot varied from 5 to several tens. Each spot had a diameter of 100 m and was arranged in a grid with a center-to-center distance of 200 pm. In total, 30,612 spots from 68 tissue sections were obtained, and each spot was measured by 26,949 gene expression. Detailed steps for collecting and handling breast cancer biopsies are known in the art, and the RNA sequencing of the barcoded libraries and subsequent analysis of tissue sections were performed using standardized protocols.
10X Genomics Visium with H&E images
The Visium dataset included two sections of invasive ductal carcinoma breast tissue. The tissue was embedded and cryosectioned, as described in the Visium Spatial Protocols (CG000240). Tissue sections, with a thickness of 10 pm, were placed on Visium Gene Expression Slides for sequencing. The sections were scanned at x 20 magnification, and the H&E images were acquired using a Nikon Ti2-E microscope. The Visium Spatial Gene Expression library was prepared following the instructions provided in the Visium Spatial Gene Expression Reagent Kits User Guide (CG000239). Previous studies were followed to segment patches of 224 x 224 pixels centered on the Visium spots. The standard quality control steps in Scanpy were performed for the gene expression data. Finally, obtained were 3,798 spots for Section 1 and 3,987 spots for Section 2, both including the measurement of 30,612 spatially resolved gene expression.
MERFISH with H&E images
The MERFISH dataset was collected by the Human Tumor Atlas Project Pilot (HTAPP). It consists of two sets of 4 and 5 metastatic breast cancer (MBC) tissue samples, along with single-nucleus gene expression data (sn-MBC) and single-cell gene expression data (sc-MBC), respectively. Both H&E images and gene expression data for each patient sample were derived from different biopsies of the same metastasis. Multiplex FISH measurements of 212 genes were performed on consecutive sections using MERFISH and included regional annotations on the H&E sections provided by expert pathologists for both datasets. The H&E images were segmented and padded to a size of 64 x 64 pixels. In total, 45,989 gene expression-image pairs were obtained for the sn-MBC dataset and 42,229 gene expressionimage pairs were obtained for the sc-MBC dataset. One scRNA-seq dataset from the adjacent slide of one patient for data augmentation evaluation.
Nanostring CosMx SMI with DAPI & IF-stained images
The NanoString dataset was generated using a 960-plex CosMx RNA panel run on a CosMx SMI prototype which contains eight different samples from five non-small cell lung cancer (NSCLC) tissues. TheLung-5-1 data files were selected with the corresponding raw morphological image files for analysis. The selected dataset consisted of fluorescence imaging data from 30 FOVs, each stained with nucleus (DAPI) and membrane markers (CD45, PanCK, and CD3) to facilitate morphology-based cell segmentation. The single-nucleus DAPI images were segmented and padded to a size of 96 x 96 pixels, and the whole-cell IF-stained images are segmented and padded to a size of 128 x 128 pixels. In total, there were 87,848 cells for the sn-NSCLC corpus and 95,255 cells for the sn-NSCLC corpus, with all measurements encompassing 957 genes.
Patch-seq with reconstructions of neuron morphology
Two of the largest Patch-seq datasets with publicly accessible sequencing data were retrieved as were files of reconstructed neuron morphology from the mouse motor cortex and the mouse visual cortex. After removing samples lacking complete neuronal morphology reconstructions, a total of 672 cells from the motor cortex dataset and 573 cells from the visual cortex dataset were retained. Both datasets provided morphological data in SWC file format, which includes 3D coordinates of neuronal compartments along with their types (e.g., soma, axon, basal dendrite, or apical dendrite). For the wholeneuron SWC files, a previous study was followed to project the reconstructed neuron morphologies onto the xy-plane and color-coded the pixels by neuronal component using the Neurom package. Finally, each SWC reconstruction of neuron morphology was saved as a 256 x 256-pixel RGB image. Due to the limited number of available gene expression-image pairs, the two datasets were merged to construct a total of 1 ,245 cells with measurements for 16,323 genes.
STcEM MERFISH with EM images
The STcEM MERFISH dataset was retrieved from a recently published study, which collected paired MERFISH and EM data from a mouse model of LPC-induced demyelinating injury. This dataset correlates large-area electron microscopy scanning with MERFISH on adjacent tissue sections. Herein,
the adjacent sections from a mouse brain were processed in parallel using harmonized MERFISH and EM protocols and spatially aligned to connect transcriptional profiles with the ultrastructure of the regions of interest. Aligning MERFISH and EM imaging data allowed for projection into a common coordinate system for analysis, resulting in 934 EM images and 1 ,502 MERFISH profiles. The EM images were then segmented and padded to a size of 256 x 256 pixels, and Hungarian algorithm was utilized to complete a one-to-one pairing for cells in MERFISH and EM sections. This resulted in 934 RNA-EM pairs, each cell measured for 286 genes.
Data pre-processing
For gene expression data, to meet the requirement of GenVinci, the gene symbol ID was first converted to Ensembl annotations from the BioMart database using the scanpy. queries. biomart annotations function. For the datasets from mouse samples, we translated mouse gene symbols to human equivalents using the mousipy package (pypi.org/project/mousipy/). The transformation involves retaining direct human orthologs when available, excluding genes with explicit no ortholog entries, and aggregating expression counts for genes associated with multiple human orthologs. When a gene does not have a direct human ortholog match, its symbol was capitalized as a proxy for a potential human counterpart. This meticulous approach ensured the accuracy of cross-species gene symbol conversion, facilitating robust analysis of gene expression across species. In addition, if not specified, genes detected in fewer than 3 cells were excluded from all the datasets. The imaging data are segmented as described in the previous section. During training, we augmented the dataset by randomly rotating the image by 0, 90, 180, or 270 degrees and flipping the image horizontally 50% of the time.
Numbered Embodiments
1 . A computer implemented method of generating a multi-dimensional image of a cell or organelle, the method comprising:
(a) providing a gene expression profile from sequencing and/or hybridization data of a plurality of genes from the cell or organelle;
(b) using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
2. The method of embodiment 1 , wherein the computer is programmed with a plurality of modules that generate the multi-dimensional image.
3. The method of embodiment 2, wherein the computer is programmed with an encoder model, a diffusion model, and a pretraining model.
4. The method of any one of embodiments 1 -3, wherein the method further comprises providing a raw image of the cell or organelle.
5. The method of any one of embodiments 1 -4, wherein the computer is programmed to normalize and bin the gene expression of each of the plurality of genes.
6. The method of embodiment 5, wherein the computer is programmed to rank each of the plurality of genes by the normalized binned gene expression of each of the plurality of genes.
7. The method of embodiment 6, wherein the computer is programmed to transform the binned expression of each of the plurality of genes to produce the image.
8. The method of embodiment 7, wherein the computer is programmed to weight the binned expression or an identification thereof of each of the plurality of genes.
9. The method of embodiment 8, the computer is programmed to mask a subset of the weighted binned expression or the identification thereof of the plurality of genes.
10. The method of embodiment 9, wherein the computer is programmed to predict a remainder of the genes or binned expression that are not part of the masked subset.
11 . The method of any one of embodiments 1 -10, wherein the computer is programmed to compress the image to reduce a dimensionality and add or reduce noise of the image.
12. The method of embodiment 11 , wherein the computer is programmed to iteratively add or reduce noise of the image.
13. The method of any one of embodiments 1 -12, wherein the computer is programmed with a diffusion model.
14. The method of any one of embodiments 1 -13, wherein the machine learning algorithm model is trained with an empirical database of single cell gene expression profiles.
15. The method of embodiment 14, wherein the computer is programmed to optimize the image based on the empirical database to produce the multi-dimensional image of the cell or organelle.
16. The method of embodiment 14 or 15, wherein the computer is programmed to optimize the image based on the empirical database comprising a plurality of sets comprising matched gene expression profiles, images, and/or text annotation.
17. The method of embodiment 16, wherein the computer is programmed to correlate the set of gene expression profiles with the image of the cell or organelle.
18. The method of embodiment 16 or 17, wherein the computer is programmed to align the plurality of gene expression profiles with the plurality of sets comprising the matched gene expression profiles, images, and/or text annotations to produce a plurality of gene expression matrices.
19. The method of embodiment 18, wherein the computer is programmed to resize each image from the plurality of sets to match the size of the image.
20. The method of embodiment 18 or 19, wherein the computer is programmed to translate the plurality of gene expression matrices into the multi-dimensional image.
21 . The method of any one of embodiments 16-20, wherein the computer is programmed to rank the multi-dimensional image based on a similarity with the plurality of sets comprising the matched gene expression profiles, images, and/or text annotations.
22. The method of embodiment 21 , wherein the computer is programmed to produce a plurality of multidimensional images, and the method further comprises selecting an image with the highest rank.
23. The method of any one of embodiments 1 -22, wherein the gene expression profiled is obtained by performing a spatial transcriptomic analysis.
24. The method of embodiment 23, wherein the spatial transcriptomic analysis comprises single nucleus RNA (snRNA) sequencing, single cell RNA (scRNA) sequencing, patch sequencing, Visium, Xenium, NanoString GeoMx, STARmap, Slide-Seq, Slide-Tags, fluorescence in situ hybridization (FISH), multiplexed error-robust fluorescence in situ hybridization (MERFISH), electron microscopy, hematoxylin and eosin (H&E) imaging, fluorescence microscopy, 4',6-diamidino-2-phenylindole (DAPI), immunofluorescence (IF), or a combination thereof.
25. The method of any one of embodiments 1 -24, wherein the organelle is a nucleus, nucleolus, ribosome, vesicle, endoplasmic reticulum, Golgi apparatus, vacuole, centriole, lysosome, or mitochondria.
26. The method of any one of embodiments 1 -25, wherein the cell is a neuron, and the method shows a morphology of the neuron.
27. The method of any one of embodiments 1 -26, wherein the method generates a multi-dimensional image of a plurality of cells or organelles.
28. The method of embodiment 27, wherein a tissue comprises the plurality of cells, and the method produces a multi-dimensional image of the tissue.
29. The method of any one of embodiments 1 -28, wherein the method produces an electron microscopy, hematoxylin and eosin, or a fluorescence microscopy multi-dimensional image.
30. The method of embodiment 29, wherein the method produces fluorescence cell morphology.
31 . A system to determine a multi-dimensional image of a cell or organelle, comprising: a storage device; and a processor communicatively coupled to the storage device, wherein the processor executes application code instructions that are stored in the storage device to cause the system to:
(a) receive a gene expression profile of a plurality of genes from the cell or organelle; and
(b) using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
32. A computer program product for determining a multi-dimensional image of a cell or organelle, comprising: a non-transitory computer-readable storage device having computer-executable program instructions embodied thereon that when executed by a computer cause the computer to determine a multi-dimensional image of a cell or an organelle, the computer-executable program instructions comprising:
(a) computer-executable program instructions to receive a gene expression profile of a plurality of genes from the cell or organelle; and
(b) computer-executable program instructions to determine the multi-dimensional image of the cell or organelle using the gene expression profile, wherein computer program product is programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
33. A software program product for determining a multi-dimensional image of a cell or organelle, comprising: a computer-executable software program comprising instructions that when executed by a computer cause the computer to determine a multi-dimensional image of a cell or an organelle, the computer-executable program instructions comprising:
(a) computer-executable software program instructions to receive a gene expression profile of a plurality of genes from the cell or organelle;
(b) computer-executable software program instructions to determine the multidimensional image of the cell or organelle using the gene expression profile, wherein software program product is programmed with a machine learning algorithm model to determine the multidimensional image of the cell or organelle.
Other Embodiments
While the invention has been described in connection with specific embodiments thereof, it will be understood that it is capable of further modifications and this application is intended to cover any variations, uses, or adaptations of the invention following, in general, the principles of the invention and including such departures from the invention that come within known or customary practice within the art to which the invention pertains and may be applied to the essential features hereinbefore set forth, and follows in the scope of the claims. Other embodiments are within the claims.
Claims
1 . A computer implemented method of generating a multi-dimensional image of a cell or organelle, the method comprising:
(a) providing a gene expression profile from sequencing and/or hybridization data of a plurality of genes from the cell or organelle;
(b) using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
2. The method of claim 1 , wherein the computer is programmed with a plurality of modules that generate the multi-dimensional image.
3. The method of claim 2, wherein the computer is programmed with an encoder model, a diffusion model, and a pretraining model.
4. The method of claim 1 , wherein the method further comprises providing a raw image of the cell or organelle.
5. The method of claim 1 , wherein the computer is programmed to normalize and bin the gene expression of each of the plurality of genes.
6. The method of claim 5, wherein the computer is programmed to rank each of the plurality of genes by the normalized binned gene expression of each of the plurality of genes.
7. The method of claim 6, wherein the computer is programmed to transform the binned expression of each of the plurality of genes to produce the image.
8. The method of claim 7, wherein the computer is programmed to weight the binned expression or an identification thereof of each of the plurality of genes.
9. The method of claim 8, the computer is programmed to mask a subset of the weighted binned expression or the identification thereof of the plurality of genes.
10. The method of claim 9, wherein the computer is programmed to predict a remainder of the genes or binned expression that are not part of the masked subset.
11 . The method of claim 1 , wherein the computer is programmed to compress the image to reduce a dimensionality and add or reduce noise of the image.
12. The method of claim 11 , wherein the computer is programmed to iteratively add or reduce noise of the image.
13. The method of claim 1 , wherein the computer is programmed with a diffusion model.
14. The method of claim 1 , wherein the machine learning algorithm model is trained with an empirical database of single cell gene expression profiles.
15. The method of claim 14, wherein the computer is programmed to optimize the image based on the empirical database to produce the multi-dimensional image of the cell or organelle.
16. The method of claim 14 or 15, wherein the computer is programmed to optimize the image based on the empirical database comprising a plurality of sets comprising matched gene expression profiles, images, and/or text annotation.
17. The method of claim 16, wherein the computer is programmed to correlate the set of gene expression profiles with the image of the cell or organelle.
18. The method of claim 16, wherein the computer is programmed to align the plurality of gene expression profiles with the plurality of sets comprising the matched gene expression profiles, images, and/or text annotations to produce a plurality of gene expression matrices.
19. The method of claim 18, wherein the computer is programmed to resize each image from the plurality of sets to match the size of the image.
20. The method of claim 18, wherein the computer is programmed to translate the plurality of gene expression matrices into the multi-dimensional image.
21 . The method of claim 1 , wherein the computer is programmed to rank the multi-dimensional image based on a similarity with the plurality of sets comprising the matched gene expression profiles, images, and/or text annotations.
22. The method of claim 21 , wherein the computer is programmed to produce a plurality of multidimensional images, and the method further comprises selecting an image with the highest rank.
23. The method of claim 1 , wherein the gene expression profiled is obtained by performing a spatial transcriptomic analysis.
24. The method of claim 23, wherein the spatial transcriptomic analysis comprises single nucleus RNA (snRNA) sequencing, single cell RNA (scRNA) sequencing, patch sequencing, Visium, Xenium, NanoString GeoMx, STARmap, Slide-Seq, Slide-Tags, fluorescence in situ hybridization (FISH), multiplexed error-robust fluorescence in situ hybridization (MERFISH), electron microscopy, hematoxylin and eosin (H&E) imaging, fluorescence microscopy, 4',6-diamidino-2-phenylindole (DAPI), immunofluorescence (IF), or a combination thereof.
25. The method of claim 1 , wherein the organelle is a nucleus, nucleolus, ribosome, vesicle, endoplasmic reticulum, Golgi apparatus, vacuole, centriole, lysosome, or mitochondria.
26. The method of claim 1 , wherein the cell is a neuron, and the method shows a morphology of the neuron.
27. The method of claim 1 , wherein the method generates a multi-dimensional image of a plurality of cells or organelles.
28. The method of claim 27, wherein a tissue comprises the plurality of cells, and the method produces a multi-dimensional image of the tissue.
29. The method of claim 1 , wherein the method produces an electron microscopy, hematoxylin and eosin, or a fluorescence microscopy multi-dimensional image.
30. The method of claim 29, wherein the method produces fluorescence cell morphology.
31 . A system to determine a multi-dimensional image of a cell or organelle, comprising: a storage device; and a processor communicatively coupled to the storage device, wherein the processor executes application code instructions that are stored in the storage device to cause the system to:
(a) receive a gene expression profile of a plurality of genes from the cell or organelle; and
(b) using the gene expression profile in a computer programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
32. A computer program product for determining a multi-dimensional image of a cell or organelle, comprising: a non-transitory computer-readable storage device having computer-executable program instructions embodied thereon that when executed by a computer cause the computer to determine a multi-dimensional image of a cell or an organelle, the computer-executable program instructions comprising:
(a) computer-executable program instructions to receive a gene expression profile of a plurality of genes from the cell or organelle; and
(b) computer-executable program instructions to determine the multi-dimensional image of the cell or organelle using the gene expression profile, wherein computer program product is programmed with a machine learning algorithm model to determine the multi-dimensional image of the cell or organelle.
33. A software program product for determining a multi-dimensional image of a cell or organelle, comprising:
a computer-executable software program comprising instructions that when executed by a computer cause the computer to determine a multi-dimensional image of a cell or an organelle, the computer-executable program instructions comprising:
(a) computer-executable software program instructions to receive a gene expression profile of a plurality of genes from the cell or organelle;
(b) computer-executable software program instructions to determine the multidimensional image of the cell or organelle using the gene expression profile, wherein software program product is programmed with a machine learning algorithm model to determine the multidimensional image of the cell or organelle.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463566420P | 2024-03-18 | 2024-03-18 | |
| US63/566,420 | 2024-03-18 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025199125A1 true WO2025199125A1 (en) | 2025-09-25 |
Family
ID=97140254
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2025/020407 Pending WO2025199125A1 (en) | 2024-03-18 | 2025-03-18 | Methods and systems for determining cellular and tissue images from gene expression |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025199125A1 (en) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20200224172A1 (en) * | 2017-09-19 | 2020-07-16 | The Broad Institute, Inc. | Methods and systems for reconstruction of developmental landscapes by optimal transport analysis |
| WO2023091970A1 (en) * | 2021-11-16 | 2023-05-25 | The General Hospital Corporation | Live-cell label-free prediction of single-cell omics profiles by microscopy |
| US20240047081A1 (en) * | 2022-07-28 | 2024-02-08 | Regents Of The University Of Michigan | Designing Chemical or Genetic Perturbations using Artificial Intelligence |
-
2025
- 2025-03-18 WO PCT/US2025/020407 patent/WO2025199125A1/en active Pending
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20200224172A1 (en) * | 2017-09-19 | 2020-07-16 | The Broad Institute, Inc. | Methods and systems for reconstruction of developmental landscapes by optimal transport analysis |
| WO2023091970A1 (en) * | 2021-11-16 | 2023-05-25 | The General Hospital Corporation | Live-cell label-free prediction of single-cell omics profiles by microscopy |
| US20240047081A1 (en) * | 2022-07-28 | 2024-02-08 | Regents Of The University Of Michigan | Designing Chemical or Genetic Perturbations using Artificial Intelligence |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Maleki et al. | Breast cancer diagnosis from histopathology images using deep neural network and XGBoost | |
| Dippel et al. | RudolfV: a foundation model by pathologists for pathologists | |
| Zhao et al. | Application of deep learning in histopathology images of breast cancer: a review | |
| CN117422704B (en) | A cancer prediction method, system and device based on multimodal data | |
| Liu et al. | A comprehensive overview of graph neural network-based approaches to clustering for spatial transcriptomics | |
| Badea et al. | Identifying transcriptomic correlates of histology using deep learning | |
| Khan et al. | Spatial transcriptomics data and analytical methods: an updated perspective | |
| CN117616505A (en) | Systems and methods for associating compounds with physiological conditions using fingerprint analysis | |
| Tavakoli | Seq2image: Sequence analysis using visualization and deep convolutional neural network | |
| CN117788369A (en) | Method for unsupervised recognition of single cell morphology map analysis based on deep learning | |
| Xun et al. | Microsnoop: A generalist tool for microscopy image representation | |
| Lee et al. | MorphNet predicts cell morphology from single-cell gene expression | |
| Wang et al. | A generalizable Hi-C foundation model for chromatin architecture, single-cell and multi-omics analysis across species | |
| Sangeetha et al. | Advanced segmentation method for integrating multi-omics data for early cancer detection | |
| Hasan et al. | Prototypical Few-Shot Learning for Histopathology Classification: Leveraging Foundation Models With Adapter Architectures | |
| Qu et al. | Enhancing understandability of omics data with shap, embedding projections and interactive visualisations | |
| Moutakanni et al. | Cell-DINO: Self-supervised image-based embeddings for cell fluorescent microscopy | |
| Yu et al. | scMapNet: Marker-based cell type annotation of scRNA-seq data via vision transfer learning with tabular-to-image transformations | |
| Thanh-Hai et al. | Diagnosis approaches for colorectal cancer using manifold learning and deep learning | |
| Wei et al. | Deep representation learning for image-based cell profiling | |
| Lei et al. | Rapidly Antibiotic Susceptibility Prediction via Deep Learning From Bacterial Fluorescence Microscopy Images | |
| Fadhil et al. | Classification of cancer microarray data based on deep learning: a review | |
| Hoque | Comparative Analysis of CNN, Vision Transformers, and Hybrid Models for Skin Lesion Classification Using the HAM10000 Dataset | |
| Carrillo Pérez | Development of advanced machine learning models for the fusion of heterogeneous biological sources in clinical decision support systems for cancer | |
| Ahsan et al. | DeepPathway: Predicting Pathway Expression from Histopathology Images |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25774482 Country of ref document: EP Kind code of ref document: A1 |